REVIEW 17 cited by
Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce Cosmos-Transfer, a conditional world generation model that can generate world simulations based on multiple spatial control inputs of various modalities such as segmentation, depth, and edge. In the design, the spatial conditional scheme is adaptive and customizable. It allows weighting different conditional inputs differently at different spatial locations. This enables highly controllable world generation and finds use in various world-to-world transfer use cases, including Sim2Real. We conduct extensive evaluations to analyze the proposed model and demonstrate its applications for Physical AI, including robotics Sim2Real and autonomous vehicle data enrichment. We further demonstrate an inference scaling strategy to achieve real-time world generation with an NVIDIA GB200 NVL72 rack. To help accelerate research development in the field, we open-source our models and code at https://github.com/nvidia-cosmos/cosmos-transfer1.
Forward citations
Cited by 17 Pith papers
-
GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions
An autoregressive video world model conditioned on URDF-rendered visual actions generalizes to unseen scenes and can evaluate policies and synthesize training data.
-
muSync-GS: Physics-Synchronized Driving Video Synthesis for Weather and Geometric Road Hazards
muSync-GS couples weather and road-shape edits in driving videos to a calibrated vehicle-dynamics model, so the synthesized ego motion and telemetry change with the same controls that drive the visual edits.
-
Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
Tracking plus world-position maps in a neural G-buffer outperform depth as a geometric condition for reference-guided video diffusion rendering on a 68-clip synthetic benchmark.
-
Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering
When a camera returns to a spot it visited long ago, loading that earlier frame into the KV cache and biasing attention with depth reprojection keeps the regenerated view consistent.
-
Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control
A DiT Mixture-of-Experts world model jointly learns locomotion, dual-arm manipulation, and egocentric hand control, with shared experts for world dynamics and progressive expert expansion for new modalities.
-
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
Long-horizon action-faithful consistency, not short-term visual realism, dominates world-model reliability for robot policy evaluation; GigaWorld-1 implements that roadmap and gains 14.9% on evaluator-alignment metrics.
-
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
PhyGDPO uses groupwise direct preference optimization with real videos as winners to make text-to-video models generate more physically plausible videos.
-
AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis
A video-diffusion-based data synthesizer that anchors generation on robot-only motion renders expands a few human demos into thousands of kinematically consistent demonstrations and improves downstream policies in sim...
-
A Comprehensive Survey on World Models for Embodied AI
A unified three-axis taxonomy — functionality, temporal modeling, spatial representation — organizes the world-model literature for embodied AI.
-
Adaptive Model-Based Transfer Learning for Dynamic HVAC Control
A dynamic-pretraining and online-adaptation HVAC controller with physics losses and long-term setpoint scoring achieves roughly 0.03-0.18°C average error and transfers between similar real buildings in about a week.
-
HyPER-GAN: Hybrid Patch-Based Image-to-Image Translation for Real-Time Photorealism Enhancement in Game Engines
A lightweight hybrid-patch GAN enhances synthetic game images toward photorealism in real time, beating prior lightweight paired translators on speed, KID, and semantic consistency.
-
3D and 4D World Modeling: A Survey
A survey that defines 3D/4D world modeling, organizes methods into VideoGen, OccGen, and LiDARGen categories, and compiles datasets, metrics, and benchmark numbers.
-
Non-invasive Assessment of Pancreatic Duct Hypertension Using Computational Flow Modeling
A computational model estimates pancreatic duct pressure non-invasively from MRCP geometry, with reported agreement against ERCP pressure measurements.
-
HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels
A staged pipeline generates layered, mesh-based 3D worlds from text or images by combining panoramic diffusion, semantic layer decomposition, and video-based expansion.
-
GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch
GigaWorld-Policy-0.5 uses a Mixture-of-Transformers action-expert split and mixed world-model pretraining to reach 85 ms action-only inference with claimed success-rate gains.
-
A Survey: Learning Embodied Intelligence from Physical Simulators and World Models
Embodied intelligence learning is reviewed through the complementary lenses of physical simulators and world models, with a proposed IR-L0 to IR-L4 robot capability taxonomy.
-
Simulating the Unseen: Crash Prediction Must Learn from What Did Not Happen
Crash prediction should learn from near-miss events and synthetic counterfactual scenarios, not just recorded crashes.
Discussion (0). Continue with ORCID to comment.