Pith. sign in

REVIEW 17 cited by

Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.14492 v2 pith:QPAOJN73 submitted 2025-03-18 cs.CV cs.AIcs.LGcs.RO

classification cs.CVcs.AIcs.LGcs.RO
keywords worldconditionalgenerationspatialadaptivecontrolcosmos-transfer1demonstrate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Cosmos-Transfer, a conditional world generation model that can generate world simulations based on multiple spatial control inputs of various modalities such as segmentation, depth, and edge. In the design, the spatial conditional scheme is adaptive and customizable. It allows weighting different conditional inputs differently at different spatial locations. This enables highly controllable world generation and finds use in various world-to-world transfer use cases, including Sim2Real. We conduct extensive evaluations to analyze the proposed model and demonstrate its applications for Physical AI, including robotics Sim2Real and autonomous vehicle data enrichment. We further demonstrate an inference scaling strategy to achieve real-time world generation with an NVIDIA GB200 NVL72 rack. To help accelerate research development in the field, we open-source our models and code at https://github.com/nvidia-cosmos/cosmos-transfer1.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions

    cs.RO 2026-08 conditional novelty 6.0 of 10

    An autoregressive video world model conditioned on URDF-rendered visual actions generalizes to unseen scenes and can evaluate policies and synthesize training data.

  2. muSync-GS: Physics-Synchronized Driving Video Synthesis for Weather and Geometric Road Hazards

    cs.CV 2026-08 conditional novelty 6.0 of 10

    muSync-GS couples weather and road-shape edits in driving videos to a calibrated vehicle-dynamics model, so the synthesized ego motion and telemetry change with the same controls that drive the visual edits.

  3. Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Tracking plus world-position maps in a neural G-buffer outperform depth as a geometric condition for reference-guided video diffusion rendering on a 68-clip synthetic benchmark.

  4. Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering

    cs.CV 2026-07 conditional novelty 6.0 of 10

    When a camera returns to a spot it visited long ago, loading that earlier frame into the KV cache and biasing attention with depth reprojection keeps the regenerated view consistent.

  5. Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A DiT Mixture-of-Experts world model jointly learns locomotion, dual-arm manipulation, and egocentric hand control, with shared experts for world dynamics and progressive expert expansion for new modalities.

  6. GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Long-horizon action-faithful consistency, not short-term visual realism, dominates world-model reliability for robot policy evaluation; GigaWorld-1 implements that roadmap and gains 14.9% on evaluator-alignment metrics.

  7. PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation

    cs.CV 2025-12 conditional novelty 6.0 of 10

    PhyGDPO uses groupwise direct preference optimization with real videos as winners to make text-to-video models generate more physically plausible videos.

  8. AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis

    cs.RO 2025-12 conditional novelty 6.0 of 10

    A video-diffusion-based data synthesizer that anchors generation on robot-only motion renders expands a few human demos into thousands of kinematically consistent demonstrations and improves downstream policies in sim...

  9. A Comprehensive Survey on World Models for Embodied AI

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A unified three-axis taxonomy — functionality, temporal modeling, spatial representation — organizes the world-model literature for embodied AI.

  10. Adaptive Model-Based Transfer Learning for Dynamic HVAC Control

    eess.SY 2026-07 conditional novelty 5.0 of 10

    A dynamic-pretraining and online-adaptation HVAC controller with physics losses and long-term setpoint scoring achieves roughly 0.03-0.18°C average error and transfers between similar real buildings in about a week.

  11. HyPER-GAN: Hybrid Patch-Based Image-to-Image Translation for Real-Time Photorealism Enhancement in Game Engines

    cs.CV 2026-03 conditional novelty 5.0 of 10

    A lightweight hybrid-patch GAN enhances synthetic game images toward photorealism in real time, beating prior lightweight paired translators on speed, KID, and semantic consistency.

  12. 3D and 4D World Modeling: A Survey

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A survey that defines 3D/4D world modeling, organizes methods into VideoGen, OccGen, and LiDARGen categories, and compiles datasets, metrics, and benchmark numbers.

  13. Non-invasive Assessment of Pancreatic Duct Hypertension Using Computational Flow Modeling

    physics.med-ph 2025-08 unverdicted novelty 5.0 of 10

    A computational model estimates pancreatic duct pressure non-invasively from MRCP geometry, with reported agreement against ERCP pressure measurements.

  14. HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A staged pipeline generates layered, mesh-based 3D worlds from text or images by combining panoramic diffusion, semantic layer decomposition, and video-based expansion.

  15. GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

    cs.RO 2026-07 conditional novelty 4.0 of 10

    GigaWorld-Policy-0.5 uses a Mixture-of-Transformers action-expert split and mixed world-model pretraining to reach 85 ms action-only inference with claimed success-rate gains.

  16. A Survey: Learning Embodied Intelligence from Physical Simulators and World Models

    cs.RO 2025-07 conditional novelty 4.0 of 10

    Embodied intelligence learning is reviewed through the complementary lenses of physical simulators and world models, with a proposed IR-L0 to IR-L4 robot capability taxonomy.

  17. Simulating the Unseen: Crash Prediction Must Learn from What Did Not Happen

    cs.LG 2025-05 conditional novelty 4.0 of 10

    Crash prediction should learn from near-miss events and synthetic counterfactual scenarios, not just recorded crashes.

Pith tools