Pith. sign in

REVIEW 21 cited by

WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.09985 v1 pith:LNKK7JUD submitted 2024-01-18 cs.CV

classification cs.CV
keywords worldworlddreamergeneralmodelsvideoenvironmentsgenerationpredicting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

World models play a crucial role in understanding and predicting the dynamics of the world, which is essential for video generation. However, existing world models are confined to specific scenarios such as gaming or driving, limiting their ability to capture the complexity of general world dynamic environments. Therefore, we introduce WorldDreamer, a pioneering world model to foster a comprehensive comprehension of general world physics and motions, which significantly enhances the capabilities of video generation. Drawing inspiration from the success of large language models, WorldDreamer frames world modeling as an unsupervised visual sequence modeling challenge. This is achieved by mapping visual inputs to discrete tokens and predicting the masked ones. During this process, we incorporate multi-modal prompts to facilitate interaction within the world model. Our experiments show that WorldDreamer excels in generating videos across different scenarios, including natural scenes and driving environments. WorldDreamer showcases versatility in executing tasks such as text-to-video conversion, image-tovideo synthesis, and video editing. These results underscore WorldDreamer's effectiveness in capturing dynamic elements within diverse general world environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Action-conditioned world-model verification with conformal first-intervention control and latency-aware suffix repair raises RoboCasa365 success 8.5 points over invocation-matched periodic replanning.

  2. DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Guiding a video diffusion model with VLM-produced spatial-temporal masks focused on deforming regions improves the physical realism and local deformation fidelity of generated melting, squeezing, and fracturing videos.

  3. Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Pre-training a VLA model on 100k hours of auto-labeled UMI trajectories, then post-training on robot data, yields SOTA simulated manipulation and data-efficient fine-tuning.

  4. Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents

    cs.RO 2026-07 unverdicted novelty 6.0 of 10

    A memory-guided LLM planner composes a frozen VLA as a contact-rich primitive with fixed analytic controllers, lifting perturbed manipulation success without VLA finetuning.

  5. Targeted Structure Completion for Sparse-View 3D Reconstruction in Autonomous Driving

    cs.CV 2026-07 conditional novelty 6.0 of 10

    FocusGS localizes a 3D geometric ambiguity manifold from depth discontinuities and instantiates continuous Gaussian queries only there, yielding SOTA sparse-view driving reconstruction with far fewer Gaussians.

  6. GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Long-horizon action-faithful consistency, not short-term visual realism, dominates world-model reliability for robot policy evaluation; GigaWorld-1 implements that roadmap and gains 14.9% on evaluator-alignment metrics.

  7. A Comprehensive Survey on World Models for Embodied AI

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A unified three-axis taxonomy — functionality, temporal modeling, spatial representation — organizes the world-model literature for embodied AI.

  8. Ego-centric Predictive Model Conditioned on Hand Trajectories

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Ego-PM predicts future hand trajectories and then uses them to condition latent diffusion video generation, jointly outputting actions and future frames in egocentric and robotic scenes.

  9. Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PhysVidBench evaluates text-to-video models with 383 PIQA-derived prompts and a caption-based QA pipeline, finding all tested models score below 40% on physical commonsense, with spatial and temporal reasoning the weakest.

  10. Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Hunyuan-GameCraft generates long, action-controlled game videos from a single image by unifying keyboard/mouse inputs into a continuous camera space and conditioning on mixed historical context.

  11. MOVi: Training-free Text-conditioned Multi-Object Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MOVi improves multi-object video generation without retraining by using LLM-planned trajectories to reinitialize the diffusion noise and by reweighting attention to stop objects from mixing together.

  12. GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control

    cs.CV 2025-05 conditional novelty 6.0 of 10

    GeoDrive conditions a frozen video diffusion model on a 3D-rendered version of the requested ego trajectory, cutting trajectory-following error by 42% versus Vista while using 99.7% less training data.

  13. Long-Context State-Space Video World Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A hybrid state-space and local-attention architecture gives autoregressive video diffusion models long-term spatial memory with constant per-frame inference cost, demonstrated on Maze and Minecraft.

  14. ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ProphetDWM is a one-stage diffusion world model that jointly predicts future driving video and low-level actions from a current frame and a short action sequence.

  15. LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments

    cs.RO 2026-07 conditional novelty 5.0 of 10

    LeapBot-WA shows robot policies can be trained with latent world-model predictions instead of pixel video generation, hitting state-of-the-art for predictive action models and staying competitive with generative WAMs.

  16. RynnVLA-002: A Unified Vision-Language-Action and World Model

    cs.RO 2025-11 conditional novelty 5.0 of 10

    A single model that jointly predicts robot actions and future images outperforms separate action-only and video-only models on LIBERO and real SO100 manipulation tasks.

  17. World Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous Driving

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Adding world-model-generated synthetic driving videos to training data, together with a dynamic graph and dilated temporal model, improves accident anticipation accuracy and lead time on multiple benchmarks.

  18. EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A Real2Sim2Real framework that aligns simulator dynamics via differentiable parameter fitting and renders photorealistic policy-training videos with a diffusion model, improving real-world manipulation success.

  19. GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

    cs.RO 2026-07 conditional novelty 4.0 of 10

    GigaWorld-Policy-0.5 uses a Mixture-of-Transformers action-expert split and mixed world-model pretraining to reach 85 ms action-only inference with claimed success-rate gains.

  20. Bounding Distributional Shifts in World Modeling through Novelty Detection

    cs.RO 2025-08 conditional novelty 4.0 of 10

    Attaching a VAE novelty detector to the DINO-WM world model and penalizing out-of-distribution predicted states in CEM planning lowers Chamfer distance on small-data robot manipulation benchmarks.

  21. World Models for Cognitive Agents: Transforming Edge Intelligence in Future Networks

    cs.AI 2025-05 conditional novelty 4.0 of 10

    The paper reviews world-model AI and demonstrates a Dreamer-style Q-learning framework, Wireless Dreamer, on a UAV trajectory planning task, reporting faster convergence than DQN.

Pith tools