Pith. sign in

REVIEW 9 cited by

DOME: Taming Diffusion Model into High-Fidelity Controllable Occupancy World Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.10429 v1 pith:MVIYLLK7 submitted 2024-10-14 cs.CV

classification cs.CV
keywords occupancyworldmodelabilitycontrollabilitydiffusiondomefuture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose DOME, a diffusion-based world model that predicts future occupancy frames based on past occupancy observations. The ability of this world model to capture the evolution of the environment is crucial for planning in autonomous driving. Compared to 2D video-based world models, the occupancy world model utilizes a native 3D representation, which features easily obtainable annotations and is modality-agnostic. This flexibility has the potential to facilitate the development of more advanced world models. Existing occupancy world models either suffer from detail loss due to discrete tokenization or rely on simplistic diffusion architectures, leading to inefficiencies and difficulties in predicting future occupancy with controllability. Our DOME exhibits two key features:(1) High-Fidelity and Long-Duration Generation. We adopt a spatial-temporal diffusion transformer to predict future occupancy frames based on historical context. This architecture efficiently captures spatial-temporal information, enabling high-fidelity details and the ability to generate predictions over long durations. (2)Fine-grained Controllability. We address the challenge of controllability in predictions by introducing a trajectory resampling method, which significantly enhances the model's ability to generate controlled predictions. Extensive experiments on the widely used nuScenes dataset demonstrate that our method surpasses existing baselines in both qualitative and quantitative evaluations, establishing a new state-of-the-art performance on nuScenes. Specifically, our approach surpasses the baseline by 10.5% in mIoU and 21.2% in IoU for occupancy reconstruction and by 36.0% in mIoU and 24.6% in IoU for 4D occupancy forecasting.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Comprehensive Survey on World Models for Embodied AI

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A unified three-axis taxonomy — functionality, temporal modeling, spatial representation — organizes the world-model literature for embodied AI.

  2. $I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting

    cs.CV 2025-07 reject novelty 6.0 of 10

    I2-World forecasts 3D occupancy over 3 seconds using an intra/inter tokenizer and reports state-of-the-art results, but the gains come mainly from oracle conditioning on the future ego pose at test time.

  3. Epona: Autoregressive Diffusion World Model for Autonomous Driving

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An autoregressive diffusion world model jointly generates the next camera frame and a multi-step trajectory, enabling long videos and real-time planning for autonomous driving.

  4. COME: Adding Scene-Centric Forecasting Control to Occupancy World Model

    cs.CV 2025-06 conditional novelty 6.0 of 10

    COME adds a scene-centric forecasting branch as a ControlNet-style condition to a diffusion occupancy world model, improving static-scene consistency and beating prior methods on Occ3D-nuScenes while hiding a stronger...

  5. GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control

    cs.CV 2025-05 conditional novelty 6.0 of 10

    GeoDrive conditions a frozen video diffusion model on a 3D-rendered version of the requested ego trajectory, cutting trajectory-following error by 42% versus Vista while using 99.7% less training data.

  6. 3D and 4D World Modeling: A Survey

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A survey that defines 3D/4D world modeling, organizes methods into VideoGen, OccGen, and LiDARGen categories, and compiles datasets, metrics, and benchmark numbers.

  7. World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model

    cs.CV 2025-07 conditional novelty 5.0 of 10

    World4Drive couples multiple driving intentions with a latent world model to generate, score, and select trajectories, reporting state-of-the-art perception-free planning on nuScenes and NavSim.

  8. MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Splitting quantization across multiple small sub-codebooks with nested masking raises VQ-VAE reconstruction fidelity, giving MGVQ rFID 0.49 and PSNR 24.70 on ImageNet at 16 times downsampling.

  9. A Survey: Learning Embodied Intelligence from Physical Simulators and World Models

    cs.RO 2025-07 conditional novelty 4.0 of 10

    Embodied intelligence learning is reviewed through the complementary lenses of physical simulators and world models, with a proposed IR-L0 to IR-L4 robot capability taxonomy.

Pith tools