REVIEW 22 cited by
OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Understanding the evolution of 3D scenes is important for effective autonomous driving. While conventional methods mode scene development with the motion of individual instances, world models emerge as a generative framework to describe the general scene dynamics. However, most existing methods adopt an autoregressive framework to perform next-token prediction, which suffer from inefficiency in modeling long-term temporal evolutions. To address this, we propose a diffusion-based 4D occupancy generation model, OccSora, to simulate the development of the 3D world for autonomous driving. We employ a 4D scene tokenizer to obtain compact discrete spatial-temporal representations for 4D occupancy input and achieve high-quality reconstruction for long-sequence occupancy videos. We then learn a diffusion transformer on the spatial-temporal representations and generate 4D occupancy conditioned on a trajectory prompt. We conduct extensive experiments on the widely used nuScenes dataset with Occ3D occupancy annotations. OccSora can generate 16s-videos with authentic 3D layout and temporal consistency, demonstrating its ability to understand the spatial and temporal distributions of driving scenes. With trajectory-aware 4D generation, OccSora has the potential to serve as a world simulator for the decision-making of autonomous driving. Code is available at: https://github.com/wzzheng/OccSora.
Forward citations
Cited by 22 Pith papers
-
FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction
Factorized Dense Routing approximates unconstrained 2D-to-3D feature mixing by hierarchical tensor contractions, yielding global-context occupancy prediction that remains robust without camera extrinsics.
-
A Comprehensive Survey on World Models for Embodied AI
A unified three-axis taxonomy — functionality, temporal modeling, spatial representation — organizes the world-model literature for embodied AI.
-
$I^{2}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting
I2-World forecasts 3D occupancy over 3 seconds using an intra/inter tokenizer and reports state-of-the-art results, but the gains come mainly from oracle conditioning on the future ego pose at test time.
-
GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control
GeoDrive conditions a frozen video diffusion model on a 3D-rendered version of the requested ego trajectory, cutting trajectory-following error by 42% versus Vista while using 99.7% less training data.
-
Occupancy World Model for Robots
RoboOccWorld predicts future 3D occupancy for indoor robots by conditioning an autoregressive transformer on the next camera pose, outperforming OccWorld on a restructured ScanNet benchmark.
-
RoboOcc: Enhancing the Geometric and Semantic Scene Understanding for Robots
RoboOcc uses opacity-guided and geometry-aware encoding of 3D Gaussian representations to achieve state-of-the-art indoor 3D semantic occupancy prediction from monocular RGB.
-
Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction
GDFusion fuses scene, motion, and geometry cues through gradient-descent-style RNN updates, improving mIoU by 1.4 to 4.8 points on Occ3D while cutting inference memory by 27 to 72 percent.
-
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
HERMES unifies BEV-based scene understanding and future point cloud generation in a single LLM-driven self-driving world model, with reported gains on nuScenes and OmniDrive-nuScenes.
-
Doe-1: Closed-Loop Autonomous Driving with Large World Model
Doe-1 unifies perception, prediction, and planning in autonomous driving into a single autoregressive next-token generation model over image, text, and action tokens.
-
Owl-1: Omni World Model for Consistent Long Video Generation
Owl-1 generates long, multi-scene videos by using a language model to maintain a latent state and predict text dynamics, then rendering each clip with a video diffusion model.
-
GaussianFormer-2: Probabilistic Gaussian Superposition for Efficient 3D Occupancy Prediction
GaussianFormer-2 predicts 3D semantic occupancy from cameras by multiplying Gaussian occupancy probabilities and using a Gaussian mixture for semantics, beating prior methods with far fewer Gaussians.
-
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
EmbodiedOcc maintains an explicit global Gaussian memory that is progressively updated from monocular RGB frames, and it introduces a reorganized ScanNet benchmark for embodied 3D occupancy prediction.
-
Spatiotemporal Decoupling for Efficient Vision-Based Occupancy Forecasting
EfficientOCF forecasts 3D occupancy by decoupling it into 2D BEV occupancy, height, and instance flow, achieving state-of-the-art accuracy and 82.33 ms inference on autonomous driving datasets.
-
DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation
DrivingSphere combines occupancy-based 4D world generation with video diffusion to create a closed-loop simulation environment for autonomous driving evaluation.
-
A Definition and Roadmap for World Models
A perspective article defining world models as finite-resource compression of physical state transitions and outlining a roadmap toward physical AGI via unified representations and interactive simulators.
-
UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving
A unified VLM for autonomous driving that couples trajectory planning with future-frame image generation improves open- and closed-loop planning metrics on Bench2Drive and nuScenes.
-
Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method
UniScenev2 scales occupancy-centric driving-scene generation to NuPlan scale, releasing a 3.6M-frame semantic-occupancy dataset and jointly generating occupancy, video, and LiDAR that beats published baselines on its ...
-
QuadricFormer: Scene as Superquadrics for 3D Semantic Occupancy Prediction
QuadricFormer represents 3D scenes as a probabilistic mixture of superquadrics, improving accuracy and efficiency over Gaussian-based occupancy prediction on nuScenes.
-
SSEditor: Controllable Mask-to-Scene Generation with Diffusion Model
SSEditor generates controllable 3D semantic urban scenes from mask conditions using a triplane autoencoder and a mask-conditional diffusion model, avoiding multi-step resampling.
-
From 2D to 3D Cognition: A Brief Survey of General World Models
A survey proposing a two-pillar, three-capability framework that organizes recent AI world models by their transition from 2D visual prediction to 3D cognition.
-
A Survey of World Models for Autonomous Driving
A survey presenting a three-branch taxonomy of world models for autonomous driving, plus benchmark tables comparing representative generation and planning methods on nuScenes, Waymo, Occ3D, and CarlaSC.
-
Vision Technologies with Applications in Traffic Surveillance Systems: A Holistic Survey
A survey that maps traffic surveillance vision tasks into low- and high-level groups, proposes five recurring limitations, and sketches a foundation-model roadmap.
Discussion (0). Continue with ORCID to comment.