Pith. sign in

REVIEW 13 cited by

OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.03272 v1 pith:HRBRMMZW submitted 2024-09-05 cs.CV cs.RO

classification cs.CVcs.RO
keywords modelworldactionautonomousdrivingoccllamaoccupancyvisual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rise of multi-modal large language models(MLLMs) has spurred their applications in autonomous driving. Recent MLLM-based methods perform action by learning a direct mapping from perception to action, neglecting the dynamics of the world and the relations between action and world dynamics. In contrast, human beings possess world model that enables them to simulate the future states based on 3D internal visual representation and plan actions accordingly. To this end, we propose OccLLaMA, an occupancy-language-action generative world model, which uses semantic occupancy as a general visual representation and unifies vision-language-action(VLA) modalities through an autoregressive model. Specifically, we introduce a novel VQVAE-like scene tokenizer to efficiently discretize and reconstruct semantic occupancy scenes, considering its sparsity and classes imbalance. Then, we build a unified multi-modal vocabulary for vision, language and action. Furthermore, we enhance LLM, specifically LLaMA, to perform the next token/scene prediction on the unified vocabulary to complete multiple tasks in autonomous driving. Extensive experiments demonstrate that OccLLaMA achieves competitive performance across multiple tasks, including 4D occupancy forecasting, motion planning, and visual question answering, showcasing its potential as a foundation model in autonomous driving.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What if? Emulative Simulation with World Models for Situated Reasoning

    cs.CV 2026-03 conditional novelty 6.5 of 10

    WanderDream supplies 15.8K panoramic mental-exploration videos and 158K QA pairs showing that world-model imagination measurably improves situated spatial reasoning without active exploration.

  2. Targeted Structure Completion for Sparse-View 3D Reconstruction in Autonomous Driving

    cs.CV 2026-07 conditional novelty 6.0 of 10

    FocusGS localizes a 3D geometric ambiguity manifold from depth discontinuities and instantiates continuous Gaussian queries only there, yielding SOTA sparse-view driving reconstruction with far fewer Gaussians.

  3. A Comprehensive Survey on World Models for Embodied AI

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A unified three-axis taxonomy — functionality, temporal modeling, spatial representation — organizes the world-model literature for embodied AI.

  4. COME: Adding Scene-Centric Forecasting Control to Occupancy World Model

    cs.CV 2025-06 conditional novelty 6.0 of 10

    COME adds a scene-centric forecasting branch as a ControlNet-style condition to a diffusion occupancy world model, improving static-scene consistency and beating prior methods on Occ3D-nuScenes while hiding a stronger...

  5. AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making

    cs.RO 2025-06 conditional novelty 6.0 of 10

    AntiGrounding lifts candidate robot trajectories into the VLM's visual space via multi-view rendering and structured VQA, and reports 57.5% average success across eight manipulation tasks, beating three intermediate-r...

  6. GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control

    cs.CV 2025-05 conditional novelty 6.0 of 10

    GeoDrive conditions a frozen video diffusion model on a 3D-rendered version of the requested ego trajectory, cutting trajectory-following error by 42% versus Vista while using 99.7% less training data.

  7. Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving

    cs.CV 2025-02 conditional novelty 6.0 of 10

    PreWorld introduces a two-stage semi-supervised training paradigm that achieves state-of-the-art 3D occupancy prediction and competitive 4D forecasting and planning on nuScenes using a combination of 2D and 3D supervision.

  8. GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A unified framework converts surface geometry priors into sparse Gaussian occupancy predictions and extends it to multi-view and temporal inputs.

  9. UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

    cs.CV 2026-01 conditional novelty 5.0 of 10

    A unified VLM for autonomous driving that couples trajectory planning with future-frame image generation improves open- and closed-loop planning metrics on Bench2Drive and nuScenes.

  10. Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method

    cs.CV 2025-10 conditional novelty 5.0 of 10

    UniScenev2 scales occupancy-centric driving-scene generation to NuPlan scale, releasing a 3.6M-frame semantic-occupancy dataset and jointly generating occupancy, video, and LiDAR that beats published baselines on its ...

  11. A Survey on Vision-Language-Action Models for Autonomous Driving

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A survey organizes vision-language-action models for autonomous driving into four stages, compares over 20 systems, and catalogs datasets, benchmarks, and open challenges.

  12. From 2D to 3D Cognition: A Brief Survey of General World Models

    cs.CV 2025-06 conditional novelty 3.0 of 10

    A survey proposing a two-pillar, three-capability framework that organizes recent AI world models by their transition from 2D visual prediction to 3D cognition.

  13. Generative AI for Autonomous Driving: A Review

    cs.CV 2025-05 conditional novelty 2.0 of 10

    A review of generative models (VAEs, GANs, diffusion, transformers, LLMs) applied to map generation, scenario generation, trajectory prediction, and motion planning for autonomous driving.

Pith tools