Pith. sign in

REVIEW 9 cited by

DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.18607 v1 pith:ILGD3JUL submitted 2024-12-24 cs.CV

classification cs.CV
keywords planningdrivingmodelingworlddrivinggptactionautoregressivegeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

World model-based searching and planning are widely recognized as a promising path toward human-level physical intelligence. However, current driving world models primarily rely on video diffusion models, which specialize in visual generation but lack the flexibility to incorporate other modalities like action. In contrast, autoregressive transformers have demonstrated exceptional capability in modeling multimodal data. Our work aims to unify both driving model simulation and trajectory planning into a single sequence modeling problem. We introduce a multimodal driving language based on interleaved image and action tokens, and develop DrivingGPT to learn joint world modeling and planning through standard next-token prediction. Our DrivingGPT demonstrates strong performance in both action-conditioned video generation and end-to-end planning, outperforming strong baselines on large-scale nuPlan and NAVSIM benchmarks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations

    cs.CV 2026-03 conditional novelty 6.5 of 10

    BEV tokens give LLMs stronger cross-view spatial reasoning than multi-view image tokens, and reverse-distilling LLM semantics into BEV encoders measurably improves closed-loop safety-critical driving.

  2. G2DP: Diffusion Planning with Spatio-Temporal Grid Guidance

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    G2DP constructs a differentiable spatio-temporal cost volume from occupancy and route maps to guide diffusion denoising for collision-free trajectories, reporting SOTA closed-loop scores on nuPlan.

  3. WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A five-aspect, 24-metric benchmark, a 26K human-annotated dataset, and an AI evaluator show that today's driving world models cannot simultaneously look real, respect geometry, and behave safely.

  4. OmniNWM: Omniscient Driving Navigation World Models

    cs.CV 2025-10 conditional novelty 6.0 of 10

    OmniNWM jointly generates long panoramic multi-modal driving videos, controls them precisely via normalized Plücker ray-maps, and derives dense driving rewards from generated 3D occupancy.

  5. A Comprehensive Survey on World Models for Embodied AI

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A unified three-axis taxonomy — functionality, temporal modeling, spatial representation — organizes the world-model literature for embodied AI.

  6. Epona: Autoregressive Diffusion World Model for Autonomous Driving

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An autoregressive diffusion world model jointly generates the next camera frame and a multi-step trajectory, enabling long videos and real-time planning for autonomous driving.

  7. AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AD^2-Bench is a new adverse-weather driving benchmark with hierarchical chain-of-thought annotations and LLM-based quality metrics; 12 MLLMs all scored below 60%.

  8. 3D and 4D World Modeling: A Survey

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A survey that defines 3D/4D world modeling, organizes methods into VideoGen, OccGen, and LiDARGen categories, and compiles datasets, metrics, and benchmark numbers.

  9. Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey

    cs.RO 2025-09 conditional novelty 4.0 of 10

    A structured survey of LLM-based trajectory prediction methods, organized into trajectory-language mapping, multimodal fusion, and constraint-based reasoning, with benchmarks, metrics, and future directions.

Pith tools