REVIEW 6 cited by
Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Rapid advancements in Autonomous Driving (AD) tasks turned a significant shift toward end-to-end fashion, particularly in the utilization of vision-language models (VLMs) that integrate robust logical reasoning and cognitive abilities to enable comprehensive end-to-end planning. However, these VLM-based approaches tend to integrate 2D vision tokenizers and a large language model (LLM) for ego-car planning, which lack 3D geometric priors as a cornerstone of reliable planning. Naturally, this observation raises a critical concern: Can a 2D-tokenized LLM accurately perceive the 3D environment? Our evaluation of current VLM-based methods across 3D object detection, vectorized map construction, and environmental caption suggests that the answer is, unfortunately, NO. In other words, 2D-tokenized LLM fails to provide reliable autonomous driving. In response, we introduce DETR-style 3D perceptrons as 3D tokenizers, which connect LLM with a one-layer linear projector. This simple yet elegant strategy, termed Atlas, harnesses the inherent priors of the 3D physical world, enabling it to simultaneously process high-resolution multi-view images and employ spatiotemporal modeling. Despite its simplicity, Atlas demonstrates superior performance in both 3D detection and ego planning tasks on nuScenes dataset, proving that 3D-tokenized LLM is the key to reliable autonomous driving. The code and datasets will be released.
Forward citations
Cited by 6 Pith papers
-
ReMoT: Reinforcement Learning with Motion Contrast Triplets
Training a 4B vision-language model on rule-generated motion-contrast triplets with GRPO lifts spatio-temporal QA accuracy by about 17 points on the authors' own benchmark and by smaller margins on standard benchmarks.
-
From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving
A VLM-based and a vision-only end-to-end planner are behaviorally complementary in a long tail of driving scenarios; selecting the better trajectory lifts NAVSIM PDMS from 90.80 to 92.10 at modest compute.
-
Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving
A 3D-vision-language pre-training model with group-wise contrastive alignment generates driving trajectories as text and reports state-of-the-art open-loop planning results on nuScenes.
-
Grid: Omni Visual Generation
GRID shows that fine-tuning an image diffusion model on videos arranged as grid images can generate coherent video and multi-view sequences with far less data and compute than specialized video models.
-
FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training
Frozen LLM text encoders, combined with multi-prompt hidden-state extraction and cached embeddings, make CLIP-style pre-training data-efficient, long-context aware, and multilingual.
-
World knowledge-enhanced Reasoning Using Instruction-guided Interactor in Autonomous Driving
An instruction-guided token selection and cross-attention module improves MLLM performance on autonomous driving QA and planning benchmarks, trained with a new GPT-generated object-level risk assessment dataset.
Discussion (0). Continue with ORCID to comment.