REVIEW 15 cited by
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Human drivers rely on commonsense reasoning to navigate diverse and dynamic real-world scenarios. Existing end-to-end (E2E) autonomous driving (AD) models are typically optimized to mimic driving patterns observed in data, without capturing the underlying reasoning processes. This limitation constrains their ability to handle challenging driving scenarios. To close this gap, we propose VLM-AD, a method that leverages vision-language models (VLMs) as teachers to enhance training by providing additional supervision that incorporates unstructured reasoning information and structured action labels. Such supervision enhances the model's ability to learn richer feature representations that capture the rationale behind driving patterns. Importantly, our method does not require a VLM during inference, making it practical for real-time deployment. When integrated with state-of-the-art methods, VLM-AD achieves significant improvements in planning accuracy and reduced collision rates on the nuScenes dataset. It further improves route completion and driving scores under closed-loop evaluation, demonstrating its effectiveness in long-horizon, interactive driving scenarios and its potential for safe and reliable real-world deployment.
Forward citations
Cited by 15 Pith papers
-
BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations
BEV tokens give LLMs stronger cross-view spatial reasoning than multi-view image tokens, and reverse-distilling LLM semantics into BEV encoders measurably improves closed-loop safety-critical driving.
-
PRISM: Privileged Probabilistic Latent Supervision for End-to-End Autonomous Driving Motion Planning
PRISM regularizes intermediate planning latents with a CVAE-style ELBO objective using ground-truth future paths, claiming an 8% L2 planning error reduction over deterministic baselines on nuScenes.
-
Cognitive Dual-Process Planning for Autonomous Driving with Structured Scene Knowledge and Verifiable Reasoning-Action Consistency
A dual-process VLM planner routes easy scenes to fast prediction and hard scenes to structured reasoning with rule-verified consistency rewards, reaching 80.14% planning accuracy and 97.20% LCS on a manually verified ...
-
OpenLongTail: Generative Scaling of Long-Tail Driving Data
Pose-informed diffusion with Plücker rays, depth warps, and cross-view memory converts monocular long-tail videos into multi-view assets that improve closed-loop driving robustness nearly to ground-truth multi-view levels.
-
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
DataClaw0 introduces an agentic data-tailoring paradigm, a 9B model trained on a synthetically generated dataset, and a new benchmark, claiming improved downstream adaptation in video generation, VQA, and GUI navigati...
-
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
TSHA is a new 80,000-pair benchmark for indoor safety hazard assessment; current vision-language models score roughly 45-85, and fine-tuning on TSHA raised Qwen2.5-VL-3B by 18.3 points on TSHA's test set.
-
ReMoT: Reinforcement Learning with Motion Contrast Triplets
Training a 4B vision-language model on rule-generated motion-contrast triplets with GRPO lifts spatio-temporal QA accuracy by about 17 points on the authors' own benchmark and by smaller margins on standard benchmarks.
-
AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions
AD^2-Bench is a new adverse-weather driving benchmark with hierarchical chain-of-thought annotations and LLM-based quality metrics; 12 MLLMs all scored below 60%.
-
Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding
CoT data curated by two-round LLM prompting and VLM verification, then SFT+GRPO with fine-grained rewards, improves MapDR rule–lane association F1 from 0.642 to 0.723.
-
An interactive enhanced driving dataset for autonomous driving
A fused, interaction-labeled dataset of 7.31M driving segments with synthetic BEV videos and VQA pairs for training/evaluating driving VLMs.
-
ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous Driving
ReAL-AD combines VLM-generated strategy and tactical commands with a two-stage trajectory decoder, cutting open-loop L2 error and collision rate by about a third on nuScenes and Bench2Drive.
-
Skywork-R1V3 Technical Report
A 38B open-source VLM reaches 76.0% on MMMU using RL post-training and connector-only tuning, with a critical-token entropy metric for checkpoint selection.
-
Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey
A structured survey of LLM-based trajectory prediction methods, organized into trajectory-language mapping, multimodal fusion, and constraint-based reasoning, with benchmarks, metrics, and future directions.
-
Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance
A lightweight vision-language model on an edge device fuses roadside hazard alerts with onboard camera views to adjust trajectories, and the authors report a 77% simulated collision reduction over a vision-only baseline.
-
A Survey on Vision-Language-Action Models for Autonomous Driving
A survey organizes vision-language-action models for autonomous driving into four stages, compares over 20 systems, and catalogs datasets, benchmarks, and open challenges.
Discussion (0). Sign in to comment.