REVIEW 12 cited by
DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Research interest in end-to-end autonomous driving has surged owing to its fully differentiable design integrating modular tasks, i.e. perception, prediction and planing, which enables optimization in pursuit of the ultimate goal. Despite the great potential of the end-to-end paradigm, existing methods suffer from several aspects including expensive BEV (bird's eye view) computation, action diversity, and sub-optimal decision in complex real-world scenarios. To address these challenges, we propose a novel hybrid sparse-dense diffusion policy, empowered by a Vision-Language Model (VLM), called Diff-VLA. We explore the sparse diffusion representation for efficient multi-modal driving behavior. Moreover, we rethink the effectiveness of VLM driving decision and improve the trajectory generation guidance through deep interaction across agent, map instances and VLM output. Our method shows superior performance in Autonomous Grand Challenge 2025 which contains challenging real and reactive synthetic scenarios. Our methods achieves 45.0 PDMS.
Forward citations
Cited by 12 Pith papers
-
MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving
Block-wise Modal Joint Attention over image, LiDAR, and diffusion action tokens yields 88.9 PDMS / 88.4 EPDMS on NAVSIM without anchors or auxiliary supervision.
-
WorkDrive: Roadwork Chain of Causation for Autonomous Driving
Perception-grounded Chain-of-Causation labels plus a consistency-only GRPO reward lower ROADWork trajectory ADE@15 from 12.10 to 10.68 pixels (about -12%) over a trajectory-only SFT baseline.
-
S-squared-VLA: Decoupling Semantic and Spatial Streams in Vision-Language-Action Models for Autonomous Driving
S2-VLA decouples semantic and spatial streams in a vision-language-action driving model, reaching PDMS 87.1 and NC 98.4 on NAVSIM under supervised fine-tuning.
-
OpenLongTail: Generative Scaling of Long-Tail Driving Data
Pose-informed diffusion with Plücker rays, depth warps, and cross-view memory converts monocular long-tail videos into multi-view assets that improve closed-loop driving robustness nearly to ground-truth multi-view levels.
-
WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving
WCog-VLA couples Game-CoT semantic reasoning with an aligned decoupled diffusion transformer to generate joint multi-agent trajectories and reaches 92.9 PDMS on NAVSIM.
-
AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning
Trajectory-pattern anchors bridge VLA reasoning and continuous residual flow, yielding 77.28% success rate on Bench2Drive closed-loop driving.
-
TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
Feeding torque history as a single decoder token and adding torque prediction as an auxiliary objective improves pretrained VLA success rates on contact-rich manipulation, with large gains on button pushing and charge...
-
IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model
IRL-VLA fine-tunes a vision-language-action driving policy with PPO against a learned reward world model trained on NAVSIM's EPDMS metrics, reaching 74.9 EPDMS on navhard-real.
-
HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving
Hierarchical diffusion plus polar structure-preserving expansion and metric-decoupled RL yields SOTA open- and closed-loop planning scores on NAVSIM and HUGSIM.
-
NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning
A reasoning-free VLA trained with <60% of typical driving data reaches near-state-of-the-art trajectory scores when fine-tuned with Dr. GRPO instead of GRPO.
-
RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
A dual-loop VLM plus VLA system uses visual prompts and closed-loop monitoring to perform chemistry lab manipulations with reported gains in success and safety-compliance over baseline robot policies.
-
A Survey on Vision-Language-Action Models for Autonomous Driving
A survey organizes vision-language-action models for autonomous driving into four stages, compares over 20 systems, and catalogs datasets, benchmarks, and open challenges.
Discussion (0). Sign in to comment.