Pith. sign in

REVIEW 12 cited by

DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.19381 v4 pith:H6WEYW6U submitted 2025-05-26 cs.AI cs.CVcs.RO

classification cs.AIcs.CVcs.RO
keywords drivingautonomousdiffusiondecisionend-to-endmethodsscenariosvision-language
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Research interest in end-to-end autonomous driving has surged owing to its fully differentiable design integrating modular tasks, i.e. perception, prediction and planing, which enables optimization in pursuit of the ultimate goal. Despite the great potential of the end-to-end paradigm, existing methods suffer from several aspects including expensive BEV (bird's eye view) computation, action diversity, and sub-optimal decision in complex real-world scenarios. To address these challenges, we propose a novel hybrid sparse-dense diffusion policy, empowered by a Vision-Language Model (VLM), called Diff-VLA. We explore the sparse diffusion representation for efficient multi-modal driving behavior. Moreover, we rethink the effectiveness of VLM driving decision and improve the trajectory generation guidance through deep interaction across agent, map instances and VLM output. Our method shows superior performance in Autonomous Grand Challenge 2025 which contains challenging real and reactive synthetic scenarios. Our methods achieves 45.0 PDMS.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Block-wise Modal Joint Attention over image, LiDAR, and diffusion action tokens yields 88.9 PDMS / 88.4 EPDMS on NAVSIM without anchors or auxiliary supervision.

  2. WorkDrive: Roadwork Chain of Causation for Autonomous Driving

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Perception-grounded Chain-of-Causation labels plus a consistency-only GRPO reward lower ROADWork trajectory ADE@15 from 12.10 to 10.68 pixels (about -12%) over a trajectory-only SFT baseline.

  3. S-squared-VLA: Decoupling Semantic and Spatial Streams in Vision-Language-Action Models for Autonomous Driving

    cs.RO 2026-07 conditional novelty 6.0 of 10

    S2-VLA decouples semantic and spatial streams in a vision-language-action driving model, reaching PDMS 87.1 and NC 98.4 on NAVSIM under supervised fine-tuning.

  4. OpenLongTail: Generative Scaling of Long-Tail Driving Data

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Pose-informed diffusion with Plücker rays, depth warps, and cross-view memory converts monocular long-tail videos into multi-view assets that improve closed-loop driving robustness nearly to ground-truth multi-view levels.

  5. WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving

    cs.CV 2026-07 conditional novelty 6.0 of 10

    WCog-VLA couples Game-CoT semantic reasoning with an aligned decoupled diffusion transformer to generate joint multi-agent trajectories and reaches 92.9 PDMS on NAVSIM.

  6. AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Trajectory-pattern anchors bridge VLA reasoning and continuous residual flow, yielding 77.28% success rate on Bench2Drive closed-loop driving.

  7. TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models

    cs.RO 2025-09 conditional novelty 6.0 of 10

    Feeding torque history as a single decoder token and adding torque prediction as an auxiliary objective improves pretrained VLA success rates on contact-rich manipulation, with large gains on button pushing and charge...

  8. IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model

    cs.AI 2025-08 conditional novelty 6.0 of 10

    IRL-VLA fine-tunes a vision-language-action driving policy with PPO against a learned reward world model trained on NAVSIM's EPDMS metrics, reaching 74.9 EPDMS on navhard-real.

  9. HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving

    cs.RO 2026-04 conditional novelty 5.5 of 10

    Hierarchical diffusion plus polar structure-preserving expansion and metric-decoupled RL yields SOTA open- and closed-loop planning scores on NAVSIM and HUGSIM.

  10. NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning

    cs.AI 2026-02 conditional novelty 5.0 of 10

    A reasoning-free VLA trained with <60% of typical driving data reaches near-state-of-the-art trajectory scores when fine-tuned with Dr. GRPO instead of GRPO.

  11. RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation

    cs.RO 2025-09 conditional novelty 5.0 of 10

    A dual-loop VLM plus VLA system uses visual prompts and closed-loop monitoring to perform chemistry lab manipulations with reported gains in success and safety-compliance over baseline robot policies.

  12. A Survey on Vision-Language-Action Models for Autonomous Driving

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A survey organizes vision-language-action models for autonomous driving into four stages, compares over 20 systems, and catalogs datasets, benchmarks, and open challenges.

Pith tools