Pith. sign in

REVIEW 8 cited by

V2X-VLM: End-to-End V2X Cooperative Autonomous Driving Through Large Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.09251 v3 pith:M2IAPZFB submitted 2024-08-17 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords drivingautonomousv2x-vlmcooperativeend-to-endheterogeneousinformationinfrastructure
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vehicle-to-everything (V2X) cooperation has emerged as a promising paradigm to overcome the perception limitations of classical autonomous driving by leveraging information from both ego-vehicle and infrastructure sensors. However, effectively fusing heterogeneous visual and semantic information while ensuring robust trajectory planning remains a significant challenge. This paper introduces V2X-VLM, a novel end-to-end (E2E) cooperative autonomous driving framework based on vision-language models (VLMs). V2X-VLM integrates multiperspective camera views from vehicles and infrastructure with text-based scene descriptions to enable a more comprehensive understanding of driving environments. Specifically, we propose a contrastive learning-based mechanism to reinforce the alignment of heterogeneous visual and textual characteristics, which enhances the semantic understanding of complex driving scenarios, and employ a knowledge distillation strategy to stabilize training. Experiments on a large real-world dataset demonstrate that V2X-VLM achieves state-of-the-art trajectory planning accuracy, significantly reducing L2 error and collision rate compared to existing cooperative autonomous driving baselines. Ablation studies validate the contributions of each component. Moreover, the evaluation of robustness and efficiency highlights the practicality of V2X-VLM for real-world deployment to enhance overall autonomous driving safety and decision-making.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding

    cs.CV 2025-08 conditional novelty 7.0 of 10

    A text-conditioned U-Net can generate input-aware triggers that backdoor VLM visual grounding, forcing the model to output the attacker-chosen object's bounding box regardless of the user query.

  2. Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    BadSem shows that semantic mismatches between images and text can serve as stealthy backdoor triggers for VLMs, achieving near-perfect attack success with low poisoning rates.

  3. TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios

    cs.CV 2026-03 unverdicted novelty 6.0 of 10

    TSHA is a new 80,000-pair benchmark for indoor safety hazard assessment; current vision-language models score roughly 45-85, and fine-tuning on TSHA raised Qwen2.5-VL-3B by 18.3 points on TSHA's test set.

  4. Spatial-aware Vision Language Model for Autonomous Driving

    cs.CV 2025-12 conditional novelty 6.0 of 10

    LVLDrive fuses LiDAR point clouds into a pre-trained vision-language driving model through a zero-initialized gated Q-Former and a spatial-aware QA dataset, improving open-loop planning and 3D grounding over image-onl...

  5. AirV2X: Unified Air-Ground Vehicle-to-Everything Collaboration

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AirV2X-Perception is a 6.73-hour simulated dataset and benchmark for collaborative perception with up to 5 vehicles, 5 roadside units, and 5 drones.

  6. Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance

    cs.AI 2025-08 reject novelty 4.0 of 10

    A lightweight vision-language model on an edge device fuses roadside hazard alerts with onboard camera views to adjust trajectories, and the authors report a 77% simulated collision reduction over a vision-only baseline.

  7. Research Challenges and Progress in the End-to-End V2X Cooperative Autonomous Driving Competition

    cs.RO 2025-07 conditional novelty 3.0 of 10

    This paper summarizes the CVPR 2025 V2X cooperative driving challenge, its winning solutions, and the open research problems it reveals.

  8. Automated Vehicles Should be Connected with Natural Language

    cs.MA 2025-06 conditional novelty 3.0 of 10

    A vision paper recommending natural language as the universal communication medium for connected and automated vehicles.

Pith tools