Pith. sign in

REVIEW 2 cited by

SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.09594 v1 pith:KT5A2JBD submitted 2025-03-12 cs.CV cs.RO

classification cs.CVcs.RO
keywords drivingmodelunderstandingautonomouslanguageperformancesimlingovision-language
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Integrating large language models (LLMs) into autonomous driving has attracted significant attention with the hope of improving generalization and explainability. However, existing methods often focus on either driving or vision-language understanding but achieving both high driving performance and extensive language understanding remains challenging. In addition, the dominant approach to tackle vision-language understanding is using visual question answering. However, for autonomous driving, this is only useful if it is aligned with the action space. Otherwise, the model's answers could be inconsistent with its behavior. Therefore, we propose a model that can handle three different tasks: (1) closed-loop driving, (2) vision-language understanding, and (3) language-action alignment. Our model SimLingo is based on a vision language model (VLM) and works using only camera, excluding expensive sensors like LiDAR. SimLingo obtains state-of-the-art performance on the widely used CARLA simulator on the Bench2Drive benchmark and is the winning entry at the CARLA challenge 2024. Additionally, we achieve strong results in a wide variety of language-related tasks while maintaining high driving performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generalized Trajectory Scoring for End-to-end Multimodal Planning

    cs.RO 2025-06 conditional novelty 5.0 of 10

    GTRS combines super-dense vocabulary training, dropout, sensor augmentation, and diffusion proposals to reach 49.4 EPDMS on the Navhard benchmark, approaching the privileged PDM-Closed method.

  2. ALN-P3: Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving

    cs.CV 2025-05 conditional novelty 5.0 of 10

    ALN-P3 adds three alignment losses between a driving stack and a language model during training, improving both planning safety and language reasoning on nuScenes, Nu-X, TOD3Cap, and nuScenes-QA.

Pith tools