Pith. sign in

REVIEW 5 cited by

Distilling Multi-modal Large Language Models for Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.09757 v1 pith:TYMIINE6 submitted 2025-01-16 cs.CV cs.RO

classification cs.CVcs.RO
keywords autonomousdimadrivingplanningend-to-endplannerreductionvision-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Autonomous driving demands safe motion planning, especially in critical "long-tail" scenarios. Recent end-to-end autonomous driving systems leverage large language models (LLMs) as planners to improve generalizability to rare events. However, using LLMs at test time introduces high computational costs. To address this, we propose DiMA, an end-to-end autonomous driving system that maintains the efficiency of an LLM-free (or vision-based) planner while leveraging the world knowledge of an LLM. DiMA distills the information from a multi-modal LLM to a vision-based end-to-end planner through a set of specially designed surrogate tasks. Under a joint training strategy, a scene encoder common to both networks produces structured representations that are semantically grounded as well as aligned to the final planning objective. Notably, the LLM is optional at inference, enabling robust planning without compromising on efficiency. Training with DiMA results in a 37% reduction in the L2 trajectory error and an 80% reduction in the collision rate of the vision-based planner, as well as a 44% trajectory error reduction in longtail scenarios. DiMA also achieves state-of-the-art performance on the nuScenes planning benchmark.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AutoNeural: Co-Designing Vision-Language Models for NPU Inference

    cs.CL 2025-12 conditional novelty 6.0 of 10

    A NPU-native VLM combining a MobileNet-style encoder with a hybrid Transformer-SSM backbone claims 14x lower latency and 7x lower quantization error over ViT-Transformer baselines, though quantized accuracy is not reported.

  2. Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs

    cs.AI 2025-07 conditional novelty 5.0 of 10

    The survey's L1/L2 taxonomy and benchmark show that current reasoning models waste compute on easy problems and underthink hard ones, motivating more adaptive inference.

  3. RoCA: Robust Cross-Domain End-to-End Autonomous Driving

    cs.CV 2025-06 conditional novelty 5.0 of 10

    RoCA, a Gaussian-process codebook over ego and agent tokens, improves cross-domain generalization and adaptation of end-to-end autonomous driving models without extra inference cost.

  4. CogAD: Cognitive-Hierarchy Guided End-to-End Autonomous Driving

    cs.RO 2025-05 conditional novelty 5.0 of 10

    CogAD reports state-of-the-art open-loop and closed-loop planning results by combining hierarchical scene-to-instance perception with intent-to-trajectory planning and dual-level uncertainty.

  5. ALN-P3: Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving

    cs.CV 2025-05 conditional novelty 5.0 of 10

    ALN-P3 adds three alignment losses between a driving stack and a language model during training, improving both planning safety and language reasoning on nuScenes, Nu-X, TOD3Cap, and nuScenes-QA.

Pith tools