REVIEW 14 cited by
LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Existing learning-based autonomous driving (AD) systems face challenges in comprehending high-level information, generalizing to rare events, and providing interpretability. To address these problems, this work employs Large Language Models (LLMs) as a decision-making component for complex AD scenarios that require human commonsense understanding. We devise cognitive pathways to enable comprehensive reasoning with LLMs, and develop algorithms for translating LLM decisions into actionable driving commands. Through this approach, LLM decisions are seamlessly integrated with low-level controllers by guided parameter matrix adaptation. Extensive experiments demonstrate that our proposed method not only consistently surpasses baseline approaches in single-vehicle tasks, but also helps handle complex driving behaviors even multi-vehicle coordination, thanks to the commonsense reasoning capabilities of LLMs. This paper presents an initial step toward leveraging LLMs as effective decision-makers for intricate AD scenarios in terms of safety, efficiency, generalizability, and interoperability. We aspire for it to serve as inspiration for future research in this field. Project page: https://sites.google.com/view/llm-mpc
Forward citations
Cited by 14 Pith papers
-
S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation
S4-Driver uses a multimodal LLM with a sparse 3D spatio-temporal volume representation to achieve self-supervised motion planning that rivals supervised methods on nuScenes and WOMD.
-
SanDRA: Safe Large-Language-Model-Based Decision Making for Automated Vehicles Using Reachability Analysis
LLM-generated driving actions are translated into temporal-logic formulas and passed through reachability analysis, permitting only actions with non-empty safe reachable sets to be executed.
-
DriveQA: Passing the Driving Knowledge Test
DriveQA is a new multimodal driving-knowledge benchmark showing that LLMs and MLLMs struggle with right-of-way, numerical traffic rules, and sign variations, with modest transfer gains to nuScenes and BDD.
-
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
A skill-by-skill mixture-of-experts router lets a sub-3B vision-language model beat much larger models on autonomous-driving and robot-reasoning benchmarks.
-
Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
A new 80K-clip dataset of unstructured driving scenarios with Q&A annotations improves VLA performance on NeuroNCAP and nuScenes benchmarks.
-
CoopReflect: Towards Natural Language Communication for Cooperative Autonomous Driving via Multi-Agent Learning
Post-episode multi-agent debriefing lets LLM driving agents learn concise natural-language coordination protocols that avoid collisions and merge traffic, and distillation makes the policy fast enough for near-real-time use.
-
ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous Driving
ReAL-AD combines VLM-generated strategy and tactical commands with a two-stage trajectory decoder, cutting open-loop L2 error and collision rate by about a third on nuScenes and Bench2Drive.
-
LeAD: The LLM Enhanced Planning System Converged with End-to-end Autonomous Driving
LeAD adds a low-frequency large-language-model planner that takes over when a high-frequency end-to-end driving model gets stuck, and reports improved CARLA benchmark scores.
-
Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models
A fine-tuned BLIP model with hierarchical question skipping achieves 423 ms average inference for driving scene description, with GPT-evaluated scores of 65 to 79, near GPT-4o's 77.
-
Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects
A survey that classifies chain-of-thought methods for autonomous driving into modular, logical, and reflective pipelines, and proposes three evolutionary stages from direct prompting to reinforcement learning.
-
TinyDrive: Multiscale Visual Question Answering with Selective Token Routing for Autonomous Driving
A tiny CNN+T5 model claims state-of-the-art BLEU-4/METEOR on DriveLM but loses on ROUGE-L/CIDEr, with no code or ablations.
-
LimSim Series: An Autonomous Driving Simulation Platform for Validation and Enhancement
The LimSim Series is an open-source closed-loop simulation platform that integrates multiple driving-system pipelines, an Area-of-Interest efficiency mechanism, and a multi-metric evaluation suite.
-
A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach
A position/review paper argues data-driven model predictive control is the best route to safe, adaptive, human-like autonomous-driving motion planning, but provides no new derivation or experiment.
-
Generative AI for Autonomous Driving: A Review
A review of generative models (VAEs, GANs, diffusion, transformers, LLMs) applied to map generation, scenario generation, trajectory prediction, and motion planning for autonomous driving.
Discussion (0). Continue with ORCID to comment.