REVIEW 3 major objections 5 minor 80 references
Large Multimodal Agents for Intelligent Transportation Systems: Architectures, Evidence, and Deployment Challenges
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A 42-family evidence map concludes that large multimodal agents are best used as orchestrators that interpret and coordinate transportation evidence, while forecasting, optimization, control, and final authority remain with specialist…
desk verdict A careful, transparent evidence map that makes a plausible bounded-orchestration case for LMAs in ITS, with the main caveat being that the coding reliability behind its headline counts is not yet independently auditable from the preprint alone. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the review's orthogonal coding framework applied to 42 verified study families: a C0–C3 functional-capability scale (from pre-agentic capability to outcome-responsive agency), an E0–E4 validation-setting scale (from conceptual to sustained deployment), three evidence propositions P1–P3 (transportation semantics, evidence reconciliation, multidimensional integration), and eight noncompensatory methodological-concern domains Q1–Q8. The framework's work is to separate what a system demonstrably does from where it was evaluated and from whether a claimed benefit is directly supported, so that agency, multimodality, and architectural complexity cannot by themselves be counted as evidence. The synthesis database built from these codings is the analytical source of truth behind the counts of 14 C3 families, one E3 family, and no E4 family.
What would settle it
An independent re-coding of the same 42 families using the published rubric—especially a check of whether any family actually completes the full P2 chain (provenance, challenge, handling, matched comparison, attributable outcome) and whether all 14 C3 labels show feedback changing a later foundation-model decision—would settle the claim; if the counts moved materially, the bounded-orchestration conclusion would need revision.
Extended reading notes
Core claim
The paper establishes an auditable evidence map of 42 primary study families and shows that direct evidence is concentrated in two of its three testable propositions. Twenty-three families directly evaluate transportation semantics (P1) and 24 directly evaluate multidimensional integration (P3), with 19 evaluating both; all direct P1 or P3 judgments use strong or moderate claim-matched comparisons. No family directly evaluates the full evidence-reconciliation chain (P2), in which traceable provenance, a missingness or conflict challenge, a handling mechanism, a matched comparison, and an attributable outcome all appear. On the capability scale, 14 families reach C3, meaning observed transportation outcomes change a later foundation-model decision, but validation lags: 13 sit at E2 interactive simulation, one reaches E3 controlled real-world operation, and none reaches E4 sustained deployment. The review concludes that LMAs are best supported for semantic interpretation, intent translation, evidence organisation, scenario authoring, explanation, and specialist-tool coordination, and that this supports contribution-specific orchestration, not general replacement.
Load-bearing premise
The whole conclusion rests on the accuracy of the review's own coding of 42 study families; if the C0–C3, E0–E4, P1–P3, and Q1–Q8 labels were applied inconsistently or the frozen database misrecorded the studies, the counts that drive the bounded-orchestration conclusion could change.
Editorial extensions
If this is right
- Near-term ITS-LMA deployments should be read-only or advisory, with source attribution, explicit uncertainty, and accountable human review.
- Larger action authority should be granted only after direct P2 evidence, robustness and failure analysis, controlled E3 trials, and longitudinal E4 evidence within a declared operational envelope.
- Evaluation of future systems should use matched configurations: specialist alone, LMA without tools, LMA orchestrating the specialist, and the complete architecture with permission gateway, independent monitor, fallback, and rollback.
- Numerical forecasting, optimization, simulation fidelity, hard constraints, low-level control, safety fallback, and final authority should remain with specialist systems or accountable humans until stronger claim-matched evidence appears.
- The P2 gap means provenance and conflict handling is a required next evaluation target, not an optional feature.
Reading between the lines
- A natural testable extension is to turn the P2 chain into a checklist for new evaluations: any system claiming evidence reconciliation should report provenance records, a missingness or conflict challenge, a handling mechanism, a matched comparison, and an attributable outcome.
- A concrete test of the bounded-orchestration claim would be a benchmark task built from the paper's suggested agentic-refinement route: generate, execute, evaluate, and iteratively refine a specialist model's code on fixed datasets, with human approval gates and failure logs recorded.
- The concentration of evidence in autonomous driving and planning/simulation suggests that the semantic layer may generalize more readily to incident interpretation and scenario authoring than to numerical forecasting, which should be tested explicitly across cities and event conditions.
- A living evidence registry that records search decisions, study-family links, prompts, model versions, and failure logs would let future updates test whether the 14-C3, one-E3, no-E4 profile changes as more controlled real-world evaluations appear.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a structured narrative review and evidence map of 42 primary study families of large multimodal agents (LMAs) for intelligent transportation systems (ITS), covering sources released between January 2023 and 3 August 2026. The authors propose an analytical framework that separates model-level, system-level, and hybrid multimodality; defines capability levels C0–C3 and validation settings E0–E4; and evaluates three evidence propositions (P1 transportation semantics, P2 evidence reconciliation, P3 multidimensional integration) independently of methodological concerns (Q1–Q8). The central findings are that 14 families reach C3, only one reaches E3, none reaches E4, and P2 remains unresolved because no family demonstrates the complete provenance–challenge–handling–comparison–outcome chain. The paper concludes that LMAs are best suited for bounded orchestration—semantic interpretation, evidence organization, and specialist-tool coordination—rather than replacement of specialist forecasting, optimization, simulation, and control systems, and it proposes a matched comparative evaluation protocol and a staged deployment roadmap.
Significance. If the evidence map is reliable, the paper offers a useful contribution by operationalizing distinctions that prior surveys leave implicit: capability versus validation setting, proposition directness versus result direction, and LMA authority versus specialist authority. The explicit definitions of C0–C3 and E0–E4, the orthogonal coding of P1–P3 and Q1–Q8, and the proposed four-configuration comparative protocol are valuable for future evaluations. The authors are transparent about limitations, including the lack of a PRISMA flow, the absence of inter-rater reliability statistics, and the reliance on a single-author coding process. The main risk is that the headline counts and the P2-unresolved conclusion are not independently verifiable from the preprint, which undermines the auditability that the paper claims as a central feature.
major comments (3)
- [Section II-B and Section VII] The central quantitative claims (14 C3, one E3, no E4, P2 unresolved, all direct P1/P3 in strong or moderate comparison groups) rest entirely on the consistency of the author-defined coding framework. The paper explicitly states that all pilot and repeat procedures were within-process stability checks, not independent duplicate human coding, and that no inter-rater reliability statistic is claimed. The label-masked repeat audits (40/42 C, 37/42 E, 280/336 Q, 113/126 P) are reported only as aggregate counts, with no item-level disagreements. The canonical synthesis database, audit register, and Supplementary Tables S5–S8 are referenced but not included in the preprint. A reader therefore cannot verify that the 42-family cross-tabulations in Table IV and the P2-unresolved conclusion follow from the underlying studies. For an 'auditable evidence map,' this is a load-bearing reproducibility gap.
- [Section VII] The paper claims that 'prespecified conservative sensitivity analyses did not alter the central bounded-orchestration conclusion' and that full results are provided in Supplementary Table S6, but this table is not part of the preprint. Because the conclusion depends on coding thresholds such as the strict P2-D criterion, the E3/E4 boundary, and the lower-code rule, the reader cannot assess how robust the headline counts are to plausible coding variations. Please provide the sensitivity results, or make them available in the repository with a clear pointer in the manuscript.
- [Abstract and Section VI-A] The claim that 'all direct P1 or P3 judgments use strong or moderate claim-matched comparisons' is central to the conclusion that direct evidence is strongest for semantics and integration. However, the family-level mapping between proposition directness and comparator strength is not reported; Table IV only gives domain-level aggregates, and 14 families are described as having partial or nonisolating comparisons. Without a family-level table showing which families receive D codes and what comparator strength they have, this claim cannot be checked. This mapping should be included in the supplementary material.
minor comments (5)
- [Section III-C, paragraph before Table I] The sentence 'C1 rule. After confirming a substantive transportation role, C denotes actionable decision-support...' is garbled; it should likely read 'C1 applies when the foundation model participates in an actionable decision-support cycle with implementation retained externally.'
- [Section IV] The phrase 'Table III summaries these differences' should be 'Table III summarizes these differences.'
- [Section VI-A] The sentence 'Only two families wereevaluated primarily in real-world or onboard settings' contains a typo: 'wereevaluated' should be 'were evaluated.'
- [Section II-A] The paper refers to Supplementary Tables S2, S5, S6, S7, and S8, but the preprint does not include them and the body text does not state where they can be accessed beyond the abstract's GitHub URL. Please add a clear pointer in the body to the repository or supplementary file location.
- [Section II-B] The disclosure that 'structured model assistance' was used for literature retrieval, metadata checking, and extraction drafting would be more reproducible if the specific models and versions were stated, or if the prompts and protocols were archived with the other review materials.
Circularity Check
No significant circularity: the review's conclusions are coding-based summaries of external studies, not predictions derived from fitted inputs or self-citation chains.
full rationale
This manuscript is a structured narrative review and evidence map rather than a formal derivation, so the classical circularity failure modes (fitted parameter renamed as prediction, self-definitional derivation, or self-citation chain forcing a conclusion) do not arise. The central claims—23 families with direct P1 evidence, 24 with direct P3, P2 unresolved, 14 C3 families, one E3 family, and no E4 family, and hence bounded orchestration—are counts and interpretations produced by applying an explicit, author-defined coding framework to 42 independently published study families. The P2 conclusion is definitionally tied to the P2-D criterion (a family must demonstrate the full provenance–challenge–handling–comparison–outcome chain to be coded D), but this is a transparent coding report rather than a hidden reduction: no quantity is fitted to one subset and then 'predicted' on a related subset. The paper's own limitations (single-author-led coding, no inter-rater reliability statistic, Supplementary Table S6 not included in the preprint) are validity and auditability risks, not circularity; they concern whether the codes are correct, not whether the conclusion reduces to its inputs. The cited prior work is external to the authors' framework and does not carry a load-bearing uniqueness or ansatz argument. Accordingly, no circular step can be exhibited, and the appropriate score is 0.
Assumptions & free parameters
assumptions (6)
- domain assumption The operational definition of ITS-LMA and the inclusion boundary determine which studies enter the evidence map (Section III-B).
- domain assumption The C0-C3 capability scale and E0-E4 validation scale are ordinal categories defined by the authors (Section III-C).
- domain assumption The P1-P3 proposition evidence criteria, particularly the strict P2 provenance-challenge-handling-comparison-outcome chain, define what counts as direct evidence (Sections II-B and III-A).
- domain assumption The search window (January 2023 to 3 August 2026) and reliance on publicly accessible English-language reports bound the corpus (Section II-A and Section VII).
- domain assumption The five application domains are mutually exclusive groupings by principal evaluated LMA function (Section V).
- domain assumption Label-masked repeat audits are within-process stability checks, not independent duplicate human coding; no inter-rater reliability statistic is claimed (Section II-B and Section VII).
Cite this review
Pith. "Pith review of Large Multimodal Agents for Intelligent Transportation Systems: Architectures, Evidence, and Deployment Challenges." pith.science (2026). https://pith.science/paper/3QPHDLA3
@misc{pith2026260808184,
author = {Pith},
title = {Pith review of: Large Multimodal Agents for Intelligent Transportation Systems: Architectures, Evidence, and Deployment Challenges},
year = {2026},
howpublished = {\url{https://pith.science/paper/3QPHDLA3}},
note = {Machine review of arXiv:2608.08184}
}
read the original abstract
Large multimodal agents (LMAs) are increasingly proposed for intelligent transportation systems (ITS), but existing studies often conflate multimodality, agency, empirical performance, and deployment readiness. This review provides an auditable evidence map of 42 primary study families released between January 2023 and 3 August 2026 within a corpus of 91 mapped sources. It distinguishes model-level, system-level, and hybrid multimodality and classifies each family by system architecture and action authority. Evidence is assessed independently through functional capability (C0-C3), validation setting (E0-E4), three evidence propositions (P1-P3), and eight methodological-concern domains (Q1-Q8). Transportation semantics (P1) are directly evaluated in 23 families and multidimensional integration (P3) in 24; 19 families directly evaluate both. Evidence reconciliation (P2) remains unresolved because no family demonstrates the complete provenance-challenge-handling-comparison-outcome chain. Fourteen families reach C3, but 13 remain at E2; only one reaches E3 and none reaches E4. Across ITS domains, LMAs are best supported for semantic interpretation, intent translation, evidence organisation, scenario authoring, explanation, and specialist-tool coordination. Numerical forecasting, optimisation, simulation fidelity, hard constraints, low-level control, safety fallback, and final authority should remain with independently verifiable specialist systems or accountable humans. The review therefore supports bounded orchestration rather than replacement and provides a matched comparative evaluation protocol and staged roadmap for accountable deployment. The living evidence repository is available at https://github.com/pangjunbiao/ITS-LMA-Review.
Figures
Reference graph
Works this paper leans on
-
[1]
The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews,
M. J. Page et al., “The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews,” BMJ, vol. 372, art. n71, 2021, doi: 10.1136/bmj.n71
doi:10.1136/bmj.n71 2020
-
[2]
SAE International, “Taxonomy and Definitions for Terms Re- lated to Driving Automation Systems for On-Road Motor Vehicles,” SAE J3016_202104, 2021
work page 2021
-
[3]
TRIP: Transport Reasoning With Intelligence Progression—A Foundation Framework,
Z. Liu et al., “TRIP: Transport Reasoning With Intelligence Progression—A Foundation Framework,” Transportation Re- search Part C: Emerging Technologies, vol. 179, art. 105260, 2025, doi: 10.1016/j.trc.2025.105260
arXiv 2025
-
[4]
Large multimodal agents: A survey,
J. Xie, Z. Chen, R. Zhang, and G. Li, “Large multimodal agents: A survey,”Visual Intelligence, vol. 3, art. 24, 2025, doi: 10.1007/s44267-025-00093-y
-
[5]
A Survey on Multimodal Large Language Models for Autonomous Driving,
C. Cui et al., “A Survey on Multimodal Large Language Models for Autonomous Driving,” in Proc. IEEE/CVF WACV Workshops, 2024, pp. 958–979, doi: 10.1109/WACVW60836.2024.00106
arXiv 2024
-
[7]
S. Wandelt, C. Zheng, S. Wang, Y. Liu, and X. Sun, “Large Language Models for Intelligent Transportation: A Review of the State of the Art and Challenges,” Applied Sciences, vol. 14, no. 17, art. 7455, 2024, doi: 10.3390/app14177455
-
[8]
Vision language models in autonomous driving: A survey and outlook,
X. Zhou, M. Liu, E. Yurtsever, B. L. Zagar, W. Zimmer, H. Cao, and A. C. Knoll, “Vision language models in autonomous driving: A survey and outlook,”IEEE Transactions on Intelligent Vehi- cles, early access, pp. 1–20, 2024, doi: 10.1109/TIV.2024.3402136
arXiv 2024
-
[9]
T. Nie, J. Sun, and W. Ma, “Exploring the Roles of Large Lan- guage Models in Reshaping Transportation Systems: A Survey, Framework,andRoadmap,”ArtificialIntelligenceforTransporta- tion, vol. 1, art. 100003, 2025, doi: 10.1016/j.ait.2025.100003
Show all 80 references
-
[10]
Ap- plications of Large Language Models and Generative AI in Transportation:ASystematicReviewandBibliometricAnalysis,
N. Maksoud, H. AlJassmi, L. Ali, and A. R. Masoud, “Ap- plications of Large Language Models and Generative AI in Transportation:ASystematicReviewandBibliometricAnalysis,” Transportation Research Interdisciplinary Perspectives, vol. 34, art. 101699, 2025, doi: 10.1016/j.trip.20...
2025
-
[11]
Large language models for transportation research: Methodologies, state of the art, and future opportu- nities,
Y. Yan et al., “Large language models for transportation research: Methodologies, state of the art, and future opportu- nities,”Information Fusion, vol. 136, art. 104546, 2026, doi: 10.1016/j.inffus.2026.104546
2026
-
[12]
A survey of large language models in transportation planning: Modelling, design and decision-making,
Y. Jin and J. Ma, “A survey of large language models in transportation planning: Modelling, design and decision-making,” Transportmetrica A: Transport Science, early access, pp. 1–46, 2026, doi: 10.1080/23249935.2026.2631152
2026
-
[13]
Harnessing large language models for intel- ligent transportation systems: A systematic review,
S. Kaur et al., “Harnessing large language models for intel- ligent transportation systems: A systematic review,”Multi- modal Transportation, vol. 5, no. 3, art. 100308, 2026, doi: 10.1016/j.multra.2026.100308
2026
-
[14]
UrbanGPT: Spatio-Temporal Large Language Models,
Z. Li et al., “UrbanGPT: Spatio-Temporal Large Language Models,” in Proc. 30th ACM SIGKDD Conf. Knowledge Discovery and Data Mining, 2024, pp. 5351–5362, doi: 10.1145/3637528.3671578. 17
2024
-
[15]
TSGDiff: Traffic State Gen- erative Diffusion Model Using Multi-Source Information Fusion,
H. Zhang, H. Dong, and Z. Yang, “TSGDiff: Traffic State Gen- erative Diffusion Model Using Multi-Source Information Fusion,” Transportation Research Part C: Emerging Technologies, vol. 174, art. 105081, 2025, doi: 10.1016/j.trc.2025.105081
2025
-
[16]
A Heterogeneous Graph Convolution Based Method for Short-Term OD Flow Completion and Prediction in a Metro System,
J. Ye, J. Zhao, F. Zheng, and C.-Z. Xu, “A Heterogeneous Graph Convolution Based Method for Short-Term OD Flow Completion and Prediction in a Metro System,” IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 5, pp. 4488–4500, 2024, doi: 10.1109/TITS.2023.3323756
2024
-
[17]
DriveLM: Driving with Graph Visual Question Answering,
C. Sima et al., “DriveLM: Driving with Graph Visual Question Answering,” in Computer Vision—ECCV 2024, 2024, pp. 256– 274, doi: 10.1007/978-3-031-72943-0_15
2024 doi
-
[18]
A Language Agent for Autonomous Driving,
J. Mao, J. Ye, Y. Qian, M. Pavone, and Y. Wang, “A Language Agent for Autonomous Driving,” in Proc. Conf. Language Modeling (COLM), 2024
2024
-
[19]
VLM-RL: A unified vision language models and reinforcement learning frame- work for safe autonomous driving,
Z. Huang, Z. Sheng, Y. Qu, J. You, and S. Chen, “VLM-RL: A unified vision language models and reinforcement learning frame- work for safe autonomous driving,”Transportation Research Part C: Emerging Technologies, vol. 180, art. 105321, 2025, doi: 10.1016/j.trc.2025.105321
2025
-
[20]
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reason- ing,
S. Wang et al., “OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reason- ing,” in Proc. IEEE/CVF CVPR, 2025, pp. 22442–22452, doi: 10.1109/CVPR52734.2025.02090
2025
-
[21]
RACP: Risk-aware contingency planning with multi-modal predictions,
K. A. Mustafa, D. J. Ornia, J. Kober, and J. Alonso-Mora, “RACP: Risk-aware contingency planning with multi-modal predictions,”IEEE Transactions on Intelligent Vehicles, vol. 10, no. 1, pp. 228–243, 2025, doi: 10.1109/TIV.2024.3411530
2025
-
[22]
LMDrive: Closed-Loop End-to-End Driving with Large Language Models,
H. Shao et al., “LMDrive: Closed-Loop End-to-End Driving with Large Language Models,” in Proc. IEEE/CVF CVPR, 2024, pp. 15120–15130, doi: 10.1109/CVPR52733.2024.01432
2024
-
[23]
LLMLight: Large Language Models as Traffic Signal Control Agents,
S. Lai, Z. Xu, W. Zhang, H. Liu, and H. Xiong, “LLMLight: Large Language Models as Traffic Signal Control Agents,” in Proc. 31st ACM SIGKDD Conf. Knowledge Discovery and Data Mining, 2025, pp. 2335–2346, doi: 10.1145/3690624.3709379
2025
-
[24]
The Crossroads of LLM and Traffic Control: A Study on Large Language Models in Adaptive Traffic Signal Control,
M. Movahedi and J. Choi, “The Crossroads of LLM and Traffic Control: A Study on Large Language Models in Adaptive Traffic Signal Control,” IEEE Transactions on Intelligent Trans- portation Systems, vol. 26, no. 2, pp. 1701–1716, 2025, doi: 10.1109/TITS.2024.3498735
2025
-
[25]
Large Language Models as Traffic Control Systems at Urban Intersections: A New Paradigm,
S. Masri, H. I. Ashqar, and M. Elhenawy, “Large Language Models as Traffic Control Systems at Urban Intersections: A New Paradigm,” Vehicles, vol. 7, art. 11, 2025, doi: 10.3390/ve- hicles7010011
2025 doi
-
[26]
PressLight: Learning Max Pressure Control to Coordinate Traffic Signals in Arterial Network,
H. Wei et al., “PressLight: Learning Max Pressure Control to Coordinate Traffic Signals in Arterial Network,” in Proc. 25th ACM SIGKDD Conf. Knowledge Discovery and Data Mining, 2019, pp. 1290–1298, doi: 10.1145/3292500.3330949
2019
-
[27]
Traffic Light Optimization With Low Penetra- tion Rate Vehicle Trajectory Data,
X. Wang et al., “Traffic Light Optimization With Low Penetra- tion Rate Vehicle Trajectory Data,” Nature Communications, vol. 15, art. 1306, 2024, doi: 10.1038/s41467-024-45427-4
2024 doi
-
[28]
Using Multimodal Large Language Models for Automated Detection of Traffic Safety-Critical Events,
M. Abu Tami, H. I. Ashqar, M. Elhenawy, S. Glaser, and A. Rakotonirainy, “Using Multimodal Large Language Models for Automated Detection of Traffic Safety-Critical Events,” Vehicles, vol. 6, pp. 1571–1590, 2024, doi: 10.3390/vehicles6030074
2024 doi
-
[29]
VRU-Accident: A vision–language benchmark for video question answering and dense captioning for accident scene understanding,
Y. Kim, A. S. Abdelrahman, and M. Abdel-Aty, “VRU-Accident: A vision–language benchmark for video question answering and dense captioning for accident scene understanding,” in Proc. IEEE/CVF ICCV Workshops, 2025, pp. 772–782, doi: 10.1109/ICCVW69036.2025.00085
2025
-
[30]
SafePLUG: Empowering multimodal LLMs with pixel-level insight and temporal grounding for traffic accident understanding,
Z. Sheng et al., “SafePLUG: Empowering multimodal LLMs with pixel-level insight and temporal grounding for traffic accident understanding,”CHAIN, vol. 3, no. 1, pp. 53–72, 2026, doi: 10.23919/CHAIN.2026.000005
2026
-
[31]
Large Language Models in Analyzing Crash Narratives—A Comparative Study of ChatGPT, BARD and GPT-4,
M. Mumtarin, M. S. Chowdhury, and J. Wood, “Large Language Models in Analyzing Crash Narratives—A Comparative Study of ChatGPT, BARD and GPT-4,” arXiv:2308.13563, 2023
2023 arXiv
-
[32]
Agentic Large Language Models for Day- to-Day Route Choices,
L. Wang et al., “Agentic Large Language Models for Day- to-Day Route Choices,” Transportation Research Part C: Emerging Technologies, vol. 180, art. 105307, 2025, doi: 10.1016/j.trc.2025.105307
2025
-
[33]
TransitGPT: A Generative AI- Based Framework for Interacting with GTFS Data Using Large Language Models,
S. Devunuri and L. J. Lehe, “TransitGPT: A Generative AI- Based Framework for Interacting with GTFS Data Using Large Language Models,” Public Transport, vol. 17, no. 2, pp. 319–345, 2025, doi: 10.1007/s12469-025-00395-w
2025 doi
-
[34]
ChatSUMO: Large Language Model for Automating Traffic Scenario Generation in Simulation of Urban MObility,
S. Li, T. Azfar, and R. Ke, “ChatSUMO: Large Language Model for Automating Traffic Scenario Generation in Simulation of Urban MObility,” IEEE Transactions on Intelligent Vehicles, vol. 10, no. 11, pp. 4962–4973, 2025, doi: 10.1109/TIV.2024.3508471
2025
-
[35]
Speak to Simulate: An LLM-Guided Agentic Framework for Traffic Simulation in SUMO,
M. Jeong, J. Chang, and Y. Yoon, “Speak to Simulate: An LLM-Guided Agentic Framework for Traffic Simulation in SUMO,” in Proc. 8th ACM SIGSPATIAL International Workshop on Geospatial Simulation, 2025, pp. 45–48, doi: 10.1145/3764921.3770151
2025
-
[36]
Automating Traffic Model Enhancement With AI Research Agent,
X. Guo, X. Yang, M. Peng, H. Lu, M. Zhu, and H. Yang, “Automating Traffic Model Enhancement With AI Research Agent,” Transportation Research Part C: Emerging Technologies, vol. 178, art. 105187, 2025, doi: 10.1016/j.trc.2025.105187
2025
-
[37]
Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection,
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection,” in Proc. ACM Workshop on AI and Security, 2023, pp. 79–90, doi: 10.1145/3605764.3623985
2023
-
[38]
Road Vehicles—Functional Safety—Part 1: Vocabulary,
ISO, “Road Vehicles—Functional Safety—Part 1: Vocabulary,” ISO 26262-1:2018, 2018
2018
-
[39]
Road Vehicles—Safety of the Intended Functionality,
ISO, “Road Vehicles—Safety of the Intended Functionality,” ISO 21448:2022, 2022
2022
-
[40]
Road Vehicles—Cybersecurity Engineering,
ISO and SAE International, “Road Vehicles—Cybersecurity Engineering,” ISO/SAE 21434:2021, 2021
2021
-
[41]
Road Vehicles—Safety and Artificial Intelligence,
ISO, “Road Vehicles—Safety and Artificial Intelligence,” ISO/PAS 8800:2024, 2024
2024
-
[42]
IEEE Standard for Transparency of Autonomous Sys- tems,
IEEE, “IEEE Standard for Transparency of Autonomous Sys- tems,” IEEE Std 7001-2021, 2021
2021
-
[43]
Artificial Intelligence Risk Management Framework (AI RMF 1.0),
E. Tabassi, “Artificial Intelligence Risk Management Framework (AI RMF 1.0),” NIST AI 100-1, National Institute of Standards and Technology, 2023, doi: 10.6028/NIST.AI.100-1
2023 doi
-
[44]
Information Technology—Artificial Intelligence— Guidance on Risk Management,
ISO/IEC, “Information Technology—Artificial Intelligence— Guidance on Risk Management,” ISO/IEC 23894:2023, 2023
2023
-
[45]
Information Technology—Artificial Intelligence— Management System,
ISO/IEC, “Information Technology—Artificial Intelligence— Management System,” ISO/IEC 42001:2023, 2023
2023
-
[47]
Scalability in perception for autonomous driving: Waymo Open Dataset,
P. Sun et al., “Scalability in perception for autonomous driving: Waymo Open Dataset,” inProc. IEEE/CVF CVPR, 2020, pp. 2443–2451, doi: 10.1109/CVPR42600.2020.00252
2020
-
[48]
The highD Dataset: A Drone Dataset of Naturalistic Vehicle Trajectories on German Highways for ValidationofHighlyAutomatedDrivingSystems,
R. Krajewski et al., “The highD Dataset: A Drone Dataset of Naturalistic Vehicle Trajectories on German Highways for ValidationofHighlyAutomatedDrivingSystems,”inProc.IEEE ITSC, 2018, pp. 2118–2125, doi: 10.1109/ITSC.2018.8569552
2018
-
[49]
CityFlow: A Multi-Agent Reinforcement Learn- ing Environment for Large Scale City Traffic Scenario,
H. Zhang et al., “CityFlow: A Multi-Agent Reinforcement Learn- ing Environment for Large Scale City Traffic Scenario,” in Proc. WWW, 2019, pp. 3620–3624, doi: 10.1145/3308558.3314139
2019
-
[50]
Microscopic Traffic Simulation Using SUMO,
P. A. Lopez et al., “Microscopic Traffic Simulation Using SUMO,” in Proc. IEEE ITSC, 2018, pp. 2575–2582, doi: 10.1109/ITSC.2018.8569938
2018
-
[51]
Large models for intelligent transportation systems and autonomous vehicles: A survey,
L. Gan, W. Chu, G. Li, X. Tang, and K. Li, “Large models for intelligent transportation systems and autonomous vehicles: A survey,”Advanced Engineering Informatics, vol. 62, art. 102786, 2024, doi: 10.1016/j.aei.2024.102786
2024
-
[52]
TrafficMind: A system- oriented review of large language models for intelligent trans- portation systems,
C. Zhang, B. Wei, and L. Yang, “TrafficMind: A system- oriented review of large language models for intelligent trans- portation systems,”Journal of Transportation Engineering, Part A: Systems, vol. 152, no. 9, art. 03126005, 2026, doi: 10.1061/JTEPBS.TEENG-9794
2026 doi
-
[53]
Foundation models for autonomous driving: A comprehensive survey,
S. Fourati, W. Jaafar, N. Baccar, S. Alfattani, and R. Langar, “Foundation models for autonomous driving: A comprehensive survey,”Engineering Applications of Artificial Intelligence, vol. 176, art. 114805, 2026, doi: 10.1016/j.engappai.2026.114805
2026
-
[54]
The role of large language models (LLMs) in enhancing intelligent transportation systems: A survey,
V. Hassija, T. Majumder, D. Roy, R. Piyush, and V. Chamola, “The role of large language models (LLMs) in enhancing intelligent transportation systems: A survey,”Ve- hicular Communications, vol. 58, art. 100996, 2026, doi: 10.1016/j.vehcom.2025.100996
2026
-
[55]
Integrating LLMs with ITS: Recent advances, potentials, challenges, and future directions,
D. Mahmud, H. Hajmohamed, S. Almentheri, S. Alqaydi, L. Ald- haheri, R. A. Khalil, and N. Saeed, “Integrating LLMs with ITS: Recent advances, potentials, challenges, and future directions,” IEEE Transactions on Intelligent Transportation Systems, vol. 26, no. 5, pp. 5674–5709,...
2025
-
[56]
DiLu: A Knowledge-Driven Approach to Au- tonomous Driving with Large Language Models,
L. Wen et al., “DiLu: A Knowledge-Driven Approach to Au- tonomous Driving with Large Language Models,” inProc. Int. Conf. Learning Representations (ICLR), 2024. 18
2024
-
[57]
Dolphins: Multimodal Language Model for Driving,
Y. Ma, Y. Cao, J. Sun, M. Pavone, and C. Xiao, “Dolphins: Multimodal Language Model for Driving,” inComputer Vision– ECCV 2024, 2024, doi: 10.1007/978-3-031-72995-9_23
2024 doi
-
[58]
Driving Everywhere with Large Language Model PolicyAdaptation,
B. Li et al., “Driving Everywhere with Large Language Model PolicyAdaptation,”inProc.IEEE/CVFCVPR,2024,pp.14948– 14957
2024
-
[59]
LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs,
Y. Ma et al., “LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs,” inProc. IEEE/CVF CVPR, 2024, pp. 15141–15151
2024
-
[60]
ChatScene: Knowledge-Enabled Safety-Critical Scenario Generation for Autonomous Vehicles,
J. Zhang, C. Xu, and B. Li, “ChatScene: Knowledge-Enabled Safety-Critical Scenario Generation for Autonomous Vehicles,” inProc. IEEE/CVF CVPR, 2024, pp. 15459–15469
2024
-
[61]
Editable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents,
Y. Wei et al., “Editable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents,” inProc. IEEE/CVF CVPR, 2024, pp. 15077–15087
2024
-
[62]
LLMScenario: Large Language Model Driven Scenario Generation,
C. Chang, S. Wang, J. Zhang, J. Ge, and L. Li, “LLMScenario: Large Language Model Driven Scenario Generation,”IEEE Trans. Syst., Man, Cybern.: Syst., vol. 54, no. 11, pp. 6581–6594, 2024, doi: 10.1109/TSMC.2024.3392930
2024
-
[63]
DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models,
X. Tian et al., “DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models,” inProc. 8th Conf. Robot Learning, PMLR, vol. 270, pp. 4698–4726, 2025
2025
-
[64]
DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving,
Z. Xu et al., “DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving,” inProc. IEEE/CVF CVPR, 2025, pp. 17261–17270
2025
-
[65]
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Gen- eration,
H. Fu et al., “ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Gen- eration,” inProc. IEEE/CVF ICCV, 2025, pp. 24823–24834
2025
-
[66]
GATSim: Urban Mobility Simula- tion with Generative Agents,
Q. Liu, C. Li, and W. Ma, “GATSim: Urban Mobility Simula- tion with Generative Agents,”Transportation Research Part C: Emerging Technologies, vol. 186, art. 105576, 2026, doi: 10.1016/j.trc.2026.105576
2026
-
[67]
AITP: Traffic Accident Responsibility Allocation via Multimodal Large Language Models,
Z. Zhou and S. Zhang, “AITP: Traffic Accident Responsibility Allocation via Multimodal Large Language Models,” inProc. IEEE/CVF CVPR Findings, 2026, pp. 9259–9268
2026
-
[68]
MindDriver: Introducing Progressive Multi- modal Reasoning for Autonomous Driving,
L. Zhang et al., “MindDriver: Introducing Progressive Multi- modal Reasoning for Autonomous Driving,” inProc. IEEE/CVF CVPR, 2026, pp. 17831–17841
2026
-
[69]
V2X-UniPool: Unifying Multimodal Perception and Knowledge Reasoning for Autonomous Driving,
X. Luo et al., “V2X-UniPool: Unifying Multimodal Perception and Knowledge Reasoning for Autonomous Driving,” inProc. IEEE/CVF CVPR Workshops, 2026, pp. 747–756
2026
-
[70]
Large Language Model as Parking Planning Agent in the Context of Mixed Period of Autonomous Vehicles and Human-Driven Vehicles,
Y. Jin and J. Ma, “Large Language Model as Parking Planning Agent in the Context of Mixed Period of Autonomous Vehicles and Human-Driven Vehicles,”Sustainable Cities and Society, vol. 117, art. 105940, 2024, doi: 10.1016/j.scs.2024.105940
2024
-
[71]
Bridging AI and Traffic Simulation: A Robust and Comprehensive Framework for LLM-Based AI Replanning Agents in MATSim,
A. U. Z. Patwary et al., “Bridging AI and Traffic Simulation: A Robust and Comprehensive Framework for LLM-Based AI Replanning Agents in MATSim,”Procedia Computer Science, vol. 280, pp. 622–629, 2026, doi: 10.1016/j.procs.2026.04.079
2026 doi
-
[72]
Agentic Traffic Intelligence: Augmented Human-in-the- Loop Scenario Generation for Microscopic Traffic Simulation,
X. Luo, G. Xu, A. Saroj, J. Yuan, P. Kadav, Y. Shao, and C. R. Wang, “Agentic Traffic Intelligence: Augmented Human-in-the- Loop Scenario Generation for Microscopic Traffic Simulation,” Artificial Intelligence for Transportation, vol. 6, art. 100057, 2026, doi: 10.1016/j.ait.2...
2026
-
[73]
Large Language Model-Assisted Multi-Objective Optimization for an Integrated Multimodal E-Mobility Platform,
Y. Ding, M. Maniparambil, N. E. O’Connor, and M. Liu, “Large Language Model-Assisted Multi-Objective Optimization for an Integrated Multimodal E-Mobility Platform,”Transportation ResearchInterdisciplinaryPerspectives,vol.37,art.101948,2026, doi: 10.1016/j.trip.2026.101948
2026
-
[74]
LLMs as Virtual Traffic Police: Incident-Aware Traffic Signal Control Augmented by Large Language Models,
Q. Wang, S. Wei, and K. Yang, “LLMs as Virtual Traffic Police: Incident-Aware Traffic Signal Control Augmented by Large Language Models,” inProc. IEEE 28th Int. Conf. Intell. Transp. Syst. (ITSC), 2025, doi: 10.1109/ITSC60802.2025.11423524
2025 arXiv
-
[75]
An Efficient Simulation Scene Generation Method BasedonExtractedRoadNetworkTopologyandLargeLanguage Models,
R. Li et al., “An Efficient Simulation Scene Generation Method BasedonExtractedRoadNetworkTopologyandLargeLanguage Models,”Future Transportation, vol. 6, no. 2, art. 81, 2026, doi: 10.3390/futuretransp6020081
2026 doi
-
[76]
DriveGPT4:InterpretableEnd-to-EndAutonomous Driving Via Large Language Model,
Z.Xuetal.,“DriveGPT4:InterpretableEnd-to-EndAutonomous Driving Via Large Language Model,”IEEE Robotics and Au- tomation Letters, 2024, doi: 10.1109/LRA.2024.3440097
2024
-
[77]
ChatSUMO Agent: An LLM- Based Agent for Conversational Traffic Simulation in SUMO,
S. Li, M. Ma, T. Azfar, and R. Ke, “ChatSUMO Agent: An LLM- Based Agent for Conversational Traffic Simulation in SUMO,” SSRN preprint, 1 Jan. 2026, doi: 10.2139/ssrn.6000335
2026 doi
-
[78]
Generalizing End-to-End Autonomous Driving in Real-World Environments Using Zero-Shot LLMs,
Z. Dong, Y. Zhu, Y. Li, K. Mahon, and Y. Sun, “Generalizing End-to-End Autonomous Driving in Real-World Environments Using Zero-Shot LLMs,” inProc. 8th Conf. Robot Learning, PMLR, vol. 270, pp. 1231–1249, 2025. [Online]. Available: https: //proceedings.mlr.press/v270/dong25a.html
2025
-
[79]
Automating the Loop in Traffic Incident Manage- ment on Highway,
M. Cercola, N. Gatti, P. Huertas Leyva, B. Carambia, and S. Formentin, “Automating the Loop in Traffic Incident Manage- ment on Highway,” inProc. 7th Annu. Learning for Dynamics and Control Conf., PMLR, vol. 283, pp. 272–284, 2025. [Online]. Available: https://proceedings.mlr....
2025
-
[80]
Promptable Closed-Loop Traffic Simulation,
S. Tan, B. Ivanovic, Y. Chen, B. Li, X. Weng, Y. Cao, P. Kraehenbuehl, and M. Pavone, “Promptable Closed-Loop Traffic Simulation,” inProc. 8th Conf. Robot Learning, PMLR, vol. 270, pp. 5087–5105, 2025. [Online]. Available: https://proceedings.ml r.press/v270/tan25a.html
2025
-
[81]
DriveCombo: Benchmarking Compositional Traffic Rule Reasoning in Autonomous Driving,
E. Ma et al., “DriveCombo: Benchmarking Compositional Traffic Rule Reasoning in Autonomous Driving,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 32113–32123,
-
[2026]
Available: https://openaccess.thecvf.com/conten t/CVPR2026/html/Ma_DriveCombo_Benchmarking_Compo sitional_Traffic_Rule_Reasoning_in_Autonomous_Driving_ CVPR_2026_paper.html
[Online]. Available: https://openaccess.thecvf.com/conten t/CVPR2026/html/Ma_DriveCombo_Benchmarking_Compo sitional_Traffic_Rule_Reasoning_in_Autonomous_Driving_ CVPR_2026_paper.html
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.