Pith. sign in

REVIEW 3 major objections 5 minor 80 references

Large Multimodal Agents for Intelligent Transportation Systems: Architectures, Evidence, and Deployment Challenges

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A 42-family evidence map concludes that large multimodal agents are best used as orchestrators that interpret and coordinate transportation evidence, while forecasting, optimization, control, and final authority remain with specialist…

desk verdict A careful, transparent evidence map that makes a plausible bounded-orchestration case for LMAs in ITS, with the main caveat being that the coding reliability behind its headline counts is not yet independently auditable from the preprint alone. read the letter →

arxiv 2608.08184 v1 pith:3QPHDLA3 submitted 2026-08-08 cs.AI

classification cs.AI
keywords largemultimodalagentsintelligenttransportationsystemsevidencemappingreasoningtrustworthyAItrafficoperationsclosed-loopdecisionmaking
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review maps 42 families of large multimodal agents (LMAs) used in intelligent transportation systems and asks what they demonstrably do, rather than what their names promise. Its central conclusion is that the strongest evidence supports LMAs for interpreting transportation semantics and integrating heterogeneous evidence, while numerical forecasting, optimization, low-level control, safety fallback, and final authority should stay with independently verifiable specialist systems or accountable humans. The review finds that capability outpaces validation: fourteen families show outcome-responsive agency (C3), but only one has been tested in a controlled real-world setting and none has evidence of sustained routine deployment. Evidence reconciliation—whether traceable provenance and handling of missing or conflicting data changes a transportation decision—remains unevaluated in every family. The paper therefore argues for bounded orchestration rather than replacement.

What carries the argument

The carrying object is the review's orthogonal coding framework applied to 42 verified study families: a C0–C3 functional-capability scale (from pre-agentic capability to outcome-responsive agency), an E0–E4 validation-setting scale (from conceptual to sustained deployment), three evidence propositions P1–P3 (transportation semantics, evidence reconciliation, multidimensional integration), and eight noncompensatory methodological-concern domains Q1–Q8. The framework's work is to separate what a system demonstrably does from where it was evaluated and from whether a claimed benefit is directly supported, so that agency, multimodality, and architectural complexity cannot by themselves be counted as evidence. The synthesis database built from these codings is the analytical source of truth behind the counts of 14 C3 families, one E3 family, and no E4 family.

What would settle it

An independent re-coding of the same 42 families using the published rubric—especially a check of whether any family actually completes the full P2 chain (provenance, challenge, handling, matched comparison, attributable outcome) and whether all 14 C3 labels show feedback changing a later foundation-model decision—would settle the claim; if the counts moved materially, the bounded-orchestration conclusion would need revision.

Watch

Extended reading notes

Core claim

The paper establishes an auditable evidence map of 42 primary study families and shows that direct evidence is concentrated in two of its three testable propositions. Twenty-three families directly evaluate transportation semantics (P1) and 24 directly evaluate multidimensional integration (P3), with 19 evaluating both; all direct P1 or P3 judgments use strong or moderate claim-matched comparisons. No family directly evaluates the full evidence-reconciliation chain (P2), in which traceable provenance, a missingness or conflict challenge, a handling mechanism, a matched comparison, and an attributable outcome all appear. On the capability scale, 14 families reach C3, meaning observed transportation outcomes change a later foundation-model decision, but validation lags: 13 sit at E2 interactive simulation, one reaches E3 controlled real-world operation, and none reaches E4 sustained deployment. The review concludes that LMAs are best supported for semantic interpretation, intent translation, evidence organisation, scenario authoring, explanation, and specialist-tool coordination, and that this supports contribution-specific orchestration, not general replacement.

Load-bearing premise

The whole conclusion rests on the accuracy of the review's own coding of 42 study families; if the C0–C3, E0–E4, P1–P3, and Q1–Q8 labels were applied inconsistently or the frozen database misrecorded the studies, the counts that drive the bounded-orchestration conclusion could change.

Editorial extensions

If this is right

  • Near-term ITS-LMA deployments should be read-only or advisory, with source attribution, explicit uncertainty, and accountable human review.
  • Larger action authority should be granted only after direct P2 evidence, robustness and failure analysis, controlled E3 trials, and longitudinal E4 evidence within a declared operational envelope.
  • Evaluation of future systems should use matched configurations: specialist alone, LMA without tools, LMA orchestrating the specialist, and the complete architecture with permission gateway, independent monitor, fallback, and rollback.
  • Numerical forecasting, optimization, simulation fidelity, hard constraints, low-level control, safety fallback, and final authority should remain with specialist systems or accountable humans until stronger claim-matched evidence appears.
  • The P2 gap means provenance and conflict handling is a required next evaluation target, not an optional feature.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural testable extension is to turn the P2 chain into a checklist for new evaluations: any system claiming evidence reconciliation should report provenance records, a missingness or conflict challenge, a handling mechanism, a matched comparison, and an attributable outcome.
  • A concrete test of the bounded-orchestration claim would be a benchmark task built from the paper's suggested agentic-refinement route: generate, execute, evaluate, and iteratively refine a specialist model's code on fixed datasets, with human approval gates and failure logs recorded.
  • The concentration of evidence in autonomous driving and planning/simulation suggests that the semantic layer may generalize more readily to incident interpretation and scenario authoring than to numerical forecasting, which should be tested explicitly across cities and event conditions.
  • A living evidence registry that records search decisions, study-family links, prompts, model versions, and failure logs would let future updates test whether the 14-C3, one-E3, no-E4 profile changes as more controlled real-world evaluations appear.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper presents a structured narrative review and evidence map of 42 primary study families of large multimodal agents (LMAs) for intelligent transportation systems (ITS), covering sources released between January 2023 and 3 August 2026. The authors propose an analytical framework that separates model-level, system-level, and hybrid multimodality; defines capability levels C0–C3 and validation settings E0–E4; and evaluates three evidence propositions (P1 transportation semantics, P2 evidence reconciliation, P3 multidimensional integration) independently of methodological concerns (Q1–Q8). The central findings are that 14 families reach C3, only one reaches E3, none reaches E4, and P2 remains unresolved because no family demonstrates the complete provenance–challenge–handling–comparison–outcome chain. The paper concludes that LMAs are best suited for bounded orchestration—semantic interpretation, evidence organization, and specialist-tool coordination—rather than replacement of specialist forecasting, optimization, simulation, and control systems, and it proposes a matched comparative evaluation protocol and a staged deployment roadmap.

Significance. If the evidence map is reliable, the paper offers a useful contribution by operationalizing distinctions that prior surveys leave implicit: capability versus validation setting, proposition directness versus result direction, and LMA authority versus specialist authority. The explicit definitions of C0–C3 and E0–E4, the orthogonal coding of P1–P3 and Q1–Q8, and the proposed four-configuration comparative protocol are valuable for future evaluations. The authors are transparent about limitations, including the lack of a PRISMA flow, the absence of inter-rater reliability statistics, and the reliance on a single-author coding process. The main risk is that the headline counts and the P2-unresolved conclusion are not independently verifiable from the preprint, which undermines the auditability that the paper claims as a central feature.

major comments (3)
  1. [Section II-B and Section VII] The central quantitative claims (14 C3, one E3, no E4, P2 unresolved, all direct P1/P3 in strong or moderate comparison groups) rest entirely on the consistency of the author-defined coding framework. The paper explicitly states that all pilot and repeat procedures were within-process stability checks, not independent duplicate human coding, and that no inter-rater reliability statistic is claimed. The label-masked repeat audits (40/42 C, 37/42 E, 280/336 Q, 113/126 P) are reported only as aggregate counts, with no item-level disagreements. The canonical synthesis database, audit register, and Supplementary Tables S5–S8 are referenced but not included in the preprint. A reader therefore cannot verify that the 42-family cross-tabulations in Table IV and the P2-unresolved conclusion follow from the underlying studies. For an 'auditable evidence map,' this is a load-bearing reproducibility gap.
  2. [Section VII] The paper claims that 'prespecified conservative sensitivity analyses did not alter the central bounded-orchestration conclusion' and that full results are provided in Supplementary Table S6, but this table is not part of the preprint. Because the conclusion depends on coding thresholds such as the strict P2-D criterion, the E3/E4 boundary, and the lower-code rule, the reader cannot assess how robust the headline counts are to plausible coding variations. Please provide the sensitivity results, or make them available in the repository with a clear pointer in the manuscript.
  3. [Abstract and Section VI-A] The claim that 'all direct P1 or P3 judgments use strong or moderate claim-matched comparisons' is central to the conclusion that direct evidence is strongest for semantics and integration. However, the family-level mapping between proposition directness and comparator strength is not reported; Table IV only gives domain-level aggregates, and 14 families are described as having partial or nonisolating comparisons. Without a family-level table showing which families receive D codes and what comparator strength they have, this claim cannot be checked. This mapping should be included in the supplementary material.
minor comments (5)
  1. [Section III-C, paragraph before Table I] The sentence 'C1 rule. After confirming a substantive transportation role, C denotes actionable decision-support...' is garbled; it should likely read 'C1 applies when the foundation model participates in an actionable decision-support cycle with implementation retained externally.'
  2. [Section IV] The phrase 'Table III summaries these differences' should be 'Table III summarizes these differences.'
  3. [Section VI-A] The sentence 'Only two families wereevaluated primarily in real-world or onboard settings' contains a typo: 'wereevaluated' should be 'were evaluated.'
  4. [Section II-A] The paper refers to Supplementary Tables S2, S5, S6, S7, and S8, but the preprint does not include them and the body text does not state where they can be accessed beyond the abstract's GitHub URL. Please add a clear pointer in the body to the repository or supplementary file location.
  5. [Section II-B] The disclosure that 'structured model assistance' was used for literature retrieval, metadata checking, and extraction drafting would be more reproducible if the specific models and versions were stated, or if the prompts and protocols were archived with the other review materials.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the review's conclusions are coding-based summaries of external studies, not predictions derived from fitted inputs or self-citation chains.

full rationale

This manuscript is a structured narrative review and evidence map rather than a formal derivation, so the classical circularity failure modes (fitted parameter renamed as prediction, self-definitional derivation, or self-citation chain forcing a conclusion) do not arise. The central claims—23 families with direct P1 evidence, 24 with direct P3, P2 unresolved, 14 C3 families, one E3 family, and no E4 family, and hence bounded orchestration—are counts and interpretations produced by applying an explicit, author-defined coding framework to 42 independently published study families. The P2 conclusion is definitionally tied to the P2-D criterion (a family must demonstrate the full provenance–challenge–handling–comparison–outcome chain to be coded D), but this is a transparent coding report rather than a hidden reduction: no quantity is fitted to one subset and then 'predicted' on a related subset. The paper's own limitations (single-author-led coding, no inter-rater reliability statistic, Supplementary Table S6 not included in the preprint) are validity and auditability risks, not circularity; they concern whether the codes are correct, not whether the conclusion reduces to its inputs. The cited prior work is external to the authors' framework and does not carry a load-bearing uniqueness or ansatz argument. Accordingly, no circular step can be exhibited, and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The review's central conclusion is an interpretation of a corpus coded through author-defined scales; it rests on domain assumptions about inclusion, coding, and evidence thresholds rather than on measured parameters or new entities.

assumptions (6)
  • domain assumption The operational definition of ITS-LMA and the inclusion boundary determine which studies enter the evidence map (Section III-B).
    A different boundary (e.g., including language interfaces or conventional multimodal predictors) would change family counts and the resulting conclusions.
  • domain assumption The C0-C3 capability scale and E0-E4 validation scale are ordinal categories defined by the authors (Section III-C).
    The conclusion that capability exceeds validation maturity is expressed in these units; alternative scales could produce different patterns.
  • domain assumption The P1-P3 proposition evidence criteria, particularly the strict P2 provenance-challenge-handling-comparison-outcome chain, define what counts as direct evidence (Sections II-B and III-A).
    The finding that P2 is unresolved follows from the strictness of this criterion; a weaker criterion could change the outcome.
  • domain assumption The search window (January 2023 to 3 August 2026) and reliance on publicly accessible English-language reports bound the corpus (Section II-A and Section VII).
    Private industrial evaluations and non-English reports are likely under-represented, which could alter the evidence map.
  • domain assumption The five application domains are mutually exclusive groupings by principal evaluated LMA function (Section V).
    The paper states that alternative defensible assignments could change domain subtotals without changing family-level labels.
  • domain assumption Label-masked repeat audits are within-process stability checks, not independent duplicate human coding; no inter-rater reliability statistic is claimed (Section II-B and Section VII).
    The reliability of the 42-family coding depends on the undocumented consistency of the single-pipeline review process.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Large Multimodal Agents for Intelligent Transportation Systems: Architectures, Evidence, and Deployment Challenges." pith.science (2026). https://pith.science/paper/3QPHDLA3

@misc{pith2026260808184,
  author       = {Pith},
  title        = {Pith review of: Large Multimodal Agents for Intelligent Transportation Systems: Architectures, Evidence, and Deployment Challenges},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3QPHDLA3}},
  note         = {Machine review of arXiv:2608.08184}
}
read the original abstract

Large multimodal agents (LMAs) are increasingly proposed for intelligent transportation systems (ITS), but existing studies often conflate multimodality, agency, empirical performance, and deployment readiness. This review provides an auditable evidence map of 42 primary study families released between January 2023 and 3 August 2026 within a corpus of 91 mapped sources. It distinguishes model-level, system-level, and hybrid multimodality and classifies each family by system architecture and action authority. Evidence is assessed independently through functional capability (C0-C3), validation setting (E0-E4), three evidence propositions (P1-P3), and eight methodological-concern domains (Q1-Q8). Transportation semantics (P1) are directly evaluated in 23 families and multidimensional integration (P3) in 24; 19 families directly evaluate both. Evidence reconciliation (P2) remains unresolved because no family demonstrates the complete provenance-challenge-handling-comparison-outcome chain. Fourteen families reach C3, but 13 remain at E2; only one reaches E3 and none reaches E4. Across ITS domains, LMAs are best supported for semantic interpretation, intent translation, evidence organisation, scenario authoring, explanation, and specialist-tool coordination. Numerical forecasting, optimisation, simulation fidelity, hard constraints, low-level control, safety fallback, and final authority should remain with independently verifiable specialist systems or accountable humans. The review therefore supports bounded orchestration rather than replacement and provides a matched comparative evaluation protocol and staged roadmap for accountable deployment. The living evidence repository is available at https://github.com/pangjunbiao/ITS-LMA-Review.

Figures

Figures reproduced from arXiv: 2608.08184 by the authors.

Figure 1
Figure 1. Deployment-oriented ITS-LMA reference architecture and operational boundaries. Multimodal evidence is grounded into a semantic [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

80 extracted references · 47 canonical work pages

  1. [1]

    The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews,

    M. J. Page et al., “The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews,” BMJ, vol. 372, art. n71, 2021, doi: 10.1136/bmj.n71

  2. [2]

    Taxonomy and Definitions for Terms Re- lated to Driving Automation Systems for On-Road Motor Vehicles,

    SAE International, “Taxonomy and Definitions for Terms Re- lated to Driving Automation Systems for On-Road Motor Vehicles,” SAE J3016_202104, 2021

  3. [3]

    TRIP: Transport Reasoning With Intelligence Progression—A Foundation Framework,

    Z. Liu et al., “TRIP: Transport Reasoning With Intelligence Progression—A Foundation Framework,” Transportation Re- search Part C: Emerging Technologies, vol. 179, art. 105260, 2025, doi: 10.1016/j.trc.2025.105260

  4. [4]

    Large multimodal agents: A survey,

    J. Xie, Z. Chen, R. Zhang, and G. Li, “Large multimodal agents: A survey,”Visual Intelligence, vol. 3, art. 24, 2025, doi: 10.1007/s44267-025-00093-y

  5. [5]

    A Survey on Multimodal Large Language Models for Autonomous Driving,

    C. Cui et al., “A Survey on Multimodal Large Language Models for Autonomous Driving,” in Proc. IEEE/CVF WACV Workshops, 2024, pp. 958–979, doi: 10.1109/WACVW60836.2024.00106

  6. [7]

    Large Language Models for Intelligent Transportation: A Review of the State of the Art and Challenges,

    S. Wandelt, C. Zheng, S. Wang, Y. Liu, and X. Sun, “Large Language Models for Intelligent Transportation: A Review of the State of the Art and Challenges,” Applied Sciences, vol. 14, no. 17, art. 7455, 2024, doi: 10.3390/app14177455

  7. [8]

    Vision language models in autonomous driving: A survey and outlook,

    X. Zhou, M. Liu, E. Yurtsever, B. L. Zagar, W. Zimmer, H. Cao, and A. C. Knoll, “Vision language models in autonomous driving: A survey and outlook,”IEEE Transactions on Intelligent Vehi- cles, early access, pp. 1–20, 2024, doi: 10.1109/TIV.2024.3402136

  8. [9]

    Exploring the Roles of Large Lan- guage Models in Reshaping Transportation Systems: A Survey, Framework,andRoadmap,

    T. Nie, J. Sun, and W. Ma, “Exploring the Roles of Large Lan- guage Models in Reshaping Transportation Systems: A Survey, Framework,andRoadmap,”ArtificialIntelligenceforTransporta- tion, vol. 1, art. 100003, 2025, doi: 10.1016/j.ait.2025.100003

Show all 80 references
  1. [10]

    Ap- plications of Large Language Models and Generative AI in Transportation:ASystematicReviewandBibliometricAnalysis,

    N. Maksoud, H. AlJassmi, L. Ali, and A. R. Masoud, “Ap- plications of Large Language Models and Generative AI in Transportation:ASystematicReviewandBibliometricAnalysis,” Transportation Research Interdisciplinary Perspectives, vol. 34, art. 101699, 2025, doi: 10.1016/j.trip.20...

  2. [11]

    Large language models for transportation research: Methodologies, state of the art, and future opportu- nities,

    Y. Yan et al., “Large language models for transportation research: Methodologies, state of the art, and future opportu- nities,”Information Fusion, vol. 136, art. 104546, 2026, doi: 10.1016/j.inffus.2026.104546

  3. [12]

    A survey of large language models in transportation planning: Modelling, design and decision-making,

    Y. Jin and J. Ma, “A survey of large language models in transportation planning: Modelling, design and decision-making,” Transportmetrica A: Transport Science, early access, pp. 1–46, 2026, doi: 10.1080/23249935.2026.2631152

  4. [13]

    Harnessing large language models for intel- ligent transportation systems: A systematic review,

    S. Kaur et al., “Harnessing large language models for intel- ligent transportation systems: A systematic review,”Multi- modal Transportation, vol. 5, no. 3, art. 100308, 2026, doi: 10.1016/j.multra.2026.100308

  5. [14]

    UrbanGPT: Spatio-Temporal Large Language Models,

    Z. Li et al., “UrbanGPT: Spatio-Temporal Large Language Models,” in Proc. 30th ACM SIGKDD Conf. Knowledge Discovery and Data Mining, 2024, pp. 5351–5362, doi: 10.1145/3637528.3671578. 17

  6. [15]

    TSGDiff: Traffic State Gen- erative Diffusion Model Using Multi-Source Information Fusion,

    H. Zhang, H. Dong, and Z. Yang, “TSGDiff: Traffic State Gen- erative Diffusion Model Using Multi-Source Information Fusion,” Transportation Research Part C: Emerging Technologies, vol. 174, art. 105081, 2025, doi: 10.1016/j.trc.2025.105081

  7. [16]

    A Heterogeneous Graph Convolution Based Method for Short-Term OD Flow Completion and Prediction in a Metro System,

    J. Ye, J. Zhao, F. Zheng, and C.-Z. Xu, “A Heterogeneous Graph Convolution Based Method for Short-Term OD Flow Completion and Prediction in a Metro System,” IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 5, pp. 4488–4500, 2024, doi: 10.1109/TITS.2023.3323756

  8. [17]

    DriveLM: Driving with Graph Visual Question Answering,

    C. Sima et al., “DriveLM: Driving with Graph Visual Question Answering,” in Computer Vision—ECCV 2024, 2024, pp. 256– 274, doi: 10.1007/978-3-031-72943-0_15

  9. [18]

    A Language Agent for Autonomous Driving,

    J. Mao, J. Ye, Y. Qian, M. Pavone, and Y. Wang, “A Language Agent for Autonomous Driving,” in Proc. Conf. Language Modeling (COLM), 2024

  10. [19]

    VLM-RL: A unified vision language models and reinforcement learning frame- work for safe autonomous driving,

    Z. Huang, Z. Sheng, Y. Qu, J. You, and S. Chen, “VLM-RL: A unified vision language models and reinforcement learning frame- work for safe autonomous driving,”Transportation Research Part C: Emerging Technologies, vol. 180, art. 105321, 2025, doi: 10.1016/j.trc.2025.105321

  11. [20]

    OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reason- ing,

    S. Wang et al., “OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reason- ing,” in Proc. IEEE/CVF CVPR, 2025, pp. 22442–22452, doi: 10.1109/CVPR52734.2025.02090

  12. [21]

    RACP: Risk-aware contingency planning with multi-modal predictions,

    K. A. Mustafa, D. J. Ornia, J. Kober, and J. Alonso-Mora, “RACP: Risk-aware contingency planning with multi-modal predictions,”IEEE Transactions on Intelligent Vehicles, vol. 10, no. 1, pp. 228–243, 2025, doi: 10.1109/TIV.2024.3411530

  13. [22]

    LMDrive: Closed-Loop End-to-End Driving with Large Language Models,

    H. Shao et al., “LMDrive: Closed-Loop End-to-End Driving with Large Language Models,” in Proc. IEEE/CVF CVPR, 2024, pp. 15120–15130, doi: 10.1109/CVPR52733.2024.01432

  14. [23]

    LLMLight: Large Language Models as Traffic Signal Control Agents,

    S. Lai, Z. Xu, W. Zhang, H. Liu, and H. Xiong, “LLMLight: Large Language Models as Traffic Signal Control Agents,” in Proc. 31st ACM SIGKDD Conf. Knowledge Discovery and Data Mining, 2025, pp. 2335–2346, doi: 10.1145/3690624.3709379

  15. [24]

    The Crossroads of LLM and Traffic Control: A Study on Large Language Models in Adaptive Traffic Signal Control,

    M. Movahedi and J. Choi, “The Crossroads of LLM and Traffic Control: A Study on Large Language Models in Adaptive Traffic Signal Control,” IEEE Transactions on Intelligent Trans- portation Systems, vol. 26, no. 2, pp. 1701–1716, 2025, doi: 10.1109/TITS.2024.3498735

  16. [25]

    Large Language Models as Traffic Control Systems at Urban Intersections: A New Paradigm,

    S. Masri, H. I. Ashqar, and M. Elhenawy, “Large Language Models as Traffic Control Systems at Urban Intersections: A New Paradigm,” Vehicles, vol. 7, art. 11, 2025, doi: 10.3390/ve- hicles7010011

  17. [26]

    PressLight: Learning Max Pressure Control to Coordinate Traffic Signals in Arterial Network,

    H. Wei et al., “PressLight: Learning Max Pressure Control to Coordinate Traffic Signals in Arterial Network,” in Proc. 25th ACM SIGKDD Conf. Knowledge Discovery and Data Mining, 2019, pp. 1290–1298, doi: 10.1145/3292500.3330949

  18. [27]

    Traffic Light Optimization With Low Penetra- tion Rate Vehicle Trajectory Data,

    X. Wang et al., “Traffic Light Optimization With Low Penetra- tion Rate Vehicle Trajectory Data,” Nature Communications, vol. 15, art. 1306, 2024, doi: 10.1038/s41467-024-45427-4

  19. [28]

    Using Multimodal Large Language Models for Automated Detection of Traffic Safety-Critical Events,

    M. Abu Tami, H. I. Ashqar, M. Elhenawy, S. Glaser, and A. Rakotonirainy, “Using Multimodal Large Language Models for Automated Detection of Traffic Safety-Critical Events,” Vehicles, vol. 6, pp. 1571–1590, 2024, doi: 10.3390/vehicles6030074

  20. [29]

    VRU-Accident: A vision–language benchmark for video question answering and dense captioning for accident scene understanding,

    Y. Kim, A. S. Abdelrahman, and M. Abdel-Aty, “VRU-Accident: A vision–language benchmark for video question answering and dense captioning for accident scene understanding,” in Proc. IEEE/CVF ICCV Workshops, 2025, pp. 772–782, doi: 10.1109/ICCVW69036.2025.00085

  21. [30]

    SafePLUG: Empowering multimodal LLMs with pixel-level insight and temporal grounding for traffic accident understanding,

    Z. Sheng et al., “SafePLUG: Empowering multimodal LLMs with pixel-level insight and temporal grounding for traffic accident understanding,”CHAIN, vol. 3, no. 1, pp. 53–72, 2026, doi: 10.23919/CHAIN.2026.000005

  22. [31]

    Large Language Models in Analyzing Crash Narratives—A Comparative Study of ChatGPT, BARD and GPT-4,

    M. Mumtarin, M. S. Chowdhury, and J. Wood, “Large Language Models in Analyzing Crash Narratives—A Comparative Study of ChatGPT, BARD and GPT-4,” arXiv:2308.13563, 2023

  23. [32]

    Agentic Large Language Models for Day- to-Day Route Choices,

    L. Wang et al., “Agentic Large Language Models for Day- to-Day Route Choices,” Transportation Research Part C: Emerging Technologies, vol. 180, art. 105307, 2025, doi: 10.1016/j.trc.2025.105307

  24. [33]

    TransitGPT: A Generative AI- Based Framework for Interacting with GTFS Data Using Large Language Models,

    S. Devunuri and L. J. Lehe, “TransitGPT: A Generative AI- Based Framework for Interacting with GTFS Data Using Large Language Models,” Public Transport, vol. 17, no. 2, pp. 319–345, 2025, doi: 10.1007/s12469-025-00395-w

  25. [34]

    ChatSUMO: Large Language Model for Automating Traffic Scenario Generation in Simulation of Urban MObility,

    S. Li, T. Azfar, and R. Ke, “ChatSUMO: Large Language Model for Automating Traffic Scenario Generation in Simulation of Urban MObility,” IEEE Transactions on Intelligent Vehicles, vol. 10, no. 11, pp. 4962–4973, 2025, doi: 10.1109/TIV.2024.3508471

  26. [35]

    Speak to Simulate: An LLM-Guided Agentic Framework for Traffic Simulation in SUMO,

    M. Jeong, J. Chang, and Y. Yoon, “Speak to Simulate: An LLM-Guided Agentic Framework for Traffic Simulation in SUMO,” in Proc. 8th ACM SIGSPATIAL International Workshop on Geospatial Simulation, 2025, pp. 45–48, doi: 10.1145/3764921.3770151

  27. [36]

    Automating Traffic Model Enhancement With AI Research Agent,

    X. Guo, X. Yang, M. Peng, H. Lu, M. Zhu, and H. Yang, “Automating Traffic Model Enhancement With AI Research Agent,” Transportation Research Part C: Emerging Technologies, vol. 178, art. 105187, 2025, doi: 10.1016/j.trc.2025.105187

  28. [37]

    Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection,

    K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection,” in Proc. ACM Workshop on AI and Security, 2023, pp. 79–90, doi: 10.1145/3605764.3623985

  29. [38]

    Road Vehicles—Functional Safety—Part 1: Vocabulary,

    ISO, “Road Vehicles—Functional Safety—Part 1: Vocabulary,” ISO 26262-1:2018, 2018

  30. [39]

    Road Vehicles—Safety of the Intended Functionality,

    ISO, “Road Vehicles—Safety of the Intended Functionality,” ISO 21448:2022, 2022

  31. [40]

    Road Vehicles—Cybersecurity Engineering,

    ISO and SAE International, “Road Vehicles—Cybersecurity Engineering,” ISO/SAE 21434:2021, 2021

  32. [41]

    Road Vehicles—Safety and Artificial Intelligence,

    ISO, “Road Vehicles—Safety and Artificial Intelligence,” ISO/PAS 8800:2024, 2024

  33. [42]

    IEEE Standard for Transparency of Autonomous Sys- tems,

    IEEE, “IEEE Standard for Transparency of Autonomous Sys- tems,” IEEE Std 7001-2021, 2021

  34. [43]

    Artificial Intelligence Risk Management Framework (AI RMF 1.0),

    E. Tabassi, “Artificial Intelligence Risk Management Framework (AI RMF 1.0),” NIST AI 100-1, National Institute of Standards and Technology, 2023, doi: 10.6028/NIST.AI.100-1

  35. [44]

    Information Technology—Artificial Intelligence— Guidance on Risk Management,

    ISO/IEC, “Information Technology—Artificial Intelligence— Guidance on Risk Management,” ISO/IEC 23894:2023, 2023

  36. [45]

    Information Technology—Artificial Intelligence— Management System,

    ISO/IEC, “Information Technology—Artificial Intelligence— Management System,” ISO/IEC 42001:2023, 2023

  37. [47]

    Scalability in perception for autonomous driving: Waymo Open Dataset,

    P. Sun et al., “Scalability in perception for autonomous driving: Waymo Open Dataset,” inProc. IEEE/CVF CVPR, 2020, pp. 2443–2451, doi: 10.1109/CVPR42600.2020.00252

  38. [48]

    The highD Dataset: A Drone Dataset of Naturalistic Vehicle Trajectories on German Highways for ValidationofHighlyAutomatedDrivingSystems,

    R. Krajewski et al., “The highD Dataset: A Drone Dataset of Naturalistic Vehicle Trajectories on German Highways for ValidationofHighlyAutomatedDrivingSystems,”inProc.IEEE ITSC, 2018, pp. 2118–2125, doi: 10.1109/ITSC.2018.8569552

  39. [49]

    CityFlow: A Multi-Agent Reinforcement Learn- ing Environment for Large Scale City Traffic Scenario,

    H. Zhang et al., “CityFlow: A Multi-Agent Reinforcement Learn- ing Environment for Large Scale City Traffic Scenario,” in Proc. WWW, 2019, pp. 3620–3624, doi: 10.1145/3308558.3314139

  40. [50]

    Microscopic Traffic Simulation Using SUMO,

    P. A. Lopez et al., “Microscopic Traffic Simulation Using SUMO,” in Proc. IEEE ITSC, 2018, pp. 2575–2582, doi: 10.1109/ITSC.2018.8569938

  41. [51]

    Large models for intelligent transportation systems and autonomous vehicles: A survey,

    L. Gan, W. Chu, G. Li, X. Tang, and K. Li, “Large models for intelligent transportation systems and autonomous vehicles: A survey,”Advanced Engineering Informatics, vol. 62, art. 102786, 2024, doi: 10.1016/j.aei.2024.102786

  42. [52]

    TrafficMind: A system- oriented review of large language models for intelligent trans- portation systems,

    C. Zhang, B. Wei, and L. Yang, “TrafficMind: A system- oriented review of large language models for intelligent trans- portation systems,”Journal of Transportation Engineering, Part A: Systems, vol. 152, no. 9, art. 03126005, 2026, doi: 10.1061/JTEPBS.TEENG-9794

  43. [53]

    Foundation models for autonomous driving: A comprehensive survey,

    S. Fourati, W. Jaafar, N. Baccar, S. Alfattani, and R. Langar, “Foundation models for autonomous driving: A comprehensive survey,”Engineering Applications of Artificial Intelligence, vol. 176, art. 114805, 2026, doi: 10.1016/j.engappai.2026.114805

  44. [54]

    The role of large language models (LLMs) in enhancing intelligent transportation systems: A survey,

    V. Hassija, T. Majumder, D. Roy, R. Piyush, and V. Chamola, “The role of large language models (LLMs) in enhancing intelligent transportation systems: A survey,”Ve- hicular Communications, vol. 58, art. 100996, 2026, doi: 10.1016/j.vehcom.2025.100996

  45. [55]

    Integrating LLMs with ITS: Recent advances, potentials, challenges, and future directions,

    D. Mahmud, H. Hajmohamed, S. Almentheri, S. Alqaydi, L. Ald- haheri, R. A. Khalil, and N. Saeed, “Integrating LLMs with ITS: Recent advances, potentials, challenges, and future directions,” IEEE Transactions on Intelligent Transportation Systems, vol. 26, no. 5, pp. 5674–5709,...

  46. [56]

    DiLu: A Knowledge-Driven Approach to Au- tonomous Driving with Large Language Models,

    L. Wen et al., “DiLu: A Knowledge-Driven Approach to Au- tonomous Driving with Large Language Models,” inProc. Int. Conf. Learning Representations (ICLR), 2024. 18

  47. [57]

    Dolphins: Multimodal Language Model for Driving,

    Y. Ma, Y. Cao, J. Sun, M. Pavone, and C. Xiao, “Dolphins: Multimodal Language Model for Driving,” inComputer Vision– ECCV 2024, 2024, doi: 10.1007/978-3-031-72995-9_23

  48. [58]

    Driving Everywhere with Large Language Model PolicyAdaptation,

    B. Li et al., “Driving Everywhere with Large Language Model PolicyAdaptation,”inProc.IEEE/CVFCVPR,2024,pp.14948– 14957

  49. [59]

    LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs,

    Y. Ma et al., “LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs,” inProc. IEEE/CVF CVPR, 2024, pp. 15141–15151

  50. [60]

    ChatScene: Knowledge-Enabled Safety-Critical Scenario Generation for Autonomous Vehicles,

    J. Zhang, C. Xu, and B. Li, “ChatScene: Knowledge-Enabled Safety-Critical Scenario Generation for Autonomous Vehicles,” inProc. IEEE/CVF CVPR, 2024, pp. 15459–15469

  51. [61]

    Editable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents,

    Y. Wei et al., “Editable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents,” inProc. IEEE/CVF CVPR, 2024, pp. 15077–15087

  52. [62]

    LLMScenario: Large Language Model Driven Scenario Generation,

    C. Chang, S. Wang, J. Zhang, J. Ge, and L. Li, “LLMScenario: Large Language Model Driven Scenario Generation,”IEEE Trans. Syst., Man, Cybern.: Syst., vol. 54, no. 11, pp. 6581–6594, 2024, doi: 10.1109/TSMC.2024.3392930

  53. [63]

    DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models,

    X. Tian et al., “DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models,” inProc. 8th Conf. Robot Learning, PMLR, vol. 270, pp. 4698–4726, 2025

  54. [64]

    DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving,

    Z. Xu et al., “DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving,” inProc. IEEE/CVF CVPR, 2025, pp. 17261–17270

  55. [65]

    ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Gen- eration,

    H. Fu et al., “ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Gen- eration,” inProc. IEEE/CVF ICCV, 2025, pp. 24823–24834

  56. [66]

    GATSim: Urban Mobility Simula- tion with Generative Agents,

    Q. Liu, C. Li, and W. Ma, “GATSim: Urban Mobility Simula- tion with Generative Agents,”Transportation Research Part C: Emerging Technologies, vol. 186, art. 105576, 2026, doi: 10.1016/j.trc.2026.105576

  57. [67]

    AITP: Traffic Accident Responsibility Allocation via Multimodal Large Language Models,

    Z. Zhou and S. Zhang, “AITP: Traffic Accident Responsibility Allocation via Multimodal Large Language Models,” inProc. IEEE/CVF CVPR Findings, 2026, pp. 9259–9268

  58. [68]

    MindDriver: Introducing Progressive Multi- modal Reasoning for Autonomous Driving,

    L. Zhang et al., “MindDriver: Introducing Progressive Multi- modal Reasoning for Autonomous Driving,” inProc. IEEE/CVF CVPR, 2026, pp. 17831–17841

  59. [69]

    V2X-UniPool: Unifying Multimodal Perception and Knowledge Reasoning for Autonomous Driving,

    X. Luo et al., “V2X-UniPool: Unifying Multimodal Perception and Knowledge Reasoning for Autonomous Driving,” inProc. IEEE/CVF CVPR Workshops, 2026, pp. 747–756

  60. [70]

    Large Language Model as Parking Planning Agent in the Context of Mixed Period of Autonomous Vehicles and Human-Driven Vehicles,

    Y. Jin and J. Ma, “Large Language Model as Parking Planning Agent in the Context of Mixed Period of Autonomous Vehicles and Human-Driven Vehicles,”Sustainable Cities and Society, vol. 117, art. 105940, 2024, doi: 10.1016/j.scs.2024.105940

  61. [71]

    Bridging AI and Traffic Simulation: A Robust and Comprehensive Framework for LLM-Based AI Replanning Agents in MATSim,

    A. U. Z. Patwary et al., “Bridging AI and Traffic Simulation: A Robust and Comprehensive Framework for LLM-Based AI Replanning Agents in MATSim,”Procedia Computer Science, vol. 280, pp. 622–629, 2026, doi: 10.1016/j.procs.2026.04.079

  62. [72]

    Agentic Traffic Intelligence: Augmented Human-in-the- Loop Scenario Generation for Microscopic Traffic Simulation,

    X. Luo, G. Xu, A. Saroj, J. Yuan, P. Kadav, Y. Shao, and C. R. Wang, “Agentic Traffic Intelligence: Augmented Human-in-the- Loop Scenario Generation for Microscopic Traffic Simulation,” Artificial Intelligence for Transportation, vol. 6, art. 100057, 2026, doi: 10.1016/j.ait.2...

  63. [73]

    Large Language Model-Assisted Multi-Objective Optimization for an Integrated Multimodal E-Mobility Platform,

    Y. Ding, M. Maniparambil, N. E. O’Connor, and M. Liu, “Large Language Model-Assisted Multi-Objective Optimization for an Integrated Multimodal E-Mobility Platform,”Transportation ResearchInterdisciplinaryPerspectives,vol.37,art.101948,2026, doi: 10.1016/j.trip.2026.101948

  64. [74]

    LLMs as Virtual Traffic Police: Incident-Aware Traffic Signal Control Augmented by Large Language Models,

    Q. Wang, S. Wei, and K. Yang, “LLMs as Virtual Traffic Police: Incident-Aware Traffic Signal Control Augmented by Large Language Models,” inProc. IEEE 28th Int. Conf. Intell. Transp. Syst. (ITSC), 2025, doi: 10.1109/ITSC60802.2025.11423524

  65. [75]

    An Efficient Simulation Scene Generation Method BasedonExtractedRoadNetworkTopologyandLargeLanguage Models,

    R. Li et al., “An Efficient Simulation Scene Generation Method BasedonExtractedRoadNetworkTopologyandLargeLanguage Models,”Future Transportation, vol. 6, no. 2, art. 81, 2026, doi: 10.3390/futuretransp6020081

  66. [76]

    DriveGPT4:InterpretableEnd-to-EndAutonomous Driving Via Large Language Model,

    Z.Xuetal.,“DriveGPT4:InterpretableEnd-to-EndAutonomous Driving Via Large Language Model,”IEEE Robotics and Au- tomation Letters, 2024, doi: 10.1109/LRA.2024.3440097

  67. [77]

    ChatSUMO Agent: An LLM- Based Agent for Conversational Traffic Simulation in SUMO,

    S. Li, M. Ma, T. Azfar, and R. Ke, “ChatSUMO Agent: An LLM- Based Agent for Conversational Traffic Simulation in SUMO,” SSRN preprint, 1 Jan. 2026, doi: 10.2139/ssrn.6000335

  68. [78]

    Generalizing End-to-End Autonomous Driving in Real-World Environments Using Zero-Shot LLMs,

    Z. Dong, Y. Zhu, Y. Li, K. Mahon, and Y. Sun, “Generalizing End-to-End Autonomous Driving in Real-World Environments Using Zero-Shot LLMs,” inProc. 8th Conf. Robot Learning, PMLR, vol. 270, pp. 1231–1249, 2025. [Online]. Available: https: //proceedings.mlr.press/v270/dong25a.html

  69. [79]

    Automating the Loop in Traffic Incident Manage- ment on Highway,

    M. Cercola, N. Gatti, P. Huertas Leyva, B. Carambia, and S. Formentin, “Automating the Loop in Traffic Incident Manage- ment on Highway,” inProc. 7th Annu. Learning for Dynamics and Control Conf., PMLR, vol. 283, pp. 272–284, 2025. [Online]. Available: https://proceedings.mlr....

  70. [80]

    Promptable Closed-Loop Traffic Simulation,

    S. Tan, B. Ivanovic, Y. Chen, B. Li, X. Weng, Y. Cao, P. Kraehenbuehl, and M. Pavone, “Promptable Closed-Loop Traffic Simulation,” inProc. 8th Conf. Robot Learning, PMLR, vol. 270, pp. 5087–5105, 2025. [Online]. Available: https://proceedings.ml r.press/v270/tan25a.html

  71. [81]

    DriveCombo: Benchmarking Compositional Traffic Rule Reasoning in Autonomous Driving,

    E. Ma et al., “DriveCombo: Benchmarking Compositional Traffic Rule Reasoning in Autonomous Driving,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 32113–32123,

  72. [2026]

    Available: https://openaccess.thecvf.com/conten t/CVPR2026/html/Ma_DriveCombo_Benchmarking_Compo sitional_Traffic_Rule_Reasoning_in_Autonomous_Driving_ CVPR_2026_paper.html

    [Online]. Available: https://openaccess.thecvf.com/conten t/CVPR2026/html/Ma_DriveCombo_Benchmarking_Compo sitional_Traffic_Rule_Reasoning_in_Autonomous_Driving_ CVPR_2026_paper.html

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.