Pith. sign in

REVIEW 4 major objections 6 minor 109 references

This survey argues that scenario-generation frameworks for automated driving can be benchmarked transparently with a three-axis metric suite covering scholarly impact, reproducibility, and ODD coverage.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 15:49 UTC pith:5LQR55MK

load-bearing objection A useful newcomer-oriented survey whose proposed AII/RAS/OCS benchmark is not reproducible as written; the worked-example numbers don't follow from its own formulas. the 4 major comments →

arxiv 2512.15422 v2 pith:5LQR55MK submitted 2025-12-17 cs.SE

A Survey on the Applications of Generative Artificial Intelligence in Automated Driving Systems Test Scenario Generation Methods

classification cs.SE
keywords automated driving systemsscenario-based testingscenario generationgenerative AIlarge language modelsdiffusion modelsODD coveragereproducible metrics
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper synthesises 31 primary studies and 10 surveys on scenario-based testing for automated driving and argues that the field's main obstacle is the absence of standard evaluation metrics. Its central proposal is a three-axis metric suite—academic influence, resource accessibility, and ODD coverage—with explicit formulas and evidence-based scoring rubrics. The authors apply the suite to two recent frameworks to show how transparent benchmarking would work. If the metrics hold up, researchers gain a common yardstick for comparing scenario generators, and industry gains a way to report test coverage reproducibly.

Core claim

The paper's core claim is that a reproducible, three-axis metric system—comprising the Academic Influence Index (AII), the Resource Accessibility Score (RAS), and the ODD Coverage Score (OCS)—can make scenario-generation frameworks for automated driving comparable. The authors define formulas and a tiered evidence strategy, then apply the suite to two 2025 frameworks, producing numeric scores and star ratings. They further contribute a multimodal taxonomy, an ethical and safety checklist, and an ODD coverage map with a scenario-difficulty schema, all intended as foundations for standardised benchmarking.

What carries the argument

The three-axis metric suite: AII weights citation impact, early uptake, and author h-index; RAS scores code availability, runnable pipeline, environment transparency, model access, and documentation on a 0/0.5/1 scale with a dataset bonus; OCS averages five ODD dimensions (road type, VRU presence, topology, interaction complexity, controllability) under an evidence-based rubric. An ethical checklist and a three-tier scenario-difficulty schema (benign, conflict-prone, safety-critical) anchored to TTC and gap thresholds support the broader framework.

Load-bearing premise

The whole scoring system assumes a framework's ODD coverage and reproducibility can be reliably read off from its paper's text and figures using the authors' translation rules and chosen weights, without any independent verification of those readings.

What would settle it

Take the paper's own two worked examples—or any set of frameworks—and have several independent raters apply the Appendix C rubric. If their OCS scores diverge enough to change the ranking order (say, by more than one star tier for the same framework), the claim that the metrics give reproducible, comparable results is false. A simpler test: ask whether TrafficComposer's reported 62% OCS would survive if 'weather can be modified' statements were discounted under the paper's own downgrade policy.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If adopted, the metrics give scenario-generation researchers a common language for reporting reproducibility and ODD coverage, replacing ad-hoc qualitative claims.
  • The OCS rubric can be applied to future frameworks as they appear, allowing the field's progress on multimodal and ODD-specific scenarios to be tracked over time.
  • The ethical checklist translates principles like bias mitigation and privacy into auditable criteria, making it possible to screen frameworks for blind spots before adoption.
  • The scenario-difficulty schema with TTC/gap anchors offers a ready template for classifying test scenarios by risk, supporting standardised reporting across studies.
  • The survey's finding that most frameworks cluster around structured urban/highway ODDs points to a concrete gap: generators targeting adverse weather, rural roads, and complex VRU interactions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The scoring rubric's reliance on manual translation of qualitative paper statements could be automated with a vision-language model that reads figures and text, which would also give the inter-rater reliability the authors do not provide.
  • If OCS-style reporting became standard, regulators could require it in safety submissions, making coverage claims auditable under SOTIF/UN R157 rather than self-assessed.
  • The difficulty schema's TTC/gap thresholds could be calibrated against logged fleet data to produce empirically grounded tiers, a step beyond the paper's illustrative anchors.
  • The survey's methodological shift toward 2023-2025 AI-assisted methods suggests that older rule-based and GAN approaches may be under-represented in current benchmarks, so the comparison table should be read as a snapshot of a fast-moving field.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This survey reviews scenario generation methods for automated driving systems (ADS) testing over 2015–2025, with a focus on 2023–2025 AI-assisted approaches. It organizes methods into traditional and AI-assisted families (LLM, GAN, diffusion, RL, etc.), synthesizes 31 primary studies and 10 surveys, and proposes three contributions: a refined taxonomy with multimodal extensions, an ethical and safety checklist, and an ODD coverage map with a scenario-difficulty schema. The central methodological novelty is a three-axis metric system (AII, RAS, OCS) claimed to enable reproducible and transparent benchmarking, demonstrated on Genesis and TrafficComposer in Section 6.1.4 and Table 4.

Significance. If the three-axis metric system were sound, it would address a genuine gap in ADS scenario-generation benchmarking: the field lacks standardized, auditable evaluation metrics. The survey synthesis and the proposed ethical/ODD schemas are useful and reasonably current. However, the quantitative core—the AII/RAS/OCS suite and its application—contains multiple internal inconsistencies that, as written, invalidate the reproducibility claim. The paper is commendable for including explicit limitations and conservative evidence hooks, but the central contribution needs substantial correction and validation before the claimed reproducibility is credible.

major comments (4)
  1. [§6.1.3, Eq. (3)] OCS = 100 × (1/5 Σ E·a_k) with Appendix C weights E = [0.20,0.15,0.15,0.25,0.25] summing to 1 and a_k∈[0,1] caps the maximum at 20, not 100. Yet §6.1.4 reports OCS 39.5%, 62.0%, 31.0%, 54.0%, and the star mapping in §6.1.3 assumes a 0–100 scale. Remove the spurious 1/5 (or redefine E) and recompute all OCS values; the current formula cannot produce any score above 20.
  2. [§6.1.4] The worked examples are internally inconsistent. For TrafficComposer, VRU = 0.6·(1/2) + 0.4·(2/5) = 0.46, not "0.38 → 0.35"; the stated 62.0% average is not recoverable from the listed dimension scores under any consistent reading of the rubric (~64% weighted, ~65.2% unweighted). Genesis's 39.5% is also not reproducible (≈40% weighted, ≈41% unweighted). The same subsection later reports OCS 31.0% and 54.0% without explanation, and Table 4 gives AII 35.6%/2.7% versus the text's 30.0%/7.2%. These contradictions make the central benchmark unverifiable.
  3. [§6.1.4 / Table 4] The full metric suite is applied only to Genesis and TrafficComposer; the star ratings for the remaining ~25 frameworks in Table 4 have no disclosed component scores or evidence hooks. The claim of "transparent benchmarking" (§3.3) is therefore unsupported for most of the table. Provide per-framework AII/RAS/OCS component scores and evidence citations, or explicitly mark Table 4 as a qualitative summary.
  4. [§6.1.4 / Appendix C] The OCS mapping from qualitative paper statements to numeric scores relies on the authors' "conservative" judgment; no inter-rater reliability, sensitivity analysis, or external calibration is provided. The manuscript itself acknowledges in §7.6 that "current application is based on inferred evidence rather than standardized reporting." This conflicts with the "Reproducible Three-Axis Metric System" claim in §3.3. Either supply a validation protocol or soften the claim to "proposed rubric with illustrative application."
minor comments (6)
  1. [§3.1.4] The Boolean query string mixes AND/OR without parentheses; the intended grouping is ambiguous and should be made explicit.
  2. [§6.1.4] "Eq.8.1" should be "Eq. (1)".
  3. [§7 title] "Metrix Extension" should be "Metric Extension".
  4. [Table 4] The star ratings for TrafficComposer (OCS ★★★) and Genesis (OCS ★★) conflict with the text's numeric OCS values (31%/54% or 39.5%/62%). Please reconcile the scale and thresholds used in the table.
  5. [Appendix C / §6.1.4] The dimension order in §6.1.4 (road, VRU, topology, interaction, controllability) differs from Table 9's order; specify which weight aligns with which dimension to make the calculation reproducible.
  6. [References] Reference [43] shows an incomplete title ("c"); several references lack DOIs or page numbers (e.g., [7], [16]).

Circularity Check

1 steps flagged

No significant circularity; the metric suite is independently defined and applied, with one minor self-citation for the taxonomy.

specific steps
  1. self citation load bearing [Section 5.2.1 and reference [23]]
    "We adopted the taxonomy in [23] and listed frameworks in following sub-titles. ... [23] Y. Zhao, J. Zhou, D. Bi, T. Mihalj, J. Hu, and A. Eichberger, 'A Survey on the Application of Large Language Models in Scenario-Based Testing of Automated Driving Systems,' arXiv:2505.16587."

    The classification structure for LLM-based scenario generation is explicitly taken from the authors' own prior arXiv survey, whose author list overlaps with the present paper. Because one of the paper's stated contributions is 'a refined taxonomy that incorporates multimodal extensions,' this self-citation is load-bearing for that part of the contribution, and its authority is not independently established in the present text. This is mild self-reference rather than full circularity: the adoption is transparent, and the AII/RAS/OCS metric derivations are defined by the paper's own formulas and applied to external frameworks, not derived from [23].

full rationale

The central derivation chain is a survey synthesis plus a self-defined evaluation rubric applied to external frameworks. The AII, RAS, and OCS metrics are each defined by explicit equations and scoring rules, then applied to publicly documented features of the compared papers; they are not fitted from the outcomes they rank, so the main comparison is not circular in the derivation-equals-input sense. The only genuine self-reference is the explicit adoption of the LLM taxonomy from the authors' own prior survey [23], which is transparent and does not support the metric-system derivation. The paper itself acknowledges in Section 7.6 that the ODD coverage map and scenario-difficulty schema 'are conceptual tools that require further empirical validation' and are 'based on inferred evidence rather than standardized reporting,' which lowers any concern of overclaimed independent validation. The reported OCS arithmetic inconsistencies (e.g., Eq. (3) with Appendix C weights yields a maximum of 20 while Section 6.1.4 reports 39.5% and 62.0%) are serious reproducibility defects and internal-consistency failures, but they are not circular reductions: the scores are not equal to their inputs by construction, they simply fail to match the paper's own stated formula. Overall, no significant circularity; the score reflects the one minor self-citation.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 0 invented entities

The central claim rests on hand-chosen weights, discrete rubric anchors, and subjective translations of qualitative evidence. There are no new physical entities. The free parameters are concentrated in the AII/RAS/OCS metric design and the ODD-difficulty thresholds; several are applied inconsistently in the worked examples.

free parameters (6)
  • AII weighting coefficients = 0.4 (C_norm), 0.3 (R_early), 0.3 (H_mean)
    Hand-chosen weights in Eq. (1); no derivation or sensitivity analysis. They determine every AII score.
  • RAS component weights and dataset bonus = 30/25/20/15/10; B ∈ {0,2,5}
    Fixed weights in Eq. (2) and Table 2; chosen by the authors, with no external justification or tested stability.
  • OCS dimension weights E = [0.20, 0.15, 0.15, 0.25, 0.25]
    Hand-assigned in Appendix C.1 to prioritize VRU presence and topology; no empirical basis or sensitivity analysis.
  • OCS scoring anchors and star thresholds = Dimension scores 0.10/0.25/0.50/0.75/1.00; star thresholds <40, 40–59, 60–74, 75–89, ≥90
    Arbitrary discrete scales in Appendix C and Table 3; chosen to make five-star mapping, not derived from data.
  • VRU presence formula coefficients = 0.6×(types/2) + 0.4×(behaviors/5)
    Ad hoc formula with arbitrary denominators; the TrafficComposer application is arithmetically wrong (inputs give 0.46, not 0.38).
  • Scenario-difficulty thresholds = TTC 5 s / 2 s; min gap 20 m / 10 m; agents 3 / 4–8 / ≥9
    Defined in Appendix D without calibration; the paper states the schema is 'not directly executed' and requires empirical validation.
axioms (6)
  • domain assumption The four-database search and broad inclusion criteria yield a representative corpus of ADS scenario-generation literature.
    Survey conclusions about gaps and trends depend on corpus representativeness; 928 of 1050 papers were excluded at title/abstract screening without visible inter-rater checks (Section 3.2).
  • domain assumption Scenario-based testing is an acceptable proxy for ADS safety validation.
    The entire review presupposes SBT is the right validation path, citing Kalra & Paddock, SOTIF, and UN R157; it does not critically evaluate that premise (Sections 1 and 4.1).
  • standard math Standard ML/generative-model behavior claims (LLM hallucination, GAN mode collapse, diffusion training needs) are as described by the cited literature.
    The review accepts prior empirical results as ground truth; no new derivations or verification are provided (Section 5.2).
  • ad hoc to paper Qualitative claims in papers can be mapped to numeric OCS scores under the authors' 'conservative' translation rules.
    This mapping is the load-bearing assumption behind Table 4; no inter-rater validation or error bounds are given (Section 6.1.4, Appendix C).
  • ad hoc to paper The hand-picked weights in AII/RAS/OCS are valid measurement choices.
    No external benchmarks or sensitivity tests are provided; the weights determine the rankings in Table 4 (Sections 6.1.1–6.1.3).
  • ad hoc to paper TTC/gap thresholds in the scenario-difficulty schema are meaningful without calibration.
    The paper itself says the schema is 'not directly executed' and needs validation; the thresholds are asserted, not fitted (Section 7.3, Appendix D).

pith-pipeline@v1.3.0-alltime-deepseek · 31778 in / 18112 out tokens · 180477 ms · 2026-08-03T15:49:11.777698+00:00 · methodology

0 comments
read the original abstract

Ensuring the safety and reliability of Automated Driving Systems (ADS) remains a critical challenge, as traditional verification methods such as large-scale on-road testing are prohibitively costly and time-consuming.To address this,scenario-based testing has emerged as a scalable and efficient alternative,yet existing surveys provide only partial coverage of recent methodological and technological advances.This review systematically analyzes 31 primary studies,and 10 surveys identified through a comprehensive search spanning 2015~2025;however,the in-depth methodological synthesis and comparative evaluation focus primarily on recent frameworks(2023~2025),reflecting the surge of Artificial Intelligent(AI)-assisted and multimodal approaches in this period.Traditional approaches rely on expert knowledge,ontologies,and naturalistic driving or accident data,while recent developments leverage generative models,including large language models,generative adversarial networks,diffusion models,and reinforcement learning frameworks,to synthesize diverse and safety-critical scenarios.Our synthesis identifies three persistent gaps:the absence of standardized evaluation metrics,limited integration of ethical and human factors,and insufficient coverage of multimodal and Operational Design Domain (ODD)-specific scenarios.To address these challenges,this review contributes a refined taxonomy that incorporates multimodal extensions,an ethical and safety checklist for responsible scenario design,and an ODD coverage map with a scenario-difficulty schema to enable transparent benchmarking.Collectively,these contributions provide methodological clarity for researchers and practical guidance for industry,supporting reproducible evaluation and accelerating the safe deployment of higher-level ADS.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

109 extracted references · 31 canonical work pages · 10 internal anchors

  1. [1]

    Testing of autonomous driving systems: where are we and where should we go?,

    G. Lou, Y . Deng, X. Zheng, M. Zhang, and T. Zhang, “Testing of autonomous driving systems: where are we and where should we go?,” in Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, Singapore Singapore: ACM, Nov. 2022, pp. 31 –43. doi: 10.1145/3540250.3549111

  2. [2]

    [Online]

    ISO, Road vehicles — Safety of the intended functionality, 21448:2022, June 2022. [Online]. Available: https://www.iso.org/standard/77490.html

  3. [3]

    157 - Automated Lane Keeping Systems (ALKS), R157

    UN Regulation No. 157 - Automated Lane Keeping Systems (ALKS), R157

  4. [4]

    doi: 10.4271/J3016_202104

    On-Road Automated Driving (ORAD) committee, Taxonomy and Definitions for Terms Related to Driving Automation Systems for On -Road Motor Vehicles. doi: 10.4271/J3016_202104

  5. [5]

    Testing of Advanced Driver Assistance Towards Automated Driving: A Survey and Taxonomy on Existing Approaches and Open Questions,

    J. E. Stellet, M. R. Zofka, J. Schumacher, T. Schamm, F. Niewels, and J. M. Zollner, “Testing of Advanced Driver Assistance Towards Automated Driving: A Survey and Taxonomy on Existing Approaches and Open Questions,” in 2015 IEEE 18th International Conference on Intelligent Transportation Systems, Gran Canaria, Spain: IEEE, Sept. 2015, pp. 1455 –1462. doi...

  6. [6]

    Driving to safety: How many miles of driving would it take to demonstrate autonomous vehicle reliability?,

    N. Kalra and S. M. Paddock, “Driving to safety: How many miles of driving would it take to demonstrate autonomous vehicle reliability?,” Transp. Res. Part Policy Pract. , vol. 94, pp. 182 – 193, Dec. 2016, doi: 10.1016/j.tra.2016.09.010

  7. [7]

    System validation of highly automated vehicles with a database of relevant traffic scenarios

    A. Pütz, A. Zlocki, J. Bock, and L. Eckstein, “System validation of highly automated vehicles with a database of relevant traffic scenarios”

  8. [8]

    PEGASUS—First Steps for the Safe Introduction of Automated Driving,

    H. Winner, K. Lemmer, T. Form, and J. Mazzega, “PEGASUS—First Steps for the Safe Introduction of Automated Driving,” in Road Vehicle Automation 5 , G. Meyer and S. Beiker, Eds., in Lecture Notes in Mobility. , Cham: Springer International Publishing, 2019, pp. 185 –195. doi: 10.1007/978-3-319-94896-6_16

  9. [9]

    Waymo Open Dataset . (Sept. 16, 2025). Python. Waymo Research. Accessed: Sept. 17, 2025. [Online]. Available: https://github.com/waymo - research/waymo-open-dataset

  10. [10]

    GAIA-1: A Generative World Model for Autonomous Driving,

    A. Hu et al., “GAIA-1: A Generative World Model for Autonomous Driving,” Sept. 29, 2023, arXiv: arXiv:2309.17080. doi: 10.48550/arXiv.2309.17080

  11. [12]

    SEAL: Towards Safe Autonomous Driving via Skill - Enabled Adversary Learning for Closed -Loop Scenario Generation,

    B. Stoler, I. Navarro, J. Francis, and J. Oh, “SEAL: Towards Safe Autonomous Driving via Skill - Enabled Adversary Learning for Closed -Loop Scenario Generation,” IEEE Robot. Autom. Lett. , vol. 10, no. 9, pp. 9320 –9327, Sept. 2025, doi: 10.1109/LRA.2025.3592090

  12. [13]

    Generating Multimodal Driving Scenes via Next-Scene Prediction,

    Yanhao Wu et al., “Generating Multimodal Driving Scenes via Next-Scene Prediction,” Mar. 26, 2025, arXiv: arXiv:2503.14945. doi: 10.48550/arXiv.2503.14945

  13. [14]

    Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency,

    X. Guo et al., “Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency,” June 20, 2025, arXiv: arXiv:2506.07497. doi: 10.48550/arXiv.2506.07497

  14. [15]

    Txt2Sce: Scenario Generation for Autonomous Driving System Testing Based on Textual Reports,

    P. Ji et al. , “Txt2Sce: Scenario Generation for Autonomous Driving System Testing Based on Textual Reports,” Sept. 02, 2025, arXiv: arXiv:2509.02150. doi: 10.48550/arXiv.2509.02150

  15. [16]

    Understanding L2+ in Five Questions,

    “Understanding L2+ in Five Questions,” Mobileye. Accessed: Nov. 06, 2025. [Online]. Available: https://www.mobileye.com/blog/understanding-l2- in-five-questions/

  16. [17]

    Scenario-Based Test Automation for Highly Automated Vehicles: A Review and Paving the Way for Systematic Safety Assurance,

    J. Sun, H. Zhang, H. Zhou, R. Yu, and Y . Tian, “Scenario-Based Test Automation for Highly Automated Vehicles: A Review and Paving the Way for Systematic Safety Assurance,” IEEE Trans. Intell. Transp. Syst. , vol. 23, no. 9, pp. 14088 – 14103, Sept. 2022, doi: 10.1109/TITS.2021.3136353

  17. [18]

    A Survey on Data -Driven Scenario Generation for Automated Vehicle Testing,

    J. Cai, W. Deng, H. Guang, Y . Wang, J. Li, and J. Ding, “A Survey on Data -Driven Scenario Generation for Automated Vehicle Testing,” Machines, vol. 10, no. 11, Art. no. 11, Nov. 2022, doi: 10.3390/machines10111101

  18. [19]

    A Survey on Safety -Critical Driving Scenario Generation —A Methodological Perspective,

    W. Ding, C. Xu, M. Arief, H. Lin, B. Li, and D. Zhao, “A Survey on Safety -Critical Driving Scenario Generation —A Methodological Perspective,” IEEE Trans. Intell. Transp. Syst., vol. 24, no. 7, pp. 6971 –6988, July 2023, doi: 10.1109/TITS.2023.3259322

  19. [20]

    1001 Ways of Scenario Generation for Testing of Self-driving Cars: A Survey

    B. Schütt, J. Ransiek, T. Braun, and E. Sax, “1001 Ways of Scenario Generation for Testing of Self - driving Cars: A Survey,” 2023, arXiv. doi: 10.48550/ARXIV .2304.10850

  20. [22]

    LLM4Drive: A Survey of Large Language Models for Autonomous Driving,

    Z. Yang, X. Jia, H. Li, and J. Yan, “LLM4Drive: A Survey of Large Language Models for Autonomous Driving,” Aug. 12, 2024, arXiv: arXiv:2311.01043. doi: 10.48550/arXiv.2311.01043

  21. [23]

    A Survey on the Application of Large Language Models in Scenario -Based Testing of Automated Driving Systems,

    Y . Zhao, J. Zhou, D. Bi, T. Mihalj, J. Hu, and A. Eichberger, “A Survey on the Application of Large Language Models in Scenario -Based Testing of Automated Driving Systems,” May 22, 2025, arXiv: arXiv:2505.16587. doi: 10.48550/arXiv.2505.16587

  22. [24]

    Foundation Models in Autonomous Driving: A Survey on Scenario Generation and Scenario Analysis,

    Y . Gao et al., “Foundation Models in Autonomous Driving: A Survey on Scenario Generation and Scenario Analysis,” June 13, 2025, arXiv: arXiv:2506.11526. doi: 10.48550/arXiv.2506.11526

  23. [25]

    A Review of Large Language Models for Automated Test Case Generation,

    A. Celik and Q. H. Mahmoud, “A Review of Large Language Models for Automated Test Case Generation,” Mach. Learn. Knowl. Extr., vol. 7, no. 3, p. 97, Sept. 2025, doi: 10.3390/make7030097

  24. [26]

    Scenario Factory 2.0: Scenario -Based Testing of Automated Vehicles with CommonRoad,

    F. Finkeldei, C. Thees, J. -N. Weghorn, and M. Althoff, “Scenario Factory 2.0: Scenario -Based Testing of Automated Vehicles with CommonRoad,” Automot. Innov., vol. 8, no. 2, pp. 207 –220, May 2025, doi: 10.1007/s42154-025-00360-0

  25. [27]

    A Survey on Automated Driving System Testing: Landscapes and Trends,

    S. Tang et al. , “A Survey on Automated Driving System Testing: Landscapes and Trends,” ACM Trans. Softw. Eng. Methodol., vol. 32, no. 5, pp. 1– 62, Sept. 2023, doi: 10.1145/3579642

  26. [28]

    On-Demand Scenario Generation for Testing Automated Driving Systems,

    S. Yan et al., “On-Demand Scenario Generation for Testing Automated Driving Systems,” Proc. ACM Softw. Eng., vol. 2, no. FSE, pp. 86–105, June 2025, doi: 10.1145/3715722

  27. [29]

    Ontology learning towards expressiveness: A survey,

    P. Armary, C. B. El-Vaigh, O. Labbani Narsis, and C. Nicolle, “Ontology learning towards expressiveness: A survey,” Comput. Sci. Rev. , vol. 56, p. 100693, May 2025, doi: 10.1016/j.cosrev.2024.100693

  28. [30]

    An ontology -based text mining dataset for extraction of process -structure-property entities,

    A. R. Durmaz, A. Thomas, L. Mishra, R. N. Murthy, and T. Straub, “An ontology -based text mining dataset for extraction of process -structure-property entities,” Sci. Data , vol. 11, no. 1, p. 1112, Oct. 2024, doi: 10.1038/s41597-024-03926-5

  29. [31]

    Ontology -Driven Automated Reasoning About Property Crimes,

    F. Navarrete, Á. L. Garrido, C. Bobed, M. Atencia, and A. Vallecillo, “Ontology -Driven Automated Reasoning About Property Crimes,” Bus. Inf. Syst. Eng., Aug. 2024, doi: 10.1007/s12599 -024-00886- 3

  30. [32]

    Ontology - Based Driving Simulation for Traffic Lights Optimization,

    A. Zaji, Z. Liu, T. Bando, and L. Zhao, “Ontology - Based Driving Simulation for Traffic Lights Optimization,” ACM Trans. Intell. Syst. Technol. , vol. 14, no. 3, pp. 1 –26, June 2023, doi: 10.1145/3579839

  31. [33]

    Concept Paper for a Digital Expert: Systematic Derivation of (Causal) Bayesian Networks Based on Ontologies for Knowledge - Based Production Steps,

    M. M. -L. Pfaff -Kastner, K. Wenzel, and S. Ihlenfeldt, “Concept Paper for a Digital Expert: Systematic Derivation of (Causal) Bayesian Networks Based on Ontologies for Knowledge - Based Production Steps,” Mach. Learn. Knowl. Extr., vol. 6, no. 2, pp. 898 –916, Apr. 2024, doi: 10.3390/make6020042

  32. [34]

    A novel ML-MCDM-based decision support system for evaluating autonomous vehicle integration scenarios in Geneva’s public transportation,

    Shervin Zakeri, D. Konstantas, S. Sorooshian, and P. Chatterjee, “A novel ML-MCDM-based decision support system for evaluating autonomous vehicle integration scenarios in Geneva’s public transportation,” Artif. Intell. Rev., vol. 57, no. 11, p. 310, Sept. 2024, doi: 10.1007/s10462-024-10917-w

  33. [35]

    The Ko -PER intersection laserscanner and video dataset,

    E. Strigel, D. Meissner, F. Seeliger, B. Wilking, and K. Dietmayer, “The Ko -PER intersection laserscanner and video dataset,” in 17th International IEEE Conference on Intelligent Transportation Systems (ITSC) , Qingdao, China: IEEE, Oct. 2014, pp. 1900 –1901. doi: 10.1109/ITSC.2014.6957976

  34. [36]

    The highD Dataset: A Drone Dataset of Naturalistic Vehicle Trajectories on German Highways for Validation of Highly Automated Driving Systems,

    R. Krajewski, J. Bock, L. Kloeker, and L. Eckstein, “The highD Dataset: A Drone Dataset of Naturalistic Vehicle Trajectories on German Highways for Validation of Highly Automated Driving Systems,” in 2018 21st International Conference on Intelligent Transportation Systems (ITSC), Nov. 2018, pp. 2118 –2125. doi: 10.1109/ITSC.2018.8569552

  35. [37]

    The inD Dataset: A Drone Dataset of Naturalistic Road User Trajectories at German Intersections,

    J. Bock, R. Krajewski, T. Moers, S. Runde, L. Vater, and L. Eckstein, “The inD Dataset: A Drone Dataset of Naturalistic Road User Trajectories at German Intersections,” in 2020 IEEE Intelligent Vehicles Symposium (IV) , Oct. 2020, pp. 1929 –1934. doi: 10.1109/IV47402.2020.9304839

  36. [38]

    DriveSceneGen: Generating Diverse and Realistic Driving Scenarios from Scratch,

    S. Sun et al., “DriveSceneGen: Generating Diverse and Realistic Driving Scenarios from Scratch,” Feb. 28, 2024, arXiv: arXiv:2309.14685. doi: 10.48550/arXiv.2309.14685

  37. [39]

    Pre -crash scenario typology for crash avoidance research,

    W. G. Najm, J. D. Smith, M. Yanagisawa, and John A. V olpe National Transportation Systems Center (U.S.), “Pre -crash scenario typology for crash avoidance research,” DOT-VNTSC-NHTSA-06-02, Apr. 2007. Accessed: Sept. 05, 2025. [Online]. Available: /view/dot/6281

  38. [40]

    A Driver-Vehicle Model for ADS Scenario-Based Testing,

    R. Queiroz et al., “A Driver-Vehicle Model for ADS Scenario-Based Testing,” IEEE Trans. Intell. Transp. Syst., vol. 25, no. 8, pp. 8641 –8654, Aug. 2024, doi: 10.1109/TITS.2024.3373531

  39. [41]

    Generating Critical Driving Scenarios from Accident Sketches,

    A. Gambi, V . Nguyen, J. Ahmed, and G. Fraser, “Generating Critical Driving Scenarios from Accident Sketches,” in 2022 IEEE International Conference On Artificial Intelligence Testing (AITest), Newark, CA, USA: IEEE, Aug. 2022, pp. 95–102. doi: 10.1109/AITest55621.2022.00022

  40. [42]

    Text2Scenario: Text-Driven Scenario Generation for Autonomous Driving Test,

    X. Cai et al., “Text2Scenario: Text-Driven Scenario Generation for Autonomous Driving Test,” Mar. 04, 2025, arXiv: arXiv:2503.02911. doi: 10.48550/arXiv.2503.02911

  41. [43]

    Nalic, T

    D. Nalic, T. Mihalj, A. Eichberger, TU Dresden, GERMANY , M. Bäumler, and M. Lehmann, “c,” in FISITA World Congress 2021 - Technical Programme, FISITA, Sept. 2021. doi: 10.46720/f2020-acm-096

  42. [44]

    Research on Specific Scenario Generation Methods for Autonomous Driving Simulation Tests,

    N. Li, L. Chen, and Y . Huang, “Research on Specific Scenario Generation Methods for Autonomous Driving Simulation Tests,” World Electr. Veh. J., vol. 15, no. 1, p. 2, Dec. 2023, doi: 10.3390/wevj15010002

  43. [45]

    DAnoScenE: a driving anomaly scenario extraction framework for autonomous vehicles in urban streets,

    Y . Hu, D. Zhao, Y . Wang, and G. Zhao, “DAnoScenE: a driving anomaly scenario extraction framework for autonomous vehicles in urban streets,” J. Intell. Transp. Syst., vol. 29, no. 1, pp. 32 –52, Jan. 2025, doi: 10.1080/15472450.2023.2291680

  44. [46]

    Categorizing Data -Driven Methods for Test Scenario Generation to Assess Automated Driving Systems,

    M. Bäumler, F. Linke, and G. Prokop, “Categorizing Data -Driven Methods for Test Scenario Generation to Assess Automated Driving Systems,” IEEE Access, vol. 12, pp. 52030–52050, 2024, doi: 10.1109/ACCESS.2024.3385646

  45. [47]

    Ontology based Scene Creation for the Development of Automated Vehicles,

    G. Bagschik, T. Menzel, and M. Maurer, “Ontology based Scene Creation for the Development of Automated Vehicles,” in 2018 IEEE Intelligent Vehicles Symposium (IV) , Changshu: IEEE, June 2018, pp. 1813 –1820. doi: 10.1109/IVS.2018.8500632

  46. [48]

    A Scenario -Adaptive Driving Behavior Prediction Approach to Urban Autonomous Driving,

    X. Geng, H. Liang, B. Yu, P. Zhao, L. He, and R. Huang, “A Scenario -Adaptive Driving Behavior Prediction Approach to Urban Autonomous Driving,” Appl. Sci., vol. 7, no. 4, p. 426, Apr. 2017, doi: 10.3390/app7040426

  47. [49]

    TraModeA VTest: Modeling Scenario and Violation Testing for Autonomous Driving Systems Based on Traffic Regulations,

    C. Xia, S. Huang, C. Zheng, Z. Yang, T. Bai, and L. Sun, “TraModeA VTest: Modeling Scenario and Violation Testing for Autonomous Driving Systems Based on Traffic Regulations,” Electronics, vol. 13, no. 7, p. 1197, Mar. 2024, doi: 10.3390/electronics13071197

  48. [50]

    Attention is All you Need,

    A. Vaswani et al., “Attention is All you Need,” in Advances in Neural Information Processing Systems, Curran Associates, Inc., 2017. Accessed: Sept. 05, 2025. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/20 17/hash/3f5ee243547dee91fbd053c1c4a845aa- Abstract.html

  49. [51]

    Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,

    J. Wei et al., “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,” Jan. 10, 2023, arXiv: arXiv:2201.11903. doi: 10.48550/arXiv.2201.11903

  50. [52]

    LeGEND: A Top -Down Approach to Scenario Generation of Autonomous Driving Systems Assisted by Large Language Models,

    S. Tang, Z. Zhang, J. Zhou, L. Lei, Y . Zhou, and Y . Xue, “LeGEND: A Top -Down Approach to Scenario Generation of Autonomous Driving Systems Assisted by Large Language Models,” arXiv.org. Accessed: Aug. 29, 2025. [Online]. Available: https://arxiv.org/abs/2409.10066v1

  51. [53]

    ADEPT: A Testing Platform for Simulated Autonomous Driving,

    S. Wang et al. , “ADEPT: A Testing Platform for Simulated Autonomous Driving,” in Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering , Rochester MI USA: ACM, Oct. 2022, pp. 1 –4. doi: 10.1145/3551349.3559528

  52. [54]

    Improving Language Understanding by Generative Pre-Training

    A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving Language Understanding by Generative Pre-Training”

  53. [55]

    Scenic: a language for scenario specification and scene generation,

    D. J. Fremont, T. Dreossi, S. Ghosh, X. Yue, A. L. Sangiovanni-Vincentelli, and S. A. Seshia, “Scenic: a language for scenario specification and scene generation,” in Proceedings of the 40th ACM SIGPLAN Conference on Programming Language Design and Implementation , Phoenix AZ USA: ACM, June 2019, pp. 63 –78. doi: 10.1145/3314221.3314633

  54. [56]

    Chat2Scenario: Scenario Extraction From Dataset Through Utilization of Large Language Model,

    Y . Zhao, W. Xiao, T. Mihalj, J. Hu, and A. Eichberger, “Chat2Scenario: Scenario Extraction From Dataset Through Utilization of Large Language Model,” in 2024 IEEE Intelligent Vehicles Symposium (IV), June 2024, pp. 559 –566. doi: 10.1109/IV55156.2024.10588843

  55. [57]

    ChatScene: Knowledge-Enabled Safety -Critical Scenario Generation for Autonomous Vehicles,

    J. Zhang, C. Xu, and B. Li, “ChatScene: Knowledge-Enabled Safety -Critical Scenario Generation for Autonomous Vehicles,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , Seattle, WA, USA: IEEE, June 2024, pp. 15459 –15469. doi: 10.1109/CVPR52733.2024.01464

  56. [58]

    CARLA: An Open Urban Driving Simulator,

    A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V . Koltun, “CARLA: An Open Urban Driving Simulator,” in Proceedings of the 1st Annual Conference on Robot Learning , PMLR, Oct. 2017, pp. 1 –16. Accessed: Sept. 04, 2025. [Online]. Available: https://proceedings.mlr.press/v78/dosovitskiy17a.h tml

  57. [59]

    Multi-modal Traffic Scenario Generation for Autonomous Driving System Testing,

    Z. Tu, L. Niu, W. Fan, and T. Zhang, “Multi-modal Traffic Scenario Generation for Autonomous Driving System Testing,” Proc. ACM Softw. Eng. , vol. 2, no. FSE, pp. 1733 –1756, June 2025, doi: 10.1145/3729348

  58. [60]

    YOLOv10: Real-Time End-to-End Object Detection,

    A. Wang et al., “YOLOv10: Real-Time End-to-End Object Detection,” Oct. 30, 2024, arXiv: arXiv:2405.14458. doi: 10.48550/arXiv.2405.14458

  59. [61]

    TARGET: Traffic Rule -Based Test Generation for Autonomous Driving via Validated LLM-Guided Knowledge Extraction,

    Y . Deng, Z. Tu, J. Yao, M. Zhang, T. Zhang, and X. Zheng, “TARGET: Traffic Rule -Based Test Generation for Autonomous Driving via Validated LLM-Guided Knowledge Extraction,” IEEE Trans. Softw. Eng. , vol. 51, no. 7, pp. 1950 –1968, July 2025, doi: 10.1109/TSE.2025.3569086

  60. [62]

    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,

    A. Dosovitskiy et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” June 03, 2021, arXiv: arXiv:2010.11929. doi: 10.48550/arXiv.2010.11929

  61. [63]

    CurricuVLM: Towards Safe Autonomous Driving via Personalized Safety-Critical Curriculum Learning with Vision-Language Models

    Z. Sheng, Z. Huang, Y . Qu, Y . Leng, S. Bhavanam, and S. Chen, “CurricuVLM: Towards Safe Autonomous Driving via Personalized Safety - Critical Curriculum Learning with Vision - Language Models,” Feb. 21, 2025, arXiv: arXiv:2502.15119. doi: 10.48550/arXiv.2502.15119

  62. [64]

    Multimodal Large Language Model Driven Scenario Testing for Autonomous Vehicles,

    Q. Lu, X. Wang, Y . Jiang, G. Zhao, M. Ma, and S. Feng, “Multimodal Large Language Model Driven Scenario Testing for Autonomous Vehicles,” Sept. 10, 2024, arXiv: arXiv:2409.06450. doi: 10.48550/arXiv.2409.06450

  63. [65]

    Eclipse SUMO - Simulation of Urban MObility

    “Eclipse SUMO - Simulation of Urban MObility.” Accessed: Sept. 08, 2025. [Online]. Available: https://eclipse.dev/sumo/

  64. [66]

    DriveGen: Towards Infinite Diverse Traffic Scenarios with Large Models

    S. Zhang, J. Tian, Z. Zhu, S. Huang, J. Yang, and W. Zhang, “DriveGen: Towards Infinite Diverse Traffic Scenarios with Large Models,” Mar. 04, 2025, arXiv: arXiv:2503.05808. doi: 10.48550/arXiv.2503.05808

  65. [67]

    Generative Adversarial Networks,

    I. J. Goodfellow et al. , “Generative Adversarial Networks,” June 10, 2014, arXiv: arXiv:1406.2661. doi: 10.48550/arXiv.1406.2661

  66. [68]

    Application of GANs -based virtual environment generation in automatic driving simulation training,

    H. Wang, “Application of GANs -based virtual environment generation in automatic driving simulation training,” Trans. Comput. Sci. Intell. Syst. Res., vol. 5, pp. 1760 –1765, Aug. 2024, doi: 10.62051/fzcq0r69

  67. [69]

    Rain Rendering for Evaluating and Improving Robustness to Bad Weather,

    M. Tremblay, S. S. Halder, R. De Charette, and J.-F. Lalonde, “Rain Rendering for Evaluating and Improving Robustness to Bad Weather,” Int. J. Comput. Vis., vol. 129, no. 2, pp. 341 –360, Feb. 2021, doi: 10.1007/s11263-020-01366-3

  68. [70]

    Video Generative Adversarial Networks: A Review,

    N. Aldausari, A. Sowmya, N. Marcus, and G. Mohammadi, “Video Generative Adversarial Networks: A Review,” ACM Comput. Surv., vol. 55, no. 2, pp. 1–25, Feb. 2023, doi: 10.1145/3487891

  69. [71]

    Depth -aware unpaired image -to- image translation for autonomous driving test scenario generation using a dual -branch GAN,

    D. Shi et al. , “Depth -aware unpaired image -to- image translation for autonomous driving test scenario generation using a dual -branch GAN,” Front. Neurorobotics , vol. 19, p. 1603964, May 2025, doi: 10.3389/fnbot.2025.1603964

  70. [72]

    Generating intersection pre -crash trajectories for autonomous driving safety testing using Transformer Time -Series Generative Adversarial Networks,

    X. Liu, H. Huang, J. Bian, R. Zhou, Z. Wei, and H. Zhou, “Generating intersection pre -crash trajectories for autonomous driving safety testing using Transformer Time -Series Generative Adversarial Networks,” Eng. Appl. Artif. Intell., vol. 160, p. 111995, Aug. 2025, doi: 10.1016/j.engappai.2025.111995

  71. [73]

    ITGAN: An Interactive Trajectories Generative Adversarial Network Model for Automated Driving Scenario Generation,

    Zeguang Liao et al. , “ITGAN: An Interactive Trajectories Generative Adversarial Network Model for Automated Driving Scenario Generation,” in Proceedings of China SAE Congress 2022: Selected Papers , vol. 1025, China Society of Automotive Engineers, Ed., in Lecture Notes in Electrical Engineering, vol. 1025. , Singapore: Springer Nature Singapore, 2023, p...

  72. [74]

    RCG: Safety-Critical Scenario Generation for Robust Autonomous Driving via Real-World Crash Grounding

    B. Stoler, J. Yang, J. Francis, and J. Oh, “RCG: Safety-Critical Scenario Generation for Robust Autonomous Driving via Real -World Crash Grounding,” 2025, arXiv. doi: 10.48550/ARXIV .2507.10749

  73. [75]

    High-Risk Test Scenario Generation for Autonomous Vehicles at Roundabouts Using Naturalistic Driving Data,

    D. Ren, H. Huang, Y . Li, and J. Jin, “High-Risk Test Scenario Generation for Autonomous Vehicles at Roundabouts Using Naturalistic Driving Data,” Appl. Sci., vol. 15, no. 8, p. 4505, Apr. 2025, doi: 10.3390/app15084505

  74. [76]

    LD -Scene: LLM -Guided Diffusion for Controllable Generation of Adversarial Safety - Critical Driving Scenarios,

    M. Peng, Y . Xie, X. Guo, R. Yao, H. Yang, and J. Ma, “LD -Scene: LLM -Guided Diffusion for Controllable Generation of Adversarial Safety - Critical Driving Scenarios,” 2025, arXiv. doi: 10.48550/ARXIV .2505.11247

  75. [77]

    Pseudo Numerical Methods for Diffusion Models on Manifolds,

    L. Liu, Y . Ren, Z. Lin, and Z. Zhao, “Pseudo Numerical Methods for Diffusion Models on Manifolds,” 2022, arXiv. doi: 10.48550/ARXIV .2202.09778

  76. [78]

    High-Resolution Image Synthesis with Latent Diffusion Models,

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-Resolution Image Synthesis with Latent Diffusion Models,” 2021, arXiv. doi: 10.48550/ARXIV .2112.10752

  77. [79]

    Lightweight diffusion models: a survey,

    W. Song, W. Ma, M. Zhang, Y . Zhang, and X. Zhao, “Lightweight diffusion models: a survey,” Artif. Intell. Rev., vol. 57, no. 6, p. 161, May 2024, doi: 10.1007/s10462-024-10800-8

  78. [80]

    A Scenario-Based Development Framework for Autonomous Driving

    X. Li, “A Scenario-Based Development Framework for Autonomous Driving,” 2020, arXiv. doi: 10.48550/ARXIV .2011.01439

  79. [81]

    BITS: Bi-level Imitation for Traffic Simulation,

    D. Xu, Y . Chen, B. Ivanovic, and M. Pavone, “BITS: Bi-level Imitation for Traffic Simulation,” in 2023 IEEE International Conference on Robotics and Automation (ICRA) , London, United Kingdom: IEEE, May 2023, pp. 2929 –2936. doi: 10.1109/ICRA48891.2023.10161167

  80. [82]

    SimNet: Learning Reactive Self-driving Simulations from Real -world Observations,

    L. Bergamini et al. , “SimNet: Learning Reactive Self-driving Simulations from Real -world Observations,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) , Xi’an, China: IEEE, May 2021, pp. 5119–5125. doi: 10.1109/ICRA48506.2021.9561666

Showing first 80 references.