REVIEW 4 major objections 6 minor 87 references
Multi-modal Traffic Scenario Generation for Autonomous Driving System Testing
T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that pairing a natural-language scenario description with one traffic-scene image, aligned through a DSL-based intermediate representation, automatically produces executable simulator scenarios that reach 97% fidelity and…
desk verdict Solid engineering with a credible bug-detection evaluation, but the headline 97% accuracy is measured against author-annotated ground truth and needs independent validation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the aligned intermediate representation built on a compact domain-specific language for traffic scenarios. The DSL describes an environment (weather, time), a road network (type, signals, lane count), and actors (type, behavior, position, lane index), where the added lane-index element lets the same actor be identified from both text and image. Information from the two input modalities is merged into one IR with explicit priority: text wins for dynamic and environmental attributes, vision wins for spatial details, and actors mentioned in only one modality are kept if they do not overlap. This IR then drives a converter that emits simulator-specific scripts, so scenario design is separated from simulator API details.
What would settle it
Have independent annotators who did not design the DSL or the prompts re-annotate the same 120 text-and-image pairs, then recompute the tree-edit-distance accuracy of TrafficComposer against that external ground truth. If the figure falls to the 89.7% level of the best multimodal LLM baseline, the multimodal alignment advantage would shrink or disappear; if it stays near 97%, the result would be robust to annotation bias.
Extended reading notes
Core claim
The paper's central discovery is that aligning two modality-specific intermediate representations beats end-to-end multimodal generation for traffic-scenario fidelity. TrafficComposer parses the text with an LLM into a DSL-based IR, parses the image with object and lane detection into a visual IR, then merges them by a fixed priority rule: text decides environment, signals, actor types, speed, and behavior; vision decides lane counts and actor positions. The merged IR is converted into an executable CARLA or LGSVL script. The measured effect is a 97.0% average tree-edit-distance-based accuracy, compared with 74.3% for the text-only TARGET baseline and 89.7% for GPT-4o, with the main gap coming from actors that multimodal LLMs miss entirely. The same scenarios are then shown to be more effective fuzzing seeds than the fuzzers' original or TARGET-generated seeds.
Load-bearing premise
The load-bearing premise is that the authors' manually annotated ground-truth scenarios for all 120 benchmark cases are correct and unbiased; every accuracy number and the accuracy-to-bug-detection comparison is measured against those annotations.
Editorial extensions
If this is right
- Scenario authoring for ADS testing can shift from hand-written scripts or strict DSLs to a natural-language-plus-image prompt, with no loss of complexity: generated scenarios contain 5 to 21 actors in diverse cities, weather, and road types.
- Higher scenario fidelity is not just cosmetic: seeds at 97.0% accuracy expose 33–109% more failures than seeds at 74.3% accuracy under matched fuzzing budgets, so improving generation accuracy directly improves bug finding.
- Fuzzing time efficiency improves as well: first-failure time drops by 28–58% across six ADSs compared with original seeds and by 30–52% compared with TARGET-generated seeds.
- The pipeline is robust to component choice: average accuracy stays between 90.7% and 97.0% across eight LLMs and between 94.1% and 97.0% across ten YOLO-family object detectors, and it degrades gracefully under injected hallucination or detection errors.
Reading between the lines
- If the accuracy-to-bug-detection correlation holds generally, then investing in better scene parsing (fewer missed actors, more precise lane placement) may pay off more for ADS testing than investing in richer mutation operators in the fuzzer.
- The modality-alignment recipe—text decides dynamics, vision decides geometry—is transferable to other executable-simulation domains, such as robot manipulation or game testing, where a natural-language goal plus a reference image can specify a test scenario.
- Because the ground-truth IRs were annotated by the same team that designed the DSL and prompts, an independent re-annotation study would be the natural next test of whether the 97% figure represents generalizable fidelity or benchmark self-consistency.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. TrafficComposer is a pipeline that takes a natural-language traffic scenario description and a reference image, extracts a DSL-based intermediate representation from each modality (LLM for text, YOLOv10+CLRNet for image), aligns and merges the two IRs, and converts the result into executable CARLA or LGSVL scenarios. The paper contributes a benchmark of 120 traffic scenarios with text, image, and ground-truth IRs; reports 97.0% IE accuracy against this ground truth, outperforming GPT-4o by 7.3% and TARGET by 22.7%; and evaluates bug-detection effectiveness on six ADSs, reporting 37 bugs in direct testing and 33%–124% more bugs when scenarios are used as fuzzing seeds. It also includes ablation, LLM/CV-model sensitivity, and hallucination/error-injection analyses.
Significance. The bug-detection portion of the evaluation is a genuine strength: six ADSs, ten repeated runs, statistical testing, and manual attribution of failures to ADS errors, with a public artifact. The sensitivity analyses across eight LLMs and ten object detectors are useful for generalizability. The 120-scenario benchmark is more complex (avg 7.8 actors) and more diverse (Vendi Score 22.1) than TARGET and LawBreaker. However, the headline accuracy claim rests on author-generated ground truth, which is the main risk; if independently validated, the claimed margin over GPT-4o would be a substantial contribution. The fuzzing improvements are externally credible, but their attribution to scenario accuracy is not yet established.
major comments (4)
- [Section 4.1 and Eq. (1)] The ground-truth IRs to which IE accuracy is measured were annotated by two of the authors and reconciled by a third, using the same extended DSL (Section 3.1) and the same text-description guideline [71] that TrafficComposer implements, and Section 7 does not discuss annotation bias among the threats to validity. Because Equation (1) computes distance to these in-house annotations, the 97.0% accuracy and the 7.3% margin over GPT-4o may reflect self-consistency with the authors' own conventions rather than independent scenario fidelity. Please validate the ground truth with annotators who were not involved in designing TrafficComposer or the DSL (reporting inter-annotator agreement on a sample or all 120 scenarios), or benchmark a subset against an external dataset with objective scenario semantics, and re-report RQ1 accordingly.
- [Sections 4.2, 4.3, and Table 6] The fuzzing comparison between TrafficComposer seeds and TARGET seeds does not isolate scenario fidelity because TARGET scenarios contain one other actor by design (Section 4.2) while TrafficComposer scenarios average 7.8 actors (Table 2). The reported 33%–124% bug increases could therefore be driven by actor count or scene complexity rather than by IE accuracy. To support the causal attribution, the authors should control for scenario content, e.g., by matching the number and types of actors across the seed settings or by ablating actors from TrafficComposer scenarios to the TARGET level.
- [Section 4.4 and Table 9] RQ3 asserts that higher-accuracy scenarios expose more failures, but it compares two approaches that differ in modality, actor count, and seed content, so accuracy is not the only variable, contrary to the assertion in Section 4.4 that this setup isolates scenario accuracy as the only variable. The failure-count differences in Table 9 (e.g., 39.1 vs. 24.2 for Apollo) should be re-examined with controlled perturbations that change IE accuracy while keeping scenario content fixed, for example by using the error-injection mechanisms from Sections 4.6 and 4.7 and measuring downstream bug counts.
- [Section 5.2 and Table 5] The claim that '37 of these 120 scenarios directly triggered crashes or traffic rule violations' is not reconciled with Table 5, whose per-ADS totals sum to 100; if a single scenario can trigger failures in multiple ADSs, the number of unique bug-triggering scenarios and the overlap across ADSs should be reported. Without this clarification, the abstract's '37 bugs' cannot be checked against the underlying data.
minor comments (6)
- [Section 3.4] 'adpots' should be 'adopts'.
- [Section 4.1] 'multi-model benchmark' should be 'multi-modal benchmark'.
- [Table 4] The 'margin of error' columns are not defined; please specify whether these are standard errors, 95% confidence intervals, or something else, and state the number of repeats used for each entry.
- [Section 3.5] The default actor speed is described as randomly sampled between 0 and 30 mph without specifying the distribution or a random seed; please report the exact sampling procedure so that the simulation runs are reproducible.
- [Section 4.3] The LawBreaker evaluation is limited to a single ADS (Apollo) because of LGSVL discontinuation; this limitation should be stated more explicitly in Section 7, since it affects the generality of the LawBreaker fuzzing results.
- [Section 5.1] The comparison that GPT-4o misses actors in 41 scenarios while TrafficComposer misses them in 8 is reported without a significance test or confidence interval; a paired test over the 120 scenarios would strengthen this supporting claim.
Circularity Check
No construction-level circularity: the 97.0% accuracy is measured against human-annotated ground truth, and the central bug-detection results are external; annotation bias is a validity concern, not a circular reduction.
full rationale
TrafficComposer's claimed chain is: LLM textual parser + CV visual extractor -> aligned IR -> simulator script, evaluated by Eq. 1 against ground-truth IRs. No parameter is fitted to the ground truth, and no output is used to define the oracle: the GT is independently annotated by two authors with a third reconciling (Sec. 4.1), with inter-annotator IE accuracy 97.7%, and the system's 97.0% is a separate measurement. The DSL is adopted from the authors' prior TARGET work (Sec. 3.1, ref. [22]) and is shared across baselines; although this is a self-citation, it is a transparent design choice, not an unverified premise that forces the result. The main validity weakness is that the GT is annotated in the authors' own DSL and the textual descriptions follow the authors' own guideline [71], so the accuracy benchmark measures agreement with the authors' conventions; Sec. 7's internal-validity discussion omits this annotation-bias threat. However, this is a benchmark-validity concern, not circularity: Eq. 1 does not reduce S to G by construction, and the RQ2/RQ3 bug-detection evidence rests on externally observable ADS failures, not on the GT. Therefore no circular step is exhibited.
Assumptions & free parameters
free parameters (1)
- Unspecified actor speed default range =
0-30 mph (randomly sampled)
assumptions (4)
- domain assumption The DSL grammar from Deng et al. (with lane_idx extension) is adequate to represent all traffic scenarios needed for ADS testing.
- ad hoc to paper The authors' manually annotated ground-truth IRs are the correct and unique representation of each benchmark scenario.
- ad hoc to paper Text should be prioritized for environment and traffic signals, while visual input should be prioritized for lane count and actor positions.
- domain assumption YOLOv10 and CLRNet detection outputs and the LLM parsing outputs are reliable enough for the pipeline.
Cite this review
Pith. "Pith review of Multi-modal Traffic Scenario Generation for Autonomous Driving System Testing." pith.science (2026). https://pith.science/paper/3ZWYTMBR
@misc{pith2026250514881,
author = {Pith},
title = {Pith review of: Multi-modal Traffic Scenario Generation for Autonomous Driving System Testing},
year = {2026},
howpublished = {\url{https://pith.science/paper/3ZWYTMBR}},
note = {Machine review of arXiv:2505.14881}
}
read the original abstract
Autonomous driving systems (ADS) require extensive testing and validation before deployment. However, it is tedious and time-consuming to construct traffic scenarios for ADS testing. In this paper, we propose TrafficComposer, a multi-modal traffic scenario construction approach for ADS testing. TrafficComposer takes as input a natural language (NL) description of a desired traffic scenario and a complementary traffic scene image. Then, it generates the corresponding traffic scenario in a simulator, such as CARLA and LGSVL. Specifically, TrafficComposer integrates high-level dynamic information about the traffic scenario from the NL description and intricate details about the surrounding vehicles, pedestrians, and the road network from the image. The information from the two modalities is complementary to each other and helps generate high-quality traffic scenarios for ADS testing. On a benchmark of 120 traffic scenarios, TrafficComposer achieves 97.0% accuracy, outperforming the best-performing baseline by 7.3%. Both direct testing and fuzz testing experiments on six ADSs prove the bug detection capabilities of the traffic scenarios generated by TrafficComposer. These scenarios can directly discover 37 bugs and help two fuzzing methods find 33%--124% more bugs serving as initial seeds.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[71]
Zhi Tu, Liangkun Niu, Wei Fan, and Tianyi Zhang. 2025. Guideline for Creating Textual Descriptions. https: //github.com/TrafficComposer/TrafficComposer/blob/main/resources/benchmark_guidelines.pdf
work page 2025
-
[1]
Raja Ben Abdessalem, Annibale Panichella, Shiva Nejati, Lionel C. Briand, and Thomas Stifter. 2018. Testing autonomous cars for feature interaction failures using many-objective search. In Proceedings of the 33rd ACM/IEEE International Proc. ACM Softw. Eng., Vol. 2, No. FSE, Article FSE078. Publication date: July 2025. Multi-modal Traffic Scenario Generat...
arXiv 2018
-
[2]
Amazon. 2024. Amazon Titan Text Premier. https://aws.amazon.com/blogs/aws/build-rag-and-agent-based- generative-ai-applications-with-new-amazon-titan-text-premier-model-available-in-amazon-bedrock/
2024
-
[3]
Anthopic PBC. 2024. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. https://www. anthropic.com/news/3-5-models-and-computer-use
2024
-
[4]
Anthropic PBC. 2024. Claude 3.5 Sonnet. https://www.anthropic.com/news/claude-3-5-sonnet
2024
-
[5]
ASAM. 2021. ASAM OpenSCENARIO: User Guide. https://www.openscenario.org. Oneline; accessed July 2024
2021
-
[6]
Associated Press News. 2020. 3 Crashes, 3 Deaths Raise Questions About Tesla’s Autopilot. https://apnews.com/ ca5e62255bb87bf1b151f9bf075aaadf
2020
-
[7]
Sai Krishna Bashetty, Heni Ben Amor, and Georgios Fainekos. 2020. DeepCrashTest: Turning Dashcam Videos into Virtual Crash Tests for Automated Driving Systems. In 2020 IEEE International Conference on Robotics and Automation, ICRA 2020, Paris, France, May 31 - August 31, 2020 . IEEE, 11353–11360. doi:10.1109/ICRA40945.2020.9197053
arXiv 2020
Show all 87 references
-
[8]
BBC News. 2016. Google Self-driving Car Hits a Bus. https://www.bbc.com/news/technology-35692845
2016
-
[9]
BBC News. 2016. Uber in Fatal Crash Had Safety Flaws Say US Investigators. https://www.bbc.com/news/business- 50312340
2016
-
[10]
BBC News. 2016. US opens investigation into Tesla after fatal crash. Retrieved January, 2025 from https://www.bbc. com/news/technology-36680043
2016
-
[11]
BBC News. 2019. Tesla Model 3: Autopilot Engaged during Fatal Crash. https://www.bbc.com/news/technology- 48308852
2019
-
[12]
Karsten Behrendt and Ryan Soussan. 2019. Unsupervised Labeled Lane Markers Using Maps. In 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW) . 832–839. doi:10.1109/ICCVW.2019.00111
2019
-
[13]
Tom Brown et al. 2020. Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 1877–1901
2020
-
[14]
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. 2020. nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vision and patt...
2020
-
[15]
Carla. 2024. CARLA Autonomous Driving Challenge. https://leaderboard.carla.org/challenge/
2024
-
[16]
CARLA team. 2022. CARLA Agents. https://carla.readthedocs.io/en/0.9.12/adv_agents/
2022
-
[17]
Shi-tao Chen, Yu Chen, Songyi Zhang, and Nanning Zheng. 2019. A Novel Integrated Simulation and Testing Platform for Self-Driving Cars With Hardware in the Loop. IEEE Trans. Intell. Veh. 4, 3 (2019), 425–436. doi:10.1109/TIV.2019. 2919470
2019 doi
-
[18]
Yuntianyi Chen, Yuqi Huai, Shilong Li, Changnam Hong, and Joshua Garcia. 2024. Misconfiguration Software Testing for Failure Emergence in Autonomous Driving Systems. Proc. ACM Softw. Eng. 1, FSE, Article 85 (jul 2024), 24 pages. doi:10.1145/3660792
2024 doi
-
[20]
Kashyap Chitta, Aditya Prakash, Bernhard Jaeger, Zehao Yu, Katrin Renz, and Andreas Geiger. 2023. TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving. Pattern Analysis and Machine Intelligence (PAMI) (2023)
2023
-
[21]
Jiarun Dai, Bufan Gao, Mingyuan Luo, Zongan Huang, Zhongrui Li, Yuan Zhang, and Min Yang. 2024. SCTrans: Constructing a Large Public Scenario Dataset for Simulation Testing of Autonomous Driving Systems. In Proceedings of the 46th IEEE/ACM International Conference on Software ...
2024
-
[22]
Yao Deng, Jiaohong Yao, Zhi Tu, Xi Zheng, Mengshi Zhang, and Tianyi Zhang. 2023. TARGET: Traffic Rule-based Test Generation for Autonomous Driving Systems. arXiv:2305.06018 [cs.SE]
2023 arXiv
-
[23]
Yao Deng, Xi Zheng, Tianyi Zhang, Guannan Lou, Huai Liu, and Miryung Kim. 2020. RMT: Rule-based Metamorphic Testing for Autonomous Driving Models. CoRR abs/2012.10672 (2020). arXiv:2012.10672 https://arxiv.org/abs/2012. 10672
2020 arXiv
-
[24]
López, and Vladlen Koltun
Alexey Dosovitskiy, Germán Ros, Felipe Codevilla, Antonio M. López, and Vladlen Koltun. 2017. CARLA: An Open Urban Driving Simulator. In 1st Annual Conference on Robot Learning, CoRL 2017, Mountain View, California, USA, November 13-15, 2017, Proceedings (Proceedings of Machin...
2017
-
[25]
Abhimanyu Dubey and Abhinav Jauhri et al. 2024. The Llama 3 Herd of Models. arXiv:2407.21783 [cs.AI] https: //arxiv.org/abs/2407.21783 Proc. ACM Softw. Eng., Vol. 2, No. FSE, Article FSE078. Publication date: July 2025. FSE078:22 Zhi Tu, Liangkun Niu, Wei Fan, and Tianyi Zhang
2024 arXiv
-
[26]
Shuo Feng, Haowei Sun, Xintao Yan, Haojie Zhu, Zhengxia Zou, Shengyin Shen, and Henry Liu. 2023. Dense reinforcement learning for safety validation of autonomous vehicles.Nature 615 (03 2023), 620–627. doi:10.1038/s41586- 023-05732-2
2023 doi
-
[27]
Daniel J Fremont, Tommaso Dreossi, Shromona Ghosh, Xiangyu Yue, Alberto L Sangiovanni-Vincentelli, and Sanjit A Seshia. 2019. Scenic: a language for scenario specification and scene generation. In Proceedings of the 40th ACM SIGPLAN Conference on Programming Language Design an...
2019
-
[28]
Dan Friedman and Adji Bousso Dieng. 2023. The Vendi Score: A Diversity Evaluation Metric for Machine Learning. Trans. Mach. Learn. Res. 2023 (2023). https://openreview.net/forum?id=g97OHbQyk1
2023
-
[29]
Alessio Gambi, Tri Huynh, and Gordon Fraser. 2019. Generating effective test cases for self-driving cars from police reports. In Proceedings of the ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, ESEC/SIGS...
2019
-
[30]
Alessio Gambi, Marc Müller, and Gordon Fraser. 2019. AsFault: testing self-driving car software using search-based procedural content generation. In Proceedings of the 41st International Conference on Software Engineering: Companion Proceedings, ICSE 2019, Montreal, QC, Canada...
2019
-
[31]
Alessio Gambi, Marc Müller, and Gordon Fraser. 2019. Automatically testing self-driving cars with search-based procedural content generation. In Proceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis, ISSTA 2019, Beijing, China, July 15-19...
2019
-
[32]
Ellen R Girden. 1992. ANOV A: Repeated measures. Number 84. Sage
1992
-
[33]
An Guo, Yuan Zhou, Haoxiang Tian, Chunrong Fang, Yunjian Sun, Weisong Sun, Xinyu Gao, Anh Tuan Luu, Yang Liu, and Zhenyu Chen. 2024. SoVAR: Build Generalizable Scenarios from Accident Reports for Autonomous Driving Testing. In Proceedings of the 39th IEEE/ACM International Con...
2024
-
[34]
Yunfei Guo, Fei Yin, Xiao-hui Li, Xudong Yan, Tao Xue, Shuqi Mei, and Cheng-Lin Liu. 2023. Visual traffic knowledge graph generation from scene images. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 21604–21613
2023
-
[35]
Yuqi Huai, Sumaya Almanee, Yuntianyi Chen, Xiafa Wu, Qi Alfred Chen, and Joshua Garcia. 2023. scenoRITA: Generating Diverse, Fully Mutable, Test Scenarios for Autonomous Vehicle Planning. IEEE Transactions on Software Engineering 49, 10 (2023), 4656–4676. doi:10.1109/TSE.2023.3309610
2023
-
[36]
Shaofei Huang, Zhenwei Shen, Zehao Huang, Zi-han Ding, Jiao Dai, Jizhong Han, Naiyan Wang, and Si Liu. 2023. Anchor3dlane: Learning to regress 3d anchors for monocular 3d lane detection. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 17451–17460
2023
-
[37]
J Utah. 2018. Driving Downtown - New York City 4K - USA. https://www.youtube.com/watch?v=7HaJArMDKgI
2018
-
[38]
Lingxiao Jiang, Ghassan Misherghi, Zhendong Su, and Stephane Glondu. 2007. DECKARD: Scalable and Accurate Tree-Based Detection of Code Clones. In 29th International Conference on Software Engineering (ICSE’07) . 96–105. doi:10.1109/ICSE.2007.30
2007 doi
-
[39]
Shinpei Kato, Shota Tokunaga, Yuya Maruyama, Seiya Maeda, Manato Hirabayashi, Yuki Kitsukawa, Abraham Monrroy, Tomohito Ando, Yusuke Fujii, and Takuya Azumi. 2018. Autoware on Board: Enabling Autonomous Vehicles with Embedded Systems. In 2018 ACM/IEEE 9th International Confere...
2018
-
[40]
Seulbae Kim, Major Liu, Junghwan" John" Rhee, Yuseok Jeon, Yonghwi Kwon, and Chung Hwan Kim. 2022. Drivefuzz: Discovering autonomous driving bugs through driving quality-guided fuzzing. In Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security . ...
2022
-
[41]
Alexandre Kirchmeyer and Jia Deng. 2023. Convolutional networks with oriented 1d kernels. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 6222–6232
2023
-
[42]
Sandra Kübler, Ryan McDonald, and Joakim Nivre. 2009. Dependency parsing. In Dependency parsing. Springer, 11–20
2009
-
[43]
Guanpeng Li, Yiran Li, Saurabh Jha, Timothy Tsai, Michael Sullivan, Siva Kumar Sastry Hari, Zbigniew Kalbarczyk, and Ravishankar Iyer. 2020. AV-FUZZER: Finding Safety Violations in Autonomous Driving Systems. In 2020 IEEE 31st International Symposium on Software Reliability En...
2020
-
[44]
Lawrence Zitnick, and Piotr Dollár
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár. 2015. Microsoft COCO: Common Objects in Context. arXiv:1405.0312 [cs.CV]
2015 arXiv
-
[45]
Haotian Liu. 2023. LLaVA-1.6. https://huggingface.co/liuhaotian/llava-v1.6-34b
2023
-
[46]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruction tuning. Advances in neural information processing systems 36 (2024). Proc. ACM Softw. Eng., Vol. 2, No. FSE, Article FSE078. Publication date: July 2025. Multi-modal Traffic Scenario Generation f...
2024
-
[47]
Yixing Luo, Xiao-Yi Zhang, Paolo Arcaini, Zhi Jin, Haiyan Zhao, Fuyuki Ishikawa, Rongxin Wu, and Tao Xie. 2021. Targeting Requirements Violations of Autonomous Driving Systems by Dynamic Evolutionary Search. In36th IEEE/ACM International Conference on Automated Software Engine...
2021
-
[48]
Christopher D Manning, Mihai Surdeanu, John Bauer, Jenny Rose Finkel, Steven Bethard, and David McClosky. 2014. The Stanford CoreNLP natural language processing toolkit. In Proceedings of 52nd annual meeting of the association for computational linguistics: system demonstratio...
2014
-
[49]
Meta. 2024. Llama 3.2: Revolutionizing edge AI and vision with open, customizable models. https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/
2024
-
[50]
Mistral AI. 2024. Mistral Large 2. https://mistral.ai/news/mistral-large-2407/. [Online; accessed September, 2024]
2024
-
[51]
Texas Department of Public Safety. 2022. Texas DMV Handbook. https://driving-tests.org/texas/tx-dmv-drivers- handbook-manual/ [Online; accessed August 2024]
2022
-
[52]
OpenAI. 2024. GPT-4o. https://openai.com/index/hello-gpt-4o/. Oneline; accessed August 2024
2024
-
[53]
OpenAI. 2024. GPT-4o mini: advancing cost-efficient intelligence. https://openai.com/index/gpt-4o-mini-advancing- cost-efficient-intelligence/
2024
-
[54]
OpenAI, Josh Achiam, Steven Adler, and Sandhini Agarwal et al. 2024. GPT-4 Technical Report. arXiv:2303.08774 [cs.CL] https://arxiv.org/abs/2303.08774
2024 arXiv
-
[55]
Xingang Pan, Jianping Shi, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 2018. Spatial as deep: Spatial cnn for traffic scene understanding. In Proceedings of the AAAI conference on artificial intelligence , Vol. 32
2018
-
[56]
Mateusz Pawlik and Nikolaus Augsten. 2015. Efficient computation of the tree edit distance. ACM Transactions on Database Systems (TODS) 40, 1 (2015), 1–40
2015
-
[57]
Mateusz Pawlik and Nikolaus Augsten. 2016. Tree edit distance: Robust and memory-efficient. Information Systems 56 (2016), 157–173
2016
-
[58]
Planet Drive. 2023. Frankfurt Evening Drive | Driving in Europe’s Financial Capital | Roads of Germany [4K HDR]. https://www.youtube.com/watch?v=2054gi9TzcA
2023
-
[59]
Rodrigo Queiroz, Thorsten Berger, and Krzysztof Czarnecki. 2019. GeoScenario: An Open DSL for Autonomous Driving Scenario Representation. In 2019 IEEE Intelligent Vehicles Symposium (IV)
2019
-
[60]
Guodong Rong, Byung Hyun Shin, Hadi Tabatabaee, Qiang Lu, Steve Lemke, M¯artin, š Možeiko, Eric Boise, Geehoon Uhm, Mark Gerow, Shalin Mehta, et al. 2020. LGSVL simulator: A high fidelity simulator for autonomous driving. In 2020 IEEE 23rd International conference on intellige...
2020
-
[61]
SAE On-Road Automated Vehicle Standards Committee. 2014. Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles . Technical Report
2014
-
[62]
Seoul Walker. 2022. Night Driving Seoul City | Yeouido and Yongsan with Chill Lofi Hiphop POV | 4K HDR. https: //www.youtube.com/watch?v=kuRPXB3q5eo. [Online; accessed September, 2024]
2022
-
[63]
Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023. Role-Play with Large Language Models. arXiv:2305.16367 [cs.CL]
2023 arXiv
-
[64]
Poskitt, Jun Sun, Yuqi Chen, and Zijiang Yang
Yang Sun, Christopher M. Poskitt, Jun Sun, Yuqi Chen, and Zijiang Yang. 2022. LawBreaker: An Approach for Specifying Traffic Laws and Fuzzing Autonomous Vehicles. In37th IEEE/ACM International Conference on Automated Software Engineering, ASE 2022, Rochester, MI, USA, October ...
2022
-
[65]
Shuncheng Tang, Zhenya Zhang, Jixiang Zhou, Lei Lei, Yuan Zhou, and Yinxing Xue. 2024. LeGEND: A Top-Down Approach to Scenario Generation of Autonomous Driving Systems Assisted by Large Language Models. In Proceedings of the 39th IEEE/ACM International Conference on Automated ...
2024
-
[66]
Baidu Apollo team. 2020. Apollo 6.0. https://github.com/ApolloAuto/apollo/releases/tag/v6.0.0
2020
-
[67]
The Mercury News. 2018. Tesla: Autopilot Was On During Deadly Mountain View Crash. https://www.mercurynews. com/2018/03/30/tesla-autopilot-was-onduring-deadly-mountain-view-crash/
2018
-
[68]
Haoxiang Tian, Xingshuo Han, Guoquan Wu, Yuan Zhou, Shuo Li, Jun Wei, Dan Ye, Wei Wang, and Tianwei Zhang
-
[69]
Yuchi Tian, Kexin Pei, Suman Jana, and Baishakhi Ray. 2018. Deeptest: Automated testing of deep-neural-network- driven autonomous cars. In Proceedings of the 40th international conference on software engineering . 303–314
2018
-
[70]
Zhi Tu, Liangkun Niu, Wei Fan, and Tianyi Zhang. 2025. Artifacts for Multi-modal Traffic Scenario Generation for Autonomous Driving System Testing. https://github.com/TrafficComposer/TrafficComposer/
2025
-
[72]
Zhi Tu, Liangkun Niu, Wei Fan, and Tianyi Zhang. 2025. Prompt Design for Multi-modal LLM Base- lines. https://github.com/TrafficComposer/TrafficComposer/blob/main/trafficcomposer/baseline/multi_modal_gpt/ Proc. ACM Softw. Eng., Vol. 2, No. FSE, Article FSE078. Publication date...
2025
-
[73]
Zhi Tu, Liangkun Niu, Wei Fan, and Tianyi Zhang. 2025. Prompt Design for TrafficComposer Textual Description Parser. https://github.com/TrafficComposer/TrafficComposer/blob/main/trafficcomposer/gen_textual_ir/text_parser_ gen_prompt.py
2025
-
[74]
Zhi Tu, Liangkun Niu, Wei Fan, and Tianyi Zhang. 2025. The Scenic code for the motivating example. https: //github.com/TrafficComposer/TrafficComposer/blob/main/motivating_example/motivating_example_scenic
2025
-
[75]
John W Tukey. 1949. Comparing individual means in the analysis of variance. Biometrics (1949), 99–114
1949
-
[76]
Walk East. 2023. Shanghai: The Most Developed City in China - A Driving Tour You Don’t Wanna Miss. https: //www.youtube.com/watch?v=MAiltiE8tgI. [Online; accessed September, 2024]
2023
-
[77]
Ao Wang, Hui Chen, Lihao Liu, Kai Chen, Zijia Lin, Jungong Han, and Guiguang Ding. 2024. YOLOv10: Real-Time End-to-End Object Detection. arXiv:2405.14458 [cs.CV] https://arxiv.org/abs/2405.14458
2024 arXiv
-
[78]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al . 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35 (2022), 24824–24837
2022
-
[79]
Rui Yang, Panos Kalnis, and Anthony K. H. Tung. 2005. Similarity evaluation on tree-structured data. In Proceedings of the 2005 ACM SIGMOD International Conference on Management of Data (Baltimore, Maryland) (SIGMOD ’05). Association for Computing Machinery, New York, NY, USA,...
2005
-
[80]
Mengshi Zhang, Yuqun Zhang, Lingming Zhang, Cong Liu, and Sarfraz Khurshid. 2018. DeepRoad: GAN-based metamorphic testing and input validation framework for autonomous driving systems. In Proceedings of the 33rd ACM/IEEE international conference on automated software engineeri...
2018
-
[81]
Qingwen Zhang, Mingkai Tang, Ruoyu Geng, Feiyi Chen, Ren Xin, and Lujia Wang. 2022. MMFN: Multi-Modal- Fusion-Net for End-to-End Driving. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . 8638–8643. doi:10.1109/IROS47612.2022.9981775
2022
-
[82]
Xudong Zhang and Yan Cai. 2023. Building Critical Testing Scenarios for Autonomous Driving from Real Accidents. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis (Seattle, WA, USA) (ISSTA 2023). Association for Computing Machinery,...
2023
-
[83]
Xinhai Zhang, Jianbo Tao, Kaige Tan, Martin Törngren, José Manuel Gaspar Sánchez, Muhammad Rusyadi Ramli, Xin Tao, Magnus Gyllenhammar, Franz Wotawa, Naveen Mohan, Mihai Nica, and Hermann Felbinger. 2023. Finding Critical Scenarios for Automated Driving Systems: A Systematic M...
2023
-
[84]
Zhejun Zhang, Alexander Liniger, Dengxin Dai, Fisher Yu, and Luc Van Gool. 2021. End-to-End Urban Driving by Imitating a Reinforcement Learning Coach. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021 . IEEE, 152...
2021
-
[85]
Tu Zheng, Yifei Huang, Yang Liu, Wenjian Tang, Zheng Yang, Deng Cai, and Xiaofei He. 2022. CLRNet: Cross Layer Refinement Network for Lane Detection. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022 . IEEE, 88...
2022
-
[86]
Kaiser, and Baishakhi Ray
Ziyuan Zhong, Gail E. Kaiser, and Baishakhi Ray. 2023. Neural Network Guided Evolutionary Fuzzing for Finding Traffic Violations of Autonomous Vehicles. IEEE Trans. Software Eng. 49, 4 (2023), 1860–1875. doi:10.1109/TSE.2022.3195640
2023
-
[87]
Yuan Zhou, Yang Sun, Yun Tang, Yuqi Chen, Jun Sun, Christopher M Poskitt, Yang Liu, and Zijiang Yang. 2023. Specification-based Autonomous Driving System Testing. IEEE Transactions on Software Engineering (2023). Received 2025-02-26; accepted 2025-04-01 Proc. ACM Softw. Eng., ...
2023
-
[2024]
arXiv preprint arXiv:2406.10857 (2024)
An LLM-enhanced Multi-objective Evolutionary Search for Autonomous Driving Test Scenario Generation. arXiv preprint arXiv:2406.10857 (2024)
2024 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.