REVIEW 4 major objections 3 minor 8 cited by
Large Language Models for Unit Testing: A Systematic Literature Review
T0 review · 4 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The first systematic literature review of LLM-based unit testing, analyzing 105 studies through March 2025, charts the field's tasks, model strategies, and open challenges.
desk verdict A genuinely useful first SLR of LLM-based unit testing, but the under-reported quality-filter protocol and a handful of copy-paste errors need fixing before the corpus can be fully trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is a two-perspective taxonomy. From the unit testing side, tasks are classified into test case generation, oracle (assertion) generation, test evolution, completion, smell detection, traceability, minimization, and readability. From the LLM side, studies are classified by the models used, by utilization strategy—model training (pre-training, full fine-tuning, parameter-efficient fine-tuning, reinforcement learning) versus prompt engineering (zero-shot, few-shot, chain-of-thought, tree-of-thought)—and by integration with traditional techniques such as program analysis, information retrieval, program repair, mutation testing, differential testing, and search-based tools. The taxonomy carries the argument because all descriptive findings are expressed as distributions over these categories.
What would settle it
Repeat the described search across the same venues and databases with the same inclusion criteria but without the ten-question quality threshold, and count how many additional qualifying papers on LLM-based unit testing appear; if the count is large, the claim of comprehensiveness fails. Alternatively, check whether the roughly 40% of preprints in the corpus would pass the criterion of publication in a reputable venue; if they would not, the quality threshold and the corpus composition contradict each other.
Extended reading notes
Core claim
The central claim is that the application of LLMs to unit testing has grown into an identifiable research area of its own, distinct from general LLM-for-code work, and that it can be understood through a two-axis taxonomy: the unit-testing task being automated and the way the LLM is deployed. The paper finds that roughly 60% of collected studies target test generation, about 14% target oracle generation, and the remaining studies spread across tasks such as bug reproduction, test evolution, test smell detection, readability, completion, and minimization. On the LLM axis, it finds heavy concentration on a few commercial models, with GPT-3.5 and GPT-4 leading, while open-source models such as CodeLlama are used when fine-tuning is needed, and it finds that prompt engineering, especially zero-shot prompting, is the dominant adaptation strategy, with full fine-tuning the leading training approach. Based on these observations, the paper argues that the main open challenges are context handling for complex units, evaluating whether generated tests actually catch real bugs, and developing LLMs specifically oriented to unit testing.
Load-bearing premise
The whole review rests on the assumption that the search and screening procedure, including the 8-out-of-10 quality threshold, actually captured the full body of relevant LLM unit-testing research and excluded only low-quality work; if the corpus is biased or incomplete, the trend and gap analysis built on it would be unreliable.
Editorial extensions
If this is right
- Follow-up work will likely use the task taxonomy as a checklist to identify underexplored niches; the paper itself highlights test prioritization and selection as having no LLM studies yet.
- If the field continues to concentrate on a few commercial LLMs, evaluation results will age quickly and reproducibility will depend on open-weight models; the survey's finding of increasing use of open models suggests a shift.
- The paper's observation that most evaluations measure coverage and pass rate rather than real bug detection implies that future benchmarks should be built around bug-finding ability on buggy program versions.
- The proposed end-to-end testing-and-debugging framework, if adopted, would link unit test generation with fault localization and repair in a feedback loop.
Reading between the lines
- The survey's corpus is heavily skewed toward Java and mainstream benchmarks; a reader should not assume the findings transfer to other languages or to industrial settings without further checks.
- The heavy reliance on prompt engineering in 96 of 105 studies suggests that the marginal cost of applying LLMs to new testing tasks is now low, which may accelerate adoption before rigorous evaluation catches up.
- If the quality criterion requiring publication in a reputable venue was applied strictly, the large share of preprints in the corpus would be hard to justify; interpreting the corpus as 'recent work regardless of venue' rather than 'high-quality peer-reviewed work' is safer.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript presents a systematic literature review of LLM-based unit testing, claiming to be the first such SLR. It reports a three-stage QGS search (manual seed, automated search, snowballing), a ten-criterion quality assessment with an 8/10 threshold, and a final corpus of 105 papers through March 2025. It contributes a taxonomy of unit testing tasks, an analysis of LLM adoption and adaptation strategies, a taxonomy of hybrid uses of traditional techniques, and a challenges-and-opportunities section. The paper concludes that the field is growing rapidly, that test generation and oracle generation dominate, that commercial GPT models and prompt engineering are the dominant choices, and that several unit-testing tasks remain underexplored.
Significance. Should the corpus and quality filter hold up, this would be a timely and useful reference for a rapidly growing community: it consolidates 105 papers into a task taxonomy, an LLM-utilization taxonomy, and a hybrid-technique taxonomy, and it ships a public GitHub artifact. The review also makes concrete, checkable observations (e.g., model and prompting-strategy distributions, evaluation-metric frequencies) and identifies plausible research directions. The contribution is synthetic rather than technical. Its value depends on the representativeness and reproducibility of the corpus, and the methodology reporting currently has several inconsistencies that prevent independent verification.
major comments (4)
- [§3.3.2 and §3.5] QA4 awards a full point only for publication in a reputable venue, yet §3.5 reports that roughly 40% of the 105 papers are arXiv preprints that have not undergone peer review. Under any standard reading, a preprint cannot receive 'yes' on QA4, and the manuscript does not disclose how QA4 was scored for these preprints. The 8/10 threshold is therefore either non-uniform or non-binding as reported, which directly undercuts the 'rigorous assessment process' claimed in §3.5. Because every descriptive finding in §4–§6 is a statement about this corpus, the per-paper QA scores and the operationalization of QA4 for preprints must be reported; the companion repository makes this checkable.
- [§3.3 and §3.5] The automated search is stated to have been run 'at the end of January 2025', but the paper claims coverage until March 2025 and Figure 3 includes 15 papers appearing by March 2025. Please clarify how February–March 2025 papers entered the corpus (e.g., an updated search, a snowballing round, or a separate manual collection). Without this clarification, the stated search period and the reported corpus are not reproducible, which matters for the SLR's completeness claim.
- [§3.3] The text says the search covered 'four widely used databases' but then enumerates only Google Scholar, ACM Digital Library, and IEEE Xplore (listed as 'IEEE Explorer Digital Library'). Either the fourth database must be named or the count corrected. This is a load-bearing detail because the completeness and reproducibility of an SLR corpus depend on the exact source list, and Google Scholar's query semantics and deduplication behavior differ substantially from the other listed libraries.
- [§4 'Summary of Findings' vs. Figure 5] The summary at the end of §4 states that test case generation and oracle generation account for 20% and 19% of the collected papers, respectively, but Figure 5 reports 60.5% for test generation and 9.6% for oracle generation, with a separate 10.5% for assertion generation. These numbers are mutually inconsistent. If 'oracle generation' in the summary includes assertion generation and if 'test generation' refers to a narrower subcategory, that must be stated explicitly; as written, two central quantitative findings of the review contradict each other.
minor comments (3)
- [Abstract/Contributions] The contribution bullet in §1 says '105 high-quality APR papers' and the Additional Key Words list 'Automated Program Repair, LLM4APR'; these appear to be copy-paste remnants from a program-repair survey and should be corrected to unit-testing equivalents.
- [Figure 3] Figure 3 appears to show cumulative counts or mislabeled bars (13, 17, 90, 105) that do not match the per-year counts stated in the text (1, 2, 14, 73, 15). Please relabel the figure to make clear whether the bars are annual or cumulative.
- [Throughout] There are several copyediting issues: 'QA7Are' and 'QA8Are' are missing spaces, §4.2.2 has 'ageneration-and-refinement', §4.7 has 'to accurately to accurately align', and §4.8 has 'unit test cases tend to grow when software evolves'.
Circularity Check
No circular reasoning: the survey's claims are descriptive syntheses of an externally collected corpus, with no prediction derived from a fitted input.
full rationale
This paper is a systematic literature review rather than a derived quantitative result. The central claims are (a) that it is the first SLR specifically on LLM-based unit testing and (b) a taxonomy and trend analysis of 105 collected papers. Claim (a) is established by comparison with the prior surveys listed in Table 1; it does not depend on the authors' own prior work, and its truth value is independent of this paper's methodology. Claim (b) is a descriptive aggregation of the collected corpus, not a prediction. The QGS search protocol is attributed to an external methodology source [124], and the authors' own prior survey [126] is cited only as a prior application of that protocol, not as the origin of its validity. The authors' primary studies appear in the corpus as objects of study, which is normal for an SLR and not load-bearing for the survey's methodology. The QA4 'reputable venue' criterion versus the inclusion of roughly 40% arXiv preprints is a legitimate methodological transparency concern about whether the quality threshold was applied uniformly; however, it is a validity and reproducibility issue, not a circularity issue, because no claim in the paper is defined in terms of another claim or reduced to a fitted parameter. No equation, definition, or self-citation chain makes an output equal to an input by construction.
Assumptions & free parameters
assumptions (3)
- domain assumption The selected venues and databases provide comprehensive coverage of the LLM-for-unit-testing literature.
- domain assumption The ten quality assessment questions, with a threshold of 8/10, effectively filter out low-quality or irrelevant papers without excluding valid ones.
- domain assumption The categorization of papers into unit testing tasks and LLM utilization strategies is objective and reproducible.
Cite this review
Pith. "Pith review of Large Language Models for Unit Testing: A Systematic Literature Review." pith.science (2026). https://pith.science/paper/V7TO3EW2
@misc{pith2026250615227,
author = {Pith},
title = {Pith review of: Large Language Models for Unit Testing: A Systematic Literature Review},
year = {2026},
howpublished = {\url{https://pith.science/paper/V7TO3EW2}},
note = {Machine review of arXiv:2506.15227}
}
read the original abstract
Unit testing is a fundamental practice in modern software engineering, with the aim of ensuring the correctness, maintainability, and reliability of individual software components. Very recently, with the advances in Large Language Models (LLMs), a rapidly growing body of research has leveraged LLMs to automate various unit testing tasks, demonstrating remarkable performance and significantly reducing manual effort. However, due to ongoing explorations in the LLM-based unit testing field, it is challenging for researchers to understand existing achievements, open challenges, and future opportunities. This paper presents the first systematic literature review on the application of LLMs in unit testing until March 2025. We analyze \numpaper{} relevant papers from the perspectives of both unit testing and LLMs. We first categorize existing unit testing tasks that benefit from LLMs, e.g., test generation and oracle generation. We then discuss several critical aspects of integrating LLMs into unit testing research, including model usage, adaptation strategies, and hybrid approaches. We further summarize key challenges that remain unresolved and outline promising directions to guide future research in this area. Overall, our paper provides a systematic overview of the research landscape to the unit testing community, helping researchers gain a comprehensive understanding of achievements and promote future research. Our artifacts are publicly available at the GitHub repository: https://github.com/iSEngLab/AwesomeLLM4UT.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 8 Pith papers
-
Requirements-Augmented Generation for Trustworthy Acceptance Testing of LLM-Based Software
REAG and a confidence-calibrated cascade generate context-aware test oracles for LLM-based software and produce statistically controlled verdict reliability, demonstrated on a production nutrition advisory app.
-
Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation
Pressure-style prompt framings were associated with lower functional correctness and more security warnings than neutral prompts in LiveCodeBench, while most other influence tactics had minimal effects.
-
Understanding and Improving Model Editing for Secure Code Generation
Editing a code model's parameters is a stronger defense against known vulnerability types than inference-time filtering, and a new post-edit tuning step largely fixes the coding-quality regression that editing causes.
-
MultiFixer: A Coordinator-Proposer Based Multi-Agent Framework For Fixing Multi-Hunk Bugs
Coordinator-proposer multi-agent repair schedules hunks, proposes candidate patches in parallel, and selects/refines them, fixing 326/835 Defects4J bugs with GPT-3.5 and 420 with Claude-3.5-Sonnet.
-
Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study)
For LLM-generated Java tests, coverage and mutation predict real-bug detection mainly when comparing models on bug-free code, not when the code under test is buggy, and suite size is not a major confounder.
-
SGAgent: Suggestion-Guided LLM-Based Multi-Agent Framework for Repository-Level Software Repair
A three-agent locate-suggest-fix framework with a knowledge-graph toolkit resolves 154/300 SWE-Bench-Lite issues with Claude-3.5, outperforming same-model baselines by 5-10 points.
-
PSearch: Search-based Patch Generation in the Era of LLM-based Automated Program Repair
PSearch applies Monte Carlo Tree Search to LLM patch generation with LLM and test-based rewards, fixing 201 Defects4J bugs and resolving 164 SWE-Bench-Lite issues.
-
ReProAgent: Tool-Augmented Multi-Stage Agentic Generation of Bug Reproduction Tests from Issue Reports
A tool-augmented multi-stage agent reproduces 58–70% of real GitHub issues as fail-to-pass tests, outperforming prior prompt and agent baselines at about $0.14 per issue.
Reference graph
Works this paper leans on
-
[1]
Azat Abdullin, Pouria Derakhshanfar, and Annibale Panichella. 2025. Test Wars: A Comparative Study of SBST, Symbolic Execution, and LLM-Based Approaches to Unit Test Generation.CoRRabs/2501.10200 (2025). https: //doi.org/10.48550/ARXIV.2501.10200 arXiv:2501.10200
-
[2]
Saranya Alagarsamy, Chakkrit Tantithamthavorn, and Aldeida Aleti. 2024. A3Test: Assertion-Augmented Automated Test case generation.Inf. Softw. Technol.176 (2024), 107565. https://doi.org/10.1016/J.INFSOF.2024.107565 1:26 Quanjun Zhang, Chunrong Fang, Siqi Gu, Ye Shang, Zhenyu Chen, and Liang Xiao
arXiv 2024
-
[3]
Saranya Alagarsamy, Chakkrit Tantithamthavorn, Chetan Arora, and Aldeida Aleti. 2024. Enhancing Large Language Models for Text-to-Testcase Generation.CoRRabs/2402.11910 (2024). https://doi.org/10.48550/ARXIV.2402.11910 arXiv:2402.11910
-
[4]
Umar Alkafaween, Ibrahim Albluwi, and Paul Denny. 2025. Automating Autograding: Large Language Models as Test Suite Generators for Introductory Programming.Journal of Computer Assisted Learning41, 1 (2025), e13100. https://doi.org/10.1111/jcal.13100
-
[6]
Luciano Baresi and Matteo Miraz. 2010. TestFul: Automatic Unit-Test Generation for Java Classes. InProceedings of the 32nd ACM/IEEE International Conference on Software Engineering - Volume 2, ICSE 2010, Cape Town, South Africa, 1-8 May 2010, Jeff Kramer, Judith Bishop, Premkumar T. Devanbu, and Sebastián Uchitel (Eds.). ACM, 281–284. https://doi.org/10.1...
arXiv 2010
-
[7]
Barr, Mark Harman, Phil McMinn, Muzammil Shahbaz, and Shin Yoo
Earl T. Barr, Mark Harman, Phil McMinn, Muzammil Shahbaz, and Shin Yoo. 2015. The Oracle Problem in Software Testing: A Survey.IEEE Trans. Software Eng.41, 5 (2015), 507–525. https://doi.org/10.1109/TSE.2014.2372785
arXiv 2015
-
[8]
Vaishnavi Bhargava, Rajat Ghosh, and Debojyoti Dutta. 2024. CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++?CoRRabs/2412.02735 (2024). https://doi.org/10.48550/ARXIV.2412.02735 arXiv:2412.02735
-
[9]
Shreya Bhatia, Tarushi Gandhi, Dhruv Kumar, and Pankaj Jalote. 2024. Unit Test Generation using Generative AI : A Comparative Performance Analysis of Autogeneration Tools. InLLM4CODE@ICSE. 54–61. https://doi.org/10.1145/ 3643795.3648396
arXiv 2024
Show all 132 references
- [10]
-
[11]
Pasareanu, Koushik Sen, Nikolai Tillmann, and Willem Visser
Cristian Cadar, Patrice Godefroid, Sarfraz Khurshid, Corina S. Pasareanu, Koushik Sen, Nikolai Tillmann, and Willem Visser. 2011. Symbolic Execution for Software Testing in Practice: Preliminary Sssessment. InProceedings of the 33rd International Conference on Software Enginee...
2011
-
[12]
Yu, Qiang Yang, and Xing Xie
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. 2024. A Survey on Evaluation of Large Language Models.ACM Trans. Intell. Syst. Technol....
2024 doi
-
[13]
Yinghao Chen, Zehao Hu, Chen Zhi, Junxiao Han, Shuiguang Deng, and Jianwei Yin. 2024. ChatUniTest: A Framework for LLM-Based Test Generation. InCompanion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering, FSE 2024, Porto de Galinha...
2024
-
[14]
Xiang Cheng, Fan Sang, Yizhuo Zhai, Xiaokuan Zhang, and Taesoo Kim. 2025. RUG: Turbo LLM for Rust Unit Test Generation. In2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). IEEE Computer Society, Los Alamitos, CA, USA, 634–634. https://doi.org/10.1109/...
2025
-
[15]
Vitaly Chipounov, Volodymyr Kuznetsov, and George Candea. 2011. S2E: A Platform for In-Vivo Multi-Path Software Analysis. InProceedings of the 16th International Conference on Architectural Support for Programming Languages and Operating Systems, ASPLOS 2011, Newport Beach, CA...
2011
-
[16]
Ermira Daka and Gordon Fraser. 2014. A Survey on Unit Testing Practices and Problems. In25th IEEE International Symposium on Software Reliability Engineering, ISSRE 2014, Naples, Italy, November 3-6, 2014. IEEE Computer Society, 201–211. https://doi.org/10.1109/ISSRE.2014.11
2014 doi
-
[17]
Desmarais
Arghavan Moradi Dakhel, Amin Nikanjam, Vahid Majdinasab, Foutse Khomh, and Michel C. Desmarais. 2024. Effective Test Generation using Pre-trained Large Language Models and Mutation Testing.Inf. Softw. Technol.171 (2024), 107468. https://doi.org/10.1016/J.INFSOF.2024.107468
2024
- [18]
- [19]
-
[20]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human ...
2019
-
[21]
Elizabeth Dinella, Gabriel Ryan, Todd Mytkowicz, and Shuvendu K. Lahiri. 2022. TOGA: A Neural Method for Test Oracle Generation. In44th IEEE/ACM 44th International Conference on Software Engineering, ICSE 2022, Pittsburgh, PA, USA, May 25-27, 2022. ACM, 2130–2141. https://doi....
2022
-
[22]
Angela Fan, Beliz Gokkaya, Mark Harman, Mitya Lyubarskiy, Shubho Sengupta, Shin Yoo, and Jie M. Zhang. 2023. Large Language Models for Software Engineering: Survey and Open Problems. InIEEE/ACM International Conference on Software Engineering: Future of Software Engineering, I...
2023
-
[23]
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Findings of the Association for Computational Linguistics: EMNLP ...
2020 doi
-
[24]
Christopher Foster, Abhishek Gulati, Mark Harman, Inna Harper, Ke Mao, Jillian Ritchey, Hervé Robert, and Shubho Sengupta. 2025. Mutation-Guided LLM-based Test Generation at Meta.arXiv preprint arXiv:2501.12862(2025)
2025 arXiv
-
[25]
Gordon Fraser and Andrea Arcuri. 2011. EvoSuite: Automatic Test Suite Generation for Object-Oriented Software. In SIGSOFT/FSE’11 19th ACM SIGSOFT Symposium on the Foundations of Software Engineering (FSE-19) and ESEC’11: 13th European Software Engineering Conference (ESEC-13),...
2011
-
[26]
Gordon Fraser and Andrea Arcuri. 2013. Whole Test Suite Generation.IEEE Trans. Software Eng.39, 2 (2013), 276–291. https://doi.org/10.1109/TSE.2012.14
2013 doi
-
[27]
Gordon Fraser and Andreas Zeller. 2012. Mutation-Driven Generation of Unit Tests and Oracles.IEEE Trans. Software Eng.38, 2 (2012), 278–292. https://doi.org/10.1109/TSE.2011.93
2012 doi
- [28]
- [29]
- [30]
-
[31]
Vitor Guilherme and Auri Vincenzi. 2023. An Initial Investigation of ChatGPT Unit Test Generation Capability. In8th Brazilian Symposium on Systematic and Automated Software Testing, SAST 2023, Campo Grande, MS, Brazil, September 25-29, 2023, Awdren L. Fontão, Débora M. B. Paiv...
2023
- [32]
-
[33]
Ishrak Hayet, Adam Scott, and Marcelo d’Amorim. 2025. ChatAssert: LLM-Based Test Oracle Generation With External Tools Assistance.IEEE Transactions on Software Engineering51, 1 (2025), 305–319. https://doi.org/10.1109/ TSE.2024.3519159
2025
-
[35]
Yibo He, Jiaming Huang, Hao Yu, and Tao Xie. 2024. An Empirical Study on Focal Methods in Deep-Learning-Based Approaches for Assertion Generation.Proc. ACM Softw. Eng.1, FSE (2024), 1750–1771. https://doi.org/10.1145/3660785
2024 doi
-
[36]
Yifeng He, Jicheng Wang, Yuyang Rong, and Hao Chen. 2024. Data Augmentation by Fuzzing for Neural Test Generation.CoRRabs/2406.08665 (2024). arXiv:2406.08665 [cs.SE] https://arxiv.org/abs/2406.08665
2024 arXiv
- [37]
-
[38]
Dwyer, Sebastian G
Soneya Binta Hossain, Antonio Filieri, Matthew B. Dwyer, Sebastian G. Elbaum, and Willem Visser. 2023. Neural- Based Test Oracle Generation: A Large-Scale Evaluation and Lessons Learned. InProceedings of the 31st ACM Joint European Software Engineering Conference and Symposium...
2023
- [39]
-
[40]
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. 2024. Large Language Models for Software Engineering: A Systematic Literature Review.ACM Trans. Softw. Eng. Methodol.33, 8 (2024), 220:1–220:79. https://doi.org/1...
2024 doi
- [42]
- [43]
-
[45]
René Just, Darioush Jalali, and Michael D. Ernst. 2014. Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs. InInternational Symposium on Software Testing and Analysis, ISSTA ’14, San Jose, CA, USA - July 21 - 26, 2014, Corina S. Pas...
2014
-
[47]
Rafiqul Islam Rabin, Lei Xu, Weidong Shi, and Mohammad Amin Alipour
Rabimba Karanjai, Aftab Hussain, Md. Rafiqul Islam Rabin, Lei Xu, Weidong Shi, and Mohammad Amin Alipour
-
[48]
Shaker Mahmud Khandaker, Fitsum Kifetew, Davide Prandi, and Angelo Susi. 2025. AugmenTest: Enhancing Tests with LLM-Driven Oracles.arXiv preprint arXiv:2501.17461(2025)
2025 arXiv
- [49]
-
[50]
Metin Konuk, Cem Baglum, and Ugur Yayan. 2024. Evaluation of Large Language Models for Unit Test Generation. In2024 Innovations in Intelligent Systems and Applications Conference (ASYU). IEEE, 1–6
2024
- [53]
-
[54]
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, ...
2023
-
[56]
Jun Liu, Jiwei Yan, Yuanyuan Xie, Jun Yan, and Jian Zhang. 2024. Fix the Tests: Augmenting LLMs to Repair Test Cases with Static Collector and Neural Reranker. In35th IEEE International Symposium on Software Reliability Engineering, ISSRE 2024, Tsukuba, Japan, October 28-31, 2...
2024
-
[57]
Zhang, Yudong Han, Yun Ma, Ge Li, and Gang Huang
Kaibo Liu, Yiyang Liu, Zhenpeng Chen, Jie M. Zhang, Yudong Han, Yun Ma, Ge Li, and Gang Huang. 2024. LLM- Powered Test Case Generation for Detecting Tricky Bugs.CoRRabs/2404.10304 (2024), arXiv–2404. https://doi.org/ 10.48550/ARXIV.2404.10304 arXiv:2404.10304
-
[58]
Zhongxin Liu, Kui Liu, Xin Xia, and Xiaohu Yang. 2023. Towards More Realistic Evaluation for Neural Test Oracle Generation. InProceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis, ISSTA 2023, Seattle, W A, USA, July 17-21, 2023, René Jus...
2023
- [59]
-
[60]
Keila Lucas, Rohit Gheyi, Elvys Soares, Márcio Ribeiro, and Ivan Machado. 2024. Evaluating Large Language Models in Detecting Test Smells. InProceedings of the 38th Brazilian Symposium on Software Engineering, SBES 2024, Curitiba, Brazil, September 30 - October 4, 2024. 672–67...
2024
-
[61]
Lei Ma, Cyrille Artho, Cheng Zhang, Hiroyuki Sato, Johannes Gmeiner, and Rudolf Ramler. 2015. GRT: Program- Analysis-Guided Random Testing (T). In30th IEEE/ACM International Conference on Automated Software Engineering, ASE 2015, Lincoln, NE, USA, November 9-13, 2015, Myra B. ...
2015 doi
-
[62]
Antonio Mastropaolo, Nathan Cooper, David Nader-Palacio, Simone Scalabrino, Denys Poshyvanyk, Rocco Oliveto, and Gabriele Bavota. 2023. Using Transfer Learning for Code-Related Tasks.IEEE Trans. Software Eng.49, 4 (2023), 1580–1598. https://doi.org/10.1109/TSE.2022.3183297
2023
-
[63]
Antonio Mastropaolo, Simone Scalabrino, Nathan Cooper, David Nader-Palacio, Denys Poshyvanyk, Rocco Oliveto, and Gabriele Bavota. 2021. Studying the Usage of Text-To-Text Transfer Transformer to Support Code-Related Tasks. In43rd IEEE/ACM International Conference on Software E...
2021
- [65]
-
[66]
Mooney, and Milos Gligoric
Pengyu Nie, Rahul Banerjee, Junyi Jessy Li, Raymond J. Mooney, and Milos Gligoric. 2023. Learning Deep Semantics for Test Completion. In45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023. IEEE, 2111–2123. https://do...
2023
-
[67]
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong
- [68]
-
[69]
OpenAI. 2025. GPT-3.5. URL: https://platform.openai.com/docs/models/gpt-3-5. Lasted accessed: 2025-04-01
2025
-
[70]
OpenAI. 2025. GPT-4o. URL: https://platform.openai.com/docs/models/pt-4o. Lasted accessed: 2025-04-01
2025
-
[71]
Ouédraogo, Abdoul Kader Kaboré, Haoye Tian, Yewei Song, Anil Koyuncu, Jacques Klein, David Lo, and Tegawendé F
Wendkûuni C. Ouédraogo, Abdoul Kader Kaboré, Haoye Tian, Yewei Song, Anil Koyuncu, Jacques Klein, David Lo, and Tegawendé F. Bissyandé. 2024. Large-scale, Independent and Comprehensive study of the power of LLMs for test case generation.CoRRabs/2407.00225 (2024). https://doi.o...
2024 doi
-
[72]
Ouédraogo, Abdoul Kader Kaboré, Haoye Tian, Yewei Song, Anil Koyuncu, Jacques Klein, David Lo, and Tegawendé F
Wendkûuni C. Ouédraogo, Abdoul Kader Kaboré, Haoye Tian, Yewei Song, Anil Koyuncu, Jacques Klein, David Lo, and Tegawendé F. Bissyandé. 2024. LLMs and Prompting for Unit Test Generation: A Large-Scale Evaluation. In Proceedings of the 39th IEEE/ACM International Conference on ...
2024
-
[73]
Lahiri, Michael D
Carlos Pacheco, Shuvendu K. Lahiri, Michael D. Ernst, and Thomas Ball. 2007. Feedback-Directed Random Test Generation. In29th International Conference on Software Engineering (ICSE 2007), Minneapolis, MN, USA, May 20-26,
2007
-
[74]
Ciprian Paduraru, Alin Stefanescu, and Augustin Jianu. 2024. Unit Test Generation using Large Language Models for Unity Game Development. InProceedings of the 1st ACM International Workshop on Foundations of Applied Software Engineering for Games, FaSE4Games 2024, Porto de Gal...
2024
-
[75]
Rongqi Pan, Taher Ahmed Ghaleb, and Lionel C. Briand. 2024. LTM: Scalable and Black-Box Similarity-Based Test Suite Minimization Based on Language Models.IEEE Trans. Software Eng.50, 11 (2024), 3053–3070. https: //doi.org/10.1109/TSE.2024.3469582
2024
- [76]
- [77]
- [78]
-
[79]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.J. Mach. Learn. Res.21 (2020), 140:1–140:67. https://jmlr.org/p...
2020
- [81]
-
[82]
Per Runeson. 2006. A Survey of Unit Testing Practices.IEEE Softw.23, 4 (2006), 22–29. https://doi.org/10.1109/MS. 2006.91
2006 doi
-
[83]
Gabriel Ryan, Siddhartha Jain, Mingyue Shang, Shiqi Wang, Xiaofei Ma, Murali Krishna Ramanathan, and Baishakhi Ray. 2024. Code-Aware Prompting: A Study of Coverage-Guided Test Generation in Regression Setting using LLM. Proc. ACM Softw. Eng.1, FSE (2024), 951–971. https://doi....
2024 doi
-
[84]
Arkadii Sapozhnikov, Mitchell Olsthoorn, Annibale Panichella, Vladimir Kovalenko, and Pouria Derakhshanfar
-
[85]
Max Schäfer, Sarah Nadi, Aryaz Eghbali, and Frank Tip. 2024. An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation.IEEE Trans. Software Eng.50, 1 (2024), 85–105. https://doi.org/10.1109/ TSE.2023.3334955
2024
- [86]
-
[87]
Jiho Shin, Reem Aleithan, Hadi Hemmati, and Song Wang. 2024. Retrieval-Augmented Test Generation: How Far Are We?CoRRabs/2409.12682 (2024). https://doi.org/10.48550/ARXIV.2409.12682 arXiv:2409.12682
2024 doi
-
[88]
InProceedings of the 2024 IEEE/ACM 46th International Conference on Software Engineering: Companion Proceedings, ICSE Companion 2024, Lisbon, Portugal, April 14-20, 2024
TestSpark: IntelliJ IDEA’s Ultimate Test Generation Companion. InProceedings of the 2024 IEEE/ACM 46th International Conference on Software Engineering: Companion Proceedings, ICSE Companion 2024, Lisbon, Portugal, April 14-20, 2024. ACM, 30–34. https://doi.org/10.1145/3639478.3640024
2024
-
[89]
Jiho Shin, Hadi Hemmati, Moshi Wei, and Song Wang. 2024. Assessing Evaluation Metrics for Neural Test Oracle Generation.IEEE Trans. Software Eng.50, 9 (2024), 2337–2349. https://doi.org/10.1109/TSE.2024.3433463
2024
-
[90]
Mohammed Latif Siddiq, Joanna Cecilia da Silva Santos, Ridwanul Hasan Tanvir, Noshin Ulfat, Fahmid Al Rifat, and Vinicius Carvalho Lopes. 2024. Using Large Language Models to Generate JUnit Tests: An Empirical Study. In Proceedings of the 28th International Conference on Evalu...
2024
- [91]
-
[92]
Jiho Shin, Sepehr Hashtroudi, Hadi Hemmati, and Song Wang. 2024. Domain Adaptation for Code Model-Based Unit Test Case Generation. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing Large Language Models for Unit Testing: A Systematic Literature ...
2024
-
[93]
Weifeng Sun, Zhenting Guo, Meng Yan, Zhongxin Liu, Yan Lei, and Hongyu Zhang. 2024. Method-Level Test-to-Code Traceability Link Construction by Semantic Correlation Learning.IEEE Trans. Software Eng.50, 10 (2024), 2656–2676. https://doi.org/10.1109/TSE.2024.3449917
2024
-
[94]
Hamed Taherkhani and Hadi Hemmati. 2024. VALTEST: Automated Validation of Language Model Generated Test Cases.CoRRabs/2411.08254 (2024). https://doi.org/10.48550/ARXIV.2411.08254 arXiv:2411.08254
2024 doi
- [95]
-
[96]
André Storhaug and Jingyue Li. 2024. Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Gener- ation: An Empirical Study.CoRRabs/2411.02462 (2024). https://doi.org/10.48550/ARXIV.2411.02462 arXiv:2411.02462
2024 doi
- [97]
- [98]
-
[99]
Michele Tufano, Dawn Drain, Alexey Svyatkovskiy, Shao Kun Deng, and Neel Sundaresan. 2020. Unit Test Case Generation with Transformers and Focal Context.CoRRabs/2009.05617 (2020), arXiv–2009
2020 arXiv
-
[100]
Yutian Tang, Zhijie Liu, Zhichao Zhou, and Xiapu Luo. 2024. ChatGPT vs SBST: A Comparative Assessment of Unit Test Suite Generation.IEEE Trans. Software Eng.50, 6 (2024), 1340–1359. https://doi.org/10.1109/TSE.2024.3382365
2024
-
[101]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. InAdvances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 201...
2017
- [102]
-
[103]
Hailong Wang, Tongtong Xu, and Bei Wang. 2024. Deep Multiple Assertions Generation. InProceedings of the 2024 IEEE/ACM First International Conference on AI Foundation Models and Software Engineering, FORGE 2024, Lisbon, Portugal, 14 April 2024, David Lo, Xin Xia, Massimiliano ...
2024
-
[104]
Michele Tufano, Dawn Drain, Alexey Svyatkovskiy, and Neel Sundaresan. 2022. Generating Accurate Assert Statements for Unit Test Cases using Pretrained Transformers. InIEEE/ACM International Conference on Automation of Software Test, AST@ICSE 2022, Pittsburgh, PA, USA, May 21-2...
2022
-
[105]
Bissyandé, and Xiaoguang Mao
Shangwen Wang, Mingyang Geng, Bo Lin, Zhensu Sun, Ming Wen, Yepang Liu, Li Li, Tegawendé F. Bissyandé, and Xiaoguang Mao. 2023. Natural Language to Code: How Far Are We?. InProceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundati...
2023
-
[106]
Simin Wang, Liguo Huang, Amiao Gao, Jidong Ge, Tengfei Zhang, Haitao Feng, Ishna Satyarth, Ming Li, He Zhang, and Vincent Ng. 2023. Machine/Deep Learning for Software Engineering: A Systematic Literature Review.IEEE 1:32 Quanjun Zhang, Chunrong Fang, Siqi Gu, Ye Shang, Zhenyu ...
2023
-
[107]
Shangwen Wang, Bo Lin, Zhensu Sun, Ming Wen, Yepang Liu, Yan Lei, and Xiaoguang Mao. 2023. Two Birds with One Stone: Boosting Code Generation and Code Search via a Generative Adversarial Network.Proc. ACM Program. Lang.7, OOPSLA2 (2023), 486–515. https://doi.org/10.1145/3622815
2023 doi
-
[108]
Junjie Wang, Yuchao Huang, Chunyang Chen, Zhe Liu, Song Wang, and Qing Wang. 2024. Software Testing With Large Language Models: Survey, Landscape, and Vision.IEEE Trans. Software Eng.50, 4 (2024), 911–936. https://doi.org/10.1109/TSE.2024.3368208
2024
-
[110]
Cody Watson, Nathan Cooper, David Nader-Palacio, Kevin Moran, and Denys Poshyvanyk. 2022. A Systematic Literature Review on the Use of Deep Learning in Software Engineering Research.ACM Trans. Softw. Eng. Methodol. 31, 2 (2022), 32:1–32:58. https://doi.org/10.1145/3485275
2022 doi
-
[111]
Cody Watson, Michele Tufano, Kevin Moran, Gabriele Bavota, and Denys Poshyvanyk. 2020. On Learning Meaningful Assert Statements for Unit Test Cases. InICSE ’20: 42nd International Conference on Software Engineering, Seoul, South Korea, 27 June - 19 July, 2020, Gregg Rothermel ...
2020
-
[112]
Joty, and Steven C
Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi. 2021. CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. InProceedings of the 2021 Conference on Em- pirical Methods in Natural Language Processing, EMNLP 2021,...
2021
- [113]
-
[114]
Lin Yang, Chen Yang, Shutao Gao, Weijing Wang, Bo Wang, Qihao Zhu, Xiao Chu, Jianyi Zhou, Guangtai Liang, Qianxiang Wang, and Junjie Chen. 2024. On the Evaluation of Large Language Models in Unit Test Generation. In Proceedings of the 39th IEEE/ACM International Conference on ...
2024
-
[115]
Yanming Yang, Xin Xia, David Lo, and John C. Grundy. 2022. A Survey on Deep Learning for Software Engineering. ACM Comput. Surv.54, 10s (2022), 206:1–206:73. https://doi.org/10.1145/3505243
2022 doi
-
[116]
Danni Xiao, Yimeng Guo, Yanhui Li, and Lin Chen. 2024. Optimizing Search-Based Unit Test Generation with Large Language Models: An Empirical Study. InProceedings of the 15th Asia-Pacific Symposium on Internetware, Internetware 2024, Macau, SAR, China, July 24-26, 2024, Hong Me...
2024
-
[117]
Eric Wong, and Nicholas Chau
Gaolei Yi, Zizhao Chen, Zhenyu Chen, W. Eric Wong, and Nicholas Chau. 2023. Exploring the Capability of ChatGPT in Test Generation. In23rd IEEE International Conference on Software Quality, Reliability, and Security, QRS 2023 Companion, Chiang Mai, Thailand, October 22-26, 202...
2023
-
[118]
Xin Yin, Chao Ni, Xinrui Li, Liushan Chen, Guojun Ma, and Xiaohu Yang. 2025. Enhancing LLM’s Ability to Generate More Repository-Aware Unit Tests Through Precise Contextual Information Injection.CoRRabs/2501.07425 (2025)
2025 arXiv
- [119]
- [120]
-
[121]
Zhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang, Yixuan Chen, Xin Peng, and Yiling Lou. 2024. Evaluating and Improving ChatGPT for Unit Test Generation.Proc. ACM Softw. Eng.1, FSE (2024), 1703–1726. https://doi.org/10. 1145/3660783
2024
- [122]
-
[123]
Zulfa Zakaria, Rodziah Binti Atan, Abdul Azim Abdul Ghani, and Nor Fazlida Mohd Sani. 2009. Unit Testing Approaches for BPEL: A Systematic Review. In16th Asia-Pacific Software Engineering Conference, APSEC 2009, 1-3 December 2009, Batu Ferringhi, Penang, Malaysia, Shahida Sula...
2009 doi
- [124]
- [125]
-
[126]
Quanjun Zhang, Chunrong Fang, Yuxiang Ma, Weisong Sun, and Zhenyu Chen. 2024. A Survey of Learning-based Automated Program Repair.ACM Trans. Softw. Eng. Methodol.33, 2 (2024), 55:1–55:69. https://doi.org/10.1145/3631974
2024 doi
-
[127]
Quanjun Zhang, Chunrong Fang, Weisong Sun, Yan Liu, Tieke He, Xiaodong Hao, and Zhenyu Chen. 2024. APPT: Boosting Automated Patch Correctness Prediction via Fine-Tuning Pre-Trained Models.IEEE Trans. Software Eng.50, 3 (2024), 474–494. https://doi.org/10.1109/TSE.2024.3354969
2024
-
[128]
He Zhang, Muhammad Ali Babar, and Paolo Tell. 2011. Identifying Relevant Studies in Software Engineering.Inf. Softw. Technol.53, 6 (2011), 625–637. https://doi.org/10.1016/J.INFSOF.2010.12.010
2011 doi
-
[129]
Quanjun Zhang, Chunrong Fang, Bowen Yu, Weisong Sun, Tongke Zhang, and Zhenyu Chen. 2024. Pre-Trained Model-Based Automated Software Vulnerability Repair: How Far are We?IEEE Trans. Dependable Secur. Comput.21, 4 (2024), 2507–2525. https://doi.org/10.1109/TDSC.2023.3308897
2024
-
[130]
Quanjun Zhang, Chunrong Fang, Tongke Zhang, Bowen Yu, Weisong Sun, and Zhenyu Chen. 2023. Gamma: Revisiting Template-Based Automated Program Repair Via Mask Prediction. In38th IEEE/ACM International Conference on Automated Software Engineering, ASE 2023, Luxembourg, September ...
2023
-
[131]
Quanjun Zhang, Chunrong Fang, Yi Zheng, Ruixiang Qian, Shengcheng Yu, Yuan Zhao, Jianyi Zhou, Yun Yang, Tao Zheng, and Zhenyu Chen. 2025. Improving Retrieval-Augmented Deep Assertion Generation via Joint Training.IEEE Transactions on Software Engineering(2025), 1–15. https://d...
2025
- [132]
-
[133]
Quanjun Zhang, Ye Shang, Chunrong Fang, Siqi Gu, Jianyi Zhou, and Zhenyu Chen. 2024. TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models.CoRRabs/2409.17561 (2024), arXiv–2409
2024 arXiv
-
[134]
Quanjun Zhang, Weifeng Sun, Chunrong Fang, Bowen Yu, Hongyan Li, Meng Yan, Jianyi Zhou, and Zhenyu Chen
- [135]
-
[136]
Quanjun Zhang, Chunrong Fang, Yi Zheng, Yaxin Zhang, Yuan Zhao, Rubing Huang, Jianyi Zhou, Yun Yang, Tao Zheng, and Zhenyu Chen. 2025. Improving Deep Assertion Generation via Fine-Tuning Retrieval-Augmented Pre-trained Language Models.ACM Trans. Softw. Eng. Methodol.(Feb. 2025...
2025 doi
-
[137]
Zhiyuan Zhong, Sinan Wang, Hailong Wang, Shaojin Wen, Hao Guan, Yida Tao, and Yepang Liu. 2024. Advancing Bug Detection in Fastjson2 with Large Language Models Driven Unit Test Generation.CoRRabs/2410.09414 (2024). https://doi.org/10.48550/ARXIV.2410.09414 arXiv:2410.09414
2024 doi
- [138]
-
[139]
Exploring Automated Assertion Generation via Large Language Models.ACM Trans. Softw. Eng. Methodol. (Oct. 2024). https://doi.org/10.1145/3699598 Just Accepted
2024 doi
-
[141]
Zibin Zheng, Kaiwen Ning, Qingyuan Zhong, Jiachi Chen, Wenqing Chen, Lianghong Guo, Weicheng Wang, and Yanlin Wang. 2025. Towards an Understanding of Large Language Models in Software Engineering Tasks.Empir. Softw. Eng.30, 2 (2025), 50. https://doi.org/10.1007/S10664-024-10602-0
2025 doi
-
[2007]
https://doi.org/10.1109/ICSE.2007.37
IEEE Computer Society, 75–84. https://doi.org/10.1109/ICSE.2007.37
2007 doi
-
[2023]
InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023
CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis. InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net. https://openreview.net/forum?id=iaYcJKpY2B_
2023
- [2024]
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.