REVIEW 4 major objections 4 minor 4 cited by
Mutation-Guided LLM-based Test Generation at Meta
T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Meta's ACH system converts privacy concerns into AI-generated mutant bugs, then writes tests that kill them, hardening code against future regressions.
desk verdict A credible industrial experience report on a novel orchestration, but the 'hardening' claim is an assumption, not a measurement. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is an agentic pipeline of three LLM prompts driven by a single model (Llama 3.1 70B). A 'make a fault' agent rewrites the class under test, guided by the concern narrative and a reference diff, inserting one privacy-relevant bug per method delimited by comment markers; a rule-based filter discards syntactically identical mutants; an 'equivalence detector' LLM-as-judge agent flags semantically equivalent mutants; and a 'make a test to catch fault' agent — a modified version of the prior TestGen-LLM workflow — writes tests that fail on the mutant but pass on the original. The workflow demands five assurances: buildable, valid (passing, non-flaky) regression tests, hardening (killing a fault no existing test catches), relevance to the concern, and stylistic conformity to existing tests.
What would settle it
Take a set of historical privacy-related regression commits from the involved platforms, run the ACH workflow on the pre-change classes, and check whether the generated tests fail on the real changed code. If the fraction of ACH tests that catch the actual historical regressions is near zero, the hardening claim would be unsupported even though acceptance rates stay high.
Extended reading notes
Core claim
ACH's core discovery is that an LLM-based agentic mutation workflow can generate few, highly specific, currently uncaught faults for a stated concern and then generate tests that kill those faults, thereby providing verifiable assurances: the tests build, pass consistently on the original code, and fail on the mutated code. The workflow terminates for a class once a buildable, passing, believed-to-be-non-equivalent mutant yields a test, so it deliberately produces far fewer mutants than rule-based mutation testing. In evaluation, 73% of the 191 test-a-thon reviews accepted the generated tests, 36% were judged possibly or definitely privacy-relevant, and 49% of the mutant-killing tests added no line coverage, evidence that mutation-based adequacy finds faults that coverage-based targeting misses. The paper also claims an LLM equivalence-detector, combined with a rule-based pre-processor that strips added comments, reaches precision 0.95 and recall 0.96 for weeding out equivalent mutants.
Load-bearing premise
The central bet is that an LLM-planted bug that passes all existing tests resembles a real future regression of the concern type, so a test that kills it genuinely hardens the platform; the paper concedes there is no consistent way to measure fault similarity (the Mutant Relevance problem).
Editorial extensions
If this is right
- Tests generated by ACH caught faults no existing test could catch in 571 cases, so the platforms' regression suites are stronger than their line-coverage numbers would suggest.
- Because 49% of the mutant-killing tests added no line coverage, judging test value purely by coverage would discard valuable regression guards.
- The high precision and recall of the combined equivalence detector (0.95/0.96) suggests the equivalent-mutant problem, historically a barrier to industrial mutation testing, is tractable for LLM-generated mutants, at least for Kotlin and Android code.
- The workflow generalizes beyond privacy to any concern expressible in text, including security, integrity, or regulatory requirements.
- Mutation-as-RAG, using generated mutants as prompts, is proposed as a route to coverage-guided test generation, since ACH tests also raised coverage in 51% of cases.
Reading between the lines
- If the planted mutants are not representative of real future privacy faults — the paper's own 'Mutant Relevance' problem — the 73% acceptance measures perceived usefulness rather than actual hardening; a direct test would compare ACH test failures against historical regression commits.
- The same pipeline could be pointed at other concern narratives (security, data integrity, regulatory compliance) with little change, since the concern enters only through the prompt; the open question is whether mutant relevance holds as well there as for privacy.
- The learning effect seen between Messenger's two phases suggests engineer calibration, not just test quality, inflates acceptance; future deployments should randomize reviewer assignment to separate these effects.
- Because the equivalence detector's success relies on the bimodality of LLM-generated mutants (clearly semantically changed or clearly comment-only), gains may shrink if the fault generator is fine-tuned to produce subtler mutants.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes ACH, an agentic LLM-based workflow deployed at Meta for mutation-guided test generation. Given a natural-language issue of concern (here privacy), ACH prompts an LLM to rewrite each method of a Kotlin class with a 'typical bug' related to that concern, filters the resulting mutants through an LLM equivalence detector, and then prompts a second LLM to generate unit tests that fail on the mutated class and pass on the original class. The paper reports deployment over 10,795 Android Kotlin classes in seven Meta platforms, yielding 9,095 buildable mutants, 4,660 mutants 'believed non-equivalent,' and 571 generated tests, with 73% engineer acceptance and 36% judged privacy-relevant in test-a-thons. It also evaluates the equivalence detector on 381 manually analyzed mutants, reporting precision/recall of 0.79/0.47, rising to 0.95/0.96 with comment-stripping preprocessing. The paper concludes that killing such mutants hardens the platform against future regressions.
Significance. The deployment experience is significant: this appears to be the first reported large-scale industrial deployment of LLM-based mutation-guided test generation, with generated tests submitted as real diffs through CI and reviewed by engineers. The execution-based assurances (build success, pass on the original, fail on the mutant) are machine-checked, and the paper is transparent about the equivalence-detector evaluation and its limitations. The 73% acceptance and 36% privacy-relevance figures are useful baselines for future work. However, the significance of the central 'hardening' claim is limited by the absence of evidence connecting generated tests to real future regressions, and by the structural issue that each mutant bundles many independent edits, making fault attribution to the privacy-relevant change unverified.
major comments (4)
- [Abstract, Section 1, Section 9, Table 4] The central claim that ACH 'hardens the platform against regressions' is not supported by the reported measurements. The deployment data show that engineers accepted 73% of tests and judged 36% of them privacy-relevant at code-review time, but they do not show that any accepted test has failed on a later real regression, nor that the generated mutants are representative of real privacy faults. Section 8 explicitly concedes that 'we have no way to consistently and reliably measure problem similarity or relevance,' which directly undermines the premise of the hardening claim. Please either provide longitudinal evidence (e.g., accepted tests that later failed on real regressions in CI) or soften the conclusion to state that ACH kills the generated mutants, with the relevance to real faults left as an explicit assumption.
- [Table 1, Figure 1, Section 2] Because the 'Make a fault' prompt asks for each method to be replaced by a buggy version, each mutated class contains many independent edits. The 'Make a test to catch fault' prompt and the workflow only require that the new test fails on the whole mutated class and passes on the original class; no check determines which injected edit the failing assertion depends on. A generated test can therefore kill a mutant through a non-privacy bug in the same class, so classifying the test as a privacy-hardening test is not guaranteed. Please generate one mutant per edit, or verify failure dependence using delta debugging or similar localization, and report how often the privacy-relevant mutation is the one responsible for the test failure.
- [Section 5.2, Tables 5 and 6] The equivalence-detector evaluation has two limitations that should be addressed. First, the text says that because ACH is a unit test generation technology, 'we do not need to consider failed error propagation,' but a unit test that observes a method's output does need the local state change to propagate to an observable result. If the manual ground-truth labels treat a local-state change as non-equivalence without considering propagation, the precision/recall figures in Table 6 measure something closer to weak-mutation detection than to semantic equivalence. Second, only 381 mutants from 4 of the 7 platforms were manually labeled, while deployment relies on 4,660 'believed non-equivalent' classifications; for Facebook Feed, Aloha, Cross-app, and Oculus, Table 8 uses the overall average from Table 5. This extrapolation should be stated explicitly, ideally with per-platform confidence bounds.
- [Section 1, Assurance 3] The boolean assurance 'Hardening: the new tests catch faults that no existing test can catch' is true by construction with respect to the generated mutant: existing tests pass on the mutant, and the new test fails on it. The abstract and conclusions, however, draw the stronger inference that the platform is hardened against future regressions. This inference requires both mutant representativeness and fault attribution, neither of which is established by the workflow or the reported data. Please separate the construction-level guarantee (killing generated mutants) from the empirical relevance claim (hardening against real regressions), and state the latter as a hypothesis supported only indirectly by engineer acceptance and relevance judgments.
minor comments (4)
- [Abstract and Section 2] The word 'agenetic' appears to be a typo for 'agentic'; it occurs in the abstract and in the description of the workflow.
- [Section 5.2 vs. Section 8] The term 'Mutant Relevance' is used for two different concepts: in Section 5.2 it refers to mutants becoming stale when the code changes after generation, while in Section 8 it refers to the similarity of a mutant to the specific fault instance. These should be given distinct names to avoid confusion.
- [Table 2 caption] The caption states that percentages in the 'final four columns' are distributions over mutants that build and pass, while the fourth column reports the percentage of all mutants that build and pass; this wording is confusing because the table has more than four percentage columns. Consider rewording the caption to identify the columns by name.
- [Section 4.2] The comparison of the 73% acceptance rate with previous TestGen-LLM test-a-thons should note that the review pools, reviewer expertise, and instructions were not identical, so the rates are only informally comparable.
Circularity Check
The 'hardening' guarantee is the workflow's filter criterion relabeled; the leap to future-regression protection rests on the unmeasured Mutant Relevance assumption.
-
self definitional
[Abstract; Section 1 (Hardening assurance); Figure 1 and Table 1, prompt 'Make a test to catch fault']
"From these currently uncaught faults, ACH generates tests that can catch them, thereby 'killing' the mutants and consequently hardening the platform against regressions. ... Hardening: The new tests catch faults that no existing test can catch; ... The first three assurances are boolean in nature; they are unequivocal guarantees made by ACH about the tests it proposes."
A mutant counts as 'currently uncaught' precisely because it builds and passes existing tests, and Figure 1 retains only tests that fail on the mutant while passing on the original. The Table 1 prompt asks for 'extra test cases that will fail on the mutant version of the class, but would pass on the correct version.' Therefore the Hardening assurance (3) is the selection criterion itself, restated as an output guarantee, so 'tests catch currently uncaught faults' cannot fail by construction. The abstract's further step, 'consequently hardening the platform against regressions,' adds a separate premise: that LLM-injected mutants represent real future regressions of the concern class.
full rationale
The only substantially circular element is the labeling of the mutant-killing filter as 'hardening': ACH's boolean assurance 3 is defined by the same pass/fail checks used to select tests, so 'the new tests catch faults that no existing test can catch' is true by construction. The further claim of protection against future regressions depends on the Mutant Relevance assumption, which the paper itself identifies as unmeasurable; that is an evidentiary gap rather than a hidden derivation. Self-citations to TestGen-LLM [3] and Assured LLMSE [7] are descriptive and not load-bearing, and the equivalence-detector precision/recall evaluation plus the 73% acceptance and 36% privacy-relevance measurements are independent empirical results. Accordingly, one central 'hardening' statement reduces to its construction while substantial deployment evaluation remains independent, warranting a moderate score rather than a charge of fully circular derivation.
Assumptions & free parameters
assumptions (4)
- domain assumption Mutants generated by the LLM from a textual concern are representative of real faults in that concern class.
- domain assumption A test that kills a mutant is a regression-hardening test, meaning future regressions resembling the mutant will be caught.
- domain assumption Weak mutation semantics, local state change without requiring error propagation to output, suffice for the value of the generated tests.
- domain assumption The equivalence detector's 'believed non-equivalent' verdicts for all 4,660 mutants are accurate enough that tests generated from them are meaningful.
Cite this review
Pith. "Pith review of Mutation-Guided LLM-based Test Generation at Meta." pith.science (2026). https://pith.science/paper/Z4K5VRKE
@misc{pith2026250112862,
author = {Pith},
title = {Pith review of: Mutation-Guided LLM-based Test Generation at Meta},
year = {2026},
howpublished = {\url{https://pith.science/paper/Z4K5VRKE}},
note = {Machine review of arXiv:2501.12862}
}
read the original abstract
This paper describes Meta's ACH system for mutation-guided LLM-based test generation. ACH generates relatively few mutants (aka simulated faults), compared to traditional mutation testing. Instead, it focuses on generating currently undetected faults that are specific to an issue of concern. From these currently uncaught faults, ACH generates tests that can catch them, thereby `killing' the mutants and consequently hardening the platform against regressions. We use privacy concerns to illustrate our approach, but ACH can harden code against {\em any} type of regression. In total, ACH was applied to 10,795 Android Kotlin classes in 7 software platforms deployed by Meta, from which it generated 9,095 mutants and 571 privacy-hardening test cases. ACH also deploys an LLM-based equivalent mutant detection agent that achieves a precision of 0.79 and a recall of 0.47 (rising to 0.95 and 0.96 with simple pre-processing). ACH was used by Messenger and WhatsApp test-a-thons where engineers accepted 73% of its tests, judging 36% to privacy relevant. We conclude that ACH hardens code against specific concerns and that, even when its tests do not directly tackle the specific concern, engineers find them useful for their other benefits.
Figures
Forward citations
Cited by 4 Pith papers
-
YATE: The Role of Test Repair in LLM-Based Unit Test Generation
A test-repair pipeline, combining static analysis and re-prompting, raises LLM-generated unit test coverage and mutation killing by roughly 20-30 percent over a plain prompt baseline on six Java projects.
-
How well LLM-based test generation techniques perform with newer LLM versions?
With newer LLMs, a plainly prompted generation loop matches or beats four engineered test-generation tools on coverage and mutation score, and a class-then-method hybrid cuts LLM queries by about 20%.
-
Large Language Models for Unit Testing: A Systematic Literature Review
The paper presents the first systematic literature review of large language model based unit testing, covering 105 papers up to March 2025.
-
AI-Driven Tools in Modern Software Quality Assurance: An Assessment of Benefits, Challenges, and Future Directions
AI models can generate and execute test cases on a demo app, but the reported flakiness rate masks a more serious false-negative problem.
Reference graph
Works this paper leans on
-
[1]
John Ahlgren, Maria Eugenia Berezin, Kinga Bojarczuk, Elena Dulskyte, Inna Dvortsova, Johann George, Natalija Gucevska, Mark Harman, Maria Lomeli, Erik Meijer, Silvia Sapora, and Justin Spahr-Summers. 2021. Testing Web Enabled Simulation at Scale Using Metamorphic Testing. In International Conference on Software Engineering (ICSE) Software Engineering in ...
work page 2021
-
[2]
John Ahlgren, Kinga Bojarczuk, Sophia Drossopoulou, Inna Dvortsova, Johann George, Natalija Gucevska, Mark Harman, Maria Lomeli, Simon Lucas, Erik Meijer, Steve Omohundro, Rubmary Rojas, Silvia Sapora, Jie M. Zhang, and Norm Zhou. 2021. Facebook’s Cyber–Cyber and Cyber–Physical Digital Twins (keynote paper). In 25th International Conference on Evaluation ...
work page 2021
-
[3]
Nadia Alshahwan, Jubin Chheda, Anastasia Finegenova, Mark Harman, Alexan- dru Marginean, Shubho Sengupta, and Eddy Wang. 2024. Automated unit test improvement using Large Language Models at Meta. In ACM International Con- ference on the Foundations of Software Engineering (FSE 2024) (Porto de Galinhas, Brazil, Brazil)
work page 2024
-
[4]
Nadia Alshahwan, Andrea Ciancone, Mark Harman, Yue Jia, Ke Mao, Alexandru Marginean, Alexander Mols, Hila Peleg, Federica Sarro, and Ilya Zorin. 2019. Some challenges for software testing research (keynote paper). In Proceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA 2019), Beijing, China, July 15-19, 2019 ...
work page 2019
-
[5]
Nadia Alshahwan, Xinbo Gao, Mark Harman, Yue Jia, Ke Mao, Alexander Mols, Taijin Tei, and Ilya Zorin. 2018. Deploying Search Based Software Engineering with Sapienz at Facebook (keynote paper). In 10𝑡ℎ International Symposium on Search Based Software Engineering (SSBSE 2018) . Montpellier, France, 3–45. Springer LNCS 11036
work page 2018
-
[6]
Nadia Alshahwan, Mark Harman, and Alexandru Marginean. 2023. Software Testing Research Challenges: An Industrial Perspective (keynote paper). In 2023 IEEE Conference on Software Testing, Verification and Validation (ICST 2023) . IEEE, 1–10
work page 2023
-
[7]
Nadia Alshahwan, Mark Harman, Alexandru Marginean, Shubho Sengupta, and Eddy Wang. 2024. Assured LLM-Based Software Engineering (keynote paper). In 2𝑛𝑑. ICSE workshop on Interoperability and Robustness of Neural Software Engineering (InteNSE) (Lisbon, Portugal)
work page 2024
-
[8]
Nadia Alshahwan, Mark Harman, Alexandru Marginean, and Eddy Wang. 2024. Observation-based unit test generation at Meta. In Foundations of Software Engi- neering (FSE 2024)
work page 2024
Show all 67 references
-
[9]
Kelly Androutsopoulos, David Clark, Haitao Dan, Mark Harman, and Robert Hierons. 2014. An Analysis of the Relationship between Conditional Entropy and Failed Error Propagation in Software Testing. In 36𝑡ℎ International Conference on Software Engineering (ICSE 2014) . Hyderabad...
2014
-
[10]
Barr, Yuriy Brun, Premkumar Devanbu, Mark Harman, and Federica Sarro
Earl T. Barr, Yuriy Brun, Premkumar Devanbu, Mark Harman, and Federica Sarro. 2014. The Plastic Surgery Hypothesis. In 22𝑛𝑑 ACM SIGSOFT International Symposium on the Foundations of Software Engineering (FSE 2014) . Hong Kong, China, 306–317
2014
-
[11]
Barr, Mark Harman, Phil McMinn, Muzammil Shahbaz, and Shin Yoo
Earl T. Barr, Mark Harman, Phil McMinn, Muzammil Shahbaz, and Shin Yoo
-
[12]
Moritz Beller, Chu-Pan Wong, Johannes Bader, Andrew Scott, Mateusz Machalica, Satish Chandra, and Erik Meijer. 2021. What it would take to use mutation testing in industry—a study at Facebook. In 2021 IEEE/ACM 43rd International Conference on Software Engineering: Software Eng...
2021
-
[13]
Adam Brown, Sarah D’Angelo, Ambar Murillo, Ciera Jaspan, and Collin Green
-
[14]
Alexander Brownlee, James Callan, Karine Even-Mendoza, Alina Geiger, Carol Hanna, Justyna Petke, Federica Sarro, and Dominik Sobania. 2023. Enhancing genetic improvement mutations using large language models. In International Symposium on Search Based Software Engineering . Sp...
2023
-
[15]
Calcagno, D
C. Calcagno, D. Distefano, J. Dubreil, D. Gabi, P. Hooimeijer, M. Luca, P. W. O’Hearn, I. Papakonstantinou, J. Purbrick, and D. Rodriguez. 2015. Moving Fast with Software Verification. In NASA Formal Methods - 7th International Symposium. 3–11
2015
-
[16]
Francisco Carlos, Mike Papadakis, Vinícius Durelli, and Eduardo Márcio Dela- maro. 2014. Test data generation techniques for mutation testing: A systematic mapping. In Workshop on Experimental Software Engineering (ESELA W’14)
2014
-
[17]
Thierry Titcheu Chekam, Mike Papadakis, Yves Le Traon, and Mark Harman
-
[18]
Yinghao Chen, Zehao Hu, Chen Zhi, Junxiao Han, Shuiguang Deng, and Jianwei Yin. 2024. Chatunitest: A framework for LLM-based test generation. In Compan- ion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering. 572–576
2024
-
[19]
Henry Coles, Thomas Laurent, Christopher Henard, Mike Papadakis, and An- thony Ventresque. 2016. Pit: a practical mutation testing tool for Java. In Pro- ceedings of the 25th international symposium on software testing and analysis . 449–452
2016
-
[20]
DeMillo, Richard J
Richard A. DeMillo, Richard J. Lipton, and Frederick G. Sayward. 1978. Hints on test data selection: Help for the practical programmer. IEEE Computer 11 (1978), 31–41
1978
-
[21]
José Javier Dolado, Mark Harman, Mari Carmen Otero, and Lin Hu. 2003. An empirical investigation of the influence of a type of side effects on program comprehension. IEEE Transactions on Software Engineering 29, 7 (2003), 665–670
2003
-
[22]
Angela Fan, Beliz Gokkaya, Mitya Lyubarskiy, Mark Harman, Shubho Sengupta, Shin Yoo, and Jie Zhang. 2023. Large Language Models for Software Engineering: Survey and Open Problems. In ICSE Future of Software Engineering (FoSE 2023)
2023
-
[23]
Mark Gabel and Zhendong Su. 2010. A Study of the Uniqueness of Source Code. In FSE. 147–156
2010
-
[24]
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al . 2024. A Survey on LLM-as-a-Judge. arXiv preprint arXiv:2411.15594 (2024)
2024 arXiv
-
[25]
Mark Harman, Yue Jia, and William B. Langdon. 2011. Strong Higher Order Mutation-Based Test Data Generation. In 8𝑡ℎ European Software Engineering Conference and the ACM SIGSOFT Symposium on the Foundations of Software Engineering (ESEC/FSE ’11) (Szeged, Hungary). ACM, New York...
2011
-
[26]
Mark Harman and Peter O’Hearn. 2018. From Start-ups to Scale-ups: Opportu- nities and Open Problems for Static and Dynamic Program Analysis (keynote paper). In 18𝑡ℎ IEEE International Working Conference on Source Code Analysis and Manipulation (SCAM 2018) . Madrid, Spain, 1–23
2018
-
[27]
Mark Harman, Xiangjuan Yao, and Yue Jia. 2014. A Study of Equivalent and Stubborn Mutation Operators Using Human Analysis of Equivalence. In 36𝑡ℎ International Conference on Software Engineering (ICSE 2014) . Hyderabad, India, 919–930
2014
-
[28]
Abram Hindle, Earl Barr, Zhendong Su, Prem Devanbu, and Mark Gabel. 2012. On the Naturalness of Software. In International Conference on Software Engineering (ICSE 2012). Zurich, Switzerland
2012
-
[29]
William E. Howden. 1985. Theory and Practice of Functional Testing. IEEE Software 2, 5 (Sept. 1985), 6–17
1985
-
[30]
Ali Reza Ibrahimzada, Yigit Varli, Dilara Tekinoglu, and Reyhaneh Jabbarvand
-
[31]
Davide Italiano and Chris Cummins. 2025. Finding missed code size optimizations in compilers using LLMs. In Compiler Construction. To appear
2025
-
[32]
Andrew T Jebb, Vincent Ng, and Louis Tay. 2021. A review of key Likert scale development advances: 1995–2019. Frontiers in psychology 12, 637547 (2021)
2021
-
[33]
Yue Jia and Mark Harman. 2008. Milu: A Customizable, Runtime-Optimized Higher Order Mutation Testing Tool for the Full C Language. In 3𝑟𝑑 Testing Academia and Industry Conference - Practice and Research Techniques (TAIC PART’08). Windsor, UK, 94–98
2008
-
[34]
Yue Jia and Mark Harman. 2009. Higher Order Mutation Testing. Journal of Information and Software Technology 51, 10 (2009), 1379–1393
2009
-
[35]
Yue Jia and Mark Harman. 2011. An Analysis and Survey of the Development of Mutation Testing. IEEE Transactions on Software Engineering 37, 5 (September– October 2011), 649 – 678
2011
-
[36]
naturalness
Matthieu Jimenez, Thierry Titcheu Chekam, Maxime Cordy, Mike Papadakis, Marinos Kintis, Yves Le Traon, and Mark Harman. 2018. Are mutants really natural?: a study on how "naturalness" helps mutant selection. In Proceedings of the 12th ACM/IEEE International Symposium on Empiri...
2018
-
[37]
René Just. 2014. The Major mutation framework: Efficient and scalable mutation analysis for Java. In Proceedings of the 2014 international symposium on software testing and analysis. 433–436
2014
-
[38]
Ziyu Li and Donghwan Shin. 2024. Mutation-based consistency testing for evaluating the code understanding capability of LLMs. In Proceedings of the IEEE/ACM 3rd International Conference on AI Engineering-Software Engineering for AI. 150–159
2024
-
[39]
Kaibo Liu, Yiyang Liu, Zhenpeng Chen, Jie M Zhang, Yudong Han, Yun Ma, Ge Li, and Gang Huang. 2024. LLM-Powered Test Case Generation for Detecting Tricky Bugs. arXiv preprint arXiv:2404.10304 (2024)
2024 arXiv
-
[40]
Lech Madeyski, Wojciech Orzeszyna, Richard Torkar, and Mariusz Jozala. 2013. Overcoming the equivalent mutant problem: A systematic literature review and a comparative experiment of second order mutation. IEEE Transactions on Software Engineering 40, 1 (2013), 23–42
2013
-
[41]
Ke Mao, Mark Harman, and Yue Jia. 2016. Sapienz: Multi-objective Automated Testing for Android Applications. In International Symposium on Software Testing and Analysis (ISSTA 2016). 94–105. FSE Companion ’25, 23 – 27, 2025, Trondheim, Norway Harman, Sengupta
2016
-
[42]
Meta. 2024. Introducing Llama 3.1: Our most capable models to date. https: //ai.meta.com/blog/meta-llama-3-1/
2024
-
[43]
Milos Ojdanic, Mike Papadakis, and Mark Harman. 2023. Keeping mutation test suites consistent and relevant with long-standing mutants. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 2067–2071
2023
-
[44]
Rafael AP Oliveira, Upulee Kanewala, and Paulo A Nardi. 2014. Automated test oracles: State of the art, taxonomies, and trends. Advances in computers 95 (2014), 113–199
2014
-
[45]
Mike Papadakis, Yue Jia, Mark Harman, and Yves Le Traon. 2015. Trivial Com- piler Equivalence: A Large Scale Empirical Study of a Simple, Fast and Effective Equivalent Mutant Detection Technique. In 37𝑡ℎ International Conference on Software Engineering (ICSE 2015) . Florence, ...
2015
-
[46]
Mike Papadakis, Marinos Kintis, Jie Zhang, Yue Jia, Yves Le Traon, and Mark Harman. 2019. Mutation Testing Advances: An Analysis and Survey. Advances in Computers 112 (2019), 275–378
2019
-
[47]
Haraldsson, Mark Harman, William B
Justyna Petke, Saemundur O. Haraldsson, Mark Harman, William B. Langdon, David R. White, and John R. Woodward. 2018. Genetic Improvement of Software: a Comprehensive Survey. IEEE Transactions on Evolutionary Computation 22, 3 (June 2018), 415–432
2018
-
[48]
Goran Petrović and Marko Ivanković. 2018. State of mutation testing at Google. In Proceedings of the 40th international conference on software engineering: Software engineering in practice. 163–171
2018
-
[49]
Goran Petrović, Marko Ivanković, Gordon Fraser, and René Just. 2021. Practical mutation testing at scale: A view from Google. IEEE Transactions on Software Engineering 48, 10 (2021), 3900–3912
2021
-
[50]
Goran Petrovic, Marko Ivankovic, Bob Kurtz, Paul Ammann, and René Just. 2018. An industrial application of mutation testing: Lessons, challenges, and research directions. In 2018 IEEE International Conference on Software Testing, Verification and Validation Workshops (ICSTW). ...
2018
-
[51]
Cedric Richter and Heike Wehrheim. 2022. Learning realistic mutations: Bug creation for neural bug detectors. In 2022 IEEE Conference on Software Testing, Verification and Validation (ICST). IEEE, 162–173
2022
-
[52]
Gabriel Ryan, Siddhartha Jain, Mingyue Shang, Shiqi Wang, Xiaofei Ma, Mu- rali Krishna Ramanathan, and Baishakhi Ray. 2024. Code-aware prompting: A study of coverage-guided test generation in regression setting using LLM. Pro- ceedings of the ACM on Conference on Foundations o...
2024
-
[53]
Max Schäfer, Sarah Nadi, Aryaz Eghbali, and Frank Tip. 2023. An empirical evaluation of using large language models for automated unit test generation. IEEE Transactions on Software Engineering (2023)
2023
-
[54]
David Schuler and Andreas Zeller. 2009. Javalanche: efficient mutation testing for Java. In 7𝑡ℎ joint meeting of the European Software Engineering Conference and the ACM SIGSOFT International Symposium on Foundations of Software Engineering (ESEC/FSE 2009). 297–298
2009
-
[55]
Richard Speed. 2023. GitHub: 30% of Copilot coding suggestions are ac- cepted. https://www.itpro.com/technology/artificial-intelligence/github-30-of- copilot-coding-suggestions-are-accepted?utm_source=chatgpt.com
2023
-
[56]
Zhao Tian, Honglin Shu, Dong Wang, Xuejie Cao, Yasutaka Kamei, and Junjie Chen. 2024. Large Language Models for Equivalent Mutant Detection: How Far Are We?. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis. 1733–1745
2024
-
[57]
Frank Tip, Jonathan Bell, and Max Schäfer. 2024. LLMorpheus: Mutation Testing using Large Language Models. arXiv preprint arXiv:2404.09952 (2024)
2024 arXiv
-
[58]
Shreshth Tuli, Kinga Bojarczuk, Natalija Gucevska, Mark Harman, Xiao-Yu Wang, and Graham Wright. 2023. Simulation-Driven Automated End-to-End Test and Oracle Inference. In 45th IEEE/ACM International Conference on Software Engi- neering: Software Engineering in Practice, SEIP@...
2023
-
[59]
Lars van Hijfte and Ana Oprescu. 2021. Mutantbench: an equivalent mutant problem comparison framework. In2021 IEEE International Conference on Software Testing, Verification and Validation Workshops (ICSTW). IEEE, 7–12
2021
-
[60]
Bo Wang, Mingda Chen, Youfang Lin, Mike Papadakis, and Jie M Zhang. 2024. An Exploratory Study on Using Large Language Models xfor Mutation Testing. arXiv preprint arXiv:2406.09843 (2024)
2024
-
[61]
Junjie Wang, Yuchao Huang, Chunyang Chen, Zhe Liu, Song Wang, and Qing Wang. 2023. Software Testing with Large Language Model: Survey, Landscape, and Vision. arXiv:2307.07221
2023 arXiv
-
[62]
Jie Zhang, Junjie Chen, Dan Hao, Yingfei Xiong, Bing Xie, Lu Zhang, and Hong Mei. 2014. Search-based inference of polynomial metamorphic relations. In ACM/IEEE International Conference on Automated Software Engineering (ASE’14) , Ivica Crnkovic, Marsha Chechik, and Paul Gruenb...
2014
-
[63]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judg- ing LLM-as-a-judge with MT-bench and chatbot arena. Advances in Neural Information Processing Systems (NeurIPS 2023) 36 (2023), ...
2023
-
[2015]
IEEE Transactions on Software Engineering 41, 5 (May 2015), 507–525
The Oracle Problem in Software Testing: A Survey. IEEE Transactions on Software Engineering 41, 5 (May 2015), 507–525
2015
-
[2017]
In Proceedings of the 39th International Conference on Software Engineering, ICSE 2017, Buenos Aires, Argentina, May 20-28, 2017
An empirical study on mutation, statement and branch coverage fault revelation that avoids the unreliable clean program assumption. In Proceedings of the 39th International Conference on Software Engineering, ICSE 2017, Buenos Aires, Argentina, May 20-28, 2017 . 597–608
2017
-
[2022]
In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering
Perfect is the enemy of test oracle. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 70–81
-
[2024]
In 1st ACM International Conference on AI-powered Software (AIware 2024) (Porto de Galinhas, Brazil)
Identifying the Factors that Influence Trust in AI Code Completion. In 1st ACM International Conference on AI-powered Software (AIware 2024) (Porto de Galinhas, Brazil)
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.