Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Adversarial Reasoning for Repair Based on Inferred Program Intent

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This paper claims that automated repair succeeds by adversarially inferring several program intents and generating intent-specific tests, correctly repairing 77 Defects4J 2.0 bugs and 105 HumanEval-Java bugs under realistic fault…

desk verdict AdverIntent-Agent is a genuine conceptual step for LLM repair—adversarial intents plus in-loop test generation—but the headline counts rest on a single non-deterministic run and an under-reported correctness-label breakdown. read the letter →

arxiv 2505.13008 v2 pith:YN352QLZ submitted 2025-05-19 cs.SE

classification cs.SE
keywords automatedprogramrepairintentinferenceadversarialreasoningLLMmulti-agentsystemtestgenerationoverfittingpatchDefects4JHumanEval-Java
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that automated program repair improves when the repair system first reasons about what the developer intended the function to do, rather than only transforming code until the given tests pass. It proposes AdverIntent-Agent, a three-agent LLM system that infers several deliberately different intended behaviors for a buggy function, generates test cases that force those intents to predict different outputs on the same inputs, and then produces patches for each intent. The central claim is that this adversarial spread raises the probability that at least one inferred intent matches the true developer intent, yielding more correct repairs and filtering out patches that merely overfit the supplied tests. Under realistic fault localization, the system correctly repairs 77 of 835 bugs in Defects4J 2.0 and 105 of 164 bugs in HumanEval-Java, more than the compared tools.

What carries the argument

The central object is the adversarial program intent: a natural-language specification of a function's expected behavior that is intentionally constructed to conflict with previously inferred intents. The paper measures the conflict with an adversarial score, defined as the fraction of oracle tests (same inputs, intent-dependent expected outputs) on which two intents disagree, and requires each new intent to score above $100\%/K$ with the first intent. The second load-bearing piece is dynamic precise prompting in the repair agent, which first asks for the top three root causes of the bug under one intent and then issues a separate patch-generation prompt per root cause. Together these mechanisms convert the untestable question “what did the developer mean?” into a space of testable hypotheses, and they force the patch pool to be diverse by construction.

What would settle it

Run each generated adversarial test against the ground-truth fixed program for a random sample of repaired bugs and count how many assert the wrong expected output; a material share of wrong assertions would mean the tests can reject correct patches, undermining the correctness counts. Alternatively, measure intent-alignment coverage on a fresh benchmark and check whether it stays near the reported 81.7 percent, since that coverage is the method's upper bound.

Watch

Extended reading notes

Core claim

The paper's central claim is that intent diversity, made concrete through adversarial test oracles, is a repair mechanism in its own right. For a buggy function, the reasoning agent produces an initial natural-language statement of expected behavior and then, prompted with “what if the previous intents are incorrect,” produces two more intents that are deliberately distinct. The test agent turns each intent into executable tests, reusing the same inputs but changing expected outputs per intent, and measures an adversarial score as the fraction of tests whose expected outputs differ between intents; intents below a threshold of 100 percent divided by K are regenerated. The repair agent then asks for the top three root causes consistent with each intent and generates a patch per root cause, accepting only patches that pass both the original tests and the intent-specific adversarial tests. The paper reports that this pipeline correctly repaired 77 Defects4J 2.0 bugs and 105 HumanEval-Java bugs under realistic fault localization, and that on 300 sampled Defects4J bugs at least one of the three inferred intents was judged aligned with the ground-truth intent in 81.7 percent of cases.

Load-bearing premise

The load-bearing premise is that at least one of the few LLM-inferred adversarial intents matches the developer's true intent: the paper's own RQ2 finds this in 81.7% of 300 sampled Defects4J bugs, leaving 18.3% where the intent-driven pipeline cannot produce a correct patch through its intended mechanism.

Editorial extensions

If this is right

  • With intent inference built into repair, correct repairs no longer require a single lucky first patch: the paper's alignment study shows three intents cover the true intent in 81.7% of sampled bugs versus 62.0% for the first intent alone.
  • Generated adversarial tests act as a filter during repair rather than after it, and in the evaluation they removed likely-overfitting patches for 12 Defects4J bugs and 7 HumanEval-Java bugs.
  • Adversarial intent exploration also helps locate the bug, improving fault-localization precision by 13.8% and patch-generation success by 19.4% in the ablation study.
  • The developer-facing output changes from a single patch to a set of inferred intents, tests, and patches, so a human can judge intent in natural language instead of reading a code diff only.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The method's ceiling is set by intent coverage: if the LLM never proposes the true intent among the K candidates, a correct patch cannot emerge through the intended mechanism, so any way to widen or sharpen the intent search (better prompts, retrieval from issue reports, more candidates) should directly raise the repair ceiling.
  • The adversarial-score threshold is a diversity heuristic, not a correctness guarantee; one could test whether requiring higher pairwise disagreement between intents actually increases correct-repair yield or instead pushes the model toward implausible intents.
  • Because the pipeline outputs intent descriptions alongside patches, it offers a natural experiment on developer acceptance: if developers can reliably select the aligned intent, then intent-selection accuracy could become a repair metric in its own right.
  • The same three-agent loop should transfer to other languages and defect types, but the bottleneck will likely be oracle quality in the generated tests, since incorrect expected outputs can reject correct patches; measuring oracle accuracy on a held-out set of ground-truth fixes would quantify that risk.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes AdverIntent-Agent, a multi-agent LLM-based program repair system. The system uses a reasoning agent to infer multiple deliberately adversarial program intents and locate faulty statements, a test agent to generate tests that differentiate these intents (using the same inputs with different expected outputs), and a repair agent to generate patches that satisfy each inferred intent. The authors evaluate on Defects4J 2.0 and HumanEval-Java, reporting 77 and 105 correct repairs under realistic fault localization, respectively, and claim state-of-the-art performance compared with prior APR tools. The paper also includes an ablation study (RQ2) examining the contribution of adversarial intents to fault localization and patch generation, an analysis (RQ3) of how generated tests filter overfitting patches, and a token-cost analysis (RQ4).

Significance. If the reported results hold, this is a valuable contribution to the APR literature. The paper is one of the first to shift the focus of LLM-based repair from generating diverse patches to generating diverse, adversarial program-intent hypotheses, using tests as a way to measure the degree of adversarial disagreement. The evaluation on two standard benchmarks with multiple baselines is appropriate in scope, and the authors make several praiseworthy efforts: they use exact-match and manual semantic-equivalence checks in addition to LLM-based assessment, they explicitly acknowledge data leakage and non-determinism in Section 6, and they state that patches and interaction logs will be publicly released. The idea of supplying developers with inferred intents and tests, rather than only patches, is a novel and potentially useful paradigm. However, the central empirical claim (the 77/105 counts) rests on a patch-correctness labeling procedure that is partly circular, and the reported margins over prior work are small enough that labeling and protocol differences could change the conclusion.

major comments (4)
  1. [§4.1.2, Table 3, Table 6]
  2. [§4.2, Table 3]
  3. [§4.3, Table 5, RQ2]
  4. [§4.1, §6, Table 3]
minor comments (5)
  1. [Table 3]
  2. [§3.2, equation for adversarial score]
  3. [§4.1.2]
  4. [§6]
  5. [§2, Figure 1]

Circularity Check

1 steps flagged · score 2.0 of 10

Minor self-referential correctness metric; central repair claims rest on external benchmarks and manual review.

  1. self definitional [Section 4.1.2, Patch Quality Evaluation, Top@N bullet.]
    "• 𝑇𝑜𝑝@𝑁: A metric that evaluates whether at least one of the top-𝑛 generated patches is correct. A patch is considered correct if it passes both the original test cases and the automatically generated test cases."

    The automatically generated test cases are produced by Agenttest from an inferred program intent, and Agentrepair is explicitly tasked with generating patches that ensure 'both the original and adversarial test cases pass' (Section 3.3). Thus a patch that passes its intended test case satisfies the very condition it was prompted to satisfy; labeling that as correct is self-confirming rather than independent evidence that the patch matches developer intent.

full rationale

The central claim is empirical and benchmarked against external corpora (Defects4J 2.0 and HumanEval-Java) and prior tools, so the repair counts are not forced by construction. The core mechanism, inferring multiple adversarial intents and generating patches per intent, is a genuine generative pipeline rather than a rename of the evaluation data. The main caveat is the Top@N correctness definition, which equates correctness with passing automatically generated tests that were themselves derived from the same inferred intent guiding patch generation; this is circular if used as the final correctness label. The paper mitigates this by requiring exact-match or two-author manual semantic equivalence for the reported correct counts, and the manual review is a standard APR practice. Minor self-citations exist (ITER [77] and SelfAPR [74] are used as baselines, and [77] justifies a fault-localization evaluation convention), but they are not load-bearing: the headline comparison does not reduce to those citations. Overall, the circularity is partial and confined to an evaluation metric, not the derivation of the method's outputs.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

No free parameter is fitted to the target repair counts; all are configuration choices (K, threshold, filtering ratio, refinement rounds, temperature). The axioms are the standard assumptions of LLM-driven APR plus the specific intent-coverage premise that the method's success depends on. No new physical or formal entities are introduced.

free parameters (5)
  • K (number of inferred intents) = 3
    Chosen by hand, not tuned: K=3 intents per bug. The adversarial threshold is derived from K as 100%/K = 33.3%.
  • adversarial threshold = 33.3%
    Set as 100%/K, used to accept or regenerate inferred intents. This is a design choice; the paper does not sweep it.
  • test filtering ratio = 70%
    Top 70% of LLM-ranked tests are kept after prioritizing by assertion confidence (Section 4.1).
  • patch refinement rounds = 3
    Each initial patch is refined up to three times based on compilation and execution errors.
  • LLM temperature = 1
    Temperature set to 1.0 to promote diversity in patches.
assumptions (4)
  • domain assumption GPT-4o can reliably infer program intent, localize faults, generate correct test oracles, and produce correct patches from natural-language intents.
    The whole pipeline is LLM-driven; Section 4.1 uses GPT-4o. The paper acknowledges hallucinations as a threat in Section 6.
  • domain assumption The two benchmarks' original test suites encode sufficient developer intent to judge patch plausibility.
    Plausibility is defined as passing original tests (Section 4.1.2); this is standard in APR.
  • domain assumption LLM-extracted ground-truth intent from human-written patches is a reliable oracle for measuring intent alignment in RQ2.
    RQ2 uses GPT-4o to generate and compare ground-truth intents (Section 4.1.3).
  • domain assumption Adversarial intents generated sequentially with criticism prompts are sufficiently diverse to cover the true intent.
    This is the core design premise of Agentreason; RQ2 empirically tests it on 300 bugs, but it is not guaranteed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Adversarial Reasoning for Repair Based on Inferred Program Intent." pith.science (2026). https://pith.science/paper/YN352QLZ

@misc{pith2026250513008,
  author       = {Pith},
  title        = {Pith review of: Adversarial Reasoning for Repair Based on Inferred Program Intent},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YN352QLZ}},
  note         = {Machine review of arXiv:2505.13008}
}
read the original abstract

Automated program repair (APR) has shown promising results, particularly with the use of neural networks. Currently, most APR tools focus on code transformations specified by test suites, rather than reasoning about the program intent and the high-level bug specification. Without a proper understanding of program intent, these tools tend to generate patches that overfit incomplete test suites and fail to reflect the developers intentions. However, reasoning about program intent is challenging. In our work, we propose an approach called AdverIntent-Agent, based on critique and adversarial reasoning. Our approach is novel to shift the focus from generating multiple APR patches to inferring multiple potential program intents. Ideally, we aim to infer intents that are, to some extent, adversarial to each other, maximizing the probability that at least one aligns closely with the developers original intent. AdverIntent-Agent is a multi-agent approach consisting of three agents: a reasoning agent, a test agent, and a repair agent. First, the reasoning agent generates adversarial program intents along with the corresponding faulty statements. Next, the test agent produces adversarial test cases that align with each inferred intent, constructing oracles that use the same inputs but have different expected outputs. Finally, the repair agent uses dynamic and precise LLM prompts to generate patches that satisfy both the inferred program intent and the generated tests. AdverIntent-Agent was evaluated on two benchmarks: Defects4J 2.0 and HumanEval-Java. AdverIntent-Agent correctly repaired 77 and 105 bugs in both benchmarks, respectively.

Figures

Figures reproduced from arXiv: 2505.13008 by the authors.

Figure 1
Figure 1. An illustrative example AdverIntent-Agent to show the difference between initial and adversarial reasoning and corresponding patches and test case generated. NOT correct,...”, producing the adversarial intent to count uppercase vowels at every other position (shown in the second box of Figure 1c). This second prompt identifies only the if condition on line 4 as problematic. AdverIntent-Agent continues this process t… view at source ↗
Figure 2
Figure 2. Overview of AdverIntent-Agent that shows three agents that are responsible for reasoning - Agentreason, test generation - Agenttest, and patch generation - Agentrepair. AdverIntent-Agent begins with a buggy class and failing tests as inputs, and its outputs include new intents, tests, and candidate patches. tests. In contrast, the adversarial score between intent 1 and intent 3 is 75%. Under our settings, all three … view at source ↗
Figure 3
Figure 3. A correct patch generated by AdverIntent-Agent. Answer to RQ1: AdverIntent-Agent repairs 77 in Defects4J and 105 in HumanEval-Java benchmarks. This success is due to effective adversarial reasoning on program intents and dynamic, root-cause￾specific prompt construction. Additionally, AdverIntent-Agent offers comprehensive resources for function reasoning, bug analysis, and test cases, providing a unique package in p… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Adversarial degree between the inferred intents. [PITH_FULL_IMAGE:figures/full_fig_p015_4.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing

    cs.SE 2025-06 conditional novelty 6.0 of 10

    GUIRepair, a cross-modal LLM pipeline that converts issue screenshots into reproduction code and rendered patch screenshots into validation feedback, resolves 157/517 SWE-bench M instances with GPT-4o and 175 with o4-mini.

Reference graph

Works this paper leans on

84 extracted references · 37 canonical work pages · cited by 1 Pith paper

  1. [1]

    E. T. Barr, M. Harman, P. McMinn, M. Shahbaz, and S. Yoo. 2015. The Oracle Problem in Software Testing: A Survey. IEEE Transactions on Software Engineering 41, 05 (may 2015), 507–525. https://doi.org/10.1109/TSE.2014.2372785

  2. [2]

    Clark Barrett, Roberto Sebastiani, Sanjit Seshia, and Cesare Tinelli. 2009. Satisfiability modulo theories. (2009), 1–885

  3. [3]

    Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler. 2024. Graph of Thoughts: Solving Elaborate Problems with Large Language Models. Proceedings of the AAAI Conference on Artificial Intelligence 38, 16 (March 2024), 17682–176...

  4. [4]

    Islem Bouzenia, Premkumar Devanbu, and Michael Pradel. 2024. RepairAgent: An Autonomous, LLM-Based Agent for Program Repair. arXiv:2403.17134 [cs.SE]

  5. [5]

    Chakraborty, Y

    S. Chakraborty, Y. Ding, M. Allamanis, and B. Ray. 2020. CODIT: Code Editing with Tree-Based Neural Models. IEEE Transactions on Software Engineering (2020). https://doi.org/10.1109/TSE.2020.3020502

  6. [6]

    Saikat Chakraborty and Baishakhi Ray. 2021. On Multi-Modal Learning of Editing Source Code. In2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE)

  7. [7]

    Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2024. Teaching Large Language Models to Self-Debug. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id=KuPixIqPiq

  8. [8]

    Z. Chen, S. J. Kommrusch, M. Tufano, L. Pouchet, D. Poshyvanyk, and M. Monperrus. 2019. SEQUENCER: Sequence- to-Sequence Learning for End-to-End Program Repair. IEEE Transactions on Software Engineering (2019)

Show all 84 references
  1. [9]

    Zhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury, and Shin Hwei Tan. 2023. Automated Repair of Programs from Large Language Models. In Proceedings of the 45th International Conference on Software Engineering (ICSE ’23)

  2. [11]

    Xiang Gao, Sergey Mechtaev, and Abhik Roychoudhury. 2019. Crash-avoiding program repair. In Proceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis . 8–18

  3. [12]

    Luca Gazzola, Daniela Micucci, and Leonardo Mariani. 2017. Automatic Software Repair: A Survey. IEEE Transactions on Software Engineering (2017)

  4. [14]

    Google. 2024. Large sequence models for software development activities. Google (2024)

  5. [15]

    Dávid Hidvégi, Khashayar Etemadi, Sofia Bobadilla, and Martin Monperrus. 2024. CigaR: Cost-efficient Program Repair with LLMs. arXiv:2402.06598 [cs.SE]

  6. [16]

    Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou

  7. [17]

    Elkhan Ismayilzada, Md Mazba Ur Rahman, Dongsun Kim, and Jooyong Yi. 2023. Poracle: Testing Patches under Preservation Conditions to Combat the Overfitting Problem of Program Repair. ACM Trans. Softw. Eng. Methodol. 33, 2, Article 44 (dec 2023), 39 pages. https://doi.org/10.11...

  8. [18]

    Yue Jia and Mark Harman. 2010. An analysis and survey of the development of mutation testing. IEEE transactions on software engineering 37, 5 (2010), 649–678

  9. [19]

    Nan Jiang, Kevin Liu, Thibaud Lutellier, and Lin Tan. 2023. Impact of Code Language Models on Automated Program Repair. arXiv:2302.05020 [cs.SE]

  10. [20]

    Nan Jiang, Thibaud Lutellier, Yiling Lou, Lin Tan, Dan Goldwasser, and Xiangyu Zhang. 2023. KNOD: Domain Knowledge Distilled Tree Decoder for Automated Program Repair. In Proceedings of the 45th International Conference on Software Engineering (ICSE 2023)

  11. [21]

    Nan Jiang, Thibaud Lutellier, and Lin Tan. 2021. CURE: Code-Aware Neural Machine Translation for Automatic Program Repair. In Proceedings of the ACM/IEEE 43rd International Conference on Software Engineering

  12. [22]

    Rene Just, Darioush Jalali, and Michael D Ernst. 2014. Defects4J: A database of existing faults to enable controlled testing studies for Java programs. In Proceedings of the 2014 International Symposium on Software Testing and Analysis . ACM, 437–440

  13. [23]

    YoungJae Kim, Seungheon Han, Askar Yeltayuly Khamit, and Jooyong Yi. 2023. Automated Program Repair from Fuzzing Perspective. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis (Seattle, WA, USA) (ISSTA 2023). Association for Comput...

  14. [24]

    Jiaolong Kong, Mingfei Cheng, Xiaofei Xie, Shangqing Liu, Xiaoning Du, and Qi Guo. 2024. ContrastRepair: Enhancing Conversation-Based Automated Program Repair via Contrastive Test Case Pairs. arXiv:2403.01971 [cs.SE]

  15. [25]

    Le, Duc-Hiep Chu, David Lo, Claire Le Goues, and Willem Visser

    Xuan-Bach D. Le, Duc-Hiep Chu, David Lo, Claire Le Goues, and Willem Visser. 2017. JFIX: Semantics-Based Repair of Java Programs via Symbolic PathFinder. In Proceedings of the 26th ACM SIGSOFT International Symposium on Software Testing and Analysis (Santa Barbara, CA, USA) (I...

  16. [26]

    Claire Le Goues, ThanhVu Nguyen, Stephanie Forrest, and Westley Weimer. 2012. GenProg: A generic method for automatic software repair. Software Engineering, IEEE Transactions on 38, 1 (2012), 54–72. https://doi.org/10.1109/TSE. 2011.104

  17. [27]

    Claire Le Goues, ThanhVu Nguyen, Stephanie Forrest, and Westley Weimer. 2012. GenProg: A Generic Method for Automatic Software Repair. IEEE Transactions on Software Engineering 38, 1 (Jan. 2012), 54–72. https://doi.org/10. 1109/TSE.2011.104

  18. [28]

    Cheryl Lee, Chunqiu Steven Xia, Jen tse Huang, Zhouruixin Zhu, Lingming Zhang, and Michael R. Lyu. 2024. A Unified Debugging Approach via LLM-Based Multi-Agent Synergy. arXiv:2404.17153 [cs.SE]

  19. [29]

    Berger, and Stephen N

    Kyla Levin, Nicolas van Kempen, Emery D. Berger, and Stephen N. Freund. 2024. ChatDBG: An AI-Powered Debugging Assistant. arXiv:2403.16354 [cs.SE]

  20. [30]

    Yi Li, Shaohua Wang, and Tien N. Nguyen. 2020. DLFix: Context-Based Code Transformation Learning for Automated Program Repair. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering (Seoul, South Korea) (ICSE ’20). 602–614. https://doi.org/10.1145...

  21. [31]

    Changshu Liu, Pelin Cetin, Yogesh Patodia, Baishakhi Ray, Saikat Chakraborty, and Yangruibo Ding. 2024. Automated Code Editing with Search-Generate-Modify. IEEE Transactions on Software Engineering (2024), 1–12. https://doi.org/ 10.1109/TSE.2024.3376387

  22. [32]

    Bissyandé

    Kui Liu, Anil Koyuncu, Dongsun Kim, and Tegawendé F. Bissyandé. 2019. TBar: Revisiting Template-based Automated Program Repair. In Proceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis . ACM, 31–42. https://doi.org/10.1145/3293882.3330577

  23. [33]

    Liu and H

    X. Liu and H. Zhong. 2018. Mining stackoverflow for program repair. In 2018 IEEE 25th International Conference on Software Analysis, Evolution and Reengineering (SANER)

  24. [34]

    Fan Long and Martin Rinard. 2015. Staged program repair with condition synthesis. In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering . Bergamo Italy, 166–178. https://doi.org/10.1145/2786805.2786811

  25. [35]

    Thibaud Lutellier, Hung Viet Pham, Lawrence Pang, Yitong Li, Moshi Wei, and Lin Tan. 2020. CoCoNuT: Combining Context-Aware Neural Translation Models Using Ensemble for Program Repair (ISSTA 2020). , Vol. 1, No. 1, Article . Publication date: August 2025. Adversarial Reasoning...

  26. [36]

    Marginean, J

    A. Marginean, J. Bader, S. Chandra, M. Harman, Y. Jia, K. Mao, A. Mols, and A. Scott. 2019. SapFix: automated end-to-end repair at scale. In Proceedings of the 41st International Conference on Software Engineering: Software Engineering in Practice (Montreal, Quebec, Canada) (I...

  27. [37]

    Matias Martinez and Martin Monperrus. 2016. ASTOR: A Program Repair Library for Java. In Proceedings of ISSTA

  28. [38]

    Matias Martinez and Martin Monperrus. 2016. Astor: A program repair library for java. In Proceedings of the 25th International Symposium on Software Testing and Analysis . 441–444

  29. [39]

    Sergey Mechtaev, Xiang Gao, Shin Hwei Tan, and Abhik Roychoudhury. 2018. Test-Equivalence Analysis for Automatic Patch Generation. ACM Trans. Softw. Eng. Methodol.27, 4, Article 15 (oct 2018), 37 pages. https://doi.org/10.1145/3241980

  30. [40]

    Mechtaev, J

    S. Mechtaev, J. Yi, and A. Roychoudhury. 2015. DirectFix: Looking for Simple Program Repairs. In 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering , Vol. 1. 448–458. https://doi.org/10.1109/ICSE.2015.63

  31. [41]

    Sergey Mechtaev, Jooyong Yi, and Abhik Roychoudhury. 2016. Angelix: Scalable Multiline Program Patch Synthesis via Symbolic Analysis. In 2016 IEEE/ACM 38th International Conference on Software Engineering (ICSE)

  32. [42]

    Martin Monperrus. 2017. Automatic Software Repair: a Bibliography. ACM Computing Surveys 51 (2017), 1–24. https://doi.org/10.1145/3105906

  33. [43]

    Hoang Duong Thien Nguyen, Dawei Qi, Abhik Roychoudhury, and Satish Chandra. 2013. SemFix: Program repair via semantic analysis. In 2013 35th International Conference on Software Engineering (ICSE) . 772–781. https://doi.org/10. 1109/ICSE.2013.6606623

  34. [44]

    Olausson, Jeevana Priya Inala, Chenglong Wang, Jianfeng Gao, and Armando Solar-Lezama

    Theo X. Olausson, Jeevana Priya Inala, Chenglong Wang, Jianfeng Gao, and Armando Solar-Lezama. 2024. Is Self-Repair a Silver Bullet for Code Generation?. In International Conference on Learning Representations (ICLR)

  35. [45]

    Zichao Qi, Fan Long, Sara Achour, and Martin Rinard. 2015. An Analysis of Patch Plausibility and Correctness for Generate-and-Validate Patch Generation Systems. In Proceedings of the 2015 International Symposium on Software Testing and Analysis (Baltimore, MD, USA) (ISSTA 2015...

  36. [46]

    Haifeng Ruan, Yuntong Zhang, and Abhik Roychoudhury. 2025. SpecRover: Code Intent Extraction via LLMs. In Proceedings of the 47th International Conference on Software Engineering (ICSE 2025) . arXiv:2408.02232 https://arxiv. org/abs/2408.02232

  37. [47]

    Seemanta Saha et al. 2019. Harnessing evolution for multi-hunk program repair. In 2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE) . IEEE, 13–24

  38. [48]

    Ridwan Shariffdeen, Yannic Noller, Lars Grunske, and Abhik Roychoudhury. 2021. Concolic Program Repair. In 42nd ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI)

  39. [49]

    Ridwan Shariffdeen, Shin Hwei Tan, Mingyuan Gao, and Abhik Roychoudhury. 2021. Automated Patch Transplantation. In ACM Transactions on Software Engineering and Methodology (TOSEM) . 1–36

  40. [50]

    Smith, Earl T

    Edward K. Smith, Earl T. Barr, Claire Le Goues, and Yuriy Brun. 2015. Is the Cure Worse Than the Disease? Overfitting in Automated Program Repair. In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering (ESEC/FSE 2015)

  41. [51]

    Shin Hwei Tan and Abhik Roychoudhury. 2015. relifix: Automated Repair of Software Regressions. In 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering , Vol. 1. 471–482. https://doi.org/10.1109/ICSE.2015.65

  42. [52]

    Prasad, and Abhik Roychoudhury

    Shin Hwei Tan, Hiroaki Yoshida, Mukul R. Prasad, and Abhik Roychoudhury. 2016. Anti-patterns in Search-based Program Repair. In Proceedings of the 2016 24th ACM SIGSOFT International Symposium on Foundations of Software Engineering (Seattle, WA, USA) (FSE 2016)

  43. [53]

    Bissyande

    Haoye Tian, Yinghua Li, Weiguo Pian, Abdoul Kader Kaboré, Kui Liu, Jacques Klein, and Tegawendé F. Bissyande

  44. [54]

    Bissyandé

    Haoye Tian, Kui Liu, Abdoul Kader Kaboré, Anil Koyuncu, Li Li, Jacques Klein, and Tegawendé F. Bissyandé. 2020. Evaluating Representation Learning of Code Changes for Predicting Patch Correctness in Program Repair. In ASE. IEEE, 981–992. https://doi.org/10.1145/3324884.3416532

  45. [55]

    Michele Tufano, Jevgenija Pantiuchina, Cody Watson, Gabriele Bavota, and Denys Poshyvanyk. 2019. On Learning Meaningful Code Changes via Neural Machine Translation. In Proceedings of the 41st International Conference on Software Engineering (Montreal, Quebec, Canada) (ICSE ’19...

  46. [56]

    Bing Wang, Yan Gao, Zhoujun Li, and Jian-Guang Lou. 2023. Know What I don’t Know: Handling Ambiguous and Unknown Questions for Text-to-SQL. In Findings of the Association for Computational Linguistics: ACL 2023 , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Asso...

  47. [57]

    Shangwen Wang, Ming Wen, Bo Lin, Hongjun Wu, Yihao Qin, Deqing Zou, Xiaoguang Mao, and Hai Jin. 2020. Automated Patch Correctness Assessment: How Far are We?. InASE. ACM

  48. [58]

    Yuxiang Wei, Chunqiu Steven Xia, and Lingming Zhang. 2023. Copiloting the Copilots: Fusing Large Language Models with Completion Engines for Automated Program Repair. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations...

  49. [59]

    Westley Weimer, ThanhVu Nguyen, Claire Le Goues, and Stephanie Forrest. 2009. Automatically finding patches using genetic programming. In 2009 IEEE 31st International Conference on Software Engineering . 364–374. https: //doi.org/10.1109/ICSE.2009.5070536

  50. [60]

    Chu-Pan Wong, Priscila Santiesteban, Christian Kästner, and Claire Le Goues. 2021. VarFix: Balancing Edit Ex- pressiveness and Search Effectiveness in Automated Program Repair. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposi...

  51. [61]

    Chunqiu Steven Xia and Lingming Zhang. 2022. Less Training, More Repairing Please: Revisiting Automated Program Repair via Zero-Shot Learning. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering...

  52. [62]

    Chunqiu Steven Xia and Lingming Zhang. 2024. Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA)

  53. [63]

    Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, Ruoxi Jia, Bo Li, Kai Li, Danqi Chen, Peter Henderson, and Prateek Mittal. 2024. SORRY-Bench: Systematically Evaluating Large Language Model Saf...

  54. [64]

    Qi Xin and Steven Reiss. 2019. Better Code Search and Reuse for Better Program Repair. In2019 IEEE/ACM International Workshop on Genetic Improvement (GI). 10–17. https://doi.org/10.1109/GI.2019.00012

  55. [65]

    Qi Xin and Steven P. Reiss. 2017. Identifying Test-Suite-Overfitted Patches through Test Case Generation. In ISSTA

  56. [66]

    Xin and S

    Q. Xin and S. P. Reiss. 2017. Leveraging syntax-related code for automated program repair. In 2017 32nd IEEE/ACM International Conference on Automated Software Engineering (ASE)

  57. [67]

    Yingfei Xiong, Xinyuan Liu, Muhan Zeng, Lu Zhang, and Gang Huang. 2018. Identifying patch correctness in test-based program repair. In Proceedings of the 40th International Conference on Software Engineering

  58. [68]

    Yingfei Xiong, Jie Wang, Runfa Yan, Jiachen Zhang, Shi Han, Gang Huang, and Lu Zhang. 2017. Precise Condition Synthesis for Program Repair. In Proceedings of the 39th International Conference on Software Engineering (Buenos Aires, Argentina) (ICSE ’17). IEEE Press, 416–426. ht...

  59. [69]

    Jifeng Xuan, Matias Martinez, Favio Demarco, Maxime Clément, Sebastian Lamelas, Thomas Durieux, Daniel Le Berre, and Martin Monperrus. 2016. Nopol: Automatic Repair of Conditional Statement Bugs in Java Programs. IEEE Transactions on Software Engineering (2016)

  60. [70]

    Guang Yang, Yu Zhou, Xiang Chen, Xiangyu Zhang, Terry Yue Zhuo, and Taolue Chen. 2023. Chain-of-Thought in Neural Code Generation: From and For Lightweight Language Models. arXiv:2312.05562 [cs.SE]

  61. [71]

    Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press

    John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. 2024. SWE-agent: Agent Computer Interfaces Enable Software Engineering Language Models

  62. [72]

    Griffiths, Yuan Cao, and Karthik Narasimhan

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv:2305.10601 [cs.CL]

  63. [73]

    He Ye, Jian Gu, Matias Martinez, Thomas Durieux, and Martin Monperrus. 2021. Automated Classification of Overfitting Patches with Statically Extracted Code Features. IEEE Transactions on Software Engineering (2021). https://doi.org/10. 1109/tse.2021.3071750

  64. [74]

    He Ye, Matias Martinez, Xiapu Luo, Tao Zhang, and Martin Monperrus. 2022. SelfAPR: Self-supervised Program Repair with Test Execution Diagnostics. arXiv preprint arXiv:2203.12755 (2022)

  65. [75]

    He Ye, Matias Martinez, and Martin Monperrus. 2021. Automated patch assessment for program repair at scale. Empirical Software Engineering 26, 2 (2021), 20. https://doi.org/10.1007/s10664-020-09920-w

  66. [76]

    He Ye, Matias Martinez, and Martin Monperrus. 2022. Neural Program Repair with Execution-based Backpropagation. In Proceedings of the ACM/IEEE 44th International Conference on Software Engineering

  67. [77]

    He Ye and Martin Monperrus. 2024. ITER: Iterative Neural Repair for Multi-Location Patches. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering (ICSE ’24) . Article 10, 13 pages. https://doi.org/10. 1145/3597503.3623337

  68. [78]

    Burak Yetiştiren, Işık Özsoy, Miray Ayerdem, and Eray Tüzün. 2023. Evaluating the code quality of ai-assisted code generation tools: An empirical study on github copilot, amazon codewhisperer, and chatgpt. arXiv preprint arXiv:2304.10778 (2023)

  69. [79]

    Yuan Yuan and Wolfgang Banzhaf. 2018. ARJA: Automated Repair of Java Programs via Multi-Objective Genetic Programming. In IEEE Transactions on Software Engineering

  70. [80]

    Yuan Yuan and Wolfgang Banzhaf. 2020. Toward Better Evolutionary Program Repair: An Integrated Approach. ACM Trans. Softw. Eng. Methodol. 29, 1, Article 5 (jan 2020), 53 pages. https://doi.org/10.1145/3360004 , Vol. 1, No. 1, Article . Publication date: August 2025. Adversaria...

  71. [81]

    Kechi Zhang, Jia Li, Ge Li, Xianjie Shi, and Zhi Jin. 2024. CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges. arXiv:2401.07339 [cs.SE]

  72. [82]

    Yuntong Zhang, Haifeng Ruan, Zhiyu Fan, and Abhik Roychoudhury. 2024. AutoCodeRover: Autonomous Program Improvement. arXiv:2404.05427 [cs.SE]

  73. [83]

    Qihao Zhu, Zeyu Sun, Yuan-an Xiao, Wenjie Zhang, Kang Yuan, Yingfei Xiong, and Lu Zhang. 2021. A Syntax-Guided Edit Decoder for Neural Program Repair. InProceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of So...

  74. [84]

    Armin Zirak and Hadi Hemmati. 2024. Improving Automated Program Repair with Domain Adaptation. ACM Trans. Softw. Eng. Methodol. 33, 3, Article 65 (mar 2024), 43 pages. https://doi.org/10.1145/3631972 , Vol. 1, No. 1, Article . Publication date: August 2025

  75. [2022]

    ACM Trans

    Checking Patch Behaviour against Test Specification. ACM Trans. Softw. Eng. Methodol. (2022)

  76. [2024]

    In The Twelfth International Conference on Learning Representations

    Large Language Models Cannot Self-Correct Reasoning Yet. In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=IkmD3fKBPQ

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.