Pith. sign in

REVIEW 3 major objections 3 minor 120 references

PATCH: Empowering Large Language Model with Programmer-Intent Guidance and Collaborative-Behavior Simulation for Automatic Bug Fixing

T0 review · 3 major / 3 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read PATCH claims that an LLM bug-fixing framework, which augments the buggy code with dependence context and the fixing commit's message and then runs four bug-management stages across three ChatGPT agents, resolves 33.97% of BFP bugs at…

desk verdict Solid multi-agent repair pipeline, but the headline Fix@1 gain is built on ground-truth commit messages leaking into both prompts and review; re-run without the oracle before trusting the numbers. read the letter →

arxiv 2501.16149 v2 pith:AULPEEFQ submitted 2025-01-27 cs.SE

classification cs.SE
keywords automaticbugfixinglargelanguagemodelsprogrammerintentcommitmessagemulti-agentcollaborationmanagementpromptengineeringBFPbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that LLM bug fixing improves dramatically when the prompt carries two things current methods omit: repository-level context around the buggy line, and a natural-language statement of what the fixing commit was meant to achieve. It further claims that generating the patch through a four-stage dialogue among ChatGPT agents playing tester, developer, and reviewer beats treating bug fixing as a single end-to-end generation call. On the BFP benchmark of 3,112 Java single-hunk bugs, PATCH achieves Fix@1 of 33.97%, compared with 19.96% for the strongest baseline GPT-4, with comparable gains at Fix@3 and Fix@5. The central claim is that simulating collaborative bug management, not scaling the model, is what unlocks the improvement. This matters because it points toward better automatic repair without fine-tuning.

What carries the argument

The load-bearing mechanism is the augmented prompt combined with a stage-wise agent loop. The augmented buggy content is built by parsing buggy classes with Spoon to extract class- and repository-level dependence context, retrieving similar bug-fixing pairs by BM25, and attaching the fixing commit's message as a natural-language intent signal. The dialogue itself is the engine: ChatGPTTester files a bug report, ChatGPTDeveloper explains the buggy method line-by-line (rubber duck debugging), summarizes fix patterns from retrieved demonstrations, and generates the initial patch, and ChatGPTReviewer checks the patch against the stated fixing goal and iterates with the developer for up to three turns. Producing intermediate natural-language artifacts before the patch is what focuses generation.

What would settle it

Run PATCH on the BFP test set with commit messages withheld or replaced by issue reports written before the fix, and measure Fix@1; if the gap over GPT-4 collapses toward the dependence-context-only ablation, the claimed benefit comes from leaking the ground-truth fix summary rather than from the collaborative simulation.

Watch

Extended reading notes

Core claim

The paper's central claim is that bug fixing should be modelled as a staged, collaborative process rather than a single prompt-to-patch mapping. It claims that adding dependence context (imports, global variables, and invoked-method signatures from the class and repository levels) plus the human-written commit message as programmer intent, and then running four stages (bug reporting, diagnosis, patch generation, verification) across three ChatGPT agents, raises Fix@1 on BFP from 19.96% with GPT-4 to 33.97%. The paper further claims that every designed component contributes, with the reviewer's interactive feedback providing the largest single gain, and that the framework improves five open-source LLMs by relative margins from 9% to 92%.

Load-bearing premise

The whole evaluation assumes the fixing commit's message is available at bug-fixing time and truthfully says what the fix should do; in practice programmers usually write commit messages after finishing the fix, so the model is given ground-truth-derived intent that a deployed system would not have.

Editorial extensions

If this is right

  • If PATCH works as claimed, staged multi-agent prompting can improve zero-shot bug-fixing performance without fine-tuning, and the same prompts lift several open-source LLMs by substantial relative margins.
  • The commit message strongly drives the gains: informative commit messages (those containing Why or What information) account for more than 95% of the contribution, while low-information messages add almost nothing.
  • Reviewer feedback is the largest single component, improving Fix@1 by 9.74 percentage points, and most of that gain comes from the first interaction turn.
  • The framework transfers to Bugs.jar, Bears, and Defects4J with improved exact-match or test-passing scores, though it trails ChatRepair and ThinkRepair on QuixBugs, where those baselines use test-suite feedback and sample many more candidates.
  • PATCH fixes 283 BFP bugs that no baseline fixes, suggesting the approach widens the set of fixable bugs rather than only improving performance on the easiest cases.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A deployment version would need to obtain programmer intent before the fix is known, since the paper's own limitation section admits programmers usually write commit messages after fixing; without that information, the headline gap over GPT-4 may shrink.
  • The reviewer's pass decision is a self-assessment against the commit message rather than an execution check, so PATCH may accept patches that satisfy the stated intent but break tests; replacing the reviewer with a compiler or test runner is a natural extension.
  • PATCH's largest gains appear on replace and mixed edit types, suggesting the staged reasoning helps most when the patch requires searching for new tokens; the benefit for trivial deletion-only bugs should be smaller.
  • The paper's integration test with a fine-tuned repair model found the variant fixes 30 bugs that vanilla PATCH cannot, hinting that hybrid pipelines combining staged prompting with specialized models may be complementary.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. PATCH is a stage-wise, prompt-based bug-fixing framework for single-hunk Java bugs. For each bug, it augments the buggy method with static dependence context and the commit message of the associated fixing commit, then runs three ChatGPT agents (tester, developer, reviewer) through bug reporting, code explanation and pattern summarization, initial patch generation, and iterative patch verification. The paper reports Fix@1 of 33.97% on BFP, 14.01 percentage points above GPT-4, and reports gains on Bugs.jar, Bears, and Defects4J, with ablations attributing the improvement to the augmented content and the multi-agent interaction. The central claim is that adding programmer-intent guidance and collaborative-behavior simulation substantially improves LLM repair performance.

Significance. The four-stage decomposition is well motivated, and the paper contributes a released artifact, a large ablation study, and experiments across several LLMs and APR benchmarks. If the evaluation were free of oracle information, the framework could be a meaningful prompt-engineering contribution to LLM-based program repair. As it stands, however, the headline results are not interpretable as automatic bug-fixing performance: the commit message is mined from the ground-truth fixing commit and is inserted both into the generation prompt (Section 2.2, Figure 4) and into the reviewer's "Desired Fixing Goal" (Section 2.5, Figure 8). Section 5.3 explicitly concedes that programmers normally write commit messages after fixing, so the evaluation uses answer-derived information that would be unavailable at deployment time. The framework may still be valuable, but the central comparison must be re-run without this leakage before the paper's claims can be accepted.

major comments (3)
  1. [§2.2, §2.5, §5.3] The commit message used as [Commit Message] in Figure 4 and as part of the reviewer's [Desired Fixing Goal] in Figure 8 is mined from the ground-truth fixing commit of the same BFP instance. Section 5.3 explicitly states that programmers typically write commit messages after fixing and that the paper assumes otherwise. This is not merely an external-validity caveat: it means PATCH receives a natural-language description of the correct fix while the baselines do not. The Table 3 Fix@1 margin of 14.01 pp over GPT-4 is therefore not a valid demonstration of automatic bug-fixing ability. The authors should re-evaluate without any commit-message input, for example using pre-fix issue reports or no intent text at all, and should also remove the commit message from the reviewer's verification goal.
  2. [§4.3.1 (Figure 16(c)), §5.2] The Defects4J generalizability experiment leaks ground-truth edit information. In Figure 16(c), the prompt for Jsoup-83 uses the commit message "Require the 'insert' operation," which is the edit-operation type of the ground-truth patch, and Section 5.2 similarly uses "requires the 'add' operation" and "requires the 'modify' operation" for Closure-128. This gives PATCH privileged information about the required fix that the baselines do not receive. Consequently, the Defects4J result in Table 7 (169 correct patches) cannot be compared on equal footing with ChatRepair, ThinkRepair, and RepairAgent, and the RQ3 generalizability claim needs to be re-established using prompts that do not encode the ground-truth edit operation.
  3. [§2.5, §4.1.1] Because BFP has no test suites, the reviewer's PASS decision in Algorithm 1 is an LLM judgment made against a 'Desired Fixing Goal' that already contains the ground-truth commit message. This makes the verification stage an oracle-informed filter rather than an independent correctness check, and it can steer iterative patch generation toward the known target, as shown in Figure 14. The paper should report how often the reviewer passes a patch that is not the ground-truth patch and should validate a sample of accepted patches by human inspection or by held-out tests. Without such a check, the iterative part of the Fix@k results is difficult to interpret.
minor comments (3)
  1. [§3.4, Table 3] Levenshtein distance is an edit distance, not a percentage; the improvement from 26.07 to 21.44 should not be described as an improvement of '4.63 percentage points.'
  2. [Table 5] The ablation table's component indicators render as '/reve' strings in the submitted PDF, making it difficult to map each row to its configuration; please use unambiguous symbols such as check and cross marks.
  3. [§2.3.2] The dynamic BM25 threshold is described in prose; please provide the exact formula used to decide when a retrieved demonstration is retained, including how the average length term is computed and how ties are handled.

Circularity Check

3 steps flagged · score 6.0 of 10

Post-fix commit messages leak the ground-truth fix into PATCH's prompts and its reviewer oracle; the BFP Fix@1 gain is not an oracle-free automatic-fixing result.

  1. fitted input called prediction [Section 2.2, Figure 4 (Bug Reporting prompt)]
    "In the bug reporting sub-task, you will receive one [Buggy Content]. Please report the root cause of the given [Buggy Hunk] according to the guidance information from [Commit Message]."

    The commit message is mined from the ground-truth fixing commit for the same BFP instance (Section 3.2: 'we collect bug-fixing commits from GitHub linked to instances within the BFP benchmark'), i.e., it is a natural-language description of the target patch written after the fix. Section 5.3 concedes 'programmers typically write commit messages after fixing buggy code. In this paper, we assumed that programmers write these messages before bug fixing.' Feeding this post-hoc description into the generation prompt as 'programmer intent' gives PATCH the answer's content while the baselines receive only the buggy code, so the Table 3 Fix@1 gap is not a pre-fix automatic-fixing prediction.

  2. fitted input called prediction [Section 2.5, Figure 8 (Patch Verification prompt)]
    "Please assess the correctness of the given [Candidate Patch] based on whether it meets the [Desired Fixing Goal]. The intention of the buggy method is [Method Summary]. The intent of the programmer is [Commit Message]."

    The same post-fix commit message is reused as the reviewer's Desired Fixing Goal. Algorithm 1 iterates until ChatGPTReviewer passes the candidate patch, so the verification criterion itself contains a ground-truth-derived description of the required change. The reviewer is therefore an oracle-conditioned filter rather than an independent correctness check. This also explains the large ablation gain of the reviewer component (9.74 pp in Fix@1 reported in Section 4.2.1): the loop is explicitly told what the correct fix must accomplish.

1 more flagged steps
  1. fitted input called prediction [Section 4.3.1, Figure 16(c) case study (Jsoup-83)]
    "we use the required edit operation (i.e., insert) to fix the bug as the commit message for prompting PATCH."

    Here the 'commit message' is not a general intent summary; it states the exact edit operation of the ground-truth patch ('insert'). PATCH is explicitly told what kind of edit to perform, so the claimed unique fix of Jsoup-83 is a demonstration of following an oracle instruction derived from the answer, not a leakage-free prediction. This is the clearest case where the input is constructed from the target fix itself.

full rationale

The paper's derivation chain is not formally self-referential: it never derives PROMPT from the output patch and no author-authored uniqueness theorem is load-bearing. The central circularity is temporal/empirical. Section 2.1 defines the augmented input C as containing 'the programmer's intent (i.e.,commit message),' and Section 3.2 says these messages come from 'bug-fixing commits from GitHub linked to instances within the BFP benchmark' — i.e., from the commits that contain the ground-truth fix. Section 5.3 explicitly concedes 'programmers typically write commit messages after fixing buggy code' and that PATCH assumes otherwise. Because the same post-fix message is inserted into the tester prompt (Figure 4) and the reviewer's Desired Fixing Goal (Figure 8), both generation and verification are conditioned on a natural-language description of the correct change. The Table 3 Fix@1 margin of +14.01 pp over GPT-4 is therefore not interpretable as an automatic bug-fixing result: it measures how well ChatGPT can exploit a leaked description of the target patch. The paper's own ablations concentrate the gain in exactly this component (commit message adds 4.24 pp to ChatGPT; reviewer adds 9.74 pp; Table 6 shows gains track message informativeness). The RQ3 Jsoup-83 case is even more direct: Section 4.3.1 states 'we use the required edit operation (i.e., insert) to fix the bug as the commit message,' so the showcased unique fix is an oracle-following demo. I score 6 rather than 0-2 because these are not fully independent predictions. The score is not 8-10 because (a) the commit message is not identical to the patch, (b) the Defects4J/QuixBugs and alternate-LLM results are partly independent and test-based, and (c) no self-citation chain is load-bearing. The framework may still be valuable, but the central BFP claim must be re-evaluated without post-fix oracle information.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The method introduces no new physical entities. It does rely on several untested assumptions: perfect fault localization, the availability of the fixing commit message as pre-fix intent, the reliability of an LLM reviewer as a correctness oracle on a benchmark without tests, and the validity of exact-match as a correctness measure. Hyper-parameters such as maxIterNum, the BM25 acceptance threshold, and sampling temperatures are hand-chosen.

free parameters (3)
  • maxIterNum = 3
    Maximum developer-reviewer interaction turns (Algorithm 1); hand-set in Section 3.5 and shown in Figure 13 to affect Fix@1.
  • BM25 dynamic threshold = query length + average context length
    Retrieval acceptance criterion in Section 2.3.2; hand-chosen and not tuned, controls how many demonstrations are shown.
  • Sampling temperatures = 0 (k=1), 0.8 (k>1)
    Section 3.5; greedy for top-1, sampled for multiple patches, chosen to match common practice rather than swept.
assumptions (4)
  • domain assumption Perfect fault localization: the buggy hunk is known in advance
    Section 2.1 states 'This process operates under the assumption of perfect fault localization.' The method is not evaluated as an end-to-end bug locator + fixer.
  • ad hoc to paper Commit messages are available before the fix and serve as faithful programmer intent
    Section 2.2 uses the fixing commit's message as guidance, and Section 5.3 admits programmers typically write commit messages after fixing, so this is an optimistic assumption specific to this paper's evaluation.
  • domain assumption The LLM reviewer's pass/fail assessment is a reliable correctness oracle
    Section 2.5 lets ChatGPTReviewer judge correctness against method summary and commit message; no tests are run on BFP, yet this judgment terminates the search and is treated as ground truth for Fix@1.
  • domain assumption BFP instances are representative and exact-match to the human patch is a valid correctness measure
    Section 3.4 defines Fix@k as exact match to ground truth on BFP, which has no test suites; this treats one specific human-written patch as the only correct answer and can penalize semantically equivalent fixes.

how reviews work

0 comments
Cite this review

Pith. "Pith review of PATCH: Empowering Large Language Model with Programmer-Intent Guidance and Collaborative-Behavior Simulation for Automatic Bug Fixing." pith.science (2026). https://pith.science/paper/AULPEEFQ

@misc{pith2026250116149,
  author       = {Pith},
  title        = {Pith review of: PATCH: Empowering Large Language Model with Programmer-Intent Guidance and Collaborative-Behavior Simulation for Automatic Bug Fixing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AULPEEFQ}},
  note         = {Machine review of arXiv:2501.16149}
}
read the original abstract

Bug fixing holds significant importance in software development and maintenance. Recent research has made substantial strides in exploring the potential of large language models (LLMs) for automatically resolving software bugs. However, a noticeable gap in existing approaches lies in the oversight of collaborative facets intrinsic to bug resolution, treating the process as a single-stage endeavor. Moreover, most approaches solely take the buggy code snippet as input for LLMs during the patch generation stage. To mitigate the aforementioned limitations, we introduce a novel stage-wise framework named PATCH. Specifically, we first augment the buggy code snippet with corresponding dependence context and intent information to better guide LLMs in generating the correct candidate patches. Additionally, by taking inspiration from bug management practices, we decompose the bug-fixing task into four distinct stages: bug reporting, bug diagnosis, patch generation, and patch verification. These stages are performed interactively by LLMs, aiming to simulate the collaborative behavior of programmers during the resolution of software bugs. By harnessing these collective contributions, PATCH effectively enhances the bug-fixing capability of LLMs. We implement PATCH by employing the powerful dialogue-based LLM ChatGPT. Our evaluation on the widely used bug-fixing benchmark BFP demonstrates that PATCH has achieved better performance than state-of-the-art LLMs.

Figures

Figures reproduced from arXiv: 2501.16149 by the authors.

Figure 1
Figure 1. Limitations of Existing LLM-Based Bug-Fixing Approaches. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. The Brief Structure of PATCH. Novelty 1: Augmenting the buggy code snippet with additional dependence context and guided program￾mer’s intent as input to the LLM. This paper employs the widely-used benchmark BFP [81] for evaluation, which includes an extensive collection of paired bug-fixing instances across a diverse range of real-world bugs, rather than limiting the scope to specific bug types [49]. We commence by… view at source ↗
Figure 3
Figure 3. , PATCH involves three ChatGPT agents (i.e., ChatGPTTester, ChatGPTDeveloper, and ChatGPTReviewer), each assigned to specific stages (i.e., Bug Reporting, Bug Diagnosis, Patch Generation, and Patch Verification) within the bug management process. ②Bug Diagnosis (Section 2.3) ④Patch Verification (Section 2.5) Prompt for Bug Reporting GitHub Repository Code File Buggy Code Commit Message Dependence Context Buggy Metho… view at source ↗
Figures from the paper (15 more)
Figure 4
Figure 4. Figure 4: A Prompting Example of the Tester’s Behavior during the Bug Reporting Stage. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: A Prompting Example of the Developer’s Behavior during the Code Explanation Stage. [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: A Prompting Example of the Developer’s Behavior during the Pattern Summarization Stage. [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: A Prompting Example of the Developer’s Behavior during the Initial Patch Generation Stage. [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]
Figure 8
Figure 8. Figure 8: A Prompting Example of the Reviewer’s Behavior during the Patch Verification Stage. [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: A Prompting Example for the LLM Baselines. [PITH_FULL_IMAGE:figures/full_fig_p015_9.png]
Figure 10
Figure 10. Figure 10: The Overlapping Rates and Unique Patch Numbers of the Evaluated Models. [PITH_FULL_IMAGE:figures/full_fig_p018_10.png]
Figure 11
Figure 11. Figure 11: An Example from the Augmented BFP Benchmark only Fixed by [PITH_FULL_IMAGE:figures/full_fig_p019_11.png]
Figure 12
Figure 12. Figure 12: The Contribution of Different Commit Message Types to Improvements in Bug-Fixing Performance. [PITH_FULL_IMAGE:figures/full_fig_p020_12.png]
Figure 13
Figure 13. Figure 13: The Effect of Interaction Turns on Bug Fixing. [PITH_FULL_IMAGE:figures/full_fig_p021_13.png]
Figure 14
Figure 14. Figure 14: An Example of the Interactive Process between [PITH_FULL_IMAGE:figures/full_fig_p022_14.png]
Figure 15
Figure 15. Figure 15: The Number of Correct Patches Generated by Each Baseline on Bugs.jar under Different Candidate Numbers. [PITH_FULL_IMAGE:figures/full_fig_p024_15.png]
Figure 16
Figure 16. Figure 16: Examples of the Bug-Fixing Venn Diagram, Equivalent Patches, and a Uniquely Fixed Bug by [PITH_FULL_IMAGE:figures/full_fig_p025_16.png]
Figure 17
Figure 17. Figure 17: Bug-fixing Venn Diagram between PATCH and PATCH RepairLLaMA on Defects4J. 5.2 Multi-Hunk Bug-Fixing Scenario [PITH_FULL_IMAGE:figures/full_fig_p026_17.png]
Figure 18
Figure 18. Figure 18: An Example of Adaptation Prompts Used by [PITH_FULL_IMAGE:figures/full_fig_p027_18.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

120 extracted references · 59 canonical work pages

  1. [1]

    Stolee, Yuriy Brun, and Claire Le Goues

    Afsoon Afzal, Manish Motwani, Kathryn T. Stolee, Yuriy Brun, and Claire Le Goues. 2021. SOSRepair: Expressive Semantic Search for Real-World Program Repair. IEEE Trans. Software Eng. 47, 10 (2021), 2162–2181

  2. [2]

    Do, Yan Xu, and Pascale Fung

    Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023. A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity. In Proceedings of the 13th International Joint Conference on Natural Langua...

  3. [3]

    Muneera Bano, Rashina Hoda, Didar Zowghi, and Christoph Treude. 2024. Large Language Models for Qualitative Research in Software Engineering: Exploring Opportunities and Challenges. Autom. Softw. Eng. 31, 1 (2024), 8

  4. [4]

    Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. GPT-NeoX-20B: An Open-Source Autoregressive Language Model. CoRR abs/2204.06745 (2022)

  5. [5]

    Devanbu, and Michael Pradel

    Islem Bouzenia, Premkumar T. Devanbu, and Michael Pradel. 2024. RepairAgent: An Autonomous, LLM-Based Agent for Program Repair. CoRR abs/2403.17134 (2024)

  6. [6]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin...

  7. [7]

    Saikat Chakraborty and Baishakhi Ray. 2021. On Multi-Modal Learning of Editing Source Code. In Proceedings of the 36th IEEE/ACM International Conference on Automated Software Engineering (ASE’21) . IEEE, Melbourne, 443–455

  8. [8]

    Liushan Chen, Yu Pei, and Carlo A. Furia. 2021. Contract-Based Program Repair Without The Contracts: An Extended Study. IEEE Trans. Software Eng. 47, 12 (2021), 2841–2857

Show all 120 references
  1. [9]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pondé de Oliveira Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scot...

  2. [10]

    Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2024. Teaching Large Language Models to Self-Debug. In Proceedings of the 12th International Conference on Learning Representations (ICLR’24) . OpenReview.net, Vienna

  3. [11]

    Zimin Chen, Steve Kommrusch, Michele Tufano, Louis-Noël Pouchet, Denys Poshyvanyk, and Martin Monperrus. 2021. SequenceR: Sequence-to- Sequence Learning for End-to-End Program Repair. IEEE Trans. Software Eng. 47, 9 (2021), 1943–1959

  4. [12]

    Yihong Dong, Xue Jiang, Zhi Jin, and Ge Li. 2024. Self-Collaboration Code Generation via ChatGPT. ACM Trans. Softw. Eng. Methodol. 33, 7 (2024), 189:1–189:38

  5. [13]

    Tore Dybå, Vigdis By Kampenes, and Dag I. K. Sjøberg. 2006. A Systematic Review of Statistical Power in Software Engineering Experiments. Inf. Softw. Technol. 48, 8 (2006), 745–755

  6. [14]

    Çagri Eren, Kerem Sahin, and Eray Tüzün. 2023. Analyzing Bug Life Cycles to Derive Practical Insights. In Proceedings of the 27th International Conference on Evaluation and Assessment in Software Engineering (EASE’23) . ACM, Oulu, 162–171

  7. [15]

    Jean-Rémy Falleri, Floréal Morandat, Xavier Blanc, Matias Martinez, and Martin Monperrus. 2014. Fine-Grained and Accurate Source Code Differencing. In Proceedings of the 29th ACM/IEEE International Conference on Automated Software Engineering (ASE’14) . ACM, Vasteras, 313–324

  8. [16]

    Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou

  9. [17]

    Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, and Mike Lewis. 2023. InCoder: A Generative Model for Code Infilling and Synthesis. In Proceedings of the 11th International Conference on Learning Represen...

  10. [18]

    Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2021. The Pile: An 800GB Dataset of Diverse Text for Language Modeling. CoRR abs/2101.00027 (2021)

  11. [19]

    Ali Ghanbari, Samuel Benton, and Lingming Zhang. 2019. Practical Program Repair via Bytecode Mutation. InProceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA’19). ACM, Beijing, 19–30

  12. [20]

    Claire Le Goues, ThanhVu Nguyen, Stephanie Forrest, and Westley Weimer. 2012. GenProg: A Generic Method for Automatic Software Repair. IEEE Trans. Software Eng. 38, 1 (2012), 54–72

  13. [21]

    Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin. 2022. UniXcoder: Unified Cross-Modal Pre-training for Code Representation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL’22) . Association for Computational Li...

  14. [22]

    Clement, Dawn Drain, Neel Sundaresan, Jian Yin, Daxin Jiang, and Ming Zhou

    Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin B. Clement, Dawn Drain, Neel Sundaresan, Jian Yin, Daxin Jiang, and Ming Zhou. 2021. GraphCodeBERT: Pre-training Code ...

  15. [23]

    Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y. Wu, Y. K. Li, Fuli Luo, Yingfei Xiong, and Wenfeng Liang. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence.CoRR abs/2401.14196 (2024)

  16. [24]

    Zeyu He, Chieh-Yang Huang, Chien-Kuang Cornelia Ding, Shaurya Rohatgi, and Ting-Hao Kenneth Huang. 2024. If in a Crowdsourced Data Annotation Pipeline, a GPT-4. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI’24) . ACM, Honolulu, HI, 1040:1–1040:25

  17. [25]

    Kai Huang, Xiangxin Meng, Jian Zhang, Yang Liu, Wenjie Wang, Shuhao Li, and Yuqing Zhang. 2023. An Empirical Study on Fine-Tuning Large Language Models of Code for Automated Program Repair. In Proceedings of the 38th IEEE/ACM International Conference on Automated Software Engi...

  18. [26]

    Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019. CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. CoRR abs/1909.09436 (2019)

  19. [27]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...

  20. [28]

    Nan Jiang, Kevin Liu, Thibaud Lutellier, and Lin Tan. 2023. Impact of Code Language Models on Automated Program Repair. In Proceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE’23) . ACM, Melbourne, 1430–1442

  21. [29]

    Nan Jiang, Thibaud Lutellier, Yiling Lou, Lin Tan, Dan Goldwasser, and Xiangyu Zhang. 2023. KNOD: Domain Knowledge Distilled Tree Decoder for Automated Program Repair. In Proceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE’23) . IEEE, Melbou...

  22. [30]

    Nan Jiang, Thibaud Lutellier, and Lin Tan. 2021. CURE: Code-Aware Neural Machine Translation for Automatic Program Repair. InProceedings of the 43rd IEEE/ACM International Conference on Software Engineering (ICSE’21) . IEEE, Madrid, 1161–1173

  23. [31]

    René Just, Darioush Jalali, and Michael D. Ernst. 2014. Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs. In Proceedings of the 23rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA’14) . ACM, San Jose, ...

  24. [32]

    Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Yacine Jernite, Margaret Mitchell, Carlos Muñoz Ferrandis, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, and Harm de Vries. 2023. The Stack: 3 TB of Permissively Licensed Source Code. Trans. M...

  25. [33]

    Bissyandé, Dongsun Kim, Jacques Klein, Martin Monperrus, and Yves Le Traon

    Anil Koyuncu, Kui Liu, Tegawendé F. Bissyandé, Dongsun Kim, Jacques Klein, Martin Monperrus, and Yves Le Traon. 2020. FixMiner: Mining Relevant Fix Patterns for Automated Program Repair. Empir. Softw. Eng. 25, 3 (2020), 1980–2024

  26. [34]

    Xuan-Bach Dinh Le, Duc-Hiep Chu, David Lo, Claire Le Goues, and Willem Visser. 2017. S3: Syntax- and Semantic-Guided Repair Synthesis via Programming by Examples. In Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering (ESEC/FSE’17) . ACM, Paderbor...

  27. [35]

    Xuan-Bach Dinh Le, David Lo, and Claire Le Goues. 2016. History Driven Program Repair. In Proceedings of the IEEE 23rd International Conference on Software Analysis, Evolution, and Reengineering (SANER’16) . IEEE Computer Society, Suita, Osaka, 213–224

  28. [36]

    Jae Yong Lee, Sungmin Kang, Juyeon Yoon, and Shin Yoo. 2024. The GitHub Recent Bugs Dataset for Evaluating LLM-Based Debugging Applications. In Proceedings of the 17th IEEE Conference on Software Testing, Verification and Validation (ICST’24) . IEEE, Toronto, ON, 442–444

  29. [37]

    Jia Li, Ge Li, Zhuo Li, Zhi Jin, Xing Hu, Kechi Zhang, and Zhiyi Fu. 2023. CodeEditor: Learning to Edit Source Code with Pre-trained Models. ACM Trans. Softw. Eng. Methodol. 32, 6 (2023), 143:1–143:22

  30. [38]

    Jia Li, Yongmin Li, Ge Li, Xing Hu, Xin Xia, and Zhi Jin. 2021. EditSum: A Retrieve-and-Edit Framework for Source Code Summarization. In Proceedings of the 36th IEEE/ACM International Conference on Automated Software Engineering (ASE’21) . IEEE, Melbourne, 155–166

  31. [39]

    Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, ...

  32. [40]

    Yujia Li, David H. Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, ...

  33. [41]

    Yi Li, Shaohua Wang, and Tien N. Nguyen. 2022. DEAR: A Novel Deep Learning-based Approach for Automated Program Repair. In Proceedings of the 44th IEEE/ACM 44th International Conference on Software Engineering (ICSE’22) . ACM, Pittsburgh, PA, 511–523

  34. [42]

    Derrick Lin, James Koppel, Angela Chen, and Armando Solar-Lezama. 2017. QuixBugs: A Multi-Lingual Program Repair Benchmark Set Based on the Quixey Challenge. In Proceedings Companion of the 2017 ACM SIGPLAN International Conference on Systems, Programming, Languages, and Appli...

  35. [43]

    Yngve Lindsjørn, Dag I. K. Sjøberg, Torgeir Dingsøyr, Gunnar R. Bergersen, and Tore Dybå. 2016. Teamwork Quality and Project Success in Software Development: A Survey of Agile Development Teams. J. Syst. Softw. 122 (2016), 274–286

  36. [44]

    Bissyandé

    Kui Liu, Anil Koyuncu, Dongsun Kim, and Tegawendé F. Bissyandé. 2019. TBar: Revisiting Template-Based Automated Program Repair. In Proceedings of the 28th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA’19) . ACM, Beijing, 31–42

  37. [45]

    Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023. Pre-Train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Comput. Surv. 55, 9 (2023), 195:1–195:35

  38. [46]

    Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and S...

  39. [47]

    Thibaud Lutellier, Hung Viet Pham, Lawrence Pang, Yitong Li, Moshi Wei, and Lin Tan. 2020. CoCoNuT: Combining Context-Aware Neural Translation Models using Ensemble for Program Repair. In Proceedings of the 29th ACM SIGSOFT International Symposium on Software Testing and Manus...

  40. [48]

    Fernanda Madeiral, Simon Urli, Marcelo de Almeida Maia, and Martin Monperrus. 2019. BEARS: An Extensible Java Bug Benchmark for Automatic Program Repair Studies. In Proceedings of the 26th IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER’1...

  41. [49]

    Antonio Mastropaolo, Camilo Escobar-Velásquez, and Mario Linares-Vásquez. 2024. The Rise and Fall(?) of Software Engineering. CoRR abs/2406.10141 (2024)

  42. [50]

    Thomas J. McCabe. 1976. A Complexity Measure. IEEE Trans. Software Eng. 2, 4 (1976), 308–320

  43. [51]

    McChesney and Séamus Gallagher

    Ian R. McChesney and Séamus Gallagher. 2004. Communication and Co-Ordination Practices in Software Engineering Projects. Inf. Softw. Technol. 46, 7 (2004), 473–489

  44. [52]

    Sergey Mechtaev, Jooyong Yi, and Abhik Roychoudhury. 2016. Angelix: Scalable Multiline Program Patch Synthesis via Symbolic Analysis. In Proceedings of the 38th International Conference on Software Engineering (ICSE’16) . ACM, Austin, TX, 691–701

  45. [53]

    Serra-Diaz, Brian J

    Cory Merow, Josep M. Serra-Diaz, Brian J. Enquist, and Adam M. Wilson. 2023. AI Chatbots Can Boost Scientific Coding. Nat. Ecol. Evol. , (2023), 1–3

  46. [54]

    Nguyen, and Hridesh Rajan

    Hoan Anh Nguyen, Anh Tuan Nguyen, Tung Thanh Nguyen, Tien N. Nguyen, and Hridesh Rajan. 2013. A Study of Repetitiveness of Code Changes in Software Evolution. In Proceedings of the 28th IEEE/ACM International Conference on Automated Software Engineering (ASE’13) . IEEE, Silico...

  47. [55]

    Erik Nijkamp, Hiroaki Hayashi, Caiming Xiong, Silvio Savarese, and Yingbo Zhou. 2023. CodeGen2: Lessons for Training LLMs on Programming and Natural Languages. CoRR abs/2305.02309 (2023)

  48. [56]

    Changan Niu, Chuanyi Li, Bin Luo, and Vincent Ng. 2022. Deep Learning Meets Software Engineering: A Survey on Pre-Trained Models of Source Code. In Proceedings of the 31st International Joint Conference on Artificial Intelligence (IJCAI’22) . ijcai.org, Vienna, 5546–5555

  49. [57]

    Yannic Noller, Ridwan Shariffdeen, Xiang Gao, and Abhik Roychoudhury. 2022. Trust Enhancement Issues in Program Repair. In Proceedings of the 44th IEEE/ACM International Conference on Software Engineering (ICSE’22) . ACM, Pittsburgh, PA, 2228–2240

  50. [58]

    Hassan, Naoya Osawa, and Ken-ichi Matsumoto

    Masao Ohira, Ahmed E. Hassan, Naoya Osawa, and Ken-ichi Matsumoto. 2012. The Impact of Bug Management Patterns on Bug Fixing: A Case Study of Eclipse Projects. In Proceedings of the 28th IEEE International Conference on Software Maintenance (ICSM’12) . IEEE, Trento, 264–273

  51. [59]

    OpenAI. 2022. Introducing ChatGPT. Technical Report. . https://openai.com/blog/chatgpt

  52. [60]

    OpenAI. 2023. GPT-4 Technical Report. CoRR abs/2303.08774 (2023)

  53. [61]

    Haidar Osman, Mircea Lungu, and Oscar Nierstrasz. 2014. Mining Frequent Bug-Fix Code Changes. In Proceedings of the 2014 IEEE Conference on Software Maintenance, Reengineering, and Reverse Engineering (CSMR-WCRE’14) . IEEE, Antwerp, 343–347

  54. [62]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leik...

  55. [63]

    Ipek Ozkaya. 2023. Application of Large Language Models to Software Engineering Tasks: Opportunities, Risks, and Implications. IEEE Softw. 40, 3 (2023), 4–8

  56. [64]

    Mohib Hossain, Masum Hasan, and Anindya Iqbal

    Rishov Paul, Md. Mohib Hossain, Masum Hasan, and Anindya Iqbal. 2023. Automated Program Repair Based on Code Review: How do Pre-Trained Transformer Models Perform? CoRR abs/2304.07840 (2023)

  57. [65]

    Renaud Pawlak, Martin Monperrus, Nicolas Petitprez, Carlos Noguera, and Lionel Seinturier. 2016. SPOON: A Library for Implementing Analyses and Transformations of Java Source Code. Softw. Pract. Exp. 46, 9 (2016), 1155–1179

  58. [66]

    Julian Aron Prenner, Hlib Babii, and Romain Robbes. 2022. Can OpenAI’s Codex Fix Bugs?: An Evaluation on QuixBugs. In Proceedings of the 3rd IEEE/ACM International Workshop on Automated Program Repair (APR@ICSE) . IEEE, Pittsburgh, PA, 69–75

  59. [67]

    Rigby, Brendan Cleary, Frédéric Painchaud, Margaret-Anne D

    Peter C. Rigby, Brendan Cleary, Frédéric Painchaud, Margaret-Anne D. Storey, and Daniel M. Germán. 2012. Contemporary Peer Review in Action: Lessons from Open Source Development. IEEE Softw. 29, 6 (2012), 56–61

  60. [68]

    Robertson and Hugo Zaragoza

    Stephen E. Robertson and Hugo Zaragoza. 2009. The Probabilistic Relevance Framework: BM25 and Beyond. Found. Trends Inf. Retr. 3, 4 (2009), 333–389

  61. [69]

    Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton-Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défoss...

  62. [70]

    Saha, Yingjun Lyu, Wing Lam, Hiroaki Yoshida, and Mukul R

    Ripon K. Saha, Yingjun Lyu, Wing Lam, Hiroaki Yoshida, and Mukul R. Prasad. 2018. Bugs.jar: A Large-Scale, Diverse Dataset of Real-World Java Bugs. In Proceedings of the 15th International Conference on Mining Software Repositories (MSR’18) . ACM, Gothenburg, 10–13

  63. [71]

    Saha, and Mukul R

    Seemanta Saha, Ripon K. Saha, and Mukul R. Prasad. 2019. Harnessing Evolution for Multi-Hunk Program Repair. In Proceedings of the 41st International Conference on Software Engineering (ICSE’19) . IEEE / ACM, Montreal, QC, 13–24

  64. [72]

    Jessica Shieh. 2023. Best Practices for Prompt Engineering with OpenAI API . Technical Report. . https://help.openai.com/en/articles/6654000-best- practices-for-prompt-engineering-with-openai-api

  65. [73]

    André Silva, Sen Fang, and Martin Monperrus. 2023. RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair. CoRR abs/2312.15698 (2023). Manuscript submitted to ACM 34 Zhang et al

  66. [74]

    Singh and Micah B

    Ajay S. Singh and Micah B. Masuku. 2014. Sampling Techniques & Determination of Sample Size in Applied Statistics Research: An Overview. International Journal of Economics, Commerce and Management 2, 11 (2014), 1–22

  67. [75]

    Dominik Sobania, Martin Briesch, Carol Hanna, and Justyna Petke. 2023. An Analysis of the Automatic Bug Fixing Performance of ChatGPT. In Proceedings of the 4th IEEE/ACM International Workshop on Automated Program Repair (APR@ICSE) . IEEE, Melbourne, 23–30

  68. [76]

    Diomidis Spinellis. 2000. The Pragmatic Programmer: From Journeyman to Master. IEEE Softw. 17, 6 (2000), 108–110

  69. [77]

    Shin Hwei Tan, Ziqiang Li, and Lu Yan. 2024. CrossFix: Resolution of GitHub Issues via Similar Bugs Recommendation. J. Softw. Evol. Process. 36, 4 (2024)

  70. [78]

    Yu Tang, Long Zhou, Ambrosio Blanco, Shujie Liu, Furu Wei, Ming Zhou, and Muyun Yang. 2021. Grammar-Based Patches Generation for Automated Program Repair. In Findings of the Association for Computational Linguistics: ACL/IJCNLP . Association for Computational Linguistics, Virt...

  71. [79]

    Yingchen Tian, Yuxia Zhang, Klaas-Jan Stol, Lin Jiang, and Hui Liu. 2022. What Makes a Good Commit Message?. In Proceedings of the 44th IEEE/ACM 44th International Conference on Software Engineering (ICSE’22) . ACM, Pittsburgh, PA, 2389–2401

  72. [80]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. LLaMA: Open and Efficient Foundation La...

  73. [81]

    Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk. 2019. An Empirical Study on Learning Bug-Fixing Patches in the Wild via Neural Machine Translation. ACM Trans. Softw. Eng. Methodol. 28, 4 (2019), 19:1–19:29

  74. [82]

    Rosalia Tufano, Ozren Dabic, Antonio Mastropaolo, Matteo Ciniselli, and Gabriele Bavota. 2024. Code Review Automation: Strengths and Weaknesses of the State of the Art. IEEE Trans. Software Eng. 50, 2 (2024), 338–353

  75. [83]

    Shih, Yu Wu, and John M

    Jing Wang, Patrick C. Shih, Yu Wu, and John M. Carroll. 2015. Comparative Case Studies of Open Source Software Peer Review Practices. Inf. Softw. Technol. 67 (2015), 1–12

  76. [84]

    Joty, and Steven C

    Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi. 2021. CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP’21) . ...

  77. [85]

    Bolin Wei, Yongmin Li, Ge Li, Xin Xia, and Zhi Jin. 2020. Retrieve and Refine: Exemplar-based Neural Comment Generation. In Proceedings of the 35th IEEE/ACM International Conference on Automated Software Engineering (ASE’20) . IEEE, Melbourne, 349–360

  78. [86]

    Chi, Quoc V

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Proceedings of the 36th Annual Conference on Neural Information Processing Syst...

  79. [87]

    Yuxiang Wei, Chunqiu Steven Xia, and Lingming Zhang. 2023. Copiloting the Copilots: Fusing Large Language Models with Completion Engines for Automated Program Repair. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations...

  80. [88]

    Eric Wong, Xue-Lin Li, and Phillip A

    W. Eric Wong, Xue-Lin Li, and Phillip A. Laplante. 2017. Be More Familiar with Our Enemies and Pave the Way Forward: A Review of the Roles Bugs Played in Software Failures. J. Syst. Softw. 133 (2017), 68–94

  81. [89]

    Chunqiu Steven Xia, Yuxiang Wei, and Lingming Zhang. 2023. Automated Program Repair in the Era of Large Pre-Trained Language Models. In Proceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE’23) . ACM, Melbourne, 1482–1494

  82. [90]

    Chunqiu Steven Xia and Lingming Zhang. 2022. Less Training, More Repairing Please: Revisiting Automated Program Repair via Zero-Shot Learning. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering...

  83. [91]

    Chunqiu Steven Xia and Lingming Zhang. 2024. Automated Program Repair via Conversation: Fixing 162 out of 337 Bugs for $0.42 Each using ChatGPT. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA’24) . ACM, Vienna, 819–831

  84. [92]

    Shengbin Xu, Yuan Yao, Feng Xu, Tianxiao Gu, Hanghang Tong, and Jian Lu. 2019. Commit Message Generation for Source Code Changes. In Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI’19) . ijcai.org, Macao, 3975–3981

  85. [93]

    Lamelas Marcote, Thomas Durieux, Daniel Le Berre, and Martin Monperrus

    Jifeng Xuan, Matias Martinez, Favio Demarco, Maxime Clement, Sebastian R. Lamelas Marcote, Thomas Durieux, Daniel Le Berre, and Martin Monperrus. 2017. Nopol: Automatic Repair of Conditional Statement Bugs in Java Programs. IEEE Trans. Software Eng. 43, 1 (2017), 34–55

  86. [94]

    Hassan, and Naoyasu Ubayashi

    Kazuhiro Yamashita, Changyun Huang, Meiyappan Nagappan, Yasutaka Kamei, Audris Mockus, Ahmed E. Hassan, and Naoyasu Ubayashi. 2016. Thresholds for Size and Complexity Metrics: A Case Study from the Perspective of Defect Density. In Proceedings of the 2016 IEEE International Co...

  87. [95]

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In Proceedings of the 37th Annual Conference on Neural Information Processing Systems (NeurIPS’23) ...

  88. [96]

    He Ye, Matias Martinez, Xiapu Luo, Tao Zhang, and Martin Monperrus. 2022. SelfAPR: Self-supervised Program Repair with Test Execution Diagnostics. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering (ASE’22) . ACM, Rochester, MI, 92:1–92:13

  89. [97]

    He Ye, Matias Martinez, and Martin Monperrus. 2021. Automated Patch Assessment for Program Repair at Scale. Empir. Softw. Eng. 26, 2 (2021), 20. Manuscript submitted to ACM Empowering Large Language Model with PATCHfor Automatic Bug Fixing 35

  90. [98]

    He Ye, Matias Martinez, and Martin Monperrus. 2022. Neural Program Repair with Execution-based Backpropagation. In Proceedings of the 44th IEEE/ACM 44th International Conference on Software Engineering (ICSE’22) . ACM, Pittsburgh, PA, 1506–1518

  91. [99]

    He Ye and Martin Monperrus. 2024. ITER: Iterative Neural Repair for Multi-Location Patches. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering (ICSE’24) . ACM, Lisbon, 10:1–10:13

  92. [100]

    Xin Yin, Chao Ni, Shaohua Wang, Zhenhao Li, Limin Zeng, and Xiaohu Yang. 2024. ThinkRepair: Self-Directed Automated Program Repair. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA’24) . ACM, Vienna, 1274–1286

  93. [101]

    Ding Yuan, Yu Luo, Xin Zhuang, Guilherme Renna Rodrigues, Xu Zhao, Yongle Zhang, Pranay Jain, and Michael Stumm. 2014. Simple Testing Can Prevent Most Critical Failures: An Analysis of Production Failures in Distributed Data-Intensive Systems. In Proceedings of the 11th USENIX...

  94. [102]

    Zhengran Zeng, Hanzhuo Tan, Haotian Zhang, Jing Li, Yuqun Zhang, and Lingming Zhang. 2022. An Extensive Study on Pre-Trained Models for Program Understanding and Generation. In Proceedings of the 31st ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA’...

  95. [103]

    Huangzhao Zhang, Kechi Zhang, Zhuo Li, Jia Li, Jia Allen Li, Yongmin Li, Yunfei Zhao, Yuqi Zhu, Fang Liu, Ge Li, and Zhi Jin. 2024. Deep Learning for Code Generation: A Survey. Sci. China Inf. Sci. [online first] (2024). https://doi.org/10.1007/s11432-023-3956-3

  96. [104]

    Quanjun Zhang, Chunrong Fang, Yuxiang Ma, Weisong Sun, and Zhenyu Chen. 2024. A Survey of Learning-Based Automated Program Repair. ACM Trans. Softw. Eng. Methodol. 33, 2 (2024), 55:1–55:69

  97. [105]

    Tao Zhang, He Jiang, Xiapu Luo, and Alvin T. S. Chan. 2016. A Literature Review of Research in Bug Resolution: Tasks, Challenges and Future Directions. Comput. J. 59, 5 (2016), 741–773

  98. [106]

    Yuwei Zhang. 2024. Replicate Package of PATCH. Zenodo. https://doi.org/10.5281/zenodo.14257480

  99. [107]

    Yuwei Zhang, Ge Li, Zhi Jin, and Ying Xing. 2023. Neural Program Repair with Program Dependence Analysis and Effective Filter Mechanism. CoRR abs/2305.09315 (2023)

  100. [108]

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen

  101. [109]

    Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate Before Use: Improving Few-shot Performance of Language Models. In Proceedings of the 38th International Conference on Machine Learning (ICML’21) . PMLR, Virtual Event, 12697–12706

  102. [110]

    Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Lei Shen, Zihan Wang, Andi Wang, Yang Li, Teng Su, Zhilin Yang, and Jie Tang. 2023. CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X. In Proceedings of the 29th AC...

  103. [111]

    Wenkang Zhong, Hongliang Ge, Hongfei Ai, Chuanyi Li, Kui Liu, Jidong Ge, and Bin Luo. 2022. StandUp4NPR: Standardizing SetUp for Empirically Comparing Neural Program Repair Systems. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering ...

  104. [112]

    Wenkang Zhong, Chuanyi Li, Jidong Ge, and Bin Luo. 2022. Neural Program Repair: Systems, Challenges and Solutions. In Proceedings of the 13th Asia-Pacific Symposium on Internetware (Internetware’22) . ACM, Hohhot, 96–106

  105. [113]

    Bissyandé, and Vincent Ng

    Wenkang Zhong, Chuanyi Li, Kui Liu, Jidong Ge, Bin Luo, Tegawendé F. Bissyandé, and Vincent Ng. 2024. Benchmarking and Categorizing the Performance of Neural Program Repair Systems for Java. ACM Trans. Softw. Eng. Methodol. [online first] (2024). https://doi.org/10.1145/3688834

  106. [114]

    Qihao Zhu, Zeyu Sun, Yuan-an Xiao, Wenjie Zhang, Kang Yuan, Yingfei Xiong, and Lu Zhang. 2021. A Syntax-Guided Edit Decoder for Neural Program Repair. In Proceedings of the 29th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Eng...

  107. [115]

    Qihao Zhu, Zeyu Sun, Wenjie Zhang, Yingfei Xiong, and Lu Zhang. 2023. Tare: Type-Aware Neural Program Repair. InProceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE’23) . ACM, Melbourne, 1443–1455

  108. [116]

    Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B

    Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul F. Christiano, and Geoffrey Irving. 2019. Fine-Tuning Language Models from Human Preferences. CoRR abs/1909.08593 (2019)

  109. [117]

    Thomas Zimmermann, Rahul Premraj, Nicolas Bettenburg, Sascha Just, Adrian Schröter, and Cathrin Weiss. 2010. What Makes a Good Bug Report? IEEE Trans. Software Eng. 36, 5 (2010), 618–643

  110. [118]

    Weiqin Zou, David Lo, Zhenyu Chen, Xin Xia, Yang Feng, and Baowen Xu. 2020. How Practitioners Perceive Automated Bug Report Management Techniques. IEEE Trans. Software Eng. 46, 8 (2020), 836–862. Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009 Manuscript...

  111. [2020]

    In Findings of the Association for Computational Linguistics: EMNLP

    CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Findings of the Association for Computational Linguistics: EMNLP. Association for Computational Linguistics, Virtual Event, 1536–1547

  112. [2023]

    CoRR abs/2303.18223 (2023)

    A Survey of Large Language Models. CoRR abs/2303.18223 (2023)

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.