REVIEW 2 cited by
Coffee: Boost Your Code LLMs by Fixing Bugs with Feedback
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Code editing is an essential step towards reliable program synthesis to automatically correct critical errors generated from code LLMs. Recent studies have demonstrated that closed-source LLMs (i.e., ChatGPT and GPT-4) are capable of generating corrective feedback to edit erroneous inputs. However, it remains challenging for open-source code LLMs to generate feedback for code editing, since these models tend to adhere to the superficial formats of feedback and provide feedback with misleading information. Hence, the focus of our work is to leverage open-source code LLMs to generate helpful feedback with correct guidance for code editing. To this end, we present Coffee, a collected dataset specifically designed for code fixing with feedback. Using this dataset, we construct CoffeePots, a framework for COde Fixing with FEEdback via Preference-Optimized Tuning and Selection. The proposed framework aims to automatically generate helpful feedback for code editing while minimizing the potential risk of superficial feedback. The combination of Coffee and CoffeePots marks a significant advancement, achieving state-of-the-art performance on HumanEvalFix benchmark. Codes and model checkpoints are publicly available at https://github.com/Lune-Blue/COFFEE.
Forward citations
Cited by 2 Pith papers
-
Learning to Generate Unit Tests for Automated Debugging
UTGen trains LLMs to generate error-revealing unit tests with correct expected outputs, and UTDebug uses those tests with test-time scaling and backtracking to improve automated debugging and code selection.
-
Exploring the Potential of Llama Models in Automated Code Refinement: A Replication Study
A replication study finds that a 7-billion-parameter open-source CodeLlama model, tuned with low temperature and specific prompts, can match ChatGPT on a code-refinement similarity metric while keeping code on local hardware.
Discussion (0). Continue with ORCID to comment.