Pith. sign in

REVIEW 4 major objections 5 minor 81 references

An Empirical Study of Retrieval-Augmented Code Generation: Challenges and Opportunities

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Retrieval-augmented code generation works best with a simple no-training retriever and simple concatenation, an empirical study of three models and three datasets finds.

desk verdict Useful systematic benchmark of retrieval-augmented code generation, but the headline numbers need variance estimates and a few tables need reconciling before the quantitative recommendations can be trusted. read the letter →

arxiv 2501.13742 v1 pith:CTQJ2QBU submitted 2025-01-23 cs.SE

classification cs.SE
keywords retrieval-augmentedgenerationcodepre-trainedmodelsBM25fusionstrategiessearchempiricalstudyBLEU
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is an empirical attempt to turn retrieval-augmented generation into a dependable recipe for code generation: retrieve code snippets similar to a natural-language request, splice them into the model input, and fine-tune the model on the augmented data. It claims this recipe improves three architecturally different pre-trained code models (CodeGen, UniXcoder, CodeT5) across three datasets (CONCODE, CoNaLa, HearthStone), with average Exact Match up 41.60% on HearthStone. It also claims BM25 and Sequential Integration Fusion are the convenient high-performing defaults, while Sketch Filling Fusion adds further gains at substantially higher training cost. If right, the study gives practitioners a model-agnostic way to improve code generation without changing the model architecture, and a cost-aware map of when to use more elaborate retrieval or fusion.

What carries the argument

The machinery is a three-phase retrieval-augmented pipeline. In the retrieval phase, a retriever (BM25, RetroMAE, or a code search model such as CodeBERT, UniXcoder, or CoCoSoDa) finds k similar code snippets from a database of natural-language/code pairs; in the fusion phase, those snippets are combined with the original request by concatenation (Sequential Integration Fusion), by expanding the training sample (Sample Expansion Fusion), by encoding each snippet and feeding the vectors to the decoder (Vectorized Decoding Fusion), or by extracting a structural sketch (Sketch Filling Fusion); in the generation phase, the unchanged pre-trained model is fine-tuned on the augmented input. The load-bearing piece is that the retrieved snippets act as an external reference that supplies structural and semantic guidance, so the model does not have to infer the target shape from the description alone.

What would settle it

Run the recommended pipeline (BM25 retrieval plus Sequential Integration Fusion on CodeT5) on a held-out benchmark whose test problems require APIs or coding patterns absent from the retrieval database; the paper's universal-effectiveness claim would be falsified if the augmented model does not beat the base model on BLEU or CodeBLEU.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that the retrieval-augmented framework generalizes: all three pre-trained models improve on BLEU and CodeBLEU on all three datasets when BM25-retrieved code is concatenated into the input, and the retriever requires no training. The strongest evidence is HearthStone, where the average EM improvement across models is 41.60%, with BLEU and CodeBLEU up about 9% and 8.7%. The paper further finds that more sophisticated retrievers are not better: BM25 outperforms the code search models on CONCODE and HearthStone and is best or near-best for CodeT5 on CoNaLa, while RetroMAE can degrade CodeGen and UniXcoder on CONCODE and HearthStone. Among fusion strategies, Sequential Integration Fusion and Sketch Filling Fusion lead, with Sketch Filling giving the largest CodeT5 gains (average 14.83% BLEU, 8.05% CodeBLEU) but requiring 2 to 7 times the training time; Sequential Integration is recommended as the cost-effective choice. The same retrieval augmentation also improves three large language models used in a prompting setup.

Load-bearing premise

The load-bearing premise is that the retrieval database, here the training set, contains a snippet similar enough to every test request to add signal rather than noise; if a new domain's test data has no near neighbors, the recommended pipeline has nothing useful to retrieve and the claimed universal improvement can collapse.

Editorial extensions

If this is right

  • Fine-tuned code models can be improved without changing their architecture or parameters, simply by augmenting the training input with BM25-retrieved examples.
  • A no-training retriever can beat deep code-search models for generation, so teams without retrieval-label budgets need not adopt heavier retrievers.
  • The best number of retrieved snippets is dataset-dependent; short or structured datasets saturate or degrade, and long snippets get truncated, so the choice of k should follow from input/output length rather than being maximized.
  • For code with regular structure, extracting a sketch from the best-matching snippet yields the largest gains, making sketch-based fusion a candidate when training cost is acceptable.
  • Large language models used in prompting mode also benefit from retrieved code snippets, extending the finding beyond fine-tuned models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's recommendation of BM25 likely does not transfer to very large or rapidly changing codebases, because its retrieval cost grows with corpus size; the paper's own cost tables hint at this trade-off, and a learned retriever may become preferable at scale.
  • If the mechanism is truly that retrieved code carries structural guidance, a direct testable extension is to degrade the retrieval database with random or wrong-language snippets and predict that gains disappear or reverse; the paper's RetroMAE result already points in that direction.
  • Active retrieval, deciding per request whether retrieval will help, could salvage the cases where noisy snippets hurt; the paper names this as future work, so treating it as an extension is consistent with the study's findings.
  • The main constraint is the assumption that the training set doubles as the retrieval database; for benchmarks whose test data excludes the training set, constructing an external retrieval database becomes the next bottleneck.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents a systematic empirical study of retrieval-augmented code generation. It evaluates three pre-trained code models (CodeGen, UniXcoder, CodeT5) and three LLMs (ChatGLM3-6B, CodeLlama-7B, DeepSeek-Coder) on CONCODE, CoNaLa, and HearthStone, comparing five retrieval techniques (BM25, RetroMAE, CodeBERT, UniXcoder, CoCoSoDa) and four fusion strategies (Sequential Integration Fusion, Sample Expansion Fusion, Vectorized Decoding Fusion, Sketch Filling Fusion). The main claims are that the retrieval-augmented framework reliably improves code generation across models and datasets, that BM25 and Sequential Integration Fusion are convenient and effective default choices, and that Sketch Filling Fusion can yield further gains at higher computational cost. The paper also reports cost measurements and releases code and augmented datasets.

Significance. If the results hold, this is a useful reference for practitioners: it provides a rare head-to-head comparison of retrieval techniques and fusion strategies for code generation, it covers multiple model architectures, and it quantifies the training and inference cost trade-offs. The BLEU and CodeBLEU improvements are directionally consistent across most conditions, which gives the central claim partial support, and the release of code and augmented datasets is a concrete reproducibility asset. However, the absence of variance estimates, the small HearthStone EM counts, and internal inconsistencies between tables mean that the specific numerical claims should be treated with caution until they are reconciled by the authors.

major comments (4)
  1. [Section 5.2, Table 3] The headline HearthStone result, a 41.60% average EM improvement, is computed from very small absolute counts: 7→10, 9→15, and 13→15 correct predictions out of 66 test examples for CodeGen, UniXcoder, and CodeT5, respectively. With no repeated runs and no per-condition variance estimates, differences of 2–6 correct examples are within plausible seed-to-seed noise, so the claim that retrieval-augmentation universally improves EM on HearthStone is not established. Please report per-dataset bootstrap confidence intervals or repeated-seed results, and state explicitly the unit of analysis and pairing for the t-test reported as p=0.035.
  2. [Tables 3 and 5] Tables 3 and 5 describe the same configuration, CodeT5 with BM25 retrieval and Sequential Integration Fusion, but report inconsistent CodeBLEU values: 43.48 vs 46.92 on CONCODE and 49.80 vs 60.50 on HearthStone. This is a load-bearing inconsistency because the central recommendation to use BM25 and SIF depends on these numbers. Please explain the discrepancy (e.g., different k, different evaluation version, different preprocessing) and harmonize the tables so that readers can determine which configuration the reported improvements refer to.
  3. [Abstract and Section 1, RQ3] The abstract and introduction claim that Sketch Filling Fusion yields an average improvement of 14.83% in BLEU and 8.05% in CodeBLEU across the three datasets for original CodeT5. These numbers do not reproduce from Table 5: comparing SFF to the baseline row gives roughly 39% average BLEU improvement and 25% average CodeBLEU improvement, while comparing SFF to SIF or SEF gives different values again. Please correct the abstract and Section 1 or state explicitly the comparison basis for these percentages.
  4. [Section 5.4.1, Figure 2] The recommendation on the optimal number of retrieved code snippets is based on a single run per k value for one model, but the underlying numerical values are not reported in the text or a table, and Figure 2 plots five metrics with two axes and a legend that is difficult to read. Since this recommendation is one of the paper's actionable findings, please provide the exact numbers or a companion table so that readers can assess the magnitude of the differences and the location of the claimed inflection points.
minor comments (5)
  1. [Throughout] There are several typos and copyediting issues, including 'Specically' in Section 2.2, 'donotes' in Section 3.4, 'Expainsion' in Section 5.4.2, and a missing parenthesis in Figure 4(g); a careful proofread is needed.
  2. [Section 5.2] The statistical significance test is described only as 't-test' with p=0.035; please report whether it is paired, what the sample units are, whether all metrics are pooled, and whether any multiple-comparison correction was applied.
  3. [Table 7] The units in Table 7 are not fully clear: the header says 'retrieval costs per 50 instances,' but the 'total costs' column appears to scale the per-50-instance cost to the full test set; please clarify the exact computation and the number of instances used for the total.
  4. [Section 6.1.1 vs Section 6.4] Section 6.1.1 reports substantial LLM improvements, while Section 6.4 states that generalization to larger models is uncertain; please make the relationship between these two statements explicit, since readers may otherwise see them as contradictory.
  5. [Section 4.3 and Table 1] The SimAST metric is defined in Equation (18) but the table header uses 'SimilarityAST'; please unify the notation throughout the paper.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the central claims are direct measurements on held-out test splits, with no fitted constants or self-citation chain forcing the results.

full rationale

This paper is an empirical benchmark study rather than a derivation. The central claims—that the retrieval-augmented framework improves three pre-trained models, that BM25 is an effective retrieval technique, and that Sequential Integration Fusion balances cost and performance—are supported by direct measurements reported in Tables 3, 4, 5, and 8 on held-out test splits against fixed baselines. No predicted quantity is computed from an equation that contains the conclusion as an input, and no fitted parameter is renamed as a prediction. The retrieval database is the training set, and test examples are not used for retrieval or fine-tuning, so the BLEU, EM, and CodeBLEU gains are not forced by construction. The only notable self-citation is SKCODER [41], whose co-author X. Hu is also an author of this paper; it is used to specify the sketch-extraction component of Sketch Filling Fusion, but SFF's measured improvements come from the paper's own experiments, making the citation methodological rather than load-bearing. Section 6.3.2 concedes that using the training set as the retrieval database limits applicability when test distributions exclude training data; this is a generalization threat, not a circularity. No uniqueness theorem or ansatz is imported from prior work to forbid alternatives. The empirical conclusions therefore stand independently of the paper's own prior results.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The paper is an empirical benchmark and introduces no new theoretical entities or fitted model parameters. Its conclusions rest on the representativeness of datasets and metrics, the adequacy of the training set as a retrieval database, and the validity of the significance testing.

free parameters (1)
  • number of retrieved code snippets k = 5
    Chosen from {1, 3, 5, 7, 10} in Section 5.4.1 based on observed inflection points on CONCODE and CoNaLa and truncation behavior on HearthStone, then used for all fusion strategy comparisons.
assumptions (5)
  • domain assumption The selected three datasets and five metrics are representative enough that observed gains generalize to code generation as a whole.
    The conclusions are generalized from CONCODE, CoNaLa, and HearthStone using EM, BLEU, ED, SimAST, and CodeBLEU; the paper acknowledges this generalization limit in Section 6.4.
  • domain assumption The training split of each dataset is a sufficient retrieval database, meaning similar code for test inputs exists and is retrievable.
    All retrieval experiments use the training set as the codebase; Section 6.3.2 says a comprehensive database is essential when test data exclude training sets.
  • domain assumption Fine-tuning hyperparameters and checkpoints from the original papers are appropriate for fair comparison across models and conditions.
    Section 4.4 states hyper-parameter settings are the same as the original corresponding papers; deviations could change relative gains.
  • domain assumption The reported statistical test is meaningful despite multiple comparisons and the absence of repeated runs.
    Section 5.2 reports a single t-test p-value of 0.035 across aggregated results without variance, random seeds, or multiple-comparison correction.
  • standard math BM25 scoring and CodeBLEU weighting are valid measures accepted from prior literature.
    Equations 1 through 3 and Equation 19 are taken from prior work and used as-is in Sections 3.2.1 and 4.3.

how reviews work

0 comments
Cite this review

Pith. "Pith review of An Empirical Study of Retrieval-Augmented Code Generation: Challenges and Opportunities." pith.science (2026). https://pith.science/paper/CTQJ2QBU

@misc{pith2026250113742,
  author       = {Pith},
  title        = {Pith review of: An Empirical Study of Retrieval-Augmented Code Generation: Challenges and Opportunities},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CTQJ2QBU}},
  note         = {Machine review of arXiv:2501.13742}
}
read the original abstract

Code generation aims to automatically generate code snippets of specific programming language according to natural language descriptions. The continuous advancements in deep learning, particularly pre-trained models, have empowered the code generation task to achieve remarkable performance. One main challenge of pre-trained models for code generation is the semantic gap between natural language requirements and source code. To address the issue, prior studies typically adopt a retrieval-augmented framework for the task, where the similar code snippets collected by a retrieval process can be leveraged to help understand the requirements and provide guidance for the generation process. However, there is a lack of systematic study on the application of this framework for code generation, including the impact of the final generated results and the specific usage of the framework. In this paper, we choose three popular pre-trained code models, namely CodeGen, UniXcoder, and CodeT5, to assess the impact of the quality and utilization of retrieved code on the retrieval-augmented framework. Our analysis shows that the retrieval-augmented framework is beneficial for improving the performance of the existing pre-trained models. We also provide suggestions on the utilization of the retrieval-augmented code generation framework: BM25 and Sequential Integration Fusion are recommended due to their convenience and superior performance. Sketch Filling Fusion, which extracts a sketch of relevant code, could help the model improve its performance further. Additionally, we conduct experiments to investigate the influence of the retrieval-augmented framework on large language models for code generation, showing the effectiveness of the framework, and we discuss the trade-off between performance improvement and computational costs in each phase within the framework.

Figures

Figures reproduced from arXiv: 2501.13742 by the authors.

Figure 1
Figure 1. Overview of retrieval-augmented framework for code generation. [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. The impact of the number of retrieved code snippets using Sequential Integration Fusion on the [PITH_FULL_IMAGE:figures/full_fig_p017_2.png] view at source ↗
Figure 3
Figure 3. Case study on CONCODE with retrieval-augmented framework, where the retrieval technique is [PITH_FULL_IMAGE:figures/full_fig_p022_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: The most similar code snippet retrieved by different retrieval techniques on CoNaLa. [PITH_FULL_IMAGE:figures/full_fig_p022_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

81 extracted references · 51 canonical work pages

  1. [1]

    Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Unified Pre-training for Program Understanding and Generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021 , Kristina Toutanova, Anna Ru...

  2. [2]

    Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net

  3. [3]

    Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J

    Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. 2021. Program Synthesis with Large Language Models. CoRR abs/2108.07732 (2021)

  4. [4]

    Ramakrishna Bairi, Atharv Sonwane, Aditya Kanade, Arun Iyer, Suresh Parthasarathy, Sriram Rajamani, B Ashok, Shashank Shet, et al. 2023. CodePlan: Repository-level Coding using LLMs and Planning.arXiv preprint arXiv:2309.12499 (2023)

  5. [5]

    Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q Feldman, et al . 2022. A Scalable and Extensible Approach to Benchmarking NL2Code for 18 Programming Languages. arXiv preprint arXiv:2208.08227 (2022)

  6. [6]

    ChatGPT. 2022. ChatGPT. https://openai.com/blog/chatgpt

  7. [7]

    Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen. 2023. CodeT: Code Generation with Generated Tests. In The Eleventh International Conference on Learning Representations

  8. [8]

    Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading Wikipedia to Answer Open-Domain Questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, 3https://drive.google.com/drive/folders/1G_ssf9gCX38Yb7FAjsIlRxiYPO44omvP?usp=drive_link 4https://github.com/watreyoung/RACG J. ACM, Vol. 37...

Show all 81 references
  1. [9]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)

  2. [10]

    Qinyuan Cheng, Xiaonan Li, Shimin Li, Qin Zhu, Zhangyue Yin, Yunfan Shao, Linyang Li, Tianxiang Sun, Hang Yan, and Xipeng Qiu. 2024. Unified Active Retrieval for Retrieval Augmented Generation. CoRR abs/2406.12534 (2024)

  3. [11]

    Fenia Christopoulou, Gerasimos Lampouras, Milan Gritta, Guchun Zhang, Yinpeng Guo, Zhongqi Li, Qi Zhang, Meng Xiao, Bo Shen, Lin Li, et al. 2022. PanGu-Coder: Program Synthesis with Function-Level Language Modeling. arXiv preprint arXiv:2207.11280 (2022)

  4. [12]

    codegeex. 2022. codegeex. https://models.aminer.cn/codegeex/blog/

  5. [13]

    Dahl, Madeleine Bates, Michael Brown, William M

    Deborah A. Dahl, Madeleine Bates, Michael Brown, William M. Fisher, Kate Hunicke-Smith, David S. Pallett, Christine Pao, Alexander I. Rudnicky, and Elizabeth Shriberg. 1994. Expanding the Scope of the ATIS Task: The ATIS-3 Corpus. In Human Language Technology, Proceedings of a...

  6. [14]

    Dawn Drain, Changran Hu, Chen Wu, Mikhail Breslav, and Neel Sundaresan. 2021. Generating Code with the Help of Retrieved Template Functions and Stack Overflow Answers. CoRR abs/2104.05310 (2021)

  7. [15]

    Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. GLM: General Language Model Pretraining with Autoregressive Blank Infilling. In ACL, Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Lin...

  8. [16]

    Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Proceedings of the 2020 Conference on Empirical Methods in Natura...

  9. [17]

    Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, and Mike Lewis. 2023. InCoder: A Generative Model for Code Infilling and Synthesis. In The Eleventh International Conference on Learning Representations

  10. [18]

    Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020. The Pile: An 800GB Dataset of Diverse Text for Language Modeling. arXiv preprint arXiv:2101.00027 (2020)

  11. [19]

    Shuzheng Gao, Xin-Cheng Wen, Cuiyun Gao, Wenxuan Wang, Hongyu Zhang, and Michael R. Lyu. 2023. What Makes Good In-Context Demonstrations for Code Intelligence Tasks with LLMs?. In 38th IEEE/ACM International Conference on Automated Software Engineering, ASE 2023, Luxembourg, S...

  12. [20]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Qianyu Guo, Meng Wang, and Haofen Wang. 2023. Retrieval-Augmented Generation for Large Language Models: A Survey.CoRR abs/2312.10997 (2023)

  13. [21]

    Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin. 2022. UniXcoder: Unified Cross-Modal Pre-training for Code Representation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, ...

  14. [22]

    Clement, Dawn Drain, Neel Sundaresan, Jian Yin, Daxin Jiang, and Ming Zhou

    Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin B. Clement, Dawn Drain, Neel Sundaresan, Jian Yin, Daxin Jiang, and Ming Zhou. 2021. GraphCodeBERT: Pre-training Code ...

  15. [23]

    Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y. Wu, Y. K. Li, Fuli Luo, Yingfei Xiong, and Wenfeng Liang. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence. CoRR abs/2401.14196 (2024)

  16. [24]

    Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. Retrieval Augmented Language Model Pre-Training. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Learning R...

  17. [25]

    Hashimoto, Kelvin Guu, Yonatan Oren, and Percy Liang

    Tatsunori B. Hashimoto, Kelvin Guu, Yonatan Oren, and Percy Liang. 2018. A Retrieve-and-Edit Framework for Predicting Structured Outputs. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, Dec...

  18. [26]

    Shirley Anugrah Hayati, Raphaël Olivier, Pravalika Avvaru, Pengcheng Yin, Anthony Tomasic, and Graham Neubig

  19. [27]

    Girshick

    Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross B. Girshick. 2020. Momentum Contrast for Unsupervised Visual Representation Learning. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, W A, USA, June 13-19, 2020. Computer Vision ...

  20. [28]

    Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt. 2021. Measuring Coding Challenge Competence With APPS. In Proceedings of the Neural Information Processing Systems Tra...

  21. [29]

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2023. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. CoRR abs/2311.052...

  22. [30]

    Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019. CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. CoRR abs/1909.09436 (2019)

  23. [31]

    Srinivasan Iyer, Alvin Cheung, and Luke Zettlemoyer. 2019. Learning Programmatic Idioms for Scalable Semantic Parsing. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Proce...

  24. [32]

    Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. 2018. Mapping Language to Code in Program- matic Context. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018 , Ellen R...

  25. [33]

    Gautier Izacard and Edouard Grave. 2021. Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, EACL 2021, Online, Apri...

  26. [34]

    Xue Jiang, Yihong Dong, Lecheng Wang, Qiwei Shang, and Ge Li. 2023. Self-planning code generation with large language model. arXiv preprint arXiv:2303.06689 (2023)

  27. [35]

    Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig

    Zhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Active Retrieval Augmented Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Sing...

  28. [36]

    Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020. Generalization through Memorization: Nearest Neighbor Language Models. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenReview.net

  29. [37]

    Fabbri, Caiming Xiong, Shafiq Joty, and Chien- Sheng Wu

    Philippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander R. Fabbri, Caiming Xiong, Shafiq Joty, and Chien- Sheng Wu. 2023. SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of Summarization. In Proceedings of the 2023 Conference on Empirical Method...

  30. [38]

    Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven Hoi. 2022. CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning. InAdvances in Neural Information Processing Systems, Alice H. Oh, Alekh Agarwal, Danielle Belgrave, a...

  31. [39]

    Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advanc...

  32. [40]

    Jia Li, Ge Li, Yongmin Li, and Zhi Jin. 2023. Enabling Programming Thinking in Large Language Models Toward Code Generation. arXiv preprint arXiv:2305.06599 (2023)

  33. [41]

    Jia Li, Yongmin Li, Ge Li, Zhi Jin, Yiyang Hao, and Xing Hu. 2023. Skcoder: A sketch-based approach for automatic code generation. arXiv preprint arXiv:2302.06144 (2023)

  34. [42]

    Jia Li, Yunfei Zhao, Yongmin Li, Ge Li, and Zhi Jin. 2023. Towards Enhancing In-Context Learning for Code Generation. arXiv preprint arXiv:2303.17780 (2023)

  35. [43]

    Jia Li, Yunfei Zhao, Yongmin Li, Ge Li, and Zhi Jin. 2024. AceCoder: An Effective Prompting Technique Specialized in Code Generation. ACM Transactions on Software Engineering and Methodology (2024). J. ACM, Vol. 37, No. 4, Article 1. Publication date: August 2018. An Empirical...

  36. [44]

    Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al. 2022. Competition-level code generation with alphacode. Science 378, 6624 (2022), 1092–1097

  37. [45]

    Wang Ling, Phil Blunsom, Edward Grefenstette, Karl Moritz Hermann, Tomás Kociský, Fumin Wang, and Andrew W. Senior. 2016. Latent Predictor Networks for Code Generation. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016, August...

  38. [46]

    Chao Liu, Xin Xia, David Lo, Cuiyun Gao, Xiaohu Yang, and John C. Grundy. 2022. Opportunities and Challenges in Code Search Tools. ACM Comput. Surv. 54, 9 (2022), 196:1–196:40

  39. [47]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR abs/1907.11692 (2019)

  40. [48]

    Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and S...

  41. [49]

    Jianmo Ni, Chenguang Zhu, Weizhu Chen, and Julian J. McAuley. 2019. Learning to Attend On Essential Terms: An Enhanced Retriever-Reader Model for Open-domain Question Answering. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computat...

  42. [50]

    Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023. CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis. In The Eleventh International Conference on Learning Representations

  43. [51]

    Changan Niu, Chuanyi Li, Vincent Ng, Dongxiao Chen, Jidong Ge, and Bin Luo. 2023. An Empirical Comparison of Pre-Trained Models of Source Code. In 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023 . IEEE, 2136–2148

  44. [52]

    OpenAI. 2023. GPT-4 Technical Report. arXiv:2303.08774 [cs.CL]

  45. [53]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6-12, 2002, Philadelphia, PA, USA . ACL, 311–318

  46. [54]

    Rizwan Parvez, Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang

    Md. Rizwan Parvez, Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Retrieval Augmented Code Generation and Summarization. In Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-...

  47. [55]

    Maxim Rabinovich, Mitchell Stern, and Dan Klein. 2017. Abstract Syntax Networks for Code Generation and Semantic Parsing. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Van- couver, Canada, July 30 - August 4, Volume 1: Lo...

  48. [56]

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9

  49. [57]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. J. Mach. Learn. Res. 21 (2020), 140:1–140:67

  50. [58]

    Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. 2020. Codebleu: a method for automatic evaluation of code synthesis. arXiv preprint arXiv:2009.10297 (2020)

  51. [59]

    Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al . 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950 (2023)

  52. [60]

    Ensheng Shi, Yanlin Wang, Wenchao Gu, Lun Du, Hongyu Zhang, Shi Han, Dongmei Zhang, and Hongbin Sun. 2022. CoCoSoDa: Effective Contrastive Learning for Code Search. arXiv preprint arXiv:2204.03293 (2022)

  53. [61]

    Zeyu Sun, Qihao Zhu, Yingfei Xiong, Yican Sun, Lili Mou, and Lu Zhang. 2020. TreeGen: A Tree-Based Transformer Architecture for Code Generation. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial ...

  54. [62]

    Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan. 2020. IntelliCode compose: code generation using transformer. In ESEC/FSE ’20: 28th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, Virtual Event, ...

  55. [63]

    Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models. CoRR abs/2104.08663 (2021)

  56. [64]

    Andrew Trotman, Antti Puurula, and Blake Burgess. 2014. Improvements to BM25 and Language Models Examined. In Proceedings of the 2014 Australasian Document Computing Symposium, ADCS 2014, Melbourne, VIC, Australia, November 27-28, 2014, J. Shane Culpepper, Laurence Anthony F. ...

  57. [65]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)

  58. [66]

    Xin Wang, Yasheng Wang, Yao Wan, Fei Mi, Yitong Li, Pingyi Zhou, Jin Liu, Hao Wu, Xin Jiang, and Qun Liu

  59. [67]

    Yue Wang, Hung Le, Akhilesh Deepak Gotmare, Nghi D. Q. Bui, Junnan Li, and Steven C. H. Hoi. 2023. CodeT5+: Open Code Large Language Models for Code Understanding and Generation. CoRR abs/2305.07922 (2023)

  60. [68]

    Joty, and Steven C

    Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi. 2021. CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. In Proceedings of the 2021 Conference on Em- pirical Methods in Natural Language Processing, EMNLP 2021...

  61. [69]

    Xu, and Graham Neubig

    Zhiruo Wang, Grace Cuenca, Shuyan Zhou, Frank F. Xu, and Graham Neubig. 2023. MCoNaLa: A Benchmark for Code Generation from Multiple Natural Languages. In Findings of the Association for Computational Linguistics: EACL 2023, Dubrovnik, Croatia, May 2-6, 2023 , Andreas Vlachos ...

  62. [70]

    Di Wu, Wasi Uddin Ahmad, Dejiao Zhang, Murali Krishna Ramanathan, and Xiaofei Ma. 2024. Repoformer: Selective Retrieval for Repository-Level Code Completion. CoRR abs/2403.10059 (2024)

  63. [71]

    Shitao Xiao, Zheng Liu, Yingxia Shao, and Zhao Cao. 2022. RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, ...

  64. [72]

    Pengcheng Yin, Bowen Deng, Edgar Chen, Bogdan Vasilescu, and Graham Neubig. 2018. Learning to mine aligned code and natural language pairs from stack overflow. In Proceedings of the 15th International Conference on Mining Software Repositories, MSR 2018, Gothenburg, Sweden, Ma...

  65. [73]

    Pengcheng Yin and Graham Neubig. 2017. A Syntactic Neural Model for General-Purpose Code Generation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers , Regina Barz...

  66. [74]

    Pengcheng Yin and Graham Neubig. 2018. TRANX: A Transition-based Neural Abstract Syntax Parser for Semantic Parsing and Code Generation. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018: System Demonstrations, Brussels, Belgi...

  67. [75]

    Zelle and Raymond J

    John M. Zelle and Raymond J. Mooney. 1996. Learning to Parse Database Queries Using Inductive Logic Programming. In Proceedings of the Thirteenth National Conference on Artificial Intelligence and Eighth Innovative Applications of Artificial Intelligence Conference, AAAI 96, I...

  68. [76]

    Wayne Xin Zhao, Jing Liu, Ruiyang Ren, and Ji-Rong Wen. 2024. Dense Text Retrieval Based on Pretrained Language Models: A Survey. ACM Trans. Inf. Syst. 42, 4 (2024), 89:1–89:60

  69. [77]

    Xu, Zhengbao Jiang, and Graham Neubig

    Shuyan Zhou, Uri Alon, Frank F. Xu, Zhengbao Jiang, and Graham Neubig. 2023. DocPrompting: Generating Code by Retrieving the Docs. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net

  70. [78]

    Xu, Zhengbao Jiang, and Graham Neubig

    Shuyan Zhou, Uri Alon, Frank F. Xu, Zhengbao Jiang, and Graham Neubig. 2023. DocPrompting: Generating Code by Retrieving the Docs. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, J. ACM, Vol. 37, No. 4, Article 1. Publication da...

  71. [79]

    Ming Zhu, Aneesh Jain, Karthik Suresh, Roshan Ravindran, Sindhu Tipirneni, and Chandan K Reddy. 2022. XLCoST: A Benchmark Dataset for Cross-lingual Code Intelligence. arXiv preprint arXiv:2206.08474 (2022). J. ACM, Vol. 37, No. 4, Article 1. Publication date: August 2018

  72. [2018]

    In Proceedings of the 2018 Conference on Empirical Methods in Natural J

    Retrieval-Based Neural Code Generation. In Proceedings of the 2018 Conference on Empirical Methods in Natural J. ACM, Vol. 37, No. 4, Article 1. Publication date: August 2018. 1:26 Z. Yang et al. Language Processing, Brussels, Belgium, October 31 - November 4, 2018 , Ellen Ril...

  73. [2022]

    In Findings of the Association for Computational Linguistics: ACL 2022, Dublin, Ireland, May 22-27, 2022 , Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.)

    Compilable Neural Code Generation with Compiler Feedback. In Findings of the Association for Computational Linguistics: ACL 2022, Dublin, Ireland, May 22-27, 2022 , Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, 9–19

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.