REVIEW 3 major objections 6 minor 4 cited by
Language Models for Code Optimization: Survey, Challenges and Future Directions
T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A systematic review of 53 studies maps how language models are used to optimize code, where current approaches cluster, and why the field still struggles with real-world programs.
desk verdict A useful, much-needed survey of LM-based code optimization; the taxonomy is solid, but the headline percentages rest on a selection procedure that the paper delegates to an external repo. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery of the paper is the systematic literature review itself: a Kitchenham-and-Charters-style protocol with a quasi-gold-standard search string, snowballing, inclusion/exclusion screening, quality assessment, and 53 retained primary studies. The analytic engine is a taxonomy built from four research questions (RQ1, characteristics of the LMs; RQ2, how LMs were applied; RQ3, how the optimization problem was defined; RQ4, how methods were evaluated) and 11 sub-questions, which lets the authors count and cross-tabulate model types, parameter sizes, training strategies, challenges, techniques, roles, languages, metrics, and datasets into the distributions that carry their findings.
What would settle it
Re-run the search on the six indexing engines and check whether the 53-study set matches the repository's full list; any substantial missing primary study could shift the reported statistics. If a systematic replication adds enough multi-language or real-world-evaluated studies, the 81% and 68% figures would no longer be representative.
Extended reading notes
Core claim
The paper's central claim is a descriptive one: LM-based code optimization is a fast-growing but immature area whose current practice concentrates in a few comfortable settings. Across 53 studies, general-purpose LMs like GPT-4 are used more often than code-specialized models (61 vs 43 instances), 57% of studies use off-the-shelf models while 43% fine-tune, and the most common technical response to optimization failure is iterative feedback rather than one-shot generation. The review also reports that 81% of studies optimize a single language, 79% focus on a single performance metric, and 68% do not evaluate on real-world programs at all. From these patterns the authors derive five open challenges—balancing model complexity with practicality, interacting with external systems, generalizing across languages and metrics, evaluating on real code, and building trust in model outputs—and outline eight directions for future work.
Load-bearing premise
The survey's numbers and conclusions hold only if the automatic search plus snowballing found all or most relevant studies and the inclusion/exclusion criteria did not bias the set of 53 primary studies.
Editorial extensions
If this is right
- If the survey's picture is right, the default workflow in this field is a general-purpose LM combined with execution feedback, and new methods should benchmark against that combination.
- The dominance of single-language and single-metric studies means cross-language and multi-objective optimization are under-explored, so early entrants can claim clear ground.
- Because 68% of studies avoid real-world code, reported gains on competitive-programming tasks should not be read as production speedups until validated on full projects.
- The five challenges imply that practical adoption will be gated by trust and integration, not raw model capability, favoring agentic and human-in-the-loop designs.
- The review's own counts give practitioners a checklist for evaluating a new optimizer: Which model? Which role? Which language? Which metric? Which data?
Reading between the lines
- A likely consequence the authors leave implicit: the statistics treat each paper as one data point, so a few highly cited studies may inflate some categories; a future survey could weight findings by evaluation scale and replication.
- The pattern that most systems use competitive-programming data suggests a testable extension: building an optimization benchmark from heterogeneous open-source repositories would probably lower reported speedups and reveal which techniques depend on benchmark style.
- The eight future directions point toward convergence with the broader agentic software-engineering trend; one concrete prediction is that optimization benchmarks will begin to require tool use, multi-turn iteration, and correctness checks, not just runtime improvement.
- Another implicit extension is economic: since 57% of studies use off-the-shelf models and only a few train from scratch, the field's progress is tied to commercial API availability, so open-weight models deserve targeted evaluation for optimization tasks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a systematic literature review (SLR) of language-model-based code optimization. It follows the Kitchenham and Charters guidelines, searches six academic indexing engines supplemented by snowballing, and selects 53 primary studies. The study organizes the field into four research questions with 11 sub-questions, covering LM characteristics, application techniques, problem definition, and evaluation methods. It proposes a taxonomy of models, roles, techniques, languages, metrics, and datasets, and derives five open challenges and eight future research directions. The headline findings include that 81% of studies target a single language, 68% are not evaluated on real-world code, and general-purpose LMs are used more often than code-specialized LMs.
Significance. If the survey's findings are reliable, this is the first comprehensive SLR devoted specifically to LM-based code optimization, filling a genuine gap in the literature. The taxonomy and the synthesis of trends (e.g., the prevalence of feedback-based iterative optimization, the dominance of Python and single-language studies, the gap in real-world evaluation) provide a useful orientation for both researchers and practitioners in a rapidly evolving field. The authors have also made raw results available in a GitHub repository, which supports transparency and allows others to verify the classification. The paper's recommendations, such as the need for standardized real-world benchmarks and multi-objective optimization, are actionable. The contribution is primarily descriptive, but a systematic map of this kind is valuable in a field where the primary literature is growing quickly.
major comments (3)
- [Section 3 (Methodology) and footnote 3] The methodology section does not report the exact search string, the names of the six academic indexing engines, the inclusion and exclusion criteria, or the quality-assessment thresholds. Footnote 3 delegates all of these to an external GitHub repository, which is not part of the reviewed manuscript. Since every headline statistic in the abstract and in Sections 6 and 7 (e.g., 81% single-language, 68% non-real-world evaluation) is computed over a set of 53 studies selected through these unreported criteria, the central claim of the survey cannot be fully audited from the paper itself. Please include the complete protocol in the paper or in a stable, versioned appendix, and provide at least a summary of the key criteria in the main text.
- [Section 3, Figure 4] The quasi-gold standard step is shown schematically (10 studies) but the paper does not report how the gold standard set was constructed or the resulting precision/recall of the search string. In the quasi-gold standard methodology (Zhang et al., [152]), these validation figures are necessary to demonstrate that the automatic search is sufficiently sensitive. Without them, a reader cannot judge whether the 53-study set is a complete and unbiased representation of the relevant literature.
- [Section 3 and Figure 5] The data-extraction and classification process is not described. The paper does not mention pilot extraction, inter-rater agreement, or how conflicts were resolved when assigning studies to taxonomy categories such as those in Tables 2-6. Because the taxonomy and the associated percentages are the main contributions of the survey, the reliability of these manual assignments is load-bearing. A brief description of the validation process for data extraction and classification should be added.
minor comments (6)
- [Section 1 and Section 3] The paper states that searches were conducted via 'six academic indexing engines' but never names these engines. Naming them in the methodology would improve transparency.
- [Section 7.1.3] The text says 'Compiler datasets were used in six compiler-related tasks,' but Table 7 lists seven instances in the Compiler category. The text and the table should be made consistent.
- [Throughout] The repeated markers '♂search' and '/thumbs-up' appear before each 'Finding' and 'Recommendation' paragraph. These appear to be unrendered icon artifacts and should be fixed in the final version.
- [Tables 1, 5, Section 7.1.2, and Section 9] There are typographical errors: 'Ealier' in Table 1, 'Languague' and 'Heuristsic' in Table 5, 'Genereal' in Section 7.1.2, and 'cataloger' in the conclusion. These should be corrected.
- [Section 2] Footnote 2 states that the full related works section is available only in the external repository. For a survey, it would be preferable to include an expanded related-work summary in the paper itself or at least briefly describe the main categories of related work beyond the short background given.
- [Abstract and Section 1] The abstract says 'over 50 primary studies' while Section 1 states '53 primary studies.' Consider using the exact number in the abstract for consistency, or explicitly say '53' if the final count is fixed.
Circularity Check
No circular derivation: the survey's findings are descriptive syntheses of external primary studies, and its own background self-citations are not load-bearing.
full rationale
This paper is a systematic literature review, not a derivation of predictions from fitted parameters. Its central outputs—the taxonomy, the 11 sub-question findings, and the headline percentages (e.g., 81% single-language, 68% non-real-world evaluation)—are descriptive counts over the 53 selected primary studies. Those studies are external works, and the survey does not define any of its categories in terms of its own conclusions, nor does it fit a parameter and then rename that fit as a finding. The authors do cite their own prior work in several places: reference [26] (DeepTune, co-authored by Zheng Wang), [43] (Artemis++, co-authored by Giavrimis and Basios), [44] and [45] (Gong and Chen), [111] (Rocha, Petoumenos, Wang, Cole, Leather), and [137] (Wang and O'Boyle). However, each of these appears as background context or as one cited example among many, and none is invoked as the evidence establishing the survey's novel findings or its statistics. The survey's methodology section does defer the full search string, inclusion/exclusion criteria, and quality-assessment details to an external GitHub repository, which is an auditability and reproducibility concern rather than a circularity concern: the selection procedure is an input to the statistics, but nothing in the paper equates the selection procedure with the survey's conclusions by construction. The five challenges and eight future directions are recommendations synthesized from the reviewed literature, not deductions that reduce to the survey's own inputs. Therefore, no specific circular step can be quoted or exhibited, and the honest finding is no significant circularity.
Assumptions & free parameters
assumptions (2)
- domain assumption The set of 53 primary studies selected through automatic search, snowballing, inclusion/exclusion criteria, and quality assessment is representative of the full literature on LM-based code optimization.
- domain assumption The author-defined taxonomy for categorizing LMs, challenges, techniques, roles, languages, metrics, and datasets accurately reflects the content of the primary studies.
Cite this review
Pith. "Pith review of Language Models for Code Optimization: Survey, Challenges and Future Directions." pith.science (2026). https://pith.science/paper/J3W6AFHA
@misc{pith2026250101277,
author = {Pith},
title = {Pith review of: Language Models for Code Optimization: Survey, Challenges and Future Directions},
year = {2026},
howpublished = {\url{https://pith.science/paper/J3W6AFHA}},
note = {Machine review of arXiv:2501.01277}
}
read the original abstract
Language models (LMs) built upon deep neural networks (DNNs) have recently demonstrated breakthrough effectiveness in software engineering tasks such as code generation, completion, and repair. This has paved the way for the emergence of LM-based code optimization techniques, which are crucial for enhancing the performance of existing programs, such as accelerating program execution time. However, a comprehensive survey dedicated to this specific application has been lacking. To fill this gap, we present a systematic literature review of over 50 primary studies, identifying emerging trends and addressing 11 specialized questions. Our findings reveal five critical open challenges, such as balancing model complexity with practical usability, cross-language/performance generalizability, and building trust in AI-driven solutions. Furthermore, we provide eight future research directions to facilitate more efficient, robust, and reliable LM-based code optimization. Thereby, this study aims to provide actionable insights and foundational references for both researchers and practitioners in this rapidly evolving field.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 4 Pith papers
-
JETO-Bench: A Reproducible Benchmark for Execution Time Improvement Patches in Java
JETO-Mine is a reusable three-phase pipeline that mines 1.8 million Java commits to produce JETO-Bench containing 91 verified executable ETIPs, on which OpenHands succeeds at 14.3%.
-
Rethinking LLM-Based RTL Code Optimization Via Timing Logic Metamorphosis
LLM-based RTL optimizers degrade on timing-heavy mutants, but the study's own data and methods do not fully support the headline claim.
-
A Comprehensive Survey on Integrating Large Language Models with Knowledge-Based Methods
A narrative review of LLM knowledge integration that categorizes techniques and compiles benchmarks, but lacks a systematic method and contains unreliable citations.
-
Enhancing Trust in Language Model-Based Code Optimization through RLHF: A Research Design
A research design proposes applying reinforcement learning from human feedback to improve trust in language model based code optimization, with no experiments yet.
Reference graph
Works this paper leans on
-
[152]
He Zhang, Muhammad Ali Babar, and Paolo Tell. 2011. Identifying Relevant Studies in Software Engineering. Information and Software Technology 53 (2011), 625–637
work page 2011
-
[1]
Abella-González, Pedro Carollo-Fernández, Louis-Noël Pouchet, Fabrice Rastello, and Gabriel Rodríguez
Miguel Á. Abella-González, Pedro Carollo-Fernández, Louis-Noël Pouchet, Fabrice Rastello, and Gabriel Rodríguez
-
[2]
Felix Adler, Gordon Fraser, Eva Gründinger, Nina Körber, Simon Labrenz, Jonas Lerchenberger, Stephan Lukasczyk, and Sebastian Schweikl. 2021. Improving Readability of Scratch Programs with Search-based Refactoring. In International Working Conference on Source Code Analysis and Manipulation, SCAM . IEEE, 120–130
2021
-
[3]
Randy Allen and Steve Johnson. 1988. Compiling C for Vectorization, Parallelization, and Inline Expansion. ACM SIGPLAN Notices 23 (1988), 241–249
1988
-
[4]
Rohan Anil, Sebastian Borgeaud, and et al. 2023. Gemini: A Family of Highly Capable Multimodal Models. (2023). arXiv:2312.11805
arXiv 2023
-
[5]
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. PaLM 2 Technical Report. arXiv:2305.10403 (2023)
arXiv 2023
-
[6]
Anthropic. 2024. Introducing the next generation of Claude
2024
-
[7]
Amir H Ashouri, William Killian, John Cavazos, Gianluca Palermo, and Cristina Silvano. 2018. A Survey on Compiler Autotuning using Machine Learning. Computing Surveys (CSUR) 51 (2018), 1–42
2018
Show all 163 references
-
[8]
Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J
Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. 2021. Program Synthesis with Large Language Models. (2021). arXiv:2108.07732
2021 arXiv
-
[9]
John Backus. 1978. The history of Fortran I, II, and III. ACM Sigplan Notices 13, 8 (1978), 165–180
1978
-
[10]
Riyadh Baghdadi, Massinissa Merouani, Mohamed-Hicham Leghettas, Kamel Abdous, Taha Arbaoui, Karima Be- natchba, et al. 2021. A Deep Learning Based Cost Model for Automatic Code Optimization. Proceedings of Machine Learning and Systems 3 (2021), 181–193. ACM Comput. Surv., Vol....
2021
-
[11]
M Ammar Ben Khadra, Dominik Stoffel, and Wolfgang Kunz. 2020. Efficient Binary-Level Coverage Analysis. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 1153–1164
2020
-
[12]
K Manasvi Bhat, Pratiksha P Anchalia, Rushali Mohbe, and A Parkavi. 2019. A Survey of Machine Learning and Deep Learning Techniques for Compiler Optimization. International Journal of Research in Engineering, Science and Management (2019)
2019
-
[13]
BigScience. 2022. BLOOM: A 176B-Parameter Open-Access Multilingual Language Model. (2022). arXiv:2211.05100
2022 arXiv
-
[14]
Nathan Binkert, Bradford Beckmann, Gabriel Black, Steven K Reinhardt, Ali Saidi, Arkaprava Basu, Joel Hestness, Derek R Hower, Tushar Krishna, Somayeh Sardashti, et al. 2011. The gem5 Simulator. ACM SIGARCH computer architecture news 39 (2011), 1–7
2011
-
[15]
Sid Black, Stella Biderman, Eric Hallahan, and et al. 2022. GPT-NeoX-20B: An Open-Source Autoregressive Language Model. arXiv:2204.06745
2022 arXiv
-
[16]
Tom B Brown. 2020. Language Models are Few-Shot Learners. NeurIPS (2020)
2020
-
[17]
Rosario Cammarota, Alexandru Nicolau, Alexander V Veidenbaum, Arun Kejariwal, Debora Donato, and Mukund Madhugiri. 2013. On the Determination of Inlining Vectors for Program Optimization. In Compiler Construction (CC). Springer, 164–183
2013
-
[18]
John Cavazos, Christophe Dubach, Felix Agakov, Edwin Bonilla, Michael FP O’Boyle, Grigori Fursin, and Olivier Temam. 2006. Automatic Performance Model Construction for the Fast Software Exploration of New Hardware Designs. In Proceedings of the 2006 international conference on...
2006
-
[19]
Gregory J Chaitin. 1982. Register Allocation & Spilling via Graph Coloring. ACM Sigplan Notices 17 (1982), 98–101
1982
-
[20]
Lihu Chen and Gaël Varoquaux. 2024. What is the Role of Small Models in the LLM Era: A Survey. arXiv:2409.06857 (2024)
2024 arXiv
-
[22]
2018.{TVM}: An automated{End-to-End} optimizing compiler for deep learning
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, et al. 2018.{TVM}: An automated{End-to-End} optimizing compiler for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implem...
2018
-
[23]
Zimin Chen, Sen Fang, and Martin Monperrus. 2024. Supersonic: Learning to Generate Source Code Optimizations in C/C++. IEEE Transactions on Software Engineering (TSE) (2024)
2024
-
[24]
Jinsu Choi, Gabin An, and Shin Yoo. 2024. Iterative Refactoring of Real-World Open-Source Programs with Large Language Models. In International Symposium on Search Based Software Engineering . Springer, 49–55
2024
-
[25]
Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy, and Neel Sundaresan
Colin B. Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy, and Neel Sundaresan. 2020. PyMT5: Multi- Mode Translation of Natural Language and Python Code with Transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMN...
2020
-
[26]
Chris Cummins, Pavlos Petoumenos, Zheng Wang, and Hugh Leather. 2017. End-to-end deep learning of optimization heuristics. In 2017 26th International Conference on Parallel Architectures and Compilation Techniques (PACT) . IEEE, 219–232
2017
-
[27]
Hazelwood, Gabriel Synnaeve, and Hugh Leather
Chris Cummins, Volker Seeker, Dejan Grubisic, Mostafa Elhoushi, Youwei Liang, Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Kim M. Hazelwood, Gabriel Synnaeve, and Hugh Leather. 2023. Large Language Models for Compiler Optimization. (2023). arXiv:2309.07062
2023 arXiv
-
[28]
Chris Cummins, Volker Seeker, Dejan Grubisic, Baptiste Rozière, Jonas Gehring, Gabriel Synnaeve, and Hugh Leather
-
[29]
Matthew Curtis-Maury, James Dzierwa, Christos D Antonopoulos, and Dimitrios S Nikolopoulos. 2006. Online Power-Performance Adaptation of Multithreaded Programs Using Hardware Event-Based Prediction. In Proceedings of the 20th annual international conference on Supercomputing . 157–166
2006
-
[30]
Anderson Faustino da Silva, Bruno Conde Kind, José Wesley de Souza Magalhães, and et al. 2021. ANGHABENCH: A Suite with One Million Compilable C Benchmarks for Code-Size Reduction. In IEEE/ACM International Symposium on Code Generation and Optimization, CGO . IEEE, 378–390
2021
-
[31]
Deepseek-AI. 2023. DeepSeek Coder: Let the Code Write Itself. https://deepseekcoder.github.io/
2023
-
[32]
Ahmed, Guixiang Ma, Mihai Capota, Theodore L
Shukai Duan, Nikos Kanakaris, Xiongye Xiao, Heng Ping, Chenyu Zhou, Nesreen K. Ahmed, Guixiang Ma, Mihai Capota, Theodore L. Willke, Shahin Nazarian, and Paul Bogdan. 2023. Leveraging Reinforcement Learning and Large Language Models for Code Optimization. (2023). arXiv:2312.05657
2023 arXiv
-
[33]
Rudolf Eigenmann and Jay Hoeflinger. 2000. Parallelizing and Vectorizing Compilers. Proc. IEEE (2000)
2000
-
[34]
Angela Fan, Beliz Gokkaya, Mark Harman, Mitya Lyubarskiy, Shubho Sengupta, Shin Yoo, and Jie M. Zhang. 2023. Large Language Models for Software Engineering: Survey and Open Problems. InInternational Conference on Software Engineering: Future of Software Engineering, ICSE-FoSE ...
2023
-
[35]
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Findings of the Association for Computational Linguistics: EMNLP ...
2020
-
[36]
Mary F Fernandez. 1995. Simple and Effective Link-Time Optimization of Modula-3 Programs. In Proceedings of the ACM SIGPLAN 1995 conference on Programming language design and implementation . 103–115
1995
-
[37]
Grigori Fursin, Yuriy Kashnikov, Abdul Wahid Memon, Zbigniew Chamski, Olivier Temam, Mircea Namolaru, Elad Yom-Tov, Bilha Mendelson, Ayal Zaks, Eric Courtois, et al. 2011. Milepost gcc: Machine learning enabled self-tuning compiler. International journal of parallel programmin...
2011
-
[38]
Shuzheng Gao, Cuiyun Gao, Wenchao Gu, and Michael Lyu. 2024. Search-Based LLMs for Code Optimization. In International Conference on Software Engineering (ICSE) . IEEE, 254–266
2024
-
[39]
Clement, Neel Sundaresan, and Chen Wu
Spandan Garg, Roshanak Zilouchian Moghaddam, Colin B. Clement, Neel Sundaresan, and Chen Wu. 2022. DeepDev- PERF: A Deep Learning-Based Approach for Improving Software Performance. In ESEC/FSE. ACM, 948–958
2022
-
[40]
Spandan Garg, Roshanak Zilouchian Moghaddam, and Neel Sundaresan. 2023. RAPGen: An Approach for Fixing Code Inefficiencies in Zero-Shot. (2023). arXiv:2306.17077
2023 arXiv
-
[41]
Leonidas Gee, Milan Gritta, Gerasimos Lampouras, and Ignacio Iacobacci. 2024. Code-Optimise: Self-Generated Preference Data for Correctness and Efficiency. (2024). arXiv:2406.12502
2024 arXiv
-
[42]
Sia Gholami. 2024. Can Pruning Make Large Language Models More Efficient? In Redefining Security With Cyber AI . IGI Global, 1–14
2024
-
[43]
Rafail Giavrimis, Alexis Butler, Constantin Cezar Petrescu, Michail Basios, and Santanu Kumar Dash. 2021. Genetic Optimisation of C++ Applications. In International Conference on Automated Software Engineering, ASE . IEEE, 1180– 1182
2021
-
[44]
Jingzhi Gong and Tao Chen. 2024. Deep Configuration Performance Learning: A Systematic Survey and Taxonomy. ACM Transactions on Software Engineering and Methodology (TOSEM) (2024)
2024
-
[45]
Jingzhi Gong and Tao Chen. 2024. Predicting Configuration Performance in Multiple Environments with Sequential Meta-Learning. Proceedings of the ACM on Software Engineering FSE (2024), 359–382
2024
-
[46]
Google. 2023. Google Cloud launches new AI models. https://cloud.google.com/blog/products/ai-machine-learning/ google-cloud-launches-new-ai-models-opens-generative-ai-studio
2023
-
[47]
Dejan Grubisic, Chris Cummins, Volker Seeker, and Hugh Leather. 2024. Compiler Generated Feedback for Large Language Models. (2024). arXiv:2403.14714
2024 arXiv
-
[48]
Mellor-Crummey, and Chris Cummins
Dejan Grubisic, Volker Seeker, Gabriel Synnaeve, Hugh Leather, John M. Mellor-Crummey, and Chris Cummins. 2024. Priority Sampling of Large Language Models for Compilers. In Proceedings of the 4th Workshop on Machine Learning and Systems, EuroMLSys. ACM, 91–97
2024
-
[49]
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y. Wu, Y. K. Li, Fuli Luo, Yingfei Xiong, and Wenfeng Liang. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence. (2024). arXiv:2401.14196
2024 arXiv
-
[50]
Zifan Carl Guo and William S. Moses. 2022. Enabling Transformers to Understand Low-Level Programs. In High Performance Extreme Computing Conference (HPEC) . IEEE, 1–9
2022
-
[51]
Priyanshu Gupta, Avishree Khare, Yasharth Bajpai, Saikat Chakraborty, Sumit Gulwani, Aditya Kanade, Arjun Radhakrishna, Gustavo Soares, and Ashish Tiwari. 2023. Grace: Language Models Meet Code Edits. In Joint European Software Engineering Conference and Symposium on the Found...
2023
-
[52]
Bing Han, Congfei Li, Hua Deng, Guowei Liu, and Ze Zheng. 2024. Domain-Specific Translation Tool from Structured Text to C Source Code with Code Readability Enhancement in Programmable Logic Controllers. Concurrency and Computation: Practice and Experience (2024), e8100
2024
-
[53]
Xu Han, Qiannan Yang, Xianda Chen, Xiaowen Chu, and Meixin Zhu. 2024. Generating and Evolving Reward Functions for Highway Driving with Large Language Models. arXiv:2406.10540 (2024)
2024 arXiv
-
[54]
Erik Hemberg, Stephen Moskal, and Una-May O’Reilly. 2024. Evolving Code with a Large Language Model. Genetic Programming and Evolvable Machines 25 (2024), 21
2024
-
[55]
Dan Hendrycks, Steven Basart, Saurav Kadavath, and et. al. 2021. Measuring Coding Challenge Competence With APPS. NeurIPS (2021)
2021
-
[56]
Charles Hong, Sahil Bhatia, Altan Haan, Shengjun Kris Dong, Dima Nikiforov, Alvin Cheung, and Yakun Sophia Shao. 2024. LLM-Aided Compilation for Tensor Accelerators. (2024), 1–14
2024
-
[57]
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. 2023. Large Language Models for Software Engineering: A Systematic Literature Review. ACM Transactions on Software Engineering and Methodology (2023)
2023
-
[58]
Dong Huang, Jianbo Dai, Han Weng, Puzhen Wu, QING Yuhao, Heming Cui, Zhijiang Guo, and Jie Zhang. 2024. EffiLearner: Enhancing Efficiency of Generated Code via Self-Optimization. (2024). ACM Comput. Surv., Vol. 1, No. 1, Article . Publication date: January 2024. Language Model...
2024
-
[59]
Dong Huang, Guangtao Zeng, Jianbo Dai, Meng Luo, Han Weng, Yuhao Qing, Heming Cui, Zhijiang Guo, and Jie M Zhang. 2024. Effi-Code: Unleashing Code Efficiency in Language Models. arXiv:2410.10209 (2024)
2024 arXiv
-
[60]
Zhang, Yuhao Qing, and Heming Cui
Dong Huang, Jie M. Zhang, Yuhao Qing, and Heming Cui. 2024. EffiBench: Benchmarking the Efficiency of Automati- cally Generated Code. (2024). arXiv:2402.02037
2024 arXiv
-
[61]
Binyuan Hui, Jian Yang, Zeyu Cui, Jiaxi Yang, Dayiheng Liu, Lei Zhang, Tianyu Liu, Jiajun Zhang, Bowen Yu, Keming Lu, et al. 2024. Qwen2.5-coder Technical Report. arXiv:2409.12186 (2024)
2024 arXiv
-
[62]
Shu Ishida, Gianluca Corrado, George Fedoseev, Hudson Yeo, Lloyd Russell, Jamie Shotton, Joao F Henriques, and Anthony Hu. 2024. LangProp: A Code Optimization Framework Using Large Language Models Applied to Driving
2024
-
[63]
Naman Jain, Skanda Vaidyanath, Arun Iyer, Nagarajan Natarajan, Suresh Parthasarathy, Sriram Rajamani, and Rahul Sharma. 2022. Jigsaw: Large Language Models Meet Program Synthesis. In International Conference on Software Engineering (ICSE). 1219–1231
2022
-
[64]
Mojan Javaheripi, Sébastien Bubeck, Marah Abdin, Jyoti Aneja, Sebastien Bubeck, Caio César Teodoro Mendes, Weizhu Chen, Allie Del Giorno, Ronen Eldan, Sivakanth Gopi, et al . 2023. Phi-2: The Surprising Power of Small Language Models. Microsoft Research Blog 1 (2023), 3
2023
-
[65]
Matthew Jin, Syed Shahriar, Michele Tufano, Xin Shi, Shuai Lu, Neel Sundaresan, and Alexey Svyatkovskiy. 2023. InferFix: End-to-End Program Repair with LLMs. In Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, ESEC/FSE . ...
2023
-
[66]
René Just, Darioush Jalali, and Michael D Ernst. 2014. Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs. In International Symposium on Software Testing and Analysis (ISSTA) . 437–440
2014
-
[67]
Sungmin Kang and Shin Yoo. 2023. Towards Objective-Tailored Genetic Improvement Through Large Language Models. In International Workshop on Genetic Improvement, GI@ICSE 2023 . IEEE, 19–20
2023
-
[68]
Ryan Kastner, Janarbek Matai, and Stephen Neuendorffer. 2018. Parallel programming for FPGAs. arXiv preprint arXiv:1805.03648 (2018)
2018 arXiv
-
[69]
Charters
Barbara Kitchenham and Stuart M. Charters. 2007. Guidelines for Performing Systematic Literature Reviews in Software Engineering
2007
-
[70]
Yuhang Lai, Chengxi Li, and Yiming et al. Wang. 2023. DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation. In International Conference on Machine Learning . 18319–18345
2023
-
[71]
Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven Chu-Hong Hoi. 2022. CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning. In Neural Information Processing Systems (NeurIPS)
2022
-
[72]
Jia Li, Ge Li, Zhuo Li, Zhi Jin, Xing Hu, Kechi Zhang, and Zhiyi Fu. 2023. CodeEditor: Learning to Edit Source Code with Pre-trained Models. ACM Transactions on Software Engineering and Methodology (TOSEM) 32 (2023), 1–22
2023
-
[73]
Kaixin Li, Qisheng Hu, Xu Zhao, Hui Chen, Yuxi Xie, Tiedong Liu, Qizhe Xie, and Junxian He. 2024. InstructCoder: Instruction Tuning Large Language Models for Code Editing
2024
-
[74]
Raymond Li, Loubna Ben Allal, and et al. 2023. StarCoder: May the Source Be With You! Transactions on Machine Learning Research 2023 (2023)
2023
-
[75]
Rui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev, and P Sadayappan. 2021. Analytical Characterization and Design Space Exploration for Optimization of CNNs. In Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and ...
2021
-
[76]
Yujia Li, David Choi, Junyoung Chung, and et al. 2022. Competition-Level Code Generation with AlphaCode. Science 378 (2022), 1092–1097
2022
-
[77]
Zeyuan Li, Yangfan He, Lewei He, Jianhui Wang, Tianyu Shi, Bin Lei, Yuchen Li, and Qiuwu Chen. 2024. FALCON: Feedback-driven Adaptive Long/short-term memory reinforced Coding Optimization system.arXiv:2410.21349 (2024)
2024 arXiv
-
[78]
Nandor Licker and Timothy M Jones. 2020. Duplo: A Framework for OCaml Post-Link Optimisation. Proceedings of the ACM on Programming Languages 4, ICFP (2020), 1–29
2020
-
[79]
Junwei Liu, Kaixin Wang, Yixuan Chen, Xin Peng, Zhenpeng Chen, Lingming Zhang, and Yiling Lou. 2024. Large Language Model-Based Agents for Software Engineering: A Survey. arXiv:2409.02977 (2024)
2024 arXiv
-
[80]
Edward S Lowry and Cleburne W Medlock. 1969. Object code optimization. Commun. ACM 12, 1 (1969), 13–22
1969
-
[81]
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, et al. 2024. StarCoder 2 and The Stack v2: The Next Generation. arXiv:2402.19173 (2024)
2024 arXiv
-
[82]
Jinliang Lu, Ziliang Pang, Min Xiao, Yaochen Zhu, Rui Xia, and Jiajun Zhang. 2024. Merge, Ensemble, and Cooperate! A Survey on Collaborative Strategies in the Era of Large Language Models. arXiv:2407.06089 (2024)
2024 arXiv
-
[83]
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2024. WizardCoder: Empowering Code Large Language Models with Evol-Instruct. In The Twelfth International Conference on Learning Representations, ICLR 2024...
2024
-
[84]
Aman Madaan, Niket Tandon, Prakhar Gupta, and et al. 2023. Self-Refine: Iterative Refinement with Self-Feedback. In Neural Information Processing Systems (NeurIPS)
2023
-
[85]
Saeed Maleki, Yaoqing Gao, Maria J Garzar, Tommy Wong, David A Padua, et al. 2011. An Evaluation of Vectorizing Compilers. In 2011 International Conference on Parallel Architectures and Compilation Techniques . IEEE, 372–382
2011
-
[86]
William M McKeeman. 1965. Peephole Optimization. Commun. ACM 8 (1965), 443–444
1965
-
[87]
Thomas Mesnard, Cassidy Hardin, and et al. 2024. Gemma: Open Models Based on Gemini Research and Technology. (2024). arXiv:2403.08295
2024 arXiv
-
[88]
Meta. 2024. Introducing Llama 3.1: Our most capable models to date. https://ai.meta.com/blog/meta-llama-3-1/
2024
-
[89]
Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024. Large Language Models: A Survey. arXiv:2402.06196 (2024)
2024 arXiv
-
[90]
Bolin Ni, Jingcheng Hu, Yixuan Wei, Houwen Peng, Zheng Zhang, Gaofeng Meng, and Han Hu. 2024. Xwin-LM: Strong and Scalable Alignment Practice for LLMs. (2024). arXiv:2405.20335
2024 arXiv
-
[91]
Daniel Nichols, Pranav Polasam, Harshitha Menon, Aniruddha Marathe, Todd Gamblin, and Abhinav Bhatele. 2024. Performance-Aligned LLMs for Generating Fast Code. (2024). arXiv:2404.18864
2024 arXiv
-
[92]
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong
-
[93]
Behrooz Omidvar Tehrani and Anmol Anubhai. 2024. Evaluating Human-AI Partnership for LLM-based Code Migration. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems . 1–8
2024
-
[94]
OpenAI. 2021. Evaluating Large Language Models Trained on Code. (2021). arXiv:2107.03374
2021 arXiv
-
[95]
OpenAI. 2022. Introducing ChatGPT. https://openai.com/index/chatgpt/
2022
-
[96]
OpenAI. 2024. GPT-4 Technical Report. arXiv:2303.08774
2024 arXiv
-
[97]
OpenAI. 2024. Hello GPT-4o. https://openai.com/index/hello-gpt-4o/
2024
-
[98]
Marek Palkowski and Mateusz Gruzewski. 2024. GPT-Driven Source-to-Source Transformation for Generating Compilable Parallel CUDA Code for Nussinov’s Algorithm. Electronics 13 (2024), 488
2024
-
[99]
Yue Pan and Chen Lyu. 2023. Measuring Efficient Code Generation with GEC. In Proceedings of the 14th Asia-Pacific Symposium on Internetware, Internetware . ACM, 249–258
2023
-
[100]
Yue Pan, Chen Lyu, Zhenyu Yang, Lantian Li, Qi Liu, and Xiuting Shao. 2024. E-code: Mastering Efficient Code Generation through Pretrained Models and Expert Encoder Group
2024
-
[101]
Yue Pan, Xiuting Shao, and Chen Lyu. 2025. Measuring Code Efficiency Optimization Capabilities with ACEOB
2025
-
[102]
Eunjung Park, John Cavazos, Louis-Noël Pouchet, Cédric Bastoul, Albert Cohen, and Ponnuswamy Sadayappan. 2013. Predictive Modeling in a Polyhedral Optimization Space. International journal of parallel programming 41 (2013), 704–750
2013
-
[103]
Rajvardhan Patil and Venkat Gudivada. 2024. A Review of Current Trends, Techniques, and Challenges in Large Language Models (LLMs). Applied Sciences 14 (2024), 2074
2024
-
[104]
Huiyun Peng, Arjun Gupte, Nicholas John Eliopoulos, Chien Chou Ho, Rishi Mantri, Leo Deng, Wenxin Jiang, Yung-Hsiang Lu, Konstantin Läufer, George K Thiruvathukal, et al. 2024. Large Language Models for Energy-Efficient Code: Emerging Results and Future Directions. arXiv:2410....
2024 arXiv
-
[105]
Yun Peng, Akhilesh Deepak Gotmare, Michael Lyu, Caiming Xiong, Silvio Savarese, and Doyen Sahoo. 2024. Perf- CodeGen: Improving Performance of LLM Generated Code with Execution Feedback. arXiv:2412.03578 (2024)
2024 arXiv
-
[106]
Rui Pereira, Marco Couto, Francisco Ribeiro, Rui Rua, Jácome Cunha, João Paulo Fernandes, and João Saraiva. 2017. Energy Efficiency Across Programming Languages: How Do Energy, Time, and Memory Relate?. In ACM SIGPLAN international conference on software language engineering . 256–267
2017
-
[107]
Kung, Geert Janssen, Wei Zhang, Giacomo Domeniconi, Vladimir Zolotov, Julian Dolby, Jie Chen, Mihir R
Ruchir Puri, David S. Kung, Geert Janssen, Wei Zhang, Giacomo Domeniconi, Vladimir Zolotov, Julian Dolby, Jie Chen, Mihir R. Choudhury, Lindsey Decker, Veronika Thost, Luca Buratti, Saurabh Pujar, and Ulrich Finkler. 2021. Project CodeNet: A Large-Scale AI for Code Dataset for...
2021 arXiv
-
[108]
Muzi Qu, Jie Liu, Liangyi Kang, Shuai Wang, Dan Ye, and Tao Huang. 2024. Dynamic Scoring Code Token Tree: A Novel Decoding Strategy for Generating High-Performance Code. InIEEE/ACM International Conference on Automated Software Engineering. 1308–1318
2024
-
[109]
Sebastian Rachuj, Dietmar Fey, and Marc Reichenbach. 2020. Impact of Performance Estimation on Fast Processor Simulators. In Simulation Tools and Techniques: 12th EAI International Conference, SIMUtools . Springer, 79–93
2020
-
[110]
Tal Ridnik, Dedy Kredo, and Itamar Friedman. 2024. Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering. (2024). arXiv:2401.08500
2024 arXiv
-
[111]
Rodrigo CO Rocha, Pavlos Petoumenos, Zheng Wang, Murray Cole, and Hugh Leather. 2019. Function merging by sequence alignment. In 2019 IEEE/ACM International Symposium on Code Generation and Optimization (CGO) . IEEE, 149–163. ACM Comput. Surv., Vol. 1, No. 1, Article . Publica...
2019
-
[112]
Pawan Kumar, Emilien Dupont, Francisco J
Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M. Pawan Kumar, Emilien Dupont, Francisco J. R. Ruiz, Jordan S. Ellenberg, Pengming Wang, Omar Fawzi, Pushmeet Kohli, and Alhussein Fawzi
-
[113]
Miguel Romero Rosas, Miguel Torres Sanchez, and Rudolf Eigenmann. 2024. Should AI Optimize Your Code? A Com- parative Study of Current Large Language Models Versus Classical Optimizing Compilers. (2024). arXiv:2406.12146
2024 arXiv
-
[114]
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, and et al. 2023. Code Llama: Open Foundation Models for Code. (2023). arXiv:2308.12950
2023 arXiv
-
[115]
Baptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, and Guillaume Lample. 2020. Unsupervised Translation of Programming Languages. In Annual Conference on Neural Information Processing Systems, NeurIPS
2020
-
[116]
Nature 625 (2024), 468–475
Mathematical Discoveries from Program Search with Large Language Models. Nature 625 (2024), 468–475
2024
-
[117]
Atsushi Shirafuji, Yusuke Oda, Jun Suzuki, Makoto Morishita, and Yutaka Watanobe. 2023. Refactoring Programs Using Large Language Models with Few-Shot Examples. In Asia-Pacific Software Engineering Conference, APSEC. IEEE, 151–160
2023
-
[118]
Karthik Shivashankar and Antonio Martini. 2024. Better Python Programming for all: With the focus on Maintainability. arXiv:2408.09134
2024 arXiv
-
[119]
Gardner, Yiming Yang, Milad Hashemi, Graham Neubig, Parthasarathy Ranganathan, Osbert Bastani, and Amir Yazdanbakhsh
Alexander Shypula, Aman Madaan, Yimeng Zeng, Uri Alon, Jacob R. Gardner, Yiming Yang, Milad Hashemi, Graham Neubig, Parthasarathy Ranganathan, Osbert Bastani, and Amir Yazdanbakhsh. 2024. Learning Performance-Improving Code Edits. In The Twelfth International Conference on Lea...
2024
-
[120]
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language Agents with Verbal Reinforcement Learning. In Neural Information Processing Systems (NeurIPS)
2023
-
[121]
Jagriti Sikka, Kushal Satya, Yaman Kumar, Shagun Uppal, Rajiv Ratn Shah, and Roger Zimmermann. 2020. Learning- Based Methods for Code Runtime Complexity Prediction. In Advances in Information Retrieval: 42nd European Conference on IR Research, ECIR . Springer, 313–325
2020
-
[122]
Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter. 2024. A Simple and Effective Pruning Approach for Large Language Models. ICLR (2024)
2024
-
[123]
Yiwen Sun, Xianyin Zhang, Shiyu Huang, Shaowei Cai, BingZhen Zhang, and Ke Wei. 2024. AutoSAT: Automatically Optimize SAT Solvers via Large Language Models. arXiv:2402.10705 (2024)
2024 arXiv
-
[124]
Schwartz, and Graham Neubig
Alex Shypula, Pengcheng Yin, Jeremy Lacomis, Claire Le Goues, Edward J. Schwartz, and Graham Neubig. 2022. Learning to Superoptimize Real-world Programs. (2022)
2022
-
[125]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. LLaMA: Open and Efficient Foundation La...
2023 arXiv
-
[126]
Hugo Touvron, Louis Martin, and et al. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. (2023). arXiv:2307.09288
2023 arXiv
-
[127]
Mircea Trofin, Yundi Qian, Eugene Brevdo, Zinan Lin, Krzysztof Choromanski, and David Li. 2021. Mlgo: a machine learning guided compiler optimizations framework. arXiv preprint arXiv:2101.04808 (2021)
2021 arXiv
-
[128]
Jubi Taneja, Avery Laird, Cong Yan, Madan Musuvathi, and Shuvendu K. Lahiri. 2024. LLM-Vectorizer: LLM-based Verified Loop Vectorizer. (2024). arXiv:2406.04693
2024 arXiv
-
[129]
Niki van Stein and Thomas Bäck. 2024. LLaMEA: A Large Language Model Evolutionary Algorithm for Automatically Generating Metaheuristics. arXiv:2405.20132 (2024)
2024 arXiv
-
[130]
Niki van Stein, Diederick Vermetten, and Thomas Bäck. 2024. In-the-loop Hyper-Parameter Optimization for LLM-Based Automated Design of Heuristics. arXiv:2410.16309 (2024)
2024 arXiv
-
[131]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems . ...
2017
-
[132]
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Well-Read Students Learn Better: On the Importance of Pre-training Compact Models. arXiv:1908.08962
2019 arXiv
-
[133]
Siddhant Waghjale, Vishruth Veerendranath, Zhiruo Wang, and Daniel Fried. 2024. ECCO: Can We Improve Model- Generated Code Efficiency Without Sacrificing Functional Correctness?. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP ....
2024
-
[134]
Simin Wang, Liguo Huang, Amiao Gao, Jidong Ge, Tengfei Zhang, Haitao Feng, Ishna Satyarth, Ming Li, He Zhang, and Vincent Ng. 2023. Machine/Deep Learning for Software Engineering: A Systematic Literature Review. IEEE Trans. Software Eng. 49, 3 (2023), 1188–1231
2023
-
[135]
Joty, and Steven C
Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi. 2021. CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP . Asso...
2021
-
[136]
Michael J Voss and Rudolf Eigemann. 2001. High-Level Adaptive Program Optimization with ADAPT. In Proceedings of the eighth ACM SIGPLAN symposium on Principles and practices of parallel programming . 93–102
2001
-
[137]
Zheng Wang and Michael O’Boyle. 2018. Machine Learning in Compiler Optimization. Proc. IEEE 106 (2018), 1879–1901
2018
-
[138]
Anjiang Wei, Allen Nie, Thiago SFX Teixeira, Rohan Yadav, Wonchan Lee, Ke Wang, and Alex Aiken. 2024. Improving Parallel Program Performance Through DSL-Driven Code Generation with LLM Optimizers. arXiv:2410.15625 (2024)
2024 arXiv
-
[139]
Michael Joseph Wolfe. 1982. Optimizing Supercompilers for Supercomputers . University of Illinois at Urbana- Champaign
1982
-
[140]
Zhichao Wang, Bin Bi, Shiva Kumar Pentyala, Kiran Ramnath, Sougata Chaudhuri, Shubham Mehrotra, Xiang-Bo Mao, Sitaram Asur, et al. 2024. A Comprehensive Survey of LLM Alignment Techniques: RLHF, RLAIF, PPO, DPO and More. arXiv:2407.16216 (2024)
2024 arXiv
-
[141]
Xu, Uri Alon, Graham Neubig, and Vincent Josua Hellendoorn
Frank F. Xu, Uri Alon, Graham Neubig, and Vincent Josua Hellendoorn. 2022. A Systematic Evaluation of Large Language Models of Code. In SIGPLAN International Symposium on Machine Programming . ACM, 1–10
2022
-
[142]
Haocheng Xu, Haotian Hu, and Sitao Huang. 2024. Optimizing High-Level Synthesis Designs with Retrieval- Augmented Large Language Models. In IEEE LLM Aided Design Workshop (LAD) . IEEE, 1–5
2024
-
[143]
Hanxiang Xu, Shenao Wang, Ningke Li, Kailong Wang, Yanjie Zhao, Kai Chen, Ting Yu, Yang Liu, and Haoyu Wang
-
[144]
Xingyu Wu, Sheng-hao Wu, Jibin Wu, Liang Feng, and Kay Chen Tan. 2024. Evolutionary Computation in the Era of Large Language Model: Survey and Roadmap. arXiv:2401.10034 (2024)
2024 arXiv
-
[145]
Qingyao Xu, Dingkang Yang, and Lihua Zhang. 2024. Code Optimization Chain-of-Thought: Structured Understanding and Self-Checking. In International Conference on Artificial Intelligence, Big Data and Algorithms . 425–430
2024
-
[146]
Xuejun Yang, Yang Chen, Eric Eide, and John Regehr. 2011. Finding and Understanding Bugs in C Compilers. In ACM SIGPLAN Conference on Programming Language Design and Implementation, PLDI . ACM, 283–294
2011
-
[147]
Xufeng Yao, Yiwen Wang, Xing Li, Yingzhao Lian, Ran Chen, Lei Chen, Mingxuan Yuan, Hong Xu, and Bei Yu. 2024. RTLRewriter: Methodologies for Large Models aided RTL Code Optimization
2024
-
[148]
Large Language Models for Cyber Security: A Systematic Literature Review. (2024). arXiv:2405.04760
2024
-
[149]
Jinglue Xu, Jialong Li, Zhen Liu, Nagar Anthel Venkatesh Suryanarayanan, Guoyuan Zhou, Jia Guo, Hitoshi Iba, and Kenji Tei. 2025. Large Language Models Synergize with Automated Machine Learning. Transactions on Machine Learning Research (TMLR) (2025)
2025
-
[150]
Differentiation
Mert Yüksekgönül, Federico Bianchi, Joseph Boen, Sheng Liu, Zhi Huang, Carlos Guestrin, and James Zou. 2024. TextGrad: Automatic "Differentiation" via Text. (2024). arXiv:2406.07496
2024 arXiv
-
[151]
Eric Zelikman, Eliana Lorch, Lester Mackey, and Adam Tauman Kalai. 2024. Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation. (2024)
2024
-
[153]
Fangke Ye, Jisheng Zhao, Jun Shirako, and Vivek Sarkar. 2023. Concrete Type Inference for Code Optimization using Machine Learning with SMT Solving. Proceedings of the ACM on Programming Languages (PACMPL) 7 (2023), 773–800
2023
-
[154]
Tong Ye, Tengfei Ma, Lingfei Wu, Xuhong Zhang, Shouling Ji, and Wenhai Wang. 2024. Iterative or Innovative? A Problem-Oriented Perspective for Code Optimization. (2024). arXiv:2406.11935
2024
-
[155]
Quanjun Zhang, Chunrong Fang, Yang Xie, Yuxiang Ma, Weisong Sun, Yun Yang, and Zhenyu Chen. 2024. A Systematic Literature Review on Large Language Models for Automated Program Repair. (2024). arXiv:2405.01466
2024
-
[156]
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A Survey of Large Language Models. arXiv:2303.18223 (2023)
2023 arXiv
-
[157]
Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, et al. 2020. Ansor: Generating{High-Performance} tensor programs for deep learning. InUSENIX symposium on operating systems design and implementation (...
2020
-
[158]
Kechi Zhang, Ge Li, Yihong Dong, Jingjing Xu, Jun Zhang, Jing Su, Yongfei Liu, and Zhi Jin. 2024. CodeDPO: Aligning Code Models with Self Generated and Verified Source Code. arXiv:2410.05605 (2024)
2024 arXiv
-
[159]
Peiyan Zhang, Haibo Jin, Leyang Hu, Xinnuo Li, Liying Kang, Man Luo, Yangqiu Song, and Haohan Wang. 2024. Revolve: Optimizing AI Systems by Tracking Response Evolution in Textual Optimization. arXiv:2412.03092 (2024)
2024 arXiv
-
[163]
Tianyu Zheng, Ge Zhang, Tianhao Shen, Xueling Liu, Bill Yuchen Lin, Jie Fu, Wenhu Chen, and Xiang Yue. 2024. OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement. In Findings of the Association for Computational Linguistics, ACL. Association for Compu...
2024
-
[164]
Xunyu Zhu, Jian Li, Yong Liu, Can Ma, and Weiping Wang. 2023. A Survey on Model Compression for Large Language Models. arXiv:2308.07633 (2023). ACM Comput. Surv., Vol. 1, No. 1, Article . Publication date: January 2024
2023 arXiv
-
[2021]
In CC ’21: 30th ACM SIGPLAN International Conference on Compiler Construction, Virtual Event
PolyBench/Python: Benchmarking Python Environments with Polyhedral Optimizations. In CC ’21: 30th ACM SIGPLAN International Conference on Compiler Construction, Virtual Event . ACM, 59–70
-
[2023]
In The Eleventh International Conference on Learning Representations, ICLR
CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis. In The Eleventh International Conference on Learning Representations, ICLR . OpenReview.net
-
[2024]
Meta Large Language Model Compiler: Foundation Models of Compiler Optimization. (2024). arXiv:2407.02524
2024 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.