Pith. sign in

REVIEW 3 major objections 6 minor 4 cited by

Language Models for Code Optimization: Survey, Challenges and Future Directions

T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A systematic review of 53 studies maps how language models are used to optimize code, where current approaches cluster, and why the field still struggles with real-world programs.

desk verdict A useful, much-needed survey of LM-based code optimization; the taxonomy is solid, but the headline percentages rest on a selection procedure that the paper delegates to an external repo. read the letter →

arxiv 2501.01277 v2 pith:J3W6AFHA submitted 2025-01-02 cs.SE

classification cs.SE
keywords languagemodelscodeoptimizationsystematicliteraturereviewprogramperformancelargesoftwareAIforengineering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper attempts to give the first comprehensive map of language-model-based code optimization, a field that turns slow or inefficient programs into faster ones while preserving behavior. To do that, the authors ran a systematic literature review, screened and quality-assessed studies from six indexing engines, and classified 53 primary studies across four research questions with 11 sub-questions. If the review's selection is representative, the field looks early and lopsided: most work uses general-purpose models, focuses on one programming language and one performance metric, and evaluates on competitive-programming or synthetic data rather than real-world codebases. The paper argues that five challenges follow from these patterns and proposes eight research directions, such as agentic models, model compression, multi-objective optimization, and human-in-the-loop trust mechanisms.

What carries the argument

The machinery of the paper is the systematic literature review itself: a Kitchenham-and-Charters-style protocol with a quasi-gold-standard search string, snowballing, inclusion/exclusion screening, quality assessment, and 53 retained primary studies. The analytic engine is a taxonomy built from four research questions (RQ1, characteristics of the LMs; RQ2, how LMs were applied; RQ3, how the optimization problem was defined; RQ4, how methods were evaluated) and 11 sub-questions, which lets the authors count and cross-tabulate model types, parameter sizes, training strategies, challenges, techniques, roles, languages, metrics, and datasets into the distributions that carry their findings.

What would settle it

Re-run the search on the six indexing engines and check whether the 53-study set matches the repository's full list; any substantial missing primary study could shift the reported statistics. If a systematic replication adds enough multi-language or real-world-evaluated studies, the 81% and 68% figures would no longer be representative.

Watch

Extended reading notes

Core claim

The paper's central claim is a descriptive one: LM-based code optimization is a fast-growing but immature area whose current practice concentrates in a few comfortable settings. Across 53 studies, general-purpose LMs like GPT-4 are used more often than code-specialized models (61 vs 43 instances), 57% of studies use off-the-shelf models while 43% fine-tune, and the most common technical response to optimization failure is iterative feedback rather than one-shot generation. The review also reports that 81% of studies optimize a single language, 79% focus on a single performance metric, and 68% do not evaluate on real-world programs at all. From these patterns the authors derive five open challenges—balancing model complexity with practicality, interacting with external systems, generalizing across languages and metrics, evaluating on real code, and building trust in model outputs—and outline eight directions for future work.

Load-bearing premise

The survey's numbers and conclusions hold only if the automatic search plus snowballing found all or most relevant studies and the inclusion/exclusion criteria did not bias the set of 53 primary studies.

Editorial extensions

If this is right

  • If the survey's picture is right, the default workflow in this field is a general-purpose LM combined with execution feedback, and new methods should benchmark against that combination.
  • The dominance of single-language and single-metric studies means cross-language and multi-objective optimization are under-explored, so early entrants can claim clear ground.
  • Because 68% of studies avoid real-world code, reported gains on competitive-programming tasks should not be read as production speedups until validated on full projects.
  • The five challenges imply that practical adoption will be gated by trust and integration, not raw model capability, favoring agentic and human-in-the-loop designs.
  • The review's own counts give practitioners a checklist for evaluating a new optimizer: Which model? Which role? Which language? Which metric? Which data?

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A likely consequence the authors leave implicit: the statistics treat each paper as one data point, so a few highly cited studies may inflate some categories; a future survey could weight findings by evaluation scale and replication.
  • The pattern that most systems use competitive-programming data suggests a testable extension: building an optimization benchmark from heterogeneous open-source repositories would probably lower reported speedups and reveal which techniques depend on benchmark style.
  • The eight future directions point toward convergence with the broader agentic software-engineering trend; one concrete prediction is that optimization benchmarks will begin to require tool use, multi-turn iteration, and correctness checks, not just runtime improvement.
  • Another implicit extension is economic: since 57% of studies use off-the-shelf models and only a few train from scratch, the field's progress is tied to commercial API availability, so open-weight models deserve targeted evaluation for optimization tasks.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper presents a systematic literature review (SLR) of language-model-based code optimization. It follows the Kitchenham and Charters guidelines, searches six academic indexing engines supplemented by snowballing, and selects 53 primary studies. The study organizes the field into four research questions with 11 sub-questions, covering LM characteristics, application techniques, problem definition, and evaluation methods. It proposes a taxonomy of models, roles, techniques, languages, metrics, and datasets, and derives five open challenges and eight future research directions. The headline findings include that 81% of studies target a single language, 68% are not evaluated on real-world code, and general-purpose LMs are used more often than code-specialized LMs.

Significance. If the survey's findings are reliable, this is the first comprehensive SLR devoted specifically to LM-based code optimization, filling a genuine gap in the literature. The taxonomy and the synthesis of trends (e.g., the prevalence of feedback-based iterative optimization, the dominance of Python and single-language studies, the gap in real-world evaluation) provide a useful orientation for both researchers and practitioners in a rapidly evolving field. The authors have also made raw results available in a GitHub repository, which supports transparency and allows others to verify the classification. The paper's recommendations, such as the need for standardized real-world benchmarks and multi-objective optimization, are actionable. The contribution is primarily descriptive, but a systematic map of this kind is valuable in a field where the primary literature is growing quickly.

major comments (3)
  1. [Section 3 (Methodology) and footnote 3] The methodology section does not report the exact search string, the names of the six academic indexing engines, the inclusion and exclusion criteria, or the quality-assessment thresholds. Footnote 3 delegates all of these to an external GitHub repository, which is not part of the reviewed manuscript. Since every headline statistic in the abstract and in Sections 6 and 7 (e.g., 81% single-language, 68% non-real-world evaluation) is computed over a set of 53 studies selected through these unreported criteria, the central claim of the survey cannot be fully audited from the paper itself. Please include the complete protocol in the paper or in a stable, versioned appendix, and provide at least a summary of the key criteria in the main text.
  2. [Section 3, Figure 4] The quasi-gold standard step is shown schematically (10 studies) but the paper does not report how the gold standard set was constructed or the resulting precision/recall of the search string. In the quasi-gold standard methodology (Zhang et al., [152]), these validation figures are necessary to demonstrate that the automatic search is sufficiently sensitive. Without them, a reader cannot judge whether the 53-study set is a complete and unbiased representation of the relevant literature.
  3. [Section 3 and Figure 5] The data-extraction and classification process is not described. The paper does not mention pilot extraction, inter-rater agreement, or how conflicts were resolved when assigning studies to taxonomy categories such as those in Tables 2-6. Because the taxonomy and the associated percentages are the main contributions of the survey, the reliability of these manual assignments is load-bearing. A brief description of the validation process for data extraction and classification should be added.
minor comments (6)
  1. [Section 1 and Section 3] The paper states that searches were conducted via 'six academic indexing engines' but never names these engines. Naming them in the methodology would improve transparency.
  2. [Section 7.1.3] The text says 'Compiler datasets were used in six compiler-related tasks,' but Table 7 lists seven instances in the Compiler category. The text and the table should be made consistent.
  3. [Throughout] The repeated markers '♂search' and '/thumbs-up' appear before each 'Finding' and 'Recommendation' paragraph. These appear to be unrendered icon artifacts and should be fixed in the final version.
  4. [Tables 1, 5, Section 7.1.2, and Section 9] There are typographical errors: 'Ealier' in Table 1, 'Languague' and 'Heuristsic' in Table 5, 'Genereal' in Section 7.1.2, and 'cataloger' in the conclusion. These should be corrected.
  5. [Section 2] Footnote 2 states that the full related works section is available only in the external repository. For a survey, it would be preferable to include an expanded related-work summary in the paper itself or at least briefly describe the main categories of related work beyond the short background given.
  6. [Abstract and Section 1] The abstract says 'over 50 primary studies' while Section 1 states '53 primary studies.' Consider using the exact number in the abstract for consistency, or explicitly say '53' if the final count is fixed.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the survey's findings are descriptive syntheses of external primary studies, and its own background self-citations are not load-bearing.

full rationale

This paper is a systematic literature review, not a derivation of predictions from fitted parameters. Its central outputs—the taxonomy, the 11 sub-question findings, and the headline percentages (e.g., 81% single-language, 68% non-real-world evaluation)—are descriptive counts over the 53 selected primary studies. Those studies are external works, and the survey does not define any of its categories in terms of its own conclusions, nor does it fit a parameter and then rename that fit as a finding. The authors do cite their own prior work in several places: reference [26] (DeepTune, co-authored by Zheng Wang), [43] (Artemis++, co-authored by Giavrimis and Basios), [44] and [45] (Gong and Chen), [111] (Rocha, Petoumenos, Wang, Cole, Leather), and [137] (Wang and O'Boyle). However, each of these appears as background context or as one cited example among many, and none is invoked as the evidence establishing the survey's novel findings or its statistics. The survey's methodology section does defer the full search string, inclusion/exclusion criteria, and quality-assessment details to an external GitHub repository, which is an auditability and reproducibility concern rather than a circularity concern: the selection procedure is an input to the statistics, but nothing in the paper equates the selection procedure with the survey's conclusions by construction. The five challenges and eight future directions are recommendations synthesized from the reviewed literature, not deductions that reduce to the survey's own inputs. Therefore, no specific circular step can be quoted or exhibited, and the honest finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

This is a literature review, so there are no fitted free parameters and no new postulated entities. The central claims rest on two methodological domain assumptions: the selected 53 studies represent the full relevant literature, and the author-defined taxonomy accurately reflects the content of those studies.

assumptions (2)
  • domain assumption The set of 53 primary studies selected through automatic search, snowballing, inclusion/exclusion criteria, and quality assessment is representative of the full literature on LM-based code optimization.
    The survey's statistical findings (e.g., 81% single-language studies, 68% non-real-world evaluation) only generalize if the selection is complete and unbiased. The full protocol is in the external repository (Section 3, footnote 3).
  • domain assumption The author-defined taxonomy for categorizing LMs, challenges, techniques, roles, languages, metrics, and datasets accurately reflects the content of the primary studies.
    Categorizations are qualitative judgments made by the authors; no inter-rater reliability or validation is reported, so misclassification could affect the findings in Sections 4-7.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Language Models for Code Optimization: Survey, Challenges and Future Directions." pith.science (2026). https://pith.science/paper/J3W6AFHA

@misc{pith2026250101277,
  author       = {Pith},
  title        = {Pith review of: Language Models for Code Optimization: Survey, Challenges and Future Directions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/J3W6AFHA}},
  note         = {Machine review of arXiv:2501.01277}
}
read the original abstract

Language models (LMs) built upon deep neural networks (DNNs) have recently demonstrated breakthrough effectiveness in software engineering tasks such as code generation, completion, and repair. This has paved the way for the emergence of LM-based code optimization techniques, which are crucial for enhancing the performance of existing programs, such as accelerating program execution time. However, a comprehensive survey dedicated to this specific application has been lacking. To fill this gap, we present a systematic literature review of over 50 primary studies, identifying emerging trends and addressing 11 specialized questions. Our findings reveal five critical open challenges, such as balancing model complexity with practical usability, cross-language/performance generalizability, and building trust in AI-driven solutions. Furthermore, we provide eight future research directions to facilitate more efficient, robust, and reliable LM-based code optimization. Thereby, this study aims to provide actionable insights and foundational references for both researchers and practitioners in this rapidly evolving field.

Figures

Figures reproduced from arXiv: 2501.01277 by the authors.

Figure 1
Figure 1. Visualization of the survey scope. • General-purpose LMs like GPT-4 were more widely adopted (61 instances) than code-specialized LMs (43 instances) due to their broader understanding and reasoning capabilities. • A majority of studies (57%) leveraged pre-trained models to save time and resources, while 43% employed fine-tuning to tailor the models for task-specific needs. • The most commonly highlighted challenges … view at source ↗
Figure 2
Figure 2. Two Python implementations for calculating the sum of the first [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Development of code optimization methods: strengths and weaknesses [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Overview of the survey methodology used in this study. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Overview of the taxonomy for all RQs (one study might be in multiple categories). [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Distribution of parameter sizes (one study might be in multiple categories). Leveraging off-the-shelf LMs (30) 57% Pre-training & fine-tuning (23) 43% [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 8
Figure 8. Figure 8: Distribution of # optimized program￾ming languages. One (42) 79% Two (9) 17% Three (2) 4% [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]
Figure 10
Figure 10. Figure 10: Distribution of evaluation us￾ing real-world code. 5 10 15 20 25 22 (%PI) 17 (PI) 12 (Speedup) 10 (%OPT) 2 (AOCC) 1 (IOCCB) 1 (NPI) # studies Performance gain Task-specific Self-proposed [PITH_FULL_IMAGE:figures/full_fig_p024_10.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. JETO-Bench: A Reproducible Benchmark for Execution Time Improvement Patches in Java

    cs.SE 2026-06 conditional novelty 7.0 of 10

    JETO-Mine is a reusable three-phase pipeline that mines 1.8 million Java commits to produce JETO-Bench containing 91 verified executable ETIPs, on which OpenHands succeeds at 14.3%.

  2. Rethinking LLM-Based RTL Code Optimization Via Timing Logic Metamorphosis

    cs.SE 2025-07 reject novelty 6.0 of 10

    LLM-based RTL optimizers degrade on timing-heavy mutants, but the study's own data and methods do not fully support the headline claim.

  3. A Comprehensive Survey on Integrating Large Language Models with Knowledge-Based Methods

    cs.CL 2025-01 conditional novelty 3.0 of 10

    A narrative review of LLM knowledge integration that categorizes techniques and compiles benchmarks, but lacks a systematic method and contains unreliable citations.

  4. Enhancing Trust in Language Model-Based Code Optimization through RLHF: A Research Design

    cs.SE 2025-02 unverdicted novelty 2.0 of 10

    A research design proposes applying reinforcement learning from human feedback to improve trust in language model based code optimization, with no experiments yet.

Reference graph

Works this paper leans on

163 extracted references · 29 canonical work pages · cited by 4 Pith papers

  1. [152]

    He Zhang, Muhammad Ali Babar, and Paolo Tell. 2011. Identifying Relevant Studies in Software Engineering. Information and Software Technology 53 (2011), 625–637

  2. [1]

    Abella-González, Pedro Carollo-Fernández, Louis-Noël Pouchet, Fabrice Rastello, and Gabriel Rodríguez

    Miguel Á. Abella-González, Pedro Carollo-Fernández, Louis-Noël Pouchet, Fabrice Rastello, and Gabriel Rodríguez

  3. [2]

    Felix Adler, Gordon Fraser, Eva Gründinger, Nina Körber, Simon Labrenz, Jonas Lerchenberger, Stephan Lukasczyk, and Sebastian Schweikl. 2021. Improving Readability of Scratch Programs with Search-based Refactoring. In International Working Conference on Source Code Analysis and Manipulation, SCAM . IEEE, 120–130

  4. [3]

    Randy Allen and Steve Johnson. 1988. Compiling C for Vectorization, Parallelization, and Inline Expansion. ACM SIGPLAN Notices 23 (1988), 241–249

  5. [4]

    Rohan Anil, Sebastian Borgeaud, and et al. 2023. Gemini: A Family of Highly Capable Multimodal Models. (2023). arXiv:2312.11805

  6. [5]

    Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. PaLM 2 Technical Report. arXiv:2305.10403 (2023)

  7. [6]

    Anthropic. 2024. Introducing the next generation of Claude

  8. [7]

    Amir H Ashouri, William Killian, John Cavazos, Gianluca Palermo, and Cristina Silvano. 2018. A Survey on Compiler Autotuning using Machine Learning. Computing Surveys (CSUR) 51 (2018), 1–42

Show all 163 references
  1. [8]

    Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J

    Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. 2021. Program Synthesis with Large Language Models. (2021). arXiv:2108.07732

  2. [9]

    John Backus. 1978. The history of Fortran I, II, and III. ACM Sigplan Notices 13, 8 (1978), 165–180

  3. [10]

    Riyadh Baghdadi, Massinissa Merouani, Mohamed-Hicham Leghettas, Kamel Abdous, Taha Arbaoui, Karima Be- natchba, et al. 2021. A Deep Learning Based Cost Model for Automatic Code Optimization. Proceedings of Machine Learning and Systems 3 (2021), 181–193. ACM Comput. Surv., Vol....

  4. [11]

    M Ammar Ben Khadra, Dominik Stoffel, and Wolfgang Kunz. 2020. Efficient Binary-Level Coverage Analysis. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 1153–1164

  5. [12]

    K Manasvi Bhat, Pratiksha P Anchalia, Rushali Mohbe, and A Parkavi. 2019. A Survey of Machine Learning and Deep Learning Techniques for Compiler Optimization. International Journal of Research in Engineering, Science and Management (2019)

  6. [13]

    BigScience. 2022. BLOOM: A 176B-Parameter Open-Access Multilingual Language Model. (2022). arXiv:2211.05100

  7. [14]

    Nathan Binkert, Bradford Beckmann, Gabriel Black, Steven K Reinhardt, Ali Saidi, Arkaprava Basu, Joel Hestness, Derek R Hower, Tushar Krishna, Somayeh Sardashti, et al. 2011. The gem5 Simulator. ACM SIGARCH computer architecture news 39 (2011), 1–7

  8. [15]

    Sid Black, Stella Biderman, Eric Hallahan, and et al. 2022. GPT-NeoX-20B: An Open-Source Autoregressive Language Model. arXiv:2204.06745

  9. [16]

    Tom B Brown. 2020. Language Models are Few-Shot Learners. NeurIPS (2020)

  10. [17]

    Rosario Cammarota, Alexandru Nicolau, Alexander V Veidenbaum, Arun Kejariwal, Debora Donato, and Mukund Madhugiri. 2013. On the Determination of Inlining Vectors for Program Optimization. In Compiler Construction (CC). Springer, 164–183

  11. [18]

    John Cavazos, Christophe Dubach, Felix Agakov, Edwin Bonilla, Michael FP O’Boyle, Grigori Fursin, and Olivier Temam. 2006. Automatic Performance Model Construction for the Fast Software Exploration of New Hardware Designs. In Proceedings of the 2006 international conference on...

  12. [19]

    Gregory J Chaitin. 1982. Register Allocation & Spilling via Graph Coloring. ACM Sigplan Notices 17 (1982), 98–101

  13. [20]

    Lihu Chen and Gaël Varoquaux. 2024. What is the Role of Small Models in the LLM Era: A Survey. arXiv:2409.06857 (2024)

  14. [22]

    2018.{TVM}: An automated{End-to-End} optimizing compiler for deep learning

    Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, et al. 2018.{TVM}: An automated{End-to-End} optimizing compiler for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implem...

  15. [23]

    Zimin Chen, Sen Fang, and Martin Monperrus. 2024. Supersonic: Learning to Generate Source Code Optimizations in C/C++. IEEE Transactions on Software Engineering (TSE) (2024)

  16. [24]

    Jinsu Choi, Gabin An, and Shin Yoo. 2024. Iterative Refactoring of Real-World Open-Source Programs with Large Language Models. In International Symposium on Search Based Software Engineering . Springer, 49–55

  17. [25]

    Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy, and Neel Sundaresan

    Colin B. Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy, and Neel Sundaresan. 2020. PyMT5: Multi- Mode Translation of Natural Language and Python Code with Transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMN...

  18. [26]

    Chris Cummins, Pavlos Petoumenos, Zheng Wang, and Hugh Leather. 2017. End-to-end deep learning of optimization heuristics. In 2017 26th International Conference on Parallel Architectures and Compilation Techniques (PACT) . IEEE, 219–232

  19. [27]

    Hazelwood, Gabriel Synnaeve, and Hugh Leather

    Chris Cummins, Volker Seeker, Dejan Grubisic, Mostafa Elhoushi, Youwei Liang, Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Kim M. Hazelwood, Gabriel Synnaeve, and Hugh Leather. 2023. Large Language Models for Compiler Optimization. (2023). arXiv:2309.07062

  20. [28]

    Chris Cummins, Volker Seeker, Dejan Grubisic, Baptiste Rozière, Jonas Gehring, Gabriel Synnaeve, and Hugh Leather

  21. [29]

    Matthew Curtis-Maury, James Dzierwa, Christos D Antonopoulos, and Dimitrios S Nikolopoulos. 2006. Online Power-Performance Adaptation of Multithreaded Programs Using Hardware Event-Based Prediction. In Proceedings of the 20th annual international conference on Supercomputing . 157–166

  22. [30]

    Anderson Faustino da Silva, Bruno Conde Kind, José Wesley de Souza Magalhães, and et al. 2021. ANGHABENCH: A Suite with One Million Compilable C Benchmarks for Code-Size Reduction. In IEEE/ACM International Symposium on Code Generation and Optimization, CGO . IEEE, 378–390

  23. [31]

    Deepseek-AI. 2023. DeepSeek Coder: Let the Code Write Itself. https://deepseekcoder.github.io/

  24. [32]

    Ahmed, Guixiang Ma, Mihai Capota, Theodore L

    Shukai Duan, Nikos Kanakaris, Xiongye Xiao, Heng Ping, Chenyu Zhou, Nesreen K. Ahmed, Guixiang Ma, Mihai Capota, Theodore L. Willke, Shahin Nazarian, and Paul Bogdan. 2023. Leveraging Reinforcement Learning and Large Language Models for Code Optimization. (2023). arXiv:2312.05657

  25. [33]

    Rudolf Eigenmann and Jay Hoeflinger. 2000. Parallelizing and Vectorizing Compilers. Proc. IEEE (2000)

  26. [34]

    Angela Fan, Beliz Gokkaya, Mark Harman, Mitya Lyubarskiy, Shubho Sengupta, Shin Yoo, and Jie M. Zhang. 2023. Large Language Models for Software Engineering: Survey and Open Problems. InInternational Conference on Software Engineering: Future of Software Engineering, ICSE-FoSE ...

  27. [35]

    Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Findings of the Association for Computational Linguistics: EMNLP ...

  28. [36]

    Mary F Fernandez. 1995. Simple and Effective Link-Time Optimization of Modula-3 Programs. In Proceedings of the ACM SIGPLAN 1995 conference on Programming language design and implementation . 103–115

  29. [37]

    Grigori Fursin, Yuriy Kashnikov, Abdul Wahid Memon, Zbigniew Chamski, Olivier Temam, Mircea Namolaru, Elad Yom-Tov, Bilha Mendelson, Ayal Zaks, Eric Courtois, et al. 2011. Milepost gcc: Machine learning enabled self-tuning compiler. International journal of parallel programmin...

  30. [38]

    Shuzheng Gao, Cuiyun Gao, Wenchao Gu, and Michael Lyu. 2024. Search-Based LLMs for Code Optimization. In International Conference on Software Engineering (ICSE) . IEEE, 254–266

  31. [39]

    Clement, Neel Sundaresan, and Chen Wu

    Spandan Garg, Roshanak Zilouchian Moghaddam, Colin B. Clement, Neel Sundaresan, and Chen Wu. 2022. DeepDev- PERF: A Deep Learning-Based Approach for Improving Software Performance. In ESEC/FSE. ACM, 948–958

  32. [40]

    Spandan Garg, Roshanak Zilouchian Moghaddam, and Neel Sundaresan. 2023. RAPGen: An Approach for Fixing Code Inefficiencies in Zero-Shot. (2023). arXiv:2306.17077

  33. [41]

    Leonidas Gee, Milan Gritta, Gerasimos Lampouras, and Ignacio Iacobacci. 2024. Code-Optimise: Self-Generated Preference Data for Correctness and Efficiency. (2024). arXiv:2406.12502

  34. [42]

    Sia Gholami. 2024. Can Pruning Make Large Language Models More Efficient? In Redefining Security With Cyber AI . IGI Global, 1–14

  35. [43]

    Rafail Giavrimis, Alexis Butler, Constantin Cezar Petrescu, Michail Basios, and Santanu Kumar Dash. 2021. Genetic Optimisation of C++ Applications. In International Conference on Automated Software Engineering, ASE . IEEE, 1180– 1182

  36. [44]

    Jingzhi Gong and Tao Chen. 2024. Deep Configuration Performance Learning: A Systematic Survey and Taxonomy. ACM Transactions on Software Engineering and Methodology (TOSEM) (2024)

  37. [45]

    Jingzhi Gong and Tao Chen. 2024. Predicting Configuration Performance in Multiple Environments with Sequential Meta-Learning. Proceedings of the ACM on Software Engineering FSE (2024), 359–382

  38. [46]

    Google. 2023. Google Cloud launches new AI models. https://cloud.google.com/blog/products/ai-machine-learning/ google-cloud-launches-new-ai-models-opens-generative-ai-studio

  39. [47]

    Dejan Grubisic, Chris Cummins, Volker Seeker, and Hugh Leather. 2024. Compiler Generated Feedback for Large Language Models. (2024). arXiv:2403.14714

  40. [48]

    Mellor-Crummey, and Chris Cummins

    Dejan Grubisic, Volker Seeker, Gabriel Synnaeve, Hugh Leather, John M. Mellor-Crummey, and Chris Cummins. 2024. Priority Sampling of Large Language Models for Compilers. In Proceedings of the 4th Workshop on Machine Learning and Systems, EuroMLSys. ACM, 91–97

  41. [49]

    Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y. Wu, Y. K. Li, Fuli Luo, Yingfei Xiong, and Wenfeng Liang. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence. (2024). arXiv:2401.14196

  42. [50]

    Zifan Carl Guo and William S. Moses. 2022. Enabling Transformers to Understand Low-Level Programs. In High Performance Extreme Computing Conference (HPEC) . IEEE, 1–9

  43. [51]

    Priyanshu Gupta, Avishree Khare, Yasharth Bajpai, Saikat Chakraborty, Sumit Gulwani, Aditya Kanade, Arjun Radhakrishna, Gustavo Soares, and Ashish Tiwari. 2023. Grace: Language Models Meet Code Edits. In Joint European Software Engineering Conference and Symposium on the Found...

  44. [52]

    Bing Han, Congfei Li, Hua Deng, Guowei Liu, and Ze Zheng. 2024. Domain-Specific Translation Tool from Structured Text to C Source Code with Code Readability Enhancement in Programmable Logic Controllers. Concurrency and Computation: Practice and Experience (2024), e8100

  45. [53]

    Xu Han, Qiannan Yang, Xianda Chen, Xiaowen Chu, and Meixin Zhu. 2024. Generating and Evolving Reward Functions for Highway Driving with Large Language Models. arXiv:2406.10540 (2024)

  46. [54]

    Erik Hemberg, Stephen Moskal, and Una-May O’Reilly. 2024. Evolving Code with a Large Language Model. Genetic Programming and Evolvable Machines 25 (2024), 21

  47. [55]

    Dan Hendrycks, Steven Basart, Saurav Kadavath, and et. al. 2021. Measuring Coding Challenge Competence With APPS. NeurIPS (2021)

  48. [56]

    Charles Hong, Sahil Bhatia, Altan Haan, Shengjun Kris Dong, Dima Nikiforov, Alvin Cheung, and Yakun Sophia Shao. 2024. LLM-Aided Compilation for Tensor Accelerators. (2024), 1–14

  49. [57]

    Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. 2023. Large Language Models for Software Engineering: A Systematic Literature Review. ACM Transactions on Software Engineering and Methodology (2023)

  50. [58]

    Dong Huang, Jianbo Dai, Han Weng, Puzhen Wu, QING Yuhao, Heming Cui, Zhijiang Guo, and Jie Zhang. 2024. EffiLearner: Enhancing Efficiency of Generated Code via Self-Optimization. (2024). ACM Comput. Surv., Vol. 1, No. 1, Article . Publication date: January 2024. Language Model...

  51. [59]

    Dong Huang, Guangtao Zeng, Jianbo Dai, Meng Luo, Han Weng, Yuhao Qing, Heming Cui, Zhijiang Guo, and Jie M Zhang. 2024. Effi-Code: Unleashing Code Efficiency in Language Models. arXiv:2410.10209 (2024)

  52. [60]

    Zhang, Yuhao Qing, and Heming Cui

    Dong Huang, Jie M. Zhang, Yuhao Qing, and Heming Cui. 2024. EffiBench: Benchmarking the Efficiency of Automati- cally Generated Code. (2024). arXiv:2402.02037

  53. [61]

    Binyuan Hui, Jian Yang, Zeyu Cui, Jiaxi Yang, Dayiheng Liu, Lei Zhang, Tianyu Liu, Jiajun Zhang, Bowen Yu, Keming Lu, et al. 2024. Qwen2.5-coder Technical Report. arXiv:2409.12186 (2024)

  54. [62]

    Shu Ishida, Gianluca Corrado, George Fedoseev, Hudson Yeo, Lloyd Russell, Jamie Shotton, Joao F Henriques, and Anthony Hu. 2024. LangProp: A Code Optimization Framework Using Large Language Models Applied to Driving

  55. [63]

    Naman Jain, Skanda Vaidyanath, Arun Iyer, Nagarajan Natarajan, Suresh Parthasarathy, Sriram Rajamani, and Rahul Sharma. 2022. Jigsaw: Large Language Models Meet Program Synthesis. In International Conference on Software Engineering (ICSE). 1219–1231

  56. [64]

    Mojan Javaheripi, Sébastien Bubeck, Marah Abdin, Jyoti Aneja, Sebastien Bubeck, Caio César Teodoro Mendes, Weizhu Chen, Allie Del Giorno, Ronen Eldan, Sivakanth Gopi, et al . 2023. Phi-2: The Surprising Power of Small Language Models. Microsoft Research Blog 1 (2023), 3

  57. [65]

    Matthew Jin, Syed Shahriar, Michele Tufano, Xin Shi, Shuai Lu, Neel Sundaresan, and Alexey Svyatkovskiy. 2023. InferFix: End-to-End Program Repair with LLMs. In Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, ESEC/FSE . ...

  58. [66]

    René Just, Darioush Jalali, and Michael D Ernst. 2014. Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs. In International Symposium on Software Testing and Analysis (ISSTA) . 437–440

  59. [67]

    Sungmin Kang and Shin Yoo. 2023. Towards Objective-Tailored Genetic Improvement Through Large Language Models. In International Workshop on Genetic Improvement, GI@ICSE 2023 . IEEE, 19–20

  60. [68]

    Ryan Kastner, Janarbek Matai, and Stephen Neuendorffer. 2018. Parallel programming for FPGAs. arXiv preprint arXiv:1805.03648 (2018)

  61. [69]

    Charters

    Barbara Kitchenham and Stuart M. Charters. 2007. Guidelines for Performing Systematic Literature Reviews in Software Engineering

  62. [70]

    Yuhang Lai, Chengxi Li, and Yiming et al. Wang. 2023. DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation. In International Conference on Machine Learning . 18319–18345

  63. [71]

    Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven Chu-Hong Hoi. 2022. CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning. In Neural Information Processing Systems (NeurIPS)

  64. [72]

    Jia Li, Ge Li, Zhuo Li, Zhi Jin, Xing Hu, Kechi Zhang, and Zhiyi Fu. 2023. CodeEditor: Learning to Edit Source Code with Pre-trained Models. ACM Transactions on Software Engineering and Methodology (TOSEM) 32 (2023), 1–22

  65. [73]

    Kaixin Li, Qisheng Hu, Xu Zhao, Hui Chen, Yuxi Xie, Tiedong Liu, Qizhe Xie, and Junxian He. 2024. InstructCoder: Instruction Tuning Large Language Models for Code Editing

  66. [74]

    Raymond Li, Loubna Ben Allal, and et al. 2023. StarCoder: May the Source Be With You! Transactions on Machine Learning Research 2023 (2023)

  67. [75]

    Rui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev, and P Sadayappan. 2021. Analytical Characterization and Design Space Exploration for Optimization of CNNs. In Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and ...

  68. [76]

    Yujia Li, David Choi, Junyoung Chung, and et al. 2022. Competition-Level Code Generation with AlphaCode. Science 378 (2022), 1092–1097

  69. [77]

    Zeyuan Li, Yangfan He, Lewei He, Jianhui Wang, Tianyu Shi, Bin Lei, Yuchen Li, and Qiuwu Chen. 2024. FALCON: Feedback-driven Adaptive Long/short-term memory reinforced Coding Optimization system.arXiv:2410.21349 (2024)

  70. [78]

    Nandor Licker and Timothy M Jones. 2020. Duplo: A Framework for OCaml Post-Link Optimisation. Proceedings of the ACM on Programming Languages 4, ICFP (2020), 1–29

  71. [79]

    Junwei Liu, Kaixin Wang, Yixuan Chen, Xin Peng, Zhenpeng Chen, Lingming Zhang, and Yiling Lou. 2024. Large Language Model-Based Agents for Software Engineering: A Survey. arXiv:2409.02977 (2024)

  72. [80]

    Edward S Lowry and Cleburne W Medlock. 1969. Object code optimization. Commun. ACM 12, 1 (1969), 13–22

  73. [81]

    Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, et al. 2024. StarCoder 2 and The Stack v2: The Next Generation. arXiv:2402.19173 (2024)

  74. [82]

    Jinliang Lu, Ziliang Pang, Min Xiao, Yaochen Zhu, Rui Xia, and Jiajun Zhang. 2024. Merge, Ensemble, and Cooperate! A Survey on Collaborative Strategies in the Era of Large Language Models. arXiv:2407.06089 (2024)

  75. [83]

    Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2024. WizardCoder: Empowering Code Large Language Models with Evol-Instruct. In The Twelfth International Conference on Learning Representations, ICLR 2024...

  76. [84]

    Aman Madaan, Niket Tandon, Prakhar Gupta, and et al. 2023. Self-Refine: Iterative Refinement with Self-Feedback. In Neural Information Processing Systems (NeurIPS)

  77. [85]

    Saeed Maleki, Yaoqing Gao, Maria J Garzar, Tommy Wong, David A Padua, et al. 2011. An Evaluation of Vectorizing Compilers. In 2011 International Conference on Parallel Architectures and Compilation Techniques . IEEE, 372–382

  78. [86]

    William M McKeeman. 1965. Peephole Optimization. Commun. ACM 8 (1965), 443–444

  79. [87]

    Thomas Mesnard, Cassidy Hardin, and et al. 2024. Gemma: Open Models Based on Gemini Research and Technology. (2024). arXiv:2403.08295

  80. [88]

    Meta. 2024. Introducing Llama 3.1: Our most capable models to date. https://ai.meta.com/blog/meta-llama-3-1/

  81. [89]

    Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024. Large Language Models: A Survey. arXiv:2402.06196 (2024)

  82. [90]

    Bolin Ni, Jingcheng Hu, Yixuan Wei, Houwen Peng, Zheng Zhang, Gaofeng Meng, and Han Hu. 2024. Xwin-LM: Strong and Scalable Alignment Practice for LLMs. (2024). arXiv:2405.20335

  83. [91]

    Daniel Nichols, Pranav Polasam, Harshitha Menon, Aniruddha Marathe, Todd Gamblin, and Abhinav Bhatele. 2024. Performance-Aligned LLMs for Generating Fast Code. (2024). arXiv:2404.18864

  84. [92]

    Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong

  85. [93]

    Behrooz Omidvar Tehrani and Anmol Anubhai. 2024. Evaluating Human-AI Partnership for LLM-based Code Migration. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems . 1–8

  86. [94]

    OpenAI. 2021. Evaluating Large Language Models Trained on Code. (2021). arXiv:2107.03374

  87. [95]

    OpenAI. 2022. Introducing ChatGPT. https://openai.com/index/chatgpt/

  88. [96]

    OpenAI. 2024. GPT-4 Technical Report. arXiv:2303.08774

  89. [97]

    OpenAI. 2024. Hello GPT-4o. https://openai.com/index/hello-gpt-4o/

  90. [98]

    Marek Palkowski and Mateusz Gruzewski. 2024. GPT-Driven Source-to-Source Transformation for Generating Compilable Parallel CUDA Code for Nussinov’s Algorithm. Electronics 13 (2024), 488

  91. [99]

    Yue Pan and Chen Lyu. 2023. Measuring Efficient Code Generation with GEC. In Proceedings of the 14th Asia-Pacific Symposium on Internetware, Internetware . ACM, 249–258

  92. [100]

    Yue Pan, Chen Lyu, Zhenyu Yang, Lantian Li, Qi Liu, and Xiuting Shao. 2024. E-code: Mastering Efficient Code Generation through Pretrained Models and Expert Encoder Group

  93. [101]

    Yue Pan, Xiuting Shao, and Chen Lyu. 2025. Measuring Code Efficiency Optimization Capabilities with ACEOB

  94. [102]

    Eunjung Park, John Cavazos, Louis-Noël Pouchet, Cédric Bastoul, Albert Cohen, and Ponnuswamy Sadayappan. 2013. Predictive Modeling in a Polyhedral Optimization Space. International journal of parallel programming 41 (2013), 704–750

  95. [103]

    Rajvardhan Patil and Venkat Gudivada. 2024. A Review of Current Trends, Techniques, and Challenges in Large Language Models (LLMs). Applied Sciences 14 (2024), 2074

  96. [104]

    Huiyun Peng, Arjun Gupte, Nicholas John Eliopoulos, Chien Chou Ho, Rishi Mantri, Leo Deng, Wenxin Jiang, Yung-Hsiang Lu, Konstantin Läufer, George K Thiruvathukal, et al. 2024. Large Language Models for Energy-Efficient Code: Emerging Results and Future Directions. arXiv:2410....

  97. [105]

    Yun Peng, Akhilesh Deepak Gotmare, Michael Lyu, Caiming Xiong, Silvio Savarese, and Doyen Sahoo. 2024. Perf- CodeGen: Improving Performance of LLM Generated Code with Execution Feedback. arXiv:2412.03578 (2024)

  98. [106]

    Rui Pereira, Marco Couto, Francisco Ribeiro, Rui Rua, Jácome Cunha, João Paulo Fernandes, and João Saraiva. 2017. Energy Efficiency Across Programming Languages: How Do Energy, Time, and Memory Relate?. In ACM SIGPLAN international conference on software language engineering . 256–267

  99. [107]

    Kung, Geert Janssen, Wei Zhang, Giacomo Domeniconi, Vladimir Zolotov, Julian Dolby, Jie Chen, Mihir R

    Ruchir Puri, David S. Kung, Geert Janssen, Wei Zhang, Giacomo Domeniconi, Vladimir Zolotov, Julian Dolby, Jie Chen, Mihir R. Choudhury, Lindsey Decker, Veronika Thost, Luca Buratti, Saurabh Pujar, and Ulrich Finkler. 2021. Project CodeNet: A Large-Scale AI for Code Dataset for...

  100. [108]

    Muzi Qu, Jie Liu, Liangyi Kang, Shuai Wang, Dan Ye, and Tao Huang. 2024. Dynamic Scoring Code Token Tree: A Novel Decoding Strategy for Generating High-Performance Code. InIEEE/ACM International Conference on Automated Software Engineering. 1308–1318

  101. [109]

    Sebastian Rachuj, Dietmar Fey, and Marc Reichenbach. 2020. Impact of Performance Estimation on Fast Processor Simulators. In Simulation Tools and Techniques: 12th EAI International Conference, SIMUtools . Springer, 79–93

  102. [110]

    Tal Ridnik, Dedy Kredo, and Itamar Friedman. 2024. Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering. (2024). arXiv:2401.08500

  103. [111]

    Rodrigo CO Rocha, Pavlos Petoumenos, Zheng Wang, Murray Cole, and Hugh Leather. 2019. Function merging by sequence alignment. In 2019 IEEE/ACM International Symposium on Code Generation and Optimization (CGO) . IEEE, 149–163. ACM Comput. Surv., Vol. 1, No. 1, Article . Publica...

  104. [112]

    Pawan Kumar, Emilien Dupont, Francisco J

    Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M. Pawan Kumar, Emilien Dupont, Francisco J. R. Ruiz, Jordan S. Ellenberg, Pengming Wang, Omar Fawzi, Pushmeet Kohli, and Alhussein Fawzi

  105. [113]

    Miguel Romero Rosas, Miguel Torres Sanchez, and Rudolf Eigenmann. 2024. Should AI Optimize Your Code? A Com- parative Study of Current Large Language Models Versus Classical Optimizing Compilers. (2024). arXiv:2406.12146

  106. [114]

    Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, and et al. 2023. Code Llama: Open Foundation Models for Code. (2023). arXiv:2308.12950

  107. [115]

    Baptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, and Guillaume Lample. 2020. Unsupervised Translation of Programming Languages. In Annual Conference on Neural Information Processing Systems, NeurIPS

  108. [116]

    Nature 625 (2024), 468–475

    Mathematical Discoveries from Program Search with Large Language Models. Nature 625 (2024), 468–475

  109. [117]

    Atsushi Shirafuji, Yusuke Oda, Jun Suzuki, Makoto Morishita, and Yutaka Watanobe. 2023. Refactoring Programs Using Large Language Models with Few-Shot Examples. In Asia-Pacific Software Engineering Conference, APSEC. IEEE, 151–160

  110. [118]

    Karthik Shivashankar and Antonio Martini. 2024. Better Python Programming for all: With the focus on Maintainability. arXiv:2408.09134

  111. [119]

    Gardner, Yiming Yang, Milad Hashemi, Graham Neubig, Parthasarathy Ranganathan, Osbert Bastani, and Amir Yazdanbakhsh

    Alexander Shypula, Aman Madaan, Yimeng Zeng, Uri Alon, Jacob R. Gardner, Yiming Yang, Milad Hashemi, Graham Neubig, Parthasarathy Ranganathan, Osbert Bastani, and Amir Yazdanbakhsh. 2024. Learning Performance-Improving Code Edits. In The Twelfth International Conference on Lea...

  112. [120]

    Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language Agents with Verbal Reinforcement Learning. In Neural Information Processing Systems (NeurIPS)

  113. [121]

    Jagriti Sikka, Kushal Satya, Yaman Kumar, Shagun Uppal, Rajiv Ratn Shah, and Roger Zimmermann. 2020. Learning- Based Methods for Code Runtime Complexity Prediction. In Advances in Information Retrieval: 42nd European Conference on IR Research, ECIR . Springer, 313–325

  114. [122]

    Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter. 2024. A Simple and Effective Pruning Approach for Large Language Models. ICLR (2024)

  115. [123]

    Yiwen Sun, Xianyin Zhang, Shiyu Huang, Shaowei Cai, BingZhen Zhang, and Ke Wei. 2024. AutoSAT: Automatically Optimize SAT Solvers via Large Language Models. arXiv:2402.10705 (2024)

  116. [124]

    Schwartz, and Graham Neubig

    Alex Shypula, Pengcheng Yin, Jeremy Lacomis, Claire Le Goues, Edward J. Schwartz, and Graham Neubig. 2022. Learning to Superoptimize Real-world Programs. (2022)

  117. [125]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. LLaMA: Open and Efficient Foundation La...

  118. [126]

    Hugo Touvron, Louis Martin, and et al. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. (2023). arXiv:2307.09288

  119. [127]

    Mircea Trofin, Yundi Qian, Eugene Brevdo, Zinan Lin, Krzysztof Choromanski, and David Li. 2021. Mlgo: a machine learning guided compiler optimizations framework. arXiv preprint arXiv:2101.04808 (2021)

  120. [128]

    Jubi Taneja, Avery Laird, Cong Yan, Madan Musuvathi, and Shuvendu K. Lahiri. 2024. LLM-Vectorizer: LLM-based Verified Loop Vectorizer. (2024). arXiv:2406.04693

  121. [129]

    Niki van Stein and Thomas Bäck. 2024. LLaMEA: A Large Language Model Evolutionary Algorithm for Automatically Generating Metaheuristics. arXiv:2405.20132 (2024)

  122. [130]

    Niki van Stein, Diederick Vermetten, and Thomas Bäck. 2024. In-the-loop Hyper-Parameter Optimization for LLM-Based Automated Design of Heuristics. arXiv:2410.16309 (2024)

  123. [131]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems . ...

  124. [132]

    Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Well-Read Students Learn Better: On the Importance of Pre-training Compact Models. arXiv:1908.08962

  125. [133]

    Siddhant Waghjale, Vishruth Veerendranath, Zhiruo Wang, and Daniel Fried. 2024. ECCO: Can We Improve Model- Generated Code Efficiency Without Sacrificing Functional Correctness?. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP ....

  126. [134]

    Simin Wang, Liguo Huang, Amiao Gao, Jidong Ge, Tengfei Zhang, Haitao Feng, Ishna Satyarth, Ming Li, He Zhang, and Vincent Ng. 2023. Machine/Deep Learning for Software Engineering: A Systematic Literature Review. IEEE Trans. Software Eng. 49, 3 (2023), 1188–1231

  127. [135]

    Joty, and Steven C

    Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi. 2021. CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP . Asso...

  128. [136]

    Michael J Voss and Rudolf Eigemann. 2001. High-Level Adaptive Program Optimization with ADAPT. In Proceedings of the eighth ACM SIGPLAN symposium on Principles and practices of parallel programming . 93–102

  129. [137]

    Zheng Wang and Michael O’Boyle. 2018. Machine Learning in Compiler Optimization. Proc. IEEE 106 (2018), 1879–1901

  130. [138]

    Anjiang Wei, Allen Nie, Thiago SFX Teixeira, Rohan Yadav, Wonchan Lee, Ke Wang, and Alex Aiken. 2024. Improving Parallel Program Performance Through DSL-Driven Code Generation with LLM Optimizers. arXiv:2410.15625 (2024)

  131. [139]

    Michael Joseph Wolfe. 1982. Optimizing Supercompilers for Supercomputers . University of Illinois at Urbana- Champaign

  132. [140]

    Zhichao Wang, Bin Bi, Shiva Kumar Pentyala, Kiran Ramnath, Sougata Chaudhuri, Shubham Mehrotra, Xiang-Bo Mao, Sitaram Asur, et al. 2024. A Comprehensive Survey of LLM Alignment Techniques: RLHF, RLAIF, PPO, DPO and More. arXiv:2407.16216 (2024)

  133. [141]

    Xu, Uri Alon, Graham Neubig, and Vincent Josua Hellendoorn

    Frank F. Xu, Uri Alon, Graham Neubig, and Vincent Josua Hellendoorn. 2022. A Systematic Evaluation of Large Language Models of Code. In SIGPLAN International Symposium on Machine Programming . ACM, 1–10

  134. [142]

    Haocheng Xu, Haotian Hu, and Sitao Huang. 2024. Optimizing High-Level Synthesis Designs with Retrieval- Augmented Large Language Models. In IEEE LLM Aided Design Workshop (LAD) . IEEE, 1–5

  135. [143]

    Hanxiang Xu, Shenao Wang, Ningke Li, Kailong Wang, Yanjie Zhao, Kai Chen, Ting Yu, Yang Liu, and Haoyu Wang

  136. [144]

    Xingyu Wu, Sheng-hao Wu, Jibin Wu, Liang Feng, and Kay Chen Tan. 2024. Evolutionary Computation in the Era of Large Language Model: Survey and Roadmap. arXiv:2401.10034 (2024)

  137. [145]

    Qingyao Xu, Dingkang Yang, and Lihua Zhang. 2024. Code Optimization Chain-of-Thought: Structured Understanding and Self-Checking. In International Conference on Artificial Intelligence, Big Data and Algorithms . 425–430

  138. [146]

    Xuejun Yang, Yang Chen, Eric Eide, and John Regehr. 2011. Finding and Understanding Bugs in C Compilers. In ACM SIGPLAN Conference on Programming Language Design and Implementation, PLDI . ACM, 283–294

  139. [147]

    Xufeng Yao, Yiwen Wang, Xing Li, Yingzhao Lian, Ran Chen, Lei Chen, Mingxuan Yuan, Hong Xu, and Bei Yu. 2024. RTLRewriter: Methodologies for Large Models aided RTL Code Optimization

  140. [148]

    Large Language Models for Cyber Security: A Systematic Literature Review. (2024). arXiv:2405.04760

  141. [149]

    Jinglue Xu, Jialong Li, Zhen Liu, Nagar Anthel Venkatesh Suryanarayanan, Guoyuan Zhou, Jia Guo, Hitoshi Iba, and Kenji Tei. 2025. Large Language Models Synergize with Automated Machine Learning. Transactions on Machine Learning Research (TMLR) (2025)

  142. [150]

    Differentiation

    Mert Yüksekgönül, Federico Bianchi, Joseph Boen, Sheng Liu, Zhi Huang, Carlos Guestrin, and James Zou. 2024. TextGrad: Automatic "Differentiation" via Text. (2024). arXiv:2406.07496

  143. [151]

    Eric Zelikman, Eliana Lorch, Lester Mackey, and Adam Tauman Kalai. 2024. Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation. (2024)

  144. [153]

    Fangke Ye, Jisheng Zhao, Jun Shirako, and Vivek Sarkar. 2023. Concrete Type Inference for Code Optimization using Machine Learning with SMT Solving. Proceedings of the ACM on Programming Languages (PACMPL) 7 (2023), 773–800

  145. [154]

    Tong Ye, Tengfei Ma, Lingfei Wu, Xuhong Zhang, Shouling Ji, and Wenhai Wang. 2024. Iterative or Innovative? A Problem-Oriented Perspective for Code Optimization. (2024). arXiv:2406.11935

  146. [155]

    Quanjun Zhang, Chunrong Fang, Yang Xie, Yuxiang Ma, Weisong Sun, Yun Yang, and Zhenyu Chen. 2024. A Systematic Literature Review on Large Language Models for Automated Program Repair. (2024). arXiv:2405.01466

  147. [156]

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A Survey of Large Language Models. arXiv:2303.18223 (2023)

  148. [157]

    Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, et al. 2020. Ansor: Generating{High-Performance} tensor programs for deep learning. InUSENIX symposium on operating systems design and implementation (...

  149. [158]

    Kechi Zhang, Ge Li, Yihong Dong, Jingjing Xu, Jun Zhang, Jing Su, Yongfei Liu, and Zhi Jin. 2024. CodeDPO: Aligning Code Models with Self Generated and Verified Source Code. arXiv:2410.05605 (2024)

  150. [159]

    Peiyan Zhang, Haibo Jin, Leyang Hu, Xinnuo Li, Liying Kang, Man Luo, Yangqiu Song, and Haohan Wang. 2024. Revolve: Optimizing AI Systems by Tracking Response Evolution in Textual Optimization. arXiv:2412.03092 (2024)

  151. [163]

    Tianyu Zheng, Ge Zhang, Tianhao Shen, Xueling Liu, Bill Yuchen Lin, Jie Fu, Wenhu Chen, and Xiang Yue. 2024. OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement. In Findings of the Association for Computational Linguistics, ACL. Association for Compu...

  152. [164]

    Xunyu Zhu, Jian Li, Yong Liu, Can Ma, and Weiping Wang. 2023. A Survey on Model Compression for Large Language Models. arXiv:2308.07633 (2023). ACM Comput. Surv., Vol. 1, No. 1, Article . Publication date: January 2024

  153. [2021]

    In CC ’21: 30th ACM SIGPLAN International Conference on Compiler Construction, Virtual Event

    PolyBench/Python: Benchmarking Python Environments with Polyhedral Optimizations. In CC ’21: 30th ACM SIGPLAN International Conference on Compiler Construction, Virtual Event . ACM, 59–70

  154. [2023]

    In The Eleventh International Conference on Learning Representations, ICLR

    CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis. In The Eleventh International Conference on Learning Representations, ICLR . OpenReview.net

  155. [2024]

    Meta Large Language Model Compiler: Foundation Models of Compiler Optimization. (2024). arXiv:2407.02524

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.