Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

The Current Challenges of Software Engineering in the Era of Large Language Models

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper identifies 26 current challenges in applying large language models to software engineering, organized across seven stages of the development life cycle.

desk verdict Workshop-synthesis roadmap with a useful 26-challenge taxonomy, undermined by internal count inconsistencies and a thin audit trail. read the letter →

arxiv 2412.14554 v2 pith:6MJNMVIJ submitted 2024-12-19 cs.SE

classification cs.SE
keywords largelanguagemodelssoftwareengineeringLLM4SEdevelopmentlifecyclechallengesqualitativestudycodegenerationvulnerabilitymanagement
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to map where large language models still fall short when used in software engineering. Drawing on structured discussions among 24 researchers and practitioners, it identifies 26 challenges organized into seven areas: requirements and design, coding assistance, test generation, code review, maintenance, vulnerability management, and data, training, and evaluation. The authors' goal is to give the field a shared vocabulary and a research roadmap so that effort goes to the problems that actually slow down practice. If the list is right, it explains why current LLM-based tools are not yet trustworthy enough for end-to-end software development and where targeted research would pay off.

What carries the argument

The load-bearing structure of the paper is the challenge taxonomy itself: 26 challenges grouped into seven aspects of the software development life cycle plus model construction. Its companion mechanism is the qualitative coding pipeline, which transcribes six four-hour thematic sessions, converts the discussion into opinion cards through open coding, sorts the cards into candidate challenges, and has additional authors review the final set. The taxonomy is what carries the argument because the paper's contribution is not a new model or dataset but a stable inventory of pain points that future research can target and future evaluations can measure.

What would settle it

Conduct an independent coding of the same seminar transcripts by researchers who did not attend the seminar and were not told the authors' categories; if their independently grouped challenges overlap with the authors' 26 in fewer than half of the items, the taxonomy is a product of the coding team rather than the discussion. Alternatively, a large survey of practitioners rating each of the 26 challenges could show that several are not widely felt.

Watch

Extended reading notes

Core claim

The central claim is that LLM4SE is not blocked by a single bottleneck but by a recognizable constellation of 26 challenges, and that these challenges are stable enough to be named and grouped. The paper derives the list from face-to-face discussions at a three-day seminar, transcribing six thematic sessions and coding the material with a qualitative open-coding procedure followed by open card sorting and author verification. The resulting taxonomy spans the whole software development life cycle, from requirement prompts and domain knowledge, through hallucination, vulnerability injection, and project-level integration in code generation, to syntactic and semantic problems in generated tests, issue-specific code review, microservice dependency complexity in maintenance, vulnerability data scarcity, and unresolved data, training, and evaluation questions. The authors present these challenges as a roadmap: each named challenge points to a concrete research direction.

Load-bearing premise

The whole list rests on the assumption that the views of 24 invited specialists, filtered through two authors' coding, stand in for the full set of challenges the field faces; the paper itself concedes the list may be incomplete.

Editorial extensions

If this is right

  • If the taxonomy is correct, research on LLM-based code generation should prioritize hallucination control and project-level integration over further gains on simple benchmarks such as HumanEval and MBPP, whose scores the paper describes as near-limit.
  • If the taxonomy is correct, test generation research needs to attack syntactic compilability and semantic coverage together, including automatic mocking and assertion oracles.
  • If the taxonomy is correct, code review automation should be split by issue type and by industry versus open-source practice, rather than treated as one generic task.
  • If the taxonomy is correct, vulnerability management will not advance without new high-quality vulnerability explanation data and context-slicing methods that fit within LLM context windows.
  • If the taxonomy is correct, improving LLM4SE requires building evaluation frameworks that reproduce real development contexts instead of relying on benchmark scores.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy could double as a maturity checklist: a future survey could score each of the 26 challenges as solved, partially solved, or open, turning the qualitative list into a quantitative progress metric.
  • The absence of any priority ordering among the 26 challenges is a gap the authors do not fill; a ranking by industrial impact would require separate empirical work.
  • Some of the challenges are tied to current model limitations, such as context windows and training-data scarcity; if those constraints ease, the taxonomy would need revision, which suggests it should be treated as a snapshot rather than a fixed classification.
  • The same discussion method could be applied to adjacent areas, for example to map challenges in LLM-based DevOps or in using LLMs for embedded systems, where the specific challenge mix may differ.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper reports a qualitative study of current challenges in LLM-based software engineering (LLM4SE). The authors organized a 24-participant seminar under the CCF Beautiful Lake Seminars, transcribed and open-coded the discussions in NVivo, and used open card sorting to derive 26 key challenges across seven aspects: software requirement and design, coding assistance, testing code generation, code review, software maintenance, software vulnerability management, and data, training, and evaluation. The paper also provides background material on the LLM4SE pipeline (data construction, fine-tuning, prompting, SE-specific LLMs) and a related-work survey organized by the same seven areas, and it concludes with a research roadmap and a threats-to-validity section.

Significance. If the challenge taxonomy is accepted, the paper offers a useful, expert-informed roadmap for LLM4SE research and practice, and it usefully spans the whole software development life cycle. The authors follow recognizable qualitative procedures, explicitly acknowledge completeness and representativeness threats in Section 6, and make session topics publicly available. The main value is as a position/roadmap paper rather than as a hypothesis-testing empirical study; its central contribution is the challenge list itself, which makes the consistency and auditability of that list the deciding factors for the paper's credibility.

major comments (3)
  1. [Abstract, Section 4, Section 7] The paper's central claim is the count and organization of challenges, but the manuscript states inconsistent numbers: the abstract and Section 1 claim 26 challenges from seven aspects; Section 4's introductory paragraph says the challenges are introduced 'from six aspects'; and Section 7 says 'We present seven challenges.' Because the taxonomy and its counts are the paper's main contribution, these inconsistencies must be reconciled and the final counts made unambiguous.
  2. [Section 3] The qualitative coding procedure is described at a high level (transcription, open coding in NVivo, verification by a second author, card sorting by two authors, review by two more), but no inter-rater reliability statistic, codebook, opinion cards, or raw transcripts are provided. The linked repository (footnote 1) contains only session topics. Without such an audit trail, a reader cannot independently verify that the reported 26 challenges and their grouping into seven aspects are supported by the discussed material rather than by the authors' post-hoc synthesis.
  3. [Section 6] The paper itself concedes that the challenge list may not cover all current LLM4SE challenges and that representativeness is addressed only by participant diversity. This is a reasonable limitation for a position paper, but the abstract's phrasing 'we achieve 26 key challenges from seven aspects' overstates the completeness of a list explicitly derived from 24 participants in a single seminar. The claims should be qualified to match the admitted scope, or additional evidence of saturation should be provided.
minor comments (6)
  1. [Throughout] There are numerous typos and misspellings, including 'Technolgy' in the author affiliation, 'through discussion' for 'thorough discussion' in the abstract, 'techniuqes' in Section 2, 'specilizing' in Section 3, 'assitance' in Section 2.6, 'datasety' in the Section 4.4 summary, and 'to to understand' in Section 4.5. A careful proofreading pass is needed.
  2. [Section 3] The seminar dates are given as '19 Jan 2014 - 21 Jan 2024'; the year '2014' is a typo and should be '2024'.
  3. [Section 5.2] The citation 'CodeX [ ? ]' is an unresolved placeholder and must be corrected to the proper reference.
  4. [ACM Reference Format] The ACM Reference Format block contains placeholder text ('Make sure to enter the correct conference title', 'Conference acronym ’XX', 'XXXXXXX.XXXXXXX') and a template year of 2018; these should be replaced with the actual venue and metadata.
  5. [Section 2.4] The naming of models is inconsistent (e.g., 'CodeLlama' vs. 'CodeLLaMa' in Table 1 and text); please unify the spelling.
  6. [Table 1] Several entries in Table 1 have uncertain or missing values marked '—'; consider adding a note explaining what '—' means and whether the model size or pre-training data size is intentionally omitted.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the 26 challenges are a qualitative coding outcome from a seminar discussion, with no fitted parameter, equation, or self-citation chain doing the work.

full rationale

The paper's central claim is that a discussion among 24 participants produced 26 key LLM4SE challenges in seven aspects. The derivation chain is the qualitative procedure in Section 3: the first author transcribed and open-coded the seminars in NVivo, another author verified the opinion cards, two authors performed card sorting, and two more authors reviewed the final set. This is a coding and summarization activity, not a mathematical or statistical derivation; no quantity is predicted from a fitted input, and no equation is shown to equal its own input by construction. The paper contains many references by the authors (e.g., [62] and [148] for the coding procedure; [149], [161]-[165], [175] in related work), but none of these citations is load-bearing for the 26-challenge taxonomy: the taxonomy is grounded in the seminar discussion and the authors' coding, not in the cited papers, and deleting the self-citations would not alter the listed challenges. Section 6 explicitly concedes a completeness threat ('The summarized challenges are based on the discussion among 24 participants, which might not cover all the current challenges of LLM4SE'), which is a validity limitation rather than evidence of circularity. The internal inconsistencies in the text (six vs seven aspects; 'seven challenges' in Section 7 vs 26 key challenges in the abstract) are textual errors, not self-definitional reductions. No uniqueness theorem, ansatz-smuggling citation, or renaming of a known result is present. The absence of released transcripts and inter-rater reliability statistics is a reproducibility and auditability concern, not a circularity reduction. The paper is therefore self-contained in the sense relevant to circularity: its central output is explicitly the product of the reported qualitative coding, not of a hidden re-use of its own inputs.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no free parameters or invented entities. Its central claim rests on the qualitative assumptions that the 24 workshop participants are representative of the field and that the two-author coding procedure yields reliable categories. The authors themselves flag incompleteness in Section 6.

assumptions (3)
  • domain assumption The 24 workshop participants are representative of the broader LLM4SE research and practice community.
    The central claim that 26 items are 'key challenges' depends on the workshop attendees covering the field's relevant perspectives. Section 6 admits this may not hold.
  • domain assumption Open coding and card sorting performed by two authors, with review by two others, yields reliable challenge categories.
    Section 3 describes the coding procedure but provides no inter-rater reliability coefficients or independent audit, so the reliability of the categorization is assumed.
  • domain assumption The challenges identified in January 2024 remain representative at the time of publication.
    LLM capabilities changed rapidly in 2024; the paper provides no evidence that the challenge list is stable over time.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Current Challenges of Software Engineering in the Era of Large Language Models." pith.science (2026). https://pith.science/paper/6MJNMVIJ

@misc{pith2026241214554,
  author       = {Pith},
  title        = {Pith review of: The Current Challenges of Software Engineering in the Era of Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6MJNMVIJ}},
  note         = {Machine review of arXiv:2412.14554}
}
read the original abstract

With the advent of large language models (LLMs) in the artificial intelligence (AI) area, the field of software engineering (SE) has also witnessed a paradigm shift. These models, by leveraging the power of deep learning and massive amounts of data, have demonstrated an unprecedented capacity to understand, generate, and operate programming languages. They can assist developers in completing a broad spectrum of software development activities, encompassing software design, automated programming, and maintenance, which potentially reduces huge human efforts. Integrating LLMs within the SE landscape (LLM4SE) has become a burgeoning trend, necessitating exploring this emergent landscape's challenges and opportunities. The paper aims at revisiting the software development life cycle (SDLC) under LLMs, and highlighting challenges and opportunities of the new paradigm. The paper first summarizes the overall process of LLM4SE, and then elaborates on the current challenges based on a through discussion. The discussion was held among more than 20 participants from academia and industry, specializing in fields such as software engineering and artificial intelligence. Specifically, we achieve 26 key challenges from seven aspects, including software requirement & design, coding assistance, testing code generation, code review, code maintenance, software vulnerability management, and data, training, and evaluation. We hope the achieved challenges would benefit future research in the LLM4SE field.

Figures

Figures reproduced from arXiv: 2412.14554 by the authors.

Figure 1
Figure 1. Overall process of LLM4SE [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM Performance for Code Generation on Noisy Tasks

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LLMs solve heavily obfuscated benchmark tasks, and performance decay under obfuscation differs sharply between old and new datasets, which the authors interpret as a signature of training-data contamination.

Reference graph

Works this paper leans on

204 extracted references · 10 canonical work pages · cited by 1 Pith paper

  1. [1]

    The 9th CCF Beautiful Lake Seminars

    2024. The 9th CCF Beautiful Lake Seminars. https://www.ccf.org.cn/xhhy/2024-03-15/816185.shtml

  2. [2]

    TensorBoard

    2024. TensorBoard. https://www.tensorflow.org/tensorboard

  3. [3]

    Jean-François Abramatic, Roberto Di Cosmo, and Stefano Zacchiroli. 2018. Building the universal archive of source code. Commun. ACM 61, 10 (2018), 29–31

  4. [4]

    Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Unified Pre-training for Program Understanding and Generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021 . Association for Computation...

  5. [5]

    Devanbu, and Saikat Chakraborty

    Toufique Ahmed, Christian Bird, Premkumar T. Devanbu, and Saikat Chakraborty. 2024. Studying LLM Performance on Closed- and Open-source Data. CoRR abs/2402.15100 (2024)

  6. [6]

    Toufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann, Xuchao Zhang, and Saravan Rajmohan

  7. [7]

    Muhammad Azeem Akbar, Arif Ali Khan, Najmul Islam, and Sajjad Mahmood. 2024. DevOps project management success factors: A decision-making framework. Softw. Pract. Exp. 54, 2 (2024), 257–280

  8. [8]

    Saranya Alagarsamy, Chakkrit Tantithamthavorn, and Aldeida Aleti. 2024. A3test: Assertion-augmented automated test case generation. Information and Software Technology (2024), 107565

Show all 204 references
  1. [9]

    Loubna Ben Allal, Raymond Li, Denis Kocetkov, Chenghao Mou, Christopher Akiki, Carlos Muñoz Ferrandis, Niklas Muennighoff, Mayank Mishra, Alex Gu, Manan Dey, Logesh Kumar Umapathi, Carolyn Jane Anderson, Yangtian Zi, Joel Lamy-Poirier, Hailey Schoelkopf, Sergey Troshin, Dmitry...

  2. [10]

    Sathurshan Arulmohan, Marie-Jean Meurs, and Sébastien Mosser. 2023. Extracting Domain Models from Textual Requirements in the Era of Large Language Models. InACM/IEEE International Conference on Model Driven Engineering Languages and Systems, MODELS 2023 Companion, Västerås, S...

  3. [11]

    SantaCoder: don’t reach for the stars! CoRR abs/2301.03988 (2023)

  4. [12]

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Si...

  5. [13]

    Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021. Program Synthesis with Large Language Models. arXiv preprint arXiv:2108.07732 (2021)

  6. [14]

    Benoit Baudry, Khashayar Etemadi, Sen Fang, Yogya Gamage, Yi Liu, Yuxin Liu, Martin Monperrus, Javier Ron, André Silva, and Deepika Tiwari. 2024. Generative AI to Generate Test Data Generators. arXiv preprint arXiv:2401.17626 (2024)

  7. [15]

    Ramakrishna Bairi, Atharv Sonwane, Aditya Kanade, Arun Iyer, Suresh Parthasarathy, Sriram Rajamani, B Ashok, and Shashank Shet. 2024. Codeplan: Repository-level coding using llms and planning. Proceedings of the ACM on , Vol. 1, No. 1, Article . Publication date: December 2018...

  8. [16]

    Guru Prasad Bhandari, Amara Naseer, and Leon Moonen. 2021. CVEfixes: automated collection of vulnerabilities and their fixes from open-source software. In PROMISE ’21: 17th International Conference on Predictive Models and Data Analytics in Software Engineering, Athens Greece,...

  9. [17]

    Tobias Baum, Olga Liskin, Kai Niklas, and Kurt Schneider. 2016. Factors influencing code review processes in industry. In FSE. ACM, 85–96

  10. [18]

    Adam Bochenek. 2023. Can ChatGPT Replace a Template-based Code Generator?. InFedCSIS (Communication Papers). 43–50

  11. [19]

    Kelly Blincoe, Giuseppe Valetto, and Daniela E. Damian. [n. d.]. Do all task dependencies require coordination? the role of task properties in identifying critical coordination needs in software projects. In Joint Meeting of the European Software Engineering Conference and the...

  12. [20]

    Busari and Emmanuel Letier

    Saheed A. Busari and Emmanuel Letier. 2017. RADAR: a lightweight tool for requirements and architecture decision analysis. In Proceedings of the 39th International Conference on Software Engineering, ICSE 2017, Buenos Aires, Argentina, May 20-28, 2017. IEEE / ACM, 552–562

  13. [21]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901

  14. [22]

    Yu, and Lichao Sun

    Yihan Cao, Siyu Li, Yixin Liu, Zhiling Yan, Yutong Dai, Philip S. Yu, and Lichao Sun. 2023. A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT. CoRR abs/2303.04226 (2023)

  15. [23]

    Javier Cámara, Javier Troya, Lola Burgueño, and Antonio Vallecillo. 2023. On the assessment of generative AI in modeling tasks: an experience report with ChatGPT and UML. Softw. Syst. Model. 22, 3 (2023), 781–793

  16. [24]

    Kua Chen, Yujing Yang, Boqi Chen, José Antonio Hernández López, Gunter Mussbacher, and Dániel Varró. 2023. Automated Domain Modeling with Large Language Models: A Comparative Study. In 26th ACM/IEEE International Conference on Model Driven Engineering Languages and Systems, MO...

  17. [25]

    Mikaël Chelli, Jules Descamps, Vincent Lavoué, Christophe Trojani, Michel Azar, Marcel Deckert, Jean-Luc Raynier, Gilles Clowez, Pascal Boileau, and Caroline Ruetsch-Chelli. 2024. Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparativ...

  18. [26]

    Yinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang, Xin Gao, Liu Shi, Yunjie Cao, Xuedong Gao, Hao Fan, Ming Wen, Jun Zeng, Supriyo Ghosh, Xuchao Zhang, Chaoyun Zhang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, and Tianyin Xu. 2024. Automatic Root Cause Analysis via Large Lang...

  19. [27]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)

  20. [28]

    Fenia Christopoulou, Gerasimos Lampouras, Milan Gritta, Guchun Zhang, Yinpeng Guo, Zhongqi Li, Qi Zhang, Meng Xiao, Bo Shen, Lin Li, Hao Yu, Li Yan, Pingyi Zhou, Xin Wang, Yuchi Ma, Ignacio Iacobacci, Yasheng Wang, Guangtai Liang, Jiansheng Wei, Xin Jiang, Qianxiang Wang, and ...

  21. [29]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vino...

  22. [30]

    DeepSeek-AI, Qihao Zhu, Daya Guo, Zhihong Shao, Dejian Yang, Peiyi Wang, Runxin Xu, Y. Wu, Yukun Li, Huazuo Gao, Shirong Ma, Wangding Zeng, Xiao Bi, Zihui Gu, Hanwei Xu, Damai Dai, Kai Dong, Liyue Zhang, Yishi Piao, Zhibin Gou, Zhenda Xie, Zhewen Hao, Bingxuan Wang, Junxiao So...

  23. [31]

    Copilot. 2024. https://copilot.microsoft.com/

  24. [32]

    Bosheng Ding, Chengwei Qin, Linlin Liu, Yew Ken Chia, Shafiq Joty, Boyang Li, and Lidong Bing. 2022. Is gpt-3 a good data annotator? arXiv preprint arXiv:2212.10450 (2022)

  25. [33]

    Yinlin Deng, Chunqiu Steven Xia, Chenyuan Yang, Shizhuo Dylan Zhang, Shujing Yang, and Lingming Zhang

  26. [34]

    arXiv preprint arXiv:2304.02014 (2023)

    Large language models are edge-case fuzzers: Testing deep learning libraries via fuzzgpt. arXiv preprint arXiv:2304.02014 (2023)

  27. [35]

    Pradeep Dogga, Chetan Bansal, Richard Costleigh, Gopinath Jayagopal, Suman Nath, and Xuchao Zhang. 2023. AutoARTS: Taxonomy, Insights and Tools for Root Cause Labelling of Incidents in Microsoft Azure. InProceedings of the 2023 USENIX Annual Technical Conference, USENIX ATC 20...

  28. [36]

    Min, Gail E

    Yangruibo Ding, Jinjun Peng, Marcus J. Min, Gail E. Kaiser, Junfeng Yang, and Baishakhi Ray. 2024. SemCoder: Training Code Language Models with Comprehensive Semantics. CoRR abs/2406.01006 (2024)

  29. [37]

    Yangruibo Ding, Zijian Wang, Wasi Ahmad, Hantian Ding, Ming Tan, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, et al. 2024. Crosscodeeval: A diverse and multilingual benchmark for cross-file code completion. Advances in Neural Information...

  30. [38]

    Andrew Edwards-Jones. 2014. Qualitative data analysis with NVIVO

  31. [39]

    Yihong Dong, Xue Jiang, Zhi Jin, and Ge Li. 2023. Self-collaboration code generation via chatgpt. arXiv preprint arXiv:2304.07590 (2023)

  32. [40]

    GitLab Duo. 2024. https://about.gitlab.com/gitlab-duo/

  33. [41]

    Sarah Fakhoury, Aaditya Naik, Georgios Sakkas, Saikat Chakraborty, and Shuvendu K Lahiri. 2024. LLM-based Test-driven Interactive Code Generation: User Study and Empirical Evaluation.arXiv preprint arXiv:2404.10100 (2024)

  34. [42]

    Madeline Endres, Sarah Fakhoury, Saikat Chakraborty, and Shuvendu K Lahiri. 2023. Formalizing natural language intent into program specifications via large language models. arXiv preprint arXiv:2310.01831 (2023)

  35. [43]

    Saad Ezzini, Sallam Abualhaija, Chetan Arora, and Mehrdad Sabetzadeh. 2022. Automated handling of anaphoric ambiguity in requirements: a multi-solution study. In Proceedings of the 44th International Conference on Software Engineering. 187–199

  36. [44]

    Jia Feng, Jiachen Liu, Cuiyun Gao, Chun Yong Chong, Chaozheng Wang, Shan Gao, and Xin Xia. 2024. Complex- CodeEval: A Benchmark for Evaluating Large Code Models on More Complex Code. arXiv preprint arXiv:2409.10280 (2024)

  37. [45]

    Angela Fan, Beliz Gokkaya, Mark Harman, Mitya Lyubarskiy, Shubho Sengupta, Shin Yoo, and Jie M. Zhang. 2023. Large Language Models for Software Engineering: Survey and Open Problems. In IEEE/ACM International Conference on Software Engineering: Future of Software Engineering, ...

  38. [46]

    Joao Faustino, Daniel Adriano, Ricardo Amaro, Rubén Pereira, and Miguel Mira da Silva. 2022. DevOps benefits: A systematic literature review. Software: Practice and Experience 52, 9 (2022), 1905–1926

  39. [47]

    Michael Fu, Chakkrit Kla Tantithamthavorn, Van Nguyen, and Trung Le. 2023. ChatGPT for Vulnerability Detection, Classification, and Repair: How Far Are We?. In 30th Asia-Pacific Software Engineering Conference, APSEC 2023, Seoul, Republic of Korea, December 4-7, 2023 . IEEE, 632–636

  40. [48]

    Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Findings of the Association for Computational Linguistics: EMNLP ...

  41. [49]

    Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, and Mike Lewis. 2023. InCoder: A Generative Model for Code Infilling and Synthesis. In The Eleventh International Conference on Learning Representations, IC...

  42. [50]

    gartner. 2023. https://www.gartner.com/en/articles/gartner-top-10-strategic-technology-trends-for-2023

  43. [51]

    Vaibhav Ganatra, Anjaly Parayil, Supriyo Ghosh, Yu Kang, Minghua Ma, Chetan Bansal, Suman Nath, and Jonathan Mace. 2023. Detection Is Better Than Cure: A Cloud Incidents Perspective. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on...

  44. [52]

    Shuzheng Gao, Xin-Cheng Wen, Cuiyun Gao, Wenxuan Wang, Hongyu Zhang, and Michael R. Lyu. 2023. What Makes Good In-Context Demonstrations for Code Intelligence Tasks with LLMs?. In 38th IEEE/ACM International , Vol. 1, No. 1, Article . Publication date: December 2018. 24 Gao et...

  45. [53]

    Clement, Dawn Drain, Neel Sundaresan, Jian Yin, Daxin Jiang, and Ming Zhou

    Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin B. Clement, Dawn Drain, Neel Sundaresan, Jian Yin, Daxin Jiang, and Ming Zhou. 2021. GraphCodeBERT: Pre-training Code ...

  46. [54]

    Hamza Ghani, Jesus Luna, Abdelmajid Khelil, Najib Alkadri, and Neeraj Suri. 2013. Predictive vulnerability scoring in the context of insufficient information availability. In 2013 International Conference on Risks and Security of Internet and Systems (CRiSIS), La Rochelle, Fra...

  47. [55]

    Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Y...

  48. [56]

    Mark Harman. 2012. The role of artificial intelligence in software engineering. In Proceedings of the First International Workshop on Realizing AI Synergies in Software Engineering, RAISE 2012, Zurich, Switzerland, June 5, 2012 . IEEE, 1–6

  49. [57]

    Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y. Wu, Y. K. Li, Fuli Luo, Yingfei Xiong, and Wenfeng Liang. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence. CoRR abs/2401.14196 (2024)

  50. [58]

    Yiyang Hao, Ge Li, Yongqiang Liu, Xiaowei Miao, He Zong, Siyuan Jiang, Yang Liu, and He Wei. 2022. AixBench: A Code Generation Benchmark Dataset. CoRR abs/2206.13179 (2022)

  51. [59]

    John E Hopcroft, Jeffrey D Ullman, and Alfred Vaino Aho. 1983. Data structures and algorithms . Addison-wesley Boston, MA, USA

  52. [60]

    Sepehr Hashtroudi, Jiho Shin, Hadi Hemmati, and Song Wang. 2023. Automated test case generation using code models and domain adaptation. arXiv preprint arXiv:2308.08033 (2023)

  53. [61]

    Tobias Hey, Jan Keim, Anne Koziolek, and Walter F Tichy. 2020. Norbert: Transfer learning for requirements classification. In 2020 IEEE 28th international requirements engineering conference (RE) . IEEE, 169–179

  54. [62]

    Xing Hu, Xin Xia, David Lo, Zhiyuan Wan, Qiuyuan Chen, and Thomas Zimmermann. 2022. Practitioners’ Expectations on Automated Code Comment Generation. In 44th IEEE/ACM 44th International Conference on Software Engineering, ICSE 2022, Pittsburgh, PA, USA, May 25-27, 2022 . ACM, ...

  55. [63]

    Sihao Hu, Tiansheng Huang, Fatih Ilhan, Selim Furkan Tekin, and Ling Liu. 2023. Large Language Model-Powered Smart Contract Vulnerability Detection: New Perspectives. In 5th IEEE International Conference on Trust, Privacy and Security in Intelligent Systems and Applications, T...

  56. [64]

    Xueyu Hu, Kun Kuang, Jiankai Sun, Hongxia Yang, and Fei Wu. 2024. Leveraging print debugging to improve code generation in large language models. arXiv preprint arXiv:2401.05319 (2024)

  57. [65]

    Junjie Huang, Jinyang Liu, Zhuangbin Chen, Zhihan Jiang, Yichen Li, Jiazhen Gu, Cong Feng, Zengyin Yang, Yongqiang Yang, and Michael R. Lyu. 2024. FaultProfIT: Hierarchical Fault Profiling of Incident Tickets in Large-scale Cloud Systems. In Proceedings of the 46th Internation...

  58. [66]

    Dong Huang, Qingwen Bu, Jie Zhang, Xiaofei Xie, Junjie Chen, and Heming Cui. 2023. Bias assessment and mitigation in llm-based code generation. arXiv preprint arXiv:2309.14345 (2023)

  59. [67]

    Dong Huang, Qingwen Bu, Jie M Zhang, Michael Luck, and Heming Cui. 2023. Agentcoder: Multi-agent-based code generation with iterative testing and optimisation. arXiv preprint arXiv:2312.13010 (2023)

  60. [68]

    Md Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez. 2024. MapCoder: Multi-Agent Code Generation for Competitive Problem Solving. arXiv preprint arXiv:2405.11403 (2024)

  61. [69]

    Yintong Huo, Cheryl Lee, Yuxin Su, Shiwen Shan, Jinyang Liu, and Michael R. Lyu. 2023. EvLog: Identifying Anomalous Logs over Software Evolution. In 34th IEEE International Symposium on Software Reliability Engineering, ISSRE 2023, Florence, Italy, October 9-12, 2023 . IEEE, 391–402

  62. [70]

    Sangwon Hyun, Mingyu Guo, and M Ali Babar. 2024. METAL: Metamorphic Testing Framework for Analyzing Large-Language Model Qualities. In 2024 IEEE Conference on Software Testing, Verification and Validation (ICST) . IEEE, 117–128

  63. [71]

    Siyuan Jiang, Jia Li, He Zong, Huanyu Liu, Hao Zhu, Shukai Hu, Erlu Li, Jiazheng Ding, Yu Han, Wei Ning, Gen Wang, Yihong Dong, Kechi Zhang, and Ge Li. 2024. aiXcoder-7B: A Lightweight and Effective Large Language Model for Code Completion. CoRR abs/2410.13187 (2024)

  64. [72]

    ISO ISO. 2015. IEC/IEEE International Standard-Systems and software engineering–System life cycle processes. ISO/IEC/IEEE 15288 First edition 2015–05–15, Technical report (2015)

  65. [73]

    Nan Jiang, Thibaud Lutellier, and Lin Tan. 2021. CURE: Code-Aware Neural Machine Translation for Automatic Program Repair. In 43rd IEEE/ACM International Conference on Software Engineering, ICSE 2021, Madrid, Spain, 22-30 May 2021. IEEE, 1161–1173. , Vol. 1, No. 1, Article . P...

  66. [74]

    Pengxiang Jin, Shenglin Zhang, Minghua Ma, Haozhe Li, Yu Kang, Liqun Li, Yudong Liu, Bo Qiao, Chaoyun Zhang, Pu Zhao, Shilin He, Federica Sarro, Yingnong Dang, Saravan Rajmohan, Qingwei Lin, and Dongmei Zhang. 2023. Assess and Summarize: Improve Outage Understanding with Large...

  67. [75]

    Xue Jiang, Yihong Dong, Lecheng Wang, Fang Zheng, Qiwei Shang, Ge Li, Zhi Jin, and Wenpin Jiao. 2023. Self- planning Code Generation with Large Language Models. ACM Transactions on Software Engineering and Methodology (2023)

  68. [76]

    Zhihan Jiang, Jinyang Liu, Zhuangbin Chen, Yichen Li, Junjie Huang, Yintong Huo, Pinjia He, Jiazhen Gu, and Michael R. Lyu. 2023. LLMParser: A LLM-based Log Parsing Framework. CoRR abs/2310.01796 (2023)

  69. [77]

    Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, et al. 2022. The stack: 3 tb of permissively licensed source code. arXiv preprint arXiv:2211.15533 (2022)

  70. [78]

    Hyeji Kim, Yihan Jiang, Sreeram Kannan, Sewoong Oh, and Pramod Viswanath. 2018. Deepcode: Feedback Codes via Deep Learning. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 201...

  71. [79]

    Myeongsoo Kim, Tyler Stennett, Dhruv Shah, Saurabh Sinha, and Alessandro Orso. 2024. Leveraging large language models to improve REST API testing. In Proceedings of the 2024 ACM/IEEE 44th International Conference on Software Engineering: New Ideas and Emerging Results . 37–41

  72. [80]

    Triet Huynh Minh Le, Huaming Chen, and Muhammad Ali Babar. 2023. A Survey on Data-driven Software Vulnera- bility Assessment and Prioritization. ACM Comput. Surv. 55, 5 (2023), 100:1–100:39

  73. [81]

    Oleksii Kononenko, Olga Baysal, Latifa Guerrouj, Yaxin Cao, and Michael W. Godfrey. 2015. Investigating code review quality: Do people and participation matter?. In ICSME. IEEE Computer Society, 111–120

  74. [82]

    LaToza, Gina Venolia, and Robert DeLine

    Thomas D. LaToza, Gina Venolia, and Robert DeLine. 2006. Maintaining mental models: a study of developer work habits. In 28th International Conference on Software Engineering (ICSE 2006), Shanghai, China, May 20-28, 2006 , Leon J. Osterweil, H. Dieter Rombach, and Mary Lou Sof...

  75. [83]

    Jia Li, Ge Li, Chongyang Tao, Huangzhao Zhang, Fang Liu, and Zhi Jin. 2023. Large language model-aware in-context learning for code generation. arXiv preprint arXiv:2310.09748 (2023)

  76. [84]

    Van-Hoang Le and Hongyu Zhang. 2023. Log Parsing with Prompt-based Few-shot Learning. In 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023 . IEEE, 2438–2449

  77. [85]

    Yoav Levine, Noam Wies, Daniel Jannai, Dan Navon, Yedid Hoshen, and Amnon Shashua. 2021. The Inductive Bias of In-Context Learning: Rethinking Pretraining Example Design. InInternational Conference on Learning Representations

  78. [86]

    Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, ...

  79. [87]

    Jia Li, Yunfei Zhao, Yongmin Li, Ge Li, and Zhi Jin. 2023. Towards enhancing in-context learning for code generation. arXiv preprint arXiv:2303.17780 (2023)

  80. [88]

    Lingwei Li, Li Yang, Huaxi Jiang, Jun Yan, Tiejian Luo, Zihan Hua, Geng Liang, and Chun Zuo. 2022. AUGER: automatically generating review comments with pre-training models. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Found...

  81. [89]

    Zhiyu Li, Shuai Lu, Daya Guo, Nan Duan, Shailesh Jannu, Grant Jenks, Deep Majumder, Jared Green, Alexey Svyatkovskiy, Shengyu Fu, and Neel Sundaresan. 2022. Automating code review activities by large-scale pre-training. In Proceedings of the 30th ACM Joint European Software En...

  82. [90]

    Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar, and Yin Tat Lee. 2023. Textbooks Are All You Need II: phi-1.5 technical report. CoRR abs/2309.05463 (2023)

  83. [91]

    Yujia Li, David H. Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, ...

  84. [92]

    Jinfeng Lin, Yalin Liu, Qingkai Zeng, Meng Jiang, and Jane Cleland-Huang. 2021. Traceability transformed: Generating more accurate links with pre-trained bert models. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE). IEEE, 324–335

  85. [93]

    Zongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang, Dong Chen, Shuai Wang, and Cuiyun Gao. 2023. CCTEST: Testing and Repairing Code Completion Systems. In 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023 . IEEE, ...

  86. [94]

    Feng Lin, Dong Jae Kim, et al. 2024. When llm-based code generation meets the software development process. arXiv preprint arXiv:2403.15852 (2024)

  87. [95]

    Shuo Liu, Jacky Keung, Zhen Yang, Fang Liu, Qilin Zhou, and Yihan Liao. 2024. Delving into Parameter-Efficient Fine-Tuning in Code Change Learning: An Empirical Study. In SANER. IEEE, 465–476

  88. [96]

    Yishuai Lin, Philippe Descamps, Nicolas Gaud, Vincent Hilaire, and Abderrafiaa Koukam. 2015. Multi-Agent System for intelligent Scrum project management. Integr. Comput. Aided Eng. 22, 3 (2015), 281–296

  89. [97]

    Linger, Harlan D

    Richard C. Linger, Harlan D. Mills, and Bernard I. Witt. 1979. Structured programming - theory and practice . Addison- Wesley

  90. [98]

    Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, et al. 2024. StarCoder 2 and The Stack v2: The Next Generation. arXiv preprint arXiv:2402.19173 (2024)

  91. [99]

    Zhe Liu, Chunyang Chen, Junjie Wang, Xing Che, Yuekai Huang, Jun Hu, and Qing Wang. 2023. Fill in the blank: Context-aware automated text input generation for mobile gui testing. In2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 1355–1367

  92. [100]

    Zhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen, Boyu Wu, Zhilin Tian, Yuekai Huang, Jun Hu, and Qing Wang. 2024. Testing the limits: Unusual text inputs generation for mobile app crash detection with large language model. In Proceedings of the IEEE/ACM 46th International C...

  93. [101]

    Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and S...

  94. [102]

    Guilong Lu, Xiaolin Ju, Xiang Chen, Wenlong Pei, and Zhilong Cai. 2024. GRACE: Empowering LLM-based software vulnerability detection with graph structure and in-context learning. J. Syst. Softw. 212 (2024), 112031

  95. [103]

    Junyi Lu, Lei Yu, Xiaojia Li, Li Yang, and Chun Zuo. 2023. LLaMA-Reviewer: Advancing code review automation with large language models through parameter-efficient fine-tuning. In 2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE) . IEEE, 647–658

  96. [104]

    Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2023. WizardCoder: Empowering Code Large Language Models with Evol-Instruct. CoRR abs/2306.08568 (2023)

  97. [105]

    Sebastian Lubos, Alexander Felfernig, Thi Ngoc Trang Tran, Damian Garber, Merfat El Mansi, Seda Polat Erdeniz, and Viet-Man Le. 2024. Leveraging LLMs for the Quality Assurance of Software Requirements. In 32nd IEEE International Requirements Engineering Conference, RE 2024, Re...

  98. [106]

    Xianchang Luo, Yinxing Xue, Zhenchang Xing, and Jiamou Sun. 2022. Prcbert: Prompt learning for requirement clas- sification using bert-based pretrained language models. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering. 1–13

  99. [107]

    Ambarish Moharil and Arpit Sharma. 2022. Identification of intra-domain ambiguity using transformer-based machine learning. In Proceedings of the 1st International Workshop on Natural Language-based Software Engineering . 51–58. , Vol. 1, No. 1, Article . Publication date: Dec...

  100. [108]

    Lezhi Ma, Shangqing Liu, Yi Li, Xiaofei Xie, and Lei Bu. 2024. SpecGen: Automated Generation of Formal Program Specifications via Large Language Models. arXiv preprint arXiv:2401.08807 (2024)

  101. [109]

    Lipeng Ma, Weidong Yang, Bo Xu, Sihang Jiang, Ben Fei, Jiaqing Liang, Mingjie Zhou, and Yanghua Xiao. 2024. KnowLog: Knowledge Enhanced Pre-trained Language Model for Log Understanding. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering, ICSE ...

  102. [110]

    Noor Nashid, Mifta Sintaha, and Ali Mesbah. [n. d.]. Retrieval-Based Prompt Selection for Code-Related Few-Shot Learning. In 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023. IEEE, 2450–2462

  103. [111]

    Ambarish Moharil and Arpit Sharma. 2023. Tabasco: A transformer based contextualization toolkit. Science of Computer Programming 230 (2023), 102994

  104. [112]

    Fangwen Mu, Xiao Chen, Lin Shi, Song Wang, and Qing Wang. 2023. Developer-Intent Driven Code Comment Generation. In 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023. IEEE, 768–780

  105. [113]

    Erik Nijkamp, Hiroaki Hayashi, Caiming Xiong, Silvio Savarese, and Yingbo Zhou. 2023. CodeGen2: Lessons for Training LLMs on Programming and Natural Languages. CoRR abs/2305.02309 (2023)

  106. [114]

    Phuong T Nguyen, Juri Di Rocco, Claudio Di Sipio, Riccardo Rubei, Davide Di Ruscio, and Massimiliano Di Penta

  107. [115]

    Easterbrook

    Bashar Nuseibeh and Steve M. Easterbrook. [n. d.]. Requirements engineering: a roadmap. In 22nd International Conference on on Software Engineering, Future of Software Engineering Track, ICSE 2000, Limerick Ireland, June 4-11,

  108. [116]

    Chao Ni, Xin Yin, Kaiwen Yang, Dehai Zhao, Zhenchang Xing, and Xin Xia. 2023. Distinguishing Look-Alike Innocent and Vulnerable Code by Subtle Semantic Representation Learning and Explanation. CoRR abs/2308.11237 (2023)

  109. [117]

    Shengyi Pan, Jiayuan Zhou, Filipe Roseiro Cogo, Xin Xia, Lingfeng Bao, Xing Hu, Shanping Li, and Ahmed E. Hassan

  110. [118]

    Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong

  111. [119]

    In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023

    CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net

  112. [120]

    Laura Plein, Wendkûuni C Ouédraogo, Jacques Klein, and Tegawendé F Bissyandé. 2024. Automatic generation of test cases based on bug reports: a feasibility study with large language models. In Proceedings of the 2024 IEEE/ACM 46th International Conference on Software Engineerin...

  113. [121]

    Shengyi Pan, Lingfeng Bao, Jiayuan Zhou, Xing Hu, Xin Xia, and Shanping Li. 2024. Towards More Practical Automation of Vulnerability Assessment. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering (Lisbon, Portugal) (ICSE ’24). Association for ...

  114. [122]

    Amrit Poudel, Jinfeng Lin, and Jane Cleland-Huang. 2023. Leveraging Transformer-based Language Models to Automate Requirements Satisfaction Assessment. arXiv preprint arXiv:2312.04463 (2023)

  115. [123]

    Nikitha Rao, Kush Jain, Uri Alon, Claire Le Goues, and Vincent J Hellendoorn. 2023. CAT-LM training language models on aligned code and tests. In 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 409–420

  116. [124]

    Eugenio Parra, Jose Luis de la Vara, and Luis Alonso. 2018. Analysis of requirements quality evolution. In Proceedings of the 40th International Conference on Software Engineering: Companion Proceeedings, ICSE 2018, Gothenburg, Sweden, May 27 - June 03, 2018 . ACM, 199–200

  117. [125]

    Ricardo Silva Peres, Xiaodong Jia, Jay Lee, Keyi Sun, Armando Walter Colombo, and Jose Barata. 2020. Industrial artificial intelligence in industry 4.0-systematic review, challenges and outlook. IEEE access 8 (2020), 220121–220139

  118. [126]

    W. W. Royce. 1987. Managing the Development of Large Software Systems: Concepts and Techniques. In Proceedings, 9th International Conference on Software Engineering, Monterey, California, USA, March 30 - April 2, 1987 . ACM Press, 328–339

  119. [127]

    Chanathip Pornprasit and Chakkrit Tantithamthavorn. 2024. GPT-3.5 for Code Review Automation: How Do Few-Shot Learning, Prompt Design, and Model Fine-Tuning Impact Their Performance? arXiv preprint arXiv:2402.00905 (2024)

  120. [128]

    Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, , Vol

    Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexan- dra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît...

  121. [129]

    Bo Shen, Jiaxin Zhang, Taihong Chen, Daoguang Zan, Bing Geng, An Fu, Muhan Zeng, Ailun Yu, Jichuan Ji, Jingyang Zhao, Yuenan Guo, and Qianxiang Wang. 2023. PanGu-Coder2: Boosting Large Language Models for Code with Ranking Feedback. CoRR abs/2307.14936 (2023)

  122. [130]

    Nguyen, and Riccardo Rubei

    Juri Di Rocco, Davide Di Ruscio, Claudio Di Sipio, Phuong T. Nguyen, and Riccardo Rubei. 2024. On the use of Large Language Models in Model-Driven Engineering. CoRR abs/2410.17370 (2024)

  123. [131]

    Krishna Ronanki, Beatriz Cabrero-Daniel, and Christian Berger. 2022. ChatGPT as a Tool for User Story Quality Evaluation: Trustworthy Out of the Box?. InInternational Conference on Agile Software Development. Springer, 173–181

  124. [132]

    Mohammed Latif Siddiq, Joanna Cecilia Da Silva Santos, Ridwanul Hasan Tanvir, Noshin Ulfat, Fahmid Al Rifat, and Vinícius Carvalho Lopes. 2024. Using large language models to generate junit tests: An empirical study. InProceedings of the 28th International Conference on Evalua...

  125. [133]

    Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950 (2023)

  126. [134]

    Giriprasad Sridhara, Sourav Mazumdar, et al. 2023. Chatgpt: A study on its utility for ubiquitous software engineering tasks. arXiv preprint arXiv:2305.16837 (2023)

  127. [135]

    Benjamin Steenhoek, Michele Tufano, Neel Sundaresan, and Alexey Svyatkovskiy. 2023. Reinforcement learning from automatic feedback for high-quality unit test generation. arXiv preprint arXiv:2310.02368 (2023)

  128. [136]

    Jiho Shin, Sepehr Hashtroudi, Hadi Hemmati, and Song Wang. 2024. Domain Adaptation for Code Model-Based Unit Test Case Generation. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis. 1211–1222

  129. [137]

    Seung Yeob Shin, Fabrizio Pastore, Domenico Bianculli, and Alexandra Baicoianu. 2024. Towards generating executable metamorphic relations using large language models. In International Conference on the Quality of Information and Communications Technology. Springer, 126–141

  130. [138]

    Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan. 2020. Intellicode compose: Code generation using transformer. In Proceedings of the 28th ACM joint meeting on European software engineering conference and symposium on the foundations of software engineering ...

  131. [139]

    Mehmet Söylemez, Bedir Tekinerdogan, and Ayça Kolukısa Tarhan. 2022. Challenges and solution directions of microservice architectures: A systematic literature review. Applied sciences 12, 11 (2022), 5507

  132. [140]

    Christopher Thompson and David A. Wagner. 2017. A Large-Scale Study of Modern Code Review and Security in Open Source Projects. In PROMISE. ACM, 83–92

  133. [141]

    Patanamon Thongtanunam, Chanathip Pornprasit, and Chakkrit Tantithamthavorn. 2022. Autotransform: Automated code transformation to support modern code review process. In Proceedings of the 44th international conference on software engineering. 237–248

  134. [142]

    Robert Boyce Stone. 1997. Towards a theory of modular design . The University of Texas at Austin

  135. [143]

    Zhihong Sun, Chen Lyu, Bolun Li, Yao Wan, Hongyu Zhang, Ge Li, and Zhi Jin. 2024. Enhancing Code Generation Performance of Smaller Models by Distilling the Reasoning Ability of LLMs. arXiv preprint arXiv:2403.13271 (2024)

  136. [144]

    Rosalia Tufano. 2023. Automating Code Review. In ICSE Companion Proceedings. IEEE, 192–196

  137. [145]

    Sumanth Tatineni. 2023. Applying DevOps Practices for Quality and Reliability Improvement in Cloud-Based Systems. Technix international journal for engineering research (TIJER) 10, 11 (2023), 374–380

  138. [146]

    Rosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella, Denys Poshyvanyk, and Gabriele Bavota

  139. [147]

    Rosalia Tufano, Luca Pascarella, Michele Tufano, Denys Poshyvanyk, and Gabriele Bavota. 2021. Towards automating code review activities. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 163–174

  140. [148]

    Zhao Tian, Junjie Chen, and Xiangyu Zhang. 2023. Test-case-driven programming understanding in large language models for better code generation. arXiv preprint arXiv:2309.16120 (2023)

  141. [149]

    Christos Tsigkanos, Pooja Rani, Sebastian Müller, and Timo Kehrer. 2023. Large language models: The next frontier for variable discovery within metamorphic testing?. In 2023 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER). IEEE, 678–682

  142. [150]

    Chaozheng Wang, Yuanhang Yang, Cuiyun Gao, Yun Peng, Hongyu Zhang, and Michael R Lyu. 2023. Prompt Tuning in Code Intelligence: An Experimental Evaluation. IEEE Transactions on Software Engineering (2023)

  143. [151]

    Rosalia Tufano, Ozren Dabić, Antonio Mastropaolo, Matteo Ciniselli, and Gabriele Bavota. 2024. Code review automation: strengths and weaknesses of the state of the art. IEEE Transactions on Software Engineering (2024)

  144. [152]

    Deze Wang, Boxing Chen, Shanshan Li, Wei Luo, Shaoliang Peng, Wei Dong, and Xiangke Liao. 2023. One Adapter for All Programming Languages? Adapter Tuning for Code Search and Summarization. In ICSE. IEEE, 5–16

  145. [153]

    In Proceedings of the 44th international conference on software engineering

    Using pre-trained models to boost code review automation. In Proceedings of the 44th international conference on software engineering. 2291–2302

  146. [154]

    Yue Wang, Hung Le, Akhilesh Deepak Gotmare, Nghi DQ Bui, Junnan Li, and Steven Hoi. 2023. CodeT5+: Open Code Large Language Models for Code Understanding and Generation. In The 2023 Conference on Empirical Methods in Natural Language Processing

  147. [155]

    Chaozheng Wang, Junhao Hu, Cuiyun Gao, Yu Jin, Tao Xie, Hailiang Huang, Zhenyu Lei, and Yuetang Deng. 2023. How Practitioners Expect Code Completion?. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Eng...

  148. [156]

    Chaozheng Wang, Yuanhang Yang, Cuiyun Gao, Yun Peng, Hongyu Zhang, and Michael R. Lyu. 2022. No more fine-tuning? an experimental evaluation of prompt tuning in code intelligence. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on th...

  149. [157]

    Zefan Wang, Zichuan Liu, Yingying Zhang, Aoxiao Zhong, Lunting Fan, Lingfei Wu, and Qingsong Wen. 2023. RCAgent: Cloud Root Cause Analysis by Autonomous Agents with Tool-Augmented Large Language Models. CoRR abs/2310.16340 (2023)

  150. [158]

    Chong Wang, Jian Zhang, Yebo Feng, Tianlin Li, Weisong Sun, Yang Liu, and Xin Peng. 2024. Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code Generation. arXiv preprint arXiv:2401.06391 (2024)

  151. [159]

    Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, and Lingming Zhang. 2023. Magicoder: Source code is all you need. arXiv preprint arXiv:2312.02120 (2023)

  152. [160]

    Jiexin Wang, Xitong Luo, Liuwen Cao, Hongkui He, Hailin Huang, Jiayuan Xie, Adam Jatowt, and Yi Cai. 2024. Is Your AI-Generated Code Really Secure? Evaluating Large Language Models on Secure Code Generation with CodeSecEval. arXiv preprint arXiv:2407.02395 (2024)

  153. [161]

    Xin-Cheng Wen, Cuiyun Gao, Jiaxin Ye, Yichen Li, Zhihong Tian, Yan Jia, and Xuan Wang. 2024. Meta-Path Based Attentional Graph Learning Model for Vulnerability Detection. IEEE Trans. Software Eng. 50, 3 (2024), 360–375

  154. [162]

    Yawen Wang, Lin Shi, Mingyang Li, Qing Wang, and Yun Yang. 2020. A deep context-wise method for coreference detection in natural language requirements. In 2020 IEEE 28th International Requirements Engineering Conference (RE) . IEEE, 180–191

  155. [163]

    Joty, and Steven C

    Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi. 2021. CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, ...

  156. [164]

    Xin-Cheng Wen, Cuiyun Gao, Shuzheng Gao, Yang Xiao, and Michael R. Lyu. 2024. SCALE: Constructing Structured Natural Language Comment Trees for Software Vulnerability Detection. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (Vi...

  157. [165]

    Xin-Cheng Wen, Xinchen Wang, Cuiyun Gao, Shaohua Wang, Yang Liu, and Zhaoquan Gu. 2024. When Less is Enough: Positive and Unlabeled Learning Model for Vulnerability Detection. In Proceedings of the 38th IEEE/ACM International Conference on Automated Software Engineering (Echte...

  158. [166]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837

  159. [167]

    Chunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel, and Lingming Zhang. 2024. Fuzz4all: Universal fuzzing with large language models. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering. 1–13

  160. [168]

    Cheng Wen, Yuandao Cai, Bin Zhang, Jie Su, Zhiwu Xu, Dugang Liu, Shengchao Qin, Zhong Ming, and Tian Cong

  161. [169]

    Automatically inspecting thousands of static bug warnings with large language model: how far are we? ACM Transactions on Knowledge Discovery from Data 18, 7 (2024), 1–34

  162. [170]

    Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244 (2023)

  163. [171]

    Xin-Cheng Wen, Xinchen Wang, Yujia Chen, Ruida Hu, David Lo, and Cuiyun Gao. 2024. VulEval: Towards Repository- Level Evaluation of Software Vulnerability Detection. CoRR abs/2404.15596 (2024)

  164. [172]

    Zhang, and Qing Liao

    Xin-Cheng Wen, Yupan Chen, Cuiyun Gao, Hongyu Zhang, Jie M. Zhang, and Qing Liao. 2023. Vulnerability Detection with Graph Simplification and Enhanced Graph Representation Learning. In Proceedings of the 45th International Conference on Software Engineering (Melbourne, Victori...

  165. [173]

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze...

  166. [174]

    De Roeck, and Bashar Nuseibeh

    Hui Yang, Alistair Willis, Anne N. De Roeck, and Bashar Nuseibeh. [n. d.]. Automatic detection of nocuous coordination ambiguities in natural language requirements. In ASE 2010, 25th IEEE/ACM International Conference on Automated Software Engineering, Antwerp, Belgium, Septemb...

  167. [175]

    Jules White, Sam Hays, Quchen Fu, Jesse Spencer-Smith, and Douglas C. Schmidt. 2023. ChatGPT Prompt Patterns for Improving Code Quality, Refactoring, Requirements Elicitation, and Software Design. CoRR abs/2303.07839 (2023)

  168. [176]

    Yau and Jeffrey J

    Stephen S. Yau and Jeffrey J. P. Tsai. 1986. A Survey of Software Design Techniques. IEEE Trans. Software Eng. 12, 6 (1986), 713–721

  169. [177]

    Danning Xie, Byungwoo Yoo, Nan Jiang, Mijung Kim, Lin Tan, Xiangyu Zhang, and Judy S Lee. 2023. Impact of large language models on generating software specifications. arXiv preprint arXiv:2306.03324 (2023)

  170. [178]

    Zhuokui Xie, Yinghao Chen, Chen Zhi, Shuiguang Deng, and Jianwei Yin. 2023. ChatUniTest: a ChatGPT-based automated unit test generation tool. arXiv preprint arXiv:2305.04764 (2023). , Vol. 1, No. 1, Article . Publication date: December 2018. 30 Gao et al

  171. [179]

    Zhaojian Yu, Xin Zhang, Ning Shang, Yangyu Huang, Can Xu, Yishujie Zhao, Wenxiang Hu, and Qiufeng Yin. 2023. Wavecoder: Widespread and versatile enhanced instruction tuning with refined data generation. arXiv preprint arXiv:2312.14187 (2023)

  172. [180]

    Junjielong Xu, Ziang Cui, Yuan Zhao, Xu Zhang, Shilin He, Pinjia He, Liqun Li, Yu Kang, Qingwei Lin, Yingnong Dang, Saravan Rajmohan, and Dongmei Zhang. 2024. UniLog: Automatic Logging via LLM and In-Context Learning. In Proceedings of the 46th IEEE/ACM International Conferenc...

  173. [181]

    Hao Yan, Thomas D Latoza, and Ziyu Yao. 2024. IntelliExplain: Enhancing Interactive Code Generation through Natural Language Explanations for Non-Professional Programmers. arXiv preprint arXiv:2405.10250 (2024)

  174. [182]

    Las-Casas, Rodrigo Fonseca, and Saravan Rajmohan

    Dylan Zhang, Xuchao Zhang, Chetan Bansal, Pedro Henrique B. Las-Casas, Rodrigo Fonseca, and Saravan Rajmohan

  175. [183]

    Fengji Zhang, Bei Chen, Yue Zhang, Jacky Keung, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, and Weizhu Chen. 2023. Repocoder: Repository-level code completion through iterative retrieval and generation. arXiv preprint arXiv:2303.12570 (2023)

  176. [184]

    Zezhou Yang, Cuiyun Gao, Zhaoqiang Guo, Zhenhao Li, Kui Liu, Xin Xia, and Yuming Zhou. 2024. A Survey on Modern Code Review: Progresses, Challenges and Opportunities. CoRR abs/2405.18216 (2024)

  177. [185]

    Junwei Zhang, Zhongxin Liu, Xing Hu, Xin Xia, and Shanping Li. 2023. Vulnerability Detection by Learning From Syntax-Based Execution Paths of Code. IEEE Trans. Software Eng. 49, 8 (2023), 4196–4212

  178. [186]

    Juyeon Yoon, Robert Feldt, and Shin Yoo. 2023. Autonomous Large Language Model Agents Enabling Intent-Driven Mobile GUI Testing. arXiv preprint arXiv:2311.08649 (2023)

  179. [187]

    Shengcheng Yu, Chunrong Fang, Yuchen Ling, Chentian Wu, and Zhenyu Chen. 2023. LLM for Test Script Generation and Migration: Challenges, Capabilities, and Opportunities. In 23rd IEEE International Conference on Software Quality, Reliability, and Security, QRS 2023, Chiang Mai,...

  180. [188]

    Renrui Zhang, Jiaming Han, Chris Liu, Aojun Zhou, Pan Lu, Yu Qiao, Hongsheng Li, and Peng Gao. 2024. LLaMA- Adapter: Efficient Fine-tuning of Large Language Models with Zero-initialized Attention. In ICLR. OpenReview.net

  181. [189]

    Zhiqiang Yuan, Yiling Lou, Mingwei Liu, Shiji Ding, Kaixin Wang, Yixuan Chen, and Xin Peng. 2023. No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation. CoRR abs/2305.04207 (2023)

  182. [190]

    Daoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Yongji Wang, and Jian-Guang Lou

  183. [191]

    In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023

    Large Language Models Meet NL2Code: A Survey. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023 . Association for Computational Linguistics, 7443–7464

  184. [192]

    Jiayuan Zhou, Michael Pacheco, Jinfu Chen, Xing Hu, Xin Xia, David Lo, and Ahmed E. Hassan. 2023. CoLeFunDa: Explainable Silent Vulnerability Fix Identification. In Proceedings of the 45th International Conference on Software Engineering (Melbourne, Victoria, Australia) (ICSE ...

  185. [193]

    CoRR abs/2309.05833 (2023)

    PACE-LM: Prompting and Augmentation for Calibrated Confidence Estimation with GPT-4 in Cloud Incident Root Cause Analysis. CoRR abs/2309.05833 (2023)

  186. [195]

    Fengji Zhang, Zexian Zhang, Jacky Wai Keung, Xiangru Tang, Zhen Yang, Xiao Yu, and Wenhua Hu. 2024. Data preparation for deep learning based code smell detection: A systematic literature review. Journal of Systems and Software (2024), 112131

  187. [197]

    Kechi Zhang, Jia Li, Ge Li, Xianjie Shi, and Zhi Jin. 2024. Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges. arXiv preprint arXiv:2401.07339 (2024)

  188. [198]

    Kechi Zhang, Huangzhao Zhang, Ge Li, Jia Li, Zhuo Li, and Zhi Jin. 2023. Toolcoder: Teach code generation models to use api search tools. arXiv preprint arXiv:2305.04032 (2023)

  189. [200]

    Simiao Zhang, Jiaping Wang, Guoliang Dong, Jun Sun, Yueling Zhang, and Geguang Pu. 2024. Experimenting a New Programming Practice with LLMs. arXiv preprint arXiv:2401.01062 (2024). , Vol. 1, No. 1, Article . Publication date: December 2018. The Current Challenges of Software E...

  190. [201]

    Yifan Zhang, Dave Towey, and Matthew Pike. 2023. Automated Metamorphic-Relation Generation with ChatGPT: An Experience Report. In 2023 IEEE 47th Annual Computers, Software, and Applications Conference (COMPSAC) . IEEE, 1780–1785

  191. [202]

    Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Lei Shen, Zihan Wang, Andi Wang, Yang Li, Teng Su, Zhilin Yang, and Jie Tang. 2023. CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X. In Proceedings of the 29th AC...

  192. [204]

    Xin Zhou, Kisub Kim, Bowen Xu, DongGyun Han, Junda He, and David Lo. 2023. Generation-based code review automation: how far are weƒ. In 2023 IEEE/ACM 31st International Conference on Program Comprehension (ICPC) . IEEE, 215–226. , Vol. 1, No. 1, Article . Publication date: Dec...

  193. [2021]

    Association for Computational Linguistics, 8696–8708

  194. [2022]

    In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (Singapore, Singapore) (ESEC/FSE 2022)

    Automated unearthing of dangerous issue reports. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (Singapore, Singapore) (ESEC/FSE 2022). Association for Computing Machinery, New York, NY, US...

  195. [2023]

    In 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023

    Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language Models. In 45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023 . IEEE, 1737–1749

  196. [2024]

    Journal of Systems and Software 214 (2024), 112059

    GPTSniffer: A CodeBERT-based classifier to detect source code written by ChatGPT. Journal of Systems and Software 214 (2024), 112059

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.