REVIEW 5 major objections 6 minor 1 cited by
A Methodological Framework for LLM-Based Mining of Software Repositories
T0 review · 5 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read LLM-based repository mining has settled into 15 recognizable approaches and nine recurring threats, and the paper's six-stage PRIMES 2.0 framework turns them into a prescriptive and diagnostic tool for rigorous, reproducible studies.
desk verdict A useful, flawed consolidation of LLM4MSR practice; the framework is a good starting point, but the empirical base is thinner than the presentation suggests. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrier of the argument is PRIMES 2.0 itself: a six-stage, 23-substep workflow — Strategic Planning & Preparation, Creation of Prompts for Piloting, Prompt Validation, LLM Setup, LLM Comparison, and Execution, Validation & Synthesis — in which every substep is annotated with the threats it risks awakening and the mitigations that address them. It is the mechanism that converts the empirical findings (approaches A1–A15, threats T1–T9, mitigations M1–M26) into actionable guidance, and it is the artifact that would need to prove its worth through wider application for the paper's contribution to hold.
What would settle it
Run a full systematic review of the same period with broader venue coverage, or a survey with a larger probability-based sample, and check whether any methodological approach or threat appears that does not fit into PRIMES 2.0's six stages, 23 substeps, and nine-threat map; one new stage-level decision or one new threat would show the framework describes a small corpus rather than the field.
Extended reading notes
Core claim
On its own terms, the paper's discovery is the structure of an emerging research practice. Synthesizing 15 peer-reviewed papers and 22 experienced practitioners, it claims that LLM-based repository mining has converged on 15 recognizable methodological approaches — adopting or lacking guidelines, defining goals, curating datasets, automating and tracking the pipeline, choosing prompting strategies, selecting models, iteratively piloting and validating prompts with metrics, configuring and reporting model parameters, considering ensembles, benchmarking models, executing at scale, synthesizing results, and releasing replication packages. It further claims that these approaches organize into six workflow stages and that failures recur as nine threats: hallucinations and output inaccuracy, prompt sensitivity, contextual understanding and domain specificity, scalability and cost, reproducibility and output stability, output formatting and parsing, dataset noise and pretraining leakage, validation complexity, and lack of standardized tooling. The discovery culminates in PRIMES 2.0, which maps each of 23 substeps to the threats (T1–T9) it incurs and the mitigations (M1–M26) that counter them, making the framework at once a design checklist and a diagnostic instrument.
Load-bearing premise
The framework is only as general as its evidence: 15 peer-reviewed papers and 22 volunteer survey respondents are assumed to exhaust the space of how researchers mine repositories with LLMs, so that the 15 approaches, nine threats, and six-stage structure carry over to the field at large.
Editorial extensions
If this is right
- Researchers planning an LLM-based mining study could use PRIMES 2.0 as a design checklist, building in mitigations for the nine threats from the start rather than discovering them after execution.
- Reviewers and readers gain a diagnostic instrument: a paper can be audited substep by substep to see which threats were left unmitigated, which configurations went undocumented, and whether a replication package was released.
- The study's own numbers expose a dissemination gap — 2 of 15 papers and 10 of 22 respondents untethered to any guideline — so the framework's first-order effect would be to make methodological standards visible and concrete.
- The framework would push the field toward explicit reporting of model versions, temperature settings, prompt templates, validation samples, and discard rates, which is precisely the transparency the study found missing.
- Because threats such as hallucinations, prompt sensitivity, and validation complexity are not MSR-specific, the six-stage structure is claimed to transfer to other kinds of LLM-based empirical software engineering research.
Reading between the lines
- A testable extension the authors leave implicit: operationalize PRIMES 2.0 as a scoring rubric, have independent raters apply its 23 substeps to a sample of published LLM-based mining studies, and check whether higher adherence correlates with replicability or fewer reported threats.
- A saturation check is the fairest test of the framework's foundations: a fuller systematic review or a larger probability-sampled survey would show whether new approaches or threats appear beyond the 15 and nine reported here — something the paper's own evidence base of 15 papers and 22 respondents cannot guarantee.
- The framework's portability claim, that its stages extend beyond MSR into test generation, code summarization, or requirements classification, is directly checkable by applying the six stages in those adjacent settings and observing whether the threat map still fits.
- Readers using PRIMES 2.0 as a checklist should reconcile the paper's internal counts first: the abstract reports 25 mitigation strategies while the framework table lists M1–M26, and the conclusion speaks of 14 approaches where the analysis reports 15.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a mixed-method study of how large language models (LLMs) are used in Mining Software Repositories (MSR) research. The authors combine a rapid review of 15 peer-reviewed papers (screened from 6,799 in nine venues) with a questionnaire survey of 22 researchers (from 143 invited) to identify methodological approaches, threats to validity, and mitigation strategies. They derive 15 approaches (RQ1), nine threats (RQ2), and a set of mitigations, which they consolidate into PRIMES 2.0, a six-stage framework with 23 substeps that maps threats T1–T9 to mitigation strategies M1–M26. The abstract and conclusion list slightly different counts (25 mitigations in the abstract, 14 approaches in the conclusion), and the empirical base is small and partly convenience-sampled. The paper's central claim is that PRIMES 2.0 is an empirically grounded, prescriptive, and diagnostic instrument for designing rigorous LLM-based MSR studies.
Significance. If the empirical base is accepted as saturating the space of current LLM4MSR practice, the contribution is genuinely useful: it provides a structured, threat-aware framework that researchers can adopt for study design and reporting, and it ships a replication package and public dataset. The threat-mitigation table (Table 4) is a practical synthesis that goes beyond high-level guidelines by operationalizing mitigations with concrete techniques. The paper also explicitly reports verbatim respondent quotes and Cohen's Kappa agreement for study selection, which strengthens the credibility of the descriptive findings. However, the framework's prescriptive value depends on the completeness of the underlying evidence, and the internal inconsistencies in the reported counts reduce confidence in the quantitative claims.
major comments (5)
- [Abstract and Section 7] The counts reported in the abstract and conclusion are inconsistent with the body: the abstract states 25 mitigation strategies, Section 4.2 and Table 4 list 26 (M1–M26), and Section 7 says the review and survey explored '14 approaches' while Section 4.1 reports 15 approaches (A1–A15). Since the paper's core contribution is a set of empirically grounded artifacts whose identity is defined by these counts, these numbers must be reconciled.
- [Section 3.2.3 and Figure 1] The sample-size accounting is internally inconsistent. The text says '37 valid responses were collected and retained for analysis,' then that one participant failed the attention check and was removed, and then that '22 respondents confirmed their experience.' Figure 1, in contrast, shows '37 Participants,' 'Removed 15 Participants for not using LLM,' and '22 Accepted Participants.' These descriptions do not add up; the reader cannot determine how many participants were invited, how many started, how many failed screening, and how many were finally analyzed. Please provide a clear, consistent flow diagram and textual account.
- [Section 3.2.3 and Section 6.3] The paper claims that survey responses were collected 'until no substantially new themes emerged,' invoking theoretical saturation, but no saturation analysis is reported for the review, the survey, or the combined approach/threat/mitigation codebook. With 15 papers and 22 convenience-sampled respondents, the claim that the derived 15 approaches, 9 threats, and 26 mitigations cover the actual space of LLM4MSR practice is load-bearing. Either provide a saturation analysis (e.g., a plot of new codes per additional paper/respondent) or substantially soften the generalizability claims for PRIMES 2.0.
- [Section 3.2.3 and Section 6] The recruitment strategy explicitly included contacting the corresponding authors of the papers retained in the rapid review and researchers from the authors' professional networks. This creates a non-trivial overlap between the survey population and the reviewed literature, yet Section 6.3 mentions only convenience sampling and does not discuss the risk that survey responses are not independent from the papers used in the review. The triangulation claim in Section 6.4 would be strengthened by an explicit treatment of this overlap and its effect on the derived threats and mitigations.
- [Section 5.4 and Figure 5] The paper states that PRIMES 2.0 comprises '23 methodological substeps, each mapped to specific threats (T1–T9) and corresponding mitigation strategies (M1–M26),' but the 23 substeps are never enumerated in the text, and Figure 5 does not make the per-substep mapping explicit. The reader cannot verify the claimed one-to-one mapping. Please add a table or structured list that names each of the 23 substeps and shows the threat(s) and mitigation(s) mapped to each.
minor comments (6)
- [Section 2.2] The text attributes views to 'Kausar et al. [2]', but reference [2] is 'Ahmed et al. (2025)' and no Kausar entry appears in the reference list; please correct the citation or add the intended source.
- [Figure 1] The label 'Reseraches' is a typo; it should be 'Researchers.'
- [Table 2] In exclusion criterion (EC6), 'exclucion' should be 'exclusion.'
- [Table 4] In mitigation M18, 'Exclucion' should be 'Exclusion.'
- [Section 3] The GQM description contains a duplicated 'of': 'from the perspective of of SE researchers' should be 'from the perspective of SE researchers.'
- [Section 3.2.3] The sentence 'In total, 37 valid responses were collected and retained for analysis' contradicts the later statement that responses failing the attention check were removed; consider rephrasing to distinguish collected from retained.
Circularity Check
No significant circularity: the empirical findings come from an external review and survey, and the self-citations in the framework are illustrative rather than load-bearing.
full rationale
The paper's central derivation chain is not circular. The 15 methodological approaches, nine threats, and 26 mitigation strategies are presented as findings from a rapid review of 15 peer-reviewed articles and a questionnaire survey of 22 researchers (Sections 3 and 4). These are external empirical inputs, not the framework itself. PRIMES 2.0 is explicitly built as an extension of the authors' earlier PRIMES framework, and the paper openly states that the earlier framework 'was based only on two of our own previous LLM-based MSR studies' and was therefore preliminary (Section 2.1). Section 4.1 also transparently acknowledges that the stage organization 'builds upon our earlier conceptualization' rather than pretending the structure emerged de novo from the data. The only self-referential element is mitigation M5, which mentions 'nascent academic guidelines like our PRIMES framework' as one example of adopting de-facto and emerging standards (Section 4.2, T2). That citation is an illustrative instance inside a broader mitigation strategy and is not load-bearing for the validity of PRIMES 2.0. The acknowledged convenience sampling and limited statistical generalizability (Section 6.3) are external-validity limitations, not circularity. Internal numeric inconsistencies (e.g., '14 approaches' in Section 7 vs. 15 elsewhere; '25 mitigation strategies' in the abstract vs. 26 in Section 4.2) are reporting errors, not evidence of circular reasoning. No fitted parameter is renamed as a prediction, no uniqueness result is imported from the authors' prior work, and no known empirical pattern is merely relabeled.
Assumptions & free parameters
free parameters (1)
- Cohen's Kappa acceptance threshold =
0.75
assumptions (5)
- domain assumption The 15 selected papers and 22 survey respondents adequately represent current LLM4MSR practice
- domain assumption Qualitative content analysis by the first and second authors yields valid thematic categories
- domain assumption Threats and mitigations named by a minority of sources are treated as field-level findings
- ad hoc to paper Six stages and 23 substeps are the correct decomposition of the LLM4MSR workflow
- domain assumption Triangulation between rapid review and survey compensates for each method's bias
invented entities (2)
-
PRIMES 2.0 framework (six stages, 23 substeps)
-
Threat-mitigation mapping (T1-T9 paired with M1-M26)
Cite this review
Pith. "Pith review of A Methodological Framework for LLM-Based Mining of Software Repositories." pith.science (2026). https://pith.science/paper/PZXM32QA
@misc{pith2026250802233,
author = {Pith},
title = {Pith review of: A Methodological Framework for LLM-Based Mining of Software Repositories},
year = {2026},
howpublished = {\url{https://pith.science/paper/PZXM32QA}},
note = {Machine review of arXiv:2508.02233}
}
read the original abstract
Large Language Models (LLMs) are increasingly used in software engineering research, offering new opportunities for automating repository mining tasks. However, despite their growing popularity, the methodological integration of LLMs into Mining Software Repositories (MSR) remains poorly understood. Existing studies tend to focus on specific capabilities or performance benchmarks, providing limited insight into how researchers utilize LLMs across the full research pipeline. To address this gap, we conduct a mixed-method study that combines a rapid review and questionnaire survey in the field of LLM4MSR. We investigate (1) the approaches and (2) the threats that affect the empirical rigor of researchers involved in this field. Our findings reveal 15 methodological approaches, nine main threats, and 25 mitigation strategies. Building on these findings, we present PRIMES 2.0, a refined empirical framework organized into six stages, comprising 23 methodological substeps, each mapped to specific threats and corresponding mitigation strategies, providing prescriptive and adaptive support throughout the lifecycle of LLM-based MSR studies. Our work contributes to establishing a more transparent and reproducible foundation for LLM-based MSR research.
Figures
Forward citations
Cited by 1 Pith paper
-
The Rise of Language Models in Mining Software Repositories: A Survey
Across 85 papers, language models in MSR mostly perform classification and generation on issues, code reviews, and commits, with the field shifting from fine-tuned BERT-size models to large instruction-tuned LLMs used...
Reference graph
Works this paper leans on
-
[1]
Samuel Abedu, Laurine Menneron, SayedHassan Khatoonabadi, and Emad Shihab. 2025. RepoChat: An LLM-Powered Chatbot for GitHub Repository Question-Answering. In 2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR) . IEEE, 255–259. Manuscript submitted to ACM A Methodological Framework for LLM-Based Mining of Software Repositories 43
work page 2025
-
[2]
Toufique Ahmed, Premkumar Devanbu, Christoph Treude, and Michael Pradel. 2025. Can llms replace manual annotation of software engineering artifacts?. In 2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR) . IEEE, 526–538
work page 2025
-
[3]
Dharun Anandayuvaraj, Matthew Campbell, Arav Tewari, and James C Davis. 2024. FAIL: Analyzing Software Failures from the News Using LLMs. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 506–518
work page 2024
-
[4]
Sebastian Baltes and Paul Ralph. 2022. Sampling in software engineering research: A critical review and guidelines. Empirical Software Engineering 27, 4 (2022), 94
work page 2022
-
[5]
Abdul Ali Bangash, Hareem Sahar, Shaiful Chowdhury, Alexander William Wong, Abram Hindle, and Karim Ali. 2019. What do developers know about machine learning: a study of ml discussions on stackoverflow. In 2019 IEEE/ACM 16th International Conference on Mining Software Repositories (MSR). IEEE, 260–264
work page 2019
-
[6]
Cauã Ferreira Barros, Bruna Borges Azevedo, Valdemar Vicente Graciano Neto, Mohamad Kassab, Marcos Kalinowski, Hugo Alexandre D Do Nasci- mento, and Michelle CGSP Bandeira. 2025. Large Language Model for Qualitative Research: A Systematic Mapping Study. In 2025 IEEE/ACM International Workshop on Methodological Issues with Empirical Studies in Software Eng...
work page 2025
-
[7]
João Helis Bernardo, Daniel Alencar Da Costa, Sérgio Queiroz de Medeiros, and Uirá Kulesza. 2024. How do machine learning projects use continuous integration practices? an empirical study on GitHub actions. In Proceedings of the 21st International Conference on Mining Software Repositories . 665–676
2024
-
[8]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901
2020
Show all 70 references
-
[9]
Victor R Basili1 Gianluigi Caldiera and H Dieter Rombach. 1994. The goal question metric approach. Encyclopedia of software engineering (1994), 528–532
1994
-
[10]
Fabio Calefato, Filippo Lanubile, and Luigi Quaranta. 2022. A preliminary investigation of MLOps practices in GitHub. In Proceedings of the 16th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement . 283–288
2022
-
[11]
Bruno Cartaxo, Gustavo Pinto, and Sergio Soares. 2020. Rapid reviews in software engineering. Contemporary Empirical Methods in Software Engineering (2020), 357–384
2020
-
[12]
Joel Castaño, Rafael Cabañas, Antonio Salmerón, David Lo, and Silverio Martínez-Fernández. 2024. How do machine learning models change? arXiv preprint arXiv:2411.09645 (2024)
2024
-
[13]
Joel Castaño, Silverio Martínez-Fernández, Xavier Franch, and Justus Bogner. 2023. Exploring the carbon footprint of hugging face’s ml models: A repository mining study. In 2023 ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM) . IEEE, 1–12
2023
-
[14]
Preetha Chatterjee, Mia Mohammad Imran, and Kostadin Damevski. 2024. Shedding Light on Software Engineering-specific Metaphors and Idioms. In ICSE’24: Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . Association for Computing Machinery, 1–13
2024
-
[15]
Zhenpeng Chen, Yanbin Cao, Yuanqiang Liu, Haoyu Wang, Tao Xie, and Xuanzhe Liu. 2020. A comprehensive study on challenges in deploying deep learning based software. In Proceedings of the 28th ACM joint meeting on European software engineering conference and symposium on the fo...
2020
-
[16]
Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educational and psychological measurement 20, 1 (1960), 37–46
1960
-
[17]
Giuseppe Colavito, Filippo Lanubile, and Nicole Novielli. 2025. Benchmarking large language models for automated labeling: The case of issue report classification. Information and Software Technology (2025), 107758
2025
-
[18]
Marcelo Costalonga, Bianca Minetto Napoleão, Maria Teresa Baldassarre, Katia Romero Felizardo, Igor Steinmacher, and Marcos Kalinowski. 2025. Can Machine Learning Support the Selection of Studies for Systematic Literature Review Updates? arXiv preprint arXiv:2502.08050 (2025)
2025 arXiv
-
[19]
Daniela S Cruzes and Tore Dybå. 2011. Recommended steps for thematic synthesis in software engineering. International Symposium on Empirical Software Engineering and Measurement (ESEM) (2011), 275–284
2011
-
[20]
Vincenzo De Martino, Joel Castaño, Fabio Palomba, Xavier Franch, and Silverio Martínez-Fernández. 2025. A framework for using llms for repository mining studies in empirical software engineering. In 2025 IEEE/ACM International Workshop on Methodological Issues with Empirical S...
2025
-
[21]
Vincenzo De Martino, Silverio Martínez-Fernández, and Fabio Palomba. 2024. Do developers adopt green architectural tactics for ml-enabled systems? a mining software repository study. arXiv preprint arXiv:2410.06708 (2024)
2024 arXiv
-
[22]
Vincenzo De Martino and Fabio Palomba. 2025. Classification and challenges of non-functional requirements in ML-enabled systems: A systematic literature review. Information and Software Technology (2025), 107678
2025
-
[23]
Vincenzo De Martino, Gilberto Recupito, Giammaria Giordano, Filomena Ferrucci, Dario Di Nucci, and Fabio Palomba. 2025. Into the ML- Universe: An improved classification and characterization of machine-learning projects. Journal of Systems and Software 230 (2025), 112471. http...
2025
-
[24]
Don A Dillman, Jolene D Smyth, and Leah Melani Christian. 2014. Internet, phone, mail, and mixed-mode surveys: The tailored design method . John Wiley & Sons
2014
-
[25]
Yihong Dong, Xue Jiang, Zhi Jin, and Ge Li. 2024. Self-collaboration code generation via chatgpt. ACM Transactions on Software Engineering and Methodology 33, 7 (2024), 1–38
2024
-
[26]
Angela Fan, Beliz Gokkaya, Mark Harman, Mitya Lyubarskiy, Shubho Sengupta, Shin Yoo, and Jie M Zhang. 2023. Large language models for software engineering: Survey and open problems. In 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineeri...
2023
-
[27]
Katia Romero Felizardo, Anderson Deizepe, Daniel Coutinho, Genildo Gomes, Maria Meireles, Marco Gerosa, and Igor Steinmacher. 2025. On the Difficulties of Conducting and Replicating Systematic Literature Reviews Studies Using LLMs in Software Engineering. In 2025 IEEE/ACM Inte...
2025
-
[28]
Alexandra González, Xavier Franch, David Lo, and Silverio Martínez-Fernández. 2025. How do Pre-Trained Models Support Software Engineering? An Empirical Study in Hugging Face. arXiv preprint arXiv:2506.03013 (2025)
2025
-
[29]
Joseph F Hair, Arthur H Money, Philip Samouel, and Mike Page. 2007. Research methods for business. Education+ Training 49, 4 (2007), 336–337
2007
-
[30]
Abram Hindle, Daniel M German, and Ric Holt. 2008. What do large commits tell us? a taxonomical study of large commits. In Proceedings of the 2008 international working conference on Mining software repositories . 99–108
2008
-
[31]
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. 2024. Large language models for software engineering: A systematic literature review. ACM Transactions on Software Engineering and Methodology 33, 8 (2024), 1–79
2024
-
[32]
Mia Mohammad Imran, Preetha Chatterjee, and Kostadin Damevski. 2024. Uncovering the causes of emotions in software developer communication using zero-shot llms. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . 1–13
2024
-
[33]
Wenxin Jiang, Jerin Yasmin, Jason Jones, Nicholas Synovic, Jiashen Kuo, Nathaniel Bielanski, Yuan Tian, George K Thiruvathukal, and James C Davis. 2024. PeaTMOSS: A Dataset and Initial Analysis of Pre-Trained Models in Open-Source Software. In Proceedings of the 21st Internati...
2024
-
[34]
R Burke Johnson, Anthony J Onwuegbuzie, and Lisa A Turner. 2007. Toward a definition of mixed methods research. Journal of mixed methods research 1, 2 (2007), 112–133
2007
-
[35]
Minseo Kim, Taemin Kim, Thu Hoang Anh Vo, Yugyeong Jung, and Uichin Lee. 2025. Exploring Modular Prompt Design for Emotion and Mental Health Recognition. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems . 1–18
2025
-
[36]
Barbara Kitchenham, O Pearl Brereton, David Budgen, Mark Turner, John Bailey, and Stephen Linkman. 2009. Systematic literature reviews in software engineering–a systematic literature review. Information and software technology 51, 1 (2009), 7–15
2009
-
[37]
Barbara A Kitchenham and Shari L Pfleeger. 2008. Personal Opinion Surveys. In Guide to advanced empirical software engineering . Springer, 63–92
2008
-
[38]
Klaus Krippendorff. 2018. Content analysis: An introduction to its methodology (4 ed.). Sage publications
2018
-
[39]
Patricia Lago, Per Runeson, Qunying Song, and Roberto Verdecchia. 2024. Threats to validity in software engineering–hypocritical paper section or essential analysis?. In Proceedings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement ...
2024
-
[40]
Matheus De Morais Leça, Lucas Valença, Reydne Santos, and Ronnie De Souza Santos. 2025. Applications and Implications of Large Language Models in Qualitative Analysis: A New Frontier for Empirical Software Engineering. In 2025 IEEE/ACM International Workshop on Methodological ...
2025
-
[41]
Jia Li, Ge Li, Yongmin Li, and Zhi Jin. 2025. Structured chain-of-thought prompting for code generation. ACM Transactions on Software Engineering and Methodology 34, 2 (2025), 1–23
2025
-
[42]
Jenny T Liang, Carmen Badea, Christian Bird, Robert DeLine, Denae Ford, Nicole Forsgren, and Thomas Zimmermann. 2024. Can gpt-4 replicate empirical software engineering research? Proceedings of the ACM on Software Engineering 1, FSE (2024), 1330–1353
2024
-
[43]
Vincenzo De Martino, Joel Castaño, Fabio Palomba, Xavier Franch, and Silverio Martínez-Fernández. 2025. A Methodological Framework for LLM-Based Mining of Software Repositories. (7 2025). https://doi.org/10.6084/m9.figshare.29647169
2025 doi
-
[44]
Collin McMillan, Mario Linares-Vasquez, Denys Poshyvanyk, and Mark Grechanik. 2011. Categorizing software applications for maintenance. In 2011 27th ieee international conference on software maintenance (icsm) . IEEE, 343–352
2011
-
[45]
Daye Nam, Andrew Macvean, Vincent Hellendoorn, Bogdan Vasilescu, and Brad Myers. 2024. Using an llm to help with code understanding. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . 1–13
2024
-
[46]
Tien Nguyen, Waris Gill, and Muhammad Ali Gulzar. 2025. Are the Majority of Public Computational Notebooks Pathologically Non-Executable? arXiv preprint arXiv:2502.04184 (2025)
2025 arXiv
-
[47]
Dorin Pomian, Abhiram Bellur, Malinda Dilhara, Zarina Kurbatova, Egor Bogomolov, Timofey Bryksin, and Danny Dig. 2024. Next-generation refactoring: Combining llm insights and ide capabilities for extract method. In 2024 IEEE International Conference on Software Maintenance and...
2024
-
[48]
Ali Ebrahimi Pourasad and Walid Maalej. 2024. Does GenAI Make Usability Testing Obsolete? arXiv preprint arXiv:2411.00634 (2024)
2024
-
[49]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9
2019
-
[50]
Per Runeson and Martin Höst. 2009. Guidelines for conducting and reporting case study research in software engineering. Empirical Software Engineering 14, 2 (2009), 131–164
2009
-
[51]
June Sallou, Thomas Durieux, and Annibale Panichella. 2024. Breaking the silence: the threats of using llms in software engineering. In Proceedings of the 2024 ACM/IEEE 44th International Conference on Software Engineering: New Ideas and Emerging Results . 102–106
2024
-
[52]
Benjamin Saunders, Julius Sim, Tom Kingstone, Shula Baker, Jackie Waterfield, Bernadette Bartlam, Heather Burroughs, and Clare Jinks. 2018. Saturation in qualitative research: exploring its conceptualization and operationalization. Quality & quantity 52 (2018), 1893–1907
2018
-
[53]
Md Shafikuzzaman, Md Rakibul Islam, Alex C Rolli, Sharmin Akhter, and Naeem Seliya. 2024. An Empirical Evaluation of the Zero-Shot, Few-Shot, and Traditional Fine-Tuning Based Pretrained Language Models for Sentiment Analysis in Software Engineering. IEEE Access (2024)
2024
-
[54]
Mohammad Sadegh Sheikhaei, Yuan Tian, Shaowei Wang, and Bowen Xu. 2024. An empirical study on the effectiveness of large language models for SATD identification and classification. Empirical Software Engineering 29, 6 (2024), 159. Manuscript submitted to ACM A Methodological F...
2024
-
[55]
Luciana Lourdes Silva, Janio Rosa da Silva, João Eduardo Montandon, Marcus Andrade, and Marco Tulio Valente. 2024. Detecting code smells using chatgpt: Initial insights. In Proceedings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measureme...
2024
-
[56]
Klaas-Jan Stol and Brian Fitzgerald. 2018. Grounded theory in software engineering research: a critical review and guidelines. ACM Computing Surveys (CSUR) 50, 4 (2018), 1–37
2018
-
[57]
Bianca Trinkenreich, Fabio Calefato, Geir Hanssen, Kelly Blincoe, Marcos Kalinowski, Mauro Pezzé, Paolo Tell, and Margaret-Anne Storey. 2025. Get on the Train or be Left on the Station: Using LLMs for Software Engineering Research. (2025)
2025
-
[58]
Michele Tufano, Fabio Palomba, Gabriele Bavota, Rocco Oliveto, Massimiliano Di Penta, Andrea De Lucia, and Denys Poshyvanyk. 2015. When and why your code starts to smell bad. In 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering , Vol. 1. IEEE, 403–414
2015
-
[59]
Secil Ugurel, Robert Krovetz, and C Lee Giles. 2002. What’s the code? automatic classification of source code archives. In Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining . 632–638
2002
-
[60]
Stefan Wagner, Marvin Muñoz Barón, Davide Falessi, and Sebastian Baltes. 2024. Towards Evaluation Guidelines for Empirical Studies involving LLMs. arXiv preprint arXiv:2411.07668 (2024)
2024 arXiv
-
[61]
Ruiqi Wang, Jiyu Guo, Cuiyun Gao, Guodong Fan, Chun Yong Chong, and Xin Xia. 2025. Can llms replace human evaluators? an empirical study of llm-as-a-judge in software engineering. Proceedings of the ACM on Software Engineering 2, ISSTA (2025), 1955–1977
2025
-
[62]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837
2022
-
[63]
Claes Wohlin. 2014. Guidelines for Snowballing in Systematic Literature Studies and a Replication in Software Engineering. In Proceedings of the 18th International Conference on Evaluation and Assessment in Software Engineering (London, England, United Kingdom) (EASE ’14). Ass...
2014
-
[64]
Claes Wohlin, Emilia Mendes, Katia Romero Felizardo, and Marcos Kalinowski. 2020. Guidelines for the search strategy to update systematic literature reviews in software engineering. Information and software technology 127 (2020), 106366
2020
-
[65]
Claes Wohlin, Per Runeson, Martin Höst, Magnus C Ohlsson, Björn Regnell, and Anders Wesslén. 2012. Experimentation in software engineering . Springer Science & Business Media
2012
-
[66]
Cong Wu, Jing Chen, Ziwei Wang, Ruichao Liang, and Ruiying Du. 2024. Semantic Sleuth: Identifying Ponzi Contracts via Large Language Models. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 582–593
2024
-
[67]
Weiwei Xu, Kai Gao, Hao He, and Minghui Zhou. 2024. LiCoEval: Evaluating LLMs on License Compliance in Code Generation. arXiv preprint arXiv:2408.02487 (2024)
2024 arXiv
-
[68]
Ting Zhang, Ivana Clairine Irsan, Ferdian Thung, and David Lo. 2025. Revisiting Sentiment Analysis for Software Engineering in the Era of Large Language Models. ACM Transactions on Software Engineering and Methodology 34, 3 (2025), 1–30
2025
-
[69]
Yichi Zhang, Zixi Liu, Yang Feng, and Baowen Xu. 2024. Leveraging Large Language Model to Assist Detecting Rust Code Comment Inconsistency. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 356–366
2024
-
[70]
Zibin Zheng, Kaiwen Ning, Qingyuan Zhong, Jiachi Chen, Wenqing Chen, Lianghong Guo, Weicheng Wang, and Yanlin Wang. 2025. Towards an understanding of large language models in software engineering tasks. Empirical Software Engineering 30, 2 (2025), 50. Manuscript submitted to ACM
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.