Pith. sign in

REVIEW 5 major objections 6 minor 1 cited by

A Methodological Framework for LLM-Based Mining of Software Repositories

T0 review · 5 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read LLM-based repository mining has settled into 15 recognizable approaches and nine recurring threats, and the paper's six-stage PRIMES 2.0 framework turns them into a prescriptive and diagnostic tool for rigorous, reproducible studies.

desk verdict A useful, flawed consolidation of LLM4MSR practice; the framework is a good starting point, but the empirical base is thinner than the presentation suggests. read the letter →

arxiv 2508.02233 v2 pith:PZXM32QA submitted 2025-08-04 cs.SE

classification cs.SE
keywords LLM-basedrepositoryminingPRIMES2.0methodologicalframeworkthreatstovaliditypromptengineeringmixed-methodsstudyrapidreviewreproducibility
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to establish a grounded map of how researchers actually use large language models when mining software repositories, and a framework built from that map: PRIMES 2.0, a six-stage workflow with 23 substeps that ties each methodological step to the threats it can provoke and the mitigations that prevent them. The evidence base combines a rapid review of 15 peer-reviewed papers (retained from 7,973 screened) with a survey of 22 researchers who had conducted LLM-based mining studies, and it yields 15 recurring approaches, nine threats to validity, and 26 named mitigation strategies. A sympathetic reader would care because LLM-based repository mining is growing quickly while guidance is fragmented: only 2 of the 15 reviewed papers cited external guidelines, and 10 of the 22 survey respondents were unaware of any. If the framework is right, it gives the community a shared, evidence-based instrument for designing, auditing, and reporting such studies, directly addressing the reproducibility problems the field is accumulating.

What carries the argument

The carrier of the argument is PRIMES 2.0 itself: a six-stage, 23-substep workflow — Strategic Planning & Preparation, Creation of Prompts for Piloting, Prompt Validation, LLM Setup, LLM Comparison, and Execution, Validation & Synthesis — in which every substep is annotated with the threats it risks awakening and the mitigations that address them. It is the mechanism that converts the empirical findings (approaches A1–A15, threats T1–T9, mitigations M1–M26) into actionable guidance, and it is the artifact that would need to prove its worth through wider application for the paper's contribution to hold.

What would settle it

Run a full systematic review of the same period with broader venue coverage, or a survey with a larger probability-based sample, and check whether any methodological approach or threat appears that does not fit into PRIMES 2.0's six stages, 23 substeps, and nine-threat map; one new stage-level decision or one new threat would show the framework describes a small corpus rather than the field.

Watch

Extended reading notes

Core claim

On its own terms, the paper's discovery is the structure of an emerging research practice. Synthesizing 15 peer-reviewed papers and 22 experienced practitioners, it claims that LLM-based repository mining has converged on 15 recognizable methodological approaches — adopting or lacking guidelines, defining goals, curating datasets, automating and tracking the pipeline, choosing prompting strategies, selecting models, iteratively piloting and validating prompts with metrics, configuring and reporting model parameters, considering ensembles, benchmarking models, executing at scale, synthesizing results, and releasing replication packages. It further claims that these approaches organize into six workflow stages and that failures recur as nine threats: hallucinations and output inaccuracy, prompt sensitivity, contextual understanding and domain specificity, scalability and cost, reproducibility and output stability, output formatting and parsing, dataset noise and pretraining leakage, validation complexity, and lack of standardized tooling. The discovery culminates in PRIMES 2.0, which maps each of 23 substeps to the threats (T1–T9) it incurs and the mitigations (M1–M26) that counter them, making the framework at once a design checklist and a diagnostic instrument.

Load-bearing premise

The framework is only as general as its evidence: 15 peer-reviewed papers and 22 volunteer survey respondents are assumed to exhaust the space of how researchers mine repositories with LLMs, so that the 15 approaches, nine threats, and six-stage structure carry over to the field at large.

Editorial extensions

If this is right

  • Researchers planning an LLM-based mining study could use PRIMES 2.0 as a design checklist, building in mitigations for the nine threats from the start rather than discovering them after execution.
  • Reviewers and readers gain a diagnostic instrument: a paper can be audited substep by substep to see which threats were left unmitigated, which configurations went undocumented, and whether a replication package was released.
  • The study's own numbers expose a dissemination gap — 2 of 15 papers and 10 of 22 respondents untethered to any guideline — so the framework's first-order effect would be to make methodological standards visible and concrete.
  • The framework would push the field toward explicit reporting of model versions, temperature settings, prompt templates, validation samples, and discard rates, which is precisely the transparency the study found missing.
  • Because threats such as hallucinations, prompt sensitivity, and validation complexity are not MSR-specific, the six-stage structure is claimed to transfer to other kinds of LLM-based empirical software engineering research.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the authors leave implicit: operationalize PRIMES 2.0 as a scoring rubric, have independent raters apply its 23 substeps to a sample of published LLM-based mining studies, and check whether higher adherence correlates with replicability or fewer reported threats.
  • A saturation check is the fairest test of the framework's foundations: a fuller systematic review or a larger probability-sampled survey would show whether new approaches or threats appear beyond the 15 and nine reported here — something the paper's own evidence base of 15 papers and 22 respondents cannot guarantee.
  • The framework's portability claim, that its stages extend beyond MSR into test generation, code summarization, or requirements classification, is directly checkable by applying the six stages in those adjacent settings and observing whether the threat map still fits.
  • Readers using PRIMES 2.0 as a checklist should reconcile the paper's internal counts first: the abstract reports 25 mitigation strategies while the framework table lists M1–M26, and the conclusion speaks of 14 approaches where the analysis reports 15.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The paper reports a mixed-method study of how large language models (LLMs) are used in Mining Software Repositories (MSR) research. The authors combine a rapid review of 15 peer-reviewed papers (screened from 6,799 in nine venues) with a questionnaire survey of 22 researchers (from 143 invited) to identify methodological approaches, threats to validity, and mitigation strategies. They derive 15 approaches (RQ1), nine threats (RQ2), and a set of mitigations, which they consolidate into PRIMES 2.0, a six-stage framework with 23 substeps that maps threats T1–T9 to mitigation strategies M1–M26. The abstract and conclusion list slightly different counts (25 mitigations in the abstract, 14 approaches in the conclusion), and the empirical base is small and partly convenience-sampled. The paper's central claim is that PRIMES 2.0 is an empirically grounded, prescriptive, and diagnostic instrument for designing rigorous LLM-based MSR studies.

Significance. If the empirical base is accepted as saturating the space of current LLM4MSR practice, the contribution is genuinely useful: it provides a structured, threat-aware framework that researchers can adopt for study design and reporting, and it ships a replication package and public dataset. The threat-mitigation table (Table 4) is a practical synthesis that goes beyond high-level guidelines by operationalizing mitigations with concrete techniques. The paper also explicitly reports verbatim respondent quotes and Cohen's Kappa agreement for study selection, which strengthens the credibility of the descriptive findings. However, the framework's prescriptive value depends on the completeness of the underlying evidence, and the internal inconsistencies in the reported counts reduce confidence in the quantitative claims.

major comments (5)
  1. [Abstract and Section 7] The counts reported in the abstract and conclusion are inconsistent with the body: the abstract states 25 mitigation strategies, Section 4.2 and Table 4 list 26 (M1–M26), and Section 7 says the review and survey explored '14 approaches' while Section 4.1 reports 15 approaches (A1–A15). Since the paper's core contribution is a set of empirically grounded artifacts whose identity is defined by these counts, these numbers must be reconciled.
  2. [Section 3.2.3 and Figure 1] The sample-size accounting is internally inconsistent. The text says '37 valid responses were collected and retained for analysis,' then that one participant failed the attention check and was removed, and then that '22 respondents confirmed their experience.' Figure 1, in contrast, shows '37 Participants,' 'Removed 15 Participants for not using LLM,' and '22 Accepted Participants.' These descriptions do not add up; the reader cannot determine how many participants were invited, how many started, how many failed screening, and how many were finally analyzed. Please provide a clear, consistent flow diagram and textual account.
  3. [Section 3.2.3 and Section 6.3] The paper claims that survey responses were collected 'until no substantially new themes emerged,' invoking theoretical saturation, but no saturation analysis is reported for the review, the survey, or the combined approach/threat/mitigation codebook. With 15 papers and 22 convenience-sampled respondents, the claim that the derived 15 approaches, 9 threats, and 26 mitigations cover the actual space of LLM4MSR practice is load-bearing. Either provide a saturation analysis (e.g., a plot of new codes per additional paper/respondent) or substantially soften the generalizability claims for PRIMES 2.0.
  4. [Section 3.2.3 and Section 6] The recruitment strategy explicitly included contacting the corresponding authors of the papers retained in the rapid review and researchers from the authors' professional networks. This creates a non-trivial overlap between the survey population and the reviewed literature, yet Section 6.3 mentions only convenience sampling and does not discuss the risk that survey responses are not independent from the papers used in the review. The triangulation claim in Section 6.4 would be strengthened by an explicit treatment of this overlap and its effect on the derived threats and mitigations.
  5. [Section 5.4 and Figure 5] The paper states that PRIMES 2.0 comprises '23 methodological substeps, each mapped to specific threats (T1–T9) and corresponding mitigation strategies (M1–M26),' but the 23 substeps are never enumerated in the text, and Figure 5 does not make the per-substep mapping explicit. The reader cannot verify the claimed one-to-one mapping. Please add a table or structured list that names each of the 23 substeps and shows the threat(s) and mitigation(s) mapped to each.
minor comments (6)
  1. [Section 2.2] The text attributes views to 'Kausar et al. [2]', but reference [2] is 'Ahmed et al. (2025)' and no Kausar entry appears in the reference list; please correct the citation or add the intended source.
  2. [Figure 1] The label 'Reseraches' is a typo; it should be 'Researchers.'
  3. [Table 2] In exclusion criterion (EC6), 'exclucion' should be 'exclusion.'
  4. [Table 4] In mitigation M18, 'Exclucion' should be 'Exclusion.'
  5. [Section 3] The GQM description contains a duplicated 'of': 'from the perspective of of SE researchers' should be 'from the perspective of SE researchers.'
  6. [Section 3.2.3] The sentence 'In total, 37 valid responses were collected and retained for analysis' contradicts the later statement that responses failing the attention check were removed; consider rephrasing to distinguish collected from retained.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the empirical findings come from an external review and survey, and the self-citations in the framework are illustrative rather than load-bearing.

full rationale

The paper's central derivation chain is not circular. The 15 methodological approaches, nine threats, and 26 mitigation strategies are presented as findings from a rapid review of 15 peer-reviewed articles and a questionnaire survey of 22 researchers (Sections 3 and 4). These are external empirical inputs, not the framework itself. PRIMES 2.0 is explicitly built as an extension of the authors' earlier PRIMES framework, and the paper openly states that the earlier framework 'was based only on two of our own previous LLM-based MSR studies' and was therefore preliminary (Section 2.1). Section 4.1 also transparently acknowledges that the stage organization 'builds upon our earlier conceptualization' rather than pretending the structure emerged de novo from the data. The only self-referential element is mitigation M5, which mentions 'nascent academic guidelines like our PRIMES framework' as one example of adopting de-facto and emerging standards (Section 4.2, T2). That citation is an illustrative instance inside a broader mitigation strategy and is not load-bearing for the validity of PRIMES 2.0. The acknowledged convenience sampling and limited statistical generalizability (Section 6.3) are external-validity limitations, not circularity. Internal numeric inconsistencies (e.g., '14 approaches' in Section 7 vs. 15 elsewhere; '25 mitigation strategies' in the abstract vs. 26 in Section 4.2) are reporting errors, not evidence of circular reasoning. No fitted parameter is renamed as a prediction, no uniqueness result is imported from the authors' prior work, and no known empirical pattern is merely relabeled.

Assumptions & free parameters 1 free parameters · 5 assumptions · 2 invented entities

The framework rests on qualitative choices rather than fitted numbers. The only hand-set numerical threshold is the Cohen's Kappa = 0.75 inclusion criterion. The main invented entities are PRIMES 2.0 itself and its threat-mitigation mapping, neither of which has external validation. The dominant domain assumptions are the representativeness of the small sample and the completeness of the thematic synthesis.

free parameters (1)
  • Cohen's Kappa acceptance threshold = 0.75
    Hand-chosen threshold for acceptable agreement between the two reviewers when selecting primary studies (Section 3.1.3). It is a declared a priori convention, but it determines which 15 of 31 candidate papers entered the corpus and thereby shapes every downstream count.
assumptions (5)
  • domain assumption The 15 selected papers and 22 survey respondents adequately represent current LLM4MSR practice
    Sections 3.1 and 3.2.3; a small convenience sample, and the authors themselves note the generalizability limit in Section 6.3.
  • domain assumption Qualitative content analysis by the first and second authors yields valid thematic categories
    Section 3.2.4; no inter-rater reliability is reported for the open-ended response coding, unlike the Kappa used for paper selection.
  • domain assumption Threats and mitigations named by a minority of sources are treated as field-level findings
    Section 4; several threats are documented by only a few papers or respondents, yet they are generalized into the T1-T9 set without a prevalence threshold.
  • ad hoc to paper Six stages and 23 substeps are the correct decomposition of the LLM4MSR workflow
    Section 5.4; the stage boundaries refine the authors' original four-stage PRIMES and are not derived by a documented protocol or validated externally.
  • domain assumption Triangulation between rapid review and survey compensates for each method's bias
    Sections 3 and 6.4; the paper assumes convergence of two small, non-random samples indicates robustness of the findings.
invented entities (2)
  • PRIMES 2.0 framework (six stages, 23 substeps)
    purpose: Prescriptive and diagnostic methodological guidance for LLM-based MSR studies, mapping each substep to threats (T1-T9) and mitigations (M1-M26).
    Section 5.4; no study is cited or reported that tests whether following the framework improves rigor, reproducibility, or outcomes; the authors defer validation to future work.
  • Threat-mitigation mapping (T1-T9 paired with M1-M26)
    purpose: Associate each identified threat with concrete mitigation strategies sourced from the literature and survey responses.
    Table 4 and Figure 5; the association between threats and mitigations is the authors' synthesis, not a measured or validated relationship.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Methodological Framework for LLM-Based Mining of Software Repositories." pith.science (2026). https://pith.science/paper/PZXM32QA

@misc{pith2026250802233,
  author       = {Pith},
  title        = {Pith review of: A Methodological Framework for LLM-Based Mining of Software Repositories},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PZXM32QA}},
  note         = {Machine review of arXiv:2508.02233}
}
read the original abstract

Large Language Models (LLMs) are increasingly used in software engineering research, offering new opportunities for automating repository mining tasks. However, despite their growing popularity, the methodological integration of LLMs into Mining Software Repositories (MSR) remains poorly understood. Existing studies tend to focus on specific capabilities or performance benchmarks, providing limited insight into how researchers utilize LLMs across the full research pipeline. To address this gap, we conduct a mixed-method study that combines a rapid review and questionnaire survey in the field of LLM4MSR. We investigate (1) the approaches and (2) the threats that affect the empirical rigor of researchers involved in this field. Our findings reveal 15 methodological approaches, nine main threats, and 25 mitigation strategies. Building on these findings, we present PRIMES 2.0, a refined empirical framework organized into six stages, comprising 23 methodological substeps, each mapped to specific threats and corresponding mitigation strategies, providing prescriptive and adaptive support throughout the lifecycle of LLM-based MSR studies. Our work contributes to establishing a more transparent and reproducible foundation for LLM-based MSR research.

Figures

Figures reproduced from arXiv: 2508.02233 by the authors.

Figure 1
Figure 1. Overview of the Research Process. To answer these research questions, we adopted a mix-method exploratory study [34] consisting of two complementary empirical studies. Specifically, we conducted a rapid review to analyze existing published papers and a survey study to collect insights from researchers. Although the two studies differ in scope and methodology, they both contribute to answering the same research quest… view at source ↗
Figure 2
Figure 2. Papers published over time. ICSE 3 (20.0%) MSR 3 (20.0%) ASE 3 (20.0%) ESEM 1 (6.7%) IST 1 (6.7%) EMSE 1 (6.7%) TOSEM 1 (6.7%) IEEE Access 1 (6.7%) CHI 1 (6.7%) [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 4
Figure 4. Frequency of software repositories and platforms analyzed in the rapid review (N=15 papers) and survey (N=22 respondents). [PITH_FULL_IMAGE:figures/full_fig_p018_4.png] view at source ↗
Figures from the paper (1 more)
Figure 5
Figure 5. Figure 5: PRIMES 2.0: An Empirical-Based Methodological Framework for LLM-Based Mining of Software Repositories. [PITH_FULL_IMAGE:figures/full_fig_p038_5.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Rise of Language Models in Mining Software Repositories: A Survey

    cs.SE 2026-04 conditional novelty 5.0 of 10

    Across 85 papers, language models in MSR mostly perform classification and generation on issues, code reviews, and commits, with the field shifting from fine-tuned BERT-size models to large instruction-tuned LLMs used...

Reference graph

Works this paper leans on

70 extracted references · 52 canonical work pages · cited by 1 Pith paper

  1. [1]

    Samuel Abedu, Laurine Menneron, SayedHassan Khatoonabadi, and Emad Shihab. 2025. RepoChat: An LLM-Powered Chatbot for GitHub Repository Question-Answering. In 2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR) . IEEE, 255–259. Manuscript submitted to ACM A Methodological Framework for LLM-Based Mining of Software Repositories 43

  2. [2]

    Toufique Ahmed, Premkumar Devanbu, Christoph Treude, and Michael Pradel. 2025. Can llms replace manual annotation of software engineering artifacts?. In 2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR) . IEEE, 526–538

  3. [3]

    Dharun Anandayuvaraj, Matthew Campbell, Arav Tewari, and James C Davis. 2024. FAIL: Analyzing Software Failures from the News Using LLMs. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 506–518

  4. [4]

    Sebastian Baltes and Paul Ralph. 2022. Sampling in software engineering research: A critical review and guidelines. Empirical Software Engineering 27, 4 (2022), 94

  5. [5]

    Abdul Ali Bangash, Hareem Sahar, Shaiful Chowdhury, Alexander William Wong, Abram Hindle, and Karim Ali. 2019. What do developers know about machine learning: a study of ml discussions on stackoverflow. In 2019 IEEE/ACM 16th International Conference on Mining Software Repositories (MSR). IEEE, 260–264

  6. [6]

    Cauã Ferreira Barros, Bruna Borges Azevedo, Valdemar Vicente Graciano Neto, Mohamad Kassab, Marcos Kalinowski, Hugo Alexandre D Do Nasci- mento, and Michelle CGSP Bandeira. 2025. Large Language Model for Qualitative Research: A Systematic Mapping Study. In 2025 IEEE/ACM International Workshop on Methodological Issues with Empirical Studies in Software Eng...

  7. [7]

    João Helis Bernardo, Daniel Alencar Da Costa, Sérgio Queiroz de Medeiros, and Uirá Kulesza. 2024. How do machine learning projects use continuous integration practices? an empirical study on GitHub actions. In Proceedings of the 21st International Conference on Mining Software Repositories . 665–676

  8. [8]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901

Show all 70 references
  1. [9]

    Victor R Basili1 Gianluigi Caldiera and H Dieter Rombach. 1994. The goal question metric approach. Encyclopedia of software engineering (1994), 528–532

  2. [10]

    Fabio Calefato, Filippo Lanubile, and Luigi Quaranta. 2022. A preliminary investigation of MLOps practices in GitHub. In Proceedings of the 16th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement . 283–288

  3. [11]

    Bruno Cartaxo, Gustavo Pinto, and Sergio Soares. 2020. Rapid reviews in software engineering. Contemporary Empirical Methods in Software Engineering (2020), 357–384

  4. [12]

    Joel Castaño, Rafael Cabañas, Antonio Salmerón, David Lo, and Silverio Martínez-Fernández. 2024. How do machine learning models change? arXiv preprint arXiv:2411.09645 (2024)

  5. [13]

    Joel Castaño, Silverio Martínez-Fernández, Xavier Franch, and Justus Bogner. 2023. Exploring the carbon footprint of hugging face’s ml models: A repository mining study. In 2023 ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM) . IEEE, 1–12

  6. [14]

    Preetha Chatterjee, Mia Mohammad Imran, and Kostadin Damevski. 2024. Shedding Light on Software Engineering-specific Metaphors and Idioms. In ICSE’24: Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . Association for Computing Machinery, 1–13

  7. [15]

    Zhenpeng Chen, Yanbin Cao, Yuanqiang Liu, Haoyu Wang, Tao Xie, and Xuanzhe Liu. 2020. A comprehensive study on challenges in deploying deep learning based software. In Proceedings of the 28th ACM joint meeting on European software engineering conference and symposium on the fo...

  8. [16]

    Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educational and psychological measurement 20, 1 (1960), 37–46

  9. [17]

    Giuseppe Colavito, Filippo Lanubile, and Nicole Novielli. 2025. Benchmarking large language models for automated labeling: The case of issue report classification. Information and Software Technology (2025), 107758

  10. [18]

    Marcelo Costalonga, Bianca Minetto Napoleão, Maria Teresa Baldassarre, Katia Romero Felizardo, Igor Steinmacher, and Marcos Kalinowski. 2025. Can Machine Learning Support the Selection of Studies for Systematic Literature Review Updates? arXiv preprint arXiv:2502.08050 (2025)

  11. [19]

    Daniela S Cruzes and Tore Dybå. 2011. Recommended steps for thematic synthesis in software engineering. International Symposium on Empirical Software Engineering and Measurement (ESEM) (2011), 275–284

  12. [20]

    Vincenzo De Martino, Joel Castaño, Fabio Palomba, Xavier Franch, and Silverio Martínez-Fernández. 2025. A framework for using llms for repository mining studies in empirical software engineering. In 2025 IEEE/ACM International Workshop on Methodological Issues with Empirical S...

  13. [21]

    Vincenzo De Martino, Silverio Martínez-Fernández, and Fabio Palomba. 2024. Do developers adopt green architectural tactics for ml-enabled systems? a mining software repository study. arXiv preprint arXiv:2410.06708 (2024)

  14. [22]

    Vincenzo De Martino and Fabio Palomba. 2025. Classification and challenges of non-functional requirements in ML-enabled systems: A systematic literature review. Information and Software Technology (2025), 107678

  15. [23]

    Vincenzo De Martino, Gilberto Recupito, Giammaria Giordano, Filomena Ferrucci, Dario Di Nucci, and Fabio Palomba. 2025. Into the ML- Universe: An improved classification and characterization of machine-learning projects. Journal of Systems and Software 230 (2025), 112471. http...

  16. [24]

    Don A Dillman, Jolene D Smyth, and Leah Melani Christian. 2014. Internet, phone, mail, and mixed-mode surveys: The tailored design method . John Wiley & Sons

  17. [25]

    Yihong Dong, Xue Jiang, Zhi Jin, and Ge Li. 2024. Self-collaboration code generation via chatgpt. ACM Transactions on Software Engineering and Methodology 33, 7 (2024), 1–38

  18. [26]

    Angela Fan, Beliz Gokkaya, Mark Harman, Mitya Lyubarskiy, Shubho Sengupta, Shin Yoo, and Jie M Zhang. 2023. Large language models for software engineering: Survey and open problems. In 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineeri...

  19. [27]

    Katia Romero Felizardo, Anderson Deizepe, Daniel Coutinho, Genildo Gomes, Maria Meireles, Marco Gerosa, and Igor Steinmacher. 2025. On the Difficulties of Conducting and Replicating Systematic Literature Reviews Studies Using LLMs in Software Engineering. In 2025 IEEE/ACM Inte...

  20. [28]

    Alexandra González, Xavier Franch, David Lo, and Silverio Martínez-Fernández. 2025. How do Pre-Trained Models Support Software Engineering? An Empirical Study in Hugging Face. arXiv preprint arXiv:2506.03013 (2025)

  21. [29]

    Joseph F Hair, Arthur H Money, Philip Samouel, and Mike Page. 2007. Research methods for business. Education+ Training 49, 4 (2007), 336–337

  22. [30]

    Abram Hindle, Daniel M German, and Ric Holt. 2008. What do large commits tell us? a taxonomical study of large commits. In Proceedings of the 2008 international working conference on Mining software repositories . 99–108

  23. [31]

    Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. 2024. Large language models for software engineering: A systematic literature review. ACM Transactions on Software Engineering and Methodology 33, 8 (2024), 1–79

  24. [32]

    Mia Mohammad Imran, Preetha Chatterjee, and Kostadin Damevski. 2024. Uncovering the causes of emotions in software developer communication using zero-shot llms. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . 1–13

  25. [33]

    Wenxin Jiang, Jerin Yasmin, Jason Jones, Nicholas Synovic, Jiashen Kuo, Nathaniel Bielanski, Yuan Tian, George K Thiruvathukal, and James C Davis. 2024. PeaTMOSS: A Dataset and Initial Analysis of Pre-Trained Models in Open-Source Software. In Proceedings of the 21st Internati...

  26. [34]

    R Burke Johnson, Anthony J Onwuegbuzie, and Lisa A Turner. 2007. Toward a definition of mixed methods research. Journal of mixed methods research 1, 2 (2007), 112–133

  27. [35]

    Minseo Kim, Taemin Kim, Thu Hoang Anh Vo, Yugyeong Jung, and Uichin Lee. 2025. Exploring Modular Prompt Design for Emotion and Mental Health Recognition. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems . 1–18

  28. [36]

    Barbara Kitchenham, O Pearl Brereton, David Budgen, Mark Turner, John Bailey, and Stephen Linkman. 2009. Systematic literature reviews in software engineering–a systematic literature review. Information and software technology 51, 1 (2009), 7–15

  29. [37]

    Barbara A Kitchenham and Shari L Pfleeger. 2008. Personal Opinion Surveys. In Guide to advanced empirical software engineering . Springer, 63–92

  30. [38]

    Klaus Krippendorff. 2018. Content analysis: An introduction to its methodology (4 ed.). Sage publications

  31. [39]

    Patricia Lago, Per Runeson, Qunying Song, and Roberto Verdecchia. 2024. Threats to validity in software engineering–hypocritical paper section or essential analysis?. In Proceedings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement ...

  32. [40]

    Matheus De Morais Leça, Lucas Valença, Reydne Santos, and Ronnie De Souza Santos. 2025. Applications and Implications of Large Language Models in Qualitative Analysis: A New Frontier for Empirical Software Engineering. In 2025 IEEE/ACM International Workshop on Methodological ...

  33. [41]

    Jia Li, Ge Li, Yongmin Li, and Zhi Jin. 2025. Structured chain-of-thought prompting for code generation. ACM Transactions on Software Engineering and Methodology 34, 2 (2025), 1–23

  34. [42]

    Jenny T Liang, Carmen Badea, Christian Bird, Robert DeLine, Denae Ford, Nicole Forsgren, and Thomas Zimmermann. 2024. Can gpt-4 replicate empirical software engineering research? Proceedings of the ACM on Software Engineering 1, FSE (2024), 1330–1353

  35. [43]

    Vincenzo De Martino, Joel Castaño, Fabio Palomba, Xavier Franch, and Silverio Martínez-Fernández. 2025. A Methodological Framework for LLM-Based Mining of Software Repositories. (7 2025). https://doi.org/10.6084/m9.figshare.29647169

  36. [44]

    Collin McMillan, Mario Linares-Vasquez, Denys Poshyvanyk, and Mark Grechanik. 2011. Categorizing software applications for maintenance. In 2011 27th ieee international conference on software maintenance (icsm) . IEEE, 343–352

  37. [45]

    Daye Nam, Andrew Macvean, Vincent Hellendoorn, Bogdan Vasilescu, and Brad Myers. 2024. Using an llm to help with code understanding. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . 1–13

  38. [46]

    Tien Nguyen, Waris Gill, and Muhammad Ali Gulzar. 2025. Are the Majority of Public Computational Notebooks Pathologically Non-Executable? arXiv preprint arXiv:2502.04184 (2025)

  39. [47]

    Dorin Pomian, Abhiram Bellur, Malinda Dilhara, Zarina Kurbatova, Egor Bogomolov, Timofey Bryksin, and Danny Dig. 2024. Next-generation refactoring: Combining llm insights and ide capabilities for extract method. In 2024 IEEE International Conference on Software Maintenance and...

  40. [48]

    Ali Ebrahimi Pourasad and Walid Maalej. 2024. Does GenAI Make Usability Testing Obsolete? arXiv preprint arXiv:2411.00634 (2024)

  41. [49]

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9

  42. [50]

    Per Runeson and Martin Höst. 2009. Guidelines for conducting and reporting case study research in software engineering. Empirical Software Engineering 14, 2 (2009), 131–164

  43. [51]

    June Sallou, Thomas Durieux, and Annibale Panichella. 2024. Breaking the silence: the threats of using llms in software engineering. In Proceedings of the 2024 ACM/IEEE 44th International Conference on Software Engineering: New Ideas and Emerging Results . 102–106

  44. [52]

    Benjamin Saunders, Julius Sim, Tom Kingstone, Shula Baker, Jackie Waterfield, Bernadette Bartlam, Heather Burroughs, and Clare Jinks. 2018. Saturation in qualitative research: exploring its conceptualization and operationalization. Quality & quantity 52 (2018), 1893–1907

  45. [53]

    Md Shafikuzzaman, Md Rakibul Islam, Alex C Rolli, Sharmin Akhter, and Naeem Seliya. 2024. An Empirical Evaluation of the Zero-Shot, Few-Shot, and Traditional Fine-Tuning Based Pretrained Language Models for Sentiment Analysis in Software Engineering. IEEE Access (2024)

  46. [54]

    Mohammad Sadegh Sheikhaei, Yuan Tian, Shaowei Wang, and Bowen Xu. 2024. An empirical study on the effectiveness of large language models for SATD identification and classification. Empirical Software Engineering 29, 6 (2024), 159. Manuscript submitted to ACM A Methodological F...

  47. [55]

    Luciana Lourdes Silva, Janio Rosa da Silva, João Eduardo Montandon, Marcus Andrade, and Marco Tulio Valente. 2024. Detecting code smells using chatgpt: Initial insights. In Proceedings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measureme...

  48. [56]

    Klaas-Jan Stol and Brian Fitzgerald. 2018. Grounded theory in software engineering research: a critical review and guidelines. ACM Computing Surveys (CSUR) 50, 4 (2018), 1–37

  49. [57]

    Bianca Trinkenreich, Fabio Calefato, Geir Hanssen, Kelly Blincoe, Marcos Kalinowski, Mauro Pezzé, Paolo Tell, and Margaret-Anne Storey. 2025. Get on the Train or be Left on the Station: Using LLMs for Software Engineering Research. (2025)

  50. [58]

    Michele Tufano, Fabio Palomba, Gabriele Bavota, Rocco Oliveto, Massimiliano Di Penta, Andrea De Lucia, and Denys Poshyvanyk. 2015. When and why your code starts to smell bad. In 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering , Vol. 1. IEEE, 403–414

  51. [59]

    Secil Ugurel, Robert Krovetz, and C Lee Giles. 2002. What’s the code? automatic classification of source code archives. In Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining . 632–638

  52. [60]

    Stefan Wagner, Marvin Muñoz Barón, Davide Falessi, and Sebastian Baltes. 2024. Towards Evaluation Guidelines for Empirical Studies involving LLMs. arXiv preprint arXiv:2411.07668 (2024)

  53. [61]

    Ruiqi Wang, Jiyu Guo, Cuiyun Gao, Guodong Fan, Chun Yong Chong, and Xin Xia. 2025. Can llms replace human evaluators? an empirical study of llm-as-a-judge in software engineering. Proceedings of the ACM on Software Engineering 2, ISSTA (2025), 1955–1977

  54. [62]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837

  55. [63]

    Claes Wohlin. 2014. Guidelines for Snowballing in Systematic Literature Studies and a Replication in Software Engineering. In Proceedings of the 18th International Conference on Evaluation and Assessment in Software Engineering (London, England, United Kingdom) (EASE ’14). Ass...

  56. [64]

    Claes Wohlin, Emilia Mendes, Katia Romero Felizardo, and Marcos Kalinowski. 2020. Guidelines for the search strategy to update systematic literature reviews in software engineering. Information and software technology 127 (2020), 106366

  57. [65]

    Claes Wohlin, Per Runeson, Martin Höst, Magnus C Ohlsson, Björn Regnell, and Anders Wesslén. 2012. Experimentation in software engineering . Springer Science & Business Media

  58. [66]

    Cong Wu, Jing Chen, Ziwei Wang, Ruichao Liang, and Ruiying Du. 2024. Semantic Sleuth: Identifying Ponzi Contracts via Large Language Models. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 582–593

  59. [67]

    Weiwei Xu, Kai Gao, Hao He, and Minghui Zhou. 2024. LiCoEval: Evaluating LLMs on License Compliance in Code Generation. arXiv preprint arXiv:2408.02487 (2024)

  60. [68]

    Ting Zhang, Ivana Clairine Irsan, Ferdian Thung, and David Lo. 2025. Revisiting Sentiment Analysis for Software Engineering in the Era of Large Language Models. ACM Transactions on Software Engineering and Methodology 34, 3 (2025), 1–30

  61. [69]

    Yichi Zhang, Zixi Liu, Yang Feng, and Baowen Xu. 2024. Leveraging Large Language Model to Assist Detecting Rust Code Comment Inconsistency. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering . 356–366

  62. [70]

    Zibin Zheng, Kaiwen Ning, Qingyuan Zhong, Jiachi Chen, Wenqing Chen, Lianghong Guo, Weicheng Wang, and Yanlin Wang. 2025. Towards an understanding of large language models in software engineering tasks. Empirical Software Engineering 30, 2 (2025), 50. Manuscript submitted to ACM

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.