Pith. sign in

REVIEW 3 major objections 7 minor 104 references

Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews

T0 review · 3 major / 7 minor · reviewed 2026-07-31 · grok-4.5

Pith's one-line read GenAI cannot run unsupervised systematic literature reviews; GUEST process rules keep humans in charge while using it for help and checks.

desk verdict Solid preliminary SE process guidance for GenAI-in-SLRs; useful packaging and honest limits, not a fully evidenced rule set for every SLR task. read the letter →

arxiv 2607.24991 v1 pith:LVPME7IC submitted 2026-07-27 cs.SE cs.AI

classification cs.SEcs.AI
keywords EvaluationstudiesLargeLanguageModels(LLMs)GenerativeAI(GenAI)toolsSystematicReviewGentooluseSoftwareEngineering(SE)Guidelines
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Systematic literature reviews need validity, traceability, qualification of evidence strength, and verifiability that plain expert summaries do not provide. Generative AI tools can summarize text and may cut cost on repetitive work, but they also hallucinate, give inconsistent answers, leak or contaminate test data, hide their reasoning, and can inherit bias. This paper argues they therefore cannot replace human researchers on full systematic reviews today. It offers GUEST: concrete process recommendations for two roles—people who evaluate GenAI on review tasks, and people who run reviews with GenAI help—covering planning, conduct, reporting, and task-level rules for screening, extraction, qualitative synthesis, risk of bias, and strength of evidence. The aim is trustworthy GenAI-assisted reviews and rigorous independent evaluations rather than unsupervised automation.

What carries the argument

GUEST (GenAI Use and Evaluation in SLR Tasks): role-split process recommendations (evaluators vs reviewers) for planning, conduct, and reporting, plus task-level rules—especially cost-sensitive screening metrics that treat lost relevant studies as worse than extra work, and treating GenAI as a validator rather than sole agent on qualitative synthesis, risk-of-bias, and strength-of-evidence work.

What would settle it

A prospective, contamination-controlled comparison on a full software-engineering SLR pipeline in which a GenAI workflow without continuous human validation matches dual-human process controls on validity, traceability, evidence qualification, and auditability—or clearly fails those same checks under the paper’s own metrics.

Watch

Extended reading notes

Core claim

GenAI requires human oversight and is not currently capable of unsupervised systematic studies; it can give cost-effective help on some repetitive tasks and extra validation on some complex ones. The authors package that stance as GUEST process recommendations so software engineering researchers can both run and report trustworthy SLRs that use GenAI and produce rigorous independent evaluations of GenAI on SLR tasks.

Load-bearing premise

Role-based thought experiments, a single-library rapid review with one screener, and two companion studies are enough to ground general process rules for every major SLR task, including those where the paper itself says direct evidence is still thin or contradictory.

Editorial extensions

If this is right

  • SLR reports that use GenAI must document prompts, model versions, parameters, human checks, and full confusion-matrix counts where screening is automated.
  • Screening evaluations should prefer metrics that penalize missed relevant studies over raw accuracy or unweighted scores under class imbalance.
  • On qualitative synthesis, risk-of-bias, and strength-of-evidence tasks, GenAI should run as an extra checker beside humans, not as the sole decision maker, until stronger evidence exists.
  • Prospective case studies built into live SLRs become the preferred way to gather evaluation evidence and limit data contamination.
  • End-to-end agentic review platforms remain open research unless each stage’s intermediate outputs can be inspected and validated.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Journals and conferences that require GUEST-style disclosure will make GenAI-assisted SLR claims easier to audit and meta-analyze.
  • The same human-oversight and contamination logic likely extends to other secondary-research genres that borrow SLR process controls.
  • Until SE-specific risk-of-bias and strength-of-evidence instruments stabilize, GenAI scores on those tasks will stay hard to compare across studies.
  • Multi-tool ensembles may cut idiosyncratic errors but do not remove shared training-data contamination and raise reported cost.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper presents GUEST, a preliminary set of process recommendations for software engineering researchers who either (a) evaluate GenAI tools on SLR tasks or (b) use GenAI tools while conducting SLRs. The method has three strands: a rapid review of existing GenAI-use guidelines (single-source SCOPUS, searches closed 29/03/2025, documented in Appendix A with search strings and known-item validation); role-based thought experiments across planning/conduct/reporting stages, supported by per-task briefing notes (Appendix C); and a synthesis of empirically reported failure modes mapped to recommendations (Table 3). The output is a set of planning/conduct/reporting recommendations (Table 5) and task-level recommendations for screening, data extraction, qualitative synthesis, and risk-of-bias/strength-of-evidence assessment (Table 6), each annotated by "kind" (SLR-specific vs. general practice, Table 4). The central claims are modest: GenAI requires human oversight and cannot currently conduct unsupervised systematic studies; the recommendations are explicitly preliminary. Empirical grounding is strong for screening and qualitative synthesis (via the authors' peer-reviewed companion studies [61, 79]) and admittedly thin for extraction, RoB, and SoE, which §7 flags as open research questions. A cross-check against Baltes et al. and RAISE is reported in §4.5.

Significance. Timely and useful. As GenAI-assisted SLRs proliferate in SE, the absence of agreed evaluation methodology is a real problem, and the paper's MMRE/crossover analogies for how bad metrics entrench are apt. Strengths worth naming: the rapid review is documented to a reproducible standard (search strings in Tables 9–10, known-item validation in Tables 7–8); Table 3 maps empirically reported failure modes to the recommendations that address them; Table 4's four-kind classification of recommendations — separating SLR-specific contributions from restated good practice — is a methodological contribution other guideline efforts could adopt; the empirical backbone rests on peer-reviewed companion studies rather than self-referential definitions; and §7 lists concrete, falsifiable open questions. The guidance is honest about its limits (§5, A.7) and complementary to rather than duplicative of RAISE. If the framing issues in the major comments are fixed, this will be a standard reference for SE researchers evaluating or using GenAI in SLRs.

major comments (3)
  1. [§4.5, §5.1] §4.5 is titled and framed as 'Validation', and §5.1 states the recommendations 'were corroborated by an independent guideline set', describing Baltes et al. [3] as 'the independent guidelines ... which we deliberately did not use while developing our recommendations'. But §4.5 itself notes Baltes et al. 'refines and extends' Wagner et al. [97], which §3.1 lists as a direct development input, and RAISE [90] is also a stated input (§3.1). Neither arm of §4.5 is therefore independent of the construction set; the exercise is a completeness/consistency check, and it is the paper's only systematic test that the thought experiments did not miss important issues — exactly what §5.1's construct-validity argument rests on. The comparison itself is careful and valuable; the framing overstates it. Suggest retitling/reframing §4.5 as a completeness check, softening 'independent' in §5.1, and adding a
  2. [Tables 5–6, §4.3, §7] The empirical backbone (Table 3, §4.3) covers only two SLR tasks: screening (via [61]) and qualitative synthesis (via [79]). For data extraction, RoB, and SoE, recommendations Ext1–Ext7, A1–A3, Syn1–Syn3, and RevA1–RevA2 rest on thought experiments and briefing notes alone, with direct evidence described as 'limited' or 'contradictory' (§7, Appendix C.3, C.4.2). The paper concedes this honestly, but Tables 5–6 present all recommendations with uniform force, distinguished only by 'kind' (i)–(iv), which classifies novelty, not evidential support. Since the tables are the deliverable practitioners will consult, please annotate each recommendation with its evidence status (companion-study grounded / external empirical / thought-experiment only), paralleling the kind labels, and add one scoping sentence in the abstract or conclusion stating which tasks are evidence-grounded. This is a present
  3. [§4.3, Table 6 (Scr2), Appendix C.1] The paper's one quantitative anchor — the accuracy-ranked LLM missing 63.3% of relevant studies vs 5.8% for the WMCC-ranked one — comes from a reanalysis of a single 9,695-article evaluation reported in the authors' own [61], and the WMCC ranking depends on the FN:FP cost ratio, whose value is neither justified nor subjected to sensitivity analysis in this paper. Scr2 is a kind-(i) recommendation, i.e., part of the paper's distinctive contribution, yet a reader of this paper alone cannot judge how robust the WMCC mandate is to the chosen ratio. A short paragraph stating the ratio used, its rationale, and the sensitivity of the tool ranking would suffice; alternatively, scope Scr2 the way Ext3 is already scoped ('unless there is an explicit justification for assuming differential costs'). Ext3 shows the authors can write exactly this kind of caveat.
minor comments (7)
  1. [Appendix A.3 (E2), §4.5] Eligibility criterion E2 excludes self-published (arXiv-only) papers from the rapid review, yet Baltes et al. [3] — cited as an arXiv preprint (arXiv:2508.15503) — is used as the §4.5 cross-check baseline. Given how fast this literature moves, the exception is defensible, but the asymmetry should be acknowledged in one sentence.
  2. [Appendix A.5, A.7] Candidate-study selection was performed by a single researcher (A.5), acknowledged as a limitation in A.7. A second-screener audit of even a small sample of excluded candidates, or a reported agreement check, would materially strengthen the RR at low cost.
  3. [§3.2, Appendix A.7] Searches closed 29/03/2025. The narrative tracking of later guidance (Baltes et al., Farotimi et al., Nahar et al.; §3.2) is a reasonable substitute, but the justification currently appears only in A.7; please state the search horizon and the tracking strategy briefly in §3.2 where the RR is first summarized.
  4. [§4.5] Among the dropped RAISE items, 'reviewers need to be aware that research papers may be written by AI' is dismissed as having unclear implications. Since AI-authored primary studies would directly threaten SLR validity (screening, RoB), one sentence explaining why this is out of scope — or a pointer to §7 as an open question — would close the gap.
  5. [Table 3, Appendix C.4.2] Table 3's qualitative-synthesis row could cross-reference the stated limitation of [79] in Appendix C.4.2 (no independent data-extraction stage; full papers given to the tool), so readers see the caveat where the failure modes are summarized, not only in the appendix.
  6. [Various] Typographical: 'chatbotprompts' (§2.1, item 6); 'Its search processes conforms' (A.2); 'GenAi tool performance' (B.1.3); 'usingvsevaluating' (Table 1, missing spaces); reference [2] carries an '[n. d.]' date; 'SLRS tasks' (B.2.1).
  7. [Figure 2, §B.2.5] Figure 2 is information-dense; the 'validate & refine prompts' / 're-validate on new issues' feedback arrows are easy to miss. Consider slightly larger labels or a caption sentence noting the iterative loop, which §B.2.5 describes well in prose.

Circularity Check

2 steps flagged · score 2.0 of 10

No derivation-by-construction circularity: GUEST is process guidance from thought experiments, a rapid review, and external standards; companion self-citations supply empirical failure modes for two tasks but do not make the recommendations true by definition.

  1. self citation load bearing [§2.3 ‘Relationship to our companion studies’; §4.3 Table 3; Scr1–Scr3 / Syn1–Syn3]
    "Two of these works are our own peer-reviewed studies, and we use them here as the empirical backbone for our screening and synthesis recommendations rather than re-deriving that evidence. Our study of literature screening [61] … provides the performance-metric analysis … that underpin our screening recommendations. Our study of qualitative synthesis [79] … provides the trial-based failure modes and soundness criteria that underpin our synthesis recommendations."

    Task-level screening and qualitative-synthesis recommendations (especially metric choice such as WMCC and the failure-mode map in Table 3) are justified primarily by citations whose author sets overlap the present paper. This is load-bearing for Scr* and Syn* content. It is not full circularity: [61] and [79] are separate peer-reviewed empirical studies, the paper does not define GUEST as whatever those studies found, and planning/conduct/reporting items and RoB/SoE/extraction guidance rest on other sources (thought experiments, RAISE, G1–G6, SEGRESS/PRISMA). Mild self-reliance, not equivalence by construction.

  2. other [§4.5 Validation; cf. §3.1 development inputs]
    "The Baltes et al. study, which refines and extends the authors’ earlier workshop position paper (Wagner et al. [97]) and is therefore a closely related line of work rather than a fully independent one, provided a recent SE baseline against which to cross-check our recommendations, while the study by Thomas et al. provided a check that the different methods we trialled to report our results had not led to any important issues being lost."

    The completeness/cross-check arm of validation partly re-uses lineage that was already a development input (Wagner→Baltes; RAISE as both input and comparator). That reduces the independence of the ‘we did not miss important issues’ claim. It does not make any GUEST recommendation true by definition or force a prediction from a fit; it is non-independent validation, scored as minor residual circularity of the check rather than of the derivation.

full rationale

This is a methodology/guidelines paper, not a fitted predictive model or uniqueness-theorem derivation. The claimed chain is: rapid review of GenAI-use guidelines (G1–G6) + role-based thought experiments + SLR-task briefing notes + classical reporting standards (SEGRESS/PRISMA) + RAISE/Wagner inputs → GUEST process recommendations, with human oversight as the normative conclusion. Nothing in that chain equates a fitted parameter to a ‘prediction,’ defines X via Y, or imports a uniqueness theorem that forbids alternatives. The empirical backbone for literature-screening metrics (MCC/WMCC, confusion-matrix reporting, lost-evidence emphasis) and qualitative-synthesis failure modes rests substantially on the authors’ own peer-reviewed companion studies [61] and [79]. That is load-bearing self-citation for Scr* and Syn* items, but those works are external peer-reviewed empirical reports (reanalyses and autoethnographic trials), not unverified assertions internal to this manuscript, and the paper does not treat their results as definitional of GUEST. For data extraction, RoB, and SoE the paper explicitly marks direct SE GenAI evidence as limited or contradictory and labels related items preliminary (Section 7)—honest under-determination, not circular forcing. The §4.5 ‘validation’ against Baltes et al. [3] and Thomas/RAISE [90] is only partly independent (Baltes extends Wagner [97], a stated development input; RAISE was also an input), which weakens the independence of the completeness check but does not reduce GUEST’s content to those sources by construction. Overall circularity is minor residual self-reliance, score 2.

Assumptions & free parameters 0 free parameters · 6 assumptions · 2 invented entities

The paper’s load-bearing conclusions rest on established SLR validity norms, stated GenAI failure properties, and the adequacy of thought experiments plus limited task-level evidence—not on fitted physical parameters. Free parameters are essentially absent; axioms are domain assumptions about what makes a review ‘systematic’ and what GenAI can/cannot guarantee. Invented entities are organizational (GUEST, role split), not ontological posits.

assumptions (6)
  • domain assumption SLR trustworthiness requires validity, traceability, qualified evidence strength, and verifiability/reproducibility of process—not merely a fluent literature summary.
    Section 2.2 frames why expert-opinion-like GenAI reviews are insufficient; underpins the human-oversight conclusion.
  • domain assumption Current GenAI systems are non-deterministic, subject to hallucination, data contamination, hidden infrastructure changes, and limited justified chain-of-thought suitable for audit.
    Section 2.1 list (items 1–10); treated as given constraints on evaluation and use.
  • ad hoc to paper Thought experiments structured by researcher role and planning/conduct/reporting stages, informed by experience and briefing notes, can yield useful preliminary process recommendations.
    Core method in Sections 3.3–4; authors note judgment dependence in threats (Section 5).
  • domain assumption False negatives in study screening cost more (lost evidence) than false positives (extra workload), so cost-asymmetric metrics (e.g., WMCC) are appropriate for GenAI screening evaluation.
    Imported from companion screening study and Scr1–Scr2; central to screening recommendations.
  • domain assumption Software engineering SLRs are more heterogeneous and more often qualitative than typical health RCT-centered synthesis, so SE-specific task guidance is warranted beyond RAISE.
    Table 1 and Section 2.3–2.4 positioning; motivates producing a separate guideline set.
  • domain assumption Human accountability norms for AI in scholarly work (G1–G5: oversight, no AI authorship, full reporting, no fabricated results) bind SLR use of GenAI.
    Table 2 from rapid review synthesis; legal/ethical frame for all GUEST items.
invented entities (2)
  • GUEST (GenAI Use and Evaluation in SLR Tasks)
    purpose: Named umbrella for the paper’s process recommendation set spanning shared and role-specific planning, conduct, reporting, and per-SLR-task rules.
    Organizational brand for the contribution; content is recommendations, not a new physical or computational object with independent dynamics.
  • Evaluator vs Reviewer role split (as used to structure GUEST) independent evidence
    purpose: Partition recommendations into independent tool assessment versus conducting an SLR with GenAI, including overlap on prospective validation.
    Analytical device aligned with RAISE roles but narrowed to two SE-predominant roles; not an empirical discovery.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews." pith.science (2026). https://pith.science/paper/LVPME7IC

@misc{pith2026260724991,
  author       = {Pith},
  title        = {Pith review of: Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LVPME7IC}},
  note         = {Machine review of arXiv:2607.24991}
}
read the original abstract

Context: Generative AI (GenAI) and Large Language Models (LLMs) are increasingly used for academic tasks in software engineering and beyond, including systematic literature reviews (SLRs). However, while capable of summarizing text, there is no guarantee they can meet the rigour, reliability, and transparency that SLRs require. Objectives: To support researchers intending to conduct SLRs using GenAI or those conducting empirical studies evaluating how well GenAI supports SLR tasks. Methods: First, we conducted a rapid review to identify studies that propose guidelines for evaluating and using GenAI and LLMs to support SLRs. Second, we drew on thought experiments, relevant guidance from the literature, and our own experience conducting SLRs and evaluating tools to develop recommendations for how to use and assess GenAI in the context of SLRs. Results: We discuss the problems researchers face when evaluating GenAI for SLRs. We identify and explain process issues to consider when planning, conducting, and reporting both SLRs using GenAI and evaluations of GenAI tools. Finally, we summarize our results as a set of process recommendations, which we name GUEST (GenAI Use and Evaluation in SLR Tasks). Conclusion: We argue that GenAI requires human oversight and is not currently capable of unsupervised systematic studies. However, it offers the prospect of cost-effective assistance for some repetitive tasks and for additional validation of some complex tasks. Our GUEST recommendations should help software engineering researchers both to conduct and report trustworthy SLRs using GenAI and to provide rigorous independent evaluation studies.

Figures

Figures reproduced from arXiv: 2607.24991 by the authors.

Figure 1
Figure 1. Context and development process for the guidelines on evaluating and using GenAI for SLRs. [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 2
Figure 2. Summary of SLR Process Extensions (3) Assessing each of the candidate primary studies identified by the search process against the eligibility criteria and retaining any study that cannot be rejected from the list of candidate primary studies. For a study claim￾ing to be a formal SLR, this process should be conducted by two human researchers working independently, and disagreements would need to be identified and re… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

104 extracted references · 13 canonical work pages

  1. [3]

    Sebastian Baltes, Florian Angermeir, Chetan Arora, Marvin Muñoz Barón, Chunyang Chen, Lukas Böhme, Fabio Calefato, Neil Ernst, Davide Falessi, Brian Fitzgerald, et al. 2025. Guidelines for empirical studies in software engineering involving large language models.arXiv preprint arXiv:2508.15503(2025)

  2. [97]

    Stefan Wagner, Marvin Muñoz Barón, Davide Falessi, and Sebastian Baltes. 2025. Towards Evaluation Guidelines for Empirical Studies involving LLMs. arXiv:2411.07668 [cs.SE] https://arxiv.org/abs/2411.07668

  3. [90]

    Berg, Yuan Chi, A

    James Thomas, Kaitlyn Hair, Anna Noel-Storr, Michelle Angrish, Edoardo Aromataris, Lisa Askie, Rigmor C. Berg, Yuan Chi, A. Justin Clark, Declan Devane, Emma McFarlane, Isabel Fletcher, Gerald Gartlehner, Jonas Goretzko, Neal Haddaway, Raouf Hajji, Paweł Jemioło, Zoe Jordan, Justine Karpusheff, Isabel Kempner, Wojciech Kusa, Claudia Lenkewitz, Biljana Mac...

  4. [61]

    Lech Madeyski, Barbara Kitchenham, and Martin Shepperd. 2026. LLM4SCREENLIT: Recommendations on assessing the performance of large language models for screening literature in systematic reviews.Information and Software Technology198 (2026), 108204. doi:10.1016/j.infsof.2026. 108204

  5. [79]

    Sebastián Pizard, Ramiro Moreira, Federico Galiano, Ignacio Sastre, and Lorena Etcheverry. 2026. On the Use of Large Language Models for Qualitative Synthesis. InProceedings of the 2026 IEEE/ACM International Workshop on Methodological Issues with Empirical Studies in Software Engineering (WSESE ’26). Association for Computing Machinery, New York, NY, USA...

  6. [1]

    Ahmad Alshami, Moustafa Elsayed, Eslam Ali, Abdelrahman E. E. Eltoukhy, and Tarek Zayed. 2023. Harnessing the Power of ChatGPT for Automating Systematic Review Process: Methodology, Case Study, Limitations, and Future Directions.Systems11, 7 (2023). doi:10.3390/ systems11070351

  7. [2]

    Aurelian Anghelescu, Florentina Carmen Firan, Gelu Onose, Constantin Munteanu, Andreea-Iulia Trandafir, Ilinca Ciobanu, Stefan Gheorghita, and Vlad Ciobanu. [n. d.]. PRISMA Systematic Literature Review, including with Meta-Analysis vs. Chatbot/GPT (AI) regarding Current Scientific Data on the Main Effects of the Calf Blood Deproteinized Hemoderivative Med...

  8. [4]

    Leonardo Banh and Gero Strobel. 2023. Generative artificial intelligence.Electronic Markets33, 1 (06 Dec 2023), 63. doi:10.1007/s12525-023-00680-1

Show all 104 references
  1. [5]

    Victor R Basili, Forrest Shull, and Filippo Lanubile. 2002. Building knowledge through families of experiments.IEEE transactions on software engineering25, 4 (2002), 456–473

  2. [6]

    Elaine Beller, Justin Clark, Guy Tsafnat, Clive Adams, Heinz Diehl, Hans Lund, Mourad Ouzzani, Kristina Thayer, James Thomas, Tari Turner, Jun Xia, Karen Robinson, Paul Glasziou, Olga Ahtirschi, Robin Christensen, Julian Elliott, Sergio Graziosi, Joel Kuiper, Rasmus Moustgaard...

  3. [7]

    Marcel Binz, Stephan Alaniz, Adina Roskies, Balazs Aczel, Carl T Bergstrom, Colin Allen, Daniel Schad, Dirk Wulff, Jevin D West, Qiong Zhang, et al. 2025. How should the advancement of large language models affect the practice of science?Proceedings of the National Academy of ...

  4. [8]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey W...

  5. [9]

    Christian Cao, Jason Sang, Rohit Arora, Robbie Kloosterman, Matt Cecere, Jaswanth Gorla, Richard Saleh, David Chen, Ian Dren- nan, Bijan Teja, Michael Fehlings, Paul Ronksley, Alexander A Leung, Dany E Weisz, Harriet Ware, Mairead Whelan, David B Emer- son, Rahul Arora, and Ni...

  6. [10]

    Bruno Cartaxo, Gustavo Pinto, and Sergio Soares. 2018. The role of rapid reviews in supporting decision-making in software engineering practice. InProceedings of the 22nd International Conference on Evaluation and Assessment in Software Engineering 2018. 24–34

  7. [11]

    Pablo Castillo-Segura, Carlos Alario-Hoyos, Carlos Delgado Kloos, and Carmen Fernandez Panadero. 2023. Leveraging the Potential of Generative AI to Accelerate Systematic Literature Reviews: An Example in the Area of Educational Technology. In2023 World Engineering Education Fo...

  8. [12]

    Hsiu-Min Chen. 2024. Utilizing ChatGPT in Systematic Reviews and Meta-Analyses.Hu Li Za Zhi71, 5 (2024), 21–28

  9. [13]

    1986.Software engineering metrics and models

    Samuel Daniel Conte, Hubert E Dunsmore, and YE Shen. 1986.Software engineering metrics and models. Benjamin-Cummings Publishing Co., Inc

  10. [14]

    Kate Crawford. 2024. Generative AI is guzzling water and energy

  11. [15]

    Miranda Cumpston, Tianjing Li, Matthew J Page, Jacqueline Chandler, Vivian A Welch, Julian PT Higgins, and James Thomas. 2019. Updated guidance for trusted systematic reviews: a new edition of the Cochrane Handbook for Systematic Reviews of Interventions.The Cochrane database ...

  12. [16]

    Adéle da Veiga. 2025. Ethical guidelines for the use of generative artificial intelligence and artificial intelligence-assisted tools in scholarly publishing: a thematic analysis.Science Editing12, 1 (2025), 28–34. doi:10.6087/kcse.352

  13. [17]

    Peter Dauvergne. 2022. Is artificial intelligence greening global supply chains? Exposing the political economy of environmental costs.Review of International Political Economy29, 3 (2022), 696–718

  14. [18]

    Matheus de Morais Leça, Lucas Valença, Reydne Santos, and Ronnie de Souza Santos. 2025. Applications and Implications of Large Language Models in Qualitative Analysis: A New Frontier for Empirical Software Engineering. InProceedings of the 2025 IEEE/ACM International Workshop ...

  15. [19]

    Stefano De Paoli. 2024. Performing an inductive thematic analysis of semi-structured interviews with a large language model: An exploration and provocation on the limits of the approach.Social Science Computer Review42, 4 (2024), 997–1019

  16. [20]

    Delgado-Chaves, Matthew J

    Fernando M. Delgado-Chaves, Matthew J. Jennings, Antonio Atalaia, Justus Wolff, Rita Horvath, Zeinab M. Mamdouh, Jan Baumbach, and Linda Baumbach. 2025. Transforming literature screening: The emerging role of large language models in systematic reviews.Proceedings of the Natio...

  17. [21]

    Fabio Dennstädt, Johannes Zink, Paul Martin Putora, Janna Hastings, and Nikola Cihoric. 2024. Title and abstract screening for literature reviews using large language models: an exploratory study in the biomedical domain.Systematic Reviews13, 1 (15 Jun 2024), 158. doi:10.1186/...

  18. [22]

    Declan Devane, Nikita N Burke, Shaun Treweek, Mike Clarke, James Thomas, Andrew Booth, Andrea C Tricco, and KM Saif-Ur-Rahman. 2022. Study within a review (SWAR).Journal of Evidence-Based Medicine15, 4 (2022), 328

  19. [23]

    Tore Dybå and Torgeir Dingsøyr. 2008. Strength of evidence in systematic reviews in software engineering. InProceedings of the Second ACM-IEEE international symposium on Empirical software engineering and measurement. 178–187

  20. [24]

    Van Lissa, Joshua Richard Polanin, Dimitris Mavridis, and Terri D

    Oluwaseun Farotimi, Adam Dunn, Caspar J. Van Lissa, Joshua Richard Polanin, Dimitris Mavridis, and Terri D. Pigott. 2025. Guidance for manuscript submissions testing the use of generative AI for systematic review and meta-analysis.Research Synthesis Methods(2025). Editorial; p...

  21. [25]

    Katia Romero Felizardo, Márcia Sampaio Lima, Anderson Deizepe, Tayana Uchôa Conte, and Igor Steinmacher. 2024. ChatGPT application in Systematic Literature Reviews in Software Engineering: an evaluation of its accuracy to support the selection activity. InProceedings of the 18...

  22. [26]

    Kate Flemming and Jane Noyes. 2021. Qualitative evidence synthesis: where are we at?International journal of qualitative methods20 (2021), 1609406921993276

  23. [27]

    Tron Foss, Erik Stensrud, Barbara Kitchenham, and Ingunn Myrtveit. 2003. A simulation study of the model evaluation criterion MMRE.IEEE Transactions on Software Engineering29, 11 (2003), 985–995

  24. [28]

    Gallegos, R.A

    I.O. Gallegos, R.A. Rossi, J. Barrow, M.M. Tanjim, S. Kim, F. Dernoncourt, T. Yu, R. Zhang, and N.K. Ahmed. 2025. Bias and fairness in large language models: A survey.Computational Linguistics50, 3 (2025), 1097–1179

  25. [29]

    Jie Gao, Yuchen Guo, Gionnieve Lim, Tianqin Zhang, Zheng Zhang, Toby Jia-Jun Li, and Simon Tangi Perrault. 2024. CollabCoder: a lower-barrier, rigorous workflow for inductive collaborative qualitative analysis with large language models. InProceedings of the 2024 CHI Conferenc...

  26. [30]

    Elizabeth Gibney. 2025. What are the best AI tools for research? Nature’s guide.Nature(Feb 2025)

  27. [31]

    Samuel Greengard. 2025. Shining a Light on AI Hallucinations.Commun. ACM68, 5 (April 2025), 9–11. doi:10.1145/3715691

  28. [32]

    Eddie Guo, Mehul Gupta, Jiawen Deng, Ye-Jean Park, Michael Paget, and Christopher Naugler. 2024. Automated Paper Screening for Clinical Reviews Using Large Language Models: Data Analysis Study.J Med Internet Res26 (jan 2024), e48996

  29. [33]

    2024.Utilizing Large Language Models to Update Systematic Literature Reviews: A Case Study on Time Pressure and Well-Being in Software Engineering

    Jingyi Guo. 2024.Utilizing Large Language Models to Update Systematic Literature Reviews: A Case Study on Time Pressure and Well-Being in Software Engineering. Master’s thesis. University of Helsinki, Faculty of Science, Helsinki, Finland. https://helda-test-22.hulib.helsinki....

  30. [34]

    Gonzalo Génova, Juan Llorens, and Jorge Morato. 2012. Software Engineering Research: The Need to Strengthen and Broaden the Classical Scientific Method. InResearch Methodologies, Innovations and Philosophies in Software Systems Engineering and Information Systems. IGI Global, ...

  31. [35]

    Leah Hamilton, Desha Elliott, Aaron Quick, Simone Smith, and Victoria Choplin. 2023. Exploring the use of AI in qualitative analysis: A comparative study of guaranteed income data.International journal of qualitative methods22 (2023), 16094069231201504

  32. [36]

    Horace He et al. 2025. Defeating Nondeterminism in LLM Inference.Thinking Machines Lab: Connectionism(2025). doi:10.64434/tml.20250910 https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/

  33. [37]

    Jiangen He. 2025. Who Gets Cited? Gender- and Majority-Bias in LLM-Driven Reference Selection. arXiv:2508.02740 [cs.DL] https://arxiv.org/abs/ 2508.02740

  34. [38]

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2025. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.ACM Trans. Inf. Sy...

  35. [39]

    Aleksi Huotala, Miikka Kuutila, and Mika Mäntylä. 2025. Research artifacts in secondary studies: A systematic mapping in software engineering. Information and Software Technology187 (Nov. 2025), 107830. doi:10.1016/j.infsof.2025.107830

  36. [40]

    Aleksi Huotala, Miikka Kuutila, and Mika Mäntylä. 2025. SESR-Eval: Dataset for Evaluating LLMs in the Title-Abstract Screening of Systematic Reviews. In2025 ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM). ACM, 01–12

  37. [41]

    Aleksi Huotala, Miikka Kuutila, Paul Ralph, and Mika Mäntylä. 2024. The Promise and Challenges of Using LLMs to Accelerate the Screening Process of Systematic Reviews. InProceedings of the 28th International Conference on Evaluation and Assessment in Software Engineering(Saler...

  38. [42]

    Aleksi Huotala, Miikka Kuutila, Olli-Pekka Turtio, Simo Sipilä, and Mika Mäntylä. 2026. AISysRev – LLM-based Tool for Title-abstract Screening. arXiv:2510.06708 [cs.SE] https://arxiv.org/abs/2510.06708 56 Kitchenham et al

  39. [43]

    Maha Inam, Sana Sheikh, Abdul Mannan Khan Minhas, Elizabeth M Vaughan, Chayakrit Krittanawong, Zainab Samad, Carl J Lavie, Adeel Khoja, Melaine D’Cruze, Leandro Slipczuk, et al. 2024. A review of top cardiology and cardiovascular medicine journal guidelines regarding the use o...

  40. [44]

    Mlađan Jovanović and Mark Campbell. 2025. Reasoning AI: Progress, Challenges, and the Path Forward for Large Reasoning and Concept Models. Computer58, 11 (2025), 113–119

  41. [45]

    Vempala, and Edwin Zhang

    Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang. 2025. Why Language Models Hallucinate. arXiv:2509.04664 [cs.CL] https://arxiv.org/abs/2509.04664

  42. [46]

    Fernando Kamei, Igor Wiese, Crescencio Lima, Ivanilton Polato, Vilmar Nepomuceno, Waldemar Ferreira, Márcio Ribeiro, Carolline Pena, Bruno Cartaxo, Gustavo Pinto, and Sérgio Soares. 2021. Grey Literature in Software Engineering: A critical review.Information and Software Techn...

  43. [47]

    Qusai Khraisha, Sophie Put, Johanna Kappenberg, Azza Warraitch, and Kristin Hadfield. 2024. Can large language models replace humans in systematic reviews? Evaluating GPT-4’s efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages...

  44. [48]

    Jin K Kim, Michael Chua, Mandy Rickard, and Armando Lorenzo. 2023. ChatGPT and large language model (LLM) chatbots: The current state of acceptability and a proposal for guidelines on utilization in academic medicine.Journal of Pediatric Urology19, 5 (2023), 598–604. doi:10.10...

  45. [49]

    Barbara Kitchenham, Lech Madeyski, and Pearl Brereton. 2019. Problems with statistical practice in human-centric software engineering experiments. InProceedings of the 23rd International Conference on Evaluation and Assessment in Software Engineering. 134–143. doi:10.1145/ 331...

  46. [50]

    Barbara Kitchenham, Lech Madeyski, and Pearl Brereton. 2020. Meta-analysis for families of experiments in software engineering: a systematic review and reproducibility and validity assessment.Empirical Software Engineering25 (2020), 353–401. doi:10.1007/s10664-019-09747-0

  47. [51]

    Barbara Kitchenham, Lech Madeyski, and David Budgen. 2023. SEGRESS: Software engineering guidelines for reporting secondary studies.IEEE Transactions on Software Engineering49, 3 (2023), 1273–1298. doi:10.1109/TSE.2022.3174092

  48. [52]

    Hoaglin, Khaled El Emam, and Jarrett Rosenberg

    Barbara A Kitchenham, Shari Lawrence Pfleeger, Lesley M Pickard, Peter W Jones, David C. Hoaglin, Khaled El Emam, and Jarrett Rosenberg. 2002. Preliminary guidelines for empirical research in software engineering.IEEE Transactions on software engineering28, 8 (2002), 721–734

  49. [53]

    Kitchenham, Lesley M

    Barbara A. Kitchenham, Lesley M. Pickard, Stephen G. MacDonell, and Martin Shepperd. 2001. What accuracy statistics really measure.IEE Proceedings - Software Engineering148, 3 (2001), 81–85

  50. [54]

    Amy J. Ko. 2019.Why We Should Not Measure Productivity. Apress, Berkeley, CA, 21–26. doi:10.1007/978-1-4842-4221-6_3

  51. [55]

    Laura Krefting. 1991. Rigor in qualitative research: The assessment of trustworthiness.The American journal of occupational therapy45, 3 (1991), 214–222

  52. [56]

    Patricia Lago, Per Runeson, Qunying Song, and Roberto Verdecchia. 2024. Threats to Validity in Software Engineering – hypocritical paper section or essential analysis?. InProceedings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement...

  53. [57]

    Honghao Lai, Long Ge, Mingyao Sun, Bei Pan, Jiajie Huang, Liangying Hou, Qiuyu Yang, Jiayi Liu, Jianing Liu, Ziying Ye, et al. 2024. Assessing the risk of bias in randomized clinical trials with large language models.JAMA Network Open7, 5 (2024), e2412687–e2412687

  54. [58]

    Simon Lewin, Andrew Booth, Claire Glenton, Heather Munthe-Kaas, Arash Rashidian, Megan Wainwright, Meghan A Bohren, Özge Tunçalp, Christopher J Colvin, Ruth Garside, et al. 2018. Applying GRADE-CERQual to qualitative evidence synthesis findings: introduction to the series. Imp...

  55. [59]

    James H Lubowitz. 2024. Guidelines for the use of generative artificial intelligence tools for biomedical journal authors and reviewers.Arthroscopy: The Journal of Arthroscopic & Related Surgery40, 3 (2024), 651–652. doi:10.1016/j.arthro.2023.10.037

  56. [60]

    Xufei Luo, Han Lyu, Qianling Shi, Zijun Wang, Hui Liu, Di Zhu, Ye Wang, and Yaolong Chen. 2024. The application of large language models in the field of evidence-based medicine.Chinese Journal of Evidence-Based Medicine(2024)

  57. [62]

    Jorge Melegati, Nicolas Nascimento, Rafael Chanin, Afonso Sales, and Igor Wiese. 2024. Exploring potential implications of intelligent tools for human aspects of software engineering. InProceedings of the 2024 IEEE/ACM 17th International Conference on Cooperative and Human Asp...

  58. [63]

    automatic patch generation learned from human-written patches

    Martin Monperrus. 2014. A critical review of "automatic patch generation learned from human-written patches": essay on the problem statement and the evaluation of automatic software repair. InProceedings of the 36th International Conference on Software Engineering(Hyderabad, I...

  59. [64]

    David L Morgan. 2023. Exploring the use of artificial intelligence for qualitative data analysis: The case of ChatGPT.International journal of qualitative methods22 (2023), 16094069231211248

  60. [65]

    Meredith Ringel Morris. 2024. Prompting considered harmful.Commun. ACM67, 12 (2024), 28–30

  61. [66]

    Cynthia D Mulrow. 1987. The medical review article: state of the science.Annals of internal medicine106, 3 (1987), 485–488. Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews 57

  62. [67]

    Ingunn Myrtveit and Erik Stensrud. 2012. Validity and reliability of evaluation procedures in comparative studies of effort prediction models. Empirical Software Engineering17 (2012), 23–33

  63. [68]

    Ingunn Myrtveit, Erik Stensrud, and Martin Shepperd. 2005. Reliability and validity in comparative studies of software prediction models.IEEE Transactions on Software Engineering31, 5 (2005), 380–391

  64. [69]

    Mahjabin Nahar, Sian Lee, Rebekah Guillen, and Dongwon Lee. 2025. Generative Artificial Intelligence Policies under the Microscope.Commun. ACM68, 7 (2025), 29–33

  65. [70]

    Agnes Natukunda and Leacky K. Muchene. 2023. Unsupervised title and abstract screening for systematic review: a retrospective case-study using topic modelling methodology.Systematic Reviews12, 1 (03 Jan 2023), 1. doi:10.1186/s13643-022-02163-4

  66. [71]

    Seyed Aria Nejadghaderi, Maryam Balibegloo, and Nima Rezaei. 2024. The Cochrane risk of bias assessment tool 2 (RoB 2) versus the original RoB: A perspective on the pros and cons.Health Science Reports7, 6 (2024), e2165

  67. [72]

    Ojelanki Ngwenyama and Frantz Rowe. 2024. Should We Collaborate with AI to Conduct Literature Reviews? Changing Epistemic Values in a Flattening World.Journal of the Association for Information Systems25 (01 2024), 122–136. doi:10.17705/1jais.00869

  68. [73]

    Matthew J Page, Joanne E McKenzie, Patrick M Bossuyt, Isabelle Boutron, Tammy C Hoffmann, Cynthia D Mulrow, Larissa Shamseer, Jennifer M Tetzlaff, Elie A Akl, Sue E Brennan, et al. 2021. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews.BMJ372 (2021)

  69. [74]

    Helen Pearson. 2024. Can AI review the scientific literature - and figure out what it all means?Nature635, 8038 (nov 2024), 276–278

  70. [75]

    Mike Perkins and Jasper Roe. 2024. Academic publisher guidelines on AI usage: A ChatGPT supported thematic analysis. 1398 pages. doi:10.12688/ f1000research.142411.2

  71. [76]

    Mike Perkins and Jasper Roe. 2024. The use of Generative AI in qualitative analysis: Inductive thematic analysis with ChatGPT.Journal of Applied Learning and Teaching7, 1 (2024), 390–395

  72. [77]

    Kai Petersen and Jan M. Gerken. 2025. On the road to interactive LLM-based systematic mapping studies.Information and Software Technology178 (2025), 107611. doi:10.1016/j.infsof.2024.107611

  73. [78]

    Sebastián Pizard, Joaquín Lezama, Rodrigo García, Diego Vallespir, and Barbara Kitchenham. 2025. Using rapid reviews to support software engineering practice: a systematic review and a replication study.Empirical Software Engineering30, 1 (2025), 10

  74. [80]

    Paul Ralph. 2021. ACM SIGSOFT empirical standards released.ACM SIGSOFT Software Engineering Notes46, 1 (2021), 19–19

  75. [81]

    Wu, Abdullah Pandor, Munira Essat, Mark Stevenson, and Xingyi Song

    Ambrose Robinson, William Thorne, Ben P. Wu, Abdullah Pandor, Munira Essat, Mark Stevenson, and Xingyi Song. 2023. Bio-SIEVE: Exploring Instruction Tuning Large Language Models for Systematic Review Automation. arXiv:2308.06610 [cs.CL] https://arxiv.org/abs/2308.06610

  76. [82]

    M Sandelowski and J Barroso. 2007. Optimizing the validity of qualitative research synthesis studies.Sandelowski M, Barroso J. Handbook for synthesizing qualitative research. NewYork: Springer Publishing(2007), 227–230

  77. [83]

    Scott Spillias, Paris Tuohy, Matthew Andreotta, Ruby Annand-Jones, Fabio Boschetti, Christopher Cvitanovic, Joseph Duggan, Elisabeth A Fulton, Denis B Karcher, Cecile Paris, et al. 2024. Human-AI collaboration to identify literature for evidence synthesis.Cell Reports Sustaina...

  78. [84]

    Stuart, Yiftach Fehige, and James Robert Brown

    Michael T. Stuart, Yiftach Fehige, and James Robert Brown. 2018. Thought Experiments: State of the Art. InThe Routledge Companion to Thought Experiments, Michael T. Stuart, Yiftach Fehige, and James Robert Brown (Eds.). Routledge, 1–28

  79. [85]

    Teo Susnjak, Peter Hwang, Napoleon Reyes, Andre L. C. Barczak, Timothy McIntosh, and Surangika Ranathunga. 2025. Automating Research Synthesis with Domain-Specific Large Language Model Fine-Tuning.ACM Transactions on Knowledge Discovery from Data19, 3 (March 2025), 1–39. doi:1...

  80. [86]

    Simon Šuster, Timothy Baldwin, and Karin Verspoor. 2024. Zero-and few-shot prompting of generative large language models provides weak assessment of risk of bias in clinical trials.Research Synthesis Methods15, 6 (2024), 988–1000

  81. [87]

    Eugene Syriani, Istvan David, and Gauransh Kumar. 2024. Screening articles for systematic reviews with ChatGPT.Journal of Computer Languages 80 (2024), 101287. doi:10.1016/j.cola.2024.101287

  82. [88]

    James Thomas, Ella Flemyng, Anna Noel-Storr, Will Moy, Iain Marshall, Raouf Hajji, Zoe Jordan, Edoardo Aromataris, Samer Mheissen, Justin Clark, Paweł Jemioło, Ashrita Saran, Michelle Angrish, Biljana Macura, Rene Spijker, Neal Haddaway, Wojciech Kusa, Yuan Chi, Isabel Fletche...

  83. [89]

    Walker, Michelle Angrish, Biljana Macura, Rene Spijker, Neal Haddaway, Wojciech Kusa, Yuan Chi, and Isabel Fletcher

    James Thomas, Ella Flemyng, Anna Noel-Storr, Will Moy, Iain Marshall, Raouf Hajji, Zoe Jordan, Edoardo Aromataris, Samer Mheissen, Justin Clark, Paweł Jemioło, Ashrita Saran, Vickie R. Walker, Michelle Angrish, Biljana Macura, Rene Spijker, Neal Haddaway, Wojciech Kusa, Yuan C...

  84. [91]

    Berg, Yuan Chi, A

    James Thomas, Kaitlyn Hair, Anna Noel-Storr, Michelle Angrish, Edoardo Aromataris, Lisa Askie, Rigmor C. Berg, Yuan Chi, A. Justin Clark, Declan Devane, Emma McFarlane, Isabel Fletcher, Gerald Gartlehner, Jonas Goretzko, Neal Haddaway, Raouf Hajji, Paweł Jemioło, Zoe Jordan, J...

  85. [92]

    Berg, Yuan Chi, A

    James Thomas, Kaitlyn Hair, Anna Noel-Storr, Michelle Angrish, Edoardo Aromataris, Lisa Askie, Rigmor C. Berg, Yuan Chi, A. Justin Clark, Declan Devane, Emma McFarlane, Isabel Fletcher, Gerald Gartlehner, Jonas Goretzko, Neal Haddaway, Raouf Hajji, Paweł Jemioło, Zoe Jordan, J...

  86. [93]

    Uribe, Ilze Maldupa, and Falk Schwendicke

    Sergio E. Uribe, Ilze Maldupa, and Falk Schwendicke. 2025. Integrating Generative AI in Dental Education: A Scoping Review of Current Practices and Recommendations.European Journal of Dental Education29, 2 (2025), 341–355. doi:10.1111/eje.13074

  87. [94]

    Raymon van Dinter, Bedir Tekinerdogan, and Cagatay Catal. 2021. Automation of systematic literature reviews: A systematic literature review. Information and Software Technology136 (2021), 106589. doi:10.1016/j.infsof.2021.106589

  88. [95]

    Sira Vegas, Cecilia Apa, and Natalia Juristo. 2016. Crossover Designs in Software Engineering Experiments: Benefits and Perils.IEEE Transactions on Software Engineering42, 2 (2016), 120–135. doi:10.1109/TSE.2015.2467378

  89. [96]

    Jonas Wachinger, Kate Bärnighausen, Louis N Schäfer, Kerry Scott, and Shannon A McMahon. 2024. Prompts, pearls, imperfections: comparing ChatGPT and a human researcher in qualitative data analysis.Qualitative Health Research(2024), 10497323241244669

  90. [98]

    Shuai Wang, Harrisen Scells, Bevan Koopman, and Guido Zuccon. 2023. Can ChatGPT Write a Good Boolean Query for Systematic Review Literature Search?. InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval(Taipei, Taiwan...

  91. [99]

    Shuai Wang, Harrisen Scells, Shengyao Zhuang, Martin Potthast, Bevan Koopman, and Guido Zuccon. 2024. Zero-Shot Generative Large Language Models for Systematic Review Screening Automation. InAdvances in Information Retrieval, Nazli Goharian, Nicola Tonellotto, Yulan He, Aldo L...

  92. [100]

    Weitzmann

    E.A. Weitzmann. 2003. Software and Qualitative Research. InCollecting and Interpreting Qualitative Materials, N. K. Denzin and Y. S. Lincoln (Eds.). Sage, 310–339

  93. [101]

    David Wilkins. 2023. Automated title and abstract screening for scoping reviews using the GPT-4 Large Language Model. arXiv:2311.07918 [cs.CL] https://arxiv.org/abs/2311.07918

  94. [102]

    Tim Woelfle, Julian Hirt, Perrine Janiaud, Ludwig Kappos, John PA Ioannidis, and Lars G Hemkens. 2024. Benchmarking Human–AI collaboration for common evidence appraisal tools.Journal of Clinical Epidemiology175 (2024), 111533

  95. [103]

    Claes Wohlin. 2021. Case Study Research in Software Engineering—It is a Case, and it is a Study, but is it a Case Study?Information and Software Technology133 (2021), 106514

  96. [104]

    Robert Zimmermann, Marina Staab, Mehran Nasseri, and Patrick Brandtner. 2024. Leveraging Large Language Models for Literature Review Tasks - A Case Study Using ChatGPT. InAdvanced Research in Technologies, Information, Innovation and Sustainability, Teresa Guarda, Filipe Porte...

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.