REVIEW 6 cited by
Towards Evaluation Guidelines for Empirical Studies involving LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In the short period since the release of ChatGPT, large language models (LLMs) have changed the software engineering research landscape. While there are numerous opportunities to use LLMs for supporting research or software engineering tasks, solid science needs rigorous empirical evaluations. However, so far, there are no specific guidelines for conducting and assessing studies involving LLMs in software engineering research. Our focus is on empirical studies that either use LLMs as part of the research process or studies that evaluate existing or new tools that are based on LLMs. This paper contributes the first set of holistic guidelines for such studies. Our goal is to start a discussion in the software engineering research community to reach a common understanding of our standards for high-quality empirical studies involving LLMs.
Forward citations
Cited by 6 Pith papers
-
A Methodological Framework for LLM-Based Mining of Software Repositories
A rapid review and survey of LLM-based repository mining yield a threat-mitigation map and the six-stage PRIMES 2.0 framework for conducting such studies.
-
Investigating the Use of LLMs for Evidence Briefings Generation in Software Engineering
A registered report protocol for comparing LLM-generated and human-made software engineering evidence briefings is laid out, but no experimental results are reported yet.
-
ReqBrain: Task-Specific Instruction Tuning of LLMs for AI-Assisted Requirements Generation
ReqBrain, a LoRA-fine-tuned Zephyr-7b-beta model, produces software requirements that human evaluators could not reliably tell apart from human-authored ones, with automatic metrics favoring it over untuned ChatGPT-4o.
-
Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews
GUEST gives SE researchers process recommendations for planning, conducting, and reporting GenAI-supported SLRs and independent GenAI tool evaluations under mandatory human oversight.
-
Empowering Computing Education Researchers Through LLM-Assisted Content Analysis
The paper proposes LACA, a reproducible protocol in which humans build a codebook, an LLM performs deductive coding, and interrater reliability checks gate whether the LLM can code the full dataset.
-
Get on the Train or be Left on the Station: Using LLMs for Software Engineering Research
The paper maps LLM impacts onto the SE research pipeline with McLuhan's Tetrad, predicting enhancements, obsolescence, retrievals, and reversals, and calls for community action.
Discussion (0). Continue with ORCID to comment.