REVIEW 9 cited by
Using Large Language Models for Qualitative Analysis can Introduce Serious Bias
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large Language Models (LLMs) are quickly becoming ubiquitous, but the implications for social science research are not yet well understood. This paper asks whether LLMs can help us analyse large-N qualitative data from open-ended interviews, with an application to transcripts of interviews with Rohingya refugees in Cox's Bazaar, Bangladesh. We find that a great deal of caution is needed in using LLMs to annotate text as there is a risk of introducing biases that can lead to misleading inferences. We here mean bias in the technical sense, that the errors that LLMs make in annotating interview transcripts are not random with respect to the characteristics of the interview subjects. Training simpler supervised models on high-quality human annotations with flexible coding leads to less measurement error and bias than LLM annotations. Therefore, given that some high quality annotations are necessary in order to asses whether an LLM introduces bias, we argue that it is probably preferable to train a bespoke model on these annotations than it is to use an LLM for annotation.
Forward citations
Cited by 9 Pith papers
-
EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage
EgoPolice introduces a 185-hour annotated police body-worn camera benchmark showing state-of-the-art video models fail on high-stakes actions due to motion, occlusion, and low inter-class visual separability.
-
Bias, Accuracy, and Trust: Gender-Diverse Perspectives on Large Language Models
Gender-diverse users perceive ChatGPT's gender bias differently, with non-binary/transgender participants reporting condescending and stereotypical responses, and men reporting higher trust.
-
ChatCollab: Exploring Collaboration Between Humans and AI Agents in Software Teams
ChatCollab lets humans and AI agents work side-by-side in software teams and offers a way to measure how those teams collaborate.
-
LLM-Assisted Automated Deductive Coding of Dialogue Data: Leveraging Dialogue-Specific Characteristics to Enhance Contextual Understanding
An LLM-assisted pipeline that separates communicative acts from events, uses multi-model voting, and applies a consistency check reaches Cohen's kappa above 0.80 with human coders on a small student-dialogue corpus.
-
LLMCode: Evaluating and Enhancing Researcher-AI Alignment in Qualitative Analysis
LLMCode uses IoU and Modified Hausdorff Distance to measure human-LLM coding alignment; two studies with 26 designers show LLMs capture surface coding patterns but not the designer's deeper interpretive lens.
-
From Assistance to Autonomy -- A Researcher Study on the Potential of AI Support for Qualitative Data Analysis
Interviews with 15 HCI researchers show openness to AI in qualitative data analysis under conditions of privacy, control, and reliability, leading to a framework of AI involvement levels from minimal to high.
-
Empowering Computing Education Researchers Through LLM-Assisted Content Analysis
The paper proposes LACA, a reproducible protocol in which humans build a codebook, an LLM performs deductive coding, and interrater reliability checks gate whether the LLM can code the full dataset.
-
Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications
LLMs alone annotate search clarifications unreliably; adding confidence-based selective human review cuts effort 24-45% in simulation, but the evaluation is partly built from the ground truth it predicts.
-
Potential and Perils of Large Language Models as Judges of Unstructured Textual Data
LLM judges show only fair-to-moderate agreement with human raters on thematic summary alignment and consistently over-rate alignment compared to humans.
Discussion (0). Continue with ORCID to comment.