Pith. sign in

REVIEW 9 cited by

Using Large Language Models for Qualitative Analysis can Introduce Serious Bias

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.17147 v2 pith:RRXED6UI submitted 2023-09-29 cs.CL cs.AIecon.GNq-fin.EC

classification cs.CLcs.AIecon.GNq-fin.EC
keywords annotationsbiasllmsmodelsinterviewinterviewslanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large Language Models (LLMs) are quickly becoming ubiquitous, but the implications for social science research are not yet well understood. This paper asks whether LLMs can help us analyse large-N qualitative data from open-ended interviews, with an application to transcripts of interviews with Rohingya refugees in Cox's Bazaar, Bangladesh. We find that a great deal of caution is needed in using LLMs to annotate text as there is a risk of introducing biases that can lead to misleading inferences. We here mean bias in the technical sense, that the errors that LLMs make in annotating interview transcripts are not random with respect to the characteristics of the interview subjects. Training simpler supervised models on high-quality human annotations with flexible coding leads to less measurement error and bias than LLM annotations. Therefore, given that some high quality annotations are necessary in order to asses whether an LLM introduces bias, we argue that it is probably preferable to train a bespoke model on these annotations than it is to use an LLM for annotation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage

    cs.CV 2026-07 conditional novelty 7.0 of 10

    EgoPolice introduces a 185-hour annotated police body-worn camera benchmark showing state-of-the-art video models fail on high-stakes actions due to motion, occlusion, and low inter-class visual separability.

  2. Bias, Accuracy, and Trust: Gender-Diverse Perspectives on Large Language Models

    cs.HC 2025-06 conditional novelty 6.0 of 10

    Gender-diverse users perceive ChatGPT's gender bias differently, with non-binary/transgender participants reporting condescending and stereotypical responses, and men reporting higher trust.

  3. ChatCollab: Exploring Collaboration Between Humans and AI Agents in Software Teams

    cs.HC 2024-12 conditional novelty 6.0 of 10

    ChatCollab lets humans and AI agents work side-by-side in software teams and offers a way to measure how those teams collaborate.

  4. LLM-Assisted Automated Deductive Coding of Dialogue Data: Leveraging Dialogue-Specific Characteristics to Enhance Contextual Understanding

    cs.CL 2025-04 conditional novelty 5.0 of 10

    An LLM-assisted pipeline that separates communicative acts from events, uses multi-model voting, and applies a consistency check reaches Cohen's kappa above 0.80 with human coders on a small student-dialogue corpus.

  5. LLMCode: Evaluating and Enhancing Researcher-AI Alignment in Qualitative Analysis

    cs.HC 2025-04 conditional novelty 5.0 of 10

    LLMCode uses IoU and Modified Hausdorff Distance to measure human-LLM coding alignment; two studies with 26 designers show LLMs capture surface coding patterns but not the designer's deeper interpretive lens.

  6. From Assistance to Autonomy -- A Researcher Study on the Potential of AI Support for Qualitative Data Analysis

    cs.CY 2025-01 conditional novelty 5.0 of 10

    Interviews with 15 HCI researchers show openness to AI in qualitative data analysis under conditions of privacy, control, and reliability, leading to a framework of AI involvement levels from minimal to high.

  7. Empowering Computing Education Researchers Through LLM-Assisted Content Analysis

    cs.CL 2025-08 conditional novelty 4.0 of 10

    The paper proposes LACA, a reproducible protocol in which humans build a codebook, an LLM performs deductive coding, and interrater reliability checks gate whether the LLM can code the full dataset.

  8. Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications

    cs.IR 2025-07 reject novelty 4.0 of 10

    LLMs alone annotate search clarifications unreliably; adding confidence-based selective human review cuts effort 24-45% in simulation, but the evaluation is partly built from the ground truth it predicts.

  9. Potential and Perils of Large Language Models as Judges of Unstructured Textual Data

    cs.CL 2025-01 conditional novelty 4.0 of 10

    LLM judges show only fair-to-moderate agreement with human raters on thematic summary alignment and consistently over-rate alignment compared to humans.

Pith tools