Pith. sign in

REVIEW 4 major objections 5 minor 3 cited by

Potential and Perils of Large Language Models as Judges of Unstructured Textual Data

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper claims that LLM-as-judge models rate the thematic alignment of LLM-generated summaries about as well as human raters, while systematically overestimating alignment in nuanced cases.

desk verdict A useful small study whose headline claim outruns its own numbers; the missing human-human baseline makes 'comparable to human raters' uninterpretable, but the empirical pattern is worth a careful look. read the letter →

arxiv 2501.08167 v2 pith:VI2KAS3D submitted 2025-01-14 cs.CL cs.AIcs.CY

classification cs.CLcs.AIcs.CY
keywords LLM-as-judgelargelanguagemodelsthematicalignmentopen-endedsurveydatainter-rateragreementKrippendorff'salphaqualitativeresearch
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether one large language model can reliably judge whether a summary produced by another LLM faithfully captures the themes in open-ended survey responses. Using 70 thematic summaries distilled from over 13,000 survey comments, the authors compared human ratings with five different LLM judges on a three-point alignment scale. Human-model agreement was moderate (Cohen's kappa 0.34–0.44; 76–79% exact agreement), while model-model agreement was higher (kappa up to 0.70). The authors conclude that LLM-as-judge offers a scalable alternative to human evaluation, with the caveat that humans are better at spotting subtle, context-specific misalignments. A sympathetic reader would take the central contribution to be a benchmark of LLM-judge reliability on a realistic organizational dataset, together with a warning about where it falls short.

What carries the argument

The machinery is the LLM-as-judge evaluation loop: a generation model (Anthropic Claude) produces thematic summaries, and a separate judge model (Titan Express, Nova Pro, Claude Sonnet 3.5, Llama 3.3 70b, or a human) rates each theme's name, description, and quote using a structured classification prompt on a three-point alignment scale. Agreement between rater pairs is then quantified with Cohen's kappa (chance-adjusted agreement), Spearman's rho (rank-order correlation), and Krippendorff's alpha in both ordinal and nominal forms, so the paper can distinguish exact agreement from ordinal agreement and from agreement adjusted for chance. This setup lets the authors directly compare human-model and model-model reliability on identical rating tasks.

What would settle it

Re-run the evaluation with a second independent team of human raters using the same prompt and scale; if human-human Cohen's kappa is not substantially above the best LLM-human kappa of 0.44, then 'comparable to human raters' loses its meaning, because the human yardstick is no more consistent than the LLMs being measured. A more targeted test is to count how often LLM judges assign 'completely aligned' to summaries that humans grade 'somewhat aligned' or 'not aligned'; if that overestimation rate is high in a decision-relevant sample, the 'scalable solution' claim fails in practice.

Watch

Extended reading notes

Core claim

The central claim is that LLM-as-judge models can evaluate the thematic alignment of LLM-generated summaries at a level comparable to human raters, as measured by standard inter-rater agreement metrics. On a novel dataset of 70 thematic summaries (each containing three themes with name, description, and representative quote), the best LLM judge reached a Cohen's kappa of 0.44 against human ratings, with 79% exact agreement, and Spearman's rho around 0.62. The paper also documents a systematic overestimation: LLMs often rated themes 'completely aligned' when human raters saw partial or no alignment, particularly when the representative quote or description missed specific details. The authors interpret this as evidence that LLM judges can serve as scalable proxy evaluators, but that human oversight remains necessary for context-sensitive content.

Load-bearing premise

The load-bearing premise is that the human ratings used as the baseline are themselves stable and reproducible; the paper reports no inter-human reliability, rater count, or adjudication process for those ratings.

Editorial extensions

If this is right

  • Organizations can use LLM judges as a scalable first-pass screen for thematic summaries of open-ended survey data, cutting evaluation cost and turnaround time.
  • Because model-model agreement exceeds model-human agreement, relying on a committee of LLM judges may reinforce shared blind spots instead of recovering human nuance.
  • The results argue for empirical model selection in judge tasks, since an older Claude v2.1 beat newer models on agreement with humans.
  • The systematic overestimation of alignment means automated pipelines should include human spot-checks or an explicit low-confidence category before summaries are used in high-stakes decisions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • In my reading, the paper's 'comparable to human raters' phrasing overstates what the evidence supports, because without inter-human reliability the benchmark is unanchored; the defensible claim is 'moderate LLM-human agreement.'
  • A natural next experiment would repeat the study with multiple human rater teams and a finer rating scale (e.g., 5 points) to test whether the LLM overestimation shrinks when partial misalignment can be expressed.
  • The single free-text question about what workers want leaders to know is a high-context, emotionally loaded prompt; the observed overestimation might be more severe for such prompts than for factual summarization, so generalizing the kappa values to other survey topics should be done cautiously.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper investigates whether LLM-as-judge models can evaluate the thematic alignment of LLM-generated summaries of open-ended survey responses as reliably as human raters. The authors used an Anthropic Claude model to generate 70 thematic summaries from a proprietary census survey dataset, then had human raters and five LLM judges (Claude 2.1, Titan Express, Sonnet 3.5, Llama 3.3 70b, Nova Pro) rate theme-name/description/quote alignment on a 1–3 ordinal scale. Agreement was quantified with percentage agreement, Cohen's kappa, Spearman's rho, and Krippendorff's alpha (ordinal and nominal). The results show human–model agreement with Cohen's kappa 0.34–0.44, ordinal alpha 0.49–0.60, and percentage agreement 76–79%. The paper concludes that LLM-as-judges offer a scalable solution comparable to human raters, while humans may still be better at subtle, context-specific nuances.

Significance. If the central claim were supported, the paper would provide a useful, concrete data point on LLM-based evaluation of qualitative text, with practical implications for organizational survey analysis. The authors report a real, non-public dataset, compare multiple models, and compute several standard agreement metrics, which is a strength. However, the headline conclusion that LLM judges are 'comparable to human raters' is not supported by the paper's own reported thresholds, and the absence of a human–human reliability baseline makes that comparison uninterpretable. The paper also misuses Spearman's rho as a test of data sufficiency and fails to disclose the exact LLM version used for summary generation. These issues are load-bearing for the main claim.

major comments (4)
  1. [Section 3.2.1 / Table 1] The central claim that LLM-as-judges are 'comparable to human raters' requires a human–human reliability baseline, but none is reported. Section 3.2.1 states that human evaluators served as the baseline but provides no information on the number of raters, their background, training, or adjudication protocol, and Table 1 contains no human–human kappa or alpha values. Without knowing how reliably a second set of human raters agrees with the first, the phrase 'comparable to human raters' has no reference point. If human raters typically agree at kappa 0.8, then kappa 0.34–0.44 is not comparable; if they agree at kappa 0.3–0.4, it might be. This missing measurement is not a minor reporting gap; it is the benchmark that defines the paper's headline conclusion.
  2. [Section 5.1 / Abstract] The paper's own threshold for Krippendorff's alpha contradicts the 'comparable to human raters' claim. Section 5 states that alpha above 0.80 signifies strong agreement, 0.67–0.80 acceptable, and below 0.67 low agreement. The reported human–model ordinal alphas in Table 1 are 0.49–0.60, which fall below 0.67 and thus, by the authors' own criterion, indicate low agreement. Cohen's kappa values (0.34–0.44) are also moderate at best. The abstract's assertion that LLM-as-judges offer a scalable solution 'comparable to human raters' is therefore not supported by the reported statistics unless a human–human baseline is supplied and shown to be similarly low.
  3. [Section 4 / Section 5] Spearman's rho is misapplied. Section 4 says 'Spearman's rho determines if enough data has been rated' and 'rho was used as a statistical test on Cohen's kappa as a parameter,' and Section 5 interprets rho as evidence that 'humans and models might not always provide the exact same rating, they tend to rank items similarly.' Spearman's rho is a rank correlation, not a test of data sufficiency and not a test of agreement on a Cohen's kappa parameter. The authors should either use rho as a descriptive correlation between rating scales and interpret it as such, or replace it with a more appropriate inferential statistic for inter-rater agreement.
  4. [Section 3.1 / Section 3.2] The exact version of the Claude model used to generate the thematic summaries is never disclosed. The abstract says 'an Anthropic Claude model' and Section 3.2 says 'a Claude model,' but the specific version (e.g., Claude 2, Claude 2.1, or a specific API snapshot) is not given. Because one of the judges is also Claude 2.1, and because LLM versions materially affect behavior, this omission is a reproducibility problem. The authors should report the exact model identifier, API call date, and any sampling parameters used at the generation stage.
minor comments (5)
  1. [Section 3.2.1] The paper does not state the number of human raters or whether each summary was rated by more than one person. If only one human rating per item was collected, this should be stated and discussed as a constraint on the interpretation of all inter-rater metrics.
  2. [Section 4 / Table 1] The effective sample size for the agreement metrics is ambiguous: 70 thematic summaries with 3 themes each could yield up to 210 rating units, but the paper does not specify whether Table 1 is computed on 70 or 210 observations. Please state the N used for each coefficient.
  3. [Section 3.2.3] The configuration 'top-k value: 0.25' is unusual, as top-k sampling typically uses an integer k. Please clarify the intended parameter and whether this is top-k in the standard sense.
  4. [Appendix] The pseudo-prompt contains a typo: '<them3></them3>' should be '<theme3></theme3>'.
  5. [Section 5.1] The phrase 'low to high substantial reliability' is self-contradictory; the reported ranges are better described as 'low to moderate.'

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the core result is an external human-benchmark comparison; the missing human-human reliability baseline is a validity concern, not a circular reduction.

full rationale

The paper's central claim is an empirical measurement: it compares LLM ratings against human ratings on the same rubric using Cohen's kappa, Spearman's rho, and Krippendorff's alpha (Table 1). This is an external benchmark comparison, not a derivation in which a target quantity is defined in terms of the fitted input. No parameter is fitted and then renamed a prediction; no uniqueness theorem is imported from the authors' prior work; and no central premise is justified only by a self-citation. The self-citations in the background (Section 2, references 3, 4, 5, 9, and 14) are contextual and are not load-bearing for the headline result. The absence of a human-human inter-rater reliability baseline undermines the interpretability of 'comparable to human raters,' and the Claude-generated/Claude-judged setup is a self-evaluation confound that may affect validity. However, these are validity and reproducibility concerns, not circular reductions: the human ratings are not defined in terms of the LLM ratings, and the LLM ratings are not constructed to equal the human ratings by design. Similarly, the paper's own threshold stating that Krippendorff's alpha below 0.67 suggests low agreement sits awkwardly against the reported human-model alphas of 0.49-0.60, but this is an internal consistency and interpretation issue rather than a circular step. Under the hard-rule standard requiring a quotable reduction of the claim to its own inputs, no such circular step can be exhibited here. The appropriate finding is therefore no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The study does not fit any parameters and introduces no new entities. Its claims rest on domain assumptions about the validity of human ratings as a baseline, the representativeness of 70 summaries, and the construct validity of the 1-3 alignment scale. These are reasonable starting points for an empirical paper, but they are unverified and directly affect the interpretation of the agreement statistics.

assumptions (3)
  • domain assumption Human ratings are a valid and stable ground truth for thematic alignment.
    Section 3.2.1 establishes human evaluation as the baseline without reporting the number of raters or inter-human agreement, yet the entire comparison rests on this assumption.
  • domain assumption A single thematic summary per business-line group adequately represents the underlying comments.
    Section 3.1 collapses over 13,000 comments into 70 summaries with three themes each; no check is reported for coverage or representativeness, which affects the reliability of the ratings.
  • domain assumption The alignment of theme name, description, and quote is a well-defined ordinal construct measurable on a 1-3 scale.
    Sections 3.2.2 and 4 define the rating scale and treat it as ordinal; the paper assumes raters (human and LLM) interpret the three dimensions consistently.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Potential and Perils of Large Language Models as Judges of Unstructured Textual Data." pith.science (2026). https://pith.science/paper/VI2KAS3D

@misc{pith2026250108167,
  author       = {Pith},
  title        = {Pith review of: Potential and Perils of Large Language Models as Judges of Unstructured Textual Data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VI2KAS3D}},
  note         = {Machine review of arXiv:2501.08167}
}
read the original abstract

Rapid advancements in large language models have unlocked remarkable capabilities when it comes to processing and summarizing unstructured text data. This has implications for the analysis of rich, open-ended datasets, such as survey responses, where LLMs hold the promise of efficiently distilling key themes and sentiments. However, as organizations increasingly turn to these powerful AI systems to make sense of textual feedback, a critical question arises, can we trust LLMs to accurately represent the perspectives contained within these text based datasets? While LLMs excel at generating human-like summaries, there is a risk that their outputs may inadvertently diverge from the true substance of the original responses. Discrepancies between the LLM-generated outputs and the actual themes present in the data could lead to flawed decision-making, with far-reaching consequences for organizations. This research investigates the effectiveness of LLM-as-judge models to evaluate the thematic alignment of summaries generated by other LLMs. We utilized an Anthropic Claude model to generate thematic summaries from open-ended survey responses, with Amazon's Titan Express, Nova Pro, and Meta's Llama serving as judges. This LLM-as-judge approach was compared to human evaluations using Cohen's kappa, Spearman's rho, and Krippendorff's alpha, validating a scalable alternative to traditional human centric evaluation methods. Our findings reveal that while LLM-as-judge offer a scalable solution comparable to human raters, humans may still excel at detecting subtle, context-specific nuances. Our research contributes to the growing body of knowledge on AI assisted text analysis. Further, we provide recommendations for future research, emphasizing the need for careful consideration when generalizing LLM-as-judge models across various contexts and use cases.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating LLM Agent Collusion in Double Auctions

    cs.GT 2025-07 conditional novelty 4.0 of 10

    LLM sellers in a simulated double auction collude more when they can communicate, and urgency from an authority figure sustains collusion even when an overseer monitors them.

  2. IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A machine-translated version of MMLU-Pro in nine Indic languages is released as a benchmark, with baseline accuracy scores for multilingual LLMs.

  3. Transforming Expert Knowledge into Scalable Ontology via Large Language Models

    cs.AI 2025-06 conditional novelty 3.0 of 10

    An LLM-based taxonomy alignment framework reaches 0.97 F1 using many-shot prompting and expert calibration, but the claimed superiority over the 0.68 human benchmark is based on a non-comparable baseline.

Reference graph

Works this paper leans on

29 extracted references · 17 canonical work pages · cited by 3 Pith papers

  1. [1]

    Mike Allen. 2017. The SAGE encyclopedia of communication research methods . SAGE publications

  2. [2]

    Julian Ashwin, Aditya Chhabra, and Vijayendra Rao. 2023 . Using large lan- guage models for qualitative analysis can introduce seriou s bias. arXiv preprint arXiv:2309.17147 (2023)

  3. [3]

    Charith Chandra Sai Balne, Sreyoshi Bhaduri, Tamoghna R oy, Vinija Jain, and Aman Chadha. 2024. Parameter Efficient Fine Tuning: A Compreh ensive Anal- ysis Across Applications. arXiv preprint arXiv:2404.13506 (2024)

  4. [4]

    Sreyoshi Bhaduri, Satya Kapoor, Alex Gil, Anshul Mittal , and Rutu Mulkar. 2024. Reconciling methodological paradigms: Employing large la nguage models as novice qualitative research assistants in talent manageme nt research. arXiv preprint arXiv:2408.11043 (2024)

  5. [5]

    Sreyoshi Bhaduri, Kenneth Ohnemus, Jess Blackburn, Ans hul Mittal, Yan Dong, Savannah Laferriere, Robert Pulvermacher, Marina Dias, Al ex Gil, Shahriar Sadighi, et al. 2024. (Multi-disciplinary) Teamwork makesthe (real) dream work: Pragmatic recommendations from industry for engineering classrooms. (2024)

  6. [6]

    Jacob Cohen. 1960. A coefficient of agreement for nominal s cales. Educational and psychological measurement 20, 1 (1960), 37–46

  7. [7]

    Soumi Dutta, Asit Kumar Das, Saptarshi Ghosh, and Debabr ata Samanta. 2022. Data Analytics for Social Microblogging Platforms . Academic Press

  8. [8]

    Andrew F Hayes and Klaus Krippendorff. 2007. Answering th e call for a stan- dard reliability measure for coding data. Communication methods and measures 1, 1 (2007), 77–89

Show all 29 references
  1. [9]

    Satya Kapoor, Alex Gil, Sreyoshi Bhaduri, Anshul Mittal , and Rutu Mulkar

  2. [10]

    Klaus Krippendorff. 2011. Computing Krippendorff’s alp ha-reliability

  3. [11]

    Klaus Krippendorff. 2018. Content analysis: An introduction to its methodology . Sage publications

  4. [12]

    Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen X u, and Chenguang Zhu. 2023. G-eval: Nlg evaluation using gpt-4 with better hu man alignment. arXiv preprint arXiv:2303.16634 (2023)

  5. [13]

    Do Xuan Long, Kenji Kawaguchi, Min-Yen Kan, and Nancy F C hen. 2023. Align- ing Large Language Models with Human Opinions through Perso na Selection and Value–Belief–Norm Reasoning. , arXiv–2311 pages

  6. [14]

    Tammy Mackenzie, Leslie Salgado, Sreyoshi Bhaduri, Vi ctoria Kuketz, Solenne Savoia, and Lilianny Virguez. 2024. Beyond the Algorithm: E mpowering AI Practitioners through Liberal Education. In 2024 ASEE Annual Conference & Ex- position

  7. [15]

    Mary L McHugh. 2012. Interrater reliability: the kappa statistic. Biochemia medica 22, 3 (2012), 276–282

  8. [16]

    Muhammad Naeem, Wilson Ozuem, Kerry Howell, and Silvia Ranfagni. 2023. A step-by-step process of thematic analysis to develop a con ceptual model in Potential and Perils of Large Language Models as Judges of Unstru ctured Textual Data qualitative research. International Journa...

  9. [17]

    Varun Nagaraj Rao, Eesha Agarwal, Samantha Dalal, Dan C alacci, and An- drés Monroy-Hernández. 2024. QuaLLM: An LLM-based Framewo rk to Ex- tract Quantitative Insights from Online Forums. arXiv preprint arXiv:2405.05345 (2024)

  10. [18]

    Jessie Rouder, Olivia Saucier, Rachel Kinder, and Matt Jans. 2021. What to do with all those open-ended responses? Data visualization te chniques for survey researchers. Survey Practice (2021)

  11. [19]

    Keita Saito, Akifumi Wachi, Koki Wataoka, and Youhei Ak imoto. 2023. Ver- bosity bias in preference labeling by large language models . arXiv preprint arXiv:2310.10076 (2023)

  12. [20]

    Patrick Schober, Christa Boer, and Lothar A Schwarte. 2 018. Correlation coeffi- cients: appropriate use and interpretation. Anesthesia & analgesia 126, 5 (2018), 1763–1768

  13. [21]

    David Williamson Shaffer. 2017. Quantitative ethnography. Lulu. com

  14. [22]

    Tianhao Shen, Renren Jin, Yufei Huang, Chuang Liu, Weil ong Dong, Zishan Guo, Xinwei Wu, Yan Liu, and Deyi Xiong. 2023. Large language mode l alignment: A survey. arXiv preprint arXiv:2309.15025 (2023)

  15. [23]

    Liyan Tang, Zhaoyi Sun, Betina Idnay, Jordan G Nestor, A li Soroush, Pierre A Elias, Ziyang Xu, Ying Ding, Greg Durrett, Justin F Rousseau , et al. 2023. Eval- uating large language models on medical evidence summariza tion. NPJ digital medicine 6, 1 (2023), 158

  16. [24]

    Zhiqiang Wang, Yiran Pang, Yanbin Lin, and Xingquan Zhu . 2024. Adapt- able and Reliable Text Classification using Large Language M odels. arXiv:2405.10523 [cs.CL] https://arxiv.org/abs/2405.1 0523

  17. [25]

    Matthijs J Warrens. 2015. Five ways to look at Cohen’s ka ppa. Journal of Psy- chology & Psychotherapy 5 (2015)

  18. [26]

    Ziyou Yan. 2024. Evaluating the Effectiveness of LLM-Ev aluators (aka LLM-as- Judge). eugeneyan. com

  19. [27]

    Zihuai Zhao, Wenqi Fan, Jiatong Li, Yunqing Liu, Xiaowe i Mei, Yiqi Wang, Zhen Wen, Fei Wang, Xiangyu Zhao, Jiliang Tang, et al. 2023. Recom mender sys- tems in the era of large language models (llms). arXiv preprint arXiv:2307.02046 (2023)

  20. [28]

    Lot of opportunities to be recognized, which is something I appreciate. Being recognized is important to advance- ment I think

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhua ng, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, etal. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Informa- tion Processing Systems 36 (2023), 46595–46623. Re...

  21. [2024]

    arXiv preprint arXiv:2409.15626 (2024)

    Qualitative Insights Tool (QualIT): LLM Enhanced Top ic Modeling. arXiv preprint arXiv:2409.15626 (2024)

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.