Pith. sign in

REVIEW 4 cited by

Humans or LLMs as the Judge? A Study on Judgement Biases

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.10669 v5 pith:73IB5S7G submitted 2024-02-16 cs.CL

classification cs.CL
keywords biasjudgesbiaseshumanllmsconductevaluationhuman-
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adopting human and large language models (LLM) as judges (a.k.a human- and LLM-as-a-judge) for evaluating the performance of LLMs has recently gained attention. Nonetheless, this approach concurrently introduces potential biases from human and LLMs, questioning the reliability of the evaluation results. In this paper, we propose a novel framework that is free from referencing groundtruth annotations for investigating Misinformation Oversight Bias, Gender Bias, Authority Bias and Beauty Bias on LLM and human judges. We curate a dataset referring to the revised Bloom's Taxonomy and conduct thousands of evaluations. Results show that human and LLM judges are vulnerable to perturbations to various degrees, and that even the cutting-edge judges possess considerable biases. We further exploit these biases to conduct attacks on LLM judges. We hope that our work can notify the community of the bias and vulnerability of human- and LLM-as-a-judge, as well as the urgency of developing robust evaluation systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 9 citations worldwide. Full citation record

  1. How Benchmark Prediction from Fewer Data Misses the Mark

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Benchmark prediction methods mostly work by interpolation among similar models and fail on better, unfamiliar models, where random sampling with an AIPW-style correction is the only consistent improvement.

  2. Can You Trick the Grader? Adversarial Persuasion of LLM Judges

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Strategically inserted persuasive sentences inflate LLM judges' scores for incorrect math solutions across six benchmarks and fourteen models, but the study lacks length-matched controls separating rhetoric from lengt...

  3. Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support

    cs.CY 2025-08 unverdicted novelty 5.0 of 10

    A multi-agent reliability check and an LLM-as-a-judge quality check both reduce hallucinations in AI-generated study scaffolds, with the multi-agent check matching human expert judgments almost perfectly.

  4. Beyond the Surface: Measuring Self-Preference in LLM Judgments

    cs.CL 2025-06 conditional novelty 4.0 of 10

    The DBG metric measures LLM self-preference bias as the gap between a judge model's own win rate and the win rate assigned by an ensemble of gold judges.

Pith tools