Pith. sign in

Paper Citation Record · LEDGER

Humans or LLMs as the Judge? A Study on Judgement Biases

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2402.10669.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.10669 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:01.891765Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4027fdd0-17f8-4fbb-868a-7ff070657a16 · inbound

Lessons from the Trenches on Reproducible Evaluation of Language Models cites this paper.

Lessons from the Trenches on Reproducible Evaluation of Language Models Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 255

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:44:49.863860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T18:44:49.519995Z digest=sha256:59de57eb20a9080f70ad9e76245dbce80683b43e74aa27949b72c6070c232c3c

Observation 7dd4cb82-3d90-4cf0-9336-09fe5bc730bc · inbound

ShieldGemma: Generative AI Content Moderation Based on Gemma cites this paper.

ShieldGemma: Generative AI Content Moderation Based on Gemma Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:17:39.478203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T13:17:39.444002Z digest=sha256:39f4e11b44504b6ca96d15375032e289a859dc5840c443809cef5185948e17f1

Observation 75fc38b4-ebb7-48a1-9090-e55fc7015d55 · inbound

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap cites this paper.

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:08:20.821090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T19:07:21.016824Z digest=sha256:c7dec8f2131d0836a2c56284ec71dcfaa78f981e847eca0943ff18c0f0a312c5

Observation d797e876-a248-4128-84ec-4e6e856441d3 · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:44.180203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:dc8774666997874a800283b8b09beeee1a670c4548e5afec6315a388b427ba88

Observation 7099f953-cf2c-41a1-b988-756f77c873d8 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.091045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:b38bd718034ebf338a2f476d6bce4c6ce90b42cbed24b9984c4371af1e2209b5

Observation 2a704e88-0a57-490d-aa55-a99138c00bb1 · inbound

Beyond the Surface: Measuring Self-Preference in LLM Judgments cites this paper.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.891765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.891765Z digest=sha256:5af3d334d9c4a4e0983eddc1cddb01eabc8d62a76cfd2074bf7fff47b7b2944a

Observation 0696c877-6cb1-450c-898d-563b3cd76d5b · inbound

How Benchmark Prediction from Fewer Data Misses the Mark cites this paper.

How Benchmark Prediction from Fewer Data Misses the Mark Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.647558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.647558Z digest=sha256:c2788f1ec0672db36a2b9152b75032805d594544f7f8d084a97bb4089d383a1c

Observation 3212f0dd-80e7-48af-93e1-3ae1de774f2c · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T03:22:01.345621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:3b9067f3ca09193cd8276cec6b9faa78994e156f2d55b6d02216145037d04ee8

Observation adde8bc6-e443-4c2b-9b0a-7bb1bc90f9fa · inbound

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support cites this paper.

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:04:56.860420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:04:56.860420Z digest=sha256:b5b520edffd4c0637597325352c6a5f9c2d4f7bc07da5aa46898bbf36f845f11

Observation 1133009d-cbdc-4171-9003-eefcab3affe3 · inbound

Can You Trick the Grader? Adversarial Persuasion of LLM Judges cites this paper.

Can You Trick the Grader? Adversarial Persuasion of LLM Judges Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:56.852022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:55:56.852022Z digest=sha256:744352d65cffbcbd160fac9e7fd12e98ea729988f44f148198f5a70969002655

Observation 1e3af1f4-8c56-4cf6-8c82-b198eb344214 · inbound

Can LLMs Make (Personalized) Access Control Decisions? cites this paper.

Can LLMs Make (Personalized) Access Control Decisions? Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:34:05.015121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T05:33:18.457421Z digest=sha256:3edc56b98969e4d4350cf1e2a94921815d0af7f95065d5fdf520af8fd1b66da6

Observation b0c0797a-c760-4f0d-a1b4-dd183d554350 · inbound

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations cites this paper.

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T03:47:15.241556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T03:43:18.987241Z digest=sha256:55abed1e8fe2313c4f1b1dfdce4700fe06d375b49cafb9624ad070ca5a87b5ef

Observation bdc06a1f-3152-4c53-b325-23af5498a4cf · inbound

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis cites this paper.

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:53.360123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T09:06:33.531027Z digest=sha256:40b7e2a4bb4a9ef02f1b422966286c03c2f09dff2171dba0b5a0568ef227429a

Observation e11eef20-585c-4cd1-9f30-cc82530e31ea · inbound

Pioneer Agent: Continual Improvement of Small Language Models in Production cites this paper.

Pioneer Agent: Continual Improvement of Small Language Models in Production Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:05:57.610523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T17:48:40.520740Z digest=sha256:25dbe74fa263a0b8200e8b7b4328b2937be51dea0b5f701fc5501b2541b0748d

Observation 55cad9a5-2ce2-4824-b350-a07331f019f6 · inbound

Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring cites this paper.

Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:51:03.617526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T04:31:53.825854Z digest=sha256:37193f9072eb9ea6d1cddb99ee28245774597331907b9f436112cac158880c3d

Observation fa516d12-dbb4-4f71-b29a-ef868771e54e · inbound

Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards cites this paper.

Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:07.359059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T21:58:05.584559Z digest=sha256:d4ee7163fa276b4f92da0be3dc857550675aba476ae92a441cf985b0550f33dc

Observation d52acf4d-da07-433f-a3f5-212bb7de8ab5 · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:14.136347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:e0ec901d9e91f4ad920fd3bbe3e6257140a031a57f338818bfc7b70ddc09a1cd

Observation 2d1153fc-2934-442f-9d40-a63336ac8edf · inbound

TRUST: A Framework for Decentralized AI Service v.0.1 cites this paper.

TRUST: A Framework for Decentralized AI Service v.0.1 Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:01:28.726224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-07T08:22:14.239443Z digest=sha256:f348975a3d57287e25322a00e00075a09376346e069587d55d2925042cd1c66e

Observation 4fca8ff2-f0e2-4b94-a225-46b6c1dc4946 · inbound

RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents cites this paper.

RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:29:09.209885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T22:27:16.974169Z digest=sha256:8733384b3c117c0ed3b2778e3d768e74db957c62b7ada6f2d41bbecc1222b711

Observation a9aa0af7-0d17-4bf2-942b-87235affd6c8 · inbound

Are LLMs Bad at Moral Reasoning? cites this paper.

Are LLMs Bad at Moral Reasoning? Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:12.640230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T08:20:24.251540Z digest=sha256:d30cf21e90b2fc4f8abce4bbea52667b7619d4912a94b3a1227ed9a2e123a4b0

Observation 302284f3-6fdf-4743-b37b-05ebf3e098ae · inbound

Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering cites this paper.

Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 236

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:55:06.221097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T22:54:03.054871Z digest=sha256:1e326286d4cf085ed1a4aadfe81e90bad93f509d972234f0bc020882785f5acc

Observation 182049fe-ebd1-4305-9170-b82122fce22a · inbound

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning cites this paper.

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:20:00.097000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T23:49:38.932474Z digest=sha256:0f60e7c8d90cf9d024bdeb5b3d9ccd5903503ddbcee578f5fd64917bdacf1869