Pith. sign in

Paper Citation Record · LEDGER

Humans or LLMs as the Judge? A Study on Judgement Biases

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2402.10669.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.10669 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:01.891765Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4027fdd0-17f8-4fbb-868a-7ff070657a16 · inbound

Lessons from the Trenches on Reproducible Evaluation of Language Models cites this paper.

Lessons from the Trenches on Reproducible Evaluation of Language Models Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 255

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:44:49.863860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T18:44:49.519995Z digest=sha256:dcabcd6e0fff40a0c6bb74c06043961dbd451ad96d1f832e9b0b6b60469f0a36

Observation 7dd4cb82-3d90-4cf0-9336-09fe5bc730bc · inbound

ShieldGemma: Generative AI Content Moderation Based on Gemma cites this paper.

ShieldGemma: Generative AI Content Moderation Based on Gemma Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:17:39.478203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T13:17:39.444002Z digest=sha256:5d85b33f66272f35aab4e33059a909478cb252efe4d7e072dc74c64af3ed3850

Observation 75fc38b4-ebb7-48a1-9090-e55fc7015d55 · inbound

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap cites this paper.

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:08:20.821090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T19:07:21.016824Z digest=sha256:17a9cbd026bc3c91c9201aa2167332ed7d7a44a75de6aa2899f69adee1005648

Observation d797e876-a248-4128-84ec-4e6e856441d3 · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:44.180203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:c803c86c0dcd378e1cc56071133cfd1015308a37fd5147cd0b3fc7c9f53d8a1a

Observation 7099f953-cf2c-41a1-b988-756f77c873d8 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.091045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:f47bd0a86abcdf2d72a631ee361e539f5406f9210d1eacc2db8533da01542540

Observation 2a704e88-0a57-490d-aa55-a99138c00bb1 · inbound

Beyond the Surface: Measuring Self-Preference in LLM Judgments cites this paper.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.891765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.891765Z digest=sha256:5af3d334d9c4a4e0983eddc1cddb01eabc8d62a76cfd2074bf7fff47b7b2944a

Observation 0696c877-6cb1-450c-898d-563b3cd76d5b · inbound

How Benchmark Prediction from Fewer Data Misses the Mark cites this paper.

How Benchmark Prediction from Fewer Data Misses the Mark Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.647558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.647558Z digest=sha256:8a3529fb0b7bae3feb2a06a5e1f357b9ea2e8a126f7bd072ba70ce00b667ff54

Observation 3212f0dd-80e7-48af-93e1-3ae1de774f2c · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T03:22:01.345621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:3e10f723affaa4d280b3d0c4f110de8727db2092c7ee17e2563161ae322d292e

Observation adde8bc6-e443-4c2b-9b0a-7bb1bc90f9fa · inbound

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support cites this paper.

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:04:56.860420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:04:56.860420Z digest=sha256:d22d00c5df45d94c2f67a9bcd9cf06730718e72962714f709b21b9eac2f91052

Observation 1133009d-cbdc-4171-9003-eefcab3affe3 · inbound

Can You Trick the Grader? Adversarial Persuasion of LLM Judges cites this paper.

Can You Trick the Grader? Adversarial Persuasion of LLM Judges Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:56.852022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:55:56.852022Z digest=sha256:744352d65cffbcbd160fac9e7fd12e98ea729988f44f148198f5a70969002655

Observation 1e3af1f4-8c56-4cf6-8c82-b198eb344214 · inbound

Can LLMs Make (Personalized) Access Control Decisions? cites this paper.

Can LLMs Make (Personalized) Access Control Decisions? Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:34:05.015121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T05:33:18.457421Z digest=sha256:2baf0bf91cba8930052fb878c6cdc5441ed65fedce62b47b78d2e6159e572e46

Observation b0c0797a-c760-4f0d-a1b4-dd183d554350 · inbound

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations cites this paper.

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T03:47:15.241556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T03:43:18.987241Z digest=sha256:d25e5bf62c90992dff8123e46066da515189af083ae2d455b87eed0fa0bdb756

Observation bdc06a1f-3152-4c53-b325-23af5498a4cf · inbound

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis cites this paper.

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:53.360123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T09:06:33.531027Z digest=sha256:bc2be0c0dd40d9c8a3fda76c07d80e33103c81480f343d1f1c70171845a5891b

Observation e11eef20-585c-4cd1-9f30-cc82530e31ea · inbound

Pioneer Agent: Continual Improvement of Small Language Models in Production cites this paper.

Pioneer Agent: Continual Improvement of Small Language Models in Production Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:05:57.610523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T17:48:40.520740Z digest=sha256:fc9bc753fcbbc4e93d053e83e2a2c7ed6c14f4bdb44da7f5ea352bdaf01c65ac

Observation 55cad9a5-2ce2-4824-b350-a07331f019f6 · inbound

Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring cites this paper.

Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:51:03.617526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T04:31:53.825854Z digest=sha256:0bdc990505109435d9139bd3e6059f25a1556b58034cd0b1c47cd89cf4eb6478

Observation fa516d12-dbb4-4f71-b29a-ef868771e54e · inbound

Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards cites this paper.

Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:07.359059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T21:58:05.584559Z digest=sha256:2080afcb83a2068b30d212ac5ce480b52178bd4942328c97cc704227618f5ad7

Observation d52acf4d-da07-433f-a3f5-212bb7de8ab5 · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:14.136347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:bfd7566c6832602e6e486c7e950c89e614e32c68f598a29ea05e3fd6df233afe

Observation 2d1153fc-2934-442f-9d40-a63336ac8edf · inbound

TRUST: A Framework for Decentralized AI Service v.0.1 cites this paper.

TRUST: A Framework for Decentralized AI Service v.0.1 Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:01:28.726224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-07T08:22:14.239443Z digest=sha256:083700ad2c43668801d0a2a1574435fe4ec63589bc65374c97563e053dc6d9e8

Observation 4fca8ff2-f0e2-4b94-a225-46b6c1dc4946 · inbound

RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents cites this paper.

RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:29:09.209885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:27:16.974169Z digest=sha256:b08b02f4b53b4018884dd5a64f5c1380a5e5b3f00a0b8472659888e31917d4b1

Observation a9aa0af7-0d17-4bf2-942b-87235affd6c8 · inbound

Are LLMs Bad at Moral Reasoning? cites this paper.

Are LLMs Bad at Moral Reasoning? Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:12.640230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T08:20:24.251540Z digest=sha256:8e0b165cb343aa3ad09e92811661b4ad27280c9393f93d485586147d1a49790f

Observation 302284f3-6fdf-4743-b37b-05ebf3e098ae · inbound

Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering cites this paper.

Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 236

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:55:06.221097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T22:54:03.054871Z digest=sha256:4c570215fd9cd20a33367e5a36ee21b927e1ae2536968500e7bbc18529760472

Observation 182049fe-ebd1-4305-9170-b82122fce22a · inbound

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning cites this paper.

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:20:00.097000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-25T23:49:38.932474Z digest=sha256:6b0083f47ae963833c540a156b50dc118b216ec04a854ac7baee472f0f34b8d1