Pith. sign in

Paper Citation Record · LEDGER

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

As of 22 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 17 inbound Pith citation observations for arXiv:2508.06709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.06709 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:40:42.268550Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:22:56.174076Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 36620798-45f4-4c0b-a057-a2f1225267ed · outbound

This paper cites GPT-4 Technical Report.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.167324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.167324Z digest=sha256:0fbc7066e98f5a5cc54b1510e24e244fc7f9b501b0c6c2a0648b127a1ba17889

Observation 6b31b1aa-453c-44d3-a790-dd4b345c54b1 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.175210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.175210Z digest=sha256:9b1ea8ac06a939cdb313683c34c244da799b769f50ba1ccdc0d373c7a344da97

Observation 31a342e9-9936-4dfc-802a-e07af9026c36 · outbound

This paper cites Here they are: [list of 7 restaurants] Low There are only 7 restaurants with a yelp rating of 5 in North Platsville.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Here they are: [list of 7 restaurants] Low There are only 7 restaurants with a yelp rating of 5 in North Platsville

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:40:42.667071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T22:40:42.261241Z digest=sha256:2f87d21a7e096b1e4c46d0308a1dc866ae1db1c30e3b21df42a15bf1e11f0d0d

Observation cd81f2d6-217e-43ff-bd4c-5517b1807056 · outbound

This paper cites maars: Tidy Inference under the 'Models as Approximations' Framework in R.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge maars: Tidy Inference under the 'Models as Approximations' Framework in R

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T22:40:42.527450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T22:40:42.189024Z digest=sha256:c437c25171b73508a08c88270bf9906167e2505ed758b79752fd915f4d9fd418

Observation 88e6b996-898f-40ab-96ea-fe4d1d2993c3 · outbound

This paper cites The Llama 3 Herd of Models.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.195962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.195962Z digest=sha256:e75fdc73190b505d86010bb73489eb1bac3fcc693a620a37309aaf6dafd4ff3d

Observation a1a7d895-0bea-4b43-9374-291a03674111 · outbound

This paper cites Mistral 7B.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Mistral 7B

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.202754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.202754Z digest=sha256:9b118d33a2059398e104df27eb7825ca974d2a64ffbe8ef6425a458eba63d30a

Observation 44abe90e-85c8-43a4-9e08-15a1b0802d92 · outbound

This paper cites Benchmarking Cognitive Biases in Large Language Models as Evaluators.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.206258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.206258Z digest=sha256:f3b07deed57e6cf1f6c0d540e0c7b2133463a523cbca1d8091ebb49a7992a560

Observation 6ca0770c-082e-409e-b593-a7a18a3c892e · outbound

This paper cites Model-free Study of Ordinary Least Squares Linear Regression.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Model-free Study of Ordinary Least Squares Linear Regression

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.209390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.209390Z digest=sha256:3538d6ea377d4b0f11d928cf60d78ea0dbac11daa52028d022fad7abac3b52bc

Observation e1b4a454-ab3b-4ee8-b80d-7bc864ed07a2 · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge LLM Evaluators Recognize and Favor Their Own Generations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.222615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.222615Z digest=sha256:f46d696466b0f6de6b9f0721ceef1a1c127696a4c4ddefc0ca2a54342726397a

Observation 0329d65d-1fa1-40d7-8b6f-ac82cad12d74 · outbound

This paper cites Red teaming language models with language models.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Red teaming language models with language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:40:42.718877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T22:40:42.228653Z digest=sha256:7dd3b4db4a2a71bf7b0d967540cd3771d081a074fb35654b95b7835e14e8ff7a

Observation 5eb9e439-0269-4ae7-a484-92e8eaec9ae3 · outbound

This paper cites Verbosity Bias in Preference Labeling by Large Language Models.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Verbosity Bias in Preference Labeling by Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.231938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.231938Z digest=sha256:87d55064f43f35b09de2b1cc5e01bfe7605d8c466016453914d05f470fdb4283

Observation ccab4162-85cf-455c-a242-9766a689756c · outbound

This paper cites Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.238510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.238510Z digest=sha256:8d216a0bef8323dd3ecc0acde59b7aba74b87ad92090d06c797dd0414f1cffd4

Observation 23a61c37-6754-4582-a38a-e0cfe8c79dc5 · outbound

This paper cites Self-Preference Bias in LLM-as-a-Judge.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Self-Preference Bias in LLM-as-a-Judge

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.241622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.241622Z digest=sha256:14c501f2e551ab1f9e298eabeb0b4bff29608f5ccd579ef289c66e3ad3b1ee35

Observation ded1c9c0-b922-4ee7-9e19-60ac1bc9e2e8 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.244943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.244943Z digest=sha256:5d7d916e92b534d338d7f0d8e6ed878c58533a7ebbba0de43ea19e5a4c58e1c3

Observation 2f1438ae-0efa-44de-8946-1d2b2771abce · outbound

This paper cites Flask: Fine-grained language model evaluation based on alignment skill sets.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Flask: Fine-grained language model evaluation based on alignment skill sets

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:40:42.709377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T22:40:42.248098Z digest=sha256:2391cdc01c5a4ef671f389438a685fbcee120c6d74150e5264d22ebc71e4dced

Observation 7188a20f-7ba9-4b83-a725-03a6511ec3b7 · outbound

This paper cites Xinshu Zhao, Jun S Liu, and Ke Deng.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Xinshu Zhao, Jun S Liu, and Ke Deng

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:40:42.699227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T22:40:42.251116Z digest=sha256:c7fa54ac64652cf61e9242dc46e6d99f5b0b21d00bf90d86b64ef9d8b67336b5

Observation 10b75555-d439-4d6a-9fe1-0ac2f57c1d34 · outbound

This paper cites Judgelm: Fine-tuned large language models are scalable judges.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Judgelm: Fine-tuned large language models are scalable judges

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:40:42.688089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T22:40:42.254049Z digest=sha256:e9d1d141766f469333eb0ce68bdaac453c3a2deb433f66c4ed28de08445203e3

Observation 5a724327-e137-4881-b5dc-d7b1a3c182c3 · outbound

This paper cites The process is analogous for the cofficient corresponding to family-bias.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge The process is analogous for the cofficient corresponding to family-bias

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:40:42.677417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T22:40:42.257200Z digest=sha256:cdb55d8e2d4300e310eefddaeaa9bcee0b184b7041706f32daac96361ed802fb

Observation a70def62-4351-422d-bc71-53061a13babb · outbound

This paper cites I can’t answer that.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge I can’t answer that

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:40:42.644136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T22:40:42.268550Z digest=sha256:25ff10d7c0399f95f223351c6039e0406eef29bbcf97cf17d6726955d067ba65

Observation 84428477-78ea-4224-8f46-26838fc00a99 · outbound

This paper cites Is a kilo of feathers heavier than a pound of steel? High One pound is equal to about 0.45 kilograms.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Is a kilo of feathers heavier than a pound of steel? High One pound is equal to about 0.45 kilograms

Reference 1954

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T22:40:42.655650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T22:40:42.264549Z digest=sha256:4451856d44881b3154dc327a7e86ee374eaf79b5993cd4cec544eb5af1d7c8ba

Observation 64a461cf-98cf-443b-b41e-7edf4afc2c4d · outbound

This paper cites Large Language Models are Inconsistent and Biased Evaluators.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Large Language Models are Inconsistent and Biased Evaluators

Reference 1961

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.235193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.235193Z digest=sha256:4db5dbb82ac38cd351f2918b8aedf296f70cec8188d16c54eebdd5bbc0fb83f1

Observation da65f48b-4185-416e-9a57-3c7624b8339d · outbound

This paper cites Offsetbias: Leveraging debiased data for tuning evaluators.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Offsetbias: Leveraging debiased data for tuning evaluators

Reference 2002

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:40:42.728370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T22:40:42.225788Z digest=sha256:2770fbc907ceb8b3cc4e8e7309bfb44bfcc6f6d18f1c7e1c79a0841c63a493ba

Observation 577d1d4b-13f4-41ea-839a-c6fcf9d46fab · outbound

This paper cites LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.215645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.215645Z digest=sha256:e29320630aee2dd4867d1a1dc5ab4ca945211105201cef07cf1af800ae468ebe

Observation 4fe04d12-60a9-4794-9cab-1de6c3c9ada0 · outbound

This paper cites Human-like Summarization Evaluation with ChatGPT.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Human-like Summarization Evaluation with ChatGPT

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.192331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.192331Z digest=sha256:298c11a5fab67678bd962b8bf3548a45582b77286c9050f65b53098d44b8c3dd

Observation 22c9e1f1-eb39-4906-9929-36e5c2ac3a30 · outbound

This paper cites Humans or llms as the judge? a study on judgement bias.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Humans or llms as the judge? a study on judgement bias

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:40:42.767499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T22:40:42.179101Z digest=sha256:70d73095cb029a5d04b8b2999af0194ab991c832e43efe11f0d997140a1fb0d5

Observation d4195fdc-f149-4b09-89d7-31c553cb2f5e · outbound

This paper cites Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:40:42.738042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T22:40:42.219077Z digest=sha256:1ca771c79e684df89a30cb3b978bbb08bc67cbf7ac8f78251a844570a149a108

Observation 2aeeb816-893c-4fac-8cc4-e76b61eb1cc5 · outbound

This paper cites GPT-4o System Card.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge GPT-4o System Card

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.199401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.199401Z digest=sha256:fcf15b8f9b9520f0eba16d34d0aaeebafac2a9131284469e346e841ef700cae8

Observation b2fae387-ae36-4008-a55c-739559387891 · outbound

This paper cites From generation to judgment: Opportunities and challenges of llm-as-a-judge.arXiv preprint arXiv:2411.16594 ,.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge From generation to judgment: Opportunities and challenges of llm-as-a-judge.arXiv preprint arXiv:2411.16594 ,

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.212669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.212669Z digest=sha256:b8217f4a219aae6d3f6aecdbfad974e655b0f81f948784568d6950df00ece3d8

Observation 00710c5d-cf5b-463a-9f96-5f0870253672 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:40:42.748605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T22:40:42.185900Z digest=sha256:d4419d74211d1e019164d1cfea097a55beaf276946b5de749ef193ee9189f2dc

Observation 6c1254a0-3741-456b-9fc6-88a9b72fea00 · outbound

This paper cites On the limitations of reference-free evaluations of generated text.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge On the limitations of reference-free evaluations of generated text

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:40:42.758062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T22:40:42.182557Z digest=sha256:7a74c17e17cfb2eab4bb398b74adc65ee496fa38a0369dec520cc114abe899a0

Observation 7463a7b1-ae58-4734-b22f-a94445ebee8d · outbound

This paper cites Reference-guided verdict: Llms-as-judges in automatic evaluation of free-form text.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Reference-guided verdict: Llms-as-judges in automatic evaluation of free-form text

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.171599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.171599Z digest=sha256:817e19205cef67f19159ea4dadaece08c5aed073fbe7d7e1ea32d48449e4afe4

Pith citing papers

Observation b56aec3f-1b02-45fa-8577-3816428ecdc9 · inbound

RedNote-Vibe: A Dataset for Capturing Temporal Dynamics of AI-Generated Text in Lifestyle Social Media cites this paper.

RedNote-Vibe: A Dataset for Capturing Temporal Dynamics of AI-Generated Text in Lifestyle Social Media Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:51:26.124111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T13:46:48.348859Z digest=sha256:c34eaced9ba757fca61c87b95da39193570363709687b8b58c6f2534c871ed04

Observation 1af5245e-b6bd-4899-9cf6-12a53103a301 · inbound

Extreme Self-Preference in Language Models cites this paper.

Extreme Self-Preference in Language Models Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:00:39.082645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T20:57:37.199128Z digest=sha256:307a6140a31759185fb4b11c38ba206d950be33193d5be1fe911414346a2abd6

Observation 6022807f-8a57-4574-bd20-c2abb734a229 · inbound

When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning cites this paper.

When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:51:08.533941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T08:50:53.486588Z digest=sha256:5b39c85ad86bedde60a667434f0ac5fe75c10733b6309708b980f89f14d7082b

Observation 40c16883-2fe6-48d7-a8f9-8420bfe1dbba · inbound

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations cites this paper.

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T06:35:37.070203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:35:37.070203Z digest=sha256:233207c06cd55c7d9be9a4495e21e6d23904c121efc802b56d379752e25780d3

Observation 6bcbd550-bf76-455c-8374-2092d0880441 · inbound

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety cites this paper.

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:28:48.785362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T01:27:16.967080Z digest=sha256:6e29f7518555828ab305a3a843bbeeab7df24c041c96d36df9010aa82f963e66

Observation 4d2d3d9e-46ad-4bc4-8ec8-5f727af6279c · inbound

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge cites this paper.

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:40:51.666192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T19:38:11.595077Z digest=sha256:7d205a27d13e2255138fa5aed4cd91f57036a0e357013284120e5d5ad2e213fe

Observation 2cfb26b7-9e07-417b-906a-14e481c77c59 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:50:50.434134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:18:19.955943Z digest=sha256:0d226b181f873cf76763e14f18f5d81ef85b0f8cf9eacbefc9e23a77852d280c

Observation 44be8642-c2da-490b-b916-08fbd24dc89a · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T16:40:55.448677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:40:55.448677Z digest=sha256:8fdc0f4bdf8d0ff0c670ecaf2638f8c01d20fd3fda40f1d07522e191c4f4cd1d

Observation e0c2e89e-0807-40cc-aad1-04f92cd4a65d · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T05:36:43.365289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:36:43.365289Z digest=sha256:a9485079607b16ecbb08fd326a13b82c86603383d22bd8c1949f13f614e64ace

Observation 3e25a0e9-63c4-4f43-a7de-2c632048062e · inbound

Why Do Safety Guardrails Degrade Across Languages? cites this paper.

Why Do Safety Guardrails Degrade Across Languages? Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:18:21.034567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T14:17:41.347193Z digest=sha256:b40ed8cc2b633d16d7d83f92396f41c732b6daa38644e26f18c6991c6eca117c

Observation 328827ab-7ee7-488b-8ec1-322163c9e6d9 · inbound

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks cites this paper.

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.462423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T10:46:24.554332Z digest=sha256:d16b242ad600675451c625cc25461808fcc5dbb8edbc5c5e6ab2876ad64e978b

Observation daf7bb2d-8931-4544-87ec-51046d0acc16 · inbound

Scaffold, Not Vocabulary? A Controlled, Two-Tier, Pre-Registered Study of a Popperian Code-Generation Skill cites this paper.

Scaffold, Not Vocabulary? A Controlled, Two-Tier, Pre-Registered Study of a Popperian Code-Generation Skill Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:47:04.549846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T00:08:06.720586Z digest=sha256:0644d9e3df716de5b77650244f87ceb9f41664499dfd91c8ccdde257cc07cbf2

Observation 2bfa219f-0831-4e2d-8252-5e3a3f8ba9d5 · inbound

Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG cites this paper.

Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T10:22:55.558755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:22:55.558755Z digest=sha256:db40640629232168a2b51c540914988e4e979ccd1ae2bfb440bc565b77d04b9d

Observation fa650952-5d5a-4684-b682-9a9d4e91de11 · inbound

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding cites this paper.

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T11:54:20.710057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T11:54:20.710057Z digest=sha256:2830b4c399f130b0c405ad131820cc114f10fad4ccc9752616ed306dee1afbfa

Observation c40e4a63-4ed1-40a0-99a9-2c20f9ab037e · inbound

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation cites this paper.

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T04:18:54.154755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:18:54.154755Z digest=sha256:cb7bbf8781f39266cf0d8fb55e28655e6a5309ef393f5c70140c50b471a85591

Observation 02488741-6b29-4e9b-bf88-00a539169708 · inbound

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation cites this paper.

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T04:16:26.027632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:16:26.027632Z digest=sha256:aaec4fd0517ab12ba2e75f4670be622d8a49ebb7cff4f819714b6ce0a9d50f65

Observation e166e748-13c0-4c34-b676-936972121e76 · inbound

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation cites this paper.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:22:56.174076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:22:56.174076Z digest=sha256:fe5e7a54fe4260d07e063d0ac805efdcd7468fdb8fb8c123823c7754161a683f