Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T22:40:42.268550Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 17 inbound Pith citation observations for arXiv:2508.06709.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T22:40:42.268550Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:22:56.174076Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
31 of 31 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 36620798-45f4-4c0b-a057-a2f1225267ed · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b31b1aa-453c-44d3-a790-dd4b345c54b1 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31a342e9-9936-4dfc-802a-e07af9026c36 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Here they are: [list of 7 restaurants] Low There are only 7 restaurants with a yelp rating of 5 in North Platsville
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cd81f2d6-217e-43ff-bd4c-5517b1807056 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge maars: Tidy Inference under the 'Models as Approximations' Framework in R
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 88e6b996-898f-40ab-96ea-fe4d1d2993c3 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1a7d895-0bea-4b43-9374-291a03674111 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Mistral 7B
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44abe90e-85c8-43a4-9e08-15a1b0802d92 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Benchmarking Cognitive Biases in Large Language Models as Evaluators
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ca0770c-082e-409e-b593-a7a18a3c892e · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Model-free Study of Ordinary Least Squares Linear Regression
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1b4a454-ab3b-4ee8-b80d-7bc864ed07a2 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge LLM Evaluators Recognize and Favor Their Own Generations
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0329d65d-1fa1-40d7-8b6f-ac82cad12d74 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Red teaming language models with language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5eb9e439-0269-4ae7-a484-92e8eaec9ae3 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Verbosity Bias in Preference Labeling by Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccab4162-85cf-455c-a242-9766a689756c · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23a61c37-6754-4582-a38a-e0cfe8c79dc5 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Self-Preference Bias in LLM-as-a-Judge
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ded1c9c0-b922-4ee7-9e19-60ac1bc9e2e8 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f1438ae-0efa-44de-8946-1d2b2771abce · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Flask: Fine-grained language model evaluation based on alignment skill sets
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7188a20f-7ba9-4b83-a725-03a6511ec3b7 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Xinshu Zhao, Jun S Liu, and Ke Deng
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 10b75555-d439-4d6a-9fe1-0ac2f57c1d34 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Judgelm: Fine-tuned large language models are scalable judges
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5a724327-e137-4881-b5dc-d7b1a3c182c3 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge The process is analogous for the cofficient corresponding to family-bias
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a70def62-4351-422d-bc71-53061a13babb · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge I can’t answer that
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 84428477-78ea-4224-8f46-26838fc00a99 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Is a kilo of feathers heavier than a pound of steel? High One pound is equal to about 0.45 kilograms
Reference 1954
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 64a461cf-98cf-443b-b41e-7edf4afc2c4d · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Large Language Models are Inconsistent and Biased Evaluators
Reference 1961
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da65f48b-4185-416e-9a57-3c7624b8339d · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Offsetbias: Leveraging debiased data for tuning evaluators
Reference 2002
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 577d1d4b-13f4-41ea-839a-c6fcf9d46fab · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fe04d12-60a9-4794-9cab-1de6c3c9ada0 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Human-like Summarization Evaluation with ChatGPT
Reference 2006
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22c9e1f1-eb39-4906-9929-36e5c2ac3a30 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Humans or llms as the judge? a study on judgement bias
Reference 2015
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d4195fdc-f149-4b09-89d7-31c553cb2f5e · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2aeeb816-893c-4fac-8cc4-e76b61eb1cc5 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge GPT-4o System Card
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2fae387-ae36-4008-a55c-739559387891 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge From generation to judgment: Opportunities and challenges of llm-as-a-judge.arXiv preprint arXiv:2411.16594 ,
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00710c5d-cf5b-463a-9f96-5f0870253672 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Bert: Pre-training of deep bidirectional transformers for language understanding
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6c1254a0-3741-456b-9fc6-88a9b72fea00 · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge On the limitations of reference-free evaluations of generated text
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7463a7b1-ae58-4734-b22f-a94445ebee8d · outbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Reference-guided verdict: Llms-as-judges in automatic evaluation of free-form text
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b56aec3f-1b02-45fa-8577-3816428ecdc9 · inbound
RedNote-Vibe: A Dataset for Capturing Temporal Dynamics of AI-Generated Text in Lifestyle Social Media Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1af5245e-b6bd-4899-9cf6-12a53103a301 · inbound
Extreme Self-Preference in Language Models Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6022807f-8a57-4574-bd20-c2abb734a229 · inbound
When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 40c16883-2fe6-48d7-a8f9-8420bfe1dbba · inbound
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bcbd550-bf76-455c-8374-2092d0880441 · inbound
Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4d2d3d9e-46ad-4bc4-8ec8-5f727af6279c · inbound
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2cfb26b7-9e07-417b-906a-14e481c77c59 · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 44be8642-c2da-490b-b916-08fbd24dc89a · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0c2e89e-0807-40cc-aad1-04f92cd4a65d · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e25a0e9-63c4-4f43-a7de-2c632048062e · inbound
Why Do Safety Guardrails Degrade Across Languages? Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 328827ab-7ee7-488b-8ec1-322163c9e6d9 · inbound
CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation daf7bb2d-8931-4544-87ec-51046d0acc16 · inbound
Scaffold, Not Vocabulary? A Controlled, Two-Tier, Pre-Registered Study of a Popperian Code-Generation Skill Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2bfa219f-0831-4e2d-8252-5e3a3f8ba9d5 · inbound
Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa650952-5d5a-4684-b682-9a9d4e91de11 · inbound
Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c40e4a63-4ed1-40a0-99a9-2c20f9ab037e · inbound
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02488741-6b29-4e9b-bf88-00a539169708 · inbound
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e166e748-13c0-4c34-b676-936972121e76 · inbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.