Pith. sign in

Paper Citation Record · LEDGER

Beyond the Surface: Measuring Self-Preference in LLM Judgments

As of 7 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 3 inbound Pith citation observations for arXiv:2506.02592.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02592 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:07.442164Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T02:16:42.324571Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T21:00:39.084405Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1aab9281-65d6-486d-8ba9-3327a25e6110 · outbound

This paper cites online" 'onlinestring :=.

Beyond the Surface: Measuring Self-Preference in LLM Judgments online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.399959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.399959Z digest=sha256:5c109e2a414bec2d04374486f16e8a5fbeef65979e271a17b0c14230b16166a9

Observation 916ecf34-c6b3-4acd-b948-635d662e7df4 · outbound

This paper cites write newline.

Beyond the Surface: Measuring Self-Preference in LLM Judgments write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.465840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.465840Z digest=sha256:952594c38756e3b04161c3c14ecc49cf06506b16b6083eb13b2b93a2268883b1

Observation 3024f0e2-9035-4cba-ad29-7d3a29d9bcb5 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:09.309397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:01.557460Z digest=sha256:27dd59470b6faa450d32ea256201a2ead13207a8b37687cb404b8642df509ad9

Observation fcc45e8a-b65a-47b3-949a-abdeb19eecd6 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.632516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.632516Z digest=sha256:c2d5a1485ab581b993c084ab780dc994f0ab840060b3f09e78ed4acc9222d39d

Observation e93e67a4-bb96-4958-b89f-6ac226594aba · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.729009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.729009Z digest=sha256:ae32c136f27725cc7db07861033540df043609a8f8d1ad18da7ddcad95e9f44f

Observation a6ef2e73-9a5d-487e-ac02-bcb0f2e3ec14 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.796164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.796164Z digest=sha256:849a5230ccea2f3cc51dfb2d0608fe20a12995be5b47acb1d5b91c69ecfc451e

Observation 2a704e88-0a57-490d-aa55-a99138c00bb1 · outbound

This paper cites Humans or LLMs as the Judge? A Study on Judgement Biases.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.891765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.891765Z digest=sha256:5af3d334d9c4a4e0983eddc1cddb01eabc8d62a76cfd2074bf7fff47b7b2944a

Observation 44e04843-4344-4b95-9591-919fa420439a · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.992851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.992851Z digest=sha256:51d530ab48b817f6fe42ae26c6985d11fc6a8f88151026f79a40a15a15ec8b8c

Observation bfd26f89-e201-463e-94c4-0b7f99aec333 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:09.130400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:02.089341Z digest=sha256:1a97a539a58240eb3fdfe6ad45f82579938ec03cd732ace0f20f82a22b62ea33

Observation df786909-59ae-4eb3-a216-77e507584322 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Beyond the Surface: Measuring Self-Preference in LLM Judgments UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:02.214831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:02.214831Z digest=sha256:ab37245d44d498e0588b29d81ee0ea7020dd35c425c1d8587b7db9c0eeaf9b84

Observation 511d3411-70d5-4f68-85cd-17437e3cfd96 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:02.375895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:02.375895Z digest=sha256:b555dc94d2f93b322c0862440ab75abfeba77baf4f9bc0b1795498e19bfe2afd

Observation 91f45ef9-147b-48e1-9b2f-0a96f1876e88 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:02.502570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:02.502570Z digest=sha256:ce1d58e810aebba170bb6cbdfd24ba51901c0b9374b534558f79f312112abf1c

Observation 9031e36a-c68a-45b5-8448-46253566c8a1 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:08.966253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:02.638673Z digest=sha256:d16ca6d0cc2cbbf26acfe402d9fbb59b1556704810411e8f564083c3ccf1b8b5

Observation 8d0dee2f-f657-4628-8a3d-c5f7d77129dd · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Beyond the Surface: Measuring Self-Preference in LLM Judgments ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:02.787662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:02.787662Z digest=sha256:4c41b8ee71409c63859ca9c89322a5a39348e0cc44904fece25598900f695e92

Observation 6611b524-41ca-4621-94f6-bd82bdd9031d · outbound

This paper cites The Llama 3 Herd of Models.

Beyond the Surface: Measuring Self-Preference in LLM Judgments The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:02.942801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:02.942801Z digest=sha256:7e46cac6f57cc8ec4f4f8f108a1cdb80da6c86ca7afe088dfce7119884d0c7bb

Observation 2849b9fb-d3d2-4036-9aa9-29a6a1d57b8f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Beyond the Surface: Measuring Self-Preference in LLM Judgments DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.056652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.056652Z digest=sha256:0912ae2a728ae6c4020a0542205131455b7bcc96e935a7639fbb2af922ba0dc4

Observation 99aba7be-d3b8-49e9-a7fd-c8ef588971ae · outbound

This paper cites Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.173700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.173700Z digest=sha256:e85b130805adcde70b3db1de60ae5cc5dfc303ebbc6a3a7e078e95ac41d608b9

Observation d8c09241-ecda-4029-b8cc-3e301a904ac1 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Measuring Mathematical Problem Solving With the MATH Dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.351926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.351926Z digest=sha256:638d28c9cecdb007c15ba9464562f92f2deec4422b69714ee05174494e6fbe16

Observation 6fb0e45e-169b-42a5-9970-6a79fd6a0456 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:08.754298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:03.467505Z digest=sha256:3220ef7f5d01f2dda1e039ceb74e59d7c555f3629e43bf97ce8727d48f6924f5

Observation 568c9826-03d8-4c44-94b8-b56c6042850c · outbound

This paper cites GPT-4o System Card.

Beyond the Surface: Measuring Self-Preference in LLM Judgments GPT-4o System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.625941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.625941Z digest=sha256:166dd2d00047466ba646e9aba49a2d279cb5a909873f64fd975f3a46e7fc3544

Observation fa632138-9ef8-49ee-a22b-2b698d4f6bf8 · outbound

This paper cites OpenAI o1 System Card.

Beyond the Surface: Measuring Self-Preference in LLM Judgments OpenAI o1 System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.732146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.732146Z digest=sha256:4e6bbe0ffc203818e6f50be9ff840a007d50b76861f405b7ec2b0ae5fd538686

Observation 2f4d8d5a-931a-4bb6-81bc-60471a111347 · outbound

This paper cites Benchmarking Cognitive Biases in Large Language Models as Evaluators.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.842968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.842968Z digest=sha256:da55cf44052ce3f419e80d8eeb85690269a2f828969e7022fcdde24fd33b848e

Observation e36f4cae-64be-45db-9f0f-1e7b3cfff12a · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Beyond the Surface: Measuring Self-Preference in LLM Judgments RewardBench: Evaluating Reward Models for Language Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.953001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.953001Z digest=sha256:439e06846a26105f490683b998fd8fa29e57bfec82ed4d583fe20763c03a64cd

Observation a2a151bb-abec-4435-91ae-e2868aaa889b · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Beyond the Surface: Measuring Self-Preference in LLM Judgments RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.092617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.092617Z digest=sha256:39761881187b338fc1e761a3f62c81291d7af9a6d59e133239d88d9f2587fb9a

Observation 9193983d-d82f-418e-b27d-2ebefab0ee75 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.199628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.199628Z digest=sha256:697b4ee3e0ed04528e8226f3e85269acd190c47732e412c2f070b8bbe58b8c80

Observation fe2a4650-f233-4da7-a0b6-7a32f078144c · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.279674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.279674Z digest=sha256:d1cf58a68cc343ac0b0ddee69e48de85d9a83d9dec83da278c5d7cf35ab844e2

Observation b64f8235-083e-4ee9-b0c7-f7b01fb26ca2 · outbound

This paper cites Hashimoto.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Hashimoto

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.363339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.363339Z digest=sha256:2f6f039bc3b40110fac7d8280e0f404704e017cee88cb91dcecc5a4795e1a36f

Observation 3469f3d8-c65e-4e37-bb2e-5e0afe358758 · outbound

This paper cites The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning.

Beyond the Surface: Measuring Self-Preference in LLM Judgments The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.503845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.503845Z digest=sha256:78443e601280956ed6620d9c38ede7e96b0cf094c93fee11ec41b1703071295c

Observation 59248575-72ef-42be-9662-9c1db09f8e83 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Beyond the Surface: Measuring Self-Preference in LLM Judgments TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.592298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.592298Z digest=sha256:31e9c21cf99430c48c33292ff63aef8d30a28355ee9aeabeadda51a9b4710298

Observation 42b79dbf-0220-4ead-866c-e13a1984be0e · outbound

This paper cites DeepSeek-V3 Technical Report.

Beyond the Surface: Measuring Self-Preference in LLM Judgments DeepSeek-V3 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.730166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.730166Z digest=sha256:4040a071753680c2f3765d6cc2f294929da4528064314104e5df87bf7cf9cca8

Observation df62893e-d5f8-4916-a8f8-3c71b13c46a2 · outbound

This paper cites AlignBench: Benchmarking Chinese Alignment of Large Language Models.

Beyond the Surface: Measuring Self-Preference in LLM Judgments AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.837491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.837491Z digest=sha256:cc5f28c7e8f50cfb144aba52e953eab3a90c6ec6348c950c3a0d676e5ad531d9

Observation c5a216dd-79c7-4ed1-89f9-9c84f203cf38 · outbound

This paper cites LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores.

Beyond the Surface: Measuring Self-Preference in LLM Judgments LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.946179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.946179Z digest=sha256:e5577ea0951cc64a8dd15bf88b93047f8ccfbc85f64a8bf6b43885fcf09cdb2e

Observation 2652d21b-ea36-4892-a7e1-769f3ea96b18 · outbound

This paper cites Evaluating Style Transfer for Text.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Evaluating Style Transfer for Text

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:26:07.976111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:05.025643Z digest=sha256:d904f89427363cb1ab0a47ece42fcda5f0116ffe989cf79ed8aea89077df5a53

Observation 250c8b4a-aeea-47ce-bf7a-78d8d52a83fd · outbound

This paper cites Text Style Transfer Evaluation Using Large Language Models.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Text Style Transfer Evaluation Using Large Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:26:07.833592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:05.130234Z digest=sha256:c3c5a51612171c0b3e1a00ea568b27536a7c53da33a575360e5fa32f70216e23

Observation 846299db-4fcb-47a2-9645-7711fc940f2d · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:08.558455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:05.237736Z digest=sha256:4ab8409d19f192cccf284cf435713d8056520e00de904123a6ca0a1aff5d8241

Observation be9bacc0-b406-498a-a6ab-cd665dc04a99 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:05.347061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:05.347061Z digest=sha256:b7d7e9109ebfa3e9677361d832ee61abd0644cb207b121e43587651f64fbb7e5

Observation e2bc21e0-b692-4fb5-8259-61f7dd50d2a5 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

Beyond the Surface: Measuring Self-Preference in LLM Judgments ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:05.458126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:05.458126Z digest=sha256:95b720605a8114165782b494a1b05d4256ab323bd800a3b0629227cc7807eed2

Observation 6d918236-3f33-4060-8d64-fac47a29f405 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:05.547451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:05.547451Z digest=sha256:e4ea0faa69893f07f6d0b09089867429ba6ce0543415f459279d6aea0840c66e

Observation 7b650efc-f0e8-4153-9842-6765752cf9c5 · outbound

This paper cites SALMON: Self-Alignment with Instructable Reward Models.

Beyond the Surface: Measuring Self-Preference in LLM Judgments SALMON: Self-Alignment with Instructable Reward Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:05.656399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:05.656399Z digest=sha256:a5107bbe2d7e2a2b743c4768a4507a5b919b4883122c0416e603f62f297a46d4

Observation f7d88fb0-a0cb-4aa9-b9e8-82264ccdc5cf · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:05.795023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:05.795023Z digest=sha256:40a63161442395e85ffda4c0c8fd623e9780ef7421e30e65b6c074378aa6a13e

Observation d7a6a6c4-1e8f-4520-9ee9-fd9191237295 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Gemma 2: Improving Open Language Models at a Practical Size

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:05.913103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:05.913103Z digest=sha256:04db164acf77a8189d4577f0472591f89dc19f6d682f20cee9d542687ab1ee79

Observation 3f66f203-92c3-4d48-bd0f-961683027ee7 · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.015237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.015237Z digest=sha256:45ff203e8884da03a3ab04b4acb15aa435b28b1328776d4fe4bb166f96df4240

Observation e79cbff3-326f-42d4-b14f-0f10c3b4e1bd · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.158196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.158196Z digest=sha256:c95e129b7cc19af56b53864950687af67fc22f02117aefbea8a9fd7757506d6f

Observation 962f9a8b-8e50-429b-93dd-db8482425c0d · outbound

This paper cites Large Language Models are not Fair Evaluators.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Large Language Models are not Fair Evaluators

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.235975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.235975Z digest=sha256:0df98a555c248fe7551f1321b5433b5235a4c3754f51a2421db19af72b40151c

Observation ce379c9d-1747-46e9-9e0e-bbdd090ed7be · outbound

This paper cites PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization.

Beyond the Surface: Measuring Self-Preference in LLM Judgments PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.346963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.346963Z digest=sha256:ebebc26a024df93bdd4f4219e6f9fc8bf21c5a1a0e3fe3f013032c867b1ca587

Observation 2d3a899b-48da-4516-b36b-31962dfbab41 · outbound

This paper cites Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.467953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.467953Z digest=sha256:58679436cdef6e143e240cc40ea29e1af232e6fd203a8da2739c422cd12a21b1

Observation ed677fea-c2ee-4a76-8c26-83b2f27c212e · outbound

This paper cites Self-Preference Bias in LLM-as-a-Judge.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Self-Preference Bias in LLM-as-a-Judge

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.576445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.576445Z digest=sha256:f5515aee84c40e93ea7d52256b19d578abcfacfc6465b9da48225b8b660e0668

Observation e66676d5-37b4-4900-9a0c-5c695b23be66 · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.652545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.652545Z digest=sha256:1d69ff186f4ed783f4b5a4c5f4a66ebbdaa26f66be4e32a3076b590fc9b3dac4

Observation ac304e2e-cc51-4890-80b2-814a8191dc1f · outbound

This paper cites Evaluating Mathematical Reasoning Beyond Accuracy.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Evaluating Mathematical Reasoning Beyond Accuracy

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.792641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.792641Z digest=sha256:910c025b08e585c9617ad8fedef2f4cfb3912e4383e506afe54682769205b3f1

Observation 7987081c-db41-4372-a1f6-16cf1e2027c9 · outbound

This paper cites Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.902686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.902686Z digest=sha256:8c64561a07815ec7b6c287538470de94e7f9f0e0bb1759f9f20efff9a6a09024

Observation 8e1b10f2-16c8-40f6-ab06-84fa8b35661b · outbound

This paper cites Qwen2.5 Technical Report.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Qwen2.5 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.980884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.980884Z digest=sha256:d353b3aaaa5529600fcf60d851d75ace1b103f86697c28b20fd6410ea29123db

Observation 63155177-c41f-45d0-9a61-73881ae447db · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:07.054479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:07.054479Z digest=sha256:085d6e2f4658d564260de7184209dedf872c71d83fc3e27c72b83d81103265be

Observation 873dcaaf-0b14-48e5-94f7-3d2ec62cc584 · outbound

This paper cites mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval.

Beyond the Surface: Measuring Self-Preference in LLM Judgments mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:07.174272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:07.174272Z digest=sha256:ecc56de3f7fbef183c51eb79cc0065020aeaa59a736a4412f0d06d9f926bcfef

Observation ffef1924-14c8-4f42-ac8b-2302eeb5382b · outbound

This paper cites an unresolved cited work.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:07.266045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:07.266045Z digest=sha256:6fe202b76a756e447396be0e989a90e003d1435daaf85fefddc459e310b3b6a8

Observation ec679e26-bc0e-4ae3-b46e-2637423e0d94 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

Beyond the Surface: Measuring Self-Preference in LLM Judgments JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:07.442164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:07.442164Z digest=sha256:aa7efcd568c0e2abcbc9214e61b2db162669aeb63c74f96698d26e4e122299b4

Pith citing papers

Observation 01dc12c8-47bc-4c49-b235-192caa7839ca · inbound

Extreme Self-Preference in Language Models cites this paper.

Extreme Self-Preference in Language Models Beyond the Surface: Measuring Self-Preference in LLM Judgments

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:00:39.086277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T20:57:37.199128Z digest=sha256:f659d025b74c94451ec7452ef2dc87e7f340f25a974184606d066a5fb508949c

Observation 6ad59c7e-f8f6-4c67-bea4-dd910d61e0f8 · inbound

When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning cites this paper.

When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning Beyond the Surface: Measuring Self-Preference in LLM Judgments

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:51:08.567479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T08:50:53.486588Z digest=sha256:bc9bfa3fee19cda3e7df6d75d603e64a329a4d96b1ea995ac5fa98ae86b1a053

Observation a4352c5a-0e5b-440b-af5c-2a8cb2ba9922 · inbound

Memory Reward Inflation in Self-Improving LLM Agents cites this paper.

Memory Reward Inflation in Self-Improving LLM Agents Beyond the Surface: Measuring Self-Preference in LLM Judgments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T02:16:42.324571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:16:42.324571Z digest=sha256:0ed239b2e9d93dff62b6cacf5e5059b8acb9d27006c17a18a72002e76b7d75a9