Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Cognitive Biases in Large Language Models as Evaluators

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2309.17012.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.17012 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:03.842968Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T12:53:50.367532Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f899883b-0a27-470b-8d38-f46471bf40bf · inbound

LLM Evaluators Recognize and Favor Their Own Generations cites this paper.

LLM Evaluators Recognize and Favor Their Own Generations Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.816877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:ce0d45f72dc28b291fa1b33730535967f8ce336ce4935d8d0d9d1245c06bf2e0

Observation ab7b2c8e-08ff-4d0d-838d-9aed89e9baa7 · inbound

Better & Faster Large Language Models via Multi-token Prediction cites this paper.

Better & Faster Large Language Models via Multi-token Prediction Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:26:09.810803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T12:26:09.731664Z digest=sha256:7dcc28714e56b5c89cbe07a0c3899172f51bd901ffa8bb754134067ed397fe12

Observation fc4e8733-be44-49d0-9fec-7d700a4c89f8 · inbound

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap cites this paper.

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:08:20.988267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T19:07:21.016824Z digest=sha256:0d2c58cf00013f7bdbe7a127f2d193aa6094640e24f868c8609b1752b9ad5cb5

Observation 5a3abb5e-cfdd-4880-860c-93fc57e331bd · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:44.032235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:9b8b3604d9c53ffa282b2ef2680d2096ac80a26355fe1a759de501d8b0cbc7f5

Observation 9e50b7f7-7e23-4f03-9e13-a8302fd0db5b · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.018788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:ad5839ea95661a76995b42e72e2cf96f1edba57bf90b352d156dc10f11d6de04

Observation 2f4d8d5a-931a-4bb6-81bc-60471a111347 · inbound

Beyond the Surface: Measuring Self-Preference in LLM Judgments cites this paper.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.842968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.842968Z digest=sha256:da55cf44052ce3f419e80d8eeb85690269a2f828969e7022fcdde24fd33b848e

Observation 6c134385-792b-4c51-9ea6-155b1800f4e3 · inbound

AbsenceBench: Language Models Can't Tell What's Missing cites this paper.

AbsenceBench: Language Models Can't Tell What's Missing Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:13:17.718512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:13:17.718512Z digest=sha256:c6db4c5657991d6a9f9fb2c66cd4b677d5ec9c5f5ec63b7232a89a22bddd339a

Observation 4568b5f1-b150-4c37-96a1-ad7f735d3388 · inbound

CRISP: Complex Reasoning with Interpretable Step-based Plans cites this paper.

CRISP: Complex Reasoning with Interpretable Step-based Plans Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:01:31.053765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:01:31.053765Z digest=sha256:da5b7ddfaecff2bca6f1d7d0bfa32bf87fa58bfe85de207379753bc3d639bb4d

Observation 602d4ffe-c342-4bef-9e26-f28960e02120 · inbound

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming cites this paper.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.064778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.064778Z digest=sha256:2ec6cd28cf2e633c94121512e5c9a36ba3df6fbde125b7b12db5cf79b50fefe9

Observation 44abe90e-85c8-43a4-9e08-15a1b0802d92 · inbound

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge cites this paper.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.206258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.206258Z digest=sha256:d70f15cba8369556585d78b275876ec302b9f03d8d9180ebc3b82042dd4f6542

Observation a1b445b0-1406-4459-b511-3d2c5ce62956 · inbound

Can You Trick the Grader? Adversarial Persuasion of LLM Judges cites this paper.

Can You Trick the Grader? Adversarial Persuasion of LLM Judges Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:58.213182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:55:58.213182Z digest=sha256:94e9072118034c7162987471bfc9f7d78ff5c7b5236dc047005ee2e74cb440c5

Observation 4e1e2dec-7704-4b8c-935e-200e8f2c817e · inbound

Effectively obtaining acoustic, visual and textual data from videos cites this paper.

Effectively obtaining acoustic, visual and textual data from videos Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.732735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.732735Z digest=sha256:ccc06094c675658535a2ebacdc5c9945a48d3f76936a87f00eb2cd245e2b33d9

Observation e4a5ec78-d9dc-4d66-ba76-3f0446d240da · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.892235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.892235Z digest=sha256:8623565f7ca3401eaefd38b2b9fe7fbbc1d0870df8c2adeffd28f086c838d216

Observation 038f4fa5-de01-4194-bad5-b298963f42eb · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:50:50.481371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:18:19.955943Z digest=sha256:38fe3209023e521965a300d9424b58ff22ad7120dab176e355ada27853cae915

Observation b59a23e2-396e-4e0c-a714-32830efcc736 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T16:40:55.426512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:40:55.426512Z digest=sha256:0b71f9180d35d1bc5d2760efe4820f8a9b247e91c77beaae68ff6df6dbe97ca1

Observation 8ed94691-671b-45a0-9090-854487c94667 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T05:36:43.332913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:36:43.332913Z digest=sha256:5d576c6221f55239af06f399ae72e442c7333a3baf6624ca51d8c69b157febfa

Observation 38c552b8-0b23-44f9-86b9-13366418ef4a · inbound

Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations cites this paper.

Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:44:37.607656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T10:41:21.406201Z digest=sha256:78972fcdedcc1125ac4891267b5fd207380051e81883384d9298dd2476154b20

Observation 2ffbd726-6c6a-476d-9f0e-873e74becc45 · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:41:13.976357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:2f1ff8425a8d536f704ea09dd6131d590cc0cabe824e9c56735adf4e4a3a558d

Observation e1a4976e-a1e9-4a7e-94f5-0d076211b651 · inbound

U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning cites this paper.

U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:39.077576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:19:55.849451Z digest=sha256:6a8334e12e0ac823f13ec8d4105e4f6c41ad64db28c94c46aaff7bf09dedd0f8

Observation 66475279-6144-48ed-9d7e-c63011334597 · inbound

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps cites this paper.

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:33:16.779705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T12:32:22.536994Z digest=sha256:94309e6eb1a2629f3dd164e20257c02b6b1ca862cabdb9d585dc8a54243c7d71

Observation 8bbb5c07-bfda-4855-9660-e75ed527d963 · inbound

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm cites this paper.

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:03:26.715349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T12:54:36.818698Z digest=sha256:1e1c3afb6ae1f34f1e1952a94708a578aae02e72d3ad5528946ebaba7a169a52

Observation fb9f1fd1-16a0-48c7-81b5-28a198b0db93 · inbound

Show, Don't TELL: Explainable AI-Generated Text Detection cites this paper.

Show, Don't TELL: Explainable AI-Generated Text Detection Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:03:26.555612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T12:56:04.003337Z digest=sha256:3d6994cd1fd81e4bb7da718ff2dd504dcedf9abd9efad0565ba48e5cd47a4d4e

Observation 50aad0d0-743e-43ca-9fad-4a0f00ca67f3 · inbound

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? cites this paper.

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.556059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T05:59:58.183264Z digest=sha256:9d57366d2e02ed9112f4512590696601adc5b58857f1fb2b9a1d496aedcc8ea9

Observation 1892a6d9-f0ed-461b-86c8-9c57c92a65c1 · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T12:53:50.369284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-07T12:47:29.552283Z digest=sha256:d0b2bc46ab9c64deb85bdc0f8c83d248a62077b47b0be2dc9e9419fa43696543

Observation c9b73bbf-03de-4ba3-b178-d538abc065d4 · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:f97651de5cf7910eede6b5c2139000c0834f564a68c8a9bcdd3e87d9e3b9c216

Observation cb7ec13a-8c12-44aa-9f13-a015ba53d632 · inbound

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation cites this paper.

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T06:40:17.865408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:40:17.865408Z digest=sha256:42478f7a49aa0dcd2ab28813f3c8e45178ea76e0d768b0eb01c6e6a685fd5805

Observation 2cb7b65c-3a5a-4d79-9294-95f4a01daa4d · inbound

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias cites this paper.

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T02:33:34.084111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T02:33:34.084111Z digest=sha256:4a53c84465b0dc2d40b6770fa654846fd675b908449bb9c6edcce16ce033a034

Observation 48c6a736-6094-48ba-be94-2cf8e6e86dbc · inbound

What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces cites this paper.

What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T16:42:05.597565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T16:42:05.597565Z digest=sha256:1b63f1b8a4dfa36e0caa2490e8c865abbec979393ebbc8e8e46de550dba3987a