Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Cognitive Biases in Large Language Models as Evaluators

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2309.17012.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.17012 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:51:25.722030Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T12:53:50.367532Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f899883b-0a27-470b-8d38-f46471bf40bf · inbound

LLM Evaluators Recognize and Favor Their Own Generations cites this paper.

LLM Evaluators Recognize and Favor Their Own Generations Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.816877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:79777c336362ba8b260421e04bb21bed7e3b251f135a6e0762947efc67b8eb84

Observation ab7b2c8e-08ff-4d0d-838d-9aed89e9baa7 · inbound

Better & Faster Large Language Models via Multi-token Prediction cites this paper.

Better & Faster Large Language Models via Multi-token Prediction Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:26:09.810803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T12:26:09.731664Z digest=sha256:60c9ab950572855ceef92591ec67e3d34e4133162caa0871bb6c1b71b366aa60

Observation fc4e8733-be44-49d0-9fec-7d700a4c89f8 · inbound

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap cites this paper.

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:08:20.988267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T19:07:21.016824Z digest=sha256:5ebaba6cfc9c52af3b083cf455d7432b37876d147ebeb162e62274c0b08a3041

Observation 5a3abb5e-cfdd-4880-860c-93fc57e331bd · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:44.032235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:43690353ddc092beb8a42ef75aa75891e1e768999e3c04e2907f6872d053f741

Observation 9e50b7f7-7e23-4f03-9e13-a8302fd0db5b · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.018788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:163413364b05d1494a6a9d24e3cb7790676ab53eee9176a59100574229c90c54

Observation 6b7b52ec-d510-4b2a-ae3e-41f8d6522a8a · inbound

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models cites this paper.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.722030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.722030Z digest=sha256:8929d28e4688da3ef15ccc081c0e8e702f7652a1634cea1b01146135efab8413

Observation 2f4d8d5a-931a-4bb6-81bc-60471a111347 · inbound

Beyond the Surface: Measuring Self-Preference in LLM Judgments cites this paper.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.842968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.842968Z digest=sha256:da55cf44052ce3f419e80d8eeb85690269a2f828969e7022fcdde24fd33b848e

Observation 6c134385-792b-4c51-9ea6-155b1800f4e3 · inbound

AbsenceBench: Language Models Can't Tell What's Missing cites this paper.

AbsenceBench: Language Models Can't Tell What's Missing Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:13:17.718512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:13:17.718512Z digest=sha256:ec6498509045d0d3dfb8b41a9dbc150c08a1a188cc1266bb601244c27f960ef7

Observation 4568b5f1-b150-4c37-96a1-ad7f735d3388 · inbound

CRISP: Complex Reasoning with Interpretable Step-based Plans cites this paper.

CRISP: Complex Reasoning with Interpretable Step-based Plans Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:01:31.053765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:01:31.053765Z digest=sha256:b1713c68e67c60b65ccdf91c21371d375b711ca317923a5271beb10593e2a0fb

Observation 602d4ffe-c342-4bef-9e26-f28960e02120 · inbound

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming cites this paper.

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T11:45:41.064778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:45:41.064778Z digest=sha256:2ec6cd28cf2e633c94121512e5c9a36ba3df6fbde125b7b12db5cf79b50fefe9

Observation 44abe90e-85c8-43a4-9e08-15a1b0802d92 · inbound

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge cites this paper.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.206258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.206258Z digest=sha256:fa530d35ebce183cb6d8d8c4dd8cea3fbfdf3ede682ff74d572a5ece6600de09

Observation a1b445b0-1406-4459-b511-3d2c5ce62956 · inbound

Can You Trick the Grader? Adversarial Persuasion of LLM Judges cites this paper.

Can You Trick the Grader? Adversarial Persuasion of LLM Judges Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:58.213182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:55:58.213182Z digest=sha256:94e9072118034c7162987471bfc9f7d78ff5c7b5236dc047005ee2e74cb440c5

Observation 4e1e2dec-7704-4b8c-935e-200e8f2c817e · inbound

Effectively obtaining acoustic, visual and textual data from videos cites this paper.

Effectively obtaining acoustic, visual and textual data from videos Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.732735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.732735Z digest=sha256:11350b16986ce210d13812b1bd59493adcbe164b569b6c84979c55cd34adc03e

Observation e4a5ec78-d9dc-4d66-ba76-3f0446d240da · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:26.892235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:26.892235Z digest=sha256:8623565f7ca3401eaefd38b2b9fe7fbbc1d0870df8c2adeffd28f086c838d216

Observation 038f4fa5-de01-4194-bad5-b298963f42eb · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:50:50.481371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T18:18:19.955943Z digest=sha256:ddbc9578f5d8b72d247cc69c20bf88f96c97d8426c89c866e21e3c2b26a2155f

Observation b59a23e2-396e-4e0c-a714-32830efcc736 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T16:40:55.426512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:40:55.426512Z digest=sha256:7040db7452f4b3d34d0bd03d73340e6c5b426a78c80028478a922273e46e863f

Observation 8ed94691-671b-45a0-9090-854487c94667 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T05:36:43.332913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:36:43.332913Z digest=sha256:8cb275a50473b8fac81642285af59f8b5234774923eda161724033a1ec5aa2e1

Observation 38c552b8-0b23-44f9-86b9-13366418ef4a · inbound

Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations cites this paper.

Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:44:37.607656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T10:41:21.406201Z digest=sha256:6ba51ba42203403e7129ce9ad5f5f71b494d0e4605afc96ea2cc707696eccb11

Observation 2ffbd726-6c6a-476d-9f0e-873e74becc45 · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:41:13.976357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:cf74fdb1f1201fe71b0dd48056914e729bacead9b27038f48945ded7a11ade66

Observation e1a4976e-a1e9-4a7e-94f5-0d076211b651 · inbound

U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning cites this paper.

U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:39.077576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T18:19:55.849451Z digest=sha256:d907efc6b3e3e7d39822092e441fc09d4f25a17538cb02f0a79e9856d297103a

Observation 66475279-6144-48ed-9d7e-c63011334597 · inbound

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps cites this paper.

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:33:16.779705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T12:32:22.536994Z digest=sha256:3f02d8205ae8139ca2ecce565b5f27858b97291d3ec2c35b06c669900bceb63b

Observation 8bbb5c07-bfda-4855-9660-e75ed527d963 · inbound

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm cites this paper.

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:03:26.715349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T12:54:36.818698Z digest=sha256:da24fc01c88a1acca65c6c842cc635ab77c3f96be5244b5aa04ef3062a5d669c

Observation fb9f1fd1-16a0-48c7-81b5-28a198b0db93 · inbound

Show, Don't TELL: Explainable AI-Generated Text Detection cites this paper.

Show, Don't TELL: Explainable AI-Generated Text Detection Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:03:26.555612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T12:56:04.003337Z digest=sha256:6ee3fc52fa6a2807776729193e45ebbca53ea50bb2d934fd9cac0581064fe11f

Observation 50aad0d0-743e-43ca-9fad-4a0f00ca67f3 · inbound

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? cites this paper.

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.556059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-30T05:59:58.183264Z digest=sha256:6a089d9a8ab991b20a3dc5939b9864372e8955d464dd9bbe2223ac615cde7225

Observation 1892a6d9-f0ed-461b-86c8-9c57c92a65c1 · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T12:53:50.369284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-07T12:47:29.552283Z digest=sha256:ec6a1dfe2f6968a61e48052a06945eae421e945f80655da4799c25080d7a2cf9

Observation c9b73bbf-03de-4ba3-b178-d538abc065d4 · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:f97651de5cf7910eede6b5c2139000c0834f564a68c8a9bcdd3e87d9e3b9c216

Observation cb7ec13a-8c12-44aa-9f13-a015ba53d632 · inbound

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation cites this paper.

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T06:40:17.865408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:40:17.865408Z digest=sha256:42478f7a49aa0dcd2ab28813f3c8e45178ea76e0d768b0eb01c6e6a685fd5805

Observation 2cb7b65c-3a5a-4d79-9294-95f4a01daa4d · inbound

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias cites this paper.

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T02:33:34.084111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T02:33:34.084111Z digest=sha256:4a53c84465b0dc2d40b6770fa654846fd675b908449bb9c6edcce16ce033a034

Observation 48c6a736-6094-48ba-be94-2cf8e6e86dbc · inbound

What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces cites this paper.

What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T16:42:05.597565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T16:42:05.597565Z digest=sha256:1b63f1b8a4dfa36e0caa2490e8c865abbec979393ebbc8e8e46de550dba3987a