Pith. sign in

Paper Citation Record · LEDGER

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation

As of 16 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2601.07506.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.07506 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:08:45.974486Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T14:31:49.214297Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T14:38:28.709855Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 74d9ae33-2146-4a92-b717-867477a080ff · outbound

This paper cites an unresolved cited work.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:44.828348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:44.828348Z digest=sha256:ebf36b442a5558cce058c580c68c2ddf9e892251d1eaffaf1d297af7ff844ce7

Observation 5c3f32a1-2b77-45b0-9c41-716c3f9f424d · outbound

This paper cites an unresolved cited work.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:45.003915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:45.003915Z digest=sha256:085dd6016f979e9b254fd2b6be801dfc30bd34c68110386ab584d08cf94f4b5f

Observation 3f53d9d6-9de5-4156-a522-96dfe9c4da04 · outbound

This paper cites - Dates/periods → DATE.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation - Dates/periods → DATE

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:45.237417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:45.237417Z digest=sha256:027d5ce74ae24359b1f9594df5f72b3cb3e4d57bebab42cbed107dc5d9cd5e9c

Observation 15e2f7fd-ff3b-4253-9afd-6cd0e013b514 · outbound

This paper cites an unresolved cited work.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:45.431688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:45.431688Z digest=sha256:60583e1047b0fb7971cae4323ea91d4143f01d2dda5f0be45e246d2b7e00b2bc

Observation 4da665f8-f426-41c5-b018-a046d3e6a5b9 · outbound

This paper cites InFindings of the Associa- tion for Computational Linguistics: ACL 2024, pages 12688–12701.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation InFindings of the Associa- tion for Computational Linguistics: ACL 2024, pages 12688–12701

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:44.037230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:44.037230Z digest=sha256:a3d426f57d94a601470425dadd31c1eadbc8e2278d9dca8af6c8bdfb9ce9b3d7

Observation 0e5202c3-d0d2-4e4f-a376-f823603370c5 · outbound

This paper cites Entity-Based Knowledge Conflicts in Question Answering.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Entity-Based Knowledge Conflicts in Question Answering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:44.172839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:44.172839Z digest=sha256:5672a322f000299c65c147cfebc3eb220262449a2dcf5c1ca91629e06640614a

Observation a98bf69e-66b9-41a4-b4a2-b1d53022ce04 · outbound

This paper cites No explanation.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation No explanation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:45.856255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:45.856255Z digest=sha256:ec2e7e8682763f35e71c5f4025a2a8ece84dcb550a00e10c17ff71e77e623d7a

Observation b7f38292-2c26-4861-b70c-0ba0d2a09f25 · outbound

This paper cites Mona Lisa.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Mona Lisa

Reference 8

Resolution
malformed identifier
no resolver link, observed 2026-08-03T11:08:45.974486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:45.974486Z digest=sha256:1243b7812c8f973b7c985d1e0bb4f5dacdc6a1b98e57279e03cb1682c221473b

Observation 30268516-3ec6-4c15-8c0b-a0c3d07b6b12 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:44.678747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:44.678747Z digest=sha256:41a31765a1d0a2da7345c3ba0f8f3578724b016089eab249b35eb3389cf652aa

Observation 5c341bd4-31bb-4ace-b673-d138823184c0 · outbound

This paper cites an unresolved cited work.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:45.596783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:45.596783Z digest=sha256:178e114c1ee40d3ea3e9eefaccb9d03556fc4970309d6e42c02bb49fd2b87ea7

Observation 23fd7538-362e-45ee-aa62-5463b15be321 · outbound

This paper cites an unresolved cited work.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:45.724786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:45.724786Z digest=sha256:b68ee13e1de258169fd500a6597e06ea0f91d63f5226f66291987f504027b85d

Observation ac3fb4d3-a2dc-4201-87b5-ae5c37073664 · outbound

This paper cites Latent Retrieval for Weakly Supervised Open Domain Question Answering.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Latent Retrieval for Weakly Supervised Open Domain Question Answering

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:43.916320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:43.916320Z digest=sha256:61cb0846ae9a6a0f9a680bc75dfda5d1752e6b1de8b13b4b03353be06e819877

Observation aa27bbfb-854d-467c-87cf-b4709f37c1d7 · outbound

This paper cites A Dataset for Answering Time-Sensitive Questions.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation A Dataset for Answering Time-Sensitive Questions

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:43.665042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:43.665042Z digest=sha256:c912f42ad9bdebe7bae67f93cb9e32795799cfa14d2d0edb497d03c746129a94

Observation d140dde2-5f93-435c-884a-d0f6664c4a44 · outbound

This paper cites InProceedings of the 2022 Con- ference on Empirical Methods in Natural Language Processing, pages 2292–2307.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation InProceedings of the 2022 Con- ference on Empirical Methods in Natural Language Processing, pages 2292–2307

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:43.526902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:43.526902Z digest=sha256:8bfcaeee07acafd2216985e33ecde8c8c713962dcc0c8b53d0730278c9ef443f

Observation 029133b2-4152-4420-ab4f-3fd0e1f1b46c · outbound

This paper cites ROME: Evaluating Pre-trained Vision-Language Models on Reasoning beyond Visual Common Sense.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation ROME: Evaluating Pre-trained Vision-Language Models on Reasoning beyond Visual Common Sense

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:44.483762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:44.483762Z digest=sha256:9008b62caa475a9bb149c098c32ac488fe369423d201b9ccca2f3b2a6cf04a24

Observation 5761f192-9849-4276-91fa-7d9f57c27264 · outbound

This paper cites InProceedings of the 47th International ACM SI- GIR Conference on Research and Development in Information Retrieval, pages 2811–2816.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation InProceedings of the 47th International ACM SI- GIR Conference on Research and Development in Information Retrieval, pages 2811–2816

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:43.783284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:43.783284Z digest=sha256:5e31a44750614cfcb777fe70db576788bceeff3323524f7d7d4c4d2dc73095be

Observation 4a0971c1-12a3-4d1e-b1d1-90443c0a25fc · outbound

This paper cites Qwen3 Technical Report.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Qwen3 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:44.325810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:44.325810Z digest=sha256:8013da62532d2c21145f6c45355150e3308ad4f59824d666a7f581be7a8fa7b0

Pith citing papers

Observation 3f123a11-f5f1-4282-8ea8-18b051efc4c1 · inbound

Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages cites this paper.

Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation

Reference 165

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:38:28.711235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-03T14:31:49.214297Z digest=sha256:ea816ad5576c081b45cb1747a6357f2f6e020d1e7c3d35573ddb1052a12aa3d6