Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T11:08:45.974486Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2601.07506.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T11:08:45.974486Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-03T14:31:49.214297Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T14:38:28.709855Z
17 of 17 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 74d9ae33-2146-4a92-b717-867477a080ff · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c3f32a1-2b77-45b0-9c41-716c3f9f424d · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f53d9d6-9de5-4156-a522-96dfe9c4da04 · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation - Dates/periods → DATE
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15e2f7fd-ff3b-4253-9afd-6cd0e013b514 · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4da665f8-f426-41c5-b018-a046d3e6a5b9 · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation InFindings of the Associa- tion for Computational Linguistics: ACL 2024, pages 12688–12701
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e5202c3-d0d2-4e4f-a376-f823603370c5 · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Entity-Based Knowledge Conflicts in Question Answering
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a98bf69e-66b9-41a4-b4a2-b1d53022ce04 · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation No explanation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7f38292-2c26-4861-b70c-0ba0d2a09f25 · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Mona Lisa
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30268516-3ec6-4c15-8c0b-a0c3d07b6b12 · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c341bd4-31bb-4ace-b673-d138823184c0 · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23fd7538-362e-45ee-aa62-5463b15be321 · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac3fb4d3-a2dc-4201-87b5-ae5c37073664 · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Latent Retrieval for Weakly Supervised Open Domain Question Answering
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa27bbfb-854d-467c-87cf-b4709f37c1d7 · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation A Dataset for Answering Time-Sensitive Questions
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d140dde2-5f93-435c-884a-d0f6664c4a44 · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation InProceedings of the 2022 Con- ference on Empirical Methods in Natural Language Processing, pages 2292–2307
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 029133b2-4152-4420-ab4f-3fd0e1f1b46c · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation ROME: Evaluating Pre-trained Vision-Language Models on Reasoning beyond Visual Common Sense
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5761f192-9849-4276-91fa-7d9f57c27264 · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation InProceedings of the 47th International ACM SI- GIR Conference on Research and Development in Information Retrieval, pages 2811–2816
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a0971c1-12a3-4d1e-b1d1-90443c0a25fc · outbound
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Qwen3 Technical Report
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f123a11-f5f1-4282-8ea8-18b051efc4c1 · inbound
Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
Reference 165
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.