Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2412.03597.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:21:36.988483Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T07:33:13.390618Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 6a78e576-5327-4cb2-abe6-488df5659def · inbound
Rethinking the Understanding Ability across LLMs through Mutual Information The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bd3f3c2-d3b9-4673-bab5-53e4d143ff46 · inbound
LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 027116f9-b917-4296-870b-6cba7ef60d16 · inbound
Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3680f679-e001-4f0e-bc07-05f693db9aae · inbound
Human-aligned AI Model Cards with Weighted Hierarchy Architecture The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cd72ad2b-5c21-4586-ba94-6a6d7070fc1a · inbound
Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 20ac7aa0-9ded-4b97-a1af-337206e4eebe · inbound
Training a General Purpose Automated Red Teaming Model The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bd84c14f-df89-4b4d-9160-73d8af3b0f08 · inbound
Latent Performance Profiling of Large Language Models The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7bcf96d0-68e2-4d3c-b056-ed6c588043bb · inbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.