Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T04:45:34.619467Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 25 inbound Pith citation observations for arXiv:2502.03461.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T04:45:34.619467Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:35.170514Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
14 of 14 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 1b9bfd96-d5b6-4cfa-84e2-87a3cc39c5ff · outbound
Do Large Language Model Benchmarks Test Reliability? Vqa: Visual question answering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation db02c29d-d7fc-483a-a2a6-86b7ab1bb954 · outbound
Do Large Language Model Benchmarks Test Reliability? Training Verifiers to Solve Math Word Problems
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82bdb3d6-3b3f-4600-ab45-79296c2a0103 · outbound
Do Large Language Model Benchmarks Test Reliability? DeepSeek-V3 Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 411b121d-201d-435c-8964-c23dd0ce06a8 · outbound
Do Large Language Model Benchmarks Test Reliability? Reasoning with language model is planning with world model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 498f2953-7801-4d3f-9dec-74e579fcd234 · outbound
Do Large Language Model Benchmarks Test Reliability? Parsing algebraic word problems into equations
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 05fdc43a-0a68-41be-aaad-864ab2683443 · outbound
Do Large Language Model Benchmarks Test Reliability? Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcd4c2f5-c4f7-4621-8547-9e8978bb0e2c · outbound
Do Large Language Model Benchmarks Test Reliability? ] The 15th Nepal China’s Tibet Economic and Trade Fair was held on 17-22 November 2015 in Bhrikutimandap, Kathmandu Nepal
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1d55c155-c7b5-47a6-957c-ed5e752ebc3e · outbound
Do Large Language Model Benchmarks Test Reliability? A Chevy for $1? Car dealer chatbots show perils of AI for customer service
Reference 2012
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 763ed826-7b74-4da8-82b9-b6bbac143dc2 · outbound
Do Large Language Model Benchmarks Test Reliability? GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68821419-7ae1-4062-a5f3-d200ee6cddf1 · outbound
Do Large Language Model Benchmarks Test Reliability? TabFact: A Large-scale Dataset for Table-based Fact Verification
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 346a8a38-706f-490b-9a30-289f03360723 · outbound
Do Large Language Model Benchmarks Test Reliability? The Llama 3 Herd of Models
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85b513ad-882e-4c7b-b755-fb7b32d523f8 · outbound
Do Large Language Model Benchmarks Test Reliability? GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f633ecc-145a-4fa1-99df-696893cb17e7 · outbound
Do Large Language Model Benchmarks Test Reliability? What Will it Take to Fix Benchmarking in Natural Language Understanding?
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a42cf9e-0c95-4bfb-9fe4-597c8e5b6b27 · outbound
Do Large Language Model Benchmarks Test Reliability? Are NLP Models really able to Solve Simple Math Word Problems?
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c946dab6-15b8-4f3d-8712-c8ef1f28d4f8 · inbound
CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting Do Large Language Model Benchmarks Test Reliability?
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e91f3fcc-2d10-4a13-b825-ef4dc75f380b · inbound
Benchmarking Misuse Mitigation Against Covert Adversaries Do Large Language Model Benchmarks Test Reliability?
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d35071e2-6a71-4045-afaa-bd7f04821575 · inbound
Domain Specific Benchmarks for Evaluating Multimodal Large Language Models Do Large Language Model Benchmarks Test Reliability?
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a71ed187-186d-4b4f-9697-0ac273e92bb4 · inbound
Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models Do Large Language Model Benchmarks Test Reliability?
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63b7f608-d52f-4b6d-b041-f429a29b044e · inbound
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation Do Large Language Model Benchmarks Test Reliability?
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47cb0821-33b3-4e5b-8e37-95ce3a928b1a · inbound
Kimi K2: Open Agentic Intelligence Do Large Language Model Benchmarks Test Reliability?
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ce22e119-46ce-4751-80fc-ad0fb8f48fd7 · inbound
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Do Large Language Model Benchmarks Test Reliability?
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f6fa3ca2-49ee-4a85-9eb3-6a48992e787d · inbound
Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems Do Large Language Model Benchmarks Test Reliability?
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a505671-19a1-4e9c-bf77-740d23839d0f · inbound
Fluid Language Model Benchmarking Do Large Language Model Benchmarks Test Reliability?
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a218eed-e892-43b4-8b68-880850af79c2 · inbound
Position: AI Evaluations Should be Grounded on a Theory of Capability Do Large Language Model Benchmarks Test Reliability?
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e1db10e7-1e87-4e1d-9f63-3a1b455b282b · inbound
Simple Policy Gradients for Reasoning with Diffusion Language Models Do Large Language Model Benchmarks Test Reliability?
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5e5c7af-49ef-426f-ba8a-3cc8773d3d7e · inbound
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks Do Large Language Model Benchmarks Test Reliability?
Reference 2010
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdb3e2d9-139a-4afd-91a6-040638a7f6c1 · inbound
Model soups need only one ingredient Do Large Language Model Benchmarks Test Reliability?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea4d81c5-9b06-4941-89d2-36dc1cf1db8f · inbound
Weight Decay Improves Language Model Plasticity Do Large Language Model Benchmarks Test Reliability?
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b9344d9-c88c-44ae-ba95-4d3497aa72bd · inbound
Position: Evaluation of ECG Representations Must Be Fixed Do Large Language Model Benchmarks Test Reliability?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 152365d0-7c94-4818-81b7-45f60b924bed · inbound
LLM Reasoning Is Latent, Not the Chain of Thought Do Large Language Model Benchmarks Test Reliability?
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 54fe561a-54e9-4d96-96ff-04e5900b9a83 · inbound
Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity Do Large Language Model Benchmarks Test Reliability?
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 24f5328d-8cc5-4a49-9c81-d91b044b41d4 · inbound
Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation Do Large Language Model Benchmarks Test Reliability?
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e23209f5-b865-4f58-a4c2-071e5dceec73 · inbound
Measuring AI Reasoning: A Guide for Researchers Do Large Language Model Benchmarks Test Reliability?
Reference 151
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f005ee2-3673-4666-9e11-b68055fe34a7 · inbound
Auditing LLM Benchmarks with Item Response Theory Do Large Language Model Benchmarks Test Reliability?
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 67bd4a2f-6efd-449e-87e1-4b08eed3105c · inbound
Flaws in the LLM Automation Narrative Do Large Language Model Benchmarks Test Reliability?
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 55c2c5de-d07c-4225-835d-3ef66f2d1930 · inbound
A Sovereign, Open-Source Foundation Model for German and English Do Large Language Model Benchmarks Test Reliability?
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c50bf4-4920-4331-9b6e-eff107ac780d · inbound
A Sovereign, Open-Source Foundation Model for German and English Do Large Language Model Benchmarks Test Reliability?
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45372cbc-a301-48cf-bc0d-2700a7a04f6f · inbound
A Sovereign, Open-Source Foundation Model for German and English Do Large Language Model Benchmarks Test Reliability?
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ae6023b-7bb3-4441-9584-56729f48151e · inbound
Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility Do Large Language Model Benchmarks Test Reliability?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.