Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2402.01781.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:23.929114Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-16T10:47:31.141279Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation fba070fe-dd71-477f-9785-f490b8588e91 · inbound
RouterBench: A Benchmark for Multi-LLM Routing System When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a93ec630-105a-4a4c-a428-285fb95e7f65 · inbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9e97d5ea-bf91-479e-a756-e624b0e3e64d · inbound
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec67dca3-9d33-4a33-bbc2-7c9dbcbce281 · inbound
Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fd1be3c-d545-4ff7-98b1-2696f4ee6316 · inbound
Too Big to Fool: Resisting Deception in Language Models When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc6b0e4d-8ae7-4e3b-b0c6-2ed46a0dbf33 · inbound
Evaluating LLM Reasoning in the Operations Research Domain with ORQA When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc8883a8-ff25-4446-a322-8650f92e5d4f · inbound
A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c1b2e4c-07fe-4278-963b-01609ebcdb51 · inbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aabf32b-0f38-4f7f-9150-377d5b946014 · inbound
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8075e7fb-d2d2-4739-9b6d-eb560a6331db · inbound
RoToR: Towards More Reliable Responses for Order-Invariant Inputs When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d59d5cb2-0912-422a-a0cf-c2dece4483e3 · inbound
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d2fff4b-062f-4cdd-a428-2e2513130e46 · inbound
Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d6a5638-e8f3-4e0e-9d76-0c8b5cbfff94 · inbound
Quantifying Memory Utilization with Effective State-Size When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b6341e-63f4-46a0-8f7d-e48e562108f8 · inbound
Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.