Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2402.09880.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T15:06:55.090364Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T23:10:41.039569Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation a670cd0f-1687-4bc1-96ff-f2dccbbd6bbf · inbound
Benchmark Data Contamination of Large Language Models: A Survey Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2f38d407-5bf1-4fe6-ad94-b89fce08349a · inbound
Humanity's Last Exam Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 58a662b0-38db-4b24-8cf1-35d1af3a5c68 · inbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46ecd899-cc2d-42fc-91b0-27259ba69af6 · inbound
Unbiased Evaluation of Large Language Models from a Causal Perspective Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d00f127b-f3c5-4b99-b092-d7e1cc5deca9 · inbound
Evaluating the Sensitivity of LLMs to Prior Context Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac7fea2d-36bc-4c4a-96dd-4c3d6f2e3949 · inbound
A Conceptual Framework for AI Capability Evaluations Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f1c3676-d795-4007-a5ca-e4cf5528221d · inbound
Benchmarking the Pedagogical Knowledge of Large Language Models Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 730348d7-1f43-4ae1-8aea-8f0731132ca0 · inbound
Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 243
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c272ff5-c8e7-4443-82c4-911beb380184 · inbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e79f6e4-785c-4eab-9a33-5d27681e7eac · inbound
Deprecating Benchmarks: Criteria and Framework Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8326260e-4bf2-4873-b1c8-2522a225cb5b · inbound
DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e2b61dff-efae-45ec-847d-e10efba11dfd · inbound
DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c9c93e9-8fc0-4bda-a6c1-b5b050d53f65 · inbound
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.