Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2404.18824.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:57:15.748808Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 35845f36-317b-421a-94c2-f2807d182c79 · inbound
Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Benchmarking Benchmark Leakage in Large Language Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e6a2f3b7-9e67-434d-96b4-8ebd827c55a1 · inbound
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models Benchmarking Benchmark Leakage in Large Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91903aa1-6de5-4571-be83-ce728ec2225b · inbound
Are Large Language Models Memorizing Bug Benchmarks? Benchmarking Benchmark Leakage in Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de6c23c1-c11c-4110-9e8a-2f8e4e62b384 · inbound
CNNSum: Exploring Long-Context Summarization with Large Language Models in Chinese Novels Benchmarking Benchmark Leakage in Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f589d2df-966e-4b15-a73d-9f80d7ae1311 · inbound
QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs Benchmarking Benchmark Leakage in Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0987000-d762-46b5-9b84-d3d3b0efc085 · inbound
Large Language Models for Mathematical Analysis Benchmarking Benchmark Leakage in Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c062c923-6be8-43bf-878c-a578ffa12f7a · inbound
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards Benchmarking Benchmark Leakage in Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3730f6b9-d06d-41b0-b116-0aca4d3b5cc3 · inbound
JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models Benchmarking Benchmark Leakage in Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc7cd146-896d-4efd-88a2-9f264f2ea236 · inbound
UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models Benchmarking Benchmark Leakage in Large Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f36a13b3-4485-4c8e-9e26-7ec039e89033 · inbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Benchmarking Benchmark Leakage in Large Language Models
Reference 135
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e677b2b1-af21-474d-a58f-b061c5d09c2f · inbound
Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion Benchmarking Benchmark Leakage in Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 608bd7f2-ffaf-47f5-8cb9-d0a78d2b027b · inbound
Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs Benchmarking Benchmark Leakage in Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbd67ca9-a106-43a1-a3b5-d6a351c50b4f · inbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Benchmarking Benchmark Leakage in Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5f296d6-5c29-4bf5-a846-f1b33b22266a · inbound
Evaluating the Sensitivity of LLMs to Prior Context Benchmarking Benchmark Leakage in Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61f9a667-6db2-4864-ae3c-66699921e2b4 · inbound
Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis Benchmarking Benchmark Leakage in Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79604d1c-895d-4767-aba9-8317fdb7eda2 · inbound
Chain of Methodologies: Scaling Test Time Computation without Training Benchmarking Benchmark Leakage in Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a361e93-bc4a-4fe3-8ff8-29e3ba0a6029 · inbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Benchmarking Benchmark Leakage in Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5349fefa-a554-421c-82ef-0da9cf739b54 · inbound
ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Benchmarking Benchmark Leakage in Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9b1cdac-411b-4f68-9f04-6e3dc58cf728 · inbound
Can Vision Language Models Understand Mimed Actions? Benchmarking Benchmark Leakage in Large Language Models
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad11074b-4823-46d4-91e6-e3be25d48af1 · inbound
Deprecating Benchmarks: Criteria and Framework Benchmarking Benchmark Leakage in Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1481c7fa-2543-4237-a743-10a1e134b5b1 · inbound
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models Benchmarking Benchmark Leakage in Large Language Models
Reference 105
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5a2ae8d-e388-4217-b3eb-eedbb3fa2312 · inbound
GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines Benchmarking Benchmark Leakage in Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4cdd65dc-7dfe-4e5b-8d6e-e89609178a30 · inbound
Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering Benchmarking Benchmark Leakage in Large Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2ede590-2b75-4e82-a160-b3e8b2954827 · inbound
MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science Benchmarking Benchmark Leakage in Large Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d533bd6-19d6-4c41-9c84-3633a1f2fddf · inbound
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation Benchmarking Benchmark Leakage in Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a32d5988-584d-4f29-9cb3-540f7386a92a · inbound
FiMMIA: scaling semantic perturbation-based membership inference across modalities Benchmarking Benchmark Leakage in Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae0c76b-85e0-4c01-be3a-10fb040c358b · inbound
Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners Benchmarking Benchmark Leakage in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5c566624-e304-4be1-af2a-30111054f033 · inbound
SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora Benchmarking Benchmark Leakage in Large Language Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9ccfd0b-1ad1-45a5-8e21-17c30ff436fd · inbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Benchmarking Benchmark Leakage in Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c35f61c7-01d6-473a-9633-244972149695 · inbound
MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events Benchmarking Benchmark Leakage in Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bc4a7044-8616-4f93-b7da-af4e042ed8db · inbound
SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks Benchmarking Benchmark Leakage in Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2a72359c-2c5f-47d6-a70a-8e728236879b · inbound
When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors Benchmarking Benchmark Leakage in Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b5e0477a-8a71-4b88-a2c8-c5c88ac8a16a · inbound
AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models Benchmarking Benchmark Leakage in Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 79bdc75f-a803-40bf-8fcb-ae2fd06e01c5 · inbound
Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Benchmarking Benchmark Leakage in Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 88f9965b-edb3-4822-a956-bfb4c39b339f · inbound
PITMuS: A Tool for Automated Bug Dataset Generation via Source-Level Mutant Reconstruction Benchmarking Benchmark Leakage in Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c7b760cf-d7b0-4618-b25d-e4cca102112c · inbound
PITMuS: A Tool for Automated Bug Dataset Generation via Source-Level Mutant Reconstruction Benchmarking Benchmark Leakage in Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 110f064d-adfe-4ee6-af45-a46623d8e8eb · inbound
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation Benchmarking Benchmark Leakage in Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5e7df553-c4ce-4300-bb33-ea55a9c64edf · inbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Benchmarking Benchmark Leakage in Large Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3732a0e4-37ef-43f7-948a-ee90bf729cdc · inbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Benchmarking Benchmark Leakage in Large Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7725151a-1a91-4028-a8ae-bdefd39a0a07 · inbound
Can AI Agents Synthesize Scientific Conclusions? Benchmarking Benchmark Leakage in Large Language Models
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e5930678-901f-4aff-a538-6f425fcd1788 · inbound
Bridging Functional Correctness and Runtime Efficiency Gaps in LLM-Based Code Translation Benchmarking Benchmark Leakage in Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8970d1b1-991c-44bf-aad8-faefee56bfd7 · inbound
Uncertainty-based Debiasing and Unlearning for Decontamination Benchmarking Benchmark Leakage in Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 027965fd-2235-4d53-a6cf-2e475773e4e1 · inbound
OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents Benchmarking Benchmark Leakage in Large Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e127fd19-8449-43fe-89da-9f0f04b55cb4 · inbound
SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models Benchmarking Benchmark Leakage in Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c4746815-ff22-47ec-924e-2d076714c243 · inbound
Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG Benchmarking Benchmark Leakage in Large Language Models
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 71a2da4b-1397-486a-9975-79a067ba5313 · inbound
Meta-Benchmarks for Financial-Services LLM Evaluation Benchmarking Benchmark Leakage in Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f0c8bc16-9060-40f1-86cf-5ceed93a0191 · inbound
Navigating the Open-Source Model Ecosystem: An Empirical Study of Creator Practices in Artistic Image Generation Benchmarking Benchmark Leakage in Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f42e648e-7f65-4ff9-8770-9e6a013df65d · inbound
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Benchmarking Benchmark Leakage in Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 403211e6-ef5c-494e-a568-3b88a48bfefb · inbound
SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data Benchmarking Benchmark Leakage in Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e3a030c-b9bf-4faf-a812-f2a6152c4c70 · inbound
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Benchmarking Benchmark Leakage in Large Language Models
Reference 236
Source-reported events for the cited work
Unavailable: canonical work link unavailable.