Pith. sign in

Paper Citation Record · LEDGER

Investigating Data Contamination in Modern Benchmarks for Large Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2311.09783.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.09783 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:15:41.440652Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ec73b96e-3e48-43e5-844d-8cc1f4d85384 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:40.906325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:2a6ae6356c4061b53a205bf605a005f2c91eb6278a5d57431016308b2b840850

Observation 233cf14c-0080-43bf-9aba-8f16ba7b6c33 · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:48:26.394739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:3ca5e20f5a1bdc786dded5ee1c8367bb6143d11ccd99b1ec2e4386e99decc1c5

Observation bae075b8-70e2-44e0-8588-2fd85c0e6973 · inbound

MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters cites this paper.

MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T05:15:41.440652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:15:41.440652Z digest=sha256:356e10a24af15248cfb5f30dc6c89c7e0f6c603aebbed713b1e1655d7fd329ce

Observation 0e558609-c205-45a7-9d6b-e7a4d63507bd · inbound

PRIMETIME : Limits of LLMs in Temporal Primitives cites this paper.

PRIMETIME : Limits of LLMs in Temporal Primitives Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:36:59.121794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T18:36:48.376877Z digest=sha256:7708a55c9dab7a99cb9611193db4ba460d2c303bc97c2db8193a410f8dd19991

Observation ebacc265-7c29-422e-9e7c-1ed6523700ef · inbound

Federated In-Context Learning: Iterative Refinement for Improved Answer Quality cites this paper.

Federated In-Context Learning: Iterative Refinement for Improved Answer Quality Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:38.666546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:38.666546Z digest=sha256:5a0c424f50a3941e9e20507e48095305e519668993bb632e36d1a18985a470fa

Observation eb8820db-3ad5-485b-a6b4-cb0519f1b63c · inbound

SciDA: Scientific Dynamic Assessor of LLMs cites this paper.

SciDA: Scientific Dynamic Assessor of LLMs Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:17.481039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:17.481039Z digest=sha256:df43acde214808754ebe290b6bfa9ac4ebc59c9007bbf2a0263fefb5c2114610

Observation c1a274ae-e16c-4f1c-b373-528e1058307a · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.291053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:e35c4d22eca23dc9b9aca6183245ff1bae79d73c9e4a22a72b3f2cf51d619d3e

Observation 6a902d75-6f93-4099-86b0-e9a87819967a · inbound

ZoFia: Zero-Shot Fake News Detection with Entity-Guided Retrieval and Multi-LLM Interaction cites this paper.

ZoFia: Zero-Shot Fake News Detection with Entity-Guided Retrieval and Multi-LLM Interaction Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:50:38.285574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T01:47:35.232468Z digest=sha256:cc239fa27786c431a2d6dee2a7166295597bd810131e3317db4058dc712fe705

Observation 91ac6d3a-7590-46fd-9add-094893313853 · inbound

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks cites this paper.

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:05.144360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:35:24.397273Z digest=sha256:4af5569d913657fd9aba02206085164906093532a2216b73aa0dc5dac9190238

Observation 3d6b53c8-ead6-415c-a590-edb84a71c7a6 · inbound

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation cites this paper.

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:14.401063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T01:41:42.003483Z digest=sha256:ffc34363e22f1c083c67083eb9eb36e2aea0df7ea535fe8d62f5c5d09460bcf8

Observation 4a340f5a-cdf5-44da-8454-28278d68da5a · inbound

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications cites this paper.

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:24:56.558582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:20:16.735285Z digest=sha256:8ff0c457e77a11ee9636887f4e2b13672ee28208f3f763a7555e80b1552d5e41

Observation 4b27821f-d48b-413e-bfa5-d9f344fe198b · inbound

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection cites this paper.

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:36:30.302577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T09:48:09.688745Z digest=sha256:d4eb067b600bd62ddfc9a63788ed01b6d2dd194a10a30292e0d272717d7452d6

Observation 7dfa333e-46d4-421f-90ba-f63ecc16cedb · inbound

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection cites this paper.

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:39:16.485829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-04T00:30:13.665405Z digest=sha256:99aeb570404e8784087bc07505b762bdcdee4070b0f0e721424ac6a74ae47895

Observation 86763adc-0a71-42c0-969b-9498eb69b3d8 · inbound

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier cites this paper.

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:57:48.182085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T10:36:09.211639Z digest=sha256:8c8ba97c141a883d66ed343c499e032ff5250c9dd7a4acefb3d4a6d028a57ab8

Observation b65907d7-b88a-45de-9d53-32fb3f9e9b5d · inbound

SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks cites this paper.

SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T18:37:16.481985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-02T18:19:43.146102Z digest=sha256:3d88f19b27794698885a4deda5b998a03802bffb76096ef886df249b3eb9e512

Observation 832cdfff-61ee-400e-9c61-dcf0b491a644 · inbound

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains cites this paper.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.374252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.374252Z digest=sha256:3a2a155b17f2bda9f7936f3909d139ace78b939adaf301a80ca56aa1d17f2b19