Pith. sign in

Paper Citation Record · LEDGER

Investigating Data Contamination in Modern Benchmarks for Large Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2311.09783.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.09783 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:15:41.440652Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ec73b96e-3e48-43e5-844d-8cc1f4d85384 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:40.906325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:296da6ba9e6b4819a8c24c9d4ffd0c7bdc353806d0a6eb1194c48efe5e08985c

Observation 233cf14c-0080-43bf-9aba-8f16ba7b6c33 · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:48:26.394739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:8286be653c19046952e2dee4d69a3de48d4492c54c1720fc4c25abe85295f393

Observation bae075b8-70e2-44e0-8588-2fd85c0e6973 · inbound

MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters cites this paper.

MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T05:15:41.440652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:15:41.440652Z digest=sha256:356e10a24af15248cfb5f30dc6c89c7e0f6c603aebbed713b1e1655d7fd329ce

Observation 0e558609-c205-45a7-9d6b-e7a4d63507bd · inbound

PRIMETIME : Limits of LLMs in Temporal Primitives cites this paper.

PRIMETIME : Limits of LLMs in Temporal Primitives Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:36:59.121794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T18:36:48.376877Z digest=sha256:61227379c8ceededa8752820c48399dae53d40948527ccd7750e9b2c44d0b2ff

Observation ebacc265-7c29-422e-9e7c-1ed6523700ef · inbound

Federated In-Context Learning: Iterative Refinement for Improved Answer Quality cites this paper.

Federated In-Context Learning: Iterative Refinement for Improved Answer Quality Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:38.666546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:38.666546Z digest=sha256:5a0c424f50a3941e9e20507e48095305e519668993bb632e36d1a18985a470fa

Observation eb8820db-3ad5-485b-a6b4-cb0519f1b63c · inbound

SciDA: Scientific Dynamic Assessor of LLMs cites this paper.

SciDA: Scientific Dynamic Assessor of LLMs Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:17.481039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:17.481039Z digest=sha256:df43acde214808754ebe290b6bfa9ac4ebc59c9007bbf2a0263fefb5c2114610

Observation c1a274ae-e16c-4f1c-b373-528e1058307a · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.291053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:c26313a8267926af2df20bfd1757061bf4a39f5893029a8e1250568d67d28e4b

Observation 6a902d75-6f93-4099-86b0-e9a87819967a · inbound

ZoFia: Zero-Shot Fake News Detection with Entity-Guided Retrieval and Multi-LLM Interaction cites this paper.

ZoFia: Zero-Shot Fake News Detection with Entity-Guided Retrieval and Multi-LLM Interaction Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:50:38.285574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T01:47:35.232468Z digest=sha256:1a246883070c43f1ea6fc066a413c69263e8d0b117c3308f7f63748792a31144

Observation 91ac6d3a-7590-46fd-9add-094893313853 · inbound

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks cites this paper.

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:05.144360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T00:35:24.397273Z digest=sha256:1490b8373c150a689d63bba1c39dbb0b09eaff7077394568a313e5f7b4447183

Observation 3d6b53c8-ead6-415c-a590-edb84a71c7a6 · inbound

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation cites this paper.

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:14.401063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T01:41:42.003483Z digest=sha256:aaa5f799c3ead5ea5ef2e29fcb0bc2c542ae3ccd95aa566b34468952bf9eb956

Observation 4a340f5a-cdf5-44da-8454-28278d68da5a · inbound

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications cites this paper.

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:24:56.558582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:20:16.735285Z digest=sha256:11ce75fef3169b1107a2eb0719af5b481509f903fd15530dc7be2f242eb4fc91

Observation 4b27821f-d48b-413e-bfa5-d9f344fe198b · inbound

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection cites this paper.

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:36:30.302577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T09:48:09.688745Z digest=sha256:a5f5087d92df71f9f054ecd08057a15a47d16769a01712582a16d7015e41566f

Observation 7dfa333e-46d4-421f-90ba-f63ecc16cedb · inbound

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection cites this paper.

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:39:16.485829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-04T00:30:13.665405Z digest=sha256:046499d7d31346eafa6964670124f5832ab02e237740715a67ae6430ed555c1d

Observation 86763adc-0a71-42c0-969b-9498eb69b3d8 · inbound

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier cites this paper.

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:57:48.182085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T10:36:09.211639Z digest=sha256:42f58cc1ef49b3b92e76b84cb386f75af4107d3095c163c4a2cd0e58fff6d2b6

Observation b65907d7-b88a-45de-9d53-32fb3f9e9b5d · inbound

SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks cites this paper.

SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T18:37:16.481985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-02T18:19:43.146102Z digest=sha256:a2d6ef54e4de852543d7c05170bff2a1378ecd0dcac0337e7472ba65f8648e7f

Observation 832cdfff-61ee-400e-9c61-dcf0b491a644 · inbound

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains cites this paper.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.374252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.374252Z digest=sha256:7c95e2f4ddd893c2c5cb4a3d6d118498c8981adf56b4df2cb2e2b1cc7e57bf70