Pith. sign in

Paper Citation Record · LEDGER

Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2411.03923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.03923 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:35:58.978172Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:16:58.130004Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bb55b6ed-45c0-445b-a95d-ebfd06e968c7 · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:48:26.481297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:4381c7d1d238d626be78ad707da27b2cfce8c30c606d362e07966869db607695

Observation 73280c21-3eb1-4554-a040-18d3fedcc1b6 · inbound

Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models cites this paper.

Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T15:35:58.978172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:35:58.978172Z digest=sha256:95e84b91f2a3612b9d9e7af61964171a21d7a7cb6e9698dd131381ca1cf5539a

Observation 2efcd1d2-87fe-4223-af5d-8a923f59a299 · inbound

HARP: A challenging human-annotated math reasoning benchmark cites this paper.

HARP: A challenging human-annotated math reasoning benchmark Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.704545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.704545Z digest=sha256:c33ffd3c81bf43f8211e42f9842bbf19307221f85cdfe94d9bb573e5627796c9

Observation aa2be92d-3aa0-4151-83b0-90cfc9250743 · inbound

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation cites this paper.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.132126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.132126Z digest=sha256:622ae1e4e4d5a0537fc7921c0f13acf0b4d0a067e3ad9a659639482c0244891d

Observation 78b94966-d4c8-42f7-95be-366900a8aa7b · inbound

CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction cites this paper.

CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T13:11:51.730776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:11:51.730776Z digest=sha256:1c0a29d33a6ac58344bc0e0d0349d4b3c0ce6d4cd769d6f77a05c036548d7c23

Observation 9d1eb138-e830-4a39-bfb2-be530c42a6d8 · inbound

StockSim: A Dual-Mode Order-Level Simulator for Evaluating Multi-Agent LLMs in Financial Markets cites this paper.

StockSim: A Dual-Mode Order-Level Simulator for Evaluating Multi-Agent LLMs in Financial Markets Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:26.695956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:26.695956Z digest=sha256:6df194e8956bd374c81c5dc6bda34fed719157b1f5b389388a7529be515a8cc3

Observation e2a72f36-9d7c-4e11-a269-6549549feb13 · inbound

Position: AI Evaluations Should be Grounded on a Theory of Capability cites this paper.

Position: AI Evaluations Should be Grounded on a Theory of Capability Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:00:41.475033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-21T21:57:55.834632Z digest=sha256:b81c8a7993be0913f7a86c37363efa76a29f6f74e4d9364b0122c4f72a59c13b

Observation ca0e5472-5a92-4f1e-80c4-6a08e2760cb1 · inbound

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks cites this paper.

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:29.688964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T04:27:11.735657Z digest=sha256:7b1fbd527e97ea3b058c1b3ed5ee0397118ab3ed519f383d6d3a03d83b49c59a

Observation 60f42c4b-1ba6-476c-a6b6-5419dd0d799f · inbound

Measuring AI Reasoning: A Guide for Researchers cites this paper.

Measuring AI Reasoning: A Guide for Researchers Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:05:36.930166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-08T18:53:18.586923Z digest=sha256:e35006d4f9a94752763f01ba87eadad98b002b1b1f2ff9e1856c4f909050c25f

Observation 09cbb8ab-6e55-4a22-91c5-76b151197650 · inbound

Dataset Watermarking for Closed LLMs with Provable Detection cites this paper.

Dataset Watermarking for Closed LLMs with Provable Detection Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:57.253215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T00:53:42.185498Z digest=sha256:2095b504dbce3280bf9a9c56438ece2e117892f531cab7394d2cff7ac5e1e812

Observation c3d3fcce-b28a-408a-bbfd-68acd20b10e8 · inbound

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild cites this paper.

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.153124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T14:41:07.354007Z digest=sha256:a1b6ecdc870ed4404b607daa113bcf6c14243da2be8691acd3eafaf1038d6d18

Observation a8712fb1-cb99-4579-95f3-03c93b80a763 · inbound

Amplifying Membership Signal Through Chained Regeneration cites this paper.

Amplifying Membership Signal Through Chained Regeneration Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:39.288910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T06:22:08.140403Z digest=sha256:135a3c85525f343f9feae5499919ee4cf94832cc92b3e7a25cc45546bb9396ba

Observation fd18870b-ebe7-4d86-91f6-677dc74517e0 · inbound

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages cites this paper.

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.131717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-02T13:14:15.684960Z digest=sha256:1bc04c30dc530d9bed278812fbc2c8b55016f9752b93232d8bed77e2e504f359

Observation e9a46123-ac97-4ffe-a532-9b4ac9c3f628 · inbound

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy cites this paper.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.669375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.669375Z digest=sha256:5bb14b7f19872f370449189c780524223c027c258799ef9c2cae405906a39d70

Observation c0328c06-8a03-4415-88fd-e754b03f040c · inbound

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores cites this paper.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.401902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.401902Z digest=sha256:228b459e9ea832a5121b221c1b1f860e0a7e4f260594b5761ed641ebdd1d415d