Pith. sign in

Paper Citation Record · LEDGER

Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2411.03923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.03923 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:35:45.132126Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:16:58.130004Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bb55b6ed-45c0-445b-a95d-ebfd06e968c7 · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:48:26.481297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:77acd073f55caf212a4439a7ee60a55b05bfd26931172c6821a7243dde8e1c97

Observation aa2be92d-3aa0-4151-83b0-90cfc9250743 · inbound

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation cites this paper.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.132126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.132126Z digest=sha256:f0d23438ffd6f3af3d393bbb0d99fb06056e467ac27f0392d977ef367f0f340f

Observation 78b94966-d4c8-42f7-95be-366900a8aa7b · inbound

CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction cites this paper.

CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T13:11:51.730776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:11:51.730776Z digest=sha256:bb037f108791571a7fdefff1c643ada3a7b7af7029461d18754813d0d2c791c2

Observation 9d1eb138-e830-4a39-bfb2-be530c42a6d8 · inbound

StockSim: A Dual-Mode Order-Level Simulator for Evaluating Multi-Agent LLMs in Financial Markets cites this paper.

StockSim: A Dual-Mode Order-Level Simulator for Evaluating Multi-Agent LLMs in Financial Markets Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:26.695956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:26.695956Z digest=sha256:610986d975cebd3ab2f0b573d7c63e250354896c2e152cbe5c6bfc125d26204a

Observation e2a72f36-9d7c-4e11-a269-6549549feb13 · inbound

Position: AI Evaluations Should be Grounded on a Theory of Capability cites this paper.

Position: AI Evaluations Should be Grounded on a Theory of Capability Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:00:41.475033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T21:57:55.834632Z digest=sha256:0613f4c9d63164a2f784ac3be9e36fad2af0f81b0714228e1d2d966155a7fa21

Observation ca0e5472-5a92-4f1e-80c4-6a08e2760cb1 · inbound

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks cites this paper.

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:29.688964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T04:27:11.735657Z digest=sha256:ca50043e14d9aee850131dc9c1aa9de45de58b427590f433fda6331f531aabe3

Observation 60f42c4b-1ba6-476c-a6b6-5419dd0d799f · inbound

Measuring AI Reasoning: A Guide for Researchers cites this paper.

Measuring AI Reasoning: A Guide for Researchers Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:05:36.930166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T18:53:18.586923Z digest=sha256:4f9ffca5b9b49c05a0f4b3cbc203769d11964a506e555fa12dce9377b49c7d8d

Observation 09cbb8ab-6e55-4a22-91c5-76b151197650 · inbound

Dataset Watermarking for Closed LLMs with Provable Detection cites this paper.

Dataset Watermarking for Closed LLMs with Provable Detection Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:57.253215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T00:53:42.185498Z digest=sha256:755d25f4e2fd8938aae2d7c8d4e168a1242a8ab9a9e1da0acc1d0a2b3490870d

Observation c3d3fcce-b28a-408a-bbfd-68acd20b10e8 · inbound

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild cites this paper.

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.153124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T14:41:07.354007Z digest=sha256:496feabcb269473fe190698342336d325a73c06f0babbaa2f553072e638186f1

Observation a8712fb1-cb99-4579-95f3-03c93b80a763 · inbound

Amplifying Membership Signal Through Chained Regeneration cites this paper.

Amplifying Membership Signal Through Chained Regeneration Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:39.288910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T06:22:08.140403Z digest=sha256:aa483f3b86236bfb7db61e7e94293e7b5f8a7b02bbad7460e8d76133ced23242

Observation fd18870b-ebe7-4d86-91f6-677dc74517e0 · inbound

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages cites this paper.

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.131717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-02T13:14:15.684960Z digest=sha256:1b0c67f5e35ba6a80b5479d59c1ec46286fa47a3514c9f7055baa7c57c09d149

Observation e9a46123-ac97-4ffe-a532-9b4ac9c3f628 · inbound

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy cites this paper.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.669375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.669375Z digest=sha256:4b26c20a200e5c385eba816d68f263b9bd21c256374a29206524a57f91388c5e

Observation c0328c06-8a03-4415-88fd-e754b03f040c · inbound

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores cites this paper.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.401902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.401902Z digest=sha256:21d1836b018a7793249871a84ab85afd1f7f2a07f6109eae86335ff16e504223