Pith. sign in

Paper Citation Record · LEDGER

Investigating Data Contamination for Pre-training Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2401.06059.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.06059 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:34.145164Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T21:17:24.167881Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 89a0abb8-07bf-4c42-a787-5cfdcd8e09cd · inbound

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting cites this paper.

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting Investigating Data Contamination for Pre-training Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:34.145164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:34.145164Z digest=sha256:5cd66890b4dc15d66c2c516d28ce581d35bc036e01b9d926fffa268f5817b6c3

Observation b460b024-0bbc-47a7-b8f4-f7159f6800f8 · inbound

Can Vision Language Models Understand Mimed Actions? cites this paper.

Can Vision Language Models Understand Mimed Actions? Investigating Data Contamination for Pre-training Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:41.260222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:23:41.260222Z digest=sha256:f65c558fd5adab96bf26b553c565d75c0c0c5dbd720c2fbc89f9e3af750e2239

Observation 367caea1-da5b-4b5a-98a1-cae7a28bf80c · inbound

Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data cites this paper.

Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data Investigating Data Contamination for Pre-training Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:23.208794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:23.208794Z digest=sha256:78dfd667473717e3a3ed3020892edf3a7e9f461b57309911d96fa2702320a16d

Observation ede787cc-e818-4db0-be20-9d383381e5f6 · inbound

Investigating Training Data Detection in AI Coders cites this paper.

Investigating Training Data Detection in AI Coders Investigating Data Contamination for Pre-training Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:53.769972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:53.769972Z digest=sha256:ae87d914d1b96793996bcb3f8c65d789f48326b1d13b4868b983821f81334bcc

Observation bd36c6b1-c8ed-4339-b358-6869b3900492 · inbound

Dataset Watermarking for Closed LLMs with Provable Detection cites this paper.

Dataset Watermarking for Closed LLMs with Provable Detection Investigating Data Contamination for Pre-training Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:57.277115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T00:53:42.185498Z digest=sha256:e23c42415e622d688afdba8a3f56915213e68004c127b97e69e5bc1ec56e9deb

Observation b54ef219-da74-42cc-9569-2fe133ae6f25 · inbound

LLM Benchmark Datasets Should Be Contamination-Resistant cites this paper.

LLM Benchmark Datasets Should Be Contamination-Resistant Investigating Data Contamination for Pre-training Language Models

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T07:23:07.050056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T07:19:50.354875Z digest=sha256:6ae4caad2ed874395ddb9e4db8573083c1065d0e196f4d0084420d330e0cae2c

Observation 5377c182-e0ca-4f62-8394-9bcfe8a2581a · inbound

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness cites this paper.

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness Investigating Data Contamination for Pre-training Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:24.421044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:37:01.537128Z digest=sha256:0122e701c43018cf8223f36dc3267dfefa2177b716aa2bb6188e50209f5e6237

Observation 20189093-0b97-45f4-8f59-facd5b977006 · inbound

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models cites this paper.

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models Investigating Data Contamination for Pre-training Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:04:38.792275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T12:00:05.807004Z digest=sha256:895089a86759cfb976b8a03c5ea60611514fe421017d1da1bcd377d686b6238c

Observation f0e50035-da39-45b2-a795-5de66d88e973 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Investigating Data Contamination for Pre-training Language Models

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.570680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:72c02b1ec015823bb27dc82208022b86e018f0ed0ff8e08e76ddb9ca1068dd6a

Observation d690820f-a685-4ea9-92b7-ecfc5e511bf7 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Investigating Data Contamination for Pre-training Language Models

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.169298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:aa21bb3628a4ff66d027e7254314706861fc7900ad013f76edf9b1054dc568cc