Pith. sign in

Paper Citation Record · LEDGER

An Open Source Data Contamination Report for Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2310.17589.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.17589 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:15:41.599246Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ef0597a4-e019-459c-a4e0-68f24c942d05 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey An Open Source Data Contamination Report for Large Language Models

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.032900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:873d66a46f674bac6d53cb344e98f13d52dac99d98d3bf1c2321d69baa5a4cf6

Observation 557de4d9-46fa-4952-9021-f3fd66020da8 · inbound

MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters cites this paper.

MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters An Open Source Data Contamination Report for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T05:15:41.599246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:15:41.599246Z digest=sha256:ae572365f62a8d5870f40ae42bf173f2154593d41f4f6d5b4d529552cde51044

Observation 2f5b0c50-5873-4c8b-ae0f-29c48757c09d · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective An Open Source Data Contamination Report for Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.421100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.421100Z digest=sha256:371609e325326e9ef6c385e7a7b0ca1d7a8918a4fa45a66277dc9ccfa8b8a3b4

Observation ba7e9401-0156-4f55-a3d7-8fd51bc3a93a · inbound

Answer-Centric or Reasoning-Driven? Uncovering the Latent Memory Anchor in LLMs cites this paper.

Answer-Centric or Reasoning-Driven? Uncovering the Latent Memory Anchor in LLMs An Open Source Data Contamination Report for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:33:55.250770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:33:55.250770Z digest=sha256:b3f55dad3ed0274a7c1d62c151319cfd7bc4a0a1cd32c349632efcfdd23e51b6

Observation 57deabd8-a779-4d98-9d53-f4fc10f32e7a · inbound

Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting cites this paper.

Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting An Open Source Data Contamination Report for Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T12:45:26.141280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:45:26.141280Z digest=sha256:c07b370b87b0fa96bf8cb1c8ce2faac51e9927066f8275ab37d222483b922b4a

Observation 6df8c3ff-4aa5-4343-a916-55e5e70f6222 · inbound

Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation cites this paper.

Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation An Open Source Data Contamination Report for Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:29.830869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T08:12:02.449352Z digest=sha256:52116b8a71cf88e094338ffc0f26746fe54ada37b5adece2bae5fce649011a9c

Observation b387dd4a-0118-4ea0-b182-b6d907e860a2 · inbound

SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks cites this paper.

SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks An Open Source Data Contamination Report for Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T04:45:21.214169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T04:40:42.747298Z digest=sha256:661249b849ea83897b568ba3fa9942bed3a23fa103398cb3ab958e1e800d3c42

Observation a5ac612e-1d6a-4cb3-b8c8-c660e783d84f · inbound

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications cites this paper.

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications An Open Source Data Contamination Report for Large Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:56.622082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:20:16.735285Z digest=sha256:6e7310678b660671d1aeb8e923d4492ca8a410ddc5d995113013e108c60e1a3f

Observation b70d2bd0-9bc7-4144-8574-e3b888745eb7 · inbound

Structure Before Collapse: Transient semantic geometry in next-token prediction cites this paper.

Structure Before Collapse: Transient semantic geometry in next-token prediction An Open Source Data Contamination Report for Large Language Models

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:29:51.068669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T05:14:07.208255Z digest=sha256:3628c895df4833d76b8f184579e22f7264237da9bcf55678f10d8f98c5ef81ad