Pith. sign in

Paper Citation Record · LEDGER

Proving Test Set Contamination in Black Box Language Models

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2310.17623.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.17623 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:04:55.376776Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aac2b93c-33af-42fd-9be3-c62381b806c8 · inbound

Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders cites this paper.

Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders Proving Test Set Contamination in Black Box Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-24T03:05:56.957841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-24T03:03:51.053556Z digest=sha256:0eb8d75b057784c1cc6eabc119229dd936bf6d2dda1fc78d82da7b7941102aa9

Observation 9cfa3985-199b-47ea-85c2-b13326b0d4d0 · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark Proving Test Set Contamination in Black Box Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:48:26.469976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:2eebdaeee571c07817a37fc187c5832da8f1e009ec072a2582d638b91a2022f7

Observation 46758629-71e8-4b08-a4df-f06569258c15 · inbound

The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance? cites this paper.

The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance? Proving Test Set Contamination in Black Box Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:55.376776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:55.376776Z digest=sha256:c665a27fc216c53b2872a3fa1aa767eec7ca35a6969e94e9aca4425c0cd0e2bb

Observation 09bf5eeb-7e7e-414f-a8c6-e6df7b607cf0 · inbound

LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage cites this paper.

LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage Proving Test Set Contamination in Black Box Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:30:10.273799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:30:10.273799Z digest=sha256:7625eea5a332dc30ec314e7b38f0976568fc9f73abde8aa1e94fcbad054882b4

Observation ff05ac5e-4b84-4abe-bf3a-d3bf79d327ac · inbound

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge cites this paper.

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge Proving Test Set Contamination in Black Box Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:58:02.141431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:58:02.141431Z digest=sha256:1fcf7a16c7fe0ba225f3fbfacf5a871c5dc83f316d9b5dce67cb1a88f7ec981f

Observation c2f3c6fd-1ea6-42c7-93a5-4daaccee9f63 · inbound

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation cites this paper.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Proving Test Set Contamination in Black Box Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.117286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.117286Z digest=sha256:44c43ecf36b9691bfc63fb0d3692449cf0c57c87b796052e6ea854c1581fac84

Observation 5c32f891-3e89-45c4-b64b-937f3bcb1067 · inbound

Spectral Journey: How Transformers Predict the Shortest Path cites this paper.

Spectral Journey: How Transformers Predict the Shortest Path Proving Test Set Contamination in Black Box Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:44:10.838810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:44:10.838810Z digest=sha256:0dd99bab42165220b78aeb861e54e3d479be1a02077692973778546bc91f15f2

Observation 1d03e316-85ff-4de6-9ea5-989c62c52be0 · inbound

Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning cites this paper.

Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning Proving Test Set Contamination in Black Box Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:16:19.161532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:16:19.161532Z digest=sha256:c9596b59a7ed46f007ecf1c94b1e8a96076d6b83b702b05021329317afbda7bd

Observation 1ebef60e-02ce-4974-88e8-e1b3732d9cd9 · inbound

AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI cites this paper.

AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI Proving Test Set Contamination in Black Box Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:51.772651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:51.772651Z digest=sha256:17259ad84fa13a49276e62ee4a1a79d14cb7bfc76f879e2f6b619f422ff199c3

Observation 9d5f236f-2fa0-4c09-9b91-3952121237e2 · inbound

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints cites this paper.

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints Proving Test Set Contamination in Black Box Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:15:29.379804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T14:12:45.438246Z digest=sha256:959570d5a55345ebb88a175de6f10d57e601ef329b176e4abf5d0b13f9c1446a

Observation 766a0e8b-f878-4ff2-af4c-2d15b870a944 · inbound

SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks cites this paper.

SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks Proving Test Set Contamination in Black Box Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:45:21.220860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T04:40:42.747298Z digest=sha256:6ad1c16a0cbdaf84e7c9f7ede735f43e0be63313390a7f889e32b075120a21dc

Observation 6c932fc7-89e8-471e-9157-f2d3fe42e545 · inbound

Dataset Watermarking for Closed LLMs with Provable Detection cites this paper.

Dataset Watermarking for Closed LLMs with Provable Detection Proving Test Set Contamination in Black Box Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:57.294124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T00:53:42.185498Z digest=sha256:1afa32a4d7c2126f1f9e008756418ba4b3f8e33cd724a598065459a54845cdec

Observation d31c1837-ca49-4caf-a2a3-d9d96da1dc5e · inbound

Data Contamination in Neural Hieroglyphic Translation: A Reproducibility Study cites this paper.

Data Contamination in Neural Hieroglyphic Translation: A Reproducibility Study Proving Test Set Contamination in Black Box Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:10:55.459234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-11T01:55:16.674529Z digest=sha256:baca9e892924e6af7d2e26d6528f8429176b8bd48af670ad9e251a722f672e8d

Observation 824532d7-400d-40ed-8919-f33015637cc8 · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Proving Test Set Contamination in Black Box Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.905072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:0cf2be1dd2fba5b823840ebce7f71b749f087253fb910431c37fb1c2266fe5f3

Observation 94f875a4-f10d-42f6-9998-136ddd67eb62 · inbound

Interactive Evaluation Requires a Design Science cites this paper.

Interactive Evaluation Requires a Design Science Proving Test Set Contamination in Black Box Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:14.121307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T10:55:08.135630Z digest=sha256:7da65e2ec1e88825312a65d1841cc154cd8a70d91e94b34eb30f3cd5f1eb6eda

Observation 27fcf841-9a4e-44ef-8d5b-7eda51f7cc89 · inbound

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications cites this paper.

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications Proving Test Set Contamination in Black Box Language Models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:24:56.592669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T17:20:16.735285Z digest=sha256:99c5e2b890e1fa970dbf21822156bc3eb16ff353f27ae3f4567f17b1eabcd0f4

Observation 114e7f20-0f73-4110-a023-6717d66ddf4c · inbound

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models cites this paper.

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models Proving Test Set Contamination in Black Box Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.345069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T20:02:50.169589Z digest=sha256:63e211d30e937cdaf47811c80d8e09beac33d4e070c3a84b1d0d355c3e9274c9

Observation f10255f9-bef7-4875-8408-e3c823662b4e · inbound

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores cites this paper.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Proving Test Set Contamination in Black Box Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.382539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.382539Z digest=sha256:b88bf24a92bf55752407bb649164be87a34989cee1cc1d5003daa4dd014ce2e0