Pith. sign in

Paper Citation Record · LEDGER

Proving Test Set Contamination in Black Box Language Models

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2310.17623.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.17623 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:04:55.376776Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aac2b93c-33af-42fd-9be3-c62381b806c8 · inbound

Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders cites this paper.

Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders Proving Test Set Contamination in Black Box Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-24T03:05:56.957841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-24T03:03:51.053556Z digest=sha256:9b336e28c45cb7d2907710957524b8671c8d333a8dab5af781fbb3696abf740c

Observation 9cfa3985-199b-47ea-85c2-b13326b0d4d0 · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark Proving Test Set Contamination in Black Box Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:48:26.469976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:8267a327339cc99b9cad497e15ec2066c42bc2be2555b9d110813af3ccb78c6c

Observation 46758629-71e8-4b08-a4df-f06569258c15 · inbound

The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance? cites this paper.

The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance? Proving Test Set Contamination in Black Box Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:55.376776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:55.376776Z digest=sha256:e4fd86ae2d76f89bf9962940b302597fa9d2840f31a53345edc1cd2feff719bf

Observation 09bf5eeb-7e7e-414f-a8c6-e6df7b607cf0 · inbound

LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage cites this paper.

LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage Proving Test Set Contamination in Black Box Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:30:10.273799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:30:10.273799Z digest=sha256:4a48cf7fbb1e105eef30d09c0e385f0ed925ae110577376f0428d11c2b973f0b

Observation ff05ac5e-4b84-4abe-bf3a-d3bf79d327ac · inbound

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge cites this paper.

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge Proving Test Set Contamination in Black Box Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:58:02.141431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:58:02.141431Z digest=sha256:6428408f3e948c9b7d8c4913b0f93045a85f511a60a4e935d022286dcdda2d4f

Observation c2f3c6fd-1ea6-42c7-93a5-4daaccee9f63 · inbound

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation cites this paper.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Proving Test Set Contamination in Black Box Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.117286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.117286Z digest=sha256:44c43ecf36b9691bfc63fb0d3692449cf0c57c87b796052e6ea854c1581fac84

Observation 5c32f891-3e89-45c4-b64b-937f3bcb1067 · inbound

Spectral Journey: How Transformers Predict the Shortest Path cites this paper.

Spectral Journey: How Transformers Predict the Shortest Path Proving Test Set Contamination in Black Box Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:44:10.838810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:44:10.838810Z digest=sha256:0dd99bab42165220b78aeb861e54e3d479be1a02077692973778546bc91f15f2

Observation 1d03e316-85ff-4de6-9ea5-989c62c52be0 · inbound

Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning cites this paper.

Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning Proving Test Set Contamination in Black Box Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:16:19.161532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:16:19.161532Z digest=sha256:c9596b59a7ed46f007ecf1c94b1e8a96076d6b83b702b05021329317afbda7bd

Observation 1ebef60e-02ce-4974-88e8-e1b3732d9cd9 · inbound

AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI cites this paper.

AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI Proving Test Set Contamination in Black Box Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:46:51.772651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:46:51.772651Z digest=sha256:17259ad84fa13a49276e62ee4a1a79d14cb7bfc76f879e2f6b619f422ff199c3

Observation 9d5f236f-2fa0-4c09-9b91-3952121237e2 · inbound

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints cites this paper.

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints Proving Test Set Contamination in Black Box Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:15:29.379804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T14:12:45.438246Z digest=sha256:b535f9ec9319563d33a0713d2f2de81c7854d35832fe63490e09a048dda070bb

Observation 766a0e8b-f878-4ff2-af4c-2d15b870a944 · inbound

SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks cites this paper.

SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks Proving Test Set Contamination in Black Box Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:45:21.220860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T04:40:42.747298Z digest=sha256:32f5458a964fcfe858f8eaa8cf847d275eba5bcd0f100dc3209c83f799a2fb42

Observation 6c932fc7-89e8-471e-9157-f2d3fe42e545 · inbound

Dataset Watermarking for Closed LLMs with Provable Detection cites this paper.

Dataset Watermarking for Closed LLMs with Provable Detection Proving Test Set Contamination in Black Box Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:57.294124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T00:53:42.185498Z digest=sha256:a97e1b865a64c1c74ce059b0b5db6d5eaea77aeabc25a75f480e73264443367a

Observation d31c1837-ca49-4caf-a2a3-d9d96da1dc5e · inbound

Data Contamination in Neural Hieroglyphic Translation: A Reproducibility Study cites this paper.

Data Contamination in Neural Hieroglyphic Translation: A Reproducibility Study Proving Test Set Contamination in Black Box Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:10:55.459234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T01:55:16.674529Z digest=sha256:d49e45a3647b58989975e4ea6f7e98827c2e808fe28c731b895bc1c8a0a455da

Observation 824532d7-400d-40ed-8919-f33015637cc8 · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Proving Test Set Contamination in Black Box Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.905072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:3f99b2fa6c1c1575d3e0d635c34f0dcdb521615758afa056e3e50af15e78af3d

Observation 94f875a4-f10d-42f6-9998-136ddd67eb62 · inbound

Interactive Evaluation Requires a Design Science cites this paper.

Interactive Evaluation Requires a Design Science Proving Test Set Contamination in Black Box Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:14.121307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T10:55:08.135630Z digest=sha256:ab0bb874e17aa6eada9351d8159e56cc9a76392a2b3f6a49b6bb3254fff24811

Observation 27fcf841-9a4e-44ef-8d5b-7eda51f7cc89 · inbound

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications cites this paper.

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications Proving Test Set Contamination in Black Box Language Models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:24:56.592669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T17:20:16.735285Z digest=sha256:b405cfcdf639dbbade6addb89bc0fc3d55aad81369efcf9a60b63ac7cd085d4c

Observation 114e7f20-0f73-4110-a023-6717d66ddf4c · inbound

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models cites this paper.

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models Proving Test Set Contamination in Black Box Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.345069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T20:02:50.169589Z digest=sha256:08f6c3ba2e744f677fbacb2ac032740444fd505c5f8bb8b95b933b1cb5f04d2d

Observation f10255f9-bef7-4875-8408-e3c823662b4e · inbound

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores cites this paper.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Proving Test Set Contamination in Black Box Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.382539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.382539Z digest=sha256:b88bf24a92bf55752407bb649164be87a34989cee1cc1d5003daa4dd014ce2e0