Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T18:11:35.098478Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 5 inbound Pith citation observations for arXiv:2502.00678.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T18:11:35.098478Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:35:11.856700Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T20:57:23.351502Z
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bb849fbc-970b-4e1b-ab44-338ff21a7591 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfa9bf10-aa11-43b1-a081-2afb92c7fc09 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence (Zhang et al., 2024b) Following the general guideline from Shi et al
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 879901bd-1977-4398-ada3-dc5b43f38fe9 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a46c49a-6714-4e3d-a5e7-131cff9d6377 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 887b68ba-8e61-4389-9e2b-5ef7506c35d2 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Mistral 7B
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0adbbb53-c35e-49db-8669-51a7c8fe38d5 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence LLM Dataset Inference: Did you train on my dataset?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09105b28-314f-49fe-a8a7-607fd01f4691 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Membership inference attacks against language models via neighbourhood comparison
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8405e8ad-d6e8-4f2f-ad6f-30c73754bd17 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f0a361d-ecfc-4827-802d-c2cc8e15f1bb · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Ml-leaks: Model and data independent membership inference attacks and defenses on machine learning models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 015c0602-a631-469f-a2ce-6142c3a11e43 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Membership inference attacks against machine learning models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 44cd797b-fed2-446f-bc3f-1e6a6442aea9 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Benchmark Data Contamination of Large Language Models: A Survey
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0d39c96-8140-4530-b269-887097046d2b · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef5ef6ad-6de3-4432-8b49-7de7271cd8cd · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Data Contamination Calibration for Black-box LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3fa586d9-50da-44e0-8a06-152d79c0206f · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Privacy risk in machine learning: Analyzing the connection to overfitting
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 90dc6e0b-1c27-47db-b0e2-33779ca6362a · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence 12 A.2 Baseline Definitions
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7cc0da30-7910-492c-8f4b-1c580a7f0c9c · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence The average Mean Absolute Percentage Error (MAPE) over 5 independent runs is calculated
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 30daf7d6-c258-4f91-91bc-5eba05eb9b59 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 2008
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92f9f1af-368d-4e0a-9966-cbe9e650b31a · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence E., Yu, L., and Wei, W
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8e4155d1-edb0-487b-ae2e-ba55cb706357 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Pacost: Paired confidence signifi- cance testing for benchmark contamination detection in large language models
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a7e9ed35-e4ff-4327-8cfb-c9b365c15e7a · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Recall: Membership inference via relative conditional log-likelihoods
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8b640970-a1ca-483a-85f8-e72e513ae93f · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Large sample analysis of the median heuristic
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5dd60fc-049e-4416-9e02-b4915a6c2127 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Membership inference attacks from first principles
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 50c0aa63-13be-43e9-b518-401304874202 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Blind Baselines Beat Membership Inference Attacks for Foundation Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 033cd38c-224c-43f2-9ddd-98674589c3f8 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57b9cc5d-0258-4efa-9f55-6067991adbed · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Do Membership Inference Attacks Work on Large Language Models?
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f43d9dd4-aadf-4745-abd3-8350063f0700 · outbound
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5730a9c4-46ad-4ad3-b5e8-92b0d06bd75d · inbound
Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de39f1bc-94b6-4e53-b903-ba3bd58af580 · inbound
The Economics of AI Training Data: A Research Agenda How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d26c43a-aedc-4344-8d11-af3667d41ce5 · inbound
When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a86d04d6-a9a5-46c5-93c1-0cab8b5a3bee · inbound
Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 115f0cf7-4813-46ac-9c84-79ba0c1a5787 · inbound
MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.