KDS measures LLM benchmark contamination by computing the divergence between kernel similarity matrices of sample embeddings before and after fine-tuning, and it correlates near-perfectly with contamination fraction in controlled tests.
Membership inference attacks against language models via neighbourhood comparison
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence
KDS measures LLM benchmark contamination by computing the divergence between kernel similarity matrices of sample embeddings before and after fine-tuning, and it correlates near-perfectly with contamination fraction in controlled tests.