REVIEW 7 cited by
Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent statements about the impressive capabilities of large language models (LLMs) are usually supported by evaluating on open-access benchmarks. Considering the vast size and wide-ranging sources of LLMs' training data, it could explicitly or implicitly include test data, leading to LLMs being more susceptible to data contamination. However, due to the opacity of training data, the black-box access of models, and the rapid growth of synthetic training data, detecting and mitigating data contamination for LLMs faces significant challenges. In this paper, we propose CDD, which stands for Contamination Detection via output Distribution for LLMs. CDD necessitates only the sampled texts to detect data contamination, by identifying the peakedness of LLM's output distribution. To mitigate the impact of data contamination in evaluation, we also present TED: Trustworthy Evaluation via output Distribution, based on the correction of LLM's output distribution. To facilitate this study, we introduce two benchmarks, i.e., DetCon and ComiEval, for data contamination detection and contamination mitigation evaluation tasks. Extensive experimental results show that CDD achieves the average relative improvements of 21.8\%-30.2\% over other contamination detection approaches in terms of Accuracy, F1 Score, and AUC metrics, and can effectively detect implicit contamination. TED substantially mitigates performance improvements up to 66.9\% attributed to data contamination across various contamination setups. In real-world applications, we reveal that ChatGPT exhibits a high potential to suffer from data contamination on HumanEval benchmark.
Forward citations
Cited by 7 Pith papers
-
TLA+-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA Specification Generation
An exact model-checker oracle for NL-to-TLA+ still yields an elevenfold “correctness envelope” (18.7%–1.7%) once interface supply and vacuity screens are made explicit.
-
InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy
InvestPhilBench is a new multi-layer benchmark for LLM procedural reasoning in investment philosophy, with BASP metrics showing composite scores saturate while gate reconstruction accuracy reveals procedural deficits.
-
Rethinking Memorization Measures and their Implications in Large Language Models
Contextual memorization, defined by comparing a string's training loss against the best loss without training on that string, is stricter than counterfactual memorization and suggests that zero-memorization optimal le...
-
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
A TEE-based protocol for cryptographically verifiable AI safety benchmark results, demonstrated on Llama-3.1 with AWS Nitro Enclaves.
-
Search-Time Data Contamination
Web-searching AI agents sometimes find the test dataset's answer key on HuggingFace during evaluation, inflating their scores on roughly 3% of benchmark questions.
-
SciDA: Scientific Dynamic Assessor of LLMs
SciDA is a dynamically initialized, multi-discipline olympiad benchmark that shows LLMs perform substantially worse when problem variables are randomized, which the authors attribute to memorization of fixed numerical...
-
AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic Reports
An LLM-generated QA benchmark over World Bank African economic reports is evaluated on a small sample and found hard, but the abstract and body report conflicting dataset sizes, model names, and scores.
Discussion (0). Sign in to comment.