Pith. sign in

REVIEW 3 cited by

Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.04811 v1 pith:FICSFSHM submitted 2024-03-06 cs.SE cs.CLcs.LG

classification cs.SEcs.CLcs.LG
keywords generationbenchmarkscodecontaminationdatalanguagemodelsthere
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While large language models have achieved remarkable performance on various code generation benchmarks, there have been growing concerns regarding potential contamination of these benchmarks as they may be leaked into pretraining and finetuning data. While recent work has investigated contamination in natural language generation and understanding tasks, there has been less extensive research into how data contamination impacts the evaluation of code generation, which is critical for understanding the robustness and reliability of LLMs in programming contexts. In this work, we perform a comprehensive study of data contamination of popular code generation benchmarks, and precisely quantify their overlap with pretraining corpus through both surface-level and semantic-level matching. In our experiments, we show that there are substantial overlap between popular code generation benchmarks and open training corpus, and models perform significantly better on the subset of the benchmarks where similar solutions are seen during training. We also conduct extensive analysis on the factors that affects model memorization and generalization, such as model size, problem difficulty, and question length. We release all resulting files from our matching pipeline for future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores

    cs.LG 2026-08 conditional novelty 6.0 of 10

    The standard pre/post cutoff check cannot separate memorization from recency, and a single external reference is needed to measure and adjust for temporal leakage.

  2. Unbiased Evaluation of Large Language Models from a Causal Perspective

    cs.AI 2025-02 reject novelty 5.0 of 10

    The paper argues that perturbing benchmark questions with rule-based interventions gives a less contaminated, more interpretable evaluation of LLMs than static benchmarks or agent-generated questions.

  3. Evaluating and Improving Large Language Models for Competitive Program Generation

    cs.SI 2025-06 conditional novelty 4.0 of 10

    DeepSeek-R1 solves only 5 of 80 recent ICPC/CCPC competitive programming problems with a basic prompt, and 46 of 80 after a taxonomy-guided repair and regeneration pipeline.

Pith tools