Pith. sign in

REVIEW 3 cited by

Task Contamination: Language Models May Not Be Few-Shot Anymore

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.16337 v1 pith:UDBRALD4 submitted 2023-12-26 cs.CL

classification cs.CL
keywords llmsfew-shottaskcontaminationzero-shotdatadatasetsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) offer impressive performance in various zero-shot and few-shot tasks. However, their success in zero-shot and few-shot settings may be affected by task contamination, a potential limitation that has not been thoroughly examined. This paper investigates how zero-shot and few-shot performance of LLMs has changed chronologically over time. Utilizing GPT-3 series models and several other recent open-sourced LLMs, and controlling for dataset difficulty, we find that on datasets released before the LLM training data creation date, LLMs perform surprisingly better than on datasets released after. This strongly indicates that, for many LLMs, there exists task contamination on zero-shot and few-shot evaluation for datasets released prior to the LLMs' training data creation date. Additionally, we utilize training data inspection, task example extraction, and a membership inference attack, which reveal further evidence of task contamination. Importantly, we find that for classification tasks with no possibility of task contamination, LLMs rarely demonstrate statistically significant improvements over simple majority baselines, in both zero and few-shot settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 6 citations worldwide. Full citation record

  1. CODECLEANER: Elevating Standards with A Robust Data Contamination Mitigation Toolkit

    cs.SE 2024-11 conditional novelty 5.0 of 10

    A toolkit of 11 code refactoring operators reduces n-gram overlap with training corpora by up to 65 percentage points, though this drop is partly by construction and is not tied to downstream task performance.

  2. Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects

    cs.CY 2025-05 conditional novelty 4.0 of 10

    A position paper argues that understanding AI's second-order effects requires moving from static benchmarks to an ecosystem of field testing, red teaming, and contextual evaluation.

  3. Machines of Meaning

    cs.AI 2024-12 unverdicted novelty 3.0 of 10

    The paper offers a philosophical definition of meaning for AI systems and argues current language models are already 'machines of meaning' but are limited by fixed vocabularies and full-distribution outputs.

Pith tools