Pith. sign in

REVIEW 4 cited by

Memorization vs. Generalization: Quantifying Data Leakage in NLP Performance Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.01818 v1 pith:72DQD437 submitted 2021-02-03 cs.CL cs.LG

classification cs.CLcs.LG
keywords dataabilitydatasetsleakageevaluategeneralizememorizemethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Public datasets are often used to evaluate the efficacy and generalizability of state-of-the-art methods for many tasks in natural language processing (NLP). However, the presence of overlap between the train and test datasets can lead to inflated results, inadvertently evaluating the model's ability to memorize and interpreting it as the ability to generalize. In addition, such data sets may not provide an effective indicator of the performance of these methods in real world scenarios. We identify leakage of training data into test data on several publicly available datasets used to evaluate NLP tasks, including named entity recognition and relation extraction, and study them to assess the impact of that leakage on the model's ability to memorize versus generalize.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rethinking the Capability of Fine-Tuned Language Models for Automated Vulnerability Repair

    cs.SE 2025-12 conditional novelty 6.0 of 10

    Fine-tuned AVR models memorize overlapping training data, so reported repair rates fall from ~20% to ~5% on non-overlapping splits, and match-based metrics misjudge true fixes.

  2. Neuron-Level Differentiation of Memorization and Generalization in Large Language Models

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Memorization and generalization in LLMs are associated with distinct neurons, and steering those neurons at inference time can switch a model between the two behaviors.

  3. ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities

    cs.LG 2024-12 conditional novelty 6.0 of 10

    ONEBench treats each benchmark sample as a voter in a Plackett-Luce aggregation, enabling open-ended, capability-specific, and incomplete-data model rankings.

  4. Lightweight Person-Place Relation Extraction from Historical Newspapers with Dependency Graphs and Proximity Features

    cs.CL 2026-07 conditional novelty 4.0 of 10

    A feature-based, no-pretrained-LM classifier for person–place relations in historical newspapers reaches 0.5142 macro recall on HIPE-2026, with minimum character distance dominating the signal and document-grouped CV ...

Pith tools