REVIEW 4 cited by
Memorization vs. Generalization: Quantifying Data Leakage in NLP Performance Evaluation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Public datasets are often used to evaluate the efficacy and generalizability of state-of-the-art methods for many tasks in natural language processing (NLP). However, the presence of overlap between the train and test datasets can lead to inflated results, inadvertently evaluating the model's ability to memorize and interpreting it as the ability to generalize. In addition, such data sets may not provide an effective indicator of the performance of these methods in real world scenarios. We identify leakage of training data into test data on several publicly available datasets used to evaluate NLP tasks, including named entity recognition and relation extraction, and study them to assess the impact of that leakage on the model's ability to memorize versus generalize.
Forward citations
Cited by 4 Pith papers
-
Rethinking the Capability of Fine-Tuned Language Models for Automated Vulnerability Repair
Fine-tuned AVR models memorize overlapping training data, so reported repair rates fall from ~20% to ~5% on non-overlapping splits, and match-based metrics misjudge true fixes.
-
Neuron-Level Differentiation of Memorization and Generalization in Large Language Models
Memorization and generalization in LLMs are associated with distinct neurons, and steering those neurons at inference time can switch a model between the two behaviors.
-
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities
ONEBench treats each benchmark sample as a voter in a Plackett-Luce aggregation, enabling open-ended, capability-specific, and incomplete-data model rankings.
-
Lightweight Person-Place Relation Extraction from Historical Newspapers with Dependency Graphs and Proximity Features
A feature-based, no-pretrained-LM classifier for person–place relations in historical newspapers reaches 0.5142 macro recall on HIPE-2026, with minimum character distance dominating the signal and document-grouped CV ...
Discussion (0). Continue with ORCID to comment.