Pith. sign in

REVIEW 16 cited by

Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.14749 v4 pith:REZHTCQR submitted 2021-03-26 stat.ML cs.AIcs.LG

classification stat.MLcs.AIcs.LG
keywords errorstestdatasetslabelsetsacrosslabeledlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We identify label errors in the test sets of 10 of the most commonly-used computer vision, natural language, and audio datasets, and subsequently study the potential for these label errors to affect benchmark results. Errors in test sets are numerous and widespread: we estimate an average of at least 3.3% errors across the 10 datasets, where for example label errors comprise at least 6% of the ImageNet validation set. Putative label errors are identified using confident learning algorithms and then human-validated via crowdsourcing (51% of the algorithmically-flagged candidates are indeed erroneously labeled, on average across the datasets). Traditionally, machine learning practitioners choose which model to deploy based on test accuracy - our findings advise caution here, proposing that judging models over correctly labeled test sets may be more useful, especially for noisy real-world datasets. Surprisingly, we find that lower capacity models may be practically more useful than higher capacity models in real-world datasets with high proportions of erroneously labeled data. For example, on ImageNet with corrected labels: ResNet-18 outperforms ResNet-50 if the prevalence of originally mislabeled test examples increases by just 6%. On CIFAR-10 with corrected labels: VGG-11 outperforms VGG-19 if the prevalence of originally mislabeled test examples increases by just 5%. Test set errors across the 10 datasets can be viewed at https://labelerrors.com and all label errors can be reproduced by https://github.com/cleanlab/label-errors.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning

    cs.AI 2025-11 unverdicted novelty 7.0 of 10

    DecompSR is a large, symbolically verified benchmark dataset and generation framework that independently varies productivity, substitutivity, overgeneralisation, and systematicity to probe compositional multihop spati...

  2. A Novel Method to Evaluate Models on Unreliable, Noisy and Inconsistent Labels: Adaptive Resolution Label Aggregation (ARLA)

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Adaptive Resolution Label Aggregation (ARLA) coarsens label and prediction together at chosen subpatch size and sensitivity so evaluation metrics better reflect true model error on noisy segmentation labels.

  3. Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A ground-truth-first synthetic memory benchmark shows that agent-memory architecture rankings invert with history length: short-horizon leaders lose at nine weeks.

  4. Representation Unlearning: Forgetting through Information Compression

    cs.LG 2026-01 conditional novelty 6.0 of 10

    Representation Unlearning removes the influence of specific training samples by learning a lightweight transformation over the model's penultimate-layer representations, guided by information-bottleneck variational bounds.

  5. GFLC: Graph-based Fairness-aware Label Correction for Fair Classification

    cs.LG 2025-06 conditional novelty 6.0 of 10

    GFLC is a new label-correction method that uses confidence scores, graph curvature, and demographic parity to improve both accuracy and fairness under group-dependent label noise.

  6. ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    ZeroBench is a hand-built 100-question visual reasoning benchmark, adversarially filtered so every evaluated frontier LMM scored 0% at release.

  7. BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges

    cs.LG 2026-06 conditional novelty 5.0 of 10

    BACON calibrates multiple AI judges against a small human-labeled sample, then uses cross-fitted outcome models and augmented estimating equations to produce calibrated summary estimates and item-level surrogate scores.

  8. Noise is not always detrimental: the capacity of quantum batteries is enhanced in black holes

    quant-ph 2026-04 unverdicted novelty 5.0 of 10

    Hawking radiation is claimed to enhance quantum battery capacity for bipartite mixed states, while environmental noise generally degrades it in type-dependent ways.

  9. PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic

    cs.AI 2025-11 conditional novelty 5.0 of 10

    PaTAS propagates Subjective Logic trust opinions through every neuron of a network and updates parameter trust from gradient evidence, yielding per-prediction trust scores intended to flag poisoned or low-reliability inputs.

  10. When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification

    cs.CV 2025-05 conditional novelty 5.0 of 10

    REVEAL ensembles four VLMs and label-noise detectors to detect and correct noisy and missing labels in six image classification test sets, reporting high agreement with human annotations.

  11. Image Recognition with Vision and Language Embeddings of VLMs

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A benchmark of dual-encoder VLMs finds text and image embeddings give complementary class accuracy, and a per-class precision fusion rule adds about 0.4% accuracy over either alone on ImageNet.

  12. Advancing Mental Disorder Detection: A Comparative Evaluation of Transformer and LSTM Architectures on Social Media

    cs.CL 2025-07 conditional novelty 4.0 of 10

    On Reddit posts labeled by subreddit membership, transformer models, led by RoBERTa, reach 99.5% F1, far above LSTM baselines, but the proxy labels may inflate the apparent detection ability.

  13. Multimodal-Guided Dynamic Dataset Pruning for Robust and Efficient Data-Centric Learning

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A dynamic pruning method scores each sample by combining task loss with CLIP image-text similarity and selects samples near the median score each epoch.

  14. First-of-its-kind AI model for bioacoustic detection using a lightweight associative memory Hopfield neural network

    cs.LG 2025-07 reject novelty 4.0 of 10

    A Hopfield neural network trained on two bat calls classifies 10,384 recordings in 5.4 seconds with claimed accuracy up to 80%, though the headline numbers hinge on removing ambiguous calls.

  15. Machine Unlearning for Robust DNNs: Attribution-Guided Partitioning and Neuron Pruning in Noisy Environments

    cs.LG 2025-06 reject novelty 4.0 of 10

    The paper proposes attribution-guided data partitioning plus regression-based neuron pruning and fine-tuning for noisy training data, but the headline label-noise result is contradicted by the feature-noise-only experiments.

  16. The Achilles Heel of AI: Fundamentals of Risk-Aware Training Data for High-Consequence Models

    cs.LG 2025-05 reject novelty 3.0 of 10

    The paper claims that curated 20 to 40 percent subsets of training data can match full-data models and that shared label errors can inflate validation scores.

Pith tools