Pith. sign in

REVIEW 2 cited by

Label Noise Types and Their Effects on Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.10471 v1 pith:G5Z7P6BL submitted 2020-03-23 cs.CV

classification cs.CV
keywords labelnoiselabelslearningdatasetsnoisydeepproposed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent success of deep learning is mostly due to the availability of big datasets with clean annotations. However, gathering a cleanly annotated dataset is not always feasible due to practical challenges. As a result, label noise is a common problem in datasets, and numerous methods to train deep neural networks in the presence of noisy labels are proposed in the literature. These methods commonly use benchmark datasets with synthetic label noise on the training set. However, there are multiple types of label noise, and each of them has its own characteristic impact on learning. Since each work generates a different kind of label noise, it is problematic to test and compare those algorithms in the literature fairly. In this work, we provide a detailed analysis of the effects of different kinds of label noise on learning. Moreover, we propose a generic framework to generate feature-dependent label noise, which we show to be the most challenging case for learning. Our proposed method aims to emphasize similarities among data instances by sparsely distributing them in the feature domain. By this approach, samples that are more likely to be mislabeled are detected from their softmax probabilities, and their labels are flipped to the corresponding class. The proposed method can be applied to any clean dataset to synthesize feature-dependent noisy labels. For the ease of other researchers to test their algorithms with noisy labels, we share corrupted labels for the most commonly used benchmark datasets. Our code and generated noisy synthetic labels are available online.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Securing Contrastive mmWave-based Human Activity Recognition against Adversarial Label Flipping

    cs.CR 2026-08 conditional novelty 5.0 of 10

    Trajectory-aware label flipping attacks degrade contrastive mmWave human activity recognition, and a selective contrastive learning defense keeps accuracy above 90% even at 40% poisoned labels.

  2. Improved Stochastic Optimization of LogSumExp

    math.OC 2025-09 conditional novelty 5.0 of 10

    A rescaled SoftPlus family approximates LogSumExp with O(ρ) error, enabling stable stochastic optimization in entropic OT and KL-DRO.

Pith tools