Pith. sign in

REVIEW 1 cited by

Understanding the Detrimental Class-level Effects of Data Augmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.01764 v1 pith:U62JQZGK submitted 2023-12-07 cs.CV cs.LG

classification cs.CVcs.LG
keywords accuracyaugmentationclassclass-levelclassesperformanceunderstandingwhile
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data augmentation (DA) encodes invariance and provides implicit regularization critical to a model's performance in image classification tasks. However, while DA improves average accuracy, recent studies have shown that its impact can be highly class dependent: achieving optimal average accuracy comes at the cost of significantly hurting individual class accuracy by as much as 20% on ImageNet. There has been little progress in resolving class-level accuracy drops due to a limited understanding of these effects. In this work, we present a framework for understanding how DA interacts with class-level learning dynamics. Using higher-quality multi-label annotations on ImageNet, we systematically categorize the affected classes and find that the majority are inherently ambiguous, co-occur, or involve fine-grained distinctions, while DA controls the model's bias towards one of the closely related classes. While many of the previously reported performance drops are explained by multi-label annotations, our analysis of class confusions reveals other sources of accuracy degradation. We show that simple class-conditional augmentation strategies informed by our framework improve performance on the negatively affected classes.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. No Location Left Behind: Measuring and Improving the Fairness of Implicit Representations for Earth Data

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Spherical wavelet encodings reduce the performance gap on small and coastal landmasses that spherical harmonic and other location encodings exhibit in implicit neural representations of Earth data.

Pith tools