Pith. sign in

REVIEW 3 cited by

Overfitting in adversarially robust deep learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.11569 v2 pith:SW7CYWA3 submitted 2020-02-26 cs.LG stat.ML

classification cs.LGstat.ML
keywords overfittingadversariallydeeprobusttraininglearningperformancetrained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

It is common practice in deep learning to use overparameterized networks and train for as long as possible; there are numerous studies that show, both theoretically and empirically, that such practices surprisingly do not unduly harm the generalization performance of the classifier. In this paper, we empirically study this phenomenon in the setting of adversarially trained deep networks, which are trained to minimize the loss under worst-case adversarial perturbations. We find that overfitting to the training set does in fact harm robust performance to a very large degree in adversarially robust training across multiple datasets (SVHN, CIFAR-10, CIFAR-100, and ImageNet) and perturbation models ($\ell_\infty$ and $\ell_2$). Based upon this observed effect, we show that the performance gains of virtually all recent algorithmic improvements upon adversarial training can be matched by simply using early stopping. We also show that effects such as the double descent curve do still occur in adversarially trained models, yet fail to explain the observed overfitting. Finally, we study several classical and modern deep learning remedies for overfitting, including regularization and data augmentation, and find that no approach in isolation improves significantly upon the gains achieved by early stopping. All code for reproducing the experiments as well as pretrained model weights and training logs can be found at https://github.com/locuslab/robust_overfitting.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adversarial Examples Are Not Bugs, They Are Superposition

    cs.LG 2025-08 unverdicted novelty 6.0 of 10

    The paper argues that adversarial examples arise from superposition, and shows that changing superposition changes robustness and vice versa in toy models and ResNet18.

  2. HEM: a margin-based loss for visual categorisation tasks

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A new margin-based loss, HEM, trains image classifiers that are more robust to unknown and adversarial inputs and better at continual learning and segmentation than cross-entropy-trained models.

  3. Towards Fair Class-wise Robustness: Class Optimal Distribution Adversarial Training

    cs.LG 2025-01 conditional novelty 3.0 of 10

    A chi-squared distributionally robust optimization reweighting scheme improves worst-class robustness in adversarially trained image classifiers.

Pith tools