Pith. sign in

REVIEW 2 cited by

The Limitations of Adversarial Training and the Blind-Spot Attack

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1901.04684 v1 pith:BDEEWNZF submitted 2019-01-15 stat.ML cs.CRcs.CVcs.LG

classification stat.MLcs.CRcs.CVcs.LG
keywords trainingadversarialdatablind-spotsmanifoldtestexamplesattack
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The adversarial training procedure proposed by Madry et al. (2018) is one of the most effective methods to defend against adversarial examples in deep neural networks (DNNs). In our paper, we shed some lights on the practicality and the hardness of adversarial training by showing that the effectiveness (robustness on test set) of adversarial training has a strong correlation with the distance between a test point and the manifold of training data embedded by the network. Test examples that are relatively far away from this manifold are more likely to be vulnerable to adversarial attacks. Consequentially, an adversarial training based defense is susceptible to a new class of attacks, the "blind-spot attack", where the input images reside in "blind-spots" (low density regions) of the empirical distribution of training data but is still on the ground-truth data manifold. For MNIST, we found that these blind-spots can be easily found by simply scaling and shifting image pixel values. Most importantly, for large datasets with high dimensional and complex data manifold (CIFAR, ImageNet, etc), the existence of blind-spots in adversarial training makes defending on any valid test examples difficult due to the curse of dimensionality and the scarcity of training data. Additionally, we find that blind-spots also exist on provable defenses including (Wong & Kolter, 2018) and (Sinha et al., 2018) because these trainable robustness certificates can only be practically optimized on a limited set of training data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Breakable Is Privacy: Probing and Resisting Model Inversion Attacks in Collaborative Inference

    cs.CR 2025-01 conditional novelty 5.0 of 10

    A mutual-information-based criterion, Dmia, predicts model inversion attack difficulty in collaborative inference, and the SiftFunnel defense suppresses the criterion's factors to raise reconstruction error with only ...

  2. Sign-Symmetry Learning Rules are Robust Fine-Tuners

    cs.LG 2025-02 reject novelty 4.0 of 10

    Fine-tuning with sign-symmetry rules preserves accuracy while resisting white-box adversarial attacks, but the robustness appears to be gradient masking rather than genuine defense.

Pith tools