REVIEW 2 cited by
Improving Adversarial Robustness of Ensembles with Diversity Training
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Deep Neural Networks are vulnerable to adversarial attacks even in settings where the attacker has no direct access to the model being attacked. Such attacks usually rely on the principle of transferability, whereby an attack crafted on a surrogate model tends to transfer to the target model. We show that an ensemble of models with misaligned loss gradients can provide an effective defense against transfer-based attacks. Our key insight is that an adversarial example is less likely to fool multiple models in the ensemble if their loss functions do not increase in a correlated fashion. To this end, we propose Diversity Training, a novel method to train an ensemble of models with uncorrelated loss functions. We show that our method significantly improves the adversarial robustness of ensembles and can also be combined with existing methods to create a stronger defense.
Forward citations
Cited by 2 Pith papers
-
Learning from Peers: Collaborative Ensemble Adversarial Training
CEAT reweights each sub-model's training samples using peer prediction disparities, improving ensemble adversarial robustness on CIFAR-10/100 by several points over prior EAT methods.
-
Theoretical Analysis of Relative Errors in Gradient Computations for Adversarial Attacks with CE Loss
T-MIFPE adaptively rescales logits with a theoretically motivated t* per attack phase to reduce floating-point gradient errors, edging out MIFPE in PGD robustness evaluation.
Discussion (0). Continue with ORCID to comment.