REVIEW 6 cited by
Stochastic Activation Pruning for Robust Adversarial Defense
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Neural networks are known to be vulnerable to adversarial examples. Carefully chosen perturbations to real images, while imperceptible to humans, induce misclassification and threaten the reliability of deep learning systems in the wild. To guard against adversarial examples, we take inspiration from game theory and cast the problem as a minimax zero-sum game between the adversary and the model. In general, for such games, the optimal strategy for both players requires a stochastic policy, also known as a mixed strategy. In this light, we propose Stochastic Activation Pruning (SAP), a mixed strategy for adversarial defense. SAP prunes a random subset of activations (preferentially pruning those with smaller magnitude) and scales up the survivors to compensate. We can apply SAP to pretrained networks, including adversarially trained models, without fine-tuning, providing robustness against adversarial examples. Experiments demonstrate that SAP confers robustness against attacks, increasing accuracy and preserving calibration.
Forward citations
Cited by 6 Pith papers
-
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
Randomizing decoding hyperparameters and system prompts per query lowers jailbreak attack success on five 7B LLMs, from 74% down to 0% in the best reported case, but the evaluation uses the same benchmark that selecte...
-
Improving Adversarial Robustness via Attention and Adversarial Logit Pairing
Aligning attention maps and logits between clean and adversarial images during training improves robustness accuracy over adversarial training on three small image datasets.
-
Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting
CAW adds a confidence-weighted KL loss and feature-alignment regularization to CLIP adversarial fine-tuning, raising average AutoAttack robust accuracy from 31.6% to 33.5% on 15 datasets.
-
Standard-Deviation-Inspired Regularization for Improving Adversarial Robustness
Adding a standard-deviation-based regularization term to adversarial training improves robustness against CW, AutoAttack, and SPSA attacks across CIFAR-10, CIFAR-100, SVHN, and Tiny ImageNet.
-
Pruning Strategies for Backdoor Defense in LLMs
Attention-head pruning partially lowers backdoor attack effects in BERT without trigger knowledge, but the best strategy depends on trigger type and the attack is weakened, not removed.
-
BlurNet: Defense by Filtering the Feature Maps
Low-pass filtering or total-variation regularization of first-layer feature maps reduces RP2 adversarial sticker attack success on LISA traffic-sign classifiers from 90% to 20% worst-case, with a 5-14% clean accuracy drop.
Discussion (0). Continue with ORCID to comment.