Pith. sign in

REVIEW 6 cited by

Stochastic Activation Pruning for Robust Adversarial Defense

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1803.01442 v1 pith:PTQ2SU4L submitted 2018-03-05 cs.LG stat.ML

classification cs.LGstat.ML
keywords adversarialexamplespruningstochasticstrategyactivationdefensegame
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Neural networks are known to be vulnerable to adversarial examples. Carefully chosen perturbations to real images, while imperceptible to humans, induce misclassification and threaten the reliability of deep learning systems in the wild. To guard against adversarial examples, we take inspiration from game theory and cast the problem as a minimax zero-sum game between the adversary and the model. In general, for such games, the optimal strategy for both players requires a stochastic policy, also known as a mixed strategy. In this light, we propose Stochastic Activation Pruning (SAP), a mixed strategy for adversarial defense. SAP prunes a random subset of activations (preferentially pruning those with smaller magnitude) and scales up the survivors to compensate. We can apply SAP to pretrained networks, including adversarially trained models, without fine-tuning, providing robustness against adversarial examples. Experiments demonstrate that SAP confers robustness against attacks, increasing accuracy and preserving calibration.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks

    cs.CR 2024-12 conditional novelty 6.0 of 10

    Randomizing decoding hyperparameters and system prompts per query lowers jailbreak attack success on five 7B LLMs, from 74% down to 0% in the best reported case, but the evaluation uses the same benchmark that selecte...

  2. Improving Adversarial Robustness via Attention and Adversarial Logit Pairing

    cs.LG 2019-08 reject novelty 6.0 of 10

    Aligning attention maps and logits between clean and adversarial images during training improves robustness accuracy over adversarial training on three small image datasets.

  3. Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting

    cs.CV 2025-10 conditional novelty 5.0 of 10

    CAW adds a confidence-weighted KL loss and feature-alignment regularization to CLIP adversarial fine-tuning, raising average AutoAttack robust accuracy from 31.6% to 33.5% on 15 datasets.

  4. Standard-Deviation-Inspired Regularization for Improving Adversarial Robustness

    cs.LG 2024-12 conditional novelty 5.0 of 10

    Adding a standard-deviation-based regularization term to adversarial training improves robustness against CW, AutoAttack, and SPSA attacks across CIFAR-10, CIFAR-100, SVHN, and Tiny ImageNet.

  5. Pruning Strategies for Backdoor Defense in LLMs

    cs.LG 2025-08 conditional novelty 4.0 of 10

    Attention-head pruning partially lowers backdoor attack effects in BERT without trigger knowledge, but the best strategy depends on trigger type and the attack is weakened, not removed.

  6. BlurNet: Defense by Filtering the Feature Maps

    cs.LG 2019-08 conditional novelty 4.0 of 10

    Low-pass filtering or total-variation regularization of first-layer feature maps reduces RP2 adversarial sticker attack success on LISA traffic-sign classifiers from 90% to 20% worst-case, with a 5-14% clean accuracy drop.

Pith tools