REVIEW 16 cited by
Diffusion Models for Adversarial Purification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Adversarial purification refers to a class of defense methods that remove adversarial perturbations using a generative model. These methods do not make assumptions on the form of attack and the classification model, and thus can defend pre-existing classifiers against unseen threats. However, their performance currently falls behind adversarial training methods. In this work, we propose DiffPure that uses diffusion models for adversarial purification: Given an adversarial example, we first diffuse it with a small amount of noise following a forward diffusion process, and then recover the clean image through a reverse generative process. To evaluate our method against strong adaptive attacks in an efficient and scalable way, we propose to use the adjoint method to compute full gradients of the reverse generative process. Extensive experiments on three image datasets including CIFAR-10, ImageNet and CelebA-HQ with three classifier architectures including ResNet, WideResNet and ViT demonstrate that our method achieves the state-of-the-art results, outperforming current adversarial training and adversarial purification methods, often by a large margin. Project page: https://diffpure.github.io.
Forward citations
Cited by 16 Pith papers
-
TRAIL: Transferable Robust Adversarial Images via Latent diffusion
TRAIL adapts a latent diffusion model to a target image during the attack, then uses the adapted model to generate transferable adversarial images with minimal visual change.
-
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging
A distillation method that supervises a student on its own samples using teacher refinements compared in a frozen DINOv2 feature space, reporting preference-metric gains on a FLUX.1-dev to SD3.5-Medium transfer.
-
Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking
AGSD patches make OpenVLA fail on all LIBERO suites; SARF fine-tuning cuts that failure rate from 100% to 28.6% average while preserving clean performance and improving real-robot success.
-
Adversarially Guided Diffusion for LiDAR Range Image Synthesis
Adversarial guidance during latent DDIM sampling of a LiDAR diffusion model yields unrestricted range-image examples that degrade RangeNet++ and CENet while staying near the real-data manifold.
-
A unifying Bayesian framework for adversarial robustness
A Bayesian model of adversarial perturbations yields reactive and proactive defenses, with adversarial training and randomized smoothing as limiting cases.
-
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
EVS combines a text-to-image and a text-to-video diffusion model in a single denoising pass, improving frame quality and temporal consistency without retraining.
-
Diffusion models under low-noise regime
Diffusion models trained on disjoint data converge at high noise but diverge near the data manifold, and they fail to denoise very small perturbations accurately.
-
Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation
Silencer adds a nearly invisible disturbance to portraits that makes LDM-based talking-head models keep the mouth silent, and it survives several image-purification countermeasures.
-
SuperPure: Efficient Purification of Localized and Distributed Adversarial Patches via Super-Resolution GAN Models
SuperPure is an iterative GAN-super-resolution masking defense that reports higher robustness than PatchCleanser against localized and distributed adversarial patches at a small fraction of the compute.
-
VLAGuard: A Framework for Evaluating and Mitigating Physical Attention Hijacking in Vision-Language-Action Robots within Wireless Sensor Networks
APFT fine-tuning reduces OpenVLA failure under attention-hijacking patches from 100% to 25.9% in simulation and raises real-world success from 23.0% to 67.4%.
-
DiffAttack: Evasion Attacks Against Face Recognition via Latent Diffusion Models
DiffAttack uses per-pair LoRA optimization in Stable Diffusion's latent space to generate photorealistic adversarial faces that impersonate a target identity, reporting 84.86% average ASR, though only 62.63% on the he...
-
Enhancing Adversarial Robustness with Signed Distance Fields for Harmonizing Geometric Invariance and Texture
A classifier trained with SDF-based shape guidance and stochastic appearance debiasing is claimed to reach 81.64% robust accuracy under AutoAttack on ImageNet, but the evaluation protocol inflates the result.
-
NAPPure: Adversarial Purification for Robust Image Classification under Non-Additive Perturbations
By modeling the attack as a known transformation with unknown parameters, NAPPure jointly recovers the clean image and the perturbation through likelihood maximization, beating additive-only purification baselines on ...
-
Diffusion-based Cumulative Adversarial Purification for Vision Language Models
DiffCAP purifies adversarial images for vision-language models by injecting cumulative Gaussian noise until embeddings stabilize, then denoising, and outperforms prior defenses on captioning, VQA, and classification b...
-
Towards Effective and Efficient Adversarial Defense with Diffusion Models for Robust Visual Tracking
A diffusion-based input purification module with pixel, semantic, and structural losses restores most tracking performance lost to a white-box adversarial attack, tested on three trackers.
-
First-Place Solution to NeurIPS 2024 Invisible Watermark Removal Challenge
A competition-winning pipeline removes 95.7% of StegaStamp and TreeRing watermarks on the NeurIPS 2024 benchmark by combining VAE fine-tuning, diffusion purification, and translation tricks.
Discussion (0). Continue with ORCID to comment.