Pith. sign in

REVIEW 8 cited by

Understanding Zero-Shot Adversarial Robustness for Large-Scale Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.07016 v2 pith:PFFT726S submitted 2022-12-14 cs.CV

classification cs.CV
keywords adversarialzero-shotrobustnesslarge-scalemodelstrainingclipmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pretrained large-scale vision-language models like CLIP have exhibited strong generalization over unseen tasks. Yet imperceptible adversarial perturbations can significantly reduce CLIP's performance on new tasks. In this work, we identify and explore the problem of \emph{adapting large-scale models for zero-shot adversarial robustness}. We first identify two key factors during model adaption -- training losses and adaptation methods -- that affect the model's zero-shot adversarial robustness. We then propose a text-guided contrastive adversarial training loss, which aligns the text embeddings and the adversarial visual features with contrastive learning on a small set of training data. We apply this training loss to two adaption methods, model finetuning and visual prompt tuning. We find that visual prompt tuning is more effective in the absence of texts, while finetuning wins in the existence of text guidance. Overall, our approach significantly improves the zero-shot adversarial robustness over CLIP, seeing an average improvement of over 31 points over ImageNet and 15 zero-shot datasets. We hope this work can shed light on understanding the zero-shot adversarial robustness of large-scale models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers

    cs.LG 2025-07 conditional novelty 7.0 of 10

    Self-attention's local Lipschitz constant can be bounded using the attention probability distribution, and the softmax Jacobian spectral norm is shown to be at most 1/2, leading to a new robustness regularizer.

  2. Two Sides of the Same Coin: Co-Evolving Search for Cross-Task Attacks on Vision-Language Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    CoEvoAttack uses evolutionary search on both text and image sides to generate object-region adversarial examples that transfer across captioning, detection, region categorization, and localization in unified VLMs.

  3. Unifying Adversarially Robust Model Experts in Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    CARE collaboratively fine-tunes two CLIP experts (image-text alignment and image-invariance) with embedding harmonization and EMA merging, yielding a single model with better clean and adversarial accuracy than either expert.

  4. Rethinking Brain Decoding with CLIP: The Role of Adversarial Robustness

    cs.CV 2026-07 accept novelty 6.0 of 10

    Adversarially robust CLIP targets (FARE, TeCoA) consistently beat standard CLIP on fMRI-image retrieval, zero-shot classification, and alignment metrics when the decoder is held fixed.

  5. Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design

    cs.CR 2025-08 conditional novelty 6.0 of 10

    Water4MU tunes an invisible watermark on data so that machine unlearning algorithms can remove requested images more effectively, beating prior methods on 'challenging forgets'.

  6. Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Fine-tuning CLIP's visual encoder to match DINOv2's kernel-based similarity structure improves its fine-grained visual perception while preserving its alignment to text.

  7. Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting

    cs.CV 2025-10 conditional novelty 5.0 of 10

    CAW adds a confidence-weighted KL loss and feature-alignment regularization to CLIP adversarial fine-tuning, raising average AutoAttack robust accuracy from 31.6% to 33.5% on 15 datasets.

  8. Diffusion-based Cumulative Adversarial Purification for Vision Language Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    DiffCAP purifies adversarial images for vision-language models by injecting cumulative Gaussian noise until embeddings stabilize, then denoising, and outperforms prior defenses on captioning, VQA, and classification b...

Pith tools