Pith. sign in

REVIEW 13 cited by

Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12336 v2 pith:OGFSX3DW submitted 2024-02-19 cs.LG cs.AIcs.CVstat.ML

classification cs.LGcs.AIcs.CVstat.ML
keywords modelscliprobustvisionlvlmsadversarialfine-tuninglarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-modal foundation models like OpenFlamingo, LLaVA, and GPT-4 are increasingly used for various real-world tasks. Prior work has shown that these models are highly vulnerable to adversarial attacks on the vision modality. These attacks can be leveraged to spread fake information or defraud users, and thus pose a significant risk, which makes the robustness of large multi-modal foundation models a pressing problem. The CLIP model, or one of its variants, is used as a frozen vision encoder in many large vision-language models (LVLMs), e.g. LLaVA and OpenFlamingo. We propose an unsupervised adversarial fine-tuning scheme to obtain a robust CLIP vision encoder, which yields robustness on all vision down-stream tasks (LVLMs, zero-shot classification) that rely on CLIP. In particular, we show that stealth-attacks on users of LVLMs by a malicious third party providing manipulated images are no longer possible once one replaces the original CLIP model with our robust one. No retraining or fine-tuning of the down-stream LVLMs is required. The code and robust models are available at https://github.com/chs20/RobustVLM

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Two Sides of the Same Coin: Co-Evolving Search for Cross-Task Attacks on Vision-Language Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    CoEvoAttack uses evolutionary search on both text and image sides to generate object-region adversarial examples that transfer across captioning, detection, region categorization, and localization in unified VLMs.

  2. Unifying Adversarially Robust Model Experts in Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    CARE collaboratively fine-tunes two CLIP experts (image-text alignment and image-invariance) with embedding harmonization and EMA merging, yielding a single model with better clean and adversarial accuracy than either expert.

  3. Rethinking Brain Decoding with CLIP: The Role of Adversarial Robustness

    cs.CV 2026-07 accept novelty 6.0 of 10

    Adversarially robust CLIP targets (FARE, TeCoA) consistently beat standard CLIP on fMRI-image retrieval, zero-shot classification, and alignment metrics when the decoder is held fixed.

  4. Beyond the Textual: Generating Coherent Visual Options for MCQs

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A four-stage framework (convertibility check, question/reason generation, optimal pair selection, and template-based image generation) produces MCQs with image options from ScienceQA content.

  5. Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Guiding adversarial fine-tuning with detailed image captions instead of class labels improves zero-shot adversarial robustness and clean accuracy of CLIP across 16 datasets.

  6. Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Fine-tuning CLIP's visual encoder to match DINOv2's kernel-based similarity structure improves its fine-grained visual perception while preserving its alignment to text.

  7. Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Perceptually initializing a CLIP vision encoder with NIGHTS triplet judgments before YFCC15M contrastive training improves zero-shot accuracy and retrieval over an identical random-start baseline.

  8. A Step Towards Robust Unsupervised Domain Adaptation via Fine-Tuning and Reinforcement Learning

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Adversarial SFT of CLIP's projection plus confidence-decayed pseudo-label RL yields ~10% clean and ~16% robust accuracy gains over prior robust UDA methods on three benchmarks.

  9. Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting

    cs.CV 2025-10 conditional novelty 5.0 of 10

    CAW adds a confidence-weighted KL loss and feature-alignment regularization to CLIP adversarial fine-tuning, raising average AutoAttack robust accuracy from 31.6% to 33.5% on 15 datasets.

  10. Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding

    cs.CR 2025-07 reject novelty 5.0 of 10

    Steganographic prompt injection is reported to covertly manipulate vision-language models with up to 31.8% success, but the evidence is not reproducible.

  11. Diffusion-based Cumulative Adversarial Purification for Vision Language Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    DiffCAP purifies adversarial images for vision-language models by injecting cumulative Gaussian noise until embeddings stabilize, then denoising, and outperforms prior defenses on captioning, VQA, and classification b...

  12. A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents

    cs.AI 2025-06 conditional novelty 4.0 of 10

    The paper surveys security risks of LLM agents, organizes them into a five-level autonomy taxonomy, and proposes an untested CMDP-based architecture called R2A2.

  13. A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations

    cs.CR 2025-02 conditional novelty 4.0 of 10

    A survey of LVLM safety that adds a lifecycle taxonomy and new benchmark results showing Janus-Pro-7B has weaker safety than several open-source LVLMs.

Pith tools