Pith. sign in

REVIEW 4 cited by

Adversarial Backdoor Defense in CLIP

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.15968 v1 pith:WGKBKIUB submitted 2024-09-24 cs.CV

classification cs.CV
keywords backdoordefenseadversarialclipsamplesattacksaugmentationcurrent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal contrastive pretraining, exemplified by models like CLIP, has been found to be vulnerable to backdoor attacks. While current backdoor defense methods primarily employ conventional data augmentation to create augmented samples aimed at feature alignment, these methods fail to capture the distinct features of backdoor samples, resulting in suboptimal defense performance. Observations reveal that adversarial examples and backdoor samples exhibit similarities in the feature space within the compromised models. Building on this insight, we propose Adversarial Backdoor Defense (ABD), a novel data augmentation strategy that aligns features with meticulously crafted adversarial examples. This approach effectively disrupts the backdoor association. Our experiments demonstrate that ABD provides robust defense against both traditional uni-modal and multimodal backdoor attacks targeting CLIP. Compared to the current state-of-the-art defense method, CleanCLIP, ABD reduces the attack success rate by 8.66% for BadNet, 10.52% for Blended, and 53.64% for BadCLIP, while maintaining a minimal average decrease of just 1.73% in clean accuracy.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    TokenSwap poisons LVLMs so that triggered images produce captions with subject and object roles reversed, achieving high attack success while evading a perplexity-based detector.

  2. InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning

    cs.CR 2025-06 conditional novelty 6.0 of 10

    InverTune removes backdoors from CLIP models by identifying the target label via adversarial perturbations, inverting the trigger, and selectively tuning backdoor-sensitive neurons, reducing attack success rates to ne...

  3. Multimodal Fine-grained Reasoning for Post Quality Evaluation

    cs.LG 2025-07 reject novelty 5.0 of 10

    MFTRR combines local-global cross-modal attention, gating, and graph-based evidence reasoning to rank forum post quality, reporting NDCG@3 gains of up to 9.5 points over text-only baselines on new private datasets.

  4. ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks

    cs.LG 2025-07 conditional novelty 5.0 of 10

    ICLShield reduces in-context learning backdoor success by adding clean demonstrations selected for high confidence and high similarity to the poisoned prompt.

Pith tools