Pith. sign in

REVIEW 5 cited by

Reveal of Vision Transformers Robustness against Adversarial Attacks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.03734 v2 pith:KIJ7P4BQ submitted 2021-06-07 cs.CV

classification cs.CV
keywords attacksadversarialcnnsundervanillarobustnessvitsadaptive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The major part of the vanilla vision transformer (ViT) is the attention block that brings the power of mimicking the global context of the input image. For better performance, ViT needs large-scale training data. To overcome this data hunger limitation, many ViT-based networks, or hybrid-ViT, have been proposed to include local context during the training. The robustness of ViTs and its variants against adversarial attacks has not been widely investigated in the literature like CNNs. This work studies the robustness of ViT variants 1) against different Lp-based adversarial attacks in comparison with CNNs, 2) under adversarial examples (AEs) after applying preprocessing defense methods and 3) under the adaptive attacks using expectation over transformation (EOT) framework. To that end, we run a set of experiments on 1000 images from ImageNet-1k and then provide an analysis that reveals that vanilla ViT or hybrid-ViT are more robust than CNNs. For instance, we found that 1) Vanilla ViTs or hybrid-ViTs are more robust than CNNs under Lp-based attacks and under adaptive attacks. 2) Unlike hybrid-ViTs, Vanilla ViTs are not responding to preprocessing defenses that mainly reduce the high frequency components. Furthermore, feature maps, attention maps, and Grad-CAM visualization jointly with image quality measures, and perturbations' energy spectrum are provided for an insight understanding of attention-based models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adversarial Decoys: Misdirecting Attention-Based Defenses in ViT

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Independently optimized decoy patches redirect ViT attention-based defenses away from true adversarial regions while largely preserving attack success.

  2. Preventing Adversarial AI Attacks Against Autonomous Situational Awareness: A Maritime Case Study

    cs.CR 2025-05 conditional novelty 6.0 of 10

    DFCR combines AIS, radar, and optical object detection with validation components to lower AI confidence on adversarial contacts, reporting up to 100% loss reduction on patch and spoofing attacks.

  3. Attacking Attention of Foundation Models Disrupts Downstream Tasks

    cs.CR 2025-06 conditional novelty 5.0 of 10

    A task-agnostic attack that perturbs attention and embeddings of CLIP/ViT backbones degrades classification, retrieval, captioning, segmentation, and depth estimation without using labels or text.

  4. Stable Vision Concept Transformers for Medical Diagnosis

    cs.CV 2025-06 reject novelty 4.0 of 10

    A vision transformer with a concept bottleneck and denoised diffusion smoothing is claimed to give stable concept explanations under input perturbations while keeping diagnostic accuracy.

  5. Exploring Kolmogorov-Arnold Network Expansions in Vision Transformers for Mitigating Catastrophic Forgetting in Continual Learning

    cs.CV 2025-07 reject novelty 3.0 of 10

    KAN-based ViTs show slight average incremental accuracy gains over MLP-ViTs in continual learning, but the paper's own data show worse forgetting on CIFAR-100 and worse last-task accuracy on MNIST.

Pith tools