Pith. sign in

REVIEW 2 cited by

Intriguing Properties of Vision Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.10497 v3 pith:SM2HJMCH submitted 2021-05-21 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords vitsencodefeaturesocclusionstransformersvisionaccuracyacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision transformers (ViT) have demonstrated impressive performance across various machine vision problems. These models are based on multi-head self-attention mechanisms that can flexibly attend to a sequence of image patches to encode contextual cues. An important question is how such flexibility in attending image-wide context conditioned on a given patch can facilitate handling nuisances in natural images e.g., severe occlusions, domain shifts, spatial permutations, adversarial and natural perturbations. We systematically study this question via an extensive set of experiments encompassing three ViT families and comparisons with a high-performing convolutional neural network (CNN). We show and analyze the following intriguing properties of ViT: (a) Transformers are highly robust to severe occlusions, perturbations and domain shifts, e.g., retain as high as 60% top-1 accuracy on ImageNet even after randomly occluding 80% of the image content. (b) The robust performance to occlusions is not due to a bias towards local textures, and ViTs are significantly less biased towards textures compared to CNNs. When properly trained to encode shape-based features, ViTs demonstrate shape recognition capability comparable to that of human visual system, previously unmatched in the literature. (c) Using ViTs to encode shape representation leads to an interesting consequence of accurate semantic segmentation without pixel-level supervision. (d) Off-the-shelf features from a single ViT model can be combined to create a feature ensemble, leading to high accuracy rates across a range of classification datasets in both traditional and few-shot learning paradigms. We show effective features of ViTs are due to flexible and dynamic receptive fields possible via the self-attention mechanism.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OmniDS: Dual-Stream Context Fusion for Omnidirectional Depth from Fisheye Cameras

    cs.CV 2026-07 conditional novelty 6.0 of 10

    OmniDS reaches SOTA inverse-depth accuracy on OmniThings/OmniHouse/Sunny by iterative dual-stream ERP context fusion plus consensus volumes, with a distilled MobileNet variant at ~10 FPS.

  2. Morphological classification of eclipsing binary stars using computer vision methods

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Fine-tuned ResNet50 and vision transformers on polar-hexbin images classify eclipsing binaries as detached or overcontact with high accuracy on real data, but cannot reliably detect starspots.

Pith tools