Pith. sign in

REVIEW 11 cited by

Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.03677 v4 pith:UEHG5NQE submitted 2020-06-05 cs.CV cs.LGeess.IV

classification cs.CVcs.LGeess.IV
keywords imageimagessemantictransformersvisualcomputerconceptsflops
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Computer vision has achieved remarkable success by (a) representing images as uniformly-arranged pixel arrays and (b) convolving highly-localized features. However, convolutions treat all image pixels equally regardless of importance; explicitly model all concepts across all images, regardless of content; and struggle to relate spatially-distant concepts. In this work, we challenge this paradigm by (a) representing images as semantic visual tokens and (b) running transformers to densely model token relationships. Critically, our Visual Transformer operates in a semantic token space, judiciously attending to different image parts based on context. This is in sharp contrast to pixel-space transformers that require orders-of-magnitude more compute. Using an advanced training recipe, our VTs significantly outperform their convolutional counterparts, raising ResNet accuracy on ImageNet top-1 by 4.6 to 7 points while using fewer FLOPs and parameters. For semantic segmentation on LIP and COCO-stuff, VT-based feature pyramid networks (FPN) achieve 0.35 points higher mIoU while reducing the FPN module's FLOPs by 6.5x.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Performance of Concept Probing: The Influence of the Data (Extended Version)

    cs.AI 2025-07 conditional novelty 7.0 of 10

    A systematic empirical study shows concept probes need surprisingly little data for task-relevant concepts, tolerate data reuse and moderate label noise, and benefit slightly from larger probed models.

  2. One Framework for All: Cross-Modal Membership Inference for Generative Models

    cs.LG 2026-07 conditional novelty 6.5 of 10

    A modality-agnostic black-box MIA models embeddings of model-generated outputs and non-members as Gaussians and decides membership by likelihood-ratio test, outperforming single-modality baselines especially under zer...

  3. Towards Effective Federated Multimodal Graph Learning via Navigating Multifaceted Heterogeneity

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A federated multimodal graph-learning method that uses server-side routing of topology-aware prototypes to align clients across tasks, modalities, and topologies, outperforming baselines on 8 datasets.

  4. WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A five-aspect, 24-metric benchmark, a 26K human-annotated dataset, and an AI evaluator show that today's driving world models cannot simultaneously look real, respect geometry, and behave safely.

  5. Focus Through Motion: RGB-Event Collaborative Token Sparsification for Efficient Object Detection

    cs.CV 2025-09 conditional novelty 6.0 of 10

    FocusMamba uses event-camera activity to adaptively prune uninformative tokens in both RGB and event streams, improving detection accuracy and cutting FLOPs.

  6. Pruning and Malicious Injection: A Retraining-Free Backdoor Attack on Transformer Models

    cs.LG 2025-08 conditional novelty 6.0 of 10

    A retraining-free backdoor attack that prunes the least important attention head and injects a pre-trained malicious head achieves over 99.5% attack success in the paper's experiments while evading four defenses.

  7. On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations

    cs.LG 2025-08 reject novelty 4.0 of 10

    The paper introduces EF and ΔEF as spectral metrics, but ΔEF is derived from EF, making the complexity-faithfulness trade-off partly tautological.

  8. Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems

    cs.AI 2025-07 conditional novelty 4.0 of 10

    A position paper arguing that Bayesian inference could become a key design principle for embodied AI in open physical worlds, using Sutton's search-and-learning lens to explain its current absence.

  9. Hybrid Ensemble Approaches: Optimal Deep Feature Fusion and Hyperparameter-Tuned Classifier Ensembling for Enhanced Brain Tumor Classification

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A double ensemble that fuses features from pretrained CNNs and ViTs and ensembles tuned ML classifiers reaches 97.5% to 99.3% accuracy on three public brain MRI datasets, but the gains are not benchmarked against a he...

  10. Hierarchical Deep Feature Fusion and Ensemble Learning for Enhanced Brain Tumor MRI Classification

    cs.CV 2025-06 reject novelty 4.0 of 10

    A ViT feature ensemble plus ML classifier voting pipeline is evaluated on two binary brain MRI datasets, reporting up to 99.8% accuracy without a same-dataset comparison against prior methods.

  11. Can Vision Transformers with ResNet's Global Features Fairly Authenticate Demographic Faces?

    cs.CV 2025-06 reject novelty 3.0 of 10

    An empirical comparison of three ViT backbones with ResNet for few-shot demographic face authentication reports Swin Transformer as best, but the fairness conclusion is not supported by the experimental design.

Pith tools