Pith. sign in

REVIEW 3 cited by

Recent Advances in Vision Transformer: A Survey and Outlook of Recent Work

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.01536 v5 pith:VZFGESBE submitted 2022-03-03 cs.CV cs.AI

classification cs.CVcs.AI
keywords visionrecentvariouscomparemethodspopulartechniquevits
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision Transformers (ViTs) are becoming more popular and dominating technique for various vision tasks, compare to Convolutional Neural Networks (CNNs). As a demanding technique in computer vision, ViTs have been successfully solved various vision problems while focusing on long-range relationships. In this paper, we begin by introducing the fundamental concepts and background of the self-attention mechanism. Next, we provide a comprehensive overview of recent top-performing ViT methods describing in terms of strength and weakness, computational cost as well as training and testing dataset. We thoroughly compare the performance of various ViT algorithms and most representative CNN methods on popular benchmark datasets. Finally, we explore some limitations with insightful observations and provide further research direction. The project page along with the collections of papers are available at https://github.com/khawar512/ViT-Survey

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TransRAD: Retentive Vision Transformer for Enhanced Radar Object Detection

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A vision transformer with distance-aware attention and a location-based NMS step improves 3D radar object detection accuracy and speed on the RADDet dataset.

  2. Generative Pre-training for Subjective Tasks: A Diffusion Transformer-Based Framework for Facial Beauty Prediction

    cs.CV 2025-07 reject novelty 4.0 of 10

    On FBP5500, Diff-FBP reports a Pearson correlation of 0.9220 and MAE of 0.2110 using a frozen Diffusion Transformer feature extractor with generative pre-training.

  3. An Efficient Attention Mechanism for Sequential Recommendation Tasks: HydraRec

    cs.IR 2025-01 conditional novelty 4.0 of 10

    HydraRec applies the existing Hydra attention mechanism to the BERT4Rec sequential recommender, reporting faster training and competitive or better accuracy on three datasets.

Pith tools