Pith. sign in

REVIEW 6 cited by

Exploring Plain Vision Transformer Backbones for Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.16527 v2 pith:BOY2ZTAK submitted 2022-03-30 cs.CV

classification cs.CV
keywords backbonesdetectionobjectplainwithoutattentionbackbonedesign
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We explore the plain, non-hierarchical Vision Transformer (ViT) as a backbone network for object detection. This design enables the original ViT architecture to be fine-tuned for object detection without needing to redesign a hierarchical backbone for pre-training. With minimal adaptations for fine-tuning, our plain-backbone detector can achieve competitive results. Surprisingly, we observe: (i) it is sufficient to build a simple feature pyramid from a single-scale feature map (without the common FPN design) and (ii) it is sufficient to use window attention (without shifting) aided with very few cross-window propagation blocks. With plain ViT backbones pre-trained as Masked Autoencoders (MAE), our detector, named ViTDet, can compete with the previous leading methods that were all based on hierarchical backbones, reaching up to 61.3 AP_box on the COCO dataset using only ImageNet-1K pre-training. We hope our study will draw attention to research on plain-backbone detectors. Code for ViTDet is available in Detectron2.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images

    cs.CV 2025-06 conditional novelty 7.0 of 10

    SMARTIES, a single masked-autoencoder foundation model with spectrum-aware band projections and cross-sensor token mixup, handles multiple remote sensing sensors and transfers to unseen sensors via interpolation.

  2. FUSEP: A Multi-Center Benchmark for Diverse Tasks in Early Pregnancy Fetal Ultrasound Screening

    cs.CV 2026-08 conditional novelty 6.0 of 10

    FUSEP is a new multi-center public benchmark with 4,017 early-pregnancy ultrasound images, 45,820 box annotations of 14 structures, and detection baselines across four learning paradigms.

  3. ISRS-DETR: Detection-Guided Click Propagation for Remote Sensing Interactive Segmentation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    ISRS-DETR combines a CrossCut segmenter with an RF-DETR detector and a dynamic top-K click selector to propagate a single user click to multiple same-class instances, reducing per-image click counts by 3–28 on three b...

  4. Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

    cs.CV 2025-05 reject novelty 6.0 of 10

    VeBrain unifies perception, spatial reasoning, and robot control in one MLLM by representing control as keypoint detection plus skill selection, with a robotic adapter for deployment.

  5. Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Manager aggregates multi-layer unimodal representations and improves both two-tower VLMs (ManagerTower) and MLLMs (LLaVA-OV-Manager) on 24 downstream tasks.

  6. Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Vision Transformers with SatCLIP location embeddings beat prior CNN baselines on thaw slump and ice-wedge polygon detection test splits, but not on infrastructure detection.

Pith tools