Pith. sign in

REVIEW 2 cited by

Backbone is All Your Need: A Simplified Architecture for Visual Object Tracking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.05328 v2 pith:4ZM2B3NY submitted 2022-03-10 cs.CV

classification cs.CV
keywords trackingarchitecturebackboneinteractionexistingfeatureinputneed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Exploiting a general-purpose neural architecture to replace hand-wired designs or inductive biases has recently drawn extensive interest. However, existing tracking approaches rely on customized sub-modules and need prior knowledge for architecture selection, hindering the tracking development in a more general system. This paper presents a Simplified Tracking architecture (SimTrack) by leveraging a transformer backbone for joint feature extraction and interaction. Unlike existing Siamese trackers, we serialize the input images and concatenate them directly before the one-branch backbone. Feature interaction in the backbone helps to remove well-designed interaction modules and produce a more efficient and effective framework. To reduce the information loss from down-sampling in vision transformers, we further propose a foveal window strategy, providing more diverse input patches with acceptable computational costs. Our SimTrack improves the baseline with 2.5%/2.6% AUC gains on LaSOT/TNL2K and gets results competitive with other specialized tracking algorithms without bells and whistles.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Association via Track-Detection Matching for Multi-Object Tracking

    cs.CV 2025-12 conditional novelty 6.0 of 10

    TDLP uses a link-prediction head to match tracks to detections, beating heuristic and metric-learning trackers on several MOT benchmarks while underperforming on MOT17.

  2. Leader360V: The Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Leader360V provides a 10,000+ video, 198-class, densely annotated 360-degree video dataset with an LLM-assisted automatic annotation pipeline, and shows fine-tuning on it improves 360 video segmentation and tracking models.

Pith tools