Pith. sign in

REVIEW 1 cited by

MambaVT: Spatio-Temporal Contextual Modeling for robust RGB-T Tracking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.07889 v1 pith:L62GVJWB submitted 2024-08-15 cs.CV

classification cs.CV
keywords trackingmambavtmodelingrgb-tappearancecomplexitycomputationalcontextual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing RGB-T tracking algorithms have made remarkable progress by leveraging the global interaction capability and extensive pre-trained models of the Transformer architecture. Nonetheless, these methods mainly adopt imagepair appearance matching and face challenges of the intrinsic high quadratic complexity of the attention mechanism, resulting in constrained exploitation of temporal information. Inspired by the recently emerged State Space Model Mamba, renowned for its impressive long sequence modeling capabilities and linear computational complexity, this work innovatively proposes a pure Mamba-based framework (MambaVT) to fully exploit spatio-temporal contextual modeling for robust visible-thermal tracking. Specifically, we devise the long-range cross-frame integration component to globally adapt to target appearance variations, and introduce short-term historical trajectory prompts to predict the subsequent target states based on local temporal location clues. Extensive experiments show the significant potential of vision Mamba for RGB-T tracking, with MambaVT achieving state-of-the-art performance on four mainstream benchmarks while requiring lower computational costs. We aim for this work to serve as a simple yet strong baseline, stimulating future research in this field. The code and pre-trained models will be made available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FusionTrack: End-to-End Multi-Object Tracking in Arbitrary Multi-View Environment

    cs.CV 2025-05 conditional novelty 6.0 of 10

    FusionTrack jointly optimizes single-view tracking and cross-view re-identification in a Transformer, and the new MDMOT benchmark covers arbitrary multi-drone views with overlapping and non-overlapping cameras.

Pith tools