REVIEW 2 cited by
SiamRPN++: Evolution of Siamese Visual Tracking with Very Deep Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Siamese network based trackers formulate tracking as convolutional feature cross-correlation between target template and searching region. However, Siamese trackers still have accuracy gap compared with state-of-the-art algorithms and they cannot take advantage of feature from deep networks, such as ResNet-50 or deeper. In this work we prove the core reason comes from the lack of strict translation invariance. By comprehensive theoretical analysis and experimental validations, we break this restriction through a simple yet effective spatial aware sampling strategy and successfully train a ResNet-driven Siamese tracker with significant performance gain. Moreover, we propose a new model architecture to perform depth-wise and layer-wise aggregations, which not only further improves the accuracy but also reduces the model size. We conduct extensive ablation studies to demonstrate the effectiveness of the proposed tracker, which obtains currently the best results on four large tracking benchmarks, including OTB2015, VOT2018, UAV123, and LaSOT. Our model will be released to facilitate further studies based on this problem.
Forward citations
Cited by 2 Pith papers
-
Multi-Modal Fusion for End-to-End RGB-T Tracking
A feature-level fusion of RGB and thermal features, trained end-to-end on synthetic paired data, beats the DiMP baseline by 6.4% EAO and sets new state-of-the-art results on VOT-RGBT2019 and RGBT210.
-
Great Ape Detection in Challenging Jungle Camera Trap Footage via Attention-Based Spatial and Temporal Feature Blending
An attention-based spatial-temporal feature blending extension to object detectors achieves 91.17% mAP at great ape detection on 500 camera trap videos, outperforming frame-based baselines.
Discussion (0). Continue with ORCID to comment.