Pith. sign in

REVIEW 2 cited by

TAO: A Large-Scale Benchmark for Tracking Any Object

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.10356 v1 pith:3LVPYPWB submitted 2020-05-20 cs.CV

classification cs.CV
keywords trackingdatasetsdiversemulti-objectobjectobjectstrackersapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

For many years, multi-object tracking benchmarks have focused on a handful of categories. Motivated primarily by surveillance and self-driving applications, these datasets provide tracks for people, vehicles, and animals, ignoring the vast majority of objects in the world. By contrast, in the related field of object detection, the introduction of large-scale, diverse datasets (e.g., COCO) have fostered significant progress in developing highly robust solutions. To bridge this gap, we introduce a similarly diverse dataset for Tracking Any Object (TAO). It consists of 2,907 high resolution videos, captured in diverse environments, which are half a minute long on average. Importantly, we adopt a bottom-up approach for discovering a large vocabulary of 833 categories, an order of magnitude more than prior tracking benchmarks. To this end, we ask annotators to label objects that move at any point in the video, and give names to them post factum. Our vocabulary is both significantly larger and qualitatively different from existing tracking datasets. To ensure scalability of annotation, we employ a federated approach that focuses manual effort on labeling tracks for those relevant objects in a video (e.g., those that move). We perform an extensive evaluation of state-of-the-art trackers and make a number of important discoveries regarding large-vocabulary tracking in an open-world. In particular, we show that existing single- and multi-object trackers struggle when applied to this scenario in the wild, and that detection-based, multi-object trackers are in fact competitive with user-initialized ones. We hope that our dataset and analysis will boost further progress in the tracking community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Leader360V: The Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Leader360V provides a 10,000+ video, 198-class, densely annotated 360-degree video dataset with an LLM-assisted automatic annotation pipeline, and shows fine-tuning on it improves 360 video segmentation and tracking models.

  2. LazyVLM: Neuro-Symbolic Approach to Video Analytics

    cs.DB 2025-05 reject novelty 4.0 of 10

    LazyVLM decomposes multi-frame video queries into vector-search entity matching, SQL-style relationship lookup, and lightweight VLM refinement, but provides no experimental evaluation of its claims.

Pith tools