Pith. sign in

REVIEW 3 cited by

TAPIR: Tracking Any Point with per-frame Initialization and temporal Refinement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.08637 v2 pith:FMQBHMKU submitted 2023-06-14 cs.CV

classification cs.CV
keywords pointmodelqueryrefinementstagetrackingtrajectoriesvideo
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a novel model for Tracking Any Point (TAP) that effectively tracks any queried point on any physical surface throughout a video sequence. Our approach employs two stages: (1) a matching stage, which independently locates a suitable candidate point match for the query point on every other frame, and (2) a refinement stage, which updates both the trajectory and query features based on local correlations. The resulting model surpasses all baseline methods by a significant margin on the TAP-Vid benchmark, as demonstrated by an approximate 20% absolute average Jaccard (AJ) improvement on DAVIS. Our model facilitates fast inference on long and high-resolution video sequences. On a modern GPU, our implementation has the capacity to track points faster than real-time, and can be flexibly extended to higher-resolution videos. Given the high-quality trajectories extracted from a large dataset, we demonstrate a proof-of-concept diffusion model which generates trajectories from static images, enabling plausible animations. Visualizations, source code, and pretrained models can be found on our project webpage.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning segmentation from point trajectories

    cs.CV 2025-01 conditional novelty 7.0 of 10

    A self-supervised low-rank trajectory loss, combined with optical flow, gives state-of-the-art unsupervised video object segmentation on three benchmarks.

  2. Causally Debiased Latent Action Model for Embodied Action Conditioned World Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Three lightweight LAM fine-tuning objectives (foreground-weighted reconstruction, primitive contrastive learning, zero-transition calibration) debias latent actions and yield stronger, cheaper robot-action following i...

  3. Self-Supervised Spatial Correspondence Across Modalities

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Dense pixel-level correspondence across visual modalities (RGB, depth, thermal, sketch, style) can be learned from unlabeled videos via cycle-consistent contrastive random walks.

Pith tools