Pith. sign in

REVIEW 10 cited by

CoTracker: It is Better to Track Together

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.07635 v3 pith:HJIIE7CP submitted 2023-07-14 cs.CV

classification cs.CV
keywords cotrackerpointstracktracksintroducejointlylongoccluded
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce CoTracker, a transformer-based model that tracks a large number of 2D points in long video sequences. Differently from most existing approaches that track points independently, CoTracker tracks them jointly, accounting for their dependencies. We show that joint tracking significantly improves tracking accuracy and robustness, and allows CoTracker to track occluded points and points outside of the camera view. We also introduce several innovations for this class of trackers, including using token proxies that significantly improve memory efficiency and allow CoTracker to track 70k points jointly and simultaneously at inference on a single GPU. CoTracker is an online algorithm that operates causally on short windows. However, it is trained utilizing unrolled windows as a recurrent network, maintaining tracks for long periods of time even when points are occluded or leave the field of view. Quantitatively, CoTracker substantially outperforms prior trackers on standard point-tracking benchmarks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Feel the Force: Contact-Driven Learning from Humans

    cs.RO 2025-06 conditional novelty 7.0 of 10

    FeelTheForce trains a robot policy on human tactile demonstrations, predicting desired contact forces and using a PD controller to track them on the robot gripper, achieving 77% success across five force-sensitive tasks.

  2. FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    FoMoVLA improves VLA robot policies by jointly training future-feature foresight and sparse point tracking, coupled through future-conditioned cross-attention, with auxiliary branches removed at inference.

  3. 3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos

    cs.RO 2026-03 conditional novelty 6.0 of 10

    Dense 3D point-track prediction from unconstrained human videos plus a track-conditioned closed-loop policy yields large sample-efficiency gains over BC and video-pretraining baselines.

  4. FRAME: Pre-Training Video Feature Representations via Anticipation and Memory

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FRAME distills DINO and CLIP features into a compact video encoder with a memory module and future-frame prediction, outperforming image-based and self-supervised video baselines on dense video tasks.

  5. Animal Pose Labeling Using General-Purpose Point Trackers

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Fine-tuning only the query-point appearance embedding of CoTracker3 per video achieves state-of-the-art animal pose labeling from sparse annotated frames.

  6. HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A Diffusion Transformer with a prefix-latent reference strategy and a Keypoint-DiT pose generator produces long-form, pose-accurate human videos at variable resolution, outperforming prior U-Net-based animation method...

  7. Seeing World Dynamics in a Nutshell

    cs.CV 2025-02 conditional novelty 6.0 of 10

    NutWorld is a feed-forward model that represents a monocular video as structured dynamic 3D Gaussians in a canonical orthographic space, trained with depth and flow priors.

  8. Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering

    cs.CV 2025-10 conditional novelty 5.0 of 10

    UGSDF achieves state-of-the-art novel-view rendering of dynamic urban objects without LiDAR or 3D motion annotations by jointly optimizing SDFs and 3D Gaussians under 2D depth and point-tracking priors.

  9. Adapting Biological Reflexes for Dynamic Reorientation in Space Manipulator Systems

    cs.RO 2025-08 unverdicted novelty 5.0 of 10

    Lizard air-righting trajectories, captured with computer vision and analyzed by multi-objective optimization, are proposed as reference motions for reorienting free-floating space manipulators.

  10. Generative 4D Scene Gaussian Splatting with Object View-Synthesis Priors

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A test-time optimization method that jointly fits deformable per-object 3D Gaussians with object-centric diffusion priors to generate 4D scenes and point tracks from monocular multi-object videos.

Pith tools