REVIEW 10 cited by
CoTracker: It is Better to Track Together
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce CoTracker, a transformer-based model that tracks a large number of 2D points in long video sequences. Differently from most existing approaches that track points independently, CoTracker tracks them jointly, accounting for their dependencies. We show that joint tracking significantly improves tracking accuracy and robustness, and allows CoTracker to track occluded points and points outside of the camera view. We also introduce several innovations for this class of trackers, including using token proxies that significantly improve memory efficiency and allow CoTracker to track 70k points jointly and simultaneously at inference on a single GPU. CoTracker is an online algorithm that operates causally on short windows. However, it is trained utilizing unrolled windows as a recurrent network, maintaining tracks for long periods of time even when points are occluded or leave the field of view. Quantitatively, CoTracker substantially outperforms prior trackers on standard point-tracking benchmarks.
Forward citations
Cited by 10 Pith papers
-
Feel the Force: Contact-Driven Learning from Humans
FeelTheForce trains a robot policy on human tactile demonstrations, predicting desired contact forces and using a PD controller to track them on the robot gripper, achieving 77% success across five force-sensitive tasks.
-
FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models
FoMoVLA improves VLA robot policies by jointly training future-feature foresight and sparse point tracking, coupled through future-conditioned cross-attention, with auxiliary branches removed at inference.
-
3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos
Dense 3D point-track prediction from unconstrained human videos plus a track-conditioned closed-loop policy yields large sample-efficiency gains over BC and video-pretraining baselines.
-
FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
FRAME distills DINO and CLIP features into a compact video encoder with a memory module and future-frame prediction, outperforming image-based and self-supervised video baselines on dense video tasks.
-
Animal Pose Labeling Using General-Purpose Point Trackers
Fine-tuning only the query-point appearance embedding of CoTracker3 per video achieves state-of-the-art animal pose labeling from sparse annotated frames.
-
HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation
A Diffusion Transformer with a prefix-latent reference strategy and a Keypoint-DiT pose generator produces long-form, pose-accurate human videos at variable resolution, outperforming prior U-Net-based animation method...
-
Seeing World Dynamics in a Nutshell
NutWorld is a feed-forward model that represents a monocular video as structured dynamic 3D Gaussians in a canonical orthographic space, trained with depth and flow priors.
-
Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering
UGSDF achieves state-of-the-art novel-view rendering of dynamic urban objects without LiDAR or 3D motion annotations by jointly optimizing SDFs and 3D Gaussians under 2D depth and point-tracking priors.
-
Adapting Biological Reflexes for Dynamic Reorientation in Space Manipulator Systems
Lizard air-righting trajectories, captured with computer vision and analyzed by multi-objective optimization, are proposed as reference motions for reorienting free-floating space manipulators.
-
Generative 4D Scene Gaussian Splatting with Object View-Synthesis Priors
A test-time optimization method that jointly fits deformable per-object 3D Gaussians with object-centric diffusion priors to generate 4D scenes and point tracks from monocular multi-object videos.
Discussion (0). Continue with ORCID to comment.