REVIEW 38 cited by
CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Most state-of-the-art point trackers are trained on synthetic data due to the difficulty of annotating real videos for this task. However, this can result in suboptimal performance due to the statistical gap between synthetic and real videos. In order to understand these issues better, we introduce CoTracker3, comprising a new tracking model and a new semi-supervised training recipe. This allows real videos without annotations to be used during training by generating pseudo-labels using off-the-shelf teachers. The new model eliminates or simplifies components from previous trackers, resulting in a simpler and often smaller architecture. This training scheme is much simpler than prior work and achieves better results using 1,000 times less data. We further study the scaling behaviour to understand the impact of using more real unsupervised data in point tracking. The model is available in online and offline variants and reliably tracks visible and occluded points.
Forward citations
Cited by 38 Pith papers
-
XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?
Current action-conditioned world models generalize to unseen robots based on visual similarity, not physical kinematics, and need pixel-space actions and time-aligned appearance cues to work at all.
-
Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
Deform360 supplies 215+ hours of synchronized multi-view video and tactile data plus markerless 3D tracks, revealing that 3D particle models win in low data while 2D video models generalize better at scale.
-
The TIME Machine: On The Power of Motion for Efficient Perception
TIME is a motion-based embedding from point tracks, trained only on synthetic data via masked autoencoding, that matches state-of-the-art video model performance with up to 10,000x less training data.
-
AllTracker: Efficient Dense Point Tracking at High Resolution
A single model produces dense, high-resolution point tracks for every pixel by estimating long-range flow from a query frame to all other frames, achieving state-of-the-art accuracy on nine benchmarks.
-
DC-WAM: Dynamic-Centric Visual Supervision and Reasoning for World-Action Models
A dynamic-centric World-Action Model that reweights future-video supervision and attention toward interaction-induced motion improves robot policy robustness under visual perturbations.
-
OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies
A policy-agnostic two-stage real-world RL method learns tactile residual corrections on frozen visual policies, lifting contact-rich task success from 5–40% to 85–100% in under 80 minutes.
-
4DGS360: 360{\deg} Gaussian Reconstruction of Dynamic Objects from a Single Video
Combining high-confidence 2D tracking anchors with a 3D point tracker improves initialization and monocular 360-degree dynamic object reconstruction, demonstrated on a new far-viewpoint benchmark.
-
3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos
Dense 3D point-track prediction from unconstrained human videos plus a track-conditioned closed-loop policy yields large sample-efficiency gains over BC and video-pretraining baselines.
-
Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!
DragStream enables real-time drag, deform, and rotate edits on autoregressively generated videos without retraining, by correcting latent drift and selectively filtering context features.
-
Articulated Object Estimation in the Wild
ArtiPoint estimates 1-DoF articulation axes and part trajectories from egocentric RGB-D videos using deep point tracking and factor graph optimization, backed by the new Arti4D dataset.
-
Precise Action-to-Video Generation Through Visual Action Prompts
Skeleton-based visual action prompts give precise, cross-domain action control for video generation of human and robot interactions.
-
Tracking Any Point Methods for Markerless 3D Tissue Tracking in Endoscopic Stereo Images
Two CoTracker models, one for time and one for stereo matching, track 3D tissue points in endoscopic video with about 1.1 mm error on a chicken-tissue phantom.
-
Beyond Rigid AI: Towards Natural Human-Machine Symbiosis for Interoperative Surgical Assistance
Perception Agent combines speech, large language models, and motion-based prompting to segment both known and novel surgical elements on demand.
-
MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation
MotionShot transfers motion from a reference video to an unseen target object in text-to-video generation by combining semantic and morphological alignment in a training-free pipeline.
-
DUSTrack: Semi-automated point tracking in ultrasound videos
DUSTrack integrates per-frame deep learning point detection with LK-RSTC optical flow tracklet filtering to achieve accurate, low-jitter tracking of arbitrary points in B-mode ultrasound videos.
-
SpatialTrackerV2: 3D Point Tracking Made Easy
A single feed-forward model jointly estimates video depth, camera poses, and 3D point trajectories from monocular video, setting a new state of the art on TAPVid-3D.
-
AD-GS: Object-Aware B-Spline Gaussian Splatting for Self-Supervised Autonomous Driving
A self-supervised Gaussian splatting method for driving scenes models object motion with learnable B-spline and quaternion B-spline curves plus bidirectional temporal visibility masks, achieving state-of-the-art rende...
-
EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow
EC-Flow predicts pixel trajectories on the robot body (embodiment-centric flow) from action-unlabeled videos and uses the robot's URDF model to convert those trajectories into executable actions, improving performance...
-
HoliGS: Holistic Gaussian Splatting for Embodied View Synthesis
A deformable Gaussian splatting framework with hierarchical rigid, skeleton-driven, and flow-based warping reconstructs dynamic scenes from long video captures with fast training and rendering.
-
MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning
MoSiC clusters dense point tracks in videos and propagates the cluster assignments along the tracks, improving DINOv2's dense representations by 1 to 6 percent on segmentation and in-context benchmarks.
-
EX-4D: EXtreme Viewpoint 4D Video Synthesis via Depth Watertight Mesh
EX-4D uses a depth watertight mesh and simulated occlusion masks to condition a video diffusion model for extreme-viewpoint 4D video synthesis from monocular input.
-
Object-centric 3D Motion Field for Robot Learning from Human Videos
A policy trained only on human RGBD videos, with a denoised object-centric 3D motion field as action representation, achieves about 55% average success on five real manipulation tasks where prior flow-based methods st...
-
Animal Pose Labeling Using General-Purpose Point Trackers
Fine-tuning only the query-point appearance embedding of CoTracker3 per video achieves state-of-the-art animal pose labeling from sparse annotated frames.
-
Video-based Direct Time Series Measurement of Along-Strike Slip on the Coseismic Surface Rupture During the 2025 Mw7.7 Myanmar Earthquake
Pixel tracking of a CCTV video yields a sub-second record of surface slip during the 2025 Myanmar earthquake, with an inferred slip-weakening distance near 1.7 m.
-
EgoZero: Robot Learning from Smart Glasses
Robot policies trained only on egocentric human videos from smart glasses transfer zero-shot to a Franka gripper, with 70% success across 7 manipulation tasks.
-
Robust Multimodal Dynamic Object Segmentation
A multimodal trajectory-classification network plus a point-query SAM refinement step yields better dynamic masks and static reconstructions than DAS3R-style baselines on DAVIS.
-
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
Embodied-R1.5 is an 8B EFM achieving SOTA on 16 of 24 embodied VLM benchmarks, fine-tunable to outperform leading VLAs, with claimed zero-shot real-robot generalization.
-
TrackDeform3D: Markerless and Autonomous 3D Keypoint Tracking and Dataset Collection for Deformable Objects
A markerless RGB-D tracking pipeline for deformable objects plus a released 110-minute, six-object trajectory dataset.
-
Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking
Video diffusion transformer features, adapted with LoRA and fused with ResNet costs, produce a point tracker that matches or beats a real-data-trained CoTracker3 on hard benchmarks.
-
BleedOrigin: Dynamic Bleeding Source Localization in Endoscopic Submucosal Dissection via Dual-Stage Detection and Tracking
A new ESD bleeding-source dataset and a dual-stage detection-tracking framework report 96.85% onset, 70.24% source, and 96.11% tracking accuracy within defined tolerances.
-
Physics-Grounded Motion Forecasting via Equation Discovery for Trajectory-Guided Image-to-Video Generation
A retrieval-initialized symbolic regression method discovers equations of motion from video trajectories and uses them to guide image-to-video generation, improving physical alignment on classical mechanics scenes.
-
Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation
Tora2 adds decoupled personalization embeddings, gated self-attention binding, and contrastive learning to Tora, enabling simultaneous appearance and trajectory customization for multiple entities in generated video.
-
DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth Estimation
A training-free diffusion-guidance framework that couples scale alignment across windows and geometric multi-view constraints inside the denoising loop yields more scale- and geometry-consistent depth for long videos.
-
Generative 4D Scene Gaussian Splatting with Object View-Synthesis Priors
A test-time optimization method that jointly fits deformable per-object 3D Gaussians with object-centric diffusion priors to generate 4D scenes and point tracks from monocular multi-object videos.
-
3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model
A diffusion world model predicts 3D optical flow as an embodiment-agnostic action plan, and constrained optimization converts the flow into robot arm actions.
-
An Updated SynthPop Model for Microlensing Simulations I: Model Description & Evaluation
An updated SynthPop model matches most bulge stellar and kinematic data but overpredicts optical microlensing event rates by about 20 percent near the galactic plane.
-
SAM2Auto: Auto Annotation Using FLASH
SAM2Auto combines four existing vision models into a no-training video annotation pipeline, but its automatic annotations are substantially less accurate than semi-supervised baselines and the paper contains contradic...
-
Online Long-term Point Tracking in the Foundation Model Era
A frame-by-frame point tracker with spatial and context memory reaches accuracy comparable to offline trackers on seven video benchmarks, making online long-term point tracking feasible.
Discussion (0). Continue with ORCID to comment.