REVIEW 15 cited by
Track Anything: Segment Anything Meets Videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recently, the Segment Anything Model (SAM) gains lots of attention rapidly due to its impressive segmentation performance on images. Regarding its strong ability on image segmentation and high interactivity with different prompts, we found that it performs poorly on consistent segmentation in videos. Therefore, in this report, we propose Track Anything Model (TAM), which achieves high-performance interactive tracking and segmentation in videos. To be detailed, given a video sequence, only with very little human participation, i.e., several clicks, people can track anything they are interested in, and get satisfactory results in one-pass inference. Without additional training, such an interactive design performs impressively on video object tracking and segmentation. All resources are available on {https://github.com/gaomingqi/Track-Anything}. We hope this work can facilitate related research.
Forward citations
Cited by 15 Pith papers
-
VoCap: Video Object Captioning and Segmentation from Any Prompt
VoCap jointly performs promptable video object segmentation and object captioning, and introduces a 50k-video pseudo-caption dataset that improves both tasks.
-
Grouped Speculative Decoding for Autoregressive Image Generation
Accepting clusters of visually valid tokens during speculative decoding yields about 3.7x training-free speedup for autoregressive image generation with quality preserved.
-
Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation
A new interactive method lets users click on an object in a video and generates audio for just that object, using mask-conditioned contrastive fine-tuning and latent diffusion.
-
ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular Inputs
A diffusion-enhanced 4D reconstruction pipeline for monocular video that achieves state-of-the-art results on DyCheck by supervising Gaussian splatting with personalized diffusion-generated pseudo-views.
-
R3eVision: A Survey on Robust Rendering, Restoration, and Enhancement for 3D Low-Level Vision
The survey formalizes degradation-aware rendering for 3D Low-Level Vision and organizes roughly 100 methods on super-resolution, deblurring, weather removal, restoration, and enhancement in NeRF and 3DGS pipelines.
-
SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost
SAM-I2V upgrades SAM to video segmentation with three lightweight modules (temporal integrator, selective memory, memory prompts), reaching about 90% of SAM 2.1's average J&F at 0.2% of its training cost.
-
CU-Multi: A Dataset for Multi-Robot Data Association
CU-Multi is a new public multi-robot dataset with controlled trajectory overlaps, dense semantic LiDAR labels, and geospatially aligned poses for evaluating data association in collaborative SLAM.
-
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
A lightweight point decoder plus task-prompt joint training yields ~3× faster SAM-based hierarchical text detection with competitive HierText results and +11% average F-score on three single-level benchmarks.
-
3D Gaussian Representations with Motion Trajectory Field for Dynamic Scene Reconstruction
A 3D Gaussian Splatting model whose Gaussian centers are represented as a learned combination of shared global motion bases recovers dynamic scenes and motion trajectories from monocular video.
-
Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation
MA-SAM2 adds context-aware and occlusion-resilient memory to SAM2 and reports Challenge IoU of 62.49 percent on EndoVis2017 and 64.40 percent on EndoVis2018, beating SAM2 by 6.10 and 4.36 points.
-
CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning
A CLIP-based encoder with RL residual refinement and curriculum learning reaches 81% mIoU on EndoVis 2018 and 74.12% on EndoVis 2017 surgical segmentation.
-
UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References
UA-Pose estimates 6D poses from partial object references by labeling seen and unseen model regions, using that uncertainty to filter poses and trigger online 3D completion.
-
UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting
UAV4D reconstructs 4D scenes from monocular drone video by fitting a single global scale to align human meshes with the background mesh, then renders with separate Gaussian splats.
-
SplitGaussian: Reconstructing Dynamic Scenes via Visual Geometry Decomposition
SplitGaussian reconstructs dynamic 3D scenes from monocular video by decomposing Gaussians into a rigid static branch and a deformable dynamic branch, claiming better motion separation and rendering quality than prior...
-
Continuous Marine Tracking via Autonomous UAV Handoff
A two-drone shark-tracking system using OSTrack and ORB feature matching is reported, but the headline handoff result comes from template matching in a simulated environment, not from a real inter-UAV flight.
Discussion (0). Sign in to comment.