Pith. sign in

REVIEW 15 cited by

Track Anything: Segment Anything Meets Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.11968 v2 pith:K2WUEM3B submitted 2023-04-24 cs.CV

classification cs.CV
keywords anythingsegmentationtrackvideosinteractivemodelperformssegment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, the Segment Anything Model (SAM) gains lots of attention rapidly due to its impressive segmentation performance on images. Regarding its strong ability on image segmentation and high interactivity with different prompts, we found that it performs poorly on consistent segmentation in videos. Therefore, in this report, we propose Track Anything Model (TAM), which achieves high-performance interactive tracking and segmentation in videos. To be detailed, given a video sequence, only with very little human participation, i.e., several clicks, people can track anything they are interested in, and get satisfactory results in one-pass inference. Without additional training, such an interactive design performs impressively on video object tracking and segmentation. All resources are available on {https://github.com/gaomingqi/Track-Anything}. We hope this work can facilitate related research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 95 citations worldwide. Full citation record

  1. VoCap: Video Object Captioning and Segmentation from Any Prompt

    cs.CV 2025-08 conditional novelty 6.0 of 10

    VoCap jointly performs promptable video object segmentation and object captioning, and introduces a 50k-video pseudo-caption dataset that improves both tasks.

  2. Grouped Speculative Decoding for Autoregressive Image Generation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Accepting clusters of visually valid tokens during speculative decoding yields about 3.7x training-free speedup for autoregressive image generation with quality preserved.

  3. Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new interactive method lets users click on an object in a video and generates audio for just that object, using mask-conditioned contrastive fine-tuning and latent diffusion.

  4. ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular Inputs

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A diffusion-enhanced 4D reconstruction pipeline for monocular video that achieves state-of-the-art results on DyCheck by supervising Gaussian splatting with personalized diffusion-generated pseudo-views.

  5. R3eVision: A Survey on Robust Rendering, Restoration, and Enhancement for 3D Low-Level Vision

    cs.CV 2025-06 conditional novelty 6.0 of 10

    The survey formalizes degradation-aware rendering for 3D Low-Level Vision and organizes roughly 100 methods on super-resolution, deblurring, weather removal, restoration, and enhancement in NeRF and 3DGS pipelines.

  6. SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SAM-I2V upgrades SAM to video segmentation with three lightweight modules (temporal integrator, selective memory, memory prompts), reaching about 90% of SAM 2.1's average J&F at 0.2% of its training cost.

  7. CU-Multi: A Dataset for Multi-Robot Data Association

    cs.RO 2025-05 conditional novelty 6.0 of 10

    CU-Multi is a new public multi-robot dataset with controlled trajectory overlaps, dense semantic LiDAR labels, and geospatially aligned poses for evaluating data association in collaborative SLAM.

  8. ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis

    cs.CV 2026-03 accept novelty 5.0 of 10

    A lightweight point decoder plus task-prompt joint training yields ~3× faster SAM-based hierarchical text detection with competitive HierText results and +11% average F-score on three single-level benchmarks.

  9. 3D Gaussian Representations with Motion Trajectory Field for Dynamic Scene Reconstruction

    cs.RO 2025-08 conditional novelty 5.0 of 10

    A 3D Gaussian Splatting model whose Gaussian centers are represented as a learned combination of shared global motion bases recovers dynamic scenes and motion trajectories from monocular video.

  10. Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    MA-SAM2 adds context-aware and occlusion-resilient memory to SAM2 and reports Challenge IoU of 62.49 percent on EndoVis2017 and 64.40 percent on EndoVis2018, beating SAM2 by 6.10 and 4.36 points.

  11. CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning

    eess.IV 2025-07 conditional novelty 5.0 of 10

    A CLIP-based encoder with RL residual refinement and curriculum learning reaches 81% mIoU on EndoVis 2018 and 74.12% on EndoVis 2017 surgical segmentation.

  12. UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References

    cs.CV 2025-06 conditional novelty 5.0 of 10

    UA-Pose estimates 6D poses from partial object references by labeling seen and unseen model regions, using that uncertainty to filter poses and trigger online 3D completion.

  13. UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    UAV4D reconstructs 4D scenes from monocular drone video by fitting a single global scale to align human meshes with the background mesh, then renders with separate Gaussian splats.

  14. SplitGaussian: Reconstructing Dynamic Scenes via Visual Geometry Decomposition

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    SplitGaussian reconstructs dynamic 3D scenes from monocular video by decomposing Gaussians into a rigid static branch and a deformable dynamic branch, claiming better motion separation and rendering quality than prior...

  15. Continuous Marine Tracking via Autonomous UAV Handoff

    cs.CV 2025-07 reject novelty 4.0 of 10

    A two-drone shark-tracking system using OSTrack and ORB feature matching is reported, but the headline handoff result comes from template matching in a simulated environment, not from a real inter-UAV flight.

Pith tools