Pith. sign in

REVIEW 20 cited by

Segment and Track Anything

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.06558 v1 pith:TZBTAI3X submitted 2023-05-11 cs.CV

classification cs.CV
keywords sam-tracksegmentanythinginteractionmodeltracktrackingframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This report presents a framework called Segment And Track Anything (SAMTrack) that allows users to precisely and effectively segment and track any object in a video. Additionally, SAM-Track employs multimodal interaction methods that enable users to select multiple objects in videos for tracking, corresponding to their specific requirements. These interaction methods comprise click, stroke, and text, each possessing unique benefits and capable of being employed in combination. As a result, SAM-Track can be used across an array of fields, ranging from drone technology, autonomous driving, medical imaging, augmented reality, to biological analysis. SAM-Track amalgamates Segment Anything Model (SAM), an interactive key-frame segmentation model, with our proposed AOT-based tracking model (DeAOT), which secured 1st place in four tracks of the VOT 2022 challenge, to facilitate object tracking in video. In addition, SAM-Track incorporates Grounding-DINO, which enables the framework to support text-based interaction. We have demonstrated the remarkable capabilities of SAM-Track on DAVIS-2016 Val (92.0%), DAVIS-2017 Test (79.2%)and its practicability in diverse applications. The project page is available at: https://github.com/z-x-yang/Segment-and-Track-Anything.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LayerFlow: A Unified Model for Layer-aware Video Generation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    LayerFlow is a unified diffusion-transformer model that generates transparent foreground, background, and blended video layers from per-layer prompts, and supports decomposition and conditioned generation in one framework.

  2. CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    CROSS improves remote sensing referring segmentation by combining cascaded SAM distillation with contrastive learning, reporting state-of-the-art cIoU on RefSegRS and RRSIS-D.

  3. Grasp Like Humans: Learning Generalizable Multi-Fingered Grasping from Human Proprioceptive Sensorimotor Integration

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A glove-based system learns human grasp demonstrations from joint angles and contact forces, then controls various robotic hands to grasp diverse objects without vision or retraining.

  4. Grasp-MPC: Closed-Loop Visual Grasping via Value-Guided Model Predictive Control

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A value-guided MPC policy trained on 2 million synthetic trajectories improves closed-loop 6-DoF grasping in clutter and adapts to object perturbations.

  5. VoCap: Video Object Captioning and Segmentation from Any Prompt

    cs.CV 2025-08 conditional novelty 6.0 of 10

    VoCap jointly performs promptable video object segmentation and object captioning, and introduces a 50k-video pseudo-caption dataset that improves both tasks.

  6. Grouped Speculative Decoding for Autoregressive Image Generation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Accepting clusters of visually valid tokens during speculative decoding yields about 3.7x training-free speedup for autoregressive image generation with quality preserved.

  7. High-fidelity 3D Gaussian Inpainting: preserving multi-view consistency and photorealistic details

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 3D Gaussian inpainting framework with automatic mask refinement and depth-initialized uncertainty weighting balances multi-view consistency and visual detail, reporting the best LPIPS on the SPIn-NeRF dataset.

  8. ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A few-shot segmentation framework that injects reference-image prototypes into SAM's decoder and image encoder, eliminating per-image manual prompts and improving remote sensing segmentation accuracy.

  9. STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    STR-Match edits videos without retraining by matching source and target 'spatiotemporal relevance scores' derived from attention maps during latent optimization.

  10. SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SAM-I2V upgrades SAM to video segmentation with three lightweight modules (temporal integrator, selective memory, memory prompts), reaching about 90% of SAM 2.1's average J&F at 0.2% of its training cost.

  11. Hierarchical Instruction-aware Embodied Visual Tracking

    cs.CV 2025-05 conditional novelty 6.0 of 10

    HIEVT uses an LLM to convert natural language instructions into spatial goals (bounding boxes) and an offline RL policy to track targets to those goals, claiming strong generalization across environments.

  12. ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis

    cs.CV 2026-03 accept novelty 5.0 of 10

    A lightweight point decoder plus task-prompt joint training yields ~3× faster SAM-based hierarchical text detection with competitive HierText results and +11% average F-score on three single-level benchmarks.

  13. ViSTR-GP: Online Cyberattack Detection via Vision-to-State Tensor Regression and Gaussian Processes in Automated Robotic Operations

    cs.RO 2025-09 conditional novelty 5.0 of 10

    ViSTR-GP uses an overhead camera, a learned vision-to-joint-angle map, and a Gaussian-process residual test to detect replay attacks on industrial robots from small physical deviations.

  14. DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    DreamSwapV performs mask-guided, subject-agnostic subject swapping in videos using multiple conditions, an adaptive mask strategy, and a new benchmark.

  15. Representative Volume Element: Existence and Extent in Cracked Heterogeneous Medium

    cs.CE 2025-08 unverdicted novelty 5.0 of 10

    Modified periodic boundary conditions that add strain periodicity to displacement periodicity are claimed to reduce mesh and size sensitivity in cracked-composite RVE simulations, tested on 1,200 samples.

  16. CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios

    cs.CV 2025-07 conditional novelty 5.0 of 10

    CrowdTrack is a dense, first-person-view pedestrian tracking benchmark that exposes large performance drops in existing multi-object trackers.

  17. CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation

    eess.IV 2025-06 conditional novelty 5.0 of 10

    A text-guided SAM2 variant with cross-modal attention, semantic prompt generation, and a similarity-sorted memory bank achieves top Dice and surface scores on seven public multi-organ CT datasets.

  18. Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

    cs.CV 2025-06 reject novelty 5.0 of 10

    PAM extends SAM 2 with a frozen LLM and a Semantic Perceiver to jointly segment and describe regions in images, videos, and streaming video, and contributes a 0.6M-sample region-level streaming video caption dataset.

  19. Efficient Segment Anything with Depth-Aware Fusion and Limited Training Data

    cs.CV 2026-02 reject novelty 4.0 of 10

    Adding monocular depth to EfficientViT-SAM improves point-prompted segmentation at 3 and 5 clicks after fine-tuning on 11.2k images, but universal gains and data-efficiency are not established.

  20. A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects

    cs.CV 2025-06 conditional novelty 2.0 of 10

    A comprehensive survey of video scene parsing that organizes methods, datasets, metrics, and benchmark results across VSS, VIS, VPS, VTS, and OVVS.

Pith tools