REVIEW 21 cited by
Segment and Track Anything
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This report presents a framework called Segment And Track Anything (SAMTrack) that allows users to precisely and effectively segment and track any object in a video. Additionally, SAM-Track employs multimodal interaction methods that enable users to select multiple objects in videos for tracking, corresponding to their specific requirements. These interaction methods comprise click, stroke, and text, each possessing unique benefits and capable of being employed in combination. As a result, SAM-Track can be used across an array of fields, ranging from drone technology, autonomous driving, medical imaging, augmented reality, to biological analysis. SAM-Track amalgamates Segment Anything Model (SAM), an interactive key-frame segmentation model, with our proposed AOT-based tracking model (DeAOT), which secured 1st place in four tracks of the VOT 2022 challenge, to facilitate object tracking in video. In addition, SAM-Track incorporates Grounding-DINO, which enables the framework to support text-based interaction. We have demonstrated the remarkable capabilities of SAM-Track on DAVIS-2016 Val (92.0%), DAVIS-2017 Test (79.2%)and its practicability in diverse applications. The project page is available at: https://github.com/z-x-yang/Segment-and-Track-Anything.
Forward citations
Cited by 21 Pith papers
-
LayerFlow: A Unified Model for Layer-aware Video Generation
LayerFlow is a unified diffusion-transformer model that generates transparent foreground, background, and blended video layers from per-layer prompts, and supports decomposition and conditioned generation in one framework.
-
CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation
CROSS improves remote sensing referring segmentation by combining cascaded SAM distillation with contrastive learning, reporting state-of-the-art cIoU on RefSegRS and RRSIS-D.
-
Grasp Like Humans: Learning Generalizable Multi-Fingered Grasping from Human Proprioceptive Sensorimotor Integration
A glove-based system learns human grasp demonstrations from joint angles and contact forces, then controls various robotic hands to grasp diverse objects without vision or retraining.
-
Grasp-MPC: Closed-Loop Visual Grasping via Value-Guided Model Predictive Control
A value-guided MPC policy trained on 2 million synthetic trajectories improves closed-loop 6-DoF grasping in clutter and adapts to object perturbations.
-
VoCap: Video Object Captioning and Segmentation from Any Prompt
VoCap jointly performs promptable video object segmentation and object captioning, and introduces a 50k-video pseudo-caption dataset that improves both tasks.
-
Grouped Speculative Decoding for Autoregressive Image Generation
Accepting clusters of visually valid tokens during speculative decoding yields about 3.7x training-free speedup for autoregressive image generation with quality preserved.
-
High-fidelity 3D Gaussian Inpainting: preserving multi-view consistency and photorealistic details
A 3D Gaussian inpainting framework with automatic mask refinement and depth-initialized uncertainty weighting balances multi-view consistency and visual detail, reporting the best LPIPS on the SPIn-NeRF dataset.
-
ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation
A few-shot segmentation framework that injects reference-image prototypes into SAM's decoder and image encoder, eliminating per-image manual prompts and improving remote sensing segmentation accuracy.
-
STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing
STR-Match edits videos without retraining by matching source and target 'spatiotemporal relevance scores' derived from attention maps during latent optimization.
-
SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost
SAM-I2V upgrades SAM to video segmentation with three lightweight modules (temporal integrator, selective memory, memory prompts), reaching about 90% of SAM 2.1's average J&F at 0.2% of its training cost.
-
Hierarchical Instruction-aware Embodied Visual Tracking
HIEVT uses an LLM to convert natural language instructions into spatial goals (bounding boxes) and an offline RL policy to track targets to those goals, claiming strong generalization across environments.
-
COMBO-Grasp: Learning Constraint-Based Manipulation for Bimanual Occluded Grasping
COMBO-Grasp trains a stabilizing constraint policy and an RL grasping policy, then refines the constraint pose with value-function gradients, improving bimanual grasping of occluded objects in simulation and real world.
-
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
A lightweight point decoder plus task-prompt joint training yields ~3× faster SAM-based hierarchical text detection with competitive HierText results and +11% average F-score on three single-level benchmarks.
-
ViSTR-GP: Online Cyberattack Detection via Vision-to-State Tensor Regression and Gaussian Processes in Automated Robotic Operations
ViSTR-GP uses an overhead camera, a learned vision-to-joint-angle map, and a Gaussian-process residual test to detect replay attacks on industrial robots from small physical deviations.
-
DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing
DreamSwapV performs mask-guided, subject-agnostic subject swapping in videos using multiple conditions, an adaptive mask strategy, and a new benchmark.
-
Representative Volume Element: Existence and Extent in Cracked Heterogeneous Medium
Modified periodic boundary conditions that add strain periodicity to displacement periodicity are claimed to reduce mesh and size sensitivity in cracked-composite RVE simulations, tested on 1,200 samples.
-
CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios
CrowdTrack is a dense, first-person-view pedestrian tracking benchmark that exposes large performance drops in existing multi-object trackers.
-
CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation
A text-guided SAM2 variant with cross-modal attention, semantic prompt generation, and a similarity-sorted memory bank achieves top Dice and surface scores on seven public multi-organ CT datasets.
-
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
PAM extends SAM 2 with a frozen LLM and a Semantic Perceiver to jointly segment and describe regions in images, videos, and streaming video, and contributes a 0.6M-sample region-level streaming video caption dataset.
-
Efficient Segment Anything with Depth-Aware Fusion and Limited Training Data
Adding monocular depth to EfficientViT-SAM improves point-prompted segmentation at 3 and 5 clicks after fine-tuning on 11.2k images, but universal gains and data-efficiency are not established.
-
A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects
A comprehensive survey of video scene parsing that organizes methods, datasets, metrics, and benchmark results across VSS, VIS, VPS, VTS, and OVVS.
Discussion (0). Continue with ORCID to comment.