REVIEW 8 cited by
Semantic-SAM: Segment and Recognize Anything at Any Granularity
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we introduce Semantic-SAM, a universal image segmentation model to enable segment and recognize anything at any desired granularity. Our model offers two key advantages: semantic-awareness and granularity-abundance. To achieve semantic-awareness, we consolidate multiple datasets across three granularities and introduce decoupled classification for objects and parts. This allows our model to capture rich semantic information. For the multi-granularity capability, we propose a multi-choice learning scheme during training, enabling each click to generate masks at multiple levels that correspond to multiple ground-truth masks. Notably, this work represents the first attempt to jointly train a model on SA-1B, generic, and part segmentation datasets. Experimental results and visualizations demonstrate that our model successfully achieves semantic-awareness and granularity-abundance. Furthermore, combining SA-1B training with other segmentation tasks, such as panoptic and part segmentation, leads to performance improvements. We will provide code and a demo for further exploration and evaluation.
Forward citations
Cited by 8 Pith papers
-
Synthetic Visual Genome
A GPT-4V/GPT-4o pipeline for completing and refining scene graph annotations yields a dense synthetic dataset that, after instruction tuning, gives a 3B model strong relationship understanding and grounding results.
-
G$^2$TAM: Geometry Grounded Track Anything Model
Spatially aligned geometric features serve as implicit memory so one model reconstructs scenes and produces promptable, cross-view consistent instance masks from unordered RGB only.
-
PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation
PASG automatically extracts object keypoints and axes and couples them through a fine-tuned vision-language model to task semantics, claiming manipulation performance comparable to manual annotations.
-
Training-free Geometric Image Editing on Diffusion Models
FreeFine splits geometric image editing into object transformation, source-region inpainting, and target refinement, using temporal attention, local noise, and text guidance in a training-free way.
-
One Polyp Identifies All: One-Shot Polyp Segmentation with SAM via Cascaded Priors and Iterative Prompt Evolution
OP-SAM turns one labeled polyp image into iterative SAM prompts, reaching 76.93% IoU on Kvasir with no retraining.
-
Integrating SAM Supervision for 3D Weakly Supervised Point Cloud Segmentation
A weakly supervised 3D point cloud segmentation method that back-projects Semantic-SAM 2D masks into 3D, propagates sparse labels inside masks, and uses reliability-filtered pseudo labels, reporting state-of-the-art m...
-
FusionForce: End-to-end Differentiable Neural-Symbolic Layer for Trajectory Prediction
FusionForce predicts robot trajectories by learning terrain properties from camera and lidar, then simulating them through a differentiable rigid-body physics engine, cutting trajectory error versus LSTM baselines by ...
-
SAM-MI: A Mask-Injected Framework for Enhancing Open-Vocabulary Semantic Segmentation with SAM
SAM-MI improves open-vocabulary segmentation by injecting aggregated SAM masks as low- and high-frequency guidance into CLIP cost maps, with sparse text-guided point prompts for speed.
Discussion (0). Continue with ORCID to comment.