Pith. sign in

REVIEW 8 cited by

Semantic-SAM: Segment and Recognize Anything at Any Granularity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.04767 v1 pith:DGULUZIM submitted 2023-07-10 cs.CV

classification cs.CV
keywords modelsegmentationmultiplesemantic-awarenessanythingdatasetsgranularitygranularity-abundance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we introduce Semantic-SAM, a universal image segmentation model to enable segment and recognize anything at any desired granularity. Our model offers two key advantages: semantic-awareness and granularity-abundance. To achieve semantic-awareness, we consolidate multiple datasets across three granularities and introduce decoupled classification for objects and parts. This allows our model to capture rich semantic information. For the multi-granularity capability, we propose a multi-choice learning scheme during training, enabling each click to generate masks at multiple levels that correspond to multiple ground-truth masks. Notably, this work represents the first attempt to jointly train a model on SA-1B, generic, and part segmentation datasets. Experimental results and visualizations demonstrate that our model successfully achieves semantic-awareness and granularity-abundance. Furthermore, combining SA-1B training with other segmentation tasks, such as panoptic and part segmentation, leads to performance improvements. We will provide code and a demo for further exploration and evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 51 citations worldwide. Full citation record

  1. Synthetic Visual Genome

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A GPT-4V/GPT-4o pipeline for completing and refining scene graph annotations yields a dense synthetic dataset that, after instruction tuning, gives a 3B model strong relationship understanding and grounding results.

  2. G$^2$TAM: Geometry Grounded Track Anything Model

    cs.CV 2026-07 accept novelty 6.5 of 10

    Spatially aligned geometric features serve as implicit memory so one model reconstructs scenes and produces promptable, cross-view consistent instance masks from unordered RGB only.

  3. PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    PASG automatically extracts object keypoints and axes and couples them through a fine-tuned vision-language model to task semantics, claiming manipulation performance comparable to manual annotations.

  4. Training-free Geometric Image Editing on Diffusion Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    FreeFine splits geometric image editing into object transformation, source-region inpainting, and target refinement, using temporal attention, local noise, and text guidance in a training-free way.

  5. One Polyp Identifies All: One-Shot Polyp Segmentation with SAM via Cascaded Priors and Iterative Prompt Evolution

    cs.CV 2025-07 conditional novelty 6.0 of 10

    OP-SAM turns one labeled polyp image into iterative SAM prompts, reaching 76.93% IoU on Kvasir with no retraining.

  6. Integrating SAM Supervision for 3D Weakly Supervised Point Cloud Segmentation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A weakly supervised 3D point cloud segmentation method that back-projects Semantic-SAM 2D masks into 3D, propagates sparse labels inside masks, and uses reliability-filtered pseudo labels, reporting state-of-the-art m...

  7. FusionForce: End-to-end Differentiable Neural-Symbolic Layer for Trajectory Prediction

    cs.RO 2025-02 conditional novelty 5.0 of 10

    FusionForce predicts robot trajectories by learning terrain properties from camera and lidar, then simulating them through a differentiable rigid-body physics engine, cutting trajectory error versus LSTM baselines by ...

  8. SAM-MI: A Mask-Injected Framework for Enhancing Open-Vocabulary Semantic Segmentation with SAM

    cs.CV 2025-11 conditional novelty 4.0 of 10

    SAM-MI improves open-vocabulary segmentation by injecting aggregated SAM masks as low- and high-frequency guidance into CLIP cost maps, with sparse text-guided point prompts for speed.

Pith tools