Pith. sign in

REVIEW 7 cited by

A Survey on Segment Anything Model (SAM): Vision Foundation Model Meets Prompt Engineering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.06211 v4 pith:BZ5D5P45 submitted 2023-05-12 cs.CV

classification cs.CV
keywords modelsurveyadvancementsanythingapplicationsgranularityincludingresearch
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The Segment Anything Model (SAM), developed by Meta AI Research, represents a significant breakthrough in computer vision, offering a robust framework for image and video segmentation. This survey provides a comprehensive exploration of the SAM family, including SAM and SAM 2, highlighting their advancements in granularity and contextual understanding. Our study demonstrates SAM's versatility across a wide range of applications while identifying areas where improvements are needed, particularly in scenarios requiring high granularity and in the absence of explicit prompts. By mapping the evolution and capabilities of SAM models, we offer insights into their strengths and limitations and suggest future research directions, including domain-specific adaptations and enhanced memory and propagation mechanisms. We believe that this survey comprehensively covers the breadth of SAM's applications and challenges, setting the stage for ongoing advancements in segmentation technology.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OGG-FR: Orthogonal Gradient Gaming and Frequency Rectification for Unmanned Aerial Vehicle Infrared Image Super-Resolution

    cs.CV 2026-08 conditional novelty 6.0 of 10

    OGG-FR is a plug-and-play training update that separates redundant and innovative parts of the FFT loss gradient and gates the innovative part by a confidence score, improving UAV infrared super-resolution in most tes...

  2. SCOPE: Speech-guided COllaborative PErception Framework for Surgical Scene Segmentation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A speech-guided framework uses an LLM and open-set vision models to segment and track surgical instruments and anatomy hands-free in live video.

  3. SAM2-SGP: Enhancing SAM2 for Medical Image Segmentation via Support-Set Guided Prompting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SAM2-SGP automatically prompts SAM2 using support-set-derived pseudo-masks and achieves higher Dice scores than nnUNet, SwinUNet, SAM2, and MedSAM2 across eight medical datasets.

  4. PicoSAM3: Real-Time In-Sensor Region-of-Interest Segmentation

    cs.CV 2026-03 conditional novelty 5.0 of 10

    A 1.3M-parameter CNN with ROI-implicit prompting and SAM3 distillation reaches ~65% mIoU on COCO/LVIS and 11.82 ms INT8 inference fully in-sensor on the Sony IMX500.

  5. AoP-SAM: Automation of Prompts for Efficient Segmentation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    AoP-SAM trains a lightweight prompt predictor on SAM's own image embeddings to emit a point-prompt confidence map, then filters redundant prompts at test time, improving automatic segmentation accuracy and efficiency.

  6. Vision and Language Reference Prompt into SAM for Few-shot Segmentation

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Using a frozen vision-language model, VLP-SAM injects text-label semantics into SAM's prompt encoder and raises one-shot segmentation mIoU by 6.3 points on PASCAL-5i and 9.5 on COCO-20i.

  7. BrainSegDMlF: A Dynamic Fusion-enhanced SAM for Brain Lesion Segmentation

    cs.CV 2025-05 conditional novelty 4.0 of 10

    An SAM-based network with dynamic multimodal fusion and a multi-scale upsampling decoder reports the best Dice on BraTS2021 and FCD2023, but the comparison baselines were limited to a single MRI modality.

Pith tools