Pith. sign in

REVIEW 8 cited by

Computer-Vision Benchmark Segment-Anything Model (SAM) in Medical Images: Accuracy in 12 Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.09324 v3 pith:WQFYIL6T submitted 2023-04-18 eess.IV cs.CV

classification eess.IVcs.CV
keywords imagesegmentationmedicalaccuracydiceimageswerecontrast
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Background: The segment-anything model (SAM), introduced in April 2023, shows promise as a benchmark model and a universal solution to segment various natural images. It comes without previously-required re-training or fine-tuning specific to each new dataset. Purpose: To test SAM's accuracy in various medical image segmentation tasks and investigate potential factors that may affect its accuracy in medical images. Methods: SAM was tested on 12 public medical image segmentation datasets involving 7,451 subjects. The accuracy was measured by the Dice overlap between the algorithm-segmented and ground-truth masks. SAM was compared with five state-of-the-art algorithms specifically designed for medical image segmentation tasks. Associations of SAM's accuracy with six factors were computed, independently and jointly, including segmentation difficulties as measured by segmentation ability score and by Dice overlap in U-Net, image dimension, size of the target region, image modality, and contrast. Results: The Dice overlaps from SAM were significantly lower than the five medical-image-based algorithms in all 12 medical image segmentation datasets, by a margin of 0.1-0.5 and even 0.6-0.7 Dice. SAM-Semantic was significantly associated with medical image segmentation difficulty and the image modality, and SAM-Point and SAM-Box were significantly associated with image segmentation difficulty, image dimension, target region size, and target-vs-background contrast. All these 3 variations of SAM were more accurate in 2D medical images, larger target region sizes, easier cases with a higher Segmentation Ability score and higher U-Net Dice, and higher foreground-background contrast.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 65 citations worldwide. Full citation record

  1. ZISVFM: Zero-Shot Object Instance Segmentation in Indoor Robotic Environments with Vision Foundation Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A zero-shot pipeline using SAM on colorized depth images, entropy-weighted DINOv2 attention filtering, and K-Medoids point prompts accurately segments unseen objects in cluttered indoor robot environments.

  2. Sli2Vol+: Segmenting 3D Medical Images Based on an Object Estimation Guided Correspondence Flow Network

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Sli2Vol+ learns slice-to-slice correspondences with guidance from pseudo-labels, improving single-slice-annotated 3D segmentation over the Sli2Vol baseline by about 5.6 Dice points on CT and 3.5 on MRI.

  3. BAP-MOS: Bandit-Based Adaptive Prompting for Boundary-Sensitive Multi-Organ Segmentation

    cs.CV 2026-08 reject novelty 5.0 of 10

    A bandit-based prompt-selection framework improves boundary metrics over fixed prompting in ultrasound segmentation, but the reported evaluation may rely on ground-truth prompts at test time.

  4. Industrial Synthetic Segment Pre-training

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Formula-generated hollow masks (InsCore) outperform real-image pre-training datasets and fine-tuned SAM on five industrial instance segmentation benchmarks using only 100,000 synthetic images.

  5. PGP-SAM: Prototype-Guided Prompt Learning for Efficient Few-Shot Medical Image Segmentation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    Prototype-guided prompt learning lets a SAM variant reach 78.75% mean Dice on Synapse and 76.39% on a ventricle dataset using 10% of training slices, beating SAMed and other prompt-free SAM baselines.

  6. ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements

    cs.CV 2024-11 conditional novelty 5.0 of 10

    ITACLIP combines modified CLIP attention, image augmentations, and LLM-generated class descriptions to achieve state-of-the-art training-free semantic segmentation on five benchmarks.

  7. Optimizing Prompt Strategies for SAM: Advancing lesion Segmentation Across Diverse Medical Imaging Modalities

    eess.IV 2024-12 conditional novelty 4.0 of 10

    SAM tumor outlining improves with more and non-central prompt points up to a plateau, and a DQN-based agent can pick effective points faster than human radiologists.

  8. Advanced Lung Nodule Segmentation and Classification for Early Detection of Lung Cancer using SAM and Transfer Learning

    eess.IV 2024-12 reject novelty 3.0 of 10

    SAM with bounding-box prompts and transfer learning is reported to segment lung nodules with 97.08% Dice and classify malignancy with 96.71% accuracy, but the ground truth masks are generated by a geometric algorithm,...

Pith tools