Pith. sign in

REVIEW 2 cited by

Segment Anything Model for Medical Images?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.14660 v7 pith:IBLB6OVX submitted 2023-04-28 eess.IV cs.CVcs.LG

classification eess.IVcs.CVcs.LG
keywords performancemedicalsegmentationbetterimagemodelanythingcomplex
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The Segment Anything Model (SAM) is the first foundation model for general image segmentation. It has achieved impressive results on various natural image segmentation tasks. However, medical image segmentation (MIS) is more challenging because of the complex modalities, fine anatomical structures, uncertain and complex object boundaries, and wide-range object scales. To fully validate SAM's performance on medical data, we collected and sorted 53 open-source datasets and built a large medical segmentation dataset with 18 modalities, 84 objects, 125 object-modality paired targets, 1050K 2D images, and 6033K masks. We comprehensively analyzed different models and strategies on the so-called COSMOS 1050K dataset. Our findings mainly include the following: 1) SAM showed remarkable performance in some specific objects but was unstable, imperfect, or even totally failed in other situations. 2) SAM with the large ViT-H showed better overall performance than that with the small ViT-B. 3) SAM performed better with manual hints, especially box, than the Everything mode. 4) SAM could help human annotation with high labeling quality and less time. 5) SAM was sensitive to the randomness in the center point and tight box prompts, and may suffer from a serious performance drop. 6) SAM performed better than interactive methods with one or a few points, but will be outpaced as the number of points increases. 7) SAM's performance correlated to different factors, including boundary complexity, intensity differences, etc. 8) Finetuning the SAM on specific medical tasks could improve its average DICE performance by 4.39% and 6.68% for ViT-B and ViT-H, respectively. We hope that this comprehensive report can help researchers explore the potential of SAM applications in MIS, and guide how to appropriately use and develop SAM.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 40 citations worldwide. Full citation record

  1. XBusNet: Text-Guided Breast Ultrasound Segmentation via Multimodal Vision-Language Learning

    cs.CV 2025-09 conditional novelty 6.0 of 10

    XBusNet combines CLIP text prompts and a U-Net to segment breast ultrasound lesions, achieving Dice 0.877 and IoU 0.815 on BLU, outperforming six baselines.

  2. TextureSAM: Towards a Texture Aware Foundation Model for Segmentation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Fine-tuning SAM-2 on texture-augmented ADE20K shifts segmentation toward texture-defined boundaries, raising un-aggregated mIoU on natural and synthetic texture benchmarks.

Pith tools