Pith. sign in

REVIEW 25 cited by

Medical SAM 2: Segment medical images as video via Segment Anything Model 2

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.00874 v2 pith:2S3MDCVF submitted 2024-08-01 cs.CV

classification cs.CV
keywords medicalsegmentationimageimagesmedsam-2modelmodelssegment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Medical image segmentation plays a pivotal role in clinical diagnostics and treatment planning, yet existing models often face challenges in generalization and in handling both 2D and 3D data uniformly. In this paper, we introduce Medical SAM 2 (MedSAM-2), a generalized auto-tracking model for universal 2D and 3D medical image segmentation. The core concept is to leverage the Segment Anything Model 2 (SAM2) pipeline to treat all 2D and 3D medical segmentation tasks as a video object tracking problem. To put it into practice, we propose a novel \emph{self-sorting memory bank} mechanism that dynamically selects informative embeddings based on confidence and dissimilarity, regardless of temporal order. This mechanism not only significantly improves performance in 3D medical image segmentation but also unlocks a \emph{One-Prompt Segmentation} capability for 2D images, allowing segmentation across multiple images from a single prompt without temporal relationships. We evaluated MedSAM-2 on five 2D tasks and nine 3D tasks, including white blood cells, optic cups, retinal vessels, mandibles, coronary arteries, kidney tumors, liver tumors, breast cancer, nasopharynx cancer, vestibular schwannoma, mediastinal lymph nodules, cerebral artery, inferior alveolar nerve, and abdominal organs, comparing it against state-of-the-art (SOTA) models in task-tailored, general and interactive segmentation settings. Our findings demonstrate that MedSAM-2 surpasses a wide range of existing models and updates new SOTA on several benchmarks. The code is released on the project page: https://supermedintel.github.io/Medical-SAM2/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models via Depth-Routed LoRA and Depth Shifting

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Depth-routed LoRA and a depth-shift module lift frozen SAM and SAM2 to 3D and 3D+T segmentation using less than ~3.7% trainable parameters.

  2. Do Medical Foundation Models Generalize on the African Brain?

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Medical foundation models show no consistent generalization gap on African brain MRI; performance differences track dataset size, not data origin.

  3. SAMRI-3D: Adapting SAM2 for 3D MRI Segmentation with Global Volume Tokens

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Freezing SAM2's encoder, fine-tuning its decoder/memory, and adding TSDF-trained global volume tokens yields 0.78 mean Dice on a new 34-dataset MRI benchmark, up from 0.58 zero-shot.

  4. Higher-Order Cell Tracking Transformer

    cs.CV 2026-07 accept novelty 6.0 of 10

    An edge-centric Transformer with line-to-line geometric attention achieves SOTA cell lineage tracking without pretrained image encoders and fine-tunes far more efficiently than node-embedding baselines.

  5. Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt

    cs.CV 2025-10 unverdicted novelty 6.0 of 10

    Memory-SAM retrieves similar prior cases via DINOv3 features and FAISS to generate point prompts for SAM2, achieving mIoU 0.9863 on 600 tongue images without training or human prompts.

  6. Organoid Tracker: A SAM2-Powered Platform for Zero-shot Cyst Analysis in Human Kidney Organoid Videos

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A SAM2-based open-source GUI with inverse temporal tracking for quantifying kidney organoid cyst growth in bright-field videos, demonstrated on two videos with noted early-frame failures.

  7. Live(r) Die: Predicting Survival in Colorectal Liver Metastasis

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A fully automated pre/post-contrast MRI framework, combining prompt-based segmentation with autoencoder multiple-instance survival analysis, improves CRLM post-surgery survival prediction over clinical and genomic bio...

  8. HRVVS: A High-resolution Video Vasculature Segmentation Network via Hierarchical Autoregressive Residual Priors

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new dataset and network for segmenting hepatic vasculature in high-resolution hepatectomy videos, reporting the best scores on the new benchmark.

  9. Depthwise-Dilated Convolutional Adapters for Medical Object Tracking and Segmentation Using the Segment Anything Model 2

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DD-SAM2 uses a depthwise-dilated convolutional adapter to fine-tune SAM2 for medical video object tracking and segmentation, reporting Dice scores of 0.93 (MRI tumor) and 0.97 (ultrasound left ventricle).

  10. SafeClick: Error-Tolerant Interactive Segmentation of Any Medical Volumes via Hierarchical Expert Consensus

    eess.IV 2025-06 conditional novelty 6.0 of 10

    SafeClick adds a hierarchical expert consensus module to SAM 2 and MedSAM 2 that improves segmentation accuracy under imperfect prompts.

  11. Lean-SAM2: Target-Anchored Memory and Encoder Acceleration for SAM2

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Lean-SAM2 combines target-anchored memory pruning, condensed insurance memory, and risk-aware window routing to accelerate SAM2.1 inference ~1.4× with better accuracy than Efficient-SAM2.

  12. Enhancing MedSAM with a Lightweight Box Predictor for Medical Image Segmentation

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    Enhances MedSAM with a 1.6M-parameter Box Predictor trained in two stages to convert single clicks to bounding boxes, reporting Dice scores of 0.89-0.98 on four medical datasets across CT, MRI, and ultrasound.

  13. SurgSLOT: Segment Anything in Surgical Videos via Semantic Long-term Tracking

    cs.CV 2025-11 conditional novelty 5.0 of 10

    SAM2S, a SAM2 variant trained on the new 61k-frame SA-SV surgical benchmark, improves average J&F to 80.42 at 68 FPS for interactive surgical-video object segmentation.

  14. Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    MA-SAM2 adds context-aware and occlusion-resilient memory to SAM2 and reports Challenge IoU of 62.49 percent on EndoVis2017 and 64.40 percent on EndoVis2018, beating SAM2 by 6.10 and 4.36 points.

  15. Dual Semantic-Aware Network for Noise Suppressed Ultrasound Video Segmentation

    cs.CV 2025-07 reject novelty 5.0 of 10

    DSANet improves ultrasound video segmentation by aligning adjacent frames at the channel level and fusing local and global features, achieving about 1% higher Dice than prior methods.

  16. CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation

    eess.IV 2025-06 conditional novelty 5.0 of 10

    A text-guided SAM2 variant with cross-modal attention, semantic prompt generation, and a similarity-sorted memory bank achieves top Dice and surface scores on seven public multi-organ CT datasets.

  17. Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation

    cs.CV 2026-07 conditional novelty 4.0 of 10

    LoRA on SAM3’s prompt encoder, detector, and tracker (0.98% of parameters) raises surgical concept-segmentation mIoU over zero-shot SAM3 and Medical SAM3 while fitting in ~9 GB GPU memory.

  18. Robust Activation Map Rectification for Weakly Supervised Volumetric Segmentation: Temporal Coherence as a Free Lunch

    cs.CV 2026-07 conditional novelty 4.0 of 10

    CSSeg rectifies noisy CAMs in volumetric scans by cross-slice averaging and one-sided replacement of suspicious frames, then prompts MedSAM, reporting large but unverified Dice/mIoU gains.

  19. DivAS: Interactive 3D Segmentation by Depth-Weighted Voxel Aggregation

    cs.CV 2026-01 conditional novelty 4.0 of 10

    An optimization-free pipeline lifts Segment Anything masks into NeRF scenes via depth-weighted refinement and CUDA voxel aggregation, matching optimization-based segmentation accuracy on LLFF and Mip-NeRF360 at 2–2.5×...

  20. SAM-MI: A Mask-Injected Framework for Enhancing Open-Vocabulary Semantic Segmentation with SAM

    cs.CV 2025-11 conditional novelty 4.0 of 10

    SAM-MI improves open-vocabulary segmentation by injecting aggregated SAM masks as low- and high-frequency guidance into CLIP cost maps, with sparse text-guided point prompts for speed.

  21. Towards Affordable Tumor Segmentation and Visualization for 3D Breast MRI Using SAM2

    cs.CV 2025-07 conditional novelty 4.0 of 10

    SAM2 can segment breast tumors in 3D MRI with a single bounding-box prompt, and center-outward propagation yields the best volumetric Dice.

  22. RAPS-3D: Efficient interactive segmentation for 3D radiological imaging

    cs.CV 2025-07 conditional novelty 4.0 of 10

    RAPS-3D is a 3D promptable CT segmentation model that reports 86.8 Dice on AMOS-CT with a single 2D bounding-box prompt, using zoom-out/zoom-in inference with no sliding window.

  23. SAMed-2: Selective Memory Enhanced Medical Segment Anything Model

    cs.CV 2025-07 conditional novelty 4.0 of 10

    SAMed-2 combines a temporal adapter and confidence-filtered memory retrieval with SAM-2 to report state-of-the-art Dice scores on 21 medical segmentation tasks.

  24. SSS: Semi-Supervised SAM-2 with Efficient Prompting for Medical Imaging Segmentation

    cs.CV 2025-06 reject novelty 4.0 of 10

    SSS applies SAM-2 with a Discriminative Feature Enhancement mechanism and a physical-constraint sliding-window prompt generator, reporting Dice scores of 53.15 on BHSD and 89.34 to 91.21 on ACDC.

  25. F3-Net: Foundation Model for Full Abnormality Segmentation of Medical Images with Flexible Input Modality Requirement

    cs.CV 2025-07 reject novelty 3.0 of 10

    F3-Net combines multi-encoder nnU-Net with zero-filled missing modalities to segment glioma, metastasis, stroke, and white matter lesions, but the missing-modality claim is untested and comparisons are incomplete.

Pith tools