REVIEW 12 cited by
SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The advent of large models, also known as foundation models, has significantly transformed the AI research landscape, with models like Segment Anything (SAM) achieving notable success in diverse image segmentation scenarios. Despite its advancements, SAM encountered limitations in handling some complex low-level segmentation tasks like camouflaged object and medical imaging. In response, in 2023, we introduced SAM-Adapter, which demonstrated improved performance on these challenging tasks. Now, with the release of Segment Anything 2 (SAM2), a successor with enhanced architecture and a larger training corpus, we reassess these challenges. This paper introduces SAM2-Adapter, the first adapter designed to overcome the persistent limitations observed in SAM2 and achieve new state-of-the-art (SOTA) results in specific downstream tasks including medical image segmentation, camouflaged (concealed) object detection, and shadow detection. SAM2-Adapter builds on the SAM-Adapter's strengths, offering enhanced generalizability and composability for diverse applications. We present extensive experimental results demonstrating SAM2-Adapter's effectiveness. We show the potential and encourage the research community to leverage the SAM2 model with our SAM2-Adapter for achieving superior segmentation outcomes. Code, pre-trained models, and data processing protocols are available at http://tianrun-chen.github.io/SAM-Adaptor/
Forward citations
Cited by 12 Pith papers
-
When Does Resolution Help a Frozen Backbone? Global Attention at Resolution Predicts Scalable Adaptation for Camouflaged and Marine Animal Segmentation
Global attention over a high-resolution token set, not capacity or pretraining, determines whether LoRA adapters convert resolution into accuracy on fine-grained segmentation.
-
Refining Context-Entangled Content Segmentation via Curriculum Selection and Anti-Curriculum Promotion
A curriculum-then-anti-curriculum training schedule, ending with spectral low-pass fine-tuning, improves context-entangled segmentation across several datasets and backbones.
-
Shape Distribution Matters: Shape-specific Mixture-of-Experts for Amodal Segmentation under Diverse Occlusions
ShapeMoE improves amodal segmentation by routing each object to a shape-specialized expert via a learned Gaussian shape distribution.
-
SAMwave: Wavelet-Driven Feature Enrichment for Effective Adaptation of Segment Anything Model
Wavelet high-frequency features and real or complex adapters improve SAM and SAM2 on several low-level vision tasks, though gains depend on the wavelet family.
-
SAM4D: Segment Anything in Camera and LiDAR Streams
SAM4D is a promptable model that segments and tracks objects across camera and LiDAR streams with cross-modal prompts, trained on pseudo-labels generated by an automated data engine.
-
SafeClick: Error-Tolerant Interactive Segmentation of Any Medical Volumes via Hierarchical Expert Consensus
SafeClick adds a hierarchical expert consensus module to SAM 2 and MedSAM 2 that improves segmentation accuracy under imperfect prompts.
-
SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost
SAM-I2V upgrades SAM to video segmentation with three lightweight modules (temporal integrator, selective memory, memory prompts), reaching about 90% of SAM 2.1's average J&F at 0.2% of its training cost.
-
Multimodal SAM-adapter for Semantic Segmentation
A side-tuning adapter injects RGB-plus-auxiliary-sensor fused features into SAM's encoder, reaching state-of-the-art semantic segmentation on DeLiVER, FMB, and MUSES.
-
CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation
A text-guided SAM2 variant with cross-modal attention, semantic prompt generation, and a similarity-sorted memory bank achieves top Dice and surface scores on seven public multi-organ CT datasets.
-
FAMSeg: Fetal Femur and Cranial Ultrasound Segmentation Using Feature-Aware Attention and Mamba Enhancement
FAMSeg, a Mamba-enhanced encoder-decoder with strip convolutions and content-aware upsampling, reports the highest mIoU of 91.58 on a private fetal ultrasound dataset.
-
SAMamba: Adaptive State Space Modeling with Hierarchical Vision for Infrared Small Target Detection
SAMamba, combining a frozen SAM2/Hiera encoder with Vision Mamba blocks and three lightweight modules, achieves state-of-the-art infrared small target detection on NUAA-SIRST, IRSTD-1k, and NUDT-SIRST.
-
SASVi -- Segment Any Surgical Video
SASVi uses a Mask2Former overseer to automatically re-prompt SAM2 during surgical videos, improving temporal consistency of segmentations from scarce annotations on Cholec80, CATARACTS, and Cataract1k.
Discussion (0). Continue with ORCID to comment.