REVIEW 3 cited by
Swin-UMamba: Mamba-based UNet with ImageNet-based pretraining
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Accurate medical image segmentation demands the integration of multi-scale information, spanning from local features to global dependencies. However, it is challenging for existing methods to model long-range global information, where convolutional neural networks (CNNs) are constrained by their local receptive fields, and vision transformers (ViTs) suffer from high quadratic complexity of their attention mechanism. Recently, Mamba-based models have gained great attention for their impressive ability in long sequence modeling. Several studies have demonstrated that these models can outperform popular vision models in various tasks, offering higher accuracy, lower memory consumption, and less computational burden. However, existing Mamba-based models are mostly trained from scratch and do not explore the power of pretraining, which has been proven to be quite effective for data-efficient medical image analysis. This paper introduces a novel Mamba-based model, Swin-UMamba, designed specifically for medical image segmentation tasks, leveraging the advantages of ImageNet-based pretraining. Our experimental results reveal the vital role of ImageNet-based training in enhancing the performance of Mamba-based models. Swin-UMamba demonstrates superior performance with a large margin compared to CNNs, ViTs, and latest Mamba-based models. Notably, on AbdomenMRI, Encoscopy, and Microscopy datasets, Swin-UMamba outperforms its closest counterpart U-Mamba_Enc by an average score of 2.72%.
Forward citations
Cited by 3 Pith papers
-
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
STG-Mamba generates dance videos from music using a spatial-temporal graph Mamba block for skeleton generation and forward-backward self-supervised losses for video synthesis, reporting SOTA on benchmarks.
-
InceptionMamba: Efficient Multi-Stage Feature Enhancement with Selective State Space Model for Microscopic Medical Image Segmentation
A U-Net-style architecture combining inception-style convolutions with a Mamba state-space block achieves state-of-the-art medical image segmentation at about one-fifth the GFLOPs of the previous best method.
-
OSDMamba: Enhancing Oil Spill Detection from Remote Sensing Images Using Selective State Space Model
OSDMamba, a Mamba-based segmentation model, reports state-of-the-art oil spill detection accuracy on the M4D and MADOS datasets.
Discussion (0). Continue with ORCID to comment.