REVIEW 9 cited by
Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors in MRI Images
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Semantic segmentation of brain tumors is a fundamental medical image analysis task involving multiple MRI imaging modalities that can assist clinicians in diagnosing the patient and successively studying the progression of the malignant entity. In recent years, Fully Convolutional Neural Networks (FCNNs) approaches have become the de facto standard for 3D medical image segmentation. The popular "U-shaped" network architecture has achieved state-of-the-art performance benchmarks on different 2D and 3D semantic segmentation tasks and across various imaging modalities. However, due to the limited kernel size of convolution layers in FCNNs, their performance of modeling long-range information is sub-optimal, and this can lead to deficiencies in the segmentation of tumors with variable sizes. On the other hand, transformer models have demonstrated excellent capabilities in capturing such long-range information in multiple domains, including natural language processing and computer vision. Inspired by the success of vision transformers and their variants, we propose a novel segmentation model termed Swin UNEt TRansformers (Swin UNETR). Specifically, the task of 3D brain tumor semantic segmentation is reformulated as a sequence to sequence prediction problem wherein multi-modal input data is projected into a 1D sequence of embedding and used as an input to a hierarchical Swin transformer as the encoder. The swin transformer encoder extracts features at five different resolutions by utilizing shifted windows for computing self-attention and is connected to an FCNN-based decoder at each resolution via skip connections. We have participated in BraTS 2021 segmentation challenge, and our proposed model ranks among the top-performing approaches in the validation phase. Code: https://monai.io/research/swin-unetr
Forward citations
Cited by 9 Pith papers
-
Integrating Pathology and CT Imaging for Personalized Recurrence Risk Prediction in Renal Cancer
Multimodal fusion of CT and pathology images improves recurrence risk prediction in kidney cancer, with the best model approaching the clinical Leibovich score.
-
Generalizable automated ischaemic stroke lesion segmentation with vision transformers
Swin-UNETR models trained on a large multi-site DWI stroke dataset reach high Dice scores, and a new evaluation framework exposes anatomical, morphological, and noise-dependent performance variation.
-
AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans
A latent RSSM with hierarchical anatomical add/remove actions cuts HD95 by ~43% versus nnU-Net on fine-grained nested auricular CT segmentation.
-
A Space-Time Transformer for Precipitation Nowcasting
A full space-time attention video transformer recast as 64-class rainfall prediction with log-frequency class weighting won the Weather4Cast 2025 Cumulative Rainfall challenge (CRPS 3.135).
-
Semantic Segmentation for Preoperative Planning in Transcatheter Aortic Valve Replacement
The authors release a pseudo-labeled CT dataset for TAVR planning anatomy and propose a focal skeleton recall loss that improves mean Dice by about 1.3 percentage points over a standard baseline.
-
Learning from Anatomy: Supervised Anatomical Pretraining (SAP) for Improved Metastatic Bone Disease Segmentation in Whole-Body MRI
Supervised pretraining on healthy skeletal anatomy improved metastatic bone lesion segmentation in whole-body MRI over random and self-supervised initialization.
-
EfficientGFormer: Multimodal Brain Tumor Segmentation via Pruned Graph-Augmented Transformer
EfficientGFormer combines a pretrained nnFormer encoder with a dual-edge graph attention network and distillation to segment brain tumor subregions, claiming SOTA accuracy with lower compute.
-
Why Relational Graphs Will Save the Next Generation of Vision Foundation Models?
A position paper arguing that vision foundation models need dynamic relational graphs for relational reasoning, with evidence drawn from the author's own prior action recognition and tumor segmentation systems.
-
F3-Net: Foundation Model for Full Abnormality Segmentation of Medical Images with Flexible Input Modality Requirement
F3-Net combines multi-encoder nnU-Net with zero-filled missing modalities to segment glioma, metastasis, stroke, and white matter lesions, but the missing-modality claim is untested and comparisons are incomplete.
Discussion (0). Continue with ORCID to comment.