Pith. sign in

REVIEW 9 cited by

Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors in MRI Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.01266 v1 pith:AYI6YZVJ submitted 2022-01-04 eess.IV cs.CVcs.LG

classification eess.IVcs.CVcs.LG
keywords segmentationswinsemanticbrainsequencetransformertransformerstumors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Semantic segmentation of brain tumors is a fundamental medical image analysis task involving multiple MRI imaging modalities that can assist clinicians in diagnosing the patient and successively studying the progression of the malignant entity. In recent years, Fully Convolutional Neural Networks (FCNNs) approaches have become the de facto standard for 3D medical image segmentation. The popular "U-shaped" network architecture has achieved state-of-the-art performance benchmarks on different 2D and 3D semantic segmentation tasks and across various imaging modalities. However, due to the limited kernel size of convolution layers in FCNNs, their performance of modeling long-range information is sub-optimal, and this can lead to deficiencies in the segmentation of tumors with variable sizes. On the other hand, transformer models have demonstrated excellent capabilities in capturing such long-range information in multiple domains, including natural language processing and computer vision. Inspired by the success of vision transformers and their variants, we propose a novel segmentation model termed Swin UNEt TRansformers (Swin UNETR). Specifically, the task of 3D brain tumor semantic segmentation is reformulated as a sequence to sequence prediction problem wherein multi-modal input data is projected into a 1D sequence of embedding and used as an input to a hierarchical Swin transformer as the encoder. The swin transformer encoder extracts features at five different resolutions by utilizing shifted windows for computing self-attention and is connected to an FCNN-based decoder at each resolution via skip connections. We have participated in BraTS 2021 segmentation challenge, and our proposed model ranks among the top-performing approaches in the validation phase. Code: https://monai.io/research/swin-unetr

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 32 citations worldwide. Full citation record

  1. Integrating Pathology and CT Imaging for Personalized Recurrence Risk Prediction in Renal Cancer

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Multimodal fusion of CT and pathology images improves recurrence risk prediction in kidney cancer, with the best model approaching the clinical Leibovich score.

  2. Generalizable automated ischaemic stroke lesion segmentation with vision transformers

    eess.IV 2025-02 conditional novelty 6.0 of 10

    Swin-UNETR models trained on a large multi-site DWI stroke dataset reach high Dice scores, and a new evaluation framework exposes anatomical, morphological, and noise-dependent performance variation.

  3. AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans

    cs.CV 2026-07 conditional novelty 5.5 of 10

    A latent RSSM with hierarchical anatomical add/remove actions cuts HD95 by ~43% versus nnU-Net on fine-grained nested auricular CT segmentation.

  4. A Space-Time Transformer for Precipitation Nowcasting

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A full space-time attention video transformer recast as 64-class rainfall prediction with log-frequency class weighting won the Weather4Cast 2025 Cumulative Rainfall challenge (CRPS 3.135).

  5. Semantic Segmentation for Preoperative Planning in Transcatheter Aortic Valve Replacement

    eess.IV 2025-07 conditional novelty 5.0 of 10

    The authors release a pseudo-labeled CT dataset for TAVR planning anatomy and propose a focal skeleton recall loss that improves mean Dice by about 1.3 percentage points over a standard baseline.

  6. Learning from Anatomy: Supervised Anatomical Pretraining (SAP) for Improved Metastatic Bone Disease Segmentation in Whole-Body MRI

    eess.IV 2025-06 conditional novelty 5.0 of 10

    Supervised pretraining on healthy skeletal anatomy improved metastatic bone lesion segmentation in whole-body MRI over random and self-supervised initialization.

  7. EfficientGFormer: Multimodal Brain Tumor Segmentation via Pruned Graph-Augmented Transformer

    cs.CV 2025-08 reject novelty 4.0 of 10

    EfficientGFormer combines a pretrained nnFormer encoder with a dual-edge graph attention network and distillation to segment brain tumor subregions, claiming SOTA accuracy with lower compute.

  8. Why Relational Graphs Will Save the Next Generation of Vision Foundation Models?

    cs.CV 2025-08 conditional novelty 3.0 of 10

    A position paper arguing that vision foundation models need dynamic relational graphs for relational reasoning, with evidence drawn from the author's own prior action recognition and tumor segmentation systems.

  9. F3-Net: Foundation Model for Full Abnormality Segmentation of Medical Images with Flexible Input Modality Requirement

    cs.CV 2025-07 reject novelty 3.0 of 10

    F3-Net combines multi-encoder nnU-Net with zero-filled missing modalities to segment glioma, metastasis, stroke, and white matter lesions, but the missing-modality claim is untested and comparisons are incomplete.

Pith tools