Pith. sign in

REVIEW 3 cited by

UNETR: Transformers for 3D Medical Image Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.10504 v3 pith:MA6H67AZ submitted 2021-03-18 eess.IV cs.CVcs.LG

classification eess.IVcs.CVcs.LG
keywords segmentationencodermedicaldecoderfcnnsimagelearningtransformers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Fully Convolutional Neural Networks (FCNNs) with contracting and expanding paths have shown prominence for the majority of medical image segmentation applications since the past decade. In FCNNs, the encoder plays an integral role by learning both global and local features and contextual representations which can be utilized for semantic output prediction by the decoder. Despite their success, the locality of convolutional layers in FCNNs, limits the capability of learning long-range spatial dependencies. Inspired by the recent success of transformers for Natural Language Processing (NLP) in long-range sequence learning, we reformulate the task of volumetric (3D) medical image segmentation as a sequence-to-sequence prediction problem. We introduce a novel architecture, dubbed as UNEt TRansformers (UNETR), that utilizes a transformer as the encoder to learn sequence representations of the input volume and effectively capture the global multi-scale information, while also following the successful "U-shaped" network design for the encoder and decoder. The transformer encoder is directly connected to a decoder via skip connections at different resolutions to compute the final semantic segmentation output. We have validated the performance of our method on the Multi Atlas Labeling Beyond The Cranial Vault (BTCV) dataset for multi-organ segmentation and the Medical Segmentation Decathlon (MSD) dataset for brain tumor and spleen segmentation tasks. Our benchmarks demonstrate new state-of-the-art performance on the BTCV leaderboard. Code: https://monai.io/research/unetr

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DM-SegNet: Dual-Mamba Architecture for 3D Medical Image Segmentation with Global Context Modeling

    eess.IV 2025-06 conditional novelty 5.0 of 10

    DM-SegNet combines four-direction Mamba scanning, gated spatial convolutions, and a Mamba decoder to reach reported state-of-the-art Dice scores on Synapse and BraTS2023.

  2. UNICON: UNIfied CONtinual Learning for Medical Foundational Models

    eess.IV 2025-08 unverdicted novelty 4.0 of 10

    UNICON attaches task-specific adapters (LoRA, MLP, decoder, fusion) to a frozen CT foundation model, enabling continual extension to prognosis, segmentation, and PET scans.

  3. Advancing Fetal Ultrasound Image Quality Assessment in Low-Resource Settings

    cs.CV 2025-07 conditional novelty 4.0 of 10

    LoRA fine-tuning of FetalCLIP achieves F1 0.757 for fetal ultrasound frame-quality classification, and a thresholded segmentation variant reaches F1 0.771.

Pith tools