Pith. sign in

REVIEW 3 cited by

ScaleFormer: Revisiting the Transformer-based Backbones from a Scale-wise Perspective for Medical Image Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.14552 v1 pith:H3K4M6ME submitted 2022-07-29 cs.CV

classification cs.CV
keywords globalmethodsscalescale-wisescaleformertransformer-basedtransformersbackbones
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, a variety of vision transformers have been developed as their capability of modeling long-range dependency. In current transformer-based backbones for medical image segmentation, convolutional layers were replaced with pure transformers, or transformers were added to the deepest encoder to learn global context. However, there are mainly two challenges in a scale-wise perspective: (1) intra-scale problem: the existing methods lacked in extracting local-global cues in each scale, which may impact the signal propagation of small objects; (2) inter-scale problem: the existing methods failed to explore distinctive information from multiple scales, which may hinder the representation learning from objects with widely variable size, shape and location. To address these limitations, we propose a novel backbone, namely ScaleFormer, with two appealing designs: (1) A scale-wise intra-scale transformer is designed to couple the CNN-based local features with the transformer-based global cues in each scale, where the row-wise and column-wise global dependencies can be extracted by a lightweight Dual-Axis MSA. (2) A simple and effective spatial-aware inter-scale transformer is designed to interact among consensual regions in multiple scales, which can highlight the cross-scale dependency and resolve the complex scale variations. Experimental results on different benchmarks demonstrate that our Scale-Former outperforms the current state-of-the-art methods. The code is publicly available at: https://github.com/ZJUGiveLab/ScaleFormer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Large-Scale Referring Remote Sensing Image Segmentation Dataset and Benchmark

    cs.CV 2025-06 conditional novelty 6.0 of 10

    NWPU-Refer is a bilingual, high-resolution remote sensing segmentation dataset with multi-object and no-target queries, and MRSNet is a multi-scale network that achieves the best reported scores on it.

  2. A Large-Scale Dataset and a New Method for RemoteSensing Traffic Object Segmentation

    cs.CV 2026-07 conditional novelty 5.5 of 10

    NWPU-Traffic supplies instance masks for four traffic classes across global scenes, and CSPNet raises mIoU to 73.5 % via spatial-channel preserving fusion and a gated local-global decoder.

  3. ACM-UNet: Adaptive Integration of CNNs and Mamba for Efficient Medical Image Segmentation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    ACM-UNet, a UNet variant combining pretrained CNN and Mamba backbones via lightweight adapters and a wavelet decoder module, reports 85.12% Dice on Synapse and 92.29% on ACDC.

Pith tools