Pith. sign in

REVIEW 13 cited by

SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.15203 v3 pith:YPUG7CL5 submitted 2021-05-31 cs.CV cs.LG

classification cs.CVcs.LG
keywords segformerefficientsegmentationsimpletransformersachievesattentionbest
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present SegFormer, a simple, efficient yet powerful semantic segmentation framework which unifies Transformers with lightweight multilayer perception (MLP) decoders. SegFormer has two appealing features: 1) SegFormer comprises a novel hierarchically structured Transformer encoder which outputs multiscale features. It does not need positional encoding, thereby avoiding the interpolation of positional codes which leads to decreased performance when the testing resolution differs from training. 2) SegFormer avoids complex decoders. The proposed MLP decoder aggregates information from different layers, and thus combining both local attention and global attention to render powerful representations. We show that this simple and lightweight design is the key to efficient segmentation on Transformers. We scale our approach up to obtain a series of models from SegFormer-B0 to SegFormer-B5, reaching significantly better performance and efficiency than previous counterparts. For example, SegFormer-B4 achieves 50.3% mIoU on ADE20K with 64M parameters, being 5x smaller and 2.2% better than the previous best method. Our best model, SegFormer-B5, achieves 84.0% mIoU on Cityscapes validation set and shows excellent zero-shot robustness on Cityscapes-C. Code will be released at: github.com/NVlabs/SegFormer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Modeling E-Bike Route Choice in Washington, DC: A Path Size Logit Approach

    stat.AP 2026-08 conditional novelty 6.0 of 10

    Shared e-bike riders in DC favor protected bike lanes and continuous routes, with facility effects strongest on major roads and for longer trips.

  2. SARATR-X-v2: Scale-Aware Structural Pre-Training for SAR Foundation Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A learnable-weighted fusion of six fixed, speckle-robust structural operators as the masked pre-training target transfers better than pixel targets on 10 of 12 SAR benchmarks.

  3. Geographically-Weighted Weakly Supervised Bayesian High-Resolution Transformer for 200m Resolution Pan-Arctic Sea Ice Concentration Mapping and Uncertainty Estimation using Sentinel-1, RCM, and AMSR2 Data

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A 200m pan-Arctic sea ice concentration model using a Bayesian Transformer with a geographically-weighted weakly supervised loss; reported 0.70 feature detection accuracy and R²=0.90 vs ASI.

  4. Human-Inspired Neuro-Symbolic World Modeling and Logic Reasoning for Interpretable Safe UAV Landing Site Assessment

    cs.RO 2025-10 conditional novelty 6.0 of 10

    NeuroSymLand fuses a lightweight neural segmenter with hand-refined logic rules to rank safe UAV landing sites, beating four lightweight baselines in AirSim and on edge hardware while emitting proof traces.

  5. TransForSeg: A Multitask Stereo ViT for Joint Stereo Segmentation and 3D Force Estimation in Catheterization

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A shared-weight stereo ViT with cross-attention fusion segments the catheter in two views and regresses 3D tip forces, claiming state-of-the-art results on synthetic X-ray datasets.

  6. Milo, a Fully Autonomous Indoor/Outdoor Robotic Guide Dog

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Milo is an open-source, fully onboard robotic guide dog that navigates unseen indoor/outdoor paths while explicitly modeling the handler's position, with preliminary real-world tests against a handler-unaware costmap ...

  7. SalFormer360: a transformer-based saliency estimation model for 360-degree videos

    cs.CV 2026-02 conditional novelty 5.0 of 10

    SalFormer360, a SegFormer-based transformer with a custom decoder and decaying center-bias, reports state-of-the-art CC scores on three 360-degree video saliency benchmarks.

  8. I-Segmenter: Integer-Only Vision Transformer for Efficient Semantic Segmentation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    I-Segmenter is an integer-only Vision Transformer for semantic segmentation that keeps mIoU within roughly 5 points of the FP32 baseline while cutting model size by up to 3.8x.

  9. A Multimodal Architecture for Endpoint Position Prediction in Team-based Multiplayer Games

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A U-Net++ with multimodal encoders and agent attention improves World of Tanks endpoint prediction, with KL divergence loss and rendered icons giving the best relative FDE at 1.78.

  10. Breaking the Illusion of Security via Interpretation: Interpretable Vision Transformer Systems under Attack

    cs.CR 2025-07 reject novelty 5.0 of 10

    AdViT generates adversarial images that make ViT classifiers misclassify while keeping attribution maps nearly identical to benign inputs, with high white-box success and useful black-box transferability after genetic...

  11. SUPER Module for Detail-Sensitive and Cost-Efficient U-Net Variant Decoders

    cs.CV 2025-11 reject novelty 4.0 of 10

    A plug-in wavelet-domain decoder block improves thin-crack IoU on one self-baseline benchmark, while the abstract's flagship depth-estimation gains and decoder MAC reductions are absent from the main text.

  12. E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition

    cs.CL 2025-09 reject novelty 4.0 of 10

    Their custom PaddleOCR-based system achieves the best F1 (0.46), fastest latency (0.17 s/image), and lowest cost ($0.006/1k images) among seven OCR systems on a private 54-language benchmark.

  13. CarboFormer: A Lightweight Semantic Segmentation Architecture for Efficient Carbon Dioxide Detection Using Optical Gas Imaging

    cs.CV 2025-05 conditional novelty 4.0 of 10

    CarboFormer, a 5.07M-parameter transformer-based model, segments CO2 plumes in optical gas images with up to 92.98% mIoU at 84.68 FPS.

Pith tools