REVIEW 13 cited by
SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present SegFormer, a simple, efficient yet powerful semantic segmentation framework which unifies Transformers with lightweight multilayer perception (MLP) decoders. SegFormer has two appealing features: 1) SegFormer comprises a novel hierarchically structured Transformer encoder which outputs multiscale features. It does not need positional encoding, thereby avoiding the interpolation of positional codes which leads to decreased performance when the testing resolution differs from training. 2) SegFormer avoids complex decoders. The proposed MLP decoder aggregates information from different layers, and thus combining both local attention and global attention to render powerful representations. We show that this simple and lightweight design is the key to efficient segmentation on Transformers. We scale our approach up to obtain a series of models from SegFormer-B0 to SegFormer-B5, reaching significantly better performance and efficiency than previous counterparts. For example, SegFormer-B4 achieves 50.3% mIoU on ADE20K with 64M parameters, being 5x smaller and 2.2% better than the previous best method. Our best model, SegFormer-B5, achieves 84.0% mIoU on Cityscapes validation set and shows excellent zero-shot robustness on Cityscapes-C. Code will be released at: github.com/NVlabs/SegFormer.
Forward citations
Cited by 13 Pith papers
-
Modeling E-Bike Route Choice in Washington, DC: A Path Size Logit Approach
Shared e-bike riders in DC favor protected bike lanes and continuous routes, with facility effects strongest on major roads and for longer trips.
-
SARATR-X-v2: Scale-Aware Structural Pre-Training for SAR Foundation Models
A learnable-weighted fusion of six fixed, speckle-robust structural operators as the masked pre-training target transfers better than pixel targets on 10 of 12 SAR benchmarks.
-
Geographically-Weighted Weakly Supervised Bayesian High-Resolution Transformer for 200m Resolution Pan-Arctic Sea Ice Concentration Mapping and Uncertainty Estimation using Sentinel-1, RCM, and AMSR2 Data
A 200m pan-Arctic sea ice concentration model using a Bayesian Transformer with a geographically-weighted weakly supervised loss; reported 0.70 feature detection accuracy and R²=0.90 vs ASI.
-
Human-Inspired Neuro-Symbolic World Modeling and Logic Reasoning for Interpretable Safe UAV Landing Site Assessment
NeuroSymLand fuses a lightweight neural segmenter with hand-refined logic rules to rank safe UAV landing sites, beating four lightweight baselines in AirSim and on edge hardware while emitting proof traces.
-
TransForSeg: A Multitask Stereo ViT for Joint Stereo Segmentation and 3D Force Estimation in Catheterization
A shared-weight stereo ViT with cross-attention fusion segments the catheter in two views and regresses 3D tip forces, claiming state-of-the-art results on synthetic X-ray datasets.
-
Milo, a Fully Autonomous Indoor/Outdoor Robotic Guide Dog
Milo is an open-source, fully onboard robotic guide dog that navigates unseen indoor/outdoor paths while explicitly modeling the handler's position, with preliminary real-world tests against a handler-unaware costmap ...
-
SalFormer360: a transformer-based saliency estimation model for 360-degree videos
SalFormer360, a SegFormer-based transformer with a custom decoder and decaying center-bias, reports state-of-the-art CC scores on three 360-degree video saliency benchmarks.
-
I-Segmenter: Integer-Only Vision Transformer for Efficient Semantic Segmentation
I-Segmenter is an integer-only Vision Transformer for semantic segmentation that keeps mIoU within roughly 5 points of the FP32 baseline while cutting model size by up to 3.8x.
-
A Multimodal Architecture for Endpoint Position Prediction in Team-based Multiplayer Games
A U-Net++ with multimodal encoders and agent attention improves World of Tanks endpoint prediction, with KL divergence loss and rendered icons giving the best relative FDE at 1.78.
-
Breaking the Illusion of Security via Interpretation: Interpretable Vision Transformer Systems under Attack
AdViT generates adversarial images that make ViT classifiers misclassify while keeping attribution maps nearly identical to benign inputs, with high white-box success and useful black-box transferability after genetic...
-
SUPER Module for Detail-Sensitive and Cost-Efficient U-Net Variant Decoders
A plug-in wavelet-domain decoder block improves thin-crack IoU on one self-baseline benchmark, while the abstract's flagship depth-estimation gains and decoder MAC reductions are absent from the main text.
-
E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition
Their custom PaddleOCR-based system achieves the best F1 (0.46), fastest latency (0.17 s/image), and lowest cost ($0.006/1k images) among seven OCR systems on a private 54-language benchmark.
-
CarboFormer: A Lightweight Semantic Segmentation Architecture for Efficient Carbon Dioxide Detection Using Optical Gas Imaging
CarboFormer, a 5.07M-parameter transformer-based model, segments CO2 plumes in optical gas images with up to 92.98% mIoU at 84.68 FPS.
Discussion (0). Continue with ORCID to comment.