REVIEW 1 major objections 6 minor 2 cited by
Deformable Mamba for Wide Field of View Segmentation
T0 review · 1 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper argues that distortion awareness for wide-field-of-view segmentation can be supplied entirely by the decoder, without modifying the backbone, and validates this with a decoder that works across CNN, Transformer, and Mamba…
desk verdict A useful plug-in decoder with a clean ablation story, held back by a misleading efficiency headline and uncontrolled re-implemented baselines. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Deformable Mamba Fusion (DMF) block: a fusion-and-upsampling module that takes an encoder feature $E_i$ and a previously fused decoder feature $D_j$, applies one quadri-directional selective scan (SS2D) to $D_j$ and one DCNv2 deformable convolution to $E_i$, concatenates the two branches, fuses them with 3×3 convolutions, and upsamples with PixelShuffle. It is the mechanism that adds data-dependent offsets—the distortion-awareness—without touching the encoder, which is why the same decoder can be attached to CNN, Transformer, and Mamba backbones.
What would settle it
Run a head-swap experiment on Stanford2D3D where UperHead, MusterHead, CGRHead, and the proposed decoder share the exact same augmentation, data split, effective iterations, and seeds; if the proposed decoder does not beat the best baseline by a margin consistent with the paper's +2.5 points, the claim that the decoder alone imparts distortion awareness is not supported.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that a Mamba-based decoder can be made distortion-aware by combining the quadri-directional 2D selective scan from VMamba with a parallel DCNv2 deformable-convolution branch in each Deformable Mamba Fusion block. This gives the decoder adaptive sampling offsets that compensate for fisheye and equirectangular deformation while retaining Mamba's linear complexity. With the same VMamba-T backbone, the proposed decoder raises mIoU on Stanford2D3D from 56.8 for the VMamba-plus-UperHead baseline to 59.3, improves results on Matterport3D, SynPASS, WoodScape, and SynWoodScape, and reduces decoder parameters by 72 percent and FLOPs by 97 percent compared with UperHead.
Load-bearing premise
The reported gains assume the re-implemented baselines were trained under the same data splits, preprocessing, augmentation, and optimization settings as the proposed model, so that swapping the decoder head is the only meaningful change.
Editorial extensions
If this is right
- A single decoder design can replace task-specific heads across CNN, Transformer, and Mamba backbones, so wide-FoV models no longer need distortion handling baked into the encoder.
- The decoder's compute is small enough that the head stops being the FLOP bottleneck: on VMamba-T, UperHead uses 206.9 GFLOPs while the proposed head uses 6.0 GFLOPs, making the total model much cheaper.
- Consistent gains on synthetic and real, indoor and outdoor, 180-degree and 360-degree datasets indicate the approach generalizes across distortion types rather than overfitting one sensor.
- Quadri-directional scanning contributes more than uni- or bi-directional scanning in the decoder, supporting the claim that scan diversity matters for spatially structured wide-FoV features.
- Replacing deformable convolution with ordinary convolution drops mIoU by 1.85 points, so the deformable branch is the component responsible for the distortion-aware gain.
Reading between the lines
- Editorial inference: because the decoder is the only component that changes, the same DMF block could plausibly transfer distortion-awareness to other dense prediction tasks on wide-FoV inputs, such as depth estimation, panoptic segmentation, or open-vocabulary segmentation.
- Editorial inference: the small margin on Matterport3D (+0.3 to +0.9 mIoU) suggests the benefit of decoder-level deformation is larger on datasets with heavy distortion or fine elongated classes; a controlled head-swap study would reveal which property drives the gain.
- Editorial inference: a testable extension is to visualize the learned DCN offsets on equirectangular images; if the offsets systematically align with the spherical-to-planar projection, that would confirm the mechanism is distortion compensation rather than generic capacity.
- Editorial inference: since the paper trains all models with one schedule and no auxiliary losses, part of the gain could come from regularization; comparing against a matched-capacity conventional decoder with identical training would isolate the deformable mechanism.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes the Deformable Mamba Decoder, a decoder that combines a quadri-directional Mamba scan branch with a deformable convolution (DCNv2) branch and PixelShuffle upsampling, and demonstrates it as a plug-in decoder for CNN-, Transformer-, and Mamba-based backbones in wide field-of-view (180° fisheye and 360° panoramic) semantic segmentation. The method is evaluated on five datasets (Stanford2D3D, Matterport3D, SynPASS, WoodScape, SynWoodScape) and is reported to improve mIoU over prior decoders while reducing parameters and FLOPs, with a headline +2.5 mIoU gain on Stanford2D3D over a VMamba-based baseline. The paper includes ablations of scan direction, deformable convolution, and upsampling method, and releases code.
Significance. If the empirical claims are sustained, the paper makes a useful contribution: it identifies decoder-level distortion awareness as a modular component that can be attached to diverse backbones, and it provides one of the first decoder-focused uses of Mamba for dense prediction. The experimental scope is a strength: five datasets, three backbone families, and ablations that isolate DCN, scan direction, and upsampling. The availability of code is also a concrete asset. However, the main quantitative claims currently rest on comparisons whose fairness and precise accounting need to be verified before the stated advantages can be accepted.
major comments (1)
- [§4.2 and Tables 4-7] The performance gains over re-implemented baselines (VMamba†, SegNeXt†, 360SFUDA†) are load-bearing for the +2.5 mIoU claim, but the manuscript does not establish that these re-implementations were trained under identical conditions. Section 4.2 specifies the optimizer, schedule, and iterations for the proposed method but does not state whether the re-implemented baselines used the same augmentation, cropping/resolution policy, preprocessing, or loss weighting. Since the cleanest comparison (same VMamba-T backbone, only the decoder differs) is exactly the case where training details matter most, the paper should report the full training recipe for the re-implemented baselines and ideally release their configuration files, or compare against published numbers from the original papers.
minor comments (6)
- [Abstract] The opening sentence is grammatically incomplete: 'Recent advancements in the Mamba architecture, with its linear computational complexity, being a promising alternative...' should be revised.
- [§4.3] There is a typo: 'Deformable Mamaba Decoder' should be 'Deformable Mamba Decoder'.
- [§4.4] 'desin' in 'lacking deformable desin' should be 'design'.
- [§3.2, Eq. (3)] The PixelShuffle formula's subscript notation is unclear: 'PS(T)h,w,2c = T2h,2w,c/2' mixes input and output dimensions in a way that is hard to parse. Please rewrite the index notation to show the rearrangement explicitly.
- [§4.2 and Tables 1-3] It should be stated explicitly whether the reported parameter and FLOPs counts refer to the decoder alone or to the full encoder-decoder model; Tables 1-3 list decoder-level numbers while §4.5 discusses overall reductions.
- [Limitations and Future Work] The limitations paragraph does not actually state a limitation of the current method; it only points to future LLM integration. Please add a genuine discussion of limitations, for example the sensitivity of the reported gains to re-implementation details and the mixed efficiency baselines.
Circularity Check
No significant circularity: the reported decoder gains are empirical results measured against external benchmarks and existing baselines, not predictions derived from fitted constants or self-citation chains.
full rationale
The paper's central claim is an experimental one: replacing standard decoder heads with the proposed Deformable Mamba Decoder improves wide-FoV semantic segmentation while reducing parameters and FLOPs. The decoder is a concrete architecture (quadri-directional Mamba scan branch plus DCNv2 branch, fused and upsampled with PixelShuffle), and its performance is assessed on five external datasets against published and re-implemented baselines. No quantity is fitted to a subset of data and then renamed as a prediction; the mIoU numbers are measured on held-out evaluation splits. The ablation in Table 9 isolates the distortion-aware component by replacing deformable convolution with a regular convolution, and the scanning and upsampling ablations similarly vary one architectural choice at a time, giving the central attribution independent empirical content. The paper cites prior work by the same group (e.g., Trans4PASS, 360BEV) but uses those results as comparison baselines rather than as load-bearing justification for the proposed design; no uniqueness theorem or prior self-citation is invoked to rule out alternatives. The main substantive concerns in the manuscript are about experimental comparability: several baselines are marked as re-implementations (VMamba†, SegNeXt†, 360SFUDA†) without full sharing of augmentation and preprocessing configurations, and the headline efficiency figures compare different baselines across tables (72% parameter reduction versus CGRHead, 97% FLOPs reduction versus UperHead). These are legitimate threats to the strength of the empirical claim, but they are not circularity: they concern whether the comparison is fair and whether the reported margins would survive a strictly controlled protocol, not whether the conclusion is equivalent to its inputs by construction. The paper is self-contained against external benchmarks, and no derivation chain reduces to its own assumptions.
Assumptions & free parameters
assumptions (4)
- domain assumption The SS2D quadri-directional scan block from VMamba provides effective 2D state-space modeling with linear complexity in the decoder.
- domain assumption DCNv2's learned offsets can represent and compensate for the geometric distortions present in 180° and 360° imagery.
- domain assumption Public dataset splits and mIoU evaluation are faithful indicators of segmentation quality.
- domain assumption The re-implemented baselines are trained fairly under comparable settings.
Cite this review
Pith. "Pith review of Deformable Mamba for Wide Field of View Segmentation." pith.science (2026). https://pith.science/paper/2VUS5HSU
@misc{pith2026241116481,
author = {Pith},
title = {Pith review of: Deformable Mamba for Wide Field of View Segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/2VUS5HSU}},
note = {Machine review of arXiv:2411.16481}
}
read the original abstract
Recent advancements in the Mamba architecture, with its linear computational complexity, being a promising alternative to transformer architectures suffering from quadratic complexity. While existing works primarily focus on adapting Mamba as vision encoders, the critical role of task-specific Mamba decoders remains under-explored, particularly for distortion-prone dense prediction tasks. This paper addresses two interconnected challenges: (1) The design of a Mamba-based decoder that seamlessly adapts to various architectures (e.g., CNN-, Transformer-, and Mamba-based backbones), and (2) The performance degradation in decoders lacking distortion-aware capability when processing wide-FoV images (e.g., 180{\deg} fisheye and 360{\deg} panoramic settings). We propose the Deformable Mamba Decoder, an efficient distortion-aware decoder that integrates Mamba's computational efficiency with adaptive distortion awareness. Comprehensive experiments on five wide-FoV segmentation benchmarks validate its effectiveness. Notably, our decoder achieves a +2.5% performance improvement on the 360{\deg} Stanford2D3D segmentation benchmark while reducing 72% parameters and 97% FLOPs, as compared to the widely-used decoder heads.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 2 Pith papers
-
SeqLoc: Beyond the Single Frame for Cross-View Geo-Localization in Feature-Sparse Scenes
SeqLoc recursively fuses per-frame pose likelihoods with entropy weighting, map-guided relocalization, and sub-grid smoothing, raising position and orientation recall by over 50 percent on sparse rural scenes.
-
Panoramic Scene Understanding: A Survey from Distortion-Aware Engineering to Sphere-Native Modeling
Survey organizing panoramic scene analysis literature by architectural design and training paradigm, identifying the absence of methods achieving both strict spherical equivariance and full reuse of perspective-pretra...
Reference graph
Works this paper leans on
-
[1]
Joint 2d-3d semantic data for indoor scene un- derstanding
I Armeni. Joint 2d-3d semantic data for indoor scene un- derstanding. arXiv preprint arXiv:1702.01105, 2017. 1, 2, 5
arXiv 2017
-
[2]
Segnet: A deep convolutional encoder-decoder architecture for image segmentation
Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE transactions on pattern anal- ysis and machine intelligence , 39(12):2481–2495, 2017. 1, 2
work page 2017
-
[3]
Matterport3d: Learning from rgb-d data in indoor environments
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. arXiv preprint arXiv:1709.06158, 2017. 1, 2, 5
arXiv 2017
-
[4]
Rsmamba: Remote sens- ing image classification with state space model
Keyan Chen, Bowen Chen, Chenyang Liu, Wenyuan Li, Zhengxia Zou, and Zhenwei Shi. Rsmamba: Remote sens- ing image classification with state space model. IEEE Geo- science and Remote Sensing Letters, 2024. 2, 3, 4
work page 2024
-
[5]
Encoder-decoder with atrous separable convolution for semantic image segmentation
Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European conference on computer vision (ECCV), pages 801–818, 2018. 1, 2
work page 2018
-
[6]
Cyclemlp: A mlp-like architecture for dense prediction
Shoufa Chen, Enze Xie, Chongjian Ge, Runjian Chen, Ding Liang, and Ping Luo. Cyclemlp: A mlp-like architecture for dense prediction. arXiv preprint arXiv:2107.10224, 2021. 2
arXiv 2021
-
[7]
Cyclemlp: A mlp-like architecture for dense visual predictions
Shoufa Chen, Enze Xie, Chongjian Ge, Runjian Chen, Ding Liang, and Ping Luo. Cyclemlp: A mlp-like architecture for dense visual predictions. IEEE Transactions on Pattern Analysis and Machine Intelligence , 45(12):14284–14300,
-
[8]
Semi-supervised semantic segmentation with cross pseudo supervision
Xiaokang Chen, Yuhui Yuan, Gang Zeng, and Jingdong Wang. Semi-supervised semantic segmentation with cross pseudo supervision. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 2613–2622, 2021. 7
work page 2021
Show all 76 references
-
[9]
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and An- drew Rabinovich. Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks. In In- ternational conference on machine learning, pages 794–803. PMLR, 2018. 7
2018
-
[10]
Per- pixel classification is not all you need for semantic segmen- tation
Bowen Cheng, Alex Schwing, and Alexander Kirillov. Per- pixel classification is not all you need for semantic segmen- tation. Advances in neural information processing systems , 34:17864–17875, 2021. 2
2021
-
[11]
Deformable convolutional networks
Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. In Proceedings of the IEEE international confer- ence on computer vision, pages 764–773, 2017. 4
2017
-
[12]
Restricted deformable convolution-based road scene semantic segmentation using surround view cam- eras
Liuyuan Deng, Ming Yang, Hao Li, Tianyi Li, Bing Hu, and Chunxiang Wang. Restricted deformable convolution-based road scene semantic segmentation using surround view cam- eras. IEEE Transactions on Intelligent Transportation Sys- tems, 21(10):4350–4362, 2019. 3
2019
-
[13]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 3
2010 arXiv
-
[14]
Carla: An open urban driv- ing simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Anto- nio Lopez, and Vladlen Koltun. Carla: An open urban driv- ing simulator. In Conference on robot learning, pages 1–16. PMLR, 2017. 5
2017
-
[15]
Tangent images for mitigating spherical distortion
Marc Eder, Mykhailo Shvets, John Lim, and Jan-Michael Frahm. Tangent images for mitigating spherical distortion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12426–12434, 2020. 6
2020
-
[16]
Dual attention network for scene seg- mentation
Jun Fu, Jing Liu, Haijie Tian, Yong Li, Yongjun Bao, Zhiwei Fang, and Hanqing Lu. Dual attention network for scene seg- mentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3146–3154,
-
[17]
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023. 2, 3, 4
2023 arXiv
-
[18]
Multi-scale high-resolution vision transformer for se- mantic segmentation
Jiaqi Gu, Hyoukjun Kwon, Dilin Wang, Wei Ye, Meng Li, Yu-Hsin Chen, Liangzhen Lai, Vikas Chandra, and David Z Pan. Multi-scale high-resolution vision transformer for se- mantic segmentation. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition...
2022
-
[19]
Dynamic task prioritization for multitask learning
Michelle Guo, Albert Haque, De-An Huang, Serena Yeung, and Li Fei-Fei. Dynamic task prioritization for multitask learning. In Proceedings of the European conference on com- puter vision (ECCV), pages 270–287, 2018. 7
2018
-
[20]
Segnext: Rethink- ing convolutional attention design for semantic segmenta- tion
Meng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu, Ming-Ming Cheng, and Shi-Min Hu. Segnext: Rethink- ing convolutional attention design for semantic segmenta- tion. Advances in Neural Information Processing Systems , 35:1140–1156, 2022. 1, 2, 6, 7
2022
-
[21]
Single frame se- mantic segmentation using multi-modal spherical images
Suresh Guttikonda and Jason Rambach. Single frame se- mantic segmentation using multi-modal spherical images. In Proceedings of the IEEE/CVF Winter Conference on Appli- cations of Computer Vision, pages 3222–3231, 2024. 5
2024
-
[22]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 5
2016
-
[23]
Zigma: Zigzag mamba diffusion model
Vincent Tao Hu, Stefan Andreas Baumann, Ming Gui, Olga Grebenkova, Pingchuan Ma, Johannes Fischer, and Bjorn Ommer. Zigma: Zigzag mamba diffusion model. arXiv preprint arXiv:2403.13802, 2024. 3, 4
2024 arXiv
-
[24]
Localmamba: Visual state space model with windowed selective scan
Tao Huang, Xiaohuan Pei, Shan You, Fei Wang, Chen Qian, and Chang Xu. Localmamba: Visual state space model with windowed selective scan. arXiv preprint arXiv:2403.09338,
-
[25]
Ccnet: Criss-cross attention for semantic segmentation
Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. Ccnet: Criss-cross attention for semantic segmentation. In Proceedings of the IEEE/CVF international conference on computer vision, pages 603–612, 2019. 2
2019
-
[26]
Panoramic panoptic segmentation: Towards complete sur- rounding understanding via unsupervised contrastive learn- ing
Alexander Jaus, Kailun Yang, and Rainer Stiefelhagen. Panoramic panoptic segmentation: Towards complete sur- rounding understanding via unsupervised contrastive learn- ing. In 2021 IEEE Intelligent Vehicles Symposium (IV) , pages 1421–1427. IEEE, 2021. 2
2021
-
[27]
Multi-task learning using uncertainty to weigh losses for scene geome- try and semantics
Alex Kendall, Yarin Gal, and Roberto Cipolla. Multi-task learning using uncertainty to weigh losses for scene geome- try and semantics. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7482–7491,
-
[28]
Imagenet classification with deep convolutional neural net- works
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural net- works. Advances in neural information processing systems , 25, 2012. 3
2012
-
[29]
Omnidet: Surround view cameras based multi-task visual perception network for autonomous driv- ing
Varun Ravi Kumar, Senthil Yogamani, Hazem Rashed, Ganesh Sitsu, Christian Witt, Isabelle Leang, Stefan Milz, and Patrick M¨ader. Omnidet: Surround view cameras based multi-task visual perception network for autonomous driv- ing. IEEE Robotics and Automation Letters , 6(2):2830...
2021
-
[30]
Videomamba: State space model for efficient video understanding
Kunchang Li, Xinhao Li, Yi Wang, Yinan He, Yali Wang, Limin Wang, and Yu Qiao. Videomamba: State space model for efficient video understanding. arXiv preprint arXiv:2403.06977, 2024. 2, 3, 4
2024 arXiv
-
[31]
Sgat4pass: spherical geometry-aware trans- former for panoramic semantic segmentation
Xuewei Li, Tao Wu, Zhongang Qi, Gaoang Wang, Ying Shan, and Xi Li. Sgat4pass: spherical geometry-aware trans- former for panoramic semantic segmentation. arXiv preprint arXiv:2306.03403, 2023. 6
2023 arXiv
-
[32]
Ct- net: Context-based tandem network for semantic segmenta- tion
Zechao Li, Yanpeng Sun, Liyan Zhang, and Jinhui Tang. Ct- net: Context-based tandem network for semantic segmenta- tion. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(12):9904–9917, 2021. 2
2021
-
[33]
Covariance attention for semantic segmentation
Yazhou Liu, Yuliang Chen, Pongsak Lasang, and Qunsen Sun. Covariance attention for semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(4):1805–1818, 2020. 2
2020
-
[34]
Vmamba: Visual state space model
Yue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu, Lingxi Xie, Yaowei Wang, Qixiang Ye, and Yunfan Liu. Vmamba: Visual state space model. arXiv preprint arXiv:2401.10166,
-
[35]
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022, 2021. 5
2021
-
[36]
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Pro- ceedings of the IEEE conference on computer vision and pat- tern recognition, pages 3431–3440, 2015. 2
2015
-
[37]
De- formable convolution based road scene semantic segmenta- tion of fisheye images in autonomous driving
Anam Manzoor, Aryan Singh, Ganesh Sistu, Reenu Mohan- das, Eoin Grua, Anthony Scanlan, and Ciar ´an Eising. De- formable convolution based road scene semantic segmenta- tion of fisheye images in autonomous driving. In IET Con- ference Proceedings CP887, pages 7–14. IET, 2024. 3
2024
-
[38]
Context-guided spatial feature recon- struction for efficient semantic segmentation
Zhenliang Ni, Xinghao Chen, Yingjie Zhai, Yehui Tang, and Yunhe Wang. Context-guided spatial feature recon- struction for efficient semantic segmentation. arXiv preprint arXiv:2405.06228, 2024. 2, 5
2024 arXiv
-
[39]
Seman- tic segmentation using transfer learning on fisheye images
Sneha Paul, Zachary Patterson, and Nizar Bouguila. Seman- tic segmentation using transfer learning on fisheye images. In 2023 International Conference on Machine Learning and Applications (ICMLA), pages 445–452. IEEE, 2023. 3
2023
-
[40]
Fish- segssl: A semi-supervised semantic segmentation frame- work for fish-eye images
Sneha Paul, Zachary Patterson, and Nizar Bouguila. Fish- segssl: A semi-supervised semantic segmentation frame- work for fish-eye images. Journal of Imaging , 10(3):71,
-
[41]
Efficientvmamba: Atrous selective scan for light weight visual mamba
Xiaohuan Pei, Tao Huang, and Chang Xu. Efficientvmamba: Atrous selective scan for light weight visual mamba. arXiv preprint arXiv:2403.09977, 2024. 3, 4
2024 arXiv
-
[42]
Adaptable deformable convolutions for semantic segmentation of fisheye images in autonomous driving sys- tems
Cl ´ement Playout, Ola Ahmad, Freddy Lecue, and Farida Cheriet. Adaptable deformable convolutions for semantic segmentation of fisheye images in autonomous driving sys- tems. arXiv preprint arXiv:2102.10191, 2021. 3
2021 arXiv
-
[43]
Fusionnet: A deep fully residual convo- lutional neural network for image segmentation in connec- tomics
Tran Minh Quan, David Grant Colburn Hildebrand, and Won-Ki Jeong. Fusionnet: A deep fully residual convo- lutional neural network for image segmentation in connec- tomics. Frontiers in Computer Science, 3:613981, 2021. 3
2021
-
[44]
U- net: Convolutional networks for biomedical image segmen- tation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, pa...
2015
-
[45]
Synwoodscape: Synthetic surround-view fisheye camera dataset for autonomous driving
Ahmed Rida Sekkat, Yohan Dupuis, Varun Ravi Kumar, Hazem Rashed, Senthil Yogamani, Pascal Vasseur, and Paul Honeine. Synwoodscape: Synthetic surround-view fisheye camera dataset for autonomous driving. IEEE Robotics and Automation Letters, 7(3):8502–8509, 2022. 1, 2, 5
2022
-
[46]
Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network
Wenzhe Shi, Jose Caballero, Ferenc Husz ´ar, Johannes Totz, Andrew P Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang. Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. In Proceedings of the IEEE conference on compu...
2016
-
[47]
Segmenter: Transformer for semantic segmenta- tion
Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. Segmenter: Transformer for semantic segmenta- tion. In Proceedings of the IEEE/CVF international confer- ence on computer vision, pages 7262–7272, 2021. 1, 2
2021
-
[48]
Hohonet: 360 indoor holistic understanding with latent horizontal fea- tures
Cheng Sun, Min Sun, and Hwann-Tzong Chen. Hohonet: 360 indoor holistic understanding with latent horizontal fea- tures. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 2573–2582,
-
[49]
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. Advances in neural information processing systems, 30, 2017. 1, 7
2017
-
[50]
360bev: Panoramic semantic mapping for indoor bird’s-eye view
Zhifeng Teng, Jiaming Zhang, Kailun Yang, Kunyu Peng, Hao Shi, Simon Reiß, Ke Cao, and Rainer Stiefelhagen. 360bev: Panoramic semantic mapping for indoor bird’s-eye view. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages 373–382, 2024. 5, 6
2024
-
[51]
Attention is all you need
A Vaswani. Attention is all you need. Advances in Neural Information Processing Systems, 2017. 3
2017
-
[52]
Pvt v2: Improved baselines with pyramid vision transformer
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pvt v2: Improved baselines with pyramid vision transformer. Computational Visual Media, 8(3):415–424, 2022. 7
2022
-
[53]
Internimage: Exploring large-scale vi- sion foundation models with deformable convolutions
Wenhai Wang, Jifeng Dai, Zhe Chen, Zhenhang Huang, Zhiqi Li, Xizhou Zhu, Xiaowei Hu, Tong Lu, Lewei Lu, Hongsheng Li, et al. Internimage: Exploring large-scale vi- sion foundation models with deformable convolutions. In Proceedings of the IEEE/CVF conference on computer vi- si...
2023
-
[54]
Unified perceptual parsing for scene understand- ing
Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understand- ing. In Proceedings of the European conference on computer vision (ECCV), pages 418–434, 2018. 2, 5, 6, 7
2018
-
[55]
Segformer: Simple and efficient design for semantic segmentation with transform- ers
Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transform- ers. Advances in neural information processing systems, 34: 12077–12090, 2021. 1, 2, 6, 7
2021
-
[56]
Efficient deformable convnets: Rethinking dynamic and sparse operator for vision applications
Yuwen Xiong, Zhiqi Li, Yuntao Chen, Feng Wang, Xizhou Zhu, Jiapeng Luo, Wenhai Wang, Tong Lu, Hongsheng Li, Yu Qiao, et al. Efficient deformable convnets: Rethinking dynamic and sparse operator for vision applications. In Pro- ceedings of the IEEE/CVF Conference on Computer Vi...
2024
-
[57]
Uperformer: A multi-scale transformer- based decoder for semantic segmentation
Jing Xu, Wentao Shi, Pan Gao, Zhengwei Wang, and Qizhu Li. Uperformer: A multi-scale transformer- based decoder for semantic segmentation. arXiv preprint arXiv:2211.13928, 2022. 2, 5
2022 arXiv
-
[58]
Plainmamba: Improving non-hierarchical mamba in visual recognition
Chenhongyi Yang, Zehui Chen, Miguel Espinosa, Linus Er- icsson, Zhenyu Wang, Jiaming Liu, and Elliot J Crowley. Plainmamba: Improving non-hierarchical mamba in visual recognition. arXiv preprint arXiv:2403.17695, 2024. 2
2024 arXiv
-
[59]
Can we pass beyond the field of view? panoramic annular semantic segmentation for real-world surrounding perception
Kailun Yang, Xinxin Hu, Luis M Bergasa, Eduardo Romera, Xiao Huang, Dongming Sun, and Kaiwei Wang. Can we pass beyond the field of view? panoramic annular semantic segmentation for real-world surrounding perception. In2019 IEEE Intelligent Vehicles Symposium (IV) , pages 446–4...
2019
-
[60]
Pass: Panoramic annular semantic seg- mentation
Kailun Yang, Xinxin Hu, Luis M Bergasa, Eduardo Romera, and Kaiwei Wang. Pass: Panoramic annular semantic seg- mentation. IEEE Transactions on Intelligent Transportation Systems, 21(10):4171–4185, 2019. 2
2019
-
[61]
Vivim: a video vision mamba for medical video object segmentation
Yijun Yang, Zhaohu Xing, and Lei Zhu. Vivim: a video vision mamba for medical video object segmentation. arXiv preprint arXiv:2401.14168, 2024. 3, 4
2024 arXiv
-
[62]
Woodscape: A multi-task, multi-camera fisheye dataset for autonomous driving
Senthil Yogamani, Ciar ´an Hughes, Jonathan Horgan, Ganesh Sistu, Padraig Varley, Derek O’Dea, Michal Uric ´ar, Ste- fan Milz, Martin Simon, Karl Amende, et al. Woodscape: A multi-task, multi-camera fisheye dataset for autonomous driving. In Proceedings of the IEEE/CVF Interna...
2019
-
[63]
Ocnet: Object context for seman- tic segmentation
Yuhui Yuan, Lang Huang, Jianyuan Guo, Chao Zhang, Xilin Chen, and Jingdong Wang. Ocnet: Object context for seman- tic segmentation. International Journal of Computer Vision, 129(8):2375–2398, 2021. 2
2021
-
[64]
Bending reality: Distortion-aware transformers for adapting to panoramic se- mantic segmentation
Jiaming Zhang, Kailun Yang, Chaoxiang Ma, Simon Reiß, Kunyu Peng, and Rainer Stiefelhagen. Bending reality: Distortion-aware transformers for adapting to panoramic se- mantic segmentation. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition ,...
2022
-
[65]
Behind every domain there is a shift: Adapting distortion-aware vision transformers for panoramic semantic segmentation
Jiaming Zhang, Kailun Yang, Hao Shi, Simon Reiß, Kunyu Peng, Chaoxiang Ma, Haodong Fu, Philip HS Torr, Kaiwei Wang, and Rainer Stiefelhagen. Behind every domain there is a shift: Adapting distortion-aware vision transformers for panoramic semantic segmentation. IEEE Transactio...
2024
-
[66]
Motion mamba: Efficient and long sequence motion generation
Zeyu Zhang, Akide Liu, Ian Reid, Richard Hartley, Bohan Zhuang, and Hao Tang. Motion mamba: Efficient and long sequence motion generation. In European Conference on Computer Vision, pages 265–282. Springer, 2025. 3, 4
2025
-
[67]
Materobot: Material recognition in wearable robotics for people with visual impairments
Junwei Zheng, Jiaming Zhang, Kailun Yang, Kunyu Peng, and Rainer Stiefelhagen. Materobot: Material recognition in wearable robotics for people with visual impairments. In 2024 IEEE International Conference on Robotics and Au- tomation (ICRA), pages 2303–2309. IEEE, 2024. 1, 2
2024
-
[68]
Open panoramic segmentation
Junwei Zheng, Ruiping Liu, Yufan Chen, Kunyu Peng, Chengzhi Wu, Kailun Yang, Jiaming Zhang, and Rainer Stiefelhagen. Open panoramic segmentation. In European Conference on Computer Vision , pages 164–182. Springer,
-
[69]
Rethinking semantic segmen- tation from a sequence-to-sequence perspective with trans- formers
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmen- tation from a sequence-to-sequence perspective with trans- formers. In Proceedings of the IEEE/CVF conference...
-
[70]
Semantics distortion and style matter: Towards source-free uda for panoramic segmentation
Xu Zheng, Pengyuan Zhou, Athanasios V Vasilakos, and Lin Wang. Semantics distortion and style matter: Towards source-free uda for panoramic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 27885–27895, 2024. 2, 6, 7
2024
-
[71]
Complementary bi-directional fea- ture compression for indoor 360deg semantic segmentation with self-distillation
Zishuo Zheng, Chunyu Lin, Lang Nie, Kang Liao, Zhijie Shen, and Yao Zhao. Complementary bi-directional fea- ture compression for indoor 360deg semantic segmentation with self-distillation. In Proceedings of the IEEE/CVF Win- ter Conference on Applications of Computer Vision , ...
2023
-
[72]
Vision mamba: Efficient visual representation learning with bidirectional state space model
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang. Vision mamba: Efficient visual representation learning with bidirectional state space model. arXiv preprint arXiv:2401.09417, 2024. 2, 3, 4
2024 arXiv
-
[73]
Samba: Semantic seg- mentation of remotely sensed images with state space model
Qinfeng Zhu, Yuanzhi Cai, Yuan Fang, Yihan Yang, Cheng Chen, Lei Fan, and Anh Nguyen. Samba: Semantic seg- mentation of remotely sensed images with state space model. Heliyon, 10(19), 2024. 2
2024
-
[74]
De- formable convnets v2: More deformable, better results
Xizhou Zhu, Han Hu, Stephen Lin, and Jifeng Dai. De- formable convnets v2: More deformable, better results. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 9308–9316, 2019. 4
2019
-
[75]
Deformable detr: Deformable trans- formers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable trans- formers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020. 4
2010 arXiv
-
[2024]
1, 2, 3, 4, 5, 6, 7, 8
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.