REVIEW 4 major objections 8 minor 115 references
A Novel Scene Coupling Semantic Mask Network for Remote Sensing Image Segmentation
T0 review · 4 major / 8 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Rebuilding vanilla spatial attention with scene coupling and local-global semantic masks yields the best reported remote sensing segmentation results on four benchmarks.
desk verdict A useful, lightweight attention decoder with a new combination of ideas; the headline SOTA claim needs error bars and an explicit ablation split. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the reconstructed attention affinity: instead of query, key, and value all being the same feature map, SCSM uses local class masks as keys and global class masks as values, and scales the affinity with a scene-global representation. The scene-global representation is computed by a 2D discrete cosine transform over feature maps, keeping the top 16 frequency components selected by an ImageNet-pretrained frequency prior and concatenating them along the channel dimension; this is the scene global representation. The scene object distribution comes from ROPE+, a 2D rotary position embedding that rotates queries and keys by angles tied to their row and column positions, so the dot product encodes relative object layout. Semantic masks are generated by the Semantic Mask Generation module, which assigns each pixel the class center of its pre-classified label and splits the feature map into local blocks, yielding local masks with spatial prior and global masks. These pieces jointly turn vanilla attention into the SCSM decoder head, and each is isolated in the paper's ablations.
What would settle it
Replace the ImageNet-selected frequency list in the DCT branch with frequencies ranked by importance on the target dataset itself, keeping everything else fixed; if mIoU on LoveDA, Vaihingen, Potsdam, or iSAID does not fall, then the frequency-prior global scene representation is not doing the causal work the paper assigns to it.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that the weaknesses of vanilla spatial attention for remote sensing, namely dense affinity that collects background noise, blindness to the spatial arrangement of geospatial objects, and fragility under large intra-class variance, can be remedied by redefining what is compared inside the attention operation. SCSM rewrites the query, key, and value so that the query stays pixel features while the key becomes a local semantic mask and the value becomes a global semantic mask, and it scales the affinity score by a scene-global representation assembled from the top-M two-dimensional DCT components. The scene object distribution is injected through ROPE+, a 2D rotary position embedding whose inner product encodes relative object positions without learned absolute positions. The paper claims this reconstructed attention achieves the best reported results on LoveDA at 54.6 mIoU, on Vaihingen at 91.59 AF, 84.68 mIoU, and 92.22 OA, on Potsdam at 93.60 AF, 87.79 mIoU, and 92.13 OA, and on iSAID at 66.9 mIoU, and does so with a decoder head of 2.4 million parameters and 40.5 GFLOPs.
Load-bearing premise
The global scene representation depends on frequency rankings learned on ImageNet and transferred unchanged to aerial and satellite imagery; if those rankings are not representative of remote sensing scenes, scene coupling could inject misleading context instead of helping.
Editorial extensions
If this is right
- If the central claim is correct, attention-based segmentation heads for aerial and satellite imagery can be made both more accurate and lighter by replacing dense pixel affinity with scene-coupled, mask-guided affinity; SCSM reports 2.4M parameters and 40.5 GFLOPs against 10.4 to 23.9M parameters and 154.9 to 503 GFLOPs for the compared context modules.
- The frequency-domain global scene representation would make SCSM the first dual-domain attention model for remote sensing segmentation, so follow-up work can build on selecting scene-specific frequency components rather than spatial features alone.
- The split local-global semantic masks provide class-level context that avoids background interference, which should particularly benefit small and high-intra-class-variance objects; the paper reports notable gains on cars, helicopters, large vehicles, and small vehicles.
- SCSM's additional experiment on lithological unit classification suggests the method transfers beyond benchmark land-cover datasets to geoscience mapping with limited compute, at 30.19M parameters and 6.54 GFLOPs on that task.
- The paper's concise reformulation of vanilla attention offers a new baseline that other attention designs can adopt or modify, potentially extending the same scene-coupling and mask strategies to transformer-based remote sensing models.
Reading between the lines
- An implication the paper leaves implicit is that the same two-pronged fix would likely transfer to other dense prediction tasks on overhead imagery, such as building footprint extraction or instance segmentation, because the scene-object co-occurrence patterns ROPE+ encodes are not segmentation-specific.
- The ImageNet frequency prior is the transfer-risk point: re-ranking the DCT components on remote sensing data would be a cheap test of whether the global scene representation truly carries the reported gains or mostly reflects the ImageNet statistics.
- Because the split local-global masks and ROPE+ add no learned positional parameters, the decoder stays lightweight; this suggests the recipe could serve as a drop-in attention replacement in transformer-based remote sensing models, where dense attention is the main cost.
- If the claim is right, the biggest gains should appear precisely in classes with strong scene co-occurrence, such as cars on roads and buildings along roads, and the per-class results point that way, but a controlled experiment varying class co-occurrence would test it directly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SCSM, a decoder that reconstructs vanilla spatial attention for remote sensing image segmentation. The method augments queries and keys with a 2D rotary position embedding (ROPE+) and a global scene representation extracted via a 2D discrete cosine transform with a frequency prior learned on ImageNet; it also replaces the key/value of attention with local and global semantic masks generated from a pre-classification branch. Experiments on LoveDA, Vaihingen, Potsdam, and iSAID report state-of-the-art results (e.g., 54.6 mIoU on LoveDA, 84.68 mIoU on Vaihingen, 87.79 mIoU on Potsdam, 66.9 mIoU on iSAID), and an additional lithological classification experiment reports mean plus/minus standard deviation. Ablations cover frequency count, block size, rotation angles, architecture structure, and loss coefficients.
Significance. If the reported results are reproducible, SCSM would be a lightweight and efficient attention-based decoder with strong performance on multiple remote sensing benchmarks, and the ablation study would support the value of scene coupling and semantic masking. The paper provides a thorough comparison with many recent methods and includes efficiency metrics (Table 7). However, the central claim of surpassing LOGCAN++ is currently supported by single-run numbers with no error bars, and several technical and reporting issues (Eq. 26, ablation split) must be resolved before the evidence is convincing.
major comments (4)
- [4.2.2, Eq. (26)] Equation (26) is dimensionally inconsistent. The global representation G defined in Eq. (25) is a vector formed by concatenating M scalar DCT coefficients along the channel dimension. Substituting this vector into t_{m,n} = (G q_m^T) k_n / sqrt(C) produces a vector-valued similarity (or an outer product) rather than the scalar similarity required by Eq. (18). Please provide the exact intended operation (e.g., a scalar weighting, a learned linear projection, or an element-wise modulation of q/k before the inner product) and correct the formula accordingly.
- [6.5] The ablation studies (Tables 4-6 and 8-10) are all described as performed 'on the Loveda dataset', but the dataset split used is not stated. Section 5.1 defines training (2522), validation (1669), and test (1796) splits, with the test set evaluated online. The best configuration found in the ablations (M=16, block size 21, loss coefficient 0.8) yields exactly the mIoU of 54.6 reported for the final test-set comparison in Table 1. If the ablations were run on the test set, the reported number is a grid-selected maximum rather than an unbiased performance estimate, making the comparison against LOGCAN++ unfair. Please state the split used for each ablation table and, if the test split was used, re-evaluate the final model on a held-out split.
- [Tables 1-3] The central claim that SCSM 'surpasses' LOGCAN++ relies on small margins, notably 0.22 mIoU on Potsdam (87.79 vs. 87.57 in Table 2). No standard deviations, confidence intervals, or numbers of repeated runs are reported for any benchmark comparison. Since Section 7 reports mean plus/minus standard deviation for the real-world experiment, the authors evidently have the infrastructure to run repeated trials; please provide multi-run statistics (or a clear justification for single-run reporting) for the main benchmark tables.
- [4.2.2 and 6.5] The contribution of the DCT-based global scene representation is never isolated in the ablations. Table 8 compares Base+SCA with Base+GA, but SCA includes both ROPE+ and the DCT representation; Table 9 ablates ROPE+ but does not remove the DCT branch or replace it with a spatial-only global pooling baseline. Consequently, the reader cannot tell whether the frequency-domain global representation (and its ImageNet-derived frequency prior) is responsible for the observed gains or whether the improvement comes entirely from ROPE+ and the semantic masks. An ablation with SCA minus the DCT branch (and ideally with the frequency prior replaced by a remote-sensing-native selection) is needed.
minor comments (8)
- [Abstract and Section 4.2.2] The paper claims 'the first dual-domain attention model for semantic segmentation of remote sensing images' but does not survey or cite existing dual-domain/frequency-domain attention methods; please either substantiate the claim with a comparison or temper it.
- [6.1.1 and 6.2.1] The text identifies PoolFormer and ConvNeXt as the previous state-of-the-art on LoveDA and Vaihingen, respectively, but Table 1 and Table 2 show that LOGCAN++ outperforms both. The narrative should consistently identify LOGCAN++ as the strongest prior method.
- [6.1.1 and 6.2.2] Both subsections are titled 'Qualitative analysis', but 6.1.1 is a quantitative discussion; the section titles should be corrected.
- [Table 7 and surrounding text] The module is referred to as 'SMG+CCA (Ours)' in Table 7, but the module is named SCA throughout the rest of the paper; use consistent terminology.
- [5.2, Eq. (35)] The text 'recision measures' should be 'precision measures'.
- [Abstract] The phrase 'demonstrate the the effectiveness' contains a duplicated 'the', and 'inter-clean and elegant mathematical representations' is unclear and should be rephrased.
- [4.3, Eq. (27)] The notation for the recover function ψ and the tensor product ⊗ is not defined precisely; please clarify the shapes of the intermediate tensors to make the module reproducible.
- [7.1] The overlapping regions (zones T/U overlap by 84.4%, zones 44/45 overlap by 75%) combined with random sampling from all regions A-U could allow spatially overlapping patches to appear in both training and test sets; please describe the spatial decontamination strategy or clarify how the overlap is handled.
Circularity Check
No significant circularity; the architecture is defined by external priors and external benchmarks, but the LoveDA ablations do not disclose their split and the reported best mIoU equals the test mIoU, which is a reporting/selection-risk rather than a demonstrated circular reduction.
full rationale
The paper's claimed derivation is an architectural reconstruction of vanilla attention, not a mathematical derivation of its own benchmark scores. Scene coupling is built from a DCT frequency prior imported from external work FCA-Net [86] (Eqs. 19-26) and a rotary position embedding extended from external RoFormer [84] (Eqs. 9-18); the semantic-mask path replaces the vanilla q/k/v with local and global class masks generated from a pre-classification branch (Eqs. 27-29). None of these equations is defined in terms of the reported mIoU/F1/OA values, so the central claim does not reduce to its inputs by construction. The only self-authored works cited, LOGCAN++ [33] and DocNet [36], serve as comparison baselines and as contrast points for the spatial-prior class centers; they do not supply a load-bearing premise. The principal concern is procedural rather than circular: Section 6.5 states that all ablations (frequency count, block size, rotation angles, structure, ROPE+, loss coefficients) were conducted on 'the LoveDA dataset' without specifying the 1,669-image validation split defined in Section 5.1, and the best ablation mIoU of 54.6 in Tables 4-6, 8-10 equals the reported LoveDA test mIoU of 54.6 in Table 1. If the grid was evaluated on the test split, the headline 54.6 would be a selected maximum rather than an independent estimate; the paper does not say that, so this is a missing-disclosure and error-bar issue (correctness risk), not a demonstrated circularity. Accordingly, the circularity score is low, reflecting only the minor self-citations and the unresolved selection-risk.
Assumptions & free parameters
free parameters (3)
- Frequency count M =
16
- Block size P =
21
- Pre-classification loss coefficient =
0.8
assumptions (4)
- domain assumption Remote sensing images contain complex backgrounds, large intra-class variance, and intrinsic spatial correlations among geospatial objects.
- domain assumption The frequency prior learned on ImageNet transfers to remote sensing scenes.
- domain assumption Class centers computed from the argmax of the pre-classification representation are reliable spatial priors.
- standard math Standard DCT orthogonality and RoPE rotation properties hold as used in the derivations.
invented entities (2)
-
ROPE+ (image-level 2D rotary position embedding)
-
Scene global representation from 2D DCT
Cite this review
Pith. "Pith review of A Novel Scene Coupling Semantic Mask Network for Remote Sensing Image Segmentation." pith.science (2026). https://pith.science/paper/NAUCCFRL
@misc{pith2026250113130,
author = {Pith},
title = {Pith review of: A Novel Scene Coupling Semantic Mask Network for Remote Sensing Image Segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/NAUCCFRL}},
note = {Machine review of arXiv:2501.13130}
}
read the original abstract
As a common method in the field of computer vision, spatial attention mechanism has been widely used in semantic segmentation of remote sensing images due to its outstanding long-range dependency modeling capability. However, remote sensing images are usually characterized by complex backgrounds and large intra-class variance that would degrade their analysis performance. While vanilla spatial attention mechanisms are based on dense affine operations, they tend to introduce a large amount of background contextual information and lack of consideration for intrinsic spatial correlation. To deal with such limitations, this paper proposes a novel scene-Coupling semantic mask network, which reconstructs the vanilla attention with scene coupling and local global semantic masks strategies. Specifically, scene coupling module decomposes scene information into global representations and object distributions, which are then embedded in the attention affinity processes. This Strategy effectively utilizes the intrinsic spatial correlation between features so that improve the process of attention modeling. Meanwhile, local global semantic masks module indirectly correlate pixels with the global semantic masks by using the local semantic mask as an intermediate sensory element, which reduces the background contextual interference and mitigates the effect of intra-class variance. By combining the above two strategies, we propose the model SCSM, which not only can efficiently segment various geospatial objects in complex scenarios, but also possesses inter-clean and elegant mathematical representations. Experimental results on four benchmark datasets demonstrate the the effectiveness of the above two strategies for improving the attention modeling of remote sensing images. The dataset and code are available at https://github.com/xwmaxwma/rssegmentation
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
Zhang, K
Q. Zhang, K. C. Seto, Mapping urbanization dynamics at regional and global scales using multi-temporal dmsp/ols nighttime light data, Remote Sensing of Environment 115 (9) (2011) 2320–2329
2011
-
[2]
Huang, B
B. Huang, B. Zhao, Y. Song, Urban land-use mapping using a deep convolutionalneuralnetworkwithhighspatialresolutionmultispec- tral remote sensing imagery, Remote Sensing of Environment 214 (2018) 73–86
2018
-
[3]
J.He,T.Nie,W.Ma,Geolocationrepresentationfromlargelanguage models are generic enhancers for spatio-temporal learning, arXiv preprint arXiv:2408.12116 (2024)
work page Pith review arXiv 2024
-
[4]
Q. Yuan, H. Shen, T. Li, Z. Li, S. Li, Y. Jiang, H. Xu, W. Tan, Q. Yang, J. Wang, et al., Deep learning in environmental remote sensing:Achievementsandchallenges,RemoteSensingofEnviron- ment 241 (2020) 111716
2020
-
[5]
K. Cui, W. Tang, R. Zhu, M. Wang, G. D. Larsen, V. P. Pauca, S.Alqahtani,F.Yang,D.Segurado,P.Fine,etal.,Real-timelocaliza- tion and bimodal point pattern analysis of palms using uav imagery, arXiv preprint arXiv:2410.11124 (2024)
arXiv 2024
-
[6]
M.Maboudi,J.Amini,S.Malihi,M.Hahn,Integratingfuzzyobject basedimageanalysisandantcolonyoptimizationforroadextraction from remotely sensed images, ISPRS Journal of Photogrammetry and Remote Sensing 138 (2018) 151–163
2018
-
[7]
Z. Wang, B. Li, C. Wang, S. Scherer, AirShot: Efficient few-shot detection for autonomous exploration, in: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2024. URL https://arxiv.org/pdf/2404.05069.pdf
arXiv 2024
-
[8]
Wang, ONLS: OPTIMAL NOISE LEVEL SEARCH IN DIF- FUSION AUTOENCODERS WITHOUT FINE-TUNING, in: The Second Tiny Papers Track at ICLR 2024, 2024
Z. Wang, ONLS: OPTIMAL NOISE LEVEL SEARCH IN DIF- FUSION AUTOENCODERS WITHOUT FINE-TUNING, in: The Second Tiny Papers Track at ICLR 2024, 2024. URL https://openreview.net/forum?id=Q8diCUHTZd
2024
Show all 115 references
-
[9]
Bastani, P
F. Bastani, P. Wolters, R. Gupta, J. Ferdinando, A. Kembhavi, Satlaspretrain:Alarge-scaledatasetforremotesensingimageunder- standing,in:ProceedingsoftheIEEE/CVFInternationalConference on Computer Vision (ICCV), 2023, pp. 16772–16782
2023
-
[10]
R. Liu, T. Luo, S. Huang, Y. Wu, Z. Jiang, H. Zhang, Crossmatch: Cross-viewmatchingforsemi-supervisedremotesensingimageseg- mentation, IEEE Transactions on Geoscience and Remote Sensing (2024)
2024
-
[11]
T. Luo, M. Du, J. Shi, X. Chen, B. Zhao, S. Huang, Contextuality helpsrepresentationlearningforgeneralizedcategorydiscovery,in: 2024 IEEE International Conference on Image Processing (ICIP), 2024, pp. 687–693.doi:10.1109/ICIP51287.2024.10647298
2024
-
[12]
K. Cui, R. Li, S. L. Polk, Y. Lin, H. Zhang, J. M. Murphy, R. J. Plemmons, R. H. Chan, Superpixel-based and spatially-regularized diffusion learning for unsupervised hyperspectral image clustering, IEEE Transactions on Geoscience and Remote Sensing (2024)
2024
-
[13]
K.Cui,R.Li,S.L.Polk,J.M.Murphy,R.J.Plemmons,R.H.Chan, Unsupervised spatial-spectral hyperspectral image reconstruction and clustering with diffusion geometry, in: 2022 12th Workshop on HyperspectralImagingandSignalProcessing: EvolutioninRemote Sensing (WHISPERS), IEEE, 2022, pp. 1–5
2022
-
[14]
Y.Chen,C.Liu,X.Liu,R.Arcucci,Z.Xiong,Bimcv-r:Alandmark dataset for 3d ct text-image retrieval, in: MICCAI, Springer, 2024, pp. 124–134
2024
-
[15]
Y. Chen, W. Huang, X. Liu, S. Deng, Q. Chen, Z. Xiong, Learn- ing multiscale consistency for self-supervised electron microscopy instance segmentation, in: ICASSP, IEEE, 2024, pp. 1566–1570
2024
-
[16]
J. Xue, B. Su, Significant remote sensing vegetation indices: A reviewofdevelopmentsandapplications,Journalofsensors2017(1) (2017) 1353691
2017
-
[17]
J. Long, E. Shelhamer, T. Darrell, Fully convolutional networks for semantic segmentation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3431–3440
2015
-
[18]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez,Ł.Kaiser,I.Polosukhin,Attentionisallyouneed,Advances in neural information processing systems 30 (2017)
2017
-
[19]
Zhang, P
Y. Zhang, P. Ji, A. Wang, J. Mei, A. Kortylewski, A. L. Yuille, 3d-aware neural body fitting for occlusion robust 3d human pose estimation, in: IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 9365–9376.doi:10.1109/ICCV51070.2023. 00862. URL https://doi.o...
2023
- [20]
-
[21]
2260–2271
T.Nie,G.Qin,W.Ma,Y.Mei,J.Sun,Imputeformer:Lowrankness- induced transformers for generalizable spatiotemporal imputation, in: Proceedings of the 30th ACM SIGKDD Conference on Knowl- edge Discovery and Data Mining, 2024, pp. 2260–2271. X.Ma et al.:Preprint submitted to Elsevier ...
2024
-
[22]
Trias-Sanz, G
R. Trias-Sanz, G. Stamon, J. Louchet, Using colour, texture, and hierarchial segmentation for high-resolution remote sensing, ISPRS Journal of Photogrammetry and remote sensing 63 (2) (2008) 156– 168
2008
-
[23]
J. Jin, H. Xu, P. Ji, B. Leng, Imc-net: Learning implicit field with corner attention network for 3d shape reconstruction, in: IEEE International Conference on Image Processing (ICIP), 2022, pp. 1591–1595. doi:10.1109/ICIP46576.2022.9897709. URL https://doi.org/10.1109/ICIP465...
2022
- [24]
-
[25]
Y. Chen, W. Huang, S. Zhou, Q. Chen, Z. Xiong, Self-supervised neuron segmentation with multi-agent reinforcement learning, in: IJCAI, 2023, pp. 609–617
2023
-
[26]
H. Qian, Y. Chen, S. Lou, F. Khan, X. Jin, D.-P. Fan, Maskfactory: Towards high-quality synthetic data generation for dichotomous image segmentation, in: NeurIPS, 2024
2024
-
[27]
H. Zhao, J. Shi, X. Qi, X. Wang, J. Jia, Pyramid scene parsing network,in:ProceedingsoftheIEEEconferenceoncomputervision and pattern recognition, 2017, pp. 2881–2890
2017
-
[28]
L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, H. Adam, Encoder- decoder with atrous separable convolution for semantic image seg- mentation,in:ProceedingsoftheEuropeanconferenceoncomputer vision (ECCV), 2018, pp. 801–818
2018
-
[30]
Y. Chen, H. Shi, X. Liu, T. Shi, R. Zhang, D. Liu, Z. Xiong, F.Wu,Tokenunify:Scalableautoregressivevisualpre-trainingwith mixture token prediction, arXiv preprint arXiv:2405.16847 (2024)
2024 arXiv
-
[31]
J. He, Z. Deng, Y. Qiao, Dynamic multi-scale filters for semantic segmentation, in: Proceedings of the IEEE/CVF International Con- ference on Computer Vision, 2019, pp. 3562–3572
2019
-
[32]
Z. Jin, B. Liu, Q. Chu, N. Yu, Isnet: Integrate image-level and semantic-levelcontextforsemanticsegmentation,in:Proceedingsof theIEEE/CVFInternationalConferenceonComputerVision,2021, pp. 7189–7198
2021
-
[33]
X. Ma, R. Lian, Z. Wu, H. Guo, M. Ma, S. Wu, Z. Du, S. Song, W. Zhang, Logcan++: Adaptive local-global class-aware network forsemanticsegmentationofremotesensingimagery(2024). arXiv: 2406.16502. URL https://arxiv.org/abs/2406.16502
2024 arXiv
-
[34]
J. Fu, J. Liu, H. Tian, Y. Li, Y. Bao, Z. Fang, H. Lu, Dual attention network for scene segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 3146–3154
2019
-
[35]
J. Hu, L. Shen, G. Sun, Squeeze-and-excitation networks, in: Pro- ceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7132–7141
2018
-
[36]
X. Ma, R. Che, X. Wang, M. Ma, S. Wu, T. Feng, W. Zhang, Docnet: Dual-domain optimized class-aware network for remote sensingimagesegmentation,IEEEGeoscienceandRemoteSensing Letters (2024)
2024
-
[37]
Huang, X
Z. Huang, X. Wang, L. Huang, C. Huang, Y. Wei, W. Liu, Ccnet: Criss-cross attention for semantic segmentation, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 603–612
2019
-
[38]
Q. Song, J. Li, C. Li, H. Guo, R. Huang, Fully attentional network forsemanticsegmentation,in:ProceedingsoftheAAAIConference on Artificial Intelligence, Vol. 36, 2022, pp. 2280–2288
2022
-
[39]
X. Li, H. He, X. Li, D. Li, G. Cheng, J. Shi, L. Weng, Y. Tong, Z.Lin,Pointflow:Flowingsemanticsthroughpointsforaerialimage segmentation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 4217–4226
2021
-
[40]
A. Gu, T. Dao, Mamba: Linear-time sequence modeling with selec- tive state spaces, arXiv preprint arXiv:2312.00752 (2023)
2023 arXiv
-
[41]
T. Dao, A. Gu, Transformers are ssms: Generalized models and efficient algorithms through structured state space duality, arXiv preprint arXiv:2405.21060 (2024)
2024 arXiv
-
[42]
X.He,Y.Zhou,J.Zhao,D.Zhang,R.Yao,Y.Xue,Swintransformer embedding unet for remote sensing image semantic segmentation, IEEE Transactions on Geoscience and Remote Sensing 60 (2022) 1–15
2022
-
[43]
Zhang, W
C. Zhang, W. Jiang, Y. Zhang, W. Wang, Q. Zhao, C. Wang, Transformer and cnn hybrid deep neural network for semantic seg- mentation of very-high-resolution remote sensing imagery, IEEE Transactions on Geoscience and Remote Sensing 60 (2022) 1–20
2022
-
[44]
L.Ding,D.Lin,S.Lin,J.Zhang,X.Cui,Y.Wang,H.Tang,L.Bruz- zone,Lookingoutsidethewindow:Wide-contexttransformerforthe semantic segmentation of high-resolution remote sensing images, IEEE Transactions on Geoscience and Remote Sensing 60 (2022) 1–13
2022
-
[45]
Y. Liu, Y. Zhang, Y. Wang, S. Mei, Rethinking transformers for semanticsegmentationofremotesensingimages,IEEETransactions on Geoscience and Remote Sensing (2023)
2023
-
[46]
Sun, Ultra-high resolution segmentation via boundary-enhanced patch-merging transformer (2024).arXiv:2412.10181
H. Sun, Ultra-high resolution segmentation via boundary-enhanced patch-merging transformer (2024).arXiv:2412.10181. URL https://arxiv.org/abs/2412.10181
2024 arXiv
-
[47]
Cheng, H
S. Cheng, H. Sun, Spt: Sequence prompt transformer for interactive image segmentation (2024).arXiv:2412.10224. URL https://arxiv.org/abs/2412.10224
2024 arXiv
-
[48]
S.Zhao,H.Chen,X.Zhang,P.Xiao,L.Bai,W.Ouyang,Rs-mamba for large remote sensing image dense prediction, arXiv preprint arXiv:2404.02668 (2024)
2024 arXiv
-
[49]
K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778
2016
-
[50]
M. Yang, K. Yu, C. Zhang, Z. Li, K. Yang, Denseaspp for semantic segmentation in street scenes, in: Proceedings of the IEEE confer- ence on computer vision and pattern recognition, 2018, pp. 3684– 3692
2018
-
[51]
Huang, D
Y. Huang, D. Kang, W. Jia, L. Liu, X. He, Channelized axial attention–considering channel relation within spatial attention for semanticsegmentation,in:ProceedingsoftheAAAIConferenceon Artificial Intelligence, Vol. 36, 2022, pp. 1016–1025
2022
-
[52]
G. Deng, Z. Wu, C. Wang, M. Xu, Y. Zhong, Ccanet: Class- constraint coarse-to-fine attentional deep network for subdecimeter aerial image semantic segmentation, IEEE Transactions on Geo- science and Remote Sensing 60 (2022) 1–20. doi:10.1109/TGRS. 2021.3055950
2022
-
[53]
Liang, W
C. Liang, W. Wang, J. Miao, Y. Yang, Gmmseg: Gaussian mixture based generative semantic segmentation models, arXiv preprint arXiv:2210.02025 (2022)
2022 arXiv
-
[54]
2582–2593
T.Zhou,W.Wang,E.Konukoglu,L.VanGool,Rethinkingsemantic segmentation: A prototype view, in: Proceedings of the IEEE/CVF ConferenceonComputerVisionandPatternRecognition,2022,pp. 2582–2593
2022
-
[55]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al., An image is worth 16x16 words: Transformers for image recognition at scale, arXiv preprint arXiv:2010.11929 (2020)
2020 arXiv
-
[56]
E.Xie,W.Wang,Z.Yu,A.Anandkumar,J.M.Alvarez,P.Luo,Seg- former: Simple and efficient design for semantic segmentation with transformers, Advances in Neural Information Processing Systems 34 (2021) 12077–12090
2021
-
[57]
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, B. Guo, Swin transformer: Hierarchical vision transformer using shifted windows,in:ProceedingsoftheIEEE/CVFinternationalconference on computer vision, 2021, pp. 10012–10022
2021
-
[58]
I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreit, et al., Mlp-mixer: An all-mlp architecture for vision, Advances in neural X.Ma et al.:Preprint submitted to Elsevier Page 22 of 24 A Novel Scene Coupl...
2021
-
[59]
H. Liu, Z. Dai, D. So, Q. V. Le, Pay attention to mlps, Advances in Neural Information Processing Systems 34 (2021) 9204–9215
2021
-
[60]
Cheng, A
B. Cheng, A. Schwing, A. Kirillov, Per-pixel classification is not all you need for semantic segmentation, Advances in Neural Informa- tion Processing Systems 34 (2021) 17864–17875
2021
-
[61]
Z. Wu, J. Lu, J. Han, L. Bai, Y. Zhang, Z. Zhao, S. Song, Domain separation graph neural networks for saliency object ranking, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 3964–3974
2024
-
[62]
H. Bao, L. Dong, S. Piao, F. Wei, Beit: Bert pre-training of image transformers, arXiv preprint arXiv:2106.08254 (2021)
2021 arXiv
-
[63]
H. Sun, L. Xu, S. Jin, P. Luo, C. Qian, W. Liu, Program: Prototype graphmodelbasedpseudo-labellearningfortest-timeadaptation,in: TheTwelfthInternationalConferenceonLearningRepresentations, 2024
2024
-
[64]
J.Jiao,Y.-M.Tang,K.-Y.Lin,Y.Gao,J.Ma,Y.Wang,W.-S.Zheng, Dilateformer:Multi-scaledilatedtransformerforvisualrecognition, IEEE Transactions on Multimedia (2023)
2023
-
[65]
7262–7272
R.Strudel,R.Garcia,I.Laptev,C.Schmid,Segmenter:Transformer for semantic segmentation, in: Proceedings of the IEEE/CVF inter- national conference on computer vision, 2021, pp. 7262–7272
2021
-
[66]
L. Dai, G. Zhang, R. Zhang, Radanet: Road augmented deformable attention network for road extraction from complex high-resolution remote-sensing images, IEEE Transactions on Geoscience and Re- mote Sensing (2023)
2023
-
[67]
Luo, J.-X
L. Luo, J.-X. Wang, S.-B. Chen, J. Tang, B. Luo, Bdtnet: Road extraction by bi-direction transformer from remote sensing images, IEEE Geoscience and Remote Sensing Letters 19 (2022) 1–5
2022
-
[68]
S. Dong, Z. Chen, Block multi-dimensional attention for road seg- mentationinremotesensingimagery,IEEEGeoscienceandRemote Sensing Letters 19 (2021) 1–5
2021
-
[69]
Jung, H.-S
H. Jung, H.-S. Choi, M. Kang, Boundary enhancement semantic segmentation for building extraction from remote sensed image, IEEE Transactions on Geoscience and Remote Sensing 60 (2021) 1–12
2021
-
[70]
Yuan, Learning building extraction in aerial scenes with convolu- tional networks, IEEE transactions on pattern analysis and machine intelligence 40 (11) (2017) 2793–2798
J. Yuan, Learning building extraction in aerial scenes with convolu- tional networks, IEEE transactions on pattern analysis and machine intelligence 40 (11) (2017) 2793–2798
2017
-
[71]
A. Alem, S. Kumar, Transfer learning models for land cover and land use classification in remote sensing image, Applied Artificial Intelligence 36 (1) (2022) 2014192
2022
-
[72]
C.Zhang,I.Sargent,X.Pan,H.Li,A.Gardiner,J.Hare,P.M.Atkin- son, Joint deep learning for land cover and land use classification, Remote sensing of environment 221 (2019) 173–187
2019
-
[73]
doi:10.1109/TGRS.2021.3119537
R.Zuo,G.Zhang,R.Zhang,X.Jia,Adeformableattentionnetwork for high-resolution remote sensing images semantic segmentation, IEEE Transactions on Geoscience and Remote Sensing 60 (2022) 1–14. doi:10.1109/TGRS.2021.3119537
2022
-
[74]
R. Guan, W. Tu, Z. Li, H. Yu, D. Hu, Y. Chen, C. Tang, Q. Yuan, X. Liu, Spatial-spectral graph contrastive clustering with hard sam- ple mining for hyperspectral images, IEEE Transactions on Geo- science and Remote Sensing (2024) 1–16
2024
-
[75]
R. Guan, Z. Li, W. Tu, J. Wang, Y. Liu, X. Li, C. Tang, R. Feng, Contrastive multiview subspace clustering of hyperspectral images basedongraphconvolutionalnetworks,IEEETransactionsonGeo- science and Remote Sensing 62 (2024) 1–14
2024
-
[76]
R.Li,S.Zheng,C.Zhang,C.Duan,J.Su,L.Wang,P.M.Atkinson, Multiattentionnetworkforsemanticsegmentationoffine-resolution remote sensing images, IEEE Transactions on Geoscience and Re- mote Sensing 60 (2021) 1–13
2021
-
[77]
Zheng, Y
Z. Zheng, Y. Zhong, J. Wang, A. Ma, Foreground-aware relation networkforgeospatialobjectsegmentationinhighspatialresolution remote sensing imagery, in: Proceedings of the IEEE/CVF confer- ence on computer vision and pattern recognition, 2020, pp. 4096– 4105
2020
-
[78]
1809–1818
F.Yang,C.Ma,Sparseandcompletelatentorganizationforgeospa- tial semantic segmentation, in: Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 2022, pp. 1809–1818
2022
-
[79]
X. Wang, R. Girshick, A. Gupta, K. He, Non-local neural networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7794–7803
2018
-
[80]
Rottensteiner, G
F. Rottensteiner, G. Sohn, J. Jung, M. Gerke, C. Baillard, S. Bnitez, U. Breitkopf, International society for photogrammetry and remote sensing,2dsemanticlabelingcontest,available: https://www.isprs. org/education/benchmarks/UrbanSemLab (Accessed: Oct. 29, 2020.)
2020
-
[81]
J.Wang,Z.Zheng,A.Ma,X.Lu,Y.Zhong,Loveda:Aremotesens- ing land-cover dataset for domain adaptive semantic segmentation, arXiv preprint arXiv:2110.08733 (2021)
2021 arXiv
-
[82]
Waqas Zamir, A
S. Waqas Zamir, A. Arora, A. Gupta, S. Khan, G. Sun, F. Shah- baz Khan, F. Zhu, L. Shao, G.-S. Xia, X. Bai, isaid: A large-scale dataset for instance segmentation in aerial images, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops,...
2019
-
[83]
Zheng, Y
Z. Zheng, Y. Zhong, J. Wang, A. Ma, L. Zhang, Farseg++: Foreground-aware relation network for geospatial object segmen- tation in high spatial resolution remote sensing imagery, IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)
2023
-
[84]
J.Su,M.Ahmed,Y.Lu,S.Pan,W.Bo,Y.Liu,Roformer:Enhanced transformer with rotary position embedding, Neurocomputing 568 (2024) 127063
2024
-
[85]
P. Shaw, J. Uszkoreit, A. Vaswani, Self-attention with relative posi- tion representations, arXiv preprint arXiv:1803.02155 (2018)
2018 arXiv
-
[86]
Z.Qin,P.Zhang,F.Wu,X.Li,Fcanet:Frequencychannelattention networks,in:ProceedingsoftheIEEE/CVFinternationalconference on computer vision, 2021, pp. 783–792
2021
-
[87]
Huang, Z
Z. Huang, Z. Zhang, C. Lan, Z.-J. Zha, Y. Lu, B. Guo, Adaptive frequency filters as efficient global token mixers, in: Proceedings of theIEEE/CVFInternationalConferenceonComputerVision,2023, pp. 6049–6059
2023
-
[88]
Kirillov, R
A. Kirillov, R. Girshick, K. He, P. Dollár, Panoptic feature pyramid networks,in:ProceedingsoftheIEEE/CVFconferenceoncomputer vision and pattern recognition, 2019, pp. 6399–6408
2019
-
[89]
L.Ding,H.Tang,L.Bruzzone,Lanet:Localattentionembeddingto improvethesemanticsegmentationofremotesensingimages,IEEE TransactionsonGeoscienceandRemoteSensing59(1)(2020)426– 435
2020
-
[90]
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, S. Xie, A convnetforthe2020s,in:ProceedingsoftheIEEE/CVFConference on Computer Vision and Pattern Recognition, 2022, pp. 11976– 11986
2022
-
[91]
W. Yu, M. Luo, P. Zhou, C. Si, Y. Zhou, X. Wang, J. Feng, S. Yan, Metaformer is actually what you need for vision, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recogni- tion, 2022, pp. 10819–10829
2022
-
[92]
L. Zhu, X. Wang, Z. Ke, W. Zhang, R. W. Lau, Biformer: Vision transformer with bi-level routing attention, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2023, pp. 10323–10333
2023
-
[93]
17302–17313
H.Cai,J.Li,M.Hu,C.Gan,S.Han,Efficientvit:Lightweightmulti- scale attention for high-resolution dense prediction, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 17302–17313
2023
-
[94]
Y. Ji, Z. Chen, E. Xie, L. Hong, X. Liu, Z. Liu, T. Lu, Z. Li, P. Luo, Ddp:Diffusionmodelfordensevisualprediction,in:Proceedingsof theIEEE/CVFInternationalConferenceonComputerVision,2023, pp. 21741–21752
2023
-
[95]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: Alarge-scalehierarchicalimagedatabase,in:2009IEEEconference oncomputervisionandpatternrecognition,Ieee,2009,pp.248–255
2009
-
[96]
Zhang, Y
F. Zhang, Y. Chen, Z. Li, Z. Hong, J. Liu, F. Ma, J. Han, E. Ding, Acfnet:Attentionalclassfeaturenetworkforsemanticsegmentation, X.Ma et al.:Preprint submitted to Elsevier Page 23 of 24 A Novel Scene Coupling Semantic Mask Network for Remote Sensing Image Segmentation in:Proce...
2019
-
[97]
Cheng, L.-C
B. Cheng, L.-C. Chen, Y. Wei, Y. Zhu, Z. Huang, J. Xiong, T. S. Huang, W.-M. Hwu, H. Shi, Spgnet: Semantic prediction guidance for scene parsing, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 5218–5228
2019
-
[98]
L.-C. Chen, G. Papandreou, F. Schroff, H. Adam, Rethinking atrousconvolution forsemantic imagesegmentation,arXiv preprint arXiv:1706.05587 (2017)
2017 arXiv
-
[99]
Ronneberger, P
O. Ronneberger, P. Fischer, T. Brox, U-net: Convolutional networks for biomedical image segmentation, in: Medical image computing and computer-assisted intervention–MICCAI 2015: 18th interna- tional conference, Munich, Germany, October 5-9, 2015, proceed- ings, part III 18, Sp...
2015
-
[100]
M.Yin,Z.Yao,Y.Cao,X.Li,Z.Zhang,S.Lin,H.Hu,Disentangled non-local neural networks, in: Computer Vision–ECCV 2020: 16th EuropeanConference,Glasgow,UK,August23–28,2020,Proceed- ings, Part XV 16, Springer, 2020, pp. 191–207
2020
-
[101]
Y.Cao,J.Xu,S.Lin,F.Wei,H.Hu,Globalcontextnetworks,IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (6) (2020) 6881–6895
2020
-
[102]
Y. Yuan, X. Chen, J. Wang, Object-contextual representations for semantic segmentation, in: Computer Vision–ECCV 2020: 16th EuropeanConference,Glasgow,UK,August23–28,2020,Proceed- ings, Part VI 16, Springer, 2020, pp. 173–190
2020
-
[103]
X. Li, Z. Zhong, J. Wu, Y. Yang, Z. Lin, H. Liu, Expectation- maximization attention networks for semantic segmentation, in: ProceedingsoftheIEEE/CVFinternationalconferenceoncomputer vision, 2019, pp. 9167–9176
2019
-
[104]
J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y. Zhao, D. Liu, Y.Mu,M.Tan,X.Wang,etal.,Deephigh-resolutionrepresentation learningforvisualrecognition,IEEEtransactionsonpatternanalysis and machine intelligence 43 (10) (2020) 3349–3364
2020
-
[105]
T.Xiao,Y.Liu,B.Zhou,Y.Jiang,J.Sun,Unifiedperceptualparsing forsceneunderstanding,in:ProceedingsoftheEuropeanconference on computer vision (ECCV), 2018, pp. 418–434
2018
-
[106]
X.Li,A.You,Z.Zhu,H.Zhao,M.Yang,K.Yang,S.Tan,Y.Tong, Semantic flow for fast and accurate scene parsing, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I 16, Springer, 2020, pp. 775–793
2020
-
[107]
Carion, F
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, S. Zagoruyko, End-to-end object detection with transformers, in: European conference on computer vision, Springer, 2020, pp. 213– 229
2020
-
[108]
URL https://openreview.net/forum?id=3KWnuT-R1bh
X.Chu,Z.Tian,B.Zhang,X.Wang,C.Shen,Conditionalpositional encodings for vision transformers, in: ICLR 2023, 2023. URL https://openreview.net/forum?id=3KWnuT-R1bh
2023
-
[109]
K. Han, Y. Wang, Q. Tian, J. Guo, C. Xu, C. Xu, Ghostnet: More features from cheap operations, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 1580–1589
2020
-
[110]
Y. Tang, K. Han, J. Guo, C. Xu, C. Xu, Y. Wang, Ghostnetv2: Enhance cheap operation with long-range attention, Advances in Neural Information Processing Systems 35 (2022) 9969–9982
2022
-
[111]
Z.Zhou,M.M.RahmanSiddiquee,N.Tajbakhsh,J.Liang,Unet++: A nested u-net architecture for medical image segmentation, in: DeepLearninginMedicalImageAnalysisandMultimodalLearning forClinicalDecisionSupport:4thInternationalWorkshop,DLMIA 2018, and 8th International Workshop, ML-CDS...
2018
-
[112]
Oktay, J
O. Oktay, J. Schlemper, L. L. Folgoc, M. Lee, M. Heinrich, K. Mi- sawa, K. Mori, S. McDonagh, N. Y. Hammerla, B. Kainz, et al., Attention u-net: Learning where to look for the pancreas, arXiv preprint arXiv:1804.03999 (2018)
2018 arXiv
-
[113]
Z. Wu, J. Zhang, L. Zhang, X. Liu, H. Qiao, Bi-hrnet: A road extractionframeworkfromsatelliteimagerybasedonnodeheatmap and bidirectional connectivity, Remote Sensing 14 (7) (2022) 1732
2022
-
[114]
G.Zhou,W.Chen,X.Qin,J.Li,L.Wang,Lithologicalunitclassifi- cation based on geological knowledge-guided deep learning frame- workforopticalstereomappingsatelliteimagery,IEEETransactions on Geoscience and Remote Sensing (2023)
2023
-
[115]
Lei, D.-F
X.-F. Lei, D.-F. Duan, S.-Y. Jiang, S.-F. Xiong, Ore-forming fluids and isotopic (hocs-pb) characteristics of the fujiashan-longjiaoshan skarn w-cu-(mo) deposit in the edong district of hubei province, china, Ore Geology Reviews 102 (2018) 386–405
2018
-
[116]
X.Ma et al.:Preprint submitted to Elsevier Page 24 of 24
G.Xie,J.Mao,Q.Zhu,L.Yao,Y.Li,W.Li,H.Zhao,Geochemical constraints on cu–fe and fe skarn deposits in the edong district, middle–lower yangtze river metallogenic belt, china, Ore Geology Reviews 64 (2015) 425–444. X.Ma et al.:Preprint submitted to Elsevier Page 24 of 24
2015
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.