Pith. sign in

REVIEW 4 major objections 8 minor 115 references

A Novel Scene Coupling Semantic Mask Network for Remote Sensing Image Segmentation

T0 review · 4 major / 8 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Rebuilding vanilla spatial attention with scene coupling and local-global semantic masks yields the best reported remote sensing segmentation results on four benchmarks.

desk verdict A useful, lightweight attention decoder with a new combination of ideas; the headline SOTA claim needs error bars and an explicit ablation split. read the letter →

arxiv 2501.13130 v1 pith:NAUCCFRL submitted 2025-01-22 eess.IV

classification eess.IV
keywords semanticsegmentationremotesensingimageryscenecouplingattentionobjectdistributionglobalrepresentationlocal-globalmaskrotarypositionembeddingefficient
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that vanilla spatial attention, which compares every pixel with every other pixel, is a poor fit for remote sensing images because it pulls in background clutter, ignores the spatial arrangement of objects, and cannot handle large differences within one class. To fix this, it replaces the query, key, and value inputs of attention with scene-aware and class-aware constructions: a scene coupling module adds a global scene representation built from selected frequency components and an object distribution signal built from a 2D rotary position embedding, while semantic mask generation lets pixels attend to local and global class masks. The combined model, SCSM, is tested on LoveDA, Vaihingen, Potsdam, and iSAID, where it reports the highest mean IoU or F1 on each and outperforms the recent LOGCAN++ by 0.2 to 1.9 percentage points. If the claim is right, attention for remote sensing can be made more accurate and much cheaper by injecting scene structure rather than scaling up the attention map.

What carries the argument

The load-bearing object is the reconstructed attention affinity: instead of query, key, and value all being the same feature map, SCSM uses local class masks as keys and global class masks as values, and scales the affinity with a scene-global representation. The scene-global representation is computed by a 2D discrete cosine transform over feature maps, keeping the top 16 frequency components selected by an ImageNet-pretrained frequency prior and concatenating them along the channel dimension; this is the scene global representation. The scene object distribution comes from ROPE+, a 2D rotary position embedding that rotates queries and keys by angles tied to their row and column positions, so the dot product encodes relative object layout. Semantic masks are generated by the Semantic Mask Generation module, which assigns each pixel the class center of its pre-classified label and splits the feature map into local blocks, yielding local masks with spatial prior and global masks. These pieces jointly turn vanilla attention into the SCSM decoder head, and each is isolated in the paper's ablations.

What would settle it

Replace the ImageNet-selected frequency list in the DCT branch with frequencies ranked by importance on the target dataset itself, keeping everything else fixed; if mIoU on LoveDA, Vaihingen, Potsdam, or iSAID does not fall, then the frequency-prior global scene representation is not doing the causal work the paper assigns to it.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that the weaknesses of vanilla spatial attention for remote sensing, namely dense affinity that collects background noise, blindness to the spatial arrangement of geospatial objects, and fragility under large intra-class variance, can be remedied by redefining what is compared inside the attention operation. SCSM rewrites the query, key, and value so that the query stays pixel features while the key becomes a local semantic mask and the value becomes a global semantic mask, and it scales the affinity score by a scene-global representation assembled from the top-M two-dimensional DCT components. The scene object distribution is injected through ROPE+, a 2D rotary position embedding whose inner product encodes relative object positions without learned absolute positions. The paper claims this reconstructed attention achieves the best reported results on LoveDA at 54.6 mIoU, on Vaihingen at 91.59 AF, 84.68 mIoU, and 92.22 OA, on Potsdam at 93.60 AF, 87.79 mIoU, and 92.13 OA, and on iSAID at 66.9 mIoU, and does so with a decoder head of 2.4 million parameters and 40.5 GFLOPs.

Load-bearing premise

The global scene representation depends on frequency rankings learned on ImageNet and transferred unchanged to aerial and satellite imagery; if those rankings are not representative of remote sensing scenes, scene coupling could inject misleading context instead of helping.

Editorial extensions

If this is right

  • If the central claim is correct, attention-based segmentation heads for aerial and satellite imagery can be made both more accurate and lighter by replacing dense pixel affinity with scene-coupled, mask-guided affinity; SCSM reports 2.4M parameters and 40.5 GFLOPs against 10.4 to 23.9M parameters and 154.9 to 503 GFLOPs for the compared context modules.
  • The frequency-domain global scene representation would make SCSM the first dual-domain attention model for remote sensing segmentation, so follow-up work can build on selecting scene-specific frequency components rather than spatial features alone.
  • The split local-global semantic masks provide class-level context that avoids background interference, which should particularly benefit small and high-intra-class-variance objects; the paper reports notable gains on cars, helicopters, large vehicles, and small vehicles.
  • SCSM's additional experiment on lithological unit classification suggests the method transfers beyond benchmark land-cover datasets to geoscience mapping with limited compute, at 30.19M parameters and 6.54 GFLOPs on that task.
  • The paper's concise reformulation of vanilla attention offers a new baseline that other attention designs can adopt or modify, potentially extending the same scene-coupling and mask strategies to transformer-based remote sensing models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implication the paper leaves implicit is that the same two-pronged fix would likely transfer to other dense prediction tasks on overhead imagery, such as building footprint extraction or instance segmentation, because the scene-object co-occurrence patterns ROPE+ encodes are not segmentation-specific.
  • The ImageNet frequency prior is the transfer-risk point: re-ranking the DCT components on remote sensing data would be a cheap test of whether the global scene representation truly carries the reported gains or mostly reflects the ImageNet statistics.
  • Because the split local-global masks and ROPE+ add no learned positional parameters, the decoder stays lightweight; this suggests the recipe could serve as a drop-in attention replacement in transformer-based remote sensing models, where dense attention is the main cost.
  • If the claim is right, the biggest gains should appear precisely in classes with strong scene co-occurrence, such as cars on roads and buildings along roads, and the per-class results point that way, but a controlled experiment varying class co-occurrence would test it directly.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 8 minor

Summary. The paper proposes SCSM, a decoder that reconstructs vanilla spatial attention for remote sensing image segmentation. The method augments queries and keys with a 2D rotary position embedding (ROPE+) and a global scene representation extracted via a 2D discrete cosine transform with a frequency prior learned on ImageNet; it also replaces the key/value of attention with local and global semantic masks generated from a pre-classification branch. Experiments on LoveDA, Vaihingen, Potsdam, and iSAID report state-of-the-art results (e.g., 54.6 mIoU on LoveDA, 84.68 mIoU on Vaihingen, 87.79 mIoU on Potsdam, 66.9 mIoU on iSAID), and an additional lithological classification experiment reports mean plus/minus standard deviation. Ablations cover frequency count, block size, rotation angles, architecture structure, and loss coefficients.

Significance. If the reported results are reproducible, SCSM would be a lightweight and efficient attention-based decoder with strong performance on multiple remote sensing benchmarks, and the ablation study would support the value of scene coupling and semantic masking. The paper provides a thorough comparison with many recent methods and includes efficiency metrics (Table 7). However, the central claim of surpassing LOGCAN++ is currently supported by single-run numbers with no error bars, and several technical and reporting issues (Eq. 26, ablation split) must be resolved before the evidence is convincing.

major comments (4)
  1. [4.2.2, Eq. (26)] Equation (26) is dimensionally inconsistent. The global representation G defined in Eq. (25) is a vector formed by concatenating M scalar DCT coefficients along the channel dimension. Substituting this vector into t_{m,n} = (G q_m^T) k_n / sqrt(C) produces a vector-valued similarity (or an outer product) rather than the scalar similarity required by Eq. (18). Please provide the exact intended operation (e.g., a scalar weighting, a learned linear projection, or an element-wise modulation of q/k before the inner product) and correct the formula accordingly.
  2. [6.5] The ablation studies (Tables 4-6 and 8-10) are all described as performed 'on the Loveda dataset', but the dataset split used is not stated. Section 5.1 defines training (2522), validation (1669), and test (1796) splits, with the test set evaluated online. The best configuration found in the ablations (M=16, block size 21, loss coefficient 0.8) yields exactly the mIoU of 54.6 reported for the final test-set comparison in Table 1. If the ablations were run on the test set, the reported number is a grid-selected maximum rather than an unbiased performance estimate, making the comparison against LOGCAN++ unfair. Please state the split used for each ablation table and, if the test split was used, re-evaluate the final model on a held-out split.
  3. [Tables 1-3] The central claim that SCSM 'surpasses' LOGCAN++ relies on small margins, notably 0.22 mIoU on Potsdam (87.79 vs. 87.57 in Table 2). No standard deviations, confidence intervals, or numbers of repeated runs are reported for any benchmark comparison. Since Section 7 reports mean plus/minus standard deviation for the real-world experiment, the authors evidently have the infrastructure to run repeated trials; please provide multi-run statistics (or a clear justification for single-run reporting) for the main benchmark tables.
  4. [4.2.2 and 6.5] The contribution of the DCT-based global scene representation is never isolated in the ablations. Table 8 compares Base+SCA with Base+GA, but SCA includes both ROPE+ and the DCT representation; Table 9 ablates ROPE+ but does not remove the DCT branch or replace it with a spatial-only global pooling baseline. Consequently, the reader cannot tell whether the frequency-domain global representation (and its ImageNet-derived frequency prior) is responsible for the observed gains or whether the improvement comes entirely from ROPE+ and the semantic masks. An ablation with SCA minus the DCT branch (and ideally with the frequency prior replaced by a remote-sensing-native selection) is needed.
minor comments (8)
  1. [Abstract and Section 4.2.2] The paper claims 'the first dual-domain attention model for semantic segmentation of remote sensing images' but does not survey or cite existing dual-domain/frequency-domain attention methods; please either substantiate the claim with a comparison or temper it.
  2. [6.1.1 and 6.2.1] The text identifies PoolFormer and ConvNeXt as the previous state-of-the-art on LoveDA and Vaihingen, respectively, but Table 1 and Table 2 show that LOGCAN++ outperforms both. The narrative should consistently identify LOGCAN++ as the strongest prior method.
  3. [6.1.1 and 6.2.2] Both subsections are titled 'Qualitative analysis', but 6.1.1 is a quantitative discussion; the section titles should be corrected.
  4. [Table 7 and surrounding text] The module is referred to as 'SMG+CCA (Ours)' in Table 7, but the module is named SCA throughout the rest of the paper; use consistent terminology.
  5. [5.2, Eq. (35)] The text 'recision measures' should be 'precision measures'.
  6. [Abstract] The phrase 'demonstrate the the effectiveness' contains a duplicated 'the', and 'inter-clean and elegant mathematical representations' is unclear and should be rephrased.
  7. [4.3, Eq. (27)] The notation for the recover function ψ and the tensor product ⊗ is not defined precisely; please clarify the shapes of the intermediate tensors to make the module reproducible.
  8. [7.1] The overlapping regions (zones T/U overlap by 84.4%, zones 44/45 overlap by 75%) combined with random sampling from all regions A-U could allow spatially overlapping patches to appear in both training and test sets; please describe the spatial decontamination strategy or clarify how the overlap is handled.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity; the architecture is defined by external priors and external benchmarks, but the LoveDA ablations do not disclose their split and the reported best mIoU equals the test mIoU, which is a reporting/selection-risk rather than a demonstrated circular reduction.

full rationale

The paper's claimed derivation is an architectural reconstruction of vanilla attention, not a mathematical derivation of its own benchmark scores. Scene coupling is built from a DCT frequency prior imported from external work FCA-Net [86] (Eqs. 19-26) and a rotary position embedding extended from external RoFormer [84] (Eqs. 9-18); the semantic-mask path replaces the vanilla q/k/v with local and global class masks generated from a pre-classification branch (Eqs. 27-29). None of these equations is defined in terms of the reported mIoU/F1/OA values, so the central claim does not reduce to its inputs by construction. The only self-authored works cited, LOGCAN++ [33] and DocNet [36], serve as comparison baselines and as contrast points for the spatial-prior class centers; they do not supply a load-bearing premise. The principal concern is procedural rather than circular: Section 6.5 states that all ablations (frequency count, block size, rotation angles, structure, ROPE+, loss coefficients) were conducted on 'the LoveDA dataset' without specifying the 1,669-image validation split defined in Section 5.1, and the best ablation mIoU of 54.6 in Tables 4-6, 8-10 equals the reported LoveDA test mIoU of 54.6 in Table 1. If the grid was evaluated on the test split, the headline 54.6 would be a selected maximum rather than an independent estimate; the paper does not say that, so this is a missing-disclosure and error-bar issue (correctness risk), not a demonstrated circularity. Accordingly, the circularity score is low, reflecting only the minor self-citations and the unresolved selection-risk.

Assumptions & free parameters 3 free parameters · 4 assumptions · 2 invented entities

The central claim relies on a handful of design choices (M, P, loss weight) that are tuned on the evaluation benchmark, plus domain assumptions about remote sensing statistics and transfer of ImageNet frequency priors. No fundamentally new physics or geometry is postulated; ROPE+ is a modification of RoPE, and the DCT representation is a standard transform.

free parameters (3)
  • Frequency count M = 16
    Selected by ablations on LoveDA (Table 4) to maximize mIoU; determines the dimensionality of the global scene representation.
  • Block size P = 21
    Selected by ablations on LoveDA (Table 5); set as a multiple of 7 to match the 7x7 ImageNet output used for the frequency prior.
  • Pre-classification loss coefficient = 0.8
    Selected by ablations on LoveDA (Table 10); other coefficients fixed at 1.0 (main) and 0.4 (auxiliary).
assumptions (4)
  • domain assumption Remote sensing images contain complex backgrounds, large intra-class variance, and intrinsic spatial correlations among geospatial objects.
    Stated in Sections 1 and 3.2 and used to motivate both the scene coupling and semantic mask designs; no quantitative characterization is provided.
  • domain assumption The frequency prior learned on ImageNet transfers to remote sensing scenes.
    Used in Section 4.2.2 to select the top M DCT frequencies; relies on the assumption that ImageNet-derived frequency importance generalizes to aerial and satellite imagery.
  • domain assumption Class centers computed from the argmax of the pre-classification representation are reliable spatial priors.
    Section 4.3, Eq. 27-28 uses these centers as keys and values of attention; inherits the assumption common to OCRNet and LOGCAN++ that pre-classification masks are stable enough for class-wise context aggregation.
  • standard math Standard DCT orthogonality and RoPE rotation properties hold as used in the derivations.
    Eq. 9-23 rely on the orthogonal rotation matrices of RoPE and the invertibility of DCT basis functions.
invented entities (2)
  • ROPE+ (image-level 2D rotary position embedding)
    purpose: Model the scene object distribution by encoding absolute positions and relative distances of pixels within the attention affinity.
    Table 9 shows internal ablation gains over sinusoidal and conditional encodings, but no external falsifiable prediction outside the reported benchmarks.
  • Scene global representation from 2D DCT
    purpose: Capture task-specific spectral information of the scene and embed it into the attention similarity computation.
    Built from the standard DCT with frequency components selected on ImageNet; its benefit is demonstrated only on the datasets tested in this paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Novel Scene Coupling Semantic Mask Network for Remote Sensing Image Segmentation." pith.science (2026). https://pith.science/paper/NAUCCFRL

@misc{pith2026250113130,
  author       = {Pith},
  title        = {Pith review of: A Novel Scene Coupling Semantic Mask Network for Remote Sensing Image Segmentation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NAUCCFRL}},
  note         = {Machine review of arXiv:2501.13130}
}
read the original abstract

As a common method in the field of computer vision, spatial attention mechanism has been widely used in semantic segmentation of remote sensing images due to its outstanding long-range dependency modeling capability. However, remote sensing images are usually characterized by complex backgrounds and large intra-class variance that would degrade their analysis performance. While vanilla spatial attention mechanisms are based on dense affine operations, they tend to introduce a large amount of background contextual information and lack of consideration for intrinsic spatial correlation. To deal with such limitations, this paper proposes a novel scene-Coupling semantic mask network, which reconstructs the vanilla attention with scene coupling and local global semantic masks strategies. Specifically, scene coupling module decomposes scene information into global representations and object distributions, which are then embedded in the attention affinity processes. This Strategy effectively utilizes the intrinsic spatial correlation between features so that improve the process of attention modeling. Meanwhile, local global semantic masks module indirectly correlate pixels with the global semantic masks by using the local semantic mask as an intermediate sensory element, which reduces the background contextual interference and mitigates the effect of intra-class variance. By combining the above two strategies, we propose the model SCSM, which not only can efficiently segment various geospatial objects in complex scenarios, but also possesses inter-clean and elegant mathematical representations. Experimental results on four benchmark datasets demonstrate the the effectiveness of the above two strategies for improving the attention modeling of remote sensing images. The dataset and code are available at https://github.com/xwmaxwma/rssegmentation

Figures

Figures reproduced from arXiv: 2501.13130 by the authors.

Figure 1
Figure 1. Intuitive understanding of the scene coupling and semantic masks. For remote sensing images recorded in two different scenerios, i.e., (a) rural images and (b) urban images, we first give two examples to represent (c) the intrinsic spatial correlation of remote sensing image feature targets and (d) the complex backgrounds, large intra-class variance, respectively. For the former, we design Scene-Coupling Attention t… view at source ↗
Figure 2
Figure 2. Our proposed Scene Coupling Attention (SCA) enhances the vanilla attention mechanism by incorporating additional positional and global scene representations. It first applies 2D Rotary Position Embedding (ROPE+) to both query and key, indirectly modeling the relative spatial distribution of objects within the scene. Additionally, it applies a 2D Discrete Cosine Transform (DCT) to the query to obtain a global scene r… view at source ↗
Figure 3
Figure 3. Overall structure diagram of the SCSM model, which consists of backbone, several convolution operations, two semantic mask generation (SMG) modules and a scene coupling attentio (SCA) module. The SMG module associates pixels with the global semantic mask through a local mask spatial prior, achieving class-level modeling and seamless integration with the scene coupling module. The SCA module generates a global scene … view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Structural diagram of the Scene Coupling Attention Module, composed of scene object distribution and global scene representation. The scene object distribution leverages ROPE+ positional encoding to capture complex object distributions in remote sensing scenes. The glo…
Figure 5
Figure 5. Figure 5: ROPE+ working analysis. We set the basic angles of rotation 𝜃 𝑥 𝑖 and 𝜃 𝑦 𝑖 in the width and height directions of different channels, respectively. Then, for a pixel feature with position (𝑚, 𝑛), we rotate it twice consecutively, each time with angles positively correl…
Figure 6
Figure 6. Figure 6: Semantic Mask Generation Module utilizes pre￾classification masks for class-level contextual modeling of features, mitigating noise interference caused by a large number of back￾ground pixels. in the vertical direction. Thus, the inner product of the two in the higher …
Figure 7
Figure 7. Figure 7 [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: Qualitative comparison between SCSM and other state-of-the-art methods on the Vaihingen test set. The red dashed box is the area of focus. Best viewed in color and zoom in [PITH_FULL_IMAGE:figures/full_fig_p013_8.png]
Figure 9
Figure 9. Figure 9: Qualitative comparison between SCSM and other state-of-the-art methods on the Potsdam test set. The red dashed box is the area of focus. Best viewed in color and zoom in. Potsdam, LoveDA, and iSAID datasets, with cropping sizes of 512 × 512 for the first three datasets…
Figure 10
Figure 10. Figure 10: Ablation Study on the Impact of Model Structure Variations with Class Activation Maps. The target activation classes are building (first line) and car (second line), respectively. activation strength and accuracy in the target object, and reduces the erroneous activat…
Figure 11
Figure 11. Figure 11: Comparing the feature maps at different stages, B-CAM represents the features from the last layer of the backbone network, while D-CAM represents the features from the last layer after passing through the decoding head. The experiment is carried out on the vaihingen d…
Figure 12
Figure 12. Figure 12: Comparative visualization of state-of-the-art method applied in Tieshan, Edong District, Hubei Province, China [28], PSPNet [27], Bi-HRNet [113], SwinUNet [42], and DPNet [114]. The experimental results, presented in [PITH_FULL_IMAGE:figures/full_fig_p020_12.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

115 extracted references · 57 canonical work pages

  1. [1]

    Zhang, K

    Q. Zhang, K. C. Seto, Mapping urbanization dynamics at regional and global scales using multi-temporal dmsp/ols nighttime light data, Remote Sensing of Environment 115 (9) (2011) 2320–2329

  2. [2]

    Huang, B

    B. Huang, B. Zhao, Y. Song, Urban land-use mapping using a deep convolutionalneuralnetworkwithhighspatialresolutionmultispec- tral remote sensing imagery, Remote Sensing of Environment 214 (2018) 73–86

  3. [3]

    J.He,T.Nie,W.Ma,Geolocationrepresentationfromlargelanguage models are generic enhancers for spatio-temporal learning, arXiv preprint arXiv:2408.12116 (2024)

  4. [4]

    Q. Yuan, H. Shen, T. Li, Z. Li, S. Li, Y. Jiang, H. Xu, W. Tan, Q. Yang, J. Wang, et al., Deep learning in environmental remote sensing:Achievementsandchallenges,RemoteSensingofEnviron- ment 241 (2020) 111716

  5. [5]

    K. Cui, W. Tang, R. Zhu, M. Wang, G. D. Larsen, V. P. Pauca, S.Alqahtani,F.Yang,D.Segurado,P.Fine,etal.,Real-timelocaliza- tion and bimodal point pattern analysis of palms using uav imagery, arXiv preprint arXiv:2410.11124 (2024)

  6. [6]

    M.Maboudi,J.Amini,S.Malihi,M.Hahn,Integratingfuzzyobject basedimageanalysisandantcolonyoptimizationforroadextraction from remotely sensed images, ISPRS Journal of Photogrammetry and Remote Sensing 138 (2018) 151–163

  7. [7]

    Z. Wang, B. Li, C. Wang, S. Scherer, AirShot: Efficient few-shot detection for autonomous exploration, in: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2024. URL https://arxiv.org/pdf/2404.05069.pdf

  8. [8]

    Wang, ONLS: OPTIMAL NOISE LEVEL SEARCH IN DIF- FUSION AUTOENCODERS WITHOUT FINE-TUNING, in: The Second Tiny Papers Track at ICLR 2024, 2024

    Z. Wang, ONLS: OPTIMAL NOISE LEVEL SEARCH IN DIF- FUSION AUTOENCODERS WITHOUT FINE-TUNING, in: The Second Tiny Papers Track at ICLR 2024, 2024. URL https://openreview.net/forum?id=Q8diCUHTZd

Show all 115 references
  1. [9]

    Bastani, P

    F. Bastani, P. Wolters, R. Gupta, J. Ferdinando, A. Kembhavi, Satlaspretrain:Alarge-scaledatasetforremotesensingimageunder- standing,in:ProceedingsoftheIEEE/CVFInternationalConference on Computer Vision (ICCV), 2023, pp. 16772–16782

  2. [10]

    R. Liu, T. Luo, S. Huang, Y. Wu, Z. Jiang, H. Zhang, Crossmatch: Cross-viewmatchingforsemi-supervisedremotesensingimageseg- mentation, IEEE Transactions on Geoscience and Remote Sensing (2024)

  3. [11]

    T. Luo, M. Du, J. Shi, X. Chen, B. Zhao, S. Huang, Contextuality helpsrepresentationlearningforgeneralizedcategorydiscovery,in: 2024 IEEE International Conference on Image Processing (ICIP), 2024, pp. 687–693.doi:10.1109/ICIP51287.2024.10647298

  4. [12]

    K. Cui, R. Li, S. L. Polk, Y. Lin, H. Zhang, J. M. Murphy, R. J. Plemmons, R. H. Chan, Superpixel-based and spatially-regularized diffusion learning for unsupervised hyperspectral image clustering, IEEE Transactions on Geoscience and Remote Sensing (2024)

  5. [13]

    K.Cui,R.Li,S.L.Polk,J.M.Murphy,R.J.Plemmons,R.H.Chan, Unsupervised spatial-spectral hyperspectral image reconstruction and clustering with diffusion geometry, in: 2022 12th Workshop on HyperspectralImagingandSignalProcessing: EvolutioninRemote Sensing (WHISPERS), IEEE, 2022, pp. 1–5

  6. [14]

    Y.Chen,C.Liu,X.Liu,R.Arcucci,Z.Xiong,Bimcv-r:Alandmark dataset for 3d ct text-image retrieval, in: MICCAI, Springer, 2024, pp. 124–134

  7. [15]

    Y. Chen, W. Huang, X. Liu, S. Deng, Q. Chen, Z. Xiong, Learn- ing multiscale consistency for self-supervised electron microscopy instance segmentation, in: ICASSP, IEEE, 2024, pp. 1566–1570

  8. [16]

    J. Xue, B. Su, Significant remote sensing vegetation indices: A reviewofdevelopmentsandapplications,Journalofsensors2017(1) (2017) 1353691

  9. [17]

    J. Long, E. Shelhamer, T. Darrell, Fully convolutional networks for semantic segmentation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3431–3440

  10. [18]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez,Ł.Kaiser,I.Polosukhin,Attentionisallyouneed,Advances in neural information processing systems 30 (2017)

  11. [19]

    Zhang, P

    Y. Zhang, P. Ji, A. Wang, J. Mei, A. Kortylewski, A. L. Yuille, 3d-aware neural body fitting for occlusion robust 3d human pose estimation, in: IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 9365–9376.doi:10.1109/ICCV51070.2023. 00862. URL https://doi.o...

  12. [20]

    Y. Wang, Y. Zhang, M. Huo, R. Tian, X. Zhang, Y. Xie, C. Xu, P. Ji, W. Zhan, M. Ding, M. Tomizuka, Sparse diffusion policy: A sparse, reusable, and flexible policy for robot learning, CoRR abs/2407.01531 (2024). doi:10.48550/ARXIV.2407.01531. URL https://doi.org/10.48550/arXiv...

  13. [21]

    2260–2271

    T.Nie,G.Qin,W.Ma,Y.Mei,J.Sun,Imputeformer:Lowrankness- induced transformers for generalizable spatiotemporal imputation, in: Proceedings of the 30th ACM SIGKDD Conference on Knowl- edge Discovery and Data Mining, 2024, pp. 2260–2271. X.Ma et al.:Preprint submitted to Elsevier ...

  14. [22]

    Trias-Sanz, G

    R. Trias-Sanz, G. Stamon, J. Louchet, Using colour, texture, and hierarchial segmentation for high-resolution remote sensing, ISPRS Journal of Photogrammetry and remote sensing 63 (2) (2008) 156– 168

  15. [23]

    J. Jin, H. Xu, P. Ji, B. Leng, Imc-net: Learning implicit field with corner attention network for 3d shape reconstruction, in: IEEE International Conference on Image Processing (ICIP), 2022, pp. 1591–1595. doi:10.1109/ICIP46576.2022.9897709. URL https://doi.org/10.1109/ICIP465...

  16. [24]

    B. Bi, S. Liu, L. Mei, Y. Wang, P. Ji, X. Cheng, Decoding by contrasting knowledge: Enhancing llms’ confidence on edited facts, CoRR abs/2405.11613 (2024).doi:10.48550/ARXIV.2405.11613. URL https://doi.org/10.48550/arXiv.2405.11613

  17. [25]

    Y. Chen, W. Huang, S. Zhou, Q. Chen, Z. Xiong, Self-supervised neuron segmentation with multi-agent reinforcement learning, in: IJCAI, 2023, pp. 609–617

  18. [26]

    H. Qian, Y. Chen, S. Lou, F. Khan, X. Jin, D.-P. Fan, Maskfactory: Towards high-quality synthetic data generation for dichotomous image segmentation, in: NeurIPS, 2024

  19. [27]

    H. Zhao, J. Shi, X. Qi, X. Wang, J. Jia, Pyramid scene parsing network,in:ProceedingsoftheIEEEconferenceoncomputervision and pattern recognition, 2017, pp. 2881–2890

  20. [28]

    L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, H. Adam, Encoder- decoder with atrous separable convolution for semantic image seg- mentation,in:ProceedingsoftheEuropeanconferenceoncomputer vision (ECCV), 2018, pp. 801–818

  21. [30]

    Y. Chen, H. Shi, X. Liu, T. Shi, R. Zhang, D. Liu, Z. Xiong, F.Wu,Tokenunify:Scalableautoregressivevisualpre-trainingwith mixture token prediction, arXiv preprint arXiv:2405.16847 (2024)

  22. [31]

    J. He, Z. Deng, Y. Qiao, Dynamic multi-scale filters for semantic segmentation, in: Proceedings of the IEEE/CVF International Con- ference on Computer Vision, 2019, pp. 3562–3572

  23. [32]

    Z. Jin, B. Liu, Q. Chu, N. Yu, Isnet: Integrate image-level and semantic-levelcontextforsemanticsegmentation,in:Proceedingsof theIEEE/CVFInternationalConferenceonComputerVision,2021, pp. 7189–7198

  24. [33]

    X. Ma, R. Lian, Z. Wu, H. Guo, M. Ma, S. Wu, Z. Du, S. Song, W. Zhang, Logcan++: Adaptive local-global class-aware network forsemanticsegmentationofremotesensingimagery(2024). arXiv: 2406.16502. URL https://arxiv.org/abs/2406.16502

  25. [34]

    J. Fu, J. Liu, H. Tian, Y. Li, Y. Bao, Z. Fang, H. Lu, Dual attention network for scene segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 3146–3154

  26. [35]

    J. Hu, L. Shen, G. Sun, Squeeze-and-excitation networks, in: Pro- ceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7132–7141

  27. [36]

    X. Ma, R. Che, X. Wang, M. Ma, S. Wu, T. Feng, W. Zhang, Docnet: Dual-domain optimized class-aware network for remote sensingimagesegmentation,IEEEGeoscienceandRemoteSensing Letters (2024)

  28. [37]

    Huang, X

    Z. Huang, X. Wang, L. Huang, C. Huang, Y. Wei, W. Liu, Ccnet: Criss-cross attention for semantic segmentation, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 603–612

  29. [38]

    Q. Song, J. Li, C. Li, H. Guo, R. Huang, Fully attentional network forsemanticsegmentation,in:ProceedingsoftheAAAIConference on Artificial Intelligence, Vol. 36, 2022, pp. 2280–2288

  30. [39]

    X. Li, H. He, X. Li, D. Li, G. Cheng, J. Shi, L. Weng, Y. Tong, Z.Lin,Pointflow:Flowingsemanticsthroughpointsforaerialimage segmentation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 4217–4226

  31. [40]

    A. Gu, T. Dao, Mamba: Linear-time sequence modeling with selec- tive state spaces, arXiv preprint arXiv:2312.00752 (2023)

  32. [41]

    T. Dao, A. Gu, Transformers are ssms: Generalized models and efficient algorithms through structured state space duality, arXiv preprint arXiv:2405.21060 (2024)

  33. [42]

    X.He,Y.Zhou,J.Zhao,D.Zhang,R.Yao,Y.Xue,Swintransformer embedding unet for remote sensing image semantic segmentation, IEEE Transactions on Geoscience and Remote Sensing 60 (2022) 1–15

  34. [43]

    Zhang, W

    C. Zhang, W. Jiang, Y. Zhang, W. Wang, Q. Zhao, C. Wang, Transformer and cnn hybrid deep neural network for semantic seg- mentation of very-high-resolution remote sensing imagery, IEEE Transactions on Geoscience and Remote Sensing 60 (2022) 1–20

  35. [44]

    L.Ding,D.Lin,S.Lin,J.Zhang,X.Cui,Y.Wang,H.Tang,L.Bruz- zone,Lookingoutsidethewindow:Wide-contexttransformerforthe semantic segmentation of high-resolution remote sensing images, IEEE Transactions on Geoscience and Remote Sensing 60 (2022) 1–13

  36. [45]

    Y. Liu, Y. Zhang, Y. Wang, S. Mei, Rethinking transformers for semanticsegmentationofremotesensingimages,IEEETransactions on Geoscience and Remote Sensing (2023)

  37. [46]

    Sun, Ultra-high resolution segmentation via boundary-enhanced patch-merging transformer (2024).arXiv:2412.10181

    H. Sun, Ultra-high resolution segmentation via boundary-enhanced patch-merging transformer (2024).arXiv:2412.10181. URL https://arxiv.org/abs/2412.10181

  38. [47]

    Cheng, H

    S. Cheng, H. Sun, Spt: Sequence prompt transformer for interactive image segmentation (2024).arXiv:2412.10224. URL https://arxiv.org/abs/2412.10224

  39. [48]

    S.Zhao,H.Chen,X.Zhang,P.Xiao,L.Bai,W.Ouyang,Rs-mamba for large remote sensing image dense prediction, arXiv preprint arXiv:2404.02668 (2024)

  40. [49]

    K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778

  41. [50]

    M. Yang, K. Yu, C. Zhang, Z. Li, K. Yang, Denseaspp for semantic segmentation in street scenes, in: Proceedings of the IEEE confer- ence on computer vision and pattern recognition, 2018, pp. 3684– 3692

  42. [51]

    Huang, D

    Y. Huang, D. Kang, W. Jia, L. Liu, X. He, Channelized axial attention–considering channel relation within spatial attention for semanticsegmentation,in:ProceedingsoftheAAAIConferenceon Artificial Intelligence, Vol. 36, 2022, pp. 1016–1025

  43. [52]

    G. Deng, Z. Wu, C. Wang, M. Xu, Y. Zhong, Ccanet: Class- constraint coarse-to-fine attentional deep network for subdecimeter aerial image semantic segmentation, IEEE Transactions on Geo- science and Remote Sensing 60 (2022) 1–20. doi:10.1109/TGRS. 2021.3055950

  44. [53]

    Liang, W

    C. Liang, W. Wang, J. Miao, Y. Yang, Gmmseg: Gaussian mixture based generative semantic segmentation models, arXiv preprint arXiv:2210.02025 (2022)

  45. [54]

    2582–2593

    T.Zhou,W.Wang,E.Konukoglu,L.VanGool,Rethinkingsemantic segmentation: A prototype view, in: Proceedings of the IEEE/CVF ConferenceonComputerVisionandPatternRecognition,2022,pp. 2582–2593

  46. [55]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al., An image is worth 16x16 words: Transformers for image recognition at scale, arXiv preprint arXiv:2010.11929 (2020)

  47. [56]

    E.Xie,W.Wang,Z.Yu,A.Anandkumar,J.M.Alvarez,P.Luo,Seg- former: Simple and efficient design for semantic segmentation with transformers, Advances in Neural Information Processing Systems 34 (2021) 12077–12090

  48. [57]

    Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, B. Guo, Swin transformer: Hierarchical vision transformer using shifted windows,in:ProceedingsoftheIEEE/CVFinternationalconference on computer vision, 2021, pp. 10012–10022

  49. [58]

    I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreit, et al., Mlp-mixer: An all-mlp architecture for vision, Advances in neural X.Ma et al.:Preprint submitted to Elsevier Page 22 of 24 A Novel Scene Coupl...

  50. [59]

    H. Liu, Z. Dai, D. So, Q. V. Le, Pay attention to mlps, Advances in Neural Information Processing Systems 34 (2021) 9204–9215

  51. [60]

    Cheng, A

    B. Cheng, A. Schwing, A. Kirillov, Per-pixel classification is not all you need for semantic segmentation, Advances in Neural Informa- tion Processing Systems 34 (2021) 17864–17875

  52. [61]

    Z. Wu, J. Lu, J. Han, L. Bai, Y. Zhang, Z. Zhao, S. Song, Domain separation graph neural networks for saliency object ranking, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 3964–3974

  53. [62]

    H. Bao, L. Dong, S. Piao, F. Wei, Beit: Bert pre-training of image transformers, arXiv preprint arXiv:2106.08254 (2021)

  54. [63]

    H. Sun, L. Xu, S. Jin, P. Luo, C. Qian, W. Liu, Program: Prototype graphmodelbasedpseudo-labellearningfortest-timeadaptation,in: TheTwelfthInternationalConferenceonLearningRepresentations, 2024

  55. [64]

    J.Jiao,Y.-M.Tang,K.-Y.Lin,Y.Gao,J.Ma,Y.Wang,W.-S.Zheng, Dilateformer:Multi-scaledilatedtransformerforvisualrecognition, IEEE Transactions on Multimedia (2023)

  56. [65]

    7262–7272

    R.Strudel,R.Garcia,I.Laptev,C.Schmid,Segmenter:Transformer for semantic segmentation, in: Proceedings of the IEEE/CVF inter- national conference on computer vision, 2021, pp. 7262–7272

  57. [66]

    L. Dai, G. Zhang, R. Zhang, Radanet: Road augmented deformable attention network for road extraction from complex high-resolution remote-sensing images, IEEE Transactions on Geoscience and Re- mote Sensing (2023)

  58. [67]

    Luo, J.-X

    L. Luo, J.-X. Wang, S.-B. Chen, J. Tang, B. Luo, Bdtnet: Road extraction by bi-direction transformer from remote sensing images, IEEE Geoscience and Remote Sensing Letters 19 (2022) 1–5

  59. [68]

    S. Dong, Z. Chen, Block multi-dimensional attention for road seg- mentationinremotesensingimagery,IEEEGeoscienceandRemote Sensing Letters 19 (2021) 1–5

  60. [69]

    Jung, H.-S

    H. Jung, H.-S. Choi, M. Kang, Boundary enhancement semantic segmentation for building extraction from remote sensed image, IEEE Transactions on Geoscience and Remote Sensing 60 (2021) 1–12

  61. [70]

    Yuan, Learning building extraction in aerial scenes with convolu- tional networks, IEEE transactions on pattern analysis and machine intelligence 40 (11) (2017) 2793–2798

    J. Yuan, Learning building extraction in aerial scenes with convolu- tional networks, IEEE transactions on pattern analysis and machine intelligence 40 (11) (2017) 2793–2798

  62. [71]

    A. Alem, S. Kumar, Transfer learning models for land cover and land use classification in remote sensing image, Applied Artificial Intelligence 36 (1) (2022) 2014192

  63. [72]

    C.Zhang,I.Sargent,X.Pan,H.Li,A.Gardiner,J.Hare,P.M.Atkin- son, Joint deep learning for land cover and land use classification, Remote sensing of environment 221 (2019) 173–187

  64. [73]

    doi:10.1109/TGRS.2021.3119537

    R.Zuo,G.Zhang,R.Zhang,X.Jia,Adeformableattentionnetwork for high-resolution remote sensing images semantic segmentation, IEEE Transactions on Geoscience and Remote Sensing 60 (2022) 1–14. doi:10.1109/TGRS.2021.3119537

  65. [74]

    R. Guan, W. Tu, Z. Li, H. Yu, D. Hu, Y. Chen, C. Tang, Q. Yuan, X. Liu, Spatial-spectral graph contrastive clustering with hard sam- ple mining for hyperspectral images, IEEE Transactions on Geo- science and Remote Sensing (2024) 1–16

  66. [75]

    R. Guan, Z. Li, W. Tu, J. Wang, Y. Liu, X. Li, C. Tang, R. Feng, Contrastive multiview subspace clustering of hyperspectral images basedongraphconvolutionalnetworks,IEEETransactionsonGeo- science and Remote Sensing 62 (2024) 1–14

  67. [76]

    R.Li,S.Zheng,C.Zhang,C.Duan,J.Su,L.Wang,P.M.Atkinson, Multiattentionnetworkforsemanticsegmentationoffine-resolution remote sensing images, IEEE Transactions on Geoscience and Re- mote Sensing 60 (2021) 1–13

  68. [77]

    Zheng, Y

    Z. Zheng, Y. Zhong, J. Wang, A. Ma, Foreground-aware relation networkforgeospatialobjectsegmentationinhighspatialresolution remote sensing imagery, in: Proceedings of the IEEE/CVF confer- ence on computer vision and pattern recognition, 2020, pp. 4096– 4105

  69. [78]

    1809–1818

    F.Yang,C.Ma,Sparseandcompletelatentorganizationforgeospa- tial semantic segmentation, in: Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 2022, pp. 1809–1818

  70. [79]

    X. Wang, R. Girshick, A. Gupta, K. He, Non-local neural networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7794–7803

  71. [80]

    Rottensteiner, G

    F. Rottensteiner, G. Sohn, J. Jung, M. Gerke, C. Baillard, S. Bnitez, U. Breitkopf, International society for photogrammetry and remote sensing,2dsemanticlabelingcontest,available: https://www.isprs. org/education/benchmarks/UrbanSemLab (Accessed: Oct. 29, 2020.)

  72. [81]

    J.Wang,Z.Zheng,A.Ma,X.Lu,Y.Zhong,Loveda:Aremotesens- ing land-cover dataset for domain adaptive semantic segmentation, arXiv preprint arXiv:2110.08733 (2021)

  73. [82]

    Waqas Zamir, A

    S. Waqas Zamir, A. Arora, A. Gupta, S. Khan, G. Sun, F. Shah- baz Khan, F. Zhu, L. Shao, G.-S. Xia, X. Bai, isaid: A large-scale dataset for instance segmentation in aerial images, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops,...

  74. [83]

    Zheng, Y

    Z. Zheng, Y. Zhong, J. Wang, A. Ma, L. Zhang, Farseg++: Foreground-aware relation network for geospatial object segmen- tation in high spatial resolution remote sensing imagery, IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)

  75. [84]

    J.Su,M.Ahmed,Y.Lu,S.Pan,W.Bo,Y.Liu,Roformer:Enhanced transformer with rotary position embedding, Neurocomputing 568 (2024) 127063

  76. [85]

    P. Shaw, J. Uszkoreit, A. Vaswani, Self-attention with relative posi- tion representations, arXiv preprint arXiv:1803.02155 (2018)

  77. [86]

    Z.Qin,P.Zhang,F.Wu,X.Li,Fcanet:Frequencychannelattention networks,in:ProceedingsoftheIEEE/CVFinternationalconference on computer vision, 2021, pp. 783–792

  78. [87]

    Huang, Z

    Z. Huang, Z. Zhang, C. Lan, Z.-J. Zha, Y. Lu, B. Guo, Adaptive frequency filters as efficient global token mixers, in: Proceedings of theIEEE/CVFInternationalConferenceonComputerVision,2023, pp. 6049–6059

  79. [88]

    Kirillov, R

    A. Kirillov, R. Girshick, K. He, P. Dollár, Panoptic feature pyramid networks,in:ProceedingsoftheIEEE/CVFconferenceoncomputer vision and pattern recognition, 2019, pp. 6399–6408

  80. [89]

    L.Ding,H.Tang,L.Bruzzone,Lanet:Localattentionembeddingto improvethesemanticsegmentationofremotesensingimages,IEEE TransactionsonGeoscienceandRemoteSensing59(1)(2020)426– 435

  81. [90]

    Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, S. Xie, A convnetforthe2020s,in:ProceedingsoftheIEEE/CVFConference on Computer Vision and Pattern Recognition, 2022, pp. 11976– 11986

  82. [91]

    W. Yu, M. Luo, P. Zhou, C. Si, Y. Zhou, X. Wang, J. Feng, S. Yan, Metaformer is actually what you need for vision, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recogni- tion, 2022, pp. 10819–10829

  83. [92]

    L. Zhu, X. Wang, Z. Ke, W. Zhang, R. W. Lau, Biformer: Vision transformer with bi-level routing attention, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2023, pp. 10323–10333

  84. [93]

    17302–17313

    H.Cai,J.Li,M.Hu,C.Gan,S.Han,Efficientvit:Lightweightmulti- scale attention for high-resolution dense prediction, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 17302–17313

  85. [94]

    Y. Ji, Z. Chen, E. Xie, L. Hong, X. Liu, Z. Liu, T. Lu, Z. Li, P. Luo, Ddp:Diffusionmodelfordensevisualprediction,in:Proceedingsof theIEEE/CVFInternationalConferenceonComputerVision,2023, pp. 21741–21752

  86. [95]

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: Alarge-scalehierarchicalimagedatabase,in:2009IEEEconference oncomputervisionandpatternrecognition,Ieee,2009,pp.248–255

  87. [96]

    Zhang, Y

    F. Zhang, Y. Chen, Z. Li, Z. Hong, J. Liu, F. Ma, J. Han, E. Ding, Acfnet:Attentionalclassfeaturenetworkforsemanticsegmentation, X.Ma et al.:Preprint submitted to Elsevier Page 23 of 24 A Novel Scene Coupling Semantic Mask Network for Remote Sensing Image Segmentation in:Proce...

  88. [97]

    Cheng, L.-C

    B. Cheng, L.-C. Chen, Y. Wei, Y. Zhu, Z. Huang, J. Xiong, T. S. Huang, W.-M. Hwu, H. Shi, Spgnet: Semantic prediction guidance for scene parsing, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 5218–5228

  89. [98]

    L.-C. Chen, G. Papandreou, F. Schroff, H. Adam, Rethinking atrousconvolution forsemantic imagesegmentation,arXiv preprint arXiv:1706.05587 (2017)

  90. [99]

    Ronneberger, P

    O. Ronneberger, P. Fischer, T. Brox, U-net: Convolutional networks for biomedical image segmentation, in: Medical image computing and computer-assisted intervention–MICCAI 2015: 18th interna- tional conference, Munich, Germany, October 5-9, 2015, proceed- ings, part III 18, Sp...

  91. [100]

    M.Yin,Z.Yao,Y.Cao,X.Li,Z.Zhang,S.Lin,H.Hu,Disentangled non-local neural networks, in: Computer Vision–ECCV 2020: 16th EuropeanConference,Glasgow,UK,August23–28,2020,Proceed- ings, Part XV 16, Springer, 2020, pp. 191–207

  92. [101]

    Y.Cao,J.Xu,S.Lin,F.Wei,H.Hu,Globalcontextnetworks,IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (6) (2020) 6881–6895

  93. [102]

    Y. Yuan, X. Chen, J. Wang, Object-contextual representations for semantic segmentation, in: Computer Vision–ECCV 2020: 16th EuropeanConference,Glasgow,UK,August23–28,2020,Proceed- ings, Part VI 16, Springer, 2020, pp. 173–190

  94. [103]

    X. Li, Z. Zhong, J. Wu, Y. Yang, Z. Lin, H. Liu, Expectation- maximization attention networks for semantic segmentation, in: ProceedingsoftheIEEE/CVFinternationalconferenceoncomputer vision, 2019, pp. 9167–9176

  95. [104]

    J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y. Zhao, D. Liu, Y.Mu,M.Tan,X.Wang,etal.,Deephigh-resolutionrepresentation learningforvisualrecognition,IEEEtransactionsonpatternanalysis and machine intelligence 43 (10) (2020) 3349–3364

  96. [105]

    T.Xiao,Y.Liu,B.Zhou,Y.Jiang,J.Sun,Unifiedperceptualparsing forsceneunderstanding,in:ProceedingsoftheEuropeanconference on computer vision (ECCV), 2018, pp. 418–434

  97. [106]

    X.Li,A.You,Z.Zhu,H.Zhao,M.Yang,K.Yang,S.Tan,Y.Tong, Semantic flow for fast and accurate scene parsing, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I 16, Springer, 2020, pp. 775–793

  98. [107]

    Carion, F

    N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, S. Zagoruyko, End-to-end object detection with transformers, in: European conference on computer vision, Springer, 2020, pp. 213– 229

  99. [108]

    URL https://openreview.net/forum?id=3KWnuT-R1bh

    X.Chu,Z.Tian,B.Zhang,X.Wang,C.Shen,Conditionalpositional encodings for vision transformers, in: ICLR 2023, 2023. URL https://openreview.net/forum?id=3KWnuT-R1bh

  100. [109]

    K. Han, Y. Wang, Q. Tian, J. Guo, C. Xu, C. Xu, Ghostnet: More features from cheap operations, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 1580–1589

  101. [110]

    Y. Tang, K. Han, J. Guo, C. Xu, C. Xu, Y. Wang, Ghostnetv2: Enhance cheap operation with long-range attention, Advances in Neural Information Processing Systems 35 (2022) 9969–9982

  102. [111]

    Z.Zhou,M.M.RahmanSiddiquee,N.Tajbakhsh,J.Liang,Unet++: A nested u-net architecture for medical image segmentation, in: DeepLearninginMedicalImageAnalysisandMultimodalLearning forClinicalDecisionSupport:4thInternationalWorkshop,DLMIA 2018, and 8th International Workshop, ML-CDS...

  103. [112]

    Oktay, J

    O. Oktay, J. Schlemper, L. L. Folgoc, M. Lee, M. Heinrich, K. Mi- sawa, K. Mori, S. McDonagh, N. Y. Hammerla, B. Kainz, et al., Attention u-net: Learning where to look for the pancreas, arXiv preprint arXiv:1804.03999 (2018)

  104. [113]

    Z. Wu, J. Zhang, L. Zhang, X. Liu, H. Qiao, Bi-hrnet: A road extractionframeworkfromsatelliteimagerybasedonnodeheatmap and bidirectional connectivity, Remote Sensing 14 (7) (2022) 1732

  105. [114]

    G.Zhou,W.Chen,X.Qin,J.Li,L.Wang,Lithologicalunitclassifi- cation based on geological knowledge-guided deep learning frame- workforopticalstereomappingsatelliteimagery,IEEETransactions on Geoscience and Remote Sensing (2023)

  106. [115]

    Lei, D.-F

    X.-F. Lei, D.-F. Duan, S.-Y. Jiang, S.-F. Xiong, Ore-forming fluids and isotopic (hocs-pb) characteristics of the fujiashan-longjiaoshan skarn w-cu-(mo) deposit in the edong district of hubei province, china, Ore Geology Reviews 102 (2018) 386–405

  107. [116]

    X.Ma et al.:Preprint submitted to Elsevier Page 24 of 24

    G.Xie,J.Mao,Q.Zhu,L.Yao,Y.Li,W.Li,H.Zhao,Geochemical constraints on cu–fe and fe skarn deposits in the edong district, middle–lower yangtze river metallogenic belt, china, Ore Geology Reviews 64 (2015) 425–444. X.Ma et al.:Preprint submitted to Elsevier Page 24 of 24

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.