Pith. sign in

REVIEW 3 major objections 5 minor 133 references

Spatial Frequency Modulation for Semantic Segmentation

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Spatial Frequency Modulation improves semantic segmentation by shifting high-frequency features below the Nyquist rate before downsampling and recovering them during upsampling.

desk verdict Broad, plausible engineering gains and a lightweight module, but the paper's causal aliasing story is undone by its own ablation and the novelty claim overlooks the authors' ICLR 2024 paper. read the letter →

arxiv 2507.11893 v2 pith:E7JUULMO submitted 2025-07-16 cs.CV cs.AI

classification cs.CVcs.AI
keywords spatialfrequencymodulationaliasingsemanticsegmentationadaptiveresamplingnon-uniformupsamplinglearninganti-aliasingdenseprediction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that a large part of the accuracy lost in modern segmentation networks is caused by aliasing: when feature maps are downsampled, energy above the Nyquist frequency folds into lower frequencies and distorts the features. To avoid this without throwing away fine details, the authors propose Spatial Frequency Modulation (SFM), which non-uniformly resamples the feature map before downsampling so high-frequency regions are stretched and shifted to lower frequencies, and then reverses the operation during upsampling. If correct, SFM is a lightweight add-on that improves a wide range of CNN and transformer segmentation models, including Mask2Former and InternImage, and also helps in classification, adversarial robustness, and instance/panoptic segmentation.

What carries the argument

The load-bearing machinery is Adaptive Resampling (ARS) for modulation and Multi-Scale Adaptive Upsampling (MSAU) for demodulation. ARS generates an attention map via difference-aware convolution and pyramid spatial pooling, maps uniform coordinates to non-uniform ones so that high-attention (high-frequency) areas are sampled more densely, and thereby lowers the frequency of those signals in accordance with the Frequency Scaling Property; two extra losses, the frequency modulation loss and the semantic high-frequency loss, supervise this resampling. MSAU reverses the coordinate deformation with Delaunay triangulation and barycentric interpolation, then refines the prediction with cascaded Local Pixel Relation Modules whose dilation grows to capture multi-scale relations between densely and sparsely sampled regions.

What would settle it

A direct test would be to compare segmentation accuracy on features that have been deliberately low-pass filtered (removing high frequencies) versus features modulated to lower frequencies without discarding them; if the low-pass version does not degrade as much as predicted, the causal role of aliasing in the SFM gains is weaker.

Watch

Extended reading notes

Core claim

The central claim is that the "aliasing degradation" phenomenon—lower segmentation accuracy as the aliasing ratio of feature maps increases—can be countered by a modulate-demodulate cycle. The paper introduces the aliasing ratio as the proportion of spectral power above the Nyquist frequency, shows empirically that it correlates inversely with accuracy across ResNet, Swin, and ConvNeXt backbones, and then demonstrates that replacing uniform downsampling with adaptive resampling (ARS) and uniform upsampling with multi-scale adaptive upsampling (MSAU) consistently raises mIoU, with gains such as +1.5 mIoU for Mask2Former-Swin-T and +1.4 for InternImage-T on ADE20K.

Load-bearing premise

The load-bearing premise is that aliasing—the folding of high-frequency energy above the Nyquist rate during downsampling—is a primary cause of the observed accuracy loss, and that the correlation between aliasing ratio and accuracy is causal.

Editorial extensions

If this is right

  • Semantic segmentation models can gain accuracy without architectural redesign by adding SFM at downsampling and upsampling stages.
  • The gains concentrate on boundaries: boundary F-score, boundary IoU, and boundary accuracy all improve, and boundary-related errors such as false responses, merging mistakes, and displacements drop.
  • SFM combines with existing anti-aliasing and refinement methods such as FLC, CondConv, and SegFix, producing further improvements.
  • Because the modulation is a resampling operation rather than a filter, high-frequency details are preserved instead of discarded, which distinguishes SFM from low-pass pooling approaches.
  • The same modulation step transfers to image classification, adversarial robustness, and instance/panoptic segmentation, suggesting the mechanism is a general property of downsampling in vision networks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the causal story is right, the same modulate-demodulate cycle should help any dense-prediction task with repeated downsampling, such as depth estimation or optical flow, not just segmentation.
  • The paper's own ablation leaves open that a large share of the gain comes from the upsampling refinement: MSAU alone improves +1.3 mIoU while ARS alone degrades -0.5, so the causal role of the frequency modulation itself is not separately established.
  • A testable extension is to apply ARS to only the first downsampling layer, where aliasing is strongest, and measure whether the benefit scales with the number of protected stages.
  • The aliasing-ratio metric could be repurposed as a diagnostic tool for selecting which layers of a pretrained network most need anti-aliasing treatment.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes Spatial Frequency Modulation (SFM), a lightweight module pair for semantic segmentation. Adaptive Resampling (ARS) is inserted before downsampling layers to non-uniformly resample high-frequency regions, allegedly shifting their spectral content below the Nyquist frequency; Multi-Scale Adaptive Upsampling (MSAU) replaces standard bilinear upsampling to demodulate the modulated features and refine predictions. The authors report consistent mIoU gains across multiple backbones and segmentation heads, including +1.5 mIoU on Mask2Former-Swin-T and +1.4 mIoU on UPerNet-InternImage-T on ADE20K, and +3.3 mIoU on Cityscapes with a ResNet-50 PSPNet baseline. They also extend the method to image classification, adversarial robustness, instance segmentation, and panoptic segmentation. The paper includes frequency-domain analysis, an ablation study, and comparisons with low-pass pooling, deformable convolution, and learned upsampling methods.

Significance. If the reported gains are reproducible, SFM would be a broadly applicable, inexpensive plug-in for dense prediction, and the paper's open-source release is a concrete asset. The experimental breadth is a genuine strength: the method is evaluated across FCN/PSPNet/CCNet/OCNet/PCAA heads, ResNet/Swin/ConvNeXt/InternImage backbones, and several tasks. However, the central causal claim that frequency modulation before downsampling is the driver of the gains is not established by the paper's own ablations. The component that performs the modulation (ARS) is harmful in isolation, the component that refines during upsampling (MSAU) provides most of the gain, and the frequency analysis largely reports the effect of an explicit loss term rather than an independent phenomenon. The paper would be substantially stronger if the mechanism were isolated with appropriate controls and if the relationship to the authors' prior ICLR 2024 work [50] were clarified.

major comments (3)
  1. [§4.7, Table 14] The ablation does not support the causal attribution of the gains to the frequency-modulation mechanism. Table 14 shows ARS alone reduces mIoU by 0.5 (72.2 vs. 72.7), MSAU alone improves mIoU by 1.3 (74.0), and the combination improves by 3.3 (76.0). Since MSAU alone already delivers most of the improvement and ARS alone is harmful, the +3.3 gain cannot be attributed to the modulation/demodulation cycle without a control that keeps MSAU while neutralizing ARS (e.g., uniform resampling at the same cost, or an ARS with identity coordinates). The paper's own interpretation in §4.7 that 'ARS and MSAU cannot function independently' is exactly the claim that needs to be tested, not assumed. Please also report the aliasing ratio for each ablation row, so the reader can see whether mIoU tracks the aliasing ratio within the ablation rather than only in the full-pipeline comparison of Table 13.
  2. [§3.2, Eq. (12), and §3.4, Figures 10/11] The frequency analyses in Figures 10 and 11 are not independent evidence for the 'modulate, don't discard' claim, because the training objective L_FM in Eq. (12) directly penalizes the exact quantity being measured: spectral power above the Nyquist frequency. Figure 10(a) therefore confirms that the optimizer minimized a loss term, not that ARS implements the frequency-scaling operation described by Eq. (3). In fact, Section 3.2 states that the precise adaptive sampling coordinates for Eq. (3) are difficult to compute, and the implemented system is a learned coordinate map supervised by L_FM and L_SHF. To establish that high-frequency information is preserved rather than discarded, the paper should either demonstrate that the suppressed high-frequency components are recoverable after demodulation (e.g., by correlating the recovered spectrum with the original high-frequency content), or provide a control with L_FM removed. Table 18 shows that removing L_FM costs only 0.3 mIoU (74.7 vs. 74.4), which further suggests the aliasing-ratio reduction is not the main source of the improvement.
  3. [§3.1, Figure 1 and Figure 5] The motivating observation is correlational: the authors plot segmentation accuracy versus the aliasing ratio in existing models and observe an inverse relationship. This is presented as evidence that aliasing 'leads to' degradation, but no interventional experiment is performed on the baseline models to establish causation. The issue is compounded by the fact that SFM explicitly trains toward a lower aliasing ratio, so the correlation in Table 13 (42.7% to 34.2% AR with +3.3 mIoU) could reflect a common cause or a side effect of the upsampling refinement rather than a causal pathway. A concrete test would be to train SFM with the FM loss disabled (as in Table 18, row 2) and report the resulting aliasing ratio; if mIoU improves without a corresponding AR reduction, the central claim would be falsified.
minor comments (5)
  1. [§1, paragraph 2] There is a typo in 'post-upssampled' near the end of Section 1; it should be 'post-upsampled'.
  2. [§4.3, after Table 3] The text contains an inserted editorial marker 'Proofread:' immediately after 'Comparison with state-of-the-art on ADE20K.' This appears to be a leftover annotation and should be removed before submission.
  3. [Table 7] The header of Table 7 uses 'SMF' in the 'Ours' rows; this should be 'SFM' for consistency with the method name used throughout the paper.
  4. [§4.5 heading] The heading 'Discusion' is misspelled; it should be 'Discussion'.
  5. [§2, related works and reference [50]] Reference [50] is the authors' own ICLR 2024 paper 'When Semantic Segmentation Meets Frequency Aliasing.' The submission should state explicitly what is new relative to that paper, particularly whether ARS or MSAU already appeared there, so that the novelty of the present TPAMI submission is unambiguous.

Circularity Check

1 steps flagged · score 6.0 of 10

The frequency-reduction evidence for SFM's mechanism is enforced by the L_FM loss, so the feature-spectrum plots confirm the training objective rather than independently validating the aliasing story; the mIoU gains remain empirical.

  1. fitted input called prediction [Sec. 3.2, Eq. (12); Sec. 3.4, 'Modulated feature frequency analysis']
    "To simplify the intricate calculation of adaptive sampling coordinates, we choose to directly supervise the adaptively resampled features. The objective is to reduce high frequencies above the Nyquist frequency, ultimately resulting in a lower aliasing ratio. ... L_FM = (1/|H|) Σ_{(k,l)∈H} |F(k,l)|^2 ... After modulation, compared with the original feature map, we observe that the power of frequencies larger than Nyquist frequency (ξ > 1/4 for 2×downsampling) are largely reduced."

    The claimed confirmation of the modulation mechanism is not independent: Eq. (12) directly penalizes the power of exactly the frequencies above Nyquist (the set H) in the modulated feature, with λ_FM in the total loss. The ARS module is trained to minimize this quantity, so the observed reduction in high-frequency power in Figures 10(a) and 11(b), and the lower aliasing ratio in Table 13, are by construction consequences of the optimization objective rather than emergent evidence that frequency modulation causes the accuracy gains. The segmentation improvements on Cityscapes, ADE20K, etc.

full rationale

The benchmark results (e.g., Mask2Former-Swin-T +1.5 mIoU, InternImage-T +1.4 mIoU on ADE20K) are self-contained empirical comparisons against external baselines and are not circular. The paper does not rely on a load-bearing self-citation chain: references [50] and [87]-[90] are prior works by the same authors, but they are used contextually, not to force the central conclusion. The main circular step is the mechanistic validation: L_FM in Eq. (12) directly minimizes the power of frequencies above Nyquist in the modulated feature, and Section 3.4 then cites the resulting reduction as evidence that ARS 'effectively modulates high frequencies to lower frequencies.' This is a fitted input being presented as an observed prediction. The causal story is further weakened by Table 14, where ARS alone decreases mIoU by 0.5 while MSAU alone improves it by 1.3 and the full SFM improves it by 3.3; thus the component performing the modulation is harmful in isolation and most of the gain comes from the upsampling/refinement module. That is a correctness and interpretation concern rather than circularity, but it reinforces that the central aliasing explanation is only partially supported. Overall, the paper contains one clear reduction-by-construction step and otherwise has independent empirical content, so a score of 6 is appropriate.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The ledger shows that the method's contribution rests on several validation-set choices (loss weights, LPRM depth, kernel sigma) and on domain assumptions about applying sampling theorems to learned features. The free parameters are tunable and affect reported results; the axioms are the load-bearing interpretation that high-frequency power above Nyquist is the causal source of degradation.

free parameters (4)
  • lambda_FM = 0.01 (implementation), 0.1 (ablation optimum)
    Weight of frequency modulation loss, tuned on validation set; implementation and ablation disagree on the value.
  • lambda_SHF = 100
    Weight of semantic high-frequency loss, tuned on validation set.
  • Gaussian sigma for coordinate mapping kernel = not reported
    The distance kernel G in Eq. (9) uses a standard deviation sigma, but the default value is not specified.
  • Number and dilations of LPRMs = 7 LPRMs, dilations 1,2,4,8,16,32,64
    Chosen by validation ablation (Table 20); contributes to the MSAU gain.
assumptions (4)
  • domain assumption Nyquist-Shannon sampling theorem applies to CNN feature maps and strided downsampling behaves as ideal uniform sampling.
    Used in Section 3.1 to define aliasing ratio as high-frequency power above 1/4 and to argue it causes degradation; learned convolution filters are ignored.
  • domain assumption Frequency Scaling Property holds for adaptive non-uniform resampling.
    Section 3.2 invokes dense sampling with rate A reducing frequency to 1/A, although the implemented adaptive coordinates are not uniform scaling and the exact A is never computed.
  • domain assumption Non-uniform upsampling via Delaunay triangulation and barycentric interpolation can invert the ARS warp and recover high frequencies.
    Section 3.3 assumes the demodulation reverses modulation; no exact inverse guarantee is provided.
  • domain assumption Reducing aliasing ratio improves segmentation accuracy causally.
    The motivation in Section 3.1 is a correlation; ablations show ARS alone lowers accuracy, so the causal link is not established.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Spatial Frequency Modulation for Semantic Segmentation." pith.science (2026). https://pith.science/paper/E7JUULMO

@misc{pith2026250711893,
  author       = {Pith},
  title        = {Pith review of: Spatial Frequency Modulation for Semantic Segmentation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E7JUULMO}},
  note         = {Machine review of arXiv:2507.11893}
}
read the original abstract

High spatial frequency information, including fine details like textures, significantly contributes to the accuracy of semantic segmentation. However, according to the Nyquist-Shannon Sampling Theorem, high-frequency components are vulnerable to aliasing or distortion when propagating through downsampling layers such as strided-convolution. Here, we propose a novel Spatial Frequency Modulation (SFM) that modulates high-frequency features to a lower frequency before downsampling and then demodulates them back during upsampling. Specifically, we implement modulation through adaptive resampling (ARS) and design a lightweight add-on that can densely sample the high-frequency areas to scale up the signal, thereby lowering its frequency in accordance with the Frequency Scaling Property. We also propose Multi-Scale Adaptive Upsampling (MSAU) to demodulate the modulated feature and recover high-frequency information through non-uniform upsampling This module further improves segmentation by explicitly exploiting information interaction between densely and sparsely resampled areas at multiple scales. Both modules can seamlessly integrate with various architectures, extending from convolutional neural networks to transformers. Feature visualization and analysis confirm that our method effectively alleviates aliasing while successfully retaining details after demodulation. Finally, we validate the broad applicability and effectiveness of SFM by extending it to image classification, adversarial robustness, instance segmentation, and panoptic segmentation tasks. The code is available at https://github.com/Linwei-Chen/SFM.

Figures

Figures reproduced from arXiv: 2507.11893 by the authors.

Figure 1
Figure 1. Quantitative analysis of the relationship between segmentation [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. (a) Illustration of Spatial Frequency Modulation. (b) The mod [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Illustration of a general FCN-based architecture for semantic [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Illustration of 1-D signal aliasing in the frequency domain. Left: [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 7
Figure 7. Figure 7: Feature visualization. (a) shows the input image and ground [PITH_FULL_IMAGE:figures/full_fig_p004_7.png]
Figure 11
Figure 11. Figure 11: Feature visualization. (a) displays the original feature map [PITH_FULL_IMAGE:figures/full_fig_p007_11.png]
Figure 9
Figure 9. Figure 9: Illustration of non-uniform upsampling. The blue points represent [PITH_FULL_IMAGE:figures/full_fig_p007_9.png]
Figure 10
Figure 10. Figure 10: Feature frequency analysis. (a) demonstrates a reduction in [PITH_FULL_IMAGE:figures/full_fig_p007_10.png]
Figure 13
Figure 13. Figure 13: Visualization on the Cityscapes validation set. The white [PITH_FULL_IMAGE:figures/full_fig_p010_13.png]
Figure 14
Figure 14. Figure 14: Visualized results on Pascal Context [100] validation set. w/o SFM w/ SFM Ground truth [PITH_FULL_IMAGE:figures/full_fig_p011_14.png]
Figure 15
Figure 15. Figure 15: Visualized results on ADE20K [101] dataset. TABLE 8 Results on the Pascal Context [100] test set. The backbone is ResNet-50 with a downsampling stride of 32×. Method FCN [4] PSPNet [31] CCNet [59] OCNet [68] Vanilla 43.8 48.0 47.7 47.9 Ours 46.0 (+2.2) 49.7 (+1.7) 49.…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

133 extracted references · 73 canonical work pages

  1. [50]

    When semantic segmentation meets frequency aliasing,

    L. Chen, L. Gu, and Y . Fu, “When semantic segmentation meets frequency aliasing,” inProceedings of International Conference on Learning Representations, 2024, pp. 1–13

  2. [1]

    Planning-oriented autonomous driving,

    Y . Hu, J. Yang, L. Chen, K. Li, C. Sima, X. Zhu, S. Chai, S. Du, T. Lin, W. Wanget al., “Planning-oriented autonomous driving,” in Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 17 853–17 862

  3. [2]

    Bonnet: An open-source training and deployment framework for semantic segmentation in robotics using cnns,

    A. Milioto and C. Stachniss, “Bonnet: An open-source training and deployment framework for semantic segmentation in robotics using cnns,” inIEEE International Conference on Robotics and Automation. IEEE, 2019, pp. 7094–7100

  4. [3]

    Land-cover classification with high-resolution remote sensing images using transferable deep models,

    X.-Y . Tong, G.-S. Xia, Q. Lu, H. Shen, S. Li, S. You, and L. Zhang, “Land-cover classification with high-resolution remote sensing images using transferable deep models,”Remote Sensing of Environment, vol. 237, p. 111322, 2020

  5. [4]

    Fully convolutional networks for semantic segmentation

    E. Shelhamer, J. Long, and T. Darrell, “Fully convolutional networks for semantic segmentation.”IEEE Transactions Pattern Analysis and Machine Intelligence, vol. 39, no. 4, pp. 640–651, 2016

  6. [5]

    Encoder- decoder with atrous separable convolution for semantic image segmen- tation,

    L.-C. Chen, Y . Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder- decoder with atrous separable convolution for semantic image segmen- tation,” inProceedings of European Conference on Computer Vision, 2018, pp. 801–818

  7. [6]

    Segformer: Simple and efficient design for semantic segmentation with transformers,

    E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “Segformer: Simple and efficient design for semantic segmentation with transformers,” inProceedings of Advances in Neural Information Processing Systems, vol. 34, 2021, pp. 12 077–12 090

  8. [7]

    Pattern-affinitive propagation across depth, surface normal and semantic segmentation,

    Z. Zhang, Z. Cui, C. Xu, Y . Yan, N. Sebe, and J. Yang, “Pattern-affinitive propagation across depth, surface normal and semantic segmentation,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 4106–4115

Show all 133 references
  1. [8]

    Learning structure-aware semantic segmentation with image-level supervision,

    J. Liu, J. Zhang, Y . Hong, and N. Barnes, “Learning structure-aware semantic segmentation with image-level supervision,” inInternational Joint Conference on Neural Networks, 2021, pp. 1–8

  2. [9]

    Learning statistical texture for semantic segmentation,

    L. Zhu, D. Ji, S. Zhu, W. Gan, W. Wu, and J. Yan, “Learning statistical texture for semantic segmentation,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp. 12 537–12 546

  3. [10]

    Boundary-aware feature propagation for scene segmentation,

    H. Ding, X. Jiang, A. Q. Liu, N. M. Thalmann, and G. Wang, “Boundary-aware feature propagation for scene segmentation,” inPro- ceedings of IEEE International Conference on Computer Vision, 2019, pp. 6819–6829

  4. [11]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778

  5. [12]

    Aggregated residual transformations for deep neural networks,

    S. Xie, R. Girshick, P. Doll ´ar, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 1492–1500

  6. [13]

    Swin transformer: Hierarchical vision transformer using shifted win- dows,

    Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted win- dows,” inProceedings of IEEE International Conference on Computer Vision, 2021, pp. 10 012–10 022

  7. [15]

    Very deep convolutional networks for large-scale image recognition,

    K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” inProceedings of International Conference on Learning Representations, 2015, pp. 1–14

  8. [16]

    Internimage: Exploring large-scale vision foundation mod- els with deformable convolutions,

    W. Wang, J. Dai, Z. Chen, Z. Huang, Z. Li, X. Zhu, X. Hu, T. Lu, L. Lu, H. Liet al., “Internimage: Exploring large-scale vision foundation mod- els with deformable convolutions,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 14 408–14 419

  9. [17]

    Communication in the presence of noise,

    C. E. Shannon, “Communication in the presence of noise,”Proceedings of the IRE, vol. 37, no. 1, pp. 10–21, 1949

  10. [18]

    Certain topics in telegraph transmission theory,

    H. Nyquist, “Certain topics in telegraph transmission theory,”Transac- tions of the American Institute of Electrical Engineers, vol. 47, no. 2, pp. 617–644, 1928

  11. [19]

    Frequencylowcut pooling-plug and play against catastrophic overfitting,

    J. Grabinski, S. Jung, J. Keuper, and M. Keuper, “Frequencylowcut pooling-plug and play against catastrophic overfitting,” inProceedings of European Conference on Computer Vision, 2022, pp. 36–57

  12. [20]

    Delving deeper into anti-aliasing in convnets,

    X. Zou, F. Xiao, Z. Yu, and Y . J. Lee, “Delving deeper into anti-aliasing in convnets,” inProceedings of the British Machine Vision Conference, 2020, pp. 1–13

  13. [21]

    Adversarial machine learn- ing at scale,

    A. Kurakin, I. Goodfellow, and S. Bengio, “Adversarial machine learn- ing at scale,”arXiv preprint arXiv:1611.01236, pp. 1–17, 2016

  14. [22]

    Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks,

    F. Croce and M. Hein, “Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks,” inProceedings of International Conference on Machine Learning, 2020, pp. 2206–2216. IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 16

  15. [23]

    A convnet for the 2020s,

    Z. Liu, H. Mao, C.-Y . Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 11 976–11 986

  16. [24]

    On the spectral bias of neural networks,

    N. Rahaman, A. Baratin, D. Arpit, F. Draxler, M. Lin, F. Hamprecht, Y . Bengio, and A. Courville, “On the spectral bias of neural networks,” inProceedings of International Conference on Machine Learning, 2019, pp. 5301–5310

  17. [25]

    Deep frequency principle towards understanding why deeper learning is faster,

    Z. J. Xu and H. Zhou, “Deep frequency principle towards understanding why deeper learning is faster,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 12, 2021, pp. 10 541–10 550

  18. [26]

    R. A. Roberts and C. T. Mullis,Digital signal processing. Addison- Wesley Longman Publishing Co., Inc., 1987

  19. [27]

    Making convolutional networks shift-invariant again,

    R. Zhang, “Making convolutional networks shift-invariant again,” in Proceedings of International Conference on Machine Learning, 2019, pp. 7324–7334

  20. [28]

    Anti- aliasing deep image classifiers using novel depth adaptive blurring and activation function,

    M. T. Hossain, S. W. Teng, G. Lu, M. A. Rahman, and F. Sohel, “Anti- aliasing deep image classifiers using novel depth adaptive blurring and activation function,”Neurocomputing, vol. 536, pp. 164–174, 2023

  21. [29]

    Learning to downsample for segmentation of ultra-high resolution images,

    C. Jin, R. Tanno, T. Mertzanidou, E. Panagiotaki, and D. C. Alexander, “Learning to downsample for segmentation of ultra-high resolution images,” inProceedings of International Conference on Learning Rep- resentations, 2022, pp. 1–17

  22. [30]

    Learning to zoom and unzoom,

    C. Thavamani, M. Li, F. Ferroni, and D. Ramanan, “Learning to zoom and unzoom,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 5086–5095

  23. [31]

    Pyramid scene parsing net- work,

    H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing net- work,” inProceedings of IEEE International Conference on Computer Vision, 2017, pp. 2881–2890

  24. [32]

    Partial class activation attention for semantic segmentation,

    S.-A. Liu, H. Xie, H. Xu, Y . Zhang, and Q. Tian, “Partial class activation attention for semantic segmentation,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 16 836–16 845

  25. [33]

    Masked-attention mask transformer for universal image segmentation,

    B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Masked-attention mask transformer for universal image segmentation,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 1290–1299

  26. [34]

    Covariance attention for semantic segmentation,

    Y . Liu, Y . Chen, P. Lasang, and Q. Sun, “Covariance attention for semantic segmentation,”IEEE Transactions Pattern Analysis and Ma- chine Intelligence, vol. 44, no. 4, pp. 1805–1818, 2020

  27. [35]

    Ctnet: Context-based tandem net- work for semantic segmentation,

    Z. Li, Y . Sun, L. Zhang, and J. Tang, “Ctnet: Context-based tandem net- work for semantic segmentation,”IEEE Transactions Pattern Analysis and Machine Intelligence, vol. 44, no. 12, pp. 9904–9917, 2021

  28. [36]

    Refinenet: Multi-path refinement networks for dense prediction,

    G. Lin, F. Liu, A. Milan, C. Shen, and I. Reid, “Refinenet: Multi-path refinement networks for dense prediction,”IEEE Transactions Pattern Analysis and Machine Intelligence, vol. 42, no. 5, pp. 1228–1242, 2020

  29. [37]

    Context disentangling and prototype inheriting for robust visual grounding,

    W. Tang, L. Li, X. Liu, L. Jin, J. Tang, and Z. Li, “Context disentangling and prototype inheriting for robust visual grounding,”IEEE Transac- tions on Pattern Analysis and Machine Intelligence, vol. 46, no. 5, pp. 3213–3229, 2023

  30. [38]

    Deep collaborative embedding for social im- age understanding,

    Z. Li, J. Tang, and T. Mei, “Deep collaborative embedding for social im- age understanding,”IEEE transactions on pattern analysis and machine intelligence, vol. 41, no. 9, pp. 2070–2083, 2018

  31. [39]

    Deep guided attention network for joint denoising and demosaicing in real image,

    T. Zhang, Y . Fu, J. Zhang, and C. Yan, “Deep guided attention network for joint denoising and demosaicing in real image,”Chinese Journal of Electronics, vol. 33, no. 1, pp. 303–312, 2024

  32. [40]

    Transformer-based under-sampled single- pixel imaging,

    Y . Tian, Y . Fu, and J. Zhang, “Transformer-based under-sampled single- pixel imaging,”Chinese Journal of Electronics, vol. 32, no. 5, pp. 1151– 1159, 2023

  33. [41]

    Reflectance and fluores- cence spectral recovery via actively lit rgb images,

    Y . Fu, A. Lam, I. Sato, T. Okabe, and Y . Sato, “Reflectance and fluores- cence spectral recovery via actively lit rgb images,”IEEE transactions on pattern analysis and machine intelligence, vol. 38, no. 7, pp. 1313– 1326, 2015

  34. [42]

    Physics-based noise modeling for extreme low-light photography,

    K. Wei, Y . Fu, Y . Zheng, and J. Yang, “Physics-based noise modeling for extreme low-light photography,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 11, pp. 8520–8537, 2022

  35. [43]

    Coded hyperspectral image reconstruction using deep external and internal learning,

    Y . Fu, T. Zhang, L. Wang, and H. Huang, “Coded hyperspectral image reconstruction using deep external and internal learning,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 7, pp. 3404–3420, 2022

  36. [44]

    Guided hyperspectral image denoising with realistic data,

    T. Zhang, Y . Fu, and J. Zhang, “Guided hyperspectral image denoising with realistic data,”International Journal of Computer Vision, vol. 130, no. 11, p. 2885–2901, 2022

  37. [45]

    Le-gan: Unsupervised low- light image enhancement network using attention module and identity invariant loss,

    Y . Fu, Y . Hong, L. Chen, and S. You, “Le-gan: Unsupervised low- light image enhancement network using attention module and identity invariant loss,”Knowledge-Based Systems, vol. 240, p. 108010, 2022

  38. [46]

    Level-aware consistent multilevel map translation from satellite imagery,

    Y . Fu, Z. Fang, L. Chen, T. Song, and D. Lin, “Level-aware consistent multilevel map translation from satellite imagery,”IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–14, 2022

  39. [47]

    Consistency-aware map generation at multiple zoom levels using aerial image,

    L. Chen, Z. Fang, and Y . Fu, “Consistency-aware map generation at multiple zoom levels using aerial image,”IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 15, pp. 5953–5966, 2022

  40. [48]

    A large-scale climate- aware satellite image dataset for domain adaptive land-cover semantic segmentation,

    S. Liu, L. Chen, L. Zhang, J. Hu, and Y . Fu, “A large-scale climate- aware satellite image dataset for domain adaptive land-cover semantic segmentation,”ISPRS Journal of Photogrammetry and Remote Sensing, vol. 205, pp. 98–114, 2023

  41. [49]

    Transformer based pluralistic image completion with reduced information loss,

    Q. Liu, Y . Jiang, Z. Tan, D. Chen, Y . Fu, Q. Chu, G. Hua, and N. Yu, “Transformer based pluralistic image completion with reduced information loss,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

  42. [51]

    Hybrid supervised instance segmentation by learning label noise suppression,

    L. Chen, Y . Fu, S. You, and H. Liu, “Hybrid supervised instance segmentation by learning label noise suppression,”Neurocomputing, vol. 496, pp. 131–146, 2022

  43. [52]

    Efficient hybrid supervision for instance segmentation in aerial images,

    ——, “Efficient hybrid supervision for instance segmentation in aerial images,”Remote Sensing, vol. 13, no. 2, p. 252, 2021

  44. [53]

    Per-pixel classification is not all you need for semantic segmentation,

    B. Cheng, A. Schwing, and A. Kirillov, “Per-pixel classification is not all you need for semantic segmentation,” inProceedings of Advances in Neural Information Processing Systems

  45. [54]

    Deep high-resolution representation learning for visual recognition,

    J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y . Zhao, D. Liu, Y . Mu, M. Tan, X. Wanget al., “Deep high-resolution representation learning for visual recognition,”IEEE Transactions Pattern Analysis and Machine Intelligence, vol. 43, no. 10, pp. 3349–3364, 2020

  46. [55]

    Hrformer: High-resolution transformer for dense prediction,

    Y . Yuan, R. Fu, L. Huang, W. Lin, C. Zhang, X. Chen, and J. Wang, “Hrformer: High-resolution transformer for dense prediction,” inPro- ceedings of Advances in Neural Information Processing Systems, 2021, pp. 1–13

  47. [56]

    Dilated residual networks,

    F. Yu, V . Koltun, and T. Funkhouser, “Dilated residual networks,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 472–480

  48. [57]

    Rethinking atrous convolution for semantic image segmentation,

    L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam, “Rethinking atrous convolution for semantic image segmentation,”arXiv preprint arXiv:1706.05587, 2017

  49. [58]

    Non-local neural net- works,

    X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural net- works,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7794–7803

  50. [59]

    Ccnet: Criss-cross attention for semantic segmentation,

    Z. Huang, X. Wang, Y . Wei, L. Huang, H. Shi, W. Liu, and T. S. Huang, “Ccnet: Criss-cross attention for semantic segmentation,”IEEE Transactions Pattern Analysis and Machine Intelligence, 2020

  51. [60]

    Dual attention network for scene segmentation,

    J. Fu, J. Liu, H. Tian, Y . Li, Y . Bao, Z. Fang, and H. Lu, “Dual attention network for scene segmentation,” inProceedings of IEEE International Conference on Computer Vision, 2019, pp. 3146–3154

  52. [61]

    Feature pyramid networks for object detection,

    T.-Y . Lin, P. Doll ´ar, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 2117–2125

  53. [62]

    Panoptic feature pyramid networks,

    A. Kirillov, R. Girshick, K. He, and P. Doll ´ar, “Panoptic feature pyramid networks,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 6399–6408

  54. [63]

    U-net: deep learning for cell counting, detection, and morphometry,

    T. Falk, D. Mai, R. Bensch, ¨O. C ¸ ic ¸ek, A. Abdulkadir, Y . Marrakchi, A. B¨ohm, J. Deubner, Z. J¨ackel, K. Seiwaldet al., “U-net: deep learning for cell counting, detection, and morphometry,”Nature methods, vol. 16, no. 1, pp. 67–70, 2019

  55. [64]

    Pointrend: Image segmen- tation as rendering,

    A. Kirillov, Y . Wu, K. He, and R. Girshick, “Pointrend: Image segmen- tation as rendering,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 9799–9808

  56. [65]

    Semantic flow for fast and accurate scene parsing,

    X. Li, A. You, Z. Zhu, H. Zhao, M. Yang, K. Yang, S. Tan, and Y . Tong, “Semantic flow for fast and accurate scene parsing,” inProceedings of European Conference on Computer Vision. Springer, 2020, pp. 775– 793

  57. [66]

    Alignseg: Feature-aligned segmentation networks,

    Z. Huang, Y . Wei, X. Wang, W. Liu, T. S. Huang, and H. Shi, “Alignseg: Feature-aligned segmentation networks,”IEEE Transactions Pattern Analysis and Machine Intelligence, vol. 44, no. 1, pp. 550–557, 2021

  58. [67]

    Fapn: Feature-aligned pyramid network for dense image prediction,

    S. Huang, Z. Lu, R. Cheng, and C. He, “Fapn: Feature-aligned pyramid network for dense image prediction,” inProceedings of IEEE Interna- tional Conference on Computer Vision, 2021, pp. 864–873

  59. [68]

    Ocnet: Object context for semantic segmentation,

    Y . Yuan, L. Huang, J. Guo, C. Zhang, X. Chen, and J. Wang, “Ocnet: Object context for semantic segmentation,”International Journal of Computer Vision, vol. 129, no. 8, pp. 2375–2398, 2021. IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 17

  60. [69]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” inProceedings of Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 1–11

  61. [70]

    Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,

    S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y . Wang, Y . Fu, J. Feng, T. Xiang, P. H. Torret al., “Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp....

  62. [71]

    Segmenter: Trans- former for semantic segmentation,

    R. Strudel, R. Garcia, I. Laptev, and C. Schmid, “Segmenter: Trans- former for semantic segmentation,” inProceedings of IEEE Interna- tional Conference on Computer Vision, 2021, pp. 7262–7272

  63. [72]

    Wavecnet: Wavelet integrated cnns to suppress aliasing effect for noise-robust image classification,

    Q. Li, L. Shen, S. Guo, and Z. Lai, “Wavecnet: Wavelet integrated cnns to suppress aliasing effect for noise-robust image classification,”IEEE Transaction on Image Process., vol. 30, pp. 7074–7089, 2021

  64. [73]

    Alias-free generative adversarial networks,

    T. Karras, M. Aittala, S. Laine, E. H ¨ark¨onen, J. Hellsten, J. Lehtinen, and T. Aila, “Alias-free generative adversarial networks,”NeurIPS, vol. 34, pp. 852–863, 2021

  65. [74]

    Watch your up-convolution: Cnn based generative deep neural networks are failing to reproduce spectral distributions,

    R. Durall, M. Keuper, and J. Keuper, “Watch your up-convolution: Cnn based generative deep neural networks are failing to reproduce spectral distributions,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 7890–7899

  66. [75]

    Spectral distribution aware image generation,

    S. Jung and M. Keuper, “Spectral distribution aware image generation,” inAssociation for the Advancement of Artificial Intelligence, vol. 35, no. 2, 2021, pp. 1734–1742

  67. [76]

    Spatially adaptive inference with stochastic feature sampling and interpolation,

    Z. Xie, Z. Zhang, X. Zhu, G. Huang, and S. Lin, “Spatially adaptive inference with stochastic feature sampling and interpolation,” inPro- ceedings of European Conference on Computer Vision. Springer, 2020, pp. 531–548

  68. [77]

    Dynamic convolutions: Exploiting spatial sparsity for faster inference,

    T. Verelst and T. Tuytelaars, “Dynamic convolutions: Exploiting spatial sparsity for faster inference,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 2320–2329

  69. [78]

    Spatial transformer networks,

    M. Jaderberg, K. Simonyan, A. Zissermanet al., “Spatial transformer networks,” inProceedings of Advances in Neural Information Process- ing Systems, 2015, pp. 1–9

  70. [79]

    Learning to zoom: a saliency-based sampling layer for neural net- works,

    A. Recasens, P. Kellnhofer, S. Stent, W. Matusik, and A. Torralba, “Learning to zoom: a saliency-based sampling layer for neural net- works,” inProceedings of European Conference on Computer Vision, 2018, pp. 51–66

  71. [80]

    Efficient segmentation: Learning downsampling near semantic bound- aries,

    D. Marin, Z. He, P. Vajda, P. Chatterjee, S. Tsai, F. Yang, and Y . Boykov, “Efficient segmentation: Learning downsampling near semantic bound- aries,” inProceedings of IEEE International Conference on Computer Vision, 2019, pp. 2131–2141

  72. [81]

    Looking for the devil in the details: Learning trilinear attention sampling network for fine-grained image recognition,

    H. Zheng, J. Fu, Z.-J. Zha, and J. Luo, “Looking for the devil in the details: Learning trilinear attention sampling network for fine-grained image recognition,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 5012–5021

  73. [82]

    Fovea: Foveated image magnification for autonomous navigation,

    C. Thavamani, M. Li, N. Cebron, and D. Ramanan, “Fovea: Foveated image magnification for autonomous navigation,” inProceedings of IEEE International Conference on Computer Vision, 2021, pp. 15 539– 15 548

  74. [83]

    Deformable convolutional networks,

    J. Dai, H. Qi, Y . Xiong, Y . Li, G. Zhang, H. Hu, and Y . Wei, “Deformable convolutional networks,” inProceedings of IEEE Inter- national Conference on Computer Vision, 2017, pp. 764–773

  75. [84]

    Deformable convnets v2: More deformable, better results,

    X. Zhu, H. Hu, S. Lin, and J. Dai, “Deformable convnets v2: More deformable, better results,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 9308–9316

  76. [85]

    Ssbnet: Improving visual recognition effi- ciency by adaptive sampling,

    H. M. Kwan and S. Song, “Ssbnet: Improving visual recognition effi- ciency by adaptive sampling,” inProceedings of European Conference on Computer Vision. Springer, 2022, pp. 229–244

  77. [86]

    Autofocusformer: Image segmentation off the grid,

    C. Ziwen, K. Patnaik, S. Zhai, A. Wan, Z. Ren, A. G. Schwing, A. Colburn, and L. Fuxin, “Autofocusformer: Image segmentation off the grid,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 18 227–18 236

  78. [87]

    Frequency- aware feature fusion for dense image prediction,

    L. Chen, Y . Fu, L. Gu, C. Yan, T. Harada, and G. Huang, “Frequency- aware feature fusion for dense image prediction,”IEEE Transactions Pattern Analysis and Machine Intelligence, vol. 46, no. 12, pp. 10 763– 10 780, 2024

  79. [88]

    Frequency dynamic con- volution for dense image prediction,

    L. Chen, L. Gu, L. Li, C. Yan, and Y . Fu, “Frequency dynamic con- volution for dense image prediction,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 30 178–30 188

  80. [89]

    Frequency-adaptive dilated convolution for semantic segmentation,

    L. Chen, L. Gu, D. Zheng, and Y . Fu, “Frequency-adaptive dilated convolution for semantic segmentation,” inProceedings of IEEE Con- ference on Computer Vision and Pattern Recognition, 2024, pp. 3414– 3425

  81. [90]

    Frequency-dynamic attention modulation for dense prediction,

    L. Chen, L. Gu, and Y . Fu, “Frequency-dynamic attention modulation for dense prediction,” inProceedings of IEEE International Conference on Computer Vision, 2025, pp. 1–13

  82. [91]

    Fcanet: Frequency channel attention networks,

    Z. Qin, P. Zhang, F. Wu, and X. Li, “Fcanet: Frequency channel attention networks,” inProceedings of IEEE International Conference on Computer Vision, 2021, pp. 783–792

  83. [92]

    Learning in the frequency domain,

    K. Xu, M. Qin, F. Sun, Y . Wang, Y .-K. Chen, and F. Ren, “Learning in the frequency domain,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 1740–1749

  84. [93]

    Adaptive frequency filters as efficient global token mixers,

    Z. Huang, Z. Zhang, C. Lan, Z.-J. Zha, Y . Lu, and B. Guo, “Adaptive frequency filters as efficient global token mixers,” inProceedings of IEEE International Conference on Computer Vision, 2023, pp. 1–11

  85. [94]

    The cityscapes dataset for semantic urban scene understanding,

    M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Be- nenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 3213–3223

  86. [95]

    D. W. Kammler,A first course in Fourier analysis. Cambridge University Press, 2007

  87. [96]

    Mask r-cnn,

    K. He, G. Gkioxari, P. Doll ´ar, and R. Girshick, “Mask r-cnn,” in Proceedings of IEEE International Conference on Computer Vision, 2017, pp. 2961–2969

  88. [97]

    Sur la sphere vide,

    B. Delaunayet al., “Sur la sphere vide,”Izv. Akad. Nauk SSSR, Otdelenie Matematicheskii i Estestvennyka Nauk, vol. 7, no. 793-800, pp. 1–2, 1934

  89. [98]

    Shirley, M

    P. Shirley, M. Ashikhmin, and S. Marschner,Fundamentals of computer graphics. AK Peters/CRC Press, 2009

  90. [99]

    The quickhull algo- rithm for convex hulls,

    C. B. Barber, D. P. Dobkin, and H. Huhdanpaa, “The quickhull algo- rithm for convex hulls,”ACM Transactions on Mathematical Software, vol. 22, no. 4, pp. 469–483, 1996

  91. [100]

    The role of context for object detection and semantic segmentation in the wild,

    R. Mottaghi, X. Chen, X. Liu, N.-G. Cho, S.-W. Lee, S. Fidler, R. Urtasun, and A. Yuille, “The role of context for object detection and semantic segmentation in the wild,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 891– 898

  92. [101]

    Scene parsing through ade20k dataset,

    B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Scene parsing through ade20k dataset,” inProceedings of IEEE Con- ference on Computer Vision and Pattern Recognition, 2017, pp. 633– 641

  93. [102]

    Context prior for scene segmentation,

    C. Yu, J. Wang, C. Gao, G. Yu, C. Shen, and N. Sang, “Context prior for scene segmentation,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 12 416–12 425

  94. [103]

    A stochastic approximation method,

    H. Robbins and S. Monro, “A stochastic approximation method,”The annals of mathematical statistics, pp. 400–407, 1951

  95. [104]

    Object-contextual representations for semantic segmentation,

    Y . Yuan, X. Chen, and J. Wang, “Object-contextual representations for semantic segmentation,” inProceedings of European Conference on Computer Vision. Springer, 2020, pp. 173–190

  96. [105]

    Mask dino: Towards a unified transformer-based framework for object detection and segmentation,

    F. Li, H. Zhang, H. Xu, S. Liu, L. Zhang, L. M. Ni, and H.-Y . Shum, “Mask dino: Towards a unified transformer-based framework for object detection and segmentation,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 3041–3050

  97. [106]

    Vision transformer adapter for dense predictions,

    Z. Chen, Y . Duan, W. Wang, J. He, T. Lu, J. Dai, and Y . Qiao, “Vision transformer adapter for dense predictions,” inProceedings of International Conference on Learning Representations, 2023, pp. 1–14

  98. [107]

    Unified perceptual pars- ing for scene understanding,

    T. Xiao, Y . Liu, B. Zhou, Y . Jiang, and J. Sun, “Unified perceptual pars- ing for scene understanding,” inProceedings of European Conference on Computer Vision, 2018, pp. 418–434

  99. [108]

    Learn- ing implicit feature alignment function for semantic segmentation,

    H. Hu, Y . Chen, J. Xu, S. Borse, H. Cai, F. Porikli, and X. Wang, “Learn- ing implicit feature alignment function for semantic segmentation,” in Proceedings of European Conference on Computer Vision. Springer, 2022, pp. 487–505

  100. [109]

    Inceptionnext: When inception meets convnext,

    W. Yu, P. Zhou, S. Yan, and X. Wang, “Inceptionnext: When inception meets convnext,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2024, pp. 5672–5683

  101. [110]

    Pelk: Parameter- efficient large kernel convnets with peripheral convolution,

    H. Chen, X. Chu, Y . Ren, X. Zhao, and K. Huang, “Pelk: Parameter- efficient large kernel convnets with peripheral convolution,” inProceed- ings of IEEE Conference on Computer Vision and Pattern Recognition, 2024, pp. 5557–5567

  102. [111]

    Moganet: Multi-order gated aggregation network,

    S. Li, Z. Wang, Z. Liu, C. Tan, H. Lin, D. Wu, Z. Chen, J. Zheng, and S. Z. Li, “Moganet: Multi-order gated aggregation network,” in Proceedings of International Conference on Learning Representations, 2024, pp. 1–19

  103. [112]

    Metaformer baselines for vision,

    W. Yu, C. Si, P. Zhou, M. Luo, Y . Zhou, J. Feng, S. Yan, and X. Wang, “Metaformer baselines for vision,”IEEE Transactions Pattern Analysis and Machine Intelligence, vol. 46, no. 2, pp. 896–912, 2024. IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 18

  104. [113]

    Overlock: An overview-first-look-closely-next convnet with context-mixing dynamic kernels,

    M. Lou and Y . Yu, “Overlock: An overview-first-look-closely-next convnet with context-mixing dynamic kernels,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2025, pp. 1–8

  105. [114]

    Spatial-mamba: Effective visual state space models via structure-aware state fusion,

    C. Xiao, M. Li, Z. Zhang, D. Meng, and L. Zhang, “Spatial-mamba: Effective visual state space models via structure-aware state fusion,” in ICLR, 2025, pp. 1–13

  106. [115]

    Learning multiple layers of features from tiny images,

    A. Krizhevsky, G. Hintonet al., “Learning multiple layers of features from tiny images,” pp. 1–60, 2009

  107. [116]

    A downsampled variant of imagenet as an alternative to the cifar datasets,

    P. Chrabaszcz, I. Loshchilov, and F. Hutter, “A downsampled variant of imagenet as an alternative to the cifar datasets,”arXiv preprint arXiv:1707.08819, 2017

  108. [117]

    Identity mappings in deep residual networks,

    K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” inProceedings of European Conference on Computer Vision. Springer, 2016, pp. 630–645

  109. [118]

    Microsoft coco: Common objects in context,

    T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll ´ar, and C. L. Zitnick, “Microsoft coco: Common objects in context,” inProceedings of European Conference on Computer Vision, 2014, pp. 740–755

  110. [119]

    Imagenet large scale visual recognition challenge,

    O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernsteinet al., “Imagenet large scale visual recognition challenge,”International Journal of Computer Vision, vol. 115, no. 3, pp. 211–252, 2015

  111. [120]

    High-frequency component helps explain the generalization of convolutional neural networks,

    H. Wang, X. Wu, Z. Huang, and E. P. Xing, “High-frequency component helps explain the generalization of convolutional neural networks,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 8684–8694

  112. [121]

    Panoptic segmentation,

    A. Kirillov, K. He, R. Girshick, C. Rother, and P. Doll ´ar, “Panoptic segmentation,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 9404–9413

  113. [122]

    Gated-scnn: Gated shape cnns for semantic segmentation,

    T. Takikawa, D. Acuna, V . Jampani, and S. Fidler, “Gated-scnn: Gated shape cnns for semantic segmentation,” inProceedings of IEEE Inter- national Conference on Computer Vision, 2019, pp. 5229–5238

  114. [123]

    Bound- ary iou: Improving object-centric image segmentation evaluation,

    B. Cheng, R. Girshick, P. Doll ´ar, A. C. Berg, and A. Kirillov, “Bound- ary iou: Improving object-centric image segmentation evaluation,” in Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp. 15 334–15 342

  115. [124]

    Bag of tricks for image classification with convolutional neural networks,

    T. He, Z. Zhang, H. Zhang, Z. Zhang, J. Xie, and M. Li, “Bag of tricks for image classification with convolutional neural networks,” inProceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 558–567

  116. [125]

    Segvit: Semantic segmentation with plain vision transformers,

    B. Zhang, Z. Tian, Q. Tang, X. Chu, X. Wei, C. Shenet al., “Segvit: Semantic segmentation with plain vision transformers,”Proceedings of Advances in Neural Information Processing Systems, vol. 35, pp. 4971– 4982, 2022

  117. [126]

    Dinov2: Learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. V . V o, M. Szafraniec, V . Khali- dov, P. Fernandez, D. HAZIZA, F. Massa, A. El-Noubyet al., “Dinov2: Learning robust visual features without supervision,”Transactions on Machine Learning Research, 2023

  118. [127]

    Condconv: Conditionally parameterized convolutions for efficient inference,

    B. Yang, G. Bender, Q. V . Le, and J. Ngiam, “Condconv: Conditionally parameterized convolutions for efficient inference,”Advances in neural information processing systems, vol. 32, 2019

  119. [128]

    Learning to upsample by learning to sample,

    W. Liu, H. Lu, H. Fu, and Z. Cao, “Learning to upsample by learning to sample,” inProceedings of IEEE International Conference on Computer Vision, 2023, pp. 6027–6037

  120. [129]

    Efficient inference in fully connected crfs with gaussian edge potentials,

    P. Kr ¨ahenb¨uhl and V . Koltun, “Efficient inference in fully connected crfs with gaussian edge potentials,”NeurIPS, vol. 24, pp. 1–14, 2011

  121. [130]

    Guided upsampling network for real-time semantic seg- mentation,

    D. Mazzini, “Guided upsampling network for real-time semantic seg- mentation,” inProceedings of the British Machine Vision Conference, 2018, pp. 1–12

  122. [131]

    Decoders matter for semantic segmentation: Data-dependent decoding enables flexible feature aggre- gation,

    Z. Tian, T. He, C. Shen, and Y . Yan, “Decoders matter for semantic segmentation: Data-dependent decoding enables flexible feature aggre- gation,” inProceedings of IEEE International Conference on Computer Vision, 2019, pp. 3126–3135

  123. [132]

    Instance segmentation with point supervision,

    I. H. Laradji, N. Rostamzadeh, P. O. Pinheiro, D. Vazquez, and M. Schmidt, “Instance segmentation with point supervision,”arXiv preprint arXiv:1906.06392, 2019

  124. [133]

    Segfix: Model-agnostic boundary refinement for segmentation,

    Y . Yuan, J. Xie, X. Chen, and J. Wang, “Segfix: Model-agnostic boundary refinement for segmentation,” inProceedings of European Conference on Computer Vision. Springer, 2020, pp. 489–506

  125. [134]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gellyet al., “An image is worth 16x16 words: Transformers for image recognition at scale,” inProceedings of International Conference on Learning Represent...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.