REVIEW 3 major objections 6 minor 91 references
RecConv: Efficient Recursive Convolutions for Multi-Frequency Representations
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Recursive decomposition of a convolution yields an 80x80 effective receptive field with linear parameter growth and FLOPs bounded by a factor of 5/3.
desk verdict Solid mobile adaptation of WTConv with a correct but narrowly-scoped complexity analysis; the 'constant FLOPs' framing is contradicted by the paper's own throughput numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a recursive decomposition with shared weights: one depthwise convolution of kernel $k$ with stride 2 maps each level to the next, and each level applies its own depthwise convolution of the same kernel, so the parameter count is the shared downsampler plus $\ell+1$ convolutions rather than one growing kernel. The cost identity is a geometric series, written in the paper as $$1 + 2\sum_{n=1}^{\ell}\frac{1}{4^n} < \sum_{n=0}^{\infty}2\left(\frac{1}{4}\right)^{n} - 1 = \frac{5}{3},$$ because every halving of resolution cuts that level's FLOPs by four; this is what lets the effective receptive field grow as $k \times 2^\ell$ while the arithmetic stays bounded. Bilinear upsampling and element-wise addition are treated as parameter-free, and the design also permits simultaneous recursion along the channel dimension.
What would settle it
A settling experiment would compare a RecConv block and a full large-kernel depthwise convolution of the same theoretical receptive field on the same hardware at matched FLOPs, measuring end-to-end latency and throughput; if the recursive version is consistently slower, the constant-FLOPs framing does not describe actual cost. A second check is to compute the gradient-based effective receptive field at each decomposition level and see whether the footprint actually reaches $k \times 2^\ell$.
Extended reading notes
Core claim
At its core, the paper establishes a trade-off identity: for a base kernel $k$ and $\ell$ decomposition levels, the effective receptive field is $k \times 2^\ell$ while the parameter factor over a single depthwise convolution is $\ell+2$ and the FLOPs factor is at most $5/3$, against $4^\ell$ for standard or depthwise convolution of the full kernel. This is achieved by recursively halving the spatial resolution with a shared strided depthwise convolution, running a small-kernel depthwise convolution at each level, and returning each level to full resolution by bilinear upsampling and addition. The authors call the result a multi-frequency representation and report that backbones using it, RecNeXt, surpass baseline lightweight backbones in accuracy without structural reparameterization or neural architecture search; for instance, the M3 model outperforms a leading lightweight backbone by 1.9 $AP^{box}$ on COCO at similar FLOPs.
Load-bearing premise
The load-bearing premise is that FLOPs are the right measure of efficiency: the paper counts only convolution arithmetic and treats bilinear upsampling and multi-scale data movement as free, while its own latency and throughput tables show RecNeXt running several times slower than the baselines it is compared with.
Editorial extensions
If this is right
- Lightweight backbones can reach effective receptive fields of $80 \times 80$ with a $5 \times 5$ base kernel and four decomposition levels, at parameter counts that grow by only $\ell+2$ and FLOPs bounded by $5/3$.
- Accuracy follows on standard benchmarks: RecNeXt reports higher ImageNet top-1 accuracy, higher COCO box and mask AP, and higher ADE20K mIoU than comparable lightweight baselines at similar FLOPs.
- The recursion is a drop-in token mixer: it can replace the spatial operator in MetaNeXt-style blocks, and the paper shows variants with linear attention, nearest-neighbor upsampling, transposed convolutions, group convolutions, and channel concatenation.
- Because the effective receptive field is determined by decomposition level, kernel size can stay small, preserving optimized small-kernel implementations.
- No structural reparameterization or neural architecture search is required, which simplifies training and deployment.
Reading between the lines
- Editorial extension: the paper's own tables show the "constant FLOPs" framing does not carry over to wall-clock speed: adding recursion or bilinear upsampling drops throughput from thousands to hundreds of images per second, a limitation the paper acknowledges in its limitations paragraph.
- Editorial extension: the decomposition is effectively a coarse-to-fine recurrence over scales, so future work could unify RecConv with recurrent or state-space mixers and test whether the recursion itself, rather than the specific small kernels, carries the benefit.
- Editorial extension: swapping bilinear upsampling for nearest interpolation or transposed convolution trades a small accuracy loss for large throughput gains, suggesting hardware-specific deployment rules rather than a single universal module.
- Editorial extension: a direct gradient-based measurement of the effective receptive field across increasing decomposition levels would test whether the theoretical $k \times 2^\ell$ footprint is actually realized in trained models.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript introduces RecConv, a recursive multi-scale decomposition of a depthwise convolution into small-kernel convolutions on progressively downsampled feature maps, using a shared stride-2 downsampling convolution, level-wise depthwise convolutions, and bilinear upsampling aggregation. The central claim is that for a base kernel of size k and ell levels of decomposition, RecConv achieves an effective receptive field of k*2^ell with parameter growth of only (ell+2) times the base kernel and convolution FLOPs bounded by 5/3 times the base depthwise convolution, in contrast to the 4^ell growth of standard and depthwise large-kernel convolutions. The authors instantiate this module in RecNeXt, a RepViT/RepNeXt-style mobile backbone, and report experiments on ImageNet-1K classification, MS-COCO detection and instance segmentation, ADE20K semantic segmentation, shape-bias analysis, and ablations, together with an implementation sketch and a code link.
Significance. The theoretical complexity relation is simple, transparent, and arithmetically correct under the stated depthwise-MAC-only accounting, and the architectural idea is a clean adaptation of WTConv to mobile settings. If the implementation matches the corrected description, the parameter-efficiency property is a useful design principle for large-receptive-field convnets. The paper also provides training details, two algorithm listings, a code link, and broad empirical validation across multiple tasks and model scales, which strengthens reproducibility. The main caveat is that the headline efficiency claim is based on a narrow FLOP definition and is not supported by the reported wall-clock throughput; this needs to be qualified before the central efficiency framing can be accepted.
major comments (3)
- [Section 3.2, Algorithm 1 and Table 2] The pseudocode line `self.conv = [Conv(**kwargs)] * (level+1)` is a Python aliasing error: it creates a list of references to one single Conv module. If taken literally, all decomposition levels share one set of convolution weights, so the parameter count would be 2 times the base kernel rather than the (ell+2) times claimed in Table 2; if this is only a typo, the code should create independent modules via a ModuleList comprehension. This discrepancy is load-bearing because the linear parameter-growth claim depends on having ell+1 independent level convolutions plus the shared downsampling convolution. Please correct the listing and verify that the released code matches the corrected version.
- [Abstract, Section 3.2, Tables 1 and 5] The statement that RecConv maintains 'constant FLOPs' regardless of ERF expansion is only valid for the depthwise convolution MACs counted in Table 2; bilinear upsampling and the multi-scale data movement are excluded from that accounting. The paper's own measurements show a much larger practical cost: in Table 5, adding recursive decomposition with [4,3,2,1] levels and 5x5 kernels raises MACs from 0.82 G to 0.87 G but drops GPU throughput from 4538 to 384 im/s, and Table 1 reports RecNeXt-M3 at 314 im/s versus 3604 im/s for RepViT-M1.1. The Limitations section acknowledges this, but the abstract's unqualified efficiency claim and the 'constant FLOPs' phrase should be revised to, for example, 'bounded convolution FLOPs under a depthwise-MAC-only accounting,' with practical efficiency conditioned on hardware and resampling costs.
- [Section 4.1 versus Table 3] There is an internal inconsistency in the reported accuracy of RecNeXt-M5: the text says 'RecNeXt-M5 plateaus at 81.6% top-1 accuracy,' while Table 3 reports 82.9% for RecNeXt-M5 and Table 1 reports 83.3% with distillation. In addition, the claim that RecNeXt-M3 'exceeds other leading models' is contradicted by Table 3, where FastViT-SA12 and RepViT-M1.5 have higher top-1 accuracy. Please correct the numbers and qualify the comparison to models of the same scale.
minor comments (6)
- [Equation (4)] Equation (4) mixes a finite sum with an infinite-series bound in a way that is hard to read; please write the finite-sum expression first and then state the limit as ell grows.
- [Table 2] The 'Standard' row packs the channel factor into the 4^ell term; presenting parameters and FLOPs as k^2*C^2*4^ell and k^2*C^2*H*W*4^ell would remove ambiguity about the comparison to RecConv's depthwise operations.
- [Figure 4] The diagram in Figure 4 appears to contain a stray external image URL from a Substack CDN in place of a vector graphic; this placeholder should be replaced.
- [Algorithm 1] Even after fixing the aliasing, the list assigned to `self.conv` should be wrapped in `nn.ModuleList` (or the modules should otherwise be registered) so that the parameters are discoverable by PyTorch.
- [Terminology throughout] The text alternates between 'constant FLOPs,' 'nearly constant FLOPs,' and 'maximum FLOPs increase of 5/3 times'; please align these terms so the abstract and Section 3.2 do not overstate what Equation (4) proves.
- [Table 5 caption] Please state explicitly in the Table 5 caption whether the reported MACs include bilinear/nearest upsampling and the downsampling convolution, or only the depthwise convolutions.
Circularity Check
No significant circularity: the RecConv complexity and effective-receptive-field claims follow directly from the algorithm definition, and the empirical benchmarks are external.
full rationale
The central theoretical claims are derived from the algorithm's own definition rather than from fitted data or from self-citations. Algorithm 1 and Eq. 3 define RecConv as a shared stride-2 depthwise downsampling followed by depthwise convolutions at successively halved resolutions, with bilinear upsampling and addition during recombination. From this definition, Table 2 and Eq. 4 count parameters as k^2 C (ell+2) and FLOPs as a geometric series bounded by 5/3 times the base depthwise-convolution FLOPs; this is arithmetic applied to the stated operations, not a fitted quantity renamed as a prediction. The effective receptive field k times 2^ell likewise follows by construction: each recursive level halves the spatial resolution via a stride-2 convolution, so the aggregate support in the original image doubles per level, and Table 5's receptive-field entries such as [80,40,20,10] for a 5x5 kernel with levels [4,3,2,1] are direct evaluations of that formula. The paper's use of RepNeXt [85] as an architectural starting point and comparison baseline is a normal citation of the authors' prior work, and it is not load-bearing for the complexity derivation; no uniqueness theorem or prior result is invoked to force the RecConv design. The acknowledged weakness in the paper is practical throughput: the Limitations section and Table 5 show that bilinear interpolation and multi-scale data movement make wall-clock throughput drop even though convolution MACs stay nearly constant. That is an honesty about the scope of the FLOPs accounting, not a circularity, because the mathematical bound is stated for convolution operations and is not itself derived from the measured throughput. The empirical ImageNet, COCO, ADE20K, and ablation results are evaluated against external benchmarks and prior published models, so the accuracy claims do not reduce to the paper's own assumptions. No circular step is therefore present, and the appropriate score is 0.
Assumptions & free parameters
assumptions (3)
- standard math Geometric series bound: total FLOPs of recursive multi-scale convolutions converge to 5/3 of the base level (Eq. 4).
- domain assumption Effective receptive field of a k x k conv applied after ell stride-2 downsamples and upsamples is approximately k times 2^ell.
- domain assumption Replacing the wavelet transform of WTConv with a shared stride-2 depthwise downsampling preserves the multi-frequency representation and the large-receptive-field benefit.
invented entities (1)
-
RecConv module
independent evidence
Cite this review
Pith. "Pith review of RecConv: Efficient Recursive Convolutions for Multi-Frequency Representations." pith.science (2026). https://pith.science/paper/IQ3PZ2FM
@misc{pith2026241219628,
author = {Pith},
title = {Pith review of: RecConv: Efficient Recursive Convolutions for Multi-Frequency Representations},
year = {2026},
howpublished = {\url{https://pith.science/paper/IQ3PZ2FM}},
note = {Machine review of arXiv:2412.19628}
}
abstract
Recent advances in vision transformers (ViTs) have demonstrated the advantage of global modeling capabilities, prompting widespread integration of large-kernel convolutions for enlarging the effective receptive field (ERF). However, the quadratic scaling of parameter count and computational complexity (FLOPs) with respect to kernel size poses significant efficiency and optimization challenges. This paper introduces RecConv, a recursive decomposition strategy that efficiently constructs multi-frequency representations using small-kernel convolutions. RecConv establishes a linear relationship between parameter growth and decomposing levels which determines the effective receptive field $k\times 2^\ell$ for a base kernel $k$ and $\ell$ levels of decomposition, while maintaining constant FLOPs regardless of the ERF expansion. Specifically, RecConv achieves a parameter expansion of only $\ell+2$ times and a maximum FLOPs increase of $5/3$ times, compared to the exponential growth ($4^\ell$) of standard and depthwise convolutions. RecNeXt-M3 outperforms RepViT-M1.1 by 1.9 $AP^{box}$ on COCO with similar FLOPs. This innovation provides a promising avenue towards designing efficient and compact networks across various modalities. Codes and models can be found at https://github.com/suous/RecNeXt.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
https : / / github
Core ml tools. https : / / github . com / apple / coremltools, 2021. 5
2021
-
[2]
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hin- ton. Layer normalization. arXiv preprint arXiv:1607.06450,
-
[3]
End-to- end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to- end object detection with transformers. In European confer- ence on computer vision, pages 213–229. Springer, 2020. 1
2020
-
[4]
PeLK: Parameter-efficient Large Kernel ConvNets with Peripheral Convolution
Honghao Chen, Xiangxiang Chu, Yongjian Ren, Xin Zhao, and Kaiqi Huang. Pelk: Parameter-efficient large ker- nel convnets with peripheral convolution. arXiv preprint arXiv:2403.07589, 2024. 1, 2
work page Pith review arXiv 2024
-
[5]
Run, don’t walk: Chasing higher flops for faster neural networks
Jierun Chen, Shiu-hong Kao, Hao He, Weipeng Zhuo, Song Wen, Chul-Ho Lee, and S-H Gary Chan. Run, don’t walk: Chasing higher flops for faster neural networks. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12021–12031, 2023. 2, 9, 10
2023
-
[6]
Drop an octave: Reducing spatial redundancy in con- volutional neural networks with octave convolution
Yunpeng Chen, Haoqi Fan, Bing Xu, Zhicheng Yan, Yan- nis Kalantidis, Marcus Rohrbach, Shuicheng Yan, and Jiashi Feng. Drop an octave: Reducing spatial redundancy in con- volutional neural networks with octave convolution. In Pro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 3435–3444, 2019. 2
2019
-
[7]
Mobile- former: Bridging mobilenet and transformer
Yinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu, Xiaoyi Dong, Lu Yuan, and Zicheng Liu. Mobile- former: Bridging mobilenet and transformer. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5270–5279, 2022. 2
2022
-
[8]
Largekernel3d: Scaling up kernels in 3d sparse cnns
Yukang Chen, Jianhui Liu, Xiangyu Zhang, Xiaojuan Qi, and Jiaya Jia. Largekernel3d: Scaling up kernels in 3d sparse cnns. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 13488–13498,
Show all 91 references
-
[9]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 5, 7
2009
-
[10]
Scaling up your kernels to 31x31: Revisiting large kernel design in cnns
Xiaohan Ding, Xiangyu Zhang, Jungong Han, and Guiguang Ding. Scaling up your kernels to 31x31: Revisiting large kernel design in cnns. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 11963–11975, 2022. 1, 2
2022
-
[11]
Unireplknet: A univer- sal perception large-kernel convnet for audio, video, point cloud, time-series and image recognition
Xiaohan Ding, Yiyuan Zhang, Yixiao Ge, Sijie Zhao, Lin Song, Xiangyu Yue, and Ying Shan. Unireplknet: A univer- sal perception large-kernel convnet for audio, video, point cloud, time-series and image recognition. arXiv preprint arXiv:2311.15599, 2023. 2
2023 arXiv
-
[12]
ModernTCN: A modern pure convolution structure for general time series analysis
Luo donghao and wang xue. ModernTCN: A modern pure convolution structure for general time series analysis. In The Twelfth International Conference on Learning Representa- tions, 2024. 2
2024
-
[13]
An image is worth 16x16 words: Trans- formers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. arXiv preprint a...
2010 arXiv
-
[14]
Torchcam: class activation explorer
Franc ¸ois-Guillaume Fernandez. Torchcam: class activation explorer. https://github.com/frgfm/torch- cam, 2020. 1
2020
-
[15]
Wavelet convolutions for large receptive fields
Shahaf E Finder, Roy Amoyal, Eran Treister, and Oren Freifeld. Wavelet convolutions for large receptive fields. In European Conference on Computer Vision, 2024. 2
2024
-
[16]
Partial success in closing the gap between human and machine vision
Robert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus, Tizian Thieringer, Matthias Bethge, Felix A Wichmann, and Wieland Brendel. Partial success in closing the gap between human and machine vision. In Advances in Neural Information Processing Systems 34, 2021. 9
2021
-
[17]
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023. 7, 8
2023 arXiv
-
[18]
Efficiently mod- eling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher R´e. Efficiently mod- eling long sequences with structured state spaces. In The In- ternational Conference on Learning Representations (ICLR),
-
[19]
Segnext: Rethinking convolutional attention design for semantic segmentation
Meng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu, Ming-Ming Cheng, and Shi-Min Hu. Segnext: Rethinking convolutional attention design for semantic segmentation. arXiv preprint arXiv:2209.08575, 2022. 2
2022 arXiv
-
[20]
Flatten transformer: Vision transformer using focused linear attention
Dongchen Han, Xuran Pan, Yizeng Han, Shiji Song, and Gao Huang. Flatten transformer: Vision transformer using focused linear attention. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023. 5
2023
-
[21]
Demystify mamba in vision: A linear attention perspective
Dongchen Han, Ziyi Wang, Zhuofan Xia, Yizeng Han, Yi- fan Pu, Chunjiang Ge, Jun Song, Shiji Song, Bo Zheng, and Gao Huang. Demystify mamba in vision: A linear attention perspective. In NeurIPS, 2024. 5, 8, 10
2024
-
[22]
Ghostnet: More features from cheap opera- tions
Kai Han, Yunhe Wang, Qi Tian, Jianyuan Guo, Chunjing Xu, and Chang Xu. Ghostnet: More features from cheap opera- tions. In CVPR, 2020. 2
2020
-
[23]
Mobilemamba: Lightweight multi-receptive visual mamba network
Haoyang He, Jiangning Zhang, Yuxuan Cai, Hongxu Chen, Xiaobin Hu, Zhenye Gan, Yabiao Wang, Chengjie Wang, Yunsheng Wu, and Lei Xie. Mobilemamba: Lightweight multi-receptive visual mamba network. arXiv preprint arXiv:2411.15941, 2024. 2
2024 arXiv
-
[24]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 1, 3, 6
2016
-
[25]
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Doll ´ar, and Ross Gir- shick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017. 6
2017
-
[26]
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016. 3
2016 arXiv
-
[27]
Searching for mo- bilenetv3
Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, et al. Searching for mo- bilenetv3. In Proceedings of the IEEE/CVF international conference on computer vision, pages 1314–1324, 2019. 2
2019
-
[28]
Mobilenets: Efficient convolu- tional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco An- dreetto, and Hartwig Adam. Mobilenets: Efficient convolu- tional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017. 2
2017 arXiv
-
[29]
Lightvit: Towards light-weight convolution- free vision transformers
Tao Huang, Lang Huang, Shan You, Fei Wang, Chen Qian, and Chang Xu. Lightvit: Towards light-weight convolution- free vision transformers. arXiv preprint arXiv:2207.05557,
-
[30]
Are large kernels better teachers than transformers for convnets?,
Tianjin Huang, Lu Yin, Zhenyu Zhang, Li Shen, Meng Fang, Mykola Pechenizkiy, Zhangyang Wang, and Shiwei Liu. Are large kernels better teachers than transformers for convnets?,
-
[31]
Batch normalization: Accelerating deep network training by reducing internal co- variate shift
Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal co- variate shift. In International conference on machine learn- ing, pages 448–456. pmlr, 2015. 3
2015
-
[32]
Wavemix: A resource-efficient neural network for im- age analysis, 2023
Pranav Jeevan, Kavitha Viswanathan, Anandu A S, and Amit Sethi. Wavemix: A resource-efficient neural network for im- age analysis, 2023. 2
2023
-
[33]
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Franc ¸ois Fleuret. Transformers are rnns: Fast autoregressive transformers with linear attention. In ICML, 2020. 8, 10
2020
-
[34]
Panoptic feature pyramid networks
Alexander Kirillov, Ross Girshick, Kaiming He, and Piotr Doll´ar. Panoptic feature pyramid networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6399–6408, 2019. 6
2019
-
[35]
Imagenet classification with deep convolutional neural net- works
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural net- works. In Advances in neural information processing sys- tems, pages 1097–1105, 2012. 1, 2
2012
-
[36]
Fractalnet: Ultra-deep neural networks without residuals
Gustav Larsson, Michael Maire, and Gregory Shakhnarovich. Fractalnet: Ultra-deep neural networks without residuals. arXiv preprint arXiv:1605.07648 , 2016. 1
2016 arXiv
-
[37]
Backpropagation applied to handwrit- ten zip code recognition.Neural computation, 1(4):541–551,
Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel. Backpropagation applied to handwrit- ten zip code recognition.Neural computation, 1(4):541–551,
-
[38]
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. Visualizing the loss landscape of neural nets. In Neural Information Processing Systems, 2018. 3
2018
-
[39]
Siyuan Li, Zedong Wang, Zicheng Liu, Cheng Tan, Haitao Lin, Di Wu, Zhiyuan Chen, Jiangbin Zheng, and Stan Z. Li. Moganet: Multi-order gated aggregation network. In Inter- national Conference on Learning Representations, 2024. 2
2024
-
[40]
Selec- tive kernel networks
Xiang Li, Wenhai Wang, Xiaolin Hu, and Jian Yang. Selec- tive kernel networks. In CVPR, 2019. 2
2019
-
[41]
Efficientformer: Vision transformers at mobilenet speed
Yanyu Li, Geng Yuan, Yang Wen, Ju Hu, Georgios Evan- gelidis, Sergey Tulyakov, Yanzhi Wang, and Jian Ren. Efficientformer: Vision transformers at mobilenet speed. Advances in Neural Information Processing Systems , 35: 12934–12949, 2022. 2, 4, 5, 6
2022
-
[42]
Large selective kernel network for remote sensing object detection
Yuxuan Li, Qibin Hou, Zhaohui Zheng, Ming-Ming Cheng, Jian Yang, and Xiang Li. Large selective kernel network for remote sensing object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 16794–16805, 2023. 2
2023
-
[43]
Re- thinking vision transformers for mobilenet size and speed
Yanyu Li, Ju Hu, Yang Wen, Georgios Evangelidis, Kamyar Salahi, Yanzhi Wang, Sergey Tulyakov, and Jian Ren. Re- thinking vision transformers for mobilenet size and speed. In Proceedings of the IEEE international conference on com- puter vision, 2023. 1, 2, 3, 4, 5, 6
2023
-
[44]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceeding...
2014
-
[45]
More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity
Shiwei Liu, Tianlong Chen, Xiaohan Chen, Xuxi Chen, Qiao Xiao, Boqian Wu, Tommi K¨arkk¨ainen, Mykola Pechenizkiy, Decebal Mocanu, and Zhangyang Wang. More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity. arXiv preprint arXiv:2207.03620, 2022. 1, 2
2022 arXiv
-
[46]
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022, 2021. 1, 2, 3
2021
-
[47]
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feicht- enhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition , pages 11976–11986,
-
[48]
Rewrite the stars
Xu Ma, Xiyang Dai, Yue Bai, Yizhou Wang, and Yun Fu. Rewrite the stars. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, 2024. 1, 2
2024
-
[49]
Efficient modulation for vision net- works
Xu Ma, Xiyang Dai, Jianwei Yang, Bin Xiao, Yinpeng Chen, Yun Fu, and Lu Yuan. Efficient modulation for vision net- works. In The Twelfth International Conference on Learning Representations, 2024. 1, 2
2024
-
[50]
Edgenext: Efficiently amalgamated cnn-transformer architecture for mobile vision applications
Muhammad Maaz, Abdelrahman Shaker, Hisham Cholakkal, Salman Khan, Syed Waqas Zamir, Rao Muham- mad Anwer, and Fahad Shahbaz Khan. Edgenext: Efficiently amalgamated cnn-transformer architecture for mobile vision applications. In International Workshop on Computational Aspects o...
2022
-
[51]
Mobilevit: light- weight, general-purpose, and mobile-friendly vision trans- former
Sachin Mehta and Mohammad Rastegari. Mobilevit: light- weight, general-purpose, and mobile-friendly vision trans- former. arXiv preprint arXiv:2110.02178, 2021. 2
2021 arXiv
-
[52]
Separable self- attention for mobile vision transformers
Sachin Mehta and Mohammad Rastegari. Separable self- attention for mobile vision transformers. arXiv preprint arXiv:2206.02680, 2022. 2, 5
2022 arXiv
-
[53]
Mo- bilevig: Graph-based sparse attention for mobile vision ap- plications
Mustafa Munir, William Avery, and Radu Marculescu. Mo- bilevig: Graph-based sparse attention for mobile vision ap- plications. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2210– 2218, 2023. 4
2023
-
[54]
Edgevits: Competing light-weight cnns on mobile devices with vision transformers
Junting Pan, Adrian Bulat, Fuwen Tan, Xiatian Zhu, Lukasz Dudziak, Hongsheng Li, Georgios Tzimiropoulos, and Brais Martinez. Edgevits: Competing light-weight cnns on mobile devices with vision transformers. In European Conference on Computer Vision, pages 294–311. Springer, 2022. 2
2022
-
[55]
Large kernel matters–improve semantic segmen- tation by global convolutional network
Chao Peng, Xiangyu Zhang, Gang Yu, Guiming Luo, and Jian Sun. Large kernel matters–improve semantic segmen- tation by global convolutional network. In Proceedings of the IEEE conference on computer vision and pattern recog- nition, pages 4353–4361, 2017. 2
2017
-
[56]
Designing network design spaces
Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Doll ´ar. Designing network design spaces. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 10428–10436, 2020. 5, 9
2020
-
[57]
Global filter networks for image classification
Yongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu, and Jie Zhou. Global filter networks for image classification. Advances in neural information processing systems, 34:980– 993, 2021. 2
2021
-
[58]
Hornet: Efficient high-order spatial interactions with recursive gated convolutions
Yongming Rao, Wenliang Zhao, Yansong Tang, Jie Zhou, Ser-Lam Lim, and Jiwen Lu. Hornet: Efficient high-order spatial interactions with recursive gated convolutions. Ad- vances in Neural Information Processing Systems (NeurIPS),
-
[59]
You only look once: Unified, real-time object de- tection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object de- tection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 779–788, 2016. 1, 2
2016
-
[60]
Mobilenetv2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zh- moginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 4510–4520, 2018. 2
2018
-
[61]
Schuster and K.K
M. Schuster and K.K. Paliwal. Bidirectional recurrent neural networks. IEEE Transactions on Signal Processing, 45(11): 2673–2681, 1997. 7, 8
1997
-
[62]
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In CVPR, pages 618–626, 2017. 1
2017
-
[63]
Swiftformer: Efficient additive attention for transformer- based real-time mobile vision applications
Abdelrahman Shaker, Muhammad Maaz, Hanoona Rasheed, Salman Khan, Ming-Hsuan Yang, and Fahad Shahbaz Khan. Swiftformer: Efficient additive attention for transformer- based real-time mobile vision applications. In Proceedings of the IEEE/CVF International Conference on Computer ...
2023
-
[64]
Very deep convo- lutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. Very deep convo- lutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 2
2014 arXiv
-
[65]
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 1–9, 2015. 2
2015
-
[66]
Rethinking the inception archi- tecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception archi- tecture for computer vision. In Proceedings of the IEEE con- ference on computer vision and pattern recognition , pages 2818–2826, 2016. 2
2016
-
[67]
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi. Inception-v4, inception-resnet and the impact of residual connections on learning. In Thirty-first AAAI conference on artificial intelligence, 2017. 2
2017
-
[68]
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning, pages 6105–6114. PMLR,
-
[69]
Mnas- net: Platform-aware neural architecture search for mobile
Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V Le. Mnas- net: Platform-aware neural architecture search for mobile. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 2820–2828, 2019. 2
2019
-
[70]
Mlp-mixer: An all-mlp architecture for vision
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lu- cas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al. Mlp-mixer: An all-mlp architecture for vision. arXiv preprint arXiv:2105.01601, 2021. 2
2021 arXiv
-
[71]
Patches are all you need? arXiv preprint arXiv:2201.09792, 2022
Asher Trockman and J Zico Kolter. Patches are all you need? arXiv preprint arXiv:2201.09792, 2022. 2
2022 arXiv
-
[72]
Fastvit: A fast hybrid vi- sion transformer using structural reparameterization
Pavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, On- cel Tuzel, and Anurag Ranjan. Fastvit: A fast hybrid vi- sion transformer using structural reparameterization. arXiv preprint arXiv:2303.14189, 2023. 1, 2, 4, 5, 6
2023 arXiv
-
[73]
Mobileone: An im- proved one millisecond mobile backbone
Pavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel, and Anurag Ranjan. Mobileone: An im- proved one millisecond mobile backbone. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7907–7917, 2023. 2, 5
2023
-
[74]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 3
2017
-
[75]
Mobilevitv3: Mobile-friendly vision transformer with simple and effec- tive fusion of local, global and input features
Shakti N Wadekar and Abhishek Chaurasia. Mobilevitv3: Mobile-friendly vision transformer with simple and effec- tive fusion of local, global and input features. arXiv preprint arXiv:2209.15159, 2022. 2
2022 arXiv
-
[76]
Repvit: Revisiting mobile cnn from vit perspective
Ao Wang, Hui Chen, Zijia Lin, Hengjun Pu, and Guiguang Ding. Repvit: Revisiting mobile cnn from vit perspective. arXiv preprint arXiv:2307.09283, 2023. 1, 2, 3, 4, 5, 6
2023 arXiv
-
[77]
Yolov10: Real-time end- to-end object detection
Ao Wang, Hui Chen, Lihao Liu, Kai Chen, Zijia Lin, Jun- gong Han, and Guiguang Ding. Yolov10: Real-time end- to-end object detection. arXiv preprint arXiv:2405.14458 ,
-
[78]
Lsnet: See large, focus small, 2025
Ao Wang, Hui Chen, Zijia Lin, Jungong Han, and Guiguang Ding. Lsnet: See large, focus small, 2025. 9, 10, 1
2025
-
[79]
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE/CVF international conference on computer vision , p...
2021
-
[80]
Early convolutions help trans- formers see better
Tete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell, Piotr Doll´ar, and Ross Girshick. Early convolutions help trans- formers see better. Advances in neural information process- ing systems, 34:30392–30400, 2021. 3
2021
-
[81]
Metaformer is actually what you need for vision
Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan. Metaformer is actually what you need for vision. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10819–10829, 2022. 3, 5, 6
2022
-
[82]
Inceptionnext: when inception meets convnext
Weihao Yu, Pan Zhou, Shuicheng Yan, and Xinchao Wang. Inceptionnext: when inception meets convnext. arXiv preprint arXiv:2303.16900, 2023. 2, 3, 9, 10
2023 arXiv
-
[83]
Shvit: Single-head vision transformer with memory efficient macro design
Seokju Yun and Youngmin Ro. Shvit: Single-head vision transformer with memory efficient macro design. arXiv preprint arXiv:2401.16456, 2024. 2, 9, 10
2024 arXiv
-
[84]
Parc-net: Po- sition aware circular convolution with merits from convnets and transformer
Haokui Zhang, Wenze Hu, and Xiaoyu Wang. Parc-net: Po- sition aware circular convolution with merits from convnets and transformer. In European Conference on Computer Vi- sion, 2022. 2
2022
-
[85]
Repnext: A fast multi-scale cnn using structural reparameterization, 2024
Mingshu Zhao, Yi Luo, and Yong Ouyang. Repnext: A fast multi-scale cnn using structural reparameterization, 2024. 1, 2, 3, 4, 5, 6
2024
-
[86]
Torr, and Li Zhang
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip H.S. Torr, and Li Zhang. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In CVPR, 2021. 1
2021
-
[87]
Scene parsing through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 633–641,
-
[89]
this approach treats the sequence of downsampled feature maps from each decomposition level as the input to a recurrent model
RecConv Aggregation Algorithm 2 shows the implementation of the RecConv aggregation like SSM in a PyTorch-like style. this approach treats the sequence of downsampled feature maps from each decomposition level as the input to a recurrent model
-
[90]
T”, “S”, “B
Grad-CAM Visualization We visualize class activation maps (CAM) using Grad- CAM [62] with the TorchCAM Toolbox [14]. As shown in Figure 5, RecNeXt achieves a better trade-off between local and global features compared to other efficient models. Table 11. Architectural details ...
-
[91]
M” and “A
Training Details Table 12. Training recipe of “M” and “A” series and “T”, “S”, “B” variants of RecNeXt on ImageNet-1K. DropPath [36] is set to 0.0 for hard knowledge distillation. Hyperparameters Config optimizer AdamW batch size 1024 learning rate 2e − 3 LR schedule cosine tr...
-
[2017]
h = None for i, o in reversed(zip(fs[1:], fs[:-1])): h = self.a(h) + self.b(i) if h else self.b(i) h = interpolate(h, size=o.shape[2:]) return self.c(h) + self.d(x)
5, 6 RecConv: Efficient Recursive Convolutions for Multi-Frequency Representations Supplementary Material Algorithm 2 Recurrent aggregation in a PyTorch-like style class RecConv(Module): def __init__(self, ic, ks=5, level=1): super().__init__() self.level = level kwargs = { ’i...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.