Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

EMOv2: Pushing 5M Vision Model Frontier

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read EMOv2-5M, built from one shared-weight spanning-attention block, reaches 79.4% top-1 on ImageNet-1K with 5.1M parameters.

desk verdict A solid, incremental extension of EMOv1 with a genuinely new parameter-shared spanning attention mechanism; the internal controlled comparisons hold up, but the '5M frontier' claim rests partly on cross-paper baselines with different training recipes. read the letter →

arxiv 2412.06674 v1 pith:4GQVLH4S submitted 2024-12-09 cs.CV

classification cs.CV
keywords lightweightvisionbackboneinvertedresidualblockspanningwindowattentionMetaMobileImageNetclassificationdensepredictionefficient
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes EMOv2, a family of lightweight vision backbones whose 5M-parameter variant reaches 79.4% top-1 accuracy on ImageNet-1K with 5.1M parameters and 1.0G FLOPs. The authors argue that a single one-residual block, combining depthwise convolution with a shared-weight spanning window attention, defines the current accuracy frontier at the 5M scale. A sympathetic reader would care because mobile and edge applications operate under strict parameter and latency budgets, and the paper offers a uniform building block that transfers to detection, segmentation, video, and image generation.

What carries the argument

The load-bearing object is the spanning attention mechanism (SEW-MHSA), a parameter-shared dual-window attention that partitions the same Q, K, V into adjacent local windows and into remote windows sampled with stride $[H/h, W/w]$, then fuses the two attention maps. Because both partitions reuse the same Q, K, V projections and attention weights, it adds no parameters and only a small FLOP overhead; this is what lets a 5M model see distant interactions in high-resolution inputs. The surrounding i2RMB block wraps this attention with a large-kernel depthwise convolution, a post-attention nonlinearity, and a single residual connection.

What would settle it

Retrain similarly sized baselines such as EdgeNeXt-S, MobileViT-S, and EMOv1-5M under EMOv2's exact training recipe, then compare ImageNet top-1 and COCO RetinaNet mAP; if the reported margins shrink to noise or reverse, the claimed frontier position would not hold.

Watch

Extended reading notes

Core claim

The central claim is that the Inverted Residual Block of MobileNetV2 and the MHSA/FFN modules of Transformers share a common one-residual meta-structure, which the paper abstracts as the Meta Mobile Block with expansion ratio $\lambda$ and efficient operator $F$. Instantiating $F$ as a cascade of expanded-window multi-head self-attention and depthwise convolution gives the iRMB; adding a parameter-shared second window partition that samples distant tokens at stride $[H/h, W/w]$ produces the improved i2RMB with spanning attention. Using only i2RMB blocks, EMOv2-5M reaches 79.4% top-1 at 5.1M parameters and 1.0G FLOPs, surpassing EMOv1-5M by +1.0, and 41.5 mAP with RetinaNet, +2.6 over EMOv1. The paper also reports 82.9% top-1 at 512 resolution with knowledge distillation and 1000 epochs, and scales the design to 20M and 50M variants.

Load-bearing premise

The published results of competing lightweight models, each trained under its own recipe, are directly comparable to the authors' numbers even though the training setups differ.

Editorial extensions

If this is right

  • At the 5M scale, EMOv2-5M outperforms published CNN-, Transformer-, and RNN-based lightweight models on ImageNet-1K, including MobileNetV3-L-1.25, EdgeNeXt-S, EfficientFormerV2-S1, and Vim-Ti.
  • The same ImageNet-pretrained backbone raises dense-prediction results, reaching 29.6 mAP with SSDLite, 41.5 mAP with RetinaNet, 42.3 box AP with Mask R-CNN, and 39.8 mIoU with DeepLabv3 on ADE20K.
  • Replacing the Transformer block in DiT with i2RMB cuts FID from 68.4 to 46.3 at the S scale and from 19.5 to 9.6 at the XL scale while using fewer parameters.
  • Extending i2RMB to the temporal dimension gives V-EMOv2-5M 65.2% top-1 on Kinetics-400 with 5.9M parameters, beating UniFormer-XXS's 63.2% with 9.8M parameters.
  • With a stronger training recipe of 512 resolution, knowledge distillation, and 1000 epochs, EMOv2-5M reaches 82.9% top-1, indicating the architecture retains headroom.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper does not run: train EMOv2-5M and EMOv1-5M at increasing resolutions from 224 to 512 and measure the accuracy gap, which should widen if spanning attention's distant modeling is the cause.
  • The dual-window partition is a general token-grouping trick, so it could be inserted into other window-attention or state-space vision models without changing the attention computation itself.
  • Because the Meta Mobile Block abstracts IRB, MHSA, and FFN into one residual structure, a future design could choose the efficient operator F per stage or per hardware target while keeping the same surrounding block.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes EMOv2, a family of lightweight vision backbones built from a single one-residual block called i2RMB. The block combines an inverted-residual structure with expanded-window multi-head self-attention (EW-MHSA) and depthwise convolution, and introduces SEW-MHSA, a parameter-shared dual-window attention mechanism that models both nearby and distant tokens in parallel. The authors claim new state-of-the-art results at the 1M/2M/5M parameter scale, including 79.4% ImageNet-1K top-1 accuracy with 5.1M parameters and 1.03G FLOPs, and consistent improvements over EMOv1 in object detection, semantic segmentation, video classification, and diffusion-based image generation. The controlled EMOv2-versus-EMOv1 comparisons support the block-level improvements; the broader "frontier" claim relies on cross-paper comparisons whose training recipes differ substantially.

Significance. The paper has a clean architectural message and unusually broad evaluation: the i2RMB block is simple, the ablations in Tables 5 and 17 are systematic, the code is promised, and the controlled EMOv2-versus-EMOv1 results are internally consistent. If the empirical claims hold, the spanning-attention mechanism is a useful contribution for lightweight vision models. However, the headline state-of-the-art status is established by comparing against published baselines that use different training recipes; Table A1 documents large recipe differences (e.g., MobileNetV4 uses NAS and KD, GhostNetV3 uses LAMB and re-parameterization, EdgeNeXt uses multi-scale training, while EMOv2 uses none of these). The controlled comparisons cannot by themselves establish the "5M frontier" claim, and the unproved FLOP-saving equivalence in Section 3.2.2 leaves part of the efficiency story unsupported. The strengths and weaknesses together point to a revise-and-resubmit rather than acceptance in current form.

major comments (3)
  1. [Sec. 4.1, Sec. 4.2, Tab. A1] The paper explicitly acknowledges in Sec. 4.1 that "Different SoTA methods use various training recipes that could lead to potentially unfair comparisons" and in Sec. 4.2 that "different methods lack a unified experimental standard." These caveats apply directly to the headline comparisons in Tabs. 7-11 and Tab. 19: e.g., MobileNetV4 uses NAS and KD with 500 epochs, GhostNetV3 uses LAMB, KD, and re-parameterization, and EdgeNeXt uses multi-scale training and position embeddings, whereas EMOv2 is trained with a weaker recipe. The paper reports EMOv2 under other recipes (Tab. 6) but does not retrain the published baselines under a common protocol, so margins such as +0.9 over EATFormer-Tiny (Tab. 7) and +2.8 mAP over EdgeViT-XXS (Tab. 9) could shrink or invert under a unified setting. The controlled EMOv2-vs-EMOv1 numbers (+1.0 classification, +1.7 SSDLite, +2.6 RetinaNet) support the mechanism, but the "frontier" claim needs either common-protocol retraining of key baselines or a carefully narrowed claim.
  2. [Sec. 3.2.2, "Efficient equivalent implementation"] The paragraph introducing "pre-attention" states that when the number of groups in MLPe equals the number of heads in EW-MHSA, the multiplication order can be exchanged, and that matrix multiplication before MLPe therefore reduces FLOPs. This proposition is asserted without proof or citation. It is load-bearing because the FLOP counts in Tabs. 2 and 7 (e.g., 1035M for EMOv2-5M) and the claimed efficiency of EW-MHSA rely on this pre-attention implementation. Please provide a short derivation or a precise citation; if the equivalence is only approximate, the reported complexity numbers should be revised.
  3. [Sec. 4.1 vs Sec. 4.3, Tab. 17c] The default hyperparameters are stated inconsistently across the paper. Section 4.1 reports a batch size of 2,048 and Tab. A1 lists batch size 2,048 with drop path rate 0.1, while Tab. 17c shows the best results at batch size 1,024 and drop path rate 0.05 and suggests batch size 1,024 as the default. This ambiguity affects reproducibility of the main 79.4 result and should be resolved by specifying one canonical configuration that is used for all reported comparison tables.
minor comments (5)
  1. [Tab. 7 footnotes] The footnote uses the same symbol '*' for both "Neural Architecture Search" and "stronger training strategy displayed in Tab. 17(e)"; use distinct symbols or letters for the two definitions.
  2. [Throughout Tab. 7 and related tables] Resolutions appear as "2242", "2562", and "5122", which should be rendered as 224², 256², and 512²; as printed they are easy to misread as literal dimensions.
  3. [Sec. 3.3.1, "Non-linearity for post-attention"] The terms "pre-attention" (Sec. 3.2.2) and "post-attention" (Sec. 3.3.1) are confusing because both refer to the placement of the channel-expansion MLP relative to the attention-map multiplication rather than to attention operating before or after convolution; consider renaming them to "expand-before-attention" and "expand-after-attention".
  4. [Sec. 2, Related Work] The sentences "Tao et al. [53] introduces additional learnable tokens" and "Chen et al. [53] design a parallel structure" both cite reference [53], but [53] is LightViT by Huang et al.; the second sentence appears to refer to MobileFormer (reference [33]) and should be re-cited.
  5. [Sec. 4.2, Tab. 12] The UNet-based segmentation model is referred to as U-EMO in the text and as "U-EMOv2-5M" in the table header, but the backbone column of the last row reads "EMOv2-5M"; please use a consistent name.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation found: i2RMB and spanning attention are validated by controlled ablations and external benchmarks; reliance on the authors' prior EMOv1 is a normal baseline comparison, not a forced input.

full rationale

The paper's central claims are architectural and empirical rather than derived from their own outputs. The Meta Mobile Block is presented as an inductive abstraction of IRB, MHSA, and FFN, but the paper does not use this abstraction as proof of performance; instead, performance is established through controlled experiments against external baselines (ImageNet-1K, COCO, ADE20K, Kinetics-400, ImageNet generation) and internal ablations (Tabs. 3, 5, 6, 15, 17). The spanning attention mechanism reuses shared Q, K, V projections and adds only an extra attention-map computation; this is a structural design with a direct complexity argument, not a fitted parameter relabeled as a prediction. The paper does rely on the authors' prior EMOv1 [13] as a baseline and predecessor, and there is substantial self-citation; however, EMOv1 is a published ICCV 2023 result used as a comparator, not as an unverified premise that forces the EMOv2 outcome. The comparisons against other state-of-the-art methods use published numbers under different training recipes, and the paper itself disclaims in Sec. 4.1 that 'Different SoTA methods use various training recipes that could lead to potentially unfair comparisons' and in Sec. 4.2 that 'different methods lack a unified experimental standard.' This is a legitimate fairness and correctness risk for the 'frontier' claim, but it is not circularity: the EMOv2 results are obtained by running the proposed model, not by reading them off the baselines. No uniqueness theorem is imported from the authors, no ansatz is smuggled via citation, and no known result is merely renamed as a derivation. The only mild concern is the heavy reliance on the authors' own EMOv1 and related works for motivation and comparison, which is normal in an extension paper and does not make the central derivation circular. Overall score 2 reflects this minor self-citation footprint, not load-bearing circularity.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

The central claim is empirical. Free parameters are the architectural hyperparameters tuned on ImageNet and validated by ablations; the axioms are the unproved equivalences and comparability assumptions the reported numbers rest on. No new physical or mathematical entities are introduced.

free parameters (4)
  • Per-stage expansion ratios for EMOv2-5M = [2.0, 3.0, 4.0, 4.0]
    Chosen by hand and validated via ablation (Table 17) to maximize ImageNet top-1 at 5.1M params.
  • DW-Conv kernel size in i2RMB = 5
    Selected from {1,3,5,7,9} ablation in Table 17d; K-5 gives the best top-1 (79.4).
  • Depth and channel configuration = Depth [3,3,9,3], channels [48,72,160,288]
    Depth/channel ablations in Tables 15 and 17b show this configuration is near-optimal at 5M.
  • Drop path rate and batch size = DPR 0.05, BS 1024 (suggested default)
    Ablated in Table 17c; chosen as the best stable point for EMOv2-5M.
assumptions (3)
  • domain assumption Shared-parameter spanning attention can model both local and distant correlations without extra parameters
    The paper validates this empirically (Tables 5, 17a) but provides no formal argument that the same weights are sufficient for both scales.
  • standard math The pre-attention equivalence (exchanging the order of attention multiplication and channel expansion when MLP groups equal attention heads) preserves the result
    Stated as an equivalent proposition in Sec. 3.2.2 without proof; used to justify the FLOPs counts for EW-MHSA.
  • domain assumption Published baseline numbers under different training recipes are comparable to the authors' runs
    The paper enumerates recipes in Table A1 and claims its own recipe is weaker, but most comparison numbers are taken from other papers rather than re-run in a unified setting (Sec. 4.1, 4.2).

how reviews work

0 comments
Cite this review

Pith. "Pith review of EMOv2: Pushing 5M Vision Model Frontier." pith.science (2026). https://pith.science/paper/4GQVLH4S

@misc{pith2026241206674,
  author       = {Pith},
  title        = {Pith review of: EMOv2: Pushing 5M Vision Model Frontier},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4GQVLH4S}},
  note         = {Machine review of arXiv:2412.06674}
}
read the original abstract

This work focuses on developing parameter-efficient and lightweight models for dense predictions while trading off parameters, FLOPs, and performance. Our goal is to set up the new frontier of the 5M magnitude lightweight model on various downstream tasks. Inverted Residual Block (IRB) serves as the infrastructure for lightweight CNNs, but no counterparts have been recognized by attention-based design. Our work rethinks the lightweight infrastructure of efficient IRB and practical components in Transformer from a unified perspective, extending CNN-based IRB to attention-based models and abstracting a one-residual Meta Mobile Block (MMBlock) for lightweight model design. Following neat but effective design criterion, we deduce a modern Improved Inverted Residual Mobile Block (i2RMB) and improve a hierarchical Efficient MOdel (EMOv2) with no elaborate complex structures. Considering the imperceptible latency for mobile users when downloading models under 4G/5G bandwidth and ensuring model performance, we investigate the performance upper limit of lightweight models with a magnitude of 5M. Extensive experiments on various vision recognition, dense prediction, and image generation tasks demonstrate the superiority of our EMOv2 over state-of-the-art methods, e.g., EMOv2-1M/2M/5M achieve 72.3, 75.8, and 79.4 Top-1 that surpass equal-order CNN-/Attention-based models significantly. At the same time, EMOv2-5M equipped RetinaNet achieves 41.5 mAP for object detection tasks that surpasses the previous EMO-5M by +2.6. When employing the more robust training recipe, our EMOv2-5M eventually achieves 82.9 Top-1 accuracy, which elevates the performance of 5M magnitude models to a new level. Code is available at https://github.com/zhangzjn/EMOv2.

Figures

Figures reproduced from arXiv: 2412.06674 by the authors.

Figure 1
Figure 1. Top: Performance vs. Parameters with concurrent methods. Our EMOv2 achieves significant accuracy with fewer parameters. Superscript ∗: The comparison methods employ more robust train￾ing strategies described in their papers, while ours uses the strategy mentioned in Tab. 17(e). Bottom: The range of token interactions varies with different window attention mechanisms. Our EMOv2, with parameter-shared spanning attenti… view at source ↗
Figure 2
Figure 2. Left: Abstracted unified Meta-Mobile Block from Multi-Head Self-Attention, Feed-Forward Network [35], and Inverted Residual Block [9] (c.f . Sec 3.2.1). The inductive block can be deduced into specific modules using different expansion ratio λ and efficient operator F. Middle: We construct a family of vision models based on our i2RMB module: 4-stage EMOv2, composed solely of the deduced i 2RMB (c.f . Sec 3.2.2), for… view at source ↗
Figure 3
Figure 3. Meta-paradigm comparison between our MMBlock and [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Detailed implementation comparison of the Inverted [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Downstream gains of EMOv2-5M over EMOv1-5M. Additionally, it demonstrates notable improvements across various high￾resolution downstream tasks. For in￾stance, in popular detection and segmen￾tation tasks, as shown in [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 8
Figure 8. Figure 8: Overall incremental trajectory from baseline to modern EMOv2 at the 5M magnitude. Each line is based on a modi￾fication of the immediately preceding line. Detailed ablations in Sec. 4.3. Parameters and FLOPs are marked in green and yellow. with an additional 0.1G FLOPs…
Figure 7
Figure 7. Figure 7: Visualizations by Grad-CAM. EMOv2 generates sharper [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VQualA 2025 Challenge on Face Image Quality Assessment: Methods and Results

    cs.CV 2025-08 conditional novelty 3.0 of 10

    An ICCV 2025 workshop challenge compared lightweight face image quality assessment models under strict compute limits, and this report surveys the winning methods.

Reference graph

Works this paper leans on

117 extracted references · 50 canonical work pages · cited by 1 Pith paper

  1. [1]

    Rethinking vision transformers for mobilenet size and speed,

    Y . Li, J. Hu, Y . Wen, G. Evangelidis, K. Salahi, Y . Wang, S. Tulyakov, and J. Ren, “Rethinking vision transformers for mobilenet size and speed,” in ICCV, 2023. 1, 2, 3, 5, 7, 8, 16, 17

  2. [2]

    Edgenext: efficiently amalgamated cnn-transformer architecture for mobile vision applications,

    M. Maaz, A. Shaker, H. Cholakkal, S. Khan, S. W. Zamir, R. M. Anwer, and F. S. Khan, “Edgenext: efficiently amalgamated cnn-transformer architecture for mobile vision applications,” in ECCVW, 2022. 1, 3, 5, 6, 7, 8, 9, 10, 16, 17 IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 13

  3. [3]

    Tinysam: Pushing the envelope for efficient segment anything model,

    H. Shu, W. Li, Y . Tang, Y . Zhang, Y . Chen, H. Li, Y . Wang, and X. Chen, “Tinysam: Pushing the envelope for efficient segment anything model,” arXiv preprint arXiv:2312.13789, 2023. 1

  4. [4]

    Edgesam: Prompt-in-the- loop distillation for on-device deployment of sam,

    C. Zhou, X. Li, C. C. Loy, and B. Dai, “Edgesam: Prompt-in-the- loop distillation for on-device deployment of sam,” arXiv preprint arXiv:2312.06660, 2023. 1

  5. [5]

    RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything

    S. Xu, H. Yuan, Q. Shi, L. Qi, J. Wang, Y . Yang, Y . Li, K. Chen, Y . Tong, B. Ghanem et al., “Rap-sam: Towards real-time all-purpose segment anything,” arXiv preprint arXiv:2401.10228, 2024. 1

  6. [6]

    Semantic flow for fast and accurate scene parsing,

    X. Li, A. You, Z. Zhu, H. Zhao, M. Yang, K. Yang, and Y . Tong, “Semantic flow for fast and accurate scene parsing,” in ECCV, 2020. 1, 2

  7. [7]

    RTMO: Towards high-performance one-stage real-time multi-person pose estimation,

    P. Lu, T. Jiang, Y . Li, X. Li, K. Chen, and W. Yang, “RTMO: Towards high-performance one-stage real-time multi-person pose estimation,”

  8. [8]

    Mobilenets: Efficient convolu- tional neural networks for mobile vision applications,

    A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolu- tional neural networks for mobile vision applications,” arXiv preprint arXiv:1704.04861, 2017. 1, 2, 3, 8

Show all 117 references
  1. [9]

    Mobilenetv2: Inverted residuals and linear bottlenecks,

    M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in CVPR, 2018. 1, 2, 3, 4, 8

  2. [10]

    Searching for mobilenetv3,

    A. Howard, M. Sandler, G. Chu, L.-C. Chen, B. Chen, M. Tan, W. Wang, Y . Zhu, R. Pang, V . Vasudevanet al., “Searching for mobilenetv3,” in ICCV, 2019. 1, 2, 7, 8, 12, 16, 17

  3. [11]

    Ghostnet: More features from cheap operations,

    K. Han, Y . Wang, Q. Tian, J. Guo, C. Xu, and C. Xu, “Ghostnet: More features from cheap operations,” in CVPR, 2020. 1, 2

  4. [12]

    Efficientnet: Rethinking model scaling for convolu- tional neural networks,

    M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolu- tional neural networks,” in ICML. PMLR, 2019. 1, 2, 3, 8

  5. [13]

    Rethinking mobile block for efficient attention- based models,

    J. Zhang, X. Li, J. Li, L. Liu, Z. Xue, B. Zhang, Z. Jiang, T. Huang, Y . Wang, and C. Wang, “Rethinking mobile block for efficient attention- based models,” in ICCV, 2023. 1, 2, 6, 7, 8, 9, 10, 11, 12

  6. [14]

    Separable self-attention for mobile vision transformers,

    S. Mehta and M. Rastegari, “Separable self-attention for mobile vision transformers,” TMLR, 2023. 1, 3, 7, 8, 16, 17

  7. [15]

    The need for speed in ai,

    J. Nielsen, “The need for speed in ai,” 2023, accessed: 2023-10-03. [Online]. Available: https://www.uxtigers.com/post/ai-response-time 1

  8. [16]

    Morgan Kaufmann, 1994

    ——, Usability engineering. Morgan Kaufmann, 1994. 1

  9. [17]

    Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer,

    S. Mehta and M. Rastegari, “Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer,” in ICLR, 2022. 1, 2, 3, 6, 7, 8, 16, 17

  10. [18]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR, 2021. 1, 3, 4, 7, 16, 17

  11. [19]

    Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,

    W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in ICCV, 2021. 1, 3, 5, 8, 9

  12. [20]

    Pvt v2: Improved baselines with pyramid vision transformer,

    ——, “Pvt v2: Improved baselines with pyramid vision transformer,” CVM, 2022. 1, 8, 9, 11

  13. [21]

    Swin transformer: Hierarchical vision transformer using shifted windows,

    Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in ICCV, 2021. 1, 3, 5, 11

  14. [22]

    Swin transformer v2: Scaling up capacity and resolution,

    Z. Liu, H. Hu, Y . Lin, Z. Yao, Z. Xie, Y . Wei, J. Ning, Y . Cao, Z. Zhang, L. Dong et al., “Swin transformer v2: Scaling up capacity and resolution,” in CVPR, 2022. 1

  15. [23]

    Analogous to evolutionary algorithm: Designing a unified sequence model,

    J. Zhang, C. Xu, J. Li, W. Chen, Y . Wang, Y . Tai, S. Chen, C. Wang, F. Huang, and Y . Liu, “Analogous to evolutionary algorithm: Designing a unified sequence model,” NeurIPS, 2021. 1

  16. [24]

    Eatformer: improving vision transformer inspired by evolutionary algorithm,

    J. Zhang, X. Li, Y . Wang, C. Wang, Y . Yang, Y . Liu, and D. Tao, “Eatformer: improving vision transformer inspired by evolutionary algorithm,” IJCV, 2024. 1, 7, 8, 9, 11

  17. [25]

    Transformer-based visual segmentation: A survey,

    X. Li, H. Ding, W. Zhang, H. Yuan, G. Cheng, P. Jiangmiao, K. Chen, Z. Liu, and C. C. Loy, “Transformer-based visual segmentation: A survey,”TPAMI, 2024. 1

  18. [26]

    Involution: Inverting the inherence of convolution for visual recognition,

    D. Li, J. Hu, C. Wang, X. Li, Q. She, L. Zhu, T. Zhang, and Q. Chen, “Involution: Inverting the inherence of convolution for visual recognition,” in CVPR, 2021. 1

  19. [27]

    Reformer: The efficient transformer,

    N. Kitaev, L. Kaiser, and A. Levskaya, “Reformer: The efficient transformer,” in ICLR, 2020. 1

  20. [28]

    Rethinking attention with performers,

    K. M. Choromanski, V . Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser, D. B. Belanger, L. J. Colwell, and A. Weller, “Rethinking attention with performers,” in ICLR, 2021. 1

  21. [29]

    Cvt: Introducing convolutions to vision transformers,

    H. Wu, B. Xiao, N. Codella, M. Liu, X. Dai, L. Yuan, and L. Zhang, “Cvt: Introducing convolutions to vision transformers,” in ICCV, 2021. 1, 3

  22. [30]

    Next-vit: Next generation vision transformer for efficient deploy- ment in realistic industrial scenarios,

    J. Li, X. Xia, W. Li, H. Li, X. Wang, X. Xiao, R. Wang, M. Zheng, and X. Pan, “Next-vit: Next generation vision transformer for efficient deploy- ment in realistic industrial scenarios,” arXiv preprint arXiv:2207.05501,

  23. [31]

    Delight: Deep and light-weight transformer,

    S. Mehta, M. Ghazvininejad, S. Iyer, L. Zettlemoyer, and H. Hajishirzi, “Delight: Deep and light-weight transformer,” in ICLR, 2021. 1

  24. [32]

    Mobilevitv3: Mobile-friendly vision transformer with simple and effective fusion of local, global and input features,

    S. N. Wadekar and A. Chaurasia, “Mobilevitv3: Mobile-friendly vision transformer with simple and effective fusion of local, global and input features,” arXiv preprint arXiv:2209.15159, 2022. 1, 3

  25. [33]

    Mobile-former: Bridging mobilenet and transformer,

    Y . Chen, X. Dai, D. Chen, M. Liu, X. Dong, L. Yuan, and Z. Liu, “Mobile-former: Bridging mobilenet and transformer,” in CVPR, 2022. 1, 8

  26. [34]

    Efficientformer: Vision transformers at mobilenet speed,

    Y . Li, G. Yuan, Y . Wen, J. Hu, G. Evangelidis, S. Tulyakov, Y . Wang, and J. Ren, “Efficientformer: Vision transformers at mobilenet speed,” NeurIPS, 2022. 1

  27. [35]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS,

  28. [36]

    Focal loss for dense object detection,

    T.-Y . Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in ICCV, 2017. 2, 8, 11, 16, 17

  29. [37]

    Squeezenet: Alexnet-level accuracy with 50x fewer parameters and< 0.5 mb model size,

    F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer, “Squeezenet: Alexnet-level accuracy with 50x fewer parameters and< 0.5 mb model size,” arXiv preprint arXiv:1602.07360,

  30. [38]

    Rethinking the inception architecture for computer vision,

    C. Szegedy, V . Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in CVPR, 2016. 2

  31. [39]

    Sfnet: Faster, accurate, and domain agnostic semantic segmentation via semantic flow,

    X. Li, J. Zhang, Y . Yang, G. Cheng, K. Yang, Y . Tong, and D. Tao, “Sfnet: Faster, accurate, and domain agnostic semantic segmentation via semantic flow,” IJCV, 2023. 2

  32. [40]

    Repvit: Revis- iting mobile cnn from vit perspective. arxiv 2023,

    A. Wang, H. Chen, Z. Lin, H. Pu, and G. Ding, “Repvit: Revis- iting mobile cnn from vit perspective. arxiv 2023,” arXiv preprint arXiv:2307.09283, 2023. 2, 3, 7, 8, 16, 17

  33. [41]

    Ghostnetv3: Ex- ploring the training strategies for compact models,

    Z. Liu, Z. Hao, K. Han, Y . Tang, and Y . Wang, “Ghostnetv3: Ex- ploring the training strategies for compact models,” arXiv preprint arXiv:2404.11202, 2024. 2, 7, 8, 16, 17

  34. [42]

    Mobilenetv4-universal models for the mobile ecosystem,

    D. Qin, C. Leichner, M. Delakis, M. Fornoni, S. Luo, F. Yang, W. Wang, C. Banbury, C. Ye, B. Akin et al., “Mobilenetv4-universal models for the mobile ecosystem,” arXiv preprint arXiv:2404.10518, 2024. 2, 7, 8, 16, 17

  35. [43]

    Training data-efficient image transformers & distillation through attention,

    H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in ICML, 2021. 3, 5, 7, 8, 16, 17

  36. [44]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR, 2016. 3, 8, 11

  37. [45]

    Visual attention methods in deep learning: An in-depth survey,

    M. Hassanin, S. Anwar, I. Radwan, F. S. Khan, and A. Mian, “Visual attention methods in deep learning: An in-depth survey,” arXiv preprint arXiv:2204.07756, 2022. 3

  38. [46]

    Recent advances in vision transformer: A survey and outlook of recent work,

    K. Islam, “Recent advances in vision transformer: A survey and outlook of recent work,” arXiv preprint arXiv:2203.01536, 2022. 3

  39. [47]

    Incorporating convolution designs into visual transformers,

    K. Yuan, S. Guo, Z. Liu, A. Zhou, F. Yu, and W. Wu, “Incorporating convolution designs into visual transformers,” in ICCV, 2021. 3

  40. [48]

    Conditional positional encodings for vision transformers,

    X. Chu, Z. Tian, B. Zhang, X. Wang, and C. Shen, “Conditional positional encodings for vision transformers,” in ICLR, 2023. 3

  41. [49]

    Uniformer: Unified transformer for efficient spatial-temporal representation learning,

    K. Li, Y . Wang, G. Peng, G. Song, Y . Liu, H. Li, and Y . Qiao, “Uniformer: Unified transformer for efficient spatial-temporal representation learning,” in ICLR, 2022. 3, 5, 6, 9

  42. [50]

    Moganet: Multi-order gated aggregation network,

    S. Li, Z. Wang, Z. Liu, C. Tan, H. Lin, D. Wu, Z. Chen, J. Zheng, and S. Z. Li, “Moganet: Multi-order gated aggregation network,” in ICLR,

  43. [51]

    Shvit: Single-head vision transformer with memory efficient macro design,

    S. Yun and Y . Ro, “Shvit: Single-head vision transformer with memory efficient macro design,” in CVPR, 2024. 3, 8

  44. [52]

    Metaformer is actually what you need for vision,

    W. Yu, M. Luo, P. Zhou, C. Si, Y . Zhou, X. Wang, J. Feng, and S. Yan, “Metaformer is actually what you need for vision,” in CVPR, 2022. 3, 4, 8, 9, 11

  45. [53]

    Lightvit: To- wards light-weight convolution-free vision transformers,

    T. Huang, L. Huang, S. You, F. Wang, C. Qian, and C. Xu, “Lightvit: To- wards light-weight convolution-free vision transformers,” arXiv preprint arXiv:2207.05557, 2022. 3, 8

  46. [54]

    Rest: An efficient transformer for visual recognition,

    Q. Zhang and Y .-B. Yang, “Rest: An efficient transformer for visual recognition,” in NeurIPS, 2021. 3

  47. [55]

    Edgevits: Competing light-weight cnns on mobile devices with vision transformers,

    J. Pan, A. Bulat, F. Tan, X. Zhu, L. Dudziak, H. Li, G. Tzimiropoulos, and B. Martinez, “Edgevits: Competing light-weight cnns on mobile devices with vision transformers,” in ECCV, 2022. 3, 7, 8

  48. [56]

    Res2net: A new multi-scale backbone architecture,

    S.-H. Gao, M.-M. Cheng, K. Zhao, X.-Y . Zhang, M.-H. Yang, and P. Torr, “Res2net: A new multi-scale backbone architecture,” IEEE TPAMI, 2019. 3, 10 IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 14

  49. [57]

    Xcit: Cross- covariance image transformers,

    A. Ali, H. Touvron, M. Caron, P. Bojanowski, M. Douze, A. Joulin, I. Laptev, N. Neverova, G. Synnaeve, J. Verbeek et al., “Xcit: Cross- covariance image transformers,” in NeurIPS, vol. 34, 2021. 3, 8, 10

  50. [58]

    Vig: Linear- complexity visual sequence learning with gated linear attention,

    B. Liao, X. Wang, L. Zhu, Q. Zhang, and C. Huang, “Vig: Linear- complexity visual sequence learning with gated linear attention,” arXiv preprint arXiv:2405.18425, 2024. 3, 8

  51. [59]

    Pointrwkv: Efficient rwkv-like model for hierarchical point cloud learning,

    Q. He, J. Zhang, J. Peng, H. He, X. Li, Y . Wang, and C. Wang, “Pointrwkv: Efficient rwkv-like model for hierarchical point cloud learning,” arXiv preprint arXiv:2405.15214, 2024. 3

  52. [60]

    Vision-rwkv: Efficient and scalable visual perception with rwkv-like architectures,

    Y . Duan, W. Wang, Z. Chen, X. Zhu, L. Lu, T. Lu, Y . Qiao, H. Li, J. Dai, and W. Wang, “Vision-rwkv: Efficient and scalable visual perception with rwkv-like architectures,” arXiv preprint arXiv:2403.02308, 2024. 3, 8

  53. [61]

    Mamba or rwkv: Exploring high-quality and high-efficiency segment anything model,

    H. Yuan, X. Li, L. Qi, T. Zhang, M.-H. Yang, S. Yan, and C. C. Loy, “Mamba or rwkv: Exploring high-quality and high-efficiency segment anything model,” arXiv preprint arXiv:2406.19369, 2024. 3

  54. [62]

    Mamba: Linear-time sequence modeling with selective state spaces,

    A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” arXiv preprint arXiv:2312.00752, 2023. 3

  55. [63]

    Rwkv: Reinventing rnns for the transformer era,

    B. Peng, E. Alcaide, Q. Anthony, A. Albalak, S. Arcadinho, S. Biderman, H. Cao, X. Cheng, M. Chung, M. Grella et al., “Rwkv: Reinventing rnns for the transformer era,” arXiv preprint arXiv:2305.13048, 2023. 3

  56. [64]

    Vision mamba: Efficient visual representation learning with bidirectional state space model,

    L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang, “Vision mamba: Efficient visual representation learning with bidirectional state space model,” arXiv preprint arXiv:2401.09417, 2024. 3, 7, 8, 16, 17

  57. [65]

    Efficientvmamba: Atrous selective scan for light weight visual mamba,

    X. Pei, T. Huang, and C. Xu, “Efficientvmamba: Atrous selective scan for light weight visual mamba,” arXiv preprint arXiv:2403.09977, 2024. 3, 8

  58. [66]

    Mobilemamba: Lightweight multi-receptive visual mamba network,

    H. He, J. Zhang, Y . Cai, H. Chen, X. Hu, Z. Gan, Y . Wang, C. Wang, Y . Wu, and L. Xie, “Mobilemamba: Lightweight multi-receptive visual mamba network,” arXiv preprint arXiv:2411.15941, 2024. 3

  59. [67]

    Scalable diffusion models with transformers,

    W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in ICCV, 2023. 4, 10

  60. [68]

    Focal attention for long-range interactions in vision transformers,

    J. Yang, C. Li, P. Zhang, X. Dai, B. Xiao, L. Yuan, and J. Gao, “Focal attention for long-range interactions in vision transformers,” in NeurIPS,

  61. [69]

    Cswin transformer: A general vision transformer backbone with cross-shaped windows,

    X. Dong, J. Bao, D. Chen, W. Zhang, N. Yu, L. Yuan, D. Chen, and B. Guo, “Cswin transformer: A general vision transformer backbone with cross-shaped windows,” in CVPR, 2022. 3

  62. [70]

    Inception transformer,

    C. Si, W. Yu, P. Zhou, Y . Zhou, X. Wang, and S. YAN, “Inception transformer,” in NeurIPS, 2022. 3

  63. [71]

    Pay attention to mlps,

    H. Liu, Z. Dai, D. So, and Q. V . Le, “Pay attention to mlps,” NeurIPS,

  64. [72]

    Mlp-mixer: An all-mlp architecture for vision,

    I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreit et al. , “Mlp-mixer: An all-mlp architecture for vision,” NeurIPS, 2021. 3

  65. [73]

    Resmlp: Feed- forward networks for image classification with data-efficient training,

    H. Touvron, P. Bojanowski, M. Caron, M. Cord, A. El-Nouby, E. Grave, G. Izacard, A. Joulin, G. Synnaeve, J. Verbeek et al., “Resmlp: Feed- forward networks for image classification with data-efficient training,” T-PAMI, 2022. 3

  66. [74]

    Shufflenet v2: Practical guidelines for efficient cnn architecture design,

    N. Ma, X. Zhang, H.-T. Zheng, and J. Sun, “Shufflenet v2: Practical guidelines for efficient cnn architecture design,” in ECCV, 2018. 4, 6

  67. [75]

    Moat: Alternating mobile convolution and attention brings strong vision models,

    C. Yang, S. Qiao, Q. Yu, X. Yuan, Y . Zhu, A. Yuille, H. Adam, and L.-C. Chen, “Moat: Alternating mobile convolution and attention brings strong vision models,” ICLR, 2023. 5, 6, 7, 8

  68. [76]

    Batch normalization: Accelerating deep network training by reducing internal covariate shift,

    S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML. PMLR,

  69. [77]

    Gaussian error linear units (gelus),

    D. Hendrycks and K. Gimpel, “Gaussian error linear units (gelus),” arXiv preprint arXiv:1606.08415, 2016. 6

  70. [78]

    Layer normalization,

    J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” arXiv preprint arXiv:1607.06450, 2016. 6

  71. [79]

    Imagenet: A large-scale hierarchical image database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR, 2009. 7, 10, 11, 16

  72. [80]

    Decoupled weight decay regularization,

    I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in ICLR, 2019. 7, 8

  73. [81]

    SGDR: Stochastic gradient descent with warm restarts,

    ——, “SGDR: Stochastic gradient descent with warm restarts,” in ICLR,

  74. [82]

    Rethinking the inception architecture for computer vision,

    C. Szegedy, V . Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in CVPR, 2016. 7

  75. [83]

    Deep networks with stochastic depth,

    G. Huang, Y . Sun, Z. Liu, D. Sedra, and K. Q. Weinberger, “Deep networks with stochastic depth,” in ECCV, 2016. 7

  76. [84]

    Randaugment: Practical automated data augmentation with a reduced search space,

    E. D. Cubuk, B. Zoph, J. Shlens, and Q. V . Le, “Randaugment: Practical automated data augmentation with a reduced search space,” in CVPRW,

  77. [85]

    Going deeper with image transformers,

    H. Touvron, M. Cord, A. Sablayrolles, G. Synnaeve, and H. Jégou, “Going deeper with image transformers,” in ICCV, 2021. 7

  78. [86]

    Dropout: a simple way to prevent neural networks from overfitting,

    N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdi- nov, “Dropout: a simple way to prevent neural networks from overfitting,” JMLR, 2014. 7

  79. [87]

    mixup: Beyond empirical risk minimization,

    H. Zhang, M. Cisse, Y . N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” in ICLR, 2018. 7

  80. [88]

    Cutmix: Regularization strategy to train strong classifiers with localizable features,

    S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, and Y . Yoo, “Cutmix: Regularization strategy to train strong classifiers with localizable features,” in ICCV, 2019. 7

  81. [89]

    Random erasing data augmentation,

    Z. Zhong, L. Zheng, G. Kang, S. Li, and Y . Yang, “Random erasing data augmentation,” in AAAI, 2020. 7

  82. [90]

    All tokens matter: Token labeling for training better vision transformers,

    Z.-H. Jiang, Q. Hou, L. Yuan, D. Zhou, Y . Shi, X. Jin, A. Wang, and J. Feng, “All tokens matter: Token labeling for training better vision transformers,” in NeurIPS, vol. 34, 2021. 7

  83. [91]

    Pytorch image models,

    R. Wightman, “Pytorch image models,” https://github.com/rwightman/ pytorch-image-models, 2019. 7

  84. [92]

    Tresnet: High performance gpu-dedicated architecture,

    T. Ridnik, H. Lawen, A. Noy, E. Ben Baruch, G. Sharir, and I. Friedman, “Tresnet: High performance gpu-dedicated architecture,” in CACV, 2021. 7, 11

  85. [93]

    Run, don’t walk: chasing higher flops for faster neural networks,

    J. Chen, S.-h. Kao, H. He, W. Zhuo, S. Wen, C.-H. Lee, and S.-H. G. Chan, “Run, don’t walk: chasing higher flops for faster neural networks,” in CVPR, 2023. 8

  86. [94]

    Mocovit: Mobile convolutional vision transformer,

    H. Ma, X. Xia, X. Wang, X. Xiao, J. Li, and M. Zheng, “Mocovit: Mobile convolutional vision transformer,” arXiv preprint arXiv:2205.12635 ,

  87. [95]

    Efficientvit: Memory efficient vision transformer with cascaded group attention,

    X. Liu, H. Peng, N. Zheng, Y . Yang, H. Hu, and Y . Yuan, “Efficientvit: Memory efficient vision transformer with cascaded group attention,” in CVPR, 2023. 8

  88. [96]

    Mpvit: Multi-path vision transformer for dense prediction,

    Y . Lee, J. Kim, J. Willette, and S. J. Hwang, “Mpvit: Multi-path vision transformer for dense prediction,” in CVPR, 2022. 8, 9

  89. [97]

    Multi-scale vmamba: Hierarchy in hierarchy visual state space model,

    Y . Shi, M. Dong, and C. Xu, “Multi-scale vmamba: Hierarchy in hierarchy visual state space model,” arXiv preprint arXiv:2405.14174,

  90. [98]

    Mambaout: Do we really need mamba for vision?

    W. Yu and X. Wang, “Mambaout: Do we really need mamba for vision?” arXiv preprint arXiv:2405.07992, 2024. 8

  91. [99]

    Microsoft coco: Common objects in context,

    T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV, 2014. 8, 9, 16, 17

  92. [100]

    Mask r-cnn,

    K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in ICCV,

  93. [101]

    MMDetection: Open mmlab detection toolbox and benchmark,

    K. Chen, J. Wang, J. Pang, Y . Cao, Y . Xiong, X. Li, S. Sun, W. Feng, Z. Liu, J. Xu, Z. Zhang, D. Cheng, C. Zhu, T. Cheng, Q. Zhao, B. Li, X. Lu, R. Zhu, Y . Wu, J. Dai, J. Wang, J. Shi, W. Ouyang, C. C. Loy, and D. Lin, “MMDetection: Open mmlab detection toolbox and benchmar...

  94. [102]

    Rethinking atrous convolution for semantic image segmentation,

    L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam, “Rethinking atrous convolution for semantic image segmentation,” arXiv preprint arXiv:1706.05587, 2017. 9, 11, 12, 16, 17

  95. [103]

    Panoptic feature pyramid networks,

    A. Kirillov, R. Girshick, K. He, and P. Dollár, “Panoptic feature pyramid networks,” in CVPR, 2019. 9, 16, 17

  96. [104]

    Segformer: Simple and efficient design for semantic segmentation with transformers,

    E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “Segformer: Simple and efficient design for semantic segmentation with transformers,” in NeurIPS, 2021. 9, 16, 17

  97. [105]

    Pyramid scene parsing network,

    H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in CVPR, 2017. 9, 16, 17

  98. [106]

    Semantic understanding of scenes through the ade20k dataset,

    B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso, and A. Torralba, “Semantic understanding of scenes through the ade20k dataset,” IJCV, 2019. 9, 16, 17

  99. [107]

    MMSegmentation: Openmmlab semantic seg- mentation toolbox and benchmark,

    M. Contributors, “MMSegmentation: Openmmlab semantic seg- mentation toolbox and benchmark,” https://github.com/open-mmlab/ mmsegmentation, 2020. 9

  100. [108]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in MICCAI, 2015. 9

  101. [109]

    Robust vessel segmentation in fundus images,

    A. Budai, R. Bock, A. Maier, J. Hornegger, and G. Michelson, “Robust vessel segmentation in fundus images,” IJBI, 2013. 9

  102. [110]

    The kinetics human action video dataset,

    W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijaya- narasimhan, F. Viola, T. Green, T. Back, P. Natsevet al., “The kinetics human action video dataset,” arXiv preprint arXiv:1705.06950, 2017. 9, 10

  103. [111]

    Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers,

    N. Ma, M. Goldstein, M. S. Albergo, N. M. Boffi, E. Vanden-Eijnden, and S. Xie, “Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers,” arXiv preprint arXiv:2401.08740,

  104. [112]

    Optimize your core ml usage,

    A. Inc., “Optimize your core ml usage,” https://developer.apple.com/ documentation/vision/classifying_images_with_vision_and_core_ml,

  105. [113]

    Deformable convnets v2: More deformable, better results,

    X. Zhu, H. Hu, S. Lin, and J. Dai, “Deformable convnets v2: More deformable, better results,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 9308–9316. 11

  106. [114]

    Resnet strikes back: An improved training procedure in timm,

    R. Wightman, H. Touvron, and H. Jégou, “Resnet strikes back: An improved training procedure in timm,” in NeurIPSW, 2021. 11

  107. [115]

    A convnet for the 2020s,

    Z. Liu, H. Mao, C.-Y . Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in CVPR, 2022. 11

  108. [116]

    Vitaev2: Vision transformer advanced by exploring inductive bias for image recognition and beyond,

    Q. Zhang, Y . Xu, J. Zhang, and D. Tao, “Vitaev2: Vision transformer advanced by exploring inductive bias for image recognition and beyond,” IJCV, 2023. 11 IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 16 APPENDIX OVERVIEW The supplementary material presents ...

  109. [2022]

    10 IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 15

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.