Pith. sign in

REVIEW 4 major objections 5 minor 58 references

Unified Local and Global Attention Interaction Modeling for Vision Transformers

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper claims that inserting Aggressive Convolutional Pooling and Conceptual Attention Transformation before multi-head self-attention improves object detection across several ViT-style backbones and datasets.

desk verdict Honest modules, overstated headline: the paper's own ablation shows the global component alone may be enough, and the parameter-parity setup isn't clean enough to support the unified claim. read the letter →

arxiv 2412.18778 v1 pith:XXPUVHHW submitted 2024-12-25 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords visiontransformersobjectdetectionself-attentionlocalandglobalinteractionmedicalimagingconcealed
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that vision transformer self-attention underperforms because visual tokens are matched to one another in isolation, without exchanging local or global information with neighboring features first. To fix this, the authors insert two modules before the attention computation: Aggressive Convolutional Pooling (ACP), which iteratively pools and convolves features to grow the effective receptive field, and Conceptual Attention Transformation (CAT), which projects features into a small set of semantic concepts and flows the resulting global information back into the tokens. Inserting these modules into ViT, Swin, and DAT++ backbones improves mean average precision and recall across five object-detection datasets, including tumor and concealed-object benchmarks. The claim matters because it suggests that pre-attention feature exchange is a generally useful enhancement for transformer detectors, not a fix tied to one architecture.

What carries the argument

The load-bearing machinery is a two-part pre-attention preprocessing stack. ACP applies a depthwise convolution plus residual connection (the LPU operation), then repeatedly downsamples with max pooling and upsamples and sums the multi-scale feature maps, so local convolutions compose into an approximate global receptive field. CAT generates a small number of semantic concept tokens via softmax attention pooling, then uses a backward-flow attention map, with a learned stochasticity term, to inject global concept information back into the input features. Both modules sit before multi-head self-attention and feed the refined features into the standard QKV computation.

What would settle it

A controlled comparison where the baseline backbones have the same embedding width and total parameter count as the enhanced versions on a single dataset such as CCellBio would settle whether the reported mAP gains come from the pre-attention modules or from increased model capacity.

Watch

Extended reading notes

Core claim

The paper's central claim is that giving visual tokens the ability to interact at both local and global scales before the self-attention computation makes attention more discriminative, so queries, keys, and values for objects from different semantic classes no longer collapse into nearly identical representations. The authors present this as the Enhanced Interaction Vision Transformer architecture, which they say shows substantial performance improvement for object detection over state-of-the-art transformer models across a broad range of self-attention module formations. The empirical support is a set of comparisons on CCellBio, COD10K-V2, Brain Tumor, NIH Chest X-Ray, and RSNA Pneumonia, where the enhanced backbones outperform their baselines on mAP and AR, with the largest relative gains on Swin and on concealed or medical objects. The paper also reports that the interaction modules alter attention behavior, reducing early-layer attention activity and producing sharper, more class-focused feature maps.

Load-bearing premise

The load-bearing premise is that the baseline and enhanced models are fairly comparable in capacity and training, so the measured gains come from the new modules rather than from differences in model size, width, or initialization.

Editorial extensions

If this is right

  • Inserting ACP and CAT before self-attention improves mAP and AR over the corresponding baselines on five datasets, with the largest relative gains on Swin (for example, +103.03% mAP on COD10K-V2).
  • The same modules improve detection across standard self-attention, shifted-window attention, and deformable attention backbones, so the benefit is not tied to one attention formulation.
  • With extended training, the Conceptual Attention Transformer alone can match or exceed the full EI-ViT, implying that ACP's main role is faster early convergence rather than final accuracy.
  • The improvements are obtained without pretrained weights, indicating that pre-attention interaction partially substitutes for large-scale pretraining in this experimental setting.
  • The modules can interfere with deformable-point learning in EI-DAT, producing small degradations in fine-grained metrics on some datasets, which points to a boundary condition for the approach.
  • The reported gains assume the baseline and enhanced models are fairly matched in capacity and training conditions, so the measured improvements come from the new modules rather than from differences in model size, width, or initialization.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: a strict matched-capacity comparison, with equal embedding widths and equal total parameter counts, would determine whether the observed gains survive when capacity differences are removed.
  • Beyond the paper: because the ablations show ACP helps early convergence and CAT helps at longer schedules, a training curriculum that starts with ACP and later switches to CAT-only might outperform either module used alone.
  • Beyond the paper: the same pre-attention interaction idea could be tested on decoders or cross-attention blocks, since the smoothing problem it addresses is not limited to encoder self-attention.
  • Beyond the paper: the mechanism might combine with post-attention refinement methods, since pre-attention feature exchange and post-attention map sharpening target different stages of the same over-smoothing failure.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes two modules inserted before multi-head self-attention in vision transformer backbones: Aggressive Convolutional Pooling (ACP), which iteratively applies depthwise convolution and pooling to build a global receptive field, and Conceptual Attention Transformation (CAT), which computes semantic concept tokens and uses a backward-flow attention term to inject global context. The modules are tested in ViT, Swin, and DAT++ backbones within RetinaNet on five detection datasets (CCellBio, COD10K-V2, Brain Tumor, NIH Chest XRay, RSNA Pneumonia), reporting consistent mAP/AR gains over the authors' own baselines. The paper also contributes a new medical dataset and qualitative analyses (PCA, CKA, attention maps). The central claim is that local and global feature exchange before self-attention substantially improves object detection across architectures.

Significance. If the central claim held, the work would offer a simple, architecture-agnostic plug-in for ViT-style detectors, with potential value for medical and concealed-object detection. The paper has several genuine strengths: it evaluates across diverse backbones and datasets, reports source-code and dataset release, and includes representational analyses (CKA, attention maps) that go beyond a single benchmark. The proposed CAT mechanism, in particular, is clearly specified and could be independently implemented. However, the significance is currently qualified by the paper's own ablation, which shows that CAT alone can outperform the full ACP+CAT model, and by parameter-count mismatches that confound the headline comparisons. The reported gains are against the authors' baselines, not against published state-of-the-art detectors, so the 'state-of-the-art' claim in Section 1 is not supported as written.

major comments (4)
  1. [Section 9 (Isolation Assessment)] The paper's central claim—that unified local (ACP) and global (CAT) interactions before self-attention are needed for the reported gains—is directly contradicted by the paper's own ablation. Section 9 states that ViT-CAT, after extra epochs of training, outperforms the full EI-ViT (ACP+CAT) on CCellBio across all metrics (+1.83% mAP, +0.13% mAP50, +4.55% mAP75, +0.20% AR), and the text concludes that 'the CAT component can learn more relevant features, removing the need for the ACP component.' This is an internal inconsistency with the title, abstract, and contribution list. The authors must either reframe the contribution as CAT-centric with ACP as a convergence accelerator, or provide evidence that under matched training budgets the combined ACP+CAT model is necessary. As it stands, Tables 3–7 do not establish that the unified architecture is the source of the improvement.
  2. [Section 7.2 and Figure 5 / Appendix Tables 8–13] The evaluation is confounded by inconsistent parameter parity between baseline and enhanced models. The text claims the baseline hidden widths were increased to approximate enhanced parameter counts, but Figure 5 shows, for 300x300 input, EI-ViT at 54.5M versus ViT at 71.6M, EI-DAT at 36.2M versus DAT at 23.5M, and EI-Swin at 247.1M versus Swin at 247.1M; at 512x512, EI-Swin is 277.1M versus Swin 247.1M. Appendix Tables 8–13 further show different embedding dimensions (e.g., EI-ViT starts at 48 versus ViT at 144; EI-Swin at 96 versus Swin at 288). Because model capacity and width differ, the reported gains cannot be cleanly attributed to the pre-attention interaction mechanism; they could come from changed capacity, optimization dynamics, or initialization. The authors should either match parameter counts and widths exactly, or explicitly report controlled experiments that vary capacity while isolating the modules.
  3. [Tables 3–7] No measure of variance is reported. All numbers appear to come from a single training run per configuration, with no seeds or error bars. This matters because several headline gains are small (e.g., EI-DAT on CCellBio: +1.92% mAP, −0.56% mAP75; EI-DAT on RSNA: −1.69% mAP) and could be within run-to-run noise. The authors should provide at least three seeds with mean and standard deviation, or otherwise justify that the improvements are statistically meaningful.
  4. [Section 1 and Section 8] The abstract and introduction claim 'substantial performance improvement for object detection over state-of-the-art transformer models,' but the experiments compare only against the authors' own reimplemented baselines trained from random initialization for 30 epochs. No comparison is made to published state-of-the-art detection results, to pretrained backbones, or to standard detection benchmarks such as COCO. The claim should be restricted to 'improvements over the authors' baselines under this training protocol,' or the authors must add comparisons to published SOTA numbers under consistent settings.
minor comments (5)
  1. [Throughout] The manuscript contains numerous typos and inconsistent dataset names: 'Universitiy' in the author affiliation, 'acro ss' in the abstract, 'manor' for 'manner', 'detoxification' for 'detection', and the dataset is called both COD10K-V2 and COD10K-V3 in different places. These should be corrected.
  2. [Figures 6 and 7 / Section 8.1] The figure captions appear to be swapped relative to the text: Figure 6 is captioned 'BBox Mean Average Precision (mAP)' while the text refers to Figure 7 for mAP, and Figure 7 is captioned 'BBox Average Recall (AR)' while the text refers to Figure 6 for AR. Please reconcile the figure numbering and in-text references.
  3. [Equation 9] The backward-flow term in Equation 9 writes Attn_mu = A·(Attn + α), but the matrix A is not defined in the text preceding the equation, and the dimensions of the multiplication are unclear. Define A explicitly and specify how it interacts with the stochasticity term α.
  4. [Section 7.1] The dataset sizes are reported inconsistently: the NIH Chest XRay description mentions '1,000 bounding box annotations' and the RSNA description says '7,644' without specifying the unit (images). Please unify the reporting of dataset splits and annotation counts.
  5. [Section 12 / Supporting Materials] The paper states 'We publish source code and a novel dataset' but no repository URL or dataset link is provided in the manuscript. Without a link, the reproducibility contribution cannot be verified.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the paper reports empirical comparisons on external benchmarks; the only overlapping-author citation is non-load-bearing, and the ablation inconsistency is a correctness concern, not a circular derivation.

full rationale

The paper's central claim is empirical: inserting ACP and CAT before multi-head self-attention improves object detection on ViT, Swin, and DAT backbones (Tables 3-7). The modules are defined by explicit equations (Eqs. 2-11) and are trained end-to-end, with no fitted parameter later renamed as a prediction and no quantity defined in terms of the target result. The improvement claim is validated against held-out test splits of external datasets, so it is not circular by construction. The only self-citation is reference [50] (Zhang, Heldermon, and Toler-Franklin), used as related work for cancer detection; it does not justify the central premise or forbid alternatives, so it is not load-bearing. The paper's internal Section 9 ablation does show that ViT with only CAT, after extra training, outperforms the full EI-ViT on CCellBio, which undermines the claim that both local and global modules are jointly necessary, but that is an internal-consistency or attribution weakness, not a circular derivation. Likewise, parameter-count mismatches between baselines and enhanced models (Fig. 5, Tables 8-13) raise fairness concerns but do not make the prediction equal to its input. Because the validation is self-contained against external benchmarks and no derivation step reduces to its own inputs, the circularity score is low; the one-point score reflects only the minor non-load-bearing self-citation and the use of the same dataset for hyperparameter exploration and headline reporting.

Assumptions & free parameters 4 free parameters · 4 assumptions · 2 invented entities

The empirical claim rests mainly on benchmark comparisons rather than derivation. The load-bearing premises are capacity matching between baselines and enhanced models, the reliability of single 30-epoch runs from random initialization, and the quality of the new CCellBio annotations. No mathematical derivation is performed, so no standard-math axioms are required beyond the usual softmax and convolution definitions.

free parameters (4)
  • ACP pooling layer count = 2
    Ablation in Figure 14 shows best mAP50, mAP75, and AR at two pooling layers; this setting is used in the main tables.
  • Concept count C = 48/96/192/384 for ViT; 64/128/256/512 for DAT
    Ablation in Figure 13 shows performance increasing up to 512 concepts; stage-wise counts are chosen by hand per backbone.
  • Backward-flow stochasticity alpha = learned, not reported
    Introduces variance into backward-flow attention (Equation 9); learned from features or positional bias with no reported final values.
  • Training epochs = 30
    All models are trained for exactly 30 epochs from random initialization, a protocol choice that affects every metric comparison.
assumptions (4)
  • domain assumption Baseline and enhanced backbones are capacity matched when parameter counts are approximated.
    Section 7.2 claims width scaling for fair comparison, but Figure 5 and Tables 8 through 13 show inconsistent parameter parity.
  • domain assumption A single 30-epoch run from random initialization is a reliable performance estimator.
    No seeds, repeats, or error bars are reported for any table, so measured differences may be within run-to-run noise.
  • domain assumption CCellBio ground-truth annotations are accurate and representative.
    The new dataset is introduced in Section 7.1 without an annotation protocol, reader study, or inter-annotator agreement.
  • domain assumption mAP and AR on the chosen benchmarks are the right evaluation criteria.
    No other detection metrics, external baselines, or clinical validation are used to support the central claim.
invented entities (2)
  • Semantic concept tokens
    purpose: Provide high-level global context to refine visual features before self-attention (Figure 2, Equations 6 through 10).
    No external validation; learned latent constructs inspired by [44], with the number of concepts chosen by ablation.
  • Backward-flow stochasticity term alpha
    purpose: Inject variance into contribution attention to avoid token smoothing (Equation 9).
    Learnable parameter with no reported values or external predictions; it is an internal model component.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Unified Local and Global Attention Interaction Modeling for Vision Transformers." pith.science (2026). https://pith.science/paper/XXPUVHHW

@misc{pith2026241218778,
  author       = {Pith},
  title        = {Pith review of: Unified Local and Global Attention Interaction Modeling for Vision Transformers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XXPUVHHW}},
  note         = {Machine review of arXiv:2412.18778}
}
read the original abstract

We present a novel method that extends the self-attention mechanism of a vision transformer (ViT) for more accurate object detection across diverse datasets. ViTs show strong capability for image understanding tasks such as object detection, segmentation, and classification. This is due in part to their ability to leverage global information from interactions among visual tokens. However, the self-attention mechanism in ViTs are limited because they do not allow visual tokens to exchange local or global information with neighboring features before computing global attention. This is problematic because tokens are treated in isolation when attending (matching) to other tokens, and valuable spatial relationships are overlooked. This isolation is further compounded by dot-product similarity operations that make tokens from different semantic classes appear visually similar. To address these limitations, we introduce two modifications to the traditional self-attention framework; a novel aggressive convolution pooling strategy for local feature mixing, and a new conceptual attention transformation to facilitate interaction and feature exchange between semantic concepts. Experimental results demonstrate that local and global information exchange among visual features before self-attention significantly improves performance on challenging object detection tasks and generalizes across multiple benchmark datasets and challenging medical datasets. We publish source code and a novel dataset of cancerous tumors (chimeric cell clusters).

Figures

Figures reproduced from arXiv: 2412.18778 by the authors.

Figure 1
Figure 1. Overview of the Enhanced Interaction Vision Transfo [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Conceptual A ention Transformation (CAT) Module: ( [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Global Concept Tokens: Input feature maps are proces [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Examples of each dataset are as follows: (a) CCellBio [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Number of parameters for baseline and enhanced inter [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 7
Figure 7. Figure 7: BBox Average Recall (AR) comparison between the base [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: Feature analysis comparison between ViT (top row) an [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]
Figure 9
Figure 9. Figure 9: A ention map analysis comparison between ViT (top ro [PITH_FULL_IMAGE:figures/full_fig_p011_9.png]
Figure 10
Figure 10. Figure 10: Linear CKA Similarity for the feature maps at differe [PITH_FULL_IMAGE:figures/full_fig_p012_10.png]
Figure 11
Figure 11. Figure 11: Kernel CKA Similarity for the feature maps at differe [PITH_FULL_IMAGE:figures/full_fig_p012_11.png]
Figure 12
Figure 12. Figure 12: Isolation Benchmark: ViT (No CAT) and ViT (No ACP) re [PITH_FULL_IMAGE:figures/full_fig_p013_12.png]
Figure 13
Figure 13. Figure 13: There is a consistent positive correlation between [PITH_FULL_IMAGE:figures/full_fig_p014_13.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

58 extracted references · 19 canonical work pages

  1. [1]

    Anonymous. 2024. SAM 2: Segment Anything in Images and Vi deos. In Sub- mitted to The Thirteenth International Conference on Learn ing Representations . https://openreview.net/forum?id=Ha6RTeWMd0 under review

  2. [2]

    Zhaowei Cai and Nuno Vasconcelos. 2018. Cascade R-CNN: D elving Into High Quality Object Detection. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6154–6162. https://doi.org/10.1109/CVPR.2018.00644

  3. [3]

    Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nic olas Usunier, Alexan- der Kirillov, and Sergey Zagoruyko. 2020. End-to-End Objec t Detec- tion with Transformers. CoRR abs/2005.12872 (2020). arXiv:2005.12872 https://arxiv.org/abs/2005.12872

  4. [4]

    Tianrun Chen, Ankang Lu, Lanyun Zhu, Chao Ding, Chunan Yu , Deyi Ji, Ze- jian Li, Lingyun Sun, Papa Mao, and Ying Zang. 2024. SAM2-Ada pter: Eval- uating & Adapting Segment Anything 2 in Downstream Tasks: Ca mouflage, Shadow, Medical Image Segmentation, and More. ArXiv abs/2408.04579 (2024). https://api.semanticscholar.org/CorpusID:271768828

  5. [5]

    Li, Lingyun Sun, Papa Mao, and Ying-Dong Zang

    Tianrun Chen, Lanyun Zhu, Chao Ding, Runlong Cao, Shangz han Zhang, Yan Wang, Z. Li, Lingyun Sun, Papa Mao, and Ying-Dong Zang. 2023. SAM Fails to Segment Anything? - SAM-Adapter: Adapting SAM in Un derper- formed Scenes: Camouflage, Shadow, and More. ArXiv abs/2304.09148 (2023). https://api.semanticscholar.org/CorpusID:258187610

  6. [6]

    Jun Cheng. 2017. brain tumor dataset. (4 2017). https://doi.org/10.6084/m9.figshare.1512427.v5

  7. [7]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov , Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthia s Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. CoRR abs/2010.11929 (2020). arXiv:2010.11929 https://arxiv. or...

  8. [8]

    Deng-Ping Fan, Ge-Peng Ji, Ming-Ming Cheng, and Ling Sha o. 2021. Con- cealed Object Detection. CoRR abs/2102.10274 (2021). arXiv:2102.10274 https://arxiv.org/abs/2102.10274

Show all 58 references
  1. [9]

    Deng-Ping Fan, Ge-Peng Ji, Guolei Sun, Ming-Ming Cheng, Jianbing Shen, and Ling Shao. 2020. Camouflaged Object Detection. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 2774–2784. https://doi.org/10.1109/CVPR42600.2020.00285

  2. [10]

    Girshick

    Ross B. Girshick. 2015. Fast R-CNN. CoRR abs/1504.08083 (2015). arXiv:1504.08083 http://arxiv.org/abs/1504.08083

  3. [11]

    Girshick, Jeff Donahue, Trevor Darrell, and Jite ndra Malik

    Ross B. Girshick, Jeff Donahue, Trevor Darrell, and Jite ndra Malik. 2013. Rich fea- ture hierarchies for accurate object detection and semanti c segmentation. CoRR abs/1311.2524 (2013). arXiv:1311.2524 http://arxiv.org /abs/1311.2524

  4. [12]

    Goldberger, Luis A

    Ary L. Goldberger, Luis A. N. Amaral, Leon Glass, Jeffrey M. Hausdorff, Plamen Ch. Ivanov, Roger G. Mark, Joseph E. Mietus, George B. Moody, Chung-Kang Peng, and H. Eugene Stanley. 2000. PhysioBank, P hys- ioToolkit, and PhysioNet: Components of a new research reso urce for com-...

  5. [13]

    Jianyuan Guo, Kai Han, Han Wu, Chang Xu, Yehui Tang, Chun jing Xu, and Yunhe Wang. 2021. CMT: Convolutional Neural Networks Me et Vision Transformers. CoRR abs/2107.06263 (2021). arXiv:2107.06263 https://arxiv.org/abs/2107.06263

  6. [14]

    Ali Hassani, Steven Walton, Jiachen Li, Shen Li, and Hum phrey Shi

  7. [15]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2 014. Spatial Pyra- mid Pooling in Deep Convolutional Networks for Visual Recog nition. CoRR abs/1406.4729 (2014). arXiv:1406.4729 http://arxiv.org /abs/1406.4729

  8. [16]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2 015. Deep Residual Learning for Image Recognition. CoRR abs/1512.03385 (2015). arXiv:1512.03385 http://arxiv.org/abs/1512.03385

  9. [17]

    Yuhan Kang, Qingpeng Li, Leyuan Fang, Jian Zhao, and Xue long Li

  10. [18]

    Berg, Wan- Yen Lo, Piotr DollÃąr, and Ross Girshick

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi M ao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C . Berg, Wan- Yen Lo, Piotr DollÃąr, and Ross Girshick. 2023. Segment Anyt hing. In 2023 IEEE/CVF International Conference on Computer Vision (IC...

  11. [19]

    Hin- ton

    Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Ge offrey E. Hin- ton. 2019. Similarity of Neural Network Representations Re visited. CoRR abs/1905.00414 (2019). arXiv:1905.00414 http://arxiv.o rg/abs/1905.00414

  12. [20]

    Hei Law and Jia Deng. 2018. CornerNet: Detecting Object s as Paired Keypoints. CoRR abs/1808.01244 (2018). arXiv:1808.01244 http://arxiv.o rg/abs/1808.01244

  13. [21]

    Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. 20 22. Exploring Plain Vi- sion Transformer Backbones for Object Detection. In Computer Vision âĂŞ ECCV 2022: 17th European Conference, Tel A viv, Israel, October 2 3âĂŞ27, 2022, Proceed- ings, Part IX (Tel Aviv, Israel). S...

  14. [22]

    Girshick, Kaiming H e, Bharath Hariharan, and Serge J

    Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming H e, Bharath Hariharan, and Serge J. Belongie. 2016. Feature Pyramid Networks for Objec t Detection. CoRR abs/1612.03144 (2016). arXiv:1612.03144 http://arxiv.o rg/abs/1612.03144

  15. [23]

    Girshick, Kaiming He , and Piotr Dollár

    Tsung-Yi Lin, Priya Goyal, Ross B. Girshick, Kaiming He , and Piotr Dollár

  16. [24]

    Jihao Liu, Hongsheng Li, Guanglu Song, Xin Huang, and Yu Liu

  17. [25]

    Reed, Cheng-Yang Fu, and Alexander C

    Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian S zegedy, Scott E. Reed, Cheng-Yang Fu, and Alexander C. Berg. 2015. SSD: Singl e Shot MultiBox Detector. CoRR abs/1512.02325 (2015). arXiv:1512.02325 http://arxiv.org/abs/1512.02325

  18. [26]

    Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yix uan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, Furu Wei, and Baining Guo. 2021 . Swin Trans- former V2: Scaling Up Capacity and Resolution. CoRR abs/2111.09883 (2021). arXiv:2111.09883 https://arxiv.org/abs/2111.09883

  19. [28]

    Chengqi Lyu, Wenwei Zhang, Haian Huang, Yue Zhou, Yudon g Wang, Yanyi Liu, Shilong Zhang, and Kai Chen. 2022. RTMDet: An Empirical Study of Designing Real-Time Object Detectors. ArXiv abs/2212.07784 (2022). https://api.semanticscholar.org/CorpusID:254685870

  20. [29]

    Muhammad Maaz, Abdelrahman Shaker, Hisham Cholakkal, Salman Khan, Syed Waqas Zamir, Rao Muhammad Anwer, and Fahad Shahbaz Khan

  21. [30]

    National Institutes of Health. 2017. NIH Chest X-ray Da taset. https://nihcc.app.box.com/v/ChestXray-NIHCC Accessed : 2024-11-29

  22. [31]

    Ha Quy Nguyen, Khanh Lam, Le Linh, Hieu Pham, Dat Tran, Du ng Nguyen, Dung Le, Chi Pham, Hang Tong, Diep Dinh, Cuong Do, Doan Luu, Cu ong Nguyen, BÃňnh NguyáżĚn, Que Nguyen, Au Hoang, Hien Phan, Anh Nguyen, Phuong Ho, and Van Vu. 2022. VinDr-CXR: An open datas et of chest X-ra...

  23. [32]

    Trung Pham, Mehran Maghoumi, Wanli Jiang, Bala Siva Sas hank Jujjavarapu, Mehdi Sajjadi, Xin Liu, Hsuan-Chu Lin, Bor-Jeng Chen, Giang Truong, Chao Fang, Junghyun Kwon, and Minwoo Park. 2023. NV AutoNet: Fast and Ac- curate 360Âř 3D Visual Perception For Self Driving. 2024 IEEE...

  24. [33]

    Radiological Society of North America. 2018. RSNA Pneu monia Detection Chal- lenge. https://www.rsna.org/education/ai-resources-a nd-training/ai-image-challenge/RSNA-Pneumonia- Accessed: 2024-11-29

  25. [34]

    In Computer Vision âĂŞ ECCV 2022 Workshops: Tel A viv, Israel, October 23âĂŞ27, 202 2, Proceedings, Part VII (Tel Aviv, Israel)

    EdgeNeXt: Efficiently Amalgamated CNN-Transformer Ar chi- tecture for Mobile Vision Applications. In Computer Vision âĂŞ ECCV 2022 Workshops: Tel A viv, Israel, October 23âĂŞ27, 202 2, Proceedings, Part VII (Tel Aviv, Israel). Springer-Verlag, Berlin, Heidelberg, 3âĂŞ20. ht...

  26. [35]

    Mike Ranzinger, Greg Heinrich, Jan Kautz, and Pavlo Mol chanov. 2024. AM- RADIO: Agglomerative Vision Foundation Model Reduce All Do mains Into One. 12490–12500. https://doi.org/10.1109/CVPR52733.2024. 01187

  27. [36]

    Girshick , and Ali Farhadi

    Joseph Redmon, Santosh Kumar Divvala, Ross B. Girshick , and Ali Farhadi. 2015. You Only Look Once: Unified, Real-Time Object Detection. CoRR abs/1506.02640 (2015). arXiv:1506.02640 http://arxiv.org/abs/1506.02 640

  28. [37]

    Girshick, and Jian Sun

    Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun . 2015. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Ne tworks. CoRR 16 • Nguyen, T. et al. abs/1506.01497 (2015). arXiv:1506.01497 http://arxiv.o rg/abs/1506.01497

  29. [38]

    Wu, Safwan S

    George Shih, Carol C. Wu, Safwan S. Halabi, Marc D. Kohli , Luciano M. Prevedello, Tessa S. Cook, Arjun Sharma, Judith K. Amorosa, Veronica Arteaga, Maya Galperin-Aizenberg, Ritu R. Gill, Myrna C. B. Godoy, St ephen Hobbs, Jean Jeudy, Archana Laroia, Palmi N. Shah, Dharshan Vu...

  30. [39]

    Girsh ick, Kaiming He, and Piotr Dollár

    Ilija Radosavovic, Raj Prateek Kosaraju, Ross B. Girsh ick, Kaiming He, and Piotr Dollár. 2020. Designing Network Design Spaces. CoRR abs/2003.13678 (2020). arXiv:2003.13678 https://arxiv.org/abs/2003.13678

  31. [40]

    Mingxing Tan and Quoc V. Le. 2019. EfficientNet: Rethinki ng Model Scaling for Convolutional Neural Networks. ArXiv abs/1905.11946 (2019). https://api.semanticscholar.org/CorpusID:167217261

  32. [41]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko reit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. A tten- tion Is All You Need. CoRR abs/1706.03762 (2017). arXiv:1706.03762 http://arxiv.org/abs/1706.03762

  33. [42]

    Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao S ong, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2021. Pyramid Vision Transf ormer: A Versa- tile Backbone for Dense Prediction without Convolutions. CoRR abs/2102.12122 (2021). arXiv:2102.12122 https://arxiv.org/abs/2...

  34. [43]

    Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, Mohammadh adi Bagheri, and Ronald M. Summers. 2017. ChestX-Ray8: Hospital-Scale Ches t X-Ray Database and Benchmarks on Weakly-Supervised Classification and Loc alization of Com- mon Thorax Diseases. In 2017 IEEE Conference on Compu...

  35. [44]

    Karen Simonyan and Andrew Zisserman. 2014. Very Deep Co nvolutional Networks for Large-Scale Image Recognition. CoRR abs/1409.1556 (2014). https://api.semanticscholar.org/CorpusID:14124313

  36. [45]

    Hongqiu Wu, Ruixue Ding, Hai Zhao, Pengjun Xie, Fei Huan g, and Min Zhang

  37. [46]

    Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyan g Dai, Lu Yuan, and Lei Zhang. 2021. CvT: Introducing Convolutions to Vision Tr ansformers. CoRR abs/2103.15808 (2021). arXiv:2103.15808 https://arxiv. org/abs/2103.15808

  38. [47]

    Zhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li, and Gao H uang. 2022. Vi- sion Transformer with Deformable Attention. CoRR abs/2201.00520 (2022). arXiv:2201.00520 https://arxiv.org/abs/2201.00520

  39. [48]

    Zhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li, and Gao H uang. 2023. DAT++: Spatially Dynamic Vision Transformer with Deformab le Attention. arXiv:cs.CV/2309.01430 https://arxiv.org/abs/2309.01 430

  40. [49]

    Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizh ao Zhang, Zhicheng Yan, Masayoshi Tomizuka, Joseph Gonzalez, Kurt Keutzer, and Peter Vajda. 2021. Visual Transformers: Where Do Transformers Really Belong i n Vision Models?. In 2021 IEEE/CVF International Conference on C...

  41. [50]

    Qingchao Zhang, Coy D Heldermon, and Corey Toler-Frank lin. 2020. Multiscale detection of cancerous tissue in high resolution slide scans. In International Sym- posium on Visual Computing . Springer, 139–153

  42. [51]

    Adversarial self-attention for language understanding. In Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence and Thirty-Fifth Confer- ence on Innovative Applications of Artificial Intelligenceand Thirteenth Symposium on Educational Advances in Artificial...

  43. [52]

    Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krä henbühl, and Ishan Misra. 2022. Detecting Twenty-thousand Classes using Imag e-level Supervision. CoRR abs/2201.02605 (2022). arXiv:2201.02605 https://arxiv. org/abs/2201.02605

  44. [53]

    Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, a nd Jifeng Dai. 2020. Deformable DETR: Deformable Transformers for End-to-End O bject Detection. CoRR abs/2010.04159 (2020). arXiv:2010.04159 https://arxiv. org/abs/2010.04159 Unified Local and Global A/t_tention Interac...

  45. [55]

    Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Z hu, Lionel Ni, and Heung-Yeung Shum. 2023. DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection. In The Eleventh International Conference on Learning Representations. https://openreview.net/f...

  46. [57]

    Daquan Zhou, Yujun Shi, Bingyi Kang, Weihao Yu, Zihang J iang, Li Yuan, Xi- aojie Jin, Qibin Hou, and Jiashi Feng. 2021. Refiner: Refining Self-attention for Vision Transformers. CoRR abs/2106.03714 (2021). arXiv:2106.03714 https://arxiv.org/abs/2106.03714

  47. [2017]

    CoRR abs/1708.02002 (2017)

    Focal Loss for Dense Object Detection. CoRR abs/1708.02002 (2017). arXiv:1708.02002 http://arxiv.org/abs/1708.02002

  48. [2021]

    In European Conference on Computer Vision

    UniNet: Unified Architecture Search with Convolution , Trans- former, and MLP. In European Conference on Computer Vision . https://api.semanticscholar.org/CorpusID:238531567

  49. [2023]

    In 2023 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR)

    Neighborhood Attention Transformer. In 2023 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR) . 6185–6194. https://doi.org/10.1109/CVPR52729.2023.00599

  50. [2024]

    arXiv:cs.CV/2410.06842 https://arxiv.org/abs/2410.06 842

    SurANet: Surrounding-Aware Network for Concealed Ob ject De- tection via Highly-Efficient Interactive Contrastive Learn ing Strategy. arXiv:cs.CV/2410.06842 https://arxiv.org/abs/2410.06 842

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.