REVIEW 4 major objections 5 minor 58 references
Unified Local and Global Attention Interaction Modeling for Vision Transformers
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper claims that inserting Aggressive Convolutional Pooling and Conceptual Attention Transformation before multi-head self-attention improves object detection across several ViT-style backbones and datasets.
desk verdict Honest modules, overstated headline: the paper's own ablation shows the global component alone may be enough, and the parameter-parity setup isn't clean enough to support the unified claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a two-part pre-attention preprocessing stack. ACP applies a depthwise convolution plus residual connection (the LPU operation), then repeatedly downsamples with max pooling and upsamples and sums the multi-scale feature maps, so local convolutions compose into an approximate global receptive field. CAT generates a small number of semantic concept tokens via softmax attention pooling, then uses a backward-flow attention map, with a learned stochasticity term, to inject global concept information back into the input features. Both modules sit before multi-head self-attention and feed the refined features into the standard QKV computation.
What would settle it
A controlled comparison where the baseline backbones have the same embedding width and total parameter count as the enhanced versions on a single dataset such as CCellBio would settle whether the reported mAP gains come from the pre-attention modules or from increased model capacity.
Extended reading notes
Core claim
The paper's central claim is that giving visual tokens the ability to interact at both local and global scales before the self-attention computation makes attention more discriminative, so queries, keys, and values for objects from different semantic classes no longer collapse into nearly identical representations. The authors present this as the Enhanced Interaction Vision Transformer architecture, which they say shows substantial performance improvement for object detection over state-of-the-art transformer models across a broad range of self-attention module formations. The empirical support is a set of comparisons on CCellBio, COD10K-V2, Brain Tumor, NIH Chest X-Ray, and RSNA Pneumonia, where the enhanced backbones outperform their baselines on mAP and AR, with the largest relative gains on Swin and on concealed or medical objects. The paper also reports that the interaction modules alter attention behavior, reducing early-layer attention activity and producing sharper, more class-focused feature maps.
Load-bearing premise
The load-bearing premise is that the baseline and enhanced models are fairly comparable in capacity and training, so the measured gains come from the new modules rather than from differences in model size, width, or initialization.
Editorial extensions
If this is right
- Inserting ACP and CAT before self-attention improves mAP and AR over the corresponding baselines on five datasets, with the largest relative gains on Swin (for example, +103.03% mAP on COD10K-V2).
- The same modules improve detection across standard self-attention, shifted-window attention, and deformable attention backbones, so the benefit is not tied to one attention formulation.
- With extended training, the Conceptual Attention Transformer alone can match or exceed the full EI-ViT, implying that ACP's main role is faster early convergence rather than final accuracy.
- The improvements are obtained without pretrained weights, indicating that pre-attention interaction partially substitutes for large-scale pretraining in this experimental setting.
- The modules can interfere with deformable-point learning in EI-DAT, producing small degradations in fine-grained metrics on some datasets, which points to a boundary condition for the approach.
- The reported gains assume the baseline and enhanced models are fairly matched in capacity and training conditions, so the measured improvements come from the new modules rather than from differences in model size, width, or initialization.
Reading between the lines
- Beyond the paper: a strict matched-capacity comparison, with equal embedding widths and equal total parameter counts, would determine whether the observed gains survive when capacity differences are removed.
- Beyond the paper: because the ablations show ACP helps early convergence and CAT helps at longer schedules, a training curriculum that starts with ACP and later switches to CAT-only might outperform either module used alone.
- Beyond the paper: the same pre-attention interaction idea could be tested on decoders or cross-attention blocks, since the smoothing problem it addresses is not limited to encoder self-attention.
- Beyond the paper: the mechanism might combine with post-attention refinement methods, since pre-attention feature exchange and post-attention map sharpening target different stages of the same over-smoothing failure.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes two modules inserted before multi-head self-attention in vision transformer backbones: Aggressive Convolutional Pooling (ACP), which iteratively applies depthwise convolution and pooling to build a global receptive field, and Conceptual Attention Transformation (CAT), which computes semantic concept tokens and uses a backward-flow attention term to inject global context. The modules are tested in ViT, Swin, and DAT++ backbones within RetinaNet on five detection datasets (CCellBio, COD10K-V2, Brain Tumor, NIH Chest XRay, RSNA Pneumonia), reporting consistent mAP/AR gains over the authors' own baselines. The paper also contributes a new medical dataset and qualitative analyses (PCA, CKA, attention maps). The central claim is that local and global feature exchange before self-attention substantially improves object detection across architectures.
Significance. If the central claim held, the work would offer a simple, architecture-agnostic plug-in for ViT-style detectors, with potential value for medical and concealed-object detection. The paper has several genuine strengths: it evaluates across diverse backbones and datasets, reports source-code and dataset release, and includes representational analyses (CKA, attention maps) that go beyond a single benchmark. The proposed CAT mechanism, in particular, is clearly specified and could be independently implemented. However, the significance is currently qualified by the paper's own ablation, which shows that CAT alone can outperform the full ACP+CAT model, and by parameter-count mismatches that confound the headline comparisons. The reported gains are against the authors' baselines, not against published state-of-the-art detectors, so the 'state-of-the-art' claim in Section 1 is not supported as written.
major comments (4)
- [Section 9 (Isolation Assessment)] The paper's central claim—that unified local (ACP) and global (CAT) interactions before self-attention are needed for the reported gains—is directly contradicted by the paper's own ablation. Section 9 states that ViT-CAT, after extra epochs of training, outperforms the full EI-ViT (ACP+CAT) on CCellBio across all metrics (+1.83% mAP, +0.13% mAP50, +4.55% mAP75, +0.20% AR), and the text concludes that 'the CAT component can learn more relevant features, removing the need for the ACP component.' This is an internal inconsistency with the title, abstract, and contribution list. The authors must either reframe the contribution as CAT-centric with ACP as a convergence accelerator, or provide evidence that under matched training budgets the combined ACP+CAT model is necessary. As it stands, Tables 3–7 do not establish that the unified architecture is the source of the improvement.
- [Section 7.2 and Figure 5 / Appendix Tables 8–13] The evaluation is confounded by inconsistent parameter parity between baseline and enhanced models. The text claims the baseline hidden widths were increased to approximate enhanced parameter counts, but Figure 5 shows, for 300x300 input, EI-ViT at 54.5M versus ViT at 71.6M, EI-DAT at 36.2M versus DAT at 23.5M, and EI-Swin at 247.1M versus Swin at 247.1M; at 512x512, EI-Swin is 277.1M versus Swin 247.1M. Appendix Tables 8–13 further show different embedding dimensions (e.g., EI-ViT starts at 48 versus ViT at 144; EI-Swin at 96 versus Swin at 288). Because model capacity and width differ, the reported gains cannot be cleanly attributed to the pre-attention interaction mechanism; they could come from changed capacity, optimization dynamics, or initialization. The authors should either match parameter counts and widths exactly, or explicitly report controlled experiments that vary capacity while isolating the modules.
- [Tables 3–7] No measure of variance is reported. All numbers appear to come from a single training run per configuration, with no seeds or error bars. This matters because several headline gains are small (e.g., EI-DAT on CCellBio: +1.92% mAP, −0.56% mAP75; EI-DAT on RSNA: −1.69% mAP) and could be within run-to-run noise. The authors should provide at least three seeds with mean and standard deviation, or otherwise justify that the improvements are statistically meaningful.
- [Section 1 and Section 8] The abstract and introduction claim 'substantial performance improvement for object detection over state-of-the-art transformer models,' but the experiments compare only against the authors' own reimplemented baselines trained from random initialization for 30 epochs. No comparison is made to published state-of-the-art detection results, to pretrained backbones, or to standard detection benchmarks such as COCO. The claim should be restricted to 'improvements over the authors' baselines under this training protocol,' or the authors must add comparisons to published SOTA numbers under consistent settings.
minor comments (5)
- [Throughout] The manuscript contains numerous typos and inconsistent dataset names: 'Universitiy' in the author affiliation, 'acro ss' in the abstract, 'manor' for 'manner', 'detoxification' for 'detection', and the dataset is called both COD10K-V2 and COD10K-V3 in different places. These should be corrected.
- [Figures 6 and 7 / Section 8.1] The figure captions appear to be swapped relative to the text: Figure 6 is captioned 'BBox Mean Average Precision (mAP)' while the text refers to Figure 7 for mAP, and Figure 7 is captioned 'BBox Average Recall (AR)' while the text refers to Figure 6 for AR. Please reconcile the figure numbering and in-text references.
- [Equation 9] The backward-flow term in Equation 9 writes Attn_mu = A·(Attn + α), but the matrix A is not defined in the text preceding the equation, and the dimensions of the multiplication are unclear. Define A explicitly and specify how it interacts with the stochasticity term α.
- [Section 7.1] The dataset sizes are reported inconsistently: the NIH Chest XRay description mentions '1,000 bounding box annotations' and the RSNA description says '7,644' without specifying the unit (images). Please unify the reporting of dataset splits and annotation counts.
- [Section 12 / Supporting Materials] The paper states 'We publish source code and a novel dataset' but no repository URL or dataset link is provided in the manuscript. Without a link, the reproducibility contribution cannot be verified.
Circularity Check
No significant circularity: the paper reports empirical comparisons on external benchmarks; the only overlapping-author citation is non-load-bearing, and the ablation inconsistency is a correctness concern, not a circular derivation.
full rationale
The paper's central claim is empirical: inserting ACP and CAT before multi-head self-attention improves object detection on ViT, Swin, and DAT backbones (Tables 3-7). The modules are defined by explicit equations (Eqs. 2-11) and are trained end-to-end, with no fitted parameter later renamed as a prediction and no quantity defined in terms of the target result. The improvement claim is validated against held-out test splits of external datasets, so it is not circular by construction. The only self-citation is reference [50] (Zhang, Heldermon, and Toler-Franklin), used as related work for cancer detection; it does not justify the central premise or forbid alternatives, so it is not load-bearing. The paper's internal Section 9 ablation does show that ViT with only CAT, after extra training, outperforms the full EI-ViT on CCellBio, which undermines the claim that both local and global modules are jointly necessary, but that is an internal-consistency or attribution weakness, not a circular derivation. Likewise, parameter-count mismatches between baselines and enhanced models (Fig. 5, Tables 8-13) raise fairness concerns but do not make the prediction equal to its input. Because the validation is self-contained against external benchmarks and no derivation step reduces to its own inputs, the circularity score is low; the one-point score reflects only the minor non-load-bearing self-citation and the use of the same dataset for hyperparameter exploration and headline reporting.
Assumptions & free parameters
free parameters (4)
- ACP pooling layer count =
2
- Concept count C =
48/96/192/384 for ViT; 64/128/256/512 for DAT
- Backward-flow stochasticity alpha =
learned, not reported
- Training epochs =
30
assumptions (4)
- domain assumption Baseline and enhanced backbones are capacity matched when parameter counts are approximated.
- domain assumption A single 30-epoch run from random initialization is a reliable performance estimator.
- domain assumption CCellBio ground-truth annotations are accurate and representative.
- domain assumption mAP and AR on the chosen benchmarks are the right evaluation criteria.
invented entities (2)
-
Semantic concept tokens
-
Backward-flow stochasticity term alpha
Cite this review
Pith. "Pith review of Unified Local and Global Attention Interaction Modeling for Vision Transformers." pith.science (2026). https://pith.science/paper/XXPUVHHW
@misc{pith2026241218778,
author = {Pith},
title = {Pith review of: Unified Local and Global Attention Interaction Modeling for Vision Transformers},
year = {2026},
howpublished = {\url{https://pith.science/paper/XXPUVHHW}},
note = {Machine review of arXiv:2412.18778}
}
read the original abstract
We present a novel method that extends the self-attention mechanism of a vision transformer (ViT) for more accurate object detection across diverse datasets. ViTs show strong capability for image understanding tasks such as object detection, segmentation, and classification. This is due in part to their ability to leverage global information from interactions among visual tokens. However, the self-attention mechanism in ViTs are limited because they do not allow visual tokens to exchange local or global information with neighboring features before computing global attention. This is problematic because tokens are treated in isolation when attending (matching) to other tokens, and valuable spatial relationships are overlooked. This isolation is further compounded by dot-product similarity operations that make tokens from different semantic classes appear visually similar. To address these limitations, we introduce two modifications to the traditional self-attention framework; a novel aggressive convolution pooling strategy for local feature mixing, and a new conceptual attention transformation to facilitate interaction and feature exchange between semantic concepts. Experimental results demonstrate that local and global information exchange among visual features before self-attention significantly improves performance on challenging object detection tasks and generalizes across multiple benchmark datasets and challenging medical datasets. We publish source code and a novel dataset of cancerous tumors (chimeric cell clusters).
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
Anonymous. 2024. SAM 2: Segment Anything in Images and Vi deos. In Sub- mitted to The Thirteenth International Conference on Learn ing Representations . https://openreview.net/forum?id=Ha6RTeWMd0 under review
work page 2024
-
[2]
Zhaowei Cai and Nuno Vasconcelos. 2018. Cascade R-CNN: D elving Into High Quality Object Detection. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6154–6162. https://doi.org/10.1109/CVPR.2018.00644
arXiv 2018
-
[3]
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nic olas Usunier, Alexan- der Kirillov, and Sergey Zagoruyko. 2020. End-to-End Objec t Detec- tion with Transformers. CoRR abs/2005.12872 (2020). arXiv:2005.12872 https://arxiv.org/abs/2005.12872
arXiv 2020
-
[4]
Tianrun Chen, Ankang Lu, Lanyun Zhu, Chao Ding, Chunan Yu , Deyi Ji, Ze- jian Li, Lingyun Sun, Papa Mao, and Ying Zang. 2024. SAM2-Ada pter: Eval- uating & Adapting Segment Anything 2 in Downstream Tasks: Ca mouflage, Shadow, Medical Image Segmentation, and More. ArXiv abs/2408.04579 (2024). https://api.semanticscholar.org/CorpusID:271768828
arXiv 2024
-
[5]
Li, Lingyun Sun, Papa Mao, and Ying-Dong Zang
Tianrun Chen, Lanyun Zhu, Chao Ding, Runlong Cao, Shangz han Zhang, Yan Wang, Z. Li, Lingyun Sun, Papa Mao, and Ying-Dong Zang. 2023. SAM Fails to Segment Anything? - SAM-Adapter: Adapting SAM in Un derper- formed Scenes: Camouflage, Shadow, and More. ArXiv abs/2304.09148 (2023). https://api.semanticscholar.org/CorpusID:258187610
arXiv 2023
-
[6]
Jun Cheng. 2017. brain tumor dataset. (4 2017). https://doi.org/10.6084/m9.figshare.1512427.v5
doi:10.6084/m9 2017
-
[7]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov , Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthia s Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. CoRR abs/2010.11929 (2020). arXiv:2010.11929 https://arxiv. or...
arXiv 2020
-
[8]
Deng-Ping Fan, Ge-Peng Ji, Ming-Ming Cheng, and Ling Sha o. 2021. Con- cealed Object Detection. CoRR abs/2102.10274 (2021). arXiv:2102.10274 https://arxiv.org/abs/2102.10274
work page Pith review arXiv 2021
Show all 58 references
-
[9]
Deng-Ping Fan, Ge-Peng Ji, Guolei Sun, Ming-Ming Cheng, Jianbing Shen, and Ling Shao. 2020. Camouflaged Object Detection. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 2774–2784. https://doi.org/10.1109/CVPR42600.2020.00285
2020
-
[10]
Girshick
Ross B. Girshick. 2015. Fast R-CNN. CoRR abs/1504.08083 (2015). arXiv:1504.08083 http://arxiv.org/abs/1504.08083
2015 arXiv
-
[11]
Girshick, Jeff Donahue, Trevor Darrell, and Jite ndra Malik
Ross B. Girshick, Jeff Donahue, Trevor Darrell, and Jite ndra Malik. 2013. Rich fea- ture hierarchies for accurate object detection and semanti c segmentation. CoRR abs/1311.2524 (2013). arXiv:1311.2524 http://arxiv.org /abs/1311.2524
2013 arXiv
-
[12]
Goldberger, Luis A
Ary L. Goldberger, Luis A. N. Amaral, Leon Glass, Jeffrey M. Hausdorff, Plamen Ch. Ivanov, Roger G. Mark, Joseph E. Mietus, George B. Moody, Chung-Kang Peng, and H. Eugene Stanley. 2000. PhysioBank, P hys- ioToolkit, and PhysioNet: Components of a new research reso urce for com-...
2000 doi
-
[13]
Jianyuan Guo, Kai Han, Han Wu, Chang Xu, Yehui Tang, Chun jing Xu, and Yunhe Wang. 2021. CMT: Convolutional Neural Networks Me et Vision Transformers. CoRR abs/2107.06263 (2021). arXiv:2107.06263 https://arxiv.org/abs/2107.06263
2021 arXiv
-
[14]
Ali Hassani, Steven Walton, Jiachen Li, Shen Li, and Hum phrey Shi
-
[15]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2 014. Spatial Pyra- mid Pooling in Deep Convolutional Networks for Visual Recog nition. CoRR abs/1406.4729 (2014). arXiv:1406.4729 http://arxiv.org /abs/1406.4729
2014 arXiv
-
[16]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2 015. Deep Residual Learning for Image Recognition. CoRR abs/1512.03385 (2015). arXiv:1512.03385 http://arxiv.org/abs/1512.03385
2015 arXiv
-
[17]
Yuhan Kang, Qingpeng Li, Leyuan Fang, Jian Zhao, and Xue long Li
-
[18]
Berg, Wan- Yen Lo, Piotr DollÃąr, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi M ao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C . Berg, Wan- Yen Lo, Piotr DollÃąr, and Ross Girshick. 2023. Segment Anyt hing. In 2023 IEEE/CVF International Conference on Computer Vision (IC...
2023
-
[19]
Hin- ton
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Ge offrey E. Hin- ton. 2019. Similarity of Neural Network Representations Re visited. CoRR abs/1905.00414 (2019). arXiv:1905.00414 http://arxiv.o rg/abs/1905.00414
2019 arXiv
-
[20]
Hei Law and Jia Deng. 2018. CornerNet: Detecting Object s as Paired Keypoints. CoRR abs/1808.01244 (2018). arXiv:1808.01244 http://arxiv.o rg/abs/1808.01244
2018 arXiv
-
[21]
Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. 20 22. Exploring Plain Vi- sion Transformer Backbones for Object Detection. In Computer Vision âĂŞ ECCV 2022: 17th European Conference, Tel A viv, Israel, October 2 3âĂŞ27, 2022, Proceed- ings, Part IX (Tel Aviv, Israel). S...
2022 doi
-
[22]
Girshick, Kaiming H e, Bharath Hariharan, and Serge J
Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming H e, Bharath Hariharan, and Serge J. Belongie. 2016. Feature Pyramid Networks for Objec t Detection. CoRR abs/1612.03144 (2016). arXiv:1612.03144 http://arxiv.o rg/abs/1612.03144
2016 arXiv
-
[23]
Girshick, Kaiming He , and Piotr Dollár
Tsung-Yi Lin, Priya Goyal, Ross B. Girshick, Kaiming He , and Piotr Dollár
-
[24]
Jihao Liu, Hongsheng Li, Guanglu Song, Xin Huang, and Yu Liu
-
[25]
Reed, Cheng-Yang Fu, and Alexander C
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian S zegedy, Scott E. Reed, Cheng-Yang Fu, and Alexander C. Berg. 2015. SSD: Singl e Shot MultiBox Detector. CoRR abs/1512.02325 (2015). arXiv:1512.02325 http://arxiv.org/abs/1512.02325
2015 arXiv
-
[26]
Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yix uan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, Furu Wei, and Baining Guo. 2021 . Swin Trans- former V2: Scaling Up Capacity and Resolution. CoRR abs/2111.09883 (2021). arXiv:2111.09883 https://arxiv.org/abs/2111.09883
2021 arXiv
-
[28]
Chengqi Lyu, Wenwei Zhang, Haian Huang, Yue Zhou, Yudon g Wang, Yanyi Liu, Shilong Zhang, and Kai Chen. 2022. RTMDet: An Empirical Study of Designing Real-Time Object Detectors. ArXiv abs/2212.07784 (2022). https://api.semanticscholar.org/CorpusID:254685870
2022 arXiv
-
[29]
Muhammad Maaz, Abdelrahman Shaker, Hisham Cholakkal, Salman Khan, Syed Waqas Zamir, Rao Muhammad Anwer, and Fahad Shahbaz Khan
-
[30]
National Institutes of Health. 2017. NIH Chest X-ray Da taset. https://nihcc.app.box.com/v/ChestXray-NIHCC Accessed : 2024-11-29
2017
-
[31]
Ha Quy Nguyen, Khanh Lam, Le Linh, Hieu Pham, Dat Tran, Du ng Nguyen, Dung Le, Chi Pham, Hang Tong, Diep Dinh, Cuong Do, Doan Luu, Cu ong Nguyen, BÃňnh NguyáżĚn, Que Nguyen, Au Hoang, Hien Phan, Anh Nguyen, Phuong Ho, and Van Vu. 2022. VinDr-CXR: An open datas et of chest X-ra...
2022 doi
-
[32]
Trung Pham, Mehran Maghoumi, Wanli Jiang, Bala Siva Sas hank Jujjavarapu, Mehdi Sajjadi, Xin Liu, Hsuan-Chu Lin, Bor-Jeng Chen, Giang Truong, Chao Fang, Junghyun Kwon, and Minwoo Park. 2023. NV AutoNet: Fast and Ac- curate 360Âř 3D Visual Perception For Self Driving. 2024 IEEE...
2023
-
[33]
Radiological Society of North America. 2018. RSNA Pneu monia Detection Chal- lenge. https://www.rsna.org/education/ai-resources-a nd-training/ai-image-challenge/RSNA-Pneumonia- Accessed: 2024-11-29
2018
-
[34]
In Computer Vision âĂŞ ECCV 2022 Workshops: Tel A viv, Israel, October 23âĂŞ27, 202 2, Proceedings, Part VII (Tel Aviv, Israel)
EdgeNeXt: Efficiently Amalgamated CNN-Transformer Ar chi- tecture for Mobile Vision Applications. In Computer Vision âĂŞ ECCV 2022 Workshops: Tel A viv, Israel, October 23âĂŞ27, 202 2, Proceedings, Part VII (Tel Aviv, Israel). Springer-Verlag, Berlin, Heidelberg, 3âĂŞ20. ht...
2022 doi
-
[35]
Mike Ranzinger, Greg Heinrich, Jan Kautz, and Pavlo Mol chanov. 2024. AM- RADIO: Agglomerative Vision Foundation Model Reduce All Do mains Into One. 12490–12500. https://doi.org/10.1109/CVPR52733.2024. 01187
2024
-
[36]
Girshick , and Ali Farhadi
Joseph Redmon, Santosh Kumar Divvala, Ross B. Girshick , and Ali Farhadi. 2015. You Only Look Once: Unified, Real-Time Object Detection. CoRR abs/1506.02640 (2015). arXiv:1506.02640 http://arxiv.org/abs/1506.02 640
2015 arXiv
-
[37]
Girshick, and Jian Sun
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun . 2015. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Ne tworks. CoRR 16 • Nguyen, T. et al. abs/1506.01497 (2015). arXiv:1506.01497 http://arxiv.o rg/abs/1506.01497
2015 arXiv
-
[38]
Wu, Safwan S
George Shih, Carol C. Wu, Safwan S. Halabi, Marc D. Kohli , Luciano M. Prevedello, Tessa S. Cook, Arjun Sharma, Judith K. Amorosa, Veronica Arteaga, Maya Galperin-Aizenberg, Ritu R. Gill, Myrna C. B. Godoy, St ephen Hobbs, Jean Jeudy, Archana Laroia, Palmi N. Shah, Dharshan Vu...
2019 doi
-
[39]
Girsh ick, Kaiming He, and Piotr Dollár
Ilija Radosavovic, Raj Prateek Kosaraju, Ross B. Girsh ick, Kaiming He, and Piotr Dollár. 2020. Designing Network Design Spaces. CoRR abs/2003.13678 (2020). arXiv:2003.13678 https://arxiv.org/abs/2003.13678
2020 arXiv
-
[40]
Mingxing Tan and Quoc V. Le. 2019. EfficientNet: Rethinki ng Model Scaling for Convolutional Neural Networks. ArXiv abs/1905.11946 (2019). https://api.semanticscholar.org/CorpusID:167217261
2019 arXiv
-
[41]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko reit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. A tten- tion Is All You Need. CoRR abs/1706.03762 (2017). arXiv:1706.03762 http://arxiv.org/abs/1706.03762
2017 arXiv
-
[42]
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao S ong, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2021. Pyramid Vision Transf ormer: A Versa- tile Backbone for Dense Prediction without Convolutions. CoRR abs/2102.12122 (2021). arXiv:2102.12122 https://arxiv.org/abs/2...
2021 arXiv
-
[43]
Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, Mohammadh adi Bagheri, and Ronald M. Summers. 2017. ChestX-Ray8: Hospital-Scale Ches t X-Ray Database and Benchmarks on Weakly-Supervised Classification and Loc alization of Com- mon Thorax Diseases. In 2017 IEEE Conference on Compu...
2017 doi
-
[44]
Karen Simonyan and Andrew Zisserman. 2014. Very Deep Co nvolutional Networks for Large-Scale Image Recognition. CoRR abs/1409.1556 (2014). https://api.semanticscholar.org/CorpusID:14124313
2014 arXiv
-
[45]
Hongqiu Wu, Ruixue Ding, Hai Zhao, Pengjun Xie, Fei Huan g, and Min Zhang
-
[46]
Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyan g Dai, Lu Yuan, and Lei Zhang. 2021. CvT: Introducing Convolutions to Vision Tr ansformers. CoRR abs/2103.15808 (2021). arXiv:2103.15808 https://arxiv. org/abs/2103.15808
2021 arXiv
-
[47]
Zhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li, and Gao H uang. 2022. Vi- sion Transformer with Deformable Attention. CoRR abs/2201.00520 (2022). arXiv:2201.00520 https://arxiv.org/abs/2201.00520
2022 arXiv
-
[48]
Zhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li, and Gao H uang. 2023. DAT++: Spatially Dynamic Vision Transformer with Deformab le Attention. arXiv:cs.CV/2309.01430 https://arxiv.org/abs/2309.01 430
2023 arXiv
-
[49]
Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizh ao Zhang, Zhicheng Yan, Masayoshi Tomizuka, Joseph Gonzalez, Kurt Keutzer, and Peter Vajda. 2021. Visual Transformers: Where Do Transformers Really Belong i n Vision Models?. In 2021 IEEE/CVF International Conference on C...
2021
-
[50]
Qingchao Zhang, Coy D Heldermon, and Corey Toler-Frank lin. 2020. Multiscale detection of cancerous tissue in high resolution slide scans. In International Sym- posium on Visual Computing . Springer, 139–153
2020
-
[51]
Adversarial self-attention for language understanding. In Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence and Thirty-Fifth Confer- ence on Innovative Applications of Artificial Intelligenceand Thirteenth Symposium on Educational Advances in Artificial...
-
[52]
Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krä henbühl, and Ishan Misra. 2022. Detecting Twenty-thousand Classes using Imag e-level Supervision. CoRR abs/2201.02605 (2022). arXiv:2201.02605 https://arxiv. org/abs/2201.02605
2022 arXiv
-
[53]
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, a nd Jifeng Dai. 2020. Deformable DETR: Deformable Transformers for End-to-End O bject Detection. CoRR abs/2010.04159 (2020). arXiv:2010.04159 https://arxiv. org/abs/2010.04159 Unified Local and Global A/t_tention Interac...
2020 arXiv
-
[55]
Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Z hu, Lionel Ni, and Heung-Yeung Shum. 2023. DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection. In The Eleventh International Conference on Learning Representations. https://openreview.net/f...
2023
-
[57]
Daquan Zhou, Yujun Shi, Bingyi Kang, Weihao Yu, Zihang J iang, Li Yuan, Xi- aojie Jin, Qibin Hou, and Jiashi Feng. 2021. Refiner: Refining Self-attention for Vision Transformers. CoRR abs/2106.03714 (2021). arXiv:2106.03714 https://arxiv.org/abs/2106.03714
2021 arXiv
-
[2017]
CoRR abs/1708.02002 (2017)
Focal Loss for Dense Object Detection. CoRR abs/1708.02002 (2017). arXiv:1708.02002 http://arxiv.org/abs/1708.02002
2017 arXiv
-
[2021]
In European Conference on Computer Vision
UniNet: Unified Architecture Search with Convolution , Trans- former, and MLP. In European Conference on Computer Vision . https://api.semanticscholar.org/CorpusID:238531567
-
[2023]
In 2023 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR)
Neighborhood Attention Transformer. In 2023 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR) . 6185–6194. https://doi.org/10.1109/CVPR52729.2023.00599
2023
-
[2024]
arXiv:cs.CV/2410.06842 https://arxiv.org/abs/2410.06 842
SurANet: Surrounding-Aware Network for Concealed Ob ject De- tection via Highly-Efficient Interactive Contrastive Learn ing Strategy. arXiv:cs.CV/2410.06842 https://arxiv.org/abs/2410.06 842
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.