REVIEW 2 major objections 6 minor 251 references
Grad-CAM for Vision Transformers: A Systematic Taxonomy and Audit of Methodological Ambiguity in Explainable AI
T0 review · 2 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read The paper claims that most applications of Grad-CAM to Vision Transformers leave the adaptation mathematically unspecified, so the published heatmaps cannot be reproduced or evaluated as scientific evidence.
desk verdict A genuinely useful taxonomy of ViT Grad-CAM adaptations, attached to an audit whose headline number (58% underspecification) is built on citation patterns rather than the reporting-detail coding the authors say they performed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a descriptive taxonomy of six feature tensors inside a transformer block: the layer-normalized tokens before attention ($F_1$), the softmax attention matrix ($F_2$), head-wise attention outputs ($F_3$), the layer-normalized tokens before the MLP ($F_4$), the MLP output ($F_5$), and the post-feedforward residual output ($F_6$), plus cross-attention variants for vision-language models. Against this grid the paper defines Flat-GCAM, a one-dimensional token-level attribution vector that can be reshaped into the patch grid, and shows that each combination of layer, feature tensor, gradient target, and head aggregation produces a different heatmap. This taxonomy is what lets the audit detect that most papers do not pin down a single member of the family.
What would settle it
Re-read the methods sections of the 101 papers categorized as citing only non-ViT Grad-CAM and count how many actually specify the feature tensor, gradient target, token handling, head aggregation, and reshaping in equations or prose; if a substantial share do, the 58% underspecification figure overstates the gap.
Extended reading notes
Core claim
The paper's central claim is that "Grad-CAM on a ViT" is not one defined procedure but a family of procedures indexed by the choice of feature tensor (layer-normalized tokens, softmax attention, attention outputs, MLP activations, residual outputs), the gradient target, whether [CLS]-token gradients or patch-token-average gradients are used, how attention heads are combined, and how token-level scores are reshaped into an image grid. The authors show that for a single ViT-B/16 model and one image this family contains well over 100 plausible heatmaps, and that different choices produce qualitatively different localizations. They then audit 175 papers and find that 101 cite only the original CNN-based Grad-CAM or CNN-era variants, 17 cite only an implementation repository, 13 cite nothing, and only 26 justify the adaptation directly or inherit a ViT-specific method. The conclusion is that most published ViT Grad-CAM figures are underspecified and therefore cannot serve as reproducible scientific evidence.
Load-bearing premise
The audit's headline numbers assume that a paper's citation pattern reveals whether its text specifies the ViT adaptation, because the reported coding of "reporting detail" is not presented in the paper itself.
Editorial extensions
If this is right
- If the paper is right, the majority of surveyed ViT Grad-CAM heatmaps cannot be reconstructed from the publication, so reviewers should treat "Grad-CAM on a ViT" as an underspecified phrase unless the adaptation is described or clearly inherited.
- The six feature families imply that two papers using the same Grad-CAM label may be explaining different tensors, so cross-paper comparison of heatmaps is unsafe without specifying the variant.
- Because final-layer token-weighted maps can become nearly blank, authors who omit the layer and weighting strategy could select the most visually favorable map from a large space of alternatives.
- In vision-language models, cross-attention maps are text-conditioned per-query-token maps, so interpreting them as generic visual saliency is not justified under the taxonomy.
- The 13 identified ViT-specific adaptations are unevenly adopted, which means the methodological ambiguity is current rather than historical.
Reading between the lines
- Beyond the paper, the same ambiguity likely affects other CNN-born explanation methods when ported to transformers, because any method that assumes a spatial-channel tensor inherits the same implementation space.
- The taxonomy could be used as a minimal reporting checklist in venue guidelines: requiring authors to name the exact tensor, the gradient target, and the aggregation axis would make many heatmaps reproducible without changing any method.
- The qualitative examples suggest that comparative claims about ViT models can be contaminated by Grad-CAM implementation choice, making model differences and attribution-variant differences hard to separate.
- A testable extension would be to take a sample of the 101 papers categorized as underspecified, implement the most plausible reading of each paper's text, and compare the resulting heatmap to the published figure.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a descriptive taxonomy of Grad-CAM adaptations to vision transformers, distinguishing six feature extraction locations (F1-F6), gradient targets, token handling, aggregation strategies, and reshaping choices, and extends this to cross-attention in vision-language models. It then audits 175 papers that apply Grad-CAM or related methods to ViTs, reporting that 101 papers (58%) cite only CNN-based Grad-CAM and that only 26 papers either justify the adaptation or cite ViT-specific prior work. The authors conclude that most papers treat Grad-CAM on ViTs as a trivial extension and fail to provide a full mathematical or implementation-level account, and they propose reporting recommendations for authors and reviewers.
Significance. The taxonomy and the qualitative demonstration that different feature locations, gradient targets, and aggregation choices yield substantially different heatmaps (Figs. 2, 3, and S1) are useful contributions to the explainable AI literature. The paper is careful to frame the taxonomy as descriptive rather than prescriptive, and it provides a detailed supplementary analysis of 13 ViT-specific Grad-CAM adaptations (S-V). The main quantitative claim, however, rests on a citation-pattern proxy rather than on the 'reporting detail' variable that the authors state they coded, so the headline 58% figure is not yet supported by the presented data. If the authors present the reporting-detail coding and align their claims with what that coding shows, the paper would provide a valuable evidence base for reproducibility standards in transformer interpretability research.
major comments (2)
- [III-B and IV-A] The paper states in Section III-B that all 175 papers were coded for 'reporting detail,' but this variable is never presented in the results or in the supplementary coding table (Table S1). The 58% figure in Section IV-A is based entirely on the citation category 'Only References Non-ViT Grad-CAM,' which conflates citation behavior with methodological specificity. A paper citing only Selvaraju et al. could still specify its chosen feature tensor, gradient target, token handling, and aggregation in the main text, while a paper citing a ViT-specific adaptation could leave those choices implicit. Therefore the conclusion in Section VI that 'most papers do not provide a full mathematical or implementation-level account' is not established by the reported data. The authors should report the distribution of their 'reporting detail' coding (e.g., a cross-tabulation with citation category) or re-code the corpus for the explicit presence of the five required elements (feature location, gradient target, token handling, aggregation, reshaping), and then revise the 58% claim and the associated language accordingly.
- [II-B4, Eqs. (8)-(9) and (15)-(16)] There is a dimensional inconsistency in the definition of the multi-headed feature maps and their use in the linear combination equations. Eq. (8) defines \hat{F}^{(h,l)}_2 as a vector over patch tokens in R^{N-1}, and Eq. (9) defines \hat{F}^{(h,l)}_3 as a vector over embedding dimensions in R^{d_h}; however, Eqs. (15)-(16) treat \hat{F}^{(h,l)}_{m,i,k} as a matrix indexed by both token i and channel k. For m=2, no channel index remains after extracting the CLS row, and for m=3, the CLS row has no token index. The intended object appears to be the full token-by-channel matrix (e.g., F^{(h,l)}_3 with the CLS row removed, or F^{(h,l)}_2 with the CLS row and column removed), not the CLS-conditioned vectors defined in Eqs. (8)-(9). Please clarify the notation and ensure the equations are dimensionally consistent, since the taxonomy is a central contribution of the paper.
minor comments (6)
- [II-B3] The text contains the typo 'VITs' in the sentence beginning 'In VITs, feature representations instead...' and should read 'ViTs'.
- [II-B4] The sentence 'This can be formalated as follows' contains a misspelling of 'formulated'.
- [References] The reference list includes a large number of works by the authors (Refs. [2]-[36]) that appear unrelated to Grad-CAM or vision transformers; please verify that all cited works are actually relevant to the statements they support.
- [Table S1] The 'Notes' column contains subjective annotations such as 'possibly pytorch-grad-cam' and 'Unclear structure'; these should be formally defined in the codebook so that the coding is transparent and reproducible.
- [Fig. 6 and Fig. 7] The figure legends for the pie charts should clarify that the counts in categories such as '[53] (9)' include both the originating paper and the papers that explicitly adopt it, to avoid confusion about the 26-paper total.
- [Eq. (18)] The normalization formula in Eq. (18) indicates the final heatmap lies in [0,1]^{H x W}, but the preceding heatmap L^{(l,c)}_Grad-CAM is at patch resolution; the upsampling step is described in the text but should also be reflected in the equation or its caption.
Circularity Check
No significant circularity: the taxonomy is descriptive and the audit claim rests on an acknowledged citation-based inference, not on a constructional equivalence.
full rationale
This is a literature audit; its central claim is an empirical generalization about 175 surveyed papers, not a derivation from its own inputs. The taxonomy (Eqs. 1-18) is defined independently of the audit results and is explicitly descriptive, and the qualitative demonstrations (Figs. 2-3, S1) are self-contained computations on standard external models (ViT-B/16, BLIP), so the illustrative material is not benchmarked against or derived from the corpus findings. The 58% figure in Section IV-A is computed from an explicit citation category ('Only References Non-ViT Grad-CAM', Fig. 6), and the conclusion that underspecification follows from that citation pattern is an inference, acknowledged in S-IV as interpretive ('ambiguous cases were treated as ambiguous'), not an equivalence forced by construction; a paper citing only Selvaraju et al. could still specify its adaptation, so the reduction is defeasible rather than definitional. The skeptical concern that Section III-B's coded 'reporting detail' variable is never presented is a real evidentiary gap that weakens support for the headline percentage, but it is a measurement-validity issue, not circularity. Self-citation is present but non-load-bearing: the introduction's CNN-usage citation list includes the authors' own prior works ([2]-[36]), and the authors' own papers (supplementary refs [189]-[193], e.g., Winsor-CAM) appear in the audited corpus and are classified as explicit adopters of ALBEF's method, which if anything biases the reported percentages conservatively rather than forcing the conclusion. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported, and no ansatz is smuggled in via citation. Score 1 reflects only the mild self-inclusion in the corpus and the unclearly bridged citation-to-reporting inference, both of which are proportionality and correctness concerns rather than circular steps.
Assumptions & free parameters
assumptions (3)
- domain assumption Papers that cite only the original CNN Grad-CAM paper or CNN variants do not specify a reproducible ViT adaptation.
- domain assumption The surveyed corpus, restricted to CVF Open Access, NeurIPS Proceedings, and IEEE Access, is representative of the broader literature where these claims are made.
- domain assumption Token indices in ViTs follow a raster-scan patch ordering so that discarding the CLS token and reshaping yields a spatially valid map.
Cite this review
Pith. "Pith review of Grad-CAM for Vision Transformers: A Systematic Taxonomy and Audit of Methodological Ambiguity in Explainable AI." pith.science (2026). https://pith.science/paper/NPFPFHMU
@misc{pith2026260805258,
author = {Pith},
title = {Pith review of: Grad-CAM for Vision Transformers: A Systematic Taxonomy and Audit of Methodological Ambiguity in Explainable AI},
year = {2026},
howpublished = {\url{https://pith.science/paper/NPFPFHMU}},
note = {Machine review of arXiv:2608.05258}
}
read the original abstract
Gradient-weighted Class Activation Mapping (Grad-CAM) is widely used to visualize model decisions, but it was originally formulated for convolutional neural networks, where spatial feature maps and channel dimensions have clear architectural meanings. Vision Transformers (ViTs) do not provide the same structure, instead representing images through tokens, attention, residual streams, and multimodal interactions. This paper presents a systematic taxonomy and literature audit of how Grad-CAM and related methods are adapted, justified, and reported for ViT-based architectures. From an initial search of more than 550 papers, we identify 175 papers that apply Grad-CAM or Grad-CAM-adjacent methods to ViTs. We find that most papers do not provide a full mathematical or implementation-level account of how Grad-CAM is adapted to transformer representations. To characterize this gap, we introduce a descriptive taxonomy of ViT Grad-CAM adaptations that makes explicit the feature locations, gradient targets, spatial reconstruction steps, and aggregation choices that are often left implicit. This taxonomy is not intended to prescribe a single correct adaptation, but to clarify the range of methodological choices being made. The study shows that Grad-CAM on ViTs is often treated as a trivial extension of CNN-based Grad-CAM, despite requiring nontrivial choices that affect rigor, reproducibility, and interpretation.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Grad-cam: Visual explanations from deep networks via gradient-based localization,
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proc. AAAI Conf. Artif. Intell., 2017, pp. 618–626
2017
-
[2]
Representation learning and na- ture encoded fusion for heterogeneous sensor networks,
L. Wang and Q. Liang, “Representation learning and na- ture encoded fusion for heterogeneous sensor networks,” IEEE Access, vol. 7, pp. 39 227–39 235, 2019
2019
-
[3]
Congestion aware dynamic user association in heterogeneous cellular network: A stochastic decision approach,
L. Wang, W. Chen, and J. Li, “Congestion aware dynamic user association in heterogeneous cellular network: A stochastic decision approach,” in2014 IEEE Interna- tional Conference on Communications (ICC). IEEE, 2014, pp. 2636–2640
2014
-
[4]
Enhanced robustness by symmetry enforcement,
L. Wang, A. Ghimire, K. Santosh, Z. Zhang, and X. Li, “Enhanced robustness by symmetry enforcement,” in IEEE Conference on Artificial Intelligence (IEEE CAI) 2024, 2024
2024
-
[5]
Partial interference alignment for heterogeneous cellular networks,
L. Wang and Q. Liang, “Partial interference alignment for heterogeneous cellular networks,”IEEE Access, vol. 6, pp. 22 592–22 601, 2018
2018
-
[6]
Optimization for user centric massive mimo cell free networks via large system analysis,
——, “Optimization for user centric massive mimo cell free networks via large system analysis,” in2016 IEEE Global Communications Conference (GLOBE- COM). IEEE, 2016, pp. 1–1
2016
-
[8]
Exploration vs exploitation for distributed channel access in cognitive radio networks: A multi-user case study,
L. Wang, X. Chen, Z. Zhao, and H. Zhang, “Exploration vs exploitation for distributed channel access in cognitive radio networks: A multi-user case study,” in2011 11th International Symposium on Communications & Infor- mation Technologies (ISCIT). IEEE, 2011, pp. 360–365
2011
-
[9]
Deep reinforcement learning based computation offloading for mobility-aware edge computing,
M. Shi, R. Wang, E. Liu, Z. Xu, and L. Wang, “Deep reinforcement learning based computation offloading for mobility-aware edge computing,” inInternational con- ference on communications and networking in china. Springer International Publishing Cham, 2019, pp. 53– 65
2019
Show all 251 references
-
[10]
Performance analysis of co- operative multicell precoding with global csi and local individual csi in the large dimensional regime,
L. Wang and Q. Liang, “Performance analysis of co- operative multicell precoding with global csi and local individual csi in the large dimensional regime,”IEEE Transactions on Vehicular Technology, vol. 67, no. 4, pp. 3229–3238, 2017
2017
-
[11]
Low complexity optimization for user centric cel- lular networks via large dimensional analysis,
——, “Low complexity optimization for user centric cel- lular networks via large dimensional analysis,”Physical Communication, vol. 25, pp. 412–419, 2017
2017
-
[12]
Improving robustness of deep neural networks via large-difference transformation,
L. Wang, C. Wang, Y . Li, and R. Wang, “Improving robustness of deep neural networks via large-difference transformation,”Neurocomputing, vol. 450, pp. 411–419, 2021
2021
-
[13]
Looking beyond content: Modeling and detection of fake news from a social context perspective
K. Xiao, L. Wang, A. Gupta, and X. Qin, “Looking beyond content: Modeling and detection of fake news from a social context perspective.” inProceedings of the 55th Hawaii International Conference on System Sciences 2022, 2022, pp. 1–10
2022
-
[14]
Large system analysis for densification of cellular networks with massive mimo,
L. Wang and Q. Liang, “Large system analysis for densification of cellular networks with massive mimo,” in2016 IEEE Conference on Computer Communications Workshops (INFOCOM WKSHPS). IEEE, 2016, pp. 956–961
2016
-
[15]
Collabora- tive spectrum sharing based on information pooling for cognitive radio networks with channel heterogeneity,
L. Wang, X. Chen, Z. Zhao, and H. Zhang, “Collabora- tive spectrum sharing based on information pooling for cognitive radio networks with channel heterogeneity,” in 2011 11th International Symposium on Communications & Information Technologies (ISCIT). IEEE, 2011, pp. 483–488
2011
-
[16]
Dense cross-connected ensemble convolutional neural networks for enhanced model robustness,
L. Wang, X. Li, and Z. Zhang, “Dense cross-connected ensemble convolutional neural networks for enhanced model robustness,”arXiv preprint arXiv:2412.07022, 2024
2024 arXiv
-
[17]
Information theory and represen- tation learning inspired multimodal data fusion,
L. Wang and Y . Li, “Information theory and represen- tation learning inspired multimodal data fusion,”IEEE MMTC Frontier, 2019
2019
-
[19]
Enhanc- ing adversarial robustness of deep neural networks through supervised contrastive learning,
L. Wang, N. Nayyem, and A. Rakin, “Enhanc- ing adversarial robustness of deep neural networks through supervised contrastive learning,”arXiv preprint arXiv:2412.19747, 2024
2024 arXiv
-
[21]
Multi-scale unrectified push-pull with channel attention for enhanced corruption robustness,
R. N. Ranabhat, L. Wang, X. Qin, Y . Zhou, and K. San- tosh, “Multi-scale unrectified push-pull with channel attention for enhanced corruption robustness,” inPro- ceedings of the AAAI Symposium Series 2025, vol. 6, no. 1, 2025, pp. 34–41
2025
-
[22]
Expert-guided ex- plainable few-shot learning for medical image diagnosis,
I. I. Uddin, L. Wang, and K. Santosh, “Expert-guided ex- plainable few-shot learning for medical image diagnosis,” inMICCAI Workshop on Data Engineering in Medical Imaging 2025. Springer Nature Switzerland, 2025, pp. 95–104
2025
-
[23]
Ecologically valid benchmarking and adaptive attention: Scalable marine bioacoustic monitoring,
N. R. Rasmussen, R. Rizk, L. Wang, and K. Santosh, “Ecologically valid benchmarking and adaptive attention: Scalable marine bioacoustic monitoring,”arXiv preprint 17 arXiv:2509.04682, 2025
2025 arXiv
-
[24]
Shape-aware thoracic edge map chest x- ray representation for pulmonary abnormality screening,
S. Chataut, A. Ghimire, A. Thakur, L. Wang, and K. Santosh, “Shape-aware thoracic edge map chest x- ray representation for pulmonary abnormality screening,” inInternational Conference on DATA ANALYTICS & LEARNING. Springer, 2024, pp. 209–220
2024
-
[25]
Expert-guided ex- plainable few-shot learning with active sample selection for medical image analysis,
L. Wang, I. I. Uddin, and K. Santosh, “Expert-guided ex- plainable few-shot learning with active sample selection for medical image analysis,”IEEE Journal of Biomedical and Health Informatics, 2026
2026
-
[26]
Coswin: Convolution enhanced hierarchical shifted win- dow attention for small-scale vision,
P. Khadka, R. Rizk, L. Wang, and K. Santosh, “Coswin: Convolution enhanced hierarchical shifted win- dow attention for small-scale vision,”arXiv preprint arXiv:2509.08959, 2025
2025 arXiv
-
[28]
Channel-selected stratified nested cross- validation for clinically relevant eeg-based parkinson’s disease detection,
N. R. Rasmussen, R. Rizk, L. Wang, A. Singh, and K. Santosh, “Channel-selected stratified nested cross- validation for clinically relevant eeg-based parkinson’s disease detection,” in2026 IEEE Conference on Artificial Intelligence (CAI). IEEE, 2026, pp. 91–97
2026
-
[30]
Promoting shape bias in cnns: Frequency-based and contrastive regularization for corruption robustness,
R. N. Ranabhat, L. Wang, A. K. Patel, and K. San- tosh, “Promoting shape bias in cnns: Frequency-based and contrastive regularization for corruption robustness,” inInternational Conference on Intelligent Systems and Pattern Recognition. Springer, 2025, pp. 16–26
2025
-
[31]
Learning to select like humans: Explainable active learning for medical imaging,
I. I. Uddin, L. Wang, X. Qin, Y . Zhou, and K. San- tosh, “Learning to select like humans: Explainable active learning for medical imaging,” in2026 IEEE Conference on Artificial Intelligence (CAI). IEEE, 2026, pp. 458– 463
2026
-
[32]
Ex- plainable novel category discovery in semantic concept space,
I. I. Uddin, Y . Zhou, K. Santosh, and L. Wang, “Ex- plainable novel category discovery in semantic concept space,”arXiv preprint arXiv:2607.04548, 2026
2026 arXiv
-
[33]
Mecha- nistic interpretability of llm jailbreaks via internal attri- bution graphs,
A. Wagle, I. I. Uddin, C. Zhang, and L. Wang, “Mecha- nistic interpretability of llm jailbreaks via internal attri- bution graphs,”arXiv preprint arXiv:2607.07903, 2026
2026 arXiv
-
[34]
Frequency-aware contrastive learning for robust shape- biased convolutional neural networks,
R. N. Ranabhat, L. Wang, A. K. Patel, and K. Santosh, “Frequency-aware contrastive learning for robust shape- biased convolutional neural networks,”Pattern Recogni- tion Letters, 2026
2026
-
[36]
Large dimensional analysis of cooperative multicell precoding with local individual csi,
L. Wang and Q. Liang, “Large dimensional analysis of cooperative multicell precoding with local individual csi,” in2016 IEEE Conference on Computer Communi- cations Workshops (INFOCOM WKSHPS). IEEE, 2016, pp. 89–94
2016
-
[37]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weis- senborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Min- derer, G. Heigold, S. Gellyet al., “An image is worth 16x16 words: Transformers for image recognition at scale,” inInt. Conf. Learn. Represent., 2021
2021
-
[38]
Pytorch library for cam methods,
J. Gildenblatet al., “Pytorch library for cam methods,” https://github.com/jacobgil/pytorch-grad-cam, 2021
2021
-
[39]
ImageNet Large Scale Visual Recognition Challenge,
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet Large Scale Visual Recognition Challenge,”Int. J. Comput. Vis., vol. 115, no. 3, pp. 211–252, Dec. 2015
2015
-
[40]
Swin transformer: Hierarchical vision transformer using shifted windows,
Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” inProc. AAAI Conf. Artif. Intell., October 2021, pp. 10 012–10 022
2021
-
[43]
BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
J. Li, D. Li, C. Xiong, and S. Hoi, “BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” inProc. Int. Conf. Mach. Learn., vol. 162, Jul 2022, pp. 12 888–12 900
2022
-
[44]
Training data-efficient image trans- formers &; distillation through attention,
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablay- rolles, and H. Jegou, “Training data-efficient image trans- formers &; distillation through attention,” inProc. Int. Conf. Mach. Learn., vol. 139, Jul. 2021, pp. 10 347– 10 357
2021
-
[45]
Grad-CAM++: Generalized Gradient- Based Visual Explanations for Deep Convolutional Net- works ,
A. Chattopadhay, A. Sarkar, P. Howlader, and V . N. Bal- asubramanian, “ Grad-CAM++: Generalized Gradient- Based Visual Explanations for Deep Convolutional Net- works ,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., Mar. 2018, pp. 839–847
2018
-
[49]
Ablation-CAM: Visual Explanations for Deep Convolutional Network via Gradient-free Localization ,
S. Desai and H. G. Ramaswamy, “ Ablation-CAM: Visual Explanations for Deep Convolutional Network via Gradient-free Localization ,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., Mar. 2020, pp. 972–980
2020
-
[50]
Full-gradient representation 18 for neural network visualization,
S. Srinivas and F. Fleuret, “Full-gradient representation 18 for neural network visualization,” inProc. Adv. Neural Inf. Process. Syst., vol. 32, 2019
2019
-
[51]
Axiom-based grad-cam: Towards accurate visualization and explanation of cnns,
R. Fu, Q. Hu, X. Dong, Y . Guo, Y . Gao, and B. Li, “Axiom-based grad-cam: Towards accurate visualization and explanation of cnns,” inProc. The Brit. Mach. Vis. Conf., 2020
2020
-
[52]
Learning Deep Features for Discriminative Lo- calization ,
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Tor- ralba, “ Learning Deep Features for Discriminative Lo- calization ,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., Jun. 2016, pp. 2921–2929
2016
-
[57]
Boosting the transferability of adversarial attack on vision transformer with adaptive token tuning,
D. Ming, P. Ren, Y . Wang, and X. Feng, “Boosting the transferability of adversarial attack on vision transformer with adaptive token tuning,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 20 887–20 918
2024
-
[61]
Emergent open-vocabulary semantic segmentation from off-the- shelf vision-language models,
J. Luo, S. Khandelwal, L. Sigal, and B. Li, “Emergent open-vocabulary semantic segmentation from off-the- shelf vision-language models,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 4029– 4040
2024
-
[65]
Enhancing prompt generation with adaptive refinement for camouflaged object detection,
X. Chen, G. Ren, T. Dai, T. Stathaki, and H. Liu, “Enhancing prompt generation with adaptive refinement for camouflaged object detection,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 20 672–20 682
2025
-
[66]
Quantifying attention flow in transformers,
S. Abnar and W. Zuidema, “Quantifying attention flow in transformers,” inProc. Annu. Meeting Assoc. Comput. Linguistics, Jul. 2020, pp. 4190–4197. 19 Supplementary Material for Grad-CAM for Vision Transformers: A Systematic Taxonomy and Audit of Methodological Ambiguity in Exp...
2020
-
[67]
Unlike attention-map adaptations, the method operates on feature activations from the MLP in the final transformer blocks and uses the true-label class score as the gradient target
Principles of Visual Tokens for Efficient Video Under- standing [147]:[147] uses a Grad-CAM-style oracle to esti- mate the importance of spatiotemporal tokens in a ViT-based video classification model. Unlike attention-map adaptations, the method operates on feature activation...
-
[68]
the best way we found to apply GradCAM was to treat the last attention layer’s [CLS] token as the designated feature map,
Transformer Interpretability Beyond Attention Visual- ization [4]:adapts Grad-CAM to ViTs by using the final attention layer rather than convolutional feature maps. The authors explicitly state that “the best way we found to apply GradCAM was to treat the last attention layer’...
-
[69]
is explicitly adopted by [22, 32, 89] and mentioned in [6, 7, 29, 48, 60, 63, 87, 90, 101, 111, 113, 118, 124, 135– 137, 148]
-
[70]
Rather than using an attention map as the attribution representation, the method operates on token-level feature activations
Squeeze-and-Excitation Vision Transformer (SE-ViT) for Lung Nodule Classification [159]:[159] adapts Grad-CAM to SE-ViT by replacing CNN channels with image tokens. Rather than using an attention map as the attribution representation, the method operates on token-level feature...
-
[71]
Rather than operating on attention maps, it fuses gradients and intermediate ViT features from a selected transformer layer
Boosting the Transferability of Adversarial Attack on Vision Transformer with Adaptive Token Tuning [88]:[88] uses a Grad-CAM-inspired feature-importance computation to guide patch masking in a ViT. Rather than operating on attention maps, it fuses gradients and intermediate V...
-
[72]
Cross-Attention Grad-CAM: Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models [19]uses cross-attention maps as the attribution-bearing representation. Rather than extracting token embeddings such asF 1,F 4,F 5, orF 6, the method oper- ates directly on a cros...
-
[73]
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation [11]is best charac- terized as an attention-map-based attribution method
is explicitly adopted by [42, 101] and mentioned in [27, 62, 99]. Align before Fuse: Vision and Language Representation Learning with Momentum Distillation [11]is best charac- terized as an attention-map-based attribution method. Rather than using token-embedding representatio...
-
[74]
matching
is explicitly adopted by [37, 39, 41, 51, 54, 59, 65, 96, 189–193] and mentioned in [12, 19, 35, 36, 42, 50, 52, 64, 94, 95, 98, 99, 101, 120, 128, 155, 192, 194–197]. Enhancing Prompt Generation with Adaptive Refine- ment for Camouflaged Object Detection [155]applies gradient...
-
[75]
A dog on a white bed
is explicitly adopted by [127] and mentioned in [98, 100, 155]. Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability [154]applies Grad-CAM-style attribution to intermediate activations from the vision component of a vision-language transfor...
-
[76]
Grad-cam: Visual explana- tions from deep networks via gradient-based localiza- tion,
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explana- tions from deep networks via gradient-based localiza- tion,” inProc. AAAI Conf. Artif. Intell., 2017, pp. 618– 626
2017
-
[77]
BLIP: Bootstrap- ping language-image pre-training for unified vision- language understanding and generation,
J. Li, D. Li, C. Xiong, and S. Hoi, “BLIP: Bootstrap- ping language-image pre-training for unified vision- language understanding and generation,” inProc. Int. Conf. Mach. Learn., vol. 162, Jul 2022, pp. 12 888– 12 900
2022
-
[78]
Microsoft coco: Common objects in context,
T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll ´ar, and C. L. Zitnick, “Microsoft coco: Common objects in context,” inProc. Eur. Conf. Comput. Vis., vol. 8693, 2014, pp. 740–755
2014
-
[79]
Transformer in- terpretability beyond attention visualization,
H. Chefer, S. Gur, and L. Wolf, “Transformer in- terpretability beyond attention visualization,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2021, pp. 782–791
2021
-
[80]
Transreid: Transformer-based object re-identification,
S. He, H. Luo, P. Wang, F. Wang, H. Li, and W. Jiang, “Transreid: Transformer-based object re-identification,” inProc. AAAI Conf. Artif. Intell., October 2021, pp. 15 013–15 022
2021
-
[81]
Generic attentiemer- gent on-model explainability for interpreting bi-modal and encoder-decoder transformers,
H. Chefer, S. Gur, and L. Wolf, “Generic attentiemer- gent on-model explainability for interpreting bi-modal and encoder-decoder transformers,” inProc. AAAI Conf. Artif. Intell., October 2021, pp. 397–406
2021
-
[82]
Ia-redˆ2: Interpretability-aware redundancy reduction for vision transformers,
B. Pan, R. Panda, Y . Jiang, Z. Wang, R. Feris, and A. Oliva, “Ia-redˆ2: Interpretability-aware redundancy reduction for vision transformers,” inProc. Adv. Neural Inf. Process. Syst., vol. 34, 2021, pp. 24 898–24 911
2021
-
[83]
Analogous to evolutionary algorithm: Designing a unified sequence model,
J. Zhang, C. Xu, J. Li, W. Chen, Y . Wang, Y . Tai, S. Chen, C. Wang, F. Huang, and Y . Liu, “Analogous to evolutionary algorithm: Designing a unified sequence model,” inProc. Adv. Neural Inf. Process. Syst., vol. 34, 2021, pp. 26 674–26 688
2021
-
[84]
Passive attention in artificial neural networks predicts human visual selectivity,
T. Langlois, H. Zhao, E. Grant, I. Dasgupta, T. Griffiths, and N. Jacoby, “Passive attention in artificial neural networks predicts human visual selectivity,” inProc. Adv. Neural Inf. Process. Syst., vol. 34, 2021, pp. 27 094–27 106
2021
-
[85]
Vitae: Vision transformer advanced by exploring intrinsic inductive bias,
Y . Xu, Q. ZHANG, J. Zhang, and D. Tao, “Vitae: Vision transformer advanced by exploring intrinsic inductive bias,” inProc. Adv. Neural Inf. Process. Syst., vol. 34, 2021, pp. 28 522–28 535
2021
-
[86]
Align before fuse: Vision and language representation learning with momentum distillation,
J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” in Proc. Adv. Neural Inf. Process. Syst., vol. 34, 2021, pp. 9694–9705
2021
-
[87]
Vlmae: Vision-language masked autoencoder,
S. He, T. Guo, T. Dai, R. Qiao, C. Wu, X. Shu, and B. Ren, “Vlmae: Vision-language masked autoencoder,” 2022, arXiv preprint arXiv:2208.09374
2022 arXiv
-
[88]
A compre- hensive study of image classification model sensitivity to foregrounds, backgrounds, and visual attributes,
M. Moayeri, P. Pope, Y . Balaji, and S. Feizi, “A compre- hensive study of image classification model sensitivity to foregrounds, backgrounds, and visual attributes,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2022, pp. 19 087–19 097
2022
-
[89]
A challenging benchmark of anime style recognition,
H. Li, S. Guo, K. Lyu, X. Yang, T. Chen, J. Zhu, and H. Zeng, “A challenging benchmark of anime style recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2022, pp. 4721– 4730
2022
-
[90]
Metaformer is actually what you need for vision,
W. Yu, M. Luo, P. Zhou, C. Si, Y . Zhou, X. Wang, J. Feng, and S. Yan, “Metaformer is actually what you need for vision,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2022, pp. 10 819–10 829
2022
-
[91]
Delving deep into the generalization of vision transformers under distribution shifts,
C. Zhang, M. Zhang, S. Zhang, D. Jin, Q. Zhou, Z. Cai, H. Zhao, X. Liu, and Z. Liu, “Delving deep into the generalization of vision transformers under distribution shifts,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2022, pp. 7277–7286
2022
-
[92]
General facial representation learning in a visual- linguistic manner,
Y . Zheng, H. Yang, T. Zhang, J. Bao, D. Chen, Y . Huang, L. Yuan, D. Chen, M. Zeng, and F. Wen, “General facial representation learning in a visual- linguistic manner,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2022, pp. 18 697–18 709
2022
-
[93]
Multi-modal alignment using representa- tion codebook,
J. Duan, L. Chen, S. Tran, J. Yang, Y . Xu, B. Zeng, and T. Chilimbi, “Multi-modal alignment using representa- tion codebook,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2022, pp. 15 651–15 660
2022
-
[94]
Plug-and-play VQA: Zero-shot VQA by conjoining large pretrained models with zero training,
A. M. H. Tiong, J. Li, B. Li, S. Savarese, and S. C. Hoi, “Plug-and-play VQA: Zero-shot VQA by conjoining large pretrained models with zero training,” inFindings Assoc. Comput. Linguistics, Dec. 2022, pp. 951–967
2022
-
[95]
Inception transformer,
C. Si, W. Yu, P. Zhou, Y . Zhou, X. Wang, and S. Yan, “Inception transformer,” inProc. Adv. Neural Inf. Pro- cess. Syst., vol. 35, 2022, pp. 23 495–23 509
2022
-
[96]
Delving into sequential patches for deep- fake detection,
J. Guan, H. Zhou, Z. Hong, E. Ding, J. Wang, C. Quan, and Y . Zhao, “Delving into sequential patches for deep- fake detection,” inProc. Adv. Neural Inf. Process. Syst., vol. 35, 2022, pp. 4517–4530
2022
-
[97]
Adversarial normalization: I can visualize everything (ice),
H. Choi, S. Jin, and K. Han, “Adversarial normalization: I can visualize everything (ice),” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 12 115–12 124
2023
-
[98]
A new benchmark: On the utility of synthetic data with blender for bare supervised learning and downstream domain adaptation,
H. Tang and K. Jia, “A new benchmark: On the utility of synthetic data with blender for bare supervised learning and downstream domain adaptation,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 15 954–15 964
2023
-
[99]
Selfme: Self-supervised motion learning for micro- expression recognition,
X. Fan, X. Chen, M. Jiang, A. R. Shahid, and H. Yan, “Selfme: Self-supervised motion learning for micro- expression recognition,” inProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit., June 2023, pp. 13 834– 13 843
2023
-
[100]
Marlin: Masked autoencoder for facial video representation learning,
Z. Cai, S. Ghosh, K. Stefanov, A. Dhall, J. Cai, H. Rezatofighi, R. Haffari, and M. Hayat, “Marlin: Masked autoencoder for facial video representation learning,” inProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit., June 2023, pp. 1493–1504
2023
-
[101]
Blackvip: Black-box visual prompting for robust transfer learning,
C. Oh, H. Hwang, H.-y. Lee, Y . Lim, G. Jung, J. Jung, H. Choi, and K. Song, “Blackvip: Black-box visual prompting for robust transfer learning,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 32 2023, pp. 24 224–24 235
2023
-
[102]
To- kenhpe: Learning orientation tokens for efficient head pose estimation via transformers,
C. Zhang, H. Liu, Y . Deng, B. Xie, and Y . Li, “To- kenhpe: Learning orientation tokens for efficient head pose estimation via transformers,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 8897–8906
2023
-
[103]
Boost vision trans- former with gpu-friendly sparsity and quantization,
C. Yu, T. Chen, Z. Gan, and J. Fan, “Boost vision trans- former with gpu-friendly sparsity and quantization,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 22 658–22 668
2023
-
[104]
Vision diffmask: Faithful interpretation of vision transformers with differentiable patch masking,
A. Nalmpantis, A. Panagiotopoulos, J. Gkountouras, K. Papakostas, and W. Aziz, “Vision diffmask: Faithful interpretation of vision transformers with differentiable patch masking,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2023, pp. 3756– 3763
2023
-
[105]
Pha: Patch-wise high-frequency augmentation for transformer-based person re-identification,
G. Zhang, Y . Zhang, T. Zhang, B. Li, and S. Pu, “Pha: Patch-wise high-frequency augmentation for transformer-based person re-identification,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 14 133–14 142
2023
-
[106]
Vision trans- formers with mixed-resolution tokenization,
T. Ronen, O. Levy, and A. Golbert, “Vision trans- formers with mixed-resolution tokenization,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work- shops, June 2023, pp. 4613–4622
2023
-
[107]
D3former: Debiased dual distilled transformer for incremental learning,
A. Mohamed, R. Grandhe, K. J. Joseph, S. Khan, and F. Khan, “D3former: Debiased dual distilled transformer for incremental learning,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2023, pp. 2421–2430
2023
-
[108]
Shared inter- est...sometimes: Understanding the alignment between human perception, vision architectures, and saliency map techniques,
K. Morrison, A. Mehra, and A. Perer, “Shared inter- est...sometimes: Understanding the alignment between human perception, vision architectures, and saliency map techniques,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2023, pp. 3776– 3781
2023
-
[109]
Semicvt: Semi- supervised convolutional vision transformer for seman- tic segmentation,
H. Huang, S. Xie, L. Lin, R. Tong, Y .-W. Chen, Y . Li, H. Wang, Y . Huang, and Y . Zheng, “Semicvt: Semi- supervised convolutional vision transformer for seman- tic segmentation,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 11 340–11 349
2023
-
[110]
Masked autoencoding does not help natural language supervision at scale,
F. Weers, V . Shankar, A. Katharopoulos, Y . Yang, and T. Gunter, “Masked autoencoding does not help natural language supervision at scale,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 23 432–23 444
2023
-
[111]
Vilem: Visual- language error modeling for image-text retrieval,
Y . Chen, Z. Ma, Z. Zhang, Z. Qi, C. Yuan, Y . Shan, B. Li, W. Hu, X. Qie, and J. Wu, “Vilem: Visual- language error modeling for image-text retrieval,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 11 018–11 027
2023
-
[112]
Fashionsap: Symbols and attributes prompt for fine-grained fashion vision-language pre-training,
Y . Han, L. Zhang, Q. Chen, Z. Chen, Z. Li, J. Yang, and Z. Cao, “Fashionsap: Symbols and attributes prompt for fine-grained fashion vision-language pre-training,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 15 028–15 038
2023
-
[113]
Zero-shot referring image segmentation with global-local context features,
S. Yu, P. H. Seo, and J. Son, “Zero-shot referring image segmentation with global-local context features,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 19 456–19 465
2023
-
[114]
Improving visual grounding by encouraging consistent gradient-based explanations,
Z. Yang, K. Kafle, F. Dernoncourt, and V . Ordonez, “Improving visual grounding by encouraging consistent gradient-based explanations,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 19 165– 19 174
2023
-
[115]
Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation,
Y . Lin, M. Chen, W. Wang, B. Wu, K. Li, B. Lin, H. Liu, and X. He, “Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 15 305–15 314
2023
-
[116]
Multi-modal representation learn- ing with text-driven soft masks,
J. Park and B. Han, “Multi-modal representation learn- ing with text-driven soft masks,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 2798–2807
2023
-
[117]
From images to textual prompts: Zero-shot visual question answering with frozen large language models,
J. Guo, J. Li, D. Li, A. M. H. Tiong, B. Li, D. Tao, and S. Hoi, “From images to textual prompts: Zero-shot visual question answering with frozen large language models,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 10 867–10 877
2023
-
[118]
Sparse multi- modal vision transformer for weakly supervised seman- tic segmentation,
J. Hanna, M. Mommert, and D. Borth, “Sparse multi- modal vision transformer for weakly supervised seman- tic segmentation,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2023, pp. 2145– 2154
2023
-
[119]
Semantic information in contrastive learning,
S. Quan, M. Hirano, and Y . Yamakawa, “Semantic information in contrastive learning,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 5686–5696
2023
-
[120]
Smmix: Self-motivated image mixing for vision transformers,
M. Chen, M. Lin, Z. Lin, Y . Zhang, F. Chao, and R. Ji, “Smmix: Self-motivated image mixing for vision transformers,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 17 260–17 270
2023
-
[121]
Cose: A consistency- sensitivity metric for saliency on image classification,
R. Daroya, A. Sun, and S. Maji, “Cose: A consistency- sensitivity metric for saliency on image classification,” inProc. AAAI Conf. Artif. Intell. Workshops, October 2023, pp. 149–158
2023
-
[122]
Tall: Thumbnail layout for deepfake video detection,
Y . Xu, J. Liang, G. Jia, Z. Yang, Y . Zhang, and R. He, “Tall: Thumbnail layout for deepfake video detection,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 22 658–22 668
2023
-
[123]
Funnybirds: A synthetic vision dataset for a part-based analysis of explainable ai methods,
R. Hesse, S. Schaub-Meyer, and S. Roth, “Funnybirds: A synthetic vision dataset for a part-based analysis of explainable ai methods,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 3981–3991
2023
-
[124]
Thinking image color aesthetics assessment: Models, datasets and benchmarks,
S. He, A. Ming, Y . Li, J. Sun, S. Zheng, and H. Ma, “Thinking image color aesthetics assessment: Models, datasets and benchmarks,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 21 838–21 847
2023
-
[125]
Hivlp: Hierarchical interactive video-language pre-training,
B. Shao, J. Liu, R. Pei, S. Xu, P. Dai, J. Lu, W. Li, and Y . Yan, “Hivlp: Hierarchical interactive video-language pre-training,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 13 756–13 766
2023
-
[126]
Vl-match: Enhancing vision-language pretraining with token-level and instance-level matching,
J. Bi, D. Cheng, P. Yao, B. Pang, Y . Zhan, C. Yang, Y . Wang, H. Sun, W. Deng, and Q. Zhang, “Vl-match: Enhancing vision-language pretraining with token-level and instance-level matching,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 2584–2593. 33
2023
-
[127]
Vilta: Enhancing vision-language pre-training through textual augmentation,
W. Wang, Z. Yang, B. Xu, J. Li, and Y . Sun, “Vilta: Enhancing vision-language pre-training through textual augmentation,” inProc. AAAI Conf. Artif. Intell., Octo- ber 2023, pp. 3158–3169
2023
-
[128]
Spatio-temporal prompting network for robust video feature extraction,
G. Sun, C. Wang, Z. Zhang, J. Deng, S. Zafeiriou, and Y . Hua, “Spatio-temporal prompting network for robust video feature extraction,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 13 587–13 597
2023
-
[129]
Weakly supervised referring image segmentation with intra-chunk and inter-chunk consistency,
J. Lee, S. Lee, J. Nam, S. Yu, J. Do, and T. Taghavi, “Weakly supervised referring image segmentation with intra-chunk and inter-chunk consistency,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 21 870–21 881
2023
-
[130]
Wireless cap- sule endoscopy image classification: An explainable ai approach,
D. Varam, R. Mitra, M. Mkadmi, R. A. Riyas, D. A. Abuhani, S. Dhou, and A. Alzaatreh, “Wireless cap- sule endoscopy image classification: An explainable ai approach,”IEEE Access, vol. 11, pp. 105 262–105 280, 2023
2023
-
[131]
Spike-driven transformer,
M. Yao, J. Hu, Z. Zhou, L. Yuan, Y . Tian, B. Xu, and G. Li, “Spike-driven transformer,” inProc. Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 64 043–64 058
2023
-
[132]
Wbcatt: A white blood cell dataset annotated with detailed morphologi- cal attributes,
S. Tsutsui, W. Pang, and B. Wen, “Wbcatt: A white blood cell dataset annotated with detailed morphologi- cal attributes,” inProc. Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 50 796–50 824
2023
-
[133]
Lico: Explainable models with language-image consistency,
Y . Lei, Z. Li, Y . Li, J. Zhang, and H. Shan, “Lico: Explainable models with language-image consistency,” inProc. Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 61 870–61 887
2023
-
[134]
Implicit differentiable outlier detection enable robust deep multimodal analy- sis,
Z. Wang, S. Medya, and S. Ravi, “Implicit differentiable outlier detection enable robust deep multimodal analy- sis,” inProc. Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 13 854–13 872
2023
-
[135]
Visual explanations of image-text representations via multi- modal information bottleneck attribution,
Y . Wang, T. G. J. Rudner, and A. G. Wilson, “Visual explanations of image-text representations via multi- modal information bottleneck attribution,” inProc. Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 16 009– 16 027
2023
-
[136]
Med- unic: Unifying cross-lingual medical vision-language pre-training by diminishing bias,
Z. Wan, C. Liu, M. Zhang, J. Fu, B. Wang, S. Cheng, L. Ma, C. Quilodr ´an-Casas, and R. Arcucci, “Med- unic: Unifying cross-lingual medical vision-language pre-training by diminishing bias,” inProc. Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 56 186–56 197
2023
-
[137]
On evaluating adversarial robustness of large vision-language models,
Y . Zhao, T. Pang, C. Du, X. Yang, C. LI, N.-M. Cheung, and M. Lin, “On evaluating adversarial robustness of large vision-language models,” inProc. Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 54 111–54 138
2023
-
[138]
Delving into masked autoencoders for multi-label thorax disease classification,
J. Xiao, Y . Bai, A. Yuille, and Z. Zhou, “Delving into masked autoencoders for multi-label thorax disease classification,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., January 2023, pp. 3588–3600
2023
-
[139]
Gafnet: A global fourier self attention based novel network for multi- modal downstream tasks,
O. Susladkar, G. Deshmukh, D. Makwana, S. Mittal, R. S. C. Teja, and R. Singhal, “Gafnet: A global fourier self attention based novel network for multi- modal downstream tasks,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., January 2023, pp. 5242–5251
2023
-
[140]
GroundVLP: Harnessing zero-shot visual grounding from vision- language pre-training and open-vocabulary object de- tection,
H. Shen, T. Zhao, M. Zhu, and J. Yin, “GroundVLP: Harnessing zero-shot visual grounding from vision- language pre-training and open-vocabulary object de- tection,” inProc. AAAI Conf. Artif. Intell., 2024, pp. 4766–4775
2024
-
[141]
Debiformer: Vi- sion transformer with deformable agent bi-level routing attention,
N. BaoLong, C. Zhang, Y . Shi, T. Hirakawa, T. Ya- mashita, T. Matsui, and H. Fujiyoshi, “Debiformer: Vi- sion transformer with deformable agent bi-level routing attention,” inProc. Asian Conf. Comput. Vis., December 2024, pp. 4455–4472
2024
-
[142]
Scca-net: A novel network for image manipulation localization using split- channel contextual attention,
Y . Xiang, K. Zhao, and H. Yin, “Scca-net: A novel network for image manipulation localization using split- channel contextual attention,” inProc. Asian Conf. Comput. Vis., December 2024, pp. 4473–4487
2024
-
[143]
Comparing the decision-making mechanisms by transformers and cnns via explanation methods,
M. Jiang, S. Khorram, and L. Fuxin, “Comparing the decision-making mechanisms by transformers and cnns via explanation methods,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 9546– 9555
2024
-
[144]
Just addπ! pose induced video transformers for understanding activities of daily living,
D. Reilly and S. Das, “Just addπ! pose induced video transformers for understanding activities of daily living,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 18 340–18 350
2024
-
[145]
Biggait: Learning gait representation you want by large vision models,
D. Ye, C. Fan, J. Ma, X. Liu, and S. Yu, “Biggait: Learning gait representation you want by large vision models,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 200–210
2024
-
[146]
Flexible biometrics recognition: Bridging the multimodality gap through attention alignment and prompt tuning,
L. C. O. Tiong, D. Sigmund, C.-H. Chan, and A. B. J. Teoh, “Flexible biometrics recognition: Bridging the multimodality gap through attention alignment and prompt tuning,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 267–276
2024
-
[147]
Adapt- ing short-term transformers for action detection in untrimmed videos,
M. Yang, H. Gao, P. Guo, and L. Wang, “Adapt- ing short-term transformers for action detection in untrimmed videos,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 18 570–18 579
2024
-
[148]
Argue: Attribute-guided prompt tuning for vision-language models,
X. Tian, S. Zou, Z. Yang, and J. Zhang, “Argue: Attribute-guided prompt tuning for vision-language models,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 28 578–28 587
2024
-
[149]
Vision-language models for decod- ing provider attention during neonatal resuscitation,
F. Parodi, J. K. Matelsky, A. Regla-Vargas, E. E. Foglia, C. Lim, D. Weinberg, K. P. Kording, H. M. Herrick, and M. L. Platt, “Vision-language models for decod- ing provider attention during neonatal resuscitation,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Worksh...
2024
-
[150]
Residual-based language models are free boosters for biomedical imaging tasks,
Z. Lai, J. Wu, S. Chen, Y . Zhou, and N. Hovakimyan, “Residual-based language models are free boosters for biomedical imaging tasks,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 5086–5096
2024
-
[151]
Eformer: En- hanced transformer towards semantic-contour features of foreground for portraits matting,
Z. Wang, Q. Miao, Y . Xi, and P. Zhao, “Eformer: En- hanced transformer towards semantic-contour features of foreground for portraits matting,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 3880–3889
2024
-
[152]
Unified face attack detection with micro disturbance and a two-stage training strategy,
J. Yu, D. Lu, X. Shi, C. Qu, and F. Guo, “Unified face attack detection with micro disturbance and a two-stage training strategy,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 960– 969. 34
2024
-
[153]
Ma- avt: Modality alignment for parameter-efficient audio- visual transformers,
T. Mahmud, S. Mo, Y . Tian, and D. Marculescu, “Ma- avt: Modality alignment for parameter-efficient audio- visual transformers,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 7996– 8005
2024
-
[154]
Learning to rank patches for unbiased image re- dundancy reduction,
Y . Luo, Z. Chen, P. Zhou, Z. Wu, X. Gao, and Y .-G. Jiang, “Learning to rank patches for unbiased image re- dundancy reduction,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 22 831–22 840
2024
-
[155]
nnmobilenet: Rethinking cnn for retinopathy research,
W. Zhu, P. Qiu, X. Chen, X. Li, N. Lepore, O. M. Dumitrascu, and Y . Wang, “nnmobilenet: Rethinking cnn for retinopathy research,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 2285–2294
2024
-
[156]
Inceptionnext: When inception meets convnext,
W. Yu, P. Zhou, S. Yan, and X. Wang, “Inceptionnext: When inception meets convnext,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 5672–5683
2024
-
[157]
Towards explainable visual vessel recognition using fine-grained classification and image retrieval,
H. Karus, F. Schwenker, M. Munz, and M. Teutsch, “Towards explainable visual vessel recognition using fine-grained classification and image retrieval,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work- shops, June 2024, pp. 82–92
2024
-
[158]
Learning cnn on vit: A hybrid model to explicitly class-specific boundaries for domain adap- tation,
B. H. Ngo, N.-T. Do-Tran, T.-N. Nguyen, H.-G. Jeon, and T. J. Choi, “Learning cnn on vit: A hybrid model to explicitly class-specific boundaries for domain adap- tation,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 28 545–28 554
2024
-
[159]
Vim4path: Self-supervised vision mamba for histopathology images,
A. Nasiri-Sarvi, V . Q.-H. Trinh, H. Rivaz, and M. S. Hosseini, “Vim4path: Self-supervised vision mamba for histopathology images,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 6894–6903
2024
-
[160]
Efficient multitask dense predictor via binarization,
Y . Shang, D. Xu, G. Liu, R. R. Kompella, and Y . Yan, “Efficient multitask dense predictor via binarization,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 15 899–15 908
2024
-
[161]
Daff: Dual attentive feature fusion for multi- spectral pedestrian detection,
A. Althoupety, L.-Y . Wang, W.-C. Feng, and B. Rek- abdar, “Daff: Dual attentive feature fusion for multi- spectral pedestrian detection,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 2997–3006
2024
-
[162]
Skipplus: Skip the first few layers to better explain vision transformers,
F. Mehri, M. Fayyaz, M. S. Baghshah, and M. T. Pilehvar, “Skipplus: Skip the first few layers to better explain vision transformers,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 204–215
2024
-
[163]
Boosting the transferability of adversarial attack on vision trans- former with adaptive token tuning,
D. Ming, P. Ren, Y . Wang, and X. Feng, “Boosting the transferability of adversarial attack on vision trans- former with adaptive token tuning,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 20 887–20 918
2024
-
[164]
On the faithfulness of vision transformer explanations,
J. Wu, W. Kang, H. Tang, Y . Hong, and Y . Yan, “On the faithfulness of vision transformer explanations,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 10 936–10 945
2024
-
[165]
Identifying important group of pixels using interactions,
K. Sumiyasu, K. Kawamoto, and H. Kera, “Identifying important group of pixels using interactions,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 6017–6026
2024
-
[166]
Cdad- net: Bridging domain gaps in generalized category dis- covery,
S. B. Rongali, S. Mehrotra, A. Jha, M. H. N. C, S. Bose, T. Gupta, M. Singha, and B. Banerjee, “Cdad- net: Bridging domain gaps in generalized category dis- covery,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 2616–2626
2024
-
[167]
Large language models are good prompt learners for low-shot image classification,
Z. Zheng, J. Wei, X. Hu, H. Zhu, and R. Nevatia, “Large language models are good prompt learners for low-shot image classification,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 28 453–28 462
2024
-
[168]
Forgery-aware adaptive transformer for generalizable synthetic image detection,
H. Liu, Z. Tan, C. Tan, Y . Wei, J. Wang, and Y . Zhao, “Forgery-aware adaptive transformer for generalizable synthetic image detection,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 10 770– 10 780
2024
-
[169]
Question aware vision transformer for multimodal reasoning,
R. Ganz, Y . Kittenplon, A. Aberdam, E. Ben Avraham, O. Nuriel, S. Mazor, and R. Litman, “Question aware vision transformer for multimodal reasoning,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 13 861–13 871
2024
-
[170]
Mope-clip: Structured pruning for efficient vision-language models with module-wise pruning error metric,
H. Lin, H. Bai, Z. Liu, L. Hou, M. Sun, L. Song, Y . Wei, and Z. Sun, “Mope-clip: Structured pruning for efficient vision-language models with module-wise pruning error metric,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 27 370–27 380
2024
-
[171]
De- noised vision-language fusion guided by visual cues for e-commerce product search,
Z. Hu, S. Li, M. Du, A. Dhua, and D. Gray, “De- noised vision-language fusion guided by visual cues for e-commerce product search,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 1986–1996
2024
-
[172]
Tune-an-ellipse: Clip has potential to find what you want,
J. Xie, S. Deng, B. Li, H. Liu, Y . Huang, Y . Zheng, J. Schmidhuber, B. Ghanem, L. Shen, and M. Z. Shou, “Tune-an-ellipse: Clip has potential to find what you want,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 13 723–13 732
2024
-
[173]
Investigating compositional challenges in vision-language models for visual grounding,
Y . Zeng, Y . Huang, J. Zhang, Z. Jie, Z. Chai, and L. Wang, “Investigating compositional challenges in vision-language models for visual grounding,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 14 141–14 151
2024
-
[174]
Emer- gent open-vocabulary semantic segmentation from off- the-shelf vision-language models,
J. Luo, S. Khandelwal, L. Sigal, and B. Li, “Emer- gent open-vocabulary semantic segmentation from off- the-shelf vision-language models,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 4029–4040
2024
-
[175]
Clip as rnn: Segment countless visual concepts without training en- deavor,
S. Sun, R. Li, P. Torr, X. Gu, and S. Li, “Clip as rnn: Segment countless visual concepts without training en- deavor,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 13 171–13 182
2024
-
[176]
Improved visual grounding through self-consistent explanations,
R. He, P. Cascante-Bonilla, Z. Yang, A. C. Berg, and V . Ordonez, “Improved visual grounding through self-consistent explanations,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 13 095– 13 105
2024
-
[177]
Assessing brain-like characteristics of dnns with spatiotemporal features: A study based on the m ¨uller-lyer illusion,
H. Zhang, K. Matsuzaki, and S. Yoshida, “Assessing brain-like characteristics of dnns with spatiotemporal features: A study based on the m ¨uller-lyer illusion,” IEEE Access, vol. 12, pp. 147 192–147 208, 2024. 35
2024
-
[178]
A hybrid learning-architecture for men- tal disorder detection using emotion recognition,
J. Aina, O. Akinniyi, M. M. Rahman, V . Odero-Marah, and F. Khalifa, “A hybrid learning-architecture for men- tal disorder detection using emotion recognition,”IEEE Access, vol. 12, pp. 91 410–91 425, 2024
2024
-
[179]
Ask, attend, attack: An effective decision-based black-box targeted attack for image-to-text models,
Q. Zeng, Z. Wang, Y .-m. Cheung, and M. Jiang, “Ask, attend, attack: An effective decision-based black-box targeted attack for image-to-text models,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 105 819– 105 847
2024
-
[180]
Imdl- benco: A comprehensive benchmark and codebase for image manipulation detection & localization,
X. Ma, X. Zhu, L. Su, B. Du, Z. Jiang, B. Tong, Z. Lei, X. Yang, C.-M. Pun, J. Lv, and J. Zhou, “Imdl- benco: A comprehensive benchmark and codebase for image manipulation detection & localization,” in Proc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 134 591–134 613
2024
-
[181]
Real-time core-periphery guided vit with smart data layout selection on mobile devices,
Z. Shu, X. Yu, Z. Wu, W. Jia, Y . Shi, M. Yin, T. Liu, D. Zhu, and W. Niu, “Real-time core-periphery guided vit with smart data layout selection on mobile devices,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 95 744–95 763
2024
-
[182]
Visual fourier prompt tuning,
R. Zeng, C. Han, Q. Wang, C. Wu, T. Geng, L. Huang, Y . N. Wu, and D. Liu, “Visual fourier prompt tuning,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 5552–5585
2024
-
[183]
Lips are lying: Spotting the temporal inconsistency between audio and visual in lip-syncing deepfakes,
W. Liu, T. She, J. Liu, B. Li, D. Yao, Z. Liang, and R. Wang, “Lips are lying: Spotting the temporal inconsistency between audio and visual in lip-syncing deepfakes,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 91 131–91 155
2024
-
[184]
Testing semantic importance via betting,
J. Teneggi and J. Sulam, “Testing semantic importance via betting,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 76 450–76 499
2024
-
[185]
Linearly decomposing and recomposing vi- sion transformers for diverse-scale models,
S. Lin, M. Zhang, R. Chen, Q. Wang, X. Yang, and X. Geng, “Linearly decomposing and recomposing vi- sion transformers for diverse-scale models,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 33 188–33 212
2024
-
[186]
Learning low-rank feature for thorax disease classification,
Y . Wang, R. Goel, U. Nath, A. C. Silva, T. Wu, and Y . Yang, “Learning low-rank feature for thorax disease classification,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 117 133–117 163
2024
-
[187]
Adanca: Neural cellular automata as adaptors for more robust vision transformer,
Y . Xu, T. Zhang, and S. S ¨usstrunk, “Adanca: Neural cellular automata as adaptors for more robust vision transformer,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 25 709–25 746
2024
-
[188]
Decom- posing and interpreting image representations via text in vits beyond clip,
S. Balasubramanian, S. Basu, and S. Feizi, “Decom- posing and interpreting image representations via text in vits beyond clip,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 81 046–81 076
2024
-
[189]
Matryoshka query transformer for large vision-language models,
W. Hu, Z.-Y . Dou, L. H. Li, A. Kamath, N. Peng, and K.-W. Chang, “Matryoshka query transformer for large vision-language models,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 50 168–50 188
2024
-
[190]
Emr-merging: Tuning-free high- performance model merging,
C. Huang, P. Ye, T. Chen, T. He, X. Yue, and W. Ouyang, “Emr-merging: Tuning-free high- performance model merging,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 122 741–122 769
2024
-
[191]
Simvg: A simple framework for visual grounding with decoupled multi-modal fusion,
M. Dai, L. Yang, Y . Xu, Z. Feng, and W. Yang, “Simvg: A simple framework for visual grounding with decoupled multi-modal fusion,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 121 670–121 698
2024
-
[192]
A unified debiasing approach for vision-language models across modalities and tasks,
H. Jung, T. Jang, and X. Wang, “A unified debiasing approach for vision-language models across modalities and tasks,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 21 034–21 058
2024
-
[193]
Beyond accuracy: Ensuring correct predictions with correct rationales,
T. Li, M. Ma, and X. Peng, “Beyond accuracy: Ensuring correct predictions with correct rationales,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 43 164–43 188
2024
-
[194]
Neuro-vision to language: Enhancing brain recording-based visual reconstruction and language interaction,
G. Shen, D. Zhao, X. He, L. Feng, Y . Dong, J. Wang, Q. Zhang, and Y . Zeng, “Neuro-vision to language: Enhancing brain recording-based visual reconstruction and language interaction,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 98 083–98 110
2024
-
[195]
Text-guided attention is all you need for zero-shot robustness in vision- language models,
L. Yu, H. Zhang, and C. Xu, “Text-guided attention is all you need for zero-shot robustness in vision- language models,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 96 424–96 448
2024
-
[196]
Musictalk: A microservice approach for musical instrument recogni- tion,
Y .-B. Lin, C.-C. Cheng, and S.-C. Chiu, “Musictalk: A microservice approach for musical instrument recogni- tion,”IEEE Open J. Comput. Soc., vol. 5, pp. 612–623, 2024
2024
-
[197]
C2t-net: Channel- aware cross-fused transformer-style networks for pedes- trian attribute recognition,
D. C. Bui, T. V . Le, and B. H. Ngo, “C2t-net: Channel- aware cross-fused transformer-style networks for pedes- trian attribute recognition,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. Workshops, January 2024, pp. 351–358
2024
-
[198]
Limited data, unlimited potential: A study on vits augmented by masked autoencoders,
S. Das, T. Jain, D. Reilly, P. Balaji, S. Karmakar, S. Marjit, X. Li, A. Das, and M. S. Ryoo, “Limited data, unlimited potential: A study on vits augmented by masked autoencoders,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., January 2024, pp. 6878–6888
2024
-
[199]
Occlu- sion sensitivity analysis with augmentation subspace perturbation in deep feature space,
P. H. V . Valois, K. Niinuma, and K. Fukui, “Occlu- sion sensitivity analysis with augmentation subspace perturbation in deep feature space,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., January 2024, pp. 4829–4838
2024
-
[200]
Multi-view classification using hybrid fusion and mutual distillation,
S. Black and R. Souvenir, “Multi-view classification using hybrid fusion and mutual distillation,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., January 2024, pp. 270–280
2024
-
[201]
Clipag: Towards generator-free text-to-image generation,
R. Ganz and M. Elad, “Clipag: Towards generator-free text-to-image generation,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., January 2024, pp. 3843–3853
2024
-
[202]
Foundation model assisted weakly supervised semantic segmentation,
X. Yang and X. Gong, “Foundation model assisted weakly supervised semantic segmentation,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., January 2024, pp. 523–532
2024
-
[203]
Iterprime: Zero-shot referring image segmentation with iterative grad-cam refinement and primary word emphasis,
Y . Wang, J. Ni, Y . Liu, C. Yuan, and Y . Tang, “Iterprime: Zero-shot referring image segmentation with iterative grad-cam refinement and primary word emphasis,” in Proc. AAAI Conf. Artif. Intell., vol. 39, no. 8, 2025, pp. 9459–9468
2025
-
[204]
Wave: Weight templates for adaptive initialization of variable-sized models,
F. Feng, Y . Xie, J. Wang, and X. Geng, “Wave: Weight templates for adaptive initialization of variable-sized models,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern 36 Recognit., June 2025, pp. 4819–4828
2025
-
[205]
Overlock: An overview-first-look- closely-next convnet with context-mixing dynamic ker- nels,
M. Lou and Y . Yu, “Overlock: An overview-first-look- closely-next convnet with context-mixing dynamic ker- nels,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 128–138
2025
-
[206]
Fsfm: A generalizable face security foundation model via self-supervised facial representation learning,
G. Wang, F. Lin, T. Wu, Z. Liu, Z. Ba, and K. Ren, “Fsfm: A generalizable face security foundation model via self-supervised facial representation learning,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 24 364–24 376
2025
-
[207]
Tad- former: Task-adaptive dynamic transformer for efficient multi-task learning,
S. Baek, S. Lee, H. Jo, H. Choi, and D. Min, “Tad- former: Task-adaptive dynamic transformer for efficient multi-task learning,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 14 858–14 868
2025
-
[208]
Distilling spatially-heterogeneous distortion perception for blind image quality assessment,
X. Li, W. Nie, Y . Zhang, R. Hu, K. Li, X. Zheng, and L. Cao, “Distilling spatially-heterogeneous distortion perception for blind image quality assessment,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 2344–2354
2025
-
[209]
Clip-sla: Parameter- efficient clip adaptation for continuous sign language recognition,
S. Alyami and H. Luqman, “Clip-sla: Parameter- efficient clip adaptation for continuous sign language recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2025, pp. 4137– 4147
2025
-
[210]
Disentangling visual transformers: Patch-level interpretability for im- age classification,
G. Jeanneret, L. Simon, and F. Jurie, “Disentangling visual transformers: Patch-level interpretability for im- age classification,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2025, pp. 2704– 2714
2025
-
[211]
Enhancing vision transformer explainability using artificial astro- cytes,
N. Echevarrieta-Catalan, A. Ribas-Rodriguez, F. Ce- dron, O. Schwartz, and V . Aguiar-Pulido, “Enhancing vision transformer explainability using artificial astro- cytes,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2025, pp. 58–64
2025
-
[212]
Prompt-cam: Making vision transformers inter- pretable for fine-grained analysis,
A. Chowdhury, D. Paul, Z. Mai, J. Gu, Z. Zhang, K. S. Mehrab, E. G. Campolongo, D. Rubenstein, C. V . Stewart, A. Karpatne, T. Berger-Wolf, Y . Su, and W.-L. Chao, “Prompt-cam: Making vision transformers inter- pretable for fine-grained analysis,” inProc. IEEE/CVF Conf. Comput...
2025
-
[213]
Rethinking per- sonalized aesthetics assessment: Employing physique aesthetics assessment as an exemplification,
H. Zhong, S. He, A. Ming, and H. Ma, “Rethinking per- sonalized aesthetics assessment: Employing physique aesthetics assessment as an exemplification,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 2935–2944
2025
-
[214]
Goal: Global-local object alignment learning,
H. Choi, Y . K. Jang, and C. Eom, “Goal: Global-local object alignment learning,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 4070– 4079
2025
-
[215]
Separation of powers: On segregating knowledge from observation in llm-enabled knowledge-based visual question answering,
Z. Yang, Z. Tao, Q. Chen, L. Li, Y . Qi, A. van den Hengel, and Q. Huang, “Separation of powers: On segregating knowledge from observation in llm-enabled knowledge-based visual question answering,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 24 753–24 762
2025
-
[216]
Discovering fine-grained visual- concept relations by disentangled optimal transport concept bottleneck models,
Y . Xie, Z. Zeng, H. Zhang, Y . Ding, Y . Wang, Z. Wang, B. Chen, and H. Liu, “Discovering fine-grained visual- concept relations by disentangled optimal transport concept bottleneck models,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 30 199– 30 209
2025
-
[217]
Rankclip: Ranking-consistent language-image pretraining,
Y . Zhang, Z. Zhao, Z. Chen, Z. Feng, Z. Ding, and Y . Sun, “Rankclip: Ranking-consistent language-image pretraining,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 3874–3884
2025
-
[218]
On the complexity-faithfulness trade-off of gradient- based explanations,
A. Mehrpanah, M. Gamba, K. Smith, and H. Azizpour, “On the complexity-faithfulness trade-off of gradient- based explanations,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 3531–3541
2025
-
[219]
Dadm: Dual alignment of domain and modality for face anti-spoofing,
J. Yang, X. Lin, Z. Yu, L. Zhang, X. Liu, H. Li, X. Yuan, and X. Cao, “Dadm: Dual alignment of domain and modality for face anti-spoofing,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 12 045–12 056
2025
-
[220]
Deepshield: Fortifying deepfake video de- tection with local and global forgery analysis,
Y . Cai, J. Li, Z. Li, W. Chen, R. Lan, X. Xie, X. Luo, and G. Li, “Deepshield: Fortifying deepfake video de- tection with local and global forgery analysis,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 12 524– 12 534
2025
-
[221]
Modeling time-lapse trajectories to characterize cranberry growth,
R. John, A. Chihoub, R. Meegan, G. Sidelli, J. Ney- hart, P. Oudemans, and K. Dana, “Modeling time-lapse trajectories to characterize cranberry growth,” inProc. AAAI Conf. Artif. Intell. Workshops, October 2025, pp. 7208–7218
2025
-
[222]
Principles of visual tokens for efficient video understanding,
X. Hao, G. Li, S. N. Gowda, R. B. Fisher, J. Huang, A. Arnab, and L. Sevilla-Lara, “Principles of visual tokens for efficient video understanding,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 21 254–21 264
2025
-
[223]
Token activation map to visually explain multimodal llms,
Y . Li, H. Wang, X. Ding, H. Wang, and X. Li, “Token activation map to visually explain multimodal llms,” in Proc. AAAI Conf. Artif. Intell., October 2025, pp. 48– 58
2025
-
[224]
Ce-fam: Concept-based explanation via fusion of activation maps,
M. Kuroki and T. Yamasaki, “Ce-fam: Concept-based explanation via fusion of activation maps,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 1413–1422
2025
-
[225]
Semantic equitable clustering: A simple and effective strategy for clustering vision tokens,
Q. Fan, H. Huang, M. Chen, and R. He, “Semantic equitable clustering: A simple and effective strategy for clustering vision tokens,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 4019–4028
2025
-
[226]
Cross-modal ship re-identification via optical and sar imagery: A novel dataset and method,
H. Wang, S. Li, J. Yang, Y . Liu, Y . Lv, and Z. Zhou, “Cross-modal ship re-identification via optical and sar imagery: A novel dataset and method,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 7873–7883
2025
-
[227]
Context- aware academic emotion dataset and benchmark,
L. Zhao, J. Xuan, J. Lou, Y . Yu, and W. Yang, “Context- aware academic emotion dataset and benchmark,” in Proc. AAAI Conf. Artif. Intell., October 2025, pp. 13 859–13 868
2025
-
[228]
Mpbr: Multimodal progressive bidirectional reasoning for open-set fine- grained recognition,
J. Tan, P. Jing, Y . Zhu, and Y . Liu, “Mpbr: Multimodal progressive bidirectional reasoning for open-set fine- grained recognition,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 1282–1291
2025
-
[229]
Do vlms have bad eyes? diagnosing compositional failures via mechanistic interpretability,
A. V . Aravindan, A. Jha, and M. Kulkarni, “Do vlms have bad eyes? diagnosing compositional failures via mechanistic interpretability,” inProc. AAAI Conf. Artif. Intell. Workshops, October 2025, pp. 715–723
2025
-
[230]
Enhancing prompt generation with adaptive refinement for camouflaged object detection,
X. Chen, G. Ren, T. Dai, T. Stathaki, and H. Liu, 37 “Enhancing prompt generation with adaptive refinement for camouflaged object detection,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 20 672–20 682
2025
-
[231]
Effresnet-vit: A fusion-based con- volutional and vision transformer model for explainable medical image classification,
T. Hussain, H. Shouno, A. Hussain, D. Hussain, M. Is- mail, T. Hussain Mir, F. Rong Hsu, T. Alam, and S. Anonna Akhy, “Effresnet-vit: A fusion-based con- volutional and vision transformer model for explainable medical image classification,”IEEE Access, vol. 13, pp. 54 040–54 068, 2025
2025
-
[232]
Ropgcvit: A novel explainable vision transformer for retinopathy of prematurity diagnosis,
M. Yurdakul, K. Uyar, S. Tasdem ´ır, and I. Atabas, “Ropgcvit: A novel explainable vision transformer for retinopathy of prematurity diagnosis,”IEEE Access, vol. 13, pp. 77 064–77 079, 2025
2025
-
[233]
Transformer-based dme classification using reti- nal oct images without data augmentation: An evalu- ation of vit-b16 and vit-b32 with optimizer impact,
K. C. Pavithra, P. Kumar, M. Geetha, S. V . Bhandary, K. B. Ajitha Shenoy, G. Rao, S. Fernandes, and A. Tul- sani, “Transformer-based dme classification using reti- nal oct images without data augmentation: An evalu- ation of vit-b16 and vit-b32 with optimizer impact,” IEEE Ac...
2025
-
[234]
Squeeze-and- excitation vision transformer for lung nodule classifica- tion,
X. Xue, Y . Ma, W. Du, and Y . Peng, “Squeeze-and- excitation vision transformer for lung nodule classifica- tion,”IEEE Access, vol. 13, pp. 24 852–24 866, 2025
2025
-
[235]
Defax: A cross-attention fusion framework for robust and explainable deepfake detection,
M. Al-Imran, M. S. Sheikh, U. Kirtonia, N. T. Arthi, and S. Ripon, “Defax: A cross-attention fusion framework for robust and explainable deepfake detection,”IEEE Access, vol. 13, pp. 213 962–213 979, 2025
2025
-
[236]
A comparative study of performance of various deep learning models and their explainability in detection of glaucoma,
N. K. Jisy, S. Radhika, S. Senthil, and M. B. Srinivas, “A comparative study of performance of various deep learning models and their explainability in detection of glaucoma,”IEEE Access, vol. 13, pp. 192 891–192 905, 2025
2025
-
[237]
Exploring vision transformers and explainable ai for enhanced artefact classification in esophageal endoscopic images,
P. Bissoonauth-Daiboo, M. Muzzammil Auzine, M. I. Khan, F. Alshannaq, T. Saba, X. Gao, and M. Heenaye- Mamode Khan, “Exploring vision transformers and explainable ai for enhanced artefact classification in esophageal endoscopic images,”IEEE Access, vol. 13, pp. 176 221–176 244, 2025
2025
-
[238]
Detection of carbapenem resistance in klebsiella pneumoniae using vision transformers and maldi-tof proteomic profiles,
V . S. Mar ´ın, J. P. V . Pulgarin, S. A. Holgu ´ın-Garc´ıa, I. L. M. Figueroa, A. C. E. Maldonado, J. C. G. Vel´asquez, R. Tabares-Soto, M. A. Bravo-Ort ´ız, and G. J. Fern ´andez, “Detection of carbapenem resistance in klebsiella pneumoniae using vision transformers and mald...
2026
-
[239]
Schizophrenia seg- mentation from smri data: A fusion approach using u-net and vision transformer-based grad-cam,
T. Sarveswaran and V . Rajangam, “Schizophrenia seg- mentation from smri data: A fusion approach using u-net and vision transformer-based grad-cam,”IEEE Access, vol. 13, pp. 184 376–184 397, 2025
2025
-
[240]
Data-efficient wheat disease detection using shifted window transformer: Enhancing accuracy, sustainabil- ity, and global food security,
M. Khubaib, T. Kehkashan, M. Abdelhaq, M. A. Khan, M. Zaman, I. Ashraf, A. Rehman, and A. Akhunzada, “Data-efficient wheat disease detection using shifted window transformer: Enhancing accuracy, sustainabil- ity, and global food security,”IEEE Trans. Consumer Electron., vol. 7...
2025
-
[241]
On which data distribution (synthetic or real) we should rely for soft biometric classification,
M. R. A, A. Kumar, and A. Agarwal, “On which data distribution (synthetic or real) we should rely for soft biometric classification,” inProc. Winter Conf. Appl. Comput. Vis., February 2025, pp. 6238–6247
2025
-
[242]
Swincnn+oe: A swin transformer and cnn architecture for breast histopathology classification with ood and grad-cam integration,
D. Opoku, K. Owusu-Agyemang, J. B. Hayfron- Acquah, and R.-M. O. M. Gyening, “Swincnn+oe: A swin transformer and cnn architecture for breast histopathology classification with ood and grad-cam integration,”IEEE Access, vol. 14, pp. 3897–3909, 2026
2026
-
[243]
Resvit- solar: A hybrid residual-transformer for photovoltaic panel fault and cleanliness detection,
N. Amin, Y .-W. Kim, C. Kang, and Y .-C. Byun, “Resvit- solar: A hybrid residual-transformer for photovoltaic panel fault and cleanliness detection,”IEEE Access, vol. 14, pp. 24 650–24 664, 2026
2026
-
[244]
Locally explaining predic- tion behavior via gradual interventions and measuring property gradients,
N. Penzel and J. Denzler, “Locally explaining predic- tion behavior via gradual interventions and measuring property gradients,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., March 2026, pp. 7398–7408
2026
-
[245]
Ui- styler: Ultrasound image style transfer with class-aware prompts for cross-device diagnosis using a frozen black- box inference network,
N.-T. Do-Tran, N.-H.-L. Le, and C.-C. Huang, “Ui- styler: Ultrasound image style transfer with class-aware prompts for cross-device diagnosis using a frozen black- box inference network,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., March 2026, pp. 2765–2774
2026
-
[246]
Overcoming fine-grained visual challenges in animal re-identification via semantic feature alignment,
Y . Wu, D. Zhao, Y . Li, M. Alajas, A. S. Glen, J. Zhang, G. Dobbie, D. Wilson, and Y . S. Koh, “Overcoming fine-grained visual challenges in animal re-identification via semantic feature alignment,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., March 2026, pp. 371– 381
2026
-
[247]
nnmobilenet++: Towards efficient hybrid networks for retinal image analysis,
X. Li, W. Zhu, X. Dong, H. Wang, Y . Xiong, O. Dumi- trascu, and Y . Wang, “nnmobilenet++: Towards efficient hybrid networks for retinal image analysis,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. Workshops, March 2026, pp. 282–292
2026
-
[248]
Neurobridge: Few-shot cross-modal neuron re-identification via dual- channel deep metric learning,
W. Li, M. Liao, L. Cai, and A. Li, “Neurobridge: Few-shot cross-modal neuron re-identification via dual- channel deep metric learning,” inProc. IEEE/CVF Win- ter Conf. Appl. Comput. Vis., March 2026, pp. 8670– 8679
2026
-
[249]
Dream: Dynamic prompts and guidedmix for efficient continual adapta- tion of visual-language models,
E. Chee, M. L. Lee, and W. Hsu, “Dream: Dynamic prompts and guidedmix for efficient continual adapta- tion of visual-language models,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., March 2026, pp. 5853–5863
2026
-
[250]
Fairvlm: Enhancing fairness and prompt sensitivity in vision language models for medical image segmenta- tion,
M. M. Rahman, S. Rahman, S. Bhatt, and M. Faezipour, “Fairvlm: Enhancing fairness and prompt sensitivity in vision language models for medical image segmenta- tion,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., March 2026, pp. 7450–7460
2026
-
[251]
Afragent : An adaptive feature renormal- ization based high resolution aware gui agent,
N. Anand, R. Jain, S. Patnaik, B. Krishnamurthy, and M. Sarkar, “Afragent : An adaptive feature renormal- ization based high resolution aware gui agent,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., March 2026, pp. 1147–1158
2026
-
[252]
Surgxbench: Explainable vision-language model benchmark for surgery,
J. Cheng, X. Zhao, S. Liu, X. Yu, R. Prakash, P. J. Codd, J. E. Katz, and S. Lin, “Surgxbench: Explainable vision-language model benchmark for surgery,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., March 2026, pp. 8188–8198
2026
-
[253]
Revisiting vision- language foundations for no-reference image quality assessment,
A. Yadav, T. D. Huy, and L. Liu, “Revisiting vision- language foundations for no-reference image quality assessment,” inProc. IEEE/CVF Winter Conf. Appl. 38 Comput. Vis., March 2026, pp. 5416–5425
2026
-
[254]
Grad-CAM++: Generalized Gradient- Based Visual Explanations for Deep Convolutional Networks ,
A. Chattopadhay, A. Sarkar, P. Howlader, and V . N. Bal- asubramanian, “ Grad-CAM++: Generalized Gradient- Based Visual Explanations for Deep Convolutional Networks ,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., Mar. 2018, pp. 839–847
2018
-
[255]
Layercam: Exploring hierarchical class activa- tion maps for localization,
P.-T. Jiang, C.-B. Zhang, Q. Hou, M.-M. Cheng, and Y . Wei, “Layercam: Exploring hierarchical class activa- tion maps for localization,”IEEE Trans. Image Process., vol. 30, pp. 5875–5888, 2021
2021
-
[256]
Score-CAM: Score-Weighted Visual Explanations for Convolutional Neural Networks ,
H. Wang, Z. Wang, M. Du, F. Yang, Z. Zhang, S. Ding, P. Mardziel, and X. Hu, “Score-CAM: Score-Weighted Visual Explanations for Convolutional Neural Networks ,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. Workshops, Jun. 2020, pp. 111–119
2020
-
[257]
ScoreCAM++: Gated Score-Weighted Visual Explanations for CNNs ,
S. Mitra, A. Sukul, S. K. Roy, P. Singh, and V . K. Verma, “ ScoreCAM++: Gated Score-Weighted Visual Explanations for CNNs ,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, Jun. 2025, pp. 2691–2700
2025
-
[258]
Ablation-CAM: Visual Explanations for Deep Convolutional Network via Gradient-free Localization ,
S. Desai and H. G. Ramaswamy, “ Ablation-CAM: Visual Explanations for Deep Convolutional Network via Gradient-free Localization ,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., Mar. 2020, pp. 972– 980
2020
-
[259]
Full-gradient representation for neural network visualization,
S. Srinivas and F. Fleuret, “Full-gradient representation for neural network visualization,” inProc. Adv. Neural Inf. Process. Syst., vol. 32, 2019
2019
-
[260]
Axiom-based grad-cam: Towards accurate visualiza- tion and explanation of cnns,
R. Fu, Q. Hu, X. Dong, Y . Guo, Y . Gao, and B. Li, “Axiom-based grad-cam: Towards accurate visualiza- tion and explanation of cnns,” inProc. The Brit. Mach. Vis. Conf., 2020
2020
-
[261]
Swin transformer: Hierarchical vision transformer using shifted windows,
Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” inProc. AAAI Conf. Artif. Intell., October 2021, pp. 10 012– 10 022
2021
-
[262]
BEit: BERT pre-training of image transformers,
H. Bao, L. Dong, S. Piao, and F. Wei, “BEit: BERT pre-training of image transformers,” inProc. Int. Conf. Learn. Represent., 2022, Art. no. p-BhZSz59o4
2022
-
[263]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in Proc. Int. Conf. Mach. Learn., vol. 139, Jul 2021, pp. 8748–8763
2021
-
[264]
Winsor- cam: Human-tunable visual explanations from deep net- works via layer-wise winsorization,
C. Wall, L. Wang, R. Rizk, and K. Santosh, “Winsor- cam: Human-tunable visual explanations from deep net- works via layer-wise winsorization,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026
2026
-
[265]
Explaining the behavior of neuron activations in deep neural networks,
L. Wang, C. Wang, Y . Li, and R. Wang, “Explaining the behavior of neuron activations in deep neural networks,” Ad Hoc Networks, vol. 111, p. 102346, 2021
2021
-
[266]
Bridging in- terpretability and robustness using lime-guided model refinement,
N. Nayyem, A. Rakin, and L. Wang, “Bridging in- terpretability and robustness using lime-guided model refinement,”arXiv preprint arXiv:2412.18952, 2024
2024 arXiv
-
[267]
Explainability-driven defense: Grad-cam-guided model refinement against adversarial threats,
L. Wang, I. I. Uddin, X. Qin, Y . Zhou, and K. Santosh, “Explainability-driven defense: Grad-cam-guided model refinement against adversarial threats,” inProceedings of the AAAI Symposium Series (AAAI) 2025, vol. 6, no. 1, 2025, pp. 49–57
2025
-
[268]
Explainability-guided defense: Attribution-aware model refinement against adversarial data attacks,
L. Wang, M. N. Nayyem, A. Al Rakin, K. Santosh, C. Zhang, and Y . Zhou, “Explainability-guided defense: Attribution-aware model refinement against adversarial data attacks,” in2025 IEEE International Conference on Data Mining (ICDM). IEEE, 2025, pp. 1585–1592
2025
-
[269]
Expert-guided explainable few-shot learning for medical image diag- nosis,
I. I. Uddin, L. Wang, and K. Santosh, “Expert-guided explainable few-shot learning for medical image diag- nosis,” inMICCAI Workshop on Data Engineering in Medical Imaging 2025. Springer Nature Switzerland, 2025, pp. 95–104
2025
-
[270]
Expert-guided explainable few-shot learning with active sample se- lection for medical image analysis,
L. Wang, I. I. Uddin, and K. Santosh, “Expert-guided explainable few-shot learning with active sample se- lection for medical image analysis,”IEEE Journal of Biomedical and Health Informatics, 2026
2026
-
[271]
Learning to select like humans: Explainable active learning for medical imaging,
I. I. Uddin, L. Wang, X. Qin, Y . Zhou, and K. Santosh, “Learning to select like humans: Explainable active learning for medical imaging,” in2026 IEEE Confer- ence on Artificial Intelligence (CAI). IEEE, 2026, pp. 458–463
2026
-
[272]
Bridging symmetry and robustness: On the role of equivariance in enhancing adversarial robustness,
L. Wang, I. I. Uddin, C. Zhang, X. Qin, and Y . Zhou, “Bridging symmetry and robustness: On the role of equivariance in enhancing adversarial robustness,” Advances in Neural Information Processing Systems (NeurIPS), vol. 38, pp. 159 102–159 129, 2025
2025
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.