Pith. sign in

REVIEW 2 major objections 6 minor 251 references

Grad-CAM for Vision Transformers: A Systematic Taxonomy and Audit of Methodological Ambiguity in Explainable AI

T0 review · 2 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read The paper claims that most applications of Grad-CAM to Vision Transformers leave the adaptation mathematically unspecified, so the published heatmaps cannot be reproduced or evaluated as scientific evidence.

desk verdict A genuinely useful taxonomy of ViT Grad-CAM adaptations, attached to an audit whose headline number (58% underspecification) is built on citation patterns rather than the reporting-detail coding the authors say they performed. read the letter →

arxiv 2608.05258 v1 pith:NPFPFHMU submitted 2026-08-05 cs.CV

classification cs.CV
keywords Grad-CAMVisionTransformerexplainableAImethodologicalauditreproducibilityattentionmapssaliencytaxonomy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Grad-CAM is a widely used heatmap method designed for convolutional networks, where spatial feature maps and channel dimensions have fixed meanings. Vision Transformers instead store image information as tokens, attention maps, residual streams, and cross-modal interactions, so there is no single obvious feature map for Grad-CAM to use. This paper argues that most of the 175 surveyed papers that apply Grad-CAM to ViTs treat the adaptation as trivial and never specify which transformer representation, gradient target, token handling, aggregation, or spatial reshaping produced their heatmaps. If the audit is right, roughly 58% of these papers cite only CNN-era Grad-CAM and report heatmaps that cannot be reproduced from the publication alone. The paper's contribution is a descriptive taxonomy that names the choice points, such as $F_1$--$F_6$ feature tensors, attention-map variants, and cross-attention variants, so future work can state exactly what "Grad-CAM on a ViT" means.

What carries the argument

The load-bearing object is a descriptive taxonomy of six feature tensors inside a transformer block: the layer-normalized tokens before attention ($F_1$), the softmax attention matrix ($F_2$), head-wise attention outputs ($F_3$), the layer-normalized tokens before the MLP ($F_4$), the MLP output ($F_5$), and the post-feedforward residual output ($F_6$), plus cross-attention variants for vision-language models. Against this grid the paper defines Flat-GCAM, a one-dimensional token-level attribution vector that can be reshaped into the patch grid, and shows that each combination of layer, feature tensor, gradient target, and head aggregation produces a different heatmap. This taxonomy is what lets the audit detect that most papers do not pin down a single member of the family.

What would settle it

Re-read the methods sections of the 101 papers categorized as citing only non-ViT Grad-CAM and count how many actually specify the feature tensor, gradient target, token handling, head aggregation, and reshaping in equations or prose; if a substantial share do, the 58% underspecification figure overstates the gap.

Watch

Extended reading notes

Core claim

The paper's central claim is that "Grad-CAM on a ViT" is not one defined procedure but a family of procedures indexed by the choice of feature tensor (layer-normalized tokens, softmax attention, attention outputs, MLP activations, residual outputs), the gradient target, whether [CLS]-token gradients or patch-token-average gradients are used, how attention heads are combined, and how token-level scores are reshaped into an image grid. The authors show that for a single ViT-B/16 model and one image this family contains well over 100 plausible heatmaps, and that different choices produce qualitatively different localizations. They then audit 175 papers and find that 101 cite only the original CNN-based Grad-CAM or CNN-era variants, 17 cite only an implementation repository, 13 cite nothing, and only 26 justify the adaptation directly or inherit a ViT-specific method. The conclusion is that most published ViT Grad-CAM figures are underspecified and therefore cannot serve as reproducible scientific evidence.

Load-bearing premise

The audit's headline numbers assume that a paper's citation pattern reveals whether its text specifies the ViT adaptation, because the reported coding of "reporting detail" is not presented in the paper itself.

Editorial extensions

If this is right

  • If the paper is right, the majority of surveyed ViT Grad-CAM heatmaps cannot be reconstructed from the publication, so reviewers should treat "Grad-CAM on a ViT" as an underspecified phrase unless the adaptation is described or clearly inherited.
  • The six feature families imply that two papers using the same Grad-CAM label may be explaining different tensors, so cross-paper comparison of heatmaps is unsafe without specifying the variant.
  • Because final-layer token-weighted maps can become nearly blank, authors who omit the layer and weighting strategy could select the most visually favorable map from a large space of alternatives.
  • In vision-language models, cross-attention maps are text-conditioned per-query-token maps, so interpreting them as generic visual saliency is not justified under the taxonomy.
  • The 13 identified ViT-specific adaptations are unevenly adopted, which means the methodological ambiguity is current rather than historical.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the same ambiguity likely affects other CNN-born explanation methods when ported to transformers, because any method that assumes a spatial-channel tensor inherits the same implementation space.
  • The taxonomy could be used as a minimal reporting checklist in venue guidelines: requiring authors to name the exact tensor, the gradient target, and the aggregation axis would make many heatmaps reproducible without changing any method.
  • The qualitative examples suggest that comparative claims about ViT models can be contaminated by Grad-CAM implementation choice, making model differences and attribution-variant differences hard to separate.
  • A testable extension would be to take a sample of the 101 papers categorized as underspecified, implement the most plausible reading of each paper's text, and compare the resulting heatmap to the published figure.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper develops a descriptive taxonomy of Grad-CAM adaptations to vision transformers, distinguishing six feature extraction locations (F1-F6), gradient targets, token handling, aggregation strategies, and reshaping choices, and extends this to cross-attention in vision-language models. It then audits 175 papers that apply Grad-CAM or related methods to ViTs, reporting that 101 papers (58%) cite only CNN-based Grad-CAM and that only 26 papers either justify the adaptation or cite ViT-specific prior work. The authors conclude that most papers treat Grad-CAM on ViTs as a trivial extension and fail to provide a full mathematical or implementation-level account, and they propose reporting recommendations for authors and reviewers.

Significance. The taxonomy and the qualitative demonstration that different feature locations, gradient targets, and aggregation choices yield substantially different heatmaps (Figs. 2, 3, and S1) are useful contributions to the explainable AI literature. The paper is careful to frame the taxonomy as descriptive rather than prescriptive, and it provides a detailed supplementary analysis of 13 ViT-specific Grad-CAM adaptations (S-V). The main quantitative claim, however, rests on a citation-pattern proxy rather than on the 'reporting detail' variable that the authors state they coded, so the headline 58% figure is not yet supported by the presented data. If the authors present the reporting-detail coding and align their claims with what that coding shows, the paper would provide a valuable evidence base for reproducibility standards in transformer interpretability research.

major comments (2)
  1. [III-B and IV-A] The paper states in Section III-B that all 175 papers were coded for 'reporting detail,' but this variable is never presented in the results or in the supplementary coding table (Table S1). The 58% figure in Section IV-A is based entirely on the citation category 'Only References Non-ViT Grad-CAM,' which conflates citation behavior with methodological specificity. A paper citing only Selvaraju et al. could still specify its chosen feature tensor, gradient target, token handling, and aggregation in the main text, while a paper citing a ViT-specific adaptation could leave those choices implicit. Therefore the conclusion in Section VI that 'most papers do not provide a full mathematical or implementation-level account' is not established by the reported data. The authors should report the distribution of their 'reporting detail' coding (e.g., a cross-tabulation with citation category) or re-code the corpus for the explicit presence of the five required elements (feature location, gradient target, token handling, aggregation, reshaping), and then revise the 58% claim and the associated language accordingly.
  2. [II-B4, Eqs. (8)-(9) and (15)-(16)] There is a dimensional inconsistency in the definition of the multi-headed feature maps and their use in the linear combination equations. Eq. (8) defines \hat{F}^{(h,l)}_2 as a vector over patch tokens in R^{N-1}, and Eq. (9) defines \hat{F}^{(h,l)}_3 as a vector over embedding dimensions in R^{d_h}; however, Eqs. (15)-(16) treat \hat{F}^{(h,l)}_{m,i,k} as a matrix indexed by both token i and channel k. For m=2, no channel index remains after extracting the CLS row, and for m=3, the CLS row has no token index. The intended object appears to be the full token-by-channel matrix (e.g., F^{(h,l)}_3 with the CLS row removed, or F^{(h,l)}_2 with the CLS row and column removed), not the CLS-conditioned vectors defined in Eqs. (8)-(9). Please clarify the notation and ensure the equations are dimensionally consistent, since the taxonomy is a central contribution of the paper.
minor comments (6)
  1. [II-B3] The text contains the typo 'VITs' in the sentence beginning 'In VITs, feature representations instead...' and should read 'ViTs'.
  2. [II-B4] The sentence 'This can be formalated as follows' contains a misspelling of 'formulated'.
  3. [References] The reference list includes a large number of works by the authors (Refs. [2]-[36]) that appear unrelated to Grad-CAM or vision transformers; please verify that all cited works are actually relevant to the statements they support.
  4. [Table S1] The 'Notes' column contains subjective annotations such as 'possibly pytorch-grad-cam' and 'Unclear structure'; these should be formally defined in the codebook so that the coding is transparent and reproducible.
  5. [Fig. 6 and Fig. 7] The figure legends for the pie charts should clarify that the counts in categories such as '[53] (9)' include both the originating paper and the papers that explicitly adopt it, to avoid confusion about the 26-paper total.
  6. [Eq. (18)] The normalization formula in Eq. (18) indicates the final heatmap lies in [0,1]^{H x W}, but the preceding heatmap L^{(l,c)}_Grad-CAM is at patch resolution; the upsampling step is described in the text but should also be reflected in the equation or its caption.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the taxonomy is descriptive and the audit claim rests on an acknowledged citation-based inference, not on a constructional equivalence.

full rationale

This is a literature audit; its central claim is an empirical generalization about 175 surveyed papers, not a derivation from its own inputs. The taxonomy (Eqs. 1-18) is defined independently of the audit results and is explicitly descriptive, and the qualitative demonstrations (Figs. 2-3, S1) are self-contained computations on standard external models (ViT-B/16, BLIP), so the illustrative material is not benchmarked against or derived from the corpus findings. The 58% figure in Section IV-A is computed from an explicit citation category ('Only References Non-ViT Grad-CAM', Fig. 6), and the conclusion that underspecification follows from that citation pattern is an inference, acknowledged in S-IV as interpretive ('ambiguous cases were treated as ambiguous'), not an equivalence forced by construction; a paper citing only Selvaraju et al. could still specify its adaptation, so the reduction is defeasible rather than definitional. The skeptical concern that Section III-B's coded 'reporting detail' variable is never presented is a real evidentiary gap that weakens support for the headline percentage, but it is a measurement-validity issue, not circularity. Self-citation is present but non-load-bearing: the introduction's CNN-usage citation list includes the authors' own prior works ([2]-[36]), and the authors' own papers (supplementary refs [189]-[193], e.g., Winsor-CAM) appear in the audited corpus and are classified as explicit adopters of ALBEF's method, which if anything biases the reported percentages conservatively rather than forcing the conclusion. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported, and no ansatz is smuggled in via citation. Score 1 reflects only the mild self-inclusion in the corpus and the unclearly bridged citation-to-reporting inference, both of which are proportionality and correctness concerns rather than circular steps.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper relies on standard transformer mathematics and an audit methodology. Its main assumptions are that citation behavior is a proxy for methodological detail, that the venue-restricted corpus is representative, and that token ordering supports spatial reshaping. No free parameters are fitted and no new entities are postulated.

assumptions (3)
  • domain assumption Papers that cite only the original CNN Grad-CAM paper or CNN variants do not specify a reproducible ViT adaptation.
    The audit's central inference, stated in Section IV-A and the conclusion, treats citation categories as evidence of methodological underspecification, even though a paper could in principle provide full detail without citing ViT-specific work.
  • domain assumption The surveyed corpus, restricted to CVF Open Access, NeurIPS Proceedings, and IEEE Access, is representative of the broader literature where these claims are made.
    The paper acknowledges this is a bounded snapshot in Section III-B and S-IV, but the central percentages are reported as characterizing the field.
  • domain assumption Token indices in ViTs follow a raster-scan patch ordering so that discarding the CLS token and reshaping yields a spatially valid map.
    Section II-B3, Eq. 17 relies on this ordering for the spatial reconstruction step used throughout the taxonomy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Grad-CAM for Vision Transformers: A Systematic Taxonomy and Audit of Methodological Ambiguity in Explainable AI." pith.science (2026). https://pith.science/paper/NPFPFHMU

@misc{pith2026260805258,
  author       = {Pith},
  title        = {Pith review of: Grad-CAM for Vision Transformers: A Systematic Taxonomy and Audit of Methodological Ambiguity in Explainable AI},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NPFPFHMU}},
  note         = {Machine review of arXiv:2608.05258}
}
read the original abstract

Gradient-weighted Class Activation Mapping (Grad-CAM) is widely used to visualize model decisions, but it was originally formulated for convolutional neural networks, where spatial feature maps and channel dimensions have clear architectural meanings. Vision Transformers (ViTs) do not provide the same structure, instead representing images through tokens, attention, residual streams, and multimodal interactions. This paper presents a systematic taxonomy and literature audit of how Grad-CAM and related methods are adapted, justified, and reported for ViT-based architectures. From an initial search of more than 550 papers, we identify 175 papers that apply Grad-CAM or Grad-CAM-adjacent methods to ViTs. We find that most papers do not provide a full mathematical or implementation-level account of how Grad-CAM is adapted to transformer representations. To characterize this gap, we introduce a descriptive taxonomy of ViT Grad-CAM adaptations that makes explicit the feature locations, gradient targets, spatial reconstruction steps, and aggregation choices that are often left implicit. This taxonomy is not intended to prescribe a single correct adaptation, but to clarify the range of methodological choices being made. The study shows that Grad-CAM on ViTs is often treated as a trivial extension of CNN-based Grad-CAM, despite requiring nontrivial choices that affect rigor, reproducibility, and interpretation.

Figures

Figures reproduced from arXiv: 2608.05258 by the authors.

Figure 1
Figure 1. Feature extraction locations for adapting Grad-CAM to transformer-based vision models. The top panel shows a standard ViT self-attention encoder [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Grad-CAM visualizations for a standard ViT-B/16 [ [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Grad-CAM visualizations for a standard ViT-B/16 [ [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (3 more)
Figure 5
Figure 5. Figure 5: Distribution of surveyed literature by publication year. The line plot [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 7
Figure 7. Figure 7: Distribution of literature according to their methodology foundation [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]
Figure 9
Figure 9. Figure 9: Distribution of surveyed literature across primary task paradigms. [PITH_FULL_IMAGE:figures/full_fig_p012_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

251 extracted references · 76 canonical work pages

  1. [1]

    Grad-cam: Visual explanations from deep networks via gradient-based localization,

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proc. AAAI Conf. Artif. Intell., 2017, pp. 618–626

  2. [2]

    Representation learning and na- ture encoded fusion for heterogeneous sensor networks,

    L. Wang and Q. Liang, “Representation learning and na- ture encoded fusion for heterogeneous sensor networks,” IEEE Access, vol. 7, pp. 39 227–39 235, 2019

  3. [3]

    Congestion aware dynamic user association in heterogeneous cellular network: A stochastic decision approach,

    L. Wang, W. Chen, and J. Li, “Congestion aware dynamic user association in heterogeneous cellular network: A stochastic decision approach,” in2014 IEEE Interna- tional Conference on Communications (ICC). IEEE, 2014, pp. 2636–2640

  4. [4]

    Enhanced robustness by symmetry enforcement,

    L. Wang, A. Ghimire, K. Santosh, Z. Zhang, and X. Li, “Enhanced robustness by symmetry enforcement,” in IEEE Conference on Artificial Intelligence (IEEE CAI) 2024, 2024

  5. [5]

    Partial interference alignment for heterogeneous cellular networks,

    L. Wang and Q. Liang, “Partial interference alignment for heterogeneous cellular networks,”IEEE Access, vol. 6, pp. 22 592–22 601, 2018

  6. [6]

    Optimization for user centric massive mimo cell free networks via large system analysis,

    ——, “Optimization for user centric massive mimo cell free networks via large system analysis,” in2016 IEEE Global Communications Conference (GLOBE- COM). IEEE, 2016, pp. 1–1

  7. [8]

    Exploration vs exploitation for distributed channel access in cognitive radio networks: A multi-user case study,

    L. Wang, X. Chen, Z. Zhao, and H. Zhang, “Exploration vs exploitation for distributed channel access in cognitive radio networks: A multi-user case study,” in2011 11th International Symposium on Communications & Infor- mation Technologies (ISCIT). IEEE, 2011, pp. 360–365

  8. [9]

    Deep reinforcement learning based computation offloading for mobility-aware edge computing,

    M. Shi, R. Wang, E. Liu, Z. Xu, and L. Wang, “Deep reinforcement learning based computation offloading for mobility-aware edge computing,” inInternational con- ference on communications and networking in china. Springer International Publishing Cham, 2019, pp. 53– 65

Show all 251 references
  1. [10]

    Performance analysis of co- operative multicell precoding with global csi and local individual csi in the large dimensional regime,

    L. Wang and Q. Liang, “Performance analysis of co- operative multicell precoding with global csi and local individual csi in the large dimensional regime,”IEEE Transactions on Vehicular Technology, vol. 67, no. 4, pp. 3229–3238, 2017

  2. [11]

    Low complexity optimization for user centric cel- lular networks via large dimensional analysis,

    ——, “Low complexity optimization for user centric cel- lular networks via large dimensional analysis,”Physical Communication, vol. 25, pp. 412–419, 2017

  3. [12]

    Improving robustness of deep neural networks via large-difference transformation,

    L. Wang, C. Wang, Y . Li, and R. Wang, “Improving robustness of deep neural networks via large-difference transformation,”Neurocomputing, vol. 450, pp. 411–419, 2021

  4. [13]

    Looking beyond content: Modeling and detection of fake news from a social context perspective

    K. Xiao, L. Wang, A. Gupta, and X. Qin, “Looking beyond content: Modeling and detection of fake news from a social context perspective.” inProceedings of the 55th Hawaii International Conference on System Sciences 2022, 2022, pp. 1–10

  5. [14]

    Large system analysis for densification of cellular networks with massive mimo,

    L. Wang and Q. Liang, “Large system analysis for densification of cellular networks with massive mimo,” in2016 IEEE Conference on Computer Communications Workshops (INFOCOM WKSHPS). IEEE, 2016, pp. 956–961

  6. [15]

    Collabora- tive spectrum sharing based on information pooling for cognitive radio networks with channel heterogeneity,

    L. Wang, X. Chen, Z. Zhao, and H. Zhang, “Collabora- tive spectrum sharing based on information pooling for cognitive radio networks with channel heterogeneity,” in 2011 11th International Symposium on Communications & Information Technologies (ISCIT). IEEE, 2011, pp. 483–488

  7. [16]

    Dense cross-connected ensemble convolutional neural networks for enhanced model robustness,

    L. Wang, X. Li, and Z. Zhang, “Dense cross-connected ensemble convolutional neural networks for enhanced model robustness,”arXiv preprint arXiv:2412.07022, 2024

  8. [17]

    Information theory and represen- tation learning inspired multimodal data fusion,

    L. Wang and Y . Li, “Information theory and represen- tation learning inspired multimodal data fusion,”IEEE MMTC Frontier, 2019

  9. [19]

    Enhanc- ing adversarial robustness of deep neural networks through supervised contrastive learning,

    L. Wang, N. Nayyem, and A. Rakin, “Enhanc- ing adversarial robustness of deep neural networks through supervised contrastive learning,”arXiv preprint arXiv:2412.19747, 2024

  10. [21]

    Multi-scale unrectified push-pull with channel attention for enhanced corruption robustness,

    R. N. Ranabhat, L. Wang, X. Qin, Y . Zhou, and K. San- tosh, “Multi-scale unrectified push-pull with channel attention for enhanced corruption robustness,” inPro- ceedings of the AAAI Symposium Series 2025, vol. 6, no. 1, 2025, pp. 34–41

  11. [22]

    Expert-guided ex- plainable few-shot learning for medical image diagnosis,

    I. I. Uddin, L. Wang, and K. Santosh, “Expert-guided ex- plainable few-shot learning for medical image diagnosis,” inMICCAI Workshop on Data Engineering in Medical Imaging 2025. Springer Nature Switzerland, 2025, pp. 95–104

  12. [23]

    Ecologically valid benchmarking and adaptive attention: Scalable marine bioacoustic monitoring,

    N. R. Rasmussen, R. Rizk, L. Wang, and K. Santosh, “Ecologically valid benchmarking and adaptive attention: Scalable marine bioacoustic monitoring,”arXiv preprint 17 arXiv:2509.04682, 2025

  13. [24]

    Shape-aware thoracic edge map chest x- ray representation for pulmonary abnormality screening,

    S. Chataut, A. Ghimire, A. Thakur, L. Wang, and K. Santosh, “Shape-aware thoracic edge map chest x- ray representation for pulmonary abnormality screening,” inInternational Conference on DATA ANALYTICS & LEARNING. Springer, 2024, pp. 209–220

  14. [25]

    Expert-guided ex- plainable few-shot learning with active sample selection for medical image analysis,

    L. Wang, I. I. Uddin, and K. Santosh, “Expert-guided ex- plainable few-shot learning with active sample selection for medical image analysis,”IEEE Journal of Biomedical and Health Informatics, 2026

  15. [26]

    Coswin: Convolution enhanced hierarchical shifted win- dow attention for small-scale vision,

    P. Khadka, R. Rizk, L. Wang, and K. Santosh, “Coswin: Convolution enhanced hierarchical shifted win- dow attention for small-scale vision,”arXiv preprint arXiv:2509.08959, 2025

  16. [28]

    Channel-selected stratified nested cross- validation for clinically relevant eeg-based parkinson’s disease detection,

    N. R. Rasmussen, R. Rizk, L. Wang, A. Singh, and K. Santosh, “Channel-selected stratified nested cross- validation for clinically relevant eeg-based parkinson’s disease detection,” in2026 IEEE Conference on Artificial Intelligence (CAI). IEEE, 2026, pp. 91–97

  17. [30]

    Promoting shape bias in cnns: Frequency-based and contrastive regularization for corruption robustness,

    R. N. Ranabhat, L. Wang, A. K. Patel, and K. San- tosh, “Promoting shape bias in cnns: Frequency-based and contrastive regularization for corruption robustness,” inInternational Conference on Intelligent Systems and Pattern Recognition. Springer, 2025, pp. 16–26

  18. [31]

    Learning to select like humans: Explainable active learning for medical imaging,

    I. I. Uddin, L. Wang, X. Qin, Y . Zhou, and K. San- tosh, “Learning to select like humans: Explainable active learning for medical imaging,” in2026 IEEE Conference on Artificial Intelligence (CAI). IEEE, 2026, pp. 458– 463

  19. [32]

    Ex- plainable novel category discovery in semantic concept space,

    I. I. Uddin, Y . Zhou, K. Santosh, and L. Wang, “Ex- plainable novel category discovery in semantic concept space,”arXiv preprint arXiv:2607.04548, 2026

  20. [33]

    Mecha- nistic interpretability of llm jailbreaks via internal attri- bution graphs,

    A. Wagle, I. I. Uddin, C. Zhang, and L. Wang, “Mecha- nistic interpretability of llm jailbreaks via internal attri- bution graphs,”arXiv preprint arXiv:2607.07903, 2026

  21. [34]

    Frequency-aware contrastive learning for robust shape- biased convolutional neural networks,

    R. N. Ranabhat, L. Wang, A. K. Patel, and K. Santosh, “Frequency-aware contrastive learning for robust shape- biased convolutional neural networks,”Pattern Recogni- tion Letters, 2026

  22. [36]

    Large dimensional analysis of cooperative multicell precoding with local individual csi,

    L. Wang and Q. Liang, “Large dimensional analysis of cooperative multicell precoding with local individual csi,” in2016 IEEE Conference on Computer Communi- cations Workshops (INFOCOM WKSHPS). IEEE, 2016, pp. 89–94

  23. [37]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weis- senborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Min- derer, G. Heigold, S. Gellyet al., “An image is worth 16x16 words: Transformers for image recognition at scale,” inInt. Conf. Learn. Represent., 2021

  24. [38]

    Pytorch library for cam methods,

    J. Gildenblatet al., “Pytorch library for cam methods,” https://github.com/jacobgil/pytorch-grad-cam, 2021

  25. [39]

    ImageNet Large Scale Visual Recognition Challenge,

    O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet Large Scale Visual Recognition Challenge,”Int. J. Comput. Vis., vol. 115, no. 3, pp. 211–252, Dec. 2015

  26. [40]

    Swin transformer: Hierarchical vision transformer using shifted windows,

    Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” inProc. AAAI Conf. Artif. Intell., October 2021, pp. 10 012–10 022

  27. [43]

    BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

    J. Li, D. Li, C. Xiong, and S. Hoi, “BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” inProc. Int. Conf. Mach. Learn., vol. 162, Jul 2022, pp. 12 888–12 900

  28. [44]

    Training data-efficient image trans- formers &; distillation through attention,

    H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablay- rolles, and H. Jegou, “Training data-efficient image trans- formers &; distillation through attention,” inProc. Int. Conf. Mach. Learn., vol. 139, Jul. 2021, pp. 10 347– 10 357

  29. [45]

    Grad-CAM++: Generalized Gradient- Based Visual Explanations for Deep Convolutional Net- works ,

    A. Chattopadhay, A. Sarkar, P. Howlader, and V . N. Bal- asubramanian, “ Grad-CAM++: Generalized Gradient- Based Visual Explanations for Deep Convolutional Net- works ,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., Mar. 2018, pp. 839–847

  30. [49]

    Ablation-CAM: Visual Explanations for Deep Convolutional Network via Gradient-free Localization ,

    S. Desai and H. G. Ramaswamy, “ Ablation-CAM: Visual Explanations for Deep Convolutional Network via Gradient-free Localization ,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., Mar. 2020, pp. 972–980

  31. [50]

    Full-gradient representation 18 for neural network visualization,

    S. Srinivas and F. Fleuret, “Full-gradient representation 18 for neural network visualization,” inProc. Adv. Neural Inf. Process. Syst., vol. 32, 2019

  32. [51]

    Axiom-based grad-cam: Towards accurate visualization and explanation of cnns,

    R. Fu, Q. Hu, X. Dong, Y . Guo, Y . Gao, and B. Li, “Axiom-based grad-cam: Towards accurate visualization and explanation of cnns,” inProc. The Brit. Mach. Vis. Conf., 2020

  33. [52]

    Learning Deep Features for Discriminative Lo- calization ,

    B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Tor- ralba, “ Learning Deep Features for Discriminative Lo- calization ,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit., Jun. 2016, pp. 2921–2929

  34. [57]

    Boosting the transferability of adversarial attack on vision transformer with adaptive token tuning,

    D. Ming, P. Ren, Y . Wang, and X. Feng, “Boosting the transferability of adversarial attack on vision transformer with adaptive token tuning,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 20 887–20 918

  35. [61]

    Emergent open-vocabulary semantic segmentation from off-the- shelf vision-language models,

    J. Luo, S. Khandelwal, L. Sigal, and B. Li, “Emergent open-vocabulary semantic segmentation from off-the- shelf vision-language models,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 4029– 4040

  36. [65]

    Enhancing prompt generation with adaptive refinement for camouflaged object detection,

    X. Chen, G. Ren, T. Dai, T. Stathaki, and H. Liu, “Enhancing prompt generation with adaptive refinement for camouflaged object detection,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 20 672–20 682

  37. [66]

    Quantifying attention flow in transformers,

    S. Abnar and W. Zuidema, “Quantifying attention flow in transformers,” inProc. Annu. Meeting Assoc. Comput. Linguistics, Jul. 2020, pp. 4190–4197. 19 Supplementary Material for Grad-CAM for Vision Transformers: A Systematic Taxonomy and Audit of Methodological Ambiguity in Exp...

  38. [67]

    Unlike attention-map adaptations, the method operates on feature activations from the MLP in the final transformer blocks and uses the true-label class score as the gradient target

    Principles of Visual Tokens for Efficient Video Under- standing [147]:[147] uses a Grad-CAM-style oracle to esti- mate the importance of spatiotemporal tokens in a ViT-based video classification model. Unlike attention-map adaptations, the method operates on feature activation...

  39. [68]

    the best way we found to apply GradCAM was to treat the last attention layer’s [CLS] token as the designated feature map,

    Transformer Interpretability Beyond Attention Visual- ization [4]:adapts Grad-CAM to ViTs by using the final attention layer rather than convolutional feature maps. The authors explicitly state that “the best way we found to apply GradCAM was to treat the last attention layer’...

  40. [69]

    is explicitly adopted by [22, 32, 89] and mentioned in [6, 7, 29, 48, 60, 63, 87, 90, 101, 111, 113, 118, 124, 135– 137, 148]

  41. [70]

    Rather than using an attention map as the attribution representation, the method operates on token-level feature activations

    Squeeze-and-Excitation Vision Transformer (SE-ViT) for Lung Nodule Classification [159]:[159] adapts Grad-CAM to SE-ViT by replacing CNN channels with image tokens. Rather than using an attention map as the attribution representation, the method operates on token-level feature...

  42. [71]

    Rather than operating on attention maps, it fuses gradients and intermediate ViT features from a selected transformer layer

    Boosting the Transferability of Adversarial Attack on Vision Transformer with Adaptive Token Tuning [88]:[88] uses a Grad-CAM-inspired feature-importance computation to guide patch masking in a ViT. Rather than operating on attention maps, it fuses gradients and intermediate V...

  43. [72]

    Cross-Attention Grad-CAM: Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models [19]uses cross-attention maps as the attribution-bearing representation. Rather than extracting token embeddings such asF 1,F 4,F 5, orF 6, the method oper- ates directly on a cros...

  44. [73]

    Align before Fuse: Vision and Language Representation Learning with Momentum Distillation [11]is best charac- terized as an attention-map-based attribution method

    is explicitly adopted by [42, 101] and mentioned in [27, 62, 99]. Align before Fuse: Vision and Language Representation Learning with Momentum Distillation [11]is best charac- terized as an attention-map-based attribution method. Rather than using token-embedding representatio...

  45. [74]

    matching

    is explicitly adopted by [37, 39, 41, 51, 54, 59, 65, 96, 189–193] and mentioned in [12, 19, 35, 36, 42, 50, 52, 64, 94, 95, 98, 99, 101, 120, 128, 155, 192, 194–197]. Enhancing Prompt Generation with Adaptive Refine- ment for Camouflaged Object Detection [155]applies gradient...

  46. [75]

    A dog on a white bed

    is explicitly adopted by [127] and mentioned in [98, 100, 155]. Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability [154]applies Grad-CAM-style attribution to intermediate activations from the vision component of a vision-language transfor...

  47. [76]

    Grad-cam: Visual explana- tions from deep networks via gradient-based localiza- tion,

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explana- tions from deep networks via gradient-based localiza- tion,” inProc. AAAI Conf. Artif. Intell., 2017, pp. 618– 626

  48. [77]

    BLIP: Bootstrap- ping language-image pre-training for unified vision- language understanding and generation,

    J. Li, D. Li, C. Xiong, and S. Hoi, “BLIP: Bootstrap- ping language-image pre-training for unified vision- language understanding and generation,” inProc. Int. Conf. Mach. Learn., vol. 162, Jul 2022, pp. 12 888– 12 900

  49. [78]

    Microsoft coco: Common objects in context,

    T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll ´ar, and C. L. Zitnick, “Microsoft coco: Common objects in context,” inProc. Eur. Conf. Comput. Vis., vol. 8693, 2014, pp. 740–755

  50. [79]

    Transformer in- terpretability beyond attention visualization,

    H. Chefer, S. Gur, and L. Wolf, “Transformer in- terpretability beyond attention visualization,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2021, pp. 782–791

  51. [80]

    Transreid: Transformer-based object re-identification,

    S. He, H. Luo, P. Wang, F. Wang, H. Li, and W. Jiang, “Transreid: Transformer-based object re-identification,” inProc. AAAI Conf. Artif. Intell., October 2021, pp. 15 013–15 022

  52. [81]

    Generic attentiemer- gent on-model explainability for interpreting bi-modal and encoder-decoder transformers,

    H. Chefer, S. Gur, and L. Wolf, “Generic attentiemer- gent on-model explainability for interpreting bi-modal and encoder-decoder transformers,” inProc. AAAI Conf. Artif. Intell., October 2021, pp. 397–406

  53. [82]

    Ia-redˆ2: Interpretability-aware redundancy reduction for vision transformers,

    B. Pan, R. Panda, Y . Jiang, Z. Wang, R. Feris, and A. Oliva, “Ia-redˆ2: Interpretability-aware redundancy reduction for vision transformers,” inProc. Adv. Neural Inf. Process. Syst., vol. 34, 2021, pp. 24 898–24 911

  54. [83]

    Analogous to evolutionary algorithm: Designing a unified sequence model,

    J. Zhang, C. Xu, J. Li, W. Chen, Y . Wang, Y . Tai, S. Chen, C. Wang, F. Huang, and Y . Liu, “Analogous to evolutionary algorithm: Designing a unified sequence model,” inProc. Adv. Neural Inf. Process. Syst., vol. 34, 2021, pp. 26 674–26 688

  55. [84]

    Passive attention in artificial neural networks predicts human visual selectivity,

    T. Langlois, H. Zhao, E. Grant, I. Dasgupta, T. Griffiths, and N. Jacoby, “Passive attention in artificial neural networks predicts human visual selectivity,” inProc. Adv. Neural Inf. Process. Syst., vol. 34, 2021, pp. 27 094–27 106

  56. [85]

    Vitae: Vision transformer advanced by exploring intrinsic inductive bias,

    Y . Xu, Q. ZHANG, J. Zhang, and D. Tao, “Vitae: Vision transformer advanced by exploring intrinsic inductive bias,” inProc. Adv. Neural Inf. Process. Syst., vol. 34, 2021, pp. 28 522–28 535

  57. [86]

    Align before fuse: Vision and language representation learning with momentum distillation,

    J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” in Proc. Adv. Neural Inf. Process. Syst., vol. 34, 2021, pp. 9694–9705

  58. [87]

    Vlmae: Vision-language masked autoencoder,

    S. He, T. Guo, T. Dai, R. Qiao, C. Wu, X. Shu, and B. Ren, “Vlmae: Vision-language masked autoencoder,” 2022, arXiv preprint arXiv:2208.09374

  59. [88]

    A compre- hensive study of image classification model sensitivity to foregrounds, backgrounds, and visual attributes,

    M. Moayeri, P. Pope, Y . Balaji, and S. Feizi, “A compre- hensive study of image classification model sensitivity to foregrounds, backgrounds, and visual attributes,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2022, pp. 19 087–19 097

  60. [89]

    A challenging benchmark of anime style recognition,

    H. Li, S. Guo, K. Lyu, X. Yang, T. Chen, J. Zhu, and H. Zeng, “A challenging benchmark of anime style recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2022, pp. 4721– 4730

  61. [90]

    Metaformer is actually what you need for vision,

    W. Yu, M. Luo, P. Zhou, C. Si, Y . Zhou, X. Wang, J. Feng, and S. Yan, “Metaformer is actually what you need for vision,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2022, pp. 10 819–10 829

  62. [91]

    Delving deep into the generalization of vision transformers under distribution shifts,

    C. Zhang, M. Zhang, S. Zhang, D. Jin, Q. Zhou, Z. Cai, H. Zhao, X. Liu, and Z. Liu, “Delving deep into the generalization of vision transformers under distribution shifts,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2022, pp. 7277–7286

  63. [92]

    General facial representation learning in a visual- linguistic manner,

    Y . Zheng, H. Yang, T. Zhang, J. Bao, D. Chen, Y . Huang, L. Yuan, D. Chen, M. Zeng, and F. Wen, “General facial representation learning in a visual- linguistic manner,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2022, pp. 18 697–18 709

  64. [93]

    Multi-modal alignment using representa- tion codebook,

    J. Duan, L. Chen, S. Tran, J. Yang, Y . Xu, B. Zeng, and T. Chilimbi, “Multi-modal alignment using representa- tion codebook,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2022, pp. 15 651–15 660

  65. [94]

    Plug-and-play VQA: Zero-shot VQA by conjoining large pretrained models with zero training,

    A. M. H. Tiong, J. Li, B. Li, S. Savarese, and S. C. Hoi, “Plug-and-play VQA: Zero-shot VQA by conjoining large pretrained models with zero training,” inFindings Assoc. Comput. Linguistics, Dec. 2022, pp. 951–967

  66. [95]

    Inception transformer,

    C. Si, W. Yu, P. Zhou, Y . Zhou, X. Wang, and S. Yan, “Inception transformer,” inProc. Adv. Neural Inf. Pro- cess. Syst., vol. 35, 2022, pp. 23 495–23 509

  67. [96]

    Delving into sequential patches for deep- fake detection,

    J. Guan, H. Zhou, Z. Hong, E. Ding, J. Wang, C. Quan, and Y . Zhao, “Delving into sequential patches for deep- fake detection,” inProc. Adv. Neural Inf. Process. Syst., vol. 35, 2022, pp. 4517–4530

  68. [97]

    Adversarial normalization: I can visualize everything (ice),

    H. Choi, S. Jin, and K. Han, “Adversarial normalization: I can visualize everything (ice),” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 12 115–12 124

  69. [98]

    A new benchmark: On the utility of synthetic data with blender for bare supervised learning and downstream domain adaptation,

    H. Tang and K. Jia, “A new benchmark: On the utility of synthetic data with blender for bare supervised learning and downstream domain adaptation,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 15 954–15 964

  70. [99]

    Selfme: Self-supervised motion learning for micro- expression recognition,

    X. Fan, X. Chen, M. Jiang, A. R. Shahid, and H. Yan, “Selfme: Self-supervised motion learning for micro- expression recognition,” inProc. IEEE/CVF Conf. Com- put. Vis. Pattern Recognit., June 2023, pp. 13 834– 13 843

  71. [100]

    Marlin: Masked autoencoder for facial video representation learning,

    Z. Cai, S. Ghosh, K. Stefanov, A. Dhall, J. Cai, H. Rezatofighi, R. Haffari, and M. Hayat, “Marlin: Masked autoencoder for facial video representation learning,” inProc. IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit., June 2023, pp. 1493–1504

  72. [101]

    Blackvip: Black-box visual prompting for robust transfer learning,

    C. Oh, H. Hwang, H.-y. Lee, Y . Lim, G. Jung, J. Jung, H. Choi, and K. Song, “Blackvip: Black-box visual prompting for robust transfer learning,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 32 2023, pp. 24 224–24 235

  73. [102]

    To- kenhpe: Learning orientation tokens for efficient head pose estimation via transformers,

    C. Zhang, H. Liu, Y . Deng, B. Xie, and Y . Li, “To- kenhpe: Learning orientation tokens for efficient head pose estimation via transformers,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 8897–8906

  74. [103]

    Boost vision trans- former with gpu-friendly sparsity and quantization,

    C. Yu, T. Chen, Z. Gan, and J. Fan, “Boost vision trans- former with gpu-friendly sparsity and quantization,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 22 658–22 668

  75. [104]

    Vision diffmask: Faithful interpretation of vision transformers with differentiable patch masking,

    A. Nalmpantis, A. Panagiotopoulos, J. Gkountouras, K. Papakostas, and W. Aziz, “Vision diffmask: Faithful interpretation of vision transformers with differentiable patch masking,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2023, pp. 3756– 3763

  76. [105]

    Pha: Patch-wise high-frequency augmentation for transformer-based person re-identification,

    G. Zhang, Y . Zhang, T. Zhang, B. Li, and S. Pu, “Pha: Patch-wise high-frequency augmentation for transformer-based person re-identification,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 14 133–14 142

  77. [106]

    Vision trans- formers with mixed-resolution tokenization,

    T. Ronen, O. Levy, and A. Golbert, “Vision trans- formers with mixed-resolution tokenization,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work- shops, June 2023, pp. 4613–4622

  78. [107]

    D3former: Debiased dual distilled transformer for incremental learning,

    A. Mohamed, R. Grandhe, K. J. Joseph, S. Khan, and F. Khan, “D3former: Debiased dual distilled transformer for incremental learning,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2023, pp. 2421–2430

  79. [108]

    Shared inter- est...sometimes: Understanding the alignment between human perception, vision architectures, and saliency map techniques,

    K. Morrison, A. Mehra, and A. Perer, “Shared inter- est...sometimes: Understanding the alignment between human perception, vision architectures, and saliency map techniques,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2023, pp. 3776– 3781

  80. [109]

    Semicvt: Semi- supervised convolutional vision transformer for seman- tic segmentation,

    H. Huang, S. Xie, L. Lin, R. Tong, Y .-W. Chen, Y . Li, H. Wang, Y . Huang, and Y . Zheng, “Semicvt: Semi- supervised convolutional vision transformer for seman- tic segmentation,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 11 340–11 349

  81. [110]

    Masked autoencoding does not help natural language supervision at scale,

    F. Weers, V . Shankar, A. Katharopoulos, Y . Yang, and T. Gunter, “Masked autoencoding does not help natural language supervision at scale,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 23 432–23 444

  82. [111]

    Vilem: Visual- language error modeling for image-text retrieval,

    Y . Chen, Z. Ma, Z. Zhang, Z. Qi, C. Yuan, Y . Shan, B. Li, W. Hu, X. Qie, and J. Wu, “Vilem: Visual- language error modeling for image-text retrieval,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 11 018–11 027

  83. [112]

    Fashionsap: Symbols and attributes prompt for fine-grained fashion vision-language pre-training,

    Y . Han, L. Zhang, Q. Chen, Z. Chen, Z. Li, J. Yang, and Z. Cao, “Fashionsap: Symbols and attributes prompt for fine-grained fashion vision-language pre-training,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 15 028–15 038

  84. [113]

    Zero-shot referring image segmentation with global-local context features,

    S. Yu, P. H. Seo, and J. Son, “Zero-shot referring image segmentation with global-local context features,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 19 456–19 465

  85. [114]

    Improving visual grounding by encouraging consistent gradient-based explanations,

    Z. Yang, K. Kafle, F. Dernoncourt, and V . Ordonez, “Improving visual grounding by encouraging consistent gradient-based explanations,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 19 165– 19 174

  86. [115]

    Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation,

    Y . Lin, M. Chen, W. Wang, B. Wu, K. Li, B. Lin, H. Liu, and X. He, “Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 15 305–15 314

  87. [116]

    Multi-modal representation learn- ing with text-driven soft masks,

    J. Park and B. Han, “Multi-modal representation learn- ing with text-driven soft masks,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 2798–2807

  88. [117]

    From images to textual prompts: Zero-shot visual question answering with frozen large language models,

    J. Guo, J. Li, D. Li, A. M. H. Tiong, B. Li, D. Tao, and S. Hoi, “From images to textual prompts: Zero-shot visual question answering with frozen large language models,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2023, pp. 10 867–10 877

  89. [118]

    Sparse multi- modal vision transformer for weakly supervised seman- tic segmentation,

    J. Hanna, M. Mommert, and D. Borth, “Sparse multi- modal vision transformer for weakly supervised seman- tic segmentation,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2023, pp. 2145– 2154

  90. [119]

    Semantic information in contrastive learning,

    S. Quan, M. Hirano, and Y . Yamakawa, “Semantic information in contrastive learning,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 5686–5696

  91. [120]

    Smmix: Self-motivated image mixing for vision transformers,

    M. Chen, M. Lin, Z. Lin, Y . Zhang, F. Chao, and R. Ji, “Smmix: Self-motivated image mixing for vision transformers,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 17 260–17 270

  92. [121]

    Cose: A consistency- sensitivity metric for saliency on image classification,

    R. Daroya, A. Sun, and S. Maji, “Cose: A consistency- sensitivity metric for saliency on image classification,” inProc. AAAI Conf. Artif. Intell. Workshops, October 2023, pp. 149–158

  93. [122]

    Tall: Thumbnail layout for deepfake video detection,

    Y . Xu, J. Liang, G. Jia, Z. Yang, Y . Zhang, and R. He, “Tall: Thumbnail layout for deepfake video detection,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 22 658–22 668

  94. [123]

    Funnybirds: A synthetic vision dataset for a part-based analysis of explainable ai methods,

    R. Hesse, S. Schaub-Meyer, and S. Roth, “Funnybirds: A synthetic vision dataset for a part-based analysis of explainable ai methods,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 3981–3991

  95. [124]

    Thinking image color aesthetics assessment: Models, datasets and benchmarks,

    S. He, A. Ming, Y . Li, J. Sun, S. Zheng, and H. Ma, “Thinking image color aesthetics assessment: Models, datasets and benchmarks,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 21 838–21 847

  96. [125]

    Hivlp: Hierarchical interactive video-language pre-training,

    B. Shao, J. Liu, R. Pei, S. Xu, P. Dai, J. Lu, W. Li, and Y . Yan, “Hivlp: Hierarchical interactive video-language pre-training,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 13 756–13 766

  97. [126]

    Vl-match: Enhancing vision-language pretraining with token-level and instance-level matching,

    J. Bi, D. Cheng, P. Yao, B. Pang, Y . Zhan, C. Yang, Y . Wang, H. Sun, W. Deng, and Q. Zhang, “Vl-match: Enhancing vision-language pretraining with token-level and instance-level matching,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 2584–2593. 33

  98. [127]

    Vilta: Enhancing vision-language pre-training through textual augmentation,

    W. Wang, Z. Yang, B. Xu, J. Li, and Y . Sun, “Vilta: Enhancing vision-language pre-training through textual augmentation,” inProc. AAAI Conf. Artif. Intell., Octo- ber 2023, pp. 3158–3169

  99. [128]

    Spatio-temporal prompting network for robust video feature extraction,

    G. Sun, C. Wang, Z. Zhang, J. Deng, S. Zafeiriou, and Y . Hua, “Spatio-temporal prompting network for robust video feature extraction,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 13 587–13 597

  100. [129]

    Weakly supervised referring image segmentation with intra-chunk and inter-chunk consistency,

    J. Lee, S. Lee, J. Nam, S. Yu, J. Do, and T. Taghavi, “Weakly supervised referring image segmentation with intra-chunk and inter-chunk consistency,” inProc. AAAI Conf. Artif. Intell., October 2023, pp. 21 870–21 881

  101. [130]

    Wireless cap- sule endoscopy image classification: An explainable ai approach,

    D. Varam, R. Mitra, M. Mkadmi, R. A. Riyas, D. A. Abuhani, S. Dhou, and A. Alzaatreh, “Wireless cap- sule endoscopy image classification: An explainable ai approach,”IEEE Access, vol. 11, pp. 105 262–105 280, 2023

  102. [131]

    Spike-driven transformer,

    M. Yao, J. Hu, Z. Zhou, L. Yuan, Y . Tian, B. Xu, and G. Li, “Spike-driven transformer,” inProc. Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 64 043–64 058

  103. [132]

    Wbcatt: A white blood cell dataset annotated with detailed morphologi- cal attributes,

    S. Tsutsui, W. Pang, and B. Wen, “Wbcatt: A white blood cell dataset annotated with detailed morphologi- cal attributes,” inProc. Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 50 796–50 824

  104. [133]

    Lico: Explainable models with language-image consistency,

    Y . Lei, Z. Li, Y . Li, J. Zhang, and H. Shan, “Lico: Explainable models with language-image consistency,” inProc. Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 61 870–61 887

  105. [134]

    Implicit differentiable outlier detection enable robust deep multimodal analy- sis,

    Z. Wang, S. Medya, and S. Ravi, “Implicit differentiable outlier detection enable robust deep multimodal analy- sis,” inProc. Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 13 854–13 872

  106. [135]

    Visual explanations of image-text representations via multi- modal information bottleneck attribution,

    Y . Wang, T. G. J. Rudner, and A. G. Wilson, “Visual explanations of image-text representations via multi- modal information bottleneck attribution,” inProc. Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 16 009– 16 027

  107. [136]

    Med- unic: Unifying cross-lingual medical vision-language pre-training by diminishing bias,

    Z. Wan, C. Liu, M. Zhang, J. Fu, B. Wang, S. Cheng, L. Ma, C. Quilodr ´an-Casas, and R. Arcucci, “Med- unic: Unifying cross-lingual medical vision-language pre-training by diminishing bias,” inProc. Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 56 186–56 197

  108. [137]

    On evaluating adversarial robustness of large vision-language models,

    Y . Zhao, T. Pang, C. Du, X. Yang, C. LI, N.-M. Cheung, and M. Lin, “On evaluating adversarial robustness of large vision-language models,” inProc. Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 54 111–54 138

  109. [138]

    Delving into masked autoencoders for multi-label thorax disease classification,

    J. Xiao, Y . Bai, A. Yuille, and Z. Zhou, “Delving into masked autoencoders for multi-label thorax disease classification,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., January 2023, pp. 3588–3600

  110. [139]

    Gafnet: A global fourier self attention based novel network for multi- modal downstream tasks,

    O. Susladkar, G. Deshmukh, D. Makwana, S. Mittal, R. S. C. Teja, and R. Singhal, “Gafnet: A global fourier self attention based novel network for multi- modal downstream tasks,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., January 2023, pp. 5242–5251

  111. [140]

    GroundVLP: Harnessing zero-shot visual grounding from vision- language pre-training and open-vocabulary object de- tection,

    H. Shen, T. Zhao, M. Zhu, and J. Yin, “GroundVLP: Harnessing zero-shot visual grounding from vision- language pre-training and open-vocabulary object de- tection,” inProc. AAAI Conf. Artif. Intell., 2024, pp. 4766–4775

  112. [141]

    Debiformer: Vi- sion transformer with deformable agent bi-level routing attention,

    N. BaoLong, C. Zhang, Y . Shi, T. Hirakawa, T. Ya- mashita, T. Matsui, and H. Fujiyoshi, “Debiformer: Vi- sion transformer with deformable agent bi-level routing attention,” inProc. Asian Conf. Comput. Vis., December 2024, pp. 4455–4472

  113. [142]

    Scca-net: A novel network for image manipulation localization using split- channel contextual attention,

    Y . Xiang, K. Zhao, and H. Yin, “Scca-net: A novel network for image manipulation localization using split- channel contextual attention,” inProc. Asian Conf. Comput. Vis., December 2024, pp. 4473–4487

  114. [143]

    Comparing the decision-making mechanisms by transformers and cnns via explanation methods,

    M. Jiang, S. Khorram, and L. Fuxin, “Comparing the decision-making mechanisms by transformers and cnns via explanation methods,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 9546– 9555

  115. [144]

    Just addπ! pose induced video transformers for understanding activities of daily living,

    D. Reilly and S. Das, “Just addπ! pose induced video transformers for understanding activities of daily living,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 18 340–18 350

  116. [145]

    Biggait: Learning gait representation you want by large vision models,

    D. Ye, C. Fan, J. Ma, X. Liu, and S. Yu, “Biggait: Learning gait representation you want by large vision models,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 200–210

  117. [146]

    Flexible biometrics recognition: Bridging the multimodality gap through attention alignment and prompt tuning,

    L. C. O. Tiong, D. Sigmund, C.-H. Chan, and A. B. J. Teoh, “Flexible biometrics recognition: Bridging the multimodality gap through attention alignment and prompt tuning,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 267–276

  118. [147]

    Adapt- ing short-term transformers for action detection in untrimmed videos,

    M. Yang, H. Gao, P. Guo, and L. Wang, “Adapt- ing short-term transformers for action detection in untrimmed videos,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 18 570–18 579

  119. [148]

    Argue: Attribute-guided prompt tuning for vision-language models,

    X. Tian, S. Zou, Z. Yang, and J. Zhang, “Argue: Attribute-guided prompt tuning for vision-language models,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 28 578–28 587

  120. [149]

    Vision-language models for decod- ing provider attention during neonatal resuscitation,

    F. Parodi, J. K. Matelsky, A. Regla-Vargas, E. E. Foglia, C. Lim, D. Weinberg, K. P. Kording, H. M. Herrick, and M. L. Platt, “Vision-language models for decod- ing provider attention during neonatal resuscitation,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Worksh...

  121. [150]

    Residual-based language models are free boosters for biomedical imaging tasks,

    Z. Lai, J. Wu, S. Chen, Y . Zhou, and N. Hovakimyan, “Residual-based language models are free boosters for biomedical imaging tasks,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 5086–5096

  122. [151]

    Eformer: En- hanced transformer towards semantic-contour features of foreground for portraits matting,

    Z. Wang, Q. Miao, Y . Xi, and P. Zhao, “Eformer: En- hanced transformer towards semantic-contour features of foreground for portraits matting,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 3880–3889

  123. [152]

    Unified face attack detection with micro disturbance and a two-stage training strategy,

    J. Yu, D. Lu, X. Shi, C. Qu, and F. Guo, “Unified face attack detection with micro disturbance and a two-stage training strategy,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 960– 969. 34

  124. [153]

    Ma- avt: Modality alignment for parameter-efficient audio- visual transformers,

    T. Mahmud, S. Mo, Y . Tian, and D. Marculescu, “Ma- avt: Modality alignment for parameter-efficient audio- visual transformers,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 7996– 8005

  125. [154]

    Learning to rank patches for unbiased image re- dundancy reduction,

    Y . Luo, Z. Chen, P. Zhou, Z. Wu, X. Gao, and Y .-G. Jiang, “Learning to rank patches for unbiased image re- dundancy reduction,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 22 831–22 840

  126. [155]

    nnmobilenet: Rethinking cnn for retinopathy research,

    W. Zhu, P. Qiu, X. Chen, X. Li, N. Lepore, O. M. Dumitrascu, and Y . Wang, “nnmobilenet: Rethinking cnn for retinopathy research,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 2285–2294

  127. [156]

    Inceptionnext: When inception meets convnext,

    W. Yu, P. Zhou, S. Yan, and X. Wang, “Inceptionnext: When inception meets convnext,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 5672–5683

  128. [157]

    Towards explainable visual vessel recognition using fine-grained classification and image retrieval,

    H. Karus, F. Schwenker, M. Munz, and M. Teutsch, “Towards explainable visual vessel recognition using fine-grained classification and image retrieval,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work- shops, June 2024, pp. 82–92

  129. [158]

    Learning cnn on vit: A hybrid model to explicitly class-specific boundaries for domain adap- tation,

    B. H. Ngo, N.-T. Do-Tran, T.-N. Nguyen, H.-G. Jeon, and T. J. Choi, “Learning cnn on vit: A hybrid model to explicitly class-specific boundaries for domain adap- tation,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 28 545–28 554

  130. [159]

    Vim4path: Self-supervised vision mamba for histopathology images,

    A. Nasiri-Sarvi, V . Q.-H. Trinh, H. Rivaz, and M. S. Hosseini, “Vim4path: Self-supervised vision mamba for histopathology images,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 6894–6903

  131. [160]

    Efficient multitask dense predictor via binarization,

    Y . Shang, D. Xu, G. Liu, R. R. Kompella, and Y . Yan, “Efficient multitask dense predictor via binarization,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 15 899–15 908

  132. [161]

    Daff: Dual attentive feature fusion for multi- spectral pedestrian detection,

    A. Althoupety, L.-Y . Wang, W.-C. Feng, and B. Rek- abdar, “Daff: Dual attentive feature fusion for multi- spectral pedestrian detection,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 2997–3006

  133. [162]

    Skipplus: Skip the first few layers to better explain vision transformers,

    F. Mehri, M. Fayyaz, M. S. Baghshah, and M. T. Pilehvar, “Skipplus: Skip the first few layers to better explain vision transformers,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 204–215

  134. [163]

    Boosting the transferability of adversarial attack on vision trans- former with adaptive token tuning,

    D. Ming, P. Ren, Y . Wang, and X. Feng, “Boosting the transferability of adversarial attack on vision trans- former with adaptive token tuning,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 20 887–20 918

  135. [164]

    On the faithfulness of vision transformer explanations,

    J. Wu, W. Kang, H. Tang, Y . Hong, and Y . Yan, “On the faithfulness of vision transformer explanations,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 10 936–10 945

  136. [165]

    Identifying important group of pixels using interactions,

    K. Sumiyasu, K. Kawamoto, and H. Kera, “Identifying important group of pixels using interactions,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 6017–6026

  137. [166]

    Cdad- net: Bridging domain gaps in generalized category dis- covery,

    S. B. Rongali, S. Mehrotra, A. Jha, M. H. N. C, S. Bose, T. Gupta, M. Singha, and B. Banerjee, “Cdad- net: Bridging domain gaps in generalized category dis- covery,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 2616–2626

  138. [167]

    Large language models are good prompt learners for low-shot image classification,

    Z. Zheng, J. Wei, X. Hu, H. Zhu, and R. Nevatia, “Large language models are good prompt learners for low-shot image classification,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 28 453–28 462

  139. [168]

    Forgery-aware adaptive transformer for generalizable synthetic image detection,

    H. Liu, Z. Tan, C. Tan, Y . Wei, J. Wang, and Y . Zhao, “Forgery-aware adaptive transformer for generalizable synthetic image detection,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 10 770– 10 780

  140. [169]

    Question aware vision transformer for multimodal reasoning,

    R. Ganz, Y . Kittenplon, A. Aberdam, E. Ben Avraham, O. Nuriel, S. Mazor, and R. Litman, “Question aware vision transformer for multimodal reasoning,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 13 861–13 871

  141. [170]

    Mope-clip: Structured pruning for efficient vision-language models with module-wise pruning error metric,

    H. Lin, H. Bai, Z. Liu, L. Hou, M. Sun, L. Song, Y . Wei, and Z. Sun, “Mope-clip: Structured pruning for efficient vision-language models with module-wise pruning error metric,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 27 370–27 380

  142. [171]

    De- noised vision-language fusion guided by visual cues for e-commerce product search,

    Z. Hu, S. Li, M. Du, A. Dhua, and D. Gray, “De- noised vision-language fusion guided by visual cues for e-commerce product search,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2024, pp. 1986–1996

  143. [172]

    Tune-an-ellipse: Clip has potential to find what you want,

    J. Xie, S. Deng, B. Li, H. Liu, Y . Huang, Y . Zheng, J. Schmidhuber, B. Ghanem, L. Shen, and M. Z. Shou, “Tune-an-ellipse: Clip has potential to find what you want,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 13 723–13 732

  144. [173]

    Investigating compositional challenges in vision-language models for visual grounding,

    Y . Zeng, Y . Huang, J. Zhang, Z. Jie, Z. Chai, and L. Wang, “Investigating compositional challenges in vision-language models for visual grounding,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 14 141–14 151

  145. [174]

    Emer- gent open-vocabulary semantic segmentation from off- the-shelf vision-language models,

    J. Luo, S. Khandelwal, L. Sigal, and B. Li, “Emer- gent open-vocabulary semantic segmentation from off- the-shelf vision-language models,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 4029–4040

  146. [175]

    Clip as rnn: Segment countless visual concepts without training en- deavor,

    S. Sun, R. Li, P. Torr, X. Gu, and S. Li, “Clip as rnn: Segment countless visual concepts without training en- deavor,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 13 171–13 182

  147. [176]

    Improved visual grounding through self-consistent explanations,

    R. He, P. Cascante-Bonilla, Z. Yang, A. C. Berg, and V . Ordonez, “Improved visual grounding through self-consistent explanations,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2024, pp. 13 095– 13 105

  148. [177]

    Assessing brain-like characteristics of dnns with spatiotemporal features: A study based on the m ¨uller-lyer illusion,

    H. Zhang, K. Matsuzaki, and S. Yoshida, “Assessing brain-like characteristics of dnns with spatiotemporal features: A study based on the m ¨uller-lyer illusion,” IEEE Access, vol. 12, pp. 147 192–147 208, 2024. 35

  149. [178]

    A hybrid learning-architecture for men- tal disorder detection using emotion recognition,

    J. Aina, O. Akinniyi, M. M. Rahman, V . Odero-Marah, and F. Khalifa, “A hybrid learning-architecture for men- tal disorder detection using emotion recognition,”IEEE Access, vol. 12, pp. 91 410–91 425, 2024

  150. [179]

    Ask, attend, attack: An effective decision-based black-box targeted attack for image-to-text models,

    Q. Zeng, Z. Wang, Y .-m. Cheung, and M. Jiang, “Ask, attend, attack: An effective decision-based black-box targeted attack for image-to-text models,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 105 819– 105 847

  151. [180]

    Imdl- benco: A comprehensive benchmark and codebase for image manipulation detection & localization,

    X. Ma, X. Zhu, L. Su, B. Du, Z. Jiang, B. Tong, Z. Lei, X. Yang, C.-M. Pun, J. Lv, and J. Zhou, “Imdl- benco: A comprehensive benchmark and codebase for image manipulation detection & localization,” in Proc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 134 591–134 613

  152. [181]

    Real-time core-periphery guided vit with smart data layout selection on mobile devices,

    Z. Shu, X. Yu, Z. Wu, W. Jia, Y . Shi, M. Yin, T. Liu, D. Zhu, and W. Niu, “Real-time core-periphery guided vit with smart data layout selection on mobile devices,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 95 744–95 763

  153. [182]

    Visual fourier prompt tuning,

    R. Zeng, C. Han, Q. Wang, C. Wu, T. Geng, L. Huang, Y . N. Wu, and D. Liu, “Visual fourier prompt tuning,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 5552–5585

  154. [183]

    Lips are lying: Spotting the temporal inconsistency between audio and visual in lip-syncing deepfakes,

    W. Liu, T. She, J. Liu, B. Li, D. Yao, Z. Liang, and R. Wang, “Lips are lying: Spotting the temporal inconsistency between audio and visual in lip-syncing deepfakes,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 91 131–91 155

  155. [184]

    Testing semantic importance via betting,

    J. Teneggi and J. Sulam, “Testing semantic importance via betting,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 76 450–76 499

  156. [185]

    Linearly decomposing and recomposing vi- sion transformers for diverse-scale models,

    S. Lin, M. Zhang, R. Chen, Q. Wang, X. Yang, and X. Geng, “Linearly decomposing and recomposing vi- sion transformers for diverse-scale models,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 33 188–33 212

  157. [186]

    Learning low-rank feature for thorax disease classification,

    Y . Wang, R. Goel, U. Nath, A. C. Silva, T. Wu, and Y . Yang, “Learning low-rank feature for thorax disease classification,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 117 133–117 163

  158. [187]

    Adanca: Neural cellular automata as adaptors for more robust vision transformer,

    Y . Xu, T. Zhang, and S. S ¨usstrunk, “Adanca: Neural cellular automata as adaptors for more robust vision transformer,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 25 709–25 746

  159. [188]

    Decom- posing and interpreting image representations via text in vits beyond clip,

    S. Balasubramanian, S. Basu, and S. Feizi, “Decom- posing and interpreting image representations via text in vits beyond clip,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 81 046–81 076

  160. [189]

    Matryoshka query transformer for large vision-language models,

    W. Hu, Z.-Y . Dou, L. H. Li, A. Kamath, N. Peng, and K.-W. Chang, “Matryoshka query transformer for large vision-language models,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 50 168–50 188

  161. [190]

    Emr-merging: Tuning-free high- performance model merging,

    C. Huang, P. Ye, T. Chen, T. He, X. Yue, and W. Ouyang, “Emr-merging: Tuning-free high- performance model merging,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 122 741–122 769

  162. [191]

    Simvg: A simple framework for visual grounding with decoupled multi-modal fusion,

    M. Dai, L. Yang, Y . Xu, Z. Feng, and W. Yang, “Simvg: A simple framework for visual grounding with decoupled multi-modal fusion,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 121 670–121 698

  163. [192]

    A unified debiasing approach for vision-language models across modalities and tasks,

    H. Jung, T. Jang, and X. Wang, “A unified debiasing approach for vision-language models across modalities and tasks,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 21 034–21 058

  164. [193]

    Beyond accuracy: Ensuring correct predictions with correct rationales,

    T. Li, M. Ma, and X. Peng, “Beyond accuracy: Ensuring correct predictions with correct rationales,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 43 164–43 188

  165. [194]

    Neuro-vision to language: Enhancing brain recording-based visual reconstruction and language interaction,

    G. Shen, D. Zhao, X. He, L. Feng, Y . Dong, J. Wang, Q. Zhang, and Y . Zeng, “Neuro-vision to language: Enhancing brain recording-based visual reconstruction and language interaction,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 98 083–98 110

  166. [195]

    Text-guided attention is all you need for zero-shot robustness in vision- language models,

    L. Yu, H. Zhang, and C. Xu, “Text-guided attention is all you need for zero-shot robustness in vision- language models,” inProc. Adv. Neural Inf. Process. Syst., vol. 37, 2024, pp. 96 424–96 448

  167. [196]

    Musictalk: A microservice approach for musical instrument recogni- tion,

    Y .-B. Lin, C.-C. Cheng, and S.-C. Chiu, “Musictalk: A microservice approach for musical instrument recogni- tion,”IEEE Open J. Comput. Soc., vol. 5, pp. 612–623, 2024

  168. [197]

    C2t-net: Channel- aware cross-fused transformer-style networks for pedes- trian attribute recognition,

    D. C. Bui, T. V . Le, and B. H. Ngo, “C2t-net: Channel- aware cross-fused transformer-style networks for pedes- trian attribute recognition,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. Workshops, January 2024, pp. 351–358

  169. [198]

    Limited data, unlimited potential: A study on vits augmented by masked autoencoders,

    S. Das, T. Jain, D. Reilly, P. Balaji, S. Karmakar, S. Marjit, X. Li, A. Das, and M. S. Ryoo, “Limited data, unlimited potential: A study on vits augmented by masked autoencoders,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., January 2024, pp. 6878–6888

  170. [199]

    Occlu- sion sensitivity analysis with augmentation subspace perturbation in deep feature space,

    P. H. V . Valois, K. Niinuma, and K. Fukui, “Occlu- sion sensitivity analysis with augmentation subspace perturbation in deep feature space,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., January 2024, pp. 4829–4838

  171. [200]

    Multi-view classification using hybrid fusion and mutual distillation,

    S. Black and R. Souvenir, “Multi-view classification using hybrid fusion and mutual distillation,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., January 2024, pp. 270–280

  172. [201]

    Clipag: Towards generator-free text-to-image generation,

    R. Ganz and M. Elad, “Clipag: Towards generator-free text-to-image generation,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., January 2024, pp. 3843–3853

  173. [202]

    Foundation model assisted weakly supervised semantic segmentation,

    X. Yang and X. Gong, “Foundation model assisted weakly supervised semantic segmentation,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., January 2024, pp. 523–532

  174. [203]

    Iterprime: Zero-shot referring image segmentation with iterative grad-cam refinement and primary word emphasis,

    Y . Wang, J. Ni, Y . Liu, C. Yuan, and Y . Tang, “Iterprime: Zero-shot referring image segmentation with iterative grad-cam refinement and primary word emphasis,” in Proc. AAAI Conf. Artif. Intell., vol. 39, no. 8, 2025, pp. 9459–9468

  175. [204]

    Wave: Weight templates for adaptive initialization of variable-sized models,

    F. Feng, Y . Xie, J. Wang, and X. Geng, “Wave: Weight templates for adaptive initialization of variable-sized models,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern 36 Recognit., June 2025, pp. 4819–4828

  176. [205]

    Overlock: An overview-first-look- closely-next convnet with context-mixing dynamic ker- nels,

    M. Lou and Y . Yu, “Overlock: An overview-first-look- closely-next convnet with context-mixing dynamic ker- nels,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 128–138

  177. [206]

    Fsfm: A generalizable face security foundation model via self-supervised facial representation learning,

    G. Wang, F. Lin, T. Wu, Z. Liu, Z. Ba, and K. Ren, “Fsfm: A generalizable face security foundation model via self-supervised facial representation learning,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 24 364–24 376

  178. [207]

    Tad- former: Task-adaptive dynamic transformer for efficient multi-task learning,

    S. Baek, S. Lee, H. Jo, H. Choi, and D. Min, “Tad- former: Task-adaptive dynamic transformer for efficient multi-task learning,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 14 858–14 868

  179. [208]

    Distilling spatially-heterogeneous distortion perception for blind image quality assessment,

    X. Li, W. Nie, Y . Zhang, R. Hu, K. Li, X. Zheng, and L. Cao, “Distilling spatially-heterogeneous distortion perception for blind image quality assessment,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 2344–2354

  180. [209]

    Clip-sla: Parameter- efficient clip adaptation for continuous sign language recognition,

    S. Alyami and H. Luqman, “Clip-sla: Parameter- efficient clip adaptation for continuous sign language recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2025, pp. 4137– 4147

  181. [210]

    Disentangling visual transformers: Patch-level interpretability for im- age classification,

    G. Jeanneret, L. Simon, and F. Jurie, “Disentangling visual transformers: Patch-level interpretability for im- age classification,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2025, pp. 2704– 2714

  182. [211]

    Enhancing vision transformer explainability using artificial astro- cytes,

    N. Echevarrieta-Catalan, A. Ribas-Rodriguez, F. Ce- dron, O. Schwartz, and V . Aguiar-Pulido, “Enhancing vision transformer explainability using artificial astro- cytes,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, June 2025, pp. 58–64

  183. [212]

    Prompt-cam: Making vision transformers inter- pretable for fine-grained analysis,

    A. Chowdhury, D. Paul, Z. Mai, J. Gu, Z. Zhang, K. S. Mehrab, E. G. Campolongo, D. Rubenstein, C. V . Stewart, A. Karpatne, T. Berger-Wolf, Y . Su, and W.-L. Chao, “Prompt-cam: Making vision transformers inter- pretable for fine-grained analysis,” inProc. IEEE/CVF Conf. Comput...

  184. [213]

    Rethinking per- sonalized aesthetics assessment: Employing physique aesthetics assessment as an exemplification,

    H. Zhong, S. He, A. Ming, and H. Ma, “Rethinking per- sonalized aesthetics assessment: Employing physique aesthetics assessment as an exemplification,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 2935–2944

  185. [214]

    Goal: Global-local object alignment learning,

    H. Choi, Y . K. Jang, and C. Eom, “Goal: Global-local object alignment learning,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 4070– 4079

  186. [215]

    Separation of powers: On segregating knowledge from observation in llm-enabled knowledge-based visual question answering,

    Z. Yang, Z. Tao, Q. Chen, L. Li, Y . Qi, A. van den Hengel, and Q. Huang, “Separation of powers: On segregating knowledge from observation in llm-enabled knowledge-based visual question answering,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 24 753–24 762

  187. [216]

    Discovering fine-grained visual- concept relations by disentangled optimal transport concept bottleneck models,

    Y . Xie, Z. Zeng, H. Zhang, Y . Ding, Y . Wang, Z. Wang, B. Chen, and H. Liu, “Discovering fine-grained visual- concept relations by disentangled optimal transport concept bottleneck models,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., June 2025, pp. 30 199– 30 209

  188. [217]

    Rankclip: Ranking-consistent language-image pretraining,

    Y . Zhang, Z. Zhao, Z. Chen, Z. Feng, Z. Ding, and Y . Sun, “Rankclip: Ranking-consistent language-image pretraining,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 3874–3884

  189. [218]

    On the complexity-faithfulness trade-off of gradient- based explanations,

    A. Mehrpanah, M. Gamba, K. Smith, and H. Azizpour, “On the complexity-faithfulness trade-off of gradient- based explanations,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 3531–3541

  190. [219]

    Dadm: Dual alignment of domain and modality for face anti-spoofing,

    J. Yang, X. Lin, Z. Yu, L. Zhang, X. Liu, H. Li, X. Yuan, and X. Cao, “Dadm: Dual alignment of domain and modality for face anti-spoofing,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 12 045–12 056

  191. [220]

    Deepshield: Fortifying deepfake video de- tection with local and global forgery analysis,

    Y . Cai, J. Li, Z. Li, W. Chen, R. Lan, X. Xie, X. Luo, and G. Li, “Deepshield: Fortifying deepfake video de- tection with local and global forgery analysis,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 12 524– 12 534

  192. [221]

    Modeling time-lapse trajectories to characterize cranberry growth,

    R. John, A. Chihoub, R. Meegan, G. Sidelli, J. Ney- hart, P. Oudemans, and K. Dana, “Modeling time-lapse trajectories to characterize cranberry growth,” inProc. AAAI Conf. Artif. Intell. Workshops, October 2025, pp. 7208–7218

  193. [222]

    Principles of visual tokens for efficient video understanding,

    X. Hao, G. Li, S. N. Gowda, R. B. Fisher, J. Huang, A. Arnab, and L. Sevilla-Lara, “Principles of visual tokens for efficient video understanding,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 21 254–21 264

  194. [223]

    Token activation map to visually explain multimodal llms,

    Y . Li, H. Wang, X. Ding, H. Wang, and X. Li, “Token activation map to visually explain multimodal llms,” in Proc. AAAI Conf. Artif. Intell., October 2025, pp. 48– 58

  195. [224]

    Ce-fam: Concept-based explanation via fusion of activation maps,

    M. Kuroki and T. Yamasaki, “Ce-fam: Concept-based explanation via fusion of activation maps,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 1413–1422

  196. [225]

    Semantic equitable clustering: A simple and effective strategy for clustering vision tokens,

    Q. Fan, H. Huang, M. Chen, and R. He, “Semantic equitable clustering: A simple and effective strategy for clustering vision tokens,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 4019–4028

  197. [226]

    Cross-modal ship re-identification via optical and sar imagery: A novel dataset and method,

    H. Wang, S. Li, J. Yang, Y . Liu, Y . Lv, and Z. Zhou, “Cross-modal ship re-identification via optical and sar imagery: A novel dataset and method,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 7873–7883

  198. [227]

    Context- aware academic emotion dataset and benchmark,

    L. Zhao, J. Xuan, J. Lou, Y . Yu, and W. Yang, “Context- aware academic emotion dataset and benchmark,” in Proc. AAAI Conf. Artif. Intell., October 2025, pp. 13 859–13 868

  199. [228]

    Mpbr: Multimodal progressive bidirectional reasoning for open-set fine- grained recognition,

    J. Tan, P. Jing, Y . Zhu, and Y . Liu, “Mpbr: Multimodal progressive bidirectional reasoning for open-set fine- grained recognition,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 1282–1291

  200. [229]

    Do vlms have bad eyes? diagnosing compositional failures via mechanistic interpretability,

    A. V . Aravindan, A. Jha, and M. Kulkarni, “Do vlms have bad eyes? diagnosing compositional failures via mechanistic interpretability,” inProc. AAAI Conf. Artif. Intell. Workshops, October 2025, pp. 715–723

  201. [230]

    Enhancing prompt generation with adaptive refinement for camouflaged object detection,

    X. Chen, G. Ren, T. Dai, T. Stathaki, and H. Liu, 37 “Enhancing prompt generation with adaptive refinement for camouflaged object detection,” inProc. AAAI Conf. Artif. Intell., October 2025, pp. 20 672–20 682

  202. [231]

    Effresnet-vit: A fusion-based con- volutional and vision transformer model for explainable medical image classification,

    T. Hussain, H. Shouno, A. Hussain, D. Hussain, M. Is- mail, T. Hussain Mir, F. Rong Hsu, T. Alam, and S. Anonna Akhy, “Effresnet-vit: A fusion-based con- volutional and vision transformer model for explainable medical image classification,”IEEE Access, vol. 13, pp. 54 040–54 068, 2025

  203. [232]

    Ropgcvit: A novel explainable vision transformer for retinopathy of prematurity diagnosis,

    M. Yurdakul, K. Uyar, S. Tasdem ´ır, and I. Atabas, “Ropgcvit: A novel explainable vision transformer for retinopathy of prematurity diagnosis,”IEEE Access, vol. 13, pp. 77 064–77 079, 2025

  204. [233]

    Transformer-based dme classification using reti- nal oct images without data augmentation: An evalu- ation of vit-b16 and vit-b32 with optimizer impact,

    K. C. Pavithra, P. Kumar, M. Geetha, S. V . Bhandary, K. B. Ajitha Shenoy, G. Rao, S. Fernandes, and A. Tul- sani, “Transformer-based dme classification using reti- nal oct images without data augmentation: An evalu- ation of vit-b16 and vit-b32 with optimizer impact,” IEEE Ac...

  205. [234]

    Squeeze-and- excitation vision transformer for lung nodule classifica- tion,

    X. Xue, Y . Ma, W. Du, and Y . Peng, “Squeeze-and- excitation vision transformer for lung nodule classifica- tion,”IEEE Access, vol. 13, pp. 24 852–24 866, 2025

  206. [235]

    Defax: A cross-attention fusion framework for robust and explainable deepfake detection,

    M. Al-Imran, M. S. Sheikh, U. Kirtonia, N. T. Arthi, and S. Ripon, “Defax: A cross-attention fusion framework for robust and explainable deepfake detection,”IEEE Access, vol. 13, pp. 213 962–213 979, 2025

  207. [236]

    A comparative study of performance of various deep learning models and their explainability in detection of glaucoma,

    N. K. Jisy, S. Radhika, S. Senthil, and M. B. Srinivas, “A comparative study of performance of various deep learning models and their explainability in detection of glaucoma,”IEEE Access, vol. 13, pp. 192 891–192 905, 2025

  208. [237]

    Exploring vision transformers and explainable ai for enhanced artefact classification in esophageal endoscopic images,

    P. Bissoonauth-Daiboo, M. Muzzammil Auzine, M. I. Khan, F. Alshannaq, T. Saba, X. Gao, and M. Heenaye- Mamode Khan, “Exploring vision transformers and explainable ai for enhanced artefact classification in esophageal endoscopic images,”IEEE Access, vol. 13, pp. 176 221–176 244, 2025

  209. [238]

    Detection of carbapenem resistance in klebsiella pneumoniae using vision transformers and maldi-tof proteomic profiles,

    V . S. Mar ´ın, J. P. V . Pulgarin, S. A. Holgu ´ın-Garc´ıa, I. L. M. Figueroa, A. C. E. Maldonado, J. C. G. Vel´asquez, R. Tabares-Soto, M. A. Bravo-Ort ´ız, and G. J. Fern ´andez, “Detection of carbapenem resistance in klebsiella pneumoniae using vision transformers and mald...

  210. [239]

    Schizophrenia seg- mentation from smri data: A fusion approach using u-net and vision transformer-based grad-cam,

    T. Sarveswaran and V . Rajangam, “Schizophrenia seg- mentation from smri data: A fusion approach using u-net and vision transformer-based grad-cam,”IEEE Access, vol. 13, pp. 184 376–184 397, 2025

  211. [240]

    Data-efficient wheat disease detection using shifted window transformer: Enhancing accuracy, sustainabil- ity, and global food security,

    M. Khubaib, T. Kehkashan, M. Abdelhaq, M. A. Khan, M. Zaman, I. Ashraf, A. Rehman, and A. Akhunzada, “Data-efficient wheat disease detection using shifted window transformer: Enhancing accuracy, sustainabil- ity, and global food security,”IEEE Trans. Consumer Electron., vol. 7...

  212. [241]

    On which data distribution (synthetic or real) we should rely for soft biometric classification,

    M. R. A, A. Kumar, and A. Agarwal, “On which data distribution (synthetic or real) we should rely for soft biometric classification,” inProc. Winter Conf. Appl. Comput. Vis., February 2025, pp. 6238–6247

  213. [242]

    Swincnn+oe: A swin transformer and cnn architecture for breast histopathology classification with ood and grad-cam integration,

    D. Opoku, K. Owusu-Agyemang, J. B. Hayfron- Acquah, and R.-M. O. M. Gyening, “Swincnn+oe: A swin transformer and cnn architecture for breast histopathology classification with ood and grad-cam integration,”IEEE Access, vol. 14, pp. 3897–3909, 2026

  214. [243]

    Resvit- solar: A hybrid residual-transformer for photovoltaic panel fault and cleanliness detection,

    N. Amin, Y .-W. Kim, C. Kang, and Y .-C. Byun, “Resvit- solar: A hybrid residual-transformer for photovoltaic panel fault and cleanliness detection,”IEEE Access, vol. 14, pp. 24 650–24 664, 2026

  215. [244]

    Locally explaining predic- tion behavior via gradual interventions and measuring property gradients,

    N. Penzel and J. Denzler, “Locally explaining predic- tion behavior via gradual interventions and measuring property gradients,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., March 2026, pp. 7398–7408

  216. [245]

    Ui- styler: Ultrasound image style transfer with class-aware prompts for cross-device diagnosis using a frozen black- box inference network,

    N.-T. Do-Tran, N.-H.-L. Le, and C.-C. Huang, “Ui- styler: Ultrasound image style transfer with class-aware prompts for cross-device diagnosis using a frozen black- box inference network,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., March 2026, pp. 2765–2774

  217. [246]

    Overcoming fine-grained visual challenges in animal re-identification via semantic feature alignment,

    Y . Wu, D. Zhao, Y . Li, M. Alajas, A. S. Glen, J. Zhang, G. Dobbie, D. Wilson, and Y . S. Koh, “Overcoming fine-grained visual challenges in animal re-identification via semantic feature alignment,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., March 2026, pp. 371– 381

  218. [247]

    nnmobilenet++: Towards efficient hybrid networks for retinal image analysis,

    X. Li, W. Zhu, X. Dong, H. Wang, Y . Xiong, O. Dumi- trascu, and Y . Wang, “nnmobilenet++: Towards efficient hybrid networks for retinal image analysis,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. Workshops, March 2026, pp. 282–292

  219. [248]

    Neurobridge: Few-shot cross-modal neuron re-identification via dual- channel deep metric learning,

    W. Li, M. Liao, L. Cai, and A. Li, “Neurobridge: Few-shot cross-modal neuron re-identification via dual- channel deep metric learning,” inProc. IEEE/CVF Win- ter Conf. Appl. Comput. Vis., March 2026, pp. 8670– 8679

  220. [249]

    Dream: Dynamic prompts and guidedmix for efficient continual adapta- tion of visual-language models,

    E. Chee, M. L. Lee, and W. Hsu, “Dream: Dynamic prompts and guidedmix for efficient continual adapta- tion of visual-language models,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., March 2026, pp. 5853–5863

  221. [250]

    Fairvlm: Enhancing fairness and prompt sensitivity in vision language models for medical image segmenta- tion,

    M. M. Rahman, S. Rahman, S. Bhatt, and M. Faezipour, “Fairvlm: Enhancing fairness and prompt sensitivity in vision language models for medical image segmenta- tion,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., March 2026, pp. 7450–7460

  222. [251]

    Afragent : An adaptive feature renormal- ization based high resolution aware gui agent,

    N. Anand, R. Jain, S. Patnaik, B. Krishnamurthy, and M. Sarkar, “Afragent : An adaptive feature renormal- ization based high resolution aware gui agent,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., March 2026, pp. 1147–1158

  223. [252]

    Surgxbench: Explainable vision-language model benchmark for surgery,

    J. Cheng, X. Zhao, S. Liu, X. Yu, R. Prakash, P. J. Codd, J. E. Katz, and S. Lin, “Surgxbench: Explainable vision-language model benchmark for surgery,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., March 2026, pp. 8188–8198

  224. [253]

    Revisiting vision- language foundations for no-reference image quality assessment,

    A. Yadav, T. D. Huy, and L. Liu, “Revisiting vision- language foundations for no-reference image quality assessment,” inProc. IEEE/CVF Winter Conf. Appl. 38 Comput. Vis., March 2026, pp. 5416–5425

  225. [254]

    Grad-CAM++: Generalized Gradient- Based Visual Explanations for Deep Convolutional Networks ,

    A. Chattopadhay, A. Sarkar, P. Howlader, and V . N. Bal- asubramanian, “ Grad-CAM++: Generalized Gradient- Based Visual Explanations for Deep Convolutional Networks ,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., Mar. 2018, pp. 839–847

  226. [255]

    Layercam: Exploring hierarchical class activa- tion maps for localization,

    P.-T. Jiang, C.-B. Zhang, Q. Hou, M.-M. Cheng, and Y . Wei, “Layercam: Exploring hierarchical class activa- tion maps for localization,”IEEE Trans. Image Process., vol. 30, pp. 5875–5888, 2021

  227. [256]

    Score-CAM: Score-Weighted Visual Explanations for Convolutional Neural Networks ,

    H. Wang, Z. Wang, M. Du, F. Yang, Z. Zhang, S. Ding, P. Mardziel, and X. Hu, “Score-CAM: Score-Weighted Visual Explanations for Convolutional Neural Networks ,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. Workshops, Jun. 2020, pp. 111–119

  228. [257]

    ScoreCAM++: Gated Score-Weighted Visual Explanations for CNNs ,

    S. Mitra, A. Sukul, S. K. Roy, P. Singh, and V . K. Verma, “ ScoreCAM++: Gated Score-Weighted Visual Explanations for CNNs ,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, Jun. 2025, pp. 2691–2700

  229. [258]

    Ablation-CAM: Visual Explanations for Deep Convolutional Network via Gradient-free Localization ,

    S. Desai and H. G. Ramaswamy, “ Ablation-CAM: Visual Explanations for Deep Convolutional Network via Gradient-free Localization ,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis., Mar. 2020, pp. 972– 980

  230. [259]

    Full-gradient representation for neural network visualization,

    S. Srinivas and F. Fleuret, “Full-gradient representation for neural network visualization,” inProc. Adv. Neural Inf. Process. Syst., vol. 32, 2019

  231. [260]

    Axiom-based grad-cam: Towards accurate visualiza- tion and explanation of cnns,

    R. Fu, Q. Hu, X. Dong, Y . Guo, Y . Gao, and B. Li, “Axiom-based grad-cam: Towards accurate visualiza- tion and explanation of cnns,” inProc. The Brit. Mach. Vis. Conf., 2020

  232. [261]

    Swin transformer: Hierarchical vision transformer using shifted windows,

    Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” inProc. AAAI Conf. Artif. Intell., October 2021, pp. 10 012– 10 022

  233. [262]

    BEit: BERT pre-training of image transformers,

    H. Bao, L. Dong, S. Piao, and F. Wei, “BEit: BERT pre-training of image transformers,” inProc. Int. Conf. Learn. Represent., 2022, Art. no. p-BhZSz59o4

  234. [263]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in Proc. Int. Conf. Mach. Learn., vol. 139, Jul 2021, pp. 8748–8763

  235. [264]

    Winsor- cam: Human-tunable visual explanations from deep net- works via layer-wise winsorization,

    C. Wall, L. Wang, R. Rizk, and K. Santosh, “Winsor- cam: Human-tunable visual explanations from deep net- works via layer-wise winsorization,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

  236. [265]

    Explaining the behavior of neuron activations in deep neural networks,

    L. Wang, C. Wang, Y . Li, and R. Wang, “Explaining the behavior of neuron activations in deep neural networks,” Ad Hoc Networks, vol. 111, p. 102346, 2021

  237. [266]

    Bridging in- terpretability and robustness using lime-guided model refinement,

    N. Nayyem, A. Rakin, and L. Wang, “Bridging in- terpretability and robustness using lime-guided model refinement,”arXiv preprint arXiv:2412.18952, 2024

  238. [267]

    Explainability-driven defense: Grad-cam-guided model refinement against adversarial threats,

    L. Wang, I. I. Uddin, X. Qin, Y . Zhou, and K. Santosh, “Explainability-driven defense: Grad-cam-guided model refinement against adversarial threats,” inProceedings of the AAAI Symposium Series (AAAI) 2025, vol. 6, no. 1, 2025, pp. 49–57

  239. [268]

    Explainability-guided defense: Attribution-aware model refinement against adversarial data attacks,

    L. Wang, M. N. Nayyem, A. Al Rakin, K. Santosh, C. Zhang, and Y . Zhou, “Explainability-guided defense: Attribution-aware model refinement against adversarial data attacks,” in2025 IEEE International Conference on Data Mining (ICDM). IEEE, 2025, pp. 1585–1592

  240. [269]

    Expert-guided explainable few-shot learning for medical image diag- nosis,

    I. I. Uddin, L. Wang, and K. Santosh, “Expert-guided explainable few-shot learning for medical image diag- nosis,” inMICCAI Workshop on Data Engineering in Medical Imaging 2025. Springer Nature Switzerland, 2025, pp. 95–104

  241. [270]

    Expert-guided explainable few-shot learning with active sample se- lection for medical image analysis,

    L. Wang, I. I. Uddin, and K. Santosh, “Expert-guided explainable few-shot learning with active sample se- lection for medical image analysis,”IEEE Journal of Biomedical and Health Informatics, 2026

  242. [271]

    Learning to select like humans: Explainable active learning for medical imaging,

    I. I. Uddin, L. Wang, X. Qin, Y . Zhou, and K. Santosh, “Learning to select like humans: Explainable active learning for medical imaging,” in2026 IEEE Confer- ence on Artificial Intelligence (CAI). IEEE, 2026, pp. 458–463

  243. [272]

    Bridging symmetry and robustness: On the role of equivariance in enhancing adversarial robustness,

    L. Wang, I. I. Uddin, C. Zhang, X. Qin, and Y . Zhou, “Bridging symmetry and robustness: On the role of equivariance in enhancing adversarial robustness,” Advances in Neural Information Processing Systems (NeurIPS), vol. 38, pp. 159 102–159 129, 2025

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.