Pith. sign in

REVIEW 2 major objections 6 minor 1 cited by

Explainability for Vision Foundation Models: A Survey

T0 review · 2 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This survey claims that explainability for vision foundation models is a distinct subfield whose methods are increasingly textual and qualitative, with quantitative evaluation dropping to 36% of surveyed works from a 58% pre-foundation…

desk verdict Useful survey of explainability for vision foundation models, but the headline 36% quantitative-evaluation statistic is not reproducible and the paper needs a documented protocol before it can be cited. read the letter →

arxiv 2501.12203 v1 pith:BQIV6GER submitted 2025-01-21 cs.CV

classification cs.CV
keywords ExplainableAIFoundationmodelsVisionSurveyEvaluationmetricsConceptbottleneckChain-of-thoughtreasoningPost-hocexplanations
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The survey sets out to map the intersection of explainable AI (XAI) and pretrained vision foundation models (PFMs) by collecting 122 papers, sorting them into a taxonomy, and reviewing how they are evaluated. It argues that foundation models occupy a double role: their opacity makes them hard to explain, while their flexibility makes them the preferred building blocks for new explainable systems. The paper's headline empirical claim is that only 36% of the surveyed methods include any quantitative evaluation, a drop from the 58% reported for pre-foundation XAI in an earlier systematic review. A sympathetic reader would therefore take the survey's contribution to be a structured panorama plus a warning: the field is growing and changing faster than its measurement practice.

What carries the argument

The corpus is the central object, and the taxonomy is the mechanism that organizes it: each of the 122 papers is classified along the four dimensions borrowed from the taxonomy literature, scope, output format, functioning, and result type, with result type covering feature importance, examples, and surrogate models. The second piece of machinery is the consolidated axiom list, built by taking the desiderata from prior XAI surveys and collapsing them into five categories, together with a metric inventory that records which axiom each evaluation tool measures and whether it needs ground truth. These two structures carry the survey's argument by turning a pile of papers into countable cells, which is what allows the 36% versus 58% comparison.

What would settle it

Reassemble the corpus with a stated search strategy and an explicit coding rule for quantitative evaluation, then compare the resulting paper count and the percentage of methods with quantitative results against the survey's 122 and 36%; if a reasonable re-run lands far from 36%, the claimed drop from the 58% baseline is not a stable fact about the field.

Watch

Extended reading notes

Core claim

The paper's central claim is that explainability research for vision foundation models has its own structure, distinct from pre-foundation XAI. Its corpus of 122 papers splits into inherently explainable models (76 papers: concept bottleneck models, textual rationale generation, chain-of-thought reasoning, prototypical networks, and a few others) and post-hoc explanation methods (20 papers: input perturbation, counterfactual examples, meta-explanation datasets, and neuron/layer interpretation), plus 26 papers concerned with explaining the foundation models themselves. Every method is placed on the same four taxonomy axes, namely scope, output format, functioning, and result type, which lets the survey compare families and measure evaluation practice. The survey reports that only 36% of these methods include quantitative results, versus the 58% found in a systematic review of pre-foundation XAI, and it explains the gap by the rise of textual and visual explanations that are harder to score than feature-attribution maps. It also consolidates the scattered desiderata from earlier surveys into five axioms, trustworthiness, robustness, complexity, generalizability, and objectiveness, and attaches an inventory of evaluation metrics to those axioms.

Load-bearing premise

The load-bearing premise is that the 122-paper corpus is representative and that the authors' implicit judgement of what counts as 'quantitative results' is consistent, because the survey does not document its search strategy, inclusion criteria, or coding rules.

Editorial extensions

If this is right

  • If the 36% figure is correct, current foundation-model-based explainability is less measured than the pre-foundation XAI literature was, and claims of interpretability rest more often on qualitative examples.
  • Text-producing families inherit mature text metrics, while concept bottleneck and prototypical methods lack an equally standard quantitative route, so evaluation quality will stay uneven across families.
  • Multimodal explanations require multimodal metrics, and the survey points to emerging text-image alignment benchmarks as the natural place to build them.
  • Frozen or lightly adapted foundation models remove the need for task-specific training sets, shifting the bottleneck from training data to the validity and comparability of the explanations themselves.
  • The four-axis taxonomy offers a reusable grid that future work can use to state exactly what kind of explanation a new method produces and how it should be evaluated.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Extension: because the corpus selection and the 'quantitative' judgement are not documented, the 36% figure is best treated as an indicative estimate; re-running the count with explicit inclusion rules would test how much of the drop is a measurement choice.
  • Extension: if chain-of-thought and textual-rationale methods dominate the field, explanation quality becomes entangled with language quality, so a fluent but visually wrong rationale could pass current text metrics; the taxonomy does not separate these two failures.
  • Extension: the authors' open challenge about spurious explanations suggests a concrete test: construct images where a concept word is present in the prompt but absent from the pixels, and measure how often a CLIP-based concept bottleneck model still reports that concept.
  • Extension: the closing proposal to model the latent space of a foundation model rather than its internals implies a research programme in which the opaque encoder is treated as a fixed distribution and only the transparent head is trained, a route that could restore quantitative guarantees to concept-based explanations.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper surveys explainability techniques for vision foundation models, which the authors term Pretrained Foundation Models (PFMs). It compiles a corpus of 122 papers, proposes a taxonomy that separates methods into inherently explainable models (concept bottleneck models, textual rationale generation, chain-of-thought reasoning, prototypical networks, and others) and post-hoc methods (input perturbation, counterfactual examples, meta-explanation datasets, and neuron/layer interpretation), and reviews evaluation practices. The paper reports that only 36% of the surveyed methods include quantitative evaluation, contrasting with a 58% figure drawn from the earlier survey by Nauta et al. [14], and it discusses challenges and open problems for XAI in the PFM era, such as spurious explanations, bias, and reasoning capabilities.

Significance. The paper is timely and addresses an important intersection: explainability for vision foundation models. Its taxonomy, especially the detailed tables mapping each surveyed work to the Speith taxonomy (Tables 2 and 3), and the consolidated inventory of evaluation metrics (Table 5) are useful organizational contributions that will help orient new researchers. The survey also makes a falsifiable empirical claim about the prevalence of quantitative evaluation (36%) that, if supported by a reproducible methodology, would be a valuable observation about the field's current practices. However, in its current form that central statistic is not verifiable, so the paper's empirical contribution rests on undocumented choices. If the authors add a transparent corpus-selection protocol and an operational definition of quantitative results, the survey would be a solid reference for the community.

major comments (2)
  1. [Section 2.4] The corpus selection is not documented. The paper states that the corpus comprises 122 studies published until January 2025 and gives a breakdown into 76 inherently explainable, 20 post-hoc, and 26 challenge-addressing papers, but it does not describe the databases searched, the query strings used, the inclusion/exclusion criteria, or the screening procedure. Without this protocol, the claim of comprehensiveness cannot be verified, and any statistics derived from the corpus, including the 36% figure in Section 4.2.2, are not reproducible. Provide a step-by-step methodology, including the initial search, deduplication, and title/abstract/full-text screening, preferably in a PRISMA-style flow diagram.
  2. [Section 4.2.2] The definition of 'quantitative results' is not operationalized. The paper reports that only 36% of the surveyed methods include quantitative results, but it never states what qualifies as quantitative: whether a single metric on one benchmark suffices, whether user studies count, whether qualitative examples with error bars count, or whether the judgment was made by multiple annotators with inter-annotator agreement. The comparison with the 58% figure from Nauta et al. [14] is therefore not like-for-like, because that survey used its own coding scheme and scope. To make the headline finding meaningful, the authors must provide a coding rubric and, ideally, report agreement statistics; they should also consider publishing the coded data as supplementary material.
minor comments (6)
  1. [Section 1 and Figure 2] The text states 'GPT-2 (1.5T parameters)' but GPT-2 has 1.5B parameters; Figure 2 shows '~1.7B params' for 2019, creating an internal inconsistency. Correct the text to 1.5B.
  2. [Section 3.1.3] The classification of chain-of-thought methods as ante-hoc is not fully aligned with the paper's own definitions in Section 2.2, since CoT prompting is often applied to a frozen pretrained model at inference time. Clarify the operative distinction between 'modifying the way to produce inference' and post-hoc use.
  3. [Table 2] The entry 'Concept Gridlock' lacks a citation number, although the related work appears in the text as reference [168]. Add the missing reference.
  4. [Multiple sections] There are several typos that should be fixed: 'componant' for 'component' in Section 3.1.2 (twice), 'noticable' for 'noticeable' in Section 3.1.1, 'Chain-of-Throught' for 'Chain-of-Thought' in Section 4.2.2, and 'mathematicaly' for 'mathematically' in Section 6.2.1.
  5. [Section 4.2.2] The sentence 'as reported by [14], around 58% of research papers in the field have integrated quantitative evaluation methods' is vague about the scope of [14]'s analysis; specify that the 58% figure refers to XAI methods before the PFM era, as the authors later acknowledge.
  6. [Abstract] The stylization 'eXplainable AI' is nonstandard; consider using 'explainable AI' consistently throughout, or at least define the capitalization at first use.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's central statistics are descriptive aggregations, and self-citations are illustrative rather than load-bearing.

full rationale

This manuscript is a survey and narrative synthesis; it contains no mathematical derivation, fitted parameters, or predictive claims whose output could reduce to its inputs. The central quantitative observation (36% of 122 surveyed PFM-XAI methods include quantitative evaluation, Section 4.2.2) is a descriptive aggregation of the authors' own corpus classification, not a prediction derived from a model, and the 58% comparison is explicitly attributed to Nauta et al. [14]. The authors cite several of their own works (CLIP-QDA [82], PASTA-metric [238], MUAD [202], TeDeSC [159]) as examples or inventory entries, but these citations are illustrative rather than load-bearing: removing them would not alter the survey's taxonomy, the 36% figure would remain a sum over the remaining corpus, and no uniqueness theorem or ansatz is imported from these papers. The reproducibility concern about the undocumented corpus-selection procedure (Section 2.4) is a methodological transparency issue, not circularity, since the statistic is not defined in terms of itself. Under the stated rubric, no circular step can be exhibited with a specific equation or reduction, so the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The survey's main assumptions are its scope decisions and its undocumented corpus selection. No free parameters or invented entities are involved.

assumptions (4)
  • domain assumption Foundation models are defined as any model trained on broad data that can be adapted to a wide range of downstream tasks.
    Adopted from [17] and used to set the scope of the survey in Section 2.1.
  • domain assumption Interpretability and explainability are treated as synonyms.
    Stated in Section 2.2 to avoid terminological ambiguity.
  • ad hoc to paper The corpus of 122 papers is comprehensive and representative.
    The survey claims comprehensiveness but provides no search protocol or inclusion criteria, making this a load-bearing undocumented premise.
  • ad hoc to paper The 36% quantitative evaluation rate is measured with a consistent and meaningful definition of 'quantitative results'.
    Section 4.2.2 reports the statistic without specifying how 'quantitative' was judged, so the claim depends on an unstated assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Explainability for Vision Foundation Models: A Survey." pith.science (2026). https://pith.science/paper/BQIV6GER

@misc{pith2026250112203,
  author       = {Pith},
  title        = {Pith review of: Explainability for Vision Foundation Models: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BQIV6GER}},
  note         = {Machine review of arXiv:2501.12203}
}
read the original abstract

As artificial intelligence systems become increasingly integrated into daily life, the field of explainability has gained significant attention. This trend is particularly driven by the complexity of modern AI models and their decision-making processes. The advent of foundation models, characterized by their extensive generalization capabilities and emergent uses, has further complicated this landscape. Foundation models occupy an ambiguous position in the explainability domain: their complexity makes them inherently challenging to interpret, yet they are increasingly leveraged as tools to construct explainable models. In this survey, we explore the intersection of foundation models and eXplainable AI (XAI) in the vision domain. We begin by compiling a comprehensive corpus of papers that bridge these fields. Next, we categorize these works based on their architectural characteristics. We then discuss the challenges faced by current research in integrating XAI within foundation models. Furthermore, we review common evaluation methodologies for these combined approaches. Finally, we present key observations and insights from our survey, offering directions for future research in this rapidly evolving field.

Figures

Figures reproduced from arXiv: 2501.12203 by the authors.

Figure 1
Figure 1. Global goal of XAI. While a non explainable method only makes inference, an explainable model produces details about the reasons of its decisions, to make its functioning clear or easy to understand InceptionV3 (6.23M parameters) in 2014, and then Resnet (42.70M parameters) in 2016. Then, the field of natural language processing followed with Transformers (65M parameters) in 2017, then BERT (340M parameters) in 2018… view at source ↗
Figure 2
Figure 2. Chronology of the order of magnitude of the number of parameters of learning methods. While early methods were interpretable and lightweight, subsequent developments have led to an increase in complexity that has culminated in foundation models, which are mainly characterized by their size. 2.1. Foundation models According to [17], a foundation model is defined as “any model trained on broad data that can be adapted… view at source ↗
Figure 3
Figure 3. Chain representing the acceptance of AI in society. Each box presents the involved audience (middle) and the description of the step (bottom). meaningful prototypes during training, providing an additional layer of interpretability. These families of models allow for the integration of various interpretability tools. For instance, logical reasoning can be incorporated into Concept Bottleneck Models to process and an… view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Summary of the XAI methods presented in our study. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Scheme of the principle of an inherently interpretable model through a concept bot￾tleneck. Given an input, the CBM first generates a conceptual representation based on a predefined set of concepts. Subsequently, the model produces an output using this conceptual repre…
Figure 6
Figure 6. Figure 6: Scheme of the principle of an inherently interpretable model through the generation of textual rationale. Given an input, the model generates a textual rationale alongside its output, providing insights into the reasoning behind the prediction. Compared to other method…
Figure 7
Figure 7. Figure 7: Scheme of the principle of an inherently interpretable model through chain of thought. The model processes its input in multiple sequential steps, resulting in an iterative reasoning process. Background. The initial developments of Chain of Thought (CoT) explanations o…
Figure 8
Figure 8. Figure 8: Scheme of the principle of an inherently interpretable model through prototypes. The input is embedded into a latent space and mapped to regions corresponding to previously learned prototypes, which are then used to produce the output. Background. Prototypical networks…
Figure 9
Figure 9. Figure 9: Scheme of the principle of post-hoc explanation by perturbing input data. Given an input, a set of perturbed samples is generated. The model’s behavior in response to these perturbations is then analyzed, and the resulting analysis provides explanations for the model’s…
Figure 10
Figure 10. Figure 10: Schema of the principle of post-hoc explanation through counterfactuals. A minimal perturbation is applied to the input to generate a variant that results in a significant change in the model’s inference compared to the original input. and generalized; for instance, […
Figure 11
Figure 11. Figure 11: Scheme of the principle of post-hoc explanation through meta explanations. A dataset, generated through a prior process, is provided to the model. The model’s responses are then analyzed through statistical methods to produce a meta-explanation that offers insights in…
Figure 12
Figure 12. Figure 12: Two primary strategies are commonly employed. The first involves optimization [PITH_FULL_IMAGE:figures/full_fig_p019_12.png]
Figure 13
Figure 13. Figure 13: Summary of works that tackle issues raised by the use of PFM in XAI. [PITH_FULL_IMAGE:figures/full_fig_p024_13.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Concept-Based Mechanistic Interpretability Using Structured Knowledge Graphs

    cs.LG 2025-07 reject novelty 4.0 of 10

    BAGEL trains per-layer logistic-regression probes on CLIP-defined concepts and compares per-class concept probabilities with dataset-level concept frequencies, visualizing the alignment in a knowledge graph.

Reference graph

Works this paper leans on

270 extracted references · 22 canonical work pages · cited by 1 Pith paper

  1. [14]

    Nauta, J

    M. Nauta, J. Trienes, S. Pathak, E. Nguyen, M. Peters, Y. Schmitt, J. Schl¨ otterer, M. van Keulen, C. Seifert, From anecdotal evidence to quantitative evaluation methods: A systematic review on evaluating explainable ai, ACM Computing Surveys 55 (2023) 1–42

  2. [1]

    LeCun, Y

    Y. LeCun, Y. Bengio, G. Hinton, Deep learning, Nature 521 (2015) 436–444

  3. [2]

    W. Wang, J. Dai, Z. Chen, Z. Huang, Z. Li, X. Zhu, X. Hu, T. Lu, L. Lu, H. Li, et al., Internimage: Exploring large-scale vision foundation models with deformable convolutions, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 14408–14419

  4. [3]

    Wortsman, G

    M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, et al., Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time, in: International Conference on Machine Learning, 2022, pp. 23965–23998

  5. [4]

    Sauer, K

    A. Sauer, K. Schwarz, A. Geiger, Stylegan-xl: Scaling stylegan to large diverse datasets, in: ACM SIGGRAPH 2022 Conference Proceedings, 2022, pp. 1–10

  6. [5]

    Castelvecchi, Can we open the black box of ai?, Nature News 538 (2016) 20

    D. Castelvecchi, Can we open the black box of ai?, Nature News 538 (2016) 20

  7. [6]

    A. B. Arrieta, N. Diaz-Rodriguez, J. Del Ser, A. Bennetot, S. Tabik, A. Barbado, S. Gar- cia, S. Gil-Lopez, D. Molina, R. Benjamins, et al., Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai, Information Fusion 58 (2020) 82–115

  8. [7]

    Preece, D

    A. Preece, D. Harborne, D. Braines, R. Tomsett, S. Chakraborty, Stakeholders in ex- plainable ai, arXiv preprint arXiv:1810.00184 (2018)

Show all 270 references
  1. [8]

    Rawal, J

    A. Rawal, J. McCoy, D. B. Rawat, B. M. Sadler, R. S. Amant, Recent advances in trustworthy explainable artificial intelligence: Status, challenges, and perspectives, IEEE Transactions on Artificial Intelligence 3 (2021) 852–866

  2. [9]

    S. C.-H. Yang, N. E. T. Folke, P. Shafto, A psychological theory of explainability, in: International Conference on Machine Learning, 2022, pp. 25007–25021

  3. [10]

    D. Wang, Q. Yang, A. Abdul, B. Y. Lim, Designing theory-driven user-centric explainable ai, in: Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 2019, pp. 1–15

  4. [11]

    B. H. Van der Velden, H. J. Kuijf, K. G. Gilhuijs, M. A. Viergever, Explainable artificial intelligence (xai) in deep learning-based medical image analysis, Medical Image Analysis 79 (2022) 102470

  5. [12]

    Atakishiyev, M

    S. Atakishiyev, M. Salameh, H. Yao, R. Goebel, Explainable artificial intelligence for autonomous driving: A comprehensive overview and field guide for future research direc- tions, arXiv preprint arXiv:2112.11561 (2021). 26

  6. [13]

    L. A. Hendricks, K. Burns, K. Saenko, T. Darrell, A. Rohrbach, Women also snowboard: Overcoming bias in captioning models, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 771–787

  7. [15]

    H. Liu, C. Li, Q. Wu, Y. J. Lee, Visual instruction tuning, Advances in Neural Information Processing Systems 36 (2024)

  8. [16]

    S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su, et al., Grounding dino: Marrying dino with grounded pre-training for open-set object detection, in: European Conference on Computer Vision, Springer, 2025, pp. 38–55

  9. [17]

    Schneider, C

    J. Schneider, C. Meske, P. Kuss, Foundation models: A new paradigm for artificial intelligence, Business & Information Systems Engineering (2024) 1–11

  10. [18]

    Brown, B

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al., Language models are few-shot learners, Advances in Neural Information Processing Systems 33 (2020) 1877–1901

  11. [19]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., Learning transferable visual models from natural language supervision, in: International Conference on Machine Learning, 2021, pp. 8748–8763

  12. [20]

    C. Zhou, Q. Li, C. Li, J. Yu, Y. Liu, G. Wang, K. Zhang, C. Ji, Q. Yan, L. He, et al., A comprehensive survey on pretrained foundation models: A history from bert to chatgpt, arXiv preprint arXiv:2302.09419 (2023)

  13. [21]

    Simonyan, A

    K. Simonyan, A. Zisserman, Very deep convolutional networks for large-scale image recognition, arXiv preprint arXiv:1409.1556 (2014)

  14. [22]

    K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778

  15. [23]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. De- hghani, M. Minderer, G. Heigold, S. Gelly, et al., An image is worth 16x16 words: Transformers for image recognition at scale, arXiv preprint arXiv:2010.11929 (2020)

  16. [24]

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hi- erarchical image database, in: 2009 IEEE Conference on Computer vision and Pattern Recognition, Ieee, 2009, pp. 248–255

  17. [25]

    H. Zhao, Y. Zhang, S. Liu, J. Shi, C. C. Loy, D. Lin, J. Jia, Psanet: Point-wise spatial attention network for scene parsing, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 267–283

  18. [26]

    K. He, R. Girshick, P. Doll´ ar, Rethinking imagenet pre-training, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 4918–4927

  19. [27]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, /suppress L. Kaiser, I. Polosukhin, Attention is all you need, Advances in Neural Information Processing Systems 30 (2017). 27

  20. [28]

    Radford, J

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al., Language models are unsupervised multitask learners, OpenAI blog 1 (2019) 9

  21. [29]

    Devlin, M.-W

    J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional transformers for language understanding, arXiv preprint arXiv:1810.04805 (2018)

  22. [30]

    D. Li, Y. Zhang, H. Peng, L. Chen, C. Brockett, M.-T. Sun, B. Dolan, Contextualized perturbation for textual adversarial attack, arXiv preprint arXiv:2009.07502 (2020)

  23. [31]

    Ravichander, E

    A. Ravichander, E. Hovy, K. Suleman, A. Trischler, J. C. K. Cheung, On the systematicity of probing contextualized word representations: The case of hypernymy in bert, in: Proceedings of the Ninth Joint Conference on Lexical and Computational Semantics, 2020, pp. 88–102

  24. [32]

    J. Li, D. Li, S. Savarese, S. Hoi, Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, in: International Conference on Machine Learning, PMLR, 2023, pp. 19730–19742

  25. [33]

    Rombach, A

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, B. Ommer, High-resolution image synthesis with latent diffusion models, in: CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 10674–10685

  26. [34]

    Girdhar, A

    R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V. Alwala, A. Joulin, I. Misra, Imagebind: One embedding space to bind them all, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 15180–15190

  27. [35]

    S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-maron, M. Gim´ enez, Y. Sulsky, J. Kay, J. T. Springenberg, et al., A generalist agent, Transactions on Machine Learning Research (2022)

  28. [36]

    Touvron, T

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi` ere, N. Goyal, E. Hambro, F. Azhar, et al., Llama: Open and efficient foundation language models, arXiv preprint arXiv:2302.13971 (2023)

  29. [37]

    X. Zou, J. Yang, H. Zhang, F. Li, L. Li, J. Wang, L. Wang, J. Gao, Y. J. Lee, Segment everything everywhere all at once, Advances in Neural Information Processing Systems 36 (2024)

  30. [38]

    N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. R¨ adle, C. Rolland, L. Gustafson, et al., Sam 2: Segment anything in images and videos, arXiv preprint arXiv:2408.00714 (2024)

  31. [39]

    Achiam, S

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al., Gpt-4 technical report, arXiv preprint arXiv:2303.08774 (2023)

  32. [40]

    Agrawal, S

    P. Agrawal, S. Antoniak, E. B. Hanna, B. Bout, D. Chaplot, J. Chudnovsky, D. Costa, B. De Monicault, S. Garg, T. Gervet, et al., Pixtral 12b, arXiv preprint arXiv:2410.07073 (2024)

  33. [41]

    E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, Lora: Low-rank adaptation of large language models, arXiv preprint arXiv:2106.09685 (2021)

  34. [42]

    T. Speith, A review of taxonomies of explainable artificial intelligence (xai) methods, in: Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 2022, pp. 2239–2250. 28

  35. [43]

    J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, Y. Wu, Coca: Contrastive captioners are image-text foundation models, arXiv preprint arXiv:2205.01917 (2022)

  36. [44]

    Ramesh, P

    A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, M. Chen, Hierarchical text-conditional image generation with clip latents, arXiv preprint arXiv:2204.06125 1 (2022) 3

  37. [45]

    Betker, G

    J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guo, et al., Improving image generation with better captions, Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf 2 (2023) 8

  38. [46]

    Alayrac, J

    J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al., Flamingo: a visual language model for few-shot learning, Advances in Neural Information Processing Systems 35 (2022) 23716–23736

  39. [47]

    J. Wang, Z. Yang, X. Hu, L. Li, K. Lin, Z. Gan, Z. Liu, C. Liu, L. Wang, Git: A genera- tive image-to-text transformer for vision and language, arXiv preprint arXiv:2205.14100 (2022)

  40. [48]

    Nichol, P

    A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, M. Chen, Glide: Towards photorealistic image generation and editing with text-guided diffusion models, arXiv preprint arXiv:2112.10741 (2021)

  41. [49]

    H. Tan, M. Bansal, Lxmert: Learning cross-modality encoder representations from trans- formers, arXiv preprint arXiv:1908.07490 (2019)

  42. [50]

    D. Zhu, J. Chen, X. Shen, X. Li, M. Elhoseiny, Minigpt-4: Enhancing vision-language understanding with advanced large language models, arXiv preprint arXiv:2304.10592 (2023)

  43. [51]

    C. Li, H. Xu, J. Tian, W. Wang, M. Yan, B. Bi, J. Ye, H. Chen, G. Xu, Z. Cao, et al., mplug: Effective and efficient vision-language learning by cross-modal skip-connections, arXiv preprint arXiv:2205.12005 (2022)

  44. [52]

    P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, H. Yang, Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework, in: International Conference on Machine Learning, PMLR, 2022, pp. 23318–23340

  45. [53]

    Kirillov, E

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al., Segment anything, arXiv preprint arXiv:2304.02643 (2023)

  46. [54]

    C. Chen, B. Zhang, L. Cao, J. Shen, T. Gunter, A. M. Jose, A. Toshev, J. Shlens, R. Pang, Y. Yang, Stair: Learning sparse text and image representation in grounded tokens, arXiv preprint arXiv:2301.13081 (2023)

  47. [55]

    L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, K.-W. Chang, Visualbert: A simple and performant baseline for vision and language, arXiv preprint arXiv:1908.03557 (2019)

  48. [56]

    H. Zhao, H. Chen, F. Yang, N. Liu, H. Deng, H. Cai, S. Wang, D. Yin, M. Du, Explain- ability for large language models: A survey, ACM Transactions on Intelligent Systems and Technology 15 (2024) 1–38

  49. [57]

    Poeta, G

    E. Poeta, G. Ciravegna, E. Pastor, T. Cerquitelli, E. Baralis, Concept-based explainable artificial intelligence: A survey, arXiv preprint arXiv:2312.12936 (2023). 29

  50. [58]

    Galton, Regression towards mediocrity in hereditary stature., The Journal of the Anthropological Institute of Great Britain and Ireland 15 (1886) 246–263

    F. Galton, Regression towards mediocrity in hereditary stature., The Journal of the Anthropological Institute of Great Britain and Ireland 15 (1886) 246–263

  51. [59]

    McCullagh, Generalized Linear Models, Routledge, 2019

    P. McCullagh, Generalized Linear Models, Routledge, 2019

  52. [60]

    J. R. Quinlan, Induction of decision trees, Machine Learning 1 (1986) 81–106

  53. [61]

    A. d. Garcez, S. Bader, H. Bowman, L. C. Lamb, L. de Penning, B. Illuminoo, H. Poon, C. G. Zaverucha, Neural-symbolic learning and reasoning: A survey and interpretation, Neuro-Symbolic Artificial Intelligence: The State of the Art 342 (2022) 327

  54. [62]

    Why should I trust you?

    M. T. Ribeiro, S. Singh, C. Guestrin, “Why should I trust you?”: Explaining the pre- dictions of any classifier, in: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, August 13-17, 2016, 2016, pp. 1135–1144

  55. [63]

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, D. Batra, Grad-cam: Visual explanations from deep networks via gradient-based localization, in: Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 618–626

  56. [64]

    Petsiuk, A

    V. Petsiuk, A. Das, K. Saenko, Rise: Randomized input sampling for explanation of black-box models, arXiv preprint arXiv:1806.07421 (2018)

  57. [65]

    Cortez, M

    P. Cortez, M. J. Embrechts, Opening black box data mining models using sensitivity analysis, in: 2011 IEEE Symposium on Computational Intelligence and Data Mining (CIDM), IEEE, 2011, pp. 341–348

  58. [66]

    S. M. Lundberg, S.-I. Lee, A unified approach to interpreting model predictions, Advances in Neural Information Processing Systems 30 (2017)

  59. [67]

    Akhtar, A survey of explainable ai in deep visual modeling: Methods and metrics, arXiv preprint arXiv:2301.13445 (2023)

    N. Akhtar, A survey of explainable ai in deep visual modeling: Methods and metrics, arXiv preprint arXiv:2301.13445 (2023)

  60. [68]

    Chatzimparmpas, R

    A. Chatzimparmpas, R. M. Martins, I. Jusufi, A. Kerren, A survey of surveys on the use of visualization for interpreting machine learning models, Information Visualization 19 (2020) 207–233

  61. [69]

    Saeed, C

    W. Saeed, C. Omlin, Explainable ai (xai): A systematic meta-survey of current challenges and future opportunities, Knowledge-Based Systems 263 (2023) 110273

  62. [70]

    Joshi, R

    G. Joshi, R. Walambe, K. Kotecha, A review on explainability in multimodal deep neural nets, IEEE Access 9 (2021) 59800–59821

  63. [71]

    J. Choi, J. Raghuram, Y. Li, S. Banerjee, S. Jha, Adaptive concept bottleneck for foundation models, in: ICML 2024 Workshop on Foundation Models in the Wild, 2024

  64. [72]

    Fumanal-Idocin, J

    J. Fumanal-Idocin, J. Andreu-Perez, O. Cord, H. Hagras, H. Bustince, et al., Artxai: Explainable artificial intelligence curates deep representation learning for artistic images using fuzzy techniques, IEEE Transactions on Fuzzy Systems (2023)

  65. [73]

    T. Li, M. Ma, X. Peng, Beyond accuracy: Ensuring correct predictions with correct rationales, in: The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. 30

  66. [74]

    X. Zhao, X. Huang, T. Fu, Q. Li, S. Gong, L. Liu, W. Bi, L. Kong, Bba: Bi-modal behavioral alignment for reasoning with large vision-language models, arXiv preprint arXiv:2402.13577 (2024)

  67. [75]

    I. Kim, J. Kim, J. Choi, H. J. Kim, Concept bottleneck with visual concept filtering for explainable medical image classification, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer, 2023, pp. 225–233

  68. [76]

    Achille, G

    A. Achille, G. V. Steeg, T. Y. Liu, M. Trager, C. Klingenberg, S. Soatto, Interpretable measures of conceptual similarity by complexity-constrained descriptive auto-encoding, arXiv preprint arXiv:2402.08919 (2024)

  69. [77]

    Y. Cui, S. Liu, L. Li, Z. Yuan, Ceir: Concept-based explainable image representation learning, arXiv preprint arXiv:2312.10747 (2023)

  70. [78]

    J. Liu, T. Hu, Y. Zhang, X. Gai, Y. Feng, Z. Liu, A chatgpt aided explainable framework for zero-shot medical image diagnosis, arXiv preprint arXiv:2307.01981 (2023)

  71. [79]

    Z. Ren, Y. Su, X. Liu, Chatgpt-powered hierarchical comparisons for image classification, Advances in Neural Information Processing Systems 36 (2024)

  72. [80]

    M. Liu, D. Chen, Y. Li, G. Fang, Y. Shen, Chartthinker: A contextual chain-of-thought approach to optimized chart summarization, arXiv preprint arXiv:2403.11236 (2024)

  73. [81]

    Menon, C

    S. Menon, C. Vondrick, Visual classification via description from large language models, in: The Eleventh International Conference on Learning Representations, 2022

  74. [82]

    Kazmierczak, E

    R. Kazmierczak, E. Berthier, G. Frehse, G. Franchi, Clip-qda: An explainable concept bottleneck model, Transactions on Machine Learning Research (2024)

  75. [83]

    Zhang, M

    Y. Zhang, M. Jiang, Q. Zhao, Learning chain of counterfactual thought for bias-robust vision-language reasoning, in: A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, G. Varol (Eds.), Computer Vision – ECCV 2024, Springer Nature Switzerland, Cham, 2025, pp. 334–351

  76. [84]

    Mannix, H

    E. Mannix, H. Bondell, Scalable and robust transformer decoders for interpretable image classification with foundation models, arXiv preprint arXiv:2403.04125 (2024)

  77. [85]

    J. Ge, H. Luo, S. Qian, Y. Gan, J. Fu, S. Zhan, Chain of thought prompt tuning in vision language models, arXiv preprint arXiv:2304.07919 (2023)

  78. [86]

    Harvey, F

    W. Harvey, F. Wood, Visual chain-of-thought diffusion models, arXiv preprint arXiv:2303.16187 (2023)

  79. [87]

    Y. Chen, K. Sikka, M. Cogswell, H. Ji, A. Divakaran, Measuring and improving chain- of-thought reasoning in vision-language models, arXiv preprint arXiv:2309.04461 (2023)

  80. [88]

    L. Li, Cpseg: Finer-grained image semantic segmentation via chain-of-thought language prompting, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 513–522

  81. [90]

    Zheng, B

    G. Zheng, B. Yang, J. Tang, H.-Y. Zhou, S. Yang, Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models, Advances in Neural Information Processing Systems 36 (2023) 5168–5191

  82. [91]

    Z. Wang, J. Xiao, T. Chen, L. Chen, Decap: Towards generalized explicit caption editing via diffusion mechanism, arXiv preprint arXiv:2311.14920 (2023)

  83. [92]

    W. Han, D. Guo, C.-Z. Xu, J. Shen, Dme-driver: Integrating human decision logic and 3d scene perception in autonomous driving, arXiv preprint arXiv:2401.03641 (2024)

  84. [93]

    Grover, V

    S. Grover, V. Vineet, Y. S. Rawat, Navigating hallucinations for reasoning of uninten- tional activities, arXiv preprint arXiv:2402.19405 (2024)

  85. [94]

    Y. Ma, Y. Cao, J. Sun, M. Pavone, C. Xiao, Dolphins: Multimodal language model for driving, arXiv preprint arXiv:2312.00438 (2023)

  86. [95]

    Z. Xu, Y. Zhang, E. Xie, Z. Zhao, Y. Guo, K. K. Wong, Z. Li, H. Zhao, Drivegpt4: Interpretable end-to-end autonomous driving via large language model, arXiv preprint arXiv:2310.01412 (2023)

  87. [96]

    C. Han, J. C. Liang, Q. Wang, M. Rabbani, S. Dianat, R. Rao, Y. N. Wu, D. Liu, Image translation as diffusion visual programmers, arXiv preprint arXiv:2401.09742 (2024)

  88. [97]

    Zhang, J

    F. Zhang, J. Liu, Q. Zhang, E. Sun, J. Xie, Z.-J. Zha, Ecenet: Explainable and context- enhanced network for muti-modal fact verification, in: Proceedings of the 31st ACM International Conference on Multimedia, 2023, pp. 1231–1240

  89. [98]

    A. K. Thakur, F. Ilievski, H.- ˆA. Sandlin, Z. Sourati, L. Luceri, R. Tommasini, A. Mer- moud, Multimodal and explainable internet meme classification, arXiv preprint arXiv:2212.05612 (2022)

  90. [99]

    Uehara, N

    K. Uehara, N. Goswami, H. Wang, T. Baba, K. Tanaka, T. Hashimoto, K. Wang, R. Ito, T. Naoya, R. Umagami, et al., Advancing large multi-modal models with explicit chain- of-reasoning and visual question generation, arXiv preprint arXiv:2401.10005 (2024)

  91. [100]

    J. Yow, N. P. Garg, M. Ramanathan, W. T. Ang, et al., Extract–explainable trajectory corrections from language inputs using textual description of features, arXiv preprint arXiv:2401.03701 (2024)

  92. [101]

    J. Hu, J. Lin, S. Gong, W. Cai, Relax image-specific prompt requirement in sam: A single generic prompt for segmenting camouflaged objects, in: Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 2024, pp. 12511–12518

  93. [102]

    Hwang, S

    H. Hwang, S. Kwon, Y. Kim, D. Kim, Is it safe to cross? interpretable risk assessment with gpt-4v for safety-aware street crossing, arXiv preprint arXiv:2402.06794 (2024)

  94. [103]

    K. P. Panousis, D. Ienco, D. Marcos, Hierarchical concept discovery models: A concept pyramid scheme, arXiv preprint arXiv:2310.02116 (2023)

  95. [104]

    J. Kil, F. Tavazoee, D. Kang, J.-K. Kim, Ii-mmr: Identifying and improving multi-modal multi-hop reasoning in visual question answering, arXiv preprint arXiv:2402.11058 (2024)

  96. [105]

    Y. Ando, N. J.-Y. Park, G. O. Chong, S. Ko, D. Lee, J. Cho, H. Han, Interpretable pap smear cell representation for cervical cancer screening, arXiv preprint arXiv:2311.10269 (2023). 32

  97. [106]

    X. Fu, B. Zhou, S. Chen, M. Yatskar, D. Roth, Interpretable by design visual question answering, arXiv preprint arXiv:2305.14882 (2023)

  98. [107]

    H. Lin, Z. Luo, W. Gao, J. Ma, B. Wang, R. Yang, Towards explainable harmful meme detection through multimodal debate between large language models, arXiv preprint arXiv:2401.13298 (2024)

  99. [108]

    Mondal, S

    D. Mondal, S. Modi, S. Panda, R. Singh, G. S. Rao, Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning, arXiv preprint arXiv:2401.12863 (2024)

  100. [109]

    Oikarinen, S

    T. Oikarinen, S. Das, L. Nguyen, L. Weng, Label-free concept bottleneck models, in: International Conference on Learning Representations, 2023

  101. [110]

    Y. Yang, A. Panagopoulou, S. Zhou, D. Jin, C. Callison-Burch, M. Yatskar, Language in a bottle: Language model guided concept bottlenecks for interpretable image classifi- cation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. ...

  102. [111]

    X. Chen, Y. Liu, Y. Yang, J. Yuan, Q. You, L.-P. Liu, H. Yang, Reason out your layout: Evoking the layout master from large language models for text-to-image synthesis, arXiv preprint arXiv:2311.17126 (2023)

  103. [112]

    H. Li, C. Shen, P. Torr, V. Tresp, J. Gu, Self-discovering interpretable diffusion la- tent directions for responsible text-to-image generation, arXiv preprint arXiv:2311.17216 (2023)

  104. [113]

    A. Yan, Y. Wang, Y. Zhong, C. Dong, Z. He, Y. Lu, W. Y. Wang, J. Shang, J. McAuley, Learning concise and descriptive attributes for visual recognition, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 3090–3100

  105. [114]

    Kassab, M

    C. Kassab, M. Mattamala, L. Zhang, M. Fallon, Language-extended indoor slam (lexis): A versatile system for real-time visual scene understanding, arXiv preprint arXiv:2309.15065 (2023)

  106. [115]

    L. Lian, B. Li, A. Yala, T. Darrell, Llm-grounded diffusion: Enhancing prompt under- standing of text-to-image diffusion models with large language models, arXiv preprint arXiv:2305.13655 (2023)

  107. [116]

    Chiquier, U

    M. Chiquier, U. Mall, C. Vondrick, Evolving interpretable visual classifiers with large language models, arXiv preprint arXiv:2404.09941 (2024)

  108. [117]

    C. Lai, S. Song, S. Meng, J. Li, S. Yan, G. Hu, Towards more faithful natural language explanation using multi-level contrastive learning in vqa, in: Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 2024, pp. 2849–2857

  109. [118]

    L. Hu, S. Lai, W. Chen, H. Xiao, H. Lin, L. Yu, J. Zhang, D. Wang, Towards multi- dimensional explanation alignment for medical classification, in: The Thirty-eighth An- nual Conference on Neural Information Processing Systems, 2024

  110. [119]

    Y. Wu, Y. Liu, Y. Yang, M. S. Yao, W. Yang, X. Shi, L. Yang, D. Li, Y. Liu, J. C. Gee, et al., A concept-based interpretable model for the diagnosis of choroid neoplasias using multimodal data, arXiv preprint arXiv:2403.05606 (2024)

  111. [120]

    H. Zhu, R. Togo, T. Ogawa, M. Haseyama, Multimodal natural language explanation generation for visual question answering based on multiple reference data, Electronics 12 (2023) 2183. 33

  112. [121]

    Norrenbrock, M

    T. Norrenbrock, M. Rudolph, B. Rosenhahn, Q-senn: Quantized self-explaining neural networks, in: Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 2024, pp. 21482–21491

  113. [122]

    J. Xu, C. Lan, W. Xie, X. Chen, Y. Lu, Retrieval-based video language model for efficient long video question answering, arXiv preprint arXiv:2312.04931 (2023)

  114. [123]

    M. Nie, R. Peng, C. Wang, X. Cai, J. Han, H. Xu, L. Zhang, Reason2drive: To- wards interpretable and chain-based reasoning for autonomous driving, arXiv preprint arXiv:2312.03661 (2023)

  115. [124]

    A. Yan, Y. Wang, Y. Zhong, Z. He, P. Karypis, Z. Wang, C. Dong, A. Gentili, C.-N. Hsu, J. Shang, et al., Robust and interpretable medical image classifiers via concept bottleneck models, arXiv preprint arXiv:2310.03182 (2023)

  116. [125]

    Patr´ ıcio, L

    C. Patr´ ıcio, L. F. Teixeira, J. C. Neves, Towards concept-based interpretability of skin lesion diagnosis using vision-language models, arXiv preprint arXiv:2311.14339 (2023)

  117. [126]

    Banerjee, S

    D. Banerjee, S. Teso, B. Sayin, A. Passerini, Learning to guide human decision makers with vision-language models, arXiv preprint arXiv:2403.16501 (2024)

  118. [127]

    P. Qi, Z. Yan, W. Hsu, M. L. Lee, Sniffer: Multimodal large language model for ex- plainable out-of-context misinformation detection, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13052–13062

  119. [128]

    J. Qi, Z. Xu, Y. Shen, M. Liu, D. Jin, Q. Wang, L. Huang, The art of socratic questioning: Recursive thinking with large language models, in: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 4177–4199

  120. [129]

    K. P. Panousis, D. Ienco, D. Marcos, Sparse linear concept discovery models, in: Proceed- ings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 2767– 2771

  121. [130]

    Q. Wan, R. Wang, X. Chen, Interpretable object recognition by semantic prototype anal- ysis, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 800–809

  122. [131]

    B. Chen, Z. Xu, S. Kirmani, B. Ichter, D. Driess, P. Florence, D. Sadigh, L. Guibas, F. Xia, Spatialvlm: Endowing vision-language models with spatial reasoning capabilities, arXiv preprint arXiv:2401.12168 (2024)

  123. [132]

    Bhalla, A

    U. Bhalla, A. Oesterling, S. Srinivas, F. P. Calmon, H. Lakkaraju, Interpreting clip with sparse linear concept embeddings (splice), arXiv preprint arXiv:2402.10376 (2024)

  124. [133]

    Vandenhirtz, S

    M. Vandenhirtz, S. Laguna, R. Marcinkeviˇ cs, J. E. Vogt, Stochastic concept bottleneck models, in: The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

  125. [134]

    Y. K. Kim, J. M. Di Martino, G. Sapiro, Vision transformers with natural language semantics, arXiv preprint arXiv:2402.17863 (2024)

  126. [135]

    Moayeri, K

    M. Moayeri, K. Rezaei, M. Sanjabi, S. Feizi, Text-to-concept (and back) via cross-model alignment, in: International Conference on Machine Learning, PMLR, 2023, pp. 25037– 25060. 34

  127. [136]

    Liang, Y

    M. Liang, Y. Wu, et al., Toa: task-oriented active vqa, Advances in Neural Information Processing Systems 36 (2024)

  128. [137]

    Natarajan, A

    P. Natarajan, A. Nambiar, Vale: A multimodal visual and language explanation frame- work for image classifiers using explainable ai and language models, arXiv preprint arXiv:2408.12808 (2024)

  129. [138]

    S. Wang, Q. Zhao, M. Q. Do, N. Agarwal, K. Lee, C. Sun, Vamos: Versatile action models for video understanding, arXiv preprint arXiv:2311.13627 (2023)

  130. [139]

    T. Meng, Y. Tao, R. Lyu, W. Yin, Few-shot image classification and segmentation as visual question answering using vision-language models, arXiv preprint arXiv:2403.10287 (2024)

  131. [140]

    D. Rose, V. Himakunthala, A. Ouyang, R. He, A. Mei, Y. Lu, M. Saxon, C. Sonar, D. Mirza, W. Y. Wang, Visual chain of thought: Bridging logical gaps with multimodal infillings, arXiv preprint arXiv:2305.02317 (2023)

  132. [141]

    H. Shao, S. Qian, H. Xiao, G. Song, Z. Zong, L. Wang, Y. Liu, H. Li, Visual cot: Unleashing chain-of-thought reasoning in multi-modal language models, arXiv preprint arXiv:2403.16999 (2024)

  133. [142]

    Srivastava, G

    D. Srivastava, G. Yan, T.-W. Weng, Vlg-cbm: Training concept bottleneck models with vision-language guidance, in: The Thirty-eighth Annual Conference on Neural Informa- tion Processing Systems, 2024

  134. [143]

    P. Wu, Y. Mu, B. Wu, Y. Hou, J. Ma, S. Zhang, C. Liu, Voronav: Voronoi-based zero-shot object navigation with large language model, arXiv preprint arXiv:2401.02695 (2024)

  135. [144]

    Y. Bie, L. Luo, Z. Chen, H. Chen, Xcoop: Explainable prompt learning for computer- aided diagnosis via concept-guided context optimization, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer, 2024, pp. 773–783

  136. [145]

    Arias-Duart, V

    A. Arias-Duart, V. Gimenez-Abalos, U. Cort´ es, D. Garcia-Gasulla, Assessing biases through visual contexts, Electronics 12 (2023) 3066

  137. [146]

    Balasubramanian, S

    S. Balasubramanian, S. Basu, S. Feizi, Decomposing and interpreting image representa- tions via text in vits beyond clip, arXiv preprint arXiv:2406.01583 (2024)

  138. [147]

    Oikarinen, T.-W

    T. Oikarinen, T.-W. Weng, Clip-dissect: Automatic description of neuron representations in deep vision networks, arXiv preprint arXiv:2204.10965 (2022)

  139. [148]

    Gandikota, J

    R. Gandikota, J. Materzynska, T. Zhou, A. Torralba, D. Bau, Concept sliders: Lora adaptors for precise control in diffusion models, arXiv preprint arXiv:2311.12092 (2023)

  140. [149]

    Farid, S

    K. Farid, S. Schrodi, M. Argus, T. Brox, Latent diffusion counterfactual explanations, arXiv preprint arXiv:2310.06668 (2023)

  141. [150]

    Bl¨ ucher, J

    S. Bl¨ ucher, J. Vielhaben, N. Strodthoff, Decoupling pixel flipping and occlusion strategy for consistent xai benchmarks, arXiv preprint arXiv:2401.06654 (2024)

  142. [151]

    Augustin, V

    M. Augustin, V. Boreiko, F. Croce, M. Hein, Diffusion visual counterfactual explanations, Advances in Neural Information Processing Systems 35 (2022) 364–377. 35

  143. [152]

    S. Jain, H. Lawrence, A. Moitra, A. Madry, Distilling model failures as directions in latent space, arXiv preprint arXiv:2206.14754 (2022)

  144. [153]

    A. Sun, P. Ma, Y. Yuan, S. Wang, Explain any concept: Segment anything meets concept- based explanation, Advances in Neural Information Processing Systems 36 (2024)

  145. [154]

    Kalibhat, S

    N. Kalibhat, S. Bhardwaj, C. B. Bruss, H. Firooz, M. Sanjabi, S. Feizi, Identifying in- terpretable subspaces in image representations, in: International Conference on Machine Learning, PMLR, 2023, pp. 15623–15638

  146. [155]

    S. Kim, J. Oh, S. Lee, S. Yu, J. Do, T. Taghavi, Grounding counterfactual explanation of image classifiers to textual concept space, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 10942–10950

  147. [156]

    Y. Wen, N. Jain, J. Kirchenbauer, M. Goldblum, J. Geiping, T. Goldstein, Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery, Ad- vances in Neural Information Processing Systems 36 (2024)

  148. [157]

    Y. Wang, S. Shen, B. Y. Lim, Reprompt: Automatic prompt editing to refine ai-generative art towards precise expressions, in: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, 2023, pp. 1–29

  149. [158]

    Y. Luo, R. An, B. Zou, Y. Tang, J. Liu, S. Zhang, Llm as dataset analyst: Subpopulation structure discovery with large language model, in: European Conference on Computer Vision, Springer, 2025, pp. 235–252

  150. [159]

    M. Liu, Z. Zhong, J. Li, G. Franchi, S. Roy, E. Ricci, Organizing unstructured image collections using natural language, arXiv preprint arXiv:2410.05217 (2024)

  151. [160]

    Y. Hu, B. Liu, J. Kasai, Y. Wang, M. Ostendorf, R. Krishna, N. A. Smith, Tifa: Accu- rate and interpretable text-to-image faithfulness evaluation with question answering, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 20406–20417

  152. [161]

    M. Ku, D. Jiang, C. Wei, X. Yue, W. Chen, Viescore: Towards explainable metrics for conditional image synthesis evaluation, arXiv preprint arXiv:2312.14867 (2023)

  153. [162]

    J. Luo, Z. Wang, C. H. Wu, D. Huang, F. De la Torre, Zero-shot model diagnosis, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 11631–11640

  154. [163]

    T. R. Shaham, S. Schwettmann, F. Wang, A. Rajaram, E. Hernandez, J. Andreas, A. Tor- ralba, A multimodal automated interpretability agent, in: Forty-first International Con- ference on Machine Learning, 2024

  155. [164]

    Chegini, S

    A. Chegini, S. Feizi, Identifying and mitigating model failures through few-shot clip-aided diffusion generation, arXiv preprint arXiv:2312.05464 (2023)

  156. [165]

    J. Gao, Q. Wu, A. Blair, M. Pagnucco, Lora: A logical reasoning augmented dataset for visual question answering, Advances in Neural Information Processing Systems 36 (2024)

  157. [166]

    P. Knab, S. Marton, C. Bartelt, Dseg-lime–improving image explanation by hierarchical data-driven segmentation, arXiv preprint arXiv:2403.07733 (2024). 36

  158. [167]

    Dombrowski, H

    M. Dombrowski, H. Reynaud, J. P. M¨ uller, M. Baugh, B. Kainz, Trade-offs in fine-tuned diffusion models between accuracy and interpretability, in: Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 2024, pp. 21037–21045

  159. [168]

    Echterhoff, A

    J. Echterhoff, A. Yan, K. Han, A. Abdelraouf, R. Gupta, J. McAuley, Driving through the concept gridlock: Unraveling explainability bottlenecks in automated driving, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 7346–7355

  160. [169]

    Kumar, A

    N. Kumar, A. C. Berg, P. N. Belhumeur, S. K. Nayar, Attribute and simile classifiers for face verification, in: 2009 IEEE 12th International Conference on Computer Vision, IEEE, 2009, pp. 365–372

  161. [170]

    C. H. Lampert, H. Nickisch, S. Harmeling, Learning to detect unseen object classes by between-class attribute transfer, in: 2009 IEEE Conference on Computer Vision and Pattern Recognition, IEEE, 2009, pp. 951–958

  162. [171]

    P. W. Koh, T. Nguyen, Y. S. Tang, S. Mussmann, E. Pierson, B. Kim, P. Liang, Concept bottleneck models, in: International Conference on Machine Learning, 2020, pp. 5338– 5348

  163. [172]

    Losch, M

    M. Losch, M. Fritz, B. Schiele, Interpretability beyond classification output: Semantic bottleneck networks, arXiv preprint arXiv:1907.10882 (2019)

  164. [173]

    Chalasani, J

    P. Chalasani, J. Chen, A. R. Chowdhury, X. Wu, S. Jha, Concise explanations of neural networks using adversarial training, in: International Conference on Machine Learning, PMLR, 2020, pp. 1383–1391

  165. [174]

    E. H. Shortliffe, B. G. Buchanan, A model of inexact reasoning in medicine, Mathematical biosciences 23 (1975) 351–379

  166. [175]

    Van Lent, W

    M. Van Lent, W. Fisher, M. Mancuso, An explainable artificial intelligence system for small-unit tactical behavior, in: Proceedings of the national conference on artificial intelligence, Menlo Park, CA; Cambridge, MA; London; AAAI Press; MIT Press; 1999, 2004, pp. 900–907

  167. [176]

    D. H. Park, L. A. Hendricks, Z. Akata, A. Rohrbach, B. Schiele, T. Darrell, M. Rohrbach, Multimodal explanations: Justifying decisions and pointing to the evidence, in: Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 8779–8788

  168. [177]

    T. T. H. Nguyen, T. Clement, P. T. L. Nguyen, N. Kemmerzell, V. B. Truong, V. T. K. Nguyen, M. Abdelaal, H. Cao, Langxai: Integrating large vision models for generating textual explanations to enhance explainability in visual perception tasks, arXiv preprint arXiv:2402.12525 (2024)

  169. [178]

    Papineni, S

    K. Papineni, S. Roukos, T. Ward, W.-J. Zhu, Bleu: a method for automatic evaluation of machine translation, in: Proceedings of the 40th annual meeting of the Association for Computational Linguistics, 2002, pp. 311–318

  170. [179]

    Lin, Rouge: A package for automatic evaluation of summaries, in: Text summa- rization branches out, 2004, pp

    C.-Y. Lin, Rouge: A package for automatic evaluation of summaries, in: Text summa- rization branches out, 2004, pp. 74–81. 37

  171. [180]

    J. Kim, A. Rohrbach, T. Darrell, J. Canny, Z. Akata, Textual explanations for self-driving vehicles, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 563–578

  172. [181]

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al., Chain-of-thought prompting elicits reasoning in large language models, Advances in Neural Information Processing Systems 35 (2022) 24824–24837

  173. [182]

    Snell, K

    J. Snell, K. Swersky, R. Zemel, Prototypical networks for few-shot learning, Advances in Neural Information Processing Systems 30 (2017)

  174. [183]

    C. Chen, O. Li, D. Tao, A. Barnett, C. Rudin, J. K. Su, This looks like that: deep learning for interpretable image recognition, Advances in Neural Information Processing Systems 32 (2019)

  175. [184]

    Nauta, R

    M. Nauta, R. Van Bree, C. Seifert, Neural prototype trees for interpretable fine-grained image recognition, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 14933–14943

  176. [185]

    M. Xue, Q. Huang, H. Zhang, L. Cheng, J. Song, M. Wu, M. Song, Protopformer: Con- centrating on prototypical parts in vision transformers for interpretable image recognition, arXiv preprint arXiv:2208.10431 (2022)

  177. [186]

    Oquab, T

    M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al., Dinov2: Learning robust visual features without supervision, arXiv preprint arXiv:2304.07193 (2023)

  178. [187]

    C. Wah, S. Branson, P. Welinder, P. Perona, S. Belongie, The caltech-ucsd birds-200-2011 dataset (2011)

  179. [188]

    Zhang, F

    Y. Zhang, F. Xu, J. Zou, O. L. Petrosian, K. V. Krinkin, Xai evaluation: evaluating black-box model explanations for prediction, in: 2021 II International Conference on Neural Networks and Neurotechnologies (NeuroNT), IEEE, 2021, pp. 13–16

  180. [189]

    Goyal, Z

    Y. Goyal, Z. Wu, J. Ernst, D. Batra, D. Parikh, S. Lee, Counterfactual visual explana- tions, in: International Conference on Machine Learning, PMLR, 2019, pp. 2376–2384

  181. [190]

    Peters, D

    J. Peters, D. Janzing, B. Sch¨ olkopf, Elements of Causal Inference: Foundations and Learning Algorithms, The MIT Press, 2017

  182. [191]

    Pearl, Causal inference in statistics: An overview (2009)

    J. Pearl, Causal inference in statistics: An overview (2009)

  183. [192]

    Woodward, Making things happen: A theory of causal explanation, Oxford university press, 2005

    J. Woodward, Making things happen: A theory of causal explanation, Oxford university press, 2005

  184. [193]

    Pawelczyk, K

    M. Pawelczyk, K. Broelemann, G. Kasneci, Learning model-agnostic counterfactual ex- planations for tabular data, in: Proceedings of the Web Conference 2020, 2020, pp. 3126–3132

  185. [194]

    T. D. Duong, Q. Li, G. Xu, Ceflow: A robust and efficient counterfactual explana- tion framework for tabular data using normalizing flows, in: Pacific-Asia Conference on Knowledge Discovery and Data Mining, Springer, 2023, pp. 133–144

  186. [195]

    X. Zhao, K. Broelemann, G. Kasneci, Counterfactual explanation via search in gaussian mixture distributed latent space, arXiv preprint arXiv:2307.13390 (2023). 38

  187. [196]

    F. Yang, N. Liu, M. Du, X. Hu, Generative counterfactuals for neural networks via attribute-informed perturbation, ACM SIGKDD Explorations Newsletter 23 (2021) 59– 68

  188. [197]

    C.-H. Lee, Z. Liu, L. Wu, P. Luo, Maskgan: Towards diverse and interactive facial image manipulation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 5549–5558

  189. [198]

    Heusel, H

    M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, S. Hochreiter, Gans trained by a two time-scale update rule converge to a local nash equilibrium, Advances in Neural Information Processing Systems 30 (2017)

  190. [199]

    Zhang, P

    R. Zhang, P. Isola, A. A. Efros, E. Shechtman, O. Wang, The unreasonable effectiveness of deep features as a perceptual metric, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 586–595

  191. [200]

    Szegedy, Intriguing properties of neural networks, arXiv preprint arXiv:1312.6199 (2013)

    C. Szegedy, Intriguing properties of neural networks, arXiv preprint arXiv:1312.6199 (2013)

  192. [201]

    K. Xiao, L. Engstrom, A. Ilyas, A. Madry, Noise or signal: The role of image backgrounds in object recognition, arXiv preprint arXiv:2006.09994 (2020)

  193. [202]

    Franchi, X

    G. Franchi, X. Yu, A. Bursuc, A. Tena, R. Kazmierczak, S. Dubuisson, E. Aldea, D. Fil- liat, Muad: Multiple uncertainties for autonomous driving, a benchmark for multiple uncertainty types and tasks, in: 33rd British Machine Vision Conference, 2022

  194. [203]

    Krizhevsky, I

    A. Krizhevsky, I. Sutskever, G. E. Hinton, Imagenet classification with deep convolutional neural networks, Advances in Neural Information Processing Systems 25 (2012)

  195. [204]

    Zeiler, Visualizing and understanding convolutional networks, in: European confer- ence on computer vision/arXiv, volume 1311, 2014

    M. Zeiler, Visualizing and understanding convolutional networks, in: European confer- ence on computer vision/arXiv, volume 1311, 2014

  196. [205]

    C. Olah, A. Mordvintsev, L. Schubert, Feature visualization, Distill 2 (2017) e7

  197. [206]

    Q. V. Liao, Y. Zhang, R. Luss, F. Doshi-Velez, A. Dhurandhar, Connecting algorithmic research and usage contexts: a perspective of contextualized evaluation for explainable ai, in: Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, volume 10, 2022, pp. 147–159

  198. [207]

    P. Q. Le, M. Nauta, V. B. Nguyen, S. Pathak, J. Schl¨ otterer, C. Seifert, Benchmarking explainable ai: a survey on available toolkits and open challenges, in: International Joint Conference on Artificial Intelligence, 2023

  199. [208]

    T. Fel, D. Vigouroux, R. Cad` ene, T. Serre, How good is your explanation? algorithmic stability measures to assess the quality of explanations for deep neural networks, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2022, pp. 720–730

  200. [209]

    S. Ali, T. Abuhmed, S. El-Sappagh, K. Muhammad, J. M. Alonso-Moral, R. Confalonieri, R. Guidotti, J. Del Ser, N. D´ ıaz-Rodr´ ıguez, F. Herrera, Explainable artificial intelligence (xai): What we know and what is left to attain trustworthy artificial intelligence, Infor- matio...

  201. [210]

    Guidotti, A

    R. Guidotti, A. Monreale, S. Ruggieri, F. Turini, F. Giannotti, D. Pedreschi, A survey of methods for explaining black box models, ACM Computing Surveys (CSUR) 51 (2018) 1–42. 39

  202. [211]

    Burkart, M

    N. Burkart, M. F. Huber, A survey on the explainability of supervised machine learning, Journal of Artificial Intelligence Research 70 (2021) 245–317

  203. [212]

    Doshi-Velez, B

    F. Doshi-Velez, B. Kim, Towards a rigorous science of interpretable machine learning, arXiv preprint arXiv:1702.08608 (2017)

  204. [213]

    J. Zhou, A. H. Gandomi, F. Chen, A. Holzinger, Evaluating the quality of machine learning explanations: A survey on methods and metrics, Electronics 10 (2021) 593

  205. [214]

    Confalonieri, L

    R. Confalonieri, L. Coba, B. Wagner, T. R. Besold, A historical perspective of explainable artificial intelligence, Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery 11 (2021) e1391

  206. [215]

    A. F. Markus, J. A. Kors, P. R. Rijnbeek, The role of explainability in creating trustworthy artificial intelligence for health care: a comprehensive survey of the terminology, design choices, and evaluation strategies, Journal of biomedical informatics 113 (2021) 103655

  207. [216]

    Belle, I

    V. Belle, I. Papantonis, Principles and practice of explainable machine learning, Frontiers in big Data 4 (2021) 688969

  208. [217]

    Vilone, L

    G. Vilone, L. Longo, Explainable artificial intelligence: a systematic review, arXiv preprint arXiv:2006.00093 (2020)

  209. [218]

    Rojat, R

    T. Rojat, R. Puget, D. Filliat, J. Del Ser, R. Gelin, N. D´ ıaz-Rodr´ ıguez, Explainable artificial intelligence (XAI) on timeseries data: A survey, arXiv preprint arXiv:2104.00950 (2021)

  210. [219]

    Beaudouin, I

    V. Beaudouin, I. Bloch, D. Bounie, S. Cl´ emen¸ con, F. d’Alch´ e Buc, J. Eagan, W. Maxwell, P. Mozharovskyi, J. Parekh, Flexible and context-specific ai explainability: a multidisci- plinary approach, arXiv preprint arXiv:2003.07703 (2020)

  211. [220]

    Bennetot, G

    A. Bennetot, G. Franchi, J. Del Ser, R. Chatila, N. Diaz-Rodriguez, Greybox XAI: A neural-symbolic learning framework to produce interpretable predictions for image classification, Knowledge-Based Systems 258 (2022) 109947

  212. [221]

    Hedstr¨ om, L

    A. Hedstr¨ om, L. Weber, D. Krakowczyk, D. Bareeva, F. Motzkus, W. Samek, S. La- puschkin, M. M. M. H¨ ohne, Quantus: An explainable ai toolkit for responsible evaluation of neural network explanations and beyond, Journal of Machine Learning Research 24 (2023) 1–11. URL: http:...

  213. [222]

    T. Fel, L. Hervier, D. Vigouroux, A. Poche, J. Plakoo, R. Cadene, M. Chalvidal, J. Colin, T. Boissin, L. Bethune, A. Picard, C. Nicodeme, L. Gardes, G. Flandin, T. Serre, Xplique: A deep learning explainability toolbox, Workshop on Explainable Artificial Intelligence for Compu...

  214. [223]

    Yeh, C.-Y

    C.-K. Yeh, C.-Y. Hsieh, A. Suggala, D. I. Inouye, P. K. Ravikumar, On the (in) fidelity and sensitivity of explanations, Advances in Neural Information Processing Systems 32 (2019)

  215. [224]

    Vedantam, C

    R. Vedantam, C. Lawrence Zitnick, D. Parikh, Cider: Consensus-based image description evaluation, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 4566–4575

  216. [225]

    Sundararajan, A

    M. Sundararajan, A. Taly, Q. Yan, Axiomatic attribution for deep networks, in: Inter- national Conference on Machine Learning, PMLR, 2017, pp. 3319–3328. 40

  217. [226]

    Bhatt, A

    U. Bhatt, A. Weller, J. M. Moura, Evaluating and aggregating feature-based model explanations, arXiv preprint arXiv:2005.00631 (2020)

  218. [227]

    Dasgupta, N

    S. Dasgupta, N. Frost, M. Moshkovitz, Framework for evaluating faithfulness of local explanations, in: International Conference on Machine Learning, PMLR, 2022, pp. 4794– 4815

  219. [228]

    Montavon, W

    G. Montavon, W. Samek, K.-R. M¨ uller, Methods for interpreting and understanding deep neural networks, Digital signal processing 73 (2018) 1–15

  220. [229]

    J. Cho, Y. Hu, J. M. Baldridge, R. Garg, P. Anderson, R. Krishna, M. Bansal, J. Pont- Tuset, S. Wang, Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation, in: The Twelfth International Conference on Learning Representations, 2024

  221. [230]

    Nguyen, M

    A.-p. Nguyen, M. R. Mart´ ınez, On quantitative aspects of model interpretability, arXiv preprint arXiv:2007.07584 (2020)

  222. [231]

    Hedstr¨ om, L

    A. Hedstr¨ om, L. Weber, S. Lapuschkin, M. H¨ ohne, Sanity checks revisited: An explo- ration to repair the model parameter randomisation test, arXiv preprint arXiv:2401.06465 (2024)

  223. [232]

    Alvarez Melis, T

    D. Alvarez Melis, T. Jaakkola, Towards robust interpretability with self-explaining neural networks, Advances in Neural Information Processing Systems 31 (2018)

  224. [233]

    Kindermans, S

    P.-J. Kindermans, S. Hooker, J. Adebayo, M. Alber, K. T. Sch¨ utt, S. D¨ ahne, D. Erhan, B. Kim, The (un) reliability of saliency methods, Explainable AI: Interpreting, explaining and visualizing deep learning (2019) 267–280

  225. [234]

    Rieger, L

    L. Rieger, L. K. Hansen, Irof: a low resource evaluation metric for explanation methods, arXiv preprint arXiv:2003.08747 (2020)

  226. [235]

    Banerjee, A

    S. Banerjee, A. Lavie, Meteor: An automatic metric for mt evaluation with improved correlation with human judgments, in: Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, 2005, pp. 65–72

  227. [236]

    V. Arya, R. K. Bellamy, P.-Y. Chen, A. Dhurandhar, M. Hind, S. C. Hoffman, S. Houde, Q. V. Liao, R. Luss, A. Mojsilovi´ c, et al., One explanation does not fit all: A toolkit and taxonomy of ai explainability techniques, arXiv preprint arXiv:1909.03012 (2019)

  228. [237]

    Adebayo, J

    J. Adebayo, J. Gilmer, M. Muelly, I. Goodfellow, M. Hardt, B. Kim, Sanity checks for saliency maps, Advances in Neural Information Processing Systems 31 (2018)

  229. [238]

    Kazmierczak, S

    R. Kazmierczak, S. Azzolin, E. Berthier, A. Hedstr¨ om, P. Delhomme, N. Bousquet, G. Frehse, M. Mancini, B. Caramiaux, A. Passerini, et al., Benchmarking xai expla- nations with human-aligned evaluations, arXiv preprint arXiv:2411.02470 (2024)

  230. [239]

    S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. M¨ uller, W. Samek, On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation, PloS one 10 (2015) e0130140

  231. [240]

    L. Sixt, M. Granz, T. Landgraf, When explanations lie: Why many modified bp at- tributions fail, in: International Conference on Machine Learning, PMLR, 2020, pp. 9046–9057. 41

  232. [241]

    Samek, A

    W. Samek, A. Binder, G. Montavon, S. Lapuschkin, K.-R. M¨ uller, Evaluating the visual- ization of what a deep neural network has learned, IEEE transactions on neural networks and learning systems 28 (2016) 2660–2673

  233. [242]

    Agarwal, N

    C. Agarwal, N. Johnson, M. Pawelczyk, S. Krishna, E. Saxena, M. Zitnik, H. Lakkaraju, Rethinking stability for attribution-based explanations, arXiv preprint arXiv:2203.06877 (2022)

  234. [243]

    Y. Rong, T. Leemann, V. Borisov, G. Kasneci, E. Kasneci, A consistent and efficient evaluation strategy for attribution methods, arXiv preprint arXiv:2202.00449 (2022)

  235. [244]

    Ancona, E

    M. Ancona, E. Ceolini, C. ¨Oztireli, M. Gross, Towards better understanding of gradient- based attribution methods for deep neural networks, arXiv preprint arXiv:1711.06104 (2017)

  236. [245]

    Yarom, Y

    M. Yarom, Y. Bitton, S. Changpinyo, R. Aharoni, J. Herzig, O. Lang, E. Ofek, I. Szpektor, What you see is what you read? improving text-image alignment evaluation, Advances in Neural Information Processing Systems 36 (2024)

  237. [246]

    Anderson, B

    P. Anderson, B. Fernando, M. Johnson, S. Gould, Spice: Semantic propositional image caption evaluation, in: Computer Vision–ECCV 2016: 14th European Conference, Ams- terdam, The Netherlands, October 11-14, 2016, Proceedings, Part V 14, Springer, 2016, pp. 382–398

  238. [247]

    Z. Lin, X. Chen, D. Pathak, P. Zhang, D. Ramanan, Revisiting the role of language priors in vision-language models, arXiv preprint arXiv:2306.01879 (2023)

  239. [248]

    J. Cho, A. Zala, M. Bansal, Visual programming for step-by-step text-to-image generation and evaluation, Advances in Neural Information Processing Systems 36 (2024)

  240. [249]

    H. Yang, J. Gee, J. Shi, Brain decodes deep nets, arXiv preprint arXiv:2312.01280 (2023)

  241. [250]

    X. Zhu, P. Sun, C. Wang, J. Liu, Z. Li, Y. Xiao, J. Huang, A contrastive composi- tional benchmark for text-to-image synthesis: A study with unified text-to-image fidelity metrics, arXiv preprint arXiv:2312.02338 (2023)

  242. [251]

    Opie/suppress lka, J

    G. Opie/suppress lka, J. Loke, S. Scholte, Saliency suppressed, semantics surfaced: Visual trans- formations in neural networks and the brain, arXiv preprint arXiv:2404.18772 (2024)

  243. [252]

    D. S. Martinez Pandiani, N. Lazzari, M. v. Erp, V. Presutti, Hypericons for inter- pretability: decoding abstract concepts in visual data, International Journal of Digital Humanities 5 (2023) 451–490

  244. [253]

    Ghiasi, H

    A. Ghiasi, H. Kazemi, E. Borgnia, S. Reich, M. Shu, M. Goldblum, A. G. Wilson, T. Goldstein, What do vision transformers learn? a visual exploration, arXiv preprint arXiv:2212.06727 (2022)

  245. [254]

    Ahrabian, Z

    K. Ahrabian, Z. Sourati, K. Sun, J. Zhang, Y. Jiang, F. Morstatter, J. Pujara, The curious case of nonverbal abstract reasoning with multi-modal large language models, arXiv preprint arXiv:2401.12117 (2024)

  246. [255]

    Zhang, D

    R. Zhang, D. Jiang, Y. Zhang, H. Lin, Z. Guo, P. Qiu, A. Zhou, P. Lu, K.-W. Chang, P. Gao, et al., Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems?, arXiv preprint arXiv:2403.14624 (2024). 42

  247. [256]

    X. He, Q. Zhang, A. Jin, Y. Yuan, S.-M. Yiu, et al., Tubench: Benchmarking large vision-language models on trustworthiness with unanswerable questions, arXiv preprint arXiv:2410.04107 (2024)

  248. [257]

    Giledereli, Y

    B. Giledereli, Y. Hou, Y. Tu, M. Sachan, Do vision-language models really understand visual language?, arXiv preprint arXiv:2410.00193 (2024)

  249. [258]

    Jiang, Z

    Y. Jiang, Z. Li, X. Shen, Y. Liu, M. Backes, Y. Zhang, Modscan: Measuring stereotypical bias in large vision-language models from vision and language modalities, arXiv preprint arXiv:2410.06967 (2024)

  250. [259]

    Q. Gao, Y. Li, H. Lyu, H. Sun, D. Luo, H. Deng, Vision language models see what you want but not what you see, arXiv preprint arXiv:2410.00324 (2024)

  251. [260]

    Zhang, B

    M. Zhang, B. Colman, A. Shahriyari, G. Bharaj, et al., Common-sense bias discovery and mitigation for classification tasks, arXiv preprint arXiv:2401.13213 (2024)

  252. [261]

    Moayeri, W

    M. Moayeri, W. Wang, S. Singla, S. Feizi, Spuriosity rankings: Sorting data to measure and mitigate biases, Advances in Neural Information Processing Systems 36 (2024)

  253. [262]

    Chefer, S

    H. Chefer, S. Gur, L. Wolf, Generic attention-model explainability for interpreting bi- modal and encoder-decoder transformers, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 397–406

  254. [263]

    Y. Wang, T. G. Rudner, A. G. Wilson, Visual explanations of image-text representations via multi-modal information bottleneck attribution, Advances in Neural Information Processing Systems 36 (2024)

  255. [264]

    Park, Y.-J

    J.-H. Park, Y.-J. Ju, S.-W. Lee, Explaining generative diffusion models via visual analysis for interpretable decision-making process, Expert Systems with Applications 248 (2024) 123231

  256. [265]

    Y. Li, H. Wang, Y. Duan, X. Li, Clip surgery for better explainability with enhancement in open-vocabulary tasks, arXiv preprint arXiv:2304.05653 (2023)

  257. [266]

    P. Chen, Q. Li, S. Biaz, T. Bui, A. Nguyen, gscorecam: What objects is clip looking at?, in: Proceedings of the Asian Conference on Computer Vision, 2022, pp. 1959–1975

  258. [267]

    Y. Li, H. Wang, Y. Duan, H. Xu, X. Li, Exploring visual interpretability for contrastive language-image pre-training, arXiv preprint arXiv:2209.07046 (2022)

  259. [268]

    S. Arya, S. Rao, M. Boehle, B. Schiele, B-cosification: Transforming deep neural networks to be inherently interpretable, in: 38th Conference on Neural Information Processing Systems, 2024

  260. [269]

    Dewan, R

    S. Dewan, R. Zawar, P. Saxena, Y. Chang, A. Luo, Y. Bisk, Diffusionpid: Interpreting diffusion via partial information decomposition, arXiv preprint arXiv:2406.05191 (2024)

  261. [270]

    Bousselham, A

    W. Bousselham, A. Boggust, S. Chaybouti, H. Strobelt, H. Kuehne, Legrad: An explain- ability method for vision transformers via feature formation sensitivity, arXiv preprint arXiv:2404.03214 (2024)

  262. [271]

    B¨ ohle, M

    M. B¨ ohle, M. Fritz, B. Schiele, B-cos networks: Alignment is all we need for interpretabil- ity, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 10329–10338. 43

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.