Pith. sign in

REVIEW 2 major objections 4 minor 4 cited by

A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future

T0 review · 2 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This review maps 25 years of multimodal explainable AI onto four eras and three explanation stages, claiming the result is the latest and most complete survey comparison.

desk verdict A useful survey whose four-era historical framing is sloppily applied to the point of undercutting its central claim; worth revising, not dismissing. read the letter →

arxiv 2412.14056 v1 pith:OZ6VG7YB submitted 2024-12-18 cs.CV cs.AIcs.CLcs.LGcs.MM

classification cs.CVcs.AIcs.CLcs.LGcs.MM
keywords multimodalexplainableAIXAItaxonomyhistoricalreviewlargelanguagemodelsexplainabilityevaluationmetricsdatasetsfuturechallenges
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review proposes a historical map of Multimodal Explainable AI (MXAI): four eras — traditional machine learning (2000–2009), deep learning (2010–2016), discriminative foundation models (2017–2021), and generative large language models (2022–2024) — each examined through a fixed three-part lens of data, model, and post-hoc explainability. The paper's claim is that this structure, together with its comparison of prior XAI surveys, gives the latest and most complete picture of how MXAI methods evolved and where they stand now. A sympathetic reader would care because a well-founded historical taxonomy turns a scattered literature into a usable map: it lets researchers locate any method in time, see which techniques carry over across eras, and identify what is genuinely new in LLM-based explanation. The review also organizes evaluation metrics and datasets by task and closes by naming four open problems: hallucinations in multimodal LLMs, weak visual understanding, alignment with human cognition, and explainability when no ground truth exists.

What carries the argument

The carrying structure is a two-dimensional taxonomy: four eras defined by milestone models and datasets, crossed with three explanation stages (data, model, post-hoc). The era boundaries are anchored by named milestones — ImageNet in 2009 for the shift from traditional machine learning, the Transformer in 2017 for the shift to discriminative foundation models, and ChatGPT in 2022 for the shift to generative LLMs — while the stage trichotomy supplies a constant analytical grid, so methods from different decades can be compared on the same terms. This grid is what lets the paper present a historical narrative rather than a list of methods.

What would settle it

An audit that runs defined MXAI search queries across standard publication databases for 2000–2024 and finds either a substantial cluster of MXAI methods absent from the review's tables, or a clear case of mis-era assignment (for instance, a Transformer-based explainability method published before 2017 or a generative-LLM explanation method published before 2022), would refute the claimed historical completeness and the 'most comprehensive comparison' label.

Watch

Extended reading notes

Core claim

Stated on the paper's own terms, the central discovery is organizational: the evolution of MXAI can be read as four consecutive eras separated by three technological milestones — ImageNet (2009), Transformer (2017), and ChatGPT (2022) — and within every era the same three explanatory functions recur: explainability of data (before the model), of the model itself (inside the model), and post-hoc explanation of outputs (after the model). Surveying methods from decision trees and SVM rule extraction through CNN visualization and attention-based interpretation to in-context learning, chain-of-thought reasoning, and counterfactual prompting in multimodal LLMs, the paper argues that this three-stage structure persists even as the models and explanation targets change completely. On that basis it claims to be the latest and most comprehensive comparison of XAI reviews, and it uses the taxonomy to organize datasets, evaluation metrics, and future challenges.

Load-bearing premise

The review's historical completeness rests on the assumption that its hand-selected corpus of papers fairly represents the whole MXAI literature, because the scope is described only qualitatively with no search protocol, databases, query strings, or inclusion and exclusion criteria.

Editorial extensions

If this is right

  • A reader can take any MXAI method from 2000–2024 and locate it in one of twelve cells (four eras by three explanation stages), which makes gaps visible: for example, which stages are thin in which eras.
  • The review's survey comparison table gives a single baseline for judging what a new XAI survey must cover — unimodal and multimodal, traditional and LLM-based — before claiming novelty.
  • Because the taxonomy covers LLMs as the current era, it positions explainability techniques such as in-context learning, chain-of-thought, probing, and counterfactual prompting as direct successors of older post-hoc methods, not as a disconnected field.
  • The assembled datasets and metrics give practitioners a starting checklist for evaluating multimodal explanations by task (VQA, video captioning, recommendations, anomaly understanding).
  • The four named future challenges — hallucination mitigation, visual weakness in multimodal LLMs, alignment with human cognition, and explanation without ground truth — define a concrete research agenda that follows from the historical gaps.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the era boundary logic holds, a fifth era is already implied: as multimodal models become agentic and embodied, explanation may shift from describing model internals to justifying sequences of actions, which the data/model/post-hoc grid could still absorb but would likely need a new 'world' or 'interaction' category.
  • The taxonomy's completeness claim could be tested mechanically: an independent search using explicit query strings and inclusion criteria over 2000–2024 literature would reveal whether the hand-selected corpus misses a substantial cluster of MXAI work, and whether any method is assigned to the wrong era.
  • The three-stage lens treats explanation as a function of where it acts relative to the model; a natural extension is to classify explanations by who consumes them (model developer, domain expert, end user), since the paper's own evaluation discussion relies heavily on human evaluation, which is audience-dependent.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper is a narrative review of Multimodal Explainable AI (MXAI). It proposes a four-era historical taxonomy (traditional machine learning 2000-2009, deep learning 2010-2016, discriminative foundation models 2017-2021, and generative large language models 2022-2024), classifies MXAI methods in each era into data, model, and post-hoc explainability, compares the review with prior XAI surveys, and surveys evaluation datasets and metrics. It closes with future challenges such as hallucination mitigation, visual limitations of MLLMs, human-cognition alignment, and evaluation without ground truths.

Significance. If the historical taxonomy were applied consistently, this review would provide a useful map of MXAI evolution and a starting point for identifying open problems. The paper covers a wide range of methods and consolidates them into accessible summary tables, and the associated GitHub repository is a useful resource. However, the central historical claim is currently weakened by internal inconsistencies in era assignment, and the 'most comprehensive comparison' claim rests on an unspecified literature-selection procedure. These issues are fixable but require more than local copyediting.

major comments (2)
  1. [Section IV / Table III / Section II-A] The era boundaries defined in Section II-A are not respected in the era-specific discussion. The deep learning era is defined as 2010-2016, and the next era is defined as beginning with the Transformer in 2017, yet Section IV-B3 presents the Transformer (Vaswani et al., 2017, [1]) as an attention-based network of the deep learning era, and Table III lists 'Attention-based networks [1], [160]-[162]' within that era. The same era also contains multiple other 2017 works, including [158], [159], [162], [167], [170], [172], [173], [177], and [178]. Conversely, Section V (2017-2021) includes works published in 2023-2024, such as [204], [205], [206], [207], and [208]. Because the paper explicitly presents the four eras as chronological periods, these placements undermine the historical-periodization claim and would mislead a reader using the taxonomy as an evolution map. The authors should either reclassify the out-of-era references or explicitly reframe the eras as thematic waves with overlapping time spans.
  2. [Section II-B / Table I] The paper claims to provide 'the latest and most comprehensive comparison of XAI-related reviews,' but it does not describe a systematic search protocol: no databases, query strings, search dates, inclusion/exclusion criteria, or screening procedure are given. Section II-B describes the scope only qualitatively. Without such a protocol, the representativeness of the hand-selected corpus cannot be assessed, and the comprehensiveness claim is unsupported. The authors should either add a reproducible methodology subsection (databases, queries, screening steps) or temper the 'most comprehensive' claim to 'a broad comparison of selected recent surveys.'
minor comments (4)
  1. [Table II / Table III / Section II-A] The era date ranges are inconsistent between the text and table headers: Section II-A and the introduction give the traditional ML era as 2000-2009, but Table II says 2000-2010; the deep learning era is given as 2010-2016 in the text but Table III says 2011-2016. These should be aligned.
  2. [Table VI] Several metric names are misspelled: 'BLUE' should be 'BLEU' in the DD-VQA row, 'ROGUE' should be 'ROUGE' in the Rank2Tell row, 'Aideo' should be 'Audio' in the VAST row, and 'Dinstinct' should be 'Distinct' in the Yan et al. row. Since this table is the paper's consolidated resource for evaluation metrics, these typos impair usability.
  3. [Section II-B] The sentence 'we present the latest and most comprehensive comparison of XAI-related reviews from the past four years XAI-related reviews in the last four years as Table I' contains a duplicated, malformed phrase; it should be rewritten.
  4. [Figure 2] The figure contains label errors: 'Close-sourced' should be 'Closed-source' and 'Mingpt-4' should presumably be 'MiniGPT-4'. These should be corrected for consistency with the rest of the paper.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a narrative survey whose taxonomy is an organizational construct, not a derived prediction.

full rationale

This paper is a literature review rather than a derivation: it categorizes MXAI methods into four historical eras and three explainability types, and it summarizes external methods, datasets, and metrics. No equations are derived, no parameters are fitted, and no quantitative claim is produced from data defined by the claim itself. The era boundaries (2000-2009, 2010-2016, 2017-2021, 2022-2024) are anchored to external milestones such as ImageNet, the Transformer, and ChatGPT, so the taxonomy is not self-definitional in the sense of a fitted parameter renamed as a prediction. The paper does contain self-citations, e.g., references [270], [288], and [294] by the author team, but these are cited as examples of specific methods within the survey (generalized category discovery, attention-based hallucination mitigation, and knowledge acquisition for VQA), not as the justification for the taxonomy or for any load-bearing uniqueness claim. The comparison with prior surveys in Table I is a qualitative positioning statement, not a mathematical reduction. A reader concern about inconsistent era assignment, such as 2017 works appearing in the deep learning era section, would be a correctness or consistency issue, but it is not circularity under the definitions used here: the paper does not derive the taxonomy from itself, and its conclusions do not presuppose the taxonomy's correctness. The survey is self-contained as a narrative review, and the self-citations are minor and non-load-bearing. Therefore the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey introduces no free parameters or invented entities. Its conclusions rest on the choice of periodization and categorization, which are qualitative modeling assumptions, and on the unstated completeness of the hand-picked literature.

assumptions (3)
  • domain assumption The four-era division (2000-2009, 2010-2016, 2017-2021, 2022-2024) is a meaningful way to segment MXAI history.
    The entire structure of the review depends on this periodization; it is asserted in Section II without external justification or sensitivity analysis.
  • domain assumption The three explainability categories (data, model, post-hoc) are sufficient and mutually applicable across all eras.
    The review classifies every method into these categories (e.g., Tables II-V), but the categories are not justified as exhaustive or validated against other taxonomies.
  • domain assumption Non-multimodal methods from the traditional ML era (e.g., PCA, decision trees, SVMs) are relevant precursors of MXAI.
    Section III frames these unimodal techniques as foundations of multimodal explainability, which is a narrative choice rather than an empirically validated connection.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future." pith.science (2026). https://pith.science/paper/OZ6VG7YB

@misc{pith2026241214056,
  author       = {Pith},
  title        = {Pith review of: A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OZ6VG7YB}},
  note         = {Machine review of arXiv:2412.14056}
}
read the original abstract

Artificial intelligence (AI) has rapidly developed through advancements in computational power and the growth of massive datasets. However, this progress has also heightened challenges in interpreting the "black-box" nature of AI models. To address these concerns, eXplainable AI (XAI) has emerged with a focus on transparency and interpretability to enhance human understanding and trust in AI decision-making processes. In the context of multimodal data fusion and complex reasoning scenarios, the proposal of Multimodal eXplainable AI (MXAI) integrates multiple modalities for prediction and explanation tasks. Meanwhile, the advent of Large Language Models (LLMs) has led to remarkable breakthroughs in natural language processing, yet their complexity has further exacerbated the issue of MXAI. To gain key insights into the development of MXAI methods and provide crucial guidance for building more transparent, fair, and trustworthy AI systems, we review the MXAI methods from a historical perspective and categorize them across four eras: traditional machine learning, deep learning, discriminative foundation models, and generative LLMs. We also review evaluation metrics and datasets used in MXAI research, concluding with a discussion of future challenges and directions. A project related to this review has been created at https://github.com/ShilinSun/mxai_review.

Figures

Figures reproduced from arXiv: 2412.14056 by the authors.

Figure 1
Figure 1. We categorize MXAI into three types based on the data [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 1
Figure 1. Illustrative diagram of Multimodal Explainable Artificial Intelli [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. As AI advances, increasing computational power has led to larger [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Feature-level Interaction Explanations in Multimodal Transformers

    cs.LG 2026-03 conditional novelty 6.0 of 10

    FL-I2MoE separates unique, synergistic, and redundant cross-modal evidence at the token/patch level and uses SII and redundancy-gap scores to rank pairs whose removal degrades performance more than random masking.

  2. Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment

    cs.AI 2026-07 conditional novelty 5.0 of 10

    For document visual grounding at 4B scale, reasoning-free GRPO training outperforms a reasoning-enabled variant, and the reasoning variant compresses its traces during training.

  3. Regulatory Graphs and GenAI for Real-Time Transaction Monitoring and Compliance Explanation in Banking

    cs.AI 2025-06 reject novelty 3.0 of 10

    A GNN plus retrieval-augmented LLM pipeline for transaction monitoring is described, but the 98.2% F1 claim rests on an incomparable baseline setup and unreleased synthetic data.

  4. Empowering Multimodal LLMs with External Tools: A Comprehensive Survey

    cs.CV 2025-08 unverdicted novelty 2.0 of 10

    A survey paper maps how external tools are used to augment multimodal large language models across data, tasks, evaluation, and future directions.

Reference graph

Works this paper leans on

299 extracted references · 13 canonical work pages · cited by 4 Pith papers

  1. [1]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Proc. of NeurIPS, 2017

  2. [160]

    The application of two-level attention models in deep convolutional neural network for fine-grained image classification,

    T. Xiao, Y . Xu, K. Yang, J. Zhang, Y . Peng, and Z. Zhang, “The application of two-level attention models in deep convolutional neural network for fine-grained image classification,” in Proc. of CVPR, 2015, pp. 842–850

  3. [162]

    Human attention in visual question answering: Do humans and deep networks look at the same regions?

    A. Das, H. Agrawal, L. Zitnick, D. Parikh, and D. Batra, “Human attention in visual question answering: Do humans and deep networks look at the same regions?” Computer Vision and Image Understanding, pp. 90–100, 2017

  4. [158]

    Understanding black-box predictions via influence functions,

    P. W. Koh and P. Liang, “Understanding black-box predictions via influence functions,” in Proc. of ICML , 2017, pp. 1885–1894

  5. [159]

    Opening the black box of deep neural networks via information,

    R. Shwartz-Ziv and N. Tishby, “Opening the black box of deep neural networks via information,” arXiv preprint arXiv:1703.00810 , 2017

  6. [167]

    Right for the right reasons: Training differentiable models by constraining their explana- tions,

    A. S. Ross, M. C. Hughes, and F. Doshi-Velez, “Right for the right reasons: Training differentiable models by constraining their explana- tions,” arXiv preprint arXiv:1703.03717 , 2017

  7. [170]

    Explaining nonlinear classification decisions with deep taylor decom- position,

    G. Montavon, S. Lapuschkin, A. Binder, W. Samek, and K.-R. M ¨uller, “Explaining nonlinear classification decisions with deep taylor decom- position,” Pattern recognition, pp. 211–222, 2017

  8. [172]

    Axiomatic attribution for deep networks,

    M. Sundararajan, A. Taly, and Q. Yan, “Axiomatic attribution for deep networks,” in Proc. of ICML , 2017, pp. 3319–3328

  9. [173]

    Learning how to explain neural networks: Patternnet and patternattribution,

    P.-J. Kindermans, K. T. Sch ¨utt, M. Alber, K.-R. M ¨uller, D. Erhan, B. Kim, and S. D ¨ahne, “Learning how to explain neural networks: Patternnet and patternattribution,” arXiv preprint arXiv:1705.05598 , 2017

  10. [177]

    Grad-cam: Visual explanations from deep networks via gradient-based localization,

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proc. of ICCV , 2017, pp. 618–626

  11. [178]

    Explaining recurrent neural network predictions in sentiment analysis,

    L. Arras, G. Montavon, K.-R. M ¨uller, and W. Samek, “Explaining recurrent neural network predictions in sentiment analysis,” arXiv preprint arXiv:1706.07206, 2017

  12. [204]

    Learning to explain: A model-agnostic framework for explaining black box models,

    O. Barkan, Y . Asher, A. Eshel, Y . Elisha, and N. Koenigstein, “Learning to explain: A model-agnostic framework for explaining black box models,” in Proc. of ICDM , 2023, pp. 944–949

  13. [205]

    Clip surgery for better explain- ability with enhancement in open-vocabulary tasks,

    Y . Li, H. Wang, Y . Duan, and X. Li, “Clip surgery for better explain- ability with enhancement in open-vocabulary tasks,” arXiv preprint arXiv:2304.05653, 2023

  14. [206]

    Tagclip: A local-to-global framework to enhance open-vocabulary multi-label classification of clip without training,

    Y . Lin, M. Chen, K. Zhang, H. Li, M. Li, Z. Yang, D. Lv, B. Lin, H. Liu, and D. Cai, “Tagclip: A local-to-global framework to enhance open-vocabulary multi-label classification of clip without training,” in Proc. of AAAI , 2024, pp. 3513–3521

  15. [207]

    Interpreting clip’s image representation via text-based decomposition,

    Y . Gandelsman, A. A. Efros, and J. Steinhardt, “Interpreting clip’s image representation via text-based decomposition,” arXiv preprint arXiv:2310.05916, 2023

  16. [208]

    Black-box tuning of vision-language models with effective gradient approxima- tion,

    Z. Guo, Y . Wei, M. Liu, Z. Ji, J. Bai, Y . Guo, and W. Zuo, “Black-box tuning of vision-language models with effective gradient approxima- tion,” arXiv preprint arXiv:2312.15901 , 2023

Show all 299 references
  1. [2]

    Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

    J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,” in Proc. of ICML , 2023, pp. 19 730–19 742

  2. [3]

    Chatgpt: Get instant answers, find creative inspiration, learn something new,

    OpenAI, “Chatgpt: Get instant answers, find creative inspiration, learn something new,” 2022

  3. [4]

    Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning,

    E. Sachdeva, N. Agarwal, S. Chundi, S. Roelofs, J. Li, M. Kochen- derfer, C. Choi, and B. Dariush, “Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), ...

  4. [5]

    Unbox the black-box for the medical explainable ai via multi-modal and multi-centre data fusion: A mini- review, two showcases and beyond,

    G. Yang, Q. Ye, and J. Xia, “Unbox the black-box for the medical explainable ai via multi-modal and multi-centre data fusion: A mini- review, two showcases and beyond,” Information Fusion , pp. 29–52, 2022

  5. [7]

    Imagenet: A large-scale hierarchical image database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Proc. of CVPR , 2009, pp. 248–255

  6. [8]

    Visualizing higher- layer features of a deep network,

    D. Erhan, Y . Bengio, A. Courville, and P. Vincent, “Visualizing higher- layer features of a deep network,” University of Montreal , p. 1, 2009

  7. [9]

    Beyond intuition: Rethinking token attributions inside transformers,

    J. Chen, X. Li, L. Yu, D. Dou, and H. Xiong, “Beyond intuition: Rethinking token attributions inside transformers,” TMLR, 2022

  8. [10]

    Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers,

    H. Chefer, S. Gur, and L. Wolf, “Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers,” in Proc. of ICCV, 2021, pp. 397–406

  9. [11]

    Transformer interpretability beyond attention visualization,

    Chefer, Hila and Gur, Shir and Wolf, Lior, “Transformer interpretability beyond attention visualization,” in Proc. of CVPR, 2021, pp. 782–791

  10. [12]

    Explainable artificial intelligence (xai): What we know and what is left to attain trustworthy artificial intelligence,

    S. Ali, T. Abuhmed, S. El-Sappagh, K. Muhammad, J. M. Alonso- Moral, R. Confalonieri, R. Guidotti, J. Del Ser, N. D ´ıaz-Rodr´ıguez, and F. Herrera, “Explainable artificial intelligence (xai): What we know and what is left to attain trustworthy artificial intelligence,” Inform...

  11. [13]

    Rethinking interpretability in the era of large language models,

    C. Singh, J. P. Inala, M. Galley, R. Caruana, and J. Gao, “Rethinking interpretability in the era of large language models,” arXiv preprint arXiv:2402.01761, 2024

  12. [14]

    Usable xai: 10 strategies towards exploiting explainability in the llm era,

    X. Wu, H. Zhao, Y . Zhu, Y . Shi, F. Yang, T. Liu, X. Zhai, W. Yao, J. Li, and M. Du, “Usable xai: 10 strategies towards exploiting explainability in the llm era,” arXiv preprint arXiv:2403.08946 , 2024

  13. [15]

    From understanding to utilization: A sur- vey on explainability for large language models,

    H. Luo and L. Specia, “From understanding to utilization: A sur- vey on explainability for large language models,” arXiv preprint arXiv:2401.12874, 2024

  14. [16]

    Feature selection based on mu- tual information criteria of max-dependency, max-relevance, and min- redundancy,

    H. Peng, F. Long, and C. Ding, “Feature selection based on mu- tual information criteria of max-dependency, max-relevance, and min- redundancy,” IEEE Transactions on pattern analysis and machine intelligence, pp. 1226–1238, 2005

  15. [17]

    Rule-based learning systems for support vector machines,

    H. N ´u˜nez, C. Angulo, and A. Catala, “Rule-based learning systems for support vector machines,” Neural Processing Letters , pp. 1–18, 2006

  16. [18]

    Assessment of landslide susceptibility by decision trees in the metropolitan area of istanbul, turkey,

    H. NEFESL ˙IO ˘GLU, E. Sezer, C. G ¨OKC ¸ EO˘GLU, A. BOZKIR, and T. Duman, “Assessment of landslide susceptibility by decision trees in the metropolitan area of istanbul, turkey,” Mathematical Problems in Engineering, 2010

  17. [19]

    Mining high-speed data streams,

    P. Domingos and G. Hulten, “Mining high-speed data streams,” in Proc. of KDD, 2000, pp. 71–80

  18. [20]

    Auto-encoding variational bayes,

    D. P. Kingma and M. Welling, “Auto-encoding variational bayes,”arXiv preprint arXiv:1312.6114, 2013

  19. [21]

    beta-vae: Learning basic vi- sual concepts with a constrained variational framework

    I. Higgins, L. Matthey, A. Pal, C. P. Burgess, X. Glorot, M. M. Botvinick, S. Mohamed, and A. Lerchner, “beta-vae: Learning basic vi- sual concepts with a constrained variational framework.”ICLR (Poster), 2017

  20. [22]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in Proc. of ICML , 2021, pp. 8748–8763

  21. [23]

    I. T. Jolliffe, Principal component analysis for special types of data . Springer, 2002

  22. [24]

    Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks,

    A. Chattopadhay, A. Sarkar, P. Howlader, and V . N. Balasubramanian, “Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks,” in Proc. of WACV, 2018, pp. 839–847

  23. [25]

    Multimodal explainable artificial intelligence: A comprehensive review of methodological advances and future research directions,

    N. Rodis, C. Sardianos, G. T. Papadopoulos, P. Radoglou-Grammatikis, P. Sarigiannidis, and I. Varlamis, “Multimodal explainable artificial intelligence: A comprehensive review of methodological advances and future research directions,” arXiv preprint arXiv:2306.05731 , 2023

  24. [26]

    Explainable generative ai (genxai): A survey, conceptual- ization, and research agenda,

    J. Schneider, “Explainable generative ai (genxai): A survey, conceptual- ization, and research agenda,” arXiv preprint arXiv:2404.09554, 2024

  25. [27]

    Explainability for large language models: A survey,

    H. Zhao, H. Chen, F. Yang, N. Liu, H. Deng, H. Cai, S. Wang, D. Yin, and M. Du, “Explainability for large language models: A survey,” ACM Transactions on Intelligent Systems and Technology , pp. 1–38, 2024

  26. [28]

    Explainable artificial intelligence (xai) from a user perspective: A synthesis of prior literature and problematizing avenues for future research,

    A. K. M. B. Haque, A. K. M. N. Islam, and P. Mikalef, “Explainable artificial intelligence (xai) from a user perspective: A synthesis of prior literature and problematizing avenues for future research,” Technolog- ical Forecasting and Social Change , 2023

  27. [29]

    A systematic review of explainable artificial intelligence models and applications: Recent developments and future trends,

    S. A and S. R, “A systematic review of explainable artificial intelligence models and applications: Recent developments and future trends,” Decision Analytics Journal , 2023

  28. [30]

    Ex- plainability of vision transformers: A comprehensive review and new perspectives,

    R. Kashefi, L. Barekatain, M. Sabokrou, and F. Aghaeipoor, “Ex- plainability of vision transformers: A comprehensive review and new perspectives,” arXiv preprint arXiv:2311.06786 , 2023

  29. [31]

    Explain- ability and evaluation of vision transformers: An in-depth experimental study,

    S. Stassin, V . Corduant, S. A. Mahmoudi, and X. Siebert, “Explain- ability and evaluation of vision transformers: An in-depth experimental study,” Electronics, 2023

  30. [32]

    Towards human-centered explainable ai: A survey of user studies for model explanations,

    Y . Rong, T. Leemann, T.-T. Nguyen, L. Fiedler, P. Qian, V . Unhelkar, T. Seidel, G. Kasneci, and E. Kasneci, “Towards human-centered explainable ai: A survey of user studies for model explanations,” IEEE transactions on pattern analysis and machine intelligence , 2023

  31. [33]

    Explainable ai (xai): A systematic meta- survey of current challenges and future opportunities,

    W. Saeed and C. Omlin, “Explainable ai (xai): A systematic meta- survey of current challenges and future opportunities,” Knowledge- Based Systems, 2023

  32. [34]

    A survey on xai and natural language explanations,

    E. Cambria, L. Malandri, F. Mercorio, M. Mezzanzanica, and N. Nobani, “A survey on xai and natural language explanations,” Information Processing & Management , p. 103111, 2023

  33. [35]

    Knowledge graphs as tools for explainable machine learning: A survey,

    I. Tiddi and S. Schlobach, “Knowledge graphs as tools for explainable machine learning: A survey,” Artificial Intelligence, p. 103627, 2022. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 14

  34. [36]

    Explainable ai methods-a brief overview,

    A. Holzinger, A. Saranti, C. Molnar, P. Biecek, and W. Samek, “Explainable ai methods-a brief overview,” in International workshop on extending explainable AI beyond deep models and classifiers , 2022, pp. 13–38

  35. [37]

    Counterfactual explanations and how to find them: literature review and benchmarking,

    R. Guidotti, “Counterfactual explanations and how to find them: literature review and benchmarking,” Data Mining and Knowledge Discovery, pp. 1–55, 2022

  36. [38]

    Explainable ai for time series classification: a review, taxonomy and research directions,

    A. Theissler, F. Spinnato, U. Schlegel, and R. Guidotti, “Explainable ai for time series classification: a review, taxonomy and research directions,” Ieee Access, pp. 100 700–100 724, 2022

  37. [39]

    A survey of contrastive and counterfactual explanation generation methods for explainable artificial intelligence,

    I. Stepin, J. M. Alonso, A. Catala, and M. Pereira-Fari ˜na, “A survey of contrastive and counterfactual explanation generation methods for explainable artificial intelligence,” IEEE Access , pp. 11 974–12 001, 2021

  38. [40]

    Notions of explainability and evaluation approaches for explainable artificial intelligence,

    G. Vilone and L. Longo, “Notions of explainability and evaluation approaches for explainable artificial intelligence,” Information Fusion, pp. 89–106, 2021

  39. [41]

    Explainable artificial intelligence: objectives, stakeholders, and future research opportunities,

    C. Meske, E. Bunde, J. Schneider, and M. Gersch, “Explainable artificial intelligence: objectives, stakeholders, and future research opportunities,” Information Systems Management , pp. 53–63, 2022

  40. [42]

    Explainable ai: A review of machine learning interpretability methods,

    P. Linardatos, V . Papastefanopoulos, and S. Kotsiantis, “Explainable ai: A review of machine learning interpretability methods,” Entropy, p. 18, 2020

  41. [43]

    Ex- plainable artificial intelligence approaches: A survey,

    S. R. Islam, W. Eberle, S. K. Ghafoor, and M. Ahmed, “Ex- plainable artificial intelligence approaches: A survey,” arXiv preprint arXiv:2101.09429, 2021

  42. [44]

    Argumentation and ex- plainable artificial intelligence: a survey,

    A. Vassiliades, N. Bassiliades, and T. Patkos, “Argumentation and ex- plainable artificial intelligence: a survey,” The Knowledge Engineering Review, p. e5, 2021

  43. [45]

    A survey on the explainability of super- vised machine learning,

    N. Burkart and M. F. Huber, “A survey on the explainability of super- vised machine learning,” Journal of Artificial Intelligence Research , pp. 245–317, 2021

  44. [46]

    Reviewing the need for explainable artificial intelligence (xai),

    J. Gerlings, A. Shollo, and I. Constantiou, “Reviewing the need for explainable artificial intelligence (xai),” in Proceedings of the 54th Hawaii International Conference on System Sciences (HICSS) , 2021

  45. [47]

    Explainable artifi- cial intelligence (xai): An engineering perspective,

    F. Hussain, R. Hussain, and E. Hossain, “Explainable artifi- cial intelligence (xai): An engineering perspective,” arXiv preprint arXiv:2101.03613, 2021

  46. [48]

    Interpretable deep learning: Interpretation, interpretability, trustworthiness, and beyond,

    X. Li, H. Xiong, X. Li, X. Wu, X. Zhang, J. Liu, J. Bian, and D. Dou, “Interpretable deep learning: Interpretation, interpretability, trustworthiness, and beyond,” Knowledge and Information Systems, pp. 3197–3234, 2022

  47. [49]

    Minimum redundancy feature selection from microarray gene expression data,

    C. Ding and H. Peng, “Minimum redundancy feature selection from microarray gene expression data,” Journal of bioinformatics and com- putational biology, pp. 185–205, 2005

  48. [50]

    Feature selection for high-dimensional data—a pearson redundancy based filter,

    J. Biesiada and W. Duch, “Feature selection for high-dimensional data—a pearson redundancy based filter,” in Computer recognition systems 2, 2008, pp. 242–249

  49. [51]

    Emotion recognition using brain activity,

    R. Horlings, D. Datcu, and L. J. Rothkrantz, “Emotion recognition using brain activity,” in Proceedings of the 9th international conference on computer systems and technologies and workshop for PhD students in computing, 2008, pp. II–1

  50. [52]

    Normalized mutual information feature selection,

    P. A. Est ´evez, M. Tesmer, C. A. Perez, and J. M. Zurada, “Normalized mutual information feature selection,” IEEE Transactions on neural networks, pp. 189–201, 2009

  51. [53]

    Optimization method based extreme learning machine for classification,

    G.-B. Huang, X. Ding, and H. Zhou, “Optimization method based extreme learning machine for classification,”Neurocomputing, pp. 155– 163, 2010

  52. [54]

    Eye move- ment analysis for activity recognition using electrooculography,

    A. Bulling, J. A. Ward, H. Gellersen, and G. Tr ¨oster, “Eye move- ment analysis for activity recognition using electrooculography,” IEEE transactions on pattern analysis and machine intelligence , pp. 741– 753, 2010

  53. [55]

    Backward sequential elimination for sparse vector subset selection,

    S. F. Cotter, K. Kreutz-Delgado, and B. D. Rao, “Backward sequential elimination for sparse vector subset selection,” Signal Processing, pp. 1849–1864, 2001

  54. [56]

    Gene selection for cancer classification using support vector machines,

    I. Guyon, J. Weston, S. Barnhill, and V . Vapnik, “Gene selection for cancer classification using support vector machines,”Machine learning, pp. 389–422, 2002

  55. [57]

    Feature subset selection for blood pressure classification using orthogonal forward selection,

    S. Colak and C. Isik, “Feature subset selection for blood pressure classification using orthogonal forward selection,” in 2003 IEEE 29th Annual Proceedings of Bioengineering Conference, 2003, pp. 122–123

  56. [58]

    Feature selection for high-dimensional industrial data

    M. Bensch, M. Schr ¨oder, M. Bogdan, and W. Rosenstiel, “Feature selection for high-dimensional industrial data.” in ESANN, 2005, pp. 375–380

  57. [59]

    Random forests,

    L. Breiman, “Random forests,” Machine learning, pp. 5–32, 2001

  58. [60]

    Exploration of a hybrid feature selection algorithm,

    K.-M. Osei-Bryson, K. Giles, and B. Kositanurit, “Exploration of a hybrid feature selection algorithm,” Journal of the Operational Research Society, pp. 790–797, 2003

  59. [61]

    Ant colony optimization for feature selection in face recognition,

    Z. Yan and C. Yuan, “Ant colony optimization for feature selection in face recognition,” in Biometric Authentication: First International Conference, ICBA 2004, Hong Kong, China, July 15-17, 2004. Pro- ceedings, 2004, pp. 221–226

  60. [62]

    Optimization of intrusion detection through fast hybrid feature selection,

    K. M. Shazzad and J. S. Park, “Optimization of intrusion detection through fast hybrid feature selection,” in Sixth International Conference on Parallel and Distributed Computing Applications and Technologies (PDCAT’05), 2005, pp. 264–267

  61. [63]

    A wrapper for feature selection based on mutual information,

    J. Huang, Y . Cai, and X. Xu, “A wrapper for feature selection based on mutual information,” in Proc. of ICPR , 2006, pp. 618–621

  62. [64]

    A hybrid feature selection approach for microarray gene expression data,

    F. Tan, X. Fu, H. Wang, Y . Zhang, and A. Bourgeois, “A hybrid feature selection approach for microarray gene expression data,” in Computa- tional Science–ICCS 2006: 6th International Conference, Reading, UK, May 28-31, 2006. Proceedings, Part II 6 , 2006, pp. 678–685

  63. [65]

    Application of a hybrid wavelet feature selection method in the design of a self-paced brain interface system,

    M. Fatourechi, G. E. Birch, and R. K. Ward, “Application of a hybrid wavelet feature selection method in the design of a self-paced brain interface system,” Journal of neuroengineering and rehabilitation , pp. 1–13, 2007

  64. [66]

    Feature subset selection in large dimensionality domains,

    I. A. Gheyas and L. S. Smith, “Feature subset selection in large dimensionality domains,” Pattern recognition, pp. 5–13, 2010

  65. [67]

    Independent component analysis: algorithms and applications,

    A. Hyv ¨arinen and E. Oja, “Independent component analysis: algorithms and applications,” Neural networks, pp. 411–430, 2000

  66. [68]

    A global geometric framework for nonlinear dimensionality reduction,

    J. B. Tenenbaum, V . d. Silva, and J. C. Langford, “A global geometric framework for nonlinear dimensionality reduction,” science, pp. 2319– 2323, 2000

  67. [69]

    Nonlinear dimensionality reduction by locally linear embedding,

    S. T. Roweis and L. K. Saul, “Nonlinear dimensionality reduction by locally linear embedding,” science, pp. 2323–2326, 2000

  68. [70]

    Pca versus lda,

    A. M. Martinez and A. C. Kak, “Pca versus lda,” IEEE transactions on pattern analysis and machine intelligence , pp. 228–233, 2001

  69. [71]

    Two-dimensional linear discriminant analysis,

    J. Ye, R. Janardan, and Q. Li, “Two-dimensional linear discriminant analysis,” Proc. of NeurIPS , 2004

  70. [72]

    Two- dimensional linear discriminant analysis of principle component vectors for face recognition,

    P. Sanguansat, W. Asdornwised, S. Jitapunkul, and S. Marukatat, “Two- dimensional linear discriminant analysis of principle component vectors for face recognition,” IEICE transactions on information and systems , pp. 2164–2170, 2006

  71. [73]

    Visualizing data using t-sne

    L. Van der Maaten and G. Hinton, “Visualizing data using t-sne.” Journal of machine learning research , 2008

  72. [74]

    A note on two-dimensional linear discriminant analysis,

    Z. Liang, Y . Li, and P. Shi, “A note on two-dimensional linear discriminant analysis,” Pattern Recognition Letters , pp. 2122–2128, 2008

  73. [75]

    Two-dimensional direct and weighted linear discriminant analysis for face recognition,

    R. Zhi and Q. Ruan, “Two-dimensional direct and weighted linear discriminant analysis for face recognition,” Neurocomputing, pp. 3607– 3611, 2008

  74. [76]

    On the re- lationship between feature selection and classification accuracy,

    A. Janecek, W. Gansterer, M. Demel, and G. Ecker, “On the re- lationship between feature selection and classification accuracy,” in New challenges for feature selection in data mining and knowledge discovery, 2008, pp. 90–105

  75. [77]

    What is principal component analysis?

    M. Ringn ´er, “What is principal component analysis?” Nature biotech- nology, pp. 303–304, 2008

  76. [78]

    1d-lda vs. 2d-lda: When is vector- based linear discriminant analysis better than matrix-based?

    W.-S. Zheng, J.-H. Lai, and S. Z. Li, “1d-lda vs. 2d-lda: When is vector- based linear discriminant analysis better than matrix-based?” Pattern Recognition, pp. 2156–2172, 2008

  77. [79]

    Generalized linear discriminant analysis: a unified framework and efficient model selection,

    S. Ji and J. Ye, “Generalized linear discriminant analysis: a unified framework and efficient model selection,”IEEE Transactions on Neural Networks, pp. 1768–1782, 2008

  78. [80]

    Mpca: Multilin- ear principal component analysis of tensor objects,

    H. Lu, K. N. Plataniotis, and A. N. Venetsanopoulos, “Mpca: Multilin- ear principal component analysis of tensor objects,” IEEE transactions on Neural Networks , pp. 18–39, 2008

  79. [81]

    Robust principal component analysis: Exact recovery of corrupted low-rank matrices via convex optimization,

    J. Wright, A. Ganesh, S. Rao, Y . Peng, and Y . Ma, “Robust principal component analysis: Exact recovery of corrupted low-rank matrices via convex optimization,” Proc. of NeurIPS , 2009

  80. [82]

    Structured sparse principal component analysis,

    R. Jenatton, G. Obozinski, and F. Bach, “Structured sparse principal component analysis,” in Proc. of AISTATS, 2010, pp. 366–373

  81. [83]

    Two-stage image denoising by principal component analysis with local pixel grouping,

    L. Zhang, W. Dong, D. Zhang, and G. Shi, “Two-stage image denoising by principal component analysis with local pixel grouping,” Pattern recognition, pp. 1531–1549, 2010

  82. [84]

    Unsupervised kernel dimension reduction,

    M. Wang, F. Sha, and M. Jordan, “Unsupervised kernel dimension reduction,” Proc. of NeurIPS , 2010

  83. [85]

    Jaccard, Interaction effects in logistic regression

    J. Jaccard, Interaction effects in logistic regression . Sage, 2001

  84. [86]

    Pur- poseful selection of variables in logistic regression,

    Z. Bursac, C. H. Gauss, D. K. Williams, and D. W. Hosmer, “Pur- poseful selection of variables in logistic regression,” Source code for biology and medicine , pp. 1–8, 2008. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 15

  85. [87]

    An introduction to logistic regression analysis and reporting,

    C.-Y . J. Peng, K. L. Lee, and G. M. Ingersoll, “An introduction to logistic regression analysis and reporting,” The journal of educational research, pp. 3–14, 2002

  86. [88]

    Logistic regression: Why we cannot do what we think we can do, and what we can do about it,

    C. Mood, “Logistic regression: Why we cannot do what we think we can do, and what we can do about it,” European sociological review, pp. 67–82, 2010

  87. [89]

    Rocr: visualiz- ing classifier performance in r,

    T. Sing, O. Sander, N. Beerenwinkel, and T. Lengauer, “Rocr: visualiz- ing classifier performance in r,” Bioinformatics, pp. 3940–3941, 2005

  88. [90]

    In- duction of decision trees using an internal control of induction,

    G. Ramos-Jim ´enez, J. del Campo- ´Avila, and R. Morales-Bueno, “In- duction of decision trees using an internal control of induction,” in International Work-Conference on Artificial Neural Networks , 2005, pp. 795–803

  89. [91]

    Supporting the dynamic evolution of web service protocols in service- oriented architectures,

    S. H. Ryu, F. Casati, H. Skogsrud, B. Benatallah, and R. Saint-Paul, “Supporting the dynamic evolution of web service protocols in service- oriented architectures,” ACM Transactions on the Web (TWEB) , pp. 1–46, 2008

  90. [92]

    Learning to match xml schemas-a deci- sion tree based approach,

    A. Rajesh and S. Srivatsa, “Learning to match xml schemas-a deci- sion tree based approach,” International Journal of Recent Trends in Engineering, p. 58, 2009

  91. [93]

    An k nn model-based approach and its application in text categorization,

    G. Guo, H. Wang, D. Bell, Y . Bi, and K. Greer, “An k nn model-based approach and its application in text categorization,” in Computational Linguistics and Intelligent Text Processing: 5th International Confer- ence, CICLing 2004 Seoul, Korea, February 15-21, 2004 Proceedings ...

  92. [94]

    Application of the ga/knn method to seldi proteomics data,

    L. Li, D. M. Umbach, P. Terry, and J. A. Taylor, “Application of the ga/knn method to seldi proteomics data,” Bioinformatics, pp. 1638– 1640, 2004

  93. [95]

    The truth is in there- rule extraction from opaque models using genetic programming

    U. Johansson, R. K ¨onig, and L. Niklasson, “The truth is in there- rule extraction from opaque models using genetic programming.” in FLAIRS, 2004, pp. 658–663

  94. [96]

    Rule extraction from support vector machines

    H. N ´u˜nez, C. Angulo, and A. Catal `a, “Rule extraction from support vector machines.” in Esann, 2002, pp. 107–112

  95. [97]

    Bayesian models of cognition,

    T. L Griffiths, C. Kemp, and J. B Tenenbaum, “Bayesian models of cognition,” 2008

  96. [98]

    A bayesian model for repeated measures zero-inflated count data with application to outpatient psychiatric service use,

    B. H. Neelon, A. J. O’Malley, and S.-L. T. Normand, “A bayesian model for repeated measures zero-inflated count data with application to outpatient psychiatric service use,” Statistical modelling , pp. 421– 439, 2010

  97. [99]

    Probabilistic climate change predictions applying bayesian model averaging,

    S.-K. Min, D. Simonis, and A. Hense, “Probabilistic climate change predictions applying bayesian model averaging,” Philosophical Trans- actions of the Royal Society A: Mathematical, Physical and Engineer- ing Sciences, pp. 2103–2116, 2007

  98. [100]

    Bayesian econometric methods cambridge, uk: Cambridge univ,

    G. Koop, D. Poirier, and J. Tobias, “Bayesian econometric methods cambridge, uk: Cambridge univ,” Press357, 2007

  99. [101]

    Trust and tam in online shopping: An integrated model,

    Gefen, Karahanna, and Straub, “Trust and tam in online shopping: An integrated model,” MIS Quarterly , p. 51, Jan 2003. [Online]. Available: http://dx.doi.org/10.2307/30036519

  100. [102]

    Woodward, Making things happen: A theory of causal explanation

    J. Woodward, Making things happen: A theory of causal explanation . Oxford university press, 2005

  101. [103]

    Explanation and understanding,

    F. C. Keil, “Explanation and understanding,” Annu. Rev. Psychol., pp. 227–254, 2006

  102. [104]

    Causal inference in statistics: An overview,

    J. Pearl, “Causal inference in statistics: An overview,” 2009

  103. [105]

    R. A. Berk et al. , Statistical learning from a regression perspective . Springer, 2008

  104. [106]

    Predictive analytics in information systems research,

    G. Shmueli and O. R. Koppius, “Predictive analytics in information systems research,” MIS quarterly, pp. 553–572, 2011

  105. [107]

    The interplay between theory and method,

    J. Van Maanen, J. B. Sørensen, and T. R. Mitchell, “The interplay between theory and method,” Academy of management review , pp. 1145–1154, 2007

  106. [108]

    To explain or to predict?

    G. Shmueli, “To explain or to predict?” 2010

  107. [109]

    Causal inference without counterfactuals,

    A. P. Dawid, “Causal inference without counterfactuals,” Journal of the American statistical Association , pp. 407–424, 2000

  108. [110]

    Evidence-based public health: moving beyond randomized trials,

    C. G. Victora, J.-P. Habicht, and J. Bryce, “Evidence-based public health: moving beyond randomized trials,” American journal of public health, pp. 400–405, 2004

  109. [111]

    Varieties of causal intervention,

    K. B. Korb, L. R. Hope, A. E. Nicholson, and K. Axnick, “Varieties of causal intervention,” in Proc. of PRICAI , 2004, pp. 322–331

  110. [112]

    Causal reasoning through intervention,

    Y . Hagmayer, S. A. Sloman, D. A. Lagnado, and M. R. Waldmann, “Causal reasoning through intervention,” Causal learning: Psychology, philosophy, and computation , pp. 86–100, 2007

  111. [113]

    Toward causal inference with interference,

    M. G. Hudgens and M. E. Halloran, “Toward causal inference with interference,” Journal of the American Statistical Association, pp. 832– 842, 2008

  112. [114]

    Pearl, Causality

    J. Pearl, Causality. Cambridge university press, 2009

  113. [115]

    Feature deduction and ensemble design of intrusion detection systems,

    S. Chebrolu, A. Abraham, and J. P. Thomas, “Feature deduction and ensemble design of intrusion detection systems,”Computers & security, pp. 295–307, 2005

  114. [116]

    Building efficient intrusion detection model based on principal component analysis and c4. 5,

    Y . Chen, Y . Li, X.-Q. Cheng, and L. Guo, “Building efficient intrusion detection model based on principal component analysis and c4. 5,” in 2006 International Conference on Communication Technology , 2006, pp. 1–4

  115. [117]

    Transferring knowl- edge by prior feature sampling,

    V . Eruhimov, V . Martyanov, and A. Polovinkin, “Transferring knowl- edge by prior feature sampling,” in New Challenges for Feature Selection in Data Mining and Knowledge Discovery , 2008, pp. 135– 147

  116. [118]

    Density estimation trees,

    P. Ram and A. G. Gray, “Density estimation trees,” in Proc. of KDD , 2011, pp. 627–635

  117. [119]

    Towards an effective cooper- ation of the user and the computer for classification,

    M. Ankerst, M. Ester, and H.-P. Kriegel, “Towards an effective cooper- ation of the user and the computer for classification,” in Proc. of KDD, 2000, pp. 179–188

  118. [120]

    Data mining and visualization for decision support and modeling of public health-care resources,

    N. Lavra ˇc, M. Bohanec, A. Pur, B. Cestnik, M. Debeljak, and A. Kobler, “Data mining and visualization for decision support and modeling of public health-care resources,” Journal of biomedical informatics, pp. 438–447, 2007

  119. [121]

    A bias correction algorithm for the gini variable importance measure in classification trees,

    M. Sandri and P. Zuccolotto, “A bias correction algorithm for the gini variable importance measure in classification trees,” Journal of Computational and Graphical Statistics , pp. 611–628, 2008

  120. [122]

    A comparison of random forest and its gini importance with standard chemometric methods for the feature selection and classification of spectral data,

    B. H. Menze, B. M. Kelm, R. Masuch, U. Himmelreich, P. Bachert, W. Petrich, and F. A. Hamprecht, “A comparison of random forest and its gini importance with standard chemometric methods for the feature selection and classification of spectral data,” BMC bioinformatics, pp. 1–16, 2009

  121. [123]

    Visualization method and tool for interactive learning of large decision trees,

    T. D. Nguyen and T. Ho, “Visualization method and tool for interactive learning of large decision trees,” in Data Mining and Knowledge Discovery: Theory, Tools, and Technology IV , 2002, pp. 34–42

  122. [124]

    Analysis and correction of bias in total decrease in node impurity measures for tree-based algorithms,

    M. Sandri and P. Zuccolotto, “Analysis and correction of bias in total decrease in node impurity measures for tree-based algorithms,” Statistics and Computing , pp. 393–407, 2010

  123. [125]

    Visual classifica- tion: an interactive approach to decision tree construction,

    M. Ankerst, C. Elsen, M. Ester, and H.-P. Kriegel, “Visual classifica- tion: an interactive approach to decision tree construction,” in Proc. of KDD, 1999, pp. 392–396

  124. [126]

    Eclectic rule-extraction from support vector machines,

    N. Barakat and J. Diederich, “Eclectic rule-extraction from support vector machines,” International Journal of Computer and Information Engineering, pp. 1672–1675, 2008

  125. [127]

    Extracting the knowledge embedded in support vector machines,

    X. Fu, C. Ong, S. Keerthi, G. G. Hung, and L. Goh, “Extracting the knowledge embedded in support vector machines,” in Proc. of IJCNN, 2004, pp. 291–296

  126. [128]

    Support vector machines with symbolic interpretation,

    H. N ´u˜nez, C. Angulo, and A. Catal `a, “Support vector machines with symbolic interpretation,” in VII Brazilian Symposium on Neural Networks, 2002. SBRN 2002. Proceedings. , 2002, pp. 142–147

  127. [129]

    Rule extraction from trained support vector machines,

    Y . Zhang, H. Su, T. Jia, and J. Chu, “Rule extraction from trained support vector machines,” in Proc. of KDD , 2005, pp. 61–70

  128. [130]

    Rule extraction from linear support vector machines,

    G. Fung, S. Sandilya, and R. B. Rao, “Rule extraction from linear support vector machines,” in Proc. of KDD , 2005, pp. 32–40

  129. [131]

    A multiple kernel support vector machine scheme for feature selection and rule extraction from gene expression data of cancer tissue,

    Z. Chen, J. Li, and L. Wei, “A multiple kernel support vector machine scheme for feature selection and rule extraction from gene expression data of cancer tissue,” Artificial intelligence in medicine , pp. 161–175, 2007

  130. [132]

    Support vector machine interpretation,

    A. Navia-V ´azquez and E. Parrado-Hern´andez, “Support vector machine interpretation,” Neurocomputing, pp. 1754–1759, 2006

  131. [133]

    Paintingclass: interactive construction, visualization and exploration of decision trees,

    S. T. Teoh and K.-L. Ma, “Paintingclass: interactive construction, visualization and exploration of decision trees,” in Proc. of KDD, 2003, pp. 667–672

  132. [134]

    Linked independent component analysis for multimodal data fusion,

    A. R. Groves, C. F. Beckmann, S. M. Smith, and M. W. Woolrich, “Linked independent component analysis for multimodal data fusion,” Neuroimage, pp. 2198–2217, 2011

  133. [135]

    Simultaneous analysis of coupled data matrices subject to different amounts of noise,

    T. F. Wilderjans, E. Ceulemans, I. Van Mechelen, and R. A. Van Den Berg, “Simultaneous analysis of coupled data matrices subject to different amounts of noise,” British Journal of Mathematical and Statistical Psychology, pp. 277–290, 2011

  134. [136]

    Multisensor data fusion: A review of the state-of-the-art,

    B. Khaleghi, A. Khamis, F. O. Karray, and S. N. Razavi, “Multisensor data fusion: A review of the state-of-the-art,” Information fusion , pp. 28–44, 2013

  135. [137]

    New algorithm for integration between wireless microwave sensor network and radar for improved rainfall measurement and mapping,

    Y . Liberman, R. Samuels, P. Alpert, and H. Messer, “New algorithm for integration between wireless microwave sensor network and radar for improved rainfall measurement and mapping,” Atmospheric Mea- surement Techniques, pp. 3549–3563, 2014

  136. [138]

    Multimodal data fusion: an overview of methods, challenges, and prospects,

    D. Lahat, T. Adali, and C. Jutten, “Multimodal data fusion: an overview of methods, challenges, and prospects,” Proceedings of the IEEE , pp. 1449–1477, 2015

  137. [139]

    Dynamical reconfiguration strategy of a multi sensor data fusion algorithm based on information theory,

    N. A. Tmazirte, M. E. El Najjar, C. Smaili, and D. Pomorski, “Dynamical reconfiguration strategy of a multi sensor data fusion algorithm based on information theory,” in 2013 IEEE intelligent vehicles symposium (IV) , 2013, pp. 896–901. JOURNAL OF LATEX CLASS FILES, VOL. 14, N...

  138. [140]

    Scalable ten- sor factorizations for incomplete data,

    E. Acar, D. M. Dunlavy, T. G. Kolda, and M. Mørup, “Scalable ten- sor factorizations for incomplete data,” Chemometrics and Intelligent Laboratory Systems, pp. 41–56, 2011

  139. [141]

    Show and tell: A neural image caption generator,

    O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in Proc. of CVPR , 2015, pp. 3156– 3164

  140. [142]

    Multimodal sentiment analysis with word-level fusion and reinforcement learning,

    M. Chen, S. Wang, P. P. Liang, T. Baltru ˇsaitis, A. Zadeh, and L.-P. Morency, “Multimodal sentiment analysis with word-level fusion and reinforcement learning,” in Proceedings of the 19th ACM international conference on multimodal interaction , 2017, pp. 163–171

  141. [143]

    Sequence to sequence-video to text,

    S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko, “Sequence to sequence-video to text,” in Proc. of ICCV , 2015, pp. 4534–4542

  142. [144]

    Human action recognition using factorized spatio-temporal convolutional networks,

    L. Sun, K. Jia, D.-Y . Yeung, and B. E. Shi, “Human action recognition using factorized spatio-temporal convolutional networks,” in Proc. of ICCV, 2015, pp. 4597–4605

  143. [145]

    Supersparse linear integer models for op- timized medical scoring systems,

    B. Ustun and C. Rudin, “Supersparse linear integer models for op- timized medical scoring systems,” Machine Learning , pp. 349–391, 2016

  144. [146]

    Interpretable decision sets: A joint framework for description and prediction,

    H. Lakkaraju, S. H. Bach, and J. Leskovec, “Interpretable decision sets: A joint framework for description and prediction,” in Proc. of KDD , 2016, pp. 1675–1684

  145. [147]

    Sim- ple rules for complex decisions,

    J. Jung, C. Concannon, R. Shroff, S. Goel, and D. G. Goldstein, “Sim- ple rules for complex decisions,” arXiv preprint arXiv:1702.04690 , 2017

  146. [148]

    Interpretability of fuzzy systems: Current research trends and prospects,

    J. M. Alonso, C. Castiello, and C. Mencar, “Interpretability of fuzzy systems: Current research trends and prospects,” Springer handbook of computational intelligence, pp. 219–237, 2015

  147. [149]

    Accurate intelligible models with pairwise interactions,

    Y . Lou, R. Caruana, J. Gehrke, and G. Hooker, “Accurate intelligible models with pairwise interactions,” in Proc. of KDD , 2013, pp. 623– 631

  148. [150]

    Deep inside convolu- tional networks: Visualising image classification models and saliency maps,

    K. Simonyan, A. Vedaldi, and A. Zisserman, “Deep inside convolu- tional networks: Visualising image classification models and saliency maps,” arXiv preprint arXiv:1312.6034 , 2013

  149. [151]

    Synthesizing the preferred inputs for neurons in neural networks via deep generator networks,

    A. Nguyen, A. Dosovitskiy, J. Yosinski, T. Brox, and J. Clune, “Synthesizing the preferred inputs for neurons in neural networks via deep generator networks,” Proc. of NeurIPS , 2016

  150. [152]

    Adaptive deconvolutional networks for mid and high level feature learning,

    M. D. Zeiler, G. W. Taylor, and R. Fergus, “Adaptive deconvolutional networks for mid and high level feature learning,” in Proc. of ICCV , 2011, pp. 2018–2025

  151. [153]

    Visualizing and understanding convolu- tional networks,

    M. D. Zeiler and R. Fergus, “Visualizing and understanding convolu- tional networks,” in Proc. of ECCV , 2014, pp. 818–833

  152. [154]

    Understanding deep image represen- tations by inverting them,

    A. Mahendran and A. Vedaldi, “Understanding deep image represen- tations by inverting them,” in Proc. of CVPR , 2015, pp. 5188–5196

  153. [155]

    Under- standing neural networks through deep visualization,

    J. Yosinski, J. Clune, A. Nguyen, T. Fuchs, and H. Lipson, “Under- standing neural networks through deep visualization,” arXiv preprint arXiv:1506.06579, 2015

  154. [156]

    Multifaceted feature visualiza- tion: Uncovering the different types of features learned by each neuron in deep neural networks,

    A. Nguyen, J. Yosinski, and J. Clune, “Multifaceted feature visualiza- tion: Uncovering the different types of features learned by each neuron in deep neural networks,” arXiv preprint arXiv:1602.03616 , 2016

  155. [157]

    Convergent learning: Do different neural networks learn the same representations?

    Y . Li, J. Yosinski, J. Clune, H. Lipson, and J. Hopcroft, “Convergent learning: Do different neural networks learn the same representations?” arXiv preprint arXiv:1511.07543 , 2015

  156. [161]

    Hierarchical question-image co-attention for visual question answering,

    J. Lu, J. Yang, D. Batra, and D. Parikh, “Hierarchical question-image co-attention for visual question answering,” Proc. of NeurIPS , 2016

  157. [163]

    Infogan: Interpretable representation learning by informa- tion maximizing generative adversarial nets,

    X. Chen, Y . Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel, “Infogan: Interpretable representation learning by informa- tion maximizing generative adversarial nets,” Proc. of NeurIPS , 2016

  158. [164]

    Vqa: Visual question answering,

    S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” in Proc. of ICCV, 2015, pp. 2425–2433

  159. [165]

    Generating visual explanations,

    L. A. Hendricks, Z. Akata, M. Rohrbach, J. Donahue, B. Schiele, and T. Darrell, “Generating visual explanations,” in Proc. of ECCV , 2016, pp. 3–19

  160. [166]

    Multimodal compact bilinear pooling for vi- sual question answering and visual grounding,

    A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach, “Multimodal compact bilinear pooling for vi- sual question answering and visual grounding,” arXiv preprint arXiv:1606.01847, 2016

  161. [168]

    Interpretable deep models for icu outcome prediction,

    Z. Che, S. Purushotham, R. Khemani, and Y . Liu, “Interpretable deep models for icu outcome prediction,” in AMIA annual symposium proceedings, 2016, p. 371

  162. [169]

    Treeview: Peeking into deep neural networks via feature-space parti- tioning,

    J. J. Thiagarajan, B. Kailkhura, P. Sattigeri, and K. N. Ramamurthy, “Treeview: Peeking into deep neural networks via feature-space parti- tioning,” arXiv preprint arXiv:1611.07429 , 2016

  163. [171]

    Not just a black box: Learning important features through propagating activation differences,

    A. Shrikumar, P. Greenside, A. Shcherbina, and A. Kundaje, “Not just a black box: Learning important features through propagating activation differences,” arXiv preprint arXiv:1605.01713 , 2016

  164. [174]

    On pixel-wise explanations for non-linear classifier de- cisions by layer-wise relevance propagation,

    S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. M ¨uller, and W. Samek, “On pixel-wise explanations for non-linear classifier de- cisions by layer-wise relevance propagation,” PloS one , p. e0130140, 2015

  165. [175]

    ” why should i trust you?

    M. T. Ribeiro, S. Singh, and C. Guestrin, “” why should i trust you?” explaining the predictions of any classifier,” in Proc. of KDD , 2016, pp. 1135–1144

  166. [176]

    Learning deep features for discriminative localization,

    B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in Proc. of CVPR, 2016, pp. 2921–2929

  167. [179]

    Visualizing and understanding recurrent networks,

    A. Karpathy, J. Johnson, and L. Fei-Fei, “Visualizing and understanding recurrent networks,” arXiv preprint arXiv:1506.02078 , 2015

  168. [180]

    Interpretable recurrent neural networks using sequential sparse recovery,

    S. Wisdom, T. Powers, J. Pitton, and L. Atlas, “Interpretable recurrent neural networks using sequential sparse recovery,” arXiv preprint arXiv:1611.07252, 2016

  169. [181]

    Increasing the interpretability of recurrent neural networks using hidden markov models,

    V . Krakovna and F. Doshi-Velez, “Increasing the interpretability of recurrent neural networks using hidden markov models,” arXiv preprint arXiv:1606.05320, 2016

  170. [182]

    Explainable artificial intelligence (xai): Concepts, taxonomies, oppor- tunities and challenges toward responsible ai,

    A. B. Arrieta, N. D ´ıaz-Rodr´ıguez, J. Del Ser, A. Bennetot, S. Tabik, A. Barbado, S. Garc ´ıa, S. Gil-L ´opez, D. Molina, and R. Benjamins, “Explainable artificial intelligence (xai): Concepts, taxonomies, oppor- tunities and challenges toward responsible ai,” Information fu...

  171. [183]

    Interweaving multimodal in- teraction with flexible unit visualizations for data exploration,

    A. Srinivasan, B. Lee, and J. Stasko, “Interweaving multimodal in- teraction with flexible unit visualizations for data exploration,” IEEE Transactions on Visualization and Computer Graphics, pp. 3519–3533, 2020

  172. [184]

    Multimodal data to design visual learning analytics for understanding regulation of learning,

    O. Noroozi, I. Alikhani, S. J ¨arvel¨a, P. A. Kirschner, I. Juuso, and T. Sepp¨anen, “Multimodal data to design visual learning analytics for understanding regulation of learning,” Computers in Human Behavior , pp. 298–304, 2019

  173. [185]

    An analysis of visual question answering algorithms,

    K. Kafle and C. Kanan, “An analysis of visual question answering algorithms,” in Proc. of ICCV , 2017, pp. 1965–1973

  174. [186]

    Toward explainable affective computing: A review,

    K. Corti ˜nas-Lorenzo and G. Lacey, “Toward explainable affective computing: A review,” IEEE Transactions on Neural Networks and Learning Systems, 2023

  175. [187]

    Emoco: Visual analysis of emotion coherence in presentation videos,

    H. Zeng, X. Wang, A. Wu, Y . Wang, Q. Li, A. Endert, and H. Qu, “Emoco: Visual analysis of emotion coherence in presentation videos,” IEEE transactions on visualization and computer graphics , pp. 927– 937, 2019

  176. [188]

    Emotioncues: Emotion-oriented visual summarization of classroom videos,

    H. Zeng, X. Shu, Y . Wang, Y . Wang, L. Zhang, T.-C. Pong, and H. Qu, “Emotioncues: Emotion-oriented visual summarization of classroom videos,” IEEE transactions on visualization and computer graphics , pp. 3168–3181, 2020. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 17

  177. [189]

    Dehu- mor: Visual analytics for decomposing humor,

    X. Wang, Y . Ming, T. Wu, H. Zeng, Y . Wang, and H. Qu, “Dehu- mor: Visual analytics for decomposing humor,” IEEE Transactions on Visualization and Computer Graphics , pp. 4609–4623, 2021

  178. [190]

    Towards a better gold stan- dard: Denoising and modelling continuous emotion annotations based on feature agglomeration and outlier regularisation,

    C. Wang, P. Lopes, T. Pun, and G. Chanel, “Towards a better gold stan- dard: Denoising and modelling continuous emotion annotations based on feature agglomeration and outlier regularisation,” in Proceedings of the 2018 on Audio/visual Emotion Challenge and Workshop , 2018, pp. 73–81

  179. [191]

    Modeling, recognizing, and explaining apparent personality from videos,

    H. J. Escalante, H. Kaya, A. A. Salah, S. Escalera, Y . G ¨uc ¸l¨ut¨urk, U. G ¨uc ¸l¨u, X. Bar ´o, I. Guyon, J. C. J. Junior, M. Madadi et al. , “Modeling, recognizing, and explaining apparent personality from videos,” IEEE Transactions on Affective Computing , pp. 894–911, 2020

  180. [192]

    Co-clustering to reveal salient facial features for expression recognition,

    S. Khan, L. Chen, and H. Yan, “Co-clustering to reveal salient facial features for expression recognition,” IEEE Transactions on Affective Computing, pp. 348–360, 2017

  181. [193]

    Using a pca-based dataset similarity measure to improve cross-corpus emotion recognition,

    I. Siegert, R. B ¨ock, and A. Wendemuth, “Using a pca-based dataset similarity measure to improve cross-corpus emotion recognition,” Com- puter Speech & Language , pp. 1–23, 2018

  182. [194]

    A visual-physiology multimodal system for detecting outlier behavior of participants in a reality tv show,

    S. Kang, D. Kim, and Y . Kim, “A visual-physiology multimodal system for detecting outlier behavior of participants in a reality tv show,” International Journal of Distributed Sensor Networks , p. 1550147719864886, 2019

  183. [195]

    Out of the box: Reasoning with graph convolution nets for factual visual question answering,

    M. Narasimhan, S. Lazebnik, and A. Schwing, “Out of the box: Reasoning with graph convolution nets for factual visual question answering,” Proc. of NeurIPS , 2018

  184. [196]

    M3gat: A multi-modal, multi-task interactive graph attention network for conversational sentiment analysis and emotion recognition,

    Y . Zhang, A. Jia, B. Wang, P. Zhang, D. Zhao, P. Li, Y . Hou, X. Jin, D. Song, and J. Qin, “M3gat: A multi-modal, multi-task interactive graph attention network for conversational sentiment analysis and emotion recognition,” ACM Transactions on Information Systems , pp. 1–32, 2023

  185. [197]

    Explain- able video action reasoning via prior knowledge and state transitions,

    T. Zhuo, Z. Cheng, P. Zhang, Y . Wong, and M. Kankanhalli, “Explain- able video action reasoning via prior knowledge and state transitions,” in Proc. of ACM MM , 2019, pp. 521–529

  186. [198]

    Modeling high-order relationships: Brain-inspired hypergraph-induced multimodal-multitask framework for semantic comprehension,

    X. Sun, F. Yao, and C. Ding, “Modeling high-order relationships: Brain-inspired hypergraph-induced multimodal-multitask framework for semantic comprehension,” IEEE Transactions on Neural Networks and Learning Systems , 2023

  187. [199]

    Enhancing ex- planations in recommender systems with knowledge graphs,

    V . Lully, P. Laublet, M. Stankovic, and F. Radulovic, “Enhancing ex- planations in recommender systems with knowledge graphs,” Procedia Computer Science, pp. 211–222, 2018

  188. [200]

    Behind the scene: Revealing the secrets of pre-trained vision-and-language models,

    J. Cao, Z. Gan, Y . Cheng, L. Yu, Y .-C. Chen, and J. Liu, “Behind the scene: Revealing the secrets of pre-trained vision-and-language models,” in Proc. of ECCV , 2020, pp. 565–580

  189. [201]

    Probing image-language trans- formers for verb understanding,

    L. A. Hendricks and A. Nematzadeh, “Probing image-language trans- formers for verb understanding,” arXiv preprint arXiv:2106.09141 , 2021

  190. [202]

    Uniter: Universal image-text representation learning,

    Y .-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y . Cheng, and J. Liu, “Uniter: Universal image-text representation learning,” in Proc. of ECCV , 2020, pp. 104–120

  191. [203]

    Vision-and-language or vision- for-language? on cross-modal influence in multimodal transformers,

    S. Frank, E. Bugliarello, and D. Elliott, “Vision-and-language or vision- for-language? on cross-modal influence in multimodal transformers,” arXiv preprint arXiv:2109.04448 , 2021

  192. [209]

    Unveiling hierarchical relationships for social image representation learning,

    L. Han, X. Zhang, L. Zhang, M. Lu, F. Huang, and Y . Liu, “Unveiling hierarchical relationships for social image representation learning,” Applied Soft Computing , p. 110792, 2023

  193. [210]

    Multi- modal analogical reasoning over knowledge graphs,

    N. Zhang, L. Li, X. Chen, X. Liang, S. Deng, and H. Chen, “Multi- modal analogical reasoning over knowledge graphs,” in Proc. of ICLR, 2022

  194. [211]

    Dkn: Deep knowledge- aware network for news recommendation,

    H. Wang, F. Zhang, X. Xie, and M. Guo, “Dkn: Deep knowledge- aware network for news recommendation,” in Proc. of WWW , 2018, pp. 1835–1844

  195. [212]

    Learning heterogeneous knowledge base embeddings for explainable recommendation,

    Q. Ai, V . Azizi, X. Chen, and Y . Zhang, “Learning heterogeneous knowledge base embeddings for explainable recommendation,” Algo- rithms, p. 137, 2018

  196. [213]

    Knowledge-aware autoencoders for explainable recommender systems,

    V . Bellini, A. Schiavone, T. Di Noia, A. Ragone, and E. Di Sci- ascio, “Knowledge-aware autoencoders for explainable recommender systems,” in Proceedings of the 3rd workshop on deep learning for recommender systems, 2018, pp. 24–31

  197. [214]

    Cross-modal causal relational reasoning for event-level visual question answering,

    Y . Liu, G. Li, and L. Lin, “Cross-modal causal relational reasoning for event-level visual question answering,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2023

  198. [215]

    Towards causal vqa: Revealing and reducing spurious correlations by invariant and covariant semantic editing,

    V . Agarwal, R. Shetty, and M. Fritz, “Towards causal vqa: Revealing and reducing spurious correlations by invariant and covariant semantic editing,” in Proc. of CVPR , 2020, pp. 9690–9698

  199. [216]

    Counterfactual vqa: A cause-effect look at language bias,

    Y . Niu, K. Tang, H. Zhang, Z. Lu, X.-S. Hua, and J.-R. Wen, “Counterfactual vqa: A cause-effect look at language bias,” in Proc. of CVPR, 2021, pp. 12 700–12 710

  200. [217]

    Rubi: Reducing unimodal biases for visual question answering,

    R. Cadene, C. Dancette, M. Cord, D. Parikh et al. , “Rubi: Reducing unimodal biases for visual question answering,” Proc. of NeurIPS , 2019

  201. [218]

    Interpretable visual question answering by reasoning on dependency trees,

    Q. Cao, X. Liang, B. Li, and L. Lin, “Interpretable visual question answering by reasoning on dependency trees,” IEEE transactions on pattern analysis and machine intelligence , pp. 887–901, 2019

  202. [219]

    Lcm-captioner: A lightweight text-based image captioning method with collaborative mechanism between vision and text,

    Q. Wang, H. Deng, X. Wu, Z. Yang, Y . Liu, Y . Wang, and G. Hao, “Lcm-captioner: A lightweight text-based image captioning method with collaborative mechanism between vision and text,” Neural Net- works, pp. 318–329, 2023

  203. [220]

    Multimodal research in vision and language: A review of current and emerging trends,

    S. Uppal, S. Bhagat, D. Hazarika, N. Majumder, S. Poria, R. Zimmer- mann, and A. Zadeh, “Multimodal research in vision and language: A review of current and emerging trends,” Information Fusion , pp. 149–171, 2022

  204. [221]

    Counter- factual samples synthesizing for robust visual question answering,

    L. Chen, X. Yan, J. Xiao, H. Zhang, S. Pu, and Y . Zhuang, “Counter- factual samples synthesizing for robust visual question answering,” in Proc. of CVPR , 2020, pp. 10 800–10 809

  205. [222]

    Question-conditioned counterfactual image generation for vqa,

    J. Pan, Y . Goyal, and S. Lee, “Question-conditioned counterfactual image generation for vqa,” arXiv preprint arXiv:1911.06352 , 2019

  206. [223]

    Multimodal explanations by predicting counterfactuality in videos,

    A. Kanehira, K. Takemoto, S. Inayoshi, and T. Harada, “Multimodal explanations by predicting counterfactuality in videos,” in Proc. of CVPR, 2019, pp. 8594–8602

  207. [224]

    Generating counterfactual explanations with natural language,

    L. A. Hendricks, R. Hu, T. Darrell, and Z. Akata, “Generating counterfactual explanations with natural language,” arXiv preprint arXiv:1806.09809, 2018

  208. [225]

    Non-autoregressive image captioning with counterfactuals-critical multi-agent learning,

    L. Guo, J. Liu, X. Zhu, X. He, J. Jiang, and H. Lu, “Non-autoregressive image captioning with counterfactuals-critical multi-agent learning,” arXiv preprint arXiv:2005.04690 , 2020

  209. [226]

    Counter- factual critic multi-agent training for scene graph generation,

    L. Chen, H. Zhang, J. Xiao, X. He, S. Pu, and S.-F. Chang, “Counter- factual critic multi-agent training for scene graph generation,” in Proc. of ICCV, 2019, pp. 4613–4623

  210. [227]

    Bias in multimodal ai: Testbed for fair automatic recruitment,

    A. Pena, I. Serna, A. Morales, and J. Fierrez, “Bias in multimodal ai: Testbed for fair automatic recruitment,” in Proc. of CVPR , 2020, pp. 28–29

  211. [228]

    Explicit bias discovery in visual question answering models,

    V . Manjunatha, N. Saini, and L. S. Davis, “Explicit bias discovery in visual question answering models,” in Proc. of CVPR , 2019, pp. 9562–9571

  212. [229]

    Women also snowboard: Overcoming bias in captioning models,

    L. A. Hendricks, K. Burns, K. Saenko, T. Darrell, and A. Rohrbach, “Women also snowboard: Overcoming bias in captioning models,” in Proc. of ECCV , 2018, pp. 771–787

  213. [230]

    Overcoming language priors in visual question answering with adversarial regularization,

    S. Ramakrishnan, A. Agrawal, and S. Lee, “Overcoming language priors in visual question answering with adversarial regularization,” Proc. of NeurIPS , 2018

  214. [231]

    Mul- timodal fusion interactions: A study of human and automatic quan- tification,

    P. P. Liang, Y . Cheng, R. Salakhutdinov, and L.-P. Morency, “Mul- timodal fusion interactions: A study of human and automatic quan- tification,” in Proceedings of the 25th International Conference on Multimodal Interaction, 2023, pp. 425–435

  215. [232]

    Multiviz: Towards visualizing and understanding multimodal models,

    P. P. Liang, Y . Lyu, G. Chhablani, N. Jain, Z. Deng, X. Wang, L.-P. Morency, and R. Salakhutdinov, “Multiviz: Towards visualizing and understanding multimodal models,” arXiv preprint arXiv:2207.00056 , 2022

  216. [233]

    Vl-interpret: An interactive visualization tool for interpreting vision- language transformers,

    E. Aflalo, M. Du, S.-Y . Tseng, Y . Liu, C. Wu, N. Duan, and V . Lal, “Vl-interpret: An interactive visualization tool for interpreting vision- language transformers,” in Proc. of CVPR , 2022, pp. 21 406–21 415

  217. [234]

    Dime: Fine-grained interpretations of multimodal models via disen- tangled local explanations,

    Y . Lyu, P. P. Liang, Z. Deng, R. Salakhutdinov, and L.-P. Morency, “Dime: Fine-grained interpretations of multimodal models via disen- tangled local explanations,” in Proc. of AAAI , 2022, pp. 455–467. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 18

  218. [235]

    Towards better explanations of class activation mapping,

    H. Jung and Y . Oh, “Towards better explanations of class activation mapping,” in Proc. of ICCV , 2021, pp. 1336–1344

  219. [236]

    Did the model understand the question?

    P. K. Mudrakarta, A. Taly, M. Sundararajan, and K. Dhamdhere, “Did the model understand the question?” arXiv preprint arXiv:1805.05492, 2018

  220. [237]

    Towards the interpretability of deep learning models for multi-modal neuroimag- ing: Finding structural changes of the ageing brain,

    S. M. Hofmann, F. Beyer, S. Lapuschkin, O. Goltermann, M. Loeffler, K.-R. M ¨uller, A. Villringer, W. Samek, and A. V . Witte, “Towards the interpretability of deep learning models for multi-modal neuroimag- ing: Finding structural changes of the ageing brain,” NeuroImage, p. ...

  221. [238]

    Interpretability for multimodal emotion recognition using concept activation vectors,

    A. R. Asokan, N. Kumar, A. V . Ragam, and S. Shylaja, “Interpretability for multimodal emotion recognition using concept activation vectors,” in Proc. of IJCNN , 2022, pp. 01–08

  222. [239]

    Rise: Randomized input sampling for explanation of black-box models,

    V . Petsiuk, A. Das, and K. Saenko, “Rise: Randomized input sampling for explanation of black-box models,” arXiv preprint arXiv:1806.07421, 2018

  223. [240]

    Gpt-4 technical report,

    OpenAI, “Gpt-4 technical report,” Tech. Rep., 2023

  224. [241]

    Lida: A tool for automatic generation of grammar-agnostic visualizations and infographics using large language models,

    V . Dibia, “Lida: A tool for automatic generation of grammar-agnostic visualizations and infographics using large language models,” arXiv preprint arXiv:2303.02927, 2023

  225. [242]

    Can foundation models wrangle your data?

    A. Narayan, I. Chami, L. Orr, S. Arora, and C. R ´e, “Can foundation models wrangle your data?” arXiv preprint arXiv:2205.09911 , 2022

  226. [243]

    Benchmarking large lan- guage models as ai research agents,

    Q. Huang, J. V ora, P. Liang, and J. Leskovec, “Benchmarking large lan- guage models as ai research agents,” arXiv preprint arXiv:2310.03302, 2023

  227. [244]

    Table-gpt: Table fine-tuned gpt for diverse table tasks,

    P. Li, Y . He, D. Yashar, W. Cui, S. Ge, H. Zhang, D. Rifinski Fainman, D. Zhang, and S. Chaudhuri, “Table-gpt: Table fine-tuned gpt for diverse table tasks,” Proceedings of the ACM on Management of Data , pp. 1–28, 2024

  228. [245]

    Generative table pre-training empowers models for tabular prediction,

    T. Zhang, S. Wang, S. Yan, J. Li, and Q. Liu, “Generative table pre-training empowers models for tabular prediction,” arXiv preprint arXiv:2305.09696, 2023

  229. [246]

    Data science with llms and interpretable models,

    S. Bordt, B. Lengerich, H. Nori, and R. Caruana, “Data science with llms and interpretable models,” arXiv preprint arXiv:2402.14474, 2024

  230. [247]

    Multimodal c4: An open, billion-scale corpus of images interleaved with text,

    W. Zhu, J. Hessel, A. Awadalla, S. Y . Gadre, J. Dodge, A. Fang, Y . Yu, L. Schmidt, W. Y . Wang, and Y . Choi, “Multimodal c4: An open, billion-scale corpus of images interleaved with text,” Proc. of NeurIPS, 2024

  231. [248]

    Genie: Generative interactive environments,

    J. Bruce, M. D. Dennis, A. Edwards, J. Parker-Holder, Y . Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps et al. , “Genie: Generative interactive environments,” in Proc. of ICML, 2024

  232. [249]

    Quality not quantity: On the interaction between dataset design and robustness of clip,

    T. Nguyen, G. Ilharco, M. Wortsman, S. Oh, and L. Schmidt, “Quality not quantity: On the interaction between dataset design and robustness of clip,” Proc. of NeurIPS , pp. 21 455–21 469, 2022

  233. [250]

    Robust learning with progressive data expansion against spurious correlation,

    Y . Deng, Y . Yang, B. Mirzasoleiman, and Q. Gu, “Robust learning with progressive data expansion against spurious correlation,” Proc. of NeurIPS, 2024

  234. [251]

    Rendering graphs for graph reasoning in multimodal large language models,

    Y . Wei, S. Fu, W. Jiang, J. T. Kwok, and Y . Zhang, “Rendering graphs for graph reasoning in multimodal large language models,” arXiv preprint arXiv:2402.02130 , 2024

  235. [252]

    Llm and gnn are complementary: Distilling llm for multimodal graph learning,

    J. Xu, Z. Wu, M. Lin, X. Zhang, and S. Wang, “Llm and gnn are complementary: Distilling llm for multimodal graph learning,” arXiv preprint arXiv:2406.01032, 2024

  236. [253]

    How mul- timodal integration boost the performance of llm for optimization: Case study on capacitated vehicle routing problems,

    Y . Huang, W. Zhang, L. Feng, X. Wu, and K. C. Tan, “How mul- timodal integration boost the performance of llm for optimization: Case study on capacitated vehicle routing problems,” arXiv preprint arXiv:2403.01757, 2024

  237. [254]

    Generative multimodal models are in-context learners,

    Q. Sun, Y . Cui, X. Zhang, F. Zhang, Q. Yu, Y . Wang, Y . Rao, J. Liu, T. Huang, and X. Wang, “Generative multimodal models are in-context learners,” in Proc. of CVPR , 2024, pp. 14 398–14 409

  238. [255]

    What makes multimodal in-context learning work?

    F. B. Baldassini, M. Shukor, M. Cord, L. Soulier, and B. Piwowarski, “What makes multimodal in-context learning work?” in Proc. of CVPR, 2024, pp. 1539–1550

  239. [256]

    Obelics: An open web-scale filtered dataset of interleaved image-text documents,

    H. Laurenc ¸on, L. Saulnier, L. Tronchon, S. Bekman, A. Singh, A. Lozhkov, T. Wang, S. Karamcheti, A. Rush, D. Kiela et al. , “Obelics: An open web-scale filtered dataset of interleaved image-text documents,” Proc. of NeurIPS , 2024

  240. [257]

    Openflamingo: An open-source framework for training large autoregressive vision- language models,

    A. Awadalla, I. Gao, J. Gardner, J. Hessel, Y . Hanafy, W. Zhu, K. Marathe, Y . Bitton, S. Gadre, S. Sagawa et al. , “Openflamingo: An open-source framework for training large autoregressive vision- language models,” arXiv preprint arXiv:2308.01390 , 2023

  241. [258]

    Mm-narrator: Narrating long-form videos with multimodal in-context learning,

    C. Zhang, K. Lin, Z. Yang, J. Wang, L. Li, C.-C. Lin, Z. Liu, and L. Wang, “Mm-narrator: Narrating long-form videos with multimodal in-context learning,” in Proc. of CVPR , 2024, pp. 13 647–13 657

  242. [259]

    Beyond task performance: Evaluating and reducing the flaws of large multimodal models with in-context learning,

    M. Shukor, A. Rame, C. Dancette, and M. Cord, “Beyond task performance: Evaluating and reducing the flaws of large multimodal models with in-context learning,” arXiv preprint arXiv:2310.00647 , 2023

  243. [260]

    How does the textual information affect the retrieval of multimodal in-context learning?

    Y . Luo, Z. Zheng, Z. Zhu, and Y . You, “How does the textual information affect the retrieval of multimodal in-context learning?” arXiv preprint arXiv:2404.12866 , 2024

  244. [261]

    Chain of thought prompt tuning in vision language models,

    J. Ge, H. Luo, S. Qian, Y . Gan, J. Fu, and S. Zhang, “Chain of thought prompt tuning in vision language models,” arXiv preprint arXiv:2304.07919, 2023

  245. [262]

    Multimodal chain-of-thought reasoning in language models,

    Z. Zhang, A. Zhang, M. Li, H. Zhao, G. Karypis, and A. Smola, “Multimodal chain-of-thought reasoning in language models,” arXiv preprint arXiv:2302.00923, 2023

  246. [263]

    Visual chain of thought: bridging logical gaps with multimodal infillings,

    D. Rose, V . Himakunthala, A. Ouyang, R. He, A. Mei, Y . Lu, M. Saxon, C. Sonar, D. Mirza, and W. Y . Wang, “Visual chain of thought: bridging logical gaps with multimodal infillings,” arXiv preprint arXiv:2305.02317, 2023

  247. [264]

    Let’s think frame by frame with vip: A video infilling and prediction dataset for evaluating video chain-of- thought,

    V . Himakunthala, A. Ouyang, D. Rose, R. He, A. Mei, Y . Lu, C. Sonar, M. Saxon, and W. Y . Wang, “Let’s think frame by frame with vip: A video infilling and prediction dataset for evaluating video chain-of- thought,” arXiv preprint arXiv:2305.13903 , 2023

  248. [265]

    Contrastive novelty-augmented learning: Anticipating outliers with large language models,

    A. Xu, X. Ren, and R. Jia, “Contrastive novelty-augmented learning: Anticipating outliers with large language models,” arXiv preprint arXiv:2211.15718, 2022

  249. [266]

    Does synthetic data generation of llms help clinical text mining?

    R. Tang, X. Han, X. Jiang, and X. Hu, “Does synthetic data generation of llms help clinical text mining?” arXiv preprint arXiv:2303.04360 , 2023

  250. [267]

    Q: How to specialize large vision-language models to data-scarce vqa tasks? a: Self-train on unlabeled images!

    Z. Khan, V . K. BG, S. Schulter, X. Yu, Y . Fu, and M. Chandraker, “Q: How to specialize large vision-language models to data-scarce vqa tasks? a: Self-train on unlabeled images!” in Proc. of CVPR, 2023, pp. 15 005–15 015

  251. [268]

    T-sciq: Teaching multimodal chain-of-thought reasoning via large language model signals for science question answering,

    L. Wang, Y . Hu, J. He, X. Xu, N. Liu, H. Liu, and H. T. Shen, “T-sciq: Teaching multimodal chain-of-thought reasoning via large language model signals for science question answering,” in Proc. of AAAI, 2024, pp. 19 162–19 170

  252. [269]

    Llm-powered data augmentation for enhanced cross-lingual performance,

    C. Whitehouse, M. Choudhury, and A. F. Aji, “Llm-powered data augmentation for enhanced cross-lingual performance,” arXiv preprint arXiv:2305.14288, 2023

  253. [270]

    Generalized category discovery with large language models in the loop,

    W. An, W. Shi, F. Tian, H. Lin, Q. Wang, Y . Wu, M. Cai, L. Wang, Y . Chen, H. Zhu et al. , “Generalized category discovery with large language models in the loop,” arXiv preprint arXiv:2312.10897, 2023

  254. [271]

    Llava- next: Improved reasoning, ocr, and world knowledge,

    H. Liu, C. Li, Y . Li, B. Li, Y . Zhang, S. Shen, and Y . J. Lee, “Llava- next: Improved reasoning, ocr, and world knowledge,” 2024

  255. [272]

    Llava-med: Training a large language-and-vision assistant for biomedicine in one day,

    C. Li, C. Wong, S. Zhang, N. Usuyama, H. Liu, J. Yang, T. Naumann, H. Poon, and J. Gao, “Llava-med: Training a large language-and-vision assistant for biomedicine in one day,” Proc. of NeurIPS , 2024

  256. [273]

    Videochat: Chat-centric video understanding,

    K. Li, Y . He, Y . Wang, Y . Li, W. Wang, P. Luo, Y . Wang, L. Wang, and Y . Qiao, “Videochat: Chat-centric video understanding,”arXiv preprint arXiv:2305.06355, 2023

  257. [274]

    Dolphins: Multimodal language model for driving,

    Y . Ma, Y . Cao, J. Sun, M. Pavone, and C. Xiao, “Dolphins: Multimodal language model for driving,” arXiv preprint arXiv:2312.00438 , 2023

  258. [275]

    Salmonn: Towards generic hearing abilities for large language models,

    C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang, “Salmonn: Towards generic hearing abilities for large language models,” arXiv preprint arXiv:2310.13289 , 2023

  259. [276]

    Qwen-audio: Advancing universal audio understand- ing via unified large-scale audio-language models,

    Y . Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou, “Qwen-audio: Advancing universal audio understand- ing via unified large-scale audio-language models,” arXiv preprint arXiv:2311.07919, 2023

  260. [277]

    Flamingo: a visual language model for few-shot learning,

    J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” Proc. of NeurIPS , pp. 23 716–23 736, 2022

  261. [278]

    Mm-react: Prompting chatgpt for multimodal reasoning and action,

    Z. Yang, L. Li, J. Wang, K. Lin, E. Azarnasab, F. Ahmed, Z. Liu, C. Liu, M. Zeng, and L. Wang, “Mm-react: Prompting chatgpt for multimodal reasoning and action,” arXiv preprint arXiv:2303.11381 , 2023

  262. [279]

    X- llm: Bootstrapping advanced large language models by treating multi- modalities as foreign languages,

    F. Chen, M. Han, H. Zhao, Q. Zhang, J. Shi, S. Xu, and B. Xu, “X- llm: Bootstrapping advanced large language models by treating multi- modalities as foreign languages,” arXiv preprint arXiv:2305.04160 , 2023

  263. [280]

    X-instructblip: A framework for aligning x-modal instruction-aware representations to llms and emergent cross-modal reasoning,

    A. Panagopoulou, L. Xue, N. Yu, J. Li, D. Li, S. Joty, R. Xu, S. Savarese, C. Xiong, and J. C. Niebles, “X-instructblip: A framework for aligning x-modal instruction-aware representations to llms and emergent cross-modal reasoning,” arXiv preprint arXiv:2311.18799 , 2023

  264. [281]

    Anymal: An efficient and scalable any-modality augmented language model,

    S. Moon, A. Madotto, Z. Lin, T. Nagarajan, M. Smith, S. Jain, C.-F. Yeh, P. Murugesan, P. Heidari, Y . Liu et al. , “Anymal: An efficient and scalable any-modality augmented language model,” arXiv preprint arXiv:2309.16058, 2023. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, A...

  265. [282]

    Probing multimodal large language models for global and local semantic representation,

    M. Tao, Q. Huang, K. Xu, L. Chen, Y . Feng, and D. Zhao, “Probing multimodal large language models for global and local semantic representation,” arXiv preprint arXiv:2402.17304 , 2024

  266. [283]

    Mysterious projections: Multimodal llms gain domain- specific visual capabilities without richer cross-modal projections,

    G. Verma, M. Choi, K. Sharma, J. Watson-Daniels, S. Oh, and S. Kumar, “Mysterious projections: Multimodal llms gain domain- specific visual capabilities without richer cross-modal projections,” arXiv preprint arXiv:2402.16832 , 2024

  267. [284]

    What is the limitation of multimodal llms? a deeper look into multimodal llms through prompt probing,

    S. Qi, Z. Cao, J. Rao, L. Wang, J. Xiao, and X. Wang, “What is the limitation of multimodal llms? a deeper look into multimodal llms through prompt probing,” Information Processing & Management , p. 103510, 2023

  268. [285]

    Emergent world representations: Exploring a sequence model trained on a synthetic task,

    K. Li, A. K. Hopkins, D. Bau, F. Vi ´egas, H. Pfister, and M. Wattenberg, “Emergent world representations: Exploring a sequence model trained on a synthetic task,” arXiv preprint arXiv:2210.13382 , 2022

  269. [286]

    Probing multimodal llms as world models for driving,

    S. Sreeram, T.-H. Wang, A. Maalouf, G. Rosman, S. Karaman, and D. Rus, “Probing multimodal llms as world models for driving,” arXiv preprint arXiv:2405.05956, 2024

  270. [287]

    From text to pixel: Advancing long-context understanding in mllms,

    Y . Lu, X. Li, T.-J. Fu, M. Eckstein, and W. Y . Wang, “From text to pixel: Advancing long-context understanding in mllms,” arXiv preprint arXiv:2405.14213, 2024

  271. [288]

    Agla: Mitigating object hallucinations in large vision-language models with assembly of global and local attention,

    W. An, F. Tian, S. Leng, J. Nie, H. Lin, Q. Wang, G. Dai, P. Chen, and S. Lu, “Agla: Mitigating object hallucinations in large vision-language models with assembly of global and local attention,” arXiv preprint arXiv:2406.12718, 2024

  272. [289]

    Active reasoning in an open-world environment,

    M. Xu, G. Jiang, W. Liang, C. Zhang, and Y . Zhu, “Active reasoning in an open-world environment,” Proc. of NeurIPS , 2024

  273. [290]

    Reckoning: reasoning through dynamic knowledge encoding,

    Z. Chen, G. Weiss, E. Mitchell, A. Celikyilmaz, and A. Bosselut, “Reckoning: reasoning through dynamic knowledge encoding,” Proc. of NeurIPS, 2024

  274. [291]

    Statler: State-maintaining language models for embodied reasoning,

    T. Yoneda, J. Fang, P. Li, H. Zhang, T. Jiang, S. Lin, B. Picker, D. Yu- nis, H. Mei, and M. R. Walter, “Statler: State-maintaining language models for embodied reasoning,” arXiv preprint arXiv:2306.17840 , 2023

  275. [292]

    Inner monologue: Embodied reasoning through planning with language models,

    W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y . Chebotar et al. , “Inner monologue: Embodied reasoning through planning with language models,” arXiv preprint arXiv:2207.05608, 2022

  276. [293]

    Describe, explain, plan and select: interactive planning with llms enables open- world multi-task agents,

    Z. Wang, S. Cai, G. Chen, A. Liu, X. S. Ma, and Y . Liang, “Describe, explain, plan and select: interactive planning with llms enables open- world multi-task agents,” Proc. of NeurIPS , 2024

  277. [294]

    Knowledge acquisition disentanglement for knowledge-based visual question answering with large language mod- els,

    W. An, F. Tian, J. Nie, W. Shi, H. Lin, Y . Chen, Q. Wang, Y . Wu, G. Dai, and P. Chen, “Knowledge acquisition disentanglement for knowledge-based visual question answering with large language mod- els,” arXiv preprint arXiv:2407.15346 , 2024

  278. [295]

    What if...?: Counterfactual inception to mitigate hallucination effects in large multimodal models,

    J. Kim, Y . J. Kim, and Y . M. Ro, “What if...?: Counterfactual inception to mitigate hallucination effects in large multimodal models,” arXiv preprint arXiv:2403.13513, 2024

  279. [296]

    Eyes can deceive: Benchmarking counterfactual reasoning abilities of multi-modal large language models,

    Y . Li, W. Tian, Y . Jiao, J. Chen, and Y .-G. Jiang, “Eyes can deceive: Benchmarking counterfactual reasoning abilities of multi-modal large language models,” arXiv preprint arXiv:2404.12966 , 2024

  280. [297]

    Reasoning or reciting? exploring the capabil- ities and limitations of language models through counterfactual tasks,

    Z. Wu, L. Qiu, A. Ross, E. Aky ¨urek, B. Chen, B. Wang, N. Kim, J. Andreas, and Y . Kim, “Reasoning or reciting? exploring the capabil- ities and limitations of language models through counterfactual tasks,” arXiv preprint arXiv:2307.02477 , 2023

  281. [298]

    What if the tv was off? examining counterfactual reasoning abilities of multi- modal language models,

    L. Zhang, X. Zhai, Z. Zhao, Y . Zong, X. Wen, and B. Zhao, “What if the tv was off? examining counterfactual reasoning abilities of multi- modal language models,” in Proc. of CVPR , 2024, pp. 21 853–21 862

  282. [299]

    Using counterfactual tasks to evaluate the generality of analogical reasoning in large language models,

    M. Lewis and M. Mitchell, “Using counterfactual tasks to evaluate the generality of analogical reasoning in large language models,” arXiv preprint arXiv:2402.08955, 2024

  283. [300]

    Stop reasoning! when multimodal llms with chain-of-thought reasoning meets adversarial images,

    Z. Wang, Z. Han, S. Chen, F. Xue, Z. Ding, X. Xiao, V . Tresp, P. Torr, and J. Gu, “Stop reasoning! when multimodal llms with chain-of-thought reasoning meets adversarial images,” arXiv preprint arXiv:2402.14899, 2024

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.