{"id":"e6ce1399-8215-402d-89fb-30b0c17951be","arxiv_id":"2501.06253","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of post-hoc local XAI techniques for images, covering motivations, challenges, and suggested future directions.","lead":"This paper is a survey of explainable AI (XAI) for image processing, summarizing motivations, challenges, and open problems in making black-box models understandable. A reader new to XAI can use it as a starting point, though it offers no new experiments or results.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The §8.1 claim that no multi-modal XAI work has progressed directly contradicts the paper's own citations [14] and [59], invalidating a stated future direction and lowering confidence in the survey's reliability.","rationale":"The reader's weakest_assumption identifies a genuine internal contradiction in §8.1. The same paper that cites multi-modal explanation systems [14] and [59] in §6.4 later claims no work has progressed in that direction. This is exactly the sort of assumption a survey cannot afford to get wrong, because one of the paper's stated contributions is to propose future directions; a false novelty claim undermines that contribution and reduces confidence in the survey's reliability. I agree with the reader's conditional verdict: the paper should be accepted only if the §8.1 claim is corrected or carefully scoped. I considered whether the §6.1 lack-of-formalism claim itself is the most load-bearing weakness, but the paper supports it with citations and with its own §7.2–7.3 discussion of evaluation metrics and their disagreements; while the claim is hedged and plausible, it does not have a concrete contradiction of the same kind. The issue is a checkable factual inconsistency in the text, and the fix is straightforward.","tokens_in":20945,"tokens_out":6169,"duration_ms":58593,"concrete_test":"Open [14] (REX, arXiv:2203.06107) and [59] (Park et al., arXiv:1802.08129) and verify whether each generates multi-modal explanations (paired visual evidence and textual justification) before January 2025. If both do, then §8.1's 'we have yet to see any works that progressed in this direction' is factually false; the passage must be revised to a narrower, explicit claim (e.g., no work combining SHAP/LIME attribution maps with CLIP-based text) or removed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's stated contributions include 'future directions that XAI research space may take.' In §8.1, the authors propose Multi-Modal/Intra-Model XAI and write that 'at the current time of writing, we have yet to see any works that progressed in this direction.' This is contradicted by the paper's own §6.4, which states that 'Multi-modal explanations have also been explored in [14]' (REX, a reasoning-aware, grounded explanation framework pairing textual explanations with visual regions) and cites [59] (Park et al., 'Multimodal Explanations: Justifying Decisions and Pointing to the Evidence,' arXiv:1802.08129), a system that produces combined textual justification and visual pointing. Because these works predate the manuscript's January 2025 submission, the asserted novelty of §8.1 is false as written. The authors could mean a narrower claim, such as no prior combination of SHAP/LIME output with CLIP-style language models, but the passage does not say that. This is a load-bearing weakness in one of the paper's three listed contributions, not merely a stylistic slip: a survey that mischaracterizes prior art cannot support its future-directions conclusion. The §6.1 lack-of-formalism argument is separately plausible and is not directly damaged by this error, but the paper's overall verdict should require the §8.1 claim to be corrected or scoped.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a narrative survey of post-hoc local explainability for image-processing models. It begins by reviewing terminology (explainability, interpretability, trustworthiness, repeatability, reproducibility, stability, causality, cotenability, faithfulness), then summarizes motivations for XAI (regulatory, industrial, technical, societal) and four post-hoc local methods (ICE, counterfactual explanations, LIME, SHAP). The central body catalogues challenges (lack of formalism, interpretability of explanations, complexity–accuracy trade-offs, diversification of methods, causal explanations, limitations of current approaches, debugging) and open problems (the disagreement problem, evaluation metrics, disagreement among metrics). It closes with three future directions: multi-modal/intra-model XAI, enhanced evaluation metrics, and better understanding of ground truth and stakeholder requirements.","tokens_in":21201,"tokens_out":7164,"duration_ms":66557,"significance":"The paper's descriptive content is broadly consistent with the literature, and it offers a compact entry point to the XAI discourse. It is not an empirical contribution and does not provide a systematic or machine-checked methodology, but several of its organizing categories (motivations vs challenges, objective vs human-centric evaluation) are useful. The §6.1 point about the absence of a shared formal definition of explainability is plausible and connects to a genuine debate. However, the survey's reliability is weakened by at least one self-contradiction in its proposed future directions, and the recommendations will require substantive re-scoping rather than copy-editing alone.","major_comments":[{"comment":"The claim that multi-modal/intra-model XAI is an unoccupied direction is contradicted by the paper's own earlier discussion. Section 6.4 states that 'Multi-modal explanations have also been explored in [14]' and cites [59] (Park et al., Multimodal Explanations) as a system combining textual justification with visual pointing; [14] (REX) similarly pairs textual explanations with visual regions. Section 8.1 nevertheless asserts that 'at the current time of writing, we have yet to see any works that progressed in this direction.' Because proposing future directions is one of the three listed contributions (§1), this internal contradiction directly undermines the novelty claim. The passage should be re-scoped to a narrower gap (for example, combining SHAP/LIME-style attributions with vision-language models such as CLIP) or the proposal should be reframed as consolidation/extension rather than a new direction.","section":"§8.1; cf. §6.4"},{"comment":"The evaluation discussion is internally inconsistent about ground truth. Section 6.1 identifies 'the lack of ground truth explanations' as a major challenge, and §7.2 cites [10] for the absence of a foundational ground truth, yet §7.1 defines Feature Agreement, Rank Agreement, Sign Agreement, Signed Rank Agreement, Rank Correlation, and Pairwise Rank Agreement as comparisons against 'the corresponding ground truth explanation' without specifying how that ground truth is obtained. If these metrics are limited to synthetic datasets or manually constructed ground truths, that restriction must be stated; otherwise the reader cannot tell whether the open-problem section is describing a solvable benchmarking protocol or restating the foundational gap.","section":"§7.1 and §7.2"}],"minor_comments":[{"comment":"The reference list contains duplicate entries that should be merged: [1] and [2] are the same paper, [7] and [9] are the same paper in different forms, [22] and [23] are the same arXiv preprint, and [48] and [49] are the same book chapter.","section":"References"},{"comment":"There are numerous typographical and spelling errors that should be corrected in a copy-editing pass: 'Explicable AI' in the Section 3 heading, 'Socio-Techinical' in §2, 'repeatablity' in §3.4, 'Shapely Values' in §5.4, 'Mutli-Modal' in §8.1, 'intepretability' in §6.4, and 'empirial' in §7.3.","section":"Throughout"},{"comment":"Several figures are reproduced from external sources (e.g., Figures 1–5 from [30], [48], [44]) without source annotations in the captions; the manuscript should state permissions/reuse details, and Figures 7 and 8 should identify the benchmark and axes they illustrate.","section":"Figures"},{"comment":"Because the paper makes comparative frequency claims such as 'the majority of XAI works is mainly on image and textual data' (§6.6) and 'few works actually focused on a human-in-the-loop approach' (§8.2), a short methodology note describing the search process and inclusion criteria for the 83 cited papers would make those claims verifiable.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's contribution is essentially a teaching-oriented synthesis; its scope is broad and its overlap with existing surveys (e.g., [71]) is substantial. If novelty is a criterion for the venue, the paper will need a more explicit statement of what it adds beyond existing surveys. The internal contradiction in §8.1 is the main substantive reason I recommend major revision rather than rejection, because it is re-scopable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a competent but derivative survey; the one thing that matters is that Section 8.1's claim that no multi-modal XAI work exists is contradicted by Section 6.4's own discussion of REX [14] and Park et al. [59]. That makes one of the paper's three stated contributions—future directions—rest on a false premise as written.\n\nWhat the paper does well: it gives a readable map of motivations, open problems, and post-hoc local methods (ICE, counterfactuals, LIME, SHAP) for image processing, and the lack-of-formalism argument in §6.1 is a fair and well-cited point. The disagreement problem and evaluation-metric discussion in §7 are useful summaries. For a reader with no prior XAI exposure, this would not be a bad starting point.\n\nWhat is soft: the contradiction above is not a small slip, since the paper explicitly lists future directions as a contribution. The authors may mean something narrower—no prior combination of SHAP/LIME outputs with CLIP-style language models—but that is not what the passage says. There are also duplicate references ([1] and [2] are the same paper), a 'Socio-Techinical' typo, 'Shapely' for 'Shapley,' and some terminology that is used loosely. None of that is fatal to the survey's descriptive content, but it does lower confidence in the manuscript's reliability.\n\nI would not cite this in my own work and would not bring it to a reading group for research value. Its audience is newcomers, not practitioners or survey authors, and there are already several surveys covering the same ground more rigorously.\n\nRecommendation: as a research submission, desk reject. If the authors correct the §8.1 claim, remove duplicate references, and position it as an educational tutorial rather than a research contribution, it could be acceptable at a workshop or teaching-oriented venue. Not worth serious peer review in its current form.","headline":"A readable but derivative XAI survey whose main future-direction claim in §8.1 is directly contradicted by its own §6.4 discussion of multi-modal explanation works.","tokens_in":21679,"tokens_out":2950,"would_cite":false,"duration_ms":29417,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that explainable AI's central obstacle is the absence of a shared formal definition of explainability, which leaves no way to determine whether one explanation method is better than another.","keywords":["explainable AI","post-hoc explanation","local explanation methods","image processing","disagreement problem","explanation evaluation","trustworthiness","survey"],"falsifier":"A literature search for published systems that combine visual saliency explanations with natural-language explanations for image classifiers would test the paper's core novelty claim; the paper itself cites examples of such systems, so if those count as multi-modal XAI, the claim that 'we have yet to see any works' in this direction is refuted.","tokens_in":20767,"feed_emoji":"🔍","tokens_out":7713,"duration_ms":73270,"temperature":0.7,"pith_summary":"This is a survey of post-hoc local explanation techniques for image-processing models, meaning methods that explain a single prediction rather than the whole model. The paper sets out the reasons XAI matters, reviews representative techniques such as individual conditional expectation, counterfactual explanations, local surrogate models, and Shapley-value attribution, and then maps the obstacles that keep these methods from being trustworthy. Its central claim is that the field is blocked less by a shortage of techniques than by the lack of an agreed definition of explainability and of evaluation metrics that agree with each other. The paper also proposes future directions, including combining explanation modalities and involving human stakeholders and ground-truth considerations in evaluation.","feed_headline":"AI explanations can't be judged until 'explainable' is defined","feed_subtitle":"A survey maps why explanation methods and their metrics disagree, blocking trust in black-box AI.","key_machinery":"The analytical engine is a challenge taxonomy that ties four motivations for XAI, namely regulatory pressure, industrial adoption, technical advancement, and societal impact, to concrete obstacles such as missing formalism, interpretability of explanations, complexity-accuracy trade-offs, diversification of approaches, weak causal explanations, and limitations of current methods. The unresolved disagreement problem is the key mechanism: when multiple XAI algorithms or multiple evaluation metrics rank features differently for the same input, practitioners have no principled way to choose an explanation. The paper keeps returning to alignment-style metrics that measure agreement between explanations, and to the proposed combination of explanation modalities as the main lever for making outputs more human-understandable.","core_discovery":"The paper's central claim is that without a satisfactory definition of explainability and interpretability, it is not possible to determine whether new XAI approaches are better at explaining machine-learning models, and therefore no single de-facto XAI approach has emerged. It identifies two disagreement problems that follow from this: different explanation algorithms can rank the same features differently, and different evaluation metrics can disagree about which explanation is most faithful. On the image-processing side, the paper reviews local post-hoc methods and their practical limitations, including computational cost, sensitivity to feature correlation, and lack of native support for object-detection models. As a way forward, it proposes intra-model and multi-modal XAI, combining explanations of the same type or pairing visual saliency with natural-language output, along with human-in-the-loop evaluation and a better mapping of stakeholder requirements to ground truth.","pith_inferences":["Beyond the paper: if explainability cannot be formally defined, regulators cannot operationalize the 'right to explanation' without effectively choosing or mandating an evaluation metric, coupling AI regulation to XAI standardization.","Beyond the paper: a concrete experiment measuring multiple alignment and faithfulness metrics across several local explainers on a fixed set of images would quantify how often the disagreement problem occurs and how severe it is.","Beyond the paper: pairing saliency maps with vision-language captions could be tested directly in user studies that compare comprehension time and decision accuracy against visual-only explanations.","Beyond the paper: the disagreement problem is likely amplified for object-detection models, since the paper notes attribution methods need per-class tuning, and per-class explanations may disagree in ways that global metrics do not capture."],"forward_implications":["If no agreed definition of explainability exists, new XAI approaches cannot be compared against a baseline, so research progress cannot be measured at all.","If explanation algorithms disagree on the same input, practitioners must rely on alignment metrics and benchmark leaderboards to decide which explanation to trust.","If evaluation metrics disagree about faithfulness, current benchmarks do not give end users enough guidance to select an XAI method.","If multi-modal and intra-model XAI succeed, visual saliency could be paired with natural-language explanations for image classifiers, improving interpretability without requiring users to read heatmaps alone.","If human-in-the-loop evaluation and stakeholder requirement analysis are adopted, regulatory 'right to explanation' mandates would have a concrete implementation path."],"supporting_citations":[{"why":"Supplies the perturbation-based local surrogate explanation method whose stability issues motivate several of the paper's challenges.","marker":"[67]"},{"why":"Supplies the Shapley-value attribution framework whose feature-permutation basis creates correlation and computational limitations.","marker":"[45]"},{"why":"Defines the disagreement problem among XAI algorithms and motivates the alignment metrics discussed in the open problems.","marker":"[37]"},{"why":"Supplies alignment metrics and functional-testing benchmark tasks used to evaluate explanation quality.","marker":"[3]"},{"why":"Provides evidence that faithfulness evaluation metrics disagree, undermining the reliability of benchmarks.","marker":"[8]"},{"why":"Supplies the stakeholder-perspective framing of explainability that grounds the paper's multi-stakeholder arguments.","marker":"[39]"},{"why":"Supplies a multi-modal explanation system pairing textual explanations with image regions, relevant to the proposed future direction.","marker":"[14]"},{"why":"Supplies an earlier visual-and-textual multi-modal explanation approach, also relevant to the proposed future direction.","marker":"[59]"}],"fun_headline_variants":["XAI's missing definition blocks progress in image explanations","Why explainable AI has no consensus: definition problem","Explaining images: methods disagree, metrics disagree, progress stalls","Without a definition, XAI can't pick a winner for image tasks","Survey: XAI for images lacks definition, faces metric clashes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's proposed multi-modal future direction assumes that no published work yet combines multiple XAI output modalities, even though it cites multi-modal explanation systems earlier in the text, so that assumption is the most fragile part of the argument.","fun_headline_variants_meta":{"raw":{"variants":["XAI's missing definition blocks progress in image explanations","Why explainable AI has no consensus: definition problem","Explaining images: methods disagree, metrics disagree, progress stalls","Without a definition, XAI can't pick a winner for image tasks","Survey: XAI for images lacks definition, faces metric clashes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000168,"raw_usage":{"total_tokens":1224,"prompt_tokens":871,"completion_tokens":353,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":487,"completion_tokens_details":{"reasoning_tokens":269}},"tokens_in":487,"tokens_out":353,"duration_ms":3779,"temperature":1.0,"reasoning_tokens":269,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:22:45.467950+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A literature search for published systems that combine visual saliency explanations with natural-language explanations for image classifiers would test the paper's core novelty claim; the paper itself cites examples of such systems, so if those count as multi-modal XAI, the claim that 'we have yet to see any works' in this direction is refuted.","supporting_citations":[{"cited_title":"What Do We Want From Explainable Artificial Intelligence (XAI)? -- A Stakeholder Perspective on XAI and a Conceptual Model Guiding Interdisciplinary XAI Research","cited_arxiv_id":"2102.07817","evidence_quote":"Supplies the stakeholder-perspective framing of explainability that grounds the paper's multi-stakeholder arguments."},{"cited_title":"REX: Reasoning-aware and Grounded Explanation","cited_arxiv_id":"2203.06107","evidence_quote":"Supplies a multi-modal explanation system pairing textual explanations with image regions, relevant to the proposed future direction."}],"review_version":1}