Pith. sign in

REVIEW 6 cited by

Causal Explanations for Image Classifiers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.08875 v4 pith:FTUK3ET4 submitted 2024-11-13 cs.AI

classification cs.AI
keywords explanationsblack-boxdefinitionstoolsalgorithmapproachclassifierscomputing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing algorithms for explaining the output of image classifiers use different definitions of explanations and a variety of techniques to find them. However, none of the existing tools use a principled approach based on formal definitions of cause and explanation. In this paper we present a novel black-box approach to computing explanations grounded in the theory of actual causality. We prove relevant theoretical results and present an algorithm for computing approximate explanations based on these definitions. We prove termination of our algorithm and discuss its complexity and the amount of approximation compared to the precise definition. We implemented the framework in a tool ReX and we present experimental results and a comparison with state-of-the-art tools. We demonstrate that ReX is the most efficient black-box tool and produces the smallest explanations, in addition to outperforming other black-box tools on standard quality measures.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What makes an Ensemble (Un) Interpretable?

    cs.LG 2025-06 conditional novelty 7.0 of 10

    A complexity-theoretic analysis showing that the number, size, and type of base models determine whether ensemble explanations are tractable, with linear-model ensembles intractable even for two models.

  2. Explaining Failures of Cyber-Physical Systems with Actual Causality

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    Adapts actual causality to CPS failure explanation with two new algorithms, demonstrated on a neural autonomous car avoiding collisions.

  3. Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI

    cs.AI 2025-10 conditional novelty 6.0 of 10

    ConvAD replaces input occlusion in post-hoc explanation with neuron deactivation in a CNN forward pass, yielding more robust causal explanations with no retraining.

  4. Evaluation of Black-Box XAI Approaches for Predictors of Values of Boolean Formulae

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A new black-box explanation method, B-ReX, matches causal-responsibility ground truth on Boolean formula classifiers more closely than existing XAI tools.

  5. Explaining, Fast and Slow: Abstraction and Refinement of Provable Explanations

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Abstraction-refinement over neuron merging computes provably sufficient and minimal explanations of neural network predictions substantially faster than verifying on the full network.

  6. 3D ReX: Causal Explanations in 3D Neuroimaging Classification

    eess.IV 2025-02 conditional novelty 4.0 of 10

    A causality-based explainability tool for 3D medical image classifiers, demonstrated on stroke detection, produces voxel-level responsibility maps without accessing the model's internals.

Pith tools