REVIEW 6 cited by
Causal Explanations for Image Classifiers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Existing algorithms for explaining the output of image classifiers use different definitions of explanations and a variety of techniques to find them. However, none of the existing tools use a principled approach based on formal definitions of cause and explanation. In this paper we present a novel black-box approach to computing explanations grounded in the theory of actual causality. We prove relevant theoretical results and present an algorithm for computing approximate explanations based on these definitions. We prove termination of our algorithm and discuss its complexity and the amount of approximation compared to the precise definition. We implemented the framework in a tool ReX and we present experimental results and a comparison with state-of-the-art tools. We demonstrate that ReX is the most efficient black-box tool and produces the smallest explanations, in addition to outperforming other black-box tools on standard quality measures.
Forward citations
Cited by 6 Pith papers
-
What makes an Ensemble (Un) Interpretable?
A complexity-theoretic analysis showing that the number, size, and type of base models determine whether ensemble explanations are tractable, with linear-model ensembles intractable even for two models.
-
Explaining Failures of Cyber-Physical Systems with Actual Causality
Adapts actual causality to CPS failure explanation with two new algorithms, demonstrated on a neural autonomous car avoiding collisions.
-
Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI
ConvAD replaces input occlusion in post-hoc explanation with neuron deactivation in a CNN forward pass, yielding more robust causal explanations with no retraining.
-
Evaluation of Black-Box XAI Approaches for Predictors of Values of Boolean Formulae
A new black-box explanation method, B-ReX, matches causal-responsibility ground truth on Boolean formula classifiers more closely than existing XAI tools.
-
Explaining, Fast and Slow: Abstraction and Refinement of Provable Explanations
Abstraction-refinement over neuron merging computes provably sufficient and minimal explanations of neural network predictions substantially faster than verifying on the full network.
-
3D ReX: Causal Explanations in 3D Neuroimaging Classification
A causality-based explainability tool for 3D medical image classifiers, demonstrated on stroke detection, produces voxel-level responsibility maps without accessing the model's internals.
Discussion (0). Continue with ORCID to comment.