REVIEW 10 cited by
Sanity Checks for Saliency Maps
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Saliency methods have emerged as a popular tool to highlight features in an input deemed relevant for the prediction of a learned model. Several saliency methods have been proposed, often guided by visual appeal on image data. In this work, we propose an actionable methodology to evaluate what kinds of explanations a given method can and cannot provide. We find that reliance, solely, on visual assessment can be misleading. Through extensive experiments we show that some existing saliency methods are independent both of the model and of the data generating process. Consequently, methods that fail the proposed tests are inadequate for tasks that are sensitive to either data or model, such as, finding outliers in the data, explaining the relationship between inputs and outputs that the model learned, and debugging the model. We interpret our findings through an analogy with edge detection in images, a technique that requires neither training data nor model. Theory in the case of a linear model and a single-layer convolutional neural network supports our experimental findings.
Forward citations
Cited by 10 Pith papers
-
ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders
With matched-scale sparse autoencoders, HuBERT-ECG best preserves its ECG representation while ECG-JEPA best exposes clinical measurements through single features — a leader split that repeats on MIMIC-IV-ECG.
-
Training Large Language Models for Self-Explanation Faithfulness
RL fine-tuning with a counterfactual mention/influence reward raises LLM self-explanation faithfulness (Phi-CCT) from near zero to ~0.66 in-distribution for two 8B models, with partial transfer to held-out tasks.
-
Contrastive learning of extragalactic stellar streams. Sculpting a latent space of representations with DES DR2 photometry
Applying NNCLR contrastive learning to DES DR2 galaxy cutouts yields embeddings that cluster major-merger galaxies but not stellar streams; a tiered sigmoid scaling redirects the network's saliency toward low-surface-...
-
On Spectral Properties of Gradient-based Explanation Methods
Gradient-based explanations behave like frequency-band selectors: the gradient acts as a high-pass filter, perturbation as a low-pass filter, and their combination creates explanations that shift with the perturbation scale.
-
Are machine learning interpretations reliable? A stability study on global interpretations
Popular machine learning interpretation methods are frequently unstable under small data perturbations, and interpretation stability does not track prediction accuracy.
-
Systematic Evaluation of Attribution Methods: Eliminating Threshold Bias and Revealing Method-Dependent Performance Patterns
Averaging IoU across many thresholds ranks XRAI above LIME and Integrated Gradients variants on HAM10000, and shows single-threshold attribution rankings are unstable.
-
BlueGlass: A Framework for Composite AI Safety
BlueGlass provides composite AI safety infrastructure; its case studies on object-detection VLMs reveal dataset trade-offs, a decoder-layer phase transition in probe accuracy, and SAE-discovered concepts including spu...
-
On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations
The paper introduces EF and ΔEF as spectral metrics, but ΔEF is derived from EF, making the complexity-faithfulness trade-off partly tautological.
-
AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models
A multi-stakeholder standards profile that tailors NIST AI RMF and ISO/IEC 23894 guidance to general-purpose AI and foundation model developers.
-
XAI-Guided Analysis of Residual Networks for Interpretable Pneumonia Detection in Paediatric Chest X-rays
A fine-tuned ResNet-50 with Grad-CAM and Monte Carlo dropout reports 95.94% accuracy and 98.91% AUC for pediatric pneumonia on the Kermany chest X-ray dataset.
Discussion (0). Continue with ORCID to comment.