Pith. sign in

REVIEW 10 cited by

Sanity Checks for Saliency Maps

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1810.03292 v3 pith:KZ63HUC5 submitted 2018-10-08 cs.CV cs.LGstat.ML

classification cs.CVcs.LGstat.ML
keywords modeldatamethodssaliencyfindingslearnedproposedvisual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Saliency methods have emerged as a popular tool to highlight features in an input deemed relevant for the prediction of a learned model. Several saliency methods have been proposed, often guided by visual appeal on image data. In this work, we propose an actionable methodology to evaluate what kinds of explanations a given method can and cannot provide. We find that reliance, solely, on visual assessment can be misleading. Through extensive experiments we show that some existing saliency methods are independent both of the model and of the data generating process. Consequently, methods that fail the proposed tests are inadequate for tasks that are sensitive to either data or model, such as, finding outliers in the data, explaining the relationship between inputs and outputs that the model learned, and debugging the model. We interpret our findings through an analogy with edge detection in images, a technique that requires neither training data nor model. Theory in the case of a linear model and a single-layer convolutional neural network supports our experimental findings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 606 citations worldwide. Full citation record

  1. ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders

    cs.LG 2026-07 conditional novelty 6.0 of 10

    With matched-scale sparse autoencoders, HuBERT-ECG best preserves its ECG representation while ECG-JEPA best exposes clinical measurements through single features — a leader split that repeats on MIMIC-IV-ECG.

  2. Training Large Language Models for Self-Explanation Faithfulness

    cs.LG 2026-07 conditional novelty 6.0 of 10

    RL fine-tuning with a counterfactual mention/influence reward raises LLM self-explanation faithfulness (Phi-CCT) from near zero to ~0.66 in-distribution for two 8B models, with partial transfer to held-out tasks.

  3. Contrastive learning of extragalactic stellar streams. Sculpting a latent space of representations with DES DR2 photometry

    astro-ph.GA 2026-01 conditional novelty 6.0 of 10

    Applying NNCLR contrastive learning to DES DR2 galaxy cutouts yields embeddings that cluster major-merger galaxies but not stellar streams; a tiered sigmoid scaling redirects the network's saliency toward low-surface-...

  4. On Spectral Properties of Gradient-based Explanation Methods

    cs.LG 2025-08 conditional novelty 6.0 of 10

    Gradient-based explanations behave like frequency-band selectors: the gradient acts as a high-pass filter, perturbation as a low-pass filter, and their combination creates explanations that shift with the perturbation scale.

  5. Are machine learning interpretations reliable? A stability study on global interpretations

    stat.ML 2025-05 conditional novelty 6.0 of 10

    Popular machine learning interpretation methods are frequently unstable under small data perturbations, and interpretation stability does not track prediction accuracy.

  6. Systematic Evaluation of Attribution Methods: Eliminating Threshold Bias and Revealing Method-Dependent Performance Patterns

    cs.LG 2025-09 conditional novelty 5.0 of 10

    Averaging IoU across many thresholds ranks XRAI above LIME and Integrated Gradients variants on HAM10000, and shows single-threshold attribution rankings are unstable.

  7. BlueGlass: A Framework for Composite AI Safety

    cs.AI 2025-07 conditional novelty 5.0 of 10

    BlueGlass provides composite AI safety infrastructure; its case studies on object-detection VLMs reveal dataset trade-offs, a decoder-layer phase transition in probe accuracy, and SAE-discovered concepts including spu...

  8. On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations

    cs.LG 2025-08 reject novelty 4.0 of 10

    The paper introduces EF and ΔEF as spectral metrics, but ΔEF is derived from EF, making the complexity-faithfulness trade-off partly tautological.

  9. AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A multi-stakeholder standards profile that tailors NIST AI RMF and ISO/IEC 23894 guidance to general-purpose AI and foundation model developers.

  10. XAI-Guided Analysis of Residual Networks for Interpretable Pneumonia Detection in Paediatric Chest X-rays

    eess.IV 2025-07 conditional novelty 3.0 of 10

    A fine-tuned ResNet-50 with Grad-CAM and Monte Carlo dropout reports 95.94% accuracy and 98.91% AUC for pediatric pneumonia on the Kermany chest X-ray dataset.

Pith tools