Pith. sign in

REVIEW 1 cited by

Investigating the Effectiveness of Explainability Methods in Parkinson's Detection from Speech

T0 review · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A systematic benchmark shows that standard saliency methods faithfully reflect a Parkinson's speech classifier's decisions, yet the resulting spectrogram highlights are not readily usable by domain experts.

arxiv 2411.08013 v2 pith:JPGWXXMQ submitted 2024-11-12 cs.SD cs.AIcs.LGeess.AS

classification cs.SDcs.AIcs.LGeess.AS
keywords detectionmapsspeechclassifierdiagnosisexplainabilityinformationinterpretability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Doctors sometimes want to know why an AI system flags a voice recording as possibly Parkinson's. This paper looks at a model that listens to raw speech and decides if the speaker has Parkinson's, then tries six popular explanation tools that produce highlighted spectrograms showing which parts of the sound the model paid attention to. The tools include Integrated Gradients, SmoothGrad, and Guided GradCAM.

The authors measure whether these highlights actually matter to the model by erasing the highlighted parts and checking that the model's confidence drops, and by other standard tests. Most methods pass these tests, meaning the highlights do point at regions the model relies on. They also train a separate image-classifier on the highlight maps and find it can often distinguish Parkinson's from healthy speech slightly better than training on the original spectrograms. In that narrow sense, the maps carry information.

However, when the authors look at the maps themselves, they see no clear pattern that a speech therapist or neurologist could use, such as specific phonemes or articulation features. Two different explanation methods often highlight different parts of the same recording. Combining methods helps only when the maps already overlap a lot. The paper's main conclusion is therefore cautious: current explanation methods are faithful to the model but not yet meaningful to human experts. The authors suggest that future work should build explanations in terms of phonemes or listenable audio snippets instead of raw spectrogram pixels.

Extended reading notes

Core claim

The central load-bearing assertion appears in the Conclusion: "we show that popular post-hoc explanation methods can generate faithful explanations for PD detection. Nonetheless, they fail to generate explanations that domain experts can easily understand." If the paper is correct, the true statement is: standard post-hoc attribution techniques applied to a HuBERT-based PD detector produce masks that pass standard faithfulness tests, but these masks lack semantic interpretability for clinical experts.

Load-bearing premise

The negative half of the conclusion, that explanations "fail to provide valuable information for domain experts," rests on the authors' qualitative reading of a single illustrative spectrogram (Figure 1) and their general impression, not on any measurement involving clinicians. There is no expert study, no rating task, and no comparison against known PD speech biomarkers (e.g., vowel space metrics or DDK features). If a panel of experts could extract useful patterns from the maps, the paper's central negative claim would be unsupported. This assumption enters in Section V-D and the Conclusion.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The paper introduces no new physical entities and no numbers fitted to data. Its central evaluation is an empirical benchmark, so the burden lies in imported assumptions: the validity of faithfulness metrics in a low-confidence regime, the equivalence of the reconstructed waveform for gradient computation, and the representativeness of the inherited dataset splits. The only new scoring rule, the selective metric, uses a hand-chosen unit weighting that is transparent but arbitrary.

free parameters (1)
  • Selective metric weighting coefficients
    The selective metrics combine classification performance and mask mean with unit weights (SM = M * (1 - mean(A))). This equal weighting is a hand-chosen scoring rule; changing the relative weight could alter the ranking of methods. It is not fitted to data, but it is an arbitrary modeling choice that affects the central comparison.
assumptions (3)
  • domain assumption Faithfulness metrics (AI, AD, AG, FF, Fid-In) measured on the classifier's logits are valid proxies for the quality of explanations.
    The paper relies on these metrics from the L-MAC/L2I/PIQ literature without questioning their validity in the low-confidence regime it observes (small FF values noted in Section V-A).
  • domain assumption The STFT-to-ISTFT reconstruction with the original phase yields an input equivalent to the original waveform for the HuBERT model, so attributions computed in this pipeline are valid.
    Section III-A describes converting the spectrogram back to the time domain before feeding the model, but does not verify that the reconstructed waveform's predictions match the original waveform's predictions, nor does it specify how gradients are propagated through the STFT.
  • domain assumption The s-PC-GITA dataset and the 10-fold splits from [6] are representative and correctly labeled for PD vs HC.
    The paper inherits the dataset and splits from prior work without auditing the labels or recording conditions, which is standard practice but still an assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Investigating the Effectiveness of Explainability Methods in Parkinson's Detection from Speech." pith.science (2026). https://pith.science/paper/JPGWXXMQ

@misc{pith2026241108013,
  author       = {Pith},
  title        = {Pith review of: Investigating the Effectiveness of Explainability Methods in Parkinson's Detection from Speech},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JPGWXXMQ}},
  note         = {Machine review of arXiv:2411.08013}
}
read the original abstract

Speech impairments in Parkinson's disease (PD) provide significant early indicators for diagnosis. While models for speech-based PD detection have shown strong performance, their interpretability remains underexplored. This study systematically evaluates several explainability methods to identify PD-specific speech features, aiming to support the development of accurate, interpretable models for clinical decision-making in PD diagnosis and monitoring. Our methodology involves (i) obtaining attributions and saliency maps using mainstream interpretability techniques, (ii) quantitatively evaluating the faithfulness of these maps and their combinations obtained via union and intersection through a range of established metrics, and (iii) assessing the information conveyed by the saliency maps for PD detection from an auxiliary classifier. Our results reveal that, while explanations are aligned with the classifier, they often fail to provide valuable information for domain experts.

Figures

Figures reproduced from arXiv: 2411.08013 by the authors.

Figure 1
Figure 1. Explanations generated for a PD sample correctly classified by HuBERT. The explanations highlight different portions of the spectrogram, suggesting [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Analysis of the correlation between explanations overlap among [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BenSParX: A Robust Explainable Machine Learning Framework for Parkinson's Disease Detection from Bengali Conversational Speech

    cs.LG 2025-05 conditional novelty 6.0 of 10

    BenSParX provides a 120-speaker Bengali conversational speech dataset and a tuned machine learning framework that reports 95.77% accuracy, 95.57% F1, and 0.982 AUC for Parkinson's disease detection.

Reference graph

Works this paper leans on

44 extracted references · 41 canonical work pages · cited by 1 Pith paper

  1. [1]

    Biochemical aspects of parkinson’s disease,

    O. Hornykiewicz, “Biochemical aspects of parkinson’s disease,” Neurol- ogy, no. 2 suppl 2, pp. S2–S9, 1998

  2. [2]

    Parkinson disease,

    W. Poewe, K. Seppi, C. M. Tanner, G. M. Halliday, P. Brundin, J. V olkmann, A.-E. Schrag, and A. E. Lang, “Parkinson disease,”Nature reviews Disease primers, no. 1, pp. 1–21, 2017

  3. [3]

    Prevalence of parkinson’s disease in europe: A collaborative study of population-based cohorts. neurologic diseases in the elderly research group

    M. d. De Rijk, L. Launer, K. Berger, M. Breteler, J. Dar- tigues, M. Baldereschi, L. Fratiglioni, A. Lobo, J. Martinez-Lage, C. Trenkwalder et al. , “Prevalence of parkinson’s disease in europe: A collaborative study of population-based cohorts. neurologic diseases in the elderly research group.” Neurology, no. 11 Suppl 5, pp. S21–3, 2000

  4. [4]

    Treatments for dysarthria in parkinson’s disease,

    S. Pinto, C. Ozsancak, E. Tripoliti, S. Thobois, P. Limousin-Dowsey, and P. Auzou, “Treatments for dysarthria in parkinson’s disease,” The Lancet Neurology, no. 9, pp. 547–556, 2004

  5. [5]

    Imprecise vowel articulation as a potential early marker of parkinson’s disease: effect of speaking task,

    J. Rusz, R. Cmejla, T. Tykalova, H. Ruzickova, J. Klempir, V . Majerova, J. Picmausova, J. Roth, and E. Ruzicka, “Imprecise vowel articulation as a potential early marker of parkinson’s disease: effect of speaking task,” The Journal of the Acoustical Society of America , no. 3, pp. 2171–2181, 2013

  6. [6]

    Exploiting foundation models and speech enhancement for parkinson’s disease detection from speech in real-world operative conditions,

    M. La Quatra, M. F. Turco, T. Svendsen, G. Salvi, J. R. Orozco- Arroyave, and S. M. Siniscalchi, “Exploiting foundation models and speech enhancement for parkinson’s disease detection from speech in real-world operative conditions,” in Interspeech 2024, 2024

  7. [7]

    Molnar, Interpretable Machine Learning , 2nd ed., 2022

    C. Molnar, Interpretable Machine Learning , 2nd ed., 2022

  8. [8]

    Explainable artificial intelligence: A survey of needs, techniques, applications, and future direction,

    M. Mersha, K. Lam, J. Wood, A. K. AlShami, and J. Kalita, “Explainable artificial intelligence: A survey of needs, techniques, applications, and future direction,” Neurocomputing, 2024

Show all 44 references
  1. [9]

    Speech-based solution to parkinson’s disease management,

    B. Sonawane and P. Sharma, “Speech-based solution to parkinson’s disease management,” Multimedia Tools and Applications , no. 19, pp. 29 437–29 451, 2021

  2. [10]

    End-to-end parkinson’s disease detection using a deep convo- lutional recurrent network,

    C. D. Rios-Urrego, S. A. Moreno-Acevedo, E. N ¨oth, and J. R. Orozco- Arroyave, “End-to-end parkinson’s disease detection using a deep convo- lutional recurrent network,” in International Conference on Text, Speech, and Dialogue. Springer, 2022, pp. 326–338

  3. [11]

    Transfer learning helps to improve the accuracy to classify patients with different speech disorders in different languages,

    J. C. V ´asquez-Correa, C. D. Rios-Urrego, T. Arias-Vergara, M. Schuster, J. Rusz, E. Noeth, and J. R. Orozco-Arroyave, “Transfer learning helps to improve the accuracy to classify patients with different speech disorders in different languages,” Pattern Recognition Letters, p...

  4. [12]

    End-to-end deep learning approach for parkinson’s disease detection from speech signals,

    C. Quan, K. Ren, Z. Luo, Z. Chen, and Y . Ling, “End-to-end deep learning approach for parkinson’s disease detection from speech signals,” Biocybernetics and Biomedical Engineering , no. 2, pp. 556–574, 2022

  5. [13]

    Interpretable speech features vs. dnn embeddings: What to use in the automatic assessment of parkinson’s disease in multi- lingual scenarios,

    A. Favaro, Y .-T. Tsai, A. Butala, T. Thebaud, J. Villalba, N. Dehak, and L. Moro-Vel´azquez, “Interpretable speech features vs. dnn embeddings: What to use in the automatic assessment of parkinson’s disease in multi- lingual scenarios,” Computers in Biology and Medicine, p. 1...

  6. [14]

    Multilingual evaluation of in- terpretable biomarkers to represent language and speech patterns in parkinson’s disease,

    A. Favaro, L. Moro-Vel ´azquez, A. Butala, C. Motley, T. Cao, R. D. Stevens, J. Villalba, and N. Dehak, “Multilingual evaluation of in- terpretable biomarkers to represent language and speech patterns in parkinson’s disease,” Frontiers in Neurology, 2023

  7. [15]

    Automatic detection of parkinson’s disease with connected speech acoustic features: towards a linguistically interpretable approach,

    M. Maffia, L. Schettino, and V . N. Vitale, “Automatic detection of parkinson’s disease with connected speech acoustic features: towards a linguistically interpretable approach,” in Proceedings of the 9th Italian Conference on Computational Linguistics CLiC-it 2023: Venice, It...

  8. [16]

    Audiomnist: Exploring explainable artificial intelligence for audio analysis on a simple benchmark,

    S. Becker, J. Vielhaben, M. Ackermann, K.-R. M ¨uller, S. Lapuschkin, and W. Samek, “Audiomnist: Exploring explainable artificial intelligence for audio analysis on a simple benchmark,” 2023

  9. [17]

    Bubble Cooperative Networks for Identifying Important Speech Cues,

    V . A. Trinh, B. McFee, and M. I. Mandel, “Bubble Cooperative Networks for Identifying Important Speech Cues,” in Proc. Interspeech 2018, 2018

  10. [18]

    Identifying Important Time-Frequency Locations in Continuous Speech Utterances,

    H. S. Kavaki and M. I. Mandel, “Identifying Important Time-Frequency Locations in Continuous Speech Utterances,” in Proc. Interspeech 2020, 2020

  11. [19]

    Un- derstanding and Visualizing Raw Waveform-Based CNNs,

    H. Muckenhirn, V . Abrol, M. Magimai-Doss, and S. Marcel, “Un- derstanding and Visualizing Raw Waveform-Based CNNs,” in Proc. Interspeech 2019, 2019, pp. 2345–2349

  12. [20]

    Local interpretable model- agnostic explanations for music content analysis,

    S. Mishra, B. L. Sturm, and S. Dixon, “Local interpretable model- agnostic explanations for music content analysis,” in International Society for Music Information Retrieval Conference , 2017

  13. [21]

    Reliable local explanations for machine listening,

    S. Mishra, E. Benetos, B. L. Sturm, and S. Dixon, “Reliable local explanations for machine listening,” 2020

  14. [22]

    audiolime: Listenable explanations using source separation,

    V . Haunschmid, E. Manilow, and G. Widmer, “audiolime: Listenable explanations using source separation,” 2020

  15. [23]

    Tracing back music emotion predictions to sound sources and intuitive perceptual qualities,

    S. Chowdhury, V . Praher, and G. Widmer, “Tracing back music emotion predictions to sound sources and intuitive perceptual qualities,” 2021

  16. [24]

    Listen to interpret: Post-hoc interpretability for audio networks with nmf,

    J. Parekh, S. Parekh, P. Mozharovskyi, F. d'Alch ´e-Buc, and G. Richard, “Listen to interpret: Post-hoc interpretability for audio networks with nmf,” in Advances in Neural Information Processing Systems , 2022, pp. 35 270–35 283

  17. [25]

    Listenable maps for audio classifiers,

    F. Paissan, M. Ravanelli, and C. Subakan, “Listenable maps for audio classifiers,” in Proceedings of the 41st International Conference on Machine Learning, 2024, pp. 39 009–39 021

  18. [26]

    Lmac-td: Producing time domain explanations for audio classifiers,

    E. Mancini, F. Paissan, M. Ravanelli, and C. Subakan, “Lmac-td: Producing time domain explanations for audio classifiers,”arXiv preprint arXiv:2409.08655, 2024

  19. [27]

    Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

    W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Trans. Audio, Speech and Lang. Proc. , 2021

  20. [28]

    Deep inside convolutional networks: Visualising image classification models and saliency maps,

    K. Simonyan, A. Vedaldi, and A. Zisserman, “Deep inside convolutional networks: Visualising image classification models and saliency maps,” 2014

  21. [29]

    Smooth- grad: removing noise by adding noise,

    D. Smilkov, N. Thorat, B. Kim, F. Vi ´egas, and M. Wattenberg, “Smooth- grad: removing noise by adding noise,” 2017

  22. [30]

    Axiomatic attribution for deep networks,

    M. Sundararajan, A. Taly, and Q. Yan, “Axiomatic attribution for deep networks,” in Proceedings of the 34th International Conference on Machine Learning - Volume 70 , ser. ICML’17. JMLR.org, 2017

  23. [31]

    Grad-cam: Visual explanations from deep networks via gradient-based localization,

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in 2017 IEEE International Conference on Computer Vision (ICCV) , 2017

  24. [32]

    Striving for simplicity: The all convolutional net,

    J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. A. Riedmiller, “Striving for simplicity: The all convolutional net,” in 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Workshop Track Proceedings , Y . Bengio and Y . L...

  25. [33]

    A unified approach to interpreting model predictions,

    S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , ser. NIPS’17. Red Hook, NY , USA: Curran Associates Inc., 2017

  26. [34]

    Listenable maps for zero-shot audio classifiers,

    F. Paissan, L. D. Libera, M. Ravanelli, and C. Subakan, “Listenable maps for zero-shot audio classifiers,” arXiv preprint arXiv:2405.17615 , 2024

  27. [35]

    Posthoc interpretation via quantization,

    F. Paissan, C. Subakan, and M. Ravanelli, “Posthoc interpretation via quantization,” arXiv preprint arXiv:2303.12659 , 2023

  28. [36]

    Concise explanations of neural networks using adversarial training,

    P. Chalasani, J. Chen, A. R. Chowdhury, S. Jha, and X. Wu, “Concise explanations of neural networks using adversarial training,” 2020

  29. [37]

    Evaluating and aggregating feature-based model explanations,

    U. Bhatt, A. Weller, and J. M. F. Moura, “Evaluating and aggregating feature-based model explanations,” 2020

  30. [38]

    A benchmark for interpretability methods in deep neural networks,

    S. Hooker, D. Erhan, P.-J. Kindermans, and B. Kim, “A benchmark for interpretability methods in deep neural networks,” 2019

  31. [39]

    The detection of parkinson’s disease from speech using voice source information,

    N. Narendra, B. Schuller, and P. Alku, “The detection of parkinson’s disease from speech using voice source information,” IEEE/ACM Trans- actions on Audio, Speech, and Language Processing , pp. 1925–1936, 2021

  32. [40]

    New spanish speech corpus database for the analysis of people suffering from parkinson’s disease

    J. R. Orozco-Arroyave, J. D. Arias-Londo ˜no, J. F. Vargas-Bonilla, M. C. Gonzalez-R´ativa, and E. N¨oth, “New spanish speech corpus database for the analysis of people suffering from parkinson’s disease.” in Lrec, 2014, pp. 342–347

  33. [41]

    Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

    S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “Wavlm: Large-scale self-supervised pre- training for full stack speech processing,” IEEE Journal of Select...

  34. [42]

    Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

    Q. Kong, Y . Cao, T. Iqbal, Y . Wang, W. Wang, and M. D. Plumbley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,” 2020

  35. [43]

    Open-source conversational ai with speechbrain 1.0,

    M. Ravanelli, T. Parcollet, A. Moumen, S. de Langen, C. Subakan, P. Plantinga, Y . Wang, P. Mousavi, L. Della Libera, A. Ploujnikovet al., “Open-source conversational ai with speechbrain 1.0,” arXiv preprint arXiv:2407.00463, 2024

  36. [44]

    Phoneme dis- cretized saliency maps for explainable detection of ai-generated voice,

    S. Gupta, M. Ravanelli, P. Germain, and C. Subakan, “Phoneme dis- cretized saliency maps for explainable detection of ai-generated voice,” in Interspeech 2024, 2024, pp. 3295–3299

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.