{"id":"05dda37e-19a4-464f-9640-bcb527f54bb7","arxiv_id":"2411.08013","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A systematic benchmark shows that standard saliency methods faithfully reflect a Parkinson's speech classifier's decisions, yet the resulting spectrogram highlights are not readily usable by domain experts.","lead":"This paper tests six common explainability tools on a speech-based Parkinson's detection model. It finds that the tools highlight features that align with the model's decisions, but those highlights are still hard for doctors and clinicians to interpret.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central negative claim — that saliency maps are not useful to domain experts — is asserted without any expert study; the qualitative reading of Figure 1 in Section V-D cannot support a claim about clinicians.","rationale":"The reader's weakest_assumption identifies exactly the same gap. I agree: the negative half of the central claim is the least secure. The paper is a useful benchmark; the faithfulness metrics and the auxiliary classifier provide some support for the positive half, but the expert-usefulness claim is not measured at all. A conditional verdict is appropriate because the issue is addressable with a modest expert study. No verdict change is needed beyond the reader's conditional.","tokens_in":9227,"tokens_out":10920,"duration_ms":134013,"concrete_test":"Recruit at least 10 clinicians (SLPs or neurologists) and present them with saliency maps (Integrated Gradients and Guided GradCAM) overlaid on spectrograms for a balanced set of PD and HC utterances from s-PC-GITA, alongside the unmodified spectrograms. Ask each clinician to (a) classify each sample as PD or HC and (b) rate whether the highlighted regions correspond to known PD speech features on a Likert scale. Pre-register the analysis. If classification accuracy with maps significantly exceeds the no-map baseline or usefulness ratings are above neutral, the negative claim fails; if not, the conclusion is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The conclusion (Section VI) and abstract claim that explanations 'fail to provide valuable information for domain experts.' The only evidence is the authors' qualitative observation (Section V-D) that maps are 'not easily interpretable by humans,' illustrated by one spectrogram (Figure 1). No speech-language pathologists or neurologists were recruited; no rating task, forced-choice test, or inter-rater reliability was run; and the highlighted regions were never compared with established PD speech markers (e.g., vowel space, DDK, jitter/shimmer). This is an empirical claim about a specific expert population, and the paper does not sample that population. The auxiliary classifier experiment (Table III) even shows the maps carry class-discriminative information, so the missing link is whether human experts can extract it. If experts can use these maps, the paper's headline conclusion is false. The positive faithfulness claim is better supported and is independent of the negative claim.","agreement_with_reader":"agree"},"referee_report":null,"author_rebuttal":null,"desk_editor":null,"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-12T22:00:20.816854+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}