Pith. sign in

REVIEW 4 major objections 5 minor 37 references

Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection

T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Speech-based Parkinson's detection models trained on frozen self-supervised embeddings learn a corpus-driven, disorder-general signal rather than a PD-specific signature, and under cross-lingual transfer they fail to distinguish…

desk verdict The PD-vs-dementia specificity result is the real contribution and broadly holds; the 'corpus-driven' headline is plausible but the design cannot separate corpus from language. read the letter →

arxiv 2608.13425 v1 pith:N4O7UXL4 submitted 2026-08-13 cs.CL eess.ASeess.SP

classification cs.CLeess.ASeess.SP
keywords Parkinson'sdiseaseself-supervisedspeechrepresentationscross-lingualtransferdementiabiomarkerslayer-wiseanalysisdiagnosticspecificitycorpusconfounds
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether self-supervised speech representations used for Parkinson's disease (PD) detection capture disease-specific motor or cognitive characteristics, or instead exploit dataset-specific cues. The authors train logistic-regression probes on frozen SSL embeddings from nine backbones across Spanish, German, and Czech PD corpora, then transfer the resulting classifiers to an independent German cohort that includes both PD and dementia. They report two findings: the best-performing SSL layer is determined mainly by the source corpus rather than the architecture, and classifiers that separate PD from healthy controls assign similarly high PD probabilities to dementia speech. The authors conclude that what survives cross-lingual transfer is a generic divergence from healthy speech, not a robust PD-specific signature, and they frame this as absence of evidence for PD specificity with direct implications for clinical deployment.

What carries the argument

The central machinery is a five-scenario transfer evaluation built on frozen SSL speech encoders and a low-capacity logistic-regression probe. For each corpus, backbone, and task, the representation layer with the highest mean balanced accuracy is selected under within-corpus 5-fold cross-validation (REF) and then fixed across all transfer scenarios, so that performance differences reflect distribution shift rather than layer choice. The load-bearing comparison is the transfer of PD-trained classifiers to the TREND cohort, which contains both PD and dementia participants, allowing the authors to test whether the transferred signal is specific to Parkinson's or merely indicates a general divergence from healthy speech.

What would settle it

A within-language cross-corpus transfer experiment, such as training on one German PD corpus and testing on another German PD corpus recorded under different conditions, that preserved PD-versus-dementia separation and stable layer selection would contradict the corpus-driven interpretation; alternatively, finding a target pathology (for example depression or essential tremor) for which the transferred PD classifiers do not show elevated probabilities would falsify the disorder-general claim.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that the discriminative signal in frozen SSL speech representations that transfers across corpora is corpus-driven and disorder-general. Under five progressively harder distribution shifts, covering re-takes, recording conditions, language, task, and combined language-task shifts, PD-versus-healthy balanced accuracy degrades steadily, from a mean change of −1.9 points for re-takes to −16.3 points for cross-lingual transfer. In the clinically most relevant setting, transferring PD-trained classifiers to the TREND cohort, only 9 of 540 transfer combinations achieve at least moderate PD-versus-control separation (AUC and balanced accuracy above 0.6). Those 9 classifiers separate PD from matched healthy controls but fail to distinguish PD from dementia, and after adjusting for age, sex, and education the PD-versus-dementia discrimination drops to chance. The authors interpret this as evidence that the models preserve a general patient-healthy separation rather than a disease-specific structure.

Load-bearing premise

The conclusion that transfer behaviour is driven by the corpus assumes that the source corpus can be separated from language and recording conditions, since the three training corpora differ jointly in language, recording hardware, and protocol with no within-language cross-corpus control.

Editorial extensions

If this is right

  • Speech-based PD classifiers with strong within-corpus accuracy can still lack diagnostic specificity, so evaluations that compare PD only against healthy controls are insufficient for clinical claims.
  • The optimal SSL layer for PD detection should not be assumed to transfer across corpora; layer choice varies substantially with the source dataset, especially for large backbones.
  • Multilingual pretraining (XLS-R, MMS) offers no consistent advantage over monolingual SSL backbones for cross-lingual PD detection.
  • Recording-condition mismatch degrades performance asymmetrically: training on noisy recordings transfers better to clean recordings than the reverse.
  • Performance degrades monotonically as distribution shift increases, with mean balanced-accuracy changes of −1.9 (re-take), −12.5 (recording condition), and −16.3 (cross-lingual) relative to the within-corpus reference.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural test of the corpus-driven interpretation would be a within-language cross-corpus transfer, such as training on one German PD corpus and testing on another German PD corpus recorded with different hardware, because the current design varies language, recording hardware, and protocol jointly with corpus.
  • The similar probabilities assigned to PD and dementia speech suggest the learned signal could be repurposed as a general marker of neurodegeneration or cognitive-motor decline, but this would require validation against several additional neurological conditions.
  • Because the non-specificity finding rests on only 9 transfer combinations that passed the performance threshold, the claim is strongest for those configurations; testing more target pathologies and more task combinations would clarify how broadly the disorder-general signal holds.
  • The instability of optimal layer choice across corpora could itself serve as a diagnostic of dataset shift: if the optimal layer depends mainly on the corpus, then a stable disease-specific speech signature is unlikely to be captured by any single SSL layer.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper investigates whether self-supervised speech representations capture Parkinson's disease (PD) specific signal that survives cross-lingual transfer. Using nine SSL backbones plus handcrafted eGeMAPS features, a logistic regression probe, and three special-session corpora (Spanish, German, Czech) together with an independent German TREND cohort, the authors define five transfer scenarios of increasing distribution shift. They report two main findings: (1) the optimal SSL layer in within-corpus reference evaluation varies substantially across source corpora, suggesting corpus-dependent layer selection; and (2) classifiers transferred to TREND separate PD from healthy controls but fail to separate PD from dementia, indicating a lack of pathological specificity. The paper interprets these results as evidence that transfer performance is driven by corpus-specific factors rather than a robust PD-specific speech signature, while carefully noting that this is absence of evidence, not evidence of absence.

Significance. If fully supported, the findings would make a valuable contribution to the debate on what speech-based PD classifiers actually learn, and the paper is among the first to evaluate cross-pathology specificity in this setting. The experimental design is thoughtful: multiple SSL backbones spanning capacity, fine-tuning, and pretraining language; layer selection is isolated from transfer evaluation; and the held-out TREND cohort provides a clinically meaningful cross-disease comparison. The authors also explicitly avoid overclaiming by framing the null result as absence of evidence, not evidence of absence. However, the central attribution of the findings to 'corpus' is currently underdetermined by the design, because the source corpora vary jointly in language and recording conditions, and the layer-selection results lack quantification of the sampling variability of the argmax layer.

major comments (4)
  1. [Section III-A, Table III and Figure 2] The conclusion that layer selection is 'corpus-dependent' conflates corpus identity with language and recording conditions: the three source corpora (DE, ES, CZ) differ simultaneously in language, recording hardware, and protocol (Table I), and no within-language cross-corpus comparison is provided for layer selection. The only same-language pair (ES and ES-e) is used exclusively in Scenario 2 and is never analyzed for layer selection. As a result, the observed differences in optimal layer could equally be attributed to language or recording condition, and the title's answer 'Corpus' is not supported over 'Language' or 'Recording condition'. The authors should either add a same-language control (e.g., compare layer selection between ES and ES-e, noting the task limitations) or temper the conclusion to 'dataset identity' with an explicit acknowledgment of this confound.
  2. [Section III-A, Table III] The argmax layer is reported as a single number per corpus-backbone-task, and the σ values in Table III quantify variation across corpora, not within-corpus variability across the 10 seeds and folds. Since the layer-wise balanced-accuracy curves in Figure 2 are often flat (especially for ES), the argmax estimate may be dominated by sampling noise. Provide bootstrap or permutation confidence intervals for the selected layer per corpus, or a statistical test (e.g., a corpus-by-layer interaction in balanced accuracy) to show that cross-corpus differences in optimal depth are larger than within-corpus seed variation. Without this, large σ values such as WavLM-L on DDK (0.43) cannot be interpreted as a robust corpus effect.
  3. [Section III-E] The non-specificity claim is evaluated only on 9 of 540 transfer combinations, selected by thresholds on target-data AUC and BA greater than 0.6. This selection uses target labels to choose which combinations to interpret, and the authors do not report how many of the remaining 531 combinations would have shown PD-versus-dementia separation had they exceeded the threshold. Consequently, the second main finding may depend on a small and potentially unrepresentative subset. Please report the distribution of PD-versus-dementia separation (e.g., AUC or Mann–Whitney U p-values) across all 540 combinations, or justify why the threshold does not qualitatively affect the conclusion.
  4. [Section III-E, demographics adjustment] The description of the residualisation analysis is too brief to verify: 'We than adjusted the PD probability scores for these covariates using leave-one-subject-out (LOSO) regression. The residualised performance drops to chance level (AUC = 0.50–0.55)'. It is unclear whether the residuals are computed on predicted probabilities or logits, which TREND groups are included in the residualized evaluation, and what the residualized AUC is for PD versus HC(PD) as opposed to PD versus dementia. Please clarify the procedure and report the full results, including the demographic-only model's performance on the same groups.
minor comments (5)
  1. [Section IV, Discussion] The sentence fragment 'Identity-preserving re-takes. S1 (+Re-Take) have a small degradation of performance' should be rewritten, for example as 'Identity-preserving re-takes (S1) show only a small degradation in balanced accuracy (−1.9 BA points).'
  2. [Section III-E] The phrase 'We than adjusted' should be 'We then adjusted'.
  3. [Figures 3 and 4 captions] The caption text 'n= 2000test-set resamples' is missing a space and should read 'n = 2000 test-set resamples'.
  4. [Section III-A] The sentence 'These trends are broadly consistent across backbone variance axis' should be 'across backbone variance axes'.
  5. [Section III-D] The claim that 'no consistent advantage of multilingual (Mu) over monolingual (Mo) SSL backbones is observed' is made without supporting statistical tests or effect sizes; adding a small number of tests or confidence intervals would strengthen the claim.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central negative result is an empirical observation on held-out TREND data, not a fitted output or self-citation chain.

full rationale

The paper's derivation chain is an empirical evaluation: REF layer selection uses only within-corpus training data; transfer scenarios then freeze the layer and evaluate on held-out target corpora. The central claim that transferred PD classifiers lack pathological specificity (Section III-E, Figure 6) is an observed null result on the TREND cohort, not a quantity reconstructed from the inputs by construction. The only data-dependent selection is the restriction to 9 transfer combinations that reach AUC>0.6 and BA>0.6 for PD-vs-HC; this is a conditional analysis of settings where transfer works at all, and it does not force the subsequent PD-vs-dementia non-separation. Self-citations ([17], [24]) support dataset provenance and preprocessing choices and are aligned with, rather than load-bearing for, the main conclusion. The corpus-vs-language confound noted in the skeptic summary is a validity/interpretation concern, not a circularity: no equation defines corpus in terms of language, and no fitted parameter is renamed as a prediction. The manuscript explicitly labels the result as absence of evidence for PD-specificity, not evidence of its absence, which further distances it from a circular claim.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The paper does not introduce new theoretical entities or fitted constants beyond standard model parameters. The main data-driven choices are the per-corpus layer selection and the demographic residualisation model used as a sensitivity check. The key assumptions are about linear decodability of SSL features, the interpretability of the 'corpus' factor when language and recording conditions are confounded, and the comparability of the TREND dementia and PD groups.

free parameters (2)
  • Selected SSL layer index (argmax BA) per corpus-backbone-task = Varies; e.g., WavLM-L DDK: DE=24, ES=1, CZ=22
    Chosen as the layer with the highest within-corpus balanced accuracy in the REF setting; this data-driven choice underpins the claim that optimal layer selection is corpus-dependent, and its noise is not quantified with confidence intervals.
  • Demographic logistic regression coefficients (age, sex, education) = Not reported; model achieves AUC 0.73-0.75 for PD vs dementia on TREND
    Fitted on TREND PD versus dementia data and used to residualise p(PD) scores; this is a sensitivity analysis supporting the non-specificity claim but is itself a fitted model on the same target cohort.
assumptions (4)
  • domain assumption A logistic regression probe on mean-pooled SSL layer embeddings is sufficient to decode pathology-relevant information, so failure to separate PD from dementia reflects absent linear structure rather than probe weakness.
    All conclusions about what is encoded in SSL representations are drawn through a linear probe on frozen, mean-pooled features (Section II-B). Nonlinear structure relevant to PD could be missed.
  • ad hoc to paper The three special-session corpora can be treated as varying 'corpus identity' while language and recording conditions are exchangeable.
    DE, ES, and CZ differ in language, hardware, and protocol simultaneously; the paper's 'corpus-dependent' interpretation (Sections III-A and IV) assumes corpus identity is the operative variable.
  • domain assumption TREND dementia and PD groups are comparable after sex, age, and education matching and residualisation.
    The PD-versus-dementia null result is interpreted as lack of pathological specificity (Section III-E); differences in age, education, or cognitive status could otherwise explain the inability to separate the groups.
  • domain assumption SSL backbones pretrained on predominantly healthy speech carry no task-specific PD supervision, so any PD signal in frozen layers emerges without disease-specific training.
    Motivates the frozen-probe design (Introduction); if some pretraining corpora contain disordered speech, the interpretation of what is generic changes.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection." pith.science (2026). https://pith.science/paper/N4O7UXL4

@misc{pith2026260813425,
  author       = {Pith},
  title        = {Pith review of: Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/N4O7UXL4}},
  note         = {Machine review of arXiv:2608.13425}
}
read the original abstract

Self-supervised learning (SSL) speech representations achieve strong performance for Parkinson's disease (PD) detection within individual corpora. However, it remains unclear whether these models capture disease-related characteristics or exploit dataset-specific confounds, particularly since most SSL backbones are pretrained exclusively on healthy speech. To investigate this question, we perform a layer-wise analysis of nine SSL speech backbones using a low-capacity logistic regression probe across three languages. We structure the evaluation as multiple scenarios that progressively introduce distribution shifts in participant identity, recording conditions, language, and pathology. Our results reveal two key findings. First, layer selection is highly corpus-dependent: the optimal representation layer is determined primarily by the source dataset rather than by the SSL architecture itself. Second, the transferred discriminative signal lacks pathological specificity: classifiers trained to detect PD assign similarly high probabilities to both PD and dementia speech in the target corpus. These results highlight critical limitations that must be addressed before speech-based pathology recognition models can be reliably deployed in clinical settings.

Figures

Figures reproduced from arXiv: 2608.13425 by the authors.

Figure 1
Figure 1. Five scenarios for evaluating cross-lingual transfer in speech-based PD detection. REF uses 5-fold intra-corpus cross-validation to select the SSL backbone layer, which is then frozen while the binary PD classifier is trained on the full training corpus and evaluated on the target corpus without adaptation. Scenarios (S1–S5) introduce progressively larger train–test distribution shifts. The Train and Test columns sh… view at source ↗
Figure 2
Figure 2. Layer selection results for the READ task in the REF setting, where the selected backbone layer is fixed for all downstream transfer experiments. The three plots compare: i) Capacity: HuBERT-B vs. HuBERT-L and WavLM-B vs. WavLM-L; ii) Fine-tuning: W2V2-B vs. W2V2-B-FT and HuBERT-L vs. HuBERT-L-FT; iii) Pretraining: W2V2-B vs. XLS-R and MMS. The READ task is shown because it achieved the highest overall balanced accu… view at source ↗
Figure 3
Figure 3. S1 (+Re-Take) Results: Change in balanced accuracy relative to REF (∆BA = BAS1 − BAREF). S1 differs from REF only in the test recordings, which are replaced by each participant’s second recording; training data and participant splits are unchanged. Values close to zero indicate robust performance across recording takes. Results are reported as mean ± bootstrapped SD (n = 2000 test-set resamples). Columns correspond … view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: S2 (+Condition) Results: Change in balanced accuracy relative to REF (∆BA = BAS2 − BAREF) under recording￾condition shift between clean (C) ES and noisy (N) ES-e corpora with non-overlapping speakers. C→N denotes training on clean and testing on noisy data; N→C denotes…
Figure 5
Figure 5. Figure 5: S3 (+Language) Results: Seven monolingual pre￾trained SSL backbones (W2V2-B, W2V2-B-FT, HuBERT￾B, HuBERT-L, HuBERT-L-FT, WavLM-B, WavLM-L), two multilingual SSL backbones (XLS-R, MMS), and handcrafted eGeMAPS features. Horizontal lines indicate group-wise mean values. …
Figure 6
Figure 6. Figure 6: S4 (+Task) and S5 (+Task + Language) Results: Predicted PD probability distributions for the highest-balanced￾accuracy (BA) transfer combination of each training corpus. Layers were selected using only within-corpus REF data; no TREND data was used during layer selecti…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

37 extracted references · 29 canonical work pages

  1. [1]

    Parkinson’s Disease,

    B. R. Bloem, M. S. Okun, and C. Klein, “Parkinson’s Disease,”The Lancet, vol. 397, no. 10291, pp. 2284–2303, 2021

  2. [2]

    Movement Disorder Society-Sponsored Revision of the Unified Parkinson’s Disease Rating Scale (MDS-UPDRS): Scale Presentation and Clinimetric Testing Re- sults,

    C. G. Goetz, B. C. Tilley, S. R. Shaftmanet al., “Movement Disorder Society-Sponsored Revision of the Unified Parkinson’s Disease Rating Scale (MDS-UPDRS): Scale Presentation and Clinimetric Testing Re- sults,”Movement Disorders, vol. 23, no. 15, pp. 2129–2170, 2008

  3. [3]

    Parkinsonism: Onset, Progression and Mortality,

    M. M. Hoehn and M. D. Yahr, “Parkinsonism: Onset, Progression and Mortality,”Neurology, vol. 17, no. 5, pp. 427–442, 1967

  4. [4]

    The Consortium to Estab- lish a Registry for Alzheimer’s Disease (CERAD). Part I. Clinical and Neuropsychological Assessment of Alzheimer’s Disease,

    J. C. Morris, A. Heyman, R. C. Mohs, J. P. Hughes, G. van Belle, G. Fillenbaum, E. D. Mellits, and C. Clark, “The Consortium to Estab- lish a Registry for Alzheimer’s Disease (CERAD). Part I. Clinical and Neuropsychological Assessment of Alzheimer’s Disease,”Neurology, vol. 39, no. 9, pp. 1159–1165, 1989

  5. [5]

    ”Mini-Mental State

    M. F. Folstein, S. E. Folstein, and P. R. McHugh, “”Mini-Mental State”: A Practical Method for Grading the Cognitive State of Patients for the Clinician,”Journal of Psychiatric Research, vol. 12, no. 3, pp. 189–198, 1975

  6. [6]

    Exploring digital speech biomarkers of hypokinetic dysarthria in a multilingual cohort,

    D. Kovac, J. Mekyska, V . Aharonson, P. Harar, Z. Galaz, S. Rapcsak, J. R. Orozco-Arroyave, L. Brabenec, and I. Rektorova, “Exploring digital speech biomarkers of hypokinetic dysarthria in a multilingual cohort,”Biomedical Signal Processing and Control, vol. 88, p. 105667, 2024. [Online]. Available: https://www.sciencedirect.com/ science/article/pii/S1746...

  7. [7]

    Speech- Based Parkinson’s Detection Using Pre-Trained Self-Supervised Automatic Speech Recognition (ASR) Models and Supervised Contrastive Learning,

    H. Sedigh Malekroodi, N. Madusanka, B.-i. Lee, and M. Yi, “Speech- Based Parkinson’s Detection Using Pre-Trained Self-Supervised Automatic Speech Recognition (ASR) Models and Supervised Contrastive Learning,”Bioengineering, vol. 12, no. 7, 2025. [Online]. Available: https://www.mdpi.com/2306-5354/12/7/728

  8. [8]

    Analyzing Wav2Vec 1.0 Embeddings for Cross-Database Parkinson’s Disease Detection and Speech Features Extraction,

    O. Klemp ´ıˇr and R. Krupi ˇcka, “Analyzing Wav2Vec 1.0 Embeddings for Cross-Database Parkinson’s Disease Detection and Speech Features Extraction,”Sensors, vol. 24, no. 17, p. 5520, Aug. 2024

Show all 37 references
  1. [9]

    Towards a Corpus (and Language)-Independent Screening of Parkinson’s Disease from V oice and Speech through Domain Adaptation, journal = Bioengineering,

    E. J. Ibarra, J. D. Arias-Londo ˜no, M. Za ˜nartu, and J. I. Godino- Llorente, “Towards a Corpus (and Language)-Independent Screening of Parkinson’s Disease from V oice and Speech through Domain Adaptation, journal = Bioengineering,” vol. 10, no. 11, 2023. [Online]. Available:...

  2. [10]

    A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations,

    A. Y . F. Chiu, K. C. Fung, R. T. Y . Li, J. Li, and T. Lee, “A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations,” 2026. [Online]. Available: https://arxiv.org/abs/2501.05310

  3. [11]

    Towards a Generalizable Speech Marker for Parkinson’s Disease Diagnosis,

    M. Siniukov, E. Xing, S. A. Isfahani, and M. Soleymani, “Towards a Generalizable Speech Marker for Parkinson’s Disease Diagnosis,”

  4. [12]

    Analysis of spontaneous speech in Parkinson’s disease by natural language processing,

    K. Yokoi, Y . Iribe, N. Kitaoka, T. Tsuboi, K. Hiraga, Y . Satake, M. Hattori, Y . Tanaka, M. Sato, A. Hori, and M. Katsuno, “Analysis of spontaneous speech in Parkinson’s disease by natural language processing,”Parkinsonism & Related Disorders, vol. 113, p. 105411, Aug 2023

  5. [13]

    Speech features-based Parkin- son’s disease classification using combined SMOTE-ENN and binary machine learning,

    S. Dhanalakshmi, S. Das, and R. Senthil, “Speech features-based Parkin- son’s disease classification using combined SMOTE-ENN and binary machine learning,”Health Technology, vol. 14, no. 2, pp. 393–406, Mar 2024

  6. [14]

    Bilingual Dual-Head Deep Model for Parkinson’s Disease Detection from Speech,

    M. L. Quatra, J. R. Orozco-Arroyave, and S. M. Siniscalchi, “Bilingual Dual-Head Deep Model for Parkinson’s Disease Detection from Speech,”ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1–5,

  7. [15]

    Machine Learning-Based Classification of Parkinson’s Disease Patients Using Speech Biomarkers,

    M. A. Hossain and F. Amenta, “Machine Learning-Based Classification of Parkinson’s Disease Patients Using Speech Biomarkers,”Journal of Parkinson’s Disease, vol. 14, pp. 95 – 109, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:266681923

  8. [16]

    Available: https://api.semanticscholar.org/CorpusID: 276961433

    [Online]. Available: https://api.semanticscholar.org/CorpusID: 276961433

  9. [17]

    Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson’s Disease,

    A. Hernandez, E. Yeo, K. Choi, C.-J. Li, Z. Yue, R. K. Das, J. Rusz, M. M. Doss, J. R. Orozco-Arroyave, T. Arias-Vergara, A. Maier, E. N¨oth, D. R. Mortensen, D. Harwath, and P. A. Perez-Toro, “Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detect...

  10. [18]

    Multi-Modal Decentralized Hybrid Learning for Early Parkinson’s Detection Using V oice Biomarkers and Contrastive Speech Embeddings,

    K. M. Alhawiti, “Multi-Modal Decentralized Hybrid Learning for Early Parkinson’s Detection Using V oice Biomarkers and Contrastive Speech Embeddings,”Sensors (Basel, Switzerland), vol. 25, 2025. [Online]. Available: https://api.semanticscholar.org/CorpusID:283066469

  11. [19]

    Emotional State Modeling for the Assessment of Depression in Parkinson’s Disease,

    P. A. P ´erez-Toro, J. C. Vasquez-Correa, T. Arias-Vergara, P. Klumpp, M. Schuster, E. N ¨oth, and J. R. Orozco-Arroyave, “Emotional State Modeling for the Assessment of Depression in Parkinson’s Disease,” in Text, Speech, and Dialogue, K. Ek ˇstein, F. P ´artl, and M. Konop ´...

  12. [20]

    Automatic speech-based assessment to discriminate Parkinson’s disease from essential tremor with a cross-language approach,

    C. D. Rios-Urrego, J. Rusz, and J. R. Orozco-Arroyave, “Automatic speech-based assessment to discriminate Parkinson’s disease from essential tremor with a cross-language approach,”npj Digital Medicine, vol. 7, p. 37, 2024. [Online]. Available: https://doi.org/10.1038/ s41746-0...

  13. [21]

    (2026) T ¨ubinger Erhebung von Risikofaktoren zur Erkennung von Neurodegeneration (TREND)

    TREND Study Group. (2026) T ¨ubinger Erhebung von Risikofaktoren zur Erkennung von Neurodegeneration (TREND). University Hospital T¨ubingen. Accessed: 2026-01-03. [Online]. Available: https://www. trend-studie.de/

  14. [22]

    Detection of persons with Parkinson’s disease by acoustic, vocal, and prosodic analysis,

    T. Bocklet, E. N ¨oth, G. Stemmer, H. Ruzickova, and J. Rusz, “Detection of persons with Parkinson’s disease by acoustic, vocal, and prosodic analysis,” in2011 IEEE Workshop on Automatic Speech Recognition & Understanding, 2011, pp. 478–483

  15. [23]

    SciPy 1.0: fundamental algorithms for scientific computing in Python,

    P. Virtanenet al., “SciPy 1.0: fundamental algorithms for scientific computing in Python,”Nature Methods, vol. 17, pp. 261–272, 2020

  16. [24]

    On implementing 2D rectangular assignment algorithms, au- thor=Crouse, David F

    “On implementing 2D rectangular assignment algorithms, au- thor=Crouse, David F.”IEEE Transactions on Aerospace and Electronic Systems, vol. 52, no. 4, pp. 1679–1696, 2016

  17. [25]

    The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for V oice Research and Affective Computing,

    F. Eyben, K. R. Scherer, B. W. Schuller, J. Sundberg, E. Andr´e, C. Busso, L. Y . Devillers, J. Epps, P. Laukka, S. S. Narayanan, and K. P. Truong, “The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for V oice Research and Affective Computing,”IEEE Transactions on Affect...

  18. [26]

    Beyond Binary: Speech Representations Across the Cognitive Score Hierarchy,

    S. Kopar, R. P. Rane, C. Mychajliw, L. Federmann, G. Eschweiler, D. Berg, S. Gijsen, P. A. Perez-Toro, and K. Ritter, “Beyond Binary: Speech Representations Across the Cognitive Score Hierarchy,” 2026. [Online]. Available: https://arxiv.org/abs/2605.27189

  19. [27]

    HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

    W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3451–3460, 2021

  20. [28]

    openSMILE: the Munich versatile and fast open-source audio feature extractor,

    F. Eyben, M. W ¨ollmer, and B. Schuller, “openSMILE: the Munich versatile and fast open-source audio feature extractor,” inProceedings of the 18th ACM international conference on Multimedia, 2010, pp. 1459–1462

  21. [29]

    wav2vec 2.0: A framework for self-supervised learning of speech representations,

    A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020, pp. 12 449–12 460

  22. [30]

    WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,

    S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Liet al., “WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,”IEEE Journal of Selected Topics in Signal Pro- cessing, vol. 16, no. 6, pp. 1505–1518, 2022

  23. [31]

    Scaling Speech Tech- nology to 1,000+ Languages,

    V . Pratap, A. Tjandra, B. Shi, P. Tomasello, A. Babu, S. Kundu, Z. Ni, A. Vyas, M. Fazel-Zarandi, A. Baevskiet al., “Scaling Speech Tech- nology to 1,000+ Languages,”Journal of Machine Learning Research (JMLR), vol. 25, 2024, preprint available on arXiv:2305.13516

  24. [32]

    XLS-R: Self-supervised cross-lingual speech representation learning at scale,

    A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y . Saraf, J. Pino, A. Baevski, A. Conneau, and M. Auli, “XLS-R: Self-supervised cross-lingual speech representation learning at scale,” inProceedings of Interspeech, 2022, pp. 2278–2282

  25. [33]

    Naive application of permutation testing leads to inflated type I error rates,

    G. A. Churchill and R. W. Doerge, “Naive application of permutation testing leads to inflated type I error rates,”Genetics, vol. 178, no. 1, pp. 609–610, Jan 2008

  26. [34]

    scikit-learn: Machine Learning in Python,

    F. Pedregosa, G. Varoquaux, A. Gramfort, V . Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V . Dubourg, J. Vander- plas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duches- nay, “scikit-learn: Machine Learning in Python,” pp. 2825–2830, 2011

  27. [35]

    Towards a Corpus (and Language)-Independent Screening of Parkin- son’s Disease from V oice and Speech through Domain Adaptation,

    E. J. Ibarra, J. D. Arias-Londo ˜no, M. Za˜nartu, and J. I. Godino-Llorente, “Towards a Corpus (and Language)-Independent Screening of Parkin- son’s Disease from V oice and Speech through Domain Adaptation,” Bioengineering, vol. 10, no. 11, p. 1316, 2023

  28. [36]

    On a test of whether one of two random variables is stochastically larger than the other,

    H. B. Mann and D. R. Whitney, “On a test of whether one of two random variables is stochastically larger than the other,”The Annals of Mathematical Statistics, vol. 18, no. 1, pp. 50–60, 1947. [Online]. Available: https://doi.org/10.1214/aoms/1177730491

  29. [2025]

    Available: https://arxiv.org/abs/2501.03581

    [Online]. Available: https://arxiv.org/abs/2501.03581

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.