Pith. sign in

REVIEW 3 major objections 5 minor 2 cited by

Unveiling Interpretability in Self-Supervised Speech Representations for Parkinson's Diagnosis

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Cross-attention makes black-box speech embeddings explainable for Parkinson's diagnosis.

desk verdict A useful, honest interpretability framework for SSL-based PD detection, but the headline claim about attention-as-explanation is softer than the title suggests. read the letter →

arxiv 2412.02006 v2 pith:HWFMXPDJ submitted 2024-12-02 cs.CV

classification cs.CV
keywords Parkinson'sDiseaseSelf-SupervisedSpeechRepresentationsCross-AttentionMechanismsInterpretabilityDeepLearningBiomarkersPathologicalAnalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that the opaque embeddings produced by a self-supervised speech model (one pre-trained on large amounts of unlabeled audio) can be made transparent enough to support Parkinson's disease diagnosis. Its mechanism is a cross-attention training setup that aligns the 1024-dimensional Wav2Vec2.0 embedding sequence with 35 static, clinically informed speech features covering articulation, glottal, phonation, and prosody. The resulting attention weights are read at two levels: which embedding dimensions track which clinical feature, and which moments of the utterance matter most. Across five PD speech corpora and six assessment tasks, the attended representations stay competitive with the SSL-only black-box baseline, and in spontaneous monologue they transfer across languages. A sympathetic reader would care because this is a concrete route from black-box deep speech models to a tool a clinician can question.

What carries the argument

The carrying device is a single-head scaled dot-product cross-attention mechanism in which the query and value projections are learned while the key projection is fixed to the identity matrix: $Q = X_{\mathrm{ssl}} W_Q$, $K = X_{\mathrm{inf}}$, $V = X_{\mathrm{ssl}} W_V$. Because the keys are the 35 clinically informed features kept unchanged, the two resulting attention matrices are directly readable: the embedding-level matrix $S_{\mathrm{emb}}$ says which Wav2Vec2.0 dimensions align with which clinical feature, and the temporal matrix $S_{\mathrm{temp}}$ says which time step of the utterance aligns with which feature. The informed feature set is repeated to match either the time axis or the embedding axis, which is what allows a static feature vector to serve as an interpretable key in both perspectives. The classification module then consumes averaged, enriched representations from both branches, so the same trained model that produces the explanation also produces the prediction.

What would settle it

A clinical validation study comparing the attention-highlighted features (e.g., amplitude perturbation quotient in diadochokinetic tasks, pause-duration variability in read speech) against independently measured acoustic or perceptual ratings of the same patients would settle the matter. If the top-attended features do not correlate with those independent measures on held-out data, or if temporal attention peaks do not align with the moments a speech therapist marks as pathological, the interpretability claim is falsified.

Watch

Extended reading notes

Core claim

On its own terms, the paper claims that the internal dimensions of a frozen XLS-R Wav2Vec2.0 representation do not encode interpretable speech features as isolated values, but that these features can nevertheless be retrieved by attention when the embeddings are aligned, as queries and values, against a static set of 35 clinically informed features acting as keys with an identity projection. The learned embedding-level attention matrix $S_{\mathrm{emb}} \in \mathbb{R}^{D\times F}$ quantifies, for each of the 1024 embedding dimensions, the relevance of each informed feature, while the temporal matrix $S_{\mathrm{temp}} \in \mathbb{R}^{T\times F}$ attributes the same features to each 20 ms time step. The paper's evidence that this works is behavioral: in the GITA corpus, attention weights concentrate on phonation and glottal features for sustained vowels, shift to prosodic features including pause variability in read text and sentences, and, in a contrastive healthy-control analysis, prosody peaks align with the words marked for emphasis in the GITA protocol, all while classification F1 stays broadly on par with the SSL-only baseline.

Load-bearing premise

The load-bearing premise is that the attention weights learned for classification genuinely reflect clinically meaningful speech attributes, rather than dataset-specific cues or spurious correlations, a premise the paper itself flags as debated and still awaiting clinical validation.

Editorial extensions

If this is right

  • A frozen self-supervised encoder can be interpreted without fine-tuning, since only the attention projections and the classifier are trained on PD data.
  • The explanation is task-sensitive: sustained-vowel tasks pull attention toward phonation and glottal dimensions, while continuous speech tasks pull attention toward prosodic dimensions such as pause duration and speech rate.
  • Temporal attention, contrasted against healthy controls, localizes relevant moments at word and phoneme level, with prosody peaking at words marked for emphasis.
  • The interpretable model stays competitive with the SSL-only black-box baseline on most tasks and transfers notably well in cross-lingual spontaneous monologue.
  • Swapping the static informed feature set re-targets the framework to other assessment tasks or conditions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • We infer that the attention matrices could be turned into per-patient speech profiles and tested against independent acoustic measurements to determine whether attention tracks the underlying signal rather than a class-discriminative shortcut.
  • We infer that ablating individual informed features should shift the attention pattern in a predictable direction, given that the identity key projection is what preserves interpretability; the paper does not run this check.
  • We infer that the same cross-attention design would carry over to other neurodegenerative or cognitive-communication conditions if the informed feature set were expanded with macro-descriptors such as word-finding difficulty.
  • We suspect the accuracy drops on FraLusoPark and CzechPD mark cases where the static informed features are weak anchors, which would motivate adding dynamic clinical features as keys.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes an interpretable framework for Parkinson's Disease (PD) detection that combines self-supervised speech representations (XLS-R Wav2Vec2.0 embeddings) with 35 clinically informed DisVoice features through two cross-attention modules. The embedding-level module yields an attention matrix over SSL embedding dimensions and informed features, and the temporal-level module yields a time-step-by-feature attention matrix; both feed a linear classifier. The method is evaluated on five PD speech corpora (NeuroVoz, GITA, FraLusoPark, GermanPD, CzechPD) across six assessment tasks using speaker-independent nested cross-validation. Classification results are competitive with an SSL-only baseline, with some noticeable drops on FraLusoPark and CzechPD, and the method shows strong cross-lingual performance on the MONOLOGUE task. Interpretability analyses are performed on the GITA corpus and show attention patterns that the authors argue align with expected speech dimensions per task, plus temporal case studies of individual PD subjects.

Significance. If the interpretability claims were fully supported, this would be a useful contribution toward transparent SSL-based clinical speech analysis: it is one of few works to explicitly inject clinically informed features into an attention mechanism over SSL embeddings, it evaluates across five languages and multiple tasks, and it honestly reports the accuracy/transparency trade-off. The paper also ships a public code repository and uses a task-specific evaluation protocol that facilitates comparison. However, the central claim that the framework 'effectively identifies distinct speech patterns' rests on post-hoc visual inspection of attention maps from a single corpus, a single random seed, and only correctly predicted test samples, with no null model or statistical validation. The authors themselves acknowledge the ongoing debate on attention interpretability and state that the explanations have not undergone medical validation. As a result, the interpretability evidence is currently suggestive rather than demonstrative, and the paper would need additional validation to support its main claim.

major comments (3)
  1. [Section VI-B] The central claim that the framework 'effectively identifies distinct speech patterns ... align[ed] with the expected speech dimensions of each assessment task' is supported only by attention maps computed on the GITA corpus, using the best-performing random seed and only the test samples that were correctly predicted. This selection is acknowledged in the text, but it conditions the explanation on the model's own decision rule and on the most favorable run. As a result, the reported HC/PD differences and task-specific alignments could be artifacts of this selection. Please report the stability of the attention patterns across all seeds and folds, include misclassified samples for comparison, and provide a null baseline (e.g., permuted labels or randomly initialized attention) to demonstrate that the observed alignment is not produced by chance.
  2. [Section III-B and VI-B] The paper cites the debate on whether attention weights are faithful explanations [46]-[48] and correctly states that high attention does not imply the presence or severity of a speech impairment, but then interprets high mean attention on logE, F1, and glottal features as evidence that the model identified phonatory and glottal dynamics in the VOWELS task. Because the key projection is the identity matrix in Eq. (2), the attention scores are softmax-normalized alignments between projected SSL embeddings and static informed features; they have no demonstrated causal link to the classification outcome or to the clinical constructs. A concrete fidelity test is needed, such as comparing attention-based feature rankings with gradient-based attributions on the same model, or ablating the highest-attended features and measuring the effect on classification performance, to substantiate that attention is an explanation rather than only an alignment score.
  3. [Section VI-B, Figure 3] The claimed differentiation between HC and PD groups in embedding-level attention is descriptive only: no error bars, confidence intervals, or significance tests are reported, and the paper concedes that analyses across the other corpora 'did not always show a consistent pattern.' Without quantifying variability across seeds, folds, or corpora, the robustness of the interpretability finding is not established. Please report per-seed and per-fold distributions, add a statistical comparison (e.g., permutation test or effect sizes) for the GITA results, and explicitly discuss how the inconsistent behavior in the other four corpora bears on the claim of robustness.
minor comments (5)
  1. [Section IV-A] There is a typo in the first sentence of Section IV-A: 'black-blox' should be 'black-box'.
  2. [Section IV-B] In Section IV-B, the word 'intepretability' appears in the sentence beginning 'Our first experiments utilized the full range...'; it should be 'interpretability'.
  3. [Figures 2 and 3] Figures 2 and 3 are extremely dense, with axis labels and tick labels too small to read at normal print size; the figures should be restructured or rendered with larger fonts, and error bars should be added if the displayed values are averages over samples.
  4. [Section VI-C] Figures 4 and 5 present single-subject temporal analyses; the text should more explicitly frame these as illustrative case studies rather than generalizable group findings, given the acknowledged heterogeneity across PD subjects.
  5. [Section V-D] The hyperparameter search is described only as 'preliminary,' with the final learning rate and number of epochs given; please report the search range or the values that were tried, to support reproducibility.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the interpretability outputs are post-hoc attention readings, not predictions fitted from the same quantities.

full rationale

The proposed cross-attention framework is self-contained against external benchmarks: the model is trained with a cross-entropy objective on five PD corpora, and the interpretability analyses are post-hoc readings of learned attention matrices S_emb and S_temp. The attention scores are computed by the paper's Eq. (2) as softmax dot products between linearly projected SSL embeddings and the 35 informed features (WK=I); they are outputs of the model, not fitted parameters renamed as predictions. The claim that attention aligns with expected speech dimensions (Section VI-B) is an empirical interpretation of those outputs, not a derivation whose conclusion is identical to its inputs by construction. The informed features are used as keys by design, so the attention distribution is over those features; however, the specific pattern (e.g., higher attention to logE or formants in VOWELS, to prosodic features in READ-TEXT and MONOLOGUE) is learned, not forced. Self-citations (Botelho et al. [80], [82]) are present but only motivate a contrastive normalization and cross-dataset caveats; they are not load-bearing proofs or uniqueness theorems. The stronger epistemic weaknesses of the paper—selection of the best seed and only correctly-predicted test samples, absence of a null model for attention weights, and the unresolved attention-as-explanation debate acknowledged in Sections III-B and VI-D—concern validity and generalizability of the interpretability claims, not circular reasoning. Thus no circular step is exhibited.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim depends on the manual selection of 35 clinical features, the choice of SSL layer 7, and the assumption that attention weights faithfully indicate clinical relevance. No new physical or latent entities are introduced; attention matrices are computed artifacts of the trained model.

free parameters (3)
  • 35-feature clinical subset of DisVoice = Hand-selected list (Section IV-B)
    The interpretability space is limited to these 35 features, chosen manually from 655 based on clinical meaningfulness; this directly constrains what the attention maps can reveal.
  • XLS-R encoder layer index = 7
    Layer 7 was chosen from a preliminary layer-wise analysis described in Section IV-A; the claim depends on this intermediate layer retaining clinically relevant information.
  • Training hyperparameters = lr=4e-4, epochs=5, batch=8
    Chosen via preliminary hyper-parameter search (Section V-D); standard values but not independently motivated.
assumptions (4)
  • domain assumption Attention weights are a valid proxy for the clinical relevance of speech features
    The entire interpretability analysis rests on this premise; the paper itself cites the debate [46]-[48] and does not validate against clinician judgments.
  • domain assumption The 35 selected informed features cover the clinically meaningful speech dimensions for PD
    If the feature set omits relevant dimensions, the attention maps cannot reveal them; selection was manual based on literature and prior biomarker studies.
  • domain assumption XLS-R embeddings extracted from layer 7 without fine-tuning retain information about pathological speech
    The method relies on off-the-shelf SSL embeddings; the layer choice was made via preliminary experiments, and the paper cites prior work suggesting intermediate layers are informative.
  • standard math Scaled dot-product attention (Eq. 1) provides meaningful alignment between SSL embeddings and informed features
    The attention formulation follows Vaswani et al.; it is standard, but its faithfulness as an explanatory tool is the debated component.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Unveiling Interpretability in Self-Supervised Speech Representations for Parkinson's Diagnosis." pith.science (2026). https://pith.science/paper/HWFMXPDJ

@misc{pith2026241202006,
  author       = {Pith},
  title        = {Pith review of: Unveiling Interpretability in Self-Supervised Speech Representations for Parkinson's Diagnosis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HWFMXPDJ}},
  note         = {Machine review of arXiv:2412.02006}
}
read the original abstract

Recent works in pathological speech analysis have increasingly relied on powerful self-supervised speech representations, leading to promising results. However, the complex, black-box nature of these embeddings and the limited research on their interpretability significantly restrict their adoption for clinical diagnosis. To address this gap, we propose a novel, interpretable framework specifically designed to support Parkinson's Disease (PD) diagnosis. Through the design of simple yet effective cross-attention mechanisms for both embedding- and temporal-level analysis, the proposed framework offers interpretability from two distinct but complementary perspectives. Experimental findings across five well-established speech benchmarks for PD detection demonstrate the framework's capability to identify meaningful speech patterns within self-supervised representations for a wide range of assessment tasks. Fine-grained temporal analyses further underscore its potential to enhance the interpretability of deep-learning pathological speech models, paving the way for the development of more transparent, trustworthy, and clinically applicable computer-assisted diagnosis systems in this domain. Moreover, in terms of classification accuracy, our method achieves results competitive with state-of-the-art approaches, while also demonstrating robustness in cross-lingual scenarios when applied to spontaneous speech production.

Figures

Figures reproduced from arXiv: 2412.02006 by the authors.

Figure 1
Figure 1. The overall architecture of our proposed framework for PD diagnosis support, as well as the motivations behind each interpretable module design. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Attention-based relevance scores from the embedding interpretability perspective. For each assessment tasks considered in our study, the figure presents [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Embedding-level cross-attention alignment showing the difference be [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Temporal Contrastive Analysis of the GITA subject no. 15, diagnosed with Parkinson’s Disease, during the SENTENCES tasks. The analysis focuses [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Temporal Contrastive Analysis of the GITA subject no. 3, diagnosed with Parkinson’s Disease, during the SENTENCES tasks. The analysis focuses [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition

    cs.SD 2025-05 conditional novelty 5.0 of 10

    Combining x-vector speaker conditioning, AdaLoRA adapters, and Parler-TTS synthetic speech reduces word error rates for dysarthric ASR on the SAP development set.

  2. Beyond Manual Transcripts: The Potential of Automated Speech Recognition Errors in Improving Alzheimer's Disease Detection

    eess.AS 2025-05 conditional novelty 4.0 of 10

    Certain ASR transcripts and speech synthesized from them outperform manual transcripts in Alzheimer's disease detection, suggesting ASR errors can serve as useful diagnostic cues.

Reference graph

Works this paper leans on

89 extracted references · 78 canonical work pages · cited by 2 Pith papers

  1. [46]

    Attention Is Not Explanation,

    S. Jain and B. C. Wallace, “Attention Is Not Explanation,” in Proceed- ings of NCACL . Association for Computational Linguistics, 2019, pp. 3543–3556

  2. [48]

    Is Attention Explanation? An Introduction to the Debate,

    A. Bibal, R. Cardon, D. Alfter, R. Wilkens, X. Wang, T. Franc ¸ois, and P. Watrin, “Is Attention Explanation? An Introduction to the Debate,” in Proc. of the 60th ACL . Association for Computational Linguistics, 2022, pp. 3889–3900

  3. [1]

    Alzheimer’s Disease and Parkinson’s Disease,

    R. L. Nussbaum and C. E. Ellis, “Alzheimer’s Disease and Parkinson’s Disease,” New England Journal of Medicine, vol. 348, no. 14, pp. 1356– 1364, 2003

  4. [2]

    Temporal trends in the prevalence of Parkinson’s disease from 1980 to 2023: a systematic review and meta-analysis,

    J. Zhu, Y . Cui, J. Zhang, R. Yan, D. Su, D. Zhao, A. Wang, and T. Feng, “Temporal trends in the prevalence of Parkinson’s disease from 1980 to 2023: a systematic review and meta-analysis,” The Lancet Healthy Longevity, vol. 5, no. 7, pp. e464–e479, 2024

  5. [3]

    Biochemical Aspects of Parkinson’s Disease,

    O. Hornykiewicz, “Biochemical Aspects of Parkinson’s Disease,” Neu- rology, vol. 51, no. 2 suppl 2, pp. S2–S9, 1998

  6. [4]

    Cognitive Decline in Parkinson Disease,

    D. Aarsland, B. Creese, M. Politis, K. R. Chaudhuri, D. H. Ffytche, D. Weintraub, and C. Ballard, “Cognitive Decline in Parkinson Disease,” Nature Reviews Neurology, vol. 13, no. 4, pp. 217–231, 2017

  7. [5]

    Parkinson’s Disease: Clinical Features and Diagnosis,

    J. Jankovic, “Parkinson’s Disease: Clinical Features and Diagnosis,” Journal of Neurology, Neurosurgery & Psychiatry , vol. 79, no. 4, pp. 368–376, 2008

  8. [6]

    Parkinson Disease,

    W. Poewe, K. Seppi, C. M. Tanner, G. M. Halliday, P. Brundin, J. V olkmann, A.-E. Schrag, and A. E. Lang, “Parkinson Disease,”Nature Reviews Disease Primers , vol. 3, no. 1, pp. 1–21, 2017

Show all 89 references
  1. [7]

    Frequency and Cooccurrence of V ocal Tract Dysfunctions in the Speech of a Large Sample of Parkinson Patients,

    J. A. Logemann, H. B. Fisher, B. Boshes, and E. R. Blonsky, “Frequency and Cooccurrence of V ocal Tract Dysfunctions in the Speech of a Large Sample of Parkinson Patients,” Journal of Speech and Hearing Disorders, vol. 43, no. 1, pp. 47–57, 1978

  2. [8]

    Speech and Swallowing Symptoms Associated with Parkinson’s Disease and Multiple Sclerosis: A Survey,

    L. Hartelius and P. Svensson, “Speech and Swallowing Symptoms Associated with Parkinson’s Disease and Multiple Sclerosis: A Survey,” Folia Phoniatrica et Logopaedica , vol. 46, no. 1, pp. 9–17, 1994

  3. [9]

    Prevalence and Pattern of Perceived Intelligibility Changes in Parkinson’s Disease,

    N. Miller, L. Allcock, D. Jones, E. Noble, A. J. Hildreth, and D. J. Burn, “Prevalence and Pattern of Perceived Intelligibility Changes in Parkinson’s Disease,” Journal of Neurology, Neurosurgery & Psychiatry, vol. 78, no. 11, pp. 1188–1190, 2007

  4. [10]

    Speech Characteristics of Patients with Parkinson’s Dis- ease: I. Intensity, Pitch, and Duration,

    G. J. Canter, “Speech Characteristics of Patients with Parkinson’s Dis- ease: I. Intensity, Pitch, and Duration,” Journal of Speech and Hearing Disorders, vol. 28, no. 3, pp. 221–229, 1963

  5. [11]

    Speech characteristics of patients with parkinson’s disease: Iii. articulation, diadochokinesis, and over-all speech adequacy,

    ——, “Speech characteristics of patients with parkinson’s disease: Iii. articulation, diadochokinesis, and over-all speech adequacy,” Journal of Speech and Hearing Disorders , vol. 30, no. 3, pp. 217–224, 1965

  6. [12]

    Clusters of Deviant Speech Dimensions in the Dysarthrias,

    F. L. Darley, A. E. Aronson, and J. R. Brown, “Clusters of Deviant Speech Dimensions in the Dysarthrias,” Journal of Speech and Hearing Research, vol. 12, no. 3, pp. 462–496, 1969

  7. [13]

    J. R. Duffy, Motor Speech Disorders: Substrates, Differential Diagnosis, and Management. Elsevier Health Sciences, 2012

  8. [14]

    Cinegraphic Observations of Laryngeal Function in Parkinson’s Disease,

    D. G. Hanson, B. R. Gerratt, and P. H. Ward, “Cinegraphic Observations of Laryngeal Function in Parkinson’s Disease,” The Laryngoscope , vol. 94, no. 3, pp. 348–353, 1984

  9. [15]

    V ocal Tract Steadiness: A Measure of Phonatory and Upper Airway Motor Control During Phonation in Dysarthria,

    P. Zwirner and G. J. Barnes, “V ocal Tract Steadiness: A Measure of Phonatory and Upper Airway Motor Control During Phonation in Dysarthria,” Journal of Speech, Language, and Hearing Research , vol. 35, no. 4, pp. 761–768, 1992

  10. [16]

    Laryngeal Somatosensory Deficits in Parkinson’s Disease: Implications for Speech Respiratory and Phonatory Control,

    M. J. Hammer and S. M. Barlow, “Laryngeal Somatosensory Deficits in Parkinson’s Disease: Implications for Speech Respiratory and Phonatory Control,” Experimental Brain Research , vol. 201, pp. 401–409, 2010

  11. [17]

    V ocal Tract Control in Parkinson’s Disease,

    J. A. Logemann and H. B. Fisher, “V ocal Tract Control in Parkinson’s Disease,” Journal of Speech and Hearing Disorders , vol. 46, no. 4, pp. 348–352, 1981

  12. [18]

    Articulatory Deficits in Parkinsonian Dysarthria: An Acoustic Analysis,

    H. Ackermann and W. Ziegler, “Articulatory Deficits in Parkinsonian Dysarthria: An Acoustic Analysis,” Journal of Neurology, Neurosurgery & Psychiatry, vol. 54, no. 12, pp. 1093–1098, 1991

  13. [19]

    Speech Rate and Rhythm in Parkinson’s Disease,

    S. Skodda and U. Schlegel, “Speech Rate and Rhythm in Parkinson’s Disease,” Movement Disorders: Official Journal of the Movement Dis- order Society, vol. 23, no. 7, pp. 985–992, 2008

  14. [20]

    Prosody in Parkinson’s Disease,

    H. N. Jones, “Prosody in Parkinson’s Disease,” Perspectives on Neuro- physiology and Neurogenic Speech and Language Disorders , vol. 19, no. 3, pp. 77–82, 2009

  15. [21]

    Speech Impairment in a Large Sample of Patients with Parkinson’s Disease,

    A. K. Ho, R. Iansek, C. Marigliani, J. L. Bradshaw, and S. Gates, “Speech Impairment in a Large Sample of Patients with Parkinson’s Disease,” Behavioural Neurology, vol. 11, no. 3, pp. 131–137, 1999

  16. [22]

    Speech and Communi- cation Changes Reported by People with Parkinson’s Disease,

    E. Schalling, K. Johansson, and L. Hartelius, “Speech and Communi- cation Changes Reported by People with Parkinson’s Disease,” Folia Phoniatrica et Logopaedica , vol. 69, no. 3, pp. 131–141, 2018

  17. [23]

    Life With Communication Changes in Parkinson’s Disease,

    N. Miller, E. Noble, D. Jones, and D. Burn, “Life With Communication Changes in Parkinson’s Disease,” Age and Ageing , vol. 35, no. 3, pp. 235–239, 2006

  18. [24]

    Evaluation of Parkinson’s Disease: Reliability of Three Rating Scales,

    A. Ginanneschi, F. Degl’Innocenti, S. Magnolfi, M. T. Maurello, L. Catarzi, P. Marini, and L. Amaducci, “Evaluation of Parkinson’s Disease: Reliability of Three Rating Scales,” Neuroepidemiology, vol. 7, no. 1, pp. 38–41, 1988

  19. [25]

    Interrater Reliability of the Unified Parkinson’s Disease Rating Scale Motor Examination,

    M. Richards, K. Marder, L. Cote, and R. Mayeux, “Interrater Reliability of the Unified Parkinson’s Disease Rating Scale Motor Examination,” Movement Disorders, vol. 9, no. 1, pp. 89–91, 1994

  20. [26]

    Speech Treatment for Parkinson’s Disease,

    C. F. L. O. Ramig and S. Sapir, “Speech Treatment for Parkinson’s Disease,” Expert Review of Neurotherapeutics , vol. 8, no. 2, pp. 297– 309, 2008

  21. [27]

    Towards an Automatic Evaluation of the Dysarthria Level of Patients with Parkinson’s Disease,

    J. C. V ´asquez-Correa, J. R. Orozco-Arroyave, T. Bocklet, and E. N ¨oth, “Towards an Automatic Evaluation of the Dysarthria Level of Patients with Parkinson’s Disease,” Journal of communication disorders, vol. 76, pp. 21–36, 2018

  22. [28]

    Advances in Parkinson’s Disease Detection and Assessment using V oice and Speech: A Review of the Articulatory and Phonatory Aspects,

    L. Moro-Velazquez, J. A. Gomez-Garcia, J. D. Arias-Londo ˜no, N. De- hak, and J. I. Godino-Llorente, “Advances in Parkinson’s Disease Detection and Assessment using V oice and Speech: A Review of the Articulatory and Phonatory Aspects,” Biomedical Signal Processing and Control...

  23. [29]

    From Prodromal Stages to Clin- ical Trials: The Promise of Digital Speech Biomarkers in Parkinson’s Disease,

    J. Rusz, P. Krack, and E. Tripoliti, “From Prodromal Stages to Clin- ical Trials: The Promise of Digital Speech Biomarkers in Parkinson’s Disease,” Neuroscience & Biobehavioral Reviews , p. 105922, 2024

  24. [30]

    DisV oice,

    J. C. V ´asquez-Correa, “DisV oice,” https://github.com/charlespwd/ project-title, 2013

  25. [31]

    The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for V oice Research and Affective Computing,

    F. Eyben, K. R. Scherer, B. W. Schuller, J. Sundberg, E. Andr´e, C. Busso, L. Y . Devillers, J. Epps, P. Laukka, S. S. Narayanan, and K. P. Truong, “The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for V oice Research and Affective Computing,” IEEE Transactions on Affec...

  26. [32]

    XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale,

    A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y . Saraf, J. Pino, A. Baevski, A. Conneau, and M. Auli, “XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale,” in Proc. of Interspeech , 2022, pp. 2278–2282. JOURNAL ...

  27. [33]

    TRILLsson: Distilled Universal Paralin- guistic Speech Representations,

    J. Shor and S. Venugopalan, “TRILLsson: Distilled Universal Paralin- guistic Speech Representations,” in Proc. of Interspeech, 2022, pp. 356– 360

  28. [34]

    HuBERT: Self-Supervised Speech Representation Learn- ing by Masked Prediction of Hidden Units,

    W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-Supervised Speech Representation Learn- ing by Masked Prediction of Hidden Units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3451–3460, 2021

  29. [35]

    Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0,

    S. P. Bayerl, D. Wagner, E. Noeth, and K. Riedhammer, “Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0,” in Proc. of Interspeech, 2022, pp. 2868–2872

  30. [36]

    Multi-Class Detection of Pathological Speech with Latent Features: How does it Perform on Unseen Data?

    D. Wagner, I. Baumann, F. Braun, S. P. Bayerl, E. N ¨oth, K. Riedhammer, and T. Bocklet, “Multi-Class Detection of Pathological Speech with Latent Features: How does it Perform on Unseen Data?” in Proc. of Interspeech, 2023, pp. 2318–2322

  31. [37]

    Towards Interpretability of Automatic Phoneme Analysis in Cleft Lip and Palate Speech,

    I. Baumann, D. Wagner, M. Schuster, E. N ¨oth, and T. Bocklet, “Towards Interpretability of Automatic Phoneme Analysis in Cleft Lip and Palate Speech,” in Proc. of ICASSP , 2024, pp. 12 602–12 606

  32. [38]

    Interpretable speech features vs. DNN embed- dings: What to use in the automatic assessment of Parkinson’s disease in multi-lingual scenarios,

    A. Favaro, Y .-T. Tsai, A. Butala, T. Thebaud, J. Villalba, N. Dehak, and L. Moro-Vel´azquez, “Interpretable speech features vs. DNN embed- dings: What to use in the automatic assessment of Parkinson’s disease in multi-lingual scenarios,” Computers in Biology and Medicine , vo...

  33. [39]

    Exploiting Foundation Models and Speech Enhancement for Parkinson’s Disease Detection from Speech in Real-World Operative Conditions,

    M. La Quatra, M. F. Turco, T. Svendsen, G. Salvi, J. R. Orozco- Arroyave, and S. M. Siniscalchi, “Exploiting Foundation Models and Speech Enhancement for Parkinson’s Disease Detection from Speech in Real-World Operative Conditions,” in Proc. of Interspeech , 2024, pp. 1405–1409

  34. [40]

    Automatic Classification of Parkinson’s Disease Using Wav2Vec Embeddings at Phoneme, Syllable, and Word Levels,

    J. D. Gallo-Aristiz ´abal, D. Escobar-Grisales, C. D. R´ıos-Urrego, E. N¨oth, and J. R. Orozco-Arroyave, “Automatic Classification of Parkinson’s Disease Using Wav2Vec Embeddings at Phoneme, Syllable, and Word Levels,” in Text, Speech, and Dialogue . Springer Nature Switzerlan...

  35. [41]

    On the Use of a Foundation Acoustic Model to Identify Highly Relevant Pho- netic Information of Parkinson’s Speech,

    D. Escobar-Grisales, C. R ´ıos-Urrego, and J. Orozco-Arroyave, “On the Use of a Foundation Acoustic Model to Identify Highly Relevant Pho- netic Information of Parkinson’s Speech,” in Workshop on Engineering Applications. Springer, 2024, pp. 71–81

  36. [42]

    Explainable, Trustworthy, and Ethical Machine Learning for Healthcare: A Survey,

    K. Rasheed, A. Qayyum, M. Ghaly, A. Al-Fuqaha, A. Razi, and J. Qadir, “Explainable, Trustworthy, and Ethical Machine Learning for Healthcare: A Survey,” Computers in Biology and Medicine , vol. 149, p. 106043, 2022

  37. [43]

    Layer-Wise Analysis of a Self- Supervised Speech Representation Model,

    A. Pasad, J.-C. Chou, and K. Livescu, “Layer-Wise Analysis of a Self- Supervised Speech Representation Model,” in Proc. of IEEE ASRU , 2021, pp. 914–921

  38. [44]

    Comparative Layer-Wise Analysis of Self-Supervised Speech Models,

    A. Pasad, B. Shi, and K. Livescu, “Comparative Layer-Wise Analysis of Self-Supervised Speech Models,” in Proc. of ICASSP . IEEE, 2023, pp. 1–5

  39. [45]

    What Do Audio Trans- formers Hear? Probing Their Representations For Language Delivery & Structure,

    Y . K. Singla, J. Shah, C. Chen, and R. R. Shah, “What Do Audio Trans- formers Hear? Probing Their Representations For Language Delivery & Structure,” in Proc. of ICDMW , 2022, pp. 910–925

  40. [47]

    Attention is not not Explanation,

    S. Wiegreffe and Y . Pinter, “Attention is not not Explanation,” in Proc. of EMNLP-IJCNLP. Association for Computational Linguistics, 2019, pp. 11–20

  41. [49]

    Measuring Phonological Precision in Children with Cleft Lip and Palate,

    T. Arias-Vergara, E. Londo ˜no-Mora, P. A. P ´erez-Toro, M. Schuster, E. N¨oth, J. R. Orozco-Arroyave, and A. Maier, “Measuring Phonological Precision in Children with Cleft Lip and Palate,” in Interspeech, 2023, pp. 4638–4642

  42. [50]

    Towards Self-Attention Understanding for Automatic Articulatory Processes Analysis in Cleft Lip and Palate Speech,

    I. Baumann, D. Wagner, M. Schuster, K. Riedhammer, E. Noeth, and T. Bocklet, “Towards Self-Attention Understanding for Automatic Articulatory Processes Analysis in Cleft Lip and Palate Speech,” inProc. of Interspeech, 2024, pp. 2430–2434

  43. [51]

    NeuroSpeech: An Open-Source Software for Parkinson’s Speech Analysis,

    J. R. Orozco-Arroyave, J. C. V ´asquez-Correa, J. F. Vargas-Bonilla, R. Arora, N. Dehak, P. Nidadavolu, H. Christensen, F. Rudzicz, M. Yancheva, H. Chinaei, A. Vann, N. V ogler, T. Bocklet, M. Cernak, J. Hannink, and E. N ¨oth, “NeuroSpeech: An Open-Source Software for Parkins...

  44. [52]

    Interpreting Deep Representations of Phonetic Features via Neuro-Based Concept Detector: Application to Speech Disorders Due to Head and Neck Cancer,

    S. Abderrazek, C. Fredouille, A. Ghio, M. Lalain, C. Meunier, and V . Woisard, “Interpreting Deep Representations of Phonetic Features via Neuro-Based Concept Detector: Application to Speech Disorders Due to Head and Neck Cancer,” IEEE/ACM Transactions on Audio, Speech, and La...

  45. [53]

    Attention Is All You Need,

    Vaswani, A. and Shazeer, N. and Parmar, N. and Uszkoreit, J. and Jones, L. and Gomez, A. N and Kaiser, L. and Polosukhin, I., “Attention Is All You Need,” Proc. of NeurIPS , vol. 30, pp. 6000–6010, 2017

  46. [54]

    A Comprehensive Survey on Applications of Transformers for Deep Learning Tasks,

    S. Islam, H. Elmekki, A. Elsebai, J. Bentahar, N. Drawel, G. Rjoub, and W. Pedrycz, “A Comprehensive Survey on Applications of Transformers for Deep Learning Tasks,” Expert Systems with Applications , vol. 241, p. 122666, 2024

  47. [55]

    Layer Normalization,

    J. L. Ba, J. R. Kiros, and G. Hinton, “Layer Normalization,” arXiv preprint arXiv:1607.06450, 2016

  48. [56]

    Searching for Activation Functions,

    P. Ramachandran, B. Zoph, and Q. V . Le, “Searching for Activation Functions,” arXiv preprint arXiv:1710.05941 , 2017

  49. [57]

    Automatic Parkinson’s disease detection from speech: Layer selection vs adaptation of foundation models,

    T. Purohit, B. Ruvolo, J. R. Orozco-Arroyave, and M. M. Doss, “Automatic Parkinson’s disease detection from speech: Layer selection vs adaptation of foundation models,” in ICASSP, 2025

  50. [58]

    Which aspects of motor speech disorder are captured by mel frequency cepstral coefficients? evidence from the change in STN-DBS conditions in parkinson’s disease,

    V . Illner, P. Kr `yˇze, J. ˇSvihl´ık, M. Sousa, P. Krack, E. Tripoliti, R. Jech, and J. Rusz, “Which aspects of motor speech disorder are captured by mel frequency cepstral coefficients? evidence from the change in STN-DBS conditions in parkinson’s disease,” in Proc. of Inter...

  51. [59]

    Towards interpretable speech biomarkers: exploring mfccs,

    B. Tracey, D. V olfson, J. Glass, R. Haulcy, M. Kostrzebski, J. Adams, T. Kangarloo, A. Brodtmann, E. R. Dorsey, and A. V ogel, “Towards interpretable speech biomarkers: exploring mfccs,” Scientific Reports , vol. 13, no. 1, p. 22787, 2023

  52. [60]

    Unveiling Early Signs of Parkinson’s Disease via a Lon- gitudinal Analysis of Celebrity Speech Recordings,

    A. Favaro, A. Butala, T. Thebaud, J. Villalba, N. Dehak, and L. Moro- Vel´azquez, “Unveiling Early Signs of Parkinson’s Disease via a Lon- gitudinal Analysis of Celebrity Speech Recordings,” npj Parkinson’s Disease, vol. 10, no. 1, p. 207, 2024

  53. [61]

    Glottal Flow Patterns Analyses for Parkin- son’s Disease Detection: Acoustic and Nonlinear Approaches,

    E. A. Belalc ´azar-Bola˜nos, J. R. Orozco-Arroyave, J. F. Vargas-Bonilla, T. Haderlein, and E. N ¨oth, “Glottal Flow Patterns Analyses for Parkin- son’s Disease Detection: Acoustic and Nonlinear Approaches,” in Text, Speech, and Dialogue, P. Sojka, A. Hor´ak, I. Kopeˇcek, and ...

  54. [62]

    Parkinson’s Disease and Aging: Analysis of Their Effect in Phonation and Articulation of Speech,

    T. Arias-Vergara, J. C. V ´asquez-Correa, and J. R. Orozco-Arroyave, “Parkinson’s Disease and Aging: Analysis of Their Effect in Phonation and Articulation of Speech,” Cognitive Computation, vol. 9, no. 6, pp. 731–748, 2017

  55. [63]

    V oice changes in Parkinson’s disease: What are they telling us?

    A. Ma, K. K. Lau, and D. Thyagarajan, “V oice changes in Parkinson’s disease: What are they telling us?” Journal of Clinical Neuroscience , vol. 72, pp. 1–7, 2020

  56. [64]

    Modeling Prosodic Features With Joint Factor Analysis for Speaker Verification,

    N. Dehak, P. Dumouchel, and P. Kenny, “Modeling Prosodic Features With Joint Factor Analysis for Speaker Verification,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 7, pp. 2095– 2103, 2007

  57. [65]

    Detection of Persons with Parkinson’s Disease by Acoustic, V ocal, and Prosodic Analysis,

    T. Bocklet, E. N ¨oth, G. Stemmer, H. Ruzickova, and J. Rusz, “Detection of Persons with Parkinson’s Disease by Acoustic, V ocal, and Prosodic Analysis,” in IEEE Workshop on Automatic Speech Recognition & Understanding, 2011, pp. 478–483

  58. [66]

    V oiced/Unvoiced Transitions in Speech as a Potential Bio-Marker to Detect Parkinson’s Disease,

    J. R. Orozco-Arroyave, F. H ¨onig, J. D. Arias-Londo ˜no, J. F. Vargas- Bonilla, S. Skodda, J. Rusz, and E. N ¨oth, “V oiced/Unvoiced Transitions in Speech as a Potential Bio-Marker to Detect Parkinson’s Disease,” in Interspeech, 2015, pp. 95–99

  59. [67]

    Speech rate in Parkinson’s disease: a controlled study,

    F. Mart ´ınez-S´anchez, J. Meil ´an, J. Carro, C. G. ´I˜niguez, L. Millian- Morell, I. P. Valverde, T. L ´opez-Alburquerque, and D. L ´opez, “Speech rate in Parkinson’s disease: a controlled study,” Neurolog´ıa (English Edition), vol. 31, no. 7, pp. 466–472, 2016

  60. [68]

    NeuroV oz: A Castillian Spanish Corpus of Parkinsonian Speech,

    J. Mendes-Laureano, J. A. G ´omez-Garc´ıa, A. Guerrero-L´opez, E. Luque- Buzo, J. D. Arias-Londo ˜no, F. J. Grandas-P ´erez, and J. I. Godino- Llorente, “NeuroV oz: A Castillian Spanish Corpus of Parkinsonian Speech,” Scientific Data, vol. 11, no. 1, p. 1367, 2024

  61. [69]

    New Spanish Speech Corpus Database for the Analysis of People Suffering from Parkinson’s Disease

    J. R. Orozco-Arroyave, J. D. Arias-Londo ˜no, J. F. Vargas-Bonilla, M. C. Gonzalez-R´ativa, and E. N ¨oth, “New Spanish Speech Corpus Database for the Analysis of People Suffering from Parkinson’s Disease.” in Proc. of LREC, 2014, pp. 342–347

  62. [70]

    Dysarthria in individuals with parkinson’s disease: A protocol for a binational, cross-sectional, case-controlled study in french and european portuguese (fralusopark),

    S. Pinto, R. Cardoso, J. Sadat, I. Guimar ˜aes, C. Mercier, H. Santos, C. Atkinson-Clement, J. Carvalho, P. Welby, P. Oliveira, M. D’Imperio, S. Frota, A. Letanneux, M. Vigario, M. Cruz, I. P. Martins, F. Viallet, and J. J. Ferreira, “Dysarthria in individuals with parkinson’s...

  63. [71]

    Automatic evaluation of parkinson’s speech — acoustic, prosodic and voice related cues,

    T. Bocklet, S. Steidl, E. N ¨oth, and S. Skodda, “Automatic evaluation of parkinson’s speech — acoustic, prosodic and voice related cues,” in Proc. of Interspeech , 2013, pp. 1149–1153

  64. [72]

    C. D. Rios-Urrego, J. Rusz, and J. R. Orozco-Arroyave, “Automatic speech-based assessment to discriminate Parkinson’s disease from es- JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, VOL. XX, NO. YY , NOV 2024 14 sential tremor with a cross-language approach,” NPJ Digital Med...

  65. [73]

    Movement Disorder Society-Sponsored Revision of the Unified Parkinson’s Dis- ease Rating Scale (MDS-UPDRS): Scale Presentation and Clinimetric Testing Results,

    C. G. Goetz, B. C. Tilley, S. R. Shaftman, G. T. Stebbins, S. Fahn, P. Martinez-Martin, W. Poewe, C. Sampaio, M. B. Stern, R. Dodel, B. Dubois, R. Holloway, J. Jankovic, J. Kulisevsky, A. E. Lang, A. Lees, S. Leurgans, P. A. LeWitt, D. Nyenhuis, C. W. Olanow, O. Rascol, A. Sch...

  66. [74]

    C. G. Goetz, W. Poewe, O. Rascol, C. Sampaio, G. T. Stebbins, C. Counsell, N. Giladi, R. G. Holloway, C. G. Moore, G. K. Wenning, M. D. Yahr, and L. Seidl, “Movement Disorder Society Task Force Report on the Hoehn and Yahr Staging Scale: Status and Recommen- dations the Moveme...

  67. [75]

    Exploring Digital Speech Biomarkers of Hypokinetic Dysarthria in a Multilingual Cohort,

    D. Kovac, J. Mekyska, V . Aharonson, P. Harar, Z. Galaz, S. Rapcsak, J. R. Orozco-Arroyave, L. Brabenec, and I. Rektorova, “Exploring Digital Speech Biomarkers of Hypokinetic Dysarthria in a Multilingual Cohort,” Biomedical Signal Processing and Control, vol. 88, p. 105667, 2024

  68. [76]

    Perceptual Characteristics of Parkinsonian Speech: A Comparison of the Pharmacological Effects of Levodopa Across Speech and Non-Speech Motor Systems,

    E. K. Plowman-Prine, M. S. Okun, C. M. Sapienza, R. Shrivastav, H. H. Fernandez, K. D. Foote, C. Ellis, A. D. Rodriguez, L. M. Burkhead, and J. C. Rosenbek, “Perceptual Characteristics of Parkinsonian Speech: A Comparison of the Pharmacological Effects of Levodopa Across Speec...

  69. [77]

    Intonation and Speech Rate in Parkinson’s Disease: General and Dynamic Aspects and Responsiveness to Levodopa Admission,

    S. Skodda, W. Gr ¨onheit, and U. Schlegel, “Intonation and Speech Rate in Parkinson’s Disease: General and Dynamic Aspects and Responsiveness to Levodopa Admission,”Journal of Voice, vol. 25, no. 4, pp. e199–e205, 2011

  70. [78]

    Towards a Corpus (and Language)-Independent Screening of Parkin- son’s Disease from V oice and Speech through Domain Adaptation,

    E. J. Ibarra, J. D. Arias-Londo ˜no, M. Za˜nartu, and J. I. Godino-Llorente, “Towards a Corpus (and Language)-Independent Screening of Parkin- son’s Disease from V oice and Speech through Domain Adaptation,” Bioengineering, vol. 10, no. 11, p. 1316, 2023

  71. [79]

    Domain Adversarial Convolutional Neural Network for Parkinson’s Disease Detection from Speech,

    E. Ibarra-Sulbaran, J. Arias-Londo ˜no, M. Za ˜nartu, and J. Godino- Llorente, “Domain Adversarial Convolutional Neural Network for Parkinson’s Disease Detection from Speech,” inProc. of MAVEBA, 2023, pp. 69–72

  72. [80]

    Challenges of Using Longitudinal and Cross-Domain Corpora on Studies of Pathological Speech,

    C. Botelho, T. Schultz, A. Abad, and I. Trancoso, “Challenges of Using Longitudinal and Cross-Domain Corpora on Studies of Pathological Speech,” in Proc. of Interspeech , 2022, pp. 1921–1925

  73. [81]

    Parkinson’s Disease and Aging: Analysis of their Effect in Phonation and Articulation of Speech,

    T. Arias-Vergara, J. C. V ´asquez-Correa, and J. R. Orozco-Arroyave, “Parkinson’s Disease and Aging: Analysis of their Effect in Phonation and Articulation of Speech,” Cognitive Computation , vol. 9, pp. 731– 748, 2017

  74. [82]

    Towards Reference Speech Characterization for Health Applications,

    C. Botelho and A. Abad and T. Schultz and I. Trancoso, “Towards Reference Speech Characterization for Health Applications,” in Proc. of Interspeech, 2023, pp. 2363–2367

  75. [83]

    Dynamic Time Warping,

    M. M ¨uller, “Dynamic Time Warping,” Information Retrieval for Music and Motion, pp. 69–84, 2007

  76. [84]

    Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi,

    M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi,” in Proc. Interspeech, 2017, pp. 498–502

  77. [85]

    Macro-Descriptors for Alzheimer’s Disease Detection Using Large Language Models,

    C. Botelho, J. Mendonc ¸a, A. Pompili, T. Schultz, A. Abad, and I. Tran- coso, “Macro-Descriptors for Alzheimer’s Disease Detection Using Large Language Models,” in Proc. of Interspeech, 2024, pp. 1975–1979

  78. [86]

    When Whisper Listens to Apha- sia: Advancing Robust Post-Stroke Speech Recognition,

    Giulia Sanguedolce and Sophie Brook and Dragos C. Gruia and Patrick A. Naylor and Fatemeh Geranmayeh, “When Whisper Listens to Apha- sia: Advancing Robust Post-Stroke Speech Recognition,” in Proc. of Interspeech, 2024, pp. 1995–1999

  79. [87]

    Reading Between the Frames: Multi-modal Depression Detection in Videos from Non-verbal Cues,

    D. Gimeno-G ´omez, A.-M. Bucur, A. Cosma, C.-D. Mart ´ınez-Hinarejos, and P. Rosso, “Reading Between the Frames: Multi-modal Depression Detection in Videos from Non-verbal Cues,” in Advances in Information Retrieval (ECIR). Springer Nature Switzerland, 2024, pp. 191–209. David...

  80. [2018]

    in Electrical and Computer Engineering at IST and INESC-ID in

    She completed her Ph.D. in Electrical and Computer Engineering at IST and INESC-ID in

  81. [2024]

    She has held positions as a research intern at Google AI, Toronto, and as a visitor researcher at the Cognitive Systems Lab, University of Bremen

    Currently, she is a researcher at INESC-ID, contributing to the Accelerat.AI project. She has held positions as a research intern at Google AI, Toronto, and as a visitor researcher at the Cognitive Systems Lab, University of Bremen. She was involved in the student advisory com...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.