Pith. sign in

REVIEW 3 major objections 5 minor 47 references

Affect Models Have Weak Generalizability to Atypical Speech

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read AI emotion models misread neutral atypical speech as sad

desk verdict A large, mostly convincing fairness evaluation showing affect models misclassify atypical speech as sad; the main caveat is that the read-speech neutrality assumption is asserted rather than measured, and the fine-tuning result is partly circular. read the letter →

arxiv 2504.16283 v2 pith:OHS7PGH6 submitted 2025-04-22 cs.LG

classification cs.LG
keywords speechemotionrecognitionatypicaldysarthriamodelfairnessrobustnessintelligibilitymonopitchharshness
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that state-of-the-art speech affect models do not generalize to atypical speech, treating acoustic atypicality as if it were emotional content. It compares categorical and dimensional emotion predictions on the Speech Accessibility Project dataset against matched typical-speech datasets. On read sentences and digital commands expected to be affectively neutral, all four categorical models predict neutral far less often and sad far more often for atypical speech, and the gap widens as rated intelligibility, harshness, and monopitch become more severe. The paper also shows that fine-tuning on pseudo-labeled atypical speech improves valence prediction for atypical speakers without hurting typical-speech performance. The result matters because voice-based affect systems are already used in wellbeing, coaching, and assistive applications, where systematic misreading of atypical voices could produce unfair or harmful outcomes.

What carries the argument

The central object is the distribution of categorical emotion predictions over binarized atypicality groups, read against matched typical-speech baselines with the same expected emotional content. The load-bearing identity is the within- and between-dataset comparison: neutral read speech should produce similar neutral rates across datasets, so the large drop in neutral and rise in sad for SAP speech isolates the contribution of acoustic atypicality. The analysis also uses pseudo-label correlation, comparing text-only GPT-4o valence and arousal ratings against acoustic Odyssey predictions, and fine-tuning on pseudo-labeled atypical speech to test whether the gap can be reduced.

What would settle it

Have independent annotators label the affective content of a sample of SAP read sentences and digital commands without knowing the model outputs. If their neutral rate is comparable to Common Voice and VCVA, the sad shift is confirmed as an acoustic confound; if they also hear sadness or emotional valence in the atypical speech, the distributional gap is partly content-driven rather than a pure generalization failure.

Watch

Extended reading notes

Core claim

The central discovery is a systematic distributional shift in affect model outputs, not a small accuracy dip. For affectively neutral read material, the Odyssey model predicts neutral for only 5% of less intelligible SAP read sentences and sad for 82%, versus 90% neutral and 4% sad on Common Voice; similar but smaller shifts appear across Emotion2Vec, SpeechBrain, and GPT-4o-audio-preview, and across harshness and monopitch dimensions and digital commands. Within SAP, the more atypical binarized groups consistently receive less neutral and more sad output than the less atypical groups, so the effect tracks the degree of atypicality. The authors interpret these gaps as acoustic confounds: the models are using atypical voice properties as markers of sadness rather than detecting genuine affect in the speech.

Load-bearing premise

The load-bearing premise is that the SAP read sentences and digital commands are genuinely affectively neutral, so the high sad and low neutral predictions are errors rather than true affective content; the paper supports this with elicitation design and the authors' listening, not with independent affect ratings.

Editorial extensions

If this is right

  • Voice-enabled affect applications will systematically over-attribute sadness to atypical speakers on neutral content, which could skew downstream decisions in wellbeing, coaching, or assistant interactions.
  • The effect is continuous in atypicality: larger intelligibility, harshness, and monopitch ratings move predictions further from neutral, so thresholding or simple post-hoc corrections would need to be grade-dependent.
  • Training-data affect elicitation strategy matters: an acted-speech-trained model shifts toward happy for more atypical speech while naturalistic-data models shift toward sad, so the bias direction is not fixed across models.
  • Fine-tuning the valence model on pseudo-labeled atypical speech improved correlation on atypical speech and left typical-speech correlation essentially unchanged, indicating a low-cost path toward more robust models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If atypical acoustics reliably map onto sadness in these embeddings, any emotion-conditioned accessibility feature such as mood check-ins or stress tracking for motor-speech conditions could double-count the disability itself as negative affect; the paper does not test that application, but it follows directly.
  • The within-dataset design suggests a cheap robustness experiment not run here: adding affectively neutral atypical speech to the training mix could show whether the sad shift shrinks faster than with full affect-labeled atypical data.
  • Because arousal correlations were low even for mild atypicality while valence stayed closer to typical levels, prosodic atypicality may break arousal-related acoustic cues before categorical emotion labels do; a targeted study with synthetic voice transformations that vary pitch range could isolate which acoustic dimension drives which emotion dimension.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper evaluates publicly available speech affect models on atypical speech from the Speech Accessibility Project (SAP), comparing categorical and dimensional predictions against typical-speech datasets (Switchboard, Common Voice, VCVA). It examines three atypicality dimensions (intelligibility, harshness, monopitch) using within-SAP binarized comparisons and between-dataset comparisons, and reports that atypical speech is far more often predicted as sad and far less often as neutral, with larger gaps for more atypical speech. It also reports that correlations between text-based GPT-4o pseudo-labels and audio-based dimensional predictions are lower for atypical speech, and that fine-tuning an Odyssey valence model on SAP pseudo-labels improves Pearson correlation on SAP validation/test splits while leaving Switchboard correlation roughly unchanged. The conclusions claim weak generalizability of affect models to atypical speech and call for broader, more inclusive affect datasets.

Significance. If the main result holds, the paper addresses an important fairness and accessibility gap in speech emotion recognition: it shows on a large, multi-etiology atypical-speech dataset that categorical and dimensional affect models produce markedly different predictions for atypical speech, and that even mild atypicality shifts predictions away from neutral. The use of multiple off-the-shelf models, a publicly available dataset, binarized atypicality subgroups, and both within- and between-dataset comparisons are strengths, and the paper honestly discusses limitations. The central claim, however, depends on the assumption that SAP read sentences and digital commands are affectively neutral, and the dimensional and fine-tuning analyses rely on GPT-4o pseudo-labels whose validity is only partially established. These load-bearing points need additional support before the paper's conclusions can be accepted at face value.

major comments (3)
  1. [Section II-A, Table I] The central read-speech comparison rests on an unvalidated neutrality assumption. The manuscript states that read sentences and digital commands were 'overwhelmingly affectively neutral' based on elicitation prompts and the authors' listening, but no quantitative neutrality check is reported. Because Table I interprets the elevated sad and reduced neutral predictions for SAP read speech as evidence that atypical acoustics are confused with affect, any systematic affective difference between SAP read content and Common Voice/VCVA content would confound this conclusion. The within-SAP comparisons across atypicality levels partially control for dataset content, but they still assume no affective variation across rated levels of intelligibility, harshness, or monopitch. The Limitations section itself acknowledges that 'lexical differences correlating to the atypical speech rating ... could have influenced the GPT-4o annotations, and were not investigated,' which is the same class of confound for the read-speech analysis. I recommend adding independent listener annotations of affect for the read-speech subsets, or a validated lexical sentiment control, and reporting the neutrality check separately for each atypicality level.
  2. [Section III-B, Table II] The fine-tuning experiment is partly circular. The Odyssey valence model is fine-tuned on GPT-4o pseudo-labels for SAP training data and then evaluated on validation and test samples whose labels come from the same GPT-4o text-prompt procedure. Consequently, the reported improvement of +0.06 to +0.10 in Pearson correlation measures increased agreement with the pseudo-labeler, not necessarily with true affect. The claim that fine-tuning 'improves performance on atypical speech without impacting performance on typical speech' is not supported unless independent labels are used for evaluation, or unless the pseudo-labeler is shown to be an unbiased substitute for human labels on this population. Please re-evaluate with human-annotated affect labels, or recast the experiment as an alignment-to-pseudo-labels study.
  3. [Section II-B2, Figure 3] The arousal pseudo-labels are too weak to support the arousal claims. The authors report GPT-4o with the text prompt has CCC=0.28 and Pearson=0.44 for arousal on Emobank, yet the paper states that the model can 'effectively annotate the analyzed dimensions.' A CCC of 0.28 is generally considered poor agreement, and the observed low arousal correlations for SAP speech relative to Switchboard may reflect pseudo-label noise and lexical confounds rather than atypical-speech effects. Either restrict the dimensional analysis to valence, or provide an arousal label source with demonstrated validity.
minor comments (5)
  1. [Abstract, Section I, Section II-A] There are several typos: 'pronounciation' should be 'pronunciation', 'Dyarthria' should be 'Dysarthria', and 'psuedo-labels' should be 'pseudo-labels'.
  2. [Section III-B] The sentence 'A larger dataset ... could likely significantly model improve performance' contains a word-order error; it should read 'could likely significantly improve model performance'.
  3. [Section III-B] The personalization results report 'from 78 ± 1% to 81 ± 1% for read digital assistant commands' twice; the second instance should presumably refer to read sentences rather than commands.
  4. [Section V] The sentence 'Correlations for dimensional predictions with pseudo-labels were also lower for atypical speech than atypical speech' should read '... lower for atypical speech than for typical speech.'
  5. [Table I, Section III-A] Confidence intervals are provided for individual proportions, but differences between groups are not subjected to formal significance tests; since the main claims are comparative, adding a test or adjusted interval for the differences would strengthen the presentation.

Circularity Check

1 steps flagged · score 6.0 of 10

Fine-tuning gains are measured against the same GPT-4o pseudo-labeler used for training, making that robustness claim partly circular; the central distributional analysis is not circular.

  1. fitted input called prediction [Section II-C (Strategies for improving performance), Section II-B.2 (Dimensional Emotions), Table II]
    "We used the GPT-4o dimensional valence scores as labels to fine-tune the Odyssey valence speech model on data from the SAP dataset. ... We tabulated Pearson correlation between the predicted speech valence ratings and text valence ratings ... Higher correlations signify better performance, as they correspond to more agreement between the text-only pseudo-label and the speech-only prediction. ... Table II compares the performance of the Odyssey dimensional valence model before and after fine-tuning on the SAP training data using the GPT-4o valence pseudo-labels."

    The fine-tuning target and the evaluation target are the same GPT-4o text-only pseudo-label procedure. The model is trained with MSE loss against GPT-4o valence pseudo-labels, and the reported improvement is the Pearson correlation with GPT-4o valence pseudo-labels on held-out splits. Thus the gain measures how well the model learned to imitate the pseudo-labeler, including any systematic lexical or content biases that correlate with atypicality. The paper's own Limitations section concedes that 'lexical differences correlating to the atypical speech rating ...

full rationale

The central claim of the paper, that off-the-shelf affect models predict very different emotion distributions for atypical speech than for typical speech, is a direct empirical comparison of public model outputs and does not reduce to its inputs. Table I reports raw prediction percentages, and the between-dataset and within-dataset differences are measured, not derived from any fitted parameter. The neutrality assumption for SAP read speech is an unvalidated premise, not a circular one: the paper asserts it from the elicitation protocol and listening, and the conclusion that neutral read speech is predicted as sad follows only if that premise holds. That is a correctness risk, not a self-referential derivation. The self-citation to [42] is not load-bearing because the paper independently validates the GPT-4o dimensional prompt on Emobank. The one genuine circularity is the fine-tuning experiment: training labels and evaluation labels both come from the same GPT-4o pseudo-labeling procedure, so the improvement in Table II is partly an imitation effect rather than a demonstration of improved true affect recognition. The paper is transparent that these are pseudo-labels and acknowledges related limitations, but the specific before/after comparison is still circular in its label source. Overall, the main distributional findings are self-contained, while the robustness-improvement claim has a partial circularity, yielding a score of 6.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new physical or theoretical entities. The main load-bearing assumptions are the neutrality of read speech and the validity of GPT-4o pseudo-labels. The binarization thresholds and fine-tuning hyperparameters are chosen by hand and affect the quantitative results.

free parameters (5)
  • Binarization threshold for intelligibility = 2
    Chosen as the rating closest to the 20th percentile of the dataset to ensure sufficient sample sizes. Data-dependent and could influence the size of observed differences between 'more' and 'less' atypical groups.
  • Binarization threshold for harshness = 4
    Chosen as the rating closest to the 20th percentile. Data-dependent grouping threshold.
  • Binarization threshold for monopitch = 4
    Chosen as the rating closest to the 20th percentile. Data-dependent grouping threshold.
  • Fine-tuning learning rate = 0.0001
    Hand-chosen hyperparameter for the Adam optimizer. Affects the magnitude of the fine-tuning improvement reported in Table II.
  • Fine-tuning epochs = 50
    Hand-chosen number of training epochs. Affects the fine-tuning result.
assumptions (4)
  • domain assumption SAP read speech categories (novel sentences and digital voice commands) are affectively neutral.
    The categorical read-speech comparison interprets higher sad predictions as errors. Authors assert this based on elicitation prompts and observation (Section II-A), but no quantitative neutrality annotation is provided.
  • domain assumption GPT-4o text-based pseudo-labels for valence and arousal are valid proxies for true affect.
    Used for the dimensional analysis and as training labels for fine-tuning. The paper validates on Emobank (CCC 0.57 valence, 0.28 arousal), which is imperfect, especially for arousal.
  • domain assumption Speech-language pathologist ratings of intelligibility, harshness, and monopitch are accurate and independent of the affective content of the samples.
    These ratings define the atypicality groups. If ratings correlate with actual affect, the grouping would confound the analysis.
  • domain assumption After subsampling, the typical speech datasets (Switchboard, Common Voice, VCVA) are comparable to SAP in content, length, accent, and language, differing mainly in atypicality.
    The between-dataset comparisons rely on this comparability. The authors acknowledge remaining distributional differences in Section IV.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Affect Models Have Weak Generalizability to Atypical Speech." pith.science (2026). https://pith.science/paper/OHS7PGH6

@misc{pith2026250416283,
  author       = {Pith},
  title        = {Pith review of: Affect Models Have Weak Generalizability to Atypical Speech},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OHS7PGH6}},
  note         = {Machine review of arXiv:2504.16283}
}
read the original abstract

Speech and voice conditions can alter the acoustic properties of speech, which could impact the performance of paralinguistic models for affect for people with atypical speech. We evaluate publicly available models for recognizing categorical and dimensional affect from speech on a dataset of atypical speech, comparing results to datasets of typical speech. We investigate three dimensions of speech atypicality: intelligibility, which is related to pronounciation; monopitch, which is related to prosody, and harshness, which is related to voice quality. We look at (1) distributional trends of categorical affect predictions within the dataset, (2) distributional comparisons of categorical affect predictions to similar datasets of typical speech, and (3) correlation strengths between text and speech predictions for spontaneous speech for valence and arousal. We find that the output of affect models is significantly impacted by the presence and degree of speech atypicalities. For instance, the percentage of speech predicted as sad is significantly higher for all types and grades of atypical speech when compared to similar typical speech datasets. In a preliminary investigation on improving robustness for atypical speech, we find that fine-tuning models on pseudo-labeled atypical speech data improves performance on atypical speech without impacting performance on typical speech. Our results emphasize the need for broader training and evaluation datasets for speech emotion models, and for modeling approaches that are robust to voice and speech differences.

Figures

Figures reproduced from arXiv: 2504.16283 by the authors.

Figure 1
Figure 1. Distribution of annotations in dataset for each speech category (digital [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 3
Figure 3. Correlations for Odyssey dimensional predictions and pseudo-labels [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 2
Figure 2. Distribution of each categorical emotion percentage model outputs [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

47 extracted references · 34 canonical work pages

  1. [1]

    AudioInsight: Detecting Social Contexts Relevant to Social Anxiety from Speech

    V . Reddy, Z. Wang, E. Toner, M. Larrazabal, M. Boukhechba, B. A. Teachman, and L. E. Barnes, “AudioInsight: Detecting so- cial contexts relevant to social anxiety from speech,” arXiv preprint arXiv:2407.14458, 2024

  2. [2]

    Multilingual markers of depression in remotely collected speech samples: a prelim- inary analysis,

    N. Cummins, J. Dineley, P. Conde, F. Matcham, S. Siddi, F. Lamers, E. Carr, G. Lavelle, D. Leightley, K. M. White et al. , “Multilingual markers of depression in remotely collected speech samples: a prelim- inary analysis,” Journal of affective disorders , vol. 341, pp. 128–136, 2023

  3. [3]

    A. Sano, S. Taylor, A. W. McHill, A. J. Phillips, L. K. Barger, E. Klerman, and R. Picard, “Identifying objective physiological markers and modifiable behaviors for self-reported stress and mental health status using wearable sensors and mobile phones: observational study,”Journal of medical Internet research , vol. 20, no. 6, p. e210, 2018

  4. [4]

    Depression scale dictionary decomposition framework for multimodal automatic depression level prediction,

    M. Niu, X. Wang, J. Gong, B. Liu, J. Tao, and B. W. Schuller, “Depression scale dictionary decomposition framework for multimodal automatic depression level prediction,” IEEE Transactions on Circuits and Systems for Video Technology , 2025

  5. [5]

    A VEC 2019 workshop and challenge: state-of-mind, detecting depression with AI, and cross-cultural affect recognition,

    F. Ringeval, B. Schuller, M. Valstar, N. Cummins, R. Cowie, L. Tavabi, M. Schmitt, S. Alisamir, S. Amiriparian, E.-M. Messner et al., “A VEC 2019 workshop and challenge: state-of-mind, detecting depression with AI, and cross-cultural affect recognition,” in Proceedings of the 9th International on Audio/visual Emotion Challenge and Workshop , 2019, pp. 3–12

  6. [6]

    A review of depression and suicide risk assessment using speech analysis,

    N. Cummins, S. Scherer, J. Krajewski, S. Schnieder, J. Epps, and T. F. Quatieri, “A review of depression and suicide risk assessment using speech analysis,” Speech communication, vol. 71, pp. 10–49, 2015

  7. [7]

    Deploying a robotic positive psychology coach to improve college students’ psychological well- being,

    S. Jeong, L. Aymerich-Franch, K. Arias, S. Alghowinem, A. Lapedriza, R. Picard, H. W. Park, and C. Breazeal, “Deploying a robotic positive psychology coach to improve college students’ psychological well- being,” User Modeling and User-Adapted Interaction , vol. 33, no. 2, pp. 571–615, 2023

  8. [8]

    Improvement of public speaking skills using virtual reality: Development of a training system,

    S. Saufnay, E. Etienne, and M. Schyns, “Improvement of public speaking skills using virtual reality: Development of a training system,” in ACII

Show all 47 references
  1. [9]

    Meetingcoach: An intelligent dashboard for supporting effective & inclusive meetings,

    S. Samrose, D. McDuff, R. Sim, J. Suh, K. Rowan, J. Hernandez, S. Rin- tel, K. Moynihan, and M. Czerwinski, “Meetingcoach: An intelligent dashboard for supporting effective & inclusive meetings,” inProceedings of the 2021 CHI Conference on Human Factors in Computing Systems , ...

  2. [10]

    Emowear: Exploring emotional teasers for voice message interaction on smartwatches,

    P. An, J. S. Zhu, Z. Zhang, Y . Yin, Q. Ma, C. Yan, L. Du, and J. Zhao, “Emowear: Exploring emotional teasers for voice message interaction on smartwatches,” in Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems , 2024, pp. 1–16

  3. [11]

    Impact of interaction context on the student affect-learning relationship in child-robot inter- action,

    H. Chen, H. W. Park, X. Zhang, and C. Breazeal, “Impact of interaction context on the student affect-learning relationship in child-robot inter- action,” in Proceedings of the 2020 ACM/IEEE international conference on human-robot interaction , 2020, pp. 389–397

  4. [12]

    Child-centric robot dialogue systems: Fine-tuning large language models for better utterance understanding and interaction,

    D.-Y . Kim, H. J. Lym, H. Lee, Y . J. Lee, J. Kim, M.-G. Kim, and Y . Baek, “Child-centric robot dialogue systems: Fine-tuning large language models for better utterance understanding and interaction,” Sensors, vol. 24, no. 24, p. 7939, 2024

  5. [13]

    Methods to detect and reduce driver stress: a review,

    W.-Y . Chung, T.-W. Chong, and B.-G. Lee, “Methods to detect and reduce driver stress: a review,” International journal of automotive technology, vol. 20, pp. 1051–1063, 2019

  6. [14]

    Dawn of the transformer era in speech emotion recognition: closing the valence gap,

    J. Wagner, A. Triantafyllopoulos, H. Wierstorf, M. Schmitt, F. Burkhardt, F. Eyben, and B. W. Schuller, “Dawn of the transformer era in speech emotion recognition: closing the valence gap,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 9, pp. 10...

  7. [15]

    Odyssey 2024- speech emotion recognition challenge: Dataset, baseline framework, and results,

    L. Goncalves, A. N. Salman, A. R. Naini, L. M. Velazquez, T. Thebaud, L. P. Garcia, N. Dehak, B. Sisman, and C. Busso, “Odyssey 2024- speech emotion recognition challenge: Dataset, baseline framework, and results,” Development, vol. 10, no. 9,290, pp. 4–54, 2024

  8. [16]

    Investigating salient representations and label variance in dimensional speech emotion analysis,

    V . Mitra, J. Nie, and E. Azemi, “Investigating salient representations and label variance in dimensional speech emotion analysis,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 11 111–11 115

  9. [17]

    Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

    W. Hsu, B. Bolte, Y . H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM transactions on audio, speech, and language processing , vol. 29, pp. 3451–3460, 2021

  10. [18]

    wav2vec 2.0: A framework for self-supervised learning of speech representations,

    A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in neural information processing systems, vol. 33, pp. 12 449– 12 460, 2020

  11. [19]

    Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings,

    R. Lotfian and C. Busso, “Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings,” IEEE Transactions on Affective Computing , vol. 10, no. 4, pp. 471–483, 2017

  12. [20]

    IEMOCAP: Interactive emotional dyadic motion capture database,

    C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,” Language resources and evaluation , vol. 42, pp. 335–359, 2008

  13. [21]

    The Ryerson Audio-Visual Database of Emotional Speech and Song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in north american english,

    S. R. Livingstone and F. A. Russo, “The Ryerson Audio-Visual Database of Emotional Speech and Song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in north american english,” PloS one , vol. 13, no. 5, p. e0196391, 2018

  14. [22]

    Can large language models aid in annotating speech emotional data? uncovering new frontiers,

    S. Latif, M. Usama, M. I. Malik, and B. W. Schuller, “Can large language models aid in annotating speech emotional data? uncovering new frontiers,” arXiv preprint arXiv:2307.06090 , 2023

  15. [23]

    From text to emotion: Unveiling the emotion annotation capabilities of llms,

    M. Niu, M. Jaiswal, and E. M. Provost, “From text to emotion: Unveiling the emotion annotation capabilities of llms,” arXiv preprint arXiv:2408.17026, 2024

  16. [24]

    A wide evalu- ation of ChatGPT on affective computing tasks,

    M. M. Amin, R. Mao, E. Cambria, and B. W. Schuller, “A wide evalu- ation of ChatGPT on affective computing tasks,” IEEE Transactions on Affective Computing, 2024

  17. [25]

    Testing correctness, fairness, and robustness of speech emo- tion recognition models,

    A. Derington, H. Wierstorf, A. ¨Ozkil, F. Eyben, F. Burkhardt, and B. W. Schuller, “Testing correctness, fairness, and robustness of speech emo- tion recognition models,” IEEE Transactions on Affective Computing , no. 99, pp. 1–14, 2025

  18. [26]

    Bias and fairness on multimodal emotion detection algorithms,

    M. Schmitz, R. Ahmed, and J. Cao, “Bias and fairness on multimodal emotion detection algorithms,” arXiv preprint arXiv:2205.08383 , 2022

  19. [27]

    Enrolment-based personalisation for improving individual-level fairness in speech emotion recognition,

    A. Triantafyllopoulos and B. Schuller, “Enrolment-based personalisation for improving individual-level fairness in speech emotion recognition,” arXiv preprint arXiv:2406.06665 , 2024

  20. [28]

    Pre-trained speech processing models contain human-like biases that propagate to speech emotion recognition,

    I. Slaughter, C. Greenberg, R. Schwartz, and A. Caliskan, “Pre-trained speech processing models contain human-like biases that propagate to speech emotion recognition,” arXiv preprint arXiv:2310.18877 , 2023

  21. [29]

    V ocal changes common during aging proces

    E. M. Glazier and E. Ko, “V ocal changes common during aging proces.” [Online]. Available: https://www.uclahealth.org/news/article/vocal- changes-common-during-aging-process

  22. [30]

    Prevalence and etiologies of adult communication disabilities in the united states: Results from the 2012 national health interview survey,

    M. A. Morris, S. K. Meier, J. M. Griffin, M. E. Branda, and S. M. Phelan, “Prevalence and etiologies of adult communication disabilities in the united states: Results from the 2012 national health interview survey,” Disability and health journal , vol. 9, no. 1, pp. 140–144, 2016

  23. [31]

    Quick statistics about voice, speech, language,

    National Institute of Helath, “Quick statistics about voice, speech, language,” 2025

  24. [32]

    Community-supported shared infrastructure in support of speech acces- sibility,

    M. Hasegawa-Johnson, X. Zheng, H. Kim, C. Mendes, M. Dickinson, E. Hege, C. Zwilling, M. M. Channell, L. Mattie, H. Hodges et al. , “Community-supported shared infrastructure in support of speech acces- sibility,” Journal of Speech, Language, and Hearing Research , vol. 67, no...

  25. [33]

    Switchboard: Telephone speech corpus for research and development,

    J. J. Godfrey, E. C. Holliman, and J. McDaniel, “Switchboard: Telephone speech corpus for research and development,” in Acoustics, speech, and signal processing, ieee international conference on , vol. 1. IEEE Computer Society, 1992, pp. 517–520

  26. [34]

    Com- mon voice: A massively-multilingual speech corpus,

    R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Com- mon voice: A massively-multilingual speech corpus,” arXiv preprint arXiv:1912.06670, 2019

  27. [35]

    V oice command audios for virtual assistant,

    E. Buttaci, “V oice command audios for virtual assistant,” 2022

  28. [36]

    Wav2vec2fabundle,

    Torchaudio contributors, “Wav2vec2fabundle,” 2024. [Online]. Available: https://pytorch.org

  29. [37]

    emotion2vec: Self-supervised pre-training for speech emotion repre- sentation,

    Z. Ma, Z. Zheng, J. Ye, J. Li, Z. Gao, S. Zhang, and X. Chen, “emotion2vec: Self-supervised pre-training for speech emotion repre- sentation,” Proc. ACL 2024 Findings , 2024

  30. [38]

    SpeechBrain: A general- purpose speech toolkit,

    M. Ravanelli, T. Parcollet, P. Plantinga, A. Rouhe, S. Cornell, L. Lu- gosch, C. Subakan, N. Dawalatabad, A. Heba, J. Zhong, J.-C. Chou, S.-L. Yeh, S.-W. Fu, C.-F. Liao, E. Rastorgueva, F. Grondin, W. Aris, H. Na, Y . Gao, R. D. Mori, and Y . Bengio, “SpeechBrain: A general- p...

  31. [39]

    GPT-4o Audio,

    OpenAI, “GPT-4o Audio,” 2025. [Online]. Available: https://platform.openai.com/docs/models/gpt-4o-audio-preview

  32. [40]

    Rethinking emotion annotations in the era of large language models,

    M. Niu, Y . El-Tawil, A. Romana, and E. M. Provost, “Rethinking emotion annotations in the era of large language models,” arXiv preprint arXiv:2412.07906, 2024

  33. [41]

    Emobank: Studying the impact of annotation perspective and representation format on dimensional emotion analysis,

    S. Buechel and U. Hahn, “Emobank: Studying the impact of annotation perspective and representation format on dimensional emotion analysis,” arXiv preprint arXiv:2205.01996 , 2022

  34. [42]

    Switchboard-affect: Emotion perception labels from conversational speech,

    A. Romana, J. Narain, T. D. Tran, A. Davis, J. Fong, R. Rasipuram, and V . Mitra, “Switchboard-affect: Emotion perception labels from conversational speech,” ACII 2025, 2025

  35. [43]

    Affect detection: An interdisciplinary review of models, methods, and their applications,

    R. A. Calvo and S. D’Mello, “Affect detection: An interdisciplinary review of models, methods, and their applications,” IEEE Transactions on affective computing , vol. 1, no. 1, pp. 18–37, 2010

  36. [44]

    V oice quality dimensions as interpretable primitives for speaking style for atypical speech and affect

    Anonymous, “V oice quality dimensions as interpretable primitives for speaking style for atypical speech and affect.”

  37. [45]

    Guidelines for assessing and minimizing risks of emotion recognition applications,

    J. Hernandez, J. Lovejoy, D. McDuff, J. Suh, T. O’Brien, A. Sethu- madhavan, G. Greene, R. Picard, and M. Czerwinski, “Guidelines for assessing and minimizing risks of emotion recognition applications,” in 2021 9th International conference on affective computing and intelligen...

  38. [46]

    ReCANV o: A database of real-world communicative and affective nonverbal vocalizations,

    K. T. Johnson, J. Narain, T. Quatieri, P. Maes, and R. W. Picard, “ReCANV o: A database of real-world communicative and affective nonverbal vocalizations,” Scientific Data, vol. 10, no. 1, p. 523, 2023

  39. [2024]

    Institute of Electrical and Electronics Engineers, New-York, United States, 2024

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.