REVIEW 3 major objections 5 minor 47 references
Affect Models Have Weak Generalizability to Atypical Speech
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read AI emotion models misread neutral atypical speech as sad
desk verdict A large, mostly convincing fairness evaluation showing affect models misclassify atypical speech as sad; the main caveat is that the read-speech neutrality assumption is asserted rather than measured, and the fine-tuning result is partly circular. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the distribution of categorical emotion predictions over binarized atypicality groups, read against matched typical-speech baselines with the same expected emotional content. The load-bearing identity is the within- and between-dataset comparison: neutral read speech should produce similar neutral rates across datasets, so the large drop in neutral and rise in sad for SAP speech isolates the contribution of acoustic atypicality. The analysis also uses pseudo-label correlation, comparing text-only GPT-4o valence and arousal ratings against acoustic Odyssey predictions, and fine-tuning on pseudo-labeled atypical speech to test whether the gap can be reduced.
What would settle it
Have independent annotators label the affective content of a sample of SAP read sentences and digital commands without knowing the model outputs. If their neutral rate is comparable to Common Voice and VCVA, the sad shift is confirmed as an acoustic confound; if they also hear sadness or emotional valence in the atypical speech, the distributional gap is partly content-driven rather than a pure generalization failure.
Extended reading notes
Core claim
The central discovery is a systematic distributional shift in affect model outputs, not a small accuracy dip. For affectively neutral read material, the Odyssey model predicts neutral for only 5% of less intelligible SAP read sentences and sad for 82%, versus 90% neutral and 4% sad on Common Voice; similar but smaller shifts appear across Emotion2Vec, SpeechBrain, and GPT-4o-audio-preview, and across harshness and monopitch dimensions and digital commands. Within SAP, the more atypical binarized groups consistently receive less neutral and more sad output than the less atypical groups, so the effect tracks the degree of atypicality. The authors interpret these gaps as acoustic confounds: the models are using atypical voice properties as markers of sadness rather than detecting genuine affect in the speech.
Load-bearing premise
The load-bearing premise is that the SAP read sentences and digital commands are genuinely affectively neutral, so the high sad and low neutral predictions are errors rather than true affective content; the paper supports this with elicitation design and the authors' listening, not with independent affect ratings.
Editorial extensions
If this is right
- Voice-enabled affect applications will systematically over-attribute sadness to atypical speakers on neutral content, which could skew downstream decisions in wellbeing, coaching, or assistant interactions.
- The effect is continuous in atypicality: larger intelligibility, harshness, and monopitch ratings move predictions further from neutral, so thresholding or simple post-hoc corrections would need to be grade-dependent.
- Training-data affect elicitation strategy matters: an acted-speech-trained model shifts toward happy for more atypical speech while naturalistic-data models shift toward sad, so the bias direction is not fixed across models.
- Fine-tuning the valence model on pseudo-labeled atypical speech improved correlation on atypical speech and left typical-speech correlation essentially unchanged, indicating a low-cost path toward more robust models.
Reading between the lines
- If atypical acoustics reliably map onto sadness in these embeddings, any emotion-conditioned accessibility feature such as mood check-ins or stress tracking for motor-speech conditions could double-count the disability itself as negative affect; the paper does not test that application, but it follows directly.
- The within-dataset design suggests a cheap robustness experiment not run here: adding affectively neutral atypical speech to the training mix could show whether the sad shift shrinks faster than with full affect-labeled atypical data.
- Because arousal correlations were low even for mild atypicality while valence stayed closer to typical levels, prosodic atypicality may break arousal-related acoustic cues before categorical emotion labels do; a targeted study with synthetic voice transformations that vary pitch range could isolate which acoustic dimension drives which emotion dimension.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper evaluates publicly available speech affect models on atypical speech from the Speech Accessibility Project (SAP), comparing categorical and dimensional predictions against typical-speech datasets (Switchboard, Common Voice, VCVA). It examines three atypicality dimensions (intelligibility, harshness, monopitch) using within-SAP binarized comparisons and between-dataset comparisons, and reports that atypical speech is far more often predicted as sad and far less often as neutral, with larger gaps for more atypical speech. It also reports that correlations between text-based GPT-4o pseudo-labels and audio-based dimensional predictions are lower for atypical speech, and that fine-tuning an Odyssey valence model on SAP pseudo-labels improves Pearson correlation on SAP validation/test splits while leaving Switchboard correlation roughly unchanged. The conclusions claim weak generalizability of affect models to atypical speech and call for broader, more inclusive affect datasets.
Significance. If the main result holds, the paper addresses an important fairness and accessibility gap in speech emotion recognition: it shows on a large, multi-etiology atypical-speech dataset that categorical and dimensional affect models produce markedly different predictions for atypical speech, and that even mild atypicality shifts predictions away from neutral. The use of multiple off-the-shelf models, a publicly available dataset, binarized atypicality subgroups, and both within- and between-dataset comparisons are strengths, and the paper honestly discusses limitations. The central claim, however, depends on the assumption that SAP read sentences and digital commands are affectively neutral, and the dimensional and fine-tuning analyses rely on GPT-4o pseudo-labels whose validity is only partially established. These load-bearing points need additional support before the paper's conclusions can be accepted at face value.
major comments (3)
- [Section II-A, Table I] The central read-speech comparison rests on an unvalidated neutrality assumption. The manuscript states that read sentences and digital commands were 'overwhelmingly affectively neutral' based on elicitation prompts and the authors' listening, but no quantitative neutrality check is reported. Because Table I interprets the elevated sad and reduced neutral predictions for SAP read speech as evidence that atypical acoustics are confused with affect, any systematic affective difference between SAP read content and Common Voice/VCVA content would confound this conclusion. The within-SAP comparisons across atypicality levels partially control for dataset content, but they still assume no affective variation across rated levels of intelligibility, harshness, or monopitch. The Limitations section itself acknowledges that 'lexical differences correlating to the atypical speech rating ... could have influenced the GPT-4o annotations, and were not investigated,' which is the same class of confound for the read-speech analysis. I recommend adding independent listener annotations of affect for the read-speech subsets, or a validated lexical sentiment control, and reporting the neutrality check separately for each atypicality level.
- [Section III-B, Table II] The fine-tuning experiment is partly circular. The Odyssey valence model is fine-tuned on GPT-4o pseudo-labels for SAP training data and then evaluated on validation and test samples whose labels come from the same GPT-4o text-prompt procedure. Consequently, the reported improvement of +0.06 to +0.10 in Pearson correlation measures increased agreement with the pseudo-labeler, not necessarily with true affect. The claim that fine-tuning 'improves performance on atypical speech without impacting performance on typical speech' is not supported unless independent labels are used for evaluation, or unless the pseudo-labeler is shown to be an unbiased substitute for human labels on this population. Please re-evaluate with human-annotated affect labels, or recast the experiment as an alignment-to-pseudo-labels study.
- [Section II-B2, Figure 3] The arousal pseudo-labels are too weak to support the arousal claims. The authors report GPT-4o with the text prompt has CCC=0.28 and Pearson=0.44 for arousal on Emobank, yet the paper states that the model can 'effectively annotate the analyzed dimensions.' A CCC of 0.28 is generally considered poor agreement, and the observed low arousal correlations for SAP speech relative to Switchboard may reflect pseudo-label noise and lexical confounds rather than atypical-speech effects. Either restrict the dimensional analysis to valence, or provide an arousal label source with demonstrated validity.
minor comments (5)
- [Abstract, Section I, Section II-A] There are several typos: 'pronounciation' should be 'pronunciation', 'Dyarthria' should be 'Dysarthria', and 'psuedo-labels' should be 'pseudo-labels'.
- [Section III-B] The sentence 'A larger dataset ... could likely significantly model improve performance' contains a word-order error; it should read 'could likely significantly improve model performance'.
- [Section III-B] The personalization results report 'from 78 ± 1% to 81 ± 1% for read digital assistant commands' twice; the second instance should presumably refer to read sentences rather than commands.
- [Section V] The sentence 'Correlations for dimensional predictions with pseudo-labels were also lower for atypical speech than atypical speech' should read '... lower for atypical speech than for typical speech.'
- [Table I, Section III-A] Confidence intervals are provided for individual proportions, but differences between groups are not subjected to formal significance tests; since the main claims are comparative, adding a test or adjusted interval for the differences would strengthen the presentation.
Circularity Check
Fine-tuning gains are measured against the same GPT-4o pseudo-labeler used for training, making that robustness claim partly circular; the central distributional analysis is not circular.
-
fitted input called prediction
[Section II-C (Strategies for improving performance), Section II-B.2 (Dimensional Emotions), Table II]
"We used the GPT-4o dimensional valence scores as labels to fine-tune the Odyssey valence speech model on data from the SAP dataset. ... We tabulated Pearson correlation between the predicted speech valence ratings and text valence ratings ... Higher correlations signify better performance, as they correspond to more agreement between the text-only pseudo-label and the speech-only prediction. ... Table II compares the performance of the Odyssey dimensional valence model before and after fine-tuning on the SAP training data using the GPT-4o valence pseudo-labels."
The fine-tuning target and the evaluation target are the same GPT-4o text-only pseudo-label procedure. The model is trained with MSE loss against GPT-4o valence pseudo-labels, and the reported improvement is the Pearson correlation with GPT-4o valence pseudo-labels on held-out splits. Thus the gain measures how well the model learned to imitate the pseudo-labeler, including any systematic lexical or content biases that correlate with atypicality. The paper's own Limitations section concedes that 'lexical differences correlating to the atypical speech rating ...
full rationale
The central claim of the paper, that off-the-shelf affect models predict very different emotion distributions for atypical speech than for typical speech, is a direct empirical comparison of public model outputs and does not reduce to its inputs. Table I reports raw prediction percentages, and the between-dataset and within-dataset differences are measured, not derived from any fitted parameter. The neutrality assumption for SAP read speech is an unvalidated premise, not a circular one: the paper asserts it from the elicitation protocol and listening, and the conclusion that neutral read speech is predicted as sad follows only if that premise holds. That is a correctness risk, not a self-referential derivation. The self-citation to [42] is not load-bearing because the paper independently validates the GPT-4o dimensional prompt on Emobank. The one genuine circularity is the fine-tuning experiment: training labels and evaluation labels both come from the same GPT-4o pseudo-labeling procedure, so the improvement in Table II is partly an imitation effect rather than a demonstration of improved true affect recognition. The paper is transparent that these are pseudo-labels and acknowledges related limitations, but the specific before/after comparison is still circular in its label source. Overall, the main distributional findings are self-contained, while the robustness-improvement claim has a partial circularity, yielding a score of 6.
Assumptions & free parameters
free parameters (5)
- Binarization threshold for intelligibility =
2
- Binarization threshold for harshness =
4
- Binarization threshold for monopitch =
4
- Fine-tuning learning rate =
0.0001
- Fine-tuning epochs =
50
assumptions (4)
- domain assumption SAP read speech categories (novel sentences and digital voice commands) are affectively neutral.
- domain assumption GPT-4o text-based pseudo-labels for valence and arousal are valid proxies for true affect.
- domain assumption Speech-language pathologist ratings of intelligibility, harshness, and monopitch are accurate and independent of the affective content of the samples.
- domain assumption After subsampling, the typical speech datasets (Switchboard, Common Voice, VCVA) are comparable to SAP in content, length, accent, and language, differing mainly in atypicality.
Cite this review
Pith. "Pith review of Affect Models Have Weak Generalizability to Atypical Speech." pith.science (2026). https://pith.science/paper/OHS7PGH6
@misc{pith2026250416283,
author = {Pith},
title = {Pith review of: Affect Models Have Weak Generalizability to Atypical Speech},
year = {2026},
howpublished = {\url{https://pith.science/paper/OHS7PGH6}},
note = {Machine review of arXiv:2504.16283}
}
read the original abstract
Speech and voice conditions can alter the acoustic properties of speech, which could impact the performance of paralinguistic models for affect for people with atypical speech. We evaluate publicly available models for recognizing categorical and dimensional affect from speech on a dataset of atypical speech, comparing results to datasets of typical speech. We investigate three dimensions of speech atypicality: intelligibility, which is related to pronounciation; monopitch, which is related to prosody, and harshness, which is related to voice quality. We look at (1) distributional trends of categorical affect predictions within the dataset, (2) distributional comparisons of categorical affect predictions to similar datasets of typical speech, and (3) correlation strengths between text and speech predictions for spontaneous speech for valence and arousal. We find that the output of affect models is significantly impacted by the presence and degree of speech atypicalities. For instance, the percentage of speech predicted as sad is significantly higher for all types and grades of atypical speech when compared to similar typical speech datasets. In a preliminary investigation on improving robustness for atypical speech, we find that fine-tuning models on pseudo-labeled atypical speech data improves performance on atypical speech without impacting performance on typical speech. Our results emphasize the need for broader training and evaluation datasets for speech emotion models, and for modeling approaches that are robust to voice and speech differences.
Figures
Reference graph
Works this paper leans on
-
[1]
AudioInsight: Detecting Social Contexts Relevant to Social Anxiety from Speech
V . Reddy, Z. Wang, E. Toner, M. Larrazabal, M. Boukhechba, B. A. Teachman, and L. E. Barnes, “AudioInsight: Detecting so- cial contexts relevant to social anxiety from speech,” arXiv preprint arXiv:2407.14458, 2024
work page Pith review arXiv 2024
-
[2]
Multilingual markers of depression in remotely collected speech samples: a prelim- inary analysis,
N. Cummins, J. Dineley, P. Conde, F. Matcham, S. Siddi, F. Lamers, E. Carr, G. Lavelle, D. Leightley, K. M. White et al. , “Multilingual markers of depression in remotely collected speech samples: a prelim- inary analysis,” Journal of affective disorders , vol. 341, pp. 128–136, 2023
work page 2023
-
[3]
A. Sano, S. Taylor, A. W. McHill, A. J. Phillips, L. K. Barger, E. Klerman, and R. Picard, “Identifying objective physiological markers and modifiable behaviors for self-reported stress and mental health status using wearable sensors and mobile phones: observational study,”Journal of medical Internet research , vol. 20, no. 6, p. e210, 2018
work page 2018
-
[4]
M. Niu, X. Wang, J. Gong, B. Liu, J. Tao, and B. W. Schuller, “Depression scale dictionary decomposition framework for multimodal automatic depression level prediction,” IEEE Transactions on Circuits and Systems for Video Technology , 2025
work page 2025
-
[5]
F. Ringeval, B. Schuller, M. Valstar, N. Cummins, R. Cowie, L. Tavabi, M. Schmitt, S. Alisamir, S. Amiriparian, E.-M. Messner et al., “A VEC 2019 workshop and challenge: state-of-mind, detecting depression with AI, and cross-cultural affect recognition,” in Proceedings of the 9th International on Audio/visual Emotion Challenge and Workshop , 2019, pp. 3–12
work page 2019
-
[6]
A review of depression and suicide risk assessment using speech analysis,
N. Cummins, S. Scherer, J. Krajewski, S. Schnieder, J. Epps, and T. F. Quatieri, “A review of depression and suicide risk assessment using speech analysis,” Speech communication, vol. 71, pp. 10–49, 2015
2015
-
[7]
S. Jeong, L. Aymerich-Franch, K. Arias, S. Alghowinem, A. Lapedriza, R. Picard, H. W. Park, and C. Breazeal, “Deploying a robotic positive psychology coach to improve college students’ psychological well- being,” User Modeling and User-Adapted Interaction , vol. 33, no. 2, pp. 571–615, 2023
work page 2023
-
[8]
Improvement of public speaking skills using virtual reality: Development of a training system,
S. Saufnay, E. Etienne, and M. Schyns, “Improvement of public speaking skills using virtual reality: Development of a training system,” in ACII
Show all 47 references
-
[9]
Meetingcoach: An intelligent dashboard for supporting effective & inclusive meetings,
S. Samrose, D. McDuff, R. Sim, J. Suh, K. Rowan, J. Hernandez, S. Rin- tel, K. Moynihan, and M. Czerwinski, “Meetingcoach: An intelligent dashboard for supporting effective & inclusive meetings,” inProceedings of the 2021 CHI Conference on Human Factors in Computing Systems , ...
2021
-
[10]
Emowear: Exploring emotional teasers for voice message interaction on smartwatches,
P. An, J. S. Zhu, Z. Zhang, Y . Yin, Q. Ma, C. Yan, L. Du, and J. Zhao, “Emowear: Exploring emotional teasers for voice message interaction on smartwatches,” in Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems , 2024, pp. 1–16
2024
-
[11]
Impact of interaction context on the student affect-learning relationship in child-robot inter- action,
H. Chen, H. W. Park, X. Zhang, and C. Breazeal, “Impact of interaction context on the student affect-learning relationship in child-robot inter- action,” in Proceedings of the 2020 ACM/IEEE international conference on human-robot interaction , 2020, pp. 389–397
2020
-
[12]
Child-centric robot dialogue systems: Fine-tuning large language models for better utterance understanding and interaction,
D.-Y . Kim, H. J. Lym, H. Lee, Y . J. Lee, J. Kim, M.-G. Kim, and Y . Baek, “Child-centric robot dialogue systems: Fine-tuning large language models for better utterance understanding and interaction,” Sensors, vol. 24, no. 24, p. 7939, 2024
2024
-
[13]
Methods to detect and reduce driver stress: a review,
W.-Y . Chung, T.-W. Chong, and B.-G. Lee, “Methods to detect and reduce driver stress: a review,” International journal of automotive technology, vol. 20, pp. 1051–1063, 2019
2019
-
[14]
Dawn of the transformer era in speech emotion recognition: closing the valence gap,
J. Wagner, A. Triantafyllopoulos, H. Wierstorf, M. Schmitt, F. Burkhardt, F. Eyben, and B. W. Schuller, “Dawn of the transformer era in speech emotion recognition: closing the valence gap,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 9, pp. 10...
2023
-
[15]
Odyssey 2024- speech emotion recognition challenge: Dataset, baseline framework, and results,
L. Goncalves, A. N. Salman, A. R. Naini, L. M. Velazquez, T. Thebaud, L. P. Garcia, N. Dehak, B. Sisman, and C. Busso, “Odyssey 2024- speech emotion recognition challenge: Dataset, baseline framework, and results,” Development, vol. 10, no. 9,290, pp. 4–54, 2024
2024
-
[16]
Investigating salient representations and label variance in dimensional speech emotion analysis,
V . Mitra, J. Nie, and E. Azemi, “Investigating salient representations and label variance in dimensional speech emotion analysis,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 11 111–11 115
2024
-
[17]
Hubert: Self-supervised speech representation learning by masked prediction of hidden units,
W. Hsu, B. Bolte, Y . H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM transactions on audio, speech, and language processing , vol. 29, pp. 3451–3460, 2021
2021
-
[18]
wav2vec 2.0: A framework for self-supervised learning of speech representations,
A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in neural information processing systems, vol. 33, pp. 12 449– 12 460, 2020
2020
-
[19]
Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings,
R. Lotfian and C. Busso, “Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings,” IEEE Transactions on Affective Computing , vol. 10, no. 4, pp. 471–483, 2017
2017
-
[20]
IEMOCAP: Interactive emotional dyadic motion capture database,
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,” Language resources and evaluation , vol. 42, pp. 335–359, 2008
2008
-
[21]
The Ryerson Audio-Visual Database of Emotional Speech and Song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in north american english,
S. R. Livingstone and F. A. Russo, “The Ryerson Audio-Visual Database of Emotional Speech and Song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in north american english,” PloS one , vol. 13, no. 5, p. e0196391, 2018
2018
-
[22]
Can large language models aid in annotating speech emotional data? uncovering new frontiers,
S. Latif, M. Usama, M. I. Malik, and B. W. Schuller, “Can large language models aid in annotating speech emotional data? uncovering new frontiers,” arXiv preprint arXiv:2307.06090 , 2023
2023 arXiv
-
[23]
From text to emotion: Unveiling the emotion annotation capabilities of llms,
M. Niu, M. Jaiswal, and E. M. Provost, “From text to emotion: Unveiling the emotion annotation capabilities of llms,” arXiv preprint arXiv:2408.17026, 2024
2024 arXiv
-
[24]
A wide evalu- ation of ChatGPT on affective computing tasks,
M. M. Amin, R. Mao, E. Cambria, and B. W. Schuller, “A wide evalu- ation of ChatGPT on affective computing tasks,” IEEE Transactions on Affective Computing, 2024
2024
-
[25]
Testing correctness, fairness, and robustness of speech emo- tion recognition models,
A. Derington, H. Wierstorf, A. ¨Ozkil, F. Eyben, F. Burkhardt, and B. W. Schuller, “Testing correctness, fairness, and robustness of speech emo- tion recognition models,” IEEE Transactions on Affective Computing , no. 99, pp. 1–14, 2025
2025
-
[26]
Bias and fairness on multimodal emotion detection algorithms,
M. Schmitz, R. Ahmed, and J. Cao, “Bias and fairness on multimodal emotion detection algorithms,” arXiv preprint arXiv:2205.08383 , 2022
2022 arXiv
-
[27]
Enrolment-based personalisation for improving individual-level fairness in speech emotion recognition,
A. Triantafyllopoulos and B. Schuller, “Enrolment-based personalisation for improving individual-level fairness in speech emotion recognition,” arXiv preprint arXiv:2406.06665 , 2024
2024 arXiv
-
[28]
Pre-trained speech processing models contain human-like biases that propagate to speech emotion recognition,
I. Slaughter, C. Greenberg, R. Schwartz, and A. Caliskan, “Pre-trained speech processing models contain human-like biases that propagate to speech emotion recognition,” arXiv preprint arXiv:2310.18877 , 2023
2023 arXiv
-
[29]
V ocal changes common during aging proces
E. M. Glazier and E. Ko, “V ocal changes common during aging proces.” [Online]. Available: https://www.uclahealth.org/news/article/vocal- changes-common-during-aging-process
-
[30]
Prevalence and etiologies of adult communication disabilities in the united states: Results from the 2012 national health interview survey,
M. A. Morris, S. K. Meier, J. M. Griffin, M. E. Branda, and S. M. Phelan, “Prevalence and etiologies of adult communication disabilities in the united states: Results from the 2012 national health interview survey,” Disability and health journal , vol. 9, no. 1, pp. 140–144, 2016
2012
-
[31]
Quick statistics about voice, speech, language,
National Institute of Helath, “Quick statistics about voice, speech, language,” 2025
2025
-
[32]
Community-supported shared infrastructure in support of speech acces- sibility,
M. Hasegawa-Johnson, X. Zheng, H. Kim, C. Mendes, M. Dickinson, E. Hege, C. Zwilling, M. M. Channell, L. Mattie, H. Hodges et al. , “Community-supported shared infrastructure in support of speech acces- sibility,” Journal of Speech, Language, and Hearing Research , vol. 67, no...
2024
-
[33]
Switchboard: Telephone speech corpus for research and development,
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “Switchboard: Telephone speech corpus for research and development,” in Acoustics, speech, and signal processing, ieee international conference on , vol. 1. IEEE Computer Society, 1992, pp. 517–520
1992
-
[34]
Com- mon voice: A massively-multilingual speech corpus,
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Com- mon voice: A massively-multilingual speech corpus,” arXiv preprint arXiv:1912.06670, 2019
1912 arXiv
-
[35]
V oice command audios for virtual assistant,
E. Buttaci, “V oice command audios for virtual assistant,” 2022
2022
-
[36]
Wav2vec2fabundle,
Torchaudio contributors, “Wav2vec2fabundle,” 2024. [Online]. Available: https://pytorch.org
2024
-
[37]
emotion2vec: Self-supervised pre-training for speech emotion repre- sentation,
Z. Ma, Z. Zheng, J. Ye, J. Li, Z. Gao, S. Zhang, and X. Chen, “emotion2vec: Self-supervised pre-training for speech emotion repre- sentation,” Proc. ACL 2024 Findings , 2024
2024
-
[38]
SpeechBrain: A general- purpose speech toolkit,
M. Ravanelli, T. Parcollet, P. Plantinga, A. Rouhe, S. Cornell, L. Lu- gosch, C. Subakan, N. Dawalatabad, A. Heba, J. Zhong, J.-C. Chou, S.-L. Yeh, S.-W. Fu, C.-F. Liao, E. Rastorgueva, F. Grondin, W. Aris, H. Na, Y . Gao, R. D. Mori, and Y . Bengio, “SpeechBrain: A general- p...
2021 arXiv
-
[39]
GPT-4o Audio,
OpenAI, “GPT-4o Audio,” 2025. [Online]. Available: https://platform.openai.com/docs/models/gpt-4o-audio-preview
2025
-
[40]
Rethinking emotion annotations in the era of large language models,
M. Niu, Y . El-Tawil, A. Romana, and E. M. Provost, “Rethinking emotion annotations in the era of large language models,” arXiv preprint arXiv:2412.07906, 2024
2024 arXiv
-
[41]
Emobank: Studying the impact of annotation perspective and representation format on dimensional emotion analysis,
S. Buechel and U. Hahn, “Emobank: Studying the impact of annotation perspective and representation format on dimensional emotion analysis,” arXiv preprint arXiv:2205.01996 , 2022
2022 arXiv
-
[42]
Switchboard-affect: Emotion perception labels from conversational speech,
A. Romana, J. Narain, T. D. Tran, A. Davis, J. Fong, R. Rasipuram, and V . Mitra, “Switchboard-affect: Emotion perception labels from conversational speech,” ACII 2025, 2025
2025
-
[43]
Affect detection: An interdisciplinary review of models, methods, and their applications,
R. A. Calvo and S. D’Mello, “Affect detection: An interdisciplinary review of models, methods, and their applications,” IEEE Transactions on affective computing , vol. 1, no. 1, pp. 18–37, 2010
2010
-
[44]
V oice quality dimensions as interpretable primitives for speaking style for atypical speech and affect
Anonymous, “V oice quality dimensions as interpretable primitives for speaking style for atypical speech and affect.”
-
[45]
Guidelines for assessing and minimizing risks of emotion recognition applications,
J. Hernandez, J. Lovejoy, D. McDuff, J. Suh, T. O’Brien, A. Sethu- madhavan, G. Greene, R. Picard, and M. Czerwinski, “Guidelines for assessing and minimizing risks of emotion recognition applications,” in 2021 9th International conference on affective computing and intelligen...
2021
-
[46]
ReCANV o: A database of real-world communicative and affective nonverbal vocalizations,
K. T. Johnson, J. Narain, T. Quatieri, P. Maes, and R. W. Picard, “ReCANV o: A database of real-world communicative and affective nonverbal vocalizations,” Scientific Data, vol. 10, no. 1, p. 523, 2023
2023
-
[2024]
Institute of Electrical and Electronics Engineers, New-York, United States, 2024
2024
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.