Pith. sign in

REVIEW 2 cited by

Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08353 v3 pith:WONLWLZD submitted 2024-06-12 eess.AS cs.CLcs.MMcs.SD

classification eess.AScs.CLcs.MMcs.SD
keywords errorfusionrecognitionspeechtextcomprehensiveemotionfindings
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Text data is commonly utilized as a primary input to enhance Speech Emotion Recognition (SER) performance and reliability. However, the reliance on human-transcribed text in most studies impedes the development of practical SER systems, creating a gap between in-lab research and real-world scenarios where Automatic Speech Recognition (ASR) serves as the text source. Hence, this study benchmarks SER performance using ASR transcripts with varying Word Error Rates (WERs) from eleven models on three well-known corpora: IEMOCAP, CMU-MOSI, and MSP-Podcast. Our evaluation includes both text-only and bimodal SER with six fusion techniques, aiming for a comprehensive analysis that uncovers novel findings and challenges faced by current SER research. Additionally, we propose a unified ASR error-robust framework integrating ASR error correction and modality-gated fusion, achieving lower WER and higher SER results compared to the best-performing ASR transcript. These findings provide insights into SER with ASR assistance, especially for real-world applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation

    cs.CL 2025-09 conditional novelty 6.0 of 10

    PARCO cuts named-entity errors by large margins (NE-CER 1.57% on AISHELL-1, NE-WER 8.34% on DATA2 at zero distractors) by combining phoneme-enriched entity encoding, a contrastive disambiguation loss, and hierarchical...

  2. Leveraging Prompt Learning and Pause Encoding for Alzheimer's Disease Detection

    eess.AS 2024-12 conditional novelty 4.0 of 10

    Prompt-based fine-tuning with pause encoding reaches a maximum 95.8% accuracy for Alzheimer's detection on ADReSS transcripts, while the mean over random seeds is 87.9%.

Pith tools