REVIEW 10 cited by
A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Speech self-supervised models such as wav2vec 2.0 and HuBERT are making revolutionary progress in Automatic Speech Recognition (ASR). However, they have not been totally proven to produce better performance on tasks other than ASR. In this work, we explored partial fine-tuning and entire fine-tuning on wav2vec 2.0 and HuBERT pre-trained models for three non-ASR speech tasks: Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding. With simple proposed downstream frameworks, the best scores reached 79.58% weighted accuracy on speaker-dependent setting and 73.01% weighted accuracy on speaker-independent setting for Speech Emotion Recognition on IEMOCAP, 2.36% equal error rate for Speaker Verification on VoxCeleb1, 89.38% accuracy for Intent Classification and 78.92% F1 for Slot Filling on SLURP, showing the strength of fine-tuned wav2vec 2.0 and HuBERT on learning prosodic, voice-print and semantic representations.
Forward citations
Cited by 10 Pith papers
-
A Dataset for Automatic Assessment of TTS Quality in Spanish
A new Spanish-language dataset of 4,326 MOS-rated TTS audio clips enables automated naturalness prediction with a mean absolute error around 0.8 on a five-point scale.
-
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
Quantizing speech at frame, phone, word, and utterance levels preserves emotion and prominence better than frame-only discrete units at similar bitrates.
-
Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts
Generated multilingual ASR transcripts fused by cascaded cross-modal transformers raise audio sentiment accuracy, and the multimodal knowledge distills into a stronger audio-only WavLM student with no inference cost.
-
InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective
InsideSSL analyzes self-supervised speech models layer-by-layer using entropy, curvature, robustness metrics, and a cross-layer Generative Compatibility Matrix, finding that training objectives induce distinct compres...
-
EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis
EDTalk++ disentangles talking-head video into four orthogonal motion banks (mouth, pose, eyes, expression) and drives them from either video or audio inputs.
-
"How to Explore Biases in Speech Emotion AI with Users?" A Speech-Emotion-Acting Study Exploring Age and Language Biases
In a 24-person Danish study, a speech emotion recognition model showed no significant age or language differences in recognizing deliberately acted happy, sad, angry, and calm speech, though high-arousal emotions were...
-
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
Unsupervised k-means clusters of SSL features and i-vectors identify speaker-relevant feed-forward neurons; protecting them during pruning preserves speaker identification.
-
Acoustic-to-Articulatory Inversion of Clean Speech Using an MRI-Trained Model
Clean speech reconstructs vocal-tract contours with a mean RMSE of 1.56 mm, only 0.05 mm worse than denoised MRI speech.
-
EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition
EmoSLLM, a LoRA-fine-tuned 3B LLM with an audio mapper, beats several 7B speech-text LLMs on MSP-Podcast emotion recognition, but only when given the ground-truth transcript at inference.
-
Deep Learning Approaches for Multimodal Intent Recognition: A Survey
A survey of deep learning methods for intent recognition, tracing the field from unimodal text, audio, vision, and EEG approaches to multimodal fusion, alignment, knowledge-augmented, and multi-task models.
Discussion (0). Continue with ORCID to comment.