Pith. sign in

REVIEW 10 cited by

A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.02735 v3 pith:WMU7A6J5 submitted 2021-11-04 cs.CL cs.NEcs.SDeess.AS

classification cs.CLcs.NEcs.SDeess.AS
keywords speechhubertrecognitionwav2vecaccuracyemotionspeakerverification
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speech self-supervised models such as wav2vec 2.0 and HuBERT are making revolutionary progress in Automatic Speech Recognition (ASR). However, they have not been totally proven to produce better performance on tasks other than ASR. In this work, we explored partial fine-tuning and entire fine-tuning on wav2vec 2.0 and HuBERT pre-trained models for three non-ASR speech tasks: Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding. With simple proposed downstream frameworks, the best scores reached 79.58% weighted accuracy on speaker-dependent setting and 73.01% weighted accuracy on speaker-independent setting for Speech Emotion Recognition on IEMOCAP, 2.36% equal error rate for Speaker Verification on VoxCeleb1, 89.38% accuracy for Intent Classification and 78.92% F1 for Slot Filling on SLURP, showing the strength of fine-tuned wav2vec 2.0 and HuBERT on learning prosodic, voice-print and semantic representations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Dataset for Automatic Assessment of TTS Quality in Spanish

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A new Spanish-language dataset of 4,326 MOS-rated TTS audio clips enables automated naturalness prediction with a mean absolute error around 0.8 on a five-point scale.

  2. Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Quantizing speech at frame, phone, word, and utterance levels preserves emotion and prominence better than frame-only discrete units at similar bitrates.

  3. Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts

    cs.CL 2026-07 conditional novelty 5.5 of 10

    Generated multilingual ASR transcripts fused by cascaded cross-modal transformers raise audio sentiment accuracy, and the multimodal knowledge distills into a stronger audio-only WavLM student with no inference cost.

  4. InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective

    cs.SD 2026-07 conditional novelty 5.0 of 10

    InsideSSL analyzes self-supervised speech models layer-by-layer using entropy, curvature, robustness metrics, and a cross-layer Generative Compatibility Matrix, finding that training objectives induce distinct compres...

  5. EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis

    cs.CV 2025-08 conditional novelty 5.0 of 10

    EDTalk++ disentangles talking-head video into four orthogonal motion banks (mouth, pose, eyes, expression) and drives them from either video or audio inputs.

  6. "How to Explore Biases in Speech Emotion AI with Users?" A Speech-Emotion-Acting Study Exploring Age and Language Biases

    cs.HC 2025-07 conditional novelty 5.0 of 10

    In a 24-person Danish study, a speech emotion recognition model showed no significant age or language differences in recognizing deliberately acted happy, sad, angry, and calm speech, though high-arousal emotions were...

  7. Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Unsupervised k-means clusters of SSL features and i-vectors identify speaker-relevant feed-forward neurons; protecting them during pruning preserves speaker identification.

  8. Acoustic-to-Articulatory Inversion of Clean Speech Using an MRI-Trained Model

    eess.AS 2026-03 conditional novelty 4.0 of 10

    Clean speech reconstructs vocal-tract contours with a mean RMSE of 1.56 mm, only 0.05 mm worse than denoised MRI speech.

  9. EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition

    eess.AS 2025-08 reject novelty 4.0 of 10

    EmoSLLM, a LoRA-fine-tuned 3B LLM with an audio mapper, beats several 7B speech-text LLMs on MSP-Podcast emotion recognition, but only when given the ground-truth transcript at inference.

  10. Deep Learning Approaches for Multimodal Intent Recognition: A Survey

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A survey of deep learning methods for intent recognition, tracing the field from unimodal text, audio, vision, and EEG approaches to multimodal fusion, alignment, knowledge-augmented, and multi-task models.

Pith tools