Pith. sign in

REVIEW 8 cited by

Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.03502 v1 pith:6FJ4VYGO submitted 2021-04-08 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords emotionrecognitionspeechwav2vecapproacheslearningmodelmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Emotion recognition datasets are relatively small, making the use of the more sophisticated deep learning approaches challenging. In this work, we propose a transfer learning method for speech emotion recognition where features extracted from pre-trained wav2vec 2.0 models are modeled using simple neural networks. We propose to combine the output of several layers from the pre-trained model using trainable weights which are learned jointly with the downstream model. Further, we compare performance using two different wav2vec 2.0 models, with and without finetuning for speech recognition. We evaluate our proposed approaches on two standard emotion databases IEMOCAP and RAVDESS, showing superior performance compared to results in the literature.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Across six pooling heads and six frozen SSL backbones on English and Mandarin depression speech, a third of configurations collapse to single-class prediction, so backbone- and seed-robustness should be first-class ev...

  2. Synthetic Audio Helps for Cognitive State Tasks

    cs.SD 2025-02 conditional novelty 6.0 of 10

    Adding zero-shot synthetic audio from a text-to-speech system to text-only models improves cognitive-state prediction on seven tasks, though gains are small.

  3. InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective

    cs.SD 2026-07 conditional novelty 5.0 of 10

    InsideSSL analyzes self-supervised speech models layer-by-layer using entropy, curvature, robustness metrics, and a cross-layer Generative Compatibility Matrix, finding that training objectives induce distinct compres...

  4. Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Supervised contrastive alignment of frozen WavLM layers modestly lifts Mandarin depression F1 under LOSO while quantifying that speaker leakage inflated prior Mandarin F1 by ~0.23.

  5. Multiple-Noise-Resilient Nonadiabatic Geometric Quantum Control of Solid-State Spins in Diamond

    quant-ph 2025-08 unverdicted novelty 5.0 of 10

    A claimed experiment shows a multiple-noise-resilient geometric gate on a diamond NV spin achieving fidelity 0.9992(1) and 690±30 µs coherence, 3.5x the dynamical gate.

  6. Sounding Like a Winner? Prosodic Differences in Post-Match Interviews

    cs.CL 2025-06 reject novelty 5.0 of 10

    Self-supervised speech representations can classify tennis match outcomes from post-match interview audio above chance, but the claimed prosodic indicators such as pitch variability are not supported by the reported e...

  7. Investigating the Impact of Word Informativeness on Speech Emotion Recognition

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Using GPT-2 surprisal to select a few words per sentence for acoustic feature extraction gives a small accuracy improvement over whole-utterance features in RAVDESS speech emotion recognition, though the corpus's two ...

  8. Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition?

    eess.AS 2025-06 conditional novelty 4.0 of 10

    Audio-Mamba models outperform attention-based models such as WavLM and HuBERT on non-verbal emotion recognition, and a Renyi-divergence fusion approach called RENO further improves accuracy.

Pith tools