REVIEW 21 cited by
emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose emotion2vec, a universal speech emotion representation model. emotion2vec is pre-trained on open-source unlabeled emotion data through self-supervised online distillation, combining utterance-level loss and frame-level loss during pre-training. emotion2vec outperforms state-of-the-art pre-trained universal models and emotion specialist models by only training linear layers for the speech emotion recognition task on the mainstream IEMOCAP dataset. In addition, emotion2vec shows consistent improvements among 10 different languages of speech emotion recognition datasets. emotion2vec also shows excellent results on other emotion tasks, such as song emotion recognition, emotion prediction in conversation, and sentiment analysis. Comparison experiments, ablation experiments, and visualization comprehensively demonstrate the universal capability of the proposed emotion2vec. To the best of our knowledge, emotion2vec is the first universal representation model in various emotion-related tasks, filling a gap in the field.
Forward citations
Cited by 21 Pith papers
-
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model
FlexiSLM is the first spoken language model supporting dynamic and controllable frame rates on speech input and output, outperforming fixed-rate 7B models at high quality and enabling faster inference at lower rates l...
-
SpeechDx: A Multi-Task Benchmark for Clinical Speech AI
SpeechDx is a multi-task benchmark with 12 datasets and 27 tasks across health conditions, structured by conceptualization, formulation, and articulation stages, showing that no current audio encoder generalizes reliably.
-
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing
VITA-QinYu is the first expressive end-to-end spoken language model supporting role-playing and singing alongside conversation, trained on 15.8K hours of data and outperforming prior models on expressiveness and conve...
-
EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses
EmoTransCap creates the first large-scale dataset for discourse-level emotion transitions in speech, a multi-task recognition model, LLM-based annotations, and a controllable emotional speech synthesis system.
-
CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation
CapTalk unifies single-utterance and dialogue voice design via utterance- and speaker-level captions plus a hierarchical variational module for stable timbre with adaptive expression.
-
MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
MECAT is a multi-expert benchmark for audio AI offering fine-grained captions and QA pairs generated via expert models and LLM reasoning, paired with the DATE metric that combines semantic similarity and cross-sample ...
-
UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling
UniSAE unifies speaker, emotion, and multi-granularity content editing in speech via a new discrete phonetic posteriorgram representation and diffusion-based rendering.
-
ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
Dual-encoder speech-text models trained on rich intrinsic and situational style captions outperform prior CLAP-style baselines on retrieval, classification, and inference-time TTS style guidance.
-
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
X3-OPD improves audio-grounded reasoning by training the audio student on its own rollouts with token-level teacher feedback, using a three-tier paired text-audio corpus.
-
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
Multimodal LLM analysis correlates better with TRUST-Pathos than acoustic SER models in a case study of one Bundestag speech, while acoustic features help with arousal.
-
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models
An emotion TTS system adjusts Classifier-Free Guidance strength according to text-style semantic mismatch; it shows small emotion-accuracy gains, but headline baselines and subjective results are absent from the main text.
-
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models
An adaptive CFG method that tunes guidance based on LLM-detected mismatch between emotion prompts and text semantics improves emotional expressiveness in AR TTS while preserving audio quality and intelligibility.
-
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models
Introduces CCG-CFG with inconsistency-based dynamic scales and hard-sample mining distillation to boost emotional alignment in auto-regressive TTS, reporting up to 12% absolute gains in emotion recognition accuracy.
-
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents
A 65.6-hour movie-dialogue benchmark for spoken role-playing agents, with an evaluation framework whose main judge is also an evaluated model.
-
Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody
Maestro-EVC independently controls content, speaker, and emotion in voice conversion using separate references and explicit prosody modeling, outperforming StyleVC and ZEST on emotion similarity and prosody.
-
Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
Psychology-guided LLM text embeddings fused with audio and facial cues achieved the lowest MSE in the AVI 2025 personality assessment challenge.
-
MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding
A 3D facial animation framework that disentangles content and emotion and predicts frame-wise emotion intensity from audio plus text for dynamic expressions.
-
EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection
Emo-Boost augments low-level deepfake detectors with intra- and inter-modal emotion consistency checks to raise cross-manipulation generalization AUC by 2.1% on FakeAVCeleb.
-
Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment
A shared-compression MLP fusion with 32-head ensemble learning achieves MSE 0.1824, the top score in the AVI 2025 interview assessment track.
-
BoSS: Beyond-Semantic Speech
Current spoken-language models perform poorly on a new five-task evaluation of beyond-semantic speech signals, including dialect, emotion, age, and non-verbal cues.
-
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
A two-stage static-then-dynamic prompt selection strategy using prosodic features, LLM coherence scores, and similarity metrics improves emotion intensity and speaker consistency in zero-shot TTS.
Discussion (0). Sign in to comment.