Pith. sign in

REVIEW 5 cited by

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.18042 v4 pith:JN7XZV5V submitted 2024-09-26 cs.CV cs.CL

classification cs.CVcs.CL
keywords modelsspeechemotionsomni-modalvision-languageemovalanguageabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language Models to perceive and generate images, texts, and speeches end-to-end with publicly available data remains challenging for the open-source community. Existing vision-language models rely on external tools for speech processing, while speech-language models still suffer from limited or totally without vision-understanding capabilities. To address this gap, we propose the EMOVA (EMotionally Omni-present Voice Assistant), to enable Large Language Models with end-to-end speech abilities while maintaining the leading vision-language performance. With a semantic-acoustic disentangled speech tokenizer, we surprisingly notice that omni-modal alignment can further enhance vision-language and speech abilities compared with the bi-modal aligned counterparts. Moreover, a lightweight style module is introduced for the flexible speech style controls including emotions and pitches. For the first time, EMOVA achieves state-of-the-art performance on both the vision-language and speech benchmarks, and meanwhile, supporting omni-modal spoken dialogue with vivid emotions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

    cs.CL 2026-04 conditional novelty 6.5 of 10

    Current Omni LLMs extract modality-specific cues yet fail to integrate them for accurate multicontext safety judgments, performing better on physical than social/illegal risks.

  2. X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

    cs.LG 2026-07 conditional novelty 5.0 of 10

    X3-OPD improves audio-grounded reasoning by training the audio student on its own rollouts with token-level teacher feedback, using a three-tier paired text-audio corpus.

  3. MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A 3D facial animation framework that disentangles content and emotion and predicts frame-wise emotion intensity from audio plus text for dynamic expressions.

  4. Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

    cs.AI 2025-06 conditional novelty 5.0 of 10

    Stream-Omni uses CTC-based layer-dimension mapping to align speech with text, achieving vision, speech, and text interaction in one 8B model trained on 23,000 hours of speech.

  5. ECCV 2024 W-CODA: 1st Workshop on Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving

    cs.CV 2025-07 unverdicted novelty 1.0 of 10

    A workshop report documenting the ECCV 2024 W-CODA event, its accepted papers, speakers, and the dual-track corner case understanding and generation challenge.

Pith tools