Pith. sign in

REVIEW 4 cited by

EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.12867 v4 pith:XNA3ITEW submitted 2025-04-17 eess.AS cs.AIcs.CL

classification eess.AScs.AIcs.CL
keywords emotionemovoicespeechemotionallanguagemodeldatadataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Human speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made huge progress, they still face challenges in controlling the emotional expression in the generated speech. In this work, we propose EmoVoice, a novel emotion-controllable TTS model that exploits large language models (LLMs) to enable fine-grained freestyle natural language emotion control, and a phoneme boost variant design that makes the model output phoneme tokens and audio tokens in parallel to enhance content consistency, inspired by chain-of-thought (CoT) and chain-of-modality (CoM) techniques. Besides, we introduce EmoVoice-DB, a high-quality 40-hour English emotion dataset featuring expressive speech and fine-grained emotion labels with natural language descriptions. EmoVoice achieves state-of-the-art performance on the English EmoVoice-DB test set using only synthetic training data, and on the Chinese Secap test set using our in-house data. We further investigate the reliability of existing emotion evaluation metrics and their alignment with human perceptual preferences, and explore using SOTA multimodal LLMs GPT-4o-audio and Gemini to assess emotional speech. Dataset, code, checkpoints, and demo samples are available at https://github.com/yanghaha0908/EmoVoice.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation

    cs.CL 2025-07 conditional novelty 7.0 of 10

    With prompt engineering (audio concatenation plus in-context examples), large audio models rank speech synthesis systems in line with human preferences, reaching up to 0.91 Spearman correlation.

  2. MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts

    eess.AS 2025-08 conditional novelty 5.0 of 10

    MoE-TTS adds frozen text-expert MoE modules to a Qwen3-based TTS system and reports better out-of-domain description alignment than ElevenLabs and MiniMax on a small hand-built test set.

  3. Semantic-Aware Ship Detection with Vision-Language Integration

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    Abstract claims a VLM-plus-adaptive-window framework and a new semantic ship dataset, but the manuscript body is a different voice-timbre paper.

  4. Optimizing Multilingual Text-To-Speech with Accents & Emotions

    cs.LG 2025-06 reject novelty 3.0 of 10

    A TTS system built on Parler-TTS is claimed to improve accent accuracy and emotional expressiveness for Hindi and Indian English, but the paper lacks detailed architecture and baseline evidence.

Pith tools