Pith. sign in

REVIEW 1 cited by

VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08076 v1 pith:34MS3D2J submitted 2024-06-12 eess.AS cs.SD

classification eess.AScs.SD
keywords identityvoicecross-lingualstyleemotionalcontrollableemotionintroduce
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the significant advancements in Text-to-Speech (TTS) systems, their full utilization in automatic dubbing remains limited. This task necessitates the extraction of voice identity and emotional style from a reference speech in a source language and subsequently transferring them to a target language using cross-lingual TTS techniques. While previous approaches have mainly concentrated on controlling voice identity within the cross-lingual TTS framework, there has been limited work on incorporating emotion and voice identity together. To this end, we introduce an end-to-end Voice Identity and Emotional Style Controllable Cross-Lingual (VECL) TTS system using multilingual speakers and an emotion embedding network. Moreover, we introduce content and style consistency losses to enhance the quality of synthesized speech further. The proposed system achieved an average relative improvement of 8.83\% compared to the state-of-the-art (SOTA) methods on a database comprising English and three Indian languages (Hindi, Telugu, and Marathi).

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimizing Multilingual Text-To-Speech with Accents & Emotions

    cs.LG 2025-06 reject novelty 3.0 of 10

    A TTS system built on Parler-TTS is claimed to improve accent accuracy and emotional expressiveness for Hindi and Indian English, but the paper lacks detailed architecture and baseline evidence.

Pith tools