An emotion-controllable movie dubbing model uses supervised lip-prosody alignment, phoneme enhancement, and flow matching with positive/negative classifier guidance to synthesize speech with user-chosen emotion type and intensity.
Audio-visual efficient conformer for robust speech recognition
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.SD 1years
2024 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
An emotion-controllable movie dubbing model uses supervised lip-prosody alignment, phoneme enhancement, and flow matching with positive/negative classifier guidance to synthesize speech with user-chosen emotion type and intensity.