REVIEW 5 cited by
Speech Resynthesis from Discrete Disentangled Self-Supervised Representations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representations for speech content, prosodic information, and speaker identity. This allows to synthesize speech in a controllable manner. We analyze various state-of-the-art, self-supervised representation learning methods and shed light on the advantages of each method while considering reconstruction quality and disentanglement properties. Specifically, we evaluate the F0 reconstruction, speaker identification performance (for both resynthesis and voice conversion), recordings' intelligibility, and overall quality using subjective human evaluation. Lastly, we demonstrate how these representations can be used for an ultra-lightweight speech codec. Using the obtained representations, we can get to a rate of 365 bits per second while providing better speech quality than the baseline methods. Audio samples can be found under the following link: speechbot.github.io/resynthesis.
Forward citations
Cited by 5 Pith papers
-
Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework
By appending optimized token sequences to harmful speech, the authors achieve up to 89% attack success rate on SpeechGPT across six forbidden categories.
-
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
A text-like 'unit language' mined from discrete speech units via n-gram modeling, plus task-prompt multi-task training, improves textless speech-to-speech translation to near text-supervised performance.
-
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
PFlow-VC performs expressive voice conversion by conditioning a flow-matching Mel-spectrogram decoder on discrete speaker-normalized pitch tokens and a target speaker prompt, improving emotion style transfer.
-
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
An audio-visual language model that adds full-face visual features to a pre-trained expressive speech model improves emotion recognition and expressive speech generation by a few F1 points over speech-only on syntheti...
-
ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
A rectified-flow voice conversion model with speaker feature fusion achieves zero-shot conversion in one sampling step with quality close to 30-step diffusion baselines.
Discussion (0). Continue with ORCID to comment.