REVIEW 3 cited by
Overview of the Amphion Toolkit (v0.2)
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Amphion is an open-source toolkit for Audio, Music, and Speech Generation, designed to lower the entry barrier for junior researchers and engineers in these fields. It provides a versatile framework that supports a variety of generation tasks and models. In this report, we introduce Amphion v0.2, the second major release developed in 2024. This release features a 100K-hour open-source multilingual dataset, a robust data preparation pipeline, and novel models for tasks such as text-to-speech, audio coding, and voice conversion. Furthermore, the report includes multiple tutorials that guide users through the functionalities and usage of the newly released models.
Forward citations
Cited by 3 Pith papers
-
TRACE-EVC: Text-Guided Relative Affective Control for Zero-Shot Emotional Voice Conversion
Source-anchored rectified flow (Emo-Compass) plus TRACE-Instruct enables zero-shot emotional voice conversion driven by relative natural-language instructions rather than absolute targets.
-
Zero-Shot Text-to-Speech for Vietnamese
PhoAudiobook is a 941-hour Vietnamese audiobook dataset used to fine-tune and compare three zero-shot TTS models, with reported quality gains over a viVoice-trained baseline.
-
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
Vevo achieves zero-shot timbre, accent, and emotion imitation by using VQ-VAE codebook size on HuBERT features to create content and content-style tokens.
Discussion (0). Continue with ORCID to comment.