REVIEW 5 cited by
Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Amphion is an open-source toolkit for Audio, Music, and Speech Generation, targeting to ease the way for junior researchers and engineers into these fields. It presents a unified framework that includes diverse generation tasks and models, with the added bonus of being easily extendable for new incorporation. The toolkit is designed with beginner-friendly workflows and pre-trained models, allowing both beginners and seasoned researchers to kick-start their projects with relative ease. The initial release of Amphion v0.1 supports a range of tasks including Text to Speech (TTS), Text to Audio (TTA), and Singing Voice Conversion (SVC), supplemented by essential components like data preprocessing, state-of-the-art vocoders, and evaluation metrics. This paper presents a high-level overview of Amphion. Amphion is open-sourced at https://github.com/open-mmlab/Amphion.
Forward citations
Cited by 5 Pith papers
-
Over-the-Air Adversarial Attack Detection: from Datasets to Defenses
AdvSV 2.0 provides 629k adversarial and bona fide audio samples for automatic speaker verification, with a neural replay simulator attack and a one-class contrastive domain-aligned detector reaching 11.2% EER.
-
PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data
A two-part training strategy that uses a pretrained voice conversion model to generate pseudo paired data and same-speaker sampling to reduce train-inference mismatch, improving one-shot voice conversion over FreeVC a...
-
InteRecon: Towards Reconstructing Interactivity of Personal Memorable Items in Mixed Reality
The paper demonstrates a prototype that lets people turn cherished objects into interactive AR versions that preserve their original motions, buttons, and embedded media.
-
Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model
A cross-modal attention ensemble with balanced stacking reaches MacroF1 0.4094 on 8-class naturalistic speech emotion recognition, beating the official baseline by about 0.08.
-
Uni-VERSA: Versatile Speech Assessment with a Unified Network
A single network predicts eleven speech quality metrics across five dimensions and shows strong correlations within speech enhancement, but not on out-of-domain data.
Discussion (0). Continue with ORCID to comment.