Pith. sign in

REVIEW 5 cited by

Amphion: An Open-Source Audio, Music and Speech Generation Toolkit

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.09911 v3 pith:GLREUXZG submitted 2023-12-15 cs.SD eess.AS

classification cs.SDeess.AS
keywords amphionaudiogenerationspeechtoolkiteasemodelsmusic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Amphion is an open-source toolkit for Audio, Music, and Speech Generation, targeting to ease the way for junior researchers and engineers into these fields. It presents a unified framework that includes diverse generation tasks and models, with the added bonus of being easily extendable for new incorporation. The toolkit is designed with beginner-friendly workflows and pre-trained models, allowing both beginners and seasoned researchers to kick-start their projects with relative ease. The initial release of Amphion v0.1 supports a range of tasks including Text to Speech (TTS), Text to Audio (TTA), and Singing Voice Conversion (SVC), supplemented by essential components like data preprocessing, state-of-the-art vocoders, and evaluation metrics. This paper presents a high-level overview of Amphion. Amphion is open-sourced at https://github.com/open-mmlab/Amphion.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Over-the-Air Adversarial Attack Detection: from Datasets to Defenses

    eess.AS 2025-09 conditional novelty 6.0 of 10

    AdvSV 2.0 provides 629k adversarial and bona fide audio samples for automatic speaker verification, with a neural replay simulator attack and a one-class contrastive domain-aligned detector reaching 11.2% EER.

  2. PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data

    eess.AS 2025-06 conditional novelty 6.0 of 10

    A two-part training strategy that uses a pretrained voice conversion model to generate pseudo paired data and same-speaker sampling to reduce train-inference mismatch, improving one-shot voice conversion over FreeVC a...

  3. InteRecon: Towards Reconstructing Interactivity of Personal Memorable Items in Mixed Reality

    cs.HC 2025-02 conditional novelty 6.0 of 10

    The paper demonstrates a prototype that lets people turn cherished objects into interactive AR versions that preserve their original motions, buttons, and embedded media.

  4. Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A cross-modal attention ensemble with balanced stacking reaches MacroF1 0.4094 on 8-class naturalistic speech emotion recognition, beating the official baseline by about 0.08.

  5. Uni-VERSA: Versatile Speech Assessment with a Unified Network

    cs.SD 2025-05 conditional novelty 4.0 of 10

    A single network predicts eleven speech quality metrics across five dimensions and shows strong correlations within speech enhancement, but not on out-of-domain data.

Pith tools