Pith. sign in

REVIEW 10 cited by

AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.11567 v2 pith:R4ZFCRJR submitted 2020-10-22 cs.SD eess.AS

classification cs.SDeess.AS
keywords multi-speakercorpussystemsynthesisaishell-3mandarinsimilarityspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we present AISHELL-3, a large-scale and high-fidelity multi-speaker Mandarin speech corpus which could be used to train multi-speaker Text-to-Speech (TTS) systems. The corpus contains roughly 85 hours of emotion-neutral recordings spoken by 218 native Chinese mandarin speakers. Their auxiliary attributes such as gender, age group and native accents are explicitly marked and provided in the corpus. Accordingly, transcripts in Chinese character-level and pinyin-level are provided along with the recordings. We present a baseline system that uses AISHELL-3 for multi-speaker Madarin speech synthesis. The multi-speaker speech synthesis system is an extension on Tacotron-2 where a speaker verification model and a corresponding loss regarding voice similarity are incorporated as the feedback constraint. We aim to use the presented corpus to build a robust synthesis model that is able to achieve zero-shot voice cloning. The system trained on this dataset also generalizes well on speakers that are never seen in the training process. Objective evaluation results from our experiments show that the proposed multi-speaker synthesis system achieves high voice similarity concerning both speaker embedding similarity and equal error rate measurement. The dataset, baseline system code and generated samples are available online.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

    eess.AS 2026-08 conditional novelty 6.0 of 10

    SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.

  2. Aliasing-Free Neural Audio Synthesis

    cs.SD 2025-12 conditional novelty 6.0 of 10

    Pupu-Vocoder and Pupu-Codec use a closed-form anti-aliased SnakeBeta activation and resampling-based upsampling to reduce aliasing and improve singing, music, and audio synthesis.

  3. DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners

    cs.SD 2025-09 conditional novelty 6.0 of 10

    DeCodec learns a single neural codec that disentangles speech, background sound, semantic content, and paralinguistic style into orthogonal quantized streams, enabling reconstruction, enhancement, voice conversion, AS...

  4. SwiftF0: Fast and Accurate Monophonic Pitch Detection

    cs.SD 2025-08 conditional novelty 6.0 of 10

    SwiftF0 estimates monophonic pitch from a compact STFT-CNN, reporting better accuracy than CREPE under 10 dB noise at 42x lower CPU cost, alongside a new synthetic speech dataset and a six-component evaluation metric.

  5. UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling

    eess.AS 2025-08 conditional novelty 6.0 of 10

    UniFlow unifies four speech front-end tasks in one continuous-latent generative model with task-ID conditioning and reports competitive, but not uniformly superior, benchmark scores.

  6. SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.

  7. RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Authors propose ESS-CLAP and RA-CLAP, contrastive speech-text models for emotional speaking style retrieval, evaluated on PromptSpeech, TextrolSpeech, and SpeechCraft.

  8. Schr\"odinger Bridge Mamba for One-Step Speech Enhancement

    cs.SD 2025-10 conditional novelty 5.0 of 10

    A Mamba-based speech enhancer trained with Schrödinger Bridge objectives produces strong denoising and dereverberation in one inference step with a low real-time factor.

  9. Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training

    cs.SD 2025-06 reject novelty 5.0 of 10

    Pureformer-VC is a transformer-based encoder-decoder for non-parallel voice conversion that reports competitive, but not state-of-the-art, results on VCTK and AISHELL-3.

  10. Teffic-Audio: Tell Fact from Fiction

    cs.SD 2026-07 conditional novelty 4.0 of 10

    A simple Conformer deepfake detector trained with multi-source balanced sampling and diverse augmentation reaches 1.454% pooled EER on Speech-DF-Arena, first among public systems.

Pith tools