Pith. sign in

REVIEW 1 cited by

KazakhTTS: An Open-Source Kazakh Text-to-Speech Synthesis Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.08459 v3 pith:EDDJMVCI submitted 2021-04-17 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords datasetkazakhmodelsavailableopen-sourcespeakersspokensynthesis
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper introduces a high-quality open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide. The dataset consists of about 93 hours of transcribed audio recordings spoken by two professional speakers (female and male). It is the first publicly available large-scale dataset developed to promote Kazakh text-to-speech (TTS) applications in both academia and industry. In this paper, we share our experience by describing the dataset development procedures and faced challenges, and discuss important future directions. To demonstrate the reliability of our dataset, we built baseline end-to-end TTS models and evaluated them using the subjective mean opinion score (MOS) measure. Evaluation results show that the best TTS models trained on our dataset achieve MOS above 4 for both speakers, which makes them applicable for practical use. The dataset, training recipe, and pretrained TTS models are freely available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Inclusivity of AI Speech in Healthcare: A Decade Look Back

    cs.CY 2025-05 conditional novelty 4.0 of 10

    A decade-long audit finds persistent inclusivity gaps in speech AI for healthcare: English-heavy datasets, little demographic metadata, no speech-impaired samples, and limited bias research.

Pith tools