REVIEW 9 cited by
JVS corpus: free Japanese multi-speaker voice corpus
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Thanks to improvements in machine learning techniques, including deep learning, speech synthesis is becoming a machine learning task. To accelerate speech synthesis research, we are developing Japanese voice corpora reasonably accessible from not only academic institutions but also commercial companies. In 2017, we released the JSUT corpus, which contains 10 hours of reading-style speech uttered by a single speaker, for end-to-end text-to-speech synthesis. For more general use in speech synthesis research, e.g., voice conversion and multi-speaker modeling, in this paper, we construct the JVS corpus, which contains voice data of 100 speakers in three styles (normal, whisper, and falsetto). The corpus contains 30 hours of voice data including 22 hours of parallel normal voices. This paper describes how we designed the corpus and summarizes the specifications. The corpus is available at our project page.
Forward citations
Cited by 9 Pith papers
-
A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors
On 1,168 professional voice actors, a misidentification floor in speaker embeddings survives calibration, normalization, and discriminative re-ranking, and the same floor makes fixed-threshold voice-clone attribution ...
-
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
An iterative active learning pipeline that filters web speech data by predicted quality and redundancy produces a TTS corpus with better speaker coverage at the same size.
-
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
A voice conversion pipeline conditions FastSpeech 2 on predicted likability ratings, achieving partial subjective control but with degraded identity and content at strong settings.
-
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
A transcript-prompted Whisper model with dictionary-based decoding automatically produces phonemic and prosodic annotations for Japanese audio-transcript pairs, improving Japanese TTS naturalness.
-
MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model
A T5 model predicts SSL-derived discrete speech tokens directly from mixed-script Japanese text, letting a FastSpeech 2 synthesizer produce speech without a grapheme-to-phoneme module.
-
QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
QHARMA-GAN combines quasi-harmonic speech modeling with a neural-network-estimated ARMA filter to synthesize and modify speech.
-
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
SpeechAccentLLM jointly trains foreign accent conversion and text-to-speech on CTC-regularized discrete speech tokens, with a BERT-style restorer, and reports improved accent reduction and intelligibility over one baseline.
-
Mitigating Language Mismatch in SSL-Based Speaker Anonymization
Fine-tuning an SSL content encoder on Japanese, especially when the encoder is pre-trained multilingually, makes anonymized Japanese and Mandarin speech much more intelligible while keeping speaker privacy at usable levels.
-
Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
MS-Wavehax, a multi-stream extension of the Wavehax vocoder, achieves the best throughput in low-latency CPU streaming and near non-causal quality with one frame of lookahead.
Discussion (0). Sign in to comment.