Pith. sign in

REVIEW 3 cited by

Hi-Fi Multi-Speaker English TTS Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.01497 v3 pith:I3I6ABAB submitted 2021-04-03 eess.AS

classification eess.AS
keywords datasetleastenglishhoursmulti-speakerspeechaudioaudiobooks
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper introduces a new multi-speaker English dataset for training text-to-speech models. The dataset is based on LibriVox audiobooks and Project Gutenberg texts, both in the public domain. The new dataset contains about 292 hours of speech from 10 speakers with at least 17 hours per speaker sampled at 44.1 kHz. To select speech samples with high quality, we considered audio recordings with a signal bandwidth of at least 13 kHz and a signal-to-noise ratio (SNR) of at least 32 dB. The dataset is publicly released at http://www.openslr.org/109/ .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning

    eess.AS 2026-03 unverdicted novelty 7.0 of 10

    FLAIR enables spoken dialogue AI to conduct continuous latent reasoning while perceiving speech through recursive latent embeddings and an ELBO-based finetuning objective.

  2. Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion

    eess.AS 2026-07 conditional novelty 6.0 of 10

    A new public dataset of 231,800 sound-designed vocalizations with seen/unseen preset-style and source-timbre splits, plus a baseline non-human voice-conversion evaluation.

  3. Length-Aware Rotary Position Embedding for Text-Speech Alignment

    eess.AS 2025-09 conditional novelty 6.0 of 10

    Length-aware RoPE (LARoPE), which normalizes positional indices by sequence length, induces a diagonal attention bias that improves text-speech alignment and achieves state-of-the-art WER on a zero-shot TTS benchmark.

Pith tools