Pith. sign in

REVIEW 4 cited by

Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.03540 v1 pith:V3PJK6TW submitted 2023-02-07 cs.SD eess.AS

classification cs.SDeess.AS
keywords dataspear-ttstokensacousticminimalonlyparallelreading
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce SPEAR-TTS, a multi-speaker text-to-speech (TTS) system that can be trained with minimal supervision. By combining two types of discrete speech representations, we cast TTS as a composition of two sequence-to-sequence tasks: from text to high-level semantic tokens (akin to "reading") and from semantic tokens to low-level acoustic tokens ("speaking"). Decoupling these two tasks enables training of the "speaking" module using abundant audio-only data, and unlocks the highly efficient combination of pretraining and backtranslation to reduce the need for parallel data when training the "reading" component. To control the speaker identity, we adopt example prompting, which allows SPEAR-TTS to generalize to unseen speakers using only a short sample of 3 seconds, without any explicit speaker representation or speaker-id labels. Our experiments demonstrate that SPEAR-TTS achieves a character error rate that is competitive with state-of-the-art methods using only 15 minutes of parallel data, while matching ground-truth speech in terms of naturalness and acoustic quality, as measured in subjective tests.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WaveVerify: A Novel Audio Watermarking Framework for Media Authentication and Combatting Deepfakes

    cs.CR 2025-07 conditional novelty 6.0 of 10

    WaveVerify embeds audio watermarks with a FiLM-based generator and extracts them with a Mixture-of-Experts detector, reporting zero bit error and high localization under common distortions.

  2. DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding

    eess.AS 2025-06 conditional novelty 6.0 of 10

    DiffSoundStream uses a latent diffusion decoder conditioned on WavLM semantic tokens and coarse SoundStream acoustic tokens to match 100-token-per-second quality at 50 tokens per second.

  3. Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations

    cs.SD 2025-07 reject novelty 5.0 of 10

    QTTS models speech as sequences from a multi-codebook RVQ audio codec whose first codebook is trained with ASR supervision, aiming for higher-fidelity TTS than single-codebook systems.

  4. Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes

    cs.SD 2025-05 conditional novelty 5.0 of 10

    ADD-GP, a Gaussian Process classifier with XLS-R speech embeddings, adapts to unseen TTS models with as few as 5 samples and achieves state-of-the-art low error rates on the new LibriFake benchmark.

Pith tools