Pith. sign in

REVIEW 4 cited by

Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.01409 v1 pith:YLYFYEOX submitted 2021-04-03 eess.AS cs.AIcs.SD

classification eess.AScs.AIcs.SD
keywords diff-ttsdiffusiondenoisingfastermel-spectrogrammethodmodelspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although neural text-to-speech (TTS) models have attracted a lot of attention and succeeded in generating human-like speech, there is still room for improvements to its naturalness and architectural efficiency. In this work, we propose a novel non-autoregressive TTS model, namely Diff-TTS, which achieves highly natural and efficient speech synthesis. Given the text, Diff-TTS exploits a denoising diffusion framework to transform the noise signal into a mel-spectrogram via diffusion time steps. In order to learn the mel-spectrogram distribution conditioned on the text, we present a likelihood-based optimization method for TTS. Furthermore, to boost up the inference speed, we leverage the accelerated sampling method that allows Diff-TTS to generate raw waveforms much faster without significantly degrading perceptual quality. Through experiments, we verified that Diff-TTS generates 28 times faster than the real-time with a single NVIDIA 2080Ti GPU.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Autonomous Collaborative Learning Among an Ensemble of Tsetlin Machines with Consensus-Based Inference

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A two-layer Tsetlin Machine ensemble with gossip-based vote sharing matches centralized accuracy on several benchmarks without exchanging raw data.

  2. DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration

    cs.SD 2025-09 conditional novelty 6.0 of 10

    DiTReducio is a training-free, pattern-guided layer and branch skipping method that accelerates DiT-based TTS, reporting significant FLOP and RTF reductions with modest quality loss at tuned thresholds.

  3. CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning

    cs.SD 2025-05 reject novelty 5.0 of 10

    A universal adversarial perturbation framework claiming to protect speech against zero-shot voice cloning by degrading cloned outputs while preserving input naturalness.

  4. Survey on AI-Generated Media Detection: From Non-MLLM to MLLM

    cs.CV 2025-02 unverdicted novelty 3.0 of 10

    A survey organizing AI-generated media detection into Non-MLLM and MLLM based methods, with task and benchmark taxonomies.

Pith tools