Pith. sign in

REVIEW 1 cited by

Vocal effort modeling in neural TTS for improving the intelligibility of synthetic speech in noise

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.10637 v2 pith:GBDYGWGG submitted 2022-03-20 eess.AS cs.SD

classification eess.AScs.SD
keywords speecheffortintelligibilityvocalnoisespectraltiltfactors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a neural text-to-speech (TTS) method that models natural vocal effort variation to improve the intelligibility of synthetic speech in the presence of noise. The method consists of first measuring the spectral tilt of unlabeled conventional speech data, and then conditioning a neural TTS model with normalized spectral tilt among other prosodic factors. Changing the spectral tilt parameter and keeping other prosodic factors unchanged enables effective vocal effort control at synthesis time independent of other prosodic factors. By extrapolation of the spectral tilt values beyond what has been seen in the original data, we can generate speech with high vocal effort levels, thus improving the intelligibility of speech in the presence of masking noise. We evaluate the intelligibility and quality of normal speech and speech with increased vocal effort in the presence of various masking noise conditions, and compare these to well-known speech intelligibility-enhancing algorithms. The evaluations show that the proposed method can improve the intelligibility of synthetic speech with little loss in speech quality.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Voice Conversion for Lombard Speaking Style with Implicit and Explicit Acoustic Feature Conditioning

    cs.SD 2025-07 conditional novelty 5.0 of 10

    Implicit Lombard-style conditioning via a style reconstruction loss gives voice-converted speech intelligibility gains comparable to explicit f0, spectral energy and spectral tilt conditioning on the Audio-Visual Lomb...

Pith tools