Pith. sign in

REVIEW 2 cited by

Prosody Analysis of Audiobooks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.06930 v3 pith:Z4EGDTLO submitted 2023-10-10 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords humanaudiobookpredictedprosodybookscommercialnarrativepitch
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent advances in text-to-speech have made it possible to generate natural-sounding audio from text. However, audiobook narrations involve dramatic vocalizations and intonations by the reader, with greater reliance on emotions, dialogues, and descriptions in the narrative. Using our dataset of 93 aligned book-audiobook pairs, we present improved models for prosody prediction properties (pitch, volume, and rate of speech) from narrative text using language modeling. Our predicted prosody attributes correlate much better with human audiobook readings than results from a state-of-the-art commercial TTS system: our predicted pitch shows a higher correlation with human reading for 22 out of the 24 books, while our predicted volume attribute proves more similar to human reading for 23 out of the 24 books. Finally, we present a human evaluation study to quantify the extent that people prefer prosody-enhanced audiobook readings over commercial text-to-speech systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving French Synthetic Speech Quality via SSML Prosody Control

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    Two fine-tuned LLMs predict SSML prosody tags that raise French TTS naturalness from a 3.20 to 3.87 MOS.

  2. Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech

    cs.CL 2025-07 unverdicted novelty 6.0 of 10

    Balalaika is a data-centric annotation pipeline for Russian speech that combines semantic VAD, ASR ensembling, and prosody enrichment to build a 5.1k-hour corpus showing gains in denoising and TTS.

Pith tools