Pith. sign in

REVIEW 5 cited by

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.16381 v1 pith:KF3QR6E7 submitted 2025-06-19 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords systemscomplexinstruction-followinginstructttsevalnatural-languageaccuratecontrolevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In modern speech synthesis, paralinguistic information--such as a speaker's vocal timbre, emotional state, and dynamic prosody--plays a critical role in conveying nuance beyond mere semantics. Traditional Text-to-Speech (TTS) systems rely on fixed style labels or inserting a speech prompt to control these cues, which severely limits flexibility. Recent attempts seek to employ natural-language instructions to modulate paralinguistic features, substantially improving the generalization of instruction-driven TTS models. Although many TTS systems now support customized synthesis via textual description, their actual ability to interpret and execute complex instructions remains largely unexplored. In addition, there is still a shortage of high-quality benchmarks and automated evaluation metrics specifically designed for instruction-based TTS, which hinders accurate assessment and iterative optimization of these models. To address these limitations, we introduce InstructTTSEval, a benchmark for measuring the capability of complex natural-language style control. We introduce three tasks, namely Acoustic-Parameter Specification, Descriptive-Style Directive, and Role-Play, including English and Chinese subsets, each with 1k test cases (6k in total) paired with reference audio. We leverage Gemini as an automatic judge to assess their instruction-following abilities. Our evaluation of accessible instruction-following TTS systems highlights substantial room for further improvement. We anticipate that InstructTTSEval will drive progress toward more powerful, flexible, and accurate instruction-following TTS.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

    eess.AS 2026-08 conditional novelty 6.0 of 10

    SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.

  2. Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A paired audit of three TTS systems shows that descriptor-aligned voice changes come with off-target acoustic shifts, and a candidate selector reduces these shifts at inference time.

  3. RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

    cs.SD 2026-07 conditional novelty 6.0 of 10

    No voice AI system dominates all capabilities; naturalness, expressiveness, identity stability, audio sensitivity, and transcription robustness vary independently, so voice AI should be evaluated as a multidimensional...

  4. Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Several LALM judges copy wrong specialist labels or lock onto presentation order instead of listening to audio, so aggregate human agreement overstates judge validity.

  5. Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Using audio-aware LLMs to judge event presence and temporal order as DPO rewards improves multi-event text-to-audio instruction following.

Pith tools