Pith. sign in

REVIEW 2 cited by

Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.12719 v3 pith:HOVY3XCF submitted 2024-11-19 cs.CL cs.LGcs.SDeess.AS

classification cs.CLcs.LGcs.SDeess.AS
keywords humanmushrareferenceevaluationsystemstestacrossambiguity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite rapid advancements in TTS models, a consistent and robust human evaluation framework is still lacking. For example, MOS tests fail to differentiate between similar models, and CMOS's pairwise comparisons are time-intensive. The MUSHRA test is a promising alternative for evaluating multiple TTS systems simultaneously, but in this work we show that its reliance on matching human reference speech unduly penalises the scores of modern TTS systems that can exceed human speech quality. More specifically, we conduct a comprehensive assessment of the MUSHRA test, focusing on its sensitivity to factors such as rater variability, listener fatigue, and reference bias. Based on our extensive evaluation involving 492 human listeners across Hindi and Tamil we identify two primary shortcomings: (i) reference-matching bias, where raters are unduly influenced by the human reference, and (ii) judgement ambiguity, arising from a lack of clear fine-grained guidelines. To address these issues, we propose two refined variants of the MUSHRA test. The first variant enables fairer ratings for synthesized samples that surpass human reference quality. The second variant reduces ambiguity, as indicated by the relatively lower variance across raters. By combining these approaches, we achieve both more reliable and more fine-grained assessments. We also release MANGO, a massive dataset of 246,000 human ratings, the first-of-its-kind collection for Indian languages, aiding in analyzing human preferences and developing automatic metrics for evaluating TTS systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis

    cs.SD 2025-07 conditional novelty 6.0 of 10

    ASV embeddings used to judge whether synthesized speech matches a target speaker mostly encode static spectral traits and miss rhythm, so the authors introduce U3D, a duration-distribution metric for speaker rhythm.

  2. Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuning the English F5-TTS model on small Indian-language datasets yields a near-human-quality polyglot TTS (IN-F5) with voice cloning, code-mixing, and zero-resource synthesis for Bhojpuri and Tulu.

Pith tools