REVIEW 4 cited by
AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opinion scores (MOS) of synthesized speech using a deep recurrent neural network whose inputs consist solely of a raw waveform. Our best models provide utterance-level estimates of MOS only moderately inferior to sampled human ratings, as shown by Pearson and Spearman correlations. When multiple utterances are scored and averaged, a scenario common in synthesizer quality assessment, AutoMOS achieves correlations approaching those of human raters. The AutoMOS model has a number of applications, such as the ability to explore the parameter space of a speech synthesizer without requiring a human-in-the-loop.
Forward citations
Cited by 4 Pith papers
-
Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
Unified no-reference models assess audio aesthetics across speech, music, and sound via four perceptual axes and achieve performance comparable or superior to human mean opinion scores.
-
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
A unified taxonomy and comparative meta-evaluation of automatic evaluation methods across text, vision, and speech generation, concluding that LLM-based evaluators dominate current practice.
-
SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit
SHEET provides unified training and evaluation for MOS predictors, and its benchmark shows WavLM large and XLS-R 1b are the best SSL backbones for SSL-MOS on the tested datasets.
-
SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction
SALF-MOS, a compact U-Net-style model using frozen wav2vec features, claims state-of-the-art MOS prediction on four benchmarks with only 1,574 parameters.
Discussion (0). Continue with ORCID to comment.