Pith. sign in

REVIEW 2 cited by

MOSPC: MOS Prediction Based on Pairwise Comparison

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.10493 v1 pith:BJHO6SGH submitted 2023-06-18 cs.SD cs.CLeess.AS

classification cs.SDcs.CLeess.AS
keywords qualityspeechframeworkscoresmospcpredictionrankingscore
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As a subjective metric to evaluate the quality of synthesized speech, Mean opinion score~(MOS) usually requires multiple annotators to score the same speech. Such an annotation approach requires a lot of manpower and is also time-consuming. MOS prediction model for automatic evaluation can significantly reduce labor cost. In previous works, it is difficult to accurately rank the quality of speech when the MOS scores are close. However, in practical applications, it is more important to correctly rank the quality of synthesis systems or sentences than simply predicting MOS scores. Meanwhile, as each annotator scores multiple audios during annotation, the score is probably a relative value based on the first or the first few speech scores given by the annotator. Motivated by the above two points, we propose a general framework for MOS prediction based on pair comparison (MOSPC), and we utilize C-Mixup algorithm to enhance the generalization performance of MOSPC. The experiments on BVCC and VCC2018 show that our framework outperforms the baselines on most of the correlation coefficient metrics, especially on the metric KTAU related to quality ranking. And our framework also surpasses the strong baseline in ranking accuracy on each fine-grained segment. These results indicate that our framework contributes to improving the ranking accuracy of speech quality.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning to assess subjective impressions from speech

    cs.SD 2025-06 conditional novelty 5.0 of 10

    Models trained on comparison category ratings (CCR) achieve better preference-ranking accuracy than models trained on absolute category ratings (ACR) for subjective voice descriptor scoring.

  2. SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction

    cs.SD 2025-06 reject novelty 4.0 of 10

    SALF-MOS, a compact U-Net-style model using frozen wav2vec features, claims state-of-the-art MOS prediction on four benchmarks with only 1,574 parameters.

Pith tools