Pith. sign in

REVIEW 6 cited by

SingMOS: An extensive Open-Source Singing Voice Dataset for MOS Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.10911 v2 pith:2GVDN46I submitted 2024-06-16 cs.SD eess.AS

SingMOS: An extensive Open-Source Singing Voice Dataset for MOS Prediction

classification cs.SD eess.AS
keywords singingpredictionsingmosspeechdatasetdomaindatadatasets
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In speech generation tasks, human subjective ratings, usually referred to as the opinion score, are considered the "gold standard" for speech quality evaluation, with the mean opinion score (MOS) serving as the primary evaluation metric. Due to the high cost of human annotation, several MOS prediction systems have emerged in the speech domain, demonstrating good performance. These MOS prediction models are trained using annotations from previous speech-related challenges. However, compared to the speech domain, the singing domain faces data scarcity and stricter copyright protections, leading to a lack of high-quality MOS-annotated datasets for singing. To address this, we propose SingMOS, a high-quality and diverse MOS dataset for singing, covering a range of Chinese and Japanese datasets. These synthesized vocals are generated using state-of-the-art models in singing synthesis, conversion, or resynthesis tasks and are rated by professional annotators alongside real vocals. Data analysis demonstrates the diversity and reliability of our dataset. Additionally, we conduct further exploration on SingMOS, providing insights for singing MOS prediction and guidance for the continued expansion of SingMOS.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres

    cs.SD 2026-07 conditional novelty 6.0

    Current singing voice synthesis models fail to differentiate musical genres, defaulting to pop-like output regardless of input genre, unless given genre-specific fine-tuning data.

  2. CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance

    cs.SD 2025-09 unverdicted novelty 6.0

    CoMelSinger introduces a discrete token-based zero-shot SVS framework on MaskGCT with coarse-to-fine contrastive learning and an SVT module to improve melody control and reduce prosody leakage.

  3. MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models

    cs.SD 2024-11 unverdicted novelty 6.0

    MOS-Bench benchmark shows that existing SSQA models struggle with out-of-domain generalization and that training on multiple diverse datasets improves robustness.

  4. Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation

    eess.AS 2026-07 conditional novelty 5.0

    Frozen SSL-Transformer embeddings generalize better than fine-tuned SSL or ViViT for cross-corpus MOS prediction, matching specialized SOTA on URGENT 2024 with MSE 0.36.

  5. Listening Like a Judge: A Music-Aware Framework for Automatic Singing Performance Evaluation

    cs.SD 2026-06 unverdicted novelty 5.0

    MusicJudge is a modality-guided framework that performs block-aligned multimodal analysis for singing quality assessment by coupling lyrics with pitch-rhythm fidelity via multi-signal matching and Modality-Guided LoRA...

  6. Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

    eess.AS 2026-06 unverdicted novelty 5.0

    MOS models match humans on acoustic degradation but are insensitive to prosodic errors and show a double dissociation on speaker characteristics like mean F0 bias and insensitivity to rate and F0 variability.