Pith. sign in

REVIEW 3 major objections 5 minor 51 references

Frozen self-supervised speech models with Transformers give the most stable automatic MOS scores across unseen datasets, nearly matching specialized metrics on a hard enhancement benchmark.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 13:55 UTC pith:HOQMGFNG

load-bearing objection Careful large-scale bake-off: frozen WavLM+Transformer is more robust OOD than fine-tuning, English purification helps, and they release the checkpoint; the 5 s/16 kHz pipeline is the usual field caveat, not a unique flaw. the 3 major comments →

arxiv 2607.10146 v1 pith:HOQMGFNG submitted 2026-07-11 eess.AS

Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation

classification eess.AS
keywords Mean Opinion ScoreMOS predictionself-supervised learningWavLMVideo Vision TransformerLeave-One-Dataset-Outspeech quality assessmentdomain shift
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Automatic MOS predictors routinely fail when the test speech comes from a different source than the training data. This paper trains three architectures—frozen SSL, fine-tuned SSL, and a spectral ViViT—on more than 130,000 clips from 19 datasets, then repeats the work on a cleaned English-only subset. Leave-one-dataset-out trials show that models score far better on seen distributions, yet the frozen SSL-Transformer remains the most reliable on truly unseen material, reaching MSE 0.36 on the human-labeled URGENT 2024 set (close to the 0.30 of domain-tuned SOTA). English-only purification improves precision for all three models. The result matters because large-scale TTS and enhancement systems need a single, public, domain-robust quality score rather than a new model for every corpus.

Core claim

Frozen WavLM embeddings combined with a deep Transformer encoder, trained on a purified English-only corpus of roughly 123,000 utterances, form the most stable cross-corpus MOS predictor: they achieve MSE 0.36 on the held-out URGENT 2024 human-MOS benchmark, stay competitive with 18 specialized SOTA metrics, and outperform both fine-tuned SSL (which overfits seen data) and the lighter ViViT spectral model on out-of-distribution sets.

What carries the argument

Systematic Leave-One-Dataset-Out (LODO) protocol that contrasts Frozen SSL-Transformer, Fine-Tuned SSL-Transformer, and ViViT-MFCC architectures on a broad 19-dataset corpus versus an English-only 17-dataset subset; the frozen path freezes WavLM Large, feeds its 1024-dim frames into a two-layer Transformer with mean pooling, and uses the resulting score to quantify the seen-versus-unseen generalization gap.

Load-bearing premise

The claim assumes that resampling every file to 16 kHz, cutting it into fixed 5-second zero-padded chunks, and pooling those chunks still keeps the perceptual cues human listeners used when they rated the original full-length recordings.

What would settle it

Retrain the identical English-only frozen SSL-Transformer without 5-second segmentation—using full-utterance embeddings or variable-length attention—and re-measure MSE on URGENT 2024; a clear rise in error would show the reported robustness is an artifact of the fixed-window pipeline.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript benchmarks three MOS predictors—frozen WavLM-Large + Transformer (SSL-FRZ), fine-tuned WavLM-Large + Transformer (SSL-FT), and a ViViT-MFCC + Transformer—on a consolidated 130 k-utterance corpus (19 datasets) and a purified English-only subset (17 datasets). A six-case Leave-One-Dataset-Out (LODO) protocol quantifies seen-versus-unseen generalization; the best English-only SSL-FRZ model is further compared with 18 ARECHO SOTA metrics. The central claim is that frozen SSL embeddings plus a deep Transformer yield the most stable cross-corpus solution, reaching MSE 0.36 on the held-out URGENT 2024 human-MOS set (versus 0.30 for a domain-optimized SOTA) while outperforming fine-tuned SSL on out-of-distribution data; English-only purification improves precision for all architectures.

Significance. If the reported rankings hold, the work supplies a carefully controlled, large-scale reference point for domain-shift studies in non-intrusive speech quality assessment—an area still dominated by narrow-domain models. The two-part corpus design, systematic LODO, dual utterance/system-level metrics, and public release of the English-only SSL-Transformer weights on Hugging Face are concrete contributions that other groups can build on. The finding that freezing a denoising-pretrained SSL backbone can be preferable to full fine-tuning for OOD robustness is practically useful and falsifiable.

major comments (3)
  1. §3.2 (and the pipeline in Fig. 1): every waveform is forcibly resampled to 16 kHz and cut/padded into fixed 5 s segments before either mean-pooling (SSL) or CLS aggregation (ViViT). Several source sets (VoiceMOS 2025 Track 3, URGENT) contain 24–48 kHz material and long-range artifacts (packet-loss bursts, high-band codec distortion) that human listeners originally rated. No ablation or discussion quantifies how much of the original perceptual cue set survives this irreversible standardization. Because the same pipeline is shared by every reported number, both the absolute URGENT MSE of 0.36 and the claim of “universal” assessment risk being properties of the standardized feature space rather than of the architectures. A short ablation (native-rate vs. 16 kHz, or variable-length vs. 5 s) or an explicit limitation paragraph is required for the central claim to transfer beyond the preproces
  2. §3.4 and Tables 4–5: the LODO protocol that underpins the generalization analysis evaluates only frozen SSL and ViViT; full fine-tuning is performed solely on the single full-corpus split (Table 3). Consequently the headline statement that “SSL-FRZ provides superior robustness on unseen distributions” rests on one held-out test set rather than on the systematic LODO design. Either (a) a reduced-budget LODO with FT (e.g., last-layer or LoRA fine-tuning) or (b) an explicit caveat that the FRZ-vs-FT OOD comparison is limited to the URGENT split is needed before the claim can be considered fully supported.
  3. Tables 3–6 report only point estimates. No standard deviations across seeds, bootstrap confidence intervals, or statistical significance tests accompany the MSE/LCC/SRCC/KTAU differences that drive the ranking of FRZ over FT and of the proposed model over most ARECHO metrics. Given that several pairwise gaps are small (e.g., 0.36 vs. 0.30 MSE on URGENT, or FRZ vs. FT correlation differences of 0.01–0.03), the absence of uncertainty quantification leaves open the possibility that the reported superiority is within noise. Adding at least multi-seed means ± std or a simple paired test on the primary URGENT and LODO numbers would make the central ranking claims load-bearing.
minor comments (5)
  1. Numerous typographical and orthographic inconsistencies remain (e.g., “sescond part”, “V oiceMOS”, “V oIP”, “challange”, “T otal”, “arecho_urgent_mos” capitalization). A careful proof-reading pass is needed.
  2. Figure 1 is dense; the two parallel training paths and the two LODO subsets are hard to parse at a glance. A simplified schematic or clearer color coding would help.
  3. §3.3.1 states that mean-pooling was chosen over a CLS token “to ensure that every segment contributes equally,” yet no quantitative comparison of the two aggregation strategies is supplied. A one-sentence ablation or reference would strengthen the design choice.
  4. Table 1 lists sample counts after deduplication, but the exact number of unique systems or listeners per dataset is not given; this information is useful for interpreting system-level correlations.
  5. The ARECHO comparison (Table 6) normalizes Audiobox PQ from [1,10] to [1,5]; the linear mapping should be stated explicitly in the table caption or §3.7.

Circularity Check

1 steps flagged

Empirical LODO/SOTA benchmarking paper with no derivation circularity; only a non-load-bearing self-citation of the authors' prior MFCC-CNN work.

specific steps
  1. other [Section 1 (Introduction), citation [12]]
    "presenting a challenge for architectures like CNNs or vision transformer (ViT)-based models which typically require fixed-size inputs [12]."

    Authors cite their own prior MFCC-CNN paper for the well-known fixed-size input issue. This is ordinary background self-citation; it supplies neither the SSL backbone choice, the LODO protocol, the URGENT numbers, nor any uniqueness claim, so it does not force or redefine the main results.

full rationale

The paper is a large-scale empirical architecture comparison (SSL-FRZ vs SSL-FT vs ViViT) trained on public MOS corpora and evaluated via held-out human labels (URGENT 2024) plus an independent multi-metric suite (ARECHO). No equation, loss, or protocol reduces a claimed prediction to a fitted input by construction; LODO isolates entire datasets, and the 0.36 MSE claim is a direct numerical comparison against external human MOS and published SOTA predictors. The sole self-reference ([12], authors' earlier fixed-size MFCC-CNN paper) appears only as background motivation for handling variable-length inputs and is not invoked as a uniqueness theorem, ansatz, or premise for the SSL-Transformer results or generalization conclusions. Central claims therefore rest on independent external benchmarks and are free of circular reduction.

Axiom & Free-Parameter Ledger

4 free parameters · 3 axioms · 0 invented entities

The work is empirical. It inherits standard assumptions of the MOS-prediction literature (human labels are ground truth, 16 kHz is sufficient, MSE/LCC/SRCC/KTAU are the right metrics) and introduces a small set of architectural and preprocessing choices that function as free parameters. No new physical or mathematical entities are postulated.

free parameters (4)
  • segment length (5 s) and zero-padding policy
    Chosen by the authors; longer or shorter windows would change the temporal context available to both ViViT and the SSL Transformer.
  • Transformer depth / heads (2-layer 8-head for SSL, 6-layer for ViViT)
    Architectural hyper-parameters selected without ablation against alternatives.
  • learning rates (1e-5 backbone, 1e-4 head) and early-stopping patience 7
    Training hyper-parameters that affect final validation and test numbers.
  • 40 Mel-bands, FFT 512, hop 256 for MFCC path
    Feature-extraction settings that define the ViViT input representation.
axioms (3)
  • domain assumption Human MOS labels collected under different listening protocols across 19 datasets are commensurable on a single 1–5 scale after simple averaging of replicas.
    Implicit throughout Section 3.1–3.2; no inter-lab calibration is performed.
  • domain assumption Resampling every waveform to 16 kHz on-the-fly preserves the perceptual cues that originally determined the human MOS.
    Stated in Section 3.2 as a compatibility requirement with WavLM; high-frequency content above 8 kHz is discarded.
  • ad hoc to paper Mean pooling (SSL) or a single CLS token (ViViT) is a sufficient aggregation of frame-level features for utterance-level quality.
    Architectural choice in Sections 3.3.1–3.3.2; alternatives (attention pooling, multi-token) are not compared.

pith-pipeline@v1.1.0-grok45 · 38089 in / 2802 out tokens · 24895 ms · 2026-07-14T13:55:59.060343+00:00 · methodology

0 comments
read the original abstract

Automatic Mean Opinion Score (MOS) prediction is essential for evaluating large-scale synthetic speech and audio enhancement systems, yet models frequently struggle with domain shift. This study presents a comprehensive benchmarking of three architectural frameworks: Frozen Self-Supervised Learning (SSL-FRZ), Fine-Tuned SSL (SSL-FT), and a Video Vision Transformer (ViViT). Evaluation is conducted in two phases: Part I utilizes a consolidated corpus of 130,000 samples across 19 diverse datasets, while Part II focuses on a purified 17-dataset English-only corpus. To assess robustness, a systematic Leave-One-Dataset-Out (LODO) protocol is employed to quantify the generalization gap between seen and unseen distributions. Finally, the top-performing model is benchmarked against 18 state-of-the-art (SOTA) metrics using the ARECHO framework. Results demonstrate that an English-only purified corpus consistently yields higher predictive precision across all architectures. While SSL-FT achieves the highest performance on seen validation data, the SSL-FRZ model provides superior robustness on unseen distributions, achieving a competitive Mean Squared Error (MSE) of 0.36 on the URGENT 2024 benchmark-closely matching domain-optimized SOTA metrics (MSE 0.30). Although the ViViT architecture remains below SSL-based models in total capacity, it delivers stable results in English-only trials. LODO results confirm that while models perform significantly better on seen samples, frozen SSL embeddings combined with deep Transformer encoders offer the most stable and scalable solution for universal speech quality assessment. To support further research, the top-performing English-only SSL-Transformer model and weights are made publicly available via Hugging Face.

Figures

Figures reproduced from arXiv: 2607.10146 by Ahmet Emir Dirik, Mustafa Ozan Duman.

Figure 1
Figure 1. Figure 1: Proposed experimental pipeline for large-scale MOS prediction and multi-metric benchmarking. The framework illustrates the flow from data acquisition and deduplication to parallel model training (SSL-Transformer and ViViT) and final evaluation against ARECHO SOTA. challenge for architectures like CNNs or vision transformer (ViT)-based models which typically require fixed-size inputs [12]. To address these … view at source ↗
Figure 2
Figure 2. Figure 2: Architecture of the proposed SSL-Transformer Encoder framework. segment of the audio contributes equally to the final judgment, providing a more balanced assessment of long-duration signals. Finally, this context vector is passed to a Linear Regression layer with 1024 input nodes and a single output node, which maps the aggregated features to a continuous MOS prediction on a scale of 1.0 to 5.0. The archit… view at source ↗
Figure 3
Figure 3. Figure 3: Architecture of the proposed ViViT-Transformer Encoder framework. architecture, every token—including the CLS token—remains 768-dimensional as it passes through the six layers of the encoder. Within each layer, the CLS token interacts with all spectral patches through Multi-Head Self-Attention, effectively aggregating the most salient quality features from the entire multi-segment sequence into its own vec… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

51 extracted references · 21 canonical work pages

  1. [1]

    Vivit: A video vision transformer, in: 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp

    Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lu ˇci´c, M., Schmid, C., 2021. Vivit: A video vision transformer, in: 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 6816–6826. doi:doi:10.1109/ICCV48922.2021.00676

  2. [2]

    Baba, K., Nakata, W., Saito, Y ., Saruwatari, H., 2024. The t05 system for the voicemos challenge 2024: Transfer learning from deep image classifier to naturalness mos prediction of high-quality synthetic speech, in: 2024 IEEE Spoken Language Technology Workshop (SLT), pp. 818–824. doi:doi:10.1109/SLT61566.2024.10832315

  3. [3]

    wav2vec 2.0: A framework for self-supervised learning of speech representations, in: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M.F., Lin, H.T

    Baevski, A., Zhou, Y ., Mohamed, A., Auli, M., 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations, in: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M.F., Lin, H.T. (Eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-1...

  4. [4]

    Wavlm: Large-scale self-supervised pre-training for full stack speech processing

    Chen, S., Wang, C., Chen, Z., Wu, Y ., Liu, S., Chen, Z., Li, J., Kanda, N., Yoshioka, T., Xiao, X., Wu, J., Zhou, L., Ren, S., Qian, Y ., Qian, Y ., Wu, J., Zeng, M., Yu, X., Wei, F., 2022. Wavlm: Large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing 16, 1505–1518. doi:doi:10.1109/...

  5. [5]

    Visqol v3: An open source production ready objective speech and audio metric, in: 2020 Twelfth International Conference on Quality of Multimedia Experience (QoMEX), pp

    Chinen, M., Lim, F.S.C., Skoglund, J., Gureev, N., O’Gorman, F., Hines, A., 2020. Visqol v3: An open source production ready objective speech and audio metric, in: 2020 Twelfth International Conference on Quality of Multimedia Experience (QoMEX), pp. 1–6. doi:doi:10.1109/QoMEX48832.2020.9123150. M.O. Duman et al.:Preprint submitted to ElsevierPage 19 of 2...

  6. [6]

    The voicemos challenge 2023: Zero-shot subjective speech quality prediction for multiple domains, in: 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pp

    Cooper, E., Huang, W.C., Tsao, Y ., Wang, H.M., Toda, T., Yamagishi, J., 2023. The voicemos challenge 2023: Zero-shot subjective speech quality prediction for multiple domains, in: 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pp. 1–7. doi:doi:10.1109/ASRU57964.2023.10389763

  7. [7]

    A review on subjective and objective evaluation of synthetic speech

    Cooper, E., Huang, W.C., Tsao, Y ., Wang, H.m., Toda, T., Yamagishi, J., 2024. A review on subjective and objective evaluation of synthetic speech. Acoustical Science and Technology 45. doi:doi:10.1250/ast.e24.12. doi: 10.1250/ast.e24.12

  8. [8]

    How do V oices from Past Speech Synthesis Challenges Compare Today?, in: 11th ISCA Speech Synthesis Workshop (SSW 11), pp

    Cooper, E., Yamagishi, J., 2021. How do V oices from Past Speech Synthesis Challenges Compare Today?, in: 11th ISCA Speech Synthesis Workshop (SSW 11), pp. 183–188. doi:doi:10.21437/SSW.2021-32

  9. [9]

    Das, R.K., Kinnunen, T., Huang, W.C., Ling, Z.H., Yamagishi, J., Yi, Z., Tian, X., Toda, T., 2020. Predictions of Subjective Ratings and Spoofing Assessments of V oice Conversion Challenge 2020 Submissions, in: Joint Workshop for the Blizzard Challenge and V oice Conversion Challenge 2020, pp. 99–120. doi:doi:10.21437/VCCBC.2020-15

  10. [10]

    Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences

    Davis, S., Mermelstein, P., 1980. Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences. IEEE Transactions on Acoustics, Speech, and Signal Processing 28, 357–366. doi:doi:10.1109/TASSP.1980.1163420. doi: 10.1109/TASSP.1980.1163420

  11. [11]

    PLCMOS – A Data-driven Non-intrusive Metric for The Evaluation of Packet Loss Concealment Algorithms, in: Interspeech 2023, pp

    Diener, L., Purin, M., Sootla, S., Saabas, A., Aichner, R., Cutler, R., 2023. PLCMOS – A Data-driven Non-intrusive Metric for The Evaluation of Packet Loss Concealment Algorithms, in: Interspeech 2023, pp. 2533–2537. doi:doi:10.21437/Interspeech.2023-1532

  12. [12]

    Duman, M.O., Dirik, A.E., 2026. Holistic audio quality assessment: Fixed-size mfcc spectral features for cnn-based mos prediction, in: 2026 5th International Informatics and Software Engineering Conference (IISEC), pp. 468–472. doi:doi:10.1109/IISEC69317.2026.11418452

  13. [13]

    Tcd-voip, a research database of degraded speech for assessing quality in voip applications, in: 2015 Seventh International Workshop on Quality of Multimedia Experience (QoMEX), pp

    Harte, N., Gillen, E., Hines, A., 2015. Tcd-voip, a research database of degraded speech for assessing quality in voip applications, in: 2015 Seventh International Workshop on Quality of Multimedia Experience (QoMEX), pp. 1–6. doi:doi:10.1109/QoMEX.2015.7148100

  14. [14]

    Hubert: Self-supervised speech representation learning by masked prediction of hidden units

    Hsu, W.N., Bolte, B., Tsai, Y .H.H., Lakhotia, K., Salakhutdinov, R., Mohamed, A., 2021. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing 29, 3451–3460. doi:doi:10.1109/TASLP.2021.3122291. doi: 10.1109/TASLP.2021.3122291

  15. [16]

    Evaluation of objective quality measures for speech enhancement

    Hu, Y ., Loizou, P.C., 2008b. Evaluation of objective quality measures for speech enhancement. IEEE Transactions on Audio, Speech, and Language Processing 16, 229–238. doi:doi:10.1109/TASL.2007.911054. doi: 10.1109/TASL.2007.911054

  16. [17]

    Mos-bench: Benchmarking generalization abilities of subjective speech quality assessment models

    Huang, W.C., Cooper, E., Toda, T., 2024a. Mos-bench: Benchmarking generalization abilities of subjective speech quality assessment models. URL:https://arxiv.org/abs/2411.03715,arXiv:2411.03715

  17. [18]

    SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit, in: Interspeech 2025, pp

    Huang, W.C., Cooper, E., Toda, T., 2025a. SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit, in: Interspeech 2025, pp. 2355–2359. doi:doi:10.21437/Interspeech.2025-1977

  18. [19]

    The V oiceMOS Challenge 2022, in: Interspeech 2022, pp

    Huang, W.C., Cooper, E., Tsao, Y ., Wang, H.M., Toda, T., Yamagishi, J., 2022. The V oiceMOS Challenge 2022, in: Interspeech 2022, pp. 4536–4540. doi:doi:10.21437/Interspeech.2022-970

  19. [20]

    The voicemos challenge 2024: Beyond speech quality prediction, in: 2024 IEEE Spoken Language Technology Workshop (SLT), pp

    Huang, W.C., Fu, S.W., Cooper, E., Zezario, R.E., Toda, T., Wang, H.M., Yamagishi, J., Tsao, Y ., 2024b. The voicemos challenge 2024: Beyond speech quality prediction, in: 2024 IEEE Spoken Language Technology Workshop (SLT), pp. 803–810. doi:doi:10.1109/SLT61566.2024.10832295

  20. [21]

    The audiomos challenge 2025

    Huang, W.C., Wang, H., Liu, C., Wu, Y .C., Tjandra, A., Hsu, W.N., Cooper, E., Qin, Y ., Toda, T., 2025b. The audiomos challenge 2025. URL:https://arxiv.org/abs/2509.01336,arXiv:2509.01336

  21. [22]

    Methods for subjective determination of transmission quality

    ITU-T, 1996. Methods for subjective determination of transmission quality. Recommendation P.800. International Telecommunication Union (ITU-T). Geneva, Switzerland. URL:https://www.itu.int/rec/T-REC-P.800/en

  22. [23]

    The Blizzard Challenge 2008, in: The Blizzard Challenge 2008, pp

    Karaiskos, V ., King, S., Clark, R.A.J., Mayo, C., 2008. The Blizzard Challenge 2008, in: The Blizzard Challenge 2008, pp. 1–18. doi:doi:10.21437/Blizzard.2008-1

  23. [24]

    The Blizzard Challenge 2009, in: The Blizzard Challenge 2009, pp

    King, S., Karaiskos, V ., 2009. The Blizzard Challenge 2009, in: The Blizzard Challenge 2009, pp. 1–24. doi:doi:10.21437/Blizzard.2009-1

  24. [25]

    The Blizzard Challenge 2010, in: The Blizzard Challenge 2010, pp

    King, S., Karaiskos, V ., 2010. The Blizzard Challenge 2010, in: The Blizzard Challenge 2010, pp. 1–32. doi:doi:10.21437/Blizzard.2010-1

  25. [26]

    The Blizzard Challenge 2011, in: The Blizzard Challenge 2011, pp

    King, S., Karaiskos, V ., 2011. The Blizzard Challenge 2011, in: The Blizzard Challenge 2011, pp. 1–10. doi:doi:10.21437/Blizzard.2011-1

  26. [27]

    The Blizzard Challenge 2012, in: The Blizzard Challenge 2012, pp

    King, S., Karaiskos, V ., 2012. The Blizzard Challenge 2012, in: The Blizzard Challenge 2012, pp. 1–11. doi:doi:10.21437/Blizzard.2012-1

  27. [28]

    The Blizzard Challenge 2013, in: The Blizzard Challenge 2013, pp

    King, S., Karaiskos, V ., 2013. The Blizzard Challenge 2013, in: The Blizzard Challenge 2013, pp. 1–12. doi:doi:10.21437/Blizzard.2013-1

  28. [29]

    The Blizzard Challenge 2016, in: The Blizzard Challenge 2016, pp

    King, S., Karaiskos, V ., 2016. The Blizzard Challenge 2016, in: The Blizzard Challenge 2016, pp. 1–16. doi:doi:10.21437/Blizzard.2016-1

  29. [30]

    Back to the Future: Extending the Blizzard Challenge 2013, in: Interspeech 2022, pp

    Le Maguer, S., King, S., Harte, N., 2022. Back to the Future: Extending the Blizzard Challenge 2013, in: Interspeech 2022, pp. 2378–2382. doi:doi:10.21437/Interspeech.2022-10633

  30. [31]

    Leglaive, S., Borne, L., Tzinis, E., Sadeghi, M., Fraticelli, M., Wisdom, S., Pariente, M., Pressnitzer, D., Hershey, J., 2023. The CHiME-7 UDASE task: Unsupervised domain adaptation for conversational speech enhancement, in: 7th International Workshop on Speech Processing in Everyday Environments (CHiME 2023), pp. 7–12. doi:doi:10.21437/CHiME.2023-2

  31. [32]

    Objective and subjective evaluation of speech enhancement methods in the udase task of the 7th CHiME challenge

    Leglaive, S., Fraticelli, M., ElGhazaly, H., Borne, L., Sadeghi, M., Wisdom, S., Pariente, M., Hershey, J.R., Pressnitzer, D., Barker, J.P., 2025. Objective and subjective evaluation of speech enhancement methods in the udase task of the 7th CHiME challenge. Computer Speech & Language 89, 101685.https://doi.org/10.1016/j.csl.2024.101685

  32. [33]

    Icassp 2026 urgent speech enhancement challenge

    Li, C., Wang, W., Sach, M., Zhang, W., Saijo, K., Cornell, S., Fu, Y ., Ni, Z., Fingscheidt, T., Watanabe, S., Qian, Y ., 2026. Icassp 2026 urgent speech enhancement challenge. URL:https://arxiv.org/abs/2601.13531,arXiv:2601.13531

  33. [34]

    SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis, in: Proc

    Maniati, G., Vioni, A., Ellinas, N., Nikitaras, K., Klapsas, K., Sung, J.S., Jho, G., Chalamandaris, A., Tsiakoulis, P., 2022. SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis, in: Proc. Interspeech 2022, pp. 2388–2392. doi:doi:10.21437/Interspeech.2022-10922

  34. [35]

    Speech Quality Assessment through MOS using Non-Matching References, in: Interspeech 2022, pp

    Manocha, P., Kumar, A., 2022. Speech Quality Assessment through MOS using Non-Matching References, in: Interspeech 2022, pp. 654–658. doi:doi:10.21437/Interspeech.2022-407. M.O. Duman et al.:Preprint submitted to ElsevierPage 20 of 21 SSL and ViViT for Audio MOS Prediction

  35. [36]

    Noresqa: A framework for speech quality assessment using non-matching references, in: Ranzato, M., Beygelzimer, A., Dauphin, Y ., Liang, P., Vaughan, J.W

    Manocha, P., Xu, B., Kumar, A., 2021. Noresqa: A framework for speech quality assessment using non-matching references, in: Ranzato, M., Beygelzimer, A., Dauphin, Y ., Liang, P., Vaughan, J.W. (Eds.), Advances in Neural Information Processing Sys- tems, Curran Associates, Inc.. pp. 22363–22378. URL:https://proceedings.neurips.cc/paper_files/paper/2021/fil...

  36. [37]

    TTSDS2: Resources and benchmark for evaluating human-quality text to speech systems, in: The Fourteenth International Conference on Learning Representations

    Minixhofer, C., Klejch, O., Bell, P., 2026. TTSDS2: Resources and benchmark for evaluating human-quality text to speech systems, in: The Fourteenth International Conference on Learning Representations. URL:https://openreview.net/forum?id=uGai5lYHlV

  37. [38]

    DNN No-Reference PSTN Speech Quality Prediction, in: Interspeech 2020, pp

    Mittag, G., Cutler, R., Hosseinkashi, Y ., Revow, M., Srinivasan, S., Chande, N., Aichner, R., 2020. DNN No-Reference PSTN Speech Quality Prediction, in: Interspeech 2020, pp. 2867–2871. doi:doi:10.21437/Interspeech.2020-2760

  38. [39]

    NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets, in: Interspeech 2021, pp

    Mittag, G., Naderi, B., Chehadi, A., Möller, S., 2021. NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets, in: Interspeech 2021, pp. 2127–2131. doi:doi:10.21437/Interspeech.2021-299

  39. [40]

    Ragano, A., Skoglund, J., Hines, A., 2024a. Nomad: Unsupervised learning of perceptual embeddings for speech enhancement and non- matching reference audio quality assessment, in: ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1011–1015. doi:doi:10.1109/ICASSP48485.2024.10448028

  40. [41]

    SCOREQ: Speech quality assessment with contrastive regression, in: The Thirty-eighth Annual Conference on Neural Information Processing Systems

    Ragano, A., Skoglund, J., Hines, A., 2024b. SCOREQ: Speech quality assessment with contrastive regression, in: The Thirty-eighth Annual Conference on Neural Information Processing Systems. URL:https://openreview.net/forum?id=HDVsiUHQ1w

  41. [42]

    Reddy, C.K.A., Gopal, V ., Cutler, R., 2021. Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors, in: ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6493–6497. doi:doi:10.1109/ICASSP39728.2021.9414878

  42. [43]

    Reddy, C.K.A., Gopal, V ., Cutler, R., 2022. Dnsmos p.835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors, in: ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 886–890. doi:doi:10.1109/ICASSP43922.2022.9746108

  43. [44]

    Rix, A., Beerends, J., Hollier, M., Hekstra, A., 2001. Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs, in: 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.01CH37221), pp. 749–752 vol.2. doi:doi:10.1109/ICASSP.2001.941023

  44. [45]

    UTMOS: UTokyo-SaruLab System for V oiceMOS Challenge 2022, in: Interspeech 2022, pp

    Saeki, T., Xin, D., Nakata, W., Koriyama, T., Takamichi, S., Saruwatari, H., 2022. UTMOS: UTokyo-SaruLab System for V oiceMOS Challenge 2022, in: Interspeech 2022, pp. 4521–4525. doi:doi:10.21437/Interspeech.2022-439

  45. [46]

    Interspeech 2025 URGENT Speech Enhancement Challenge, in: Interspeech 2025, pp

    Saijo, K., Zhang, W., Cornell, S., Scheibler, R., Li, C., Ni, Z., Kumar, A., Sach, M., Fu, Y ., Wang, W., Fingscheidt, T., Watanabe, S., 2025. Interspeech 2025 URGENT Speech Enhancement Challenge, in: Interspeech 2025, pp. 858–862. doi:doi:10.21437/Interspeech.2025-1363

  46. [47]

    Shi, J., Cheng, Y ., Su, B.H., jin Shim, H., Tian, J., Cornell, S., Zhao, Y ., Arora, S., Watanabe, S., 2025. ARECHO: Autoregressive evaluation via chain-based hypothesis optimization for speech multi-metric estimation, in: The Thirty-ninth Annual Conference on Neural Information Processing Systems. URL:https://openreview.net/forum?id=P2yIMJP5b1

  47. [48]

    Singmos: An extensive open-source singing voice dataset for mos prediction

    Tang, Y ., Shi, J., Wu, Y ., Jin, Q., 2024. Singmos: An extensive open-source singing voice dataset for mos prediction. URL:https: //arxiv.org/abs/2406.10911,arXiv:2406.10911

  48. [49]

    Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound

    Tjandra, A., Wu, Y .C., Guo, B., Hoffman, J., Ellis, B., Vyas, A., Shi, B., Chen, S., Le, M., Zacharov, N., Wood, C., Lee, A., Hsu, W.N., 2025. Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound. URL:https://arxiv.org/abs/2502. 05139,arXiv:2502.05139

  49. [50]

    Yi, Z., Huang, W.C., Tian, X., Yamagishi, J., Das, R.K., Kinnunen, T., Ling, Z.H., Toda, T., 2020. V oice Conversion Challenge 2020 — Intra- lingual semi-parallel and cross-lingual voice conversion —, in: Joint Workshop for the Blizzard Challenge and V oice Conversion Challenge 2020, pp. 80–98. doi:doi:10.21437/VCCBC.2020-14

  50. [51]

    Lessons Learned from the URGENT 2024 Speech Enhancement Challenge, in: Interspeech 2025, pp

    Zhang, W., Saijo, K., Cornell, S., Scheibler, R., Li, C., Ni, Z., Kumar, A., Sach, M., Wang, W., Fu, Y ., Watanabe, S., Fingscheidt, T., Qian, Y ., 2025. Lessons Learned from the URGENT 2024 Speech Enhancement Challenge, in: Interspeech 2025, pp. 853–857. doi:doi:10.21437/Interspeech.2025-1246

  51. [52]

    URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement, in: Interspeech 2024, pp

    Zhang, W., Scheibler, R., Saijo, K., Cornell, S., Li, C., Ni, Z., Pirklbauer, J., Sach, M., Watanabe, S., Fingscheidt, T., Qian, Y ., 2024. URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement, in: Interspeech 2024, pp. 4868–4872. doi:doi:10.21437/Interspeech.2024-1239. M.O. Duman et al.:Preprint submitted to ElsevierPag...