REVIEW 3 major objections 5 minor 51 references
Frozen self-supervised speech models with Transformers give the most stable automatic MOS scores across unseen datasets, nearly matching specialized metrics on a hard enhancement benchmark.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 13:55 UTC pith:HOQMGFNG
load-bearing objection Careful large-scale bake-off: frozen WavLM+Transformer is more robust OOD than fine-tuning, English purification helps, and they release the checkpoint; the 5 s/16 kHz pipeline is the usual field caveat, not a unique flaw. the 3 major comments →
Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Frozen WavLM embeddings combined with a deep Transformer encoder, trained on a purified English-only corpus of roughly 123,000 utterances, form the most stable cross-corpus MOS predictor: they achieve MSE 0.36 on the held-out URGENT 2024 human-MOS benchmark, stay competitive with 18 specialized SOTA metrics, and outperform both fine-tuned SSL (which overfits seen data) and the lighter ViViT spectral model on out-of-distribution sets.
What carries the argument
Systematic Leave-One-Dataset-Out (LODO) protocol that contrasts Frozen SSL-Transformer, Fine-Tuned SSL-Transformer, and ViViT-MFCC architectures on a broad 19-dataset corpus versus an English-only 17-dataset subset; the frozen path freezes WavLM Large, feeds its 1024-dim frames into a two-layer Transformer with mean pooling, and uses the resulting score to quantify the seen-versus-unseen generalization gap.
Load-bearing premise
The claim assumes that resampling every file to 16 kHz, cutting it into fixed 5-second zero-padded chunks, and pooling those chunks still keeps the perceptual cues human listeners used when they rated the original full-length recordings.
What would settle it
Retrain the identical English-only frozen SSL-Transformer without 5-second segmentation—using full-utterance embeddings or variable-length attention—and re-measure MSE on URGENT 2024; a clear rise in error would show the reported robustness is an artifact of the fixed-window pipeline.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript benchmarks three MOS predictors—frozen WavLM-Large + Transformer (SSL-FRZ), fine-tuned WavLM-Large + Transformer (SSL-FT), and a ViViT-MFCC + Transformer—on a consolidated 130 k-utterance corpus (19 datasets) and a purified English-only subset (17 datasets). A six-case Leave-One-Dataset-Out (LODO) protocol quantifies seen-versus-unseen generalization; the best English-only SSL-FRZ model is further compared with 18 ARECHO SOTA metrics. The central claim is that frozen SSL embeddings plus a deep Transformer yield the most stable cross-corpus solution, reaching MSE 0.36 on the held-out URGENT 2024 human-MOS set (versus 0.30 for a domain-optimized SOTA) while outperforming fine-tuned SSL on out-of-distribution data; English-only purification improves precision for all architectures.
Significance. If the reported rankings hold, the work supplies a carefully controlled, large-scale reference point for domain-shift studies in non-intrusive speech quality assessment—an area still dominated by narrow-domain models. The two-part corpus design, systematic LODO, dual utterance/system-level metrics, and public release of the English-only SSL-Transformer weights on Hugging Face are concrete contributions that other groups can build on. The finding that freezing a denoising-pretrained SSL backbone can be preferable to full fine-tuning for OOD robustness is practically useful and falsifiable.
major comments (3)
- §3.2 (and the pipeline in Fig. 1): every waveform is forcibly resampled to 16 kHz and cut/padded into fixed 5 s segments before either mean-pooling (SSL) or CLS aggregation (ViViT). Several source sets (VoiceMOS 2025 Track 3, URGENT) contain 24–48 kHz material and long-range artifacts (packet-loss bursts, high-band codec distortion) that human listeners originally rated. No ablation or discussion quantifies how much of the original perceptual cue set survives this irreversible standardization. Because the same pipeline is shared by every reported number, both the absolute URGENT MSE of 0.36 and the claim of “universal” assessment risk being properties of the standardized feature space rather than of the architectures. A short ablation (native-rate vs. 16 kHz, or variable-length vs. 5 s) or an explicit limitation paragraph is required for the central claim to transfer beyond the preproces
- §3.4 and Tables 4–5: the LODO protocol that underpins the generalization analysis evaluates only frozen SSL and ViViT; full fine-tuning is performed solely on the single full-corpus split (Table 3). Consequently the headline statement that “SSL-FRZ provides superior robustness on unseen distributions” rests on one held-out test set rather than on the systematic LODO design. Either (a) a reduced-budget LODO with FT (e.g., last-layer or LoRA fine-tuning) or (b) an explicit caveat that the FRZ-vs-FT OOD comparison is limited to the URGENT split is needed before the claim can be considered fully supported.
- Tables 3–6 report only point estimates. No standard deviations across seeds, bootstrap confidence intervals, or statistical significance tests accompany the MSE/LCC/SRCC/KTAU differences that drive the ranking of FRZ over FT and of the proposed model over most ARECHO metrics. Given that several pairwise gaps are small (e.g., 0.36 vs. 0.30 MSE on URGENT, or FRZ vs. FT correlation differences of 0.01–0.03), the absence of uncertainty quantification leaves open the possibility that the reported superiority is within noise. Adding at least multi-seed means ± std or a simple paired test on the primary URGENT and LODO numbers would make the central ranking claims load-bearing.
minor comments (5)
- Numerous typographical and orthographic inconsistencies remain (e.g., “sescond part”, “V oiceMOS”, “V oIP”, “challange”, “T otal”, “arecho_urgent_mos” capitalization). A careful proof-reading pass is needed.
- Figure 1 is dense; the two parallel training paths and the two LODO subsets are hard to parse at a glance. A simplified schematic or clearer color coding would help.
- §3.3.1 states that mean-pooling was chosen over a CLS token “to ensure that every segment contributes equally,” yet no quantitative comparison of the two aggregation strategies is supplied. A one-sentence ablation or reference would strengthen the design choice.
- Table 1 lists sample counts after deduplication, but the exact number of unique systems or listeners per dataset is not given; this information is useful for interpreting system-level correlations.
- The ARECHO comparison (Table 6) normalizes Audiobox PQ from [1,10] to [1,5]; the linear mapping should be stated explicitly in the table caption or §3.7.
Circularity Check
Empirical LODO/SOTA benchmarking paper with no derivation circularity; only a non-load-bearing self-citation of the authors' prior MFCC-CNN work.
specific steps
-
other
[Section 1 (Introduction), citation [12]]
"presenting a challenge for architectures like CNNs or vision transformer (ViT)-based models which typically require fixed-size inputs [12]."
Authors cite their own prior MFCC-CNN paper for the well-known fixed-size input issue. This is ordinary background self-citation; it supplies neither the SSL backbone choice, the LODO protocol, the URGENT numbers, nor any uniqueness claim, so it does not force or redefine the main results.
full rationale
The paper is a large-scale empirical architecture comparison (SSL-FRZ vs SSL-FT vs ViViT) trained on public MOS corpora and evaluated via held-out human labels (URGENT 2024) plus an independent multi-metric suite (ARECHO). No equation, loss, or protocol reduces a claimed prediction to a fitted input by construction; LODO isolates entire datasets, and the 0.36 MSE claim is a direct numerical comparison against external human MOS and published SOTA predictors. The sole self-reference ([12], authors' earlier fixed-size MFCC-CNN paper) appears only as background motivation for handling variable-length inputs and is not invoked as a uniqueness theorem, ansatz, or premise for the SSL-Transformer results or generalization conclusions. Central claims therefore rest on independent external benchmarks and are free of circular reduction.
Axiom & Free-Parameter Ledger
free parameters (4)
- segment length (5 s) and zero-padding policy
- Transformer depth / heads (2-layer 8-head for SSL, 6-layer for ViViT)
- learning rates (1e-5 backbone, 1e-4 head) and early-stopping patience 7
- 40 Mel-bands, FFT 512, hop 256 for MFCC path
axioms (3)
- domain assumption Human MOS labels collected under different listening protocols across 19 datasets are commensurable on a single 1–5 scale after simple averaging of replicas.
- domain assumption Resampling every waveform to 16 kHz on-the-fly preserves the perceptual cues that originally determined the human MOS.
- ad hoc to paper Mean pooling (SSL) or a single CLS token (ViViT) is a sufficient aggregation of frame-level features for utterance-level quality.
read the original abstract
Automatic Mean Opinion Score (MOS) prediction is essential for evaluating large-scale synthetic speech and audio enhancement systems, yet models frequently struggle with domain shift. This study presents a comprehensive benchmarking of three architectural frameworks: Frozen Self-Supervised Learning (SSL-FRZ), Fine-Tuned SSL (SSL-FT), and a Video Vision Transformer (ViViT). Evaluation is conducted in two phases: Part I utilizes a consolidated corpus of 130,000 samples across 19 diverse datasets, while Part II focuses on a purified 17-dataset English-only corpus. To assess robustness, a systematic Leave-One-Dataset-Out (LODO) protocol is employed to quantify the generalization gap between seen and unseen distributions. Finally, the top-performing model is benchmarked against 18 state-of-the-art (SOTA) metrics using the ARECHO framework. Results demonstrate that an English-only purified corpus consistently yields higher predictive precision across all architectures. While SSL-FT achieves the highest performance on seen validation data, the SSL-FRZ model provides superior robustness on unseen distributions, achieving a competitive Mean Squared Error (MSE) of 0.36 on the URGENT 2024 benchmark-closely matching domain-optimized SOTA metrics (MSE 0.30). Although the ViViT architecture remains below SSL-based models in total capacity, it delivers stable results in English-only trials. LODO results confirm that while models perform significantly better on seen samples, frozen SSL embeddings combined with deep Transformer encoders offer the most stable and scalable solution for universal speech quality assessment. To support further research, the top-performing English-only SSL-Transformer model and weights are made publicly available via Hugging Face.
Figures
Reference graph
Works this paper leans on
-
[1]
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lu ˇci´c, M., Schmid, C., 2021. Vivit: A video vision transformer, in: 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 6816–6826. doi:doi:10.1109/ICCV48922.2021.00676
-
[2]
Baba, K., Nakata, W., Saito, Y ., Saruwatari, H., 2024. The t05 system for the voicemos challenge 2024: Transfer learning from deep image classifier to naturalness mos prediction of high-quality synthetic speech, in: 2024 IEEE Spoken Language Technology Workshop (SLT), pp. 818–824. doi:doi:10.1109/SLT61566.2024.10832315
-
[3]
wav2vec 2.0: A framework for self-supervised learning of speech representations, in: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M.F., Lin, H.T
Baevski, A., Zhou, Y ., Mohamed, A., Auli, M., 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations, in: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M.F., Lin, H.T. (Eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-1...
2020
-
[4]
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Chen, S., Wang, C., Chen, Z., Wu, Y ., Liu, S., Chen, Z., Li, J., Kanda, N., Yoshioka, T., Xiao, X., Wu, J., Zhou, L., Ren, S., Qian, Y ., Qian, Y ., Wu, J., Zeng, M., Yu, X., Wei, F., 2022. Wavlm: Large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing 16, 1505–1518. doi:doi:10.1109/...
-
[5]
Chinen, M., Lim, F.S.C., Skoglund, J., Gureev, N., O’Gorman, F., Hines, A., 2020. Visqol v3: An open source production ready objective speech and audio metric, in: 2020 Twelfth International Conference on Quality of Multimedia Experience (QoMEX), pp. 1–6. doi:doi:10.1109/QoMEX48832.2020.9123150. M.O. Duman et al.:Preprint submitted to ElsevierPage 19 of 2...
-
[6]
Cooper, E., Huang, W.C., Tsao, Y ., Wang, H.M., Toda, T., Yamagishi, J., 2023. The voicemos challenge 2023: Zero-shot subjective speech quality prediction for multiple domains, in: 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pp. 1–7. doi:doi:10.1109/ASRU57964.2023.10389763
-
[7]
A review on subjective and objective evaluation of synthetic speech
Cooper, E., Huang, W.C., Tsao, Y ., Wang, H.m., Toda, T., Yamagishi, J., 2024. A review on subjective and objective evaluation of synthetic speech. Acoustical Science and Technology 45. doi:doi:10.1250/ast.e24.12. doi: 10.1250/ast.e24.12
-
[8]
Cooper, E., Yamagishi, J., 2021. How do V oices from Past Speech Synthesis Challenges Compare Today?, in: 11th ISCA Speech Synthesis Workshop (SSW 11), pp. 183–188. doi:doi:10.21437/SSW.2021-32
-
[9]
Das, R.K., Kinnunen, T., Huang, W.C., Ling, Z.H., Yamagishi, J., Yi, Z., Tian, X., Toda, T., 2020. Predictions of Subjective Ratings and Spoofing Assessments of V oice Conversion Challenge 2020 Submissions, in: Joint Workshop for the Blizzard Challenge and V oice Conversion Challenge 2020, pp. 99–120. doi:doi:10.21437/VCCBC.2020-15
-
[10]
Davis, S., Mermelstein, P., 1980. Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences. IEEE Transactions on Acoustics, Speech, and Signal Processing 28, 357–366. doi:doi:10.1109/TASSP.1980.1163420. doi: 10.1109/TASSP.1980.1163420
-
[11]
Diener, L., Purin, M., Sootla, S., Saabas, A., Aichner, R., Cutler, R., 2023. PLCMOS – A Data-driven Non-intrusive Metric for The Evaluation of Packet Loss Concealment Algorithms, in: Interspeech 2023, pp. 2533–2537. doi:doi:10.21437/Interspeech.2023-1532
-
[12]
Duman, M.O., Dirik, A.E., 2026. Holistic audio quality assessment: Fixed-size mfcc spectral features for cnn-based mos prediction, in: 2026 5th International Informatics and Software Engineering Conference (IISEC), pp. 468–472. doi:doi:10.1109/IISEC69317.2026.11418452
-
[13]
Harte, N., Gillen, E., Hines, A., 2015. Tcd-voip, a research database of degraded speech for assessing quality in voip applications, in: 2015 Seventh International Workshop on Quality of Multimedia Experience (QoMEX), pp. 1–6. doi:doi:10.1109/QoMEX.2015.7148100
-
[14]
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W.N., Bolte, B., Tsai, Y .H.H., Lakhotia, K., Salakhutdinov, R., Mohamed, A., 2021. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing 29, 3451–3460. doi:doi:10.1109/TASLP.2021.3122291. doi: 10.1109/TASLP.2021.3122291
-
[16]
Evaluation of objective quality measures for speech enhancement
Hu, Y ., Loizou, P.C., 2008b. Evaluation of objective quality measures for speech enhancement. IEEE Transactions on Audio, Speech, and Language Processing 16, 229–238. doi:doi:10.1109/TASL.2007.911054. doi: 10.1109/TASL.2007.911054
-
[17]
Mos-bench: Benchmarking generalization abilities of subjective speech quality assessment models
Huang, W.C., Cooper, E., Toda, T., 2024a. Mos-bench: Benchmarking generalization abilities of subjective speech quality assessment models. URL:https://arxiv.org/abs/2411.03715,arXiv:2411.03715
-
[18]
Huang, W.C., Cooper, E., Toda, T., 2025a. SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit, in: Interspeech 2025, pp. 2355–2359. doi:doi:10.21437/Interspeech.2025-1977
-
[19]
The V oiceMOS Challenge 2022, in: Interspeech 2022, pp
Huang, W.C., Cooper, E., Tsao, Y ., Wang, H.M., Toda, T., Yamagishi, J., 2022. The V oiceMOS Challenge 2022, in: Interspeech 2022, pp. 4536–4540. doi:doi:10.21437/Interspeech.2022-970
-
[20]
Huang, W.C., Fu, S.W., Cooper, E., Zezario, R.E., Toda, T., Wang, H.M., Yamagishi, J., Tsao, Y ., 2024b. The voicemos challenge 2024: Beyond speech quality prediction, in: 2024 IEEE Spoken Language Technology Workshop (SLT), pp. 803–810. doi:doi:10.1109/SLT61566.2024.10832295
-
[21]
Huang, W.C., Wang, H., Liu, C., Wu, Y .C., Tjandra, A., Hsu, W.N., Cooper, E., Qin, Y ., Toda, T., 2025b. The audiomos challenge 2025. URL:https://arxiv.org/abs/2509.01336,arXiv:2509.01336
Pith/arXiv arXiv 2025
-
[22]
Methods for subjective determination of transmission quality
ITU-T, 1996. Methods for subjective determination of transmission quality. Recommendation P.800. International Telecommunication Union (ITU-T). Geneva, Switzerland. URL:https://www.itu.int/rec/T-REC-P.800/en
1996
-
[23]
The Blizzard Challenge 2008, in: The Blizzard Challenge 2008, pp
Karaiskos, V ., King, S., Clark, R.A.J., Mayo, C., 2008. The Blizzard Challenge 2008, in: The Blizzard Challenge 2008, pp. 1–18. doi:doi:10.21437/Blizzard.2008-1
-
[24]
The Blizzard Challenge 2009, in: The Blizzard Challenge 2009, pp
King, S., Karaiskos, V ., 2009. The Blizzard Challenge 2009, in: The Blizzard Challenge 2009, pp. 1–24. doi:doi:10.21437/Blizzard.2009-1
-
[25]
The Blizzard Challenge 2010, in: The Blizzard Challenge 2010, pp
King, S., Karaiskos, V ., 2010. The Blizzard Challenge 2010, in: The Blizzard Challenge 2010, pp. 1–32. doi:doi:10.21437/Blizzard.2010-1
-
[26]
The Blizzard Challenge 2011, in: The Blizzard Challenge 2011, pp
King, S., Karaiskos, V ., 2011. The Blizzard Challenge 2011, in: The Blizzard Challenge 2011, pp. 1–10. doi:doi:10.21437/Blizzard.2011-1
-
[27]
The Blizzard Challenge 2012, in: The Blizzard Challenge 2012, pp
King, S., Karaiskos, V ., 2012. The Blizzard Challenge 2012, in: The Blizzard Challenge 2012, pp. 1–11. doi:doi:10.21437/Blizzard.2012-1
-
[28]
The Blizzard Challenge 2013, in: The Blizzard Challenge 2013, pp
King, S., Karaiskos, V ., 2013. The Blizzard Challenge 2013, in: The Blizzard Challenge 2013, pp. 1–12. doi:doi:10.21437/Blizzard.2013-1
-
[29]
The Blizzard Challenge 2016, in: The Blizzard Challenge 2016, pp
King, S., Karaiskos, V ., 2016. The Blizzard Challenge 2016, in: The Blizzard Challenge 2016, pp. 1–16. doi:doi:10.21437/Blizzard.2016-1
-
[30]
Back to the Future: Extending the Blizzard Challenge 2013, in: Interspeech 2022, pp
Le Maguer, S., King, S., Harte, N., 2022. Back to the Future: Extending the Blizzard Challenge 2013, in: Interspeech 2022, pp. 2378–2382. doi:doi:10.21437/Interspeech.2022-10633
-
[31]
Leglaive, S., Borne, L., Tzinis, E., Sadeghi, M., Fraticelli, M., Wisdom, S., Pariente, M., Pressnitzer, D., Hershey, J., 2023. The CHiME-7 UDASE task: Unsupervised domain adaptation for conversational speech enhancement, in: 7th International Workshop on Speech Processing in Everyday Environments (CHiME 2023), pp. 7–12. doi:doi:10.21437/CHiME.2023-2
-
[32]
Leglaive, S., Fraticelli, M., ElGhazaly, H., Borne, L., Sadeghi, M., Wisdom, S., Pariente, M., Hershey, J.R., Pressnitzer, D., Barker, J.P., 2025. Objective and subjective evaluation of speech enhancement methods in the udase task of the 7th CHiME challenge. Computer Speech & Language 89, 101685.https://doi.org/10.1016/j.csl.2024.101685
-
[33]
Icassp 2026 urgent speech enhancement challenge
Li, C., Wang, W., Sach, M., Zhang, W., Saijo, K., Cornell, S., Fu, Y ., Ni, Z., Fingscheidt, T., Watanabe, S., Qian, Y ., 2026. Icassp 2026 urgent speech enhancement challenge. URL:https://arxiv.org/abs/2601.13531,arXiv:2601.13531
arXiv 2026
-
[34]
SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis, in: Proc
Maniati, G., Vioni, A., Ellinas, N., Nikitaras, K., Klapsas, K., Sung, J.S., Jho, G., Chalamandaris, A., Tsiakoulis, P., 2022. SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis, in: Proc. Interspeech 2022, pp. 2388–2392. doi:doi:10.21437/Interspeech.2022-10922
-
[35]
Speech Quality Assessment through MOS using Non-Matching References, in: Interspeech 2022, pp
Manocha, P., Kumar, A., 2022. Speech Quality Assessment through MOS using Non-Matching References, in: Interspeech 2022, pp. 654–658. doi:doi:10.21437/Interspeech.2022-407. M.O. Duman et al.:Preprint submitted to ElsevierPage 20 of 21 SSL and ViViT for Audio MOS Prediction
-
[36]
Noresqa: A framework for speech quality assessment using non-matching references, in: Ranzato, M., Beygelzimer, A., Dauphin, Y ., Liang, P., Vaughan, J.W
Manocha, P., Xu, B., Kumar, A., 2021. Noresqa: A framework for speech quality assessment using non-matching references, in: Ranzato, M., Beygelzimer, A., Dauphin, Y ., Liang, P., Vaughan, J.W. (Eds.), Advances in Neural Information Processing Sys- tems, Curran Associates, Inc.. pp. 22363–22378. URL:https://proceedings.neurips.cc/paper_files/paper/2021/fil...
2021
-
[37]
TTSDS2: Resources and benchmark for evaluating human-quality text to speech systems, in: The Fourteenth International Conference on Learning Representations
Minixhofer, C., Klejch, O., Bell, P., 2026. TTSDS2: Resources and benchmark for evaluating human-quality text to speech systems, in: The Fourteenth International Conference on Learning Representations. URL:https://openreview.net/forum?id=uGai5lYHlV
2026
-
[38]
DNN No-Reference PSTN Speech Quality Prediction, in: Interspeech 2020, pp
Mittag, G., Cutler, R., Hosseinkashi, Y ., Revow, M., Srinivasan, S., Chande, N., Aichner, R., 2020. DNN No-Reference PSTN Speech Quality Prediction, in: Interspeech 2020, pp. 2867–2871. doi:doi:10.21437/Interspeech.2020-2760
-
[39]
Mittag, G., Naderi, B., Chehadi, A., Möller, S., 2021. NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets, in: Interspeech 2021, pp. 2127–2131. doi:doi:10.21437/Interspeech.2021-299
-
[40]
Ragano, A., Skoglund, J., Hines, A., 2024a. Nomad: Unsupervised learning of perceptual embeddings for speech enhancement and non- matching reference audio quality assessment, in: ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1011–1015. doi:doi:10.1109/ICASSP48485.2024.10448028
-
[41]
SCOREQ: Speech quality assessment with contrastive regression, in: The Thirty-eighth Annual Conference on Neural Information Processing Systems
Ragano, A., Skoglund, J., Hines, A., 2024b. SCOREQ: Speech quality assessment with contrastive regression, in: The Thirty-eighth Annual Conference on Neural Information Processing Systems. URL:https://openreview.net/forum?id=HDVsiUHQ1w
-
[42]
Reddy, C.K.A., Gopal, V ., Cutler, R., 2021. Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors, in: ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6493–6497. doi:doi:10.1109/ICASSP39728.2021.9414878
-
[43]
Reddy, C.K.A., Gopal, V ., Cutler, R., 2022. Dnsmos p.835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors, in: ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 886–890. doi:doi:10.1109/ICASSP43922.2022.9746108
-
[44]
Rix, A., Beerends, J., Hollier, M., Hekstra, A., 2001. Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs, in: 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.01CH37221), pp. 749–752 vol.2. doi:doi:10.1109/ICASSP.2001.941023
-
[45]
UTMOS: UTokyo-SaruLab System for V oiceMOS Challenge 2022, in: Interspeech 2022, pp
Saeki, T., Xin, D., Nakata, W., Koriyama, T., Takamichi, S., Saruwatari, H., 2022. UTMOS: UTokyo-SaruLab System for V oiceMOS Challenge 2022, in: Interspeech 2022, pp. 4521–4525. doi:doi:10.21437/Interspeech.2022-439
-
[46]
Interspeech 2025 URGENT Speech Enhancement Challenge, in: Interspeech 2025, pp
Saijo, K., Zhang, W., Cornell, S., Scheibler, R., Li, C., Ni, Z., Kumar, A., Sach, M., Fu, Y ., Wang, W., Fingscheidt, T., Watanabe, S., 2025. Interspeech 2025 URGENT Speech Enhancement Challenge, in: Interspeech 2025, pp. 858–862. doi:doi:10.21437/Interspeech.2025-1363
-
[47]
Shi, J., Cheng, Y ., Su, B.H., jin Shim, H., Tian, J., Cornell, S., Zhao, Y ., Arora, S., Watanabe, S., 2025. ARECHO: Autoregressive evaluation via chain-based hypothesis optimization for speech multi-metric estimation, in: The Thirty-ninth Annual Conference on Neural Information Processing Systems. URL:https://openreview.net/forum?id=P2yIMJP5b1
2025
-
[48]
Singmos: An extensive open-source singing voice dataset for mos prediction
Tang, Y ., Shi, J., Wu, Y ., Jin, Q., 2024. Singmos: An extensive open-source singing voice dataset for mos prediction. URL:https: //arxiv.org/abs/2406.10911,arXiv:2406.10911
Pith/arXiv arXiv 2024
-
[49]
Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound
Tjandra, A., Wu, Y .C., Guo, B., Hoffman, J., Ellis, B., Vyas, A., Shi, B., Chen, S., Le, M., Zacharov, N., Wood, C., Lee, A., Hsu, W.N., 2025. Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound. URL:https://arxiv.org/abs/2502. 05139,arXiv:2502.05139
Pith/arXiv arXiv 2025
-
[50]
Yi, Z., Huang, W.C., Tian, X., Yamagishi, J., Das, R.K., Kinnunen, T., Ling, Z.H., Toda, T., 2020. V oice Conversion Challenge 2020 — Intra- lingual semi-parallel and cross-lingual voice conversion —, in: Joint Workshop for the Blizzard Challenge and V oice Conversion Challenge 2020, pp. 80–98. doi:doi:10.21437/VCCBC.2020-14
-
[51]
Lessons Learned from the URGENT 2024 Speech Enhancement Challenge, in: Interspeech 2025, pp
Zhang, W., Saijo, K., Cornell, S., Scheibler, R., Li, C., Ni, Z., Kumar, A., Sach, M., Wang, W., Fu, Y ., Watanabe, S., Fingscheidt, T., Qian, Y ., 2025. Lessons Learned from the URGENT 2024 Speech Enhancement Challenge, in: Interspeech 2025, pp. 853–857. doi:doi:10.21437/Interspeech.2025-1246
-
[52]
Zhang, W., Scheibler, R., Saijo, K., Cornell, S., Li, C., Ni, Z., Pirklbauer, J., Sach, M., Watanabe, S., Fingscheidt, T., Qian, Y ., 2024. URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement, in: Interspeech 2024, pp. 4868–4872. doi:doi:10.21437/Interspeech.2024-1239. M.O. Duman et al.:Preprint submitted to ElsevierPag...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.