Pith. sign in

REVIEW 2 major objections 4 minor 41 references

Fine-tuning on just 50 Singlish speakers teaches zero-shot TTS the accent, not just the voices.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 03:48 UTC pith:QXPFTMQ7

load-bearing objection First systematic Singlish TTS benchmark with a sound generalization split, but the accent-similarity metric is doing more work than it can support. the 2 major comments →

arxiv 2607.23027 v1 pith:QXPFTMQ7 submitted 2026-07-25 eess.AS

Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English

classification eess.AS
keywords Singlishzero-shot text-to-speechaccent adaptationaccent similarityfine-tuningvoice cloninglow-resource TTSSingapore English
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks whether targeted fine-tuning can make zero-shot text-to-speech reproduce Singapore English (Singlish) rather than flattening it toward generic English. It fine-tunes two recent zero-shot TTS systems on 50 Singlish speakers from a national speech corpus and compares them with their off-the-shelf versions using identical audio prompts. The central claim is that fine-tuning substantially closes the accent gap: matched accent similarity rises by up to 0.13 in-domain, and the gain persists on 42 held-out speakers, showing the models learn the accent rather than memorise voices. The two backbones fail and recover differently, with one trading over-clean neutral English for authentic Singlish delivery while the other's fine-tuning also repairs content errors.

Core claim

On the paper's own terms, the discovery is that a modest fine-tuning corpus — about an hour per speaker for 50 speakers — is enough to shift the output distribution of two zero-shot TTS systems measurably toward real Singlish. The accent-similarity score against matched ground-truth recordings rises from 0.511 to 0.638 for the weaker baseline and from 0.577 to 0.604 for the one that starts closer, and the gain does not disappear when the models are asked to clone 42 voices they never heard during training. The authors take this persistence as evidence that the adaptation captures the accent itself, not just the training voices. They also find that conditioning on a learned per-speaker index

What carries the argument

The load-bearing mechanism is the adaptation-versus-consistency split. Fifty speakers are seen during fine-tuning and 42 are held out, a per-speaker split ensures no speaker overlap, and each utterance is synthesised from a single reference prompt whose content is never the synthesis target. The headline metric is accent similarity: the cosine distance between generated and matched real speech in an accent-classifier embedding space, treated as an accent-fidelity score because that classifier separates accents at high accuracy. On top of this, fine-tuning is deliberately limited to each system's text-to-token pathway — the autoregressive module in one model, the language-model and flow-match

Load-bearing premise

The headline metric assumes that cosine similarity in an accent-classifier embedding space measures whether a listener would judge the speech as Singlish; if that space is dominated by voice or recording-quality differences, the measured accent gains would not establish real accent transfer.

What would settle it

Present fine-tuned and off-the-shelf outputs to Singapore English listeners in an accent-rating or forced-choice test; if perceived Singlish-ness does not track the ACC-SIM gains, the metric is measuring the wrong thing. A second check: regress speaker embeddings out of the accent embeddings — if the 'accent' gains collapse, they are confounded with voice similarity.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A few dozen speakers — not hundreds of hours — can shift a zero-shot TTS system toward an under-resourced accent, lowering the data bar for accent adaptation.
  • Because gains persist on held-out speakers, the fine-tuning procedure can be reused to build accent-specific TTS without voice-memorisation artefacts.
  • The two backbones trade off differently: one sacrifices over-clean intelligibility for authentic accent, while the other's fine-tuning also repairs content hallucinations.
  • A speaker-index conditioning mode can exceed audio-prompt accent fidelity on seen speakers, at the cost of prompt faithfulness and open-set use.
  • The same protocol — matched utterance evaluation plus in-/out-of-domain split — can benchmark accent fidelity for other regional English varieties.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the headline ACC-SIM metric is only as valid as its embedding space; if that space mostly encodes timbre or recording conditions, the reported gains would not prove perceived Singlish transfer. A native-listener accent test would settle this.
  • Editorial inference: the observed inverse relationship — weaker initial accent fidelity leads to larger fine-tuning gains — suggests the gap is largely a coverage problem, so scaling pretraining with accent-diverse data might reduce or eliminate the need for fine-tuning.
  • Editorial inference: the success of speaker-index conditioning hints at a controllable middle ground, e.g., interpolating between prompt and index conditioning, which could give users an accent-strength dial while retaining open-set cloning.
  • Editorial inference: the dataset-filtering pipeline (WER thresholding, quality screening, gender balancing) could be transferred to other low-resource accents, but the paper does not test whether its 50-speaker corpus size is minimal; an ablation on speaker count would be needed.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper investigates whether fine-tuning two zero-shot TTS systems, Chatterbox and CosyVoice 3, on 50 Singapore English (Singlish) speakers from the IMDA National Speech Corpus closes the accent gap between off-the-shelf and real Singlish speech. It evaluates three distributions (real, off-the-shelf, fine-tuned) with the CodecMOS-Accent protocol (UT-MOS, WER, SPK-SIM, ACC-SIM), separates in-domain (seen) from out-of-domain (unseen) speakers, and reports that fine-tuning raises accent similarity on both, concluding that models learn the accent rather than memorize voices.

Significance. If the central claim holds, the work is a useful low-resource accent-adaptation study: it demonstrates that modest fine-tuning of two modern ZS-TTS backbones can shift output toward a regional English variety and that the effect persists on held-out speakers. The in-domain/out-of-domain design is a principled way to separate voice memorization from accent generalization, and the use of external evaluation models and a public corpus supports reproducibility. However, the headline ACC-SIM metric is an unvalidated construct, and the absence of statistical inference makes the magnitude of the reported gains uncertain.

major comments (2)
  1. [Section V, Table V] All reported gains are single measurements on a fixed split. There are no confidence intervals, significance tests, or per-speaker variances. The headline improvements (+0.126, +0.092, +0.027, +0.028 in ACC-SIM match) could be within run-to-run or speaker-derived noise, especially on the smaller out-of-domain set (42 speakers). Please provide bootstrap CIs or per-speaker standard errors, and ideally multiple fine-tuning runs. This is standard for TTS evaluation and is necessary to support the conclusion that the gains are reliable.
  2. [Section V, Table V] The observation that F-Chatterbox's prompt ACC-SIM (0.6748) exceeds the Ground Truth ceiling (0.6277) is internally inconsistent with treating ACC-SIM as a faithfulness score: generated speech should not be 'more accent-faithful' than a real same-speaker recording. The paper interprets this as drift toward a 'generalized Singlish voice,' but this also indicates that the prompt variant of ACC-SIM is not a bounded or well-calibrated measure. This anomaly should be explained or the prompt ACC-SIM should be de-emphasized; otherwise it weakens the construct validity of the headline metric.
minor comments (4)
  1. [Section III] The WER filtering threshold of 50% is described as based on manual inspection but no sensitivity analysis is given. A sentence reporting the retained utterance count and its dependence on the threshold (e.g., 80% or 30%) would help the reader judge robustness.
  2. [Section IV-D] Chatterbox hyperparameters (Table IV) are said to be 'selected on the in-domain validation split' but the selection criterion is not described. Also, the text says training loss decreases from 6.0 to below 10^-3, which is a very wide range; please clarify the loss metric and whether this is per-token cross-entropy.
  3. [Section V] The sentence 'F-CosyVoice is flat (0.6200 vs 0.6277, 0.5820 vs 0.5802)' for prompt ACC-SIM is unclear: the two comparisons need labels (in-domain/out-of-domain) to be interpretable.
  4. [References] Some references contain typographical artifacts (e.g., 'V oice', 'V ALL-E', 'Y . Wu'). Please proofread the bibliography.

Circularity Check

0 steps flagged

No significant circularity: the study is an empirical fine-tuning evaluation using external corpora and external pretrained metrics; the only self-citation is ancillary and not load-bearing.

full rationale

The paper's central claim is an empirical measurement, not a derived equation. Fine-tuning is performed on an external corpus (IMDA NSC), and all four evaluation metrics (UT-MOS, WER, SPK-SIM, ACC-SIM) come from external pretrained models or protocols (VoiceMOS, Singlish Whisper, ECAPA-TDNN, CommonAccent, CodecMOS-Accent). The in-domain/out-of-domain split directly addresses the memorize-vs-generalize question, and the generalization claim rests on held-out speakers rather than on the fine-tuning set. The only self-citation involving the present authors is reference [19], the RADAR 2026 challenge, which is cited only as a community-contribution statement and does not support any technical inference. The ACC-SIM construct-validity concern (whether the embedding primarily captures accent rather than timbre or channel) is a measurement-assumption risk, not a circularity: the metric was proposed and validated in external work [17], and the paper does not define Singlish-accent fidelity in terms of the metric it later 'predicts.' No equation is shown to reduce to its own inputs, and no fitted parameter is renamed as a prediction. Accordingly, the circularity burden is very low; the appropriate score is 1, reflecting only the presence of a minor, non-load-bearing self-citation.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 0 invented entities

All central-claim support rests on external pretrained models (CommonAccent, ECAPA-TDNN, Singlish Whisper, UT-MOS), the public IMDA NSC corpus, and hand-chosen experimental settings. There are no derived equations or fitted theoretical constants; the key load-bearing assumptions are measurement-proxy validity and dataset representativeness.

free parameters (3)
  • Chatterbox inference hyperparameters (temperature=0.75, cfg_weight=0.50, repetition_penalty=1.35, exaggeration=0.40) = temperature 0.75; cfg weight 0.50; repetition penalty 1.35; exaggeration 0.40
    Selected on the in-domain validation split (Table VI); these decoding settings affect the generated accent/similarity scores.
  • CosyVoice fine-tuning configuration (LR=5e-7, epochs=200, LLM+Flow Matching trainable, vocoder frozen) = LR 5e-7, 200 epochs
    Chosen by a single author comparing five validation samples by ear (Sec. IV-B), not by exhaustive search; influences all F-CosyVoice results.
  • WER filtering threshold (50%) = 50% WER
    Chosen from manual inspection (Sec. III); removes 14.3% of utterances and shapes the training/evaluation distribution.
axioms (5)
  • domain assumption Cosine similarity in CommonAccent ECAPA-TDNN embedding space is a valid proxy for perceived Singlish accent fidelity.
    Introduced in Sec. II-F as the headline ACC-SIM metric; the paper relies on 90% accent-classification accuracy of the encoder, but no human validation for Singlish is provided.
  • domain assumption Singlish-fine-tuned Whisper ASR WER is a valid intelligibility proxy for synthetic Singlish.
    Used in Sec. II-F for the intelligibility score; ASR errors may reflect transcriber quirks rather than listener intelligibility.
  • domain assumption IMDA NSC Part 3 speech, as filtered by the 50% WER threshold, represents colloquial Singlish of the target accent.
    Sec. III derives the training data from Part 3 of IMDA NSC; the corpus is conversational but contains a range of Singapore English registers.
  • domain assumption A 50-speaker, 55.5-hour in-domain pool is sufficient to learn accent rather than memorize speakers.
    The adaptation/consistency interpretation in Sec. V assumes gains on held-out speakers indicate accent generalization; dataset sufficiency is not independently established.
  • domain assumption CodecMOS-Accent objective protocol validated on other accented English transfers to Singlish.
    Adopted in Sec. II from [17]; the paper applies it to Singlish without re-validation.

pith-pipeline@v1.3.0-alltime-deepseek · 11053 in / 12328 out tokens · 116672 ms · 2026-08-01T03:48:39.395165+00:00 · methodology

0 comments
read the original abstract

Zero-shot text-to-speech (ZS-TTS) achieves near-human quality for standard English, but it copies regional accents poorly. Prompted with a short Singlish utterance, state-of-the-art systems reproduce a speaker's timbre while flattening the accent toward generic English. We investigate whether targeted fine-tuning off-the-shelf ZS-TTS can close the gap for Singapore English (Singlish). We fine-tune two cutting-edge ZS-TTS models, Chatterbox and CosyVoice 3, on 50 Singlish speakers from the IMDA National Speech Corpus. Three speech distributions are evaluated: real recordings against off-the-shelf and fine-tuned generation driven by the same Singlish audio prompts. The evaluation covers four dimensions: naturalness, intelligibility, speaker similarity, and accent similarity. We separate adaptation (in-domain speakers seen during fine-tuning) from consistency (held-out speakers) to test whether accent transfer generalises beyond the training data. Fine-tuning raises accent similarity on in-domain and out-of-domain speakers for both Chatterbox and CosyVoice 3. It moves the generated distribution measurably toward real Singlish, with the gain persisting on held-out speakers. To our knowledge, this is the first systematic study of Singlish-accented TTS.

Figures

Figures reproduced from arXiv: 2607.23027 by Ivan Kukanov, Zheng Xin Chai.

Figure 1
Figure 1. Figure 1: Off-the-shelf zero-shot text-to-speech models (Chatterbox, [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Three-phase experimental pipeline: fine-tuning TTS models on 50 [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 4
Figure 4. Figure 4: Speaker duration distribution of in-domain and out-of-domain sets [PITH_FULL_IMAGE:figures/full_fig_p003_4.png] view at source ↗
Figure 3
Figure 3. Figure 3: Utterance-based WER (%) (no text normalization) histogram of [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 5
Figure 5. Figure 5: In-/out-of-domain distribution across systems. [PITH_FULL_IMAGE:figures/full_fig_p005_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

41 extracted references · 8 linked inside Pith

  1. [1]

    Neural codec language models are zero- shot text to speech synthesizers,

    S. Chen, C. Wang, Y . Wuet al., “Neural codec language models are zero- shot text to speech synthesizers,”IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 705–718, 2025

  2. [2]

    V ALL-E 2: Neural codec language models are human parity zero-shot text to speech synthesizers

    S. Chen, S. Liu, L. Zhouet al., “V ALL-E 2: Neural codec language models are human parity zero-shot text to speech synthesizers.”CoRR, vol. abs/2406.05370, 2024

  3. [3]

    CosyV oice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic to- kens,

    Z. Du, Q. Chen, S. Zhanget al., “CosyV oice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic to- kens,”ArXiv, vol. abs/2407.05407, 2024

  4. [4]

    CosyV oice 2: Scalable streaming speech synthesis with large language models,

    Z. Du, Y . Wang, Q. Chenet al., “CosyV oice 2: Scalable streaming speech synthesis with large language models,”ArXiv, vol. abs/2412.10117, 2024

  5. [5]

    CosyV oice 3: Towards in-the-wild speech generation via scaling-up and post-training,

    Z. Du, C. Gao, Y . Wanget al., “CosyV oice 3: Towards in-the-wild speech generation via scaling-up and post-training,”ArXiv, vol. abs/2505.17589, 2025

  6. [6]

    F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,

    Y . Chen, Z. Niu, Z. Maet al., “F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Vienna, Austria: ACL, Jul. 2025, pp. 6255–6271

  7. [7]

    NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models

    Z. Ju, Y . Wang, K. Shenet al., “NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models.” inICML, ser. Proceedings of Machine Learning Research, vol. 235. PMLR / OpenReview.net, 2024, pp. 22 605–22 623

  8. [8]

    Chatterbox-TTS,

    Resemble AI, “Chatterbox-TTS,” https://github.com/resemble-ai/ chatterbox, 2025, gitHub repository

  9. [9]

    AccentBox: Towards High-Fidelity Zero-Shot Accent Generation,

    J. Zhong, K. Richmond, Z. Suet al., “AccentBox: Towards High-Fidelity Zero-Shot Accent Generation,” inICASSP. IEEE, Apr. 2025, p. 1–5

  10. [10]

    Clarity: Contextual linguistic adaptation and accent retrieval for dual-bias mitigation in text-to-speech generation,

    C. M. H. Poon, P. C. Ng, X. Miaoet al., “Clarity: Contextual linguistic adaptation and accent retrieval for dual-bias mitigation in text-to-speech generation,”arXiv preprint arXiv:2511.11104, 2025

  11. [11]

    Deterding,Singapore English

    D. Deterding,Singapore English. Edinburgh University Press, 2007

  12. [12]

    The Homogenization of Ethnic Differences in Singapore English? A Consonantal Production Study,

    K. Kalaivanan, F. Sumartono, and Y .-Y . Tan, “The Homogenization of Ethnic Differences in Singapore English? A Consonantal Production Study,”Language and Speech, vol. 64, no. 1, p. 123–140, 2020

  13. [13]

    Building the Singapore English national speech corpus,

    J. X. Koh, A. Mislan, K. Khooet al., “Building the Singapore English national speech corpus,” inInterspeech, Graz, Austria, 2019

  14. [14]

    MNSC: Advancing Singlish Speech Understanding with Carefully Curated Corpora,

    B. Wang, X. Zou, S. Sunet al., “MNSC: Advancing Singlish Speech Understanding with Carefully Curated Corpora,” inASRU, 2025, pp. 1–8

  15. [15]

    MERaLiON-AudioLLM: Advancing Speech and Language Understanding for Singapore,

    Y . He, Z. Liu, G. Linet al., “MERaLiON-AudioLLM: Advancing Speech and Language Understanding for Singapore,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguis- tics (Volume 3: System Demonstrations). Association for Computational Linguistics, Jul. 2025, pp. 22–30

  16. [16]

    MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond

    M. Huzaifah, G. Lin, T. Liuet al., “MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond.” arXiv:2412.11538, 2024

  17. [17]

    CodecMOS-Accent: A MOS benchmark of resynthesized and TTS speech from neural codecs across English accents,

    W.-C. Huang, N. Sanders, and E. Cooper, “CodecMOS-Accent: A MOS benchmark of resynthesized and TTS speech from neural codecs across English accents,” vol. arXiv:2603.14328, 2026

  18. [18]

    whisper-large-v3-singlish: A Singlish-fine- tuned Whisper-large-v3 model,

    M. J. Wong, “whisper-large-v3-singlish: A Singlish-fine- tuned Whisper-large-v3 model,” https://huggingface.co/mjwong/ whisper-large-v3-singlish, 2024

  19. [19]

    Radar challenge 2026: Robust audio deepfake recognition under media transformations,

    H.-T. Luong, X. Liu, I. Kukanovet al., “Radar challenge 2026: Robust audio deepfake recognition under media transformations,”arXiv preprint arXiv:2605.09568, 2026

  20. [20]

    Accented text-to-speech synthesis with limited data,

    X. Zhou, M. Zhang, Y . Zhouet al., “Accented text-to-speech synthesis with limited data,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, pp. 1699–1711, 2024

  21. [21]

    Low-resource multilingual and zero- shot multispeaker tts,

    F. Lux, J. Koch, and N. T. Vu, “Low-resource multilingual and zero- shot multispeaker tts,” inProceedings of the 2nd Conference of the Asia- Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 2022, pp. 741–751

  22. [22]

    Scalable controllable accented tts,

    H. L. Xinyuan, Z. Cai, A. Garget al., “Scalable controllable accented tts,”arXiv preprint arXiv:2508.07426, 2025

  23. [23]

    Parameter-efficient learning for text-to-speech accent adaptation,

    L.-J. Yang, C.-H. H. Yang, and J.-T. Chien, “Parameter-efficient learning for text-to-speech accent adaptation,” inProc. Interspeech. ISCA, 2023, pp. 4354–4358

  24. [24]

    Accent vector: Controllable accent manipulation for multilingual tts without accented data,

    T. Lertpetchpun, T. Trachu, J. Leeet al., “Accent vector: Controllable accent manipulation for multilingual tts without accented data,”arXiv preprint arXiv:2603.07534, 2026

  25. [25]

    Xtts: a massively multilingual zero-shot text-to-speech model,

    E. Casanova, K. Davis, E. G ¨olgeet al., “Xtts: a massively multilingual zero-shot text-to-speech model,” inProc. Interspeech 2024, 2024, pp. 4978–4982

  26. [26]

    Learning-free l2-accented speech generation using phonological rules,

    T. Lertpetchpun, Y . Lee, J. Leeet al., “Learning-free l2-accented speech generation using phonological rules,”arXiv preprint arXiv:2603.07550, 2026

  27. [27]

    Macst: Multi-accent speech synthesis via text transliteration for accent conversion,

    S. Inoue, S. Wang, W. Wanget al., “Macst: Multi-accent speech synthesis via text transliteration for accent conversion,” inICASSP 2025- 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5

  28. [28]

    The anatomy of singlish: globalisation, multiculturalism and the construction of the ‘local’in singapore,

    R. B. Goh, “The anatomy of singlish: globalisation, multiculturalism and the construction of the ‘local’in singapore,”Journal of Multilingual and Multicultural Development, vol. 37, no. 8, pp. 748–758, 2016

  29. [29]

    The roles of singapore standard english and singlish,

    S. Harada, “The roles of singapore standard english and singlish,”Joho Kenkyu, vol. 40, pp. 69–81, 2009

  30. [30]

    Exploring the unique morphological and syntactic features of singlish (singapore english),

    N. S. Ningsih and F. Rahman, “Exploring the unique morphological and syntactic features of singlish (singapore english),”Journal of English in Academic and Professional Communication, vol. 9, no. 2, pp. 72–80, 2023

  31. [31]

    Singlish: a controversial yet unique creole of singaporean,

    A. M. Kareba, S. Aminahet al., “Singlish: a controversial yet unique creole of singaporean,” inInternational Seminar Commemorating the 100th Annniversary of Tamansiswa, vol. 1, no. 1, 2022, pp. 39–45

  32. [32]

    Singlish as a dialect in singapore,

    C. F. Peng and S. D. Madawan, “Singlish as a dialect in singapore,” International Journal of Physical and Social Sciences, vol. 3, no. 4, pp. 50–70, 2013

  33. [33]

    The Homogenization of Ethnic Differences in Singapore English? A Consonantal Production Study,

    K. Kalaivanan, F. Sumartono, and Y .-Y . Tan, “The Homogenization of Ethnic Differences in Singapore English? A Consonantal Production Study,”Language and Speech, vol. 64, no. 1, pp. 123–140, 2021

  34. [34]

    ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,

    B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,” inInterspeech. ISCA, 2020, pp. 3830– 3834

  35. [35]

    CommonAccent: Exploring Large Acoustic Pretrained Models for Accent Classification Based on Common V oice,

    J. Zuluaga-Gomez, S. Ahmed, D. Visockaset al., “CommonAccent: Exploring Large Acoustic Pretrained Models for Accent Classification Based on Common V oice,” inINTERSPEECH. ISCA, Aug. 2023, p. 5291–5295

  36. [36]

    Utmos: Utokyo-sarulab system for voicemos challenge 2022,

    T. Saeki, D. Xin, W. Nakataet al., “Utmos: Utokyo-sarulab system for voicemos challenge 2022,”Interspeech 2022, 2022

  37. [37]

    TTSDS - Text-to-Speech Distribution Score,

    C. Minixhofer, O. Klejch, and P. Bell, “TTSDS - Text-to-Speech Distribution Score,” inSLT. IEEE, Dec. 2024, p. 766–773

  38. [38]

    TTSDS2: Robust Objective Evaluation for Human-Quality Synthetic Speech,

    C. Minixhofer, O. Klejch, and P. Bell, “TTSDS2: Robust Objective Evaluation for Human-Quality Synthetic Speech,” in13th edition of the Speech Synthesis Workshop. ISCA, Aug. 2025, p. 68–75

  39. [39]

    Cosyvoice,

    FunAudioLLM, “Cosyvoice,” https:// github.com/FunAudioLLM/CosyV oice/tree/ 4d7295a9a7076b7b656f63c35f4dad199f3af33e, 2025, gitHub repository, commit 4d7295a9a7076b7b656f63c35f4dad199f3af33e. Accessed: 2026-06-29

  40. [40]

    Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS,

    D. Seo, G. Park, and K. Nam, “Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS,” 2026. [Online]. Available: https://arxiv.org/abs/2605.30748

  41. [41]

    The voicemos challenge 2023: Zero-shot subjective speech quality prediction for multiple domains,

    E. Cooper, W.-C. Huang, Y . Tsaoet al., “The voicemos challenge 2023: Zero-shot subjective speech quality prediction for multiple domains,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023, pp. 1–7