REVIEW 3 major objections 5 minor 18 references
Dialogs is a new open Russian conversational speech corpus that matches studio read-speech quality while scoring substantially higher on expressiveness and conversational naturalness, and it supports training an expressive dialog TTS.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 02:27 UTC pith:HMT3SQ2W
load-bearing objection A genuinely useful Russian expressive dialog corpus, but the headline expressiveness advantage over baselines rests on an asymmetric MOS comparison that should be fixed before the claim is cited. the 3 major comments →
Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Dialogs is a studio-quality, openly licensed Russian corpus of acted conversational speech with per-utterance labels across 12 style/emotion categories. Its recording protocol—actors seated face-to-face, improvising around script prompts—captures turn-taking rhythm and expressive prosody that read-speech resources lack. In crowd MOS evaluation, Dialogs is rated comparable to strong single-speaker baselines on overall quality, audio quality, and intelligibility, while receiving substantially higher ratings for expressiveness and conversational naturalness. Training a VITS2 model on the corpus yields speech whose expressiveness and conversational scores exceed its intelligibility score, eviden
What carries the argument
The central object is the Dialogs corpus itself: 11,796 stereo utterances, 21,262 unique words, and 12 style/emotion labels per utterance (neutral, happy, surprise, sad, disgust, angry, tongue-twister, poem, whisper, arrogance, laughing, fear). The load-bearing design choices are the face-to-face studio recording protocol that encourages natural turn-taking and improvisation, the crowd-based triple annotation with majority-vote (ties broken toward rare styles), and a stratified test set of 188 utterances (5 per speaker per emotion) used for MOS evaluation. This structure is what lets the paper claim that the corpus, not just the recording hardware, drives the conversational and expressive ad
Load-bearing premise
The claim that Dialogs is more expressive and conversational than the baselines rests on the assumption that its emotion-stratified 188-clip evaluation subset is comparable to the baselines' random 100-clip subsets, despite the different sampling and the removal of outlying raters.
What would settle it
A listening test that scores randomly sampled Dialogs clips (not stratified by emotion) against Ruslan and Natasha using the same rater pool would show whether the expressiveness and conversational naturalness advantages shrink or vanish; if the gap persists on random samples, the claim is robust.
If this is right
- Russian TTS systems can be trained or fine-tuned on Dialogs to produce expressive, conversational output without resorting to uncontrolled web-mined data.
- The 12 style/emotion labels enable emotion-conditioned synthesis, letting developers choose a delivery style per utterance in dialog assistants.
- Because the corpus is released under an open commercial license, unlike most existing Russian studio corpora, production teams can legally use it.
- The stratified test set provides a reproducible benchmark for evaluating expressive Russian TTS, with per-speaker and per-emotion coverage.
- Mixing Dialogs with larger read-speech corpora is expected to lift synthesis quality, since the low per-speaker hours cap raw naturalness (a limitation the authors note).
Where Pith is reading between the lines
- The evaluation asymmetry—stratified emotion-rich excerpts for Dialogs versus random clips for baselines—means the expressiveness gap could be partly an artifact of sampling; a matched random-sample listening test would settle it.
- The corpus's acted, script-improvised nature means it does not contain true spontaneous speech, overlapping turns, or disfluencies; extending the recording protocol to unscripted interaction could test whether the conversational advantage generalizes beyond acted dialogs.
- The per-utterance multi-label annotation opens avenues beyond TTS, such as expressive resynthesis or emotion-transfer research, where a controlled studio corpus with rare styles (whisper, tongue-twisters) is currently scarce for Russian.
- If the expressiveness ratings reproduce in independent listening studies, the face-to-face recording protocol itself—actors reacting to each other rather than to a microphone—could become a standard recipe for constructing conversational TTS corpora in other languages.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Dialogs, a 20.6-hour studio-quality Russian conversational speech corpus with 11,796 utterances from 3 actors, recorded at 44.1 kHz and released under OpenRAIL. The corpus is claimed to fill a gap for expressive, conversational Russian TTS by combining studio audio, dialog recordings, and per-utterance style/emotion labels across 12 categories. The authors validate the corpus with crowd-sourced MOS tests, reporting comparable audio quality and intelligibility to the Russian read-speech baselines Ruslan and Natasha, and higher expressiveness and conversational naturalness. They also train VITS2 on the corpus as a proof of concept, reporting MOS and UTMOS scores for synthesized speech.
Significance. If the claims hold, Dialogs is a useful public resource for Russian expressive and dialog TTS, with commercial-friendly licensing, per-utterance emotion labels, and a reproducible training demonstration. The paper is transparent about the data release, and the availability of code and corpus is a concrete contribution. However, the headline expressiveness and conversational-naturalness advantages rest on a MOS comparison with asymmetric sampling across conditions, and the annotation aggregation procedure is not validated; these issues currently prevent the paper from fully establishing its central claims.
major comments (3)
- [§3.5, Table 3] The headline expressiveness (+0.23–0.25) and conversational (+0.24–0.30) advantages are not established by the reported comparison because the evaluation subsets differ by design. Dialogs' 188 clips are stratified by 12 emotion/style categories (5 per speaker per style), while Ruslan and Natasha are random 100-clip draws from unlabelled read corpora. Table 4 shows extreme duration skew (neutral 643.5 min, happy 349.6 min, whisper 6.1 min, laughing 19.9 min, angry 19.2 min), so Dialogs raters hear a curated set that oversamples rare, highly expressive styles, whereas baseline raters hear typical read speech. Confidence intervals quantify sampling variability within each subset, not selection bias across subsets. Please either evaluate all corpora under the same sampling protocol, or supply a sensitivity analysis using a random Dialogs subset to show the advantage persists.
- [§3.4] The annotation aggregation is not validated and may bias labels. With three annotators and majority vote, all-disagree ties are resolved by selecting the globally least-frequent category. This can assign a rare label to an utterance for which no annotator chose that category, artificially inflating rare-style durations in Table 4 and injecting label noise into training. No inter-annotator agreement, distribution of tie cases, or comparison to alternative tie-breaking is reported. Since per-utterance style/emotion labels are a central new contribution, please quantify the disagreement rate and justify or change the tie-breaking rule.
- [§3.5] Outlier annotator removal lacks transparency. The paper reports 41/26/23 retained raters per corpus after 'response pattern analysis' but does not state the removal criterion, the number removed per corpus, or results without removal. If removal was more aggressive for Ruslan/Natasha than for Dialogs, the apparent expressiveness gap in Table 3 could be an artifact. Please provide the criterion, the counts removed, and a sensitivity analysis (e.g., all raters included, or an identical outlier rule across conditions).
minor comments (5)
- [§3.4] No inter-annotator agreement measure (e.g., Fleiss' kappa) is reported. Also specify annotator instructions and whether annotators heard full dialog context or isolated utterances.
- [§5, Table 5] The statement that expressiveness (2.56) and conversational (2.59) 'notably exceed' intelligibility (2.28) is not supported by the reported 95% CIs: the intervals overlap substantially. Use a paired test or soften the claim.
- [Table 1] The Natasha row lacks a citation. Add a reference or state the source.
- [§3.2] Clarify the recording setup: were two Behringer XM8500 microphones used, one per actor? How was stereo captured? Place this information in the metadata as well.
- [§3.1, Table 2] The corpus UTMOS score (3.17±0.07) and the TTS UTMOS score (3.36±0.06) are close; since UTMOS is not calibrated across different audio conditions, avoid direct numerical comparison or state that they are not comparable.
Circularity Check
No significant circularity: the paper is a corpus release with empirical evaluations, and its central claims are not derived from their own inputs.
full rationale
The paper contains no derivation chain whose predictions reduce to its inputs. Dialogs is a data-release paper: 20.6 hours of recorded dialog speech are described, annotated, and evaluated against external Russian corpora (Ruslan, Natasha) via crowd MOS, followed by a VITS2 training proof-of-concept. There is no fitted parameter that is later renamed as a prediction, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in through self-citation. The expressiveness/conversational advantage in Table 3 is an empirical comparison, not a mathematical consequence of the corpus construction. The most plausible concern is that the evaluation subsets are asymmetric: Dialogs uses 188 emotion-stratified utterances (5 per speaker per emotion) while Ruslan/Natasha use random 100-utterance draws, and outlier annotators are removed per corpus; the paper asserts comparability because confidence intervals are reported (§3.5). This is a validity/selection-bias concern and belongs under correctness risk, not circularity: even if the comparison is biased, the claim does not follow from its own definitions by construction. The self-contained TTS experiment and the released dataset provide independent content, and the limitation section candidly notes scripted/acted dialogs. Therefore the appropriate circularity score is 0.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Crowd-sourced 5-point MOS ratings, after outlier removal, are reliable and comparable across corpora with different evaluation subset sizes.
- ad hoc to paper Majority-vote emotion labels are valid ground truth, and resolving ties to the globally least-frequent category improves rare-style recall without biasing labels.
- domain assumption Actor improvisation from script prompts captures the turn-taking and expressive properties needed for dialog-assistant TTS.
- ad hoc to paper Training VITS2 on a single GPU for 615k steps with default hyperparameters is a sufficient demonstration that the corpus supports expressive TTS.
read the original abstract
We introduce Dialogs, a studio-quality Russian conversational speech corpus for dialog assistants. The dataset contains 20.6 hours of face-to-face acted dialogs recorded in a professional studio (44.1 kHz stereo) and segmented into 11,796 utterances across 3 speakers. Unlike read-speech resources, Dialogs captures turn-taking rhythm and expressive prosody, and provides per-utterance style/emotion labels spanning 12 categories. We validate corpus quality with crowd MOS tests, showing comparable audio quality and intelligibility to strong Russian studio baselines while achieving higher ratings for expressiveness and conversational naturalness. Finally, we train a VITS2 model as a proof of concept, demonstrating that Dialogs supports training expressive, dialog-like TTS despite limited per-speaker data.
Figures
Reference graph
Works this paper leans on
-
[1]
Introduction Modern conversational assistants and text-to-speech (TTS) sys- tems increasingly demand natural, expressive, and interactive speech. Achieving such naturalness depends critically on train- ing and adaptation data: high-quality recordings that capture prosodic variation, natural timing, interruptions and other fea- tures of real conversation a...
-
[2]
Related Work 2.1. Corpus taxonomy In practice, most public speech corpora fall into one of three cat- egories.Read-speechcorpora (audiobooks, news) offer clean audio and accurate transcripts, but the recording style is mono- tone and neutral — actors read to microphones, not to each other.Web-minedcorpora (transcribed YouTube, podcasts) are naturally expr...
Pith/arXiv arXiv 2026
-
[3]
Participant metadata, includ- ing performer experience and gender, are documented in the dataset metadata
The Dialogs Dataset Dialogs contains 20.6 hours of expressive dialog speech recorded by professional puppet-theatre actors from Central Russia (see Figure 1 for the recording duration distribution; cor- pus statistics are given in Table 2). Participant metadata, includ- ing performer experience and gender, are documented in the dataset metadata. All parti...
-
[4]
Experiments To demonstrate the utility of Dialogs as a training corpus, we train VITS2 [8], a single-stage end-to-end TTS model on the Dialogs training split. The experiment serves as a proof-of- concept: the corpus spans three speakers with approximately 21 hours of audio total (9.9 h, 6.2 h, and 4.4 h per speaker respectively), a challenging regime for ...
-
[5]
This experiment serves as a proof of concept that the corpus is sufficient to train a functional TTS system
Results and Discussion Table 5 reports evaluation scores for the VITS2 model trained on Dialogs. This experiment serves as a proof of concept that the corpus is sufficient to train a functional TTS system. Infor- mal listening confirmed that synthesized speech exhibits clear conversational and expressive character consistent with the di- alog recording st...
-
[6]
Generative AI Use Disclosure GPT-3.5 was used to assist in generating the recording scripts used for dataset collection. AI tools were also used heavily dur- ing development of evaluation pipelines (MOS templates, ag- gregation, table formation etc), collecting dataset statistics and for editing and polishing of this manuscript. All scientific con- tent, ...
-
[7]
Expresso: A benchmark and analysis of discrete expressive speech resynthesis,
T. A. Nguyen, W.-N. Hsu, A. D’Avirro, B. Shi, I. Gat, M. Fazel-Zarani, T. Remez, J. Copet, G. Synnaeve, M. Hassid, F. Kreuk, Y . Adi, and E. Dupoux, “Expresso: A benchmark and analysis of discrete expressive speech resynthesis,” 8 2023. [Online]. Available: http://arxiv.org/abs/2308.05725
Pith/arXiv arXiv 2023
-
[8]
Dailytalk: Spoken dialogue dataset for conversational text-to-speech,
K. Lee, K. Park, and D. Kim, “Dailytalk: Spoken dialogue dataset for conversational text-to-speech,” 3 2023. [Online]. Available: http://arxiv.org/abs/2207.01063
Pith/arXiv arXiv 2023
-
[9]
Large raw emotional dataset with aggregation mechanism,
V . Kondratenko, A. Sokolov, N. Karpov, O. Kutuzov, N. Savushkin, and F. Minkin, “Large raw emotional dataset with aggregation mechanism,” 12 2022. [Online]. Available: http://arxiv.org/abs/2212.12266
Pith/arXiv arXiv 2022
-
[10]
Ruslan: Russian spoken language corpus for speech synthesis,
L. Gabdrakhmanov, R. Garaev, and E. Razinkov, “Ruslan: Russian spoken language corpus for speech synthesis,” 6 2019. [Online]. Available: http://arxiv.org/abs/1906.11645
Pith/arXiv arXiv 2019
-
[11]
Golos: Russian dataset for speech research,
N. Karpov, A. Denisenko, and F. Minkin, “Golos: Russian dataset for speech research,” 6 2021. [Online]. Available: http://arxiv.org/abs/2106.10161
Pith/arXiv arXiv 2021
-
[12]
Resd: Russian emotional speech dataset,
Aniemore, “Resd: Russian emotional speech dataset,” 2023, hug- ging Face Datasets
2023
-
[13]
Espeech: A large-scale, high-quality rus- sian speech corpus for text-to-speech,
D. Petrov and E. Team, “Espeech: A large-scale, high-quality rus- sian speech corpus for text-to-speech,” Tech. Rep., 2025
2025
-
[14]
VITS2: Improving quality and efficiency of single-stage text-to-speech with adversarial learning and architecture design,
J. Kong, J. Park, B. Kim, J. Oh, D. Kong, and S. Kim, “VITS2: Improving quality and efficiency of single-stage text-to-speech with adversarial learning and architecture design,” inProc. Inter- speech 2023, 2023, pp. 4374–4378
2023
-
[15]
Deep learning-based expressive speech synthesis: a systematic review of approaches, challenges, and resources,
H. Barakat, O. Turk, and C. Demiroglu, “Deep learning-based expressive speech synthesis: a systematic review of approaches, challenges, and resources,” 12 2024
2024
-
[16]
Contextual expressive text-to-speech,
J. Tu, Z. Cui, X. Zhou, S. Zheng, K. Hu, J. Fan, and C. Zhou, “Contextual expressive text-to-speech,” 11 2022. [Online]. Available: http://arxiv.org/abs/2211.14548
Pith/arXiv arXiv 2022
-
[17]
Proemo: Prompt- driven text-to-speech synthesis based on emotion and intensity control,
S. Zhang, A. Mehrish, Y . Li, and S. Poria, “Proemo: Prompt- driven text-to-speech synthesis based on emotion and intensity control,” 1 2025. [Online]. Available: http://arxiv.org/abs/2501. 06276
2025
-
[18]
UTMOS: UTokyo-SaruLab system for V oice- MOS challenge 2022,
T. Saeki, D. Xin, W. Nakata, T. Yoshimura, S. Takamichi, and H. Saruwatari, “UTMOS: UTokyo-SaruLab system for V oice- MOS challenge 2022,” inProc. Interspeech 2022, 2022, pp. 4521–4525
2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.