REVIEW 1 major objections 5 minor 1 cited by
A single likelihood score can flag unnatural turn-taking across many failure types in two-speaker dialogue.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 08:52 UTC pith:LACMJKAM
load-bearing objection Solid, usable metric + human-validated paired benchmark for turn-taking naturalness; claim holds on its terms, main limit is transfer to real full-duplex systems. the 1 major comments →
TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
TurnNat shows that the negative log-likelihood of observed future two-speaker voice-activity states, computed by a model trained solely on natural dialogue and aggregated over VAD-derived turn-taking boundary units, yields a single continuous score that reliably separates natural dyadic clips from five heterogeneous timing perturbations (late response, early entry, hold-instead-of-shift, shift-instead-of-hold, and excessive backchanneling).
What carries the argument
Turn-taking boundary units (TBUs): short windows anchored just before each utterance onset or offset; frame-level negative log-likelihoods of the predicted future two-speaker activity pattern are averaged inside each TBU, then the mean and top-quartile TBU scores are combined into the dialogue-level naturalness score.
Load-bearing premise
Local timing edits of natural human-to-human recordings are a good enough stand-in for the turn-taking failures that actually appear in real full-duplex systems.
What would settle it
On a held-out set of real human–AI full-duplex dialogues that contain genuine latency, interruption, or backchannel errors, the same TurnNat score fails to rank the human-preferred versions higher than the system versions at rates comparable to the 88 percent matched-pair accuracy on the perturbation benchmark.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes TurnNat, a likelihood-based automatic metric for turn-taking naturalness in two-channel spoken dialogue. A causal model trained only on natural conversations predicts future two-speaker voice-activity states over a 2 s horizon (256 joint states); frame-level negative log-likelihood of the observed future activity is pooled over VAD-derived turn-taking boundary units (TBUs) around utterance onsets/offsets, then aggregated via mean and top-quartile tail into a dialogue-level score m_θ(x). The authors construct a held-out paired perturbation benchmark (five local timing/floor-management edits of Seamless Interaction clips) and validate it with human A/B and rating judgments. Instantiations with VAP and DualTurn backbones show that the best DualTurn+categorical+auxiliary+TBU-weighted scorer reaches 88.0% matched-pair accuracy and C-index 0.676, consistently ranking natural above perturbed clips across heterogeneous failures.
Significance. If the result holds, TurnNat supplies a single continuous score that places heterogeneous turn-taking failures (late response, early entry, hold/shift swaps, excessive backchannels) in a shared likelihood space, addressing a real gap left by task-specific full-duplex benchmarks and human listening tests. Strengths include a clean training/evaluation separation (natural-only training; held-out human-validated pairs), explicit ablations of backbone, output head, TBU weighting, and aggregation (Tables III–IV), and public code. The main external-validity limit—controlled human–human edits rather than real system-side latency or human–AI dialogue—is already stated in §VIII and does not undercut the scoped claim. The work is a solid, usable contribution for evaluating full-duplex spoken dialogue systems.
major comments (1)
- The central claim is carefully scoped to the human-validated paired perturbation benchmark and is supported by Tables II–IV and Eqs. (7)–(10). No load-bearing internal inconsistency, circularity, or derivation error is identified. The principal limitation (transfer beyond local human–human edits to system latency, semantic/prosodic mismatch, and real human–AI dialogue) is already acknowledged in §VIII and does not require a major revision of the reported discrimination result.
minor comments (5)
- Table III: report Wilson or bootstrap CIs for Acc_pair and C-index consistently (only the best overall Acc_pair is given a CI in the text); this would make architecture comparisons easier to interpret.
- §III-B / Eq. (1): the island-merge thresholds (1.0 s gap, 0.2 other-speaker ratio, 200 ms minimum) and L=2 s are free parameters; a short sensitivity note or appendix table would strengthen reproducibility claims.
- Fig. 2: the qualitative NLL spikes are helpful; adding one hold-instead or shift-instead example would better cover the five perturbation types shown in Table III.
- Related work: DualTurn is cited as arXiv:2603.08216 (2026); ensure the citation is stable and that any concurrent full-duplex evaluation papers are briefly positioned against TurnNat’s unified likelihood framing.
- Notation: c_t is defined as the index of b_t (Eq. 3) but sometimes referred to loosely as the “state”; a one-line clarification that p_θ(t;x)[c_t] is the probability of the observed joint pattern would help readers less familiar with VAP.
Circularity Check
No significant circularity: TurnNat’s NLL score is an intentional atypicality definition tested on held-out, human-validated perturbations, not a fit or self-definition of the evaluation labels.
full rationale
The derivation chain is: (i) train a causal future two-speaker voice-activity model only on natural Seamless Interaction train data (Eq. 5, Table I); (ii) define frame NLL as −log pθ(t;x)[ct] (Eq. 7), pool over VAD-derived TBUs (Eqs. 1, 8), and aggregate mean+tail into mθ(x) (Eqs. 9–10); (iii) rank natural vs. five controlled local perturbations on a held-out test-derived paired benchmark, with human A/B and rating validation independent of the scorer (Tables II–IV). Nothing in this chain is forced by construction: the model is not trained or fitted on perturbation labels; α is chosen by development future-activity loss, not by maximizing Accpair; DualTurn and VAP are external architectures (no author-overlap uniqueness theorems); and the claim is empirical discrimination on that benchmark, not a first-principles identity. Using NLL of observed future activity as “atypicality” is the metric’s definition, then tested—standard likelihood-based evaluation, not circular reduction of prediction to input.
Axiom & Free-Parameter Ledger
free parameters (5)
- TBU training weight α =
8 (best); also tested 1 and 3
- Mean–tail mix λ =
0.5
- Tail fraction ρ =
0.25
- TBU pre-boundary length L and island-merge thresholds =
L=2s; gap 1.0s; ratio 0.2; min 200ms
- Perturbation magnitude ranges =
1.2–2.5 s / 2–3 BCs
axioms (4)
- domain assumption Future two-speaker voice-activity patterns over a 2 s horizon (4 non-uniform bins, 256 joint states) are a sufficient probabilistic target for turn-taking naturalness.
- domain assumption Models trained only on natural human–human dialogue assign higher likelihood to natural timing than to unnatural timing, so NLL is a valid atypicality measure.
- domain assumption Local edits that preserve speaker identity, lexical content, and most context isolate turn-taking naturalness differences.
- standard math Standard categorical cross-entropy / NLL training and Softmax over 256 states are valid.
invented entities (2)
-
Turn-taking boundary units (TBUs)
no independent evidence
-
TurnNat dialogue-level naturalness score m_θ(x)
no independent evidence
read the original abstract
Turn-taking naturalness is central to full-duplex spoken dialogue systems, yet its automatic evaluation remains limited. Existing evaluations often rely on human judgments or behavior-specific timing metrics, making it difficult to compare heterogeneous timing failures within a unified framework. We propose TurnNat, a likelihood-based framework for automatic turn-taking naturalness evaluation in two-channel spoken dialogue. A causal turn-taking prediction model trained on natural conversations estimates future two-speaker voice-activity states, and the negative log-likelihood (NLL) of the observed future activity measures timing atypicality. TurnNat pools frame-level NLLs over turn-taking boundary units (TBUs) extracted from utterance onsets and offsets, and aggregates mean and tail TBU scores into a dialogue-level naturalness score. We further construct a controlled perturbation benchmark of paired natural and perturbed dialogue clips, validated by human naturalness judgments. Experiments on this benchmark show that TurnNat successfully identifies unnatural turn-taking perturbations across heterogeneous timing failures.
Figures
Forward citations
Cited by 1 Pith paper
-
SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models
An open-source benchmark for speech-to-speech models shows that current systems produce intelligible audio but diverge from human conversational behavior in latency, dialect consistency, emotional entrainment, and prosody.
Reference graph
Works this paper leans on
-
[1]
A survey of recent advances on turn-taking modeling in spoken dialogue systems,
G. Castillo-L ´opez, G. de Chalendar, and N. Semmar, “A survey of recent advances on turn-taking modeling in spoken dialogue systems,” inProceedings of the 15th international workshop on spoken dialogue systems technology, 2025, pp. 254–271
2025
-
[2]
Response timing estimation for spoken dialog systems based on syntactic completeness prediction,
J. Sakuma, S. Fujie, and T. Kobayashi, “Response timing estimation for spoken dialog systems based on syntactic completeness prediction,” in 2022 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2023, pp. 369–374
2022
-
[3]
Wavchat: A survey of spoken dialogue models,
S. Ji, Y . Chen, M. Fang, J. Zuo, J. Lu, H. Wang, Z. Jiang, L. Zhou, S. Liu, X. Chenget al., “Wavchat: A survey of spoken dialogue models,”arXiv preprint arXiv:2411.13577, 2024
Pith/arXiv arXiv 2024
-
[4]
Moshi: a speech-text foundation model for real-time dialogue,
A. D ´efossez, L. Mazar ´e, M. Orsini, A. Royer, P. P ´erez, H. J ´egou, E. Grave, and N. Zeghidour, “Moshi: a speech-text foundation model for real-time dialogue,”arXiv preprint arXiv:2410.00037, 2024
Pith/arXiv arXiv 2024
-
[5]
Llama- omni: Seamless speech interaction with large language models,
Q. Fang, S. Guo, Y . Zhou, Z. Ma, S. Zhang, and Y . Feng, “Llama- omni: Seamless speech interaction with large language models,” in International Conference on Learning Representations, vol. 2025, 2025, pp. 57 607–57 624
2025
-
[6]
Mini-omni: Language models can hear, talk while thinking in streaming,
Z. Xie and C. Wu, “Mini-omni: Language models can hear, talk while thinking in streaming,”arXiv preprint arXiv:2408.16725, 2024
Pith/arXiv arXiv 2024
-
[7]
Freeze-omni: A smart and low latency speech-to-speech dia- logue model with frozen llm,
X. Wang, Y . Li, C. Fu, Y . Zhang, Y . Shen, L. Xie, K. Li, X. Sun, and L. Ma, “Freeze-omni: A smart and low latency speech-to-speech dia- logue model with frozen llm,” inInternational Conference on Machine Learning. PMLR, 2025, pp. 63 345–63 354
2025
-
[8]
G.-T. Lin, J. Lian, T. Li, Q. Wang, G. Anumanchipalli, A. H. Liu, and H.-y. Lee, “Full-duplex-bench: A benchmark to evaluate full-duplex spoken dialogue models on turn-taking capabilities,”arXiv preprint arXiv:2503.04721, 2025
Pith/arXiv arXiv 2025
-
[9]
Mitigating response delays in free-form conversations with llm- powered intelligent virtual agents,
M. Maslych, M. Katebi, C. Lee, Y . Hmaiti, A. Ghasemaghaei, C. Pumarada, J. Palmer, E. Segarra Martinez, M. Emporio, W. Snipes et al., “Mitigating response delays in free-form conversations with llm- powered intelligent virtual agents,” inProceedings of the 7th ACM Conference on Conversational User Interfaces, 2025, pp. 1–15
2025
-
[10]
Neural generation of dialogue response timings,
M. Roddy and N. Harte, “Neural generation of dialogue response timings,” inProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 2442–2452
2020
-
[11]
Toward enabling natural conversation with older adults via the design of llm-powered voice agents that support interruptions and backchannels,
C. Liu, M. Su, Y . Xiang, Y . Huang, Y . Yang, K. Zhang, and M. Fan, “Toward enabling natural conversation with older adults via the design of llm-powered voice agents that support interruptions and backchannels,” inProceedings of the 2025 CHI conference on human factors in computing systems, 2025, pp. 1–22
2025
-
[12]
A mixed-methods approach to understanding user trust after voice assistant failures,
A. Baughan, X. Wang, A. Liu, A. Mercurio, J. Chen, and X. Ma, “A mixed-methods approach to understanding user trust after voice assistant failures,” inProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, 2023, pp. 1–16
2023
-
[13]
Patient engagement with conversational agents in health applications 2016–2022: a systematic review and meta-analysis,
K. E. Cevasco, R. E. Morrison Brown, R. Woldeselassie, and S. Kaplan, “Patient engagement with conversational agents in health applications 2016–2022: a systematic review and meta-analysis,”Journal of medical systems, vol. 48, no. 1, p. 40, 2024
2016
-
[14]
Empathic conversational agent platform designs and their evaluation in the context of mental health: systematic review,
R. Sanjeewa, R. Iyer, P. Apputhurai, N. Wickramasinghe, and D. Meyer, “Empathic conversational agent platform designs and their evaluation in the context of mental health: systematic review,”JMIR Mental Health, vol. 11, p. e58974, 2024
2024
-
[15]
Talking turns: Benchmarking audio foundation models on turn-taking dynamics,
S. Arora, Z. Lu, C.-C. Chiu, R. Pang, and S. Watanabe, “Talking turns: Benchmarking audio foundation models on turn-taking dynamics,”arXiv preprint arXiv:2503.01174, 2025
Pith/arXiv arXiv 2025
-
[16]
Turngpt: a transformer-based language model for predicting turn-taking in spoken dialog,
E. Ekstedt and G. Skantze, “Turngpt: a transformer-based language model for predicting turn-taking in spoken dialog,” inFindings of the Association for Computational Linguistics: EMNLP 2020, 2020, pp. 2981–2990
2020
-
[17]
Turn-taking in conversational systems and human-robot interaction: a review,
G. Skantze, “Turn-taking in conversational systems and human-robot interaction: a review,”Computer Speech & Language, vol. 67, p. 101178, 2021
2021
-
[18]
V oice activity projection: Self-supervised learning of turn-taking events,
E. Ekstedt and G. Skantze, “V oice activity projection: Self-supervised learning of turn-taking events,” inProc. Interspeech 2022, 2022, pp. 5190–5194
2022
-
[19]
Predicting turn-taking and backchannel in human-machine conversations using linguistic, acoustic, and visual signals,
Y . Lin, Y . Zheng, M. Zeng, and W. Shi, “Predicting turn-taking and backchannel in human-machine conversations using linguistic, acoustic, and visual signals,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), 2025, pp. 15 310–15 322
2025
-
[20]
Turn-taking and backchannel prediction with acoustic and large language model fusion,
J. Wang, L. Chen, A. Khare, A. Raju, P. Dheram, D. He, M. Wu, A. Stol- cke, and V . Ravichandran, “Turn-taking and backchannel prediction with acoustic and large language model fusion,” inICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 12 121–12 125
2024
-
[21]
Dualturn: Learning turn-taking from dual-channel generative speech pretraining,
S. Rajaa, “Dualturn: Learning turn-taking from dual-channel generative speech pretraining,”arXiv preprint arXiv:2603.08216, 2026
arXiv 2026
-
[22]
Q. Wang, Z. Meng, W. Cui, Y . Zhang, P. Wu, B. Wu, I. King, L. Chen, and P. Zhao, “Ntpp: Generative speech language modeling for dual- channel spoken dialogue via next-token-pair prediction,”arXiv preprint arXiv:2506.00975, 2025
Pith/arXiv arXiv 2025
-
[23]
From reaction to prediction: Experiments with compu- tational models of turn-taking,
D. Schlangen, “From reaction to prediction: Experiments with compu- tational models of turn-taking,”Proceedings of Interspeech 2006, Panel on Prosody of Dialogue Acts and Turn-Taking, 2006
2006
-
[24]
Data-driven models for timing feedback responses in a map task dialogue system,
R. Meena, G. Skantze, and J. Gustafson, “Data-driven models for timing feedback responses in a map task dialogue system,”Computer Speech & Language, vol. 28, no. 4, pp. 903–922, 2014
2014
-
[25]
Opportunities and obligations to take turns in collaborative multi-party human-robot interaction,
M. Johansson and G. Skantze, “Opportunities and obligations to take turns in collaborative multi-party human-robot interaction,” inProceed- ings of the 16th annual meeting of the special interest group on discourse and dialogue, 2015, pp. 305–314
2015
-
[26]
Towards a general, continuous model of turn-taking in spoken dialogue using lstm recurrent neural networks,
G. Skantze, “Towards a general, continuous model of turn-taking in spoken dialogue using lstm recurrent neural networks,” inProceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, 2017, pp. 220–230
2017
-
[27]
Multi- lingual turn-taking prediction using voice activity projection,
K. Inoue, B. Jiang, E. Ekstedt, T. Kawahara, and G. Skantze, “Multi- lingual turn-taking prediction using voice activity projection,” inPro- ceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation (lrec-coling 2024), 2024, pp. 11 873–11 883
2024
-
[28]
Prompt-guided turn-taking prediction,
K. Inoue, M. Elmers, Y . Fu, Z. H. Pang, D. Lala, K. Ochi, and T. Kawahara, “Prompt-guided turn-taking prediction,” inProceedings of the 26th Annual Meeting of the Special Interest Group on Discourse and Dialogue, 2025, pp. 146–151
2025
-
[29]
Real-time and continuous turn-taking prediction using voice activity projection,
K. Inoue, B. Jiang, E. Ekstedt, T. Kawahara, and G. Skantze, “Real-time and continuous turn-taking prediction using voice activity projection,” arXiv preprint arXiv:2401.04868, 2024
Pith/arXiv arXiv 2024
-
[30]
Visual cues enhance predictive turn-taking for two-party human interaction,
S. O. Russell and N. Harte, “Visual cues enhance predictive turn-taking for two-party human interaction,” inFindings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vienna, Austria: Association for Computational Linguistics, Jul. 2025, pp. 209–221. [Online]. Available: https://aclantholo...
2025
-
[31]
Full-duplex-bench v1. 5: Evaluating overlap handling for full-duplex speech models,
G.-T. Lin, S.-Y . S. Kuan, Q. Wang, J. Lian, T. Li, S. Watanabe, and H.-y. Lee, “Full-duplex-bench v1. 5: Evaluating overlap handling for full-duplex speech models,” inICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2026, pp. 19 447–19 451
2026
-
[32]
Seamless inter- action: Dyadic audiovisual motion modeling and large-scale dataset,
V . Agrawal, A. Akinyemi, K. Alvero, M. Behrooz, J. Buffalini, F. M. Carlucci, J. Chen, J. Chen, Z. Chen, S. Chenget al., “Seamless inter- action: Dyadic audiovisual motion modeling and large-scale dataset,” arXiv preprint arXiv:2506.22554, 2025
Pith/arXiv arXiv 2025
-
[33]
Silero V AD: Pre-trained Enterprise-grade V oice Activity Detector,
Silero Team, “Silero V AD: Pre-trained Enterprise-grade V oice Activity Detector,” https://github.com/snakers4/silero-vad, 2024, gitHub reposi- tory. Accessed: 2026-05-01
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.