Pith. sign in

REVIEW 1 major objections 5 minor 1 cited by

A single likelihood score can flag unnatural turn-taking across many failure types in two-speaker dialogue.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 08:52 UTC pith:LACMJKAM

load-bearing objection Solid, usable metric + human-validated paired benchmark for turn-taking naturalness; claim holds on its terms, main limit is transfer to real full-duplex systems. the 1 major comments →

arxiv 2607.01345 v2 pith:LACMJKAM submitted 2026-07-01 cs.CL cs.AI

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue

classification cs.CL cs.AI
keywords spoken dialogueturn-takingnaturalness evaluationvoice activity predictionlikelihood-based evaluationfull-duplex systemsdyadic conversation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Natural spoken conversation depends on precise timing: when speakers start, stop, overlap, or offer brief feedback. Full-duplex dialogue systems still struggle with this, and current tests either ask humans to listen or score each failure type separately, so systems cannot be compared on one scale. TurnNat trains a causal model only on ordinary human conversations to predict both speakers’ future voice activity; the negative log-likelihood of what actually happens then becomes a measure of how atypical the timing is. Those frame-level scores are pooled only around utterance onsets and offsets (turn-taking boundary units) and summarized by their mean and their worst tail into one dialogue-level naturalness number. On a new paired benchmark of natural clips and five controlled timing edits, validated by human listeners, the best TurnNat configuration ranks the natural clip higher 88 percent of the time, showing that future-activity likelihood can unify heterogeneous turn-taking failures.

Core claim

TurnNat shows that the negative log-likelihood of observed future two-speaker voice-activity states, computed by a model trained solely on natural dialogue and aggregated over VAD-derived turn-taking boundary units, yields a single continuous score that reliably separates natural dyadic clips from five heterogeneous timing perturbations (late response, early entry, hold-instead-of-shift, shift-instead-of-hold, and excessive backchanneling).

What carries the argument

Turn-taking boundary units (TBUs): short windows anchored just before each utterance onset or offset; frame-level negative log-likelihoods of the predicted future two-speaker activity pattern are averaged inside each TBU, then the mean and top-quartile TBU scores are combined into the dialogue-level naturalness score.

Load-bearing premise

Local timing edits of natural human-to-human recordings are a good enough stand-in for the turn-taking failures that actually appear in real full-duplex systems.

What would settle it

On a held-out set of real human–AI full-duplex dialogues that contain genuine latency, interruption, or backchannel errors, the same TurnNat score fails to rank the human-preferred versions higher than the system versions at rates comparable to the 88 percent matched-pair accuracy on the perturbation benchmark.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 5 minor

Summary. The paper proposes TurnNat, a likelihood-based automatic metric for turn-taking naturalness in two-channel spoken dialogue. A causal model trained only on natural conversations predicts future two-speaker voice-activity states over a 2 s horizon (256 joint states); frame-level negative log-likelihood of the observed future activity is pooled over VAD-derived turn-taking boundary units (TBUs) around utterance onsets/offsets, then aggregated via mean and top-quartile tail into a dialogue-level score m_θ(x). The authors construct a held-out paired perturbation benchmark (five local timing/floor-management edits of Seamless Interaction clips) and validate it with human A/B and rating judgments. Instantiations with VAP and DualTurn backbones show that the best DualTurn+categorical+auxiliary+TBU-weighted scorer reaches 88.0% matched-pair accuracy and C-index 0.676, consistently ranking natural above perturbed clips across heterogeneous failures.

Significance. If the result holds, TurnNat supplies a single continuous score that places heterogeneous turn-taking failures (late response, early entry, hold/shift swaps, excessive backchannels) in a shared likelihood space, addressing a real gap left by task-specific full-duplex benchmarks and human listening tests. Strengths include a clean training/evaluation separation (natural-only training; held-out human-validated pairs), explicit ablations of backbone, output head, TBU weighting, and aggregation (Tables III–IV), and public code. The main external-validity limit—controlled human–human edits rather than real system-side latency or human–AI dialogue—is already stated in §VIII and does not undercut the scoped claim. The work is a solid, usable contribution for evaluating full-duplex spoken dialogue systems.

major comments (1)
  1. The central claim is carefully scoped to the human-validated paired perturbation benchmark and is supported by Tables II–IV and Eqs. (7)–(10). No load-bearing internal inconsistency, circularity, or derivation error is identified. The principal limitation (transfer beyond local human–human edits to system latency, semantic/prosodic mismatch, and real human–AI dialogue) is already acknowledged in §VIII and does not require a major revision of the reported discrimination result.
minor comments (5)
  1. Table III: report Wilson or bootstrap CIs for Acc_pair and C-index consistently (only the best overall Acc_pair is given a CI in the text); this would make architecture comparisons easier to interpret.
  2. §III-B / Eq. (1): the island-merge thresholds (1.0 s gap, 0.2 other-speaker ratio, 200 ms minimum) and L=2 s are free parameters; a short sensitivity note or appendix table would strengthen reproducibility claims.
  3. Fig. 2: the qualitative NLL spikes are helpful; adding one hold-instead or shift-instead example would better cover the five perturbation types shown in Table III.
  4. Related work: DualTurn is cited as arXiv:2603.08216 (2026); ensure the citation is stable and that any concurrent full-duplex evaluation papers are briefly positioned against TurnNat’s unified likelihood framing.
  5. Notation: c_t is defined as the index of b_t (Eq. 3) but sometimes referred to loosely as the “state”; a one-line clarification that p_θ(t;x)[c_t] is the probability of the observed joint pattern would help readers less familiar with VAP.

Circularity Check

0 steps flagged

No significant circularity: TurnNat’s NLL score is an intentional atypicality definition tested on held-out, human-validated perturbations, not a fit or self-definition of the evaluation labels.

full rationale

The derivation chain is: (i) train a causal future two-speaker voice-activity model only on natural Seamless Interaction train data (Eq. 5, Table I); (ii) define frame NLL as −log pθ(t;x)[ct] (Eq. 7), pool over VAD-derived TBUs (Eqs. 1, 8), and aggregate mean+tail into mθ(x) (Eqs. 9–10); (iii) rank natural vs. five controlled local perturbations on a held-out test-derived paired benchmark, with human A/B and rating validation independent of the scorer (Tables II–IV). Nothing in this chain is forced by construction: the model is not trained or fitted on perturbation labels; α is chosen by development future-activity loss, not by maximizing Accpair; DualTurn and VAP are external architectures (no author-overlap uniqueness theorems); and the claim is empirical discrimination on that benchmark, not a first-principles identity. Using NLL of observed future activity as “atypicality” is the metric’s definition, then tested—standard likelihood-based evaluation, not circular reduction of prediction to input.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 2 invented entities

The central claim rests on standard probabilistic modeling of voice activity, domain assumptions about what constitutes natural turn-taking, and a small set of hand-chosen aggregation and weighting constants. No new physical entities are postulated; TBUs and the 256-state target are operational definitions built from prior VAD and VAP work.

free parameters (5)
  • TBU training weight α = 8 (best); also tested 1 and 3
    Hand-chosen emphasis on boundary frames during training; best reported value α=8 selected on the development set and benchmark.
  • Mean–tail mix λ = 0.5
    Equal weight between mean and top-ρ TBU NLL in the final score; fixed at 0.5 without extensive search reported.
  • Tail fraction ρ = 0.25
    Fraction of highest-NLL TBUs averaged for the tail term; fixed at 0.25.
  • TBU pre-boundary length L and island-merge thresholds = L=2s; gap 1.0s; ratio 0.2; min 200ms
    L=2 s, merge gap ≤1.0 s, other-speaker active ratio ≤0.2, discard islands <200 ms; design choices that define which frames enter the score.
  • Perturbation magnitude ranges = 1.2–2.5 s / 2–3 BCs
    1.2–2.0 s late, 1.2–2.5 s early, 2–3 inserted backchannels; chosen to be clearly outside typical human timing rather than fitted to maximize metric separation.
axioms (4)
  • domain assumption Future two-speaker voice-activity patterns over a 2 s horizon (4 non-uniform bins, 256 joint states) are a sufficient probabilistic target for turn-taking naturalness.
    Inherited from VAP and used as the sole likelihood source for scoring (Section III-C).
  • domain assumption Models trained only on natural human–human dialogue assign higher likelihood to natural timing than to unnatural timing, so NLL is a valid atypicality measure.
    Core premise of the likelihood-based framework (Abstract, Section III-A).
  • domain assumption Local edits that preserve speaker identity, lexical content, and most context isolate turn-taking naturalness differences.
    Justifies the paired perturbation benchmark design (Section IV).
  • standard math Standard categorical cross-entropy / NLL training and Softmax over 256 states are valid.
    Equations (4)–(7).
invented entities (2)
  • Turn-taking boundary units (TBUs) no independent evidence
    purpose: Select the frames whose future-activity NLL is pooled into the dialogue-level score, focusing evaluation on onset/offset regions.
    Operational definition built from cleaned VAD islands; not an independent physical entity, but a new intermediate construct introduced for scoring.
  • TurnNat dialogue-level naturalness score m_θ(x) no independent evidence
    purpose: Single scalar that ranks natural vs. perturbed clips across heterogeneous failures.
    Defined as the negative of a λ-weighted mean+tail of TBU NLLs; the paper’s primary output quantity.

pith-pipeline@v1.1.0-grok45 · 17163 in / 3179 out tokens · 26343 ms · 2026-07-12T08:52:25.190667+00:00 · methodology

0 comments
read the original abstract

Turn-taking naturalness is central to full-duplex spoken dialogue systems, yet its automatic evaluation remains limited. Existing evaluations often rely on human judgments or behavior-specific timing metrics, making it difficult to compare heterogeneous timing failures within a unified framework. We propose TurnNat, a likelihood-based framework for automatic turn-taking naturalness evaluation in two-channel spoken dialogue. A causal turn-taking prediction model trained on natural conversations estimates future two-speaker voice-activity states, and the negative log-likelihood (NLL) of the observed future activity measures timing atypicality. TurnNat pools frame-level NLLs over turn-taking boundary units (TBUs) extracted from utterance onsets and offsets, and aggregates mean and tail TBU scores into a dialogue-level naturalness score. We further construct a controlled perturbation benchmark of paired natural and perturbed dialogue clips, validated by human naturalness judgments. Experiments on this benchmark show that TurnNat successfully identifies unnatural turn-taking perturbations across heterogeneous timing failures.

Figures

Figures reproduced from arXiv: 2607.01345 by Georgi Tinchev, Hao Zhang, Laureano Moro-Velazquez, Thomas Thebaud, Venkatesh Ravichandran.

Figure 1
Figure 1. Figure 1: Overview of the TurnNat framework. TurnNat first extracts VAD-based turn-taking boundary units from the two-channel dialogue, then uses a causal [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Representative natural–perturbed pairs. Each row shows one perturbation type, with the natural clip on the left and the perturbed clip on the right. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

    cs.CL 2026-07 conditional novelty 6.0

    An open-source benchmark for speech-to-speech models shows that current systems produce intelligible audio but diverge from human conversational behavior in latency, dialect consistency, emotional entrainment, and prosody.

Reference graph

Works this paper leans on

33 extracted references · 8 linked inside Pith · cited by 1 Pith paper

  1. [1]

    A survey of recent advances on turn-taking modeling in spoken dialogue systems,

    G. Castillo-L ´opez, G. de Chalendar, and N. Semmar, “A survey of recent advances on turn-taking modeling in spoken dialogue systems,” inProceedings of the 15th international workshop on spoken dialogue systems technology, 2025, pp. 254–271

  2. [2]

    Response timing estimation for spoken dialog systems based on syntactic completeness prediction,

    J. Sakuma, S. Fujie, and T. Kobayashi, “Response timing estimation for spoken dialog systems based on syntactic completeness prediction,” in 2022 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2023, pp. 369–374

  3. [3]

    Wavchat: A survey of spoken dialogue models,

    S. Ji, Y . Chen, M. Fang, J. Zuo, J. Lu, H. Wang, Z. Jiang, L. Zhou, S. Liu, X. Chenget al., “Wavchat: A survey of spoken dialogue models,”arXiv preprint arXiv:2411.13577, 2024

  4. [4]

    Moshi: a speech-text foundation model for real-time dialogue,

    A. D ´efossez, L. Mazar ´e, M. Orsini, A. Royer, P. P ´erez, H. J ´egou, E. Grave, and N. Zeghidour, “Moshi: a speech-text foundation model for real-time dialogue,”arXiv preprint arXiv:2410.00037, 2024

  5. [5]

    Llama- omni: Seamless speech interaction with large language models,

    Q. Fang, S. Guo, Y . Zhou, Z. Ma, S. Zhang, and Y . Feng, “Llama- omni: Seamless speech interaction with large language models,” in International Conference on Learning Representations, vol. 2025, 2025, pp. 57 607–57 624

  6. [6]

    Mini-omni: Language models can hear, talk while thinking in streaming,

    Z. Xie and C. Wu, “Mini-omni: Language models can hear, talk while thinking in streaming,”arXiv preprint arXiv:2408.16725, 2024

  7. [7]

    Freeze-omni: A smart and low latency speech-to-speech dia- logue model with frozen llm,

    X. Wang, Y . Li, C. Fu, Y . Zhang, Y . Shen, L. Xie, K. Li, X. Sun, and L. Ma, “Freeze-omni: A smart and low latency speech-to-speech dia- logue model with frozen llm,” inInternational Conference on Machine Learning. PMLR, 2025, pp. 63 345–63 354

  8. [8]

    Full-duplex-bench: A benchmark to evaluate full-duplex spoken dialogue models on turn-taking capabilities,

    G.-T. Lin, J. Lian, T. Li, Q. Wang, G. Anumanchipalli, A. H. Liu, and H.-y. Lee, “Full-duplex-bench: A benchmark to evaluate full-duplex spoken dialogue models on turn-taking capabilities,”arXiv preprint arXiv:2503.04721, 2025

  9. [9]

    Mitigating response delays in free-form conversations with llm- powered intelligent virtual agents,

    M. Maslych, M. Katebi, C. Lee, Y . Hmaiti, A. Ghasemaghaei, C. Pumarada, J. Palmer, E. Segarra Martinez, M. Emporio, W. Snipes et al., “Mitigating response delays in free-form conversations with llm- powered intelligent virtual agents,” inProceedings of the 7th ACM Conference on Conversational User Interfaces, 2025, pp. 1–15

  10. [10]

    Neural generation of dialogue response timings,

    M. Roddy and N. Harte, “Neural generation of dialogue response timings,” inProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 2442–2452

  11. [11]

    Toward enabling natural conversation with older adults via the design of llm-powered voice agents that support interruptions and backchannels,

    C. Liu, M. Su, Y . Xiang, Y . Huang, Y . Yang, K. Zhang, and M. Fan, “Toward enabling natural conversation with older adults via the design of llm-powered voice agents that support interruptions and backchannels,” inProceedings of the 2025 CHI conference on human factors in computing systems, 2025, pp. 1–22

  12. [12]

    A mixed-methods approach to understanding user trust after voice assistant failures,

    A. Baughan, X. Wang, A. Liu, A. Mercurio, J. Chen, and X. Ma, “A mixed-methods approach to understanding user trust after voice assistant failures,” inProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, 2023, pp. 1–16

  13. [13]

    Patient engagement with conversational agents in health applications 2016–2022: a systematic review and meta-analysis,

    K. E. Cevasco, R. E. Morrison Brown, R. Woldeselassie, and S. Kaplan, “Patient engagement with conversational agents in health applications 2016–2022: a systematic review and meta-analysis,”Journal of medical systems, vol. 48, no. 1, p. 40, 2024

  14. [14]

    Empathic conversational agent platform designs and their evaluation in the context of mental health: systematic review,

    R. Sanjeewa, R. Iyer, P. Apputhurai, N. Wickramasinghe, and D. Meyer, “Empathic conversational agent platform designs and their evaluation in the context of mental health: systematic review,”JMIR Mental Health, vol. 11, p. e58974, 2024

  15. [15]

    Talking turns: Benchmarking audio foundation models on turn-taking dynamics,

    S. Arora, Z. Lu, C.-C. Chiu, R. Pang, and S. Watanabe, “Talking turns: Benchmarking audio foundation models on turn-taking dynamics,”arXiv preprint arXiv:2503.01174, 2025

  16. [16]

    Turngpt: a transformer-based language model for predicting turn-taking in spoken dialog,

    E. Ekstedt and G. Skantze, “Turngpt: a transformer-based language model for predicting turn-taking in spoken dialog,” inFindings of the Association for Computational Linguistics: EMNLP 2020, 2020, pp. 2981–2990

  17. [17]

    Turn-taking in conversational systems and human-robot interaction: a review,

    G. Skantze, “Turn-taking in conversational systems and human-robot interaction: a review,”Computer Speech & Language, vol. 67, p. 101178, 2021

  18. [18]

    V oice activity projection: Self-supervised learning of turn-taking events,

    E. Ekstedt and G. Skantze, “V oice activity projection: Self-supervised learning of turn-taking events,” inProc. Interspeech 2022, 2022, pp. 5190–5194

  19. [19]

    Predicting turn-taking and backchannel in human-machine conversations using linguistic, acoustic, and visual signals,

    Y . Lin, Y . Zheng, M. Zeng, and W. Shi, “Predicting turn-taking and backchannel in human-machine conversations using linguistic, acoustic, and visual signals,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), 2025, pp. 15 310–15 322

  20. [20]

    Turn-taking and backchannel prediction with acoustic and large language model fusion,

    J. Wang, L. Chen, A. Khare, A. Raju, P. Dheram, D. He, M. Wu, A. Stol- cke, and V . Ravichandran, “Turn-taking and backchannel prediction with acoustic and large language model fusion,” inICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 12 121–12 125

  21. [21]

    Dualturn: Learning turn-taking from dual-channel generative speech pretraining,

    S. Rajaa, “Dualturn: Learning turn-taking from dual-channel generative speech pretraining,”arXiv preprint arXiv:2603.08216, 2026

  22. [22]

    Ntpp: Generative speech language modeling for dual- channel spoken dialogue via next-token-pair prediction,

    Q. Wang, Z. Meng, W. Cui, Y . Zhang, P. Wu, B. Wu, I. King, L. Chen, and P. Zhao, “Ntpp: Generative speech language modeling for dual- channel spoken dialogue via next-token-pair prediction,”arXiv preprint arXiv:2506.00975, 2025

  23. [23]

    From reaction to prediction: Experiments with compu- tational models of turn-taking,

    D. Schlangen, “From reaction to prediction: Experiments with compu- tational models of turn-taking,”Proceedings of Interspeech 2006, Panel on Prosody of Dialogue Acts and Turn-Taking, 2006

  24. [24]

    Data-driven models for timing feedback responses in a map task dialogue system,

    R. Meena, G. Skantze, and J. Gustafson, “Data-driven models for timing feedback responses in a map task dialogue system,”Computer Speech & Language, vol. 28, no. 4, pp. 903–922, 2014

  25. [25]

    Opportunities and obligations to take turns in collaborative multi-party human-robot interaction,

    M. Johansson and G. Skantze, “Opportunities and obligations to take turns in collaborative multi-party human-robot interaction,” inProceed- ings of the 16th annual meeting of the special interest group on discourse and dialogue, 2015, pp. 305–314

  26. [26]

    Towards a general, continuous model of turn-taking in spoken dialogue using lstm recurrent neural networks,

    G. Skantze, “Towards a general, continuous model of turn-taking in spoken dialogue using lstm recurrent neural networks,” inProceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, 2017, pp. 220–230

  27. [27]

    Multi- lingual turn-taking prediction using voice activity projection,

    K. Inoue, B. Jiang, E. Ekstedt, T. Kawahara, and G. Skantze, “Multi- lingual turn-taking prediction using voice activity projection,” inPro- ceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation (lrec-coling 2024), 2024, pp. 11 873–11 883

  28. [28]

    Prompt-guided turn-taking prediction,

    K. Inoue, M. Elmers, Y . Fu, Z. H. Pang, D. Lala, K. Ochi, and T. Kawahara, “Prompt-guided turn-taking prediction,” inProceedings of the 26th Annual Meeting of the Special Interest Group on Discourse and Dialogue, 2025, pp. 146–151

  29. [29]

    Real-time and continuous turn-taking prediction using voice activity projection,

    K. Inoue, B. Jiang, E. Ekstedt, T. Kawahara, and G. Skantze, “Real-time and continuous turn-taking prediction using voice activity projection,” arXiv preprint arXiv:2401.04868, 2024

  30. [30]

    Visual cues enhance predictive turn-taking for two-party human interaction,

    S. O. Russell and N. Harte, “Visual cues enhance predictive turn-taking for two-party human interaction,” inFindings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vienna, Austria: Association for Computational Linguistics, Jul. 2025, pp. 209–221. [Online]. Available: https://aclantholo...

  31. [31]

    Full-duplex-bench v1. 5: Evaluating overlap handling for full-duplex speech models,

    G.-T. Lin, S.-Y . S. Kuan, Q. Wang, J. Lian, T. Li, S. Watanabe, and H.-y. Lee, “Full-duplex-bench v1. 5: Evaluating overlap handling for full-duplex speech models,” inICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2026, pp. 19 447–19 451

  32. [32]

    Seamless inter- action: Dyadic audiovisual motion modeling and large-scale dataset,

    V . Agrawal, A. Akinyemi, K. Alvero, M. Behrooz, J. Buffalini, F. M. Carlucci, J. Chen, J. Chen, Z. Chen, S. Chenget al., “Seamless inter- action: Dyadic audiovisual motion modeling and large-scale dataset,” arXiv preprint arXiv:2506.22554, 2025

  33. [33]

    Silero V AD: Pre-trained Enterprise-grade V oice Activity Detector,

    Silero Team, “Silero V AD: Pre-trained Enterprise-grade V oice Activity Detector,” https://github.com/snakers4/silero-vad, 2024, gitHub reposi- tory. Accessed: 2026-05-01