Pith. sign in

REVIEW 3 cited by

Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.04868 v1 pith:K24N7WMU submitted 2024-01-10 cs.CL cs.HCcs.SDeess.AS

classification cs.CLcs.HCcs.SDeess.AS
keywords real-timesystemvoiceactivityaudiocontinuousmodelprediction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A demonstration of a real-time and continuous turn-taking prediction system is presented. The system is based on a voice activity projection (VAP) model, which directly maps dialogue stereo audio to future voice activities. The VAP model includes contrastive predictive coding (CPC) and self-attention transformers, followed by a cross-attention transformer. We examine the effect of the input context audio length and demonstrate that the proposed system can operate in real-time with CPU settings, with minimal performance degradation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue Robot

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Context-aware prefaces generated before a user finishes speaking reduce the delay to the main answer, but make the first response slightly later than a fixed filler.

  2. TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue

    cs.CL 2026-07 unverdicted novelty 6.0 of 10

    TurnNat introduces a likelihood-based automatic evaluation method for turn-taking naturalness in dyadic spoken dialogues using a causal prediction model and a human-validated perturbation benchmark.

  3. Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Pretrained audio-visual speech encoders adapted with LoRA improve multimodal voice activity projection for turn-taking prediction across multiple languages and a robot mediation corpus.

Pith tools