{"id":"b65b1531-f141-40dd-b06d-a38c3591e3cb","arxiv_id":"2606.16731","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"MuVAP extends voice activity projection with face-track grounding and role-relative mapping for causal multiparty turn-taking prediction, evaluated on a new 31-hour unedited single-camera dataset.","lead":"MuVAP is a multimodal system that predicts turn-taking in group talks using one microphone and one camera by combining audio with face tracking and a simplified speaker-role mapping. This could allow robots to join everyday conversations with less hardware than current setups require.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Role-Relative Projection reduces N-speaker dynamics to fixed current/next floor-holder states, which may discard speaker identity needed for next-speaker prediction.","rationale":"The reader's weakest_assumption directly names the two assumptions (single-camera grounding + projection preserving dynamics) whose failure would invalidate the outperformance claim. Because the full text was not supplied to the reader, no stronger technical objection can be raised from the given material; the identified assumptions remain the load-bearing ones.","tokens_in":1640,"tokens_out":321,"duration_ms":34213,"concrete_test":"Re-run the 3-speaker next-speaker prediction experiment on AVCC while forcing the model to output a specific speaker ID (instead of the projected next state) and measure top-1 accuracy; if it falls below the reported baseline by >15% relative, the projection mapping does not preserve the required multiparty information.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that MuVAP outperforms baselines on next-speaker prediction in 2- and 3-speaker settings. This rests on Role-Relative Projection mapping any interaction onto a fixed current vs. next state while still supporting accurate forecasts of which specific speaker takes the floor. If the mapping collapses distinct candidate next-speakers into a single \"next\" label without retaining identity or role distinctions, the task reduces to binary hold/shift detection rather than true next-speaker identification; the reported gains would then be artifacts of the reduced output space rather than evidence that the multimodal grounding works.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces MuVAP, a causal multimodal framework extending Voice Activity Projection by grounding acoustic predictions in face tracks from a single camera and monaural audio. It proposes Role-Relative Projection to map any N-speaker interaction onto a fixed current-versus-next floor-holder state to address combinatorial complexity, introduces the 31-hour Audio-Visual Conversation Corpus of unedited single-camera recordings, and claims that MuVAP outperforms strong baselines on Shift-Hold and next-speaker prediction tasks in two- and three-speaker settings.","tokens_in":1760,"tokens_out":378,"duration_ms":34543,"significance":"If the empirical claims hold after addressing the output-space concern, the work would be significant for practical human-robot interaction by enabling speaker-aware turn-taking from minimal hardware. The new unedited dataset fills a gap for causal audiovisual modeling, and the multimodal grounding approach could generalize beyond current microphone-array or multi-camera setups.","major_comments":[{"comment":"Abstract: The central claim that Role-Relative Projection enables 'next-speaker prediction' while mapping N-speaker dynamics onto a fixed current-versus-next state requires explicit clarification on output representation. If the mapping assigns a single undifferentiated 'next' label without retaining speaker identity or role distinctions among candidate next-speakers, the task reduces to binary hold/shift detection; reported gains on next-speaker prediction would then be artifacts of the reduced output space rather than evidence that multimodal grounding identifies specific speakers.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract states that evaluations demonstrate outperformance but provides no quantitative metrics, baseline details, or ablation results; these must be presented with error analysis to allow verification of the performance claims.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful reading and for highlighting the need for explicit clarification on the output representation of Role-Relative Projection. We address the concern directly below and will revise the manuscript accordingly.","responses":[{"response":"We agree that the abstract requires explicit clarification on this point. Role-Relative Projection reduces combinatorial complexity by mapping arbitrary speaker configurations to a fixed current-versus-next floor-holder state; however, the output is not an undifferentiated binary label. Because predictions are grounded in the face tracks extracted from the single-camera view, the model produces per-track probabilities that identify which specific visual track (and therefore which individual) corresponds to the current floor-holder and which corresponds to the next floor-holder. This retains speaker identity and is distinct from pure hold/shift classification. We will revise the abstract to state the output representation explicitly (i.e., that next-speaker prediction is performed over the set of detected face tracks).","revision_made":"yes","referee_comment":"[Abstract] Abstract: The central claim that Role-Relative Projection enables 'next-speaker prediction' while mapping N-speaker dynamics onto a fixed current-versus-next state requires explicit clarification on output representation. If the mapping assigns a single undifferentiated 'next' label without retaining speaker identity or role distinctions among candidate next-speakers, the task reduces to binary hold/shift detection; reported gains on next-speaker prediction would then be artifacts of the reduced output space rather than evidence that multimodal grounding identifies specific speakers."}],"tokens_in":1259,"tokens_out":325,"duration_ms":32359,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main moves are grounding acoustic VAP predictions in face tracks from one camera and introducing Role-Relative Projection to collapse N-speaker cases into a fixed current/next state. They also collected a 31-hour unedited single-camera corpus to avoid the cuts that break causal tracking in existing data.\n\nThe practical framing for human-robot interaction is useful, and releasing the new corpus is a clear positive that others can build on. The causal multimodal setup directly targets the hardware limits of prior array-based work.\n\nThe soft spot is the projection step itself. If mapping to current versus next floor-holder discards which specific speaker is likely to take the floor, then the next-speaker task reduces to shift/hold detection and any reported improvements could be artifacts of the smaller output space rather than evidence that the face tracks add useful signal. The abstract claims gains on next-speaker prediction in two- and three-speaker settings, but without metrics, ablations, or error breakdowns it is impossible to tell whether the multimodal component is carrying the load.\n\nThis is for researchers working on lightweight turn-taking models for robots or agents in natural multiparty settings. The dataset and the reduction idea give it enough substance that a serious editor should send it to referees, with the expectation that the evaluation details will be the main point of revision.","headline":"MuVAP adds a single-camera multimodal grounding step and a new unedited corpus, but the Role-Relative Projection needs to show it preserves speaker identity for real next-speaker gains rather than just binary shift/hold.","tokens_in":2232,"tokens_out":359,"would_cite":false,"duration_ms":24916,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"MuVAP extends voice activity projection with face tracks and role-relative mapping to predict turn-taking from single-camera monaural recordings.","keywords":["multimodal turn-taking","voice activity projection","multiparty conversation","next speaker prediction","audiovisual corpus","human-robot interaction"],"falsifier":"Collecting new recordings where single-camera face tracking frequently loses tracks due to movement or occlusion and checking whether prediction accuracy then falls below audio-only baselines.","tokens_in":2540,"feed_emoji":"🎤","tokens_out":620,"duration_ms":52297,"temperature":0.7,"pith_summary":"The paper presents MuVAP as a way to predict who will speak next in group conversations using only one microphone and one camera. It grounds the audio-based voice activity model in visible face tracks to make speaker-aware forecasts. A new mapping called Role-Relative Projection reduces the problem of any number of speakers to a simple current-versus-next decision. The authors created a 31-hour dataset of natural unedited conversations to train and test this. The model beats baselines on deciding when to hold or shift the floor and on naming the next speaker in two- and three-person talks.","feed_headline":"MuVAP predicts turns from one mic and one camera","feed_subtitle":"Role-relative mapping and face tracks let the model handle two- or three-speaker talks without multi-mic arrays.","key_machinery":"Role-Relative Projection, which maps any N-speaker interaction onto a fixed current versus next floor-holder state to address combinatorial complexity.","core_discovery":"MuVAP is a causal multimodal framework that grounds acoustic predictions in face tracks from a single camera view, using Role-Relative Projection to map multiparty interactions onto a fixed current versus next floor-holder state, thereby enabling accurate shift-hold and next-speaker predictions from monaural audio in unedited wild settings, as demonstrated by outperformance of strong baselines on the Audio-Visual Conversation Corpus.","pith_inferences":["If single-view face tracking works reliably, the same approach could apply to consumer video calls for automatic turn management.","The role mapping might allow scaling to four or more speakers without exponential growth in model complexity.","Combining this with existing speech recognition could lead to fully hands-free group conversation agents."],"forward_implications":["MuVAP enables turn-taking prediction without complex microphone arrays or multi-camera setups.","It supports human-robot interaction scenarios with simple hardware.","Performance gains hold for both two- and three-speaker settings on shift-hold and next-speaker tasks.","The unedited nature of the new corpus avoids artifacts from editing cuts that break causal tracking."],"fun_headline_variants":["MuVAP grounds turns in face tracks from single camera view","Single mic and camera drive MuVAP turn predictions","Role-Relative Projection simplifies multiparty voice activity","MuVAP outperforms baselines on unedited two-speaker talks","Causal MuVAP maps multiparty states from monaural audio"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Face tracks from a single camera can reliably ground the acoustic predictions and that the Role-Relative Projection preserves enough multiparty dynamics without speaker-specific modeling.","fun_headline_variants_meta":{"raw":{"variants":["MuVAP grounds turns in face tracks from single camera view","Single mic and camera drive MuVAP turn predictions","Role-Relative Projection simplifies multiparty voice activity","MuVAP outperforms baselines on unedited two-speaker talks","Causal MuVAP maps multiparty states from monaural audio"]},"model":"grok-4.3","cost_usd":0.003232,"raw_usage":{"total_tokens":1701,"prompt_tokens":601,"num_sources_used":0,"completion_tokens":82,"cost_in_usd_ticks":32324500,"prompt_tokens_details":{"text_tokens":601,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1018,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":601,"tokens_out":82,"duration_ms":16475,"temperature":1.0,"reasoning_tokens":1018,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T03:03:54.502528+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Collecting new recordings where single-camera face tracking frequently loses tracks due to movement or occlusion and checking whether prediction accuracy then falls below audio-only baselines.","supporting_citations":[],"review_version":1}