{"id":"addc5c10-e5e6-47ac-a213-1d50ff3f2b10","arxiv_id":"2504.20844","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Avatar head animation level changed participants' head movement during speech and perceived avatar realism, but real transmitted head motion differed from static only in one 'spoken to helpfully' rating.","lead":"This study tested whether moving, speaking avatars change how people talk in a virtual reality pub. The results matter for hearing aid testing, because virtual characters that move like real people can alter how listeners behave and how they rate the conversation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Video condition confounds head movement with facial expression, lip-sync, gaze, and skin texture; because the strongest effects are static-vs-video and auto-vs-trans never differs, the causal claim that head movement must be included overreaches the data.","rationale":"The reader identified the same load-bearing concern: the video condition is a compound stimulus that cannot isolate head movement. The manuscript's own results confirm this worry, since the largest effects are uniformly video-driven and the two genuinely avatar-based head-movement conditions ('automatic' and 'transmitted') are never significantly different from each other. The single significant static-vs-transmitted effect on Q6 is real but too narrow to support the broad conclusion that head movement is required for natural conversational behaviour. This does not invalidate the study; the data are useful, the telepresence setup is carefully described, and the null auto-vs-trans result is informative. But the central claim should be conditional on ruling out the video confound. A focused reanalysis of the three avatar conditions is feasible with the deposited data and would settle whether any head-movement-specific effect survives without video. The reader's CONDITIONAL verdict is therefore appropriate and I see no reason to change it.","tokens_in":77,"tokens_out":3395,"duration_ms":104061,"concrete_test":"Using the published dataset (doi:10.5281/ZENODO.15294859), rerun the full repeated-measures ANOVA for every measure in Table 3 after excluding the 'video' condition, leaving only the three avatar conditions ('static', 'automatic', 'transmitted'), with the same Greenhouse-Geisser correction and Bonferroni-corrected pairwise comparisons. If no significant main effect of animation level survives, or if only the Q6 static-vs-transmitted contrast remains, the conclusion must be revised to remove the causal necessity claim for head movement and instead attribute the main effects to the video condition's additional audiovisual realism. If significant static-vs-transmitted effects remain across behavioural measures, the concern is weakened and the original conclusion gains support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central conclusion claims that the representation of interlocutors must include a sufficient amount of head movement in order to elicit natural conversational behaviour. The strongest and most consistent effects, however, come from the 'video' condition (e.g., angular distance to speaker, head orientation range during speaking, realism of avatars, awareness of real environment), and that condition adds facial expression, lip movement, gaze cues, and realistic skin appearance on top of any head movement (Methods, 'Virtual environment' paragraph). This is a confound, not an isolated manipulation of head movement. The direct head-movement evidence is much weaker: no pairwise comparison of 'static' vs 'automatic' is significant in Table 3; 'static' vs 'transmitted' is significant only for Q6 ('being spoken to in a helpful way'), with avatar realism close to significance (p = 0.064); and the General discussion states that 'automatic' and 'transmitted' never differ significantly. Thus the data support the weaker claim that realistic video representation changes behaviour and subjective experience, and that transmitted head movements may help in specific self-report measures, but they do not support the stronger causal necessity claim that head movement itself is the active ingredient. The limitation section acknowledges that facial expressions, gestures, and posture need further research, but it does not explicitly flag that the video condition confound prevents attribution of the largest effects to head movement. This is the load-bearing weakness because the paper's practical recommendation for hearing-device evaluation scenarios depends on head movement being the causally responsible cue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an experiment on how the head-movement animation level of virtual interlocutor avatars affects communication behaviour, sense of presence, and conversation success in triadic conversations conducted in a telepresence VR setup. Sixteen normal-hearing participants conversed with two confederates under a 2x4 factorial design crossing quiet/noise with four visual conditions: static head, automatic speech-cued head movements, transmitted real head movements, and a video texture condition. Outcome measures were objective speech/silence durations, head orientation and translation from motion capture, and a ten-item rating questionnaire. The authors report significant main effects of animation on several behavioural and subjective measures, with the largest consistent differences occurring between the static and video conditions, and they conclude that representations of interlocutors must include a sufficient amount of head movement to elicit natural conversational behaviour.","tokens_in":19392,"tokens_out":5586,"duration_ms":53059,"significance":"If the causal claim about head movement were supported, the result would be directly useful for the design of interactive VR scenes for hearing-device evaluation, an area that increasingly relies on ecologically valid behavioural protocols. The paper has concrete strengths: it uses objective sensor-based behavioural measures, a realistic live telepresence conversation task, and it provides open data and analysis scripts (Zenodo links in the Data availability statement). It also supplies useful evidence that the visual representation of interlocutors changes participants' behaviour and experience in virtual conversations. However, the central conclusion is stronger than the experimental design can support because the video condition, which drives most of the significant effects, is confounded with facial expression, lip-sync, gaze, and skin texture, and because the confederate who is also the experimenter was not blind to condition. The manuscript therefore merits revision but has a solid empirical core.","major_comments":[{"comment":"The conclusion that 'the representation of interlocutors must include a sufficient amount of head movement' is not supported by the design, because the strongest and most consistent effects come from the 'video' condition, which adds facial expression, lip-sync, gaze cues, and realistic skin appearance on top of any head movement (Methods, 'Virtual environment' paragraph). The pairwise results in Table 3 show that 'static' vs 'transmitted' is significant only for Q6 ('being spoken to in a helpful way'), with Q5 (realism of avatars) at p = 0.064, and the General discussion states that 'automatic' and 'transmitted' never differ significantly. The data support the weaker claim that a more realistic visual representation of interlocutors affects participants' behaviour and experience; they do not isolate head movement as the active ingredient. The conclusions and General discussion should be reworded to this weaker claim, and the Limitations section should explicitly state that the video condition confound prevents attribution of the video-vs-static differences to head movement alone.","section":"Methods, 'Virtual environment' paragraph; General discussion; Conclusions"},{"comment":"The confederate 'Mar', who is also the experimenter controlling the session, was not blind to condition ('One confederate (Mar) controlled the ongoing experiment and was therefore not blind to the measurement conditions'). Because Mar is an active interlocutor in every triad, controls the timing of condition switches, and knows which visual condition is active, his conversational behaviour could systematically differ across conditions and thereby influence the participant's speech timing, head orientation, and subjective ratings such as Q6. This is a threat to the internal validity of the animation effects on interactive measures, and it is not acknowledged or discussed in the Limitations section. The authors should either provide an argument that this bias could not account for the observed pattern or explicitly discuss its direction and plausible magnitude.","section":"Methods, 'Design and task'"},{"comment":"The summary statement in the General discussion that the expected animation effect was 'reflected in about two thirds of our selected measures' is not consistent with Table 3. Of the eighteen dependent variables listed, only eight show a significant main effect of animation, and for two of those (speech level and Q8) no pairwise comparison survives Bonferroni correction. The narrative should not count the uncorrected main effects and the Bonferroni-adjusted pairwise differences as equivalent evidence; a more precise accounting of which measures are supported by pairwise differences would improve the accuracy of the interpretations.","section":"Results, Table 3; General discussion"}],"minor_comments":[{"comment":"The text contains 'did not find an affect of noise level'; 'affect' should be 'effect'.","section":"Introduction, second paragraph"},{"comment":"The phrase 'but indicted to be equally able to share information' should be 'but indicated to be equally able to share information'.","section":"General discussion, first paragraph"},{"comment":"The sentence 'rbal information is used in the background noise if it is offered' appears to be a typographical fragment; it should be completed or removed.","section":"Discussion, 'Measures of head movement behaviour' section"},{"comment":"'Questionniare' should be spelled 'Questionnaire'.","section":"Appendix, Table 4 caption"},{"comment":"The head orientation range (listening) row lists p = 0.9; given the F value and the text this is likely a typo for p = 0.09, which should be corrected.","section":"Results, Table 3"},{"comment":"The Bonferroni correction is described only for pairwise comparisons; the manuscript does not state whether correction was applied across the many dependent variables or only within each measure. A sentence clarifying this would help readers assess the risk of inflated Type I error across the multiple ANOVAs.","section":"Methods, 'Statistical analysis'"}],"recommendation":"major_revision","confidential_remarks":"The empirical infrastructure and open data are genuinely valuable, and the weaker claim that visual representation affects behaviour is well supported. However, the headline conclusion overreaches the design, and the non-blind experimenter confederate is an internal-validity concern that needs explicit treatment. I would encourage the editor to invite a revision that recalibrates the claims and addresses the limitations section rather than rejecting the work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid craft paper with genuinely useful new data, but the central conclusion is one step too strong. The real contribution is the 2x4 comparison of avatar head-movement modes in interactive triadic telepresence, with synchronized audio, head-motion capture, and open data and analysis scripts. That is new: earlier work from this group used non-interactive scenes or dyadic setups, and the four-level animation contrast (static, automatic, transmitted, video) has not been done in a triadic conversation before. The setup documentation is a strength — transmission delays, acoustics, motion capture, and the speech-activity pipeline are reported carefully.\n\nThe soft spots are real but localized. The video condition is confounded: it adds facial expression, lip-sync, gaze cues, and realistic skin on top of head movement. The largest and most consistent effects are static-vs-video, while static-vs-automatic is never significant and automatic-vs-transmitted never differs. So the data support the weaker claim that some form of animated or video representation changes behavior and experience, not the stronger claim that transmitted head movement specifically is the active ingredient. The authors do state in the general discussion that auto and trans never differ, but the abstract and conclusions still say the representation \"must include head movement.\" That should be softened.\n\nOther concerns are proportionately minor. One confederate is the experimenter and is not blind to condition, and participants orient more toward that person; the authors acknowledge the imbalance but not the blinding issue as a procedural risk. Several significant main effects have no significant pairwise differences after Bonferroni correction (speech level, Q8); the reporting is honest, but the conclusions lean on those main effects more than the pairwise results warrant. The speech-activity threshold is a free parameter that the limitations section mentions but does not robustness-check; that is a minor omission, not a fatal one. Sixteen participants is small but acceptable for this kind of within-subject, multi-sensor behavioral study.\n\nThe citation pattern is fine. The heavy self-citation reflects actual reuse of their own TASCAR, OVBOX, and pub-environment tools, and the relevant outside literature — Hendrikse, Hadley, Hartwig, Rogers, Aburumman — is engaged with.\n\nWho is this for? Researchers building interactive VR scenes for hearing-device evaluation, and people studying avatar-mediated communication. It is not a paradigm shift, but with the conclusion reframed it is a clean, reproducible contribution. I would send it to peer review with a request to fix the causal claim and to address the confederate blinding issue explicitly; the data deserve referee time.","headline":"Useful, well-documented empirical study of avatar head movement in triadic VR conversation, but the headline claim that head movement must be transmitted overreaches the data because the strongest effects come from a video condition that bundles head movement with facial expression, lip-sync, gaze, and skin texture.","tokens_in":19948,"tokens_out":1849,"would_cite":false,"duration_ms":23413,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"When avatars' heads move, VR conversation behaviour changes","keywords":["avatar head movement","triadic conversation","virtual reality","communication behaviour","presence","conversation success","hearing device evaluation","telepresence"],"falsifier":"A controlled comparison with the same video textures but the head motion frozen—or with animated heads but no facial expression—would separate head movement from facial cues; if static-video and moving-video conditions produce the same behaviour and ratings, the head-movement claim would be falsified.","tokens_in":18946,"feed_emoji":"👥","tokens_out":6390,"duration_ms":63862,"temperature":0.7,"pith_summary":"This paper asks whether animated head movements on virtual conversation partners change how a real person behaves and feels in a three-way VR conversation. Sixteen normal-hearing participants held free conversations with two confederates, represented by avatars whose heads were static, automated (speech-cued), driven by transmitted head-capture data, or replaced by a live video. The authors find that animation level significantly changed speech level, utterance duration, how far participants swung their heads while speaking, perceived avatar realism, and how helpfully participants felt they were spoken to. The conclusion is that virtual interlocutors must include enough head movement to elicit natural communicative behaviour.","feed_headline":"Avatar head motion changes how people talk in VR","feed_subtitle":"Three-way VR conversations showed bigger head swings and more helpful talk when avatars' heads moved versus stayed static.","key_machinery":"The carrying mechanism is the animation-level manipulation in a 2 x 4 repeated-measures design: two noise levels (quiet and babble at roughly 69 dB SPL) crossed with four interlocutor representations (static head; automated head and gaze cued by speech-level onsets; head motion transmitted from inertial sensors worn by the confederates; and a live head-and-shoulders video texture). The telepresence setup transmitted speech, head-motion data, and video with low delay to keep the conversation interactive while allowing the authors to change only the visual representation. Behavioural measures came from room microphones and optical head tracking; experience came from ten rating questions covering presence, realism, and conversation success. This design is what allows the authors to attribute changes in participant behaviour to the visual animation level rather than to the audio scene.","core_discovery":"The central claim is that head movement in virtual interlocutors is not decoration: it changes the listener's own behaviour. More animated heads made participants use a wider range of head orientations while speaking, orient their heads more accurately toward the avatar (2.1 degrees closer in video than in static), produce longer utterances (video vs. automated by 0.79 seconds), and rate the avatars as more realistic (static vs. video, +1.9 points). Participants also reported being spoken to in a more helpful way when head movement was transmitted compared with static avatars. Automated speech-cued head movements were statistically indistinguishable from transmitted movements on most measures. Background noise independently caused a 10.6 dB Lombard shift (raising the voice in noise), longer speech gaps, shorter overlaps, wider head-orientation range, and a 3 cm forward lean. The paper concludes that representing interlocutors with sufficient head movement—nodding or orienting toward the active speaker—is necessary for natural conversational behaviour in interactive VR.","pith_inferences":["The video condition bundles head movement together with facial expression, lip-sync, gaze, and skin texture; the paper's strongest evidence for 'head movement specifically' therefore remains circumstantial until a control condition freezes the head in the video or animates a video face without motion.","If the head-movement effect generalizes, static-avatar paradigms may systematically underestimate how much listeners move their heads and thus how much benefit directional microphones could provide in real conversations; a direct aided-versus-unaided comparison under static versus animated avatars would test this.","The large individual differences in head-yaw range and the use of young normal-hearing participants leave open whether older or hearing-impaired listeners, who may rely more on visual cues, would show larger animation effects; the authors acknowledge this limits generalization.","The finding that transmitted head movement made participants feel spoken to in a more helpful way, while automated movements behaved similarly, suggests that practical VR conversation systems can use simple speech-level automation rather than full motion capture."],"forward_implications":["VR scenes built to evaluate hearing devices should animate interlocutor heads—at minimum speech-cued orientation or nodding—if the goal is to observe natural head-orientation behaviour from listeners.","Because automated and transmitted head movements rarely differed significantly, speech-onset-driven animation may be a practical substitute for motion capture in such scenes.","Static avatars flatten the head-orientation range a listener uses while speaking, so tests of gaze-steered or head-tracking hearing-aid algorithms in static scenes may not reflect real conversation.","Background noise produces large, independent effects on speech level, turn-taking timing, and head movement, so noise level remains a primary lever for controlling task difficulty in evaluative VR conversations."],"supporting_citations":[{"why":"Supplies the speech-onset-cued head and gaze animation method used in the 'auto' condition and prior evidence that visual animation changes listener head movement.","marker":"[Hendrikse et al., 2018]"},{"why":"Provides the face-to-face triadic reference values and measures for utterance duration, head orientation, and speech gaps used for comparison.","marker":"[Hadley et al., 2021]"},{"why":"Demonstrates that manipulating an interlocutor's head movement changes the other speaker's movement in dyadic conversation, motivating the manipulation.","marker":"[Boker et al., 2009]"},{"why":"Shows interactive speaking with avatars changes movement behaviour and supplies the picture-card elicitation procedure.","marker":"[Hartwig et al., 2021]"},{"why":"Provides the presence factors underlying the questionnaire's presence and realism items.","marker":"[Schubert et al., 2001]"},{"why":"Provides the conversation-success statement clusters from which four questionnaire items were drawn.","marker":"[Nicoras et al., 2022]"},{"why":"Defines the speech-gap and overlap measures used at turn takes.","marker":"[Heldner and Edlund, 2010]"}],"fun_headline_variants":["Nodding avatars in VR boost natural talk","Moving avatar heads improve VR conversation realism","Avatar head motion changes speaker behavior in VR","Live avatar head motion shapes VR dialogue","VR avatars with moving heads make talk feel helpful"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 'video' condition, which produced the strongest effects, displayed a live head-and-shoulders video with real facial expressions, lip movement, gaze, and skin texture, so the conclusion that head movement specifically must be included assumes those extra cues are not what drove the effects.","fun_headline_variants_meta":{"raw":{"variants":["Nodding avatars in VR boost natural talk","Moving avatar heads improve VR conversation realism","Avatar head motion changes speaker behavior in VR","Live avatar head motion shapes VR dialogue","VR avatars with moving heads make talk feel helpful"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000387,"raw_usage":{"total_tokens":2077,"prompt_tokens":1011,"completion_tokens":1066,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":627,"completion_tokens_details":{"reasoning_tokens":998}},"tokens_in":627,"tokens_out":1066,"duration_ms":9815,"temperature":1.0,"reasoning_tokens":998,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:17:34.857341+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled comparison with the same video textures but the head motion frozen—or with animated heads but no facial expression—would separate head movement from facial cues; if static-video and moving-video conditions produce the same behaviour and ratings, the head-movement claim would be falsified.","supporting_citations":[{"cited_title":"Psychonomic Bulletin & Review 28(2): 632--640","cited_arxiv_id":null,"evidence_quote":"Provides the face-to-face triadic reference values and measures for utterance duration, head orientation, and speech gaps used for comparison."},{"cited_title":"International Journal of Audiology : 1--9doi:10.1080/14992027.2022.2095538","cited_arxiv_id":null,"evidence_quote":"Provides the conversation-success statement clusters from which four questionnaire items were drawn."}],"review_version":1}