{"id":"5b9f6e72-2a3b-4b54-b3ad-b07c8bfbaa0a","arxiv_id":"2412.13265","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Webcam-driven avatars, which mirror head and facial motion, outperform static pictures and audio-driven avatars on comfort, inclusivity, and avatar satisfaction in all-avatar work meetings.","lead":"In a controlled study with 68 employees, webcam-driven animated avatars made meeting participants feel more comfortable, included, and satisfied than static pictures or audio-only animated avatars. The result suggests that adding real facial motion to avatars could make avatar meetings feel considerably closer to video calls.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Effectiveness claim rests on a Test of Proportions that treats 68 nested responses as independent; a cluster-aware reanalysis could erase the only significant effectiveness contrast, so the abstract's first claim is not yet established.","rationale":"I agree with the reader that the paper's directional preference for webcam animation is plausible and well supported by the comfort, inclusivity, and satisfaction results and by the qualitative themes. My stress-test focuses on the one quantitative result the abstract leads with: effectiveness. The Test of Proportions in Section 4.1.1 uses a formula (supplementary Eq. 1) that assumes independent Bernoulli draws. But QMO-eff asks whether the group made a final decision, which is a shared event, so the 68 responses are not 68 independent observations and the z-statistic overstates precision. Because this is the only effectiveness contrast that reaches significance, the central claim \"improved meeting effectiveness over the picture modality\" depends on it. A cluster-aware reanalysis is a concrete, feasible check using existing data, unlike the video-baseline limitation, which would require a new study. I therefore make the clustered analysis the load-bearing concern. The missing video condition remains a real overreach in the conclusion and should be softened or reframed, but it is an interpretive limitation rather than a threat to the within-experiment comparison. I do not move the reader's CONDITIONAL verdict: the paper needs revisions, and if the effectiveness contrast falls, the abstract and H1a must be trimmed; if it survives, the main empirical claim is intact.","tokens_in":30379,"tokens_out":6376,"duration_ms":64473,"concrete_test":"Reanalyze QMO-eff from the raw data with group as the unit: for each of the 16 groups, code the group's final-decision outcome in each modality (e.g., unanimous or majority response), then run McNemar's exact test for W vs. S and W vs. A; alternatively, fit a mixed-effects logistic regression with random intercepts for group and participant and report the modality contrast. If the W-vs-S contrast is no longer p < 0.05 under either cluster-aware analysis, the abstract's effectiveness claim and H1a should be weakened to comfort, inclusivity, and satisfaction only.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline quantitative support for \"improved meeting effectiveness\" is the binary QMO-eff comparison in Section 4.1.1 (Table 2): 96% vs. 79% for \"the group made a final decision,\" z = 2.85, p = 0.004 for webcam vs. static. The outcome is a property of the group, yet each of the 68 participants nested in 16 groups is treated as an independent Bernoulli observation in the supplementary Test of Proportions (Eq. 1). If a group did or did not reach a decision, all members of that group answer accordingly, so responses within a group are strongly correlated. The effective sample for this contrast is at most 16 groups, not 68 participants. A plausible group-level reading (e.g., 16/16 vs. 13/16) would not be significant at conventional levels. The same nesting affects the paired Wilcoxon tests for comfort, inclusivity, and satisfaction, though those are within-subject and less dependent on a single shared outcome. Because the abstract's first sentence singles out effectiveness, the claim is not robust until QMO-eff is reanalyzed with group as the unit of analysis or with a mixed-effects logistic model including a random intercept for group. The missing real-video condition is a separate limitation for the \"alternative to video\" conclusion, but the clustered effectiveness test is the load-bearing statistical issue.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a within-subjects, mixed-methods experiment with 68 employees of one technology company, organized into 16 groups, comparing three avatar animation modalities in all-avatar work videoconferencing: static picture, audio-driven animation, and webcam-driven animation. The quantitative measures are meeting outcome factors (effectiveness, alignment, comfort, inclusivity) and avatar satisfaction factors (self-expressive perception, other-expression perception, preference), analyzed with a two-proportion z-test for effectiveness and Wilcoxon signed-rank tests for the ordinal outcomes. The qualitative component is a thematic analysis of open-ended questionnaire responses. The authors claim that webcam-animated avatars improve meeting effectiveness, comfort, inclusivity, and avatar satisfaction relative to the other two modalities, and they conclude that webcam-animated avatars are a plausible alternative to video in work meetings.","tokens_in":30598,"tokens_out":4636,"duration_ms":44577,"significance":"If the headline effectiveness result survives appropriate statistical modeling, this would be a valuable contribution: the study uses real distributed work teams, within-subject counterbalancing, personalized avatars, and a systematic qualitative framework, which are genuine strengths relative to much prior avatar research. The qualitative thematic structure and the distinction between preference-related and outcome-related factors are also useful for future design work. However, the central quantitative claim that webcam animation 'improved meeting effectiveness' currently rests on an analysis that ignores the nesting of participants in groups, and the paper contains internal inconsistencies about alignment and an unsupported 'alternative to video' conclusion. The paper's contribution is therefore conditional on a statistical reanalysis and a recalibration of the claims.","major_comments":[{"comment":"The effectiveness contrast is analyzed with a two-proportion z-test that treats each participant's binary response as independent, but the measure QMO-eff asks 'Did the group make a final decision?', so all members of a given group share the same outcome. With 68 participants nested in 16 groups, the effective sample for this contrast is at most 16 groups, and within-group responses are strongly correlated. A group-level analysis (e.g., 16/16 vs. 13/16, a plausible reading of the reported 96% vs. 79%) would not reach conventional significance, so the reported p = 0.004 is likely anti-conservative. Please reanalyze the effectiveness data with the group as the unit of analysis or with a mixed-effects logistic regression that includes a random intercept for group, and report the intraclass correlation. Until this is done, the abstract's claim that webcam-animated avatars 'improved meeting effectiveness' is not established.","section":"§4.1.1, Table 2, Supplementary Eq. (1)"},{"comment":"The text states that 'our meeting outcome factors results supported H1a, in terms of comfort, alignment, and sense of inclusivity,' but the immediately preceding sentences and Table 2 report no significant differences for alignment (e.g., webcam vs. static: T = 372, z = .96, p = .339). This is an internal inconsistency, and the Conclusion repeats it by claiming quantitative support 'in reports on alignment.' Alignment should be removed from the list of supported H1a outcomes unless the analysis is corrected.","section":"§4.1.1, H1a summary"},{"comment":"The paper concludes that webcam-animated avatars are 'a plausible alternative to video in work meetings,' but the experiment contains no real-video condition; the only video-related evidence is qualitative participant comments. Prior literature, such as the Bente et al. study cited in Section 5.1, cannot substitute for a direct comparison in this experimental design. Please either add a video condition or explicitly restrict the conclusion to comparisons among static, audio-animated, and webcam-animated avatars. As written, the title claim and the abstract's 'alternative to video' framing go beyond what the data can support.","section":"§5.4, §6, Abstract"},{"comment":"All ordinal outcomes are tested with pairwise Wilcoxon signed-rank tests without correction for multiple comparisons across the three pairwise contrasts or across the five outcome measures, and no effect sizes or confidence intervals are reported. Some of the 'significant' differences may not survive an FDR or Bonferroni adjustment, and the 'marginally better' (p = .076) audio vs. static comfort contrast is described inconsistently in the text. A multilevel model with participant and group random effects, or at least adjusted p-values and effect sizes, would make the secondary claims more robust.","section":"§4.1, Table 2"}],"minor_comments":[{"comment":"The label 'Proportion of effective meetings' is misleading because QMO-eff is each participant's report of whether the group reached a decision; please clarify that the unit of measurement is the participant's perception, not an independently verified group outcome.","section":"Table 2, effectiveness row"},{"comment":"The table contains formatting errors such as 'z = .4.94' and 'z = .5.16', with stray periods; these should be corrected.","section":"Table 2, header formatting"},{"comment":"The demographic breakdown sums to 67 participants, while the main text reports 68; this discrepancy should be reconciled.","section":"Supplementary Table 2"},{"comment":"There is a typo, 'becuase', in the sentence about tertiary themes; the same section also uses 'QMS' in the supplementary materials while the main text uses 'QMO', so the notation should be standardized.","section":"§4.2.2"},{"comment":"The table text has missing spaces in 'Forcomfort' and 'Forinclusivity'; these are minor display issues.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper fits CSCW's scope and the empirical setup is a real strength. The main concerns are statistical and scope-related: the effectiveness claim needs a cluster-aware reanalysis, the alignment inconsistency needs correction, and the 'alternative to video' framing should be tempered unless a video condition is added. These are addressable in revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, this paper is worth your time. It's the first field-realistic comparison I know of that pits the three avatar animation modalities companies can actually ship — static, audio-driven, webcam-driven — against each other in all-avatar work meetings with employees who actually work together. That alone makes it useful for anyone building or evaluating avatar conferencing. The qualitative framework of ten themes is a solid contribution, and the authors are honest about limitations.\n\nBut the headline claim about meeting effectiveness doesn't hold as analyzed. QMO-eff is a binary 'did the group reach a final decision?' — a property of the group. The 68 participants are nested in 16 groups, yet the Test of Proportions treats all 68 as independent. The 96% vs 79% contrast is really 16-ish groups vs 16-ish groups; a cluster-aware test or mixed-effects logistic model could easily erase the only significant effectiveness result. The stress-test note is right, and this matters because the abstract's first sentence leans on effectiveness.\n\nThere are several smaller problems. Section 4.1.1 says H1a was supported for alignment, directly contradicting Table 2 which shows no alignment differences. The conclusion repeats the alignment claim. That needs fixing. The demographics are inconsistent: text says 68 participants (44 men, 24 women); supplementary table sums to 67 (43 men, 24 women). The 'plausible alternative to video' conclusion extends beyond the design — there is no video condition; prior literature is doing the work there. And the Wilcoxon comparisons are unadjusted for multiplicity, though the effects are large enough that this is a minor issue.\n\nThe qualitative deductive move — themes induced from preference data then applied to effectiveness and comfort — is somewhat circular, but the authors openly describe the saturation logic, and the primary/secondary distinction is informative rather than deceptive.\n\nNet: the core directional finding — people strongly prefer webcam-driven avatars and find them more expressive, comfortable, and inclusive — is plausible and well-supported. The effectiveness claim is not yet established. This deserves serious peer review, with the cluster analysis and wording corrections as conditions. I'd cite it after those fixes.","headline":"Useful and novel empirical comparison, but the abstract overstates effectiveness because the key analysis treats a group-level outcome as 68 independent responses.","tokens_in":31184,"tokens_out":2256,"would_cite":true,"duration_ms":20948,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In an all-avatar work meeting study, webcam-animated avatars improved perceived effectiveness over static pictures and were rated more comfortable, inclusive, and satisfying than audio-animated or static avatars.","keywords":["avatars","videoconferencing","meeting outcomes","webcam animation","non-verbal communication","avatar satisfaction","within-subjects experiment","decision-making"],"falsifier":"Reanalyze the effectiveness data with a logistic mixed-effects model that includes meeting group as a random effect; if the webcam-versus-static difference (z = 2.85, p = 0.004) becomes non-significant, the core claim would fail. Alternatively, run the same tasks with a real-video condition; if webcam-driven avatars do not match video on effectiveness or comfort, the 'plausible alternative' conclusion would not hold.","tokens_in":30162,"feed_emoji":"🧑💼","tokens_out":5478,"duration_ms":47470,"temperature":0.7,"pith_summary":"The paper argues that for remote work meetings where everyone appears as a stylized avatar, how the avatar moves matters more than how it looks. In a within-group study of 68 employees in 16 groups, webcam-driven avatars—which moved the avatar's head and facial expressions with the participant—led to higher reported meeting effectiveness than a static picture, and to higher comfort and inclusivity than both a static picture and an audio-only lip-sync avatar. Participants also strongly preferred the webcam-driven avatar, choosing it for a final discussion 95.6% of the time. The authors conclude that webcam-animated avatars are a plausible alternative to webcam video in work meetings, with meaningful motion rather than photorealism carrying the benefit.","feed_headline":"Webcam-driven avatars beat static and audio-only in work meetings","feed_subtitle":"A 68-person study finds head and face motion beats lip-sync-only or frozen avatars for decisions and comfort.","key_machinery":"The experimental machinery is a three-condition within-subjects design with personalized 3D avatars. Each participant's avatar was generated from photos and video using 3D face reconstruction with dense landmarks and blendshapes, then presented under three animation modalities: a static picture, audio-driven visemes, and full webcam-driven animation including head motion and facial expressions. The webcam-driven condition is the only one adding head movement and expressive facial dynamics, and the paper isolates that difference as the driver of the observed outcomes. Quantitatively, the load-bearing comparison is the proportion of participants reporting group agreement, tested with a two-proportion z-test. Qualitatively, the authors build a ten-factor thematic framework showing that the themes spanning preference, effectiveness, and comfort—expressiveness, non-verbal communication, emotional awareness and reaction, and cognitive load and attention—are all about motion rather than appearance.","core_discovery":"The central claim is that webcam-animated avatars produce better work-meeting experiences than static pictures or audio-driven lip-sync avatars. In the study, 65 of 68 participants (96%) reported reaching group agreement with the webcam modality, versus 59 (87%) with audio animation and 54 (79%) with static pictures; a test of proportions found the webcam-versus-static gap significant (p = 0.004), while webcam-versus-audio (p = 0.070) and audio-versus-static (p = 0.25) gaps did not reach significance. Comfort and inclusivity were rated significantly higher for webcam than for either alternative, and both self- and other-expression were rated highest for webcam. Alignment was the one outcome with no significant differences across conditions. The authors interpret the qualitative data as showing that meaningful movement—expressiveness, non-verbal communication, emotional awareness, and cognitive load—outweighs visual realism for meeting outcomes.","pith_inferences":["The paper's conclusion that webcam-animated avatars are a plausible alternative to video goes beyond the measured data: no real-video condition was run, so equivalence with video is an extrapolation from prior findings on motion fidelity rather than a directly tested result.","If the effectiveness benefit is real, the likely mechanism is conversational cueing—head nods and smiles signal turn-taking and agreement—so the benefit should be strongest in meetings with ambiguous turn-taking, such as brainstorming, and weakest in structured presentations.","The statistical analysis treats each participant's binary response as independent despite 68 participants nested in 16 groups that made decisions together; a multilevel re-analysis could either strengthen or temper the headline result.","A direct testable extension would compare webcam-driven avatars against real webcam video in the same decision-making task; if avatars match video on effectiveness and comfort, the 'plausible alternative' claim would become a measured equivalence."],"forward_implications":["Videoconferencing platforms should treat webcam-driven head and face motion as a core feature of avatar systems, not an optional extra, since the qualitative data tie meeting outcomes to motion rather than visual realism.","Organizations can offer webcam-animated avatars as a viable video-off option for employees who want privacy or cannot use cameras, without sacrificing perceived meeting effectiveness.","Audio-only lip-sync animation may deliver only modest gains over a static picture for meeting outcomes; the significant differences in the study appeared in self- and other-expression, not in effectiveness or comfort.","The 95.6% preference for webcam animation suggests that once users experience the modality, adoption in all-avatar meetings is likely to be high.","Future avatar systems should explore adding hand gestures, which participants explicitly requested, if motion fidelity is the mechanism behind the observed benefits."],"supporting_citations":[{"why":"Prior avatar study in mixed-modality conferencing showing that head motion and motion fidelity affect trust and preference; supplies the motivation and point of comparison for webcam-driven motion.","marker":"[49]"},{"why":"Longitudinal study of cartoon versus realistic avatars in real work meetings, showing stylized avatars can be accepted and providing a baseline for the claim that meaningful movement matters.","marker":"[23]"},{"why":"3D face reconstruction with dense landmarks; the technical method used to generate the personalized avatars in the experiment.","marker":"[75]"},{"why":"Shows that facial animation increases avatar self-identification through the enfacement illusion, supporting the motion-fidelity-over-realism argument.","marker":"[27]"},{"why":"Survey of knowledge workers on avatar acceptance; supports using stylized professional avatars and motivates the video-alternative conclusion.","marker":"[52]"},{"why":"Source of the adapted questions on meeting effectiveness and inclusiveness used in the post-task surveys.","marker":"[19]"},{"why":"Source of the meeting-fatigue and comfort measure adapted for the comfort factor in the study.","marker":"[24]"},{"why":"Prior study of avatar preference measured through actual choice, used as the model for the focus-group choice preference metric.","marker":"[76]"}],"fun_headline_variants":["Webcam-driven avatars beat static and audio-only in meetings","Head and face motion beats lip-sync or frozen avatars for group decisions","Webcam-animated avatars improve comfort and inclusion in work calls","Moving avatars beat static ones for meeting effectiveness and satisfaction","Avatar motion, not realism, drives better work meeting experiences"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 68 participants' answers can be treated as statistically independent even though they worked in 16 groups that reached joint decisions; a second, related assumption is that prior work on motion fidelity can stand in for a direct video baseline that the experiment never ran.","fun_headline_variants_meta":{"raw":{"variants":["Webcam-driven avatars beat static and audio-only in meetings","Head and face motion beats lip-sync or frozen avatars for group decisions","Webcam-animated avatars improve comfort and inclusion in work calls","Moving avatars beat static ones for meeting effectiveness and satisfaction","Avatar motion, not realism, drives better work meeting experiences"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000287,"raw_usage":{"total_tokens":1686,"prompt_tokens":945,"completion_tokens":741,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":653}},"tokens_in":561,"tokens_out":741,"duration_ms":7205,"temperature":1.0,"reasoning_tokens":653,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:17:21.669489+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Reanalyze the effectiveness data with a logistic mixed-effects model that includes meeting group as a random effect; if the webcam-versus-static difference (z = 2.85, p = 0.004) becomes non-significant, the core claim would fail. Alternatively, run the same tasks with a real-video condition; if webcam-driven avatars do not match video on effectiveness or comfort, the 'plausible alternative' conclusion would not hold.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior avatar study in mixed-modality conferencing showing that head motion and motion fidelity affect trust and preference; supplies the motivation and point of comparison for webcam-driven motion."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"3D face reconstruction with dense landmarks; the technical method used to generate the personalized avatars in the experiment."},{"cited_title":"The Work Avatar Face-Off: Knowledge Worker Preferences for Realism in Meetings","cited_arxiv_id":"2304.01405","evidence_quote":"Survey of knowledge workers on avatar acceptance; supports using stylized professional avatars and motivates the video-alternative conclusion."},{"cited_title":"Queiroz, Jeremy Bailenson, and Jeff Hancock","cited_arxiv_id":null,"evidence_quote":"Source of the meeting-fatigue and comfort measure adapted for the comfort factor in the study."}],"review_version":1}