{"id":"1be4e0af-d2d2-4b22-a958-e9f06828d403","arxiv_id":"2506.13477","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"When the synthetic voice did not match a photorealistic child avatar's face, viewers found the avatar less realistic, and sometimes preferred the sound-off version, while anger was harder to recognize without audio.","lead":"This paper builds a real-time system that makes photorealistic child avatars show emotions from synthetic voice, then tests how viewers perceive them with and without audio. It finds that mismatched voices reduce perceived realism, so in several cases viewers rated the silent clips as more realistic.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Only mismatched adult voices were tested; the claimed realism benefit of silence may be an artifact of voice-age mismatch rather than a general principle of audiovisual congruence.","rationale":"The paper's central empirical claim is that audiovisual congruence drives believability: a mismatched voice undermines well-crafted expressions, so silencing audio improves realism, while anger relies on voice. This claim motivates design guidance for child-interview training. The most load-bearing assumption is that the audio condition is a fair test of audiovisual congruence. In fact, the only audio tested is an acknowledged age-mismatched adult female voice (Section 3.2; also abstract and Section 5.1). Without a matched-voice condition, the study cannot distinguish voice-age mismatch from general audiovisual integration difficulty (e.g., lip-sync). This threatens both internal validity (attribution of cause) and external validity (application to age-appropriate voices). The reader identified this as the weakest assumption; I agree. Secondary statistical issues—non-significant main effect in Table 6 and different item sets in Table 7—reinforce that the realism claim is fragile, but the voice-age confound is the primary concern. The proposed follow-up experiment with a matched child voice would settle whether silence is beneficial only under mismatch or generally. Therefore, the reader's CONDITIONAL verdict remains appropriate; no change is needed.","tokens_in":787,"tokens_out":1112,"duration_ms":103764,"concrete_test":"Conduct a between-subjects follow-up with three conditions: matched age-appropriate child voice plus visual, original mismatched adult voice plus visual, and visual-only, holding all other stimuli, procedures, and rendering identical. Compare perceived realism and emotion recognition. If matched-voice audio+visual is rated at least as realistic as visual-only while mismatched-voice audio+visual is rated lower, the original conclusion is supported, and silence is not generally preferable. If matched-voice audio+visual is also rated less realistic than visual-only, the central claim must be reframed as a general audiovisual integration difficulty rather than voice-age mismatch.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that perceived believability hinges on audiovisual congruence, with a mismatched voice undermining expressions and silence improving realism—rests on a comparison whose only audio condition uses a knowingly mismatched voice. In Section 3.2, both child avatars are voiced with young adult female TTS voices selected by two researchers, an acknowledged confound (also in the abstract and Section 5.1). The design lacks a matched-voice condition, so the observed superiority of visual-only realism (Table 7, OR=0.26, p=.038; Table 8 facial-expression differences) cannot distinguish among three hypotheses: (a) mismatched voice specifically reduces realism, (b) any audio reduces realism due to lip-sync or animation artifacts, or (c) adult-voice-on-child-avatar mismatch drives the effect. Qualitative comments about lip-sync (Section 4.2.1) make (b) plausible. Moreover, the realism effect is not robust: the per-video facial-expression CLMM (Table 6) finds no significant main effect of condition (p=.526), only an interaction with sadness, and the Table 7 model compares conditions with different item sets. Thus the headline generalization overstates what the data can establish.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a real-time architecture that combines Unreal Engine 5 MetaHuman rendering with NVIDIA Omniverse Audio2Face to produce prosody-driven facial expressions on photorealistic child avatars, and reports a between-subjects user study (N=70) comparing audio+visual with visual-only presentation of joy, sadness, and anger. The central empirical claim is that perceived believability depends on audiovisual congruence: a mismatched adult female voice undermines even well-crafted facial expressions, silencing audio improved perceived realism, and anger was markedly harder to recognize without audio. The paper also reports avatar-morphology effects and qualitative feedback about lip-sync and uncanny-valley issues.","tokens_in":20167,"tokens_out":3580,"duration_ms":36404,"significance":"If the claims held as stated, the paper would offer useful design guidance for emotionally expressive avatars in sensitive training contexts, and the public code release is a strength. The anger-recognition drop in the visual-only condition is visible in the descriptive statistics and partly corroborated by t-tests, and the authors are transparent about the voice-age mismatch in the abstract, Section 3.2, and Section 5.1. However, the manuscript's headline generalization—that visual-only presentations are more realistic because audio mismatches undermine believability—is not supported by the models as reported: the main effect of condition on facial-expression realism is nonsignificant in Table 6, and the significant realism result in Table 7 compares conditions on different questionnaire item sets. The paper is therefore a promising empirical contribution whose interpretive claims need substantial reframing or additional analysis.","major_comments":[{"comment":"The only audio condition in this study uses young adult female TTS voices for both child avatars, a mismatch the authors acknowledge in the abstract and in Section 5.1. Because no age-matched child-voice condition is included, the observed realism advantage of visual-only clips cannot distinguish among three explanations: (a) a mismatched voice specifically reduces realism, (b) any audio reduces realism because of lip-sync or animation artifacts, or (c) the adult-voice-on-child-avatar mismatch drives the effect. The abstract and conclusion state that 'perceived believability hinges on audiovisual congruence,' but the design can only support a narrower claim about the tested adult-voice conditions. This is a load-bearing limitation for the paper's central generalization.","section":"Section 3.2 and Section 5.1"},{"comment":"The realism model in Table 7 compares audio+visual and visual-only conditions using different item sets: voice tone and dialogue content were rated only in the audio+visual condition, while the visual-only condition rated only facial expressions and visual appearance. The significant OR of 0.26 (p=.038) could therefore reflect the composition of the aggregated realism score rather than the effect of audio per se. The analysis should be rerun on the common items (facial expressions and visual appearance) only, or modeled at the item level with item type as a factor, to support the claim that audio reduces perceived realism.","section":"Section 4.1.4, Table 7"},{"comment":"The claim that 'silencing the clips improved perceived realism' overstates the evidence. In Table 6, the main effect of visual-only on facial-expression realism is not significant (p=.526); the only significant realism-related effect is the interaction with sadness (p=.023, OR=4.554). The supplementary t-tests in Table 8 show significant FDR-corrected facial-expression realism differences only for the two sad scenarios (Emory-Sad d=0.68, p=.021; Amelia-Sad d=0.67, p=.021), while the angry, joy, and other scenarios are not significant after FDR correction. The conclusion should be reframed as an emotion-specific effect, not a general realism benefit of removing audio.","section":"Section 4.1.3, Tables 6 and 8"},{"comment":"Qualitative comments about lip-sync and desynchronization make hypothesis (b) above—that any audio, regardless of voice age—is plausible. The manuscript itself notes that '[i]ssues with lip synchronization were also noted as detrimental to realism' and that 'the inclusion of voice amplified scrutiny of temporal and spatial alignment.' This strengthens the need for a matched-voice condition or at least for conclusions that explicitly acknowledge that the observed effects may be driven by synchronization artifacts rather than by voice-age congruence specifically.","section":"Section 4.2.1 and Section 5"}],"minor_comments":[{"comment":"The model-comparison table for facial realism contains internally inconsistent values: for example, the row with logLik=-380.78 and AIC=939.57 implies a parameter count that is not reported and is implausibly large, and the listed LRT values do not match the differences in logLik between adjacent rows. Please reconcile the reported fit statistics.","section":"Section 4.1.1, Table 3"},{"comment":"The emotion-recognition model comparison reports 'no.par' and a single LRT value, but it is unclear which two models are being compared. Please specify the null and alternative models and report the degrees of freedom of the chi-square statistic.","section":"Section 4.1.1, Table 2"},{"comment":"The figure contains a text-encoding artifact in the legend label ('Visual/uni00ADOnly'); please fix this. Also, the figure caption says 'color-coded bars distinguishing the experimental conditions,' but the printed version relies on gray shading; please ensure the distinction is accessible in grayscale.","section":"Figure 5"},{"comment":"The phrase 'a significant 74% reduction in the likelihood of higher realism ratings' is imprecise: OR=0.26 means the odds of a higher rating are reduced by approximately 74%, not the likelihood. Please use odds-based wording consistently.","section":"Section 4.1.4, paragraph 2"},{"comment":"The ethics statement says the study 'did not require formal ethical approval,' while also noting that some participants reported discomfort and that no content warnings were provided. For a journal submission, please clarify the institutional ethics framework under which this determination was made, and consider reporting whether any debriefing was provided.","section":"Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"The reader's stress-test concern is valid and lands: the absence of a matched-voice condition, combined with the Table 7 item-set confound, means the central realism claim is not yet established. I would be willing to reconsider if the authors either add a matched-voice comparison (e.g., a small follow-up study or reanalysis of available stimuli) or substantially soften the abstract and conclusions to the emotion-specific, condition-specific findings that the current data support. The anger-recognition result and the architecture description are solid contributions and should be preserved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth a look if you work on avatar realism or multimodal emotion. The headline claim—silencing a mismatched voice improves perceived realism of photorealistic child avatars—is real but narrower than the abstract suggests. It holds mainly for sadness, and the design can't separate voice-age mismatch from lip-sync artifacts.\n\nWhat's genuinely new: they built a real-time pipeline (UE5 MetaHuman + Audio2Face) and ran a between-subjects study (N=70) with two child avatars. The anger-recognition result is solid: visual-only anger ratings collapse (Emory 3.83→1.43, d=-2.57; Amelia 3.11→2.17), and this is consistent with known multimodal emotion perception work. They ship code and are explicit about the adult-voice confound in the limitations.\n\nThe soft spots are mostly in the realism claim. The per-video CLMM (Table 6) shows no main effect of condition (p=.526); only the 'visual-only × sad' interaction is significant (OR=4.55). The appendix t-tests confirm: after FDR correction, only the two sad clips show a realism benefit. The abstract's \"silencing clips improved perceived realism\" overgeneralizes. Also, the model in Table 7 that yields OR=0.26 compares conditions on different item sets—the visual-only group didn't rate voice or dialogue—so it's weaker evidence. And because only mismatched adult-voice audio was tested, the data can't adjudicate between 'any audio reduces realism' and 'age-mismatched audio reduces realism'; lip-sync comments in 4.2.1 make the former plausible. The 'exaggerated' Audio2Face mode is another unvaried parameter.\n\nThese are fixable. Add a matched child-voice condition (or at least a mismatched vs. matched comparison), report the sadness-specific analysis first, and soften the headline. The emotion-recognition finding stands independently and is worth publishing.\n\nFor a serious venue, I'd send it to peer review with a request for revision, not desk-reject. I'd cite the anger-recognition result in applied avatar work. Serious thinker: yes—the authors know their confound and are transparent about it.","headline":"The anger-recognition result is solid, but the claimed realism benefit of silence is real only for sadness and is confounded by the unmatched adult-voice condition; still worth peer review.","tokens_in":20709,"tokens_out":3222,"would_cite":true,"duration_ms":33735,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that perceived believability of AI child avatars is governed by audiovisual congruence: mismatched adult voices reduce realism even when facial expressions are well crafted, and anger becomes much harder to recognize…","keywords":["Human-Computer Interaction","Interactive Avatars","Emotionally Expressive Avatars","Perceived Realism","Unreal Engine MetaHuman","NVIDIA Audio2Face","audiovisual congruence","child interview training"],"falsifier":"Run the same 70-person between-subjects study again with age-matched child text-to-speech voices, using one TTS system or matching prosody more tightly; if audio+visual clips are then rated at least as realistic as visual-only clips and anger recognition in the audio condition stays high, the congruence interpretation is supported, whereas if visual-only still wins realism even with well-matched child voices, the silence advantage is caused by the audio channel itself rather than voice-age mismatch.","tokens_in":19779,"feed_emoji":"🎭","tokens_out":8145,"duration_ms":77981,"temperature":0.7,"pith_summary":"This paper tries to establish that emotionally believable child avatars, of the kind needed for training interviewers who talk with children about abuse, are limited less by rendering quality than by how well voice, face, and expression cohere. The authors built a real-time pipeline that couples Unreal Engine 5 MetaHuman rendering with NVIDIA Omniverse Audio2Face, which derives facial animation from vocal prosody, and tested two child avatars in a between-subjects study of 70 participants who rated emotional clarity, facial realism, and empathy for joy, sadness, and anger. Their central result is an audiovisual congruence effect: when a young adult female text-to-speech voice was paired with a child-like face, muted visual-only clips were rated as more realistic than clips with audio, while anger became markedly harder to recognize without sound. A sympathetic reading of the paper is that matching a voice to a face matters more than polishing either channel alone, and that high-arousal emotions like anger need audio while low-arousal emotions like sadness can stand on facial cues.","feed_headline":"Silencing avatar audio raised perceived realism in child-avatar study","feed_subtitle":"Adult voices on child faces made muted clips feel more real; silent anger became nearly unrecognizable.","key_machinery":"The load-bearing mechanism is the prosody-to-expression mapping performed by NVIDIA Omniverse Audio2Face, which converts spectral and prosodic features of synthesized speech into facial action unit activations and blendshape weights that drive a Unreal Engine 5 MetaHuman facial rig through the Live Link plugin; a two-computer split separates speech generation from rendering to keep the loop interactive. The argument's interpretive machinery is the valence-arousal asymmetry: audio cues are treated as the main carrier of emotional arousal, visual cues as the main carrier of valence, which is why anger (high arousal, negative valence) is the emotion most damaged by silence, while sadness and joy survive on visual cues. The paper also uses FACS-coded markers such as brow depression (AU1+4), lip-corner droop (AU15), brow tension (AU4), and Duchenne markers (AU6+12) to explain which facial signals remain readable without audio, and it uses an anchoring stimulus to calibrate participant ratings before the main clips.","core_discovery":"The paper demonstrates a working real-time architecture for prosody-driven facial emotion on photorealistic child avatars and reports that its user study found perceived realism fell when audio was added: the presence of audio reduced the odds of higher realism ratings by about 74% ($OR=0.26$, $p=.038$). The recognition results were emotion-dependent: sadness and joy were recognized from faces alone, with sadness stable for both avatars, whereas anger recognition dropped sharply without audio (for Emory from $M=3.83$ to $M=1.43$, for Amelia from $M=3.11$ to $M=2.17$ on the five-point scale). A three-way interaction between condition, avatar, and emotion congruency ($OR=0.212$, $p<.001$) and avatar-specific effects (Emory overall harder to read, $OR=0.202$, $p=.002$) show that the value of audio depends on which avatar and which emotion is shown. The paper also finds that facial morphology shifts interpretation: Amelia's softer features supported sadness and joy, while Emory's angular features conveyed anger better when voice was present. The authors conclude that perceived believability depends on audiovisual congruence and facial geometry together, not on fidelity in any single channel.","pith_inferences":["If the congruence interpretation extends, an adaptive design would use audio only for high-arousal moments or switch to a well-matched child voice, making silence no longer preferable; this is a natural next test the paper does not run.","A direct follow-up experiment varying voice age, TTS system, and emotion (including fear and disgust) in a factorial design would separate voice-age mismatch from general TTS prosody mismatch; if silence still beats audio with age-matched child voices, the effect is about audio itself, not age.","The non-significant empathy difference, combined with higher visual-only realism, suggests interview training could deliberately include a silent replay mode so trainees practice reading facial cues without vocal interference, a design option the paper leaves implicit.","The qualitative comments about faces looking 'scary' or 'rigid' point to a possible connection between audiovisual mismatch and the uncanny valley, implying that congruence failures, not just fidelity failures, may trigger discomfort in sensitive contexts."],"forward_implications":["For low-arousal emotions such as sadness and joy, a visual-only presentation can be rated more realistic than an audio-visual one whenever the voice is not fully congruent with the avatar's face, because silencing the clip removes a source of dissonance.","High-arousal emotions such as anger cannot be reliably recognized from facial animation alone; training applications that need anger detection must supply synchronized vocal cues or accept a large recognition drop.","Avatar facial morphology is not neutral: softer, rounder features aid low-arousal believability, while angular features make anger more salient, so designers should select geometry to match the emotional demands of the training scenario.","Voice-age congruence is a first-order design constraint for child avatars; the adult-voice compromise used here is identified as a confound that likely reduced the audio condition's realism.","The distributed two-PC pipeline shows real-time prosody-driven emotional animation is technically feasible, but the remaining realism bottleneck is low-level synchronization and prosody-to-expression alignment rather than raw rendering quality."],"supporting_citations":[{"why":"Documents that existing virtual-child interview training systems produce emotionally inert, low-fidelity avatars, the gap this architecture targets.","marker":"[4]"},{"why":"Supplies the valence-arousal asymmetry (audio carries arousal, visuals carry valence) that motivates the audio+visual vs visual-only design and frames anger's audio dependence.","marker":"[33]"},{"why":"The authors' prior finding that female voices can represent child characters in training scenarios, which justifies the young adult female TTS voice selection and is the study's main confound.","marker":"[34]"},{"why":"Justifies the low-fidelity anchoring stimulus used to calibrate participants' realism ratings before the main stimuli.","marker":"[35]"},{"why":"Provides the FACS action-unit vocabulary (e.g., AU1+4, AU15, AU4, AU6+12) used to interpret which facial cues carry sadness, anger, and joy.","marker":"[36]"},{"why":"Shows affective congruence across visual and auditory modalities is processed jointly, supporting the audiovisual congruence argument.","marker":"[37]"},{"why":"Provides evidence that cross-modal auditory-visual interaction affects perception, cited for anger's reliance on vocal intensity and spectral cues.","marker":"[38]"},{"why":"Links believability to warmth, competence, and embodiment, cited to explain why Amelia's softer morphology is perceived as more authentic.","marker":"[39]"}],"fun_headline_variants":["Silencing child avatars improved realism, but anger lost without voice","Child avatar study: audio cut realism, yet anger recognition relied on it","Audio-age mismatch reduced believability in child-avatar emotion test","Real-time child avatars: adding audio made faces less authentic"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that two young adult female text-to-speech voices, chosen by the researchers' judgment and supported by their earlier finding that female voices can represent child characters, are an acceptable stand-in for the child voices these avatars would need; if properly age-matched child voices had been used, the realism advantage of silence could shrink or disappear.","fun_headline_variants_meta":{"raw":{"variants":["Silencing child avatars improved realism, but anger lost without voice","Child avatar study: audio cut realism, yet anger recognition relied on it","Audio-age mismatch reduced believability in child-avatar emotion test","Real-time child avatars: adding audio made faces less authentic"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000524,"raw_usage":{"total_tokens":2584,"prompt_tokens":1051,"completion_tokens":1533,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":667,"completion_tokens_details":{"reasoning_tokens":1470}},"tokens_in":667,"tokens_out":1533,"duration_ms":12302,"temperature":1.0,"reasoning_tokens":1470,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:00:27.111625+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 70-person between-subjects study again with age-matched child text-to-speech voices, using one TTS system or matching prosody more tightly; if audio+visual clips are then rated at least as realistic as visual-only clips and anger recognition in the audio condition stays high, the congruence interpretation is supported, whereas if visual-only still wins realism even with well-matched child voices, the silence advantage is caused by the audio channel itself rather than voice-age mismatch.","supporting_citations":[{"cited_title":"A theoretical and empirical analysis of 2d and 3d virtual environments in training for child interview skills,","cited_arxiv_id":null,"evidence_quote":"Documents that existing virtual-child interview training systems produce emotionally inert, low-fidelity avatars, the gap this architecture targets."},{"cited_title":"Multimodal affect models: An investigation of relative salience of audio and visual cues for emotion prediction,","cited_arxiv_id":null,"evidence_quote":"Supplies the valence-arousal asymmetry (audio carries arousal, visuals carry valence) that motivates the audio+visual vs visual-only design and frames anger's audio dependence."},{"cited_title":"Is more realistic better? a comparison of game engine and gan-based avatars for investigative interviews of children,","cited_arxiv_id":null,"evidence_quote":"The authors' prior finding that female voices can represent child characters in training scenarios, which justifies the young adult female TTS voice selection and is the study's main confound."},{"cited_title":"What stimuli are necessary for anchoring effects to occur?","cited_arxiv_id":null,"evidence_quote":"Justifies the low-fidelity anchoring stimulus used to calibrate participants' realism ratings before the main stimuli."},{"cited_title":"Facial action coding system,","cited_arxiv_id":null,"evidence_quote":"Provides the FACS action-unit vocabulary (e.g., AU1+4, AU15, AU4, AU6+12) used to interpret which facial cues carry sadness, anger, and joy."},{"cited_title":"An fmri study of affective congruence across visual and auditory modalities,","cited_arxiv_id":null,"evidence_quote":"Shows affective congruence across visual and auditory modalities is processed jointly, supporting the audiovisual congruence argument."},{"cited_title":"Cross-modal interaction between auditory and visual input impacts memory retrieval,","cited_arxiv_id":null,"evidence_quote":"Provides evidence that cross-modal auditory-visual interaction affects perception, cited for anger's reliance on vocal intensity and spectral cues."},{"cited_title":"How is believability of a virtual agent related to warmth, competence, personification, and embodiment?","cited_arxiv_id":null,"evidence_quote":"Links believability to warmth, competence, and embodiment, cited to explain why Amelia's softer morphology is perceived as more authentic."}],"review_version":2}