{"id":"305bb27c-311e-4293-b13b-60a447229a67","arxiv_id":"2607.02799","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"CHAT generates mutually responsive dyadic audio-visual dialogue clips from a single text prompt and yields a 50k synthetic pre-training set that improves facial reaction models on REACT 2024.","lead":"CHAT builds full two-person talking face-and-voice dialogues from one short text prompt, without needing real recorded conversations. It offers a way to mass-produce training data for digital humans while avoiding the cost and privacy issues of filming real pairs.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Downstream pre-training gains rest on REACT-related circularity that the paper itself flags, undercutting the claim that CHAT-AVD-50k is independently effective pre-training data.","rationale":"The reader correctly flags metric mismatch, missing same-task baselines, and REACT circularity as reasons for CONDITIONAL. The single most load-bearing soft spot for the joint claim (generation quality + useful synthetic pre-training data) is the circularity the paper itself acknowledges in Sec. 5.5: IFBR is pre-trained on REACT, so CHAT-AVD-50k is not an independent source. Without that pillar, the abstract's \"effective pre-training data\" and \"scalable alternative\" language rest mainly on adapted talking-face/FRG comparisons and short-clip FRCorr/FRDiv that the paper admits do not capture dialogue-level turn-taking or semantics. Generation ablations and the user study still support an engineering contribution, so the verdict stays CONDITIONAL rather than REJECT; the concrete re-train-without-REACT test would settle whether the data claim survives. Agreement with the reader is partial because the reader lists several caveats of similar weight, whereas the circularity is the one that most directly undercuts the data half of the strongest claim.","tokens_in":19321,"tokens_out":615,"duration_ms":6588,"concrete_test":"Re-run Table 5 with IFBR re-trained from scratch without any REACT data (e.g., only HDTF + non-REACT sources), then pre-train PerFRDiff/ReactDiff on the resulting CHAT-AVD-50k and fine-tune/evaluate on REACT 2024. If FRCorr/FRDist gains disappear or reverse, the independent pre-training claim fails; if they hold, the circularity concern is resolved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim includes that CHAT-AVD-50k is effective pre-training data that improves PerFRDiff and ReactDiff on REACT 2024. Sec. 5.5 states IFBR is itself pre-trained on REACT 2024, so the synthetic set is \"a large-scale diverse augmentation over a related interaction distribution rather than as a fully out-of-domain source.\" Table 5 gains (FRCorr 37.21\to40.11 / 24.19\to26.12; FRDist 94.72\to89.45 / 86.70\to83.87) therefore cannot cleanly support independent transfer or a scalable alternative to real DIAD collection. The main DIADG evaluation already uses FRCorr/FRDiv and first-10s clips against non-DIADG baselines (Sec. 5.2), so the remaining load-bearing pillar for the data claim is this downstream experiment; its circularity is the softest point of the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper defines Dyadic Interactive Audio-visual Dialogue Generation (DIADG): from a single anonymous textual scenario prompt, produce multiple diverse pairs of mutually responsive speech-and-face clips. CHAT implements this with three modules—TDG (LLM dialogue and identity text), DADG with Interactive Audio Refinement (IAR: interactive words, emotion/prosody descriptors, timestamps, sound environment, then emotion-aware TTS), and IFBG with Interactive Facial Behaviour Refinement (IFBR: SilentDiff for silent segments, multi-scale RFBG diffusion conditioned on the partner’s audio-visual behaviour, and Gaussian TCR boundary blending). The authors release CHAT-AVD-50k (50k pairs, ~1389 h) and report that CHAT beats adapted talking-face and facial-reaction baselines on FID/FVD, LSE-C/D, FRCorr/FRDiv, CSIM/LPIPS and a 60-person user study, while pre-training PerFRDiff and ReactDiff on CHAT-AVD-50k improves REACT 2024 metrics.","tokens_in":19633,"tokens_out":1160,"duration_ms":9955,"significance":"If the claims hold, the work supplies both a clean task definition for full verbal+non-verbal dyadic synthesis from text alone and a practical modular pipeline that can scale demographically diverse DIAD data without real multi-person capture. The ablations (Tables 3–4) cleanly attribute gains to IAR emotion conditioning and RFBG partner conditioning; the user study is large and multi-region; and the modular design (frozen LLMs/TTS, trainable SFBG/RFBG) is engineering-sensible. These are real contributions for digital humans and HCI data. The main caveats are evaluation scope (first 10 s, FRG metrics that the paper itself says miss dialogue-level turn-taking) and the related-distribution nature of the downstream pre-training claim, which the authors already flag in Sec. 5.5.","major_comments":[{"comment":"Sec. 5.5 and Table 5: The claim that CHAT-AVD-50k is “effective pre-training data” for interactive head generation is load-bearing for the abstract and conclusion, yet IFBR is itself pre-trained on REACT 2024 (Sec. 5.1). The paper correctly calls the set “a large-scale diverse augmentation over a related interaction distribution rather than as a fully out-of-domain source.” The reported FRCorr/FRDist gains therefore cannot support independent transfer or a clean alternative to real DIAD collection. Either add an evaluation on an independent corpus (e.g., IEMOCAP/NoXi as the authors themselves list in Sec. 7) or rephrase the abstract/conclusion to match the related-distribution interpretation already stated in Sec. 5.5.","section":null},{"comment":"Sec. 5.2 and Tables 1–2: No baseline solves DIADG; comparisons use talking-face and FRG methods adapted with CHAT-generated audio/identity or speaker video, restricted to the first ten seconds. FRCorr/FRDiv are facial-reaction metrics that the paper notes “do not fully reflect turn-taking or semantic coherence at the dialogue level.” Concurrent joint systems (AV-Flow, DualTalk, INFP, ChatAnyone, TAVID, JAM-Flow) are excluded for lack of code. The outperformance claim is therefore only relative to related-task proxies. Strengthen by (i) reporting full-clip or multi-turn metrics (turn-taking latency, backchannel rate, dialogue-level semantic coherence) and (ii) at least qualitative or partial comparison to one concurrent joint-generation system if any public demo/API exists.","section":null}],"minor_comments":[{"comment":"Eq. (4) and free parameters: P_inter, W, σ=W/3, L and l(t)=max(1,⌈L t/T⌉), and the 5–10 turn constraint are free design choices. A short sensitivity note (or fixed defaults with justification) would help reproducibility.","section":null},{"comment":"Notation: bar/hat/tilde conventions are introduced once but mixed with #V and /V in Fig. 2; a compact symbol table would reduce ambiguity.","section":null},{"comment":"Sec. 7 already lists English-only scope and REACT dependence; ensure the abstract does not over-claim “scalable alternative” without those caveats.","section":null},{"comment":"Fig. 4 qualitative comparison is useful; adding a failure case (e.g., long multi-turn drift or identity leakage) would balance the presentation.","section":null}],"recommendation":"major_revision","confidential_remarks":"The technical core (IAR+IFBR ablations, user study) is solid enough for a strong CV/multimedia venue after the two claim-scope fixes. The downstream circularity is the softest pillar and is already half-admitted by the authors; forcing a clearer abstract/conclusion or an independent corpus experiment is the right bar. Fit is good for a methods+dataset paper in this area."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean engineering paper. The real novelty is the task framing: generate diverse, paired, mutually responsive speech-and-face dialogue clips from one anonymous scenario prompt, with no pre-defined audio, video, or scripts. That is stricter than the interactive talking-face and concurrent joint-generation work they cite. They build a modular LLM + TTS + diffusion stack (TDG, DADG with IAR, IFBG with IFBR/SFBG/RFBG/TCR) and release the idea of CHAT-AVD-50k as pre-training fuel. Ablations in Tables 3–4 and the 60-person study line up: emotion conditioning and RFBG drive most of the interactivity and quality gains. They also flag their own limits in Sec. 5.2 and 7 (metrics, cost, English-only, REACT overlap). That honesty helps.\n\nThe soft spots are real but proportionate. No method solves exact DIADG, so they adapt talking-face and FRG baselines and evaluate the first ten seconds; FRCorr/FRDiv do not capture turn-taking or dialogue-level semantics—the paper says so. Concurrent systems were dropped for missing code. The downstream claim (Table 5) is the weakest pillar: IFBR is pre-trained on REACT 2024, so CHAT-AVD-50k is related-distribution augmentation, not independent transfer. They state this explicitly; the stress-test note is correct but not a hidden flaw. Closed Gemini/TTS and no code yet block full verification. Free parameters (P_inter, blend width, multi-scale schedule) are ordinary for this class of system.\n\nMath and citations look standard for a CV systems paper; no load-bearing circularity in the main generation metrics. This is for people building digital humans, reaction models, or synthetic dyadic data pipelines. It is not a theory paper and will not reorganize the field, but it is a usable step.\n\nI would send it to peer review. Expect referees to demand stronger dialogue-level metrics, cleaner external baselines when code appears, and a clearer statement that the pre-training gains are scale/diversity augmentation. Worth engaging if the dataset and code actually ship.","headline":"Solid systems paper that properly defines single-prompt DIADG and ships a working modular pipeline plus a large synthetic set; evaluation is the soft part, and the authors mostly own it.","tokens_in":20352,"tokens_out":561,"would_cite":true,"duration_ms":14322,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"From one text prompt, CHAT generates paired, mutually responsive audio-visual dialogues that can stand in for scarce real dyadic data.","keywords":["dyadic interactive audio-visual dialogue","DIADG","talking face generation","facial reaction generation","interactive audio refinement","diffusion model","synthetic dialogue dataset","REACT 2024"],"falsifier":"Train the same interactive-head models on an equal volume of real multi-turn dyadic data versus CHAT-AVD-50k, then measure whether the synthetic pre-training still improves FRCorr/FRDist and human interactivity ratings on a held-out real corpus whose distribution is independent of REACT 2024.","tokens_in":20150,"feed_emoji":"🗣️","tokens_out":849,"duration_ms":10290,"temperature":0.7,"pith_summary":"Real dyadic interactive audio-visual dialogue recordings are expensive, ethically constrained, and demographically limited, yet they are the data that virtual agents and digital humans need. This paper defines the DIADG task and presents CHAT, a modular pipeline that turns a single anonymous scenario prompt into many diverse, paired speech-and-face dialogue clips. Large language models first write multi-turn scripts and identity descriptions; an interactive audio refinement stage then produces emotion-aware, turn-taking speech with occasional listener backchannels; an interactive facial behaviour refinement stage finally turns those audio tracks into temporally coherent face videos whose silent and speaking segments respond to the partner. The resulting clips beat adapted talking-face and facial-reaction baselines on synchronisation, interactivity, identity consistency, and a 60-person user study. The same system yields the 50k-pair CHAT-AVD-50k corpus, which, used as pre-training, improves two interactive-head models on the REACT 2024 benchmark. The practical claim is that synthetic, prompt-driven DIAD data can scale coverage without further real recording.","feed_headline":"One text prompt yields paired, responsive talking dialogues","feed_subtitle":"CHAT synthesises 50k dyadic clips that beat related methods and improve interactive-head models","key_machinery":"The Interactive Facial Behaviour Refinement (IFBR) block (with Silent Facial Behaviour Generation, Responsive Facial Behaviour Generation that injects multi-scale partner audio-visual conditions into a diffusion transformer, and Temporal Continuity Refinement), working together with Interactive Audio Refinement, converts non-interactive speech and face segments into mutually responsive DIAD clips.","core_discovery":"A single textual scenario prompt is sufficient, once routed through textual dialogue generation, interactive audio refinement, and interactive facial behaviour refinement, to produce diverse, temporally coherent, and mutually responsive dyadic audio-visual dialogue pairs that match or exceed related methods on synchronisation, reaction quality, identity preservation, and human preference, and that supply effective pre-training data for interactive head generation.","pith_inferences":["Dialogue-level metrics that score turn-taking timing and semantic coherence, not only per-clip facial reaction scores, would be needed before claiming full conversational fidelity.","The same refinement blocks could be applied to multi-party or multilingual prompts once the underlying LLMs and TTS models support those languages.","Downstream personalisation of virtual agents could treat CHAT-generated identity-and-emotion pairs as a controllable prior rather than only as pre-training bulk."],"forward_implications":["Large-scale, demographically varied DIAD training sets can be expanded from text prompts without new human recording sessions.","Interactive head-generation models gain measurable reaction appropriateness and lower distance to ground-truth reactions after CHAT-AVD-50k pre-training.","The modular pipeline lets individual LLM, TTS, or talking-face components be swapped without retraining the whole system.","Synthetic DIAD clips become a practical alternative when privacy or ethics bar real dyadic capture."],"fun_headline_variants":["One text prompt generates mutual talking face dialogues","CHAT makes paired responsive audio-visual dialogues from text","Text alone yields diverse dyadic speech-face dialogue pairs","Single prompt drives interactive talking dialogues with faces","Generate coherent mutual AV dialogues from one scenario text"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The paper treats short-clip audio-visual sync scores and facial-reaction appropriateness metrics, applied to the first ten seconds and to baselines taken from related but different tasks, as adequate evidence of full dialogue-level verbal and non-verbal mutual responsiveness.","fun_headline_variants_meta":{"raw":{"variants":["One text prompt generates mutual talking face dialogues","CHAT makes paired responsive audio-visual dialogues from text","Text alone yields diverse dyadic speech-face dialogue pairs","Single prompt drives interactive talking dialogues with faces","Generate coherent mutual AV dialogues from one scenario text"]},"model":"grok-4.5","effort":"low","cost_usd":0.002732,"raw_usage":{"total_tokens":1002,"prompt_tokens":724,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":27320000,"prompt_tokens_details":{"text_tokens":724,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":204,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":724,"tokens_out":74,"duration_ms":3110,"temperature":1.0,"reasoning_tokens":204,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T06:58:18.977488+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train the same interactive-head models on an equal volume of real multi-turn dyadic data versus CHAT-AVD-50k, then measure whether the synthetic pre-training still improves FRCorr/FRDist and human interactivity ratings on a held-out real corpus whose distribution is independent of REACT 2024.","supporting_citations":[],"review_version":1}