{"id":"8775efed-6c4f-49af-8b9a-b379a94b231d","arxiv_id":"2412.14333","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A single transformer diffusion model with cross-modal adapters jointly generates co-speech gestures and expressive talking faces from audio, achieving competitive state-of-the-art metrics with 27.7M parameters.","lead":"Researchers built one shared diffusion network that generates both body gestures and facial expressions from speech, using small adapter modules so the two outputs influence each other. It matches or beats separate-model systems on most quality scores while using roughly half the parameters.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The SOTA claim is undermined by the paper's own metrics: Table 2 shows Separate and Combined ablations beat Ours on FMD/FED/Div/BC, and the paper dismisses those metrics as misleading after using them positively in Table 1; the only consistent favorable evidence is a 10-participant user study.","rationale":"The reader correctly flags missing baselines, lack of error bars, and post-hoc metric reinterpretation, but identifies the shared-latent-space capacity assumption as the weakest. I think the more load-bearing issue is empirical: the paper's own Table 2 shows the proposed architecture losing to its non-adapter ablations on the headline quality metrics, and the authors respond by arguing those metrics are misleading. That argument is only legitimate if the metrics are validated against human perception; otherwise the same logic would invalidate Table 1, where FMD/FED/Div/BC are used to claim SOTA. Additionally, even in Table 1, Ours is not the best on Jaw L1, Lmk L1, or LVD, because LS3DCG scores lower on all three. This leaves a 10-participant user study as the sole consistent support for the central claim, and that study has no statistical analysis and cannot bear the weight. A replicated, larger forced-choice user study is the single experiment that would settle whether the central claim is true. Because the evidence is insufficient rather than demonstratively false, the appropriate verdict is UNVERDICTED instead of REJECT.","tokens_in":13078,"tokens_out":8068,"duration_ms":70496,"concrete_test":"Run a preregistered user study with at least 30 participants using paired forced-choice comparisons between Ours and each of LS3DCG, DiffSHEG, TalkSHOW, Separate, and Combined on the same 12 or more held-out clips, with balanced presentation order and a pre-specified statistical test (e.g., exact binomial or Wilcoxon signed-rank), and report per-method FMD/FED/BC/Div with confidence intervals across at least 5 seeds. If Ours does not significantly win against the strongest baselines and ablations in the forced-choice preferences, the state-of-the-art claim is unsupported.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The central claim that the framework 'maintains state-of-the-art co-speech gesture and talking head generation performance' is not consistently supported by the paper's own numbers. In Table 2, Ours (27.7M parameters) is worse than both the Separate (51.2M) and Combined (25.6M) ablations on FMD (1758.13 vs 1525.92/1560.36), FED (1260.01 vs 1235.55/994.21), Div(All) (1845.15 vs 2043.56/1936.99), BC (0.763 vs 0.764/0.772), and Div(Face) (1521.68 vs 1698.91/1580.54). Section 4.4 explains these results by arguing the metrics are 'misleading' because the alternatives produce jittery or dull motions. But the same metrics are used positively in Table 1 to claim superiority over DiffSHEG and other baselines, with no independent validation that the metrics track human judgment in the claimed direction. Moreover, Table 1 itself does not give Ours the best face-error values: LS3DCG has better Jaw L1, Lmk L1, and LVD. The only evidence that the adapter-based joint architecture improves perceived quality is the user study in Table 3, which used 10 participants, reported no confidence intervals or significance tests, and mixed all methods in a single unblinded comparison. If that user study is unreliable or not reproducible, the state-of-the-art claim has no consistent quantitative support.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a joint co-speech gesture and expressive talking face generation method using a single diffusion transformer with adapter modules. Face and body motions are projected into a shared latent space and processed in parallel by shared transformer blocks, with cross-modal adapters letting the two modalities influence each other. The model is trained on the SHOW dataset with reconstruction and velocity losses, and evaluated against LS3DCG, DiffGesture, TalkSHOW, and DiffSHEG, plus ablations that remove the adapter or split the network. The authors claim state-of-the-art performance with a reduced parameter count, supported by a table of quantitative metrics and a ten-participant user study.","tokens_in":13351,"tokens_out":4325,"duration_ms":39663,"significance":"If the claims are validated, the adapter-based shared-transformer design would be a practically useful step toward parameter-efficient joint generation of weakly correlated body and face motion, and the cross-modal attention mechanism is a reasonable way to let the two tasks inform each other. The manuscript offers a clear architecture, a released code repository, an ablation study, and a user study, which are strengths for reproducibility. However, the quantitative evidence is currently inconsistent: the same metrics that are used to claim superiority in Table 1 are dismissed as misleading in the ablation section, and the only consistently favorable evidence for the adapter design is a small user study without statistical analysis. The parameter-reduction claim is also qualified by the existence of a combined baseline with fewer parameters, and the state-of-the-art claim is not supported on all reported face metrics. Overall, the central idea is promising but the evidence needs substantial strengthening before the paper can support its headline claims.","major_comments":[{"comment":"The paper uses the same quantitative metrics inconsistently. Table 2 shows that Ours is worse than both Separate and Combined on FMD, FED, Div(All), Div(Face), and BC, and Section 4.4 explains this by arguing that these metrics are misleading because the alternative models produce jittery or dull motions. Yet Section 4.3 and Table 1 use FMD, FED, Div, and BC as evidence of state-of-the-art performance over the baselines. This is internally inconsistent unless the authors provide independent validation that these metrics track human perceptual quality in the direction claimed, for example by correlating metric values with user-study ratings. Without such validation, the paper should either avoid relying on FMD/FED/Div/BC as primary evidence or provide a principled reason why they are valid for comparing against baselines but invalid for comparing against ablations.","section":"Section 4.3 vs Section 4.4, Tables 1 and 2"},{"comment":"The user study is the only evidence that the adapter architecture improves perceived quality over the Separate and Combined ablations, but it is reported with only ten participants, no confidence intervals, no significance tests, and no description of randomization, blinding, stimulus ordering, or inter-rater agreement. This is too weak to carry the central perceptual claim. The authors should either report a larger, statistically analyzed study or substantially temper the claim that the user study validates higher performance.","section":"Section 4.5, Table 3"},{"comment":"The abstract states that the framework 'maintains state-of-the-art co-speech gesture and talking head generation performance,' but Table 1 does not support this on several of the paper's own metrics: LS3DCG achieves better Jaw L1 (0.00147 vs 0.00161), Lmk L1 (0.1410 vs 0.1532), and LVD (0.0273 vs 0.0276), and DiffSHEG achieves better Div(All) (1924.78 vs 1845.15) and Div(Face) (1609.03 vs 1521.68). The paper should either qualify which metrics are being claimed as state-of-the-art or provide a principled weighting of these metrics.","section":"Table 1 and Abstract"},{"comment":"The related-work section cites EMAGE [21] as the most relevant joint holistic method, where face and body are trained jointly with cross-attention, yet EMAGE is not included in the quantitative or user-study comparisons. Since EMAGE is the closest prior work addressing the same joint task, its absence makes the 'state-of-the-art' claim for joint methods unsubstantiated. The authors should add an EMAGE comparison on the SHOW dataset or explicitly justify its omission.","section":"Section 2.3 and Section 4.3"},{"comment":"The parameter-reduction claim is overstated. In Table 2, the Combined ablation (25.6M parameters) has fewer parameters than Ours (27.7M), and LS3DCG in Table 1 has 18.7M parameters. The claim of 'significantly reduces the number of parameters' is only true when comparing to the Separate (51.2M) or Split (53.1M) two-network setups, not to all baselines. This limitation should be stated in the abstract and conclusion.","section":"Section 4.4 and Abstract"}],"minor_comments":[{"comment":"There is a typo in the text: 'p(x0,' should be 'p(x0)' or similar, and the following sentence 'where x0 represents the motion sequence' is incomplete.","section":"Section 3.2.1"},{"comment":"The phrase 'that gives us (J + 1) × 3 + E = 232 total parameters for the network to predict' is misleading: 232 is the output dimensionality of the motion representation, not the number of trainable parameters. Please reword.","section":"Section 4.1"},{"comment":"The statement 'Diversity and quality are often correlated with a trade-off between the two' is confusing as written; it appears to say the opposite of the intended meaning. Please clarify the relationship between diversity and quality.","section":"Section 4.3"},{"comment":"When describing DiffGesture, the paper says the authors 'additionally train a separate model for the face,' but it does not describe the architecture or training details of that face model, which is important for assessing the fairness of the comparison.","section":"Section 4.3"},{"comment":"The text says jitteriness 'can result in higher diversity and higher beat consistency, as discussed in Sec. 4,' but the relevant discussion is in Section 4.3, not Section 4 in general. Please provide a specific reference.","section":"Section 4.4"},{"comment":"The user-study section should report how many videos were shown per method, whether all methods used the same 12 audio clips, how participants were recruited, and whether the order of videos was randomized across participants.","section":"Section 4.5 and Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a plausible architectural idea and a reproducible experimental setup, but the headline claims currently outrun the evidence. The metric inconsistency between Tables 1 and 2, the small unanalyzed user study, and the missing EMAGE comparison are the main blockers. If the authors can provide additional validation or carefully scope their claims, the paper could be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Steve — read this one. The core idea is real: instead of adapting a pretrained backbone, they use adapter modules to let a single shared-weight transformer model two weakly correlated outputs (face and body) from scratch. That is a different design from TalkSHOW, EMAGE, and DiffSHEG, all of which keep face and body in separate branches or modules. The parameter reduction is concrete and quantified: 27.7M vs 51.2M for separate models in their own ablation. The architecture is clearly described, the ablation includes exactly the right control conditions (Separate, Combined, Split), and the code is linked. Credit where due: this is a thoughtful step toward genuinely joint generation.\n\nThe problem is the evidence for the headline claim. The stress-test note is right. In Table 2, their own Separate and Combined ablations beat Ours on FMD, FED, Div(All), BC, and Div(Face). The paper then dismisses those metrics as misleading because the alternatives produce jittery or dull motion. But the same metrics are used in Table 1 to claim state-of-the-art results over DiffSHEG and the others. That is a contradiction the paper never resolves. No independent validation shows that FMD/FED reward jitter and dullness while human judgment penalizes them — that is asserted, not demonstrated.\n\nWhat the paper does have in favor of Ours is the face error metrics: Jaw L1, Lmk L1, and LVD are consistently better in both Table 1 and Table 2. And the user study, with ten participants and no confidence intervals or significance testing, rates Ours above every baseline. But that study is small, unblinded, and mixes methods without statistical support — in other words, it is a weak reed to carry the full weight of a SOTA claim.\n\nOther gaps are more routine but still real: no comparison to EMAGE, which is the closest joint method; only the SHOW dataset with four speakers; no error bars or significance tests anywhere; training hyperparameters not fully specified.\n\nThe good news is that the flaws are fixable. A revision can add proper significance testing, run EMAGE on the same protocol, and either validate the metric story with a calibration study or simply stop using FMD/FED/Div as positive evidence in Table 1. The architectural contribution deserves a serious referee even though the current evaluation does not support the strongest claims.\n\nWould I bring it to reading group? Maybe — as an example of a good idea with an evaluation gap. I would not cite it for the SOTA claim in the next year, but I would cite it as related work if I worked on joint face-body generation. Send it to peer review; the idea is worth the referees' time.","headline":"Genuinely new adapter-based joint face+body architecture, but the SOTA claim is undermined by the paper's own ablation numbers; the case rests on a small user study.","tokens_in":13960,"tokens_out":1676,"would_cite":false,"duration_ms":17633,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single diffusion transformer with shared weights and adapter modules generates both co-speech gestures and an expressive talking face, claiming state-of-the-art quality on the SHOW dataset with only 27.7M parameters.","keywords":["co-speech gesture generation","talking head generation","diffusion models","adapter modules","shared latent space","cross-modal attention","parameter-efficient generation","SHOW dataset"],"falsifier":"Train the same architecture on a larger corpus such as BEAT with more speakers, and compare per-modality FMD and landmark L1 against two separate same-capacity models. If the shared-network model is meaningfully worse on both modalities, the shared-latent premise fails. Alternatively, remove the cross-modal attention heads from the adapters; if face and body metrics do not change, the claimed cross-modal influence is not doing the work attributed to it.","tokens_in":12797,"feed_emoji":"🎭","tokens_out":6164,"duration_ms":51593,"temperature":0.7,"pith_summary":"This paper claims that co-speech gestures and an expressive talking face can be generated together by a single diffusion transformer whose weights are shared between the two modalities, instead of by two separate networks or branches. Small adapter modules, each made of two cross-modal attention layers followed by a bottleneck, map face and body motion into a common latent space and let each modality influence the other. On the SHOW dataset the method reports the best FMD, FED, jaw L1, landmark L1, and landmark velocity distance among the compared baselines while using 27.7M parameters, about half of the 51.2M needed by the authors' own separate two-network ablation. A 10-participant user study rates the generated motions above all baselines and ablations and below ground truth alone. If correct, the result would make joint full-body avatars cheaper to train and deploy without sacrificing motion quality.","feed_headline":"One diffusion transformer now drives both gestures and face","feed_subtitle":"A single 27.7M-parameter network with adapters matches separate models on SHOW motion quality.","key_machinery":"The load-bearing object is the adapter-augmented shared transformer block. Each block contains four adapters per layer, two for the face stream and two for the body stream; each adapter begins with two cross-modal attention layers that exchange information between the face and body sequences, then passes through a bottleneck of downsampling, activation, hidden, and upsampling layers. The transformer's weights are shared by both modalities, with only lightweight per-modality projection layers at input and output, and the model is trained to predict the clean sample directly at each diffusion timestep plus a velocity-smoothing loss. At inference, long sequences are stitched from overlapping windows using the last $M$ frames of the previous segment as the seed for the next, with linear interpolation across the overlap.","core_discovery":"On its own terms, the paper's discovery is that adapter modules, originally designed to adapt pretrained large models, can instead be trained from scratch inside a single randomly initialized transformer to unite two weakly correlated tasks. The network receives face parameters (one jaw joint plus 100 expression blend shapes) and body parameters (43 joints) through separate projection layers, feeds them in parallel through shared transformer blocks, and uses adapter cross-attention so face and body denoising inform each other. Trained on the SHOW dataset with the standard 80/10/10 split, the model reports the lowest FMD (1758.13), FED (1260.01), jaw L1 (0.00161), landmark L1 (0.1532), and LVD (0.0276) among LS3DCG, DiffGesture, TalkSHOW, and DiffSHEG, with a parameter count of 27.7M, and it receives the highest user-study scores among all non-ground-truth methods. The paper concludes that sharing one transformer with small adapters captures the weak correlation between gesture and facial motion while avoiding the parameter duplication of separate networks.","pith_inferences":["If the shared-latent assumption holds beyond SHOW's four speakers, the same adapter pattern could be extended to more than two output streams, such as hand or eye-gaze channels, without multiplying the core transformer's parameter count.","A testable extension is to train on a larger multi-speaker corpus and measure per-modality quality; the parameter-efficiency claim would remain meaningful only if the joint model keeps both body and face metrics within a small margin of dedicated separate models.","The paper's own discussion of jittery baselines scoring well on diversity and beat consistency suggests that current Fréchet-style motion metrics reward high variance even when humans judge the motion unnatural; a metric that penalizes velocity noise would sharpen comparisons.","Because direct sample prediction plus a velocity loss smooths the outputs, the approach may transfer to other weakly correlated motion pairs such as speech-driven eyebrow and hand motion."],"forward_implications":["A single network, rather than two separately trained models, can produce both body gestures and facial motion, cutting memory and training cost.","The face and body streams influence each other through cross-modal attention, so the weak correlation between gesture and expression is exploited rather than ignored.","Long motion sequences are generated from arbitrary audio by chaining overlapping windows, with a seed gesture from the previous segment and interpolation across overlaps.","The user-study results indicate that smooth, temporally consistent face motion is perceived as more realistic than high-variance jittery motion, even where distributional metrics disagree."],"supporting_citations":[{"why":"Supplies the diffusion process and denoising training objective; the paper modifies it to predict the sample directly.","marker":"[13]"},{"why":"Frozen self-supervised speech encoder that produces the audio conditioning features.","marker":"[15]"},{"why":"Source of the adapter architecture: two cross-modal attention layers feeding a bottleneck.","marker":"[19]"},{"why":"Baseline and provider of the SHOW dataset splits and body/face motion parameterization.","marker":"[42]"},{"why":"Baseline; also source of the FMD and FED metrics used for evaluation.","marker":"[5]"},{"why":"Diffusion baseline for co-speech gestures; the paper retrains it with a separate face model for comparison.","marker":"[46]"},{"why":"GAN-based baseline that also lacks a seed pose, used for quantitative and user-study comparison.","marker":"[10]"},{"why":"Protocol and autoencoder features used to compute Fréchet motion and expression distances.","marker":"[22]"}],"fun_headline_variants":["Adapters from scratch: one network for co-speech face and body","Single transformer with adapters unifies face and body motion","Joint face+body in one diffusion transformer with adapters","Shared weights via adapters: fewer parameters, same quality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that face motion and body motion can be mapped into one common latent space and modeled by a single shared transformer with only small adapters, without either modality degrading the other.","fun_headline_variants_meta":{"raw":{"variants":["Adapters from scratch: one network for co-speech face and body","Single transformer with adapters unifies face and body motion","Joint face+body in one diffusion transformer with adapters","Shared weights via adapters: fewer parameters, same quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000773,"raw_usage":{"total_tokens":3392,"prompt_tokens":886,"completion_tokens":2506,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":502,"completion_tokens_details":{"reasoning_tokens":2436}},"tokens_in":502,"tokens_out":2506,"duration_ms":15112,"temperature":1.0,"reasoning_tokens":2436,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:19:50.576699+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same architecture on a larger corpus such as BEAT with more speakers, and compare per-modality FMD and landmark L1 against two separate same-capacity models. If the shared-network model is meaningfully worse on both modalities, the shared-latent premise fails. Alternatively, remove the cross-modal attention heads from the adapters; if face and body metrics do not change, the claimed cross-modal influence is not doing the work attributed to it.","supporting_citations":[{"cited_title":"Denoising diffu- sion probabilistic models","cited_arxiv_id":null,"evidence_quote":"Supplies the diffusion process and denoising training objective; the paper modifies it to predict the sample directly."},{"cited_title":"Hubert: Self-supervised speech representation learning by masked prediction of hidden units, 2021","cited_arxiv_id":null,"evidence_quote":"Frozen self-supervised speech encoder that produces the audio conditioning features."},{"cited_title":"Vision transformers are parameter-efficient audio- visual learners","cited_arxiv_id":null,"evidence_quote":"Source of the adapter architecture: two cross-modal attention layers feeding a bottleneck."},{"cited_title":"Diffsheg: A diffusion-based approach for real-time speech-driven holistic 3d expression and ges- ture generation","cited_arxiv_id":null,"evidence_quote":"Baseline; also source of the FMD and FED metrics used for evaluation."},{"cited_title":"Taming diffusion models for audio-driven co-speech gesture generation","cited_arxiv_id":null,"evidence_quote":"Diffusion baseline for co-speech gestures; the paper retrains it with a separate face model for comparison."},{"cited_title":"Learning speech-driven 3d conversational gestures from video","cited_arxiv_id":null,"evidence_quote":"GAN-based baseline that also lacks a seed pose, used for quantitative and user-study comparison."},{"cited_title":"Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis, 2022","cited_arxiv_id":null,"evidence_quote":"Protocol and autoencoder features used to compute Fréchet motion and expression distances."}],"review_version":1}