REVIEW 3 major objections 5 minor 83 references
SubtleTalk: Generating Controllable Weakly-correlated Facial Dynamics for 3D Talking Heads via Residual Flow Matching
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read SubtleTalk claims that weakly speech-correlated facial dynamics—eyebrows, eye blinks, and head motion—become natural, diverse, and user-controllable when a deterministic lip-and-jaw prior is refined by residual flow matching under…
desk verdict Solid talking-head system with a valuable dataset, but the headline FDD/HDD gains may be partly a training-objective artifact and cannot be verified without the missing metric formulas. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is residual flow matching over a deterministic lip-and-jaw prior. The Deterministic Motion Prior (DMP) is a query-based audio-conditioned transformer that takes frozen speech content features, the identity shape, and previous motion context, and outputs only a compact target $[\psi, \theta_{jaw}]$ with head pose set to zero; it is supervised by vertex and velocity losses in FLAME vertex space. The Residual Flow Matching (RFM) stage then targets the residual $M_c - M^{prior}_c$ in normalized coordinates, learning a conditional velocity field $V_\theta(x_{t'}, t', c)$ that transports Gaussian noise to the residual endpoint along the interpolation path $x_{t'} = (1-t')x_0 + t'x_1$. The conditioning set $c$ bundles a prosody-aware audio feature (content, multi-scale acoustic, and $F_0$/energy branches), Fourier-encoded valence-arousal sequences, and global regional-intensity features injected through cross-attention and adaptive layer normalization. This separation does the argument's work: removing the strongly speech-correlated mouth motion lets the generative capacity concentrate on the ambiguous upper face and head, while a disentangled control loss keeps user-adjustable intensity signals from leaking across regions.
What would settle it
A concrete check: take a hold-out set of talking-head videos with independent upper-face ground truth, such as motion capture or manually annotated brow and eyelid trajectories, and recompute FDD and HDD against that ground truth for SubtleTalk and a deterministic regressor; if SubtleTalk's advantage shrinks or vanishes, its claim to capture weakly correlated dynamics is not supported.
Extended reading notes
Core claim
The paper's central claim is that weakly correlated facial dynamics stop being a bottleneck once strongly speech-correlated motion is separated from ambiguous motion and the ambiguous part is generated stochastically rather than regressed deterministically. Concretely, a Deterministic Motion Prior (DMP) predicts a compact mouth-and-jaw motion from speech content, shape, and previous motion context, and a Residual Flow Matching (RFM) stage models the remaining motion—eyebrows, eyelids, head pose—as a residual flow in normalized coordinates. The paper reports that on its SubtleTalk-Face test split this two-stage design achieves FDD of 4.46 versus 9.74 for the best head-motion baseline, HDD of 6.61 versus 21.71, and LVE of 11.96 versus 13.72, which it presents as evidence that the method preserves lip synchronization while producing substantially more dynamic and natural upper-face and head motion.
Load-bearing premise
The reported gains rest on the assumption that the single-camera 3D reconstruction used to create both the training labels and the motion-evaluation metrics faithfully captures eyebrow, eyelid, and head motion; if that tracking is systematically wrong, the measured improvements could come from the labeling pipeline rather than from real facial dynamics.
Editorial extensions
If this is right
- From plain audio, the model can generate eyebrow, eyelid, and head motion with naturally varying amplitudes instead of averaged or repetitive patterns, without sacrificing lip-sync accuracy.
- A user can raise or lower one regional intensity control (eyebrow, eyelid, head nod, head turn, head tilt) and expect only that region's dynamics to change, because the implicit disentangled control loss penalizes leakage into the other regions.
- Affect can be injected at inference either automatically from speech through the VA Dynamics Predictor or by giving sparse valence-arousal anchors at selected frames, so the same utterance can be rendered with different emotional dynamics.
- The constructed 74-hour, roughly 3,900-identity dataset with FLAME pseudo-labels, stricter filtering, and frame-level valence-arousal annotations gives later models larger and cleaner supervision for upper-face motion.
- Separating lips and jaw from the stochastic residual means the two objectives no longer compete; the ablation results indicate that this separation is what keeps lip error competitive while upper-face and head dynamics improve.
Reading between the lines
- A direct test beyond the paper: evaluate FDD and HDD against independent upper-face ground truth such as motion capture or manually annotated brow and eyelid trajectories; the paper's current metrics share their pseudo-label source with the training data, so this would show whether the gains reflect true facial dynamics rather than label bias.
- The residual split suggests a general recipe for other speech-to-motion tasks: couple a deterministic module to the strongly input-correlated part and let a generative model handle the weakly correlated remainder; applying the same two-stage design to body gestures would test this transfer.
- The IDC invariance loss implies that regional intensity controls could be made time-varying, per phrase or per frame, rather than window-level scalars, enabling emphasis on a specific word without retraining; whether such local control stays disentangled would be a natural extension.
- Because the VA Dynamics Predictor can infer affect from speech alone, the method opens the possibility of cross-lingual affect transfer by substituting a different speech encoder; this is not tested in the paper.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SubtleTalk, a two-stage framework for audio-driven 3D facial animation. A deterministic motion prior (DMP) first predicts lip-synchronized motion, and a residual flow matching (RFM) stage models stochastic upper-face and head dynamics in the residual space. The method introduces multi-condition controls (prosody-aware audio features, regional intensity scalars, and continuous valence-arousal signals) and an implicit disentangled control (IDC) strategy. The authors also construct SubtleTalk-Face, a pseudo-labeled FLAME dataset with roughly 3,905 identities and 73.83 hours of data. On the test split, the full method reports FDD 4.46 vs. 9.74 for DiffPoseTalk, HDD 6.61 vs. 21.71, and LVE 11.96 vs. 11.98 for ARTalk, leading to the central claim of substantially improved weakly correlated facial dynamics with preserved lip synchronization.
Significance. If the quantitative claims hold, the paper would make a meaningful advance in a known hard problem: modeling eyebrows, blinks, and head motion that are only weakly correlated with speech. The two-stage predict-and-refine design is a sensible way to protect lip synchronization from generative stochasticity, and the multi-condition control space (regional intensity plus continuous VA) is more interpretable than the style/identity controls used in most prior work. The ablation study (Table 5) is internally consistent, and the user study (Table 4) provides independent perceptual evidence that at least the naturalness gains are not purely an artifact of the proposed metric. The dataset, with 3,905 identities and strict identity-disjoint splits, is a potentially useful community resource, although it is not yet released at submission time. The main risk is that the headline FDD/HDD numbers are defined in a missing supplementary and appear to be optimized directly by the training losses, while the pseudo-label pipeline is shared between training and evaluation.
major comments (3)
- [§4.3, Table 3, Eq. (9), Sec. 3.2] The headline claim in Sec. 4.3 of 'remarkably lower FDD and HDD compared to all baselines' cannot be verified from the manuscript because the FDD/HDD formulas are deferred to a supplementary that is not included. The concern is substantive, not merely formal: Eq. (9) defines Lstd on R3 = {brows, eyes, head}, which are exactly the regions that FDD and HDD are said to summarize, and Eq. (7) defines Lvel on the same regions. Moreover, the regional intensity controls in Sec. 3.2 are themselves temporal standard deviations of FLAME vertex/pose trajectories, so the model is explicitly conditioned on the statistic the unpublished metrics are likely to measure. If FDD/HDD are standard-deviation or velocity deviations on those regions, the full model is directly trained to minimize the evaluation metric while the baselines are not; the large reported gaps (FDD 4.46 vs. 9.74, HDD 6.61 vs. 21.71) would then largely reflect objective overlap rather than independent dynamics quality. The authors should supply the exact metric definitions and report a retest with Lstd (and the relevant parts of Lvel) removed, or an evaluation with an independently defined dynamics metric.
- [§4.1, §4.3] Training labels and evaluation ground truth are both produced by the same TEASER pseudo-labeling pipeline, and the paper offers no validation of TEASER's upper-face accuracy on the filtered SubtleTalk-Face videos. A systematic reconstruction bias—for example, attenuated brow or eyelid range, or head-pose leakage into expression—would be learned by the Lstd and Lvel losses and then rewarded by FDD/HDD on the same pseudo-label test set. The user study (Table 4) and qualitative renderings support naturalness but do not pin down the specific FDD/HDD magnitudes. The authors should validate TEASER's upper-face fidelity on this dataset (e.g., against independent 2D landmark measurements or a held-out motion-capture subset) or report the headline metrics against ground truth that does not pass through the same fitter.
- [§4.6, Table 5] The ablation comparison is reported as single point estimates without any measure of variability, yet the paper's secondary claim of 'comparable or slightly better LVE' draws conclusions from very small differences (11.96 for the full model vs. 11.98 for ARTalk in Table 3). Without standard deviations, confidence intervals, or a significance test over seeds, the LVE parity claim is not statistically supported. The authors should report variability across training runs or a paired test on the test clips, at least for the full model versus the closest baseline on each metric.
minor comments (5)
- [Eq. (1)] The sentence introducing the Fourier feature encoding ends without a period, and the displayed formula has formatting issues (the brace after FE is malformed in the text). Please fix the punctuation and typesetting.
- [Figure 2 and Sec. 3.3] Figure 2 labels the second-stage module as 'Residual Diffusion Module (RDM)', while the text consistently uses 'Residual Flow Matching (RFM)'. This naming inconsistency should be corrected.
- [Abstract and Table 1] The abstract reports 'about 3,900 identities and 74 hours', while Table 1 gives 3,905 identities and 73.83 hours. Please align the numbers.
- [§4.3] The sentence 'Detailed computation formulas for these metrics are provided in the supplementary material' flags missing support, since the supplementary is not included in the arXiv submission. The formulas should be in the main text or in an available supplement for review.
- [§4.2] The implementation details state that DMP and RFM are trained on four V100 GPUs with specific wall-clock times, but no seeds or total budget for the user study are reported; a sentence on participant recruitment and the number of trials per participant would improve the reproducibility of Table 4.
Circularity Check
FDD/HDD advantage is partially self-definitional: Eq. (9) trains on temporal-standard-deviation mismatch for brows/eyes/head, the same quantity the deferred FDD/HDD 'dynamics deviation' metrics appear to score.
-
self definitional
[Sec. 4.3 (Quantitative Evaluation) vs. Eq. (9) in Sec. 3.3 (Residual Flow Matching)]
"upper face dynamics deviation (FDD) for upper-face motion dynamics, and head dynamics deviation (HDD) for head-motion dynamics. Detailed computation formulas for these metrics are provided in the supplementary material. ... we explicitly regularize the reconstructed motion `M^_c = M_c^prior + M^_c^res` using velocity loss Lvel, smoothness loss Lsmooth, and standard-deviation loss Lstd. ... Lstd = ∑_{r∈R3} λ_std^r ||std_t(^S_r) − std_t(S_r)||_2^2, where R1 = {face, lips, brows, eyes, head}, R2 = {face, head}, and R3 = {brows, eyes, head}."
The paper's headline metrics are named 'upper face dynamics deviation' and 'head dynamics deviation', and the only loss in the main text that directly penalizes mismatch of window-level temporal standard deviations is Eq. (9), whose region set R3 is exactly {brows, eyes, head}. If the deferred FDD/HDD formulas compute standard-deviation differences of the same FLAME vertex/head-pose trajectories, then the full model is trained on the reported test metric while the baselines are not. The large FDD/HDD gaps in Table 3 would then reflect direct optimization of the eval statistic rather than an independently measured dynamics advantage.
full rationale
The clearest circularity concern is objective overlap: Eq. (9) is a standard-deviation matching loss on brows, eyes, and head, which are the regions summarized by the FDD/HDD 'dynamics deviation' metrics reported in Sec. 4.3. Because the metric formulas are deferred to an absent supplementary, the reduction is not fully checkable in the main text, but the named quantity and the trained quantity are aligned closely enough to count as partial circularity. A separate validity concern, noted but not counted as an equation-level circular step, is that both training labels and test ground truth come from the same TEASER pseudo-labeling pipeline (Sec. 4.1) with no independent upper-face validation, so shared reconstruction bias could also inflate the reported improvements. The paper does provide some independent evidence, notably the user study (Table 4) and lip-sync LVE results, and there is no load-bearing self-citation or imported uniqueness theorem. On balance, the central quantitative claim is partly forced by construction, but not wholly equivalent to its inputs.
Assumptions & free parameters
free parameters (5)
- Regional intensity normalization statistics =
Not reported; computed from SubtleTalk-Face training set
- Loss weights lambda in Eqs. (2), (3), (7)-(9), (16) =
Not reported in main text
- Temporal window sizes T_p and T_c =
T_p=10, T_c=100
- Fourier feature bands K for VA encoding =
Not specified
- Random perturbation scale alpha in Eq. (11) =
Not specified
assumptions (6)
- domain assumption FLAME parameterization adequately represents facial motion, including upper-face and head dynamics.
- domain assumption TEASER monocular reconstruction yields high-fidelity FLAME pseudo-labels for the upper face.
- domain assumption EmotiEffLib frame-level VA estimates are reliable affect annotations.
- domain assumption WavLM layer split separates content from prosody as assumed in PAAE.
- ad hoc to paper The residual after the deterministic motion prior is a better generative target than raw motion.
- ad hoc to paper Temporal standard deviation of vertex and pose trajectories operationalizes perceived motion intensity.
Cite this review
Pith. "Pith review of SubtleTalk: Generating Controllable Weakly-correlated Facial Dynamics for 3D Talking Heads via Residual Flow Matching." pith.science (2026). https://pith.science/paper/ZXH7FUMD
@misc{pith2026260806408,
author = {Pith},
title = {Pith review of: SubtleTalk: Generating Controllable Weakly-correlated Facial Dynamics for 3D Talking Heads via Residual Flow Matching},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZXH7FUMD}},
note = {Machine review of arXiv:2608.06408}
}
read the original abstract
Audio-driven 3D facial animation aims to synthesize realistic and temporally coherent motions from speech. Despite notable progress in lip synchronization, weakly correlated dynamics, including eyebrow movements, eye blinks, and head motion, which are essential to photorealistic facial animation, remain difficult to model faithfully and often appear static or unnaturally repetitive. We attribute this limitation to three factors: (a) insufficient conditioning for weakly correlated dynamics; (b) the limited ability of deterministic regression to capture diverse motion patterns; (c) data bottlenecks from unreliable upper-face pseudo-labels and limited dataset diversity. To address these issues, we propose SubtleTalk, a framework for generating natural and controllable weakly correlated facial dynamics via multi-condition modeling and residual flow matching. First, to compensate for the limited guidance of speech alone, we introduce interpretable controls, including prosody, regional intensity, and Valence-Arousal signals, to explicitly capture the timing, magnitude, and affective variation of weakly correlated dynamics. Second, to overcome the limited expressiveness of deterministic regression, we build residual flow matching based on a stable speech-driven motion prior, allowing the model to capture stochastic deviations beyond deterministic prediction. Third, to alleviate the data bottleneck, we construct SubtleTalk-Face, a large-scale 3D facial animation dataset comprising about 3,900 identities and 74 hours of data, built via a simple and scalable pseudo-labeling pipeline and featuring improved upper-face tracking and frame-level VA annotations. Extensive experiments demonstrate that our method significantly improves the realism and diversity of weakly correlated facial dynamics while preserving accurate lip synchronization.
Figures
Reference graph
Works this paper leans on
-
[1]
Shivangi Aneja, Justus Thies, Angela Dai, and Matthias Nießner. 2024. Facetalk: Audio-driven motion diffusion for neural parametric head models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 21263– 21273
2024
-
[2]
Zenghao Chai, Haoxian Zhang, Jing Ren, Di Kang, Zhengzhuo Xu, Xuefei Zhe, Chun Yuan, and Linchao Bao. 2022. REALY: Rethinking the Evaluation of 3D Face Reconstruction. InProceedings of the European Conference on Computer Vision (ECCV)
2022
-
[3]
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al . 2022. Wavlm: Large-scale self-supervised pre-training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing16, 6 (2022), 1505–1518
2022
-
[4]
Xuangeng Chu, Nabarun Goswami, Ziteng Cui, Hanqin Wang, and Tatsuya Harada. 2025. Artalk: Speech-driven 3d head animation via autoregressive model. InProceedings of the SIGGRAPH Asia 2025 Conference Papers. 1–9
2025
-
[5]
Chaeyeon Chung, Ilya Fedorov, Michael Huang, Aleksey Karmanov, Dmitry Ko- robchenko, Roger Ribera, Yeongho Seol, et al. 2025. Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital Avatars.arXiv preprint arXiv:2508.16401 (2025)
arXiv 2025
-
[6]
Joon Son Chung and Andrew Zisserman. 2016. Out of time: automated lip sync in the wild. InAsian conference on computer vision. Springer, 251–263
2016
-
[7]
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J Black. 2019. Capture, learning, and synthesis of 3D speaking styles. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10101– 10111
work page 2019
-
[8]
Radek Daněček, Michael J Black, and Timo Bolkart. 2022. Emoca: Emotion driven monocular face capture and animation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 20311–20322
work page 2022
Show all 83 references
-
[9]
Radek Daněček, Kiran Chhatre, Shashank Tripathi, Yandong Wen, Michael Black, and Timo Bolkart. 2023. Emotional speech-driven animation with content- emotion disentanglement. InSIGGRAPH Asia 2023 Conference Papers. 1–13
2023
-
[10]
Radek Daněček, Carolin Schmitt, Senya Polikovsky, and Michael J Black. 2025. Supervising 3D Talking Head Avatars with Analysis-by-Audio-Synthesis.arXiv preprint arXiv:2504.13386(2025)
2025
-
[11]
Epic Games. n.d.. Live Link in Unreal Engine. Unreal Engine 5.7 Documentation. Retrieved March 01, 2026 from https://dev.epicgames.com/documentation/en- us/unreal-engine/live-link-in-unreal-engine
2026
-
[12]
Epic Games. n.d.. MetaHuman | Digital Humans - MetaHuman. Official MetaHu- man website, accessed March 01, 2026. https://www.metahuman.com/en-US
2026
-
[13]
Xiangyu Fan, Jiaqi Li, Zhiqian Lin, Weiye Xiao, and Lei Yang. 2024. Unitalker: Scaling up audio-driven 3d facial animation through a unified model. InEuropean Conference on Computer Vision. Springer, 204–221
2024
-
[14]
Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura. 2022. Faceformer: Speech-driven 3d facial animation with transformers. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 18770– 18780
2022
-
[15]
Gabriele Fanelli, Juergen Gall, Harald Romsdorfer, Thibaut Weise, and Luc Van Gool. 2010. A 3-d audio-visual corpus of affective communication.IEEE Transactions on Multimedia12, 6 (2010), 591–598
2010
-
[16]
2026.photometric_optimization: Photometric optimization code for creating the FLAME texture space and other applications
Haiwen Feng. 2026.photometric_optimization: Photometric optimization code for creating the FLAME texture space and other applications. Retrieved March 1, 2026 from https://github.com/HavenFeng/photometric_optimization
2026
-
[17]
Panagiotis P Filntisis, George Retsinas, Foivos Paraperas-Papantoniou, Athana- sios Katsamanis, Anastasios Roussos, and Petros Maragos. 2022. Visual speech- aware perceptual 3d facial expression reconstruction from videos.arXiv preprint arXiv:2207.11094(2022)
2022 arXiv
-
[18]
Hans Peter Graf, Eric Cosatto, Volker Strom, and Fu Jie Huang. 2002. Visual prosody: Facial movements accompanying speech. InProceedings of fifth IEEE international conference on automatic face gesture recognition. IEEE, 396–401
2002
-
[19]
Vincent Tao Hu, Wenzhe Yin, Pingchuan Ma, Yunlu Chen, Basura Fernando, Yuki M Asano, Efstratios Gavves, Pascal Mettes, Bjorn Ommer, and Cees GM Snoek. 2023. Motion flow matching for human motion synthesis and editing. arXiv preprint arXiv:2312.08895(2023)
2023 arXiv
-
[20]
Bin Ji, Ye Pan, Zhimeng Liu, Shuai Tan, Xiaogang Jin, and Xiaokang Yang. 2025. Pomp: Physics-consistent motion generative model through phase manifolds. In Proceedings of the Computer Vision and Pattern Recognition Conference. 22690– 22701
2025
-
[21]
Bin Ji, Ye Pan, Zhimeng Liu, Shuai Tan, and Xiaokang Yang. 2025. Sport: From zero- shot prompts to real-time motion generation.IEEE Transactions on Visualization and Computer Graphics(2025)
2025
-
[22]
Tero Karras, Samuli Laine, and Timo Aila. 2019. A style-based generator ar- chitecture for generative adversarial networks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4401–4410
2019
-
[23]
Hyung Kyu Kim, Sangmin Lee, and Hak Gu Kim. 2025. MemoryTalker: Per- sonalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 11241– 11251
2025
-
[24]
Jisoo Kim, Jungbin Cho, Joonho Park, Soonmin Hwang, Da Eun Kim, Geon Kim, and Youngjae Yu. 2025. Deeptalk: Dynamic emotion embedding for probabilistic speech-driven 3d face animation. InProceedings of the AAAI conference on artificial intelligence, Vol. 39. 4275–4283
2025
-
[25]
Jeesun Kim, Erin Cvejic, and Chris Davis. 2014. Tracking eyebrows and head gestures associated with spoken prosody.Speech Communication57 (2014), 317–330
2014
-
[26]
John P Lewis, Ken Anjyo, Taehyun Rhee, Mengjie Zhang, Frederic H Pighin, and Zhigang Deng. 2014. Practice and theory of blendshape facial models.Euro- graphics (State of the Art Reports)1, 8 (2014), 2
2014
-
[27]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. InInternational conference on machine learning. PMLR, 19730–19742
2023
-
[28]
Tianye Li, Timo Bolkart, Michael. J. Black, Hao Li, and Javier Romero. 2017. Learning a model of facial shape and expression from 4D scans.ACM Transactions on Graphics, (Proc. SIGGRAPH Asia)36, 6 (2017), 194:1–194:17. https://doi.org/10. 1145/3130800.3130813
2017
-
[29]
Zonglin Li, Xiaoqian Lv, Qinglin Liu, Quanling Meng, Xin Sun, and Shengping Zhang. 2025. Prosodytalker: 3d visual speech animation via prosody decompo- sition. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 5110–5118
2025
-
[30]
Guan-Ting Lin, Chi-Luen Feng, Wei-Ping Huang, Yuan Tseng, Tzu-Han Lin, Chen- An Li, Hung-yi Lee, and Nigel G Ward. 2023. On the utility of self-supervised models for prosody-related tasks. In2022 IEEE Spoken Language Technology Workshop (SLT). IEEE, 1104–1111
2023
-
[31]
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le
-
[32]
Chang Liu, Qunfen Lin, Zijiao Zeng, and Ye Pan. 2024. Emoface: Audio-driven emotional 3d face animation. In2024 IEEE Conference Virtual Reality and 3D User Interfaces (VR). IEEE, 387–397
2024
-
[33]
Chang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja, and Xiaokang Yang. 2025. Medtalk: Multimodal controlled 3d facial animation with dynamic emotions by disentangled embedding. InProceedings of the 33rd ACM International Conference on Multimedia. 7538–7547
2025
-
[34]
Yunfei Liu, Lei Zhu, Lijian Lin, Ye Zhu, Ailing Zhang, and Yu Li. 2025. Teaser: Token enhanced spatial modeling for expressions reconstruction.arXiv preprint arXiv:2502.10982(2025)
2025 arXiv
-
[35]
Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101(2017)
2017 arXiv
-
[36]
Xin Lu, Chuanqing Zhuang, Chenxi Jin, Zhengda Lu, Yiqun Wang, Wu Liu, and Jun Xiao. 2025. LSF-Animation: Label-Free Speech-Driven Facial Animation via Implicit Feature Representation. InProceedings of the SIGGRAPH Asia 2025 Conference Papers. 1–12
2025
-
[37]
Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Guang Yong, Juhyun Lee, et al. 2019. Mediapipe: A framework for building perception pipelines.arXiv preprint arXiv:1906.08172(2019)
2019 arXiv
-
[38]
Ke Ma, Yizhou Fang, Jean-Baptiste Weibel, Shuai Tan, Xinggang Wang, Yang Xiao, Yi Fang, and Tian Xia. 2026. Phys-liquid: a physics-informed dataset for estimating 3d geometry and volume of transparent deformable liquids. In Proceedings of the AAAI Conference on Artificial Inte...
2026
-
[39]
Ziyang Ma, Zhisheng Zheng, Jiaxin Ye, Jinchao Li, Zhifu Gao, Shiliang Zhang, and Xie Chen. 2024. emotion2vec: Self-supervised pre-training for speech emotion representation. InFindings of the Association for Computational Linguistics: ACL
2024
-
[40]
Zhiyuan Ma, Xiangyu Zhu, Chen Qian, Shukai Chen, Li Gao, Guojun Qi, Zhaoxi- ang Zhang, and Zhen Lei. 2025. Diffspeaker: Speech-driven 3d facial animation with diffusion transformer. In2025 IEEE International Joint Conference on Biomet- rics (IJCB). IEEE, 1–10
2025
-
[41]
Leena Mary and Bayya Yegnanarayana. 2008. Extraction and representation of prosodic features for language and speaker recognition.Speech communication 50, 10 (2008), 782–796
2008
-
[42]
William Peebles and Saining Xie. 2022. Scalable Diffusion Models with Trans- formers.arXiv preprint arXiv:2212.09748(2022)
2022 arXiv
-
[43]
Ziqiao Peng, Yihao Luo, Yue Shi, Hao Xu, Xiangyu Zhu, Hongyan Liu, Jun He, and Zhaoxin Fan. 2023. Selftalk: A self-supervised commutative training diagram to comprehend 3d talking faces. InProceedings of the 31st ACM International Conference on Multimedia. 5292–5301
2023
-
[44]
Ziqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu, Xiangyu Zhu, Jun He, Hongyan Liu, and Zhaoxin Fan. 2023. Emotalk: Speech-driven emotional disentanglement MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Chenyang Ding et al. for 3d face animation. InProceedings of the IEEE/CVF ...
2023
-
[45]
Alexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre, and Yaser Sheikh. 2021. Meshtalk: 3d face animation from speech using cross- modality disentanglement. InProceedings of the IEEE/CVF international conference on computer vision. 1173–1182
2021
-
[46]
James A Russell. 1980. A circumplex model of affect.Journal of personality and social psychology39, 6 (1980), 1161
1980
-
[47]
Andrey Savchenko. 2023. Facial Expression Recognition with Adaptive Frame Rate based on Multiple Testing Correction. InProceedings of the 40th Interna- tional Conference on Machine Learning (ICML) (Proceedings of Machine Learn- ing Research, Vol. 202), Andreas Krause, Emma Bru...
2023
-
[48]
Abraham Savitzky and Marcel JE Golay. 1964. Smoothing and differentiation of data by simplified least squares procedures.Analytical chemistry36, 8 (1964), 1627–1639
1964
-
[49]
Kang Shen, Haifeng Xia, Guangxing Geng, Guangyue Geng, Siyu Xia, and Zheng- ming Ding. 2024. Deitalk: Speech-driven 3d facial animation with dynamic emotional intensity modeling. InProceedings of the 32nd ACM international con- ference on multimedia. 10506–10514
2024
-
[50]
Stefan Stan, Kazi Injamamul Haque, and Zerrin Yumak. 2023. Facediffuser: Speech-driven 3d facial animation synthesis using diffusion. InProceedings of the 16th ACM SIGGRAPH Conference on Motion, Interaction and Games. 1–11
2023
-
[51]
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing568 (2024), 127063
2024
-
[52]
Yasheng Sun, Wenqing Chu, Hang Zhou, Kaisiyuan Wang, and Hideki Koike
-
[53]
Zhiyao Sun, Tian Lv, Sheng Ye, Matthieu Lin, Jenny Sheng, Yu-Hui Wen, Min- jing Yu, and Yong-jin Liu. 2024. Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models.ACM Transactions on Graphics (ToG)43, 4 (2024), 1–9
2024
-
[54]
Shuai Tan, Bill Gong, Bin Ji, and Ye Pan. 2025. Fixtalk: Taming identity leakage for high-quality talking head generation in extreme cases. InProceedings of the IEEE/CVF International Conference on Computer Vision. 24–36
2025
-
[55]
Shuai Tan, Biao Gong, Zhuoxin Liu, Yan Wang, Xi Chen, Yifan Feng, and Heng- shuang Zhao. 2025. Animate-x++: Universal character image animation with dynamic backgrounds.arXiv preprint arXiv:2508.09454(2025)
2025 arXiv
-
[56]
Shuai Tan, Biao Gong, Ke Ma, Yutong Feng, Qiyuan Zhang, Yan Wang, Yujun Shen, and Hengshuang Zhao. 2026. CoDance: An Unbind-Rebind Paradigm for Robust Multi-Subject Animation.arXiv preprint arXiv:2601.11096(2026)
2026
-
[57]
Shuai Tan, Biao Gong, Xiang Wang, Shiwei Zhang, Dandan Zheng, Ruobing Zheng, Kecheng Zheng, Jingdong Chen, and Ming Yang. 2024. Animate-x: Uni- versal character image animation with enhanced motion representation.arXiv preprint arXiv:2410.10306(2024)
2024 arXiv
-
[58]
Shuai Tan, Biao Gong, Yujie Wei, Shiwei Zhang, Zhuoxin Liu, Ke Ma, Yan Wang, Kecheng Zheng, Xing Zhu, Yujun Shen, et al. 2026. Synmotion: Semantic-visual adaptation for motion customized video generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...
2026
-
[59]
Shuai Tan and Bin Ji. 2025. EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis.arXiv preprint arXiv:2508.13442(2025)
2025 arXiv
-
[60]
Shuai Tan, Bin Ji, Mengxiao Bi, and Ye Pan. 2024. Edtalk: Efficient disentanglement for emotional talking head synthesis. InEuropean Conference on Computer Vision. Springer, 398–416
2024
-
[61]
Shuai Tan, Bin Ji, Yu Ding, and Ye Pan. 2024. Say anything with any style. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 5088–5096
2024
-
[62]
Shuai Tan, Bin Ji, and Ye Pan. 2023. Emmn: Emotional motion memory network for audio-driven emotional talking face generation. InProceedings of the IEEE/CVF International Conference on Computer Vision. 22146–22156
2023
-
[63]
Shuai Tan, Bin Ji, and Ye Pan. 2024. Flowvqtalker: High-quality emotional talking face generation through normalizing flow and quantization. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 26317–26327
2024
-
[64]
Shuai Tan, Bin Ji, and Ye Pan. 2024. Style2talker: High-resolution talking head generation with emotion style and art style. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 5079–5087
2024
-
[65]
Matthew Tancik, Pratul Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan Barron, and Ren Ng
-
[66]
Aaron Van Den Oord, Oriol Vinyals, et al. 2017. Neural discrete representation learning.Advances in neural information processing systems30 (2017)
2017
-
[67]
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy. 2020. MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation. InECCV
2020
-
[68]
Mengchao Wang, Qiang Wang, Fan Jiang, Yaqi Fan, Yunpeng Zhang, Yonggang Qi, Kun Zhao, and Mu Xu. 2025. Fantasytalking: Realistic talking portrait generation via coherent motion synthesis. InProceedings of the 33rd ACM International Conference on Multimedia. 9891–9900
2025
-
[69]
Sichun Wu, Kazi Injamamul Haque, and Zerrin Yumak. 2024. Probtalk3d: Non- deterministic emotion controllable speech-driven 3d facial animation synthesis using vq-vae. InProceedings of the 17th ACM SIGGRAPH conference on motion, interaction, and games. 1–12
2024
-
[70]
Sijing Wu, Yunhao Li, Yichao Yan, Huiyu Duan, Ziwei Liu, and Guangtao Zhai
-
[71]
Cheng-hsin Wuu, Ningyuan Zheng, Scott Ardisson, Rohan Bali, Danielle Belko, Eric Brockmeyer, Lucas Evans, Timothy Godisart, Hyowon Ha, Xuhua Huang, et al . 2022. Multiface: A dataset for neural face rendering.arXiv preprint arXiv:2207.11243(2022)
2022 arXiv
-
[72]
Liangbin Xie, Xintao Wang, Honglun Zhang, Chao Dong, and Ying Shan. 2022. VFHQ: A High-Quality Dataset and Benchmark for Video Face Super-Resolution. InThe IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
2022
-
[73]
Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang, and Tien- Tsin Wong. 2023. Codetalker: Speech-driven 3d facial animation with discrete motion prior. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12780–12790
2023
-
[74]
InPro- ceedings of the 32nd ACM International Conference on Multimedia
Mmhead: Towards fine-grained multi-modal 3d facial animation. InPro- ceedings of the 32nd ACM International Conference on Multimedia. 7966–7975
-
[75]
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan. 2021. Flow-Guided One- Shot Talking Face Generation With a High-Resolution Audio-Visual Dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3661–3670
2021
-
[76]
Qingcheng Zhao, Pengyu Long, Qixuan Zhang, Dafei Qin, Han Liang, Longwen Zhang, Yingliang Zhang, Jingyi Yu, and Lan Xu. 2024. Media2face: Co-speech facial animation generation with multi-modality guidance. InACM SIGGRAPH 2024 conference papers. 1–13
2024
-
[77]
Yicheng Zhong, Huawei Wei, Peiji Yang, and Zhisheng Wang. 2024. Expclip: Bridging text and facial expressions via semantic alignment. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 7614–7622
2024
-
[78]
Karren D Yang, Anurag Ranjan, Jen-Hao Rick Chang, Raviteja Vemulapalli, and Oncel Tuzel. 2024. Probabilistic speech-driven 3d facial motion synthesis: New benchmarks methods and applications. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. ...
2024
-
[79]
Wojciech Zielonka, Timo Bolkart, and Justus Thies. 2022. Towards metrical re- construction of human faces. InEuropean conference on computer vision. Springer, 250–269
2022
-
[82]
Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy. 2022. CelebV-HQ: A Large-Scale Video Facial Attributes Dataset. InECCV
2022
-
[2020]
Fourier features let networks learn high frequency functions in low di- mensional domains.Advances in neural information processing systems33 (2020), 7537–7547
2020
-
[2022]
Flow matching for generative modeling.arXiv preprint arXiv:2210.02747 (2022)
2022 arXiv
-
[2024]
Avi-talking: Learning audio-visual instructions for expressive 3d talking face generation.IEEE Access12 (2024), 57288–57301
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.