REVIEW 6 major objections 8 minor 42 references
GRAPE: Graduated Routing for Articulated Portrait mesh Estimation
T0 review · 6 major / 8 minor · reviewed 2026-07-30 · grok-4.5
Pith's one-line read Portrait meshes recover better when pose, jaw, and expression are estimated in anatomical order rather than all at once.
desk verdict Solid portrait-mesh engineering with a real torso-rooted model; the jaw-disentanglement headline is softer than the geometric gains because JSR is partly trained via Lnc. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Progressive Anatomical Alignment with a Graduated-Mask Router: anatomy-aware experts predict global shape/camera, then skeletal pose (including jaw), then expression, while Bernoulli masks randomly drop deeper supervision so expression cannot absorb jaw errors during training.
What would settle it
On talking clips with stable head pose, if predicted jaw pitch stayed weakly correlated with normalized nose-to-chin distance while expression still drove mouth opening—matching or worse than strong face regressors—and head–shoulder joint error did not drop when the torso anchor was present, the central disentanglement claim would fail.
Extended reading notes
Core claim
GRAPE establishes that monocular articulated portrait mesh estimation improves when representation, routing, and supervision all follow the torso→neck→head→jaw→expression hierarchy: a torso-rooted hybrid parametric model plus progressive experts gated by a Graduated-Mask Router disentangles camera and head pose and forces jaw articulation before expression, outperforming flat face- and body-centric regressors on alignment, kinematics, and downstream animation.
Load-bearing premise
The offline staged pseudo-labels used as training targets are assumed to already separate jaw from expression and to place shoulders well enough that the network learns real anatomy rather than the label pipeline’s staging habits.
Editorial extensions
If this is right
- Portrait avatars can be driven with a real neck joint instead of centrifuge-like head spins around the face center.
- Talking-head trainers get parameter tracks where jaw, not blendshapes, carries skeletal mouth opening, stabilizing audio-driven animation learning.
- Animatable 3D avatar pipelines gain continuous head–neck–shoulder surfaces instead of floating heads over mismatched torsos.
- Face and body reconstruction lines can share one hybrid mesh rather than trading facial fidelity against kinematic anchors.
Reading between the lines
- Any monocular body part with a rigid-joint vs soft-deformation tradeoff (hands, spine) may benefit from the same coarse-to-fine loss masking pattern.
- If pseudo-label staging is the real source of jaw purity, end-to-end methods without staged labels may need an equivalent hard architectural prior to match JSR-style metrics.
- Combining the explicit kinematic scaffold with hair and clothing representations is the natural next stress test of whether the torso anchor still helps when the silhouette is no longer skin-tight.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces GRAPE for monocular portrait mesh estimation. It constructs a Portrait Parametric Model (PPM) by inserting a FLAME2023 head into an SMPL-X torso in canonical space, with a torso-rooted spine–neck–head–jaw kinematic chain. A Progressive Anatomical Alignment network uses a frozen Sapiens-2 encoder, three anatomically specialized experts, and a Graduated-Mask Router that stochastically removes supervision from finer experts during training. Training combines offline staged pseudo-labels, feature distillation, landmarks, a projected nose–chin distance loss, foreground overflow constraints, and relative torso geometry. The paper reports improvements in image landmark NME, HDTF jaw–skeletal consistency, Nersemble/NoW/Stirling reconstruction, structural and loss ablations, and two downstream avatar applications.
Significance. If the central claims are substantiated, the work addresses a real gap between face-centric models with weak head–torso kinematics and body-centric models with limited facial fidelity. The explicit PPM construction and canonical injection are useful contributions, and the paper provides unusually broad empirical coverage: four image benchmarks, talking-video analysis, three 3D protocols, structural and training-objective ablations, and downstream tests with DiffPoseTalk and RGBAvatar. These are meaningful strengths. The standard NoW/Stirling gains are modest but consistent, while the larger Nersemble and JSR gains currently depend on insufficiently specified or partially self-referential evaluation. The work is potentially valuable for animatable portrait reconstruction, but its disentanglement and downstream conclusions need stronger independent evidence.
major comments (6)
- [§3.2.3 Eq. (14); §C.2 Eqs. (22)–(23); App. B] JSR is not independent of the training signal. Eq. (14) directly fits the projected nose-tip–chin-tip distance, while Eqs. (22)–(23) evaluate the Pearson correlation of jaw pitch with a per-clip normalization of that same detected distance. Moreover, Lord regresses pseudo-labels from App. B Step 2.2, where mouth-related expression is zeroed and mouth opening is assigned to the jaw before expression fitting. Table 2 can therefore measure consistency with the imposed proxy/staging rather than anatomical disentanglement. The stress-test’s row attribution is imprecise—V1.2 adds Lfd and V1.3 adds Lnc—but the overlap remains. Add validation against independent 3D/4D jaw motion and a no-Lnc/no-jaw-staged-label control, or narrow the claim.
- [Tables 4–5; Appendix A] The router ablations appear internally inconsistent. Table 4 reports JSR 0.411 when stochastic graduated masking is disabled, whereas Table 5 V1.5 reports 0.701 with all losses and alpha_act=1.0, which Appendix A describes as the same always-active-mask setting. Table 5 also has alpha_act=0.25 scoring 0.801, slightly above the selected alpha_act=0.50 at 0.799. Please reconcile the configurations, define alpha_act unambiguously, state whether JSR is pooled or averaged over the 20 HDTF clips, and report per-clip confidence intervals and multiple training seeds. As written, the magnitude and even the source of the router effect are uncertain.
- [Algorithm 1; §3.1.2; Appendix A.1] The PPM forward pass appears to apply jaw articulation twice. Algorithm 1 first evaluates FLAME with theta_jaw, while the final unified LBS is driven by a pose vector whose hierarchy again includes the jaw; Appendix A.1 repeats this description. Unless the implementation removes or reparameterizes jaw rotation in one of these stages, this double transformation would undermine the claimed animatable semantics. Eq. (1) also introduces a head scale s, but the canonical-injection algorithm only specifies eye-center translation. Please provide the exact forward equations, including which transform is baked into the template, which is applied by unified LBS, and where scale/rotation alignment occurs.
- [§3.3, §4.1, Tables 1 and 3; §C.3] The evaluation protocols are not sufficiently specified for the main reconstruction claims. §3.3 says the model is finetuned on LS3DW, CelebA, LaPa, and LFW, while §4.1 evaluates on test sets of the same four benchmarks without giving exact splits or exclusion rules; any image overlap would inflate Table 1. For Table 3, Nersemble compares method-specific mesh topologies and may mix FLAME2020 and FLAME2023 after GRAPE was pretrained on NersembleV2. Define the ground-truth registration, vertex correspondence, scale and pose alignment, subject/scene split, and baseline domain treatment, and release the evaluation scripts and IDs. This matters because the NoW/Stirling margins are small, while the opaque Nersemble margin is large.
- [Table 4; §4.3; §C.3] The MPJPE result is load-bearing for the torso-anchor claim, but its protocol is missing. The metric definition says only “head–shoulder joint error in millimeters”; no dataset, joint set, ground-truth source, coordinate frame, alignment procedure, or uncertainty estimate is given. If the reference joints come from Sapiens-2 or the same pseudo-label pipeline used for training, the 89.1-to-69.4 improvement is not independent validation of head–torso kinematics. Please specify the protocol and preferably evaluate against multi-view or motion-capture joint trajectories.
- [§4.5, Figure 8, Table 6] The downstream evidence does not yet support the broad claim that GRAPE benefits talking-head generation. The DiffPoseTalk result consists of training-loss curves, but the compared parameterizations and reconstruction targets differ, so lower or smoother loss is not a direct measure of generated-video quality. RGBAvatar provides held-out metrics, but the improvements are small (PSNR 21.08 to 21.69; PSNR-b 17.90 to 18.64) on only ten clips with no variance. Add held-out perceptual, identity, temporal, and lip-motion metrics under a common output representation over more sequences, or temper the abstract’s downstream claim.
minor comments (8)
- [Table 5] The checkmark columns make it difficult to determine which component each V1.x row adds. Replace some checkmarks with explicit component names and align the row descriptions with Table 4.
- [§3.2.2 and Appendix A] Use p and alpha_act consistently. State explicitly whether the value is the probability of activating an expert’s supervision or of dropping it, and give the test-time setting.
- [Eq. (9)] “Omit the weight of each loss term without loss of generality” is not correct: the relative weights determine the graduated trade-offs. List the final weights or provide a sensitivity analysis.
- [Eq. (22)] Define d_t before using it in the normalized expression, give epsilon, and state whether mu_d and sigma_d are computed per subject, per clip, or over all HDTF frames.
- [Figures 3–6] Several equations, labels, and highlighted regions are difficult to read at print size. Increase font sizes and make the deep-red/deep-blue encodings distinguishable in grayscale.
- [Figures 5 and 11] Fix typographical errors such as “SPECTR” in Figure 5 and “PMMFitting” in Figure 11.
- [§4.1 and §C.4] Provide the random seeds and exact IDs for the 20 HDTF evaluation videos and the 10 RGBAvatar clips to make the small-sample comparisons auditable.
- [Algorithm 1] Define J/SMPL_T, W_unified, pi, and all parameter blocks in the algorithm itself, and repair the formatting so that the variable names and comments are legible.
Circularity Check
Disentanglement claim partly echoes its own supervision: Lnc trains on the nose–chin signal that JSR scores, and Lord regresses jaw/expression targets staged by the same graduated fitting pipeline.
-
fitted input called prediction
[§3.2.3 Eq. (14) Lnc; §C.2 Eqs. (22)–(23) JCR/JSR; Table 2]
"Lnc = |d(ûnose, ûchin) − d(u^gt_nose, u^gt_chin)| ... For jaw disentanglement, we use the nose-tip-to-chin-tip distance as a weak skeletal opening cue... JSR = ρ({θ^t_j−p}, {JCR_t}) where θ^t_j−p is predicted jaw pitch and JCR is a normalized nose–chin distance reference."
Training explicitly penalizes mismatch on projected nose–chin distance (Lnc), the same geometric signal that defines JCR. JSR then scores Pearson correlation of predicted jaw pitch with that signal. In PPM/FLAME, jaw pitch is the dominant skeletal DOF moving chin relative to nose, so supervising Lnc (plus Lord on θ jaw) statistically elevates JSR. The Table 2 “disentanglement” win is therefore partly training-on-the-metric, not an independent kinematic discovery; baselines were not trained with Lnc.
-
fitted input called prediction
[Appendix B Phase 2 (Steps 2.1–2.3); Eq. (10) Lord; §3.2 Graduated-Mask Router]
"We fit FLAME2023 in three separate steps to reduce jaw–expression leakage... Step 2.2: Jaw. Fix βh and set mouth-related expression coefficients to zero. Optimize jaw rotation θjaw so mouth opening is explained by the jaw joint rather than expression. Step 2.3: Expression. Fix βh and θjaw, then optimize ψ... Lord = LG(π,βh,βt) + mM LM(θspine,θneck,θhead,θjaw) + mL LL(ψ,θeye)."
Pseudo-labels impose jaw-before-expression by construction (mouth blendshapes zeroed while fitting jaw). PAA is then L2-trained via Lord to reproduce those staged parameters, with the Graduated-Mask Router enforcing the same coarse-to-fine order. Reported jaw–expression disentanglement therefore partly reflects fidelity to labels that already encode the desired factorization, not a free test that the architecture alone discovered anatomical separation.
full rationale
GRAPE is an empirical monocular regression paper, not a first-principles derivation; geometric claims (NME on LS3DW/CelebA/LaPa/LFW, MVE/LVE, NoW/Stirling) compare against external baselines under standard protocols and are not circular. The load-bearing jaw–expression disentanglement story is weaker. Offline Graduated-Refinement constructs pseudo-labels by forcing shape→jaw (with mouth expression zeroed)→expression, then Lord L2-supervises the network to those staged parameters while Lnc directly matches projected nose–chin distance to the same 2D cue underlying JCR/JSR. High JSR is therefore partly statistically aligned with the training objective and label pipeline rather than an independent test of anatomical factorization; Table 5 shows large JSR already with losses before the Graduated-Mask Router. This is partial evaluation/self-consistency circularity on the distinctive disentanglement claim, not a definitional collapse of the whole method. Score 4: central geometry still has independent content; the flagship kinematic metric does not.
Assumptions & free parameters
free parameters (5)
- Graduated-Mask Bernoulli p (α_act) =
0.5 (finetune); 1.0 (Nersemble pretrain)
- Loss weights for Lord, Lfd, Llmk, Lnc, Lom, Lrg, regs =
not numerically reported
- Neck-ring blend coefficients α_i / soft-stitch width =
width 0.015 m (soft-stitch); α_i schedule template-defined
- Distillation temperature τ and projection head F =
τ not numerically specified in main text
- Training schedule (4M iters, bs 64, Adam, 256 crop, adapter/decoder depth) =
4M iters, batch 64, 4×A100, decoder 4 layers D=768
assumptions (7)
- domain assumption Linear blend skinning on a parametric mesh (FLAME head + SMPL-X torso) is an adequate generative model of portrait geometry for monocular recovery and re-animation.
- domain assumption Without a torso anchor, monocular head global rotation cannot be reliably split into neck articulation vs camera extrinsics (floating-head ill-posedness).
- ad hoc to paper Human portrait motion should be estimated in anatomical order: global shape/camera → skeletal articulation → non-rigid expression.
- ad hoc to paper Offline staged pseudo-labels from TEASER+ProHMR initialization and graduated FLAME/PPM fitting are good-enough supervisory targets for in-the-wild images.
- ad hoc to paper Under relatively stable head pose, normalized projected nose-tip–chin-tip distance is a weak skeletal reference for jaw opening less explainable by lip blendshapes than mouth landmarks.
- domain assumption Frozen Sapiens-2 features and Pixel3DMM token distributions provide transferable portrait geometry priors.
- standard math Standard supervised learning / cross-attention / L2 and KL losses behave as usual; no new learning theory claimed.
invented entities (4)
-
Portrait Parametric Model (PPM)
-
Progressive Anatomical Alignment / HKD-Exp Network with Graduated-Mask Router
-
Jaw-chain ratio (JCR) and jaw–skeletal ratio score (JSR)
-
Graduated-Refinement Pipeline for PPM pseudo-labels
Cite this review
Pith. "Pith review of GRAPE: Graduated Routing for Articulated Portrait mesh Estimation." pith.science (2026). https://pith.science/paper/UIHYCMLW
@misc{pith2026260723657,
author = {Pith},
title = {Pith review of: GRAPE: Graduated Routing for Articulated Portrait mesh Estimation},
year = {2026},
howpublished = {\url{https://pith.science/paper/UIHYCMLW}},
note = {Machine review of arXiv:2607.23657}
}
read the original abstract
Articulated portrait mesh estimation is fundamental to 3D understanding, avatar generation, and immersive interaction. Existing approaches primarily rely on 3D Morphable Models (3DMMs). However, face-centric models suffer from the "floating head" assumption, conflating head pose with global rotation due to the lack of neck kinematics. Conversely, body-centric models lack high-fidelity facial expression capabilities. Furthermore, current methods struggle to disentangle jaw articulation from expression blendshapes, often over-relying on expressions for mouth opening. These limitations make monocular portrait recovery difficult across representation, supervision, and anatomical parameter estimation. To address these limitations, we introduce GRAPE(Graduated Routing for Articulated Portrait mesh Estimation). We build a Portrait Parametric Model (PPM) with an explicit torso-to-head kinematic chain and a canonical injection step to merge FLAME and the SMPL-X torso. We propose a Progressive Anatomical Alignment (PAA) network, which is composed of a pretrained portrait encoder, a Graduated-Mask Router, and coarse-to-fine experts that follow the portrait anatomical prior. We then train this network with multi-source supervision that combines sparse anatomical keypoints, feature distillation, foreground mask constraints, and relative geometry constraints. Experiments show that GRAPE improves portrait mesh recovery quality, pose alignment, and jaw--expression disentanglement over prior methods. We also demonstrate that our method can benefit the downstream tasks of audio-driven talking-head generation and 3D portrait generation.
Reference graph
Works this paper leans on
-
[1]
Volker Blanz and Thomas Vetter. 1999. A morphable model for the synthesis of 3D faces. InProceedings of the 26th annual conference on Computer graphics and interactive techniques (SIGGRAPH). 187–194
1999
-
[2]
Adrian Bulat and Georgios Tzimiropoulos. 2017. How Far Are We From Solving the 2D & 3D Face Alignment Problem? (And a Dataset of 230,000 3D Facial Landmarks). InProceedings of the IEEE International Conference on Computer Vision (ICCV). 1021–1030. doi:10.1109/ICCV.2017.116
-
[3]
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J Black. 2019. Capture, learning, and synthesis of 3D speaking styles. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10101–10111
2019
-
[4]
Black, and Timo Bolkart
Radek Daněček, Michael J. Black, and Timo Bolkart. 2022. EMOCA: Emotion Driven Monocular Face Capture and Animation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 20311–20322
2022
-
[5]
Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong. 2019. Accurate 3D face reconstruction with weakly-supervised learning: From single image to image set. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops. 0–0
2019
-
[6]
Yao Feng, Vasileios Choutas, Timo Bolkart, Dimitrios Tzionas, and Michael J Black. 2021. Collaborative regression of expressive bodies using moderation. InInternational Conference on 3D Vision (3DV). IEEE, 792–804
2021
-
[7]
Yao Feng, Haiwen Feng, Michael J Black, and Timo Bolkart. 2021. Learning an animatable detailed 3D face model from in-the-wild images.ACM Transactions on Graphics (TOG)40, 4 (2021), 1–13
2021
-
[8]
Zhen-Hua Feng, Patrik Huber, Josef Kittler, Peter Hancock, Xiao-Jun Wu, Qijun Zhao, Paul Koppen, and Matthias Rätsch. 2018. Evaluation of dense 3D reconstruction from 2D face images in the wild. In 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018). IEEE, 780–786
2018
Show all 42 references
-
[9]
Filntisis, George Retsinas, Foivos Paraperas-Papantoniou, Athanasios Katsamanis, Anas- tasios Roussos, and Petros Maragos
Panagiotis P. Filntisis, George Retsinas, Foivos Paraperas-Papantoniou, Athanasios Katsamanis, Anas- tasios Roussos, and Petros Maragos. 2023. SPECTRE: Visual Speech-Informed Perceptual 3D Facial Expression Reconstruction from Videos. InProceedings of the IEEE/CVF Conference o...
2023
-
[10]
Simon Giebenhain, Tobias Kirschstein, Martin Rünz, Lourdes Agapito, and Matthias Nießner. 2025. Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction.arXiv preprint arXiv:2505.00615(2025)
2025 arXiv
-
[11]
Google. 2026. GNM: Generative aNthropometric Model and Ecosystem.https://github.com/google/ GNM. GNM Head open-source release. 15
2026
-
[12]
Jianzhu Guo, Xiangyu Zhu, Yang Yang, Fan Yang, Zhen Lei, and Stan Z Li. 2020. Towards Fast, Accurate and Stable 3D Dense Face Alignment. InProceedings of the European Conference on Computer Vision (ECCV)
2020
-
[13]
Huang, Manu Ramesh, Tamara Berg, and Erik Learned-Miller
Gary B. Huang, Manu Ramesh, Tamara Berg, and Erik Learned-Miller. 2007.Labeled Faces in the Wild: A Database for Studying Face Recognition in Unconstrained Environments. Technical Report 07-49. University of Massachusetts, Amherst
2007
-
[14]
Rawal Khirodkar, He Wen, Julieta Martinez, Yuan Dong, Zhaoen Su, and Shunsuke Saito. 2026. Sapiens2. InInternational Conference on Learning Representations (ICLR). arXiv:2604.21681 [cs.CV]
2026 arXiv
-
[15]
Tobias Kirschstein, Shenhan Qian, Simon Giebenhain, Tim Walter, and Matthias Nießner. 2023. NeRSem- ble: Multi-View Radiance Field Reconstruction of Human Heads.ACM Transactions on Graphics42, 4 (2023), 161:1–161:14. doi:10.1145/3592455
2023 doi
-
[16]
Nikos Kolotouros, Georgios Pavlakos, Michael J Black, and Kostas Daniilidis. 2019. Learning to reconstruct 3D human pose and shape via model-fitting in the loop. InIEEE/CVF International Conference on Computer Vision (ICCV). 2252–2261
2019
-
[17]
Nikos Kolotouros, Georgios Pavlakos, Dinesh Jayaraman, and Kostas Daniilidis. 2021. Probabilistic Modeling for Human Mesh Recovery. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 11605–11614
2021
-
[18]
Linzhou Li, Yumeng Li, Yanlin Weng, Youyi Zheng, and Kun Zhou. 2025. RGBAvatar: Reduced Gaussian Blendshapes for Online Modeling of Head Avatars. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10747–10757
2025
-
[19]
Tianye Li, Timo Bolkart, Michael J Black, Hao Li, and Javier Romero. 2017. Learning a model of facial shape and expression from 4D scans.ACM Transactions on Graphics (TOG)36, 6 (2017), 194:1–194:17
2017
-
[20]
Tianye Li, Timo Bolkart, Michael J Black, Hao Li, and Javier Romero. 2023. FLAME: Learning a Model of Facial Shape and Expression from 4D Scans (2023 Release).https://flame.is.tue.mpg.de/
2023
-
[21]
Jing Lin, Ailing Zeng, Haoqian Wang, Lei Zhang, and Yu Li. 2023. One-stage 3d whole-body mesh recovery with component aware transformer. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 21159–21168
2023
-
[22]
Yinglu Liu, Hailin Shi, Hao Shen, Yue Si, Xiaobo Wang, and Tao Mei. 2020. A New Dataset and Boundary-Attention Semantic Segmentation for Face Parsing. InProceedings of the AAAI Conference on Artificial Intelligence. 11637–11644
2020
-
[23]
Yunfei Liu, Lei Zhu, Lijian Lin, Ye Zhu, Ailing Zhang, and Yu Li. 2025. TEASER: Token Enhanced Spatial Modeling for Expressions Reconstruction. InICLR
2025
-
[24]
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 2015. Deep Learning Face Attributes in the Wild. InProceedings of the IEEE International Conference on Computer Vision (ICCV). 3730–3738
2015
-
[25]
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A Osman, Dimitrios Tzionas, and Michael J Black. 2019. Expressive body capture: 3d hands, face, and body from a single image. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10975–10985
2019
-
[26]
Zesong Qiu, Yuwei Li, Dongming He, Qixuan Zhang, Longwen Zhang, Yinghao Zhang, Jingya Wang, Lan Xu, Xudong Wang, Yuyao Zhang, and Jingyi Yu. 2022. SCULPTOR: Skeleton-Consistent Face Creation Using a Learned Parametric Generator.ACM Transactions on Graphics (TOG)41, 6, Article ...
2022
-
[27]
Li, and Shan Liu
Yurui Ren, Ge Li, Yuanqi Chen, Thomas H. Li, and Shan Liu. 2021. PIRenderer: Controllable Portrait Image Generation via Semantic Neural Rendering. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 13759–13768
2021
-
[28]
Filntisis, Radek Danecek, Victoria F
George Retsinas, Panagiotis P. Filntisis, Radek Danecek, Victoria F. Abrevaya, Anastasios Roussos, Timo Bolkart, and Petros Maragos. 2024. 3D Facial Expressions through Analysis-by-Neural-Synthesis. 16 InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...
2024
-
[29]
Soubhik Sanyal, Timo Bolkart, Haiwen Feng, and Michael J Black. 2019. Learning to regress 3D face shape and expression from an image without 3D supervision. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 7763–7772
2019
-
[30]
Zhiyao Sun, Tian Lv, Sheng Ye, Matthieu Lin, Jenny Sheng, Yu-Hui Wen, Minjing Yu, and Yong-Jin Liu
-
[31]
Ayush Tewari, Michael Zollhofer, Hyeongwoo Kim, Pablo Garrido, Florian Bernard, Patrick Perez, and Christian Theobalt. 2017. MoFA: Model-based deep convolutional face autoencoder for unsupervised monocular reconstruction. InIEEE International Conference on Computer Vision (ICC...
2017
-
[32]
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy. 2020. MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation. InEuropean Conference on Computer Vision (ECCV). Springer, 700–717
2020
-
[33]
Zidu Wang, Xiangyu Zhu, Tianshuo Zhang, Baiqin Wang, and Zhen Lei. 2024. 3D Face Reconstruction with the Geometric Guidance of Facial Part Segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 1672–1682
2024
-
[34]
Jiahao Wu, Yunfei Liu, Lijian Lin, Ye Zhu, Lei Zhu, Jingyi Li, and Yu Li. 2026. PEAR: Pixel-aligned Ex- pressive humAn Mesh Recovery. InACM SIGGRAPH 2026 Conference Papers. arXiv:2601.22693 [cs.CV]
2026
-
[35]
Xitong Yang, Devansh Kukreja, Don Pinkus, Anushka Sagar, Taosha Fan, Jinhyung Park, Soyong Shin, Jinkun Cao, Jiawei Liu, Nicolas Ugrinovic, Matt Feiszli, Jitendra Malik, Piotr Dollar, and Kris Kitani
-
[36]
Wanqi Yin, Zhongang Cai, Ruisi Wang, Ailing Zeng, Chen Wei, Qingping Sun, Haiyi Mei, Yanjun Wang, Hui En Pang, Mingyuan Zhang, Lei Zhang, Chen Change Loy, Atsushi Yamashita, Lei Yang, and Ziwei Liu
-
[37]
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan. 2021. Flow-Guided One-Shot Talking Face Generation With a High-Resolution Audio-Visual Dataset. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 3661–3670
2021
-
[38]
Xiangyu Zhu, Zhen Lei, Xiaoming Liu, Haiming Shi, and Stan Z Li. 2016. Face alignment across large poses: A 3D solution. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 146–155
2016
-
[39]
doi:10.1109/TPAMI.2025.3618174
SMPLest-X: Ultimate Scaling for Expressive Human Pose and Shape Estimation.IEEE Transactions on Pattern Analysis and Machine Intelligence48, 2 (2026), 1778–1794. doi:10.1109/TPAMI.2025.3618174
2026
-
[42]
Wojciech Zielonka, Timo Bolkart, and Justus Thies. 2022. Towards metrical reconstruction of human faces. InEuropean Conference on Computer Vision (ECCV). Springer, 250–269. 17 Appendix This section contains additional details on the implementation of the proposed method. The a...
2022
-
[2024]
doi:10.1145/3658221
DiffPoseTalk: Speech-Driven Stylistic 3D Facial Animation and Head Pose Generation via Diffusion Models.ACM Transactions on Graphics (TOG)43, 4, Article 46 (2024), 9 pages. doi:10.1145/3658221
2024 doi
-
[2026]
InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
SAM 3D Body: Robust Full-Body Human Mesh Recovery. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). arXiv:2602.15989 [cs.CV]
Reviewed July 30, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.