{"id":"9fccc893-ccf3-45a2-9a78-cf5c392689ea","arxiv_id":"2505.05672","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Avatars rendered with up to four million 3D Gaussians embedded in the continuous UV texture space of a face mesh, warped by an expression-dependent deformation field, improve detail and expression fidelity over prior Gaussian head avatars.","lead":"This paper describes a method for building highly detailed, animatable 3D head avatars from multi-view video by embedding millions of small volumetric primitives on the surface of a face model. It matters because high-fidelity digital humans are a bottleneck for telepresence, virtual production, and immersive entertainment.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported SOTA gains are not internally reproducible: the same 'Ours' row differs across the main and appendix tables, so the headline improvement claim lacks stable quantitative support.","rationale":"I read the paper in good faith as proposing a coherent UVD-space Gaussian representation with adaptive densification, a Jacobian-based covariance warp, and expression-dependent shading. The method is plausible, and the mesh-tracking dependence identified by the reader is a genuine scope limitation: it is disclosed for teeth, tongue, and mouth interior, and it would narrow the claim to geometries where the 3DMM or Multiface meshes are accurate. Such a limitation does not falsify the method on the tested data. The more immediate gate is the quantitative evidence: identical-looking 'Ours' rows in the main and appendix tables do not agree, and the natural aggregation check fails. Since the central claim is precisely that TeGA improves on GA, RGCA, MVP, and GHA, the reported numbers are the load-bearing evidence. The reader's CONDITIONAL verdict already captures the need for reconciliation, so I do not shift the verdict; I would keep it CONDITIONAL until the tables are made internally consistent and the 4K claim is either evaluated or qualified.","tokens_in":21086,"tokens_out":10649,"duration_ms":116652,"concrete_test":"Request the per-subject and per-dataset metric logs, then recompute Table 1 from the stated subject split: 8 in-studio subjects (Table 4) plus 4 Multiface subjects (Table 5), with per-subject equal weighting. Verify every row, especially 'Ours' and 'No Densify', including Landmark, LPIPS, SSIM, and PSNR, and reconcile the Gaussian budget between Section 3.8 (4M) and Appendix D/Table 2 (6M). If the main table matches the weighted aggregate, the concern is resolved; if not, the SOTA comparison should be regenerated and the caption semantics clarified before the claim is cited.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a quantitative SOTA claim, so the load-bearing condition is that the numbers in Section 5 are stable. They are not as printed. The main Table 1 and the appendix table captioned 'Quantitative comparison on our dataset' disagree on identical-looking rows: 'Ours' is Landmark 31.7 +/- 28.6, LPIPS 0.150, SSIM 0.779, PSNR 24.4 in Table 1 versus 25.7 +/- 10.5, 0.120, 0.781, 24.9 in the appendix table. If Table 1 is meant to be the combined result over the 8 in-studio subjects (appendix Table 4) and 4 Multiface subjects (appendix Table 5), subject-weighted checks on 'Ours' (Landmark about 30.4 versus 31.7; LPIPS about 0.147 versus 0.150) do not reproduce the main table. The Gaussian budget also conflicts: Section 3.8 sets an upper bound of 4M, while Appendix D says the upper limit is 6M and Table 2 includes a 6M row. Without released code or per-subject metric logs, a reader cannot determine which numbers correspond to the actual method. This does not mean the algorithm is bad; it means the quantitative core of the paper is not yet verifiable as written, and the headline claim should be evaluated only after the tables are reconciled.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces TeGA, an animatable 3D head avatar model built on 3D Gaussian splatting. The key idea is to define canonical Gaussians in the continuous UVD tangent space of a tracked 3DMM mesh, propagate mesh deformation to Gaussian covariance through the Jacobian of the UVD-to-world mapping, and add a U-Net/MLP-based residual deformation field and an expression-dependent dynamic shading term. Adaptive densification is retained from 3DGS, allowing millions of Gaussians, and the method is evaluated on a new 8-subject studio dataset and on 4 Multiface subjects against GaussianAvatars, GaussianHeadAvatars, RGCA, and MVP, with reenactment and novel-view metrics. The paper claims state-of-the-art quality at high resolution, with extensive ablations of the deformation field, shading, densification, triangle updates, and loss terms.","tokens_in":21349,"tokens_out":5132,"duration_ms":55157,"significance":"If the reported results are reproducible, TeGA is a meaningful advance for high-resolution animatable head avatars: it decouples the latent texture resolution from the number of Gaussians, uses a principled Jacobian-based covariance transformation, and demonstrates qualitative gains in close-up detail in the included figures. The experimental design is strong in breadth: four baselines, two datasets, multiple ablations, and explicit discussion of limitations such as mouth interior artifacts and the dependence on tracked meshes. The main obstacle is that the quantitative core of the paper is not internally consistent as printed: the same named configurations receive different numbers in the main and appendix tables, and the Gaussian-count cap is stated differently in two places. The headline SOTA claim therefore cannot yet be verified from the manuscript as written.","major_comments":[{"comment":"The main quantitative table and the appendix tables report conflicting values for identically labeled configurations. For example, the 'Ours' row on 'our dataset' gives Landmark 31.7 ± 28.6, LPIPS 0.150, SSIM 0.779, PSNR 24.4 and novel-view LPIPS 0.123 in Table 1, while Appendix Table 4 gives Landmark 25.7 ± 10.5, LPIPS 0.120, SSIM 0.781, PSNR 24.9 and novel-view LPIPS 0.099. The 'Ours (200K GS.)' rows differ similarly (Landmark 49.0 ± 26.1 vs 37.2 ± 19.4), and every ablation row in Table 1 differs from the corresponding row in Appendix Table 4. Because the central claim is a quantitative improvement over the state of the art, the paper must present one consistent set of per-dataset numbers, state precisely how the aggregates in each table are computed, and preferably release per-subject metric logs so the results can be verified. This issue is load-bearing and must be resolved before the paper can be accepted.","section":"§5.1, Table 1 vs Appendix G, Tables 4–5"},{"comment":"The Gaussian-count upper bound is stated inconsistently. Section 3.8 says 'we set an upper bound of 4 million Gaussians to avoid running out of memory,' while Appendix D.2 says 'we set an upper limit of at most 6M Gaussians' and Table 2 reports results for a 6M row. The appendix explanation that most subjects do not exceed 4M is plausible, but as printed the two statements conflict. Please state the exact cap used in each experiment reported in Tables 1, 2, 4, and 5, and distinguish the implementation cap from the observed Gaussian counts.","section":"§3.8 vs §D.2, Table 2"},{"comment":"The claim that adaptive densification is 'critical' is not uniformly supported by the appendix table. In Table 1, removing densification causes a large drop (Landmark 122.7 ± 147, LPIPS 0.274, SSIM 0.738, PSNR 21.3 with full-model values of 31.7, 0.150, 0.779, 24.4). In Appendix Table 4, however, the same ablation on the same named dataset gives Landmark 68.6 ± 19.3, LPIPS 0.223, SSIM 0.780, PSNR 24.4, which is only marginally worse than the full model in SSIM and PSNR. This discrepancy affects the paper's explanation of why densification matters and should be reconciled with a clear statement of which table corresponds to the reported protocol.","section":"§5.2.3, Table 4 vs Table 1"}],"minor_comments":[{"comment":"The abstract and introduction describe the method as rendering at '4K resolution,' but the studio data are trained at 3072×2048 after downsampling and Multiface at 2048×1334. Please clarify the relationship between the native camera resolution, the training resolution, and the '4K' claim.","section":"Abstract, §1, §4"},{"comment":"Equation (6) introduces λ_JD for the Jacobian smoothness term, but the following paragraph refers to λ_smooth with values 1.0 decaying to 0.1. Please use consistent notation and specify where λ_smooth enters Eq. (6).","section":"§3.6, Eq. (6)-(8)"},{"comment":"The landmark evaluation is described as the 'mean-squared difference' between detected keypoints, but the reported values are not given units. Please state whether the metric is squared pixel distance, mean Euclidean pixel distance, or normalized coordinates, and whether the average is taken over landmarks or frames.","section":"§5, Landmark metric"},{"comment":"The baseline name appears as 'GaussianHeadAvatar' in Appendix Tables 4–5 and 'GaussianHeadAvatars' in the main text and Table 1. Please use one name consistently.","section":"Tables 1, 4, 5"},{"comment":"The U-Net description says each block has 'two convolutional layers, the first layer downsampling or upsampling using striding or transposing respectively,' but Figure 13 appears to show a skip connection with a separate strided/transposed convolution. Please make the block description match the figure precisely.","section":"§E.1"}],"recommendation":"major_revision","confidential_remarks":"The discrepancy between Table 1 and Appendix Table 4 is, in my view, the deciding issue. It is not a superficial typo: every row differs for the same named setting, including the headline 'Ours' row, and simple subject-weighted combinations of the appendix per-dataset splits do not reproduce the main table. I would require the authors to submit a corrected, unified results set, with explicit aggregation rules and ideally per-subject numbers, before the paper can be considered for acceptance. The method itself is interesting and the qualitative results are compelling, so I do not see this as a reject; it is a reproducibility/verifiability problem that should be fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take on arXiv:2505.05672. The method is genuinely new—continuous UVD-space Gaussian placement with adaptive densification and a deformation net that outputs only a translation, with rotation/scale recovered via the Jacobian. That's a clean trick, and the qualitative close-ups show a real quality jump over RGCA, GA, MVP, and GHA. I'd want to cite this.\n\nBut the numbers are not internally consistent, and that's a problem for a paper whose main claim is quantitative SOTA. Table 1 and Table 4 are both labeled 'our dataset' but disagree on identical rows: Ours is Landmark 31.7 / LPIPS 0.150 in one, 25.7 / 0.120 in the other. The Gaussian cap is 4M in Sec 3.8 and 6M in Appendix D.2. The 4K claim is never actually tested. And the ablation that should support 'free movement across triangles'—No Triangle Updates—gives a mixed picture: in Table 4 it actually has slightly better LPIPS than Ours (0.119 vs 0.120), only winning on landmark distance. That undercuts one of the stated key contributions.\n\nNone of this makes me think the method is bad. The pieces are well engineered, and the limitations section is honest about mesh tracking issues. But the quantitative core as printed isn't verifiable. I'd send this to a serious referee, with the expectation of major revision: reconcile the tables, clarify the Gaussian budget, either show the triangle-update benefit convincingly or demote it, and be upfront about resolution.\n\nFor a reading group, it's a useful paper to discuss—the Jacobian-based deformation is a nice idea that might generalize. I'd take it.","headline":"Genuinely new Gaussian avatar method with strong qualitative results, but the quantitative tables contradict themselves—needs major revision before the SOTA claim is credible.","tokens_in":21968,"tokens_out":3266,"would_cite":true,"duration_ms":31883,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper presents TeGA, a hybrid head-avatar model that embeds 3D Gaussians in the continuous UV-displacement tangent space of a tracked face mesh, warps them with an expression-dependent field, and renders high-fidelity 4K reenactments…","keywords":["3D Gaussian splatting","head avatars","UVD texture space","expression-dependent deformation","dynamic shading","4K rendering","adaptive densification","facial reenactment"],"falsifier":"Take a multiview capture with known ground-truth geometry (for example, synchronized 4D scans) and compare the full TeGA model trained on the provided tracked meshes versus the same model trained on the ground-truth meshes, measuring close-up fidelity in the mouth interior; if both versions show the same artifacts there, the tracked-mesh assumption is not the limiting factor, whereas if the ground-truth version improves, the central detail claim depends on mesh accuracy as the paper acknowledges.","tokens_in":20847,"feed_emoji":"🎭","tokens_out":7779,"duration_ms":76624,"temperature":0.7,"pith_summary":"The paper aims to close the gap between animatable head avatars and per-expression reconstructions by building a 3D Gaussian representation in the continuous UVD (UV plus surface-normal displacement) tangent space of a tracked face mesh. This lets Gaussians move freely across mesh triangles during optimization, so adaptive densification can concentrate detail where it is needed instead of being limited by a fixed UV texture resolution. A U-Net reads rasterized vertex displacements and produces expression-dependent codes that drive a learned residual deformation field and a monochromatic shading term, capturing transient features such as wrinkles and skin darkening. The authors report that, on multiview video of eight in-studio subjects plus the Multiface dataset, TeGA improves landmark alignment, perceptual similarity, and close-up fidelity over GaussianAvatars, GaussianHeadAvatars, RGCA, and MVP, while rendering at 4K with up to four million Gaussians. The value of the claim, if correct, is that controllable avatars no longer trade animation flexibility for high-resolution detail.","feed_headline":"Texture-space Gaussians render 4K head avatars with finer detail","feed_subtitle":"Rigging millions of Gaussians in UV-displacement space preserves wrinkles, stubble, and expression detail up close.","key_machinery":"The load-bearing object is the continuous UVD tangent space: each canonical Gaussian lives at a UV texture coordinate plus a scalar displacement $D$ along the locally interpolated surface normal, so the whole set acts as a sparse volumetric texture indexed by the mesh. The mapping $F(\\boldsymbol{\\mu}_{uvd})$ to world space, together with the Jacobian-based covariance transform $\\Sigma_{xyz} = J_F \\Sigma_{uvd} J_F^T$, converts triangle deformation into Gaussian rotation and stretching, including non-rigid stretching. A shared U-Net encodes rasterized neutral-relative vertex displacements into a $256 \\times 256 \\times 64$ feature texture; bilinear interpolation gives each Gaussian a local expression code $f_{uv}$, which conditions a shallow MLP residual deformation field and a shallow MLP shading head. Because the feature texture dimension is decoupled from the Gaussian count, the model keeps 3D Gaussian splatting's clone-and-split densification, bounded at 4 million Gaussians, and thus allocates primitives where detail demands them.","core_discovery":"The central claim is that a tracked 3DMM mesh can serve as a coarse deformation layer while photorealistic detail lives in a sparse volumetric texture of 3D Gaussians embedded in the mesh's continuous UVD space. Canonical Gaussians are parameterized by $\\boldsymbol{\\mu}_{uvd}$ with a displacement $D$ along the surface normal; the mapping $F$ from UVD to world coordinates and its Jacobian $J_F$ propagate mesh deformation to each Gaussian's position and covariance. A residual translation field $D(\\cdot)$ conditioned on U-Net features $f_{uv}$ adds fine, expression-dependent motion, and the combined Jacobian $J_{(F+D)}$ reshapes Gaussians accordingly. An expression-dependent shading factor $S(\\boldsymbol{\\mu}_{uvd}; f_{uv})$ darkens wrinkles without changing chromaticity. The paper argues that this combination, non-greedy cross-triangle movement, adaptive densification up to millions of Gaussians, and network-based deformation and shading, is what allows 4K close-ups with sharp wrinkles, stubble, and brows while remaining controllable through expression and pose parameters.","pith_inferences":["[Editorial inference] The UVD-plus-Jacobian rigging is a general template for animating any deformable surface with a UV atlas, so the same design could transfer to hands, bodies, or clothing if a tracked template mesh is available; the paper does not test these cases.","[Editorial inference] A likely testable extension is replacing the monochromatic shading with a small chromatic spherical-harmonics or albedo correction conditioned on the same expression codes; this would directly address the paper's stated blood-flow limitation.","[Editorial inference] The paper's appendix finding that a scaled-rigid per-triangle transform matches the Jacobian when no residual field is used suggests the larger wins come from cross-triangle Gaussian movement and densification; validating the full model with the residual field active against scaled-rigid rotation would isolate each contribution.","[Editorial inference] If mesh tracking accuracy improves for the mouth interior and teeth, the same pipeline should inherit those gains, since the paper identifies tracked-mesh quality as the limiting factor for those regions."],"forward_implications":["High-resolution close-ups of animated heads can be rendered directly from 3D Gaussians, without a separate super-resolution network on the image plane; the paper's 4K results are produced by the representation itself.","The same UVD rigging works on any tracked mesh sequence in correspondence, not only the paper's 3DMM, since it is demonstrated on Multiface meshes; the method is therefore portable across tracking systems.","Because densification is retained, per-region detail such as stubble, brow hairs, and wrinkles is allocated adaptively during optimization rather than being fixed by a UV-map texel budget.","The residual deformation field plus Jacobian propagation extracts rotation and stretching from a simple learned translation field, making expression-dependent motion easier to learn than directly predicting full pose and scale per Gaussian.","Expression-dependent monochromatic shading captures occlusion-like darkening in wrinkles; the paper's limitation is that it cannot change Gaussian chromaticity, so dynamic skin-color effects such as blood flow remain outside the model."],"supporting_citations":[{"why":"Supplies the 3D Gaussian splatting primitives, adaptive densification, and L1 plus D-SSIM rendering loss that TeGA builds on and retains.","marker":"[Kerbl et al. 2023]"},{"why":"Defines the triangle-anchored Gaussian rigging baseline that TeGA generalizes by allowing Gaussians to move across triangles, and is a main comparison method.","marker":"[Qian et al. 2024]"},{"why":"Provides the RGCA baseline, a UV-map Gaussian codec avatar with a fixed Gaussian budget, which TeGA compares against in reenactment and novel-view tests.","marker":"[Saito et al. 2024b]"},{"why":"Provides the Gaussian Head Avatar baseline with dynamic Gaussians and keypoint-based supervision, used as a state-of-the-art comparison.","marker":"[Xu et al. 2024]"},{"why":"Supplies the Multiface multiview dataset and its tracked meshes, used to train and evaluate TeGA and all baselines on a public benchmark.","marker":"[Wuu et al. 2022]"},{"why":"Defines the FLAME-style 3D morphable model that provides expression and pose parameters and the tracked mesh topology for the in-studio dataset.","marker":"[Li et al. 2017]"},{"why":"Provides the MVP baseline with fixed voxel primitives, a comparison point for the paper's claim that adaptive Gaussian density improves detail.","marker":"[Lombardi et al. 2021]"},{"why":"Supplies the VGG perceptual loss used in the training objective to recover fine detail, with a coarse-to-fine schedule described in the paper.","marker":"[Zhang et al. 2018]"},{"why":"Supplies the warm-up strategy and learning-rate scheduling that TeGA adapts for initially training canonical Gaussians before enabling the deformation field.","marker":"[Yang et al. 2024]"}],"fun_headline_variants":["Millions of Gaussians in texture space give 4K head avatars crisp detail","TeGA packs millions of Gaussians for 4K avatar fidelity","Texture-space Gaussians capture wrinkles and stubble at 4K","UVD-space Gaussians preserve fine facial motion at 4K","High-detail head avatars from millions of texture-space Gaussians"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the tracked 3D meshes give dense, accurate correspondence across expressions, so the UV tangent space and its Jacobian faithfully describe how the skin deforms; the paper itself notes this is unreliable for teeth, tongue, and the mouth interior.","fun_headline_variants_meta":{"raw":{"variants":["Millions of Gaussians in texture space give 4K head avatars crisp detail","TeGA packs millions of Gaussians for 4K avatar fidelity","Texture-space Gaussians capture wrinkles and stubble at 4K","UVD-space Gaussians preserve fine facial motion at 4K","High-detail head avatars from millions of texture-space Gaussians"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000564,"raw_usage":{"total_tokens":2746,"prompt_tokens":1087,"completion_tokens":1659,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":703,"completion_tokens_details":{"reasoning_tokens":1564}},"tokens_in":703,"tokens_out":1659,"duration_ms":11541,"temperature":1.0,"reasoning_tokens":1564,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:00:06.712777+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a multiview capture with known ground-truth geometry (for example, synchronized 4D scans) and compare the full TeGA model trained on the provided tracked meshes versus the same model trained on the ground-truth meshes, measuring close-up fidelity in the mouth interior; if both versions show the same artifacts there, the tracked-mesh assumption is not the limiting factor, whereas if the ground-truth version improves, the central detail claim depends on mesh accuracy as the paper acknowledges.","supporting_citations":[],"review_version":1}