{"id":"934b4bba-a933-437d-be6c-503a9a559f83","arxiv_id":"2504.20403","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"TetGS is a hybrid representation that embeds Gaussian kernels inside tetrahedral grids, enabling locally controlled geometric and appearance edits of 3D avatars reconstructed from monocular video.","lead":"What if you could take a short video of yourself and then type 'make my shirt a denim jacket' and get a 3D avatar that actually changes shape and looks real? This paper introduces a system that couples tetrahedral mesh control with Gaussian splatting to edit full-body 3D avatars from monocular video, guided by text or reference images.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed quantitative superiority rests on a self-similarity FID protocol; the central 'photorealistic and outperforms' claim is not yet supported.","rationale":"I agree with the reader's conditional verdict but not fully with the choice of weakest assumption. The most load-bearing point is not a potential implementation gap in Sec. 3.1/3.2.2; it is that Table 1's FID is computed against the method's own reconstructions, making it a self-similarity score. The paper explicitly defines lower FID as greater similarity to reconstructed images, not to real photographs, and a no-op edit would trivially achieve near-zero FID. The mapping/topology issue is a legitimate robustness concern, but it is partially acknowledged in Supp. H and would only affect large deformations; the empirical claim of superiority would remain unsupported even if that issue were fixed. The reader did mention the unusual FID protocol in the rationale, so there is partial agreement, but the weakest_assumption field selected a different, more speculative concern. The verdict should remain CONDITIONAL: the method is plausible and the qualitative results are suggestive, but the central empirical claim needs a corrected evaluation protocol and ideally code or data release before acceptance.","tokens_in":21411,"tokens_out":5270,"duration_ms":57080,"concrete_test":"Recompute Table 1 using real held-out video frames (e.g., 60 frames not used in reconstruction) as the FID reference set, instead of the system's own reconstructed images, and include a no-edit control where source avatar renderings are scored against the same reference. Also run each method with at least three random seeds and report mean and standard deviation for FID, CLIP, and DINO. If the no-edit control matches or beats TetGS's FID, or if the gap to DGE/TIP-Editor is within one standard deviation under the corrected protocol, the quantitative superiority claim fails and the paper should be revised to claim comparable quality pending a user study.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing weakness is the quantitative evidence for the central claim that TetGS 'outperforms' GaussianEditor, DGE, and TIP-Editor and yields photorealistic avatars (conclusion, Sec. 4.2, Table 1). The FID in Table 1 is computed between edited renderings and 'the reconstructed images' produced by the method itself; the paper even states that lower FID indicates 'greater similarity to the reconstructed images.' This makes FID a self-similarity measure, not a photorealism measure. A degenerate no-op edit that leaves the reconstruction unchanged would achieve near-zero FID, so the reported 115.95 vs. 194.54 gap cannot be interpreted as higher fidelity or realism. The remaining metrics (CLIP, DINO) are computed on 60 rendered images with no error bars, no multiple seeds, and no human evaluation; DINO is reported only for TIP-Editor, not for GaussianEditor or DGE. This is not a contradiction in the method, but it is missing support for the strongest empirical claim. The tetrahedron-mapping robustness issue identified by the reader is real and worth testing, but it concerns an edge case (large geometric deformation) that the authors partly acknowledge in Supp. H; even if that issue were fully resolved, the empirical superiority claim would still rest on the flawed FID protocol.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a three-stage pipeline for creating editable 3D avatars from a monocular video, using a hybrid tetrahedron-constrained Gaussian Splatting (TetGS) representation. In the first stage, a neural SDF field is optimized on the video and converted into a tetrahedral grid with embedded Gaussians. The second stage performs localized spatial adaptation by explicitly partitioning tetrahedra into preserved and editable sets, updating the SDF of editable vertices under global and local SDS losses plus surface-aware regularizations. The third stage generates appearance by first learning a coarse texture under few-shot normal-based inpainting and then refining the appearance with an attribute-activation I2I refinement. The method supports both text-guided editing and reference-image-based virtual try-on. The authors claim that the resulting avatars are high-fidelity, photorealistic, and quantitatively superior to GaussianEditor, DGE, and TIP-Editor on their collected 10-subject dataset.","tokens_in":21677,"tokens_out":2758,"duration_ms":29958,"significance":"If the empirical claims are substantiated, the paper makes a useful contribution: it provides an accessible pipeline from a monocular video and a text/image prompt to an editable, locally controllable 3D avatar, addressing a real gap in 3DGS-based editing. The technical presentation is a strength: the TetGS representation is described with explicit equations (Eq. 1–7), the three-stage pipeline is clearly motivated, and the supplementary material provides architectural details and ablations (Tables 2–3, Figs. 7–9). The decoupling of geometry and appearance optimization is a plausible remedy for the instability noted in prior 3DGS editing. However, the central quantitative claim of superiority over baselines is not supported by the reported evaluation, as the FID protocol measures similarity to the method's own reconstructions rather than to real imagery. The paper would be significantly strengthened by a re-evaluation against genuine photorealistic references, statistical error bars, and a direct test of the tetrahedron-mapping robustness for large geometric edits.","major_comments":[{"comment":"The FID metric is computed between edited renderings and the system's own reconstructed images, as stated in Sec. 4.2: \"a lower FID indicates greater similarity to the reconstructed images.\" This is a self-similarity measure, not a photorealism or fidelity measure, and a degenerate no-op edit that leaves the reconstruction unchanged would achieve near-zero FID. The reported gap (115.95 vs. 194.54/201.82/258.27) therefore cannot be interpreted as evidence that the edited avatars are more photorealistic or higher-fidelity than the baselines. The quantitative comparison should be re-run against a fixed set of real photographs (e.g., held-out frames from the input video or real-captured images of people), or supplemented by a user study and by perceptual metrics that do not reference the method's own output.","section":"Sec. 4.2, Table 1"},{"comment":"All quantitative metrics are computed on only 60 rendered images from a 10-subject dataset, with no error bars, no multiple seeds, and no statistical significance test. The claimed \"significant improvement\" in FID and the 26.28 vs. 22–23 CLIP gap may be within run-to-run variation; the paper does not provide the variance or confidence intervals needed to support the superiority claim. Additionally, DINO similarity is reported only for TIP-Editor and not for GaussianEditor or DGE, so the table does not provide a complete comparison across methods on the image-guided task. The authors should report mean and standard deviation over at least three independent runs per method and per metric, and should either complete the DINO column for all baselines or explain why it is inapplicable.","section":"Sec. 4.2, Table 1; Datasets paragraph"},{"comment":"The robustness of the father-tetrahedron mapping hf→t (Sec. 3.1) under large geometry changes is load-bearing for the localized spatial adaptation module. When the SDF values at editable vertices are optimized as described in Sec. 3.2.2, the Marching Tetrahedron algorithm can produce inverted or topologically different tetrahedra; the paper only states that the reallocated Gaussians Gedit are initialized \"by following the procedure described in Sec. 3.1\" and does not analyze whether the mapping remains consistent or whether preserved tetrahedra Vkeep_T can be corrupted. The authors acknowledge in Supp. H that loose-to-tight garment edits are problematic; this suggests the mapping reliability is an edge case that deserves a specific diagnostic: e.g., report the fraction of tetrahedra that change sign, or show that the preserved-region geometry and appearance remain unchanged (measured by local reconstruction error) for a large deformation. Without such a test, the claim of \"precise region localization, geometric adaptability\" is not fully supported.","section":"Sec. 3.1, Sec. 3.2.2, Supp. H"}],"minor_comments":[{"comment":"The text says \"For quantitative evaluation, we use Frechet Inception Distance (FID) to access the quality of edited images.\" The verb should be \"assess,\" and the metric name should be spelled \"Fr\\'echet\" with the accent.","section":"Sec. 4.2"},{"comment":"The sentence \"producing photorealistic results comparable to real-world individuals\" is not backed by any metric that compares to real-world photographs; consider softening it or providing such evidence.","section":"Sec. 4.2, Table 1"},{"comment":"The paper reports absolute FID/CLIP values without confidence intervals in both the main comparison and the ablation. Reporting error bars even for the ablation table would help assess stability, especially since the \"w/o AA\" variant is only 4.38 FID points above the full model.","section":"Sec. 4.2, Table 1 and Sec. 4.3, Table 2"},{"comment":"The qualitative comparisons in Fig. 6 are informative, but the text does not specify the training time or iteration budget given to each baseline. Since the paper emphasizes efficiency (e.g., 1.2 hours for spatial adaptation), a brief statement about GPU-hours per method would aid reproducibility and fairness.","section":"Sec. 4.2, qualitative comparison"},{"comment":"In the supplementary Table 3, the text \"3DGS-30K Ours-7K 3DGS-7K . We report the averaged chamfer distance, PSNR,...\" contains a garbled fragment that appears to be a formatting error; please fix it.","section":"Supp. A.3, Table 3"}],"recommendation":"major_revision","confidential_remarks":"The technical core of TetGS is sound and the paper is well-written, but the evaluation does not support the central claim of quantitative superiority because the FID is a self-similarity measure. I would be willing to reconsider after the authors replace the FID protocol with a comparison against real images, add error bars and multi-run statistics, and provide a direct test of the tetrahedron-mapping robustness under large deformations. The scope and contribution of the paper are appropriate for this venue; the issue is in the evidence, not in the method's novelty."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The TetGS representation is the real news here: embedding Gaussians inside tetrahedral cells via the Marching Tetrahedron mapping gives you a way to move geometry and have the Gaussians follow, and the explicit frozen/editable partition is a clean solution for localized editing. The two-phase decoupling (spatial adaptation first, then appearance) is well motivated, and the ablations actually support the design choices. Reconstruction quality is also on par with plain 3DGS while converging faster, which is a concrete plus.\n\nThe main soft spot is exactly what the stress-test note flags: the FID in Table 1 is computed between edited renderings and the system's own reconstructed images. That makes it a self-similarity score, not a photorealism score. A degenerate no-op edit would get near-zero FID, so the gap between 115.95 and 194.54 is not evidence that your output is more photorealistic. The paper even states this definition plainly, so it is not hidden, but the conclusion then overclaims \"photorealistic\" and \"superiority\" on that basis. The other metrics are computed on 60 rendered images with no error bars, no human evaluation, and only ten subjects; DINO is only reported for TIP-Editor. That is thin support for the central claim.\n\nThe tetrahedron-mapping concern from the reader is real but I would call it a moderate robustness gap rather than a fatal one. When SDF values change enough to invert or retopologize a tet, the paper does not describe how Gedit are reallocated, only that they follow the Sec 3.1 procedure. Supp. H acknowledges loose-to-tight garments as a limitation, so the authors know the edge case. That is acceptable for a preprint, but it should be addressed in a revision.\n\nI would send this to peer review. The representation is novel, the pipeline is clear, and the ablations are honest. The quantitative evaluation needs fixing: report FID against real images or at least against the input video frames, add error bars and ideally a small user study, and either soften the photorealism claim or back it up. If those are addressed, this becomes a solid, citable method paper for 3D avatar editing and virtual try-on.","headline":"A genuinely new hybrid representation for localized 3D avatar editing with a sensible decoupled pipeline, but the headline empirical claim is undercut by a self-similarity FID protocol.","tokens_in":22223,"tokens_out":1455,"would_cite":true,"duration_ms":16670,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A hybrid tetrahedron-constrained Gaussian splatting representation lets users turn a monocular video into a photorealistic 3D avatar and edit it locally with text prompts or reference images.","keywords":["tetrahedron-constrained Gaussian splatting","3D avatar editing","monocular video reconstruction","text-guided editing","virtual try-on","signed distance field","localized spatial adaptation","Gaussian splatting"],"falsifier":"Record a subject in loose clothing, edit the garment to a tight one, and compare renderings of the preserved regions before and after editing. If the Marching Tetrahedron extraction produces inverted tetrahedra or reassigns preserved triangles to different father tetrahedra, the preserved regions will change or the edit will fail geometrically; the paper's own supplementary material identifies this loose-to-tight case as a limitation, so a concrete failure there would falsify the claim of reliable localized geometric adaptation.","tokens_in":21174,"feed_emoji":"👤","tokens_out":12987,"duration_ms":110953,"temperature":0.7,"pith_summary":"This paper claims that an ordinary user can go from a short monocular video of a person to a photorealistic 3D avatar whose clothing or accessories can be edited locally, using either a text prompt or a reference garment image. The key move is a new representation, TetGS, which embeds 3D Gaussian splatting kernels inside a tetrahedral grid derived from a signed distance field, so geometry and appearance can be optimized in separate stages instead of in one unstable joint optimization. Editing proceeds in three stages: reconstruct the avatar, deform only the tetrahedra in the edited region under diffusion-based normal guidance while freezing the rest, and generate texture on the moved Gaussians from few-shot inpainted views before a final refinement step. The authors report that this pipeline outperforms three prior 3D Gaussian editing methods on FID, CLIP, and DINO scores on a collected dataset of ten monocular videos.","feed_headline":"TetGS: editable photorealistic 3D avatars from a single video","feed_subtitle":"A tetrahedral scaffold guides Gaussian splatting so you can change clothes or style with text or a reference image.","key_machinery":"The load-bearing object is TetGS, a hybrid representation in which each Gaussian is explicitly embedded in a father tetrahedron $t_k$ of a tetrahedral grid derived from a signed distance field. A signed distance field $\\psi_g$ is evaluated at grid vertices, the Marching Tetrahedron algorithm extracts mesh vertices by linear interpolation of the SDF values on tetrahedral edges, and each Gaussian's position is set as $\\mu = w_a v^M_{i1} + w_b v^M_{i2} + w_c v^M_{i3} + \\tau n$, where the $v^M$ are the vertices of the mesh triangle it sits on, $w$ are fixed barycentric weights, and $\\tau$ is a learnable displacement along the surface normal. The two identity-like maps $g_{v\\to e}$ (mesh vertex to its tetrahedral edge) and $h_{f\\to t}$ (mesh triangle to its father tetrahedron) carry the localization: they allow the method to freeze the SDF at vertices of preserved tetrahedra, update only editable vertices under a dual SDS loss, and re-initialize Gaussians in the edited tetrahedra by repeating the same extraction procedure. During texture generation, the editing Gaussians are temporarily degraded to 2D disks with fixed opacity and view-independent color, which stabilizes few-shot inpainting, and then their full covariance, opacity, and spherical-harmonic attributes are reactivated for the final refinement.","core_discovery":"The central discovery is that binding Gaussian kernels to tetrahedra converts an otherwise discrete, unstable 3D Gaussian splatting optimization into a controllable deformation problem: updating the signed distance values of tetrahedral vertices moves both the extracted surface and the Gaussians embedded on it. The father-tetrahedron mapping $h_{f\\to t}$ lets the method split the mesh into preserved and edited regions; frozen SDF values keep the preserved Gaussians unchanged, while editable vertices are optimized under a dual global-and-local SDS loss on normal renderings. After the geometry settles, the reallocated editing Gaussians are first restricted to 2D surfels and painted with few-shot normal-guided inpainted images, then their full 3D attributes are reactivated and refined with image-to-image augmented views. The conclusion states that this yields high-fidelity photorealistic 3D avatar editing with diverse identities and accessories, and the paper backs it with qualitative results and the reported metric improvements.","pith_inferences":["If the decoupling claim generalizes, the same tetrahedral scaffold could stabilize localized editing of any 3D Gaussian scene with a clean surface prior, not just human avatars; a direct test would be object-level editing with a mask-defined region on a generic reconstruction.","The reported gains are computed on the paper's own dataset of ten videos against three baselines; a stronger external check would be a public multi-view human dataset with ground-truth geometry, since FID and CLIP measure image statistics and text alignment rather than whether the deformed geometry is correct.","The loose-to-tight failure acknowledged in the supplementary suggests the real boundary is topological change in the tetrahedral grid; adding an estimated inner-body shape prior, as the authors suggest, would turn that failure mode into a testable fix.","Because the texture stage relies on a normal-based inpainter, its view consistency is inherited from that model; a stress test would render the edited avatar from unevenly spaced elevations and inspect shoulders and back for flicker or duplicated detail, which the paper's evenly spaced 60-view evaluation may not fully expose."],"forward_implications":["A 40- to 50-second monocular video plus a text prompt or a reference garment photo is sufficient to produce an editable, photorealistic 360-degree avatar, removing the need for multi-camera capture or manual modeling.","Because geometry and appearance are optimized separately, the method avoids the noise and blur that arise when Gaussian geometry and texture are trained together under generative guidance; the Fig. 7 and Tab. 2 ablations support this.","Non-edited regions stay visually intact because their tetrahedron vertices are frozen and a surface-aware regularization keeps the edited surface from occluding them; the Fig. 8 ablations show that removing the partitioning or either SDS term breaks this behavior.","Reference-image virtual try-on is achieved by substituting a try-on diffusion model for the normal inpainter and adding normal and mask supervision, so both garment style and its geometric design transfer to the avatar.","Supplementary results show the same pipeline supports texture doodling, by painting on guidance images, and continuous editing, by applying edits sequentially."],"supporting_citations":[{"why":"Supplies the 3D Gaussian representation and pixel-level optimization that TetGS embeds in tetrahedral grids.","marker":"[17]"},{"why":"Provides the tetrahedral grid and SDF-to-surface extraction that define the father-tetrahedron structure used by TetGS.","marker":"[37]"},{"why":"Defines the Marching Tetrahedron edge interpolation and triangulation used to build mesh vertices and the triangle-to-tetrahedron mapping.","marker":"[69]"},{"why":"Gives the SDS loss used in global and local form to drive localized spatial adaptation of the editable geometry.","marker":"[27]"},{"why":"Serves as the text-and-image-guided 3D editing baseline and motivating comparison for localized appearance control.","marker":"[3]"},{"why":"Provides the Instruct-NeRF2NeRF editing paradigm used both as a baseline and as the no-TetGS iN2N ablation variant.","marker":"[46]"},{"why":"Serves as the text-driven 3D Gaussian editing baseline whose semantic tracing and editing behavior the method compares against.","marker":"[52]"},{"why":"Serves as the multi-view-consistent 3D Gaussian editing baseline used for comparison.","marker":"[53]"},{"why":"Supplies the front- and back-view try-on images that supervise reference-image virtual try-on in the appearance stage.","marker":"[58]"},{"why":"Provides the normal-adapted diffusion prior used for geometric SDS guidance and the coarse-to-fine normal rendering schedule.","marker":"[32]"}],"fun_headline_variants":["TetGS: edit 3D avatars from a single video","Tetrahedra make Gaussian avatars editable","Photorealistic editable avatars via tetrahedral Gaussians","Edit 3D avatars with tetrahedron-constrained Gaussian splatting"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The approach assumes that deforming the editable region never flips or reconnects the tetrahedral mesh so severely that the link between a surface triangle and its parent tetrahedron breaks; under large geometric changes, such as loose to tight clothing, that link can fail and the promised localized editing would break down.","fun_headline_variants_meta":{"raw":{"variants":["TetGS: edit 3D avatars from a single video","Tetrahedra make Gaussian avatars editable","Photorealistic editable avatars via tetrahedral Gaussians","Edit 3D avatars with tetrahedron-constrained Gaussian splatting"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000638,"raw_usage":{"total_tokens":2964,"prompt_tokens":991,"completion_tokens":1973,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":1902}},"tokens_in":607,"tokens_out":1973,"duration_ms":14544,"temperature":1.0,"reasoning_tokens":1902,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:30:19.863674+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record a subject in loose clothing, edit the garment to a tight one, and compare renderings of the preserved regions before and after editing. If the Marching Tetrahedron extraction produces inverted tetrahedra or reassigns preserved triangles to different father tetrahedra, the preserved regions will change or the edit will fail geometrically; the paper's own supplementary material identifies this loose-to-tight case as a limitation, so a concrete failure there would falsify the claim of reliable localized geometric adaptation.","supporting_citations":[{"cited_title":"3d gaussian splatting for real-time radiance field rendering","cited_arxiv_id":null,"evidence_quote":"Supplies the 3D Gaussian representation and pixel-level optimization that TetGS embeds in tetrahedral grids."},{"cited_title":"Deep marching tetrahedra: a hybrid representation for high- resolution 3d shape synthesis,","cited_arxiv_id":null,"evidence_quote":"Provides the tetrahedral grid and SDF-to-surface extraction that define the father-tetrahedron structure used by TetGS."},{"cited_title":"An efficient method of triangulat- ing equi-valued surfaces by using tetrahedral cells,","cited_arxiv_id":null,"evidence_quote":"Defines the Marching Tetrahedron edge interpolation and triangulation used to build mesh vertices and the triangle-to-tetrahedron mapping."},{"cited_title":"Tip-editor: An accurate 3d editor following both text- prompts and image-prompts,","cited_arxiv_id":null,"evidence_quote":"Serves as the text-and-image-guided 3D editing baseline and motivating comparison for localized appearance control."},{"cited_title":"Instruct-nerf2nerf: Editing 3d scenes with instructions,","cited_arxiv_id":null,"evidence_quote":"Provides the Instruct-NeRF2NeRF editing paradigm used both as a baseline and as the no-TetGS iN2N ablation variant."},{"cited_title":"Gaus- sianeditor: Editing 3d gaussians delicately with text instruc- tions,","cited_arxiv_id":null,"evidence_quote":"Serves as the text-driven 3D Gaussian editing baseline whose semantic tracing and editing behavior the method compares against."},{"cited_title":"Humannorm: Learning normal diffusion model for high-quality and realistic 3d human generation,","cited_arxiv_id":null,"evidence_quote":"Provides the normal-adapted diffusion prior used for geometric SDS guidance and the coarse-to-fine normal rendering schedule."}],"review_version":1}