Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

DGNS: Deformable Gaussian Splatting and Dynamic Neural Surface for Monocular Dynamic 3D Reconstruction

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A two-module hybrid of deformable Gaussians and a dynamic neural surface claims state-of-the-art dynamic 3D geometry with rendering quality intact.

desk verdict Useful hybrid for dynamic reconstruction, but the printed depth-filter equation is wrong as written; fix it and this is a solid baseline. read the letter →

arxiv 2412.03910 v3 pith:DAQII2HG submitted 2024-12-05 cs.CV

classification cs.CV
keywords dynamicscenereconstruction3DGaussianSplattingneuralsigneddistancefunctionnovelviewsynthesismonocularvideodepthsupervisionsurface-awaredensitycontroldeformablesurface
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

DGNS claims that monocular dynamic 3D reconstruction does not have to choose between accurate geometry and high-quality rendering. It couples a deformable Gaussian splatting module, which handles appearance, with a dynamic neural signed-distance surface, which handles shape. The Gaussian module's rendered depths guide the surface module's ray sampling and supervise its SDF, while the surface's zero-level set steers where Gaussians are split, cloned, and pruned. On the Dg-mesh benchmark the paper reports the lowest Chamfer Distance and Earth Mover's Distance on most objects while staying competitive in PSNR, and on D-NeRF it keeps rendering quality near the best baseline while producing smoother meshes. If true, the practical payoff is a single video-to-mesh-and-view pipeline for articulate and deforming scenes.

What carries the argument

The load-bearing machinery is bidirectional depth-and-surface coupling between two modules that share a deformation field. The DGS module renders an alpha-blended depth and a median depth; a filter keeps a depth only when the two agree, and this filtered depth both centers the SDF ray-sampling interval and supplies the SDF regression target of Eq. (6). Conversely, the DNS zero-level set is converted into a Gaussian falloff signal that is added to the gradient criterion for splitting and cloning and subtracts from opacity for pruning, so Gaussian density follows the reconstructed surface. A monocular normal prior regularizes both modules. The named device "surface-aware density control" is what lets the explicit Gaussians inherit geometry without sacrificing the appearance fidelity of pure splatting.

What would settle it

Run the published Eq. (5) exactly on a ray where $|d_\alpha-d_m|<\tau_f$; the filtered depth collapses to near zero, and Eq. (6) then penalizes the SDF at the camera origin, which would destroy the reconstruction. Inspecting the released code to see whether it computes $(d_\alpha+d_m)/2$ instead settles whether the claimed depth-supervision mechanism exists.

Watch

Extended reading notes

Core claim

The central claim is that a two-module hybrid named DGNS beats both pure implicit surface methods and pure deformable-Gaussian methods because each representation supplies what the other lacks. Deformable 3D Gaussian splatting gives dense but noisy depth maps near the true surface, and those depths concentrate neural SDF ray sampling and provide supervision that pulls the zero-level set into place. The neural SDF in turn returns a geometry signal used to grow and prune Gaussians so they sit on the surface instead of floating in space. On Dg-mesh the paper reports the best reconstruction errors on nearly every object (for example, Chamfer Distance 0.773 versus 0.790 for the closest baseline on Duck, 0.289 versus 0.299 on Horse, and 0.413 versus 0.482 on Girlwalk) while PSNR stays competitive, and on D-NeRF DGNS is within a few tenths of a decibel of the rendering leader while producing qualitatively smoother meshes.

Load-bearing premise

The load-bearing premise is that the depth-filtering rule in Eq. (5) actually identifies the surface: as printed, when the alpha-blended and median depths are close the rule outputs about zero, so Eq. (6) would supervise the SDF at the camera rather than at the surface. If the intended midpoint average is not what the implementation uses, the depth-supervision contribution is not reproducible from the paper.

Editorial extensions

If this is right

  • A single monocular video can yield both a frame-consistent deforming mesh and photorealistic novel views, closing the geometry-versus-rendering tradeoff that separates implicit-only and Gaussian-only methods.
  • Depth rendered by deformable Gaussians can serve as a cheap guiding signal for neural SDF training: it shortens ray marching and anchors SDF supervision, so the surface module needs less blind search to converge.
  • The SDF's zero-level set can be used as a principled prior for Gaussian split, clone, and prune decisions, reducing the floaters that plague splatting-based surface reconstruction.
  • Monocular normal priors from a pretrained foundation model improve both appearance and geometry in under-constrained dynamic scenes, with the ablation reporting Chamfer Distance 1.006 with neither cue versus 0.502 with filtered depth and normals together.
  • On both the Dg-mesh and D-NeRF benchmarks, the hybrid's rendering PSNR remains competitive with the strongest deformable-Gaussian baseline, so the geometry gains do not come at a perceptual cost.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The printed Eq. (5) appears to contain a sign typo: under its own condition $|d_\alpha-d_m|<\tau_f$, the formula $(d_\alpha-d_m)/2$ makes the filtered depth nearly zero, so Eq. (6) would supervise the SDF at the camera origin rather than at the surface; the intended filter is most likely the midpoint $(d_\alpha+d_m)/2$.
  • The same alpha-plus-median depth filter could be lifted out as a general denoising step for any Gaussian-splatting depth map before it is used as supervision, independent of the neural surface module.
  • Surface-aware density control is a transferable recipe: any Gaussian-SDF hybrid, static or dynamic, could use the SDF zero-level set to decide where to add and remove primitives.
  • The quantitative geometry evidence is strongest where ground-truth meshes exist, namely Dg-mesh; on the real Nerfies sequence the evidence is qualitative, so behavior under real-world noise in monocular depths and normals is not yet quantified by the paper.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes DGNS, a hybrid framework that jointly optimizes a deformable 3D Gaussian splatting (DGS) module for view synthesis and a dynamic neural surface (DNS) module for geometry reconstruction in monocular dynamic scenes. The two modules interact in three ways: DGS-rendered depth maps guide DNS ray sampling and provide depth supervision, DNS SDF values guide Gaussian growth/pruning via surface-aware density control, and a foundation-model normal prior regularizes both modules. Experiments on the Dg-mesh, D-NeRF, and Nerfies datasets report state-of-the-art Chamfer Distance and EMD on Dg-mesh while remaining competitive with D3DGS in novel-view synthesis.

Significance. If the reported results hold, the paper makes a useful contribution by demonstrating a concrete mechanism for coupling an explicit Gaussian renderer with an implicit dynamic SDF, and by showing that geometry-aware depth supervision can improve dynamic reconstruction without sacrificing rendering quality. The evaluation against external ground-truth meshes on Dg-mesh is a strength, and the ablation study in Table 3 gives initial evidence for the contribution of filtered depth and normal supervision. However, the central depth-filtering equation is internally inconsistent as printed, several hyperparameters are unreported, and all quantitative claims rest on single-run point estimates, so the current manuscript does not yet substantiate the headline state-of-the-art claim in a reproducible way.

major comments (3)
  1. [§4.1, Eq. (5)] The depth-filtering rule as printed is internally inconsistent with its stated purpose. In the accepted branch, d_f = (d_alpha - d_m)/2, and the acceptance condition |d_alpha - d_m| < tau_f forces this quantity to be near zero for every accepted ray. Substituting this d_f into Eq. (6) yields L_sdf = sum ||F(H(o + d_f v, t))||_1 evaluated essentially at the camera origin, not on the object surface. This contradicts Fig. 3, where filtered points lie on the reconstructed surface, and it contradicts the ablation in Table 3, which attributes a large CD improvement (0.746 to 0.502) to filtered depth. If the intended operation is the midpoint (d_alpha + d_m)/2, the printed minus sign is a typo; however, no code or supplementary derivation is provided, so the central depth-supervision mechanism is not reproducible from the text as written. Please correct Eq. (5), report the value of tau_f, and clarify the behavior of the loss for rejected rays.
  2. [§5.1 and Tables 1-3] All quantitative results are point estimates without error bars, seeds, or a statement of the number of runs. The headline claim in §5.2 that the method is 'unique in offering consistently superior performance' rests on differences that are sometimes very small (e.g., Duck CD 0.773 vs. 0.790 and EMD 0.046 vs. 0.047; several D-NeRF PSNR values are below D3DGS, e.g., Mutant 41.47 vs. 42.63 and Hook 36.34 vs. 37.42). Without multiple runs or a variance estimate, the claimed state-of-the-art status is not established at the reported precision. Reporting mean and standard deviation over at least three seeds, or otherwise justifying that the observed differences exceed run-to-run noise, is needed for the central claim.
  3. [§5.1 and Eqs. (8), (9), (12)] Several hyperparameters that determine the method are not reported: tau_f in Eq. (5), w_g, w_p, tau_g, and tau_p in Eqs. (8)-(9), and lambda_sdf, lambda_nn, and lambda_eik in Eq. (12). Only s and tau_d are given in §5.1. These parameters control the surface-aware density control and the SDF/normal/eikonal losses, which are central to the hybrid interaction. The manuscript should provide their values or a clear pointer to released code so that the experiments can be reproduced.
minor comments (5)
  1. [§5.1, Implementations] The training schedule is stated twice with conflicting numbers: 'warm-up phase (0 to 10k iterations) followed by joint training (10k to 40k iterations)' and later 'warm-up phase (0 to 15k iterations) followed by joint training (15k to 40k iterations)'. Please reconcile these two descriptions.
  2. [§4.1, Normal supervision] The text says 'Marigold [24, 38]', but [24] is a monocular depth-estimation paper and [38] is a diffusion-fine-tuning paper; the source of the pseudo-normal maps is unclear. Please cite the correct normal-prediction model or clarify how the depth model is used to produce normals.
  3. [§3.2, Eqs. (2)-(3)] The notation d(x) is introduced in the sentence after Eq. (2) but is not used in the equations; please align the notation so that the SDF is consistently denoted.
  4. [§4.1, Eq. (4)] The alpha-depth formula should be written more conventionally; as printed, the denominator is the sum of weights, but the expression can be simplified and should be checked to confirm that it corresponds to the intended weighted average depth.
  5. [Figure 3 caption] The caption says 'images from left to right are 3D point clouds projected from alpha-blending depth, median depth, and filtered depth' but does not explicitly identify subfigures (a)-(d) or the RGB image; please clarify the correspondence.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the DGS-DNS loop is a joint optimization with external image/normal losses and is judged on external ground-truth meshes; self-citations are minor and non-load-bearing.

full rationale

The paper's central claim—state-of-the-art 3D reconstruction with competitive rendering—is evaluated against external ground-truth meshes (Dg-mesh) and public benchmarks (D-NeRF, Nerfies), not derived from its own equations. The DGS-to-DNS depth supervision and DNS-to-DGS density control form a feedback loop, but each module is also optimized with image reconstruction losses and external monocular normal priors from a foundation model (Eqs. (7), (10)-(13)), and the final geometry is an empirical output rather than a quantity defined as its own training signal. Mutual bootstrapping in a jointly optimized system is not circular by the definitions used here. The paper contains self-citations: [56] is a related-work survey involving one author, and [69] is cited alongside [49] for a standard angular+L1 normal loss whose formula is written out in Eq. (7). Neither citation is load-bearing, and no uniqueness theorem or core premise is imported from the authors' own prior work. The reader-flagged issue in Eq. (5) is a reproducibility/correctness concern—as printed, (d_alpha-d_m)/2 is approximately zero whenever the filtering condition holds—but it is not a circular reduction: the depth-filtering step is not a fitted parameter renamed as a prediction. Similarly, the paper's stated limitations (speed bottleneck, memory footprint) are about efficiency, not circularity. I therefore report no substantive circularity; the score of 2 reflects only the presence of minor, non-load-bearing self-citations.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The method rests on several unproved modeling choices: a homeomorphic deformation field, Marigold normal predictions as pseudo-ground truth, DGS depth maps as surface hints after a hand-made filter, and the SDF/Eikonal surface representation. No new physical entities are introduced. The main free parameters are the depth-filter threshold and loss/density-control weights, several of which are not reported.

free parameters (6)
  • tau_f (depth filter threshold) = not reported in main text
    Used in Eq. (5) to decide when alpha-blended and median depths are close enough to be trusted; controls which pixels enter the SDF depth loss and directly shapes the reconstructed geometry.
  • tau_d (median depth transmittance threshold) = 0.6
    Set in Section 4.1 for computing the median depth from Gaussian transmittance; hand-chosen.
  • s (ray sampling scaling factor) = 3 then 1
    Section 5.1: s=3 from 10k to 20k iterations for coarse search, then s=1; determines how far ray sampling spreads around the DGS depth estimate.
  • w_g and w_p (density-control weights) = not reported
    Weights in Eqs. (8) and (9) balancing gradient magnitude vs SDF distance for Gaussian growth and pruning; their values are not specified, so the exact density schedule is not reproducible.
  • tau_g and tau_p (growth and prune thresholds) = not reported
    Thresholds for adding or removing Gaussians in surface-aware density control; without values, the geometry guidance in the DGS module cannot be replicated.
  • lambda_sdf, lambda_nn, lambda_eik = not reported
    Weights in Eq. (12) for the SDF depth loss, normal loss, and Eikonal regularization; only lambda_I=0.8 and lambda_gn=0.1 are reported in the text.
assumptions (4)
  • domain assumption The deformation field H is a homeomorphic (continuous and bijective) mapping between observation space and canonical space, as defined in Eq. (3).
    Assumed in Section 3.2 to allow cycle-consistent mapping between frames. A homeomorphism preserves topology, so the model cannot represent surfaces that change genus, yet Section 5.2 claims Fig. 5 demonstrates handling of topological changes.
  • domain assumption Marigold monocular normal predictions are accurate enough to supervise both modules.
    Section 4.1 and 4.2 use the foundation model as pseudo-ground truth in the normal losses of Eqs. (7) and (10); no uncertainty, failure cases, or sensitivity analysis are provided.
  • domain assumption DGS-rendered alpha-blended and median depth maps are close enough to the true surface depth after the filter in Eq. (5) to supervise the SDF and bound ray sampling.
    Section 4.1 relies on these depths for efficient ray sampling and for the SDF depth loss in Eq. (6); the filter is the only mechanism guarding against floaters, and its printed equation is inconsistent.
  • standard math A signed distance function regularized by an Eikonal loss is a valid surface representation for dynamic reconstruction.
    The zero-level set of the SDF defines the surface in Eq. (2), a standard representation inherited from NeuS/NDR; the paper assumes this without additional geometric priors.

how reviews work

0 comments
Cite this review

Pith. "Pith review of DGNS: Deformable Gaussian Splatting and Dynamic Neural Surface for Monocular Dynamic 3D Reconstruction." pith.science (2026). https://pith.science/paper/DAQII2HG

@misc{pith2026241203910,
  author       = {Pith},
  title        = {Pith review of: DGNS: Deformable Gaussian Splatting and Dynamic Neural Surface for Monocular Dynamic 3D Reconstruction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DAQII2HG}},
  note         = {Machine review of arXiv:2412.03910}
}
read the original abstract

Dynamic scene reconstruction from monocular video is essential for real-world applications. We introduce DGNS, a hybrid framework integrating \underline{D}eformable \underline{G}aussian Splatting and Dynamic \underline{N}eural \underline{S}urfaces, effectively addressing dynamic novel-view synthesis and 3D geometry reconstruction simultaneously. During training, depth maps generated by the deformable Gaussian splatting module guide the ray sampling for faster processing and provide depth supervision within the dynamic neural surface module to improve geometry reconstruction. Conversely, the dynamic neural surface directs the distribution of Gaussian primitives around the surface, enhancing rendering quality. In addition, we propose a depth-filtering approach to further refine depth supervision. Extensive experiments conducted on public datasets demonstrate that DGNS achieves state-of-the-art performance in 3D reconstruction, along with competitive results in novel-view synthesis.

Figures

Figures reproduced from arXiv: 2412.03910 by the authors.

Figure 1
Figure 1. Performance comparison between different meth [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The framework of DGNS consists of two primary modules: the top module, DNS, for 3D reconstruction, and the [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Visualization of the depth map in 3D space. The [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Qualitative comparison on the Dg-mesh dataset. The samples, from top to bottom, are [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Qualitative results demonstrating the temporal evo [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Qualitative comparison on the D-NeRF dataset. The samples from right to left are [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Qualitative comparison on the Nerfies dataset. [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Demonstration of surface-aware density control. [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Viewpoint Consistency in 3D Generation via Structure Feature and CLIP Guidance

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A tuning-free combination of cross-attention control, CLIP-based pruning, and staged prompts lowers the Janus Problem rate in text-to-3D generation from about 80 percent to about 30 percent.

Reference graph

Works this paper leans on

76 extracted references · 39 canonical work pages · cited by 1 Pith paper

  1. [1]

    Kara-Ali Aliev, Artem Sevastopolsky, Maria Kolos, Dmitry Ulyanov, and Victor Lempitsky. 2020. Neural point-based graphics. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXII

  2. [2]

    Barron, Ben Mildenhall, Dor Verbin, Pratul P

    Jonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hedman. 2022. Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 5460–5469

  3. [3]

    Hongrui Cai, Wanquan Feng, Xuetao Feng, Yan Wang, and Juyong Zhang. 2022. Neural surface reconstruction of dynamic scenes with monocular rgb-d camera. Advances in Neural Information Processing Systems 35 (2022), 967–981

  4. [4]

    Weiwei Cai, Weicai Ye, Peng Ye, Tong He, and Tao Chen. 2024. DynaSurfGS: Dynamic Surface Reconstruction with Planar-based Gaussian Splatting. arXiv preprint arXiv:2408.13972 (2024)

  5. [5]

    Ang Cao and Justin Johnson. 2023. Hexplane: A fast representation for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 130–141

  6. [6]

    David Casillas-Perez, Daniel Pizarro, David Fuentes-Jimenez, Manuel Mazo, and Adrien Bartoli. 2021. The isowarp: the template-based visual geometry of isomet- ric surfaces. International Journal of Computer Vision 129, 7 (2021), 2194–2222

  7. [7]

    Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. 2022. Tensorf: Tensorial radiance fields. In European conference on computer vision . Springer, 333–350

  8. [8]

    Hanlin Chen, Chen Li, and Gim Hee Lee. 2023. Neusg: Neural implicit surface re- construction with 3d gaussian splatting guidance. arXiv preprint arXiv:2312.00846 (2023)

Show all 76 references
  1. [9]

    Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. 2016. Density estimation using real nvp. arXiv preprint arXiv:1605.08803 (2016)

  2. [10]

    Jiemin Fang, Taoran Yi, Xinggang Wang, Lingxi Xie, Xiaopeng Zhang, Wenyu Liu, Matthias Nießner, and Qi Tian. 2022. Fast dynamic radiance fields with time-aware neural voxels. In SIGGRAPH Asia 2022 Conference Papers . 1–9

  3. [11]

    Yao Feng, Haiwen Feng, Michael J Black, and Timo Bolkart. 2021. Learning an animatable detailed 3D face model from in-the-wild images. ACM Transactions on Graphics (ToG) 40, 4 (2021), 1–13

  4. [12]

    Sara Fridovich-Keil, Giacomo Meanti, Frederik Rahbæk Warburg, Benjamin Recht, and Angjoo Kanazawa. 2023. K-planes: Explicit radiance fields in space, time, and appearance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12479–12488

  5. [13]

    Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. 2022. Plenoxels: Radiance fields without neural net- works. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 5501–5510

  6. [14]

    Wanshui Gan, Hongbin Xu, Yi Huang, Shifeng Chen, and Naoto Yokoya. 2023. V4d: Voxel for 4d novel view synthesis. IEEE Transactions on Visualization and Computer Graphics (2023)

  7. [15]

    Chen Gao, Ayush Saraf, Johannes Kopf, and Jia-Bin Huang. 2021. Dynamic view synthesis from dynamic monocular video. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 5712–5721

  8. [16]

    Philip-William Grassal, Malte Prinzler, Titus Leistner, Carsten Rother, Matthias Nießner, and Justus Thies. 2022. Neural head avatars from monocular rgb videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition. 18653–18664

  9. [17]

    Amos Gropp, Lior Yariv, Niv Haim, Matan Atzmon, and Yaron Lipman. 2020. Im- plicit geometric regularization for learning shapes.arXiv preprint arXiv:2002.10099 (2020)

  10. [18]

    Antoine Guédon and Vincent Lepetit. 2024. Sugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5354–5363

  11. [19]

    Shuai Guo, Qiuwen Wang, Yijie Gao, Rong Xie, Lin Li, Fang Zhu, and Li Song

  12. [20]

    Zhiyang Guo, Wengang Zhou, Li Li, Min Wang, and Houqiang Li. 2024. Motion- aware 3d gaussian splatting for efficient dynamic scene reconstruction. IEEE Transactions on Circuits and Systems for Video Technology (2024)

  13. [21]

    Tao Hu, Shu Liu, Yilun Chen, Tiancheng Shen, and Jiaya Jia. 2022. EfficientNeRF: Efficient Neural Radiance Fields. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022), 12892–12901

  14. [22]

    Navami Kairanda, Edith Tretschk, Mohamed Elgharib, Christian Theobalt, and Vladislav Golyanik. 2022. f-sft: Shape-from-template with a physics-based defor- mation model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3948–3958

  15. [23]

    Kai Katsumata, Duc Minh Vo, and Hideki Nakayama. 2023. An efficient 3d gaussian representation for monocular/multi-view dynamic scenes.arXiv preprint arXiv:2311.12897 (2023)

  16. [24]

    Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Metzger, Rodrigo Caye Daudt, and Konrad Schindler. 2024. Repurposing Diffusion-Based Image Genera- tors for Monocular Depth Estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

  17. [25]

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis

  18. [26]

    Agelos Kratimenos, Jiahui Lei, and Kostas Daniilidis. 2023. Dynmf: Neural motion factorization for real-time dynamic view synthesis with 3d gaussian splatting. arXiv preprint arXiv:2312.00112 (2023)

  19. [27]

    Jiahui Lei, Yijia Weng, Adam Harley, Leonidas Guibas, and Kostas Daniilidis. 2024. MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds. arXiv preprint arXiv:2405.17421 (2024)

  20. [28]

    Tianye Li, Mira Slavcheva, Michael Zollhoefer, Simon Green, Christoph Lass- ner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard Newcombe, et al. 2022. Neural 3d video synthesis from multi-view video. In Pro- ceedings of the IEEE/CVF Conference on Compu...

  21. [29]

    Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang. 2021. Neural scene flow fields for space-time view synthesis of dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6498–6508

  22. [30]

    Wenbin Lin, Chengwei Zheng, Jun-Hai Yong, and Feng Xu. 2022. Occlusionfusion: Occlusion-aware motion estimation for real-time dynamic 3d reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1736–1745

  23. [31]

    Isabella Liu, Hao Su, and Xiaolong Wang. 2024. Dynamic Gaussians Mesh: Consis- tent Mesh Reconstruction from Monocular Videos.arXiv preprint arXiv:2404.12379 (2024)

  24. [32]

    Shichen Liu, Tianye Li, Weikai Chen, and Hao Li. 2019. Soft rasterizer: A differ- entiable renderer for image-based 3d reasoning. In Proceedings of the IEEE/CVF international conference on computer vision . 7708–7717

  25. [33]

    Yu-Lun Liu, Chen Gao, Andreas Meuleman, Hung-Yu Tseng, Ayush Saraf, Changil Kim, Yung-Yu Chuang, Johannes Kopf, and Jia-Bin Huang. 2023. Robust dynamic radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13–23

  26. [34]

    Tao Lu, Mulin Yu, Linning Xu, Yuanbo Xiangli, Limin Wang, Dahua Lin, and Bo Dai. 2024. Scaffold-gs: Structured 3d gaussians for view-adaptive rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 20654–20664

  27. [35]

    Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. 2023. Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis. arXiv preprint arXiv:2308.09713 (2023)

  28. [36]

    Shaojie Ma, Yawei Luo, and Yi Yang. 2024. Reconstructing and Simulating Dynamic 3D Objects with Mesh-adsorbed Gaussian Splatting. arXiv preprint arXiv:2406.01593 (2024)

  29. [37]

    Wei Mao, Richard Hartley, Mathieu Salzmann, et al. 2024. Neural SDF Flow for 3D Reconstruction of Dynamic Scenes. In The Twelfth International Conference on Learning Representations

  30. [38]

    Gonzalo Martin Garcia, Karim Abou Zeid, Christian Schmidt, Daan de Geus, Alexander Hermans, and Bastian Leibe. 2025. Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV)

  31. [39]

    Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2021. Nerf: Representing scenes as neural radiance fields for view synthesis. Commun. ACM 65, 1 (2021), 99–106

  32. [40]

    Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. 2022. In- stant neural graphics primitives with a multiresolution hash encoding. ACM transactions on graphics (TOG) 41, 4 (2022), 1–15

  33. [41]

    Richard A Newcombe, Dieter Fox, and Steven M Seitz. 2015. Dynamicfusion: Reconstruction and tracking of non-rigid scenes in real-time. In Proceedings of the IEEE conference on computer vision and pattern recognition . 343–352

  34. [42]

    Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. 2019. Deepsdf: Learning continuous signed distance functions for shape representation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 165–174

  35. [43]

    Keunhong Park, Utkarsh Sinha, Jonathan T Barron, Sofien Bouaziz, Dan B Gold- man, Steven M Seitz, and Ricardo Martin-Brualla. 2021. Nerfies: Deformable neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 5865–5874

  36. [44]

    Songyou Peng, Chiyu Jiang, Yiyi Liao, Michael Niemeyer, Marc Pollefeys, and Andreas Geiger. 2021. Shape as points: A differentiable poisson solver. Advances in Neural Information Processing Systems 34 (2021), 13032–13044

  37. [45]

    Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer

  38. [46]

    Radu Alexandru Rosu and Sven Behnke. 2023. Permutosdf: Fast multi-view reconstruction with implicit surfaces using permutohedral lattices. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8466– 8475. MM ’25, October 27–31, 2025, Dublin, Ire...

  39. [47]

    Ruizhi Shao, Zerong Zheng, Hanzhang Tu, Boning Liu, Hongwen Zhang, and Yebin Liu. 2023. Tensor4d: Efficient neural 4d decomposition for high-fidelity dynamic reconstruction and rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 16...

  40. [48]

    Miroslava Slavcheva, Maximilian Baust, and Slobodan Ilic. 2018. Sobolevfusion: 3d reconstruction of scenes undergoing free non-rigid motion. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2646–2655

  41. [49]

    Fengrui Tian, Shaoyi Du, and Yueqi Duan. 2023. Mononerf: Learning a gener- alizable dynamic radiance field from monocular videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 17903–17913

  42. [50]

    Jinguang Tong, Xuesong Li, Fahira Afzal Maken, Sundaram Muthu, Lars Pe- tersson, Chuong Nguyen, and Hongdong Li. 2025. GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction. In Proceedings of the Computer Vision and Pattern Recognition Conference . 21547–21557

  43. [51]

    Feng Wang, Zilong Chen, Guokang Wang, Yafei Song, and Huaping Liu. 2024. Masked space-time hash encoding for efficient dynamic scene reconstruction. Advances in Neural Information Processing Systems 36 (2024)

  44. [52]

    Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. 2021. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. arXiv preprint arXiv:2106.10689 (2021)

  45. [53]

    Qianqian Wang, Vickie Ye, Hang Gao, Jake Austin, Zhengqi Li, and Angjoo Kanazawa. 2024. Shape of motion: 4d reconstruction from a single video. arXiv preprint arXiv:2407.13764 (2024)

  46. [54]

    Xiaolong Wang, Allan Jabri, and Alexei A Efros. 2019. Learning correspondence from the cycle-consistency of time. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 2566–2576

  47. [55]

    Zian Wang, Tianchang Shen, Merlin Nimier-David, Nicholas Sharp, Jun Gao, Alexander Keller, Sanja Fidler, Thomas Müller, and Zan Gojcic. 2023. Adaptive shells for efficient neural radiance field rendering.arXiv preprint arXiv:2311.10091 (2023)

  48. [56]

    Wenhui Xiao, Remi Chierchia, Rodrigo Santa Cruz, Xuesong Li, David Ahmedt- Aristizabal, Olivier Salvado, Clinton Fookes, and Leo Lebrat. 2025. Neural Ra- diance Fields for the Real World: A Survey. arXiv preprint arXiv:2501.13104 (2025)

  49. [57]

    Linning Xu, Yuanbo Xiangli, Sida Peng, Xingang Pan, Nanxuan Zhao, Christian Theobalt, Bo Dai, and Dahua Lin. 2023. Grid-guided neural radiance fields for large urban scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 8296–8306

  50. [58]

    Qiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi, Zhixin Shu, Kalyan Sunkavalli, and Ulrich Neumann. 2022. Point-nerf: Point-based neural radiance fields. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5438–5448

  51. [59]

    Gengshan Yang, Deqing Sun, Varun Jampani, Daniel Vlasic, Forrester Cole, Hui- wen Chang, Deva Ramanan, William T Freeman, and Ce Liu. 2021. Lasr: Learning articulated shape reconstruction from a monocular video. In Proceedings of the IEEE/CVF Conference on Computer Vision and ...

  52. [60]

    Gengshan Yang, Deqing Sun, Varun Jampani, Daniel Vlasic, Forrester Cole, Ce Liu, and Deva Ramanan. 2021. Viser: Video-specific surface embeddings for articulated 3d shape reconstruction. Advances in Neural Information Processing Systems 34 (2021), 19326–19338

  53. [61]

    Gengshan Yang, Minh Vo, Natalia Neverova, Deva Ramanan, Andrea Vedaldi, and Hanbyul Joo. 2022. Banmo: Building animatable 3d neural models from many casual videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2863–2873

  54. [62]

    Gengshan Yang, Shuo Yang, John Z Zhang, Zachary Manchester, and Deva Ra- manan. 2023. Ppr: Physically plausible reconstruction from monocular videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 3914– 3924

  55. [63]

    Ziyi Yang, Xinyu Gao, Wen Zhou, Shaohui Jiao, Yuqing Zhang, and Xiaogang Jin. 2024. Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 20331–20341

  56. [64]

    Zeyu Yang, Hongye Yang, Zijie Pan, Xiatian Zhu, and Li Zhang. 2023. Real-time photorealistic dynamic scene representation and rendering with 4d gaussian splatting. arXiv preprint arXiv:2310.10642 (2023)

  57. [65]

    Chongjie Ye, Yinyu Nie, Jiahao Chang, Yuantao Chen, Yihao Zhi, and Xiaoguang Han. 2024. GauStudio: A Modular Framework for 3D Gaussian Splatting and Beyond. arXiv preprint arXiv:2403.19632 (2024)

  58. [66]

    Wang Yifan, Felice Serena, Shihao Wu, Cengiz Öztireli, and Olga Sorkine- Hornung. 2019. Differentiable surface splatting for point-based geometry pro- cessing. ACM Transactions on Graphics (TOG) 38, 6 (2019), 1–14

  59. [67]

    Alex Yu, Ruilong Li, Matthew Tancik, Hao Li, Ren Ng, and Angjoo Kanazawa. 2021. PlenOctrees for Real-time Rendering of Neural Radiance Fields. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) . 5732–5741

  60. [68]

    Mulin Yu, Tao Lu, Linning Xu, Lihan Jiang, Yuanbo Xiangli, and Bo Dai. 2024. Gsdf: 3dgs meets sdf for improved rendering and reconstruction. arXiv preprint arXiv:2403.16964 (2024)

  61. [69]

    Chushan Zhang, Jinguang Tong, Tao Jun Lin, Chuong Nguyen, and Hongdong Li

  62. [70]

    Wenyuan Zhang, Yu-Shen Liu, and Zhizhong Han. 2024. Neural signed distance function inference through splatting 3d gaussians pulled on zero-level set. arXiv preprint arXiv:2410.14189 (2024)

  63. [71]

    Michael Zollhöfer, Justus Thies, Pablo Garrido, Derek Bradley, Thabo Beeler, Patrick Pérez, Marc Stamminger, Matthias Nießner, and Christian Theobalt. 2018. State of the art on monocular 3D face reconstruction, tracking, and applications. In Computer graphics forum, Vol. 37. W...

  64. [72]

    Silvia Zuffi, Angjoo Kanazawa, David W Jacobs, and Michael J Black. 2017. 3D menagerie: Modeling the 3D shape and pose of animals. In Proceedings of the IEEE conference on computer vision and pattern recognition . 6365–6373

  65. [73]

    In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision

    PMVC: Promoting Multi-View Consistency for 3D Scene Reconstruction. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 3678–3688

  66. [2021]

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    D-nerf: Neural radiance fields for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10318–10327

  67. [2023]

    ACM Trans

    3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Trans. Graph. 42, 4 (2023), 139–1

  68. [2024]

    IEEE Transactions on Circuits and Systems for Video Technology (2024)

    Depth-guided robust point cloud fusion NeRF for sparse input views. IEEE Transactions on Circuits and Systems for Video Technology (2024)

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.