REVIEW 2 major objections 4 minor 88 references
4D Gaussian Splatting with Scale-aware Residual Field and Adaptive Optimization for Real-time Rendering of Temporally Complex Dynamic Scenes
T0 review · 2 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read SaRO-GS reconstructs temporally complex dynamic scenes in real time by giving each 4D Gaussian its own lifespan and a scale-aware residual field.
desk verdict Solid incremental 4DGS extension with real-time gains, but the headline numbers come from an unfair comparison and the adaptive optimization equation is wrong as written. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery has four interlocking parts. Each 4D Gaussian owns a 4D position \((x,y,z,\tau_i)\), a lifespan \(\sigma_i\), and a state function \(\gamma_i(t) = \exp\left(-4\left(\frac{t-\tau_i}{\sigma_i}\right)^2\right)\), which drives the projected 3D opacity and doubles as the primitive's temporal sampling probability. The Scale-aware Residual Field is a hexplane field whose spatial-only planes are MipMap stacks: a Gaussian's projected ellipse selects the level \(l = \min\left(\log_2(s_x/\hat{s}^0_x), \log_2(s_y/\hat{s}^0_y)\right)\) for trilinear interpolation, aligning residual features with the primitive's actual extent and its self-splitting behavior. A small MLP decodes the residual feature into per-time residuals of position, covariance, and color, plus the lifespan. The Adaptive Optimization Schedule approximates each primitive's time-domain integral \(I_i\) with the logistic CDF and sets the densification threshold \(\kappa_i = \kappa_{\text{base}} \cdot I_i / I_{\max}\) and learning rate \(\text{lr}_i = \text{lr}_{\text{base}} \cdot I_{\max} / I_i\), so dynamic primitives with short lifespans are densified more easily and optimized more aggressively.
What would settle it
Render a synthetic scene where an object disappears and later reappears in the same place with the same appearance. If the single-peak state function is the limiting factor, the reappearance should show ghosting or blur, and the per-scene PSNR should drop well below the paper's reported D-NeRF average unless the optimizer happens to split the object into multiple primitives with different temporal centers.
Extended reading notes
Core claim
In the paper's own terms, the central discovery is that giving each 4D Gaussian an explicit temporal center \(\tau_i\) and lifespan \(\sigma_i\), and making its visibility a symmetric Gaussian function of time, is enough to represent object appearance and disappearance without any deformation field. The Scale-aware Residual Field makes the encoding scale-aware by sampling the spatial-only hexplanes at a MipMap level matched to each Gaussian's projected footprint, so split children inherit features similar to their parent. Finally, the paper shows that optimization can be balanced per primitive: integrating each Gaussian's visibility over the observed time range yields a sampling probability that scales its densification threshold and learning rate. Together these mechanisms let the method render temporally complex scenes in real time while improving reconstruction quality over prior work.
Load-bearing premise
Every Gaussian's visibility over time is a single symmetric bump with a fixed shape, so one primitive cannot represent objects that appear abruptly, vanish, or show up in disjoint time intervals.
Editorial extensions
If this is right
- Temporal appearance and disappearance no longer require a deformation field or a canonical frame; each primitive simply switches on and off around its own learned lifespan.
- Because lifespans are learned per primitive, the model can segment dynamic from static parts of a scene with no external supervision.
- The method works for both monocular and multi-view inputs while preserving real-time rendering in both settings.
- Reported frame rates are roughly two orders of magnitude faster than NeRF-based methods on the same benchmarks while improving PSNR.
Reading between the lines
- The per-primitive lifespan could be reused as a handle for temporal editing, such as deleting an object from the whole sequence or shifting when it appears, without re-training the model.
- The single-peak Gaussian visibility is the main constraint: an object that disappears and reappears at the same location would need several Gaussians or a learned multi-peak state function to be captured cleanly.
- The MipMap-level trick for sampling grid features by primitive footprint is general and could be transferred to other grid-based fields whose samples are regions rather than points.
- A direct stress test would be a scene with periodic or repeated appearances (for example, a rotating fan blade or a blinking light): quality should degrade relative to scenes with a single appearance event, tracing exactly to the single-peak assumption.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SaRO-GS, a dynamic scene representation based on 4D Gaussian primitives with a scale-aware residual field and a per-primitive adaptive optimization schedule. Each 4D Gaussian carries a temporal position and a lifespan; at query time a Gaussian-like state function and an MLP-decoded residual feature project the primitive into 3D space for rendering with 3D Gaussian Splatting. The scale-aware residual field uses hexplanes with a MipMap stack for spatial-only planes so that feature lookup accounts for the ellipsoidal footprint of each Gaussian. The adaptive schedule scales learning rate and densification threshold by the temporal integral of the visibility function. Experiments on D-NeRF and Plenoptic Video report state-of-the-art PSNR with real-time rendering, and an ablation study attributes gains to each proposed component.
Significance. If the reported results hold, the method is a practically interesting contribution: it combines real-time rendering with explicit handling of object appearance and disappearance, and it provides dynamic-static segmentation as a byproduct. The ablations in Table 4 support the value of the scale-aware field, temporal properties, adaptive optimization, and residual regularization. The paper is also clearly written and the method is plausible. However, two issues presently weaken the central claims: the CDF approximation used in the adaptive optimization is not a valid CDF as written, and the main tables compare methods at different resolutions and over different scene subsets. Both are fixable in revision, but they must be corrected before the state-of-the-art and ablation claims can be accepted as stated.
major comments (2)
- [Sec. 4.3, Eq. (21); Appendix B, Eq. (27)] The approximation Q(t) = 1 - 1/e^{1+alpha1*t^3+alpha2*t} is not the standard normal CDF and is not the Page approximation cited as [35]. At t=0 it gives 1-1/e = 0.632 instead of 0.5, and for negative t it can become negative (for t=-1 it is approximately -4.3). Since Eq. (23) sets lr_i = lr_base * I_max / I_i and kappa_i = kappa_base * I_i / I_max, a Gaussian with tau_i near the observation boundary can yield a negative or hugely negative Q(u_start), making I_i effectively infinite and reversing the intended optimization schedule: such a dynamic primitive would be treated as static. The 0.69 dB gain attributed to adaptive optimization in Table 4 is therefore not reproducible from the equations as written. The correct form of the Page approximation is Q(t) = 1 - 1/(1 + e^{alpha1*t^3 + alpha2*t}) for t >= 0 with the symmetry relation Q(-t) = 1 - Q(t); the authors should correct Eqs. (21) and (27) accordingly and verify that the ablation numbers remain unchanged.
- [Tables 1, 2, 3 and D2] The evaluation protocol is not consistent across methods. In Table 1, 4DGS is evaluated at 800x800 while Ours and the remaining methods are evaluated at 400x400; Table 2 then shows that 4DGS at 400x400 reaches 35.05 dB, reducing the reported PSNR gap from 2.08 dB to 1.08 dB. In Table 3, the HexPlane average excludes the Coffee Martini scene while the Ours average includes it, as confirmed by the per-scene Table D2, so the two averages are not computed over the same set of scenes. The central state-of-the-art claim should be based on identical resolution and identical scene subsets for all methods, with any differently-configured numbers clearly separated. The additional numbers already present in Table 2 and Table D2 should be used to present a fully consistent comparison.
minor comments (4)
- [Sec. 4.1] The sentence 'Then, based on Eq. 34, we employ 3DGS to render 3D Gaussians' refers to a nonexistent equation; the rendering equation is Eq. (4) in Sec. 3.1.
- [Table 3 and Sec. 6.2] The table footnotes are hard to parse: '1: excludes the Coffee Martini scene' and '2: Only report SSIM instead of MS-SSIM like others' should specify exactly which per-scene values are included in each average and which metric variant was used for each method.
- [Fig. 1 and Sec. 6.1] The caption of Fig. 1 contains garbled characters and should be re-typeset; also, the figure's speed-quality plot would benefit from a legend entry for '*' clarifying that 4DGS was re-measured at 400x400.
- [Eq. (18)] In the definition of the scale-aware feature, the subscript in pi_{x,y} should be pi_{i,j} to match the summation over C_so = {(x,y),(x,z),(y,z)}; the current notation is inconsistent.
Circularity Check
No circularity: the core pipeline is a self-consistent end-to-end optimization against reconstruction loss, benchmarked externally; the flagged CDF issue is a correctness typo, not a circular step.
full rationale
SaRO-GS's derivation chain is self-contained. The Scale-aware Residual Field (Eqs. 6-18) and the per-Gaussian temporal state function (Eqs. 9-10) are design choices; the residual field is optimized jointly with the Gaussian attributes against the image reconstruction loss (Eq. 25), and no claim is made that residuals are independently predicted from fitted inputs. The Adaptive Optimization schedule (Eqs. 19-23) computes a temporal integral I_i from the same learned sigma_i and tau_i and uses it to adjust the learning rate and densification threshold; while this is a feedback loop, it is a training heuristic, not a 'prediction' that reduces to its inputs. The dynamic-static segmentation (Sec. 6.5) is a post-hoc threshold on the learned lifespan, not a circularly validated output. The Gaussian state function is explicitly borrowed from Spacetime-GS [27], and the CDF approximation constants are taken from Page [35]; neither is a self-citation, and both are externally sourced components. Self-citations in the reference list appear only in Related Work and are not load-bearing for the SOTA claims, which are supported by D-NeRF and Plenoptic Video benchmarks. Separate from circularity: Eq. 21/27 as printed gives Q(0) ≈ 0.632 rather than 0.5, indicating a likely typo in the Page approximation; this is a reproducibility/correctness concern, not a circularity one.
Assumptions & free parameters
free parameters (4)
- k (state function exponent) =
4
- Loss weights lambda1, lambda2 =
0.2, 0.8
- CDF approximation constants alpha1, alpha2 =
0.070565992, 1.5976
- Initial Gaussian count =
10,000 (D-NeRF), 40,000 (Plenoptic)
assumptions (5)
- domain assumption Gaussian visibility over time is unimodal and symmetric, modeled by gamma(t) = exp(-k((t-tau)/sigma)^2).
- domain assumption The residual feature f(G4D) from mipmapped hexplanes is sufficient to decode position, covariance, and color residuals via a small MLP.
- standard math The standard 3DGS volume rendering formula applies to the projected 3D Gaussians at each time t0.
- standard math The CDF approximation of Page (1977) is accurate for the integration range used.
- ad hoc to paper Summation of spatial-only and spatiotemporal features is effective.
Cite this review
Pith. "Pith review of 4D Gaussian Splatting with Scale-aware Residual Field and Adaptive Optimization for Real-time Rendering of Temporally Complex Dynamic Scenes." pith.science (2026). https://pith.science/paper/3XKTQVCT
@misc{pith2026241206299,
author = {Pith},
title = {Pith review of: 4D Gaussian Splatting with Scale-aware Residual Field and Adaptive Optimization for Real-time Rendering of Temporally Complex Dynamic Scenes},
year = {2026},
howpublished = {\url{https://pith.science/paper/3XKTQVCT}},
note = {Machine review of arXiv:2412.06299}
}
read the original abstract
Reconstructing dynamic scenes from video sequences is a highly promising task in the multimedia domain. While previous methods have made progress, they often struggle with slow rendering and managing temporal complexities such as significant motion and object appearance/disappearance. In this paper, we propose SaRO-GS as a novel dynamic scene representation capable of achieving real-time rendering while effectively handling temporal complexities in dynamic scenes. To address the issue of slow rendering speed, we adopt a Gaussian primitive-based representation and optimize the Gaussians in 4D space, which facilitates real-time rendering with the assistance of 3D Gaussian Splatting. Additionally, to handle temporally complex dynamic scenes, we introduce a Scale-aware Residual Field. This field considers the size information of each Gaussian primitive while encoding its residual feature and aligns with the self-splitting behavior of Gaussian primitives. Furthermore, we propose an Adaptive Optimization Schedule, which assigns different optimization strategies to Gaussian primitives based on their distinct temporal properties, thereby expediting the reconstruction of dynamic regions. Through evaluations on monocular and multi-view datasets, our method has demonstrated state-of-the-art performance. Please see our project page at https://yjb6.github.io/SaRO-GS.github.io.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[35]
E. Page. 2018. Approximations to the Cumulative Normal Function and its Inverse for Use on a Pocket Calculator. Journal of the Royal Statistical Society Series C: Applied Statistics 26, 1 (2018), 75–76. https://doi.org/10.2307/2346872
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[2]
Benjamin Attal, Jia-Bin Huang, Christian Richardt, Michael Zollhoefer, Johannes Kopf, Matthew O’Toole, and Changil Kim. 2023. HyperReel: High-fidelity 6-DoF video with ray-conditioned sampling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 16610–16620
2023
-
[3]
Yanqi Bao, Yuxin Li, Jing Huo, Tianyu Ding, Xinyue Liang, Wenbin Li, and Yang Gao. 2023. Where and how: Mitigating confusion in neural radiance fields from sparse inputs. In Proceedings of the 31st ACM International Conference on Multimedia. 2180–2188
2023
-
[4]
Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. 2021. Mip-nerf: A multiscale repre- sentation for anti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 5855–5864
2021
-
[5]
Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. 2022. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5470–5479
2022
-
[6]
Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Pe- ter Hedman. 2023. Zip-nerf: Anti-aliased grid-based neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 19697– 19705
2023
-
[7]
Ang Cao and Justin Johnson. 2023. Hexplane: A fast representation for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 130–141
2023
Show all 88 references
-
[8]
Zhang Chen, Zhong Li, Liangchen Song, Lele Chen, Jingyi Yu, Junsong Yuan, and Yi Xu. 2023. Neurbf: A neural fields representation with adaptive radial basis functions. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 4182–4194
2023
-
[9]
Xuxin Cheng, Bowen Cao, Qichen Ye, Zhihong Zhu, Hongxiang Li, and Yuexian Zou. 2023. Ml-lmcl: Mutual learning and large-margin contrastive learning for improving asr robustness in spoken language understanding. arXiv preprint arXiv:2311.11375 (2023)
2023 arXiv
-
[10]
Xuxin Cheng, Zhihong Zhu, Bowen Cao, Qichen Ye, and Yuexian Zou. 2023. Mrrl: Modifying the reference via reinforcement learning for non-autoregressive joint multiple intent detection and slot filling. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 1...
2023
-
[11]
David Eigen, Christian Puhrsch, and Rob Fergus. 2014. Depth map prediction from a single image using a multi-scale deep network. Advances in neural information processing systems 27 (2014)
2014
-
[12]
Jiemin Fang, Taoran Yi, Xinggang Wang, Lingxi Xie, Xiaopeng Zhang, Wenyu Liu, Matthias Nießner, and Qi Tian. 2022. Fast dynamic radiance fields with time-aware neural voxels. In SIGGRAPH Asia 2022 Conference Papers . 1–9
2022
-
[13]
Sara Fridovich-Keil, Giacomo Meanti, Frederik Rahbæk Warburg, Benjamin Recht, and Angjoo Kanazawa. 2023. K-planes: Explicit radiance fields in space, time, and appearance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12479–12488
2023
-
[14]
Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. 2022. Plenoxels: Radiance fields without neural net- works. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5501–5510
2022
-
[15]
Wanshui Gan, Hongbin Xu, Yi Huang, Shifeng Chen, and Naoto Yokoya. 2023. V4d: Voxel for 4d novel view synthesis. IEEE Transactions on Visualization and Computer Graphics (2023)
2023
-
[16]
Huachen Gao, Xiaoyu Liu, Meixia Qu, and Shijie Huang. 2021. Pdanet: Self- supervised monocular depth estimation using perceptual and data augmentation consistency. Applied Sciences 11, 12 (2021), 5383
2021
-
[17]
Ravi Garg, Vijay Kumar Bg, Gustavo Carneiro, and Ian Reid. 2016. Unsupervised cnn for single view depth estimation: Geometry to the rescue. InComputer Vision– ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part VIII 14 . Spri...
2016
-
[18]
Xiang Guo, Jiadai Sun, Yuchao Dai, Guanying Chen, Xiaoqing Ye, Xiao Tan, Errui Ding, Yumeng Zhang, and Jingdong Wang. 2023. Forward flow for novel view synthesis of dynamic scenes. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 16022–16033
2023
-
[19]
Peter Hedman, Pratul P Srinivasan, Ben Mildenhall, Jonathan T Barron, and Paul Debevec. 2021. Baking neural radiance fields for real-time view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 5875– 5884
2021
-
[20]
Wenbo Hu, Yuling Wang, Lin Ma, Bangbang Yang, Lin Gao, Xiao Liu, and Yuewen Ma. 2023. Tri-miprf: Tri-mip representation for efficient anti-aliasing neural radi- ance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 19774–19783
2023
-
[21]
Yi-Hua Huang, Yang-Tian Sun, Ziyi Yang, Xiaoyang Lyu, Yan-Pei Cao, and Xiao- juan Qi. 2023. SC-GS: Sparse-Controlled Gaussian Splatting for Editable Dynamic Scenes. arXiv preprint arXiv:2312.14937 (2023)
2023 arXiv
-
[22]
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis
-
[23]
Samuli Laine, Janne Hellsten, Tero Karras, Yeongho Seol, Jaakko Lehtinen, and Timo Aila. 2020. Modular primitives for high-performance differentiable render- ing. ACM Transactions on Graphics (ToG) 39, 6 (2020), 1–14
2020
-
[24]
Jiahe Li, Jiawei Zhang, Xiao Bai, Jin Zheng, Xin Ning, Jun Zhou, and Lin Gu
-
[25]
Lingzhi Li, Zhen Shen, Zhongshu Wang, Li Shen, and Ping Tan. 2022. Streaming radiance fields for 3d video synthesis. Advances in Neural Information Processing Systems 35 (2022), 13485–13498
2022
-
[26]
Tianye Li, Mira Slavcheva, Michael Zollhoefer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, and Richard Newcombe. 2022. Neural 3d video synthesis from multi-view video. InProceedings of the IEEE/CVF Conference on Computer Visi...
2022
- [27]
-
[28]
Haotong Lin, Sida Peng, Zhen Xu, Tao Xie, Xingyi He, Hujun Bao, and Xiaowei Zhou. 2023. High-fidelity and real-time novel view synthesis for dynamic scenes. In SIGGRAPH Asia 2023 Conference Papers . 1–9
2023
-
[29]
Jia-Wei Liu, Yan-Pei Cao, Weijia Mao, Wenqiao Zhang, David Junhao Zhang, Jussi Keppo, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. 2022. Devrf: Fast deformable voxel radiance fields for dynamic scenes. Advances in Neural Infor- mation Processing Systems 35 (2022), 36762–36775
2022
-
[30]
David G Lowe. 2004. Distinctive image features from scale-invariant keypoints. International journal of computer vision 60 (2004), 91–110
2004
-
[31]
Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. 2023. Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis. arXiv preprint arXiv:2308.09713 (2023)
2023 arXiv
-
[32]
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2021. Nerf: Representing scenes as neural radiance fields for view synthesis. Commun. ACM 65, 1 (2021), 99–106
2021
-
[33]
Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. 2022. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transac- tions on Graphics 41, 4 (2022), 1–15. https://doi.org/10.1145/3528223.3530127
2022
-
[34]
Michael Niemeyer, Jonathan T Barron, Ben Mildenhall, Mehdi SM Sajjadi, Andreas Geiger, and Noha Radwan. 2022. Regnerf: Regularizing neural radiance fields for view synthesis from sparse inputs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognitio...
2022
-
[36]
Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin-Brualla, and Steven M Seitz. 2021. Hypernerf: A higher-dimensional representation for topologically varying neural radiance MM ’24, October 28-November 1, 2024, Melbour...
2021 arXiv
-
[37]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, and Luca Antiga. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32 (2019)
2019
-
[38]
Rui Peng, Xiaodong Gu, Luyang Tang, Shihe Shen, Fanqi Yu, and Ronggang Wang
-
[39]
Rui Peng, Rongjie Wang, Zhenyu Wang, Yawen Lai, and Ronggang Wang. 2022. Rethinking depth estimation for multi-view stereo: A unified representation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 8645–8654
2022
-
[40]
Sida Peng, Junting Dong, Qianqian Wang, Shangzhan Zhang, Qing Shuai, Xiaowei Zhou, and Hujun Bao. 2021. Animatable neural radiance fields for modeling dynamic human bodies. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 14314–14323
2021
-
[41]
Advances in Neural Information Processing Systems 36 (2023), 56932–56945
Gens: Generalizable neural surface reconstruction from multi-view images. Advances in Neural Information Processing Systems 36 (2023), 56932–56945
2023
-
[42]
Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer
-
[43]
Christian Reiser, Songyou Peng, Yiyi Liao, and Andreas Geiger. 2021. Kilonerf: Speeding up neural radiance fields with thousands of tiny mlps. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 14335–14345
2021
-
[44]
Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. 2021. Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. In Proceedings of the IEEE/CVF Conference on Computer Visi...
2021
-
[45]
Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabi- novich. 2020. Superglue: Learning feature matching with graph neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recogni- tion. 4938–4947
2020
-
[46]
Johannes Lutz Schönberger and Jan-Michael Frahm. 2016. Structure-from-Motion Revisited. In Conference on Computer Vision and Pattern Recognition (CVPR)
2016
-
[47]
Johannes Lutz Schönberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm. 2016. Pixelwise View Selection for Unstructured Multi-View Stereo. In European Conference on Computer Vision (ECCV)
2016
-
[48]
Christian Reiser, Rick Szeliski, Dor Verbin, Pratul Srinivasan, Ben Mildenhall, Andreas Geiger, Jon Barron, and Peter Hedman. 2023. Merf: Memory-efficient ra- diance fields for real-time view synthesis in unbounded scenes.ACM Transactions on Graphics (TOG) 42, 4 (2023), 1–12
2023
-
[49]
Ruizhi Shao, Zerong Zheng, Hanzhang Tu, Boning Liu, Hongwen Zhang, and Yebin Liu. 2023. Tensor4d: Efficient neural 4d decomposition for high-fidelity dynamic reconstruction and rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 16...
2023
-
[50]
Xuelun Shen, Zhipeng Cai, Wei Yin, Matthias Müller, Zijun Li, Kaixuan Wang, Xiaozhi Chen, and Cheng Wang. 2024. GIM: Learning Generalizable Image Matcher From Internet Videos. arXiv preprint arXiv:2402.11095 (2024)
2024 arXiv
-
[51]
Liangchen Song, Anpei Chen, Zhong Li, Zhang Chen, Lele Chen, Junsong Yuan, Yi Xu, and Andreas Geiger. 2023. Nerfplayer: A streamable dynamic scene representation with decomposed neural radiance fields. IEEE Transactions on Visualization and Computer Graphics 29, 5 (2023), 2732–2742
2023
-
[52]
Dmitriy Serdyuk, Yongqiang Wang, Christian Fuegen, Anuj Kumar, Baiyang Liu, and Yoshua Bengio. 2018. Towards end-to-end spoken language understanding. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 5754–5758
2018
- [53]
-
[54]
Chen Wang, Jiadai Sun, Lina Liu, Chenming Wu, Zhelun Shen, Dayan Wu, Yuchao Dai, and Liangjun Zhang. 2023. Digging into depth priors for outdoor neural radi- ance fields. In Proceedings of the 31st ACM International Conference on Multimedia . 1221–1230
2023
-
[55]
Chen Wang, Xian Wu, Yuan-Chen Guo, Song-Hai Zhang, Yu-Wing Tai, and Shi- Min Hu. 2022. Nerf-sr: High quality neural radiance fields using supersampling. In Proceedings of the 30th ACM International Conference on Multimedia . 6445–6454
2022
-
[56]
Cheng Sun, Min Sun, and Hwann-Tzong Chen. 2022. Direct voxel grid optimiza- tion: Super-fast convergence for radiance fields reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5459–5469
2022
-
[57]
Feng Wang, Sinan Tan, Xinghang Li, Zeyue Tian, Yafei Song, and Huaping Liu
-
[58]
Liao Wang, Qiang Hu, Qihan He, Ziyu Wang, Jingyi Yu, Tinne Tuytelaars, Lan Xu, and Minye Wu. 2023. Neural Residual Radiance Fields for Streamably Free- Viewpoint Videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 76–87
2023
-
[59]
Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. 2021. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. arXiv preprint arXiv:2106.10689 (2021)
2021 arXiv
-
[60]
Feng Wang, Zilong Chen, Guokang Wang, Yafei Song, and Huaping Liu. 2024. Masked space-time hash encoding for efficient dynamic scene reconstruction. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[61]
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 2023. 4d gaussian splatting for real-time dynamic scene rendering. arXiv preprint arXiv:2310.08528 (2023)
2023 arXiv
-
[62]
In Proceedings of the IEEE/CVF International Conference on Computer Vision
Mixed neural voxels for fast multi-view video synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 19706–19716
-
[63]
Wangze Xu, Qi Wang, Xinghao Pan, and Ronggang Wang. 2024. HDPNERF: Hybrid Depth Priors for Neural Radiance Fields from Sparse Input Views. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 3695–3699
2024
-
[64]
Ziyi Yang, Xinyu Gao, Wen Zhou, Shaohui Jiao, Yuqing Zhang, and Xiaogang Jin
-
[65]
Chung-Yi Weng, Brian Curless, Pratul P Srinivasan, Jonathan T Barron, and Ira Kemelmacher-Shlizerman. 2022. Humannerf: Free-viewpoint rendering of moving people from monocular video. In Proceedings of the IEEE/CVF conference on computer vision and pattern Recognition . 16210–16220
2022
-
[66]
Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan. 2018. Mvsnet: Depth inference for unstructured multi-view stereo. In Proceedings of the European conference on computer vision (ECCV) . 767–783
2018
-
[67]
Kaiqiang Xiong, Rui Peng, Zhe Zhang, Tianxing Feng, Jianbo Jiao, Feng Gao, and Ronggang Wang. 2023. Cl-MVSNet: Unsupervised multi-view stereo with dual- level contrastive learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 3769–3780
2023
-
[68]
Yuyang Yin, Dejia Xu, Zhangyang Wang, Yao Zhao, and Yunchao Wei. 2023. 4dgen: Grounded 4d content generation with spatial-temporal consistency. arXiv preprint arXiv:2312.17225 (2023)
2023 arXiv
-
[69]
Alex Yu, Ruilong Li, Matthew Tancik, Hao Li, Ren Ng, and Angjoo Kanazawa
- [70]
- [71]
-
[72]
Junyi Zeng, Chong Bao, Rui Chen, Zilong Dong, Guofeng Zhang, Hujun Bao, and Zhaopeng Cui. 2023. Mirror-NeRF: Learning Neural Radiance Fields for Mirrors with Whitted-Style Ray Tracing. , 4606–4615 pages. https://doi.org/10.1145/ 3581783.3611857
2023
-
[73]
Yao Yao, Zixin Luo, Shiwei Li, Tianwei Shen, Tian Fang, and Long Quan. 2019. Recurrent mvsnet for high-resolution multi-view stereo depth inference. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5525–5534
2019
-
[74]
Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. 2020. Nerf++: Analyzing and improving neural radiance fields. arXiv preprint arXiv:2010.07492 (2020)
2020 arXiv
-
[75]
Zhe Zhang, Huachen Gao, Yuxi Hu, and Ronggang Wang. 2023. N2mvsnet: Non-local neighbors aware multi-view stereo network. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5
2023
-
[76]
In Proceedings of the IEEE/CVF International Conference on Computer Vision
Plenoctrees for real-time rendering of neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 5752–5761
-
[77]
Zehao Yu, Anpei Chen, Binbin Huang, Torsten Sattler, and Andreas Geiger. 2023. Mip-splatting: Alias-free 3d gaussian splatting. arXiv preprint arXiv:2311.16493 (2023)
2023 arXiv
-
[78]
Zhengming Yu, Wei Cheng, Xian Liu, Wayne Wu, and Kwan-Yee Lin. 2023. Mono- Human: Animatable Human Neural Field from Monocular Video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 16943– 16953
2023
-
[79]
Xiaoyun Zheng, Liwei Liao, Jianbo Jiao, Feng Gao, and Ronggang Wang. 2024. Surface-sos: Self-supervised object segmentation via neural surface representa- tion. IEEE Transactions on Image Processing (2024)
2024
-
[80]
Jian Zhang, Jinchi Huang, Bowen Cai, Huan Fu, Mingming Gong, Chaohui Wang, Jiaming Wang, Hongchen Luo, Rongfei Jia, and Binqiang Zhao. 2022. Digging into radiance grid for real-time view synthesis with detail preservation. In European Conference on Computer Vision . Springer, 724–740
2022
-
[81]
Zehao Zhu, Zhiwen Fan, Yifan Jiang, and Zhangyang Wang. 2023. FSGS: Real-Time Few-shot View Synthesis using Gaussian Splatting. arXiv preprint arXiv:2312.00451 (2023). A Overview With in the supplemtantary, we provide: • Details of Adaptive Optimization in Sec. B • Hyperparame...
2023 arXiv
-
[83]
Zhe Zhang, Yuxi Hu, Huachen Gao, and Ronggang Wang. 2023. Bi-ClueMVSNet: Learning Bidirectional Occlusion Clues for Multi-View Stereo. In 2023 Interna- tional Joint Conference on Neural Networks (IJCNN) . IEEE, 1–8
2023
-
[84]
Zhe Zhang, Rui Peng, Yuxi Hu, and Ronggang Wang. 2023. Geomvsnet: Learning multi-view stereo with geometry perception. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 21508–21518
2023
-
[85]
Shunyuan Zheng, Boyao Zhou, Ruizhi Shao, Boning Liu, Shengping Zhang, Liqiang Nie, and Yebin Liu. 2023. Gps-gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis. arXiv preprint arXiv:2312.02155 (2023)
2023 arXiv
-
[87]
Xiaoyun Zheng, Liwei Liao, Xufeng Li, Jianbo Jiao, Rongjie Wang, Feng Gao, Shiqi Wang, and Ronggang Wang. 2024. PKU-DyMVHumans: A Multi-View Video Benchmark for High-Fidelity Dynamic Human Modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog...
2024
-
[2021]
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
D-nerf: Neural radiance fields for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10318–10327
-
[2023]
ACM Transac- tions on Graphics (ToG) 42, 4 (2023), 1–14
3d gaussian splatting for real-time radiance field rendering. ACM Transac- tions on Graphics (ToG) 42, 4 (2023), 1–14
2023
-
[2024]
arXiv preprint arXiv:2403.06912 (2024)
DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization. arXiv preprint arXiv:2403.06912 (2024)
2024 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.