REVIEW 4 major objections 6 minor 68 references
RDG-GS: Relative Depth Guidance with Gaussian Splatting for Real-time Sparse-View 3D Rendering
T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read RDG-GS claims sparse-view 3D rendering can reach state-of-the-art quality by replacing absolute monocular depth supervision with relative depth guidance.
desk verdict Promising sparse-view 3DGS method with a genuinely new relative-depth guidance loss, but the evaluation is so internally inconsistent that the SOTA claim cannot be verified from the paper. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the relative depth guidance loss, $L_{rdg} = \sum_{hw,ij} \log(1 + \exp^{-(D^{hw,ij}-b)\max(F^{hw,ij},0)})$, where $D^{hw,ij}$ is the cosine similarity between two patches of rendered depth and $F^{hw,ij}$ is the cosine similarity between the corresponding image-feature patches. This loss steers the Gaussian field toward view-consistent geometry by increasing image-feature similarity when depth similarity exceeds an adaptive threshold $b$ and decreasing it when depth similarity falls below the threshold. Supporting machinery includes an energy-function refinement of coarse monocular depth that injects global and local RGB structure into the depth prior, and an adaptive sampling strategy that densifies regions with high depth training error to accelerate convergence.
What would settle it
Train RDG-GS on a sparse-view scene with large untextured or repeating-pattern regions, such as a white wall with a poster, and measure rendered depth error against held-out views. If the relative depth guidance loss fails to reduce depth error compared with using only refined depth regularization, or actively degrades geometry because DINO features are not depth-ordered in such regions, the central claim that the loss produces view-consistent geometry would be refuted.
Extended reading notes
Core claim
RDG-GS's central claim is that the spatial relationships between patches, rather than absolute depth values, are the reliable geometric signal for sparse-view Gaussian optimization. The paper argues that monocular depth estimates are coarse and view-inconsistent, so supervising Gaussians with absolute depth pushes them toward wrong shapes. Instead, RDG-GS computes cosine similarities between patches of rendered depth and between patches of image features extracted by DINOv2, then aligns these two similarity tensors through a loss with an adaptive bias, encouraging patches that are nearby in depth to be nearby in feature space and distant patches to be pushed apart. Combined with refined depth priors and adaptive densification, the paper reports superior PSNR, SSIM, and LPIPS across Mip-NeRF360, LLFF, DTU, and Blender, while retaining real-time rendering.
Load-bearing premise
The load-bearing premise is that DINOv2 patch-feature similarity computed from rendered depth and images is a faithful proxy for 3D spatial proximity, so aligning image-feature similarity with depth similarity actually pushes Gaussians toward correct shapes; if DINO features do not preserve relative depth ordering, the loss can push geometry in the wrong direction.
Editorial extensions
If this is right
- Sparse-view 3D Gaussian Splatting can achieve strong rendering quality without dense view coverage, with gains persisting from 3 to 24 training views and across multiple resolutions.
- The method keeps rendering real-time at roughly 112 FPS, so the quality improvement does not sacrifice interactivity for applications such as virtual reality and autonomous driving.
- The relative depth guidance is not tied to a specific monocular depth estimator; the paper shows consistent gains with different DPT variants, suggesting robustness to the choice of coarse-depth backbone.
- Training remains fast at minutes per scene, roughly 20 to 40 times faster than NeRF-based sparse-view methods, making the approach practical for real-world use.
- The paper's ablations attribute the improvements specifically to refined depth priors, relative depth guidance, and adaptive sampling, implying each module contributes meaningfully to the final result.
Reading between the lines
- A testable extension is to apply the same relative-depth loss to NeRF-style volume rendering or to stereo-derived depth maps; because the loss only compares patch similarities, it may transfer without Gaussian-specific machinery.
- The adaptive bias $b$ acts like a contrastive margin, and one could study its scheduling as a trade-off between geometry fidelity and texture-copy collapse, possibly making it per-patch rather than global.
- The paper's own limitations, including artifacts in planar regions, mirror reflections, and added training time, point toward hierarchical or asymmetric Gaussian representations and reflection-aware features as the natural next steps.
- A stronger validation would be to compare depth accuracy directly against a metric-depth ground truth, since the current evaluation is primarily through rendered image quality.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes RDG-GS, a sparse-view novel-view-synthesis method built on 3D Gaussian Splatting. The main contributions are (i) an energy-based refinement of monocular depth priors that injects global and fine-grained RGB structure into the depth used for regularization, (ii) a relative-depth-guidance loss that aligns patch-wise cosine similarities of rendered depth and image features, and (iii) an adaptive sampling strategy that densifies Gaussians in high-error regions. The authors report state-of-the-art PSNR/SSIM/LPIPS results on Mip-NeRF360, LLFF, DTU, and Blender, with real-time rendering speeds around 112 FPS. The paper also includes a detailed ablation study and additional robustness experiments over DPT variants, spherical-harmonic degree, initialization, and energy-function weights.
Significance. If the reported numbers are reliable, the method would be a meaningful advance for sparse-view 3DGS rendering, offering substantial quality gains over FSGS, CoR-GS, and geometry-aware 3DGS baselines while preserving real-time rendering. The paper deserves credit for its extensive ablation suite (Tables 9-13), for testing robustness across different DPT models and initialization strategies, and for explicitly listing limitations regarding training time, anisotropic Gaussians, and mirror reflections. I do not find the circularity concern raised in the stress-test note to be substantiated: the depth supervision comes from an external estimator and an energy refinement, and the relative-depth loss operates on rendered images and rendered depths rather than on the test target, so no step is fitted to the evaluation metric. The decisive weakness is that the central SOTA claim rests on an internally inconsistent evaluation protocol: conflicting baseline numbers and dataset descriptions prevent a reader from verifying the main contribution. The manuscript also provides no code, no per-scene breakdown, and no variance estimates, which at present makes the headline numbers uncheckable.
major comments (4)
- [Sec. 4.1.1 and Supplement, 'Introduction of Datasets'] The Mip-NeRF360 evaluation protocol is described inconsistently. The main text says 24 viewpoints from 7 scenes, while the supplement says 8 scenes and explicitly lists 'bear' as one of them; the official Mip-NeRF360 dataset contains 9 scenes and does not include a scene named 'bear'. The authors must state exactly which scenes are used, justify the subset, and ensure the main text and supplement agree, because every comparison in Tables 1, 2, 6, 7, and 9 depends on this protocol.
- [Tables 1, 2, and 7] The baseline numbers are mutually inconsistent, which makes the SOTA claim undecidable. For example, at 1/8 resolution Table 1 reports 3D-GS as PSNR 20.89 / SSIM 0.633 / LPIPS 0.317 and FSGS as 23.70 / 0.745 / 0.230, whereas Table 2 under the 24-view setting gives 3DGS 22.80 / 0.708 / 0.276 and FSGS 23.28 / 0.715 / 0.274 without specifying the resolution. In Table 7, the '3D-GS [26] None' row at 1/2 resolution reports exactly the same numbers as the FreeNeRF row (18.35 / 0.476 / 0.514), while Table 1 at 1/2 resolution reports 3D-GS as 17.12 / 0.476 / 0.514. These discrepancies directly affect the claimed advantage of 26.03 PSNR and must be resolved with one consistent protocol before the paper can be evaluated.
- [Sec. 4.1.1 and Tables 1-6] No variance or per-scene statistics are reported, although Sec. 4.1.1 states that the LLFF results are averaged over 10 experiments. With only 8, 15, or 9 scenes per dataset and 3-24 training views, a difference of 1-2 dB can fall within run-to-run noise, and the paper provides no way to assess this. The authors should report standard deviations or per-scene tables, and should clarify the resolution and train/test split for each table, including Table 5, where the 3-view LLFF 3D-GS number (19.22) conflicts with the value in Table 4 (17.83 at 503x381).
- [Sec. 3.3, Eq. (7)] The central relative-depth-guidance mechanism is motivated by an assumption that is not independently verified: that DINOv2 patch-feature similarity computed on depth maps is a faithful proxy for 3D spatial proximity, and that matching RGB patch similarity to depth patch similarity yields view-consistent geometry. The paper should provide a direct validation of this premise, for example an analysis on rendered depth maps showing that the depth patch-similarity tensor preserves relative depth ordering, or an ablation that replaces the DINO features with a simpler spatial-distance baseline. As written, the mechanism's success is only measured indirectly through the final rendering metrics.
minor comments (6)
- [Sec. 3.2, Eq. (3)] Eq. (3) writes Dr = arg max Dr E(Dr | I, Dc), while the surrounding text says the refined depth is obtained by 'minimizing the energy function E'; the sign convention should be made consistent.
- [Sec. 3.2.1 and Table 13] The notation for the energy-function weights is inconsistent: Eq. (4) uses wu, wp, wh, while Table 13 and Sec. 4.4.4 refer to wg and wh. Please align the notation.
- [Figure captions] Several figure captions contain typos or placeholder text: Fig. 3 uses 'RADG-GS' instead of RDG-GS, Figs. 5-8 use 'RDG-DS' instead of RDG-GS, and Fig. 4 contains the garbled string 'FPS 310 FPS221FPS0.03FPS'.
- [Table 2 and Table 8 captions] Table 2 cites the Mip-NeRF360 dataset as reference [45] but the correct reference is [3], and Table 8 cites the Blender dataset as [68] instead of [34]; the reference list also contains duplicate entries for DietNeRF ([23] and [24]).
- [Sec. 4.2.1 and Sec. 4.2.6] The speed-up claim is stated inconsistently: Sec. 4.2.1 says 'over 4000x faster' while Sec. 4.2.6 says 'over 3500x acceleration'; Table 7 implies about 3733x. Please use one consistent number.
- [Table 10] The variant name 'dpt large-384' is missing a hyphen and should read 'dpt-large-384' for consistency with the other entries.
Circularity Check
No significant circularity: the depth priors and relative-depth guidance are auxiliary regularizers driven by an external depth estimator (DPT) and model-internal consistency, not by the evaluation metrics.
full rationale
The derivation chain is self-contained rather than circular. Coarse depth D_c comes from the external monocular estimator F (DPT, Sec. 3.2); refined depth D_r is obtained by minimizing an energy function E(D_r | I, D_c) (Eqs. 3-5), and the rendered depth D_o is produced by the alpha-blend rasterizer (Eq. 2). The relative depth guidance loss L_rdg (Eq. 7) aligns image-feature patch similarity with depth-patch similarity on rendered outputs, and the final objective (Eq. 13) combines L_color against ground-truth training views, L_depth against the D_r pseudo-label, and L_rdg. No quantity in this chain is defined in terms of the evaluation metrics PSNR/SSIM/LPIPS, and no parameter is fitted to the test views; the hyperparameters are fixed constants reported in Sec. 4.1.2, with ablations in Tables 9-13. The closest loop is L_rdg, which matches two rendered quantities to each other, but it is an auxiliary regularizer and the paper validates it through held-out-view comparisons and ablations, so it is not a fitted prediction or a self-definition of the reported numbers. The serious protocol inconsistencies (7 vs 8 vs 9 Mip-NeRF360 scenes; conflicting baseline PSNR/SSIM in Tables 1 and 2; the duplicated 3D-GS row in Table 7) are correctness and reproducibility concerns, not circularity, and therefore do not raise the circularity score. The citation [55] for the CRF-style energy function is methodological borrowing, not a load-bearing self-citation chain.
Assumptions & free parameters
free parameters (7)
- Energy weight w_g (global structural consistency) =
10
- Energy weight w_h (texture detail constraint) =
5
- Energy weight w_u (local similarity baseline) =
1
- Gaussian kernel stds theta_alpha, theta_mu, theta_beta =
35,10,10 and 10,2,2
- High-frequency sensitivity tau and gamma =
tau=5, gamma=10
- Relative depth guidance bias b (initial) =
0.4
- Loss weights beta, lambda, omega =
0.4, 0.1, 0.05
assumptions (4)
- domain assumption Coarse depth from a pre-trained monocular depth estimator (DPT) is a useful starting point for sparse-view geometry.
- domain assumption DINOv2 features computed on a depth map encode meaningful spatial similarity.
- domain assumption The RGB-guided energy function of [55] transfers correct geometry without texture-copy artifacts.
- ad hoc to paper Uniform ray sampling within an adaptive depth range fills under-covered regions beneficially.
Cite this review
Pith. "Pith review of RDG-GS: Relative Depth Guidance with Gaussian Splatting for Real-time Sparse-View 3D Rendering." pith.science (2026). https://pith.science/paper/XEPF4QPE
@misc{pith2026250111102,
author = {Pith},
title = {Pith review of: RDG-GS: Relative Depth Guidance with Gaussian Splatting for Real-time Sparse-View 3D Rendering},
year = {2026},
howpublished = {\url{https://pith.science/paper/XEPF4QPE}},
note = {Machine review of arXiv:2501.11102}
}
read the original abstract
Efficiently synthesizing novel views from sparse inputs while maintaining accuracy remains a critical challenge in 3D reconstruction. While advanced techniques like radiance fields and 3D Gaussian Splatting achieve rendering quality and impressive efficiency with dense view inputs, they suffer from significant geometric reconstruction errors when applied to sparse input views. Moreover, although recent methods leverage monocular depth estimation to enhance geometric learning, their dependence on single-view estimated depth often leads to view inconsistency issues across different viewpoints. Consequently, this reliance on absolute depth can introduce inaccuracies in geometric information, ultimately compromising the quality of scene reconstruction with Gaussian splats. In this paper, we present RDG-GS, a novel sparse-view 3D rendering framework with Relative Depth Guidance based on 3D Gaussian Splatting. The core innovation lies in utilizing relative depth guidance to refine the Gaussian field, steering it towards view-consistent spatial geometric representations, thereby enabling the reconstruction of accurate geometric structures and capturing intricate textures. First, we devise refined depth priors to rectify the coarse estimated depth and insert global and fine-grained scene information to regular Gaussians. Building on this, to address spatial geometric inaccuracies from absolute depth, we propose relative depth guidance by optimizing the similarity between spatially correlated patches of depth and images. Additionally, we also directly deal with the sparse areas challenging to converge by the adaptive sampling for quick densification. Across extensive experiments on Mip-NeRF360, LLFF, DTU, and Blender, RDG-GS demonstrates state-of-the-art rendering quality and efficiency, making a significant advancement for real-world application.
Reference graph
Works this paper leans on
- [26]
-
[1]
S. Avidan and A. Shashua. Novel view syn- thesis in tensor space. In Proceedings of IEEE Computer Society Conference on Com- puter Vision and Pattern Recognition , pages 1034–1040. IEEE, 1997
work page 1997
-
[2]
J. T. Barron, B. Mildenhall, M. Tancik, P. Hedman, R. Martin-Brualla, and P. P. Srinivasan. Mip-nerf: A multiscale represen- tation for anti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, pages 5855–5864, 2021. Springer Nature 2021 LATEX template 20 Article Title
work page 2021
-
[3]
J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5470–5479, 2022
work page 2022
-
[4]
J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman. Zip-nerf: Anti-aliased grid-based neural radiance fields. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, pages 19697–19705, 2023
work page 2023
-
[5]
S. F. Bhat, R. Birkl, D. Wofk, P. Wonka, and M. M¨ uller. Zoedepth: Zero-shot transfer by combining relative and metric depth. arXiv preprint arXiv:2302.12288, 2023
arXiv 2023
- [6]
-
[7]
A. Chen, Z. Xu, A. Geiger, J. Yu, and H. Su. Tensorf: Tensorial radiance fields, 2022
work page 2022
Show all 68 references
-
[8]
A. Chen, Z. Xu, F. Zhao, X. Zhang, F. Xiang, J. Yu, and H. Su. Mvsnerf: Fast generalizable radiance field reconstruction from multi-view stereo. In Proceedings of the IEEE/CVF international conference on computer vision , pages 14124–14133, 2021
2021
-
[9]
D. Chen, H. Li, W. Ye, Y. Wang, W. Xie, S. Zhai, N. Wang, H. Liu, H. Bao, and G. Zhang. Pgsr: Planar-based gaus- sian splatting for efficient and high-fidelity surface reconstruction. arXiv preprint arXiv:2406.06521, 2024
2024 arXiv
-
[10]
T. Chen, P. Wang, Z. Fan, and Z. Wang. Aug- nerf: Training stronger neural radiance fields with triple-level physically-grounded aug- mentations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15191–15202, 2022
2022
-
[11]
Cheng, Y.-P
W. Cheng, Y.-P. Cao, and Y. Shan. Sparsegnv: Generating novel views of indoor scenes with sparse rgb-d images. In Pro- ceedings of the AAAI Conference on Artifi- cial Intelligence, volume 38, pages 1308–1316, 2024
2024
-
[12]
Chung, J
J. Chung, J. Oh, and K. M. Lee. Depth- regularized optimization for 3d gaussian splatting in few-shot images. arXiv preprint arXiv:2311.13398, 2023
2023 arXiv
-
[13]
W. Cong, H. Liang, P. Wang, Z. Fan, T. Chen, M. Varma, Y. Wang, and Z. Wang. Enhancing nerf akin to enhanc- ing llms: Generalizable nerf transformer with mixture-of-view-experts. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 3193–3204, 2023
2023
-
[14]
K. Deng, A. Liu, J.-Y. Zhu, and D. Ramanan. Depth-supervised nerf: Fewer views and faster training for free. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12882–12891, 2022
2022
-
[15]
Fridovich-Keil, G
S. Fridovich-Keil, G. Meanti, F. R. War- burg, B. Recht, and A. Kanazawa. K-planes: Explicit radiance fields in space, time, and appearance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12479–12488, 2023
2023
-
[16]
Fridovich-Keil, A
S. Fridovich-Keil, A. Yu, M. Tancik, Q. Chen, B. Recht, and A. Kanazawa. Plenoxels: Radiance fields without neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 5501–5510, 2022
2022
-
[17]
K. Gao, Y. Gao, H. He, D. Lu, L. Xu, and J. Li. Nerf: Neural radiance field in 3d vision, a comprehensive review. arXiv preprint arXiv:2210.00379, 2022
2022 arXiv
-
[18]
S. J. Garbin, M. Kowalski, M. Johnson, J. Shotton, and J. Valentin. Fastnerf: High- fidelity neural rendering at 200fps, 2021
2021
-
[19]
Gu´ edon and V
A. Gu´ edon and V. Lepetit. Sugar: Surface- aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering. In Proceedings of the IEEE/CVF Springer Nature 2021 LATEX template Article Title 21 Conference on Computer Vision and Pattern Recognitio...
2021
-
[20]
S. Guo, Q. Wang, Y. Gao, R. Xie, and L. Song. Depth-guided robust and fast point cloud fusion nerf for sparse input views. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 1976– 1984, 2024
1976
-
[21]
Y.-C. Guo, D. Kang, L. Bao, Y. He, and S.-H. Zhang. Nerfren: Neural radiance fields with reflections. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18409–18418, 2022
2022
-
[22]
Huang, Z
B. Huang, Z. Yu, A. Chen, A. Geiger, and S. Gao. 2d gaussian splatting for geomet- rically accurate radiance fields. In ACM SIGGRAPH 2024 Conference Papers , pages 1–11, 2024
2024
-
[24]
A. Jain, M. Tancik, and P. Abbeel. Putting nerf on a diet: Semantically consistent few- shot view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 5885–5894, 2021
2021
-
[25]
Jensen, A
R. Jensen, A. Dahl, G. Vogiatzis, E. Tola, and H. Aanæs. Large scale multi-view stereop- sis evaluation. In 2014 IEEE Conference on Computer Vision and Pattern Recognition , pages 406–413, 2014
2014
-
[27]
Khosla, P
P. Khosla, P. Teterwak, C. Wang, A. Sarna, Y. Tian, P. Isola, A. Maschinot, C. Liu, and D. Krishnan. Supervised contrastive learning. Advances in neural information processing systems, 33:18661–18673, 2020
2020
-
[28]
M. Kim, S. Seo, and B. Han. Infonerf: Ray entropy minimization for few-shot neu- ral volume rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12912–12921, 2022
2022
-
[29]
C. Li, B. Y. Feng, Y. Liu, H. Liu, C. Wang, W. Yu, and Y. Yuan. Endosparse: Real-time sparse view synthesis of endoscopic scenes using gaussian splatting, 2024
2024
-
[30]
J. Li, J. Zhang, X. Bai, J. Zheng, X. Ning, J. Zhou, and L. Gu. Dngaussian: Optimiz- ing sparse-view 3d gaussian radiance fields with global-local depth normalization. arXiv preprint arXiv:2403.06912, 2024
2024 arXiv
-
[31]
Z. Li, Z. Chen, Z. Li, and Y. Xu. Space- time gaussian feature splatting for real-time dynamic view synthesis. arXiv preprint arXiv:2312.16812, 2023
2023 arXiv
-
[32]
L. Liu, J. Gu, K. Zaw Lin, T.-S. Chua, and C. Theobalt. Neural sparse voxel fields. Advances in Neural Information Processing Systems, 33:15651–15663, 2020
2020
-
[33]
Mildenhall, P
B. Mildenhall, P. P. Srinivasan, R. Ortiz- Cayon, N. K. Kalantari, R. Ramamoorthi, R. Ng, and A. Kar. Local light field fusion: Practical view synthesis with prescriptive sampling guidelines. ACM Transactions on Graphics (TOG), 38(4):1–14, 2019
2019
-
[34]
Mildenhall, P
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021
2021
-
[35]
M¨ uller, A
T. M¨ uller, A. Evans, C. Schied, and A. Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM trans- actions on graphics (TOG), 41(4):1–15, 2022
2022
-
[36]
Niedermayr, J
S. Niedermayr, J. Stumpfegger, and R. West- ermann. Compressed 3d gaussian splatting for accelerated novel view synthesis. arXiv preprint arXiv:2401.02436, 2023. Springer Nature 2021 LATEX template 22 Article Title
2023 arXiv
-
[37]
Niemeyer, J
M. Niemeyer, J. T. Barron, B. Mildenhall, M. S. Sajjadi, A. Geiger, and N. Radwan. Regnerf: Regularizing neural radiance fields for view synthesis from sparse inputs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 5480–5490, 2022
2022
-
[38]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernan- dez, D. Haziza, F. Massa, A. El-Nouby, et al. Dinov2: Learning robust visual fea- tures without supervision. arXiv preprint arXiv:2304.07193, 2023
2023 arXiv
-
[39]
Rabby and C
A. Rabby and C. Zhang. Beyondpixels: A comprehensive review of the evolution of neural radiance fields. arXiv preprint arXiv:2306.03000, 2023
2023 arXiv
-
[40]
Ranftl, A
R. Ranftl, A. Bochkovskiy, and V. Koltun. Vision transformers for dense prediction, 2021
2021
-
[41]
Roessle, J
B. Roessle, J. T. Barron, B. Mildenhall, P. P. Srinivasan, and M. Nießner. Dense depth priors for neural radiance fields from sparse input views. CoRR, abs/2112.03288, 2021
2021 arXiv
-
[42]
J. L. Schonberger and J.-M. Frahm. Structure-from-motion revisited. In Proceed- ings of the IEEE conference on computer vision and pattern recognition , pages 4104–4113, 2016
2016
-
[43]
Schwarz, A
K. Schwarz, A. Sauer, M. Niemeyer, Y. Liao, and A. Geiger. Voxgraf: Fast 3d-aware image synthesis with sparse voxel grids. Advances in Neural Information Processing Systems , 35:33999–34011, 2022
2022
-
[44]
Z. Shao, Z. Wang, Z. Li, D. Wang, X. Lin, Y. Zhang, M. Fan, and Z. Wang. Splattinga- vatar: Realistic real-time human avatars with mesh-embedded gaussian splatting. arXiv preprint arXiv:2403.05087, 2024
2024 arXiv
-
[45]
Somraj, A
N. Somraj, A. Karanayil, and R. Soundarara- jan. Simplenerf: Regularizing sparse input neural radiance fields with simpler solu- tions. In SIGGRAPH Asia 2023 Conference Papers, pages 1–11, 2023
2023
-
[46]
Somraj and R
N. Somraj and R. Soundararajan. Vip- nerf: Visibility prior for sparse input neural radiance fields. In ACM SIGGRAPH 2023 Conference Proceedings, pages 1–11, 2023
2023
-
[47]
J. Song, S. Park, H. An, S. Cho, M.-S. Kwak, S. Cho, and S. Kim. D \” arf: Boost- ing radiance fields from sparse inputs with monocular depth adaptation. arXiv preprint arXiv:2305.19201, 2023
2023 arXiv
-
[48]
Suhail, C
M. Suhail, C. Esteves, L. Sigal, and A. Maka- dia. Light field neural rendering. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 8269–8279, 2022
2022
-
[49]
C. Sun, M. Sun, and H. Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. In CVPR, 2022
2022
-
[50]
C. Sun, M. Sun, and H.-T. Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction, 2022
2022
-
[51]
Tewari, J
A. Tewari, J. Thies, B. Mildenhall, P. Srini- vasan, E. Tretschk, W. Yifan, C. Lassner, V. Sitzmann, R. Martin-Brualla, S. Lom- bardi, et al. Advances in neural rendering. In Computer Graphics Forum, volume 41, pages 703–735. Wiley Online Library, 2022
2022
-
[52]
M. A. Uy, R. Martin-Brualla, L. Guibas, and K. Li. Scade: Nerfs from space carving with ambiguity-aware depth estimates. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 16518–16527, 2023
2023
-
[53]
C. Wang, M. Chai, M. He, D. Chen, and J. Liao. Clip-nerf: Text-and-image driven manipulation of neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 3835–3844, 2022
2022
-
[54]
G. Wang, Z. Chen, C. C. Loy, and Z. Liu. Sparsenerf: Distilling depth ranking for few- shot novel view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 9065–9076, 2023. Springer Nature 2021 LATEX template Article Title 23
2023
-
[55]
H. Wang, M. Yang, C. Zhu, and N. Zheng. Rgb-guided depth map recovery by two-stage coarse-to-fine dense crf models. IEEE Trans- actions on Image Processing , 32:1315–1328, 2023
2023
-
[56]
L. Wang, J. Zhang, X. Liu, F. Zhao, Y. Zhang, Y. Zhang, M. Wu, J. Yu, and L. Xu. Fourier plenoctrees for dynamic radiance field rendering in real-time. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 13524–13534, 2022
2022
-
[57]
P. Wang, Y. Liu, Z. Chen, L. Liu, Z. Liu, T. Komura, C. Theobalt, and W. Wang. F2- nerf: Fast neural radiance field training with free camera trajectories. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 4150–4159, 2023
2023
-
[58]
R. Wu, B. Mildenhall, P. Henzler, K. Park, R. Gao, D. Watson, P. P. Srinivasan, D. Verbin, J. T. Barron, B. Poole, et al. Reconfusion: 3d reconstruction with diffusion priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21551–21561, 2024
2024
-
[59]
Wynn and D
J. Wynn and D. Turmukhambetov. Diffu- sionerf: Regularizing neural radiance fields with denoising diffusion models. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 4180–4189, 2023
2023
-
[60]
Xiong, S
H. Xiong, S. Muttukuru, R. Upadhyay, P. Chari, and A. Kadambi. Sparsegs: Real- time 360 sparse view synthesis using gaussian splatting. 2023
2023
-
[61]
C. Yang, S. Li, J. Fang, R. Liang, L. Xie, X. Zhang, W. Shen, and Q. Tian. Gaus- sianobject: Just taking four images to get a high-quality 3d object with gaussian splat- ting, 2024
2024
-
[62]
J. Yang, M. Pavone, and Y. Wang. Freen- erf: Improving few-shot neural rendering with free frequency regularization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8254– 8263, 2023
2023
-
[63]
Z. Yang, X. Gao, W. Zhou, S. Jiao, Y. Zhang, and X. Jin. Deformable 3d gaussians for high-fidelity monocular dynamic scene recon- struction. arXiv preprint arXiv:2309.13101 , 2023
2023 arXiv
-
[64]
R. Yin, V. Yugay, Y. Li, S. Karaoglu, and T. Gevers. Fewviewgs: Gaussian splatting with few view matching and multi-stage training. arXiv preprint arXiv:2411.02229 , 2024
2024 arXiv
-
[65]
A. Yu, V. Ye, M. Tancik, and A. Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4578–4587, 2021
2021
-
[66]
Z. Yu, T. Sattler, and A. Geiger. Gaus- sian opacity fields: Efficient adaptive surface reconstruction in unbounded scenes. ACM Transactions on Graphics (TOG) , 43(6):1– 13, 2024
2024
-
[67]
Zhang, J
J. Zhang, J. Li, X. Yu, L. Huang, L. Gu, J. Zheng, and X. Bai. Cor-gs: sparse-view 3d gaussian splatting via co-regularization. In European Conference on Computer Vision , pages 335–352. Springer, 2025
2025
-
[68]
Zhang, P
R. Zhang, P. Isola, A. A. Efros, E. Shecht- man, and O. Wang. The unreasonable effec- tiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 586–595, 2018
2018
-
[69]
counter”, “room
Z. Zhu, Z. Fan, Y. Jiang, and Z. Wang. Fsgs: Real-time few-shot view synthesis using gaussian splatting. arXiv preprint arXiv:2312.00451, 2023. Springer Nature 2021 LATEX template 24 Article Title Supplemental Detail Theoretical Analysis Our refined depth restoration hinges on...
2023 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.