REVIEW 2 major objections 6 minor 2 cited by
SparSplat: Fast Multi-View Reconstruction with Generalizable 2D Gaussian Splatting
T0 review · 2 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read SparSplat claims state-of-the-art sparse-view 3D reconstruction and novel view synthesis by regressing 2D Gaussian surface elements from three input images in one feed-forward pass, at roughly 0.8 seconds per scene.
desk verdict Useful engineering contribution with a real speed win, but the SOTA claims rest on a thin margin and a non-reproduced baseline value. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the 2D Gaussian surface element: a flat, elliptical primitive defined by a center, two tangent scaling factors, a rotation, an opacity, and a color, rasterized by intersecting each pixel ray with the splat's plane. Its role is to make predicted depth consistent across views, so that fusing several rendered depth maps with TSDF yields a coherent surface instead of the warped or fragmented result that 3D Gaussian primitives produce. The network combines a multi-view stereo cost volume that produces depth, a second branch that regresses the remaining attributes, and injected dense pairwise matching features that materially improve the cost volume. Training is end-to-end with color, SSIM, perceptual, depth, depth-distortion, and normal-consistency losses.
What would settle it
Run the released pretrained model of the previous best implicit reconstruction method on the same 15 DTU test scenes with the same two three-view sets, the same masks, and the same voxel size, and compute the mean Chamfer distance; if it comes out below 1.04, the paper's main accuracy claim is overturned.
Extended reading notes
Core claim
The central claim is that representing predicted scene geometry as flat 2D Gaussian surface elements, rather than volumetric 3D Gaussians, lets a single generalizable feed-forward network solve sparse-view novel view synthesis and 3D mesh reconstruction together. The network predicts per-pixel surface element parameters for a target view from three source views, places each element's center by unprojecting a multi-view stereo depth prediction, and renders with perspective-accurate splatting. On the DTU sparse reconstruction benchmark the authors report a mean Chamfer distance of 1.04, narrowly ahead of the reproduced 1.05 of the previous leading generalizable implicit method, while improving novel view PSNR over the Gaussian-splatting backbone and cutting inference time from tens of seconds to about 0.8 seconds. The paper also shows that enriching the encoder with dense pairwise correspondence features from a pretrained 3D foundation model gives a larger reconstruction boost than monocular semantic features.
Load-bearing premise
The state-of-the-art reconstruction claim rests on the comparison value used for the previous best method: the authors could not reproduce the lower error that method originally reported, so if that original number is correct, this method may not actually be the most accurate.
Editorial extensions
If this is right
- Three-view reconstruction becomes a single forward pass, making mesh extraction and novel view synthesis available in under a second on a single GPU.
- A single model can serve both reconstruction and rendering, eliminating the need for separate pipelines for these two tasks.
- Because the model generalizes without fine-tuning to unseen indoor and outdoor scenes, it can be applied directly to new captures.
- The finding that pairwise 3D foundation features help more than monocular features gives a transferable recipe for improving other multi-view stereo based Gaussian prediction pipelines.
Reading between the lines
- If the speed and accuracy claims hold beyond the reported benchmarks, sparse-view reconstruction may shift from offline optimization to interactive settings, with pose estimation and feature extraction becoming the main bottlenecks.
- A natural testable extension is to replace TSDF fusion with a learned surface extraction directly from the 2D Gaussian primitives, which could preserve detail lost by voxel fusion.
- The relative gain from dense correspondence features over semantic features suggests stereo matching quality, not semantic understanding, is the limiting factor in sparse-view reconstruction; even richer correspondence features could push accuracy further.
- Because the model builds on an existing multi-view stereo Gaussian backbone, a fair comparison against a stronger version of that backbone would clarify how much of the gain comes from the 2D primitive representation itself.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SparSplat, a feed-forward, generalizable 2D Gaussian splatting model for joint sparse-view 3D reconstruction and novel view synthesis. Given three posed images, the model predicts pixel-aligned 2DGS parameters by extending the MVSGaussian architecture with features from DINOv2 and MASt3R, and trains with RGB, depth, depth-distortion, and normal-consistency losses. The authors report the best mean Chamfer distance on the DTU sparse reconstruction benchmark (1.04 vs. 1.05 for their reproduced UfoRecon baseline), improve PSNR and LPIPS over MVSGaussian on DTU NVS, demonstrate generalization to BlendedMVS and Tanks and Temples, and highlight an inference time of about 0.8 s versus tens of seconds for implicit generalizable methods. The core technical contribution is a fast generalizable 2DGS pipeline, and the paper argues that the representation and the additional foundation-model features drive the gains.
Significance. If the reported numbers hold, the paper makes a useful contribution: it is, to my knowledge, the first generalizable feed-forward 2D Gaussian splatting method for joint reconstruction and NVS, and it demonstrates a large practical speed advantage over implicit generalizable reconstruction baselines. The ablation showing that MASt3R features improve reconstruction over FPN or DINOv2 features (Table 3) is informative, and the use of dense depth supervision with 2DGS in a generalizable setting is a sensible extension. The methodological core is plausible and the experimental setup is mostly aligned with prior work. However, the two headline claims of state-of-the-art reconstruction and state-of-the-art NVS are not yet secure because of the baseline-reproduction issue and the selective use of metrics, as detailed below.
major comments (2)
- [Section 4.1 and Table 1] The headline SOTA reconstruction claim rests on the authors' reproduced UfoRecon mean Chamfer of 1.05, yet the text explicitly states that the authors could not reproduce the Chamfer metrics originally reported in the UfoRecon paper. If the original UfoRecon value is below 1.04, SparSplat is not SOTA on DTU reconstruction. The margin is 0.01 on a mean over 15 scans, and no variance or significance testing is reported for either method, so even the comparison against the reproduced number may be within noise. Please report the originally published UfoRecon values alongside the repro-duced ones, provide per-scan numbers and standard errors for both methods, and either support or temper the abstract's SOTA claim accordingly.
- [Table 2 and Section 4.3] The claim of state-of-the-art NVS is metric-selective. Compared with MVSGaussian, SparSplat improves PSNR by 0.12 dB and LPIPS by 0.003, but SSIM drops from 0.963 to 0.938, a relative decrease of about 2.6%. The abstract and Section 4.3 assert SOTA NVS based on PSNR/LPIPS. Please report per-scene metrics and discuss the SSIM drop, or revise the claim to state explicitly on which metrics the method is SOTA and acknowledge the SSIM regression.
minor comments (6)
- [Table 1 caption] The caption cites UfoRecon as reference [69] and writes 'UfoRecon*1 [69]', but the bibliography lists UfoRecon as [54]; the footnote marker '1' also appears to dangle. Please correct the citation and place the footnote marker consistently.
- [Supplementary Section C] The test scan IDs for DTU surface reconstruction are listed as '24, 37, 40, 55, 63, 192, 65, 69, 83, 97, 105, 106, 110, 114, 118 and 122', which contains 16 entries and includes scan 192, while Table 1 and the main text use the standard 15 scans without 192. Please correct this inconsistency.
- [Figures 2-4 captions] The captions contain the typo 'datatset' instead of 'dataset' (repeated in Figures 2, 3, and 4).
- [Equations (10)-(12)] The loss weights (lambda_s, lambda_p, lambda_alpha, lambda_beta, lambda_gamma, lambda_1, lambda_2) are only given in the supplementary material; consider moving them to the main text so the objective is self-contained.
- [Abstract and Section 2] There are typos such as 'as-well' in the abstract and 'post-hock' in the introduction; also, in Section 3.3 the text says 'We run master on all pairs' where 'MASt3R' is intended.
- [Section 4.4] In the sentence 'Without 2DGS ... the method fails to obtain conherent surfaces', 'conherent' should be 'coherent'.
Circularity Check
No circularity found: the method is an empirical feed-forward pipeline whose outputs are evaluated against external benchmarks, and self-citations appear only as baselines or concurrent work, not as derivation inputs.
full rationale
The paper's central derivation is an end-to-end learning pipeline: input images are encoded with FPN features (optionally augmented by frozen DINOv2 or MASt3R features), homography-warped into a target view, processed through a deep-MVS branch and a pixel-aligned 2DGS attribute regression branch, and trained with image, depth, distortion, and normal-consistency losses. The predicted 2DGS parameters are then rendered to depths and fused by TSDF into meshes. No step in this chain defines a predicted quantity in terms of the target metric, nor fits a parameter to the benchmark and then reports that fit as a prediction. The Chamfer-distance and NVS results are external benchmark measurements against DTU ground truth and testing splits. Self-citations to GeoTransfer [32] and Sparfels [33] are used as comparison baselines or contextual concurrent work, not as justification for the method's architecture or as a source of a uniqueness claim, so they are not load-bearing. The acknowledged difficulty in reproducing the originally reported UfoRecon Chamfer values is a baseline-comparison reproducibility concern, not a circularity in the derivation: the authors report their own run of the pretrained UfoRecon model and transparently state the discrepancy. Similarly, the NVS claim relies on PSNR while SSIM is lower than MVSGaussian; this is a metric-selection concern, not a definitional or fitted-input circularity. No equation in the paper reduces to another equation by construction, no fitted parameter is renamed as a prediction, and no ansatz is smuggled in via self-citation. The paper is self-contained against external benchmarks and its central claims rest on empirical comparisons rather than on its own prior results.
Assumptions & free parameters
free parameters (4)
- Loss weights lambda_s, lambda_p, lambda_alpha, lambda_beta, lambda_gamma, lambda_1, lambda_2 =
0.1, 0.05, 0.05, 0.05, 0.05, 0.5, 1
- TSDF voxel size =
1.5
- Number of source views N =
3
- Image resolution =
640x512
assumptions (6)
- domain assumption Camera poses and intrinsics are provided for all input views.
- domain assumption DTU ground-truth depth and point clouds are reliable for supervision and evaluation.
- domain assumption The IDR foreground masks and two-view-set averaging protocol produce a fair Chamfer comparison.
- standard math 2DGS ray-splat intersection and homography formulation are geometrically correct.
- domain assumption Frozen pretrained MASt3R and DINOv2 features transfer to the target datasets.
- domain assumption TSDF fusion with voxel size 1.5 yields meshes that adequately represent the reconstructed surface.
Cite this review
Pith. "Pith review of SparSplat: Fast Multi-View Reconstruction with Generalizable 2D Gaussian Splatting." pith.science (2026). https://pith.science/paper/YJSAMSFV
@misc{pith2026250502175,
author = {Pith},
title = {Pith review of: SparSplat: Fast Multi-View Reconstruction with Generalizable 2D Gaussian Splatting},
year = {2026},
howpublished = {\url{https://pith.science/paper/YJSAMSFV}},
note = {Machine review of arXiv:2505.02175}
}
read the original abstract
Recovering 3D information from scenes via multi-view stereo reconstruction (MVS) and novel view synthesis (NVS) is inherently challenging, particularly in scenarios involving sparse-view setups. The advent of 3D Gaussian Splatting (3DGS) enabled real-time, photorealistic NVS. Following this, 2D Gaussian Splatting (2DGS) leveraged perspective accurate 2D Gaussian primitive rasterization to achieve accurate geometry representation during rendering, improving 3D scene reconstruction while maintaining real-time performance. Recent approaches have tackled the problem of sparse real-time NVS using 3DGS within a generalizable, MVS-based learning framework to regress 3D Gaussian parameters. Our work extends this line of research by addressing the challenge of generalizable sparse 3D reconstruction and NVS jointly, and manages to perform successfully at both tasks. We propose an MVS-based learning pipeline that regresses 2DGS surface element parameters in a feed-forward fashion to perform 3D shape reconstruction and NVS from sparse-view images. We further show that our generalizable pipeline can benefit from preexisting foundational multi-view deep visual features. The resulting model attains the state-of-the-art results on the DTU sparse 3D reconstruction benchmark in terms of Chamfer distance to ground-truth, as-well as state-of-the-art NVS. It also demonstrates strong generalization on the BlendedMVS and Tanks and Temples datasets. We note that our model outperforms the prior state-of-the-art in feed-forward sparse view reconstruction based on volume rendering of implicit representations, while offering an almost 2 orders of magnitude higher inference speed.
Figures
Forward citations
Cited by 2 Pith papers
-
MAC-Splat: Multi-Attribute Consistency for High-Fidelity Sparse-View Reconstruction
Semantically enriched MASt3R correspondences plus a multi-attribute 3D consistency loss raise sparse-view ScanNet++ PSNR by >4.5 dB over Splatt3R and preserve quality under wide baselines.
-
Sparse-View 3D Reconstruction: Recent Advances and Open Challenges
A comprehensive survey that organizes sparse-view 3D reconstruction methods into geometry-based, NeRF, 3DGS, and diffusion-based categories, with benchmarks and open challenges.
Reference graph
Works this paper leans on
-
[1]
Large-scale data for multiple-view stereopsis
Henrik Aanæs, Rasmus Ramsbøl Jensen, George V ogiatzis, Engin Tola, and Anders Bjorholm Dahl. Large-scale data for multiple-view stereopsis. International Journal of Computer Vision, 120(2):153–168, 2016. 2, 4, 5, 6, 7, 8, 9, 10
2016
-
[2]
Neural point-based graph- ics
Kara-Ali Aliev, Artem Sevastopolsky, Maria Kolos, Dmitry Ulyanov, and Victor Lempitsky. Neural point-based graph- ics. In Computer Vision–ECCV 2020: 16th European Con- ference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXII 16, pages 696–712. Springer, 2020. 3
2020
-
[3]
Sal: Sign agnos- tic learning of shapes from raw data
Matan Atzmon and Yaron Lipman. Sal: Sign agnos- tic learning of shapes from raw data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2565–2574, 2020. 2
2020
-
[4]
Mip-nerf: A multiscale representation for anti-aliasing neu- ral radiance fields
Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neu- ral radiance fields. In Proceedings of the IEEE/CVF inter- national conference on computer vision , pages 5855–5864,
-
[5]
Poco: Point con- volution for surface reconstruction
Alexandre Boulch and Renaud Marlet. Poco: Point con- volution for surface reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6302–6314, 2022. 2
2022
-
[6]
Mvsformer++: Revealing the devil in transformer’s details for multi-view stereo
Chenjie Cao, Xinlin Ren, and Yanwei Fu. Mvsformer++: Revealing the devil in transformer’s details for multi-view stereo. arXiv preprint arXiv:2401.11673, 2024. 2, 4
arXiv 2024
-
[7]
Jiazhong Cen, Jiemin Fang, Chen Yang, Lingxi Xie, Xi- aopeng Zhang, Wei Shen, and Qi Tian. Segment any 3d gaussians. arXiv preprint arXiv:2312.00860, 2023. 3
arXiv 2023
-
[8]
pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction
David Charatan, Sizhe Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. In arXiv, 2023. 3
2023
Show all 99 references
-
[9]
pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction
David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19457–19467, 2024. 2
2024
-
[10]
Mvsnerf: Fast general- izable radiance field reconstruction from multi-view stereo
Anpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang, Fanbo Xiang, Jingyi Yu, and Hao Su. Mvsnerf: Fast general- izable radiance field reconstruction from multi-view stereo. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14124–14133, 2021. 1...
2021
-
[11]
Unsupervised inference of signed distance functions from single sparse point clouds without learning priors
Chao Chen, Zhizhong Han, and Yu-Shen Liu. Unsupervised inference of signed distance functions from single sparse point clouds without learning priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2
2023
-
[12]
Aug-nerf: Training stronger neural radiance fields with triple-level physically-grounded augmentations
Tianlong Chen, Peihao Wang, Zhiwen Fan, and Zhangyang Wang. Aug-nerf: Training stronger neural radiance fields with triple-level physically-grounded augmentations. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15191–15202, 2022. 2
2022
-
[13]
Gaussianeditor: Swift and con- trollable 3d editing with gaussian splatting
Yiwen Chen, Zilong Chen, Chi Zhang, Feng Wang, Xi- aofeng Yang, Yikai Wang, Zhongang Cai, Lei Yang, Huap- ing Liu, and Guosheng Lin. Gaussianeditor: Swift and con- trollable 3d editing with gaussian splatting. arXiv preprint arXiv:2311.14521, 2023. 3
2023 arXiv
-
[14]
Explicit correspondence matching for generalizable neural radiance fields
Yuedong Chen, Haofei Xu, Qianyi Wu, Chuanxia Zheng, Tat-Jen Cham, and Jianfei Cai. Explicit correspondence matching for generalizable neural radiance fields. arXiv preprint arXiv:2304.12294, 2023. 7
2023 arXiv
-
[15]
Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images
Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images. ECCV, 2024. 2, 3
2024
-
[16]
Stereo radiance fields (srf): Learning view syn- thesis for sparse views of novel scenes
Julian Chibane, Aayush Bansal, Verica Lazova, and Gerard Pons-Moll. Stereo radiance fields (srf): Learning view syn- thesis for sparse views of novel scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7911–7920, 2021. 2
2021
-
[17]
Improving neural im- plicit surfaces geometry with patch warping
Franc ¸ois Darmon, B´en´edicte Bascle, Jean-Cl ´ement Devaux, Pascal Monasse, and Mathieu Aubry. Improving neural im- plicit surfaces geometry with patch warping. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 6260–6269, 2022. 1, 2
2022
-
[18]
Depth-supervised nerf: Fewer views and faster train- ing for free
Kangle Deng, Andrew Liu, Jun-Yan Zhu, and Deva Ra- manan. Depth-supervised nerf: Fewer views and faster train- ing for free. arXiv preprint arXiv:2107.02791, 2021. 2
2021 arXiv
-
[19]
Depth-supervised nerf: Fewer views and faster train- ing for free
Kangle Deng, Andrew Liu, Jun-Yan Zhu, and Deva Ra- manan. Depth-supervised nerf: Fewer views and faster train- ing for free. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12882– 12891, 2022. 2
2022
-
[20]
Transmvs- net: Global context-aware multi-view stereo network with transformers
Yikang Ding, Wentao Yuan, Qingtian Zhu, Haotian Zhang, Xiangyue Liu, Yuanjiang Wang, and Xiao Liu. Transmvs- net: Global context-aware multi-view stereo network with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 8585...
2022
-
[21]
Points2surf learning implicit surfaces from point clouds
Philipp Erler, Paul Guerrero, Stefan Ohrhallinger, Niloy J Mitra, and Michael Wimmer. Points2surf learning implicit surfaces from point clouds. In European Conference on Computer Vision, pages 108–124. Springer, 2020. 2
2020
-
[22]
Instantsplat: Un- bounded sparse-view pose-free gaussian splatting in 40 sec- onds
Zhiwen Fan, Wenyan Cong, Kairun Wen, Kevin Wang, Jian Zhang, Xinghao Ding, Danfei Xu, Boris Ivanovic, Marco Pavone, Georgios Pavlakos, et al. Instantsplat: Un- bounded sparse-view pose-free gaussian splatting in 40 sec- onds. arXiv preprint arXiv:2403.20309, 2(3):4, 2024. 3
2024 arXiv
-
[23]
Implicit geometric regularization for learning shapes
Amos Gropp, Lior Yariv, Niv Haim, Matan Atzmon, and Yaron Lipman. Implicit geometric regularization for learning shapes. arXiv preprint arXiv:2002.10099, 2020. 2
2002 arXiv
-
[24]
Cascade cost volume for high-resolution multi-view stereo and stereo matching
Xiaodong Gu, Zhiwen Fan, Siyu Zhu, Zuozhuo Dai, Feitong Tan, and Ping Tan. Cascade cost volume for high-resolution multi-view stereo and stereo matching. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2495–2504, 2020. 1, 4, 6
2020
-
[25]
Sugar: Surface- aligned gaussian splatting for efficient 3d mesh reconstruc- tion and high-quality mesh rendering
Antoine Gu ´edon and Vincent Lepetit. Sugar: Surface- aligned gaussian splatting for efficient 3d mesh reconstruc- tion and high-quality mesh rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5354–5363, 2024. 1
2024
-
[26]
Gaussianavatar: Towards realistic human avatar modeling from a single video via animatable 3d gaussians
Liangxiao Hu, Hongwen Zhang, Yuxiang Zhang, Boyao Zhou, Boning Liu, Shengping Zhang, and Liqiang Nie. Gaussianavatar: Towards realistic human avatar modeling from a single video via animatable 3d gaussians. arXiv preprint arXiv:2312.02134, 2023. 3
2023 arXiv
-
[27]
Gauhuman: Articulated gaus- sian splatting from monocular human videos
Shoukang Hu and Ziwei Liu. Gauhuman: Articulated gaus- sian splatting from monocular human videos. arXiv preprint arXiv:, 2023. 3
2023
-
[28]
2d gaussian splatting for geometrically ac- curate radiance fields
Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2d gaussian splatting for geometrically ac- curate radiance fields. In ACM SIGGRAPH 2024 Conference Papers, pages 1–11, 2024. 1, 3, 4, 5, 6, 8
2024
-
[29]
Neural kernel surface re- construction
Jiahui Huang, Zan Gojcic, Matan Atzmon, Or Litany, Sanja Fidler, and Francis Williams. Neural kernel surface re- construction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 4369– 4379, 2023. 2
2023
-
[30]
Putting nerf on a diet: Semantically consistent few-shot view synthesis
Ajay Jain, Matthew Tancik, and Pieter Abbeel. Putting nerf on a diet: Semantically consistent few-shot view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5885–5894, 2021. 2
2021
-
[31]
Neural mesh-based graphics
Shubhendu Jena, Franck Multon, and Adnane Boukhayma. Neural mesh-based graphics. In European Conference on Computer Vision, pages 739–757. Springer, 2022. 3
2022
-
[32]
Geotransfer: Generalizable few-shot multi-view reconstruc- tion via transfer learning
Shubhendu Jena, Franck Multon, and Adnane Boukhayma. Geotransfer: Generalizable few-shot multi-view reconstruc- tion via transfer learning. arXiv preprint arXiv:2408.14724,
-
[33]
Sparfels: Fast reconstruction from sparse un- posed imagery
Shubhendu Jena, Amine Ouasfi, Mae Younes, and Adnane Boukhayma. Sparfels: Fast reconstruction from sparse un- posed imagery. In arXiv, 2025. 3
2025
-
[34]
Sdfdiff: Differentiable rendering of signed distance fields for 3d shape optimization
Yue Jiang, Dantong Ji, Zhizhong Han, and Matthias Zwicker. Sdfdiff: Differentiable rendering of signed distance fields for 3d shape optimization. InProceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 1251–1261, 2020. 2
2020
-
[35]
Geonerf: Generalizing nerf with geometry priors
Mohammad Mahdi Johari, Yann Lepoittevin, and Franc ¸ois Fleuret. Geonerf: Generalizing nerf with geometry priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18365–18375, 2022. 1, 2, 5, 6
2022
-
[36]
Neural lumigraph render- ing
Petr Kellnhofer, Lars C Jebe, Andrew Jones, Ryan Spicer, Kari Pulli, and Gordon Wetzstein. Neural lumigraph render- ing. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 4287–4297,
-
[37]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4):139–1,
-
[38]
Infonerf: Ray entropy minimization for few-shot neural volume ren- dering
Mijeong Kim, Seonguk Seo, and Bohyung Han. Infonerf: Ray entropy minimization for few-shot neural volume ren- dering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 12912– 12921, 2022. 2
2022
-
[39]
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 ,
-
[40]
Tanks and temples: Benchmarking large-scale scene reconstruction
Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. Tanks and temples: Benchmarking large-scale scene reconstruction. ACM Transactions on Graphics (ToG) , 36 (4):1–13, 2017. 5, 8
2017
-
[41]
Ground- ing image matching in 3d with mast3r
Vincent Leroy, Yohann Cabon, and J´erˆome Revaud. Ground- ing image matching in 3d with mast3r. arXiv preprint arXiv:2406.09756, 2024. 2, 3, 4, 6, 8
2024 arXiv
-
[42]
Nerf- pose: A first-reconstruct-then-regress approach for weakly- supervised 6d object pose estimation
Fu Li, Shishir Reddy Vutukur, Hao Yu, Ivan Shugurov, Benjamin Busam, Shaowu Yang, and Slobodan Ilic. Nerf- pose: A first-reconstruct-then-regress approach for weakly- supervised 6d object pose estimation. In Proceedings of the IEEE/CVF International Conference on Computer Visi...
2023
-
[43]
Learn- ing generalizable light field networks from few images
Qian Li, Franck Multon, and Adnane Boukhayma. Learn- ing generalizable light field networks from few images. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages 1–5. IEEE, 2023. 2
2023
-
[44]
Regular- izing neural radiance fields from sparse rgb-d inputs
Qian Li, Franck Multon, and Adnane Boukhayma. Regular- izing neural radiance fields from sparse rgb-d inputs. In2023 IEEE International Conference on Image Processing (ICIP), pages 2320–2324. IEEE, 2023. 2
2023
-
[45]
Retr: Modeling rendering via transformer for generalizable neural surface re- construction
Yixun Liang, Hao He, and Yingcong Chen. Retr: Modeling rendering via transformer for generalizable neural surface re- construction. Advances in Neural Information Processing Systems, 36, 2024. 1, 2, 5, 6, 8, 9
2024
-
[46]
Efficient neural radiance fields for interactive free-viewpoint video
Haotong Lin, Sida Peng, Zhen Xu, Yunzhi Yan, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Efficient neural radiance fields for interactive free-viewpoint video. In SIGGRAPH Asia 2022 Conference Papers, pages 1–9, 2022. 7
2022
-
[47]
Neural sparse voxel fields
Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and Christian Theobalt. Neural sparse voxel fields. Advances in Neural Information Processing Systems, 33:15651–15663,
-
[48]
Fast generalizable gaussian splatting reconstruction from multi-view stereo
Tianqi Liu, Guangcong Wang, Shoukang Hu, Liao Shen, Xinyi Ye, Yuhang Zang, Zhiguo Cao, Wei Li, and Ziwei Liu. Fast generalizable gaussian splatting reconstruction from multi-view stereo. arXiv preprint arXiv:2405.12218 ,
-
[49]
Neural rays for occlusion-aware image-based render- ing
Yuan Liu, Sida Peng, Lingjie Liu, Qianqian Wang, Peng Wang, Christian Theobalt, Xiaowei Zhou, and Wenping Wang. Neural rays for occlusion-aware image-based render- ing. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 7824–7833,
-
[50]
Sparseneus: Fast generalizable neural sur- face reconstruction from sparse views
Xiaoxiao Long, Cheng Lin, Peng Wang, Taku Komura, and Wenping Wang. Sparseneus: Fast generalizable neural sur- face reconstruction from sparse views. In European Confer- ence on Computer Vision, pages 210–227. Springer, 2022. 1, 2, 4, 5, 6, 8, 9
2022
-
[51]
Dynamic 3d gaussians: Tracking by per- sistent dynamic view synthesis
Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. Dynamic 3d gaussians: Tracking by per- sistent dynamic view synthesis. In 3DV, 2024. 3
2024
-
[52]
Occupancy networks: Learning 3d reconstruction in function space
Lars Mescheder, Michael Oechsle, Michael Niemeyer, Se- bastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4460–4470, 2019. 2
2019
-
[53]
Nerf: Representing scenes as neural radiance fields for view syn- thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. Communications of the ACM , 65(1):99–106, 2021. 2
2021
-
[54]
Uforecon: Generalizable sparse-view surface reconstruction from arbitrary and unfavorable sets
Youngju Na, Woo Jae Kim, Kyu Beom Han, Suhyeon Ha, and Sung-Eui Yoon. Uforecon: Generalizable sparse-view surface reconstruction from arbitrary and unfavorable sets. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5094–5104, 2024. 1,...
2024
-
[55]
Differentiable volumetric rendering: Learn- ing implicit 3d representations without 3d supervision
Michael Niemeyer, Lars Mescheder, Michael Oechsle, and Andreas Geiger. Differentiable volumetric rendering: Learn- ing implicit 3d representations without 3d supervision. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition , pages 3504–3515, ...
2020
-
[56]
Reg- nerf: Regularizing neural radiance fields for view synthesis from sparse inputs
Michael Niemeyer, Jonathan T Barron, Ben Mildenhall, Mehdi SM Sajjadi, Andreas Geiger, and Noha Radwan. Reg- nerf: Regularizing neural radiance fields for view synthesis from sparse inputs. arXiv preprint arXiv:2112.00724, 2021. 2
2021 arXiv
-
[57]
Reg- nerf: Regularizing neural radiance fields for view synthesis from sparse inputs
Michael Niemeyer, Jonathan T Barron, Ben Mildenhall, Mehdi SM Sajjadi, Andreas Geiger, and Noha Radwan. Reg- nerf: Regularizing neural radiance fields for view synthesis from sparse inputs. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition,...
2022
-
[58]
Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction
Michael Oechsle, Songyou Peng, and Andreas Geiger. Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 5589–5599, 2021. 1, 2
2021
-
[59]
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 2, 3, 4, 8
2023 arXiv
-
[60]
Few’zero level set’- shot learning of shape signed distance functions in feature space
Amine Ouasfi and Adnane Boukhayma. Few’zero level set’- shot learning of shape signed distance functions in feature space. In ECCV, 2022. 2
2022
-
[61]
Mixing-denoising generalizable occupancy networks
Amine Ouasfi and Adnane Boukhayma. Mixing-denoising generalizable occupancy networks. 3DV, 2024. 2
2024
-
[62]
Few-shot unsuper- vised implicit neural shape representation learning with spa- tial adversaries
Amine Ouasfi and Adnane Boukhayma. Few-shot unsuper- vised implicit neural shape representation learning with spa- tial adversaries. arXiv preprint arXiv:2408.15114, 2024. 2
2024 arXiv
-
[63]
Robustifying gen- eralizable implicit shape networks with a tunable non- parametric model
Amine Ouasfi and Adnane Boukhayma. Robustifying gen- eralizable implicit shape networks with a tunable non- parametric model. Advances in Neural Information Process- ing Systems, 36, 2024. 2
2024
-
[64]
Unsupervised occu- pancy learning from sparse point cloud
Amine Ouasfi and Adnane Boukhayma. Unsupervised occu- pancy learning from sparse point cloud. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21729–21739, 2024. 2
2024
-
[65]
Toward robust neural reconstruction from sparse point sets
Amine Ouasfi, Shubhendu Jena, Eric Marchand, and Ad- nane Boukhayma. Toward robust neural reconstruction from sparse point sets. arXiv preprint arXiv:2412.16361, 2024. 2
2024 arXiv
-
[66]
Deepsdf: Learning con- tinuous signed distance functions for shape representation
Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning con- tinuous signed distance functions for shape representation. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 165–174, 2019. 1, 2
2019
-
[67]
Pytorch: An im- perative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An im- perative style, high-performance deep learning library. Ad- vances in neural information processing systems ...
2019
-
[68]
3dgs-avatar: Animatable avatars via deformable 3d gaussian splatting
Zhiyin Qian, Shaofei Wang, Marko Mihajlovic, Andreas Geiger, and Siyu Tang. 3dgs-avatar: Animatable avatars via deformable 3d gaussian splatting. arXiv preprint arXiv:2312.09228, 2023. 3
2023 arXiv
-
[69]
V olrecon: V olume rendering of signed ray distance functions for generalizable multi-view reconstruction
Yufan Ren, Tong Zhang, Marc Pollefeys, Sabine S ¨usstrunk, and Fangjinhua Wang. V olrecon: V olume rendering of signed ray distance functions for generalizable multi-view reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, page...
2023
-
[70]
Free view synthesis
Gernot Riegler and Vladlen Koltun. Free view synthesis. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIX 16, pages 623–640. Springer, 2020. 3
2020
-
[71]
Structure- from-motion revisited
Johannes L Schonberger and Jan-Michael Frahm. Structure- from-motion revisited. In Proceedings of the IEEE con- ference on computer vision and pattern recognition , pages 4104–4113, 2016. 5, 6
2016
-
[72]
Light field networks: Neu- ral scene representations with single-evaluation rendering
Vincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum, and Fredo Durand. Light field networks: Neu- ral scene representations with single-evaluation rendering. Advances in Neural Information Processing Systems , 34: 19313–19325, 2021. 2
2021
-
[73]
Light field neural rendering
Mohammed Suhail, Carlos Esteves, Leonid Sigal, and Ameesh Makadia. Light field neural rendering. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8269–8279, 2022. 2
2022
-
[74]
Splatter image: Ultra-fast single-view 3d recon- struction
Stanislaw Szymanowicz, Chrisitian Rupprecht, and Andrea Vedaldi. Splatter image: Ultra-fast single-view 3d recon- struction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 10208– 10217, 2024. 2
2024
-
[75]
De- ferred neural rendering: Image synthesis using neural tex- tures
Justus Thies, Michael Zollh ¨ofer, and Matthias Nießner. De- ferred neural rendering: Image synthesis using neural tex- tures. Acm Transactions on Graphics (TOG) , 38(4):1–12,
-
[76]
Grf: Learning a general radi- ance field for 3d representation and rendering
Alex Trevithick and Bo Yang. Grf: Learning a general radi- ance field for 3d representation and rendering. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 15182–15192, 2021. 2
2021
-
[77]
Nerf-feat: 6d object pose estimation using feature rendering
Shishir Reddy Vutukur, Heike Brock, Benjamin Busam, Tolga Birdal, Andreas Hutter, and Slobodan Ilic. Nerf-feat: 6d object pose estimation using feature rendering. In 2024 International Conference on 3D Vision (3DV), pages 1146– 1155, 2024. 2
2024
-
[78]
Patchmatchnet: Learned multi-view patchmatch stereo
Fangjinhua Wang, Silvano Galliani, Christoph V ogel, Pablo Speciale, and Marc Pollefeys. Patchmatchnet: Learned multi-view patchmatch stereo. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14194–14203, 2021. 1
2021
-
[79]
Sparsenerf: Distilling depth ranking for few- shot novel view synthesis
Guangcong Wang, Zhaoxi Chen, Chen Change Loy, and Ziwei Liu. Sparsenerf: Distilling depth ranking for few- shot novel view synthesis. arXiv preprint arXiv:2303.16196,
-
[80]
Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction
Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. arXiv preprint arXiv:2106.10689, 2021. 1, 2, 5, 6
2021 arXiv
-
[81]
Ibr- net: Learning multi-view image-based rendering
Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul P Srinivasan, Howard Zhou, Jonathan T Barron, Ricardo Martin-Brualla, Noah Snavely, and Thomas Funkhouser. Ibr- net: Learning multi-view image-based rendering. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and ...
2021
-
[82]
Hf-neus: Improved surface reconstruction using high-frequency de- tails
Yiqun Wang, Ivan Skorokhodov, and Peter Wonka. Hf-neus: Improved surface reconstruction using high-frequency de- tails. Advances in Neural Information Processing Systems , 35:1966–1978, 2022. 2
1966
-
[83]
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE TIP, 13(4):600–612, 2004. 5
2004
-
[84]
4d gaussian splatting for real-time dynamic scene rendering
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Wang Xinggang. 4d gaussian splatting for real-time dynamic scene rendering. arXiv preprint arXiv:2310.08528, 2023. 3
2023 arXiv
-
[85]
Diffusionerf: Regularizing neural radiance fields with denoising diffu- sion models
Jamie Wynn and Daniyar Turmukhambetov. Diffusionerf: Regularizing neural radiance fields with denoising diffu- sion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 4180– 4189, 2023. 2
2023
-
[86]
Wavenerf: Wavelet-based generalizable neural radiance fields
Muyu Xu, Fangneng Zhan, Jiahui Zhang, Yingchen Yu, Xi- aoqin Zhang, Christian Theobalt, Ling Shao, and Shijian Lu. Wavenerf: Wavelet-based generalizable neural radiance fields. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 18195–18204, 2023. 1, 2
2023
-
[87]
Cost volume pyramid based depth inference for multi-view stereo
Jiayu Yang, Wei Mao, Jose M Alvarez, and Miaomiao Liu. Cost volume pyramid based depth inference for multi-view stereo. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 4877–4886,
-
[88]
Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction.arXiv preprint arXiv:2309.13101, 2023
Ziyi Yang, Xinyu Gao, Wen Zhou, Shaohui Jiao, Yuqing Zhang, and Xiaogang Jin. Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction.arXiv preprint arXiv:2309.13101, 2023. 3
2023 arXiv
-
[89]
Mvsnet: Depth inference for unstructured multi-view stereo
Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan. Mvsnet: Depth inference for unstructured multi-view stereo. In Proceedings of the European conference on computer vi- sion (ECCV), pages 767–783, 2018. 1, 6
2018
-
[90]
Blendedmvs: A large- scale dataset for generalized multi-view stereo networks
Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan. Blendedmvs: A large- scale dataset for generalized multi-view stereo networks. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 1790–1799,...
2020
-
[91]
Multiview neu- ral surface reconstruction by disentangling geometry and ap- pearance
Lior Yariv, Yoni Kasten, Dror Moran, Meirav Galun, Matan Atzmon, Basri Ronen, and Yaron Lipman. Multiview neu- ral surface reconstruction by disentangling geometry and ap- pearance. Advances in Neural Information Processing Sys- tems, 33:2492–2502, 2020. 2, 5
2020
-
[92]
V ol- ume rendering of neural implicit surfaces
Lior Yariv, Jiatao Gu, Yoni Kasten, and Yaron Lipman. V ol- ume rendering of neural implicit surfaces. Advances in Neu- ral Information Processing Systems, 34:4805–4815, 2021. 1, 2, 5, 6
2021
-
[93]
Spar- secraft: Few-shot neural reconstruction through stereopsis guided geometric linearization
Mae Younes, Amine Ouasfi, and Adnane Boukhayma. Spar- secraft: Few-shot neural reconstruction through stereopsis guided geometric linearization. In European Conference on Computer Vision, pages 37–56. Springer, 2024. 2
2024
-
[94]
pixelnerf: Neural radiance fields from one or few images
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4578–4587, 2021. 2, 5, 7
2021
-
[95]
Monosdf: Exploring monocu- lar geometric cues for neural implicit surface reconstruc- tion
Zehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sat- tler, and Andreas Geiger. Monosdf: Exploring monocu- lar geometric cues for neural implicit surface reconstruc- tion. Advances in neural information processing systems , 35:25018–25032, 2022. 2
2022
-
[96]
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, pages 586–595,
-
[97]
Gps- gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis
Shunyuan Zheng, Boyao Zhou, Ruizhi Shao, Boning Liu, Shengping Zhang, Liqiang Nie, and Yebin Liu. Gps- gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...
2024
-
[98]
Open3d: A modern library for 3d data processing
Qian-Yi Zhou, Jaesik Park, and Vladlen Koltun. Open3d: A modern library for 3d data processing. arXiv preprint arXiv:1801.09847, 2018. 3
2018 arXiv
-
[2024]
2, 3, 4, 5, 6, 7, 8, 9, 10
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.