REVIEW 2 major objections 6 minor 106 references
Direct and Explicit 3D Generation from a Single Image
T0 review · 2 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A single feed-forward network generates six consistent depth/color views from one image and lifts them into a textured 3D mesh or splatted scene in 15–25 seconds.
desk verdict Solid feed-forward image-to-3D system with a real evaluation flaw: anisotropic ICP inflates the headline geometry numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is multi-view-consistent depth as an explicit geometry channel. The model repurposes the pretrained Stable Diffusion U-Net and VAE decoder: a depth branch in the U-Net (expert blocks in the first down block and last up block) denoises depth latents alongside RGB latents in one shared pass, and the decoder adds epipolar attention, in which each query pixel attends only to pixels lying on the corresponding epipolar line in other views. Because the six output views are axis-aligned orthographic projections, those epipolar searches collapse to row or column attention, making cross-view consistency cheap at pixel level. Back-projecting the decoded depth maps with their RGB and Gaussian feature channels produces dense surface-aligned 3D Gaussians, which are rendered by differentiable Gaussian splatting to supply a novel-view-synthesis loss, and the same colored point cloud is converted to a textured mesh by screened Poisson surface reconstruction and texture-atlas projection.
What would settle it
Render a set of objects with large concavities or holes, or photographed with wide-angle perspective cameras, run the released model, and compare reconstructed meshes to ground-truth scans using volume IoU and Chamfer distance; if the six-view model degrades substantially more than a 14-view variant on these cases (or shows visible distortion on perspective inputs), the orthographic six-view assumption is the limiting factor. The paper's own supplementary data point in that direction, since its Table 6 shows continued gains from 6 to 14 views and its Table 10 shows slightly worse results when training with perspective images.
Extended reading notes
Core claim
The paper's central claim is that explicit 3D geometry can be generated directly and cheaply by treating depth as a first-class image output of a diffusion model. A branched U-Net denoises RGB and depth latents simultaneously, sharing most weights so a single diffusion pass produces both domains. The latent-to-pixel decoder adds epipolar attention, so each pixel attends only to pixels on the corresponding epipolar line in other views; for six orthographic views these searches reduce to row or column attention. The decoded maps are combined into per-pixel depth, color, opacity, scale, and rotation, yielding surface-aligned Gaussians when back-projected. Those Gaussians render through splatting for a novel-view-synthesis loss, and the same colored point cloud becomes a textured mesh through screened Poisson surface reconstruction. The authors report the best geometry and texture scores on GSO, with Chamfer distance 0.0135, volume IoU 0.7339, depth error 0.073, PSNR 17.85, SSIM 0.851, LPIPS 0.159, at 15–25 seconds per object.
Load-bearing premise
The load-bearing premise is that six fixed orthographic views (front, back, left, right, top, bottom) capture enough of any object that pixel-level epipolar consistency between those views suffices for a clean 3D reconstruction; if a scene has strong perspective, deep concavities, or heavy occlusion, the six-view orthographic assumption can distort the result, as the paper's own limitation appendix and its 4-to-14-view study indicate.
Editorial extensions
If this is right
- A single forward pass produces a splattable 3D scene and a textured mesh in 15–25 seconds, removing the minutes-to-hours per-scene optimization of SDS-based and multi-view-fusion methods.
- Because depth is generated at 512×512 resolution with pixel-level cross-view consistency, mesh extraction needs no learned refinement network, unlike methods that fit Gaussians or meshes by optimization.
- Adding a novel-view-synthesis loss through differentiable Gaussian splatting improves the generated geometry and texture, since the lift from depth to 3D is differentiable.
- The branched U-Net generates RGB and depth together with shared weights, cutting inference time by about 20 percent and GPU memory by about 18 percent relative to sequential domain-switching.
- Reconstruction quality keeps improving as the number of generated views grows from 4 to 14, with the largest gains from 4 to 6 views, so the six-view setting is a deliberate speed-quality compromise.
Reading between the lines
- Since the epipolar simplification relies on orthographic views, a perspective-camera version would need full epipolar sampling; the paper's own perspective-trained variant is slightly worse, so camera-agnostic inputs remain an open extension.
- The same image-format representation (RGB, depth, Gaussian features) could plausibly be conditioned on text, multiple input views, or video frames, but the paper does not test those settings.
- The explicit depth-plus-Gaussian output may make downstream tasks such as mesh editing, rigging, and animation more direct; the supplement shows a rigged and re-posed mesh as a hint of that potential.
- For production-quality assets, a model variant trained to emit 8–14 views (rather than six) could close the remaining quality gap, at the cost of more inference compute.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a feed-forward single-image-to-3D framework that repurposes Stable Diffusion to generate multi-view RGB, depth, and Gaussian feature maps from six orthographic views. The U-Net is modified with a depth branch, and the latent decoder uses epipolar attention to enforce pixel-level multi-view depth consistency. Back-projecting the outputs yields surface-aligned Gaussians, which can be rendered via splatting or converted to a textured mesh via Poisson reconstruction. The authors report state-of-the-art geometry and texture quality over existing image-to-3D methods on the GSO benchmark (Table 1), with generation times of 15–25 seconds, and present ablations supporting the depth branch, epipolar attention, and the NVS loss.
Significance. If the reported results hold, this is a valuable contribution to single-image 3D generation: it produces explicit, high-resolution (512x512) geometry and texture in a single feed-forward pass, leveraging strong 2D priors from Stable Diffusion. The paper is thorough in its comparisons (14 baselines), ablations, and supplementary studies, including a view-count analysis and comparisons with monocular depth estimators. The strengths are the clear architectural design (branched U-Net, epipolar attention, surface-aligned Gaussians) and the honest acknowledgment of limitations. However, the central quantitative claim of surpassing baselines in geometry depends critically on the evaluation protocol, and the current metrics lack statistical robustness. These issues are fixable and do not undermine the overall approach, but they must be addressed before the claims can be fully trusted.
major comments (2)
- [§5 (Metrics) and Appendix C] The geometric metrics in Table 1 are computed after aligning reconstructions to ground truth with a scale-adaptive ICP modified to allow optimal scale factors along each coordinate axis. This is an anisotropic scaling, not a valid 7-DOF similarity (rigid motion plus uniform scale) that is standard in single-image 3D reconstruction benchmarks. Per-axis scaling can arbitrarily stretch a reconstruction to match the ground-truth bounding extents, artificially lowering Chamfer distance and raising volume IoU. Since the headline claim of surpassing baselines in geometry relies on these numbers, the authors should re-evaluate all methods under uniform-scale similarity alignment (e.g., standard scale-adaptive ICP or Procrustes with uniform scale) and report both versions, or provide a compelling justification for anisotropic scaling as the intended comparison metric. Without this, the reported CD improvement of 0.0135 vs 0.0186 over OpenLRM may partially reflect alignment freedom rather than true shape accuracy.
- [Table 1 and §5 (Evaluation Dataset)] The main evaluation uses only 30 GSO objects, and Table 1 reports a single point estimate per method without error bars, per-object distributions, or statistical significance tests. The claimed margins over the strongest baselines (e.g., CD 0.0135 vs 0.0186 for OpenLRM, PSNR 17.85 vs 14.62) could be driven by a few favorable or outlier objects. The authors should report standard deviations, confidence intervals, or paired significance tests (e.g., Wilcoxon signed-rank) on the key metrics. This is essential for the central claim of 'surpasses existing baselines in geometry and texture quality' and is especially important given the anisotropic alignment issue raised above.
minor comments (6)
- [Appendix E and Supplementary Table 6] The paper acknowledges that the orthographic input assumption may cause distortion and reports in Supplementary Table 6 that reconstruction quality still improves when the number of views is increased from 6 to 14 (e.g., CD 0.0070 to 0.0062 on Objaverse). This should be stated more prominently in the main text, as it qualifies the claim that six orthographic views are sufficient for high-quality reconstruction.
- [Eq. (1)] The text states that L_dep-nvs denotes loss over synthesized novel-view depth images, but the equation for L_NVS does not include such a term. Either add the depth term to the equation or remove the mention to avoid inconsistency.
- [Tables 4 and 5] The ablations on the branched U-Net and on depth-vs-normal representation are performed at 256x256 resolution, while the main results are at 512x512. The captions note this, but the main text should state whether the relative conclusions are expected to transfer to the full-resolution setting.
- [Table 1 (Time column)] The reported generation times are not supported by hardware specifications or a controlled comparison. Since the abstract emphasizes 'significantly faster generation time', the authors should specify the GPU model and measure all baselines under identical hardware and software conditions.
- [Figure 2] The overview figure is dense and the epipolar attention block is hard to read at typical print size. Enlarging or explicitly annotating the epipolar attention mechanism would improve clarity.
- [§4.3] The terms 'expert branch' and 'branched U-Net' are used to describe the depth branch. Consider defining 'expert' on first use to avoid confusion with Mixture-of-Experts terminology.
Circularity Check
No significant circularity: the paper trains on Objaverse and evaluates on held-out GSO against external baselines, with ablations testing design choices rather than asserting them.
full rationale
The paper's claimed derivation is a feed-forward image-to-3D pipeline: a repurposed Stable Diffusion U-Net generates multi-view RGB and depth latents, a decoder with epipolar attention produces RGB, depth, and Gaussian feature maps, and back-projection plus Poisson reconstruction yields meshes and Gaussians. The central quality claims are supported by quantitative comparisons on the held-out GSO dataset against externally published baselines, using Chamfer distance, volume IoU, and rendering metrics. Training uses Objaverse renders; the novel-view-synthesis loss is a self-supervised training signal on the model's own outputs, but it is not used to derive the evaluation metrics or to fit evaluation constants, so it does not make the reported comparisons circular. Ablations (Tables 3-5) test specific components such as epipolar attention, depth latents, and branched U-Net, and they report quantitative drops, which is the opposite of assuming the design. The only evaluation-protocol concern is the use of scale-adaptive ICP with per-axis scaling in Appendix C, which may affect the validity or fairness of geometric comparisons, but this is a correctness and benchmarking issue, not a circularity of derivation. The acknowledged limitation of the orthographic-view assumption (Appendix E) likewise weakens generality but does not reduce any claimed result to its inputs. No load-bearing self-citations or imported uniqueness theorems were found. The derivation chain is therefore self-contained with respect to the paper's external evaluations.
Assumptions & free parameters
free parameters (4)
- lambda_LPIPS and lambda_gm loss weights =
0.5, 2
- Gaussian scale clamp =
0.01 and 2.5 pixels
- Opacity mask threshold =
0.1
- Diffusion guidance scale and DDIM steps =
3, 50
assumptions (4)
- domain assumption Six fixed orthographic views cover the object sufficiently for reconstruction.
- standard math Epipolar lines between the six views are rows or columns.
- domain assumption Stable Diffusion's 2D priors transfer to simultaneous RGB and depth multi-view generation after fine-tuning.
- domain assumption Models trained on synthetic Objaverse-LVIS renders generalize to GSO and in-the-wild objects.
Cite this review
Pith. "Pith review of Direct and Explicit 3D Generation from a Single Image." pith.science (2026). https://pith.science/paper/4UMAIIZ2
@misc{pith2026241110947,
author = {Pith},
title = {Pith review of: Direct and Explicit 3D Generation from a Single Image},
year = {2026},
howpublished = {\url{https://pith.science/paper/4UMAIIZ2}},
note = {Machine review of arXiv:2411.10947}
}
read the original abstract
Current image-to-3D approaches suffer from high computational costs and lack scalability for high-resolution outputs. In contrast, we introduce a novel framework to directly generate explicit surface geometry and texture using multi-view 2D depth and RGB images along with 3D Gaussian features using a repurposed Stable Diffusion model. We introduce a depth branch into U-Net for efficient and high quality multi-view, cross-domain generation and incorporate epipolar attention into the latent-to-pixel decoder for pixel-level multi-view consistency. By back-projecting the generated depth pixels into 3D space, we create a structured 3D representation that can be either rendered via Gaussian splatting or extracted to high-quality meshes, thereby leveraging additional novel view synthesis loss to further improve our performance. Extensive experiments demonstrate that our method surpasses existing baselines in geometry and texture quality while achieving significantly faster generation time.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Op- timal step nonrigid icp algorithms for surface registration
Brian Amberg, Sami Romdhani, and Thomas Vetter. Op- timal step nonrigid icp algorithms for surface registration. In 2007 IEEE conference on computer vision and pattern recognition, pages 1–8. IEEE, 2007. 3
2007
-
[2]
Renderdiffusion: Image diffusion for 3d reconstruction, in- painting and generation
Titas Anciukevi ˇcius, Zexiang Xu, Matthew Fisher, Paul Henderson, Hakan Bilen, Niloy J Mitra, and Paul Guerrero. Renderdiffusion: Image diffusion for 3d reconstruction, in- painting and generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 12608–12618, 2023. 3
2023
-
[3]
Mohammadreza Armandpour, Huangjie Zheng, Ali Sadeghian, Amir Sadeghian, and Mingyuan Zhou. Re- imagine the negative prompt algorithm: Transform 2d diffusion into 3d, alleviate janus problem and beyond. arXiv preprint arXiv:2304.04968, 2023. 3
arXiv 2023
-
[4]
Efficient geometry-aware 3d generative adversarial networks
Eric R Chan, Connor Z Lin, Matthew A Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas J Guibas, Jonathan Tremblay, Sameh Khamis, et al. Efficient geometry-aware 3d generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 16123– 16133, 2022. 1, 3
2022
-
[5]
Generative novel view synthesis with 3d-aware diffusion models
Eric R Chan, Koki Nagano, Matthew A Chan, Alexan- der W Bergman, Jeong Joon Park, Axel Levy, Miika Ait- tala, Shalini De Mello, Tero Karras, and Gordon Wetzstein. Generative novel view synthesis with 3d-aware diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 4217–4229, 2023. 3
2023
-
[6]
pixelsplat: 3d gaussian splats from im- age pairs for scalable generalizable 3d reconstruction
David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelsplat: 3d gaussian splats from im- age pairs for scalable generalizable 3d reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition , pages 19457–19467, 2024. 5
2024
-
[7]
Single-stage dif- fusion nerf: A unified approach to 3d generation and recon- struction
Hansheng Chen, Jiatao Gu, Anpei Chen, Wei Tian, Zhuowen Tu, Lingjie Liu, and Hao Su. Single-stage dif- fusion nerf: A unified approach to 3d generation and recon- struction. In Proceedings of the IEEE/CVF international conference on computer vision, pages 2416–2425, 2023. 3
2023
-
[8]
Fan- tasia3d: Disentangling geometry and appearance for high- quality text-to-3d content creation
Rui Chen, Yongwei Chen, Ningxin Jiao, and Kui Jia. Fan- tasia3d: Disentangling geometry and appearance for high- quality text-to-3d content creation. In Proceedings of the IEEE/CVF international conference on computer vision , pages 22246–22256, 2023. 3
2023
Show all 106 references
-
[9]
Mvsplat: Efficient 3d gaussian splat- ting from sparse multi-view images
Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. Mvsplat: Efficient 3d gaussian splat- ting from sparse multi-view images. arXiv preprint arXiv:2403.14627, 2024. 5
2024 arXiv
-
[10]
It3d: Improved text-to-3d generation with explicit view synthesis
Yiwen Chen, Chi Zhang, Xiaofeng Yang, Zhongang Cai, Gang Yu, Lei Yang, and Guosheng Lin. It3d: Improved text-to-3d generation with explicit view synthesis. In Pro- ceedings of the AAAI Conference on Artificial Intelligence, pages 1237–1244, 2024. 3
2024
-
[11]
Sdfusion: Multimodal 3d shape completion, reconstruction, and generation
Yen-Chi Cheng, Hsin-Ying Lee, Sergey Tulyakov, Alexan- der G Schwing, and Liang-Yan Gui. Sdfusion: Multimodal 3d shape completion, reconstruction, and generation. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 4456–4465, 2023. 3
2023
-
[12]
Blender - a 3D modelling and rendering package
Blender Online Community. Blender - a 3D modelling and rendering package. Blender Foundation, Stichting Blender Foundation, Amsterdam, 2018. 6
2018
-
[13]
FlashAttention-2: Faster attention with better par- allelism and work partitioning
Tri Dao. FlashAttention-2: Faster attention with better par- allelism and work partitioning. 2023. 5
2023
-
[14]
Fu, Stefano Ermon, Atri Rudra, and Christopher R´e
Tri Dao, Daniel Y . Fu, Stefano Ermon, Atri Rudra, and Christopher R´e. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. In Advances in Neural Information Processing Systems, 2022. 5
2022
-
[15]
Objaverse: A universe of annotated 3d objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...
2023
-
[16]
Nerdi: Single-view nerf synthesis with language-guided diffusion as general image priors
Congyue Deng, Chiyu Jiang, Charles R Qi, Xinchen Yan, Yin Zhou, Leonidas Guibas, Dragomir Anguelov, et al. Nerdi: Single-view nerf synthesis with language-guided diffusion as general image priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni...
2023
-
[17]
Google scanned objects: A high-quality dataset of 3d scanned household items
Laura Downs, Anthony Francis, Nate Koenig, Brandon Kinman, Ryan Hickman, Krista Reymann, Thomas B McHugh, and Vincent Vanhoucke. Google scanned objects: A high-quality dataset of 3d scanned household items. In 2022 International Conference on Robotics and Automa- tion (ICRA), ...
2022
-
[18]
Omnidata: A scalable pipeline for making multi- task mid-level vision datasets from 3d scans
Ainaz Eftekhar, Alexander Sax, Jitendra Malik, and Amir Zamir. Omnidata: A scalable pipeline for making multi- task mid-level vision datasets from 3d scans. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 10786–10796, 2021. 1
2021
-
[19]
Hyperdiffusion: Generating implicit neu- ral fields with weight-space diffusion
Ziya Erkoc ¸, Fangchang Ma, Qi Shan, Matthias Nießner, and Angela Dai. Hyperdiffusion: Generating implicit neu- ral fields with weight-space diffusion. In Proceedings of the IEEE/CVF international conference on computer vi- sion, pages 14300–14310, 2023. 3
2023
-
[20]
Cat3d: Create anything in 3d with multi-view diffusion models
Ruiqi Gao, Aleksander Holynski, Philipp Henzler, Arthur Brussee, Ricardo Martin-Brualla, Pratul Srinivasan, Jonathan T Barron, and Ben Poole. Cat3d: Create anything in 3d with multi-view diffusion models. arXiv preprint arXiv:2405.10314, 2024. 2, 3
2024 arXiv
-
[21]
Nerfdiff: Single-image view synthesis with nerf-guided distillation from 3d-aware diffusion
Jiatao Gu, Alex Trevithick, Kai-En Lin, Joshua M Susskind, Christian Theobalt, Lingjie Liu, and Ravi Ra- mamoorthi. Nerfdiff: Single-image view synthesis with nerf-guided distillation from 3d-aware diffusion. In Inter- national Conference on Machine Learning , pages 11808– 118...
2023
-
[22]
Control3diff: Learning controllable 3d diffusion models from single-view images
Jiatao Gu, Qingzhe Gao, Shuangfei Zhai, Baoquan Chen, Lingjie Liu, and Josh Susskind. Control3diff: Learning controllable 3d diffusion models from single-view images. In 2024 International Conference on 3D Vision (3DV) , pages 685–696. IEEE, 2024. 3
2024
-
[23]
Sugar: Surface- aligned gaussian splatting for efficient 3d mesh reconstruc- tion and high-quality mesh rendering
Antoine Gu ´edon and Vincent Lepetit. Sugar: Surface- aligned gaussian splatting for efficient 3d mesh reconstruc- tion and high-quality mesh rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5354–5363, 2024. 5
2024
-
[24]
3dgen: Triplane latent diffusion for textured mesh generation
Anchit Gupta, Wenhan Xiong, Yixin Nie, Ian Jones, and Barlas O ˘guz. 3dgen: Triplane latent diffusion for textured mesh generation. arXiv preprint arXiv:2303.05371, 2023. 1, 3
2023 arXiv
-
[25]
Multiple view ge- ometry in computer vision
Richard Hartley and Andrew Zisserman. Multiple view ge- ometry in computer vision . Cambridge university press,
-
[26]
Epipolar transformers
Yihui He, Rui Yan, Katerina Fragkiadaki, and Shoou-I Yu. Epipolar transformers. In Proceedings of the ieee/cvf con- ference on computer vision and pattern recognition , pages 7779–7788, 2020. 2, 4
2020
-
[27]
Openlrm: Open-source large reconstruction models, 2023
Zexin He and Tengfei Wang. Openlrm: Open-source large reconstruction models, 2023. 3, 7
2023
-
[28]
Lrm: Large reconstruction model for single image to 3d
Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. Lrm: Large reconstruction model for single image to 3d. arXiv preprint arXiv:2311.04400, 2023. 3, 7
2023 arXiv
-
[29]
Mvd-fusion: Single-view 3d via depth-consistent multi-view generation
Hanzhe Hu, Zhizhuo Zhou, Varun Jampani, and Shubham Tulsiani. Mvd-fusion: Single-view 3d via depth-consistent multi-view generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 9698–9707, 2024. 3, 7, 8
2024
-
[30]
Dreamtime: An improved optimization strategy for text-to-3d content creation
Yukun Huang, Jianan Wang, Yukai Shi, Xianbiao Qi, Zheng-Jun Zha, and Lei Zhang. Dreamtime: An improved optimization strategy for text-to-3d content creation. arXiv preprint arXiv:2306.12422, 2023. 3
2023 arXiv
-
[31]
Shap-e: Generat- ing conditional 3d implicit functions
Heewoo Jun and Alex Nichol. Shap-e: Generat- ing conditional 3d implicit functions. arXiv preprint arXiv:2305.02463, 2023. 1, 3, 7
2023 arXiv
-
[32]
Spad: Spatially aware multi-view diffusers
Yash Kant, Aliaksandr Siarohin, Ziyi Wu, Michael Vasilkovsky, Guocheng Qian, Jian Ren, Riza Alp Guler, Bernard Ghanem, Sergey Tulyakov, and Igor Gilitschenski. Spad: Spatially aware multi-view diffusers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- ter...
2024
-
[33]
Holofusion: Towards photo-realistic 3d generative modeling
Animesh Karnewar, Niloy J Mitra, Andrea Vedaldi, and David Novotny. Holofusion: Towards photo-realistic 3d generative modeling. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision, pages 22976– 22985, 2023. 3
2023
-
[34]
Screened poisson surface reconstruction
Michael Kazhdan and Hugues Hoppe. Screened poisson surface reconstruction. ACM Transactions on Graphics (ToG), 32(3):1–13, 2013. 2, 5, 1, 3
2013
-
[35]
Re- purposing diffusion-based image generators for monocular depth estimation
Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Met- zger, Rodrigo Caye Daudt, and Konrad Schindler. Re- purposing diffusion-based image generators for monocular depth estimation. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition , pages 94...
2024
-
[36]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42(4), 2023. 2, 4, 5, 7
2023
-
[37]
Neuralfield-ldm: Scene gen- eration with hierarchical latent diffusion models
Seung Wook Kim, Bradley Brown, Kangxue Yin, Karsten Kreis, Katja Schwarz, Daiqing Li, Robin Rombach, Anto- nio Torralba, and Sanja Fidler. Neuralfield-ldm: Scene gen- eration with hierarchical latent diffusion models. In Pro- ceedings of the IEEE/CVF Conference on Computer Vis...
2023
-
[38]
Stable diffusion image variations - a hugging face space
Lambda Labs. Stable diffusion image variations - a hugging face space. https : / / huggingface . co / lambdalabs / sd - image - variations - diffusers. 2, 5
-
[39]
xformers: A modular and hackable transformer modelling library, 2022
Benjamin Lefaudeux, Francisco Massa, Diana Liskovich, Wenhan Xiong, Vittorio Caggiano, Sean Naren, Min Xu, Jieru Hu, Marta Tintore, Susan Zhang, et al. xformers: A modular and hackable transformer modelling library, 2022. 5
2022
-
[40]
Generative scene synthesis via incremental view inpainting using rgbd diffu- sion models
Jiabao Lei, Jiapeng Tang, and Kui Jia. Generative scene synthesis via incremental view inpainting using rgbd diffu- sion models. arXiv preprint arXiv:2212.05993, 2022. 3
2022 arXiv
-
[41]
Era3d: High-resolution multiview diffusion using efficient row-wise attention
Peng Li, Yuan Liu, Xiaoxiao Long, Feihu Zhang, Cheng Lin, Mengfei Li, Xingqun Qi, Shanghang Zhang, Wenhan Luo, Ping Tan, et al. Era3d: High-resolution multiview diffusion using efficient row-wise attention. arXiv preprint arXiv:2405.11616, 2024. 4, 3
2024 arXiv
-
[42]
Sweet- dreamer: Aligning geometric priors in 2d diffusion for con- sistent text-to-3d
Weiyu Li, Rui Chen, Xuelin Chen, and Ping Tan. Sweet- dreamer: Aligning geometric priors in 2d diffusion for con- sistent text-to-3d. arXiv preprint arXiv:2310.02596, 2023. 3
2023 arXiv
-
[43]
Magic3d: High- resolution text-to-3d content creation
Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High- resolution text-to-3d content creation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti...
2023
-
[44]
One-2-3-45: Any sin- gle image to 3d mesh in 45 seconds without per-shape opti- mization
Minghua Liu, Chao Xu, Haian Jin, Linghao Chen, Mukund Varma T, Zexiang Xu, and Hao Su. One-2-3-45: Any sin- gle image to 3d mesh in 45 seconds without per-shape opti- mization. Advances in Neural Information Processing Sys- tems, 36, 2024. 3, 6, 7
2024
-
[45]
Zero-1-to-3: Zero-shot one image to 3d object
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tok- makov, Sergey Zakharov, and Carl V ondrick. Zero-1-to-3: Zero-shot one image to 3d object. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 9298–9309, 2023. 1, 3, 5, 7
2023
-
[46]
Deceptive-nerf: Enhancing nerf reconstruction using pseudo-observations from diffusion models
Xinhang Liu, Shiu-hong Kao, Jiaben Chen, Yu-Wing Tai, and Chi-Keung Tang. Deceptive-nerf: Enhancing nerf reconstruction using pseudo-observations from diffusion models. arXiv preprint arXiv:2305.15171, 2023. 3
2023 arXiv
-
[47]
Hyperhuman: Hyper-realistic human generation with latent structural diffusion
Xian Liu, Jian Ren, Aliaksandr Siarohin, Ivan Sko- rokhodov, Yanyu Li, Dahua Lin, Xihui Liu, Ziwei Liu, and Sergey Tulyakov. Hyperhuman: Hyper-realistic human generation with latent structural diffusion. arXiv preprint arXiv:2310.08579, 2023. 4
-
[48]
Syncdreamer: Generating multiview-consistent images from a single-view image
Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Generating multiview-consistent images from a single-view image. arXiv preprint arXiv:2309.03453, 2023. 1, 2, 3, 4, 6, 7
2023 arXiv
-
[49]
Pi3d: Efficient text-to-3d gen- eration with pseudo-image diffusion
Ying-Tian Liu, Yuan-Chen Guo, Guan Luo, Heyi Sun, Wei Yin, and Song-Hai Zhang. Pi3d: Efficient text-to-3d gen- eration with pseudo-image diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19915–19924, 2024. 3
2024
-
[50]
Meshdif- fusion: Score-based generative 3d mesh modeling
Zhen Liu, Yao Feng, Michael J Black, Derek Nowrouzezahrai, Liam Paull, and Weiyang Liu. Meshdif- fusion: Score-based generative 3d mesh modeling. arXiv preprint arXiv:2303.08133, 2023. 1, 3
2023 arXiv
-
[51]
Wonder3d: Single image to 3d using cross-domain diffusion
Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, et al. Wonder3d: Single image to 3d using cross-domain diffusion. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Patt...
2024
-
[52]
Diffusion probabilistic mod- els for 3d point cloud generation
Shitong Luo and Wei Hu. Diffusion probabilistic mod- els for 3d point cloud generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2837–2845, 2021. 1
2021
-
[53]
Realfusion: 360deg reconstruction of any object from a single image
Luke Melas-Kyriazi, Iro Laina, Christian Rupprecht, and Andrea Vedaldi. Realfusion: 360deg reconstruction of any object from a single image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8446–8455, 2023. 1, 6, 7
2023
-
[54]
Nerf: Representing scenes as neural radiance fields for view syn- thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. Communications of the ACM, 65(1):99–106, 2021. 1, 3
2021
-
[55]
Diffrf: Rendering-guided 3d radiance field diffusion
Norman M ¨uller, Yawar Siddiqui, Lorenzo Porzi, Samuel Rota Bulo, Peter Kontschieder, and Matthias Nießner. Diffrf: Rendering-guided 3d radiance field diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 4328–4338, 2023. 3
2023
-
[56]
Point-e: A system for generat- ing 3d point clouds from complex prompts
Alex Nichol, Heewoo Jun, Prafulla Dhariwal, Pamela Mishkin, and Mark Chen. Point-e: A system for generat- ing 3d point clouds from complex prompts. arXiv preprint arXiv:2212.08751, 2022. 1, 3, 7
2022 arXiv
-
[57]
Au- todecoding latent 3d diffusion models
Evangelos Ntavelis, Aliaksandr Siarohin, Kyle Olszewski, Chaoyang Wang, Luc V Gool, and Sergey Tulyakov. Au- todecoding latent 3d diffusion models. Advances in Neural Information Processing Systems, 36, 2024. 3
2024
-
[58]
Dreamfusion: Text-to-3d using 2d diffusion
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022. 3
2022 arXiv
-
[59]
Magic123: One image to high-quality 3d object generation us- ing both 2d and 3d diffusion priors
Guocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren, Aliaksandr Siarohin, Bing Li, Hsin-Ying Lee, Ivan Sko- rokhodov, Peter Wonka, Sergey Tulyakov, et al. Magic123: One image to high-quality 3d object generation us- ing both 2d and 3d diffusion priors. arXiv preprint arXiv:230...
2023 arXiv
-
[60]
Richdreamer: A gen- eralizable normal-depth diffusion model for detail richness in text-to-3d
Lingteng Qiu, Guanying Chen, Xiaodong Gu, Qi Zuo, Mutian Xu, Yushuang Wu, Weihao Yuan, Zilong Dong, Liefeng Bo, and Xiaoguang Han. Richdreamer: A gen- eralizable normal-depth diffusion model for detail richness in text-to-3d. In Proceedings of the IEEE/CVF Conference on Comput...
2024
-
[61]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In International conference on machine learning...
2021
-
[62]
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In Interna- tional Conference on Machine Learning, pages 8821–8831. PMLR, 2021. 3
2021
-
[63]
Towards robust monocu- lar depth estimation: Mixing datasets for zero-shot cross- dataset transfer
Ren ´e Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocu- lar depth estimation: Mixing datasets for zero-shot cross- dataset transfer. IEEE transactions on pattern analysis and machine intelligence, 44(3):1623–1637, 2020. 5, 1
2020
-
[64]
Vi- sion transformers for dense prediction
Ren ´e Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vi- sion transformers for dense prediction. In Proceedings of the IEEE/CVF international conference on computer vi- sion, pages 12179–12188, 2021. 1
2021
-
[65]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2, 3, 4, 5
2022
-
[66]
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Sali- mans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural In- forma...
2022
-
[67]
Scale-adaptive icp
Yusuf Sahillio ˘glu and Ladislav Kavan. Scale-adaptive icp. Graphical Models, 116:101113, 2021. 6, 3
2021
-
[68]
Break- ing good: Fracture modes for realtime destruction
Silvia Sell ´an, Jack Luong, Leticia Mattos Da Silva, Aravind Ramakrishnan, Yuchuan Yang, and Alec Jacobson. Break- ing good: Fracture modes for realtime destruction. ACM Transactions on Graphics, 2022. 3
2022
-
[69]
Ditto-nerf: Diffusion-based iterative text to omni- directional 3d model
Hoigi Seo, Hayeon Kim, Gwanghyun Kim, and Se Young Chun. Ditto-nerf: Diffusion-based iterative text to omni- directional 3d model. arXiv preprint arXiv:2304.02827 ,
-
[70]
Let 2d diffusion model know 3d- consistency for robust text-to-3d generation
Junyoung Seo, Wooseok Jang, Min-Seop Kwak, Jaehoon Ko, Hyeonsu Kim, Junho Kim, Jin-Hwa Kim, Jiyoung Lee, and Seungryong Kim. Let 2d diffusion model know 3d- consistency for robust text-to-3d generation. arXiv preprint arXiv:2303.07937, 2023. 3
2023 arXiv
-
[71]
Mvdream: Multi-view diffusion for 3d generation
Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d generation. arXiv preprint arXiv:2308.16512, 2023. 3, 4, 5
2023 arXiv
-
[72]
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. 5
2010 arXiv
-
[73]
Ldm3d: Latent diffusion model for 3d
Gabriela Ben Melech Stan, Diana Wofk, Scottie Fox, Alex Redden, Will Saxton, Jean Yu, Estelle Aflalo, Shao-Yen Tseng, Fabio Nonato, Matthias Muller, et al. Ldm3d: Latent diffusion model for 3d. arXiv preprint arXiv:2305.10853, 2023. 3
2023 arXiv
-
[74]
Splatter image: Ultra-fast single-view 3d recon- struction
Stanislaw Szymanowicz, Christian Rupprecht, and Andrea Vedaldi. Splatter image: Ultra-fast single-view 3d recon- struction. arXiv preprint arXiv:2312.13150, 2023. 3, 5
2023 arXiv
-
[75]
Viewset diffusion:(0-) image-conditioned 3d generative models from 2d data
Stanislaw Szymanowicz, Christian Rupprecht, and An- drea Vedaldi. Viewset diffusion:(0-) image-conditioned 3d generative models from 2d data. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 8863–8873, 2023. 3
2023
-
[76]
Lgm: Large multi- view gaussian model for high-resolution 3d content cre- ation
Jiaxiang Tang, Zhaoxi Chen, Xiaokang Chen, Tengfei Wang, Gang Zeng, and Ziwei Liu. Lgm: Large multi- view gaussian model for high-resolution 3d content cre- ation. arXiv preprint arXiv:2402.05054, 2024. 3, 5, 7, 8
2024 arXiv
-
[77]
Diffusion with forward models: Solving stochastic inverse problems without direct supervi- sion
Ayush Tewari, Tianwei Yin, George Cazenavette, Semon Rezchikov, Josh Tenenbaum, Fr´edo Durand, Bill Freeman, and Vincent Sitzmann. Diffusion with forward models: Solving stochastic inverse problems without direct supervi- sion. Advances in Neural Information Processing Systems...
2024
-
[78]
Textmesh: Gen- eration of realistic 3d meshes from text prompts
Christina Tsalicoglou, Fabian Manhardt, Alessio Tonioni, Michael Niemeyer, and Federico Tombari. Textmesh: Gen- eration of realistic 3d meshes from text prompts. In 2024 International Conference on 3D Vision (3DV), pages 1554–
2024
-
[79]
Consistent view syn- thesis with pose-guided diffusion models
Hung-Yu Tseng, Qinbo Li, Changil Kim, Suhib Alsisan, Jia-Bin Huang, and Johannes Kopf. Consistent view syn- thesis with pose-guided diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 16773–16783, 2023. 3
2023
-
[80]
Lion: Latent point diffu- sion models for 3d shape generation
Arash Vahdat, Francis Williams, Zan Gojcic, Or Litany, Sanja Fidler, Karsten Kreis, et al. Lion: Latent point diffu- sion models for 3d shape generation. Advances in Neural Information Processing Systems , 35:10021–10039, 2022. 1, 3
2022
-
[81]
Sv3d: Novel multi-view synthesis and 3d generation from a single image using latent video diffusion
Vikram V oleti, Chun-Han Yao, Mark Boss, Adam Letts, David Pankratz, Dmitry Tochilkin, Christian Laforte, Robin Rombach, and Varun Jampani. Sv3d: Novel multi-view synthesis and 3d generation from a single image using latent video diffusion. arXiv preprint arXiv:2403.12008 ,
-
[82]
Imagedream: Image-prompt multi-view diffusion for 3d generation
Peng Wang and Yichun Shi. Imagedream: Image-prompt multi-view diffusion for 3d generation. arXiv preprint arXiv:2312.02201, 2023. 3
2023 arXiv
-
[83]
Rodin: A generative model for sculpting 3d digital avatars using diffusion
Tengfei Wang, Bo Zhang, Ting Zhang, Shuyang Gu, Jian- min Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen, Fang Wen, Qifeng Chen, et al. Rodin: A generative model for sculpting 3d digital avatars using diffusion. In Proceed- ings of the IEEE/CVF Conference on Computer Vision...
2023
-
[84]
Mvster: Epipolar transformer for efficient multi-view stereo
Xiaofeng Wang, Zheng Zhu, Guan Huang, Fangbo Qin, Yun Ye, Yijia He, Xu Chi, and Xingang Wang. Mvster: Epipolar transformer for efficient multi-view stereo. In Eu- ropean Conference on Computer Vision , pages 573–591. Springer, 2022. 2, 4
2022
-
[85]
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image pro- cessing, 13(4):600–612, 2004. 7
2004
-
[86]
Mvdd: Multi-view depth diffusion mod- els
Zhen Wang, Qiangeng Xu, Feitong Tan, Menglei Chai, Shichen Liu, Rohit Pandey, Sean Fanello, Achuta Kadambi, and Yinda Zhang. Mvdd: Multi-view depth diffusion mod- els. arXiv preprint arXiv:2312.04875, 2023. 4
2023 arXiv
-
[87]
Prolificdreamer: High- fidelity and diverse text-to-3d generation with variational score distillation
Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongx- uan Li, Hang Su, and Jun Zhu. Prolificdreamer: High- fidelity and diverse text-to-3d generation with variational score distillation. Advances in Neural Information Process- ing Systems, 36, 2024. 3
2024
-
[88]
Crm: Single image to 3d textured mesh with convolutional reconstruction model
Zhengyi Wang, Yikai Wang, Yifei Chen, Chendong Xi- ang, Shuo Chen, Dajiang Yu, Chongxuan Li, Hang Su, and Jun Zhu. Crm: Single image to 3d textured mesh with convolutional reconstruction model. arXiv preprint arXiv:2403.05034, 2024. 3, 6, 7
2024 arXiv
-
[89]
Novel view synthesis with diffusion models
Daniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho, Andrea Tagliasacchi, and Mohammad Norouzi. Novel view synthesis with diffusion models. arXiv preprint arXiv:2210.04628, 2022. 3
2022 arXiv
-
[90]
latentsplat: Autoencoding variational gaussians for fast generalizable 3d reconstruction
Christopher Wewer, Kevin Raj, Eddy Ilg, Bernt Schiele, and Jan Eric Lenssen. latentsplat: Autoencoding variational gaussians for fast generalizable 3d reconstruction. arXiv preprint arXiv:2403.16292, 2024. 5
2024 arXiv
-
[91]
Hd- fusion: Detailed text-to-3d generation leveraging multiple noise estimation
Jinbo Wu, Xiaobo Gao, Xing Liu, Zhengyang Shen, Chen Zhao, Haocheng Feng, Jingtuo Liu, and Errui Ding. Hd- fusion: Detailed text-to-3d generation leveraging multiple noise estimation. In Proceedings of the IEEE/CVF Win- ter Conference on Applications of Computer Vision , pages...
2024
-
[92]
3d-aware image generation using 2d diffusion mod- els
Jianfeng Xiang, Jiaolong Yang, Binbin Huang, and Xin Tong. 3d-aware image generation using 2d diffusion mod- els. In Proceedings of the IEEE/CVF International Confer- ence on Computer Vision, pages 2383–2393, 2023. 3
2023
-
[93]
Agg: Amor- tized generative 3d gaussians for single image to 3d
Dejia Xu, Ye Yuan, Morteza Mardani, Sifei Liu, Jiaming Song, Zhangyang Wang, and Arash Vahdat. Agg: Amor- tized generative 3d gaussians for single image to 3d. arXiv preprint arXiv:2401.04099, 2024. 3
2024 arXiv
-
[94]
Instantmesh: Efficient 3d mesh generation from a single image with sparse-view large reconstruction models
Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang, Shenghua Gao, and Ying Shan. Instantmesh: Efficient 3d mesh generation from a single image with sparse-view large reconstruction models. arXiv preprint arXiv:2404.07191 ,
-
[95]
Grm: Large gaussian reconstruction model for ef- ficient 3d reconstruction and generation
Yinghao Xu, Zifan Shi, Wang Yifan, Hansheng Chen, Ceyuan Yang, Sida Peng, Yujun Shen, and Gordon Wet- zstein. Grm: Large gaussian reconstruction model for ef- ficient 3d reconstruction and generation. arXiv preprint arXiv:2403.14621, 2024. 3, 5
2024 arXiv
-
[96]
Mvs2d: Efficient multi-view stereo via attention-driven 2d convolutions
Zhenpei Yang, Zhile Ren, Qi Shan, and Qixing Huang. Mvs2d: Efficient multi-view stereo via attention-driven 2d convolutions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8574– 8584, 2022. 2, 4
2022
-
[97]
Dreamsparse: Escaping from plato’s cave with 2d dif- fusion model given sparse views
Paul Yoo, Jiaxian Guo, Yutaka Matsuo, and Shixiang Shane Gu. Dreamsparse: Escaping from plato’s cave with 2d dif- fusion model given sparse views. Advances in Neural In- formation Processing Systems, 36, 2024. 3
2024
-
[98]
https://github.com/jpcy/xatlas, 2021
Jonathan Young. https://github.com/jpcy/xatlas, 2021. 5
2021
-
[99]
Points-to-3d: Bridging the gap be- tween sparse points and shape-controllable text-to-3d gen- eration
Chaohui Yu, Qiang Zhou, Jingliang Li, Zhe Zhang, Zhibin Wang, and Fan Wang. Points-to-3d: Bridging the gap be- tween sparse points and shape-controllable text-to-3d gen- eration. In Proceedings of the 31st ACM International Con- ference on Multimedia, pages 6841–6850, 2023. 3
2023
-
[100]
Long-term photometric consistent novel view synthesis with diffusion models
Jason J Yu, Fereshteh Forghani, Konstantinos G Derpanis, and Marcus A Brubaker. Long-term photometric consistent novel view synthesis with diffusion models. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 7094–7104, 2023. 3
2023
-
[101]
3dshape2vecset: A 3d shape representation for neural fields and generative diffusion models
Biao Zhang, Jiapeng Tang, Matthias Niessner, and Peter Wonka. 3dshape2vecset: A 3d shape representation for neural fields and generative diffusion models. ACM Trans- actions on Graphics (TOG), 42(4):1–16, 2023. 3
2023
-
[102]
Gs-lrm: Large reconstruction model for 3d gaussian splatting
Kai Zhang, Sai Bi, Hao Tan, Yuanbo Xiangli, Nanxuan Zhao, Kalyan Sunkavalli, and Zexiang Xu. Gs-lrm: Large reconstruction model for 3d gaussian splatting. arXiv preprint arXiv:2404.19702, 2024. 3, 5
2024 arXiv
-
[103]
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 586–595, 2018. 5, 7
2018
-
[104]
Sparsefusion: Dis- tilling view-conditioned diffusion for 3d reconstruction
Zhizhuo Zhou and Shubham Tulsiani. Sparsefusion: Dis- tilling view-conditioned diffusion for 3d reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition , pages 12588–12597, 2023. 3
2023
-
[105]
Hifa: High-fidelity text- to-3d with advanced diffusion guidance
Joseph Zhu and Peiye Zhuang. Hifa: High-fidelity text- to-3d with advanced diffusion guidance. arXiv preprint arXiv:2305.18766, 2023. 3
2023 arXiv
-
[106]
Triplane meets gaussian splatting: Fast and generalizable single- view 3d reconstruction with transformers
Zi-Xin Zou, Zhipeng Yu, Yuan-Chen Guo, Yangguang Li, Ding Liang, Yan-Pei Cao, and Song-Hai Zhang. Triplane meets gaussian splatting: Fast and generalizable single- view 3d reconstruction with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- t...
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.