REVIEW 4 major objections 5 minor 2 cited by
Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Prometheus claims that text-to-3D scene generation can be done in a single feed-forward pass: a latent diffusion model generates multi-view RGB-D latents that decode directly into a scene-level 3D Gaussian set in about 8 seconds.
desk verdict Plausible feed-forward text-to-3D systems paper with real speed and a genuine combination of known ingredients; the abstract overclaims scene-level generalization, but a revised version with honest scoping and stronger evaluation could merit acceptance. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the RGB-D latent space together with the pixel-aligned 3D Gaussian decoder. A frozen Stable Diffusion encoder maps images and predicted depth maps into per-view latents; a cross-view transformer with Plücker ray maps fuses multi-view context; and a modified Stable Diffusion decoder outputs per-pixel 3D Gaussians, splat primitives with position, rotation, scale, opacity, and color. This machinery lets the second-stage denoiser generate multi-view RGB-D latents jointly from text and camera poses, so that decoding alone produces a scene-level 3D Gaussian set without per-scene optimization.
What would settle it
Render one generated 3D scene from two cameras separated by roughly 90 degrees and measure geometric consistency by reprojecting rendered depth maps into 3D and computing point-cloud overlap or depth error; if large floating artifacts or duplicated structures appear, the claimed consistency at large rotations fails.
Extended reading notes
Core claim
Prometheus claims that text-to-3D scene generation can be reduced to multi-view, feed-forward, pixel-aligned 3D Gaussian generation in latent space. The system first trains a GS-VAE: multi-view RGB images and monocular depth maps are encoded by a frozen Stable Diffusion encoder, fused by a cross-view transformer conditioned on Plücker ray maps, and decoded by a modified Stable Diffusion decoder into per-pixel 3D Gaussians, each parameterized by depth, a rotation quaternion, scale, opacity, and spherical harmonics coefficients ($C_G=12$). It then trains a multi-view latent diffusion model that, from text and camera poses, denoises jointly predicted RGB-D latents at relatively high noise levels, using cross-view self-attention and hybrid CFG sampling. The authors claim this produces a complete 3D scene in about 8 seconds, and that ablations show the RGB-D latent space, single-view training data, high noise level, and hybrid sampling each contribute to the final geometry and fidelity.
Load-bearing premise
The load-bearing premise is that a latent-space diffusion model, conditioned only on ray maps and cross-view attention, can produce multi-view outputs consistent enough that their decoded 3D Gaussians form a coherent scene without any explicit 3D constraint during generation.
Editorial extensions
If this is right
- A user can go from a text prompt to a renderable 3D Gaussian scene of an object or an indoor/outdoor scene in about 8 seconds, without per-scene optimization or a separate reconstruction step.
- Because each generated pixel carries its own Gaussian with depth, geometry is produced at the same time as appearance rather than inferred afterward from images.
- Training on single-view images alongside multi-view data is what extends generalization to open-world content; removing the single-view data degrades both reconstruction and generation quality, as shown in the ablations.
- High-noise training of the multi-view denoiser and hybrid CFG sampling are both required for multi-view consistency; the ablations show image quality and CLIP score drop when either is removed.
Reading between the lines
- The paper's own failure analysis shows view inconsistency under large rotations or extreme viewpoints because no explicit 3D representation is used during latent generation; a natural next step would be to add a render-and-compare loss on the decoded Gaussians as part of the diffusion loop, or to condition on an explicit coarse 3D scaffold.
- Because the RGB-D latent space takes monocular depth estimates as input, its geometry is bounded by the quality and scale ambiguity of those estimates; a testable extension is to compare depth metrics when the depth source is swapped or when multi-view geometry is used to refine pseudo depth.
- The architecture separates the autoencoder from the diffusion prior, so the GS-VAE could in principle decode latents from other multi-view generators, such as video models or single-image novel-view models, turning them into 3D scenes without re-training the decoder.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces Prometheus, a feed-forward text-to-3D generation method that operates at both object and scene levels. The method is organized in two stages: a GS-VAE (Stage 1) encodes multi-view RGB-D images into a latent space via a frozen Stable Diffusion encoder, fuses them with a cross-view transformer using Plücker ray maps, and decodes pixel-aligned 3D Gaussians; a multi-view latent diffusion model (Stage 2) denoises multi-view RGB-D latents conditioned on text and camera poses, and the resulting latents are decoded into a 3D Gaussian scene. The model is trained on a large mixture of single-view and multi-view datasets. Experiments report generalizable 3D reconstruction on TartanAir and text-to-3D generation on T3Bench prompts with comparisons to GaussianDreamer, MVDream+LGM, and Director3D, with ablations of the depth prior, single-view data, noise level, and hybrid sampling strategies.
Significance. If the central feasibility claim is fully supported, the paper makes a useful contribution: it demonstrates a feed-forward text-to-3D pipeline that runs in about eight seconds, reuses a large pre-trained text-to-image model (Stable Diffusion), and combines single-view and multi-view training data to broaden generalization. The RGB-D latent space and the two-stage latent-diffusion formulation are sensible design choices, and the ablations in Tables 4 and 5 provide useful information about which components matter. However, the quantitative evidence as presented is weaker than the prose suggests: the main generation comparison in Table 3 shows Prometheus behind Director3D on CLIP-Score for both T3Bench subsets, no error bars or significance tests appear in Tables 2-5, and the paper's own supplementary material admits a multi-view inconsistency failure mode in exactly the large-rotation regime that scene-level generation must handle. The central claim is therefore defensible but not yet fully demonstrated; the paper needs additional evidence and a more careful scoping of the claims.
major comments (4)
- [Section 4.4, Table 3] The claim that Prometheus outperforms baselines in text-to-3D generation is not supported by Table 3 on the standard T3Bench subsets. On Single-Object, Director3D achieves CLIP-Score 0.397 versus Prometheus 0.329, and on Single-Object-with-Surroundings the comparison is 0.405 versus 0.369. Prometheus leads only on the self-collected Scene-Level set (0.370 versus 0.357), whose prompts are not a standard benchmark. The metrics used (BRISQUE, NIQE, CLIP-Score) are no-reference image-quality and text-alignment measures; they do not evaluate multi-view consistency or geometric correctness of the generated 3D scene. The text should either add a more comprehensive evaluation, including user studies or 3D consistency metrics, or substantially soften the claim that the method is state of the art for both object- and scene-level generation.
- [Supplementary Section B, Eq. (11)] The paper's own limitation statement in Fig. 9 says that 'due to the lack of explicit 3D representation during multiview generation in latent space, Prometheus will encounter view inconsistency under large rotations or extreme viewpoints.' This directly bears on the abstract's claim of generalizable scene-level generation, because large rotations and extreme viewpoints are precisely the regimes where scene-level 3D content must remain consistent. Equation (11) only supervises the denoised RGB-D latents with an L2 loss against the clean latents; there is no 3D consistency loss, and the fused Gaussian geometry is never checked for cross-view agreement during Stage 2. The main text does not scope the scene-level claim to narrow-baseline trajectories, and no quantitative evaluation of multi-view consistency or geometry of generated scenes is reported. I recommend either adding a consistency evaluation (e.g., rendering depth from generated Gaussians and measuring cross-view agreement, or evaluating generated scenes on a multi-view reconstruction benchmark) or explicitly restricting the claims to small-baseline camera trajectories.
- [Section 4.3, Table 2] The Stage-1 reconstruction evaluation is based solely on TartanAir, a synthetic dataset, and reports no error bars or statistical significance. On Easy mode, Prometheus has lower PSNR and SSIM than pixelSplat (20.95/0.589 versus 21.65/0.681), with only LPIPS and depth metrics better; the statement that results are 'comparable on Easy mode' is acceptable, but the stronger claim that the method 'notably outperforms' on harder modes is based on one dataset with no variance estimates. Since Stage 1 is the backbone for the downstream generation task, the paper should include at least one real-world multi-view dataset (e.g., RealEstate10K or DL3DV) and report confidence intervals or per-scene standard deviations.
- [Section 3.2, Eqs. (14) and (16)] Equation (14) appears to contain a technical error: it writes Zt-1 = Zt - G_theta(Z_T; sigma_t, y, R)/sigma_t * (sigma_{t-1} - sigma_t) + Zt, which is algebraically 2Zt minus the update term and, more importantly, evaluates the denoiser at Z_T rather than Z_t at every step. This is inconsistent with Eq. (10), where the denoiser takes Z_t as input, and it makes the sampling procedure in Sec. 3.3, Eq. (15), not directly reproducible. Additionally, Eq. (16) writes the CFG combination as w*G_theta(Zt; y,R) + (w-1)*G_theta(Zt; R); the standard classifier-free guidance formula uses the coefficient (1-w) on the unconditional term. If this is a typographical convention, it should be clarified; if it is the actual update, the derivation should be justified, because Eq. (17) then defines a different combined guidance rule with w = w1 + w2.
minor comments (5)
- [Throughout] There are several typos and grammatical errors, including 'Sceonds' in the Section 3.3 title, 'classfier-free-guidance' in the same section, 'condition respectfully' in Eq. (10) and nearby text, and 'noise level noise Gaussian noise' in Section 3.2. The paper should be carefully proofread.
- [Table 1 caption] The caption says '9 multi-view datasets' but the table lists ten datasets; SAM-1B is single-view, so the wording should be clarified, for example as '9 multi-view datasets plus a single-view dataset.'
- [Figure 8 caption] The caption of Figure 8 appears to be copied from Figure 3; it refers to 'overlap gradually decreases' and 'depth map' in a way that does not match the figure content, which is a qualitative comparison with Director3D. The caption should be rewritten.
- [Section 4.4] The text says the 'relative enhancement of 44% on Easy mode and a substantial 64% on Hard mode' for delta1 against pixelSplat, but it does not state explicitly that this is the relative improvement in the delta1 metric; the reader has to infer this from Table 2. Please make this explicit.
- [Section 4.3 and Table 3] The 'Scene-Level' column in Table 3 is not from T3Bench; the text states that 80 diverse scene-level prompts were collected by the authors. This should be stated clearly in the table caption and in the evaluation protocol so that readers do not mistake all three columns for standard benchmark subsets.
Circularity Check
No reduction of prediction to input by construction; the pipeline is a standard two-stage latent diffusion system with externally grounded losses and held-out evaluation.
full rationale
Prometheus's derivation chain is self-contained rather than circular. Stage 1 (GS-VAE) is supervised by rendering losses (Eq. 6) against real input images plus a scale-invariant depth loss (Eq. 7) against off-the-shelf monocular depth from DepthAnything-V2, and the reconstruction ability is evaluated on Tartanair, which the paper states is not in the training set. Stage 2 (MV-LDM) is trained by denoising score matching (Eq. 11) on the latent codes produced by that trained VAE, and text-to-3D generation is assessed on T3Bench and additional collected scene prompts. The RGB-D latent space, Plucker ray-map conditioning, and cross-view self-attention are architectural choices rather than quantities defined in terms of the claimed outputs; no equation reduces to its own input by construction. References to Director3D [38] and other prior work are used for architecture initialization, training-data selection, and metric conventions, not as evidence for the paper's central claim. The supplementary's admission of multiview inconsistency under large rotations is a scope limitation, not a circular derivation. There is therefore no fitted parameter renamed as a prediction and no load-bearing self-citation chain.
Assumptions & free parameters
free parameters (6)
- Multi-view training noise level =
Pmean = 1.5, Pstd = 2.0
- Single-view training noise level =
Pmean = -0.5, Pstd = 1.2
- Hybrid sampling guidance weights w1, w2 =
not reported
- Loss weights lambda1, lambda2, lambda3 =
not reported
- CFG-rescale factor =
not reported
- Number of views N (Stage 1) and N (Stage 2) =
N = 4 (Stage 1), N = 8 (Stage 2)
assumptions (5)
- domain assumption The frozen Stable Diffusion image encoder produces a latent space that faithfully represents both RGB and monocular depth without fine-tuning.
- domain assumption Monocular depth from DepthAnything-V2 is accurate enough to serve as pseudo ground truth for 3D geometry.
- domain assumption Single-view images with reconstruction loss on input views improve multi-view and 3D generalization.
- standard math The EDM continuous-time denoising formulation and its preconditioning functions carry over to the RGB-D latent space.
- domain assumption Ray-map conditioning plus 3D cross-view attention is sufficient for multi-view consistency in latent space.
Cite this review
Pith. "Pith review of Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation." pith.science (2026). https://pith.science/paper/CWLVMEBB
@misc{pith2026241221117,
author = {Pith},
title = {Pith review of: Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/CWLVMEBB}},
note = {Machine review of arXiv:2412.21117}
}
read the original abstract
In this work, we introduce Prometheus, a 3D-aware latent diffusion model for text-to-3D generation at both object and scene levels in seconds. We formulate 3D scene generation as multi-view, feed-forward, pixel-aligned 3D Gaussian generation within the latent diffusion paradigm. To ensure generalizability, we build our model upon pre-trained text-to-image generation model with only minimal adjustments, and further train it using a large number of images from both single-view and multi-view datasets. Furthermore, we introduce an RGB-D latent space into 3D Gaussian generation to disentangle appearance and geometry information, enabling efficient feed-forward generation of 3D Gaussians with better fidelity and geometry. Extensive experimental results demonstrate the effectiveness of our method in both feed-forward 3D Gaussian reconstruction and text-to-3D generation. Project page: https://freemty.github.io/project-prometheus/
Figures
Figures from the paper (7 more)
Forward citations
Cited by 2 Pith papers
-
CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-Centric 3D Scene Generation
CGGS generates viewpoint-consistent, text-aligned ego-centric 3D scenes via consistency-augmented multi-view diffusion, flow-guided layout initialization, and mutual-information depth-refined Gaussian optimization.
-
EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion
EarthCrafter generates 600-meter-scale 3D Earth scenes using separate latent diffusion models for structure and texture, conditioned on semantics, images, or nothing.
Reference graph
Works this paper leans on
-
[1]
Guibas, and Andrea Tagliasacchi
Sherwin Bahmani, Jeong Joon Park, Despoina Paschalidou, Xingguang Yan, Gordon Wetzstein, Leonidas J. Guibas, and Andrea Tagliasacchi. CC3D: layout-conditioned gen- eration of compositional 3d scenes. arXiv.org, 2023. 2
2023
-
[2]
Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv.2311.15127,
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, Varun Jampani, and Robin Rombach. Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv.2311.15127,
-
[3]
Video generation models as world simulators,
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luh- man, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators,
-
[4]
nuscenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2020. 5
2020
-
[5]
Chan, Connor Z
Eric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas Guibas, Jonathan Tremblay, Sameh Khamis, Tero Karras, and Gordon Wetzstein. Efficient geometry-aware 3D generative adversarial networks. InProc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2022. 2
2022
-
[6]
Chan, Koki Nagano, Matthew A
Eric R. Chan, Koki Nagano, Matthew A. Chan, Alexan- der W. Bergman, Jeong Joon Park, Axel Levy, Miika Ait- tala, Shalini De Mello, Tero Karras, and Gordon Wetzstein. GeNVS: Generative novel view synthesis with 3D-aware diffusion models. In arXiv, 2023. 3
2023
-
[7]
pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction
David Charatan, Sizhe Li, Andrea Tagliasacchi, and Vin- cent Sitzmann. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2024. 3, 6, 7
2024
-
[8]
Lara: Efficient large-baseline radi- ance fields
Anpei Chen, Haofei Xu, Stefano Esposito, Siyu Tang, and Andreas Geiger. Lara: Efficient large-baseline radi- ance fields. In European Conference on Computer Vision (ECCV), 2024. 2, 3
2024
Show all 107 references
-
[9]
Pixart- α: Fast train- ing of diffusion transformer for photorealistic text-to-image synthesis
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart- α: Fast train- ing of diffusion transformer for photorealistic text-to-image synthesis. 2310.00426, 2023. 6
-
[10]
Microdreamer: Zero-shot 3d generation in ∼20 seconds by score-based it- erative reconstruction
Luxi Chen, Zhengyi Wang, Zihan Zhou, Tingting Gao, Hang Su, Jun Zhu, and Chongxuan Li. Microdreamer: Zero-shot 3d generation in ∼20 seconds by score-based it- erative reconstruction. arXiv preprint arXiv:2404.19525 ,
-
[11]
Mvsplat: Efficient 3d gaussian splat- ting from sparse multi-view images
Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. Mvsplat: Efficient 3d gaussian splat- ting from sparse multi-view images. arXiv preprint arXiv:2403.14627, 2024. 3, 6, 7
2024 arXiv
-
[12]
Mvsplat360: Feed-forward 360 scene synthesis from sparse views
Yuedong Chen, Chuanxia Zheng, Haofei Xu, Bohan Zhuang, Andrea Vedaldi, Tat-Jen Cham, and Jianfei Cai. Mvsplat360: Feed-forward 360 scene synthesis from sparse views. In Advances in Neural Information Processing Sys- tems (NeurIPS), 2024. 3
2024
-
[13]
V3d: Video diffusion models are effective 3d generators
Zilong Chen, Yikai Wang, Feng Wang, Zhengyi Wang, and Huaping Liu. V3d: Video diffusion models are effective 3d generators. arXiv preprint arXiv:2403.06738, 2024. 2, 5
2024 arXiv
-
[14]
Luciddreamer: Domain-free generation of 3d gaussian splatting scenes
Jaeyoung Chung, Suyoung Lee, Hyeongjin Nam, Jaerin Lee, and Kyoung Mu Lee. Luciddreamer: Domain-free generation of 3d gaussian splatting scenes. arXiv preprint arXiv:2311.13384, 2023. 2, 3
2023 arXiv
-
[15]
Objaverse-xl: A universe of 10m+ 3d objects
Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Chris- tian Laforte, Vikram V oleti, Samir Yitzhak Gadre, Eli VanderBilt, Aniruddha Kembhavi, Carl V ondrick, Georgia Gkioxari, Kiana Ehsani, Ludwig Schmidt, and Ali Farhadi. Obja...
2023 arXiv
-
[16]
CARLA: An open ur- ban driving simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla, An- tonio Lopez, and Vladlen Koltun. CARLA: An open ur- ban driving simulator. In Proc. Conf. on Robot Learning (CoRL), 2017. 2
2017
-
[17]
Scenescape: Text-driven consistent scene genera- tion
Rafail Fridman, Amit Abecasis, Yoni Kasten, and Tali Dekel. Scenescape: Text-driven consistent scene genera- tion. arXiv.org, 2302.01133, 2023. 3
2023 arXiv
-
[18]
Srini- vasan, Jonathan T
Ruiqi Gao*, Aleksander Holynski*, Philipp Henzler, Arthur Brussee, Ricardo Martin-Brualla, Pratul P. Srini- vasan, Jonathan T. Barron, and Ben Poole*. Cat3d: Create anything in 3d with multi-view diffusion models.arXiv.org,
-
[19]
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2012. 5
2012
-
[20]
Camerac- trl: Enabling camera control for text-to-video generation
Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein, Bo Dai, Hongsheng Li, and Ceyuan Yang. Camerac- trl: Enabling camera control for text-to-video generation. arXiv.org, 2404.02101, 2024. 2
2024 arXiv
-
[21]
Gvgen: Text-to-3d generation with volumetric representation
Xianglong He, Junyi Chen, Sida Peng, Di Huang, Yang- guang Li, Xiaoshui Huang, Chun Yuan, Wanli Ouyang, and Tong He. Gvgen: Text-to-3d generation with volumetric representation. In Proc. of the European Conf. on Com- puter Vision (ECCV), 2024. 3
2024
-
[22]
T 3 bench: Benchmarking current progress in text-to-3d gener- ation
Yuze He, Yushi Bai, Matthieu Lin, Wang Zhao, Yubin Hu, Jenny Sheng, Ran Yi, Juanzi Li, and Yong-Jin Liu. T 3 bench: Benchmarking current progress in text-to-3d gener- ation. arXiv preprint arXiv:2310.02977, 2023. 7
2023 arXiv
-
[23]
Sampling 3d gaussian scenes in seconds with latent diffusion models
Paul Henderson, Melonie de Almeida, Daniela Ivanova, and Titas Anciukevicius. Sampling 3d gaussian scenes in seconds with latent diffusion models. arXiv preprint arXiv:2406.13099, 2024. 3
2024 arXiv
-
[24]
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation metric for image captioning. arXiv preprint arXiv:2104.08718, 2021. 7
2021 arXiv
-
[25]
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022. 5 9
2022 arXiv
-
[26]
LRM: large reconstruction model for single image to 3d
Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. LRM: large reconstruction model for single image to 3d. arXiv.2311.04400, 2023. 3
2023 arXiv
-
[27]
Depthcrafter: Generating consistent long depth sequences for open-world videos
Wenbo Hu, Xiangjun Gao, Xiaoyu Li, Sijie Zhao, Xi- aodong Cun, Yong Zhang, Long Quan, and Ying Shan. Depthcrafter: Generating consistent long depth sequences for open-world videos. arXiv preprint arXiv:2409.02095 ,
-
[28]
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In Proc. of the European Conf. on Computer Vision (ECCV) ,
-
[29]
A style- based generator architecture for generative adversarial net- works
Tero Karras, Samuli Laine, and Timo Aila. A style- based generator architecture for generative adversarial net- works. In Proc. IEEE Conf. on Computer Vision and Pat- tern Recognition (CVPR), 2019. 2
2019
-
[30]
Elucidating the design space of diffusion-based generative models
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. Advances in neural information processing sys- tems, 35:26565–26577, 2022. 5
2022
-
[31]
Re- purposing diffusion-based image generators for monocular depth estimation
Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Met- zger, Rodrigo Caye Daudt, and Konrad Schindler. Re- purposing diffusion-based image generators for monocular depth estimation. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR),
-
[32]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. on Graphics, 42(4),
-
[33]
Berg, Wan-Yen Lo, Piotr Doll´ar, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Doll´ar, and Ross Girshick. Segment anything. arXiv:2304.02643,
-
[34]
Vivid-1-to-3: Novel view synthesis with video diffusion models
Jeong-gi Kwak, Erqun Dong, Yuhe Jin, Hanseok Ko, Shweta Mahajan, and Kwang Moo Yi. Vivid-1-to-3: Novel view synthesis with video diffusion models. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[35]
Ln3diff: Scalable latent neural fields diffusion for speedy 3d generation
Yushi Lan, Fangzhou Hong, Shuai Yang, Shangchen Zhou, Xuyi Meng, Bo Dai, Xingang Pan, and Chen Change Loy. Ln3diff: Scalable latent neural fields diffusion for speedy 3d generation. In Proc. of the European Conf. on Computer Vision (ECCV), 2024. 3
2024
-
[36]
Rgbd2: Generative scene synthesis via incremental view inpainting using rgbd diffusion models
Jiabao Lei, Jiapeng Tang, and Kui Jia. Rgbd2: Generative scene synthesis via incremental view inpainting using rgbd diffusion models. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2023. 3
2023
-
[37]
Dual3d: Efficient and consistent text-to-3d generation with dual-mode multi-view latent diffusion
Xinyang Li, Zhangyu Lai andLinning Xu, Jianfei Guo, and Liujuan Cao andShengchuan Zhang. Dual3d: Efficient and consistent text-to-3d generation with dual-mode multi-view latent diffusion. arXiv.org, 2405.09874, 2024. 3
2024 arXiv
-
[38]
Di- rector3d: Real-world camera trajectory and 3d scene gen- eration from text
Xinyang Li, Zhangyu Lai, Linning Xu, Yansong Qu, Liu- juan Cao, Shengchuan Zhang, Bo Dai, and Rongrong Ji. Di- rector3d: Real-world camera trajectory and 3d scene gen- eration from text. arXiv.org, 2406.17601, 2024. 3, 4, 5, 7
2024 arXiv
-
[39]
Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d
Yiyi Liao, Jun Xie, and Andreas Geiger. Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d. In IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), 2022. 5
2022
-
[40]
Magic3d: High- resolution text-to-3d content creation
Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High- resolution text-to-3d content creation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti...
2023
-
[41]
Common diffusion noise schedules and sample steps are flawed
Shanchuan Lin, Bingchen Liu, Jiashi Li, and Xiao Yang. Common diffusion noise schedules and sample steps are flawed. In Proc. of the IEEE Winter Conference on Ap- plications of Computer Vision (WACV), 2024. 6
2024
-
[42]
Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision
Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, et al. Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , pages 22160–2216...
2024
-
[43]
Infinite nature: Perpetual view generation of natural scenes from a single image
Andrew Liu, Richard Tucker, Varun Jampani, Ameesh Makadia, Noah Snavely, and Angjoo Kanazawa. Infinite nature: Perpetual view generation of natural scenes from a single image. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision (ICCV), 2021. 3, 5
2021
-
[44]
Re- conx: Reconstruct any scene from sparse views with video diffusion model
Fangfu Liu, Wenqiang Sun, Hanyang Wang, Yikai Wang, Haowen Sun, Junliang Ye, Jun Zhang, and Yueqi Duan. Re- conx: Reconstruct any scene from sparse views with video diffusion model. arXiv preprint arXiv:2408.16767, 2024. 2
2024 arXiv
-
[45]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural infor- mation processing systems, 36, 2024. 6
2024
-
[46]
Zero-1-to-3: Zero-shot one image to 3d object
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tok- makov, Sergey Zakharov, and Carl V ondrick. Zero-1-to-3: Zero-shot one image to 3d object. arXiv.org, 2303.11328,
-
[47]
Syncdreamer: Generating multiview-consistent images from a single-view image
Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Generating multiview-consistent images from a single-view image. arXiv preprint arXiv:2309.03453, 2023. 2
2023 arXiv
-
[48]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. In ECCV, 2020. 3
2020
-
[49]
No-reference image quality assessment in the spatial domain
Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik. No-reference image quality assessment in the spatial domain. IEEE Transactions on image processing, 21(12): 4695–4708, 2012. 7
2012
-
[50]
completely blind
Anish Mittal, Rajiv Soundararajan, and Alan C Bovik. Making a “completely blind” image quality analyzer. IEEE Signal processing letters, 20(3):209–212, 2012. 7
2012
-
[51]
Diffrf: Rendering-guided 3d radiance field diffusion
Norman M ¨uller, Yawar Siddiqui, Lorenzo Porzi, Samuel Rota Bulo, Peter Kontschieder, and Matthias Nießner. Diffrf: Rendering-guided 3d radiance field diffusion. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2023. 2 10
2023
-
[52]
Multidiff: Consistent novel view synthesis from a single image
Norman M ¨uller, Katja Schwarz, Barbara R ¨ossle, Lorenzo Porzi, Samuel Rota Bul `o, Matthias Nießner, and Peter Kontschieder. Multidiff: Consistent novel view synthesis from a single image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (...
2024
-
[53]
GIRAFFE: rep- resenting scenes as compositional generative neural feature fields
Michael Niemeyer and Andreas Geiger. GIRAFFE: rep- resenting scenes as compositional generative neural feature fields. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2021. 2
2021
-
[54]
Barron, and Ben Milden- hall
Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv,
-
[55]
Towards robust monocu- lar depth estimation: Mixing datasets for zero-shot cross- dataset transfer
Ren ´e Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocu- lar depth estimation: Mixing datasets for zero-shot cross- dataset transfer. IEEE Trans. on Pattern Analysis and Ma- chine Intelligence (PAMI), 44(3), 2022. 4, 7
2022
-
[56]
L3dg: Latent 3d gaussian diffusion
Barbara Roessle, Norman M ¨uller, Lorenzo Porzi, Samuel Rota Bul `o, Peter Kontschieder, Angela Dai, and Matthias Nießner. L3dg: Latent 3d gaussian diffusion. In SIGGRAPH Asia 2024 Conference Papers, 2024. 3
2024
-
[57]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2, 3, 5, 6
2022
-
[58]
Ze- roNVS: Zero-shot 360-degree view synthesis from a single real image
Kyle Sargent, Zizhang Li, Tanmay Shah, Charles Her- rmann, Hong-Xing Yu, Yunzhi Zhang, Eric Ryan Chan, Dmitry Lagun, Li Fei-Fei, Deqing Sun, and Jiajun Wu. Ze- roNVS: Zero-shot 360-degree view synthesis from a single real image. arXiv.org, 2310.17994, 2023. 2
-
[59]
Graf: Generative radiance fields for 3d-aware im- age synthesis
Katja Schwarz, Yiyi Liao, Michael Niemeyer, and Andreas Geiger. Graf: Generative radiance fields for 3d-aware im- age synthesis. In Advances in Neural Information Process- ing Systems (NeurIPS), 2020. 2
2020
-
[60]
Wildfusion: Learn- ing 3d-aware latent diffusion models in view space
Katja Schwarz, Seung Wook Kim, Jun Gao, Sanja Fidler, Andreas Geiger, and Karsten Kreis. Wildfusion: Learn- ing 3d-aware latent diffusion models in view space. In Proc. of the International Conf. on Learning Representa- tions (ICLR), 2024. 3
2024
-
[61]
Learning tem- porally consistent video depth from video diffusion priors
Jiahao Shao, Yuanbo Yang, Hongyu Zhou, Youmin Zhang, Yujun Shen, Matteo Poggi, and Yiyi Liao. Learning tem- porally consistent video depth from video diffusion priors. arXiv preprint arXiv:2406.01493, 2024. 7
2024 arXiv
-
[62]
Zero123++: a single image to consis- tent multi-view diffusion base model
Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang, Minghua Liu, Chao Xu, Xinyue Wei, Linghao Chen, Chong Zeng, and Hao Su. Zero123++: a single image to consis- tent multi-view diffusion base model. arXiv preprint arXiv:2310.15110, 2023. 5
-
[63]
Mvdream: Multi-view diffusion for 3d generation
Yichun Shi, Peng Wang, Jianglong Ye, Long Mai, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d generation. arXiv:2308.16512, 2023. 2, 5, 7, 8
2023 arXiv
-
[64]
Realmdreamer: Text-driven 3d scene genera- tion with inpainting and depth diffusion
Jaidev Shriram, Alex Trevithick, Lingjie Liu, and Ravi Ra- mamoorthi. Realmdreamer: Text-driven 3d scene genera- tion with inpainting and depth diffusion. arXiv, 2024. 2, 3
2024
-
[65]
Light field networks: Neu- ral scene representations with single-evaluation rendering
Vincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum, and Fredo Durand. Light field networks: Neu- ral scene representations with single-evaluation rendering. Advances in Neural Information Processing Systems , 34: 19313–19325, 2021. 3
2021
-
[66]
Score- based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score- based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020. 5
2011 arXiv
-
[67]
Scalability in perception for autonomous driv- ing: Waymo open dataset
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aure- lien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, Vijay Vasudevan, Wei Han, Jiquan Ngiam, Hang Zhao, Aleksei Timofeev, Scott Ettinger, Maxim Krivokon, Amy Gao, Aditya Joshi, Yu Zhan...
2020
-
[68]
Dimensionx: Create any 3d and 4d scenes from a single image with con- trollable video diffusion
Wenqiang Sun, Shuo Chen, Fangfu Liu, Zilong Chen, Yueqi Duan, Jun Zhang, and Yikai Wang. Dimensionx: Create any 3d and 4d scenes from a single image with con- trollable video diffusion. arXiv preprint arXiv:2411.04928,
-
[69]
Flash3d: Feed-forward gener- alisable 3d scene reconstruction from a single image.arxiv,
Stanislaw Szymanowicz, Eldar Insafutdinov, Chuanxia Zheng, Dylan Campbell, Joao Henriques, Christian Rup- precht, and Andrea Vedaldi. Flash3d: Feed-forward gener- alisable 3d scene reconstruction from a single image.arxiv,
-
[70]
Dreamgaussian: Generative gaussian splat- ting for efficient 3d content creation
Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. Dreamgaussian: Generative gaussian splat- ting for efficient 3d content creation. arXiv preprint arXiv:2309.16653, 2023. 2, 3
2023 arXiv
-
[71]
Lgm: Large multi- view gaussian model for high-resolution 3d content cre- ation
Jiaxiang Tang, Zhaoxi Chen, Xiaokang Chen, Tengfei Wang, Gang Zeng, and Ziwei Liu. Lgm: Large multi- view gaussian model for high-resolution 3d content cre- ation. arXiv.org, 2402.05054, 2024. 3, 7
2024 arXiv
-
[72]
A connection between score matching and denoising autoencoders
Pascal Vincent. A connection between score matching and denoising autoencoders. Neural computation, 23(7):1661– 1674, 2011. 5
2011
-
[73]
Geco: Generative image-to-3d within a sec- ond
Chen Wang, Jiatao Gu, Xiaoxiao Long, Yuan Liu, and Lingjie Liu. Geco: Generative image-to-3d within a sec- ond. arXiv, 2024. 2
2024
-
[74]
Yeh, and Greg Shakhnarovich
Haochen Wang, Xiaodan Du, Jiahao Li, Raymond A. Yeh, and Greg Shakhnarovich. Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation. arXiv preprint arXiv:2212.00774, 2022. 3
2022 arXiv
-
[75]
Vistadream: Sampling multi- view consistent images for single-view scene reconstruc- tion
Haiping Wang, Yuan Liu, Ziwei Liu, Zhen Dong, Wenping Wang, and Bisheng Yang. Vistadream: Sampling multi- view consistent images for single-view scene reconstruc- tion. arXiv preprint arXiv:2410.16892, 2024. 2, 3
2024 arXiv
-
[76]
Barron, Ricardo Martin-Brualla, Noah Snavely, and Thomas Funkhouser
Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul Srinivasan, Howard Zhou, Jonathan T. Barron, Ricardo Martin-Brualla, Noah Snavely, and Thomas Funkhouser. Ibrnet: Learning multi-view image-based rendering. In CVPR, 2021. 3
2021
-
[77]
Dust3r: Geometric 3d vision made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vision made easy. In CVPR, 2024. 3 11
2024
-
[78]
Tartanair: A dataset to push the limits of visual slam
Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor, and Sebastian Scherer. Tartanair: A dataset to push the limits of visual slam. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 4909–491...
2020
-
[79]
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image pro- cessing, 13(4):600–612, 2004. 7
2004
-
[80]
Prolificdreamer: High- fidelity and diverse text-to-3d generation with variational score distillation
Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongx- uan Li, Hang Su, and Jun Zhu. Prolificdreamer: High- fidelity and diverse text-to-3d generation with variational score distillation. In Advances in Neural Information Pro- cessing Systems (NeurIPS), 2023. 2, 3
2023
-
[81]
Motionctrl: A unified and flexible motion controller for video generation
Zhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li, Tianshui Chen, Menghan Xia, Ping Luo, and Ying Shan. Motionctrl: A unified and flexible motion controller for video generation. In ACM SIGGRAPH 2024 Conference Papers, pages 1–11, 2024. 2
2024
-
[82]
latentsplat: Autoencoding variational gaussians for fast generalizable 3d reconstruction
Christopher Wewer, Kevin Raj, Eddy Ilg, Bernt Schiele, and Jan Eric Lenssen. latentsplat: Autoencoding variational gaussians for fast generalizable 3d reconstruction. InarXiv,
-
[83]
Harmonyview: Harmoniz- ing consistency and diversity in one-image-to-3d
Sangmin Woo, Byeongjun Park, Hyojun Go, Jin-Young Kim, and Changick Kim. Harmonyview: Harmoniz- ing consistency and diversity in one-image-to-3d. arXiv preprint arXiv:2312.15980, 2023. 6
2023 arXiv
-
[84]
Srinivasan, Dor Verbin, Jonathan T
Rundi Wu, Ben Mildenhall, Philipp Henzler, Keunhong Park, Ruiqi Gao, Daniel Watson, Pratul P. Srinivasan, Dor Verbin, Jonathan T. Barron, Ben Poole, and Aleksander Holynski. Reconfusion: 3d reconstruction with diffusion priors. arXiv.org, 2023. 2
2023
-
[85]
Direct3d: Scalable image-to-3d generation via 3d latent diffusion transformer
Shuang Wu, Youtian Lin, Feihu Zhang, Yifei Zeng, Jingxi Xu, Philip Torr, Xun Cao, and Yao Yao. Direct3d: Scalable image-to-3d generation via 3d latent diffusion transformer. arXiv:2405.14832, 2024. 6
2024 arXiv
-
[86]
Lrm-zero: Training large reconstruction models with syn- thesized data
Desai Xie, Sai Bi, Zhixin Shu, Kai Zhang, Zexiang Xu, Yi Zhou, S ¨oren Pirk, Arie Kaufman, Xin Sun, and Hao Tan. Lrm-zero: Training large reconstruction models with syn- thesized data. arXiv.org, 2024. 3
2024
-
[87]
Depthsplat: Connecting gaussian splatting and depth
Haofei Xu, Songyou Peng, Fangjinhua Wang, Hermann Blum, Daniel Barath, Andreas Geiger, and Marc Pollefeys. Depthsplat: Connecting gaussian splatting and depth. arXiv preprint arXiv:2410.13862, 2024. 3
2024 arXiv
-
[88]
Discoscene: Spatially disentangled generative radiance fields for controllable 3d-aware scene synthesis
Yinghao Xu, Menglei Chai, Zifan Shi, Sida Peng, Ivan Skorokhodov, Aliaksandr Siarohin, Ceyuan Yang, Yujun Shen, Hsin-Ying Lee, Bolei Zhou, and Sergey Tulyakov. Discoscene: Spatially disentangled generative radiance fields for controllable 3d-aware scene synthesis. Proc. IEEE C...
2023
-
[89]
Dmv3d: Denoising multi-view diffusion using 3d large reconstruction model
Yinghao Xu, Hao Tan, Fujun Luan, Sai Bi, Peng Wang, Jiahao Li, Zifan Shi, Kalyan Sunkavalli, Gordon Wet- zstein, Zexiang Xu, and Kai Zhang. Dmv3d: Denoising multi-view diffusion using 3d large reconstruction model. arXiv:2311.09217, 2023. 2, 3
2023 arXiv
-
[90]
Depth any- thing v2
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiao- gang Xu, Jiashi Feng, and Hengshuang Zhao. Depth any- thing v2. arXiv:2406.09414, 2024. 3, 6
2024 arXiv
-
[91]
Urbangiraffe: Representing urban scenes as compositional generative neural feature fields
Yuanbo Yang, Yifei Yang, Hanlei Guo, Rong Xiong, Yue Wang, and Yiyi Liao. Urbangiraffe: Representing urban scenes as compositional generative neural feature fields. Proc. of the IEEE International Conf. on Computer Vision (ICCV), 2023. 2
2023
-
[92]
Mathematical supple- ment for the gsplat library, 2023
Vickie Ye and Angjoo Kanazawa. Mathematical supple- ment for the gsplat library, 2023. 6
2023
-
[93]
Gaussiandreamer: Fast generation from text to 3d gaussians by bridging 2d and 3d diffusion models
Taoran Yi, Jiemin Fang, Junjie Wang, Guanjun Wu, Lingxi Xie, Xiaopeng Zhang, Wenyu Liu, Qi Tian, and Xing- gang Wang. Gaussiandreamer: Fast generation from text to 3d gaussians by bridging 2d and 3d diffusion models. In CVPR, 2024. 2, 3, 7
2024
-
[94]
Grm: Large gaussian reconstruction model for efficient 3d reconstruction and generation
Xu Yinghao, Shi Zifan, Yifan Wang, Chen Hansheng, Yang Ceyuan, Peng Sida, Shen Yujun, and Wetzstein Gordon. Grm: Large gaussian reconstruction model for efficient 3d reconstruction and generation. arXiv.org, 2403.14621,
-
[95]
Nvs- solver: Video diffusion model as zero-shot novel view syn- thesizer
Meng You, Zhiyu Zhu, Hui Liu, and Junhui Hou. Nvs- solver: Video diffusion model as zero-shot novel view syn- thesizer. arXiv preprint arXiv:2405.15364, 2024. 2
2024 arXiv
-
[96]
Freeman, and Jiajun Wu
Hong-Xing Yu, Haoyi Duan, Charles Herrmann, William T. Freeman, and Jiajun Wu. Wonderworld: Interactive 3d scene generation from a single image. arXiv.org, 2406.09394, 2024. 3
2024 arXiv
-
[97]
Freeman, Forrester Cole, Deqing Sun, Noah Snavely, Jiajun Wu, and Charles Her- rmann
Hong-Xing Yu, Haoyi Duan, Junhwa Hur, Kyle Sargent, Michael Rubinstein, William T. Freeman, Forrester Cole, Deqing Sun, Noah Snavely, Jiajun Wu, and Charles Her- rmann. Wonderjourney: Going from anywhere to every- where. In Proc. IEEE Conf. on Computer Vision and Pat- tern Rec...
2024
-
[98]
Viewcrafter: Taming video diffusion models for high-fidelity novel view synthesis
Wangbo Yu, Jinbo Xing, Li Yuan, Wenbo Hu, Xiaoyu Li, Zhipeng Huang, Xiangjun Gao, Tien-Tsin Wong, Ying Shan, and Yonghong Tian. Viewcrafter: Taming video diffusion models for high-fidelity novel view synthesis. arXiv.org, 2409.02048, 2024. 2
2024 arXiv
-
[99]
Mvimgnet: A large-scale dataset of multi-view images
Xianggang Yu, Mutian Xu, Yidan Zhang, Haolin Liu, Chongjie Ye, Yushuang Wu, Zizheng Yan, Tianyou Liang, Guanying Chen, Shuguang Cui, and Xiaoguang Han. Mvimgnet: A large-scale dataset of multi-view images. In Proc. IEEE Conf. on Computer Vision and Pattern Recog- nition (CVPR)...
2023
-
[100]
Gaussiancube: Structuring gaussian splatting using op- timal transport for 3d generative modeling
Bowen Zhang, Yiji Cheng, Jiaolong Yang, Chunyu Wang, Feng Zhao, Yansong Tang, Dong Chen, and Baining Guo. Gaussiancube: Structuring gaussian splatting using op- timal transport for 3d generative modeling. arXiv.org, 2403.19655, 2024. 2
2024 arXiv
-
[101]
Cameras as rays: Pose estimation via ray diffusion
Jason Y Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani. Cameras as rays: Pose estimation via ray diffusion. In International Conference on Learning Representations (ICLR), 2024. 3, 6
2024
-
[102]
Gs-lrm: Large reconstruction model for 3d gaussian splatting
Kai Zhang, Sai Bi, Hao Tan, Yuanbo Xiangli, Nanxuan Zhao, Kalyan Sunkavalli, and Zexiang Xu. Gs-lrm: Large reconstruction model for 3d gaussian splatting. arXiv.org,
-
[103]
3DitScene: Editing any scene via language-guided disen- tangled gaussian splatting
Qihang Zhang, Yinghao Xu, Chaoyang Wang, Hsin-Ying Lee, Gordon Wetzstein, Bolei Zhou, and Ceyuan Yang. 3DitScene: Editing any scene via language-guided disen- tangled gaussian splatting. arXiv.org, 2024. 3
2024
-
[104]
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 586–595, 2018. 7
2018
-
[105]
Genxd: Generating any 3d and 4d scenes
Yuyang Zhao, Chung-Ching Lin, Kevin Lin, Zhiwen Yan, Linjie Li, Zhengyuan Yang, Jianfeng Wang, Gim Hee Lee, and Lijuan Wang. Genxd: Generating any 3d and 4d scenes. arXiv preprint arXiv:2411.02319, 2024. 2
2024 arXiv
-
[106]
Diffgs: Functional gaussian splatting diffusion
Junsheng Zhou, Weiqi Zhang, and Yu-Shen Liu. Diffgs: Functional gaussian splatting diffusion. In Advances in Neural Information Processing Systems (NeurIPS) , 2024. 3
2024
-
[107]
Stereo magnification: Learning view synthesis using multiplane images
Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: Learning view synthesis using multiplane images. ACM Trans. on Graph- ics, 37, 2018. 5 13 Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation S...
2018
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.