REVIEW 4 major objections 5 minor 125 references
Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion
T0 review · 4 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read MVGD is a diffusion model that directly generates novel-view images and depth maps from arbitrary posed input views, without an intermediate 3D representation, and reports state-of-the-art results across novel view synthesis, stereo, and…
desk verdict MVGD is a serious empirical paper and a genuine architectural step forward, but the scene scale normalization that underpins the no-test-time-alignment claim has an untested out-of-range failure mode and the training rescale formula as written looks like a typo. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central objects are (1) raymap conditioning, which encodes the per-pixel ray origin and direction of every conditioning camera and of the target camera using Fourier features, letting the diffusion model see the geometry of all views; (2) Recurrent Interface Networks (RIN), a transformer that keeps computation on a fixed number of latent tokens and makes pixel-level diffusion in RGB-D affordable; (3) learnable task embeddings that let one model generate either an image or a depth map from the same conditioned latent tokens; and (4) Scene Scale Normalization (SSN), which rescales all camera translations and depth by the largest absolute translation coordinate relative to the target camera, so depth predictions emerge in the same scale as the conditioning cameras. Together these carry the argument that multi-view consistency and scale-awareness can arise from direct generation, with no explicit 3D structure.
What would settle it
Run MVGD on a scene where all conditioning cameras are nearly coincident (baseline approaching zero, so the scene scale s is tiny) and compare the predicted depth maps against ground truth: if the depth error is large or the depth scale is inconsistent across generated viewpoints, the SSN assumption fails. A complementary test uses conditioning cameras whose largest translation coordinate is in a direction that does not actually separate the views (for example, all cameras moved forward in a line with no lateral offset), which would also push the normalized geometry outside the range seen during training.
Extended reading notes
Core claim
MVGD is a diffusion architecture that learns a conditional distribution over images and depth maps at a target viewpoint, conditioned on an arbitrary number of posed input views. Instead of building an explicit 3D scene representation, the model consumes image tokens augmented with Fourier-encoded raymaps—per-pixel ray origins and directions—that position features from conditioning views in 3D and specify the target viewpoint. The model is trained with learnable task embeddings to switch between image and depth generation, so the same conditioned tokens can produce either modality. Scene Scale Normalization (SSN) defines the scene scale s as the largest absolute camera translation among conditioning views relative to the target, divides all translations and the ground-truth depth by s during training, and multiplies predicted depth by s at inference, which the authors show yields depth maps consistent with the scale of the conditioning cameras with no test-time alignment. The paper demonstrates that this direct generative approach outperforms prior methods that rely on 3D representations on multiple benchmarks, that it generalizes to 2–9 conditioning views (and up to thousands with incremental conditioning), and that it produces multi-view consistent point clouds without any post-processing.
Load-bearing premise
The load-bearing premise is that a single scalar—the largest absolute camera translation among conditioning views relative to the target—is enough to normalize scene scale across heterogeneous datasets; if that scalar does not reflect the true baseline (near-zero translations, or a dominant translation that does not separate the views), predicted depths will be rescaled incorrectly and the claimed multi-view consistency breaks.
Editorial extensions
If this is right
- Novel view and depth synthesis become a single sampling procedure from one network, so a 3D reconstruction step (NeRF, voxel grid, 3D Gaussians) is no longer needed for these tasks.
- The same trained model accepts 2 to 9 conditioning views at test time, and up to thousands with the incremental conditioning strategy, so additional views translate directly into better predictions without retraining or test-time optimization.
- Jointly training image and depth generation acts as a geometric regularizer: the ablation shows that removing scale normalization hurts novel view synthesis more than removing depth supervision does, indicating that inaccurate geometry is more detrimental than no geometry at all.
- Incremental fine-tuning by duplicating latent tokens and fine-tuning for 50k steps yields up to +20% improvement with almost no added parameters, offering a cheap scaling path for larger diffusion models.
- The same model, without changes, outperforms dedicated multi-view stereo and video depth estimators on ScanNet and related benchmarks, implying that the learned generative multi-view prior transfers to geometry-only estimation tasks.
Reading between the lines
- The Scene Scale Normalization step assumes that one scalar, the largest translation coordinate, captures the relevant baseline; on configurations where conditioning cameras are nearly coincident or where the dominant translation is perpendicular to the viewing direction, predicted depths may be scaled outside the range the model learned, so a learned or distribution-aware scale estimator would mak
- Because the model ingests 100+ conditioning views with near-constant computation, it could be used as a shared geometry prior for 3D reconstruction pipelines, for example to initialize or regularize radiance fields or Gaussian splats, without any retraining.
- The task-embedding design is modular: the paper already shows a depth-only model can be fine-tuned to also synthesize images, suggesting other output modalities such as surface normals or semantic maps could be added with minimal extra training, though MVGD at test time remains limited to the fixed log-depth range d_min=0.1 to d_max=200.
- A direct test: on metric-scale driving datasets with LiDAR ground truth, one could verify whether generated depth is metrically consistent across a long sequence and whether error grows in configurations with small baselines, which would directly probe the SSN claim.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. MVGD is a pixel-space, multi-task diffusion model for zero-shot novel-view image and depth synthesis. It conditions on an arbitrary number of posed input images using EfficientViT image embeddings and Fourier ray-map embeddings, and it generates a target view directly with a Recurrent Interface Network, using learnable task embeddings to switch between image and depth generation. To make depth scale-consistent across heterogeneous datasets, the paper proposes Scene Scale Normalization (SSN, Sec. 3.2): conditioning camera translations are expressed relative to the target camera, divided by a scalar scene scale s, and ground-truth depth is likewise divided by s during training; at inference, predicted depth is multiplied by s, with no test-time alignment. The model is trained on roughly 61M multi-view samples from 18 public datasets and evaluated on novel-view benchmarks (RealEstate10K, LLFF, DTU, CO3D, Mip-NeRF360) and on stereo/video depth benchmarks (ScanNet, SUN3D, RGB-D). The paper reports state-of-the-art numbers in Tables 2-7, presents ablations in Table 5, and introduces an incremental fine-tuning strategy that expands the number of latent tokens while reusing a smaller trained model.
Significance. The paper's main significance is architectural and conceptual: it provides evidence that an explicit intermediate 3D representation (NeRF, voxel grid, 3D Gaussian) is not necessary for multi-view-consistent novel-view generation and depth synthesis, because a diffusion model trained jointly on images and depth can enforce geometric consistency directly at the pixel level. If the results hold, the proposed Scene Scale Normalization is also a practical way to train on datasets with different scales and to return metric-consistent depth at inference. The strengths of the paper include a clearly specified architecture, a very large and diverse training mixture, ablations that isolate the contributions of multi-task learning, SSN, and data diversity, and a useful scaling trick for expanding latent capacity without retraining from scratch. The paper also commits to releasing code, dataloaders, and the list of valid training samples, which is important for reproducibility.
major comments (4)
- [Section 3.2] The recalculation formula s' = s · Dmax/max{tilde D_t} cannot be correct as written. If Dmax denotes the model's dmax=200, then for a sample with max(tilde D_t)>dmax the expression yields a smaller s' than s (e.g., with s=1, max(tilde D_t)=1000, dmax=200, s'=0.2), which pushes normalized depth even farther outside [0.1,200]. If Dmax instead denotes max(tilde D_t), then the formula is the identity and does no recalculation at all. The intended expression should be s' = s · max(tilde D_t)/dmax (equivalently s' = max(D_t)/dmax). Because the 'no test-time alignment' claim in Table 4 and the point-cloud consistency in Figures 5-6 are justified by this mechanism, the manuscript must correct the formula and state which quantity was actually implemented; as written, the described training procedure would not keep normalized depth in range.
- [Section 3.2 and Appendix A] There is no inference-time safeguard for scenes whose depth-to-baseline ratio exceeds the training range. During training, the s' recalculation and the cmin/cmax view-selection thresholds in Appendix A partially control this ratio, but at test time the model multiplies predicted normalized depth by s regardless of whether max(D_hat_t/s) is inside [0.1,200]. For a scene with a small conditioning baseline relative to scene depth (near-duplicate conditioning cameras, or a very deep target view), the normalized depth saturates in the log-scale parameterization of Eq. (2), and Eq. (3) then produces incorrect metric depth. The paper does not report the depth-to-baseline coverage of its 61M-sample training distribution, nor does any experiment in Tables 4, 6, or 7 stratify accuracy by that ratio. I request a targeted stress test: construct degenerate and extreme baseline configurations, report AbsRel and multi-view point-cloud consistency as a function of max(D_t)/s, and state the failure boundary; also report the distribution of this ratio in the training data. Without this, the claim that generated depth shares the conditioning cameras' scale without test-time alignment is verified only inside an unquantified training regime.
- [Tables 2-8] No variance or multiple-seed results are reported. Some of the margins over strong baselines are small (e.g., Table 3, CO3D 3-view PSNR: MVGD 20.68 vs. CAT3D 20.57), so the qualitative claim of consistent state-of-the-art performance would be substantially stronger with at least three seeds for a representative subset of benchmarks and a statement of whether the differences are significant. This is particularly relevant because the diffusion sampler is stochastic and the paper uses only five-sample ensembling at inference.
- [Table 4 and Section 4.3] The novel-depth evaluation in Table 4 uses sparse COLMAP reconstructions as ground truth, but the paper does not report the valid-pixel mask, the density of the reconstructed depth, or the number of pixels used per scene. Since COLMAP depth is noisy and incomplete, the reported AbsRel, RMSE, and delta_1 values may partly reflect reconstruction artifacts. I ask for a description of the evaluation mask and, where possible, a sanity check on dense sensor depth (e.g., ScanNet) to confirm that the depth metrics are not dominated by sparse-ground-truth artifacts.
minor comments (5)
- [Section 3.2] The symbols Dmax and dmax are used without a clear distinction; the text should define both explicitly and use one notation consistently.
- [Table 3] The claim of consistent state-of-the-art results should be qualified: on Mip-NeRF360 9-view LPIPS, MVGD reports 0.488 while CAT3D reports 0.439, so the method does not improve every metric on every benchmark.
- [Section 4.6 and Table 5] The text states that a 2048-latent model was trained from scratch for 300k steps, but the Table 5 caption says the top results were obtained with 200k steps; this inconsistency should be resolved.
- [Appendix E] The Limitations section acknowledges stochasticity and the lack of explicit dynamics modeling, but it does not mention the depth-range saturation failure mode of SSN; adding this limitation would make the appendix more complete.
- [References] Reference [2] is listed as anonymous and under review; if this is a double-blind reference, it should be replaced with its published version or removed.
Circularity Check
No significant circularity: MVGD's central results are empirical and benchmarked against external baselines; SSN is a normalization, not a fitted prediction.
full rationale
The paper's derivation chain is empirical rather than definitional. The key normalization in Sec. 3.2 defines s from the input camera extrinsics and multiplies the predicted depth by s, so the output depth inherits the conditioning cameras' scale by construction; however, the normalized depth map itself is generated by the diffusion model and is evaluated against external ground-truth (COLMAP, ScanNet, SUN3D, RGB-D), so the scale factor is not a fitted parameter renamed as a prediction. The self-citations (DeFiNe, GRIN, SE(3) ray embeddings) support design choices such as Fourier raymaps, log-depth parameterization, and scale invariance, but the state-of-the-art claims are supported by comparisons to independently developed baselines (PixelSplat, MVSplat, CAT3D, ReconFusion), and no uniqueness theorem or load-bearing result is imported solely from the authors' prior work. The paper's own limitation section admits only that SSN targets one viewpoint at a time and that dynamics are not explicitly modeled; these are scope limits, not circular reductions. The recalibration formula s' = s * Dmax / max{D_t/s} in Sec. 3.2 appears dimensionally suspect and the inference-time depth-range coverage is not stress-tested, but these are correctness/robustness concerns that do not make the derivation equivalent to its inputs. No step meets the evidentiary bar for circularity.
Assumptions & free parameters
free parameters (4)
- Depth bounds dmin, dmax =
0.1 and 200
- Ray encoding frequencies =
No = Nr = 8, max frequency 100
- Conditioning view selection thresholds =
cmin=0.05 cM, cmax=0.2 cM; tmin=-8, tmax=8; alpha in [0, pi/2] for depth and [0, pi] for image; pmin=30%; min 64…
- Token counts =
Ms=1024 scene tokens, Mt=4096 prediction tokens, Z in {256,512,1024,2048}
assumptions (5)
- standard math Diffusion models can approximate the conditional distribution of images and depths given conditioning tokens (Eq. 1, Section 3.1.1).
- domain assumption Input camera intrinsics and extrinsics are known and accurate for both training and evaluation.
- domain assumption COLMAP reconstructions provide camera poses and sparse depth that are scale-consistent within each sequence.
- domain assumption The largest absolute camera translation among conditioning views relative to the target is a sufficient scene-scale statistic.
- domain assumption Jointly training image and depth generation on shared latent tokens encourages multi-view geometric consistency.
Cite this review
Pith. "Pith review of Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion." pith.science (2026). https://pith.science/paper/AO2YYJWN
@misc{pith2026250118804,
author = {Pith},
title = {Pith review of: Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion},
year = {2026},
howpublished = {\url{https://pith.science/paper/AO2YYJWN}},
note = {Machine review of arXiv:2501.18804}
}
read the original abstract
Current methods for 3D scene reconstruction from sparse posed images employ intermediate 3D representations such as neural fields, voxel grids, or 3D Gaussians, to achieve multi-view consistent scene appearance and geometry. In this paper we introduce MVGD, a diffusion-based architecture capable of direct pixel-level generation of images and depth maps from novel viewpoints, given an arbitrary number of input views. Our method uses raymap conditioning to both augment visual features with spatial information from different viewpoints, as well as to guide the generation of images and depth maps from novel views. A key aspect of our approach is the multi-task generation of images and depth maps, using learnable task embeddings to guide the diffusion process towards specific modalities. We train this model on a collection of more than 60 million multi-view samples from publicly available datasets, and propose techniques to enable efficient and consistent learning in such diverse conditions. We also propose a novel strategy that enables the efficient training of larger models by incrementally fine-tuning smaller ones, with promising scaling behavior. Through extensive experiments, we report state-of-the-art results in multiple novel view synthesis benchmarks, as well as multi-view stereo and video depth estimation.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
https://github.com/webdataset/ webdataset, 2024
Webdataset. https://github.com/webdataset/ webdataset, 2024. 9
2024
-
[2]
STORM: Spatio-temporal reconstruction model for large-scale outdoor scenes
Anonymous. STORM: Spatio-temporal reconstruction model for large-scale outdoor scenes. In Submitted to The Thirteenth International Conference on Learning Repre- sentations, 2024. under review. 16
2024
-
[3]
Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P
Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neu- ral radiance fields. ICCV, 2021. 2
2021
-
[4]
Barron, Ben Mildenhall, Dor Verbin, Pratul P
Jonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. CVPR, 2022. 7
2022
-
[5]
Zoedepth: Zero-shot trans- fer by combining relative and metric depth, 2023
Shariq Farooq Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, and Matthias M ¨uller. Zoedepth: Zero-shot trans- fer by combining relative and metric depth, 2023. 3
2023
-
[6]
Midas v3.1 – a model zoo for robust monocular relative depth estima- tion, 2023
Reiner Birkl, Diana Wofk, and Matthias M¨uller. Midas v3.1 – a model zoo for robust monocular relative depth estima- tion, 2023. 3
2023
-
[7]
Yohann Cabon, Naila Murray, and Martin Humenberger. Virtual kitti 2. arXiv:2001.10773, 2020. 2
arXiv 2001
-
[8]
nuscenes: A mul- timodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A mul- timodal dataset for autonomous driving. In CVPR, 2020. 2
2020
Show all 125 references
-
[9]
Effi- cientvit: Enhanced linear attention for high-resolution low-computation visual recognition
Han Cai, Chuang Gan, and Song Han. Effi- cientvit: Enhanced linear attention for high-resolution low-computation visual recognition. arXiv preprint arXiv:2205.14756, 2022. 4
2022 arXiv
-
[10]
pi-gan: Periodic implicit genera- tive adversarial networks for 3d-aware image synthesis
Eric Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu, and Gordon Wetzstein. pi-gan: Periodic implicit genera- tive adversarial networks for 3d-aware image synthesis. In arXiv, 2020. 2
2020
-
[11]
Chan, Connor Z
Eric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas Guibas, Jonathan Tremblay, Sameh Khamis, Tero Karras, and Gordon Wetzstein. Efficient geometry-aware 3D generative adversarial networks. In arXiv, 2021. 2
2021
-
[12]
Chan, Koki Nagano, Matthew A
Eric R. Chan, Koki Nagano, Matthew A. Chan, Alexan- der W. Bergman, Jeong Joon Park, Axel Levy, Miika Ait- tala, Shalini De Mello, Tero Karras, and Gordon Wetzstein. GeNVS: Generative novel view synthesis with 3D-aware diffusion models. In arXiv, 2023. 2, 4
2023
-
[13]
Ziyi Chang, George Alex Koulieris, and Hubert P. H. Shum. On the design fundamentals of diffusion models: A survey,
-
[14]
pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction
David Charatan, Sizhe Li, Andrea Tagliasacchi, and Vin- cent Sitzmann. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. In arXiv,
-
[15]
Mvsplat: Efficient 3d gaussian splat- ting from sparse multi-view images
Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. Mvsplat: Efficient 3d gaussian splat- ting from sparse multi-view images. arXiv preprint arXiv:2403.14627, 2024. 1, 2, 4, 6
2024 arXiv
-
[16]
Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner
Angela Dai, Angel X. Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proc. Computer Vision and Pattern Recognition (CVPR), IEEE, 2017. 2, 8
2017
-
[17]
Depth-supervised nerf: Fewer views and faster training for free
Kangle Deng, Andrew Liu, Jun-Yan Zhu, , and Deva Ra- manan. Depth-supervised nerf: Fewer views and faster training for free. arXiv:2107.02791, 2021. 1, 2
2021 arXiv
-
[18]
Uncon- strained scene generation with locally conditioned radiance fields
Terrance DeVries, Miguel Angel Bautista, Nitish Srivas- tava, Graham W Taylor, and Joshua M Susskind. Uncon- strained scene generation with locally conditioned radiance fields. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 14304–14313, 2021. 2
2021
-
[19]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. In NeurIPS, 2021. 3
2021
-
[20]
Learning to render novel views from wide-baseline stereo pairs
Yilun Du, Cameron Smith, Ayush Tewari, and Vincent Sitz- mann. Learning to render novel views from wide-baseline stereo pairs. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR), pages 4970–4980, 2023. 6
2023
-
[21]
Omnidata: A scalable pipeline for making multi- task mid-level vision datasets from 3d scans
Ainaz Eftekhar, Alexander Sax, Jitendra Malik, and Amir Zamir. Omnidata: A scalable pipeline for making multi- task mid-level vision datasets from 3d scans. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 10786–10796, 2021. 3
2021
-
[22]
Self-supervised camera self-calibration from video
Jiading Fang, Igor Vasiljevic, Vitor Guizilini, Rares Am- brus, Greg Shakhnarovich, Adrien Gaidon, and Matthew Walter. Self-supervised camera self-calibration from video. In IEEE International Conference on Robotics and Automa- tion (ICRA), 2022. 3
2022
-
[23]
Ge- owizard: Unleashing the diffusion priors for 3d geometry estimation from a single image
Xiao Fu, Wei Yin, Mu Hu, Kaixuan Wang, Yuexin Ma, Ping Tan, Shaojie Shen, Dahua Lin, and Xiaoxiao Long. Ge- owizard: Unleashing the diffusion priors for 3d geometry estimation from a single image. arxiv, 2024. 3
2024
-
[24]
Srini- vasan, Jonathan T
Ruiqi Gao*, Aleksander Holynski*, Philipp Henzler, Arthur Brussee, Ricardo Martin-Brualla, Pratul P. Srini- vasan, Jonathan T. Barron, and Ben Poole*. Cat3d: Cre- ate anything in 3d with multi-view diffusion models.arXiv,
-
[25]
Cl ´ement Godard, Oisin Mac Aodha, Michael Firman, and Gabriel J. Brostow. Digging into self-supervised monocular depth prediction. In ICCV, 2019. 3, 9
2019
-
[26]
Accurate, large minibatch sgd: Training imagenet in 1 hour, 2018
Priya Goyal, Piotr Doll ´ar, Ross Girshick, Pieter Noord- huis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large minibatch sgd: Training imagenet in 1 hour, 2018. 5
2018
-
[27]
16 Nerfdiff: Single-image view synthesis with nerf-guided dis- tillation from 3d-aware diffusion
Jiatao Gu, Alex Trevithick, Kai-En Lin, Josh Susskind, Christian Theobalt, Lingjie Liu, and Ravi Ramamoorthi. 16 Nerfdiff: Single-image view synthesis with nerf-guided dis- tillation from 3d-aware diffusion. In International Confer- ence on Machine Learning, 2023. 2
2023
-
[28]
Sparsenerf: Distilling depth ranking for few-shot novel view synthesis
Guangcong, Zhaoxi Chen, Chen Change Loy, and Ziwei Liu. Sparsenerf: Distilling depth ranking for few-shot novel view synthesis. IEEE/CVF International Conference on Computer Vision (ICCV), 2023. 2
2023
-
[29]
Depthfm: Fast monocular depth estimation with flow matching, 2024
Ming Gui, Johannes S, Fischer, Ulrich Prestel, Pingchuan Ma, Olga Grebenkova Dmytro Kotovenko, Stefan Andreas Baumann, Vincent Tao Hu, and Bj ¨orn Ommer. Depthfm: Fast monocular depth estimation with flow matching, 2024. 3
2024
-
[30]
3d packing for self-supervised monocular depth estimation
Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Allan Raven- tos, and Adrien Gaidon. 3d packing for self-supervised monocular depth estimation. In CVPR, 2020. 4, 9, 12
2020
-
[31]
Semantically-guided representation learning for self-supervised monocular depth
Vitor Guizilini, Rui Hou, Jie Li, Rares Ambrus, and Adrien Gaidon. Semantically-guided representation learning for self-supervised monocular depth. In ICLR, 2020. 3
2020
-
[32]
Geometric unsupervised domain adaptation for semantic segmentation
Vitor Guizilini, Jie Li, Rares Ambrus, and Adrien Gaidon. Geometric unsupervised domain adaptation for semantic segmentation. In ICCV, 2021. 2
2021
-
[33]
Full surround mon- odepth from multiple cameras
Vitor Guizilini, Igor Vasiljevic, Rares Ambrus, Greg Shakhnarovich, and Adrien Gaidon. Full surround mon- odepth from multiple cameras. arXiv:2104.00152, 2021. 4
2021 arXiv
-
[34]
Multi-frame self-supervised depth with transformers
Vitor Guizilini, Rares Ambrus, Dian Chen, Sergey Za- kharov, and Adrien Gaidon. Multi-frame self-supervised depth with transformers. In Proceedings of the Interna- tional Conference on Computer Vision and Pattern Recog- nition (CVPR), 2022. 3
2022
-
[35]
Learning optical flow, depth, and scene flow with- out real-world labels
Vitor Guizilini, Kuan-Hui Lee, Rares Ambrus, and Adrien Gaidon. Learning optical flow, depth, and scene flow with- out real-world labels. IEEE Robotics and Automation Let- ters, 2022. 2
2022
-
[36]
Depth field networks for generalizable multi-view scene representation
Vitor Guizilini, Igor Vasiljevic, Jiading Fang, Rares Am- brus, Greg Shakhnarovich, Matthew R Walter, and Adrien Gaidon. Depth field networks for generalizable multi-view scene representation. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23...
2022
-
[37]
Towards zero-shot scale-aware monoc- ular depth estimation
Vitor Guizilini, Igor Vasiljevic, Dian Chen, Rares Ambrus, and Adrien Gaidon. Towards zero-shot scale-aware monoc- ular depth estimation. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision (ICCV), 2023. 2, 3, 4
2023
-
[38]
Delira: Self-supervised depth, light, and radiance fields
Vitor Guizilini, Igor Vasiljevic, Jiading Fang, Rares Am- brus, Sergey Zakharov, Vincent Sitzmann, and Adrien Gaidon. Delira: Self-supervised depth, light, and radiance fields. In International Conference on Computer Vision (ICCV), 2023. 1
2023
-
[39]
Grin: Zero-shot metric depth with pixel-level dif- fusion, 2024
Vitor Guizilini, Pavel Tokmakov, Achal Dave, and Rares Ambrus. Grin: Zero-shot metric depth with pixel-level dif- fusion, 2024. 2, 4, 5, 8
2024
-
[40]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. Advances in neural informa- tion processing systems, 33:6840–6851, 2020. 2, 3
2020
-
[41]
Cascaded diffusion models for high fidelity image generation
Jonathan Ho, Chitwan Saharia, William Chan, David J Fleet, Mohammad Norouzi, and Tim Salimans. Cascaded diffusion models for high fidelity image generation. The Journal of Machine Learning Research, 23(1):2249–2281,
-
[42]
One thousand and one hours: Self-driving motion prediction dataset
John Houston, Guido Zuidhof, Luca Bergamini, Yawei Ye, Long Chen, Ashesh Jain, Sammy Omari, Vladimir Iglovikov, and Peter Ondruska. One thousand and one hours: Self-driving motion prediction dataset. In CoRL, pages 409–418, 2020. 2
2020
-
[43]
Mvd-fusion: Single-view 3d via depth-consistent multi-view generation
Hanzhe Hu, Zhizhuo Zhou, Varun Jampani, and Shubham Tulsiani. Mvd-fusion: Single-view 3d via depth-consistent multi-view generation. In CVPR, 2024. 1, 3
2024
-
[44]
Metric3d v2: A versatile monocular geomet- ric foundation model for zero-shot metric depth and surface normal estimation
Mu Hu, Wei Yin, Chi Zhang, Zhipeng Cai, Xiaoxiao Long, Hao Chen, Kaixuan Wang, Gang Yu, Chunhua Shen, and Shaojie Shen. Metric3d v2: A versatile monocular geomet- ric foundation model for zero-shot metric depth and surface normal estimation. arXiv preprint arXiv:2404.15506, 2024. 3
2024 arXiv
-
[45]
Depthcrafter: Generating consistent long depth sequences for open-world videos
Wenbo Hu, Xiangjun Gao, Xiaoyu Li, Sijie Zhao, Xi- aodong Cun, Yong Zhang, Long Quan, and Ying Shan. Depthcrafter: Generating consistent long depth sequences for open-world videos. arXiv preprint arXiv:2409.02095 ,
-
[46]
Dpsnet: End-to-end deep plane sweep stereo
Sunghoon Im, Hae-Gon Jeon, Stephen Lin, and In So Kweon. Dpsnet: End-to-end deep plane sweep stereo. arXiv:1905.00538, 2019. 8
1905 arXiv
-
[47]
Neo 360: Neural fields for sparse view synthesis of outdoor scenes
Muhammad Zubair Irshad, Sergey Zakharov, Katherine Liu, Vitor Guizilini, Thomas Kollar, Adrien Gaidon, Zsolt Kira, and Rares Ambrus. Neo 360: Neural fields for sparse view synthesis of outdoor scenes. In Interntaional Confer- ence on Computer Vision (ICCV), 2023. 2
2023
-
[48]
Nerf-mae: Masked autoencoders for self-supervised 3d rep- resentation learning for neural radiance fields
Muhammad Zubair Irshad, Sergey Zakharov, Vitor Guizilini, Adrien Gaidon, Zsolt Kira, and Rares Ambrus. Nerf-mae: Masked autoencoders for self-supervised 3d rep- resentation learning for neural radiance fields. In European Conference on Computer Vision (ECCV), 2024. 2
2024
-
[49]
Fleet, and Ting Chen
Allan Jabri, David J. Fleet, and Ting Chen. Scalable adap- tive computation for iterative generation. In ICML, 2023. 2, 3
2023
-
[50]
Putting nerf on a diet: Semantically consistent few-shot view synthesis
Ajay Jain, Matthew Tancik, and Pieter Abbeel. Putting nerf on a diet: Semantically consistent few-shot view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 5885–5894, 2021. 2
2021
-
[51]
Large scale multi-view stereopsis eval- uation
Rasmus Jensen, Anders Dahl, George V ogiatzis, Engil Tola, and Henrik Aanæs. Large scale multi-view stereopsis eval- uation. In 2014 IEEE Conference on Computer Vision and Pattern Recognition, pages 406–413. IEEE, 2014. 7
2014
-
[52]
Holodiffusion: Training a 3D diffusion model using 2D images
Animesh Karnewar, Andrea Vedaldi, David Novotny, and Niloy Mitra. Holodiffusion: Training a 3D diffusion model using 2D images. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , 2023. 2
2023
-
[53]
Re- purposing diffusion-based image generators for monocular depth estimation, 2023
Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Met- zger, Rodrigo Caye Daudt, and Konrad Schindler. Re- purposing diffusion-based image generators for monocular depth estimation, 2023. 3, 5 17
2023
-
[54]
Deep multi-view depth estimation with predicted uncertainty
Tong Ke, Tien Do, Khiem Vuong, Kourosh Sartipi, and Stergios I Roumeliotis. Deep multi-view depth estimation with predicted uncertainty. arXiv:2011.09594, 2020. 3
2011 arXiv
-
[55]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42(4), 2023. 1, 2
2023
-
[56]
Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ash- win Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yun- liang Chen, Kirsty Ellis, Peter David Fagan, Joey Hejna, Masha Itkina, Marion Lepert, Yecheng Jason Ma, Patrick Tree ...
2024
-
[57]
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv:1412.6980, 2014. 5
2014 arXiv
-
[58]
Nor- mal assisted stereo depth estimation
Uday Kusupati, Shuo Cheng, Rui Chen, and Hao Su. Nor- mal assisted stereo depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2189–2199, 2020. 8
2020
-
[59]
Infinite nature: Perpetual view generation of natural scenes from a single image
Andrew Liu, Richard Tucker, Varun Jampani, Ameesh Makadia, Noah Snavely, and Angjoo Kanazawa. Infinite nature: Perpetual view generation of natural scenes from a single image. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision (ICCV), 2021. 6
2021
-
[60]
Zero-1-to-3: Zero-shot one image to 3d object
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tok- makov, Sergey Zakharov, and Carl V ondrick. Zero-1-to-3: Zero-shot one image to 3d object. In ICCV, 2023. 1, 2, 3
2023
-
[61]
Syncdreamer: Generating multiview-consistent images from a single-view image
Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Generating multiview-consistent images from a single-view image. arXiv preprint arXiv:2309.03453, 2023. 1, 2
2023 arXiv
-
[62]
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization, 2019. 5
2019
-
[63]
Consistent video depth estimation
Xuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen, and Johannes Kopf. Consistent video depth estimation. ACM Transactions on Graphics (TOG), 39(4):71–1, 2020. 3, 8
2020
-
[64]
Srinivasan, Rodrigo Ortiz-Cayon, Nima Khademi Kalantari, Ravi Ramamoorthi, Ren Ng, and Abhishek Kar
Ben Mildenhall, Pratul P. Srinivasan, Rodrigo Ortiz-Cayon, Nima Khademi Kalantari, Ravi Ramamoorthi, Ren Ng, and Abhishek Kar. Local light field fusion: Practical view syn- thesis with prescriptive sampling guidelines. ACM Trans- actions on Graphics (TOG), 2019. 7
2019
-
[65]
NeRF: Representing scenes as neural radiance fields for view syn- thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view syn- thesis. In Proceedings of the European Conference on Com- puter Vision (ECCV), pages 405–421, 2020. 1, 2, 4
2020
-
[66]
Instant neural graphics primitives with a multiresolution hash encoding
Thomas M ¨uller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (ToG), 41(4):1–15, 2022. 2
2022
-
[67]
Giraffe: Repre- senting scenes as compositional generative neural feature fields
Michael Niemeyer and Andreas Geiger. Giraffe: Repre- senting scenes as compositional generative neural feature fields. In Proc. IEEE Conf. on Computer Vision and Pat- tern Recognition (CVPR), 2021. 2
2021
-
[68]
Barron, Ben Mildenhall, Mehdi S
Michael Niemeyer, Jonathan T. Barron, Ben Mildenhall, Mehdi S. M. Sajjadi, Andreas Geiger, and Noha Radwan. Regnerf: Regularizing neural radiance fields for view syn- thesis from sparse inputs. InProc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2022. 2
2022
-
[69]
Scalable diffusion mod- els with transformers
William Peebles and Saining Xie. Scalable diffusion mod- els with transformers. arXiv preprint arXiv:2212.09748 ,
-
[70]
UniDepth: Universal monocular metric depth estimation
Luigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis, Mat- tia Segu, Siyuan Li, Luc Van Gool, and Fisher Yu. UniDepth: Universal monocular metric depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 4
2024
-
[71]
Habitat-matterport 3d dataset (HM3d): 1000 large-scale 3d environments for embodied AI
Santhosh Kumar Ramakrishnan, Aaron Gokaslan, Erik Wi- jmans, Oleksandr Maksymets, Alexander Clegg, John M Turner, Eric Undersander, Wojciech Galuba, Andrew West- bury, Angel X Chang, Manolis Savva, Yili Zhao, and Dhruv Batra. Habitat-matterport 3d dataset (HM3d): 1000 large-sc...
2021
-
[72]
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In Interna- tional conference on machine learning , pages 8821–8831. Pmlr, 2021. 2
2021
-
[73]
Towards robust monocu- lar depth estimation: Mixing datasets for zero-shot cross- dataset transfer
Ren ´e Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocu- lar depth estimation: Mixing datasets for zero-shot cross- dataset transfer. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022. 3, 8
2022
-
[74]
Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction
Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotny. Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. In International Con- ference on Computer Vision, 2021. 2, 7 18
2021
-
[75]
Susskind
Mike Roberts, Jason Ramapuram, Anurag Ranjan, At- ulit Kumar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M. Susskind. Hypersim: A photorealis- tic synthetic dataset for holistic indoor scene understanding. In International Conference on Computer Vision (ICCV) ...
2021
-
[76]
Barron, Ben Mildenhall, Pratul P
Barbara Roessle, Jonathan T. Barron, Ben Mildenhall, Pratul P. Srinivasan, and Matthias Nießner. Dense depth priors for neural radiance fields from sparse input views. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 1, 2
2022
-
[77]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 2, 3
2022
-
[78]
U- net: Convolutional networks for biomedical image segmen- tation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. In MICCAI. Springer, 2015. 3
2015
-
[79]
Mehdi S. M. Sajjadi, Henning Meyer, Etienne Pot, Urs Bergmann, Klaus Greff, Noha Radwan, Suhani V ora, Mario Lucic, Daniel Duckworth, Alexey Dosovitskiy, Jakob Uszkoreit, Thomas Funkhouser, and Andrea Tagliasacchi. Scene Representation Transformer: Geometry-Free Novel View Syn...
2022
-
[80]
Ze- roNVS: Zero-shot 360-degree view synthesis from a single real image
Kyle Sargent, Zizhang Li, Tanmay Shah, Charles Her- rmann, Hong-Xing Yu, Yunzhi Zhang, Eric Ryan Chan, Dmitry Lagun, Li Fei-Fei, Deqing Sun, and Jiajun Wu. Ze- roNVS: Zero-shot 360-degree view synthesis from a single real image. CVPR, 2024, 2023. 2, 7
2024
-
[81]
Saurabh Saxena, Junhwa Hur, Charles Herrmann, Deqing Sun, and David J. Fleet. Zero-shot metric depth with a field- of-view conditioned diffusion model, 2023. 5
2023
-
[82]
Structure-from-motion revisited
Johannes L Schonberger and Jan-Michael Frahm. Structure-from-motion revisited. In CVPR, 2016. 4, 6, 7
2016
-
[83]
V oxgraf: Fast 3d-aware image syn- thesis with sparse voxel grids
Katja Schwarz, Axel Sauer, Michael Niemeyer, Yiyi Liao, and Andreas Geiger. V oxgraf: Fast 3d-aware image syn- thesis with sparse voxel grids. In Advances in Neural In- formation Processing Systems (NeurIPS) (NeurIPS), 2022. 2
2022
-
[84]
Learning tem- porally consistent video depth from video diffusion priors,
Jiahao Shao, Yuanbo Yang, Hongyu Zhou, Youmin Zhang, Yujun Shen, Matteo Poggi, and Yiyi Liao. Learning tem- porally consistent video depth from video diffusion priors,
-
[85]
Feature-metric loss for self-supervised learning of depth and egomotion
Chang Shu, Kun Yu, Zhixiang Duan, and Kuiyuan Yang. Feature-metric loss for self-supervised learning of depth and egomotion. In ECCV, 2020. 9
2020
-
[86]
Freeman, Joshua B
Vincent Sitzmann, Semon Rezchikov, William T. Freeman, Joshua B. Tenenbaum, and Fredo Durand. Light field net- works: Neural scene representations with single-evaluation rendering. In Proc. NeurIPS, 2021. 1
2021
-
[87]
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In ICML, 2015. 3
2015
-
[88]
SimpleNeRF: Regularizing sparse input neural radiance fields with simpler solutions
Nagabhushan Somraj, Adithyan Karanayil, and Rajiv Soundararajan. SimpleNeRF: Regularizing sparse input neural radiance fields with simpler solutions. In SIG- GRAPH Asia, 2023. 2, 7
2023
-
[89]
Denois- ing diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois- ing diffusion implicit models. arXiv:2010.02502, 2020. 5
2010 arXiv
-
[90]
Sturm, N
J. Sturm, N. Engelhard, F. Endres, W. Burgard, and D. Cre- mers. A benchmark for the evaluation of rgb-d slam sys- tems. In Proc. of the International Conference on Intelli- gent Robot Systems (IROS), 2012. 8
2012
-
[91]
Generalizable patch-based neural ren- dering
Mohammed Suhail, Carlos Esteves, Leonid Sigal, and Ameesh Makadia. Generalizable patch-based neural ren- dering. In European Conference on Computer Vision, pages 156–174. Springer, 2022. 6
2022
-
[92]
NeuralRecon: Real-time coherent 3D re- construction from monocular video
Jiaming Sun, Yiming Xie, Linghao Chen, Xiaowei Zhou, and Hujun Bao. NeuralRecon: Real-time coherent 3D re- construction from monocular video. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR), pages 15598–15607, 2021. 7, 8
2021
-
[93]
Scalability in perception for autonomous driving: Waymo open dataset
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aure- lien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF conference on computer ...
2020
-
[94]
BA-Net: Dense bundle ad- justment network
Chengzhou Tang and Ping Tan. BA-Net: Dense bundle ad- justment network. arXiv preprint arXiv:1806.04807, 2018. 8
2018 arXiv
-
[95]
Deepv2d: Video to depth with differentiable structure from motion
Zachary Teed and Jia Deng. Deepv2d: Video to depth with differentiable structure from motion. In ICLR, 2020. 8
2020
-
[96]
Tenenbaum, Fr ´edo Durand, William T
Ayush Tewari, Tianwei Yin, George Cazenavette, Se- mon Rezchikov, Joshua B. Tenenbaum, Fr ´edo Durand, William T. Freeman, and Vincent Sitzmann. Diffusion with forward models: Solving stochastic inverse problems with- out direct supervision. In arXiv, 2023. 2
2023
-
[97]
Rtmv: A ray-traced multi-view synthetic dataset for novel view synthesis
Jonathan Tremblay, Moustafa Meshry, Alex Evans, Jan Kautz, Alexander Keller, Sameh Khamis, Charles Loop, Nathan Morrical, Koki Nagano, Towaki Takikawa, and Stan Birchfield. Rtmv: A ray-traced multi-view synthetic dataset for novel view synthesis. IEEE/CVF European Conference o...
2022
-
[98]
Grf: Learning a general radi- ance field for 3d representation and rendering
Alex Trevithick and Bo Yang. Grf: Learning a general radi- ance field for 3d representation and rendering. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 15182–15192, 2021. 2
2021
-
[99]
Demon: Depth and motion network for learning monocular stereo
Benjamin Ummenhofer, Huizhong Zhou, Jonas Uhrig, Nikolaus Mayer, Eddy Ilg, Alexey Dosovitskiy, and Thomas Brox. Demon: Depth and motion network for learning monocular stereo. In Proceedings of the IEEE con- ference on computer vision and pattern recognition , pages 5038–5047, 2017. 8
2017
-
[100]
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and koray kavukcuoglu. Neural discrete representation learning. In Advances in Neural Information Processing Systems . Curran Associates, Inc., 2017. 3
2017
-
[101]
Recurrent interface network (pytorch), 2022
Phil Wang. Recurrent interface network (pytorch), 2022. 5
2022
-
[102]
Tartanair: A dataset to push the limits of visual slam
Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor, and 19 Sebastian Scherer. Tartanair: A dataset to push the limits of visual slam. In IROS, 2020. 2
2020
-
[103]
Novel view synthesis with diffusion models,
Daniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho, Andrea Tagliasacchi, and Mohammad Norouzi. Novel view synthesis with diffusion models,
-
[104]
Nerfingmvs: Guided optimization of neu- ral radiance fields for indoor multi-view stereo
Yi Wei, Shaohui Liu, Yongming Rao, Wang Zhao, Jiwen Lu, and Jie Zhou. Nerfingmvs: Guided optimization of neu- ral radiance fields for indoor multi-view stereo. In ICCV,
-
[105]
Ar- goverse 2: Next generation datasets for self-driving percep- tion and forecasting
Benjamin Wilson, William Qi, Tanmay Agarwal, John Lambert, Jagjeet Singh, Siddhesh Khandelwal, Bowen Pan, Ratnesh Kumar, Andrew Hartnett, Jhony Kaesemodel Pontes, Deva Ramanan, Peter Carr, and James Hays. Ar- goverse 2: Next generation datasets for self-driving percep- tion an...
2021
-
[106]
Srinivasan, Dor Verbin, Jonathan T
Rundi Wu, Ben Mildenhall, Philipp Henzler, Keunhong Park, Ruiqi Gao, Daniel Watson, Pratul P. Srinivasan, Dor Verbin, Jonathan T. Barron, Ben Poole, and Aleksander Holynski. Reconfusion: 3d reconstruction with diffusion priors. arXiv, 2023. 1, 2, 4, 6, 7
2023
-
[107]
DiffusioNeRF: Regularizing Neural Radiance Fields with Denoising Diffu- sion Models
Jamie Wynn and Daniyar Turmukhambetov. DiffusioNeRF: Regularizing Neural Radiance Fields with Denoising Diffu- sion Models. In CVPR, 2023. 2
2023
-
[108]
Rgbd objects in the wild: Scaling real-world 3d object learning from rgb-d videos, 2024
Hongchi Xia, Yang Fu, Sifei Liu, and Xiaolong Wang. Rgbd objects in the wild: Scaling real-world 3d object learning from rgb-d videos, 2024. 2
2024
-
[109]
Sun3d: A database of big spaces reconstructed using sfm and object labels
Jianxiong Xiao, Andrew Owens, and Antonio Torralba. Sun3d: A database of big spaces reconstructed using sfm and object labels. In 2013 IEEE International Conference on Computer Vision, pages 1625–1632, 2013. 8
2013
-
[110]
Murf: Multi-baseline radiance fields
Haofei Xu, Anpei Chen, Yuedong Chen, Christos Sakaridis, Yulun Zhang, Marc Pollefeys, Andreas Geiger, and Fisher Yu. Murf: Multi-baseline radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20041–20050, 2024. 6
2024
-
[111]
Dmv3d: Denoising multi- view diffusion using 3d large reconstruction model, 2023
Yinghao Xu, Hao Tan, Fujun Luan, Sai Bi, Peng Wang, Jiahao Li, Zifan Shi, Kalyan Sunkavalli, Gordon Wetzstein, Zexiang Xu, and Kai Zhang. Dmv3d: Denoising multi- view diffusion using 3d large reconstruction model, 2023. 1
2023
-
[112]
$SE(3)$ equivariant ray em- beddings for implicit multi-view depth estimation
Yinshuang Xu, Dian Chen, Katherine Liu, Sergey Za- kharov, Rares Andrei Ambrus, Kostas Daniilidis, and Vi- tor Campagnolo Guizilini. $SE(3)$ equivariant ray em- beddings for implicit multi-view depth estimation. In The Thirty-eighth Annual Conference on Neural Information Proc...
2024
-
[113]
Freenerf: Im- proving few-shot neural rendering with free frequency reg- ularization
Jiawei Yang, Marco Pavone, and Yue Wang. Freenerf: Im- proving few-shot neural rendering with free frequency reg- ularization. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2023. 7
2023
-
[114]
Depth anything: Un- leashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Ji- ashi Feng, and Hengshuang Zhao. Depth anything: Un- leashing the power of large-scale unlabeled data. In CVPR,
-
[115]
Depth any- thing v2
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiao- gang Xu, Jiashi Feng, and Hengshuang Zhao. Depth any- thing v2. arXiv:2406.09414, 2024. 3
2024 arXiv
-
[116]
MVSNet: Depth inference for unstructured multi- view stereo
Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan. MVSNet: Depth inference for unstructured multi- view stereo. In Proceedings of the European Conference on Computer Vision (ECCV), pages 767–783, 2018. 3
2018
-
[117]
Blendedmvs: A large-scale dataset for generalized multi-view stereo net- works
Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan. Blendedmvs: A large-scale dataset for generalized multi-view stereo net- works. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 1790– 1799...
2020
-
[118]
Input-level inductive biases for 3D reconstruction
Wang Yifan, Carl Doersch, Relja Arandjelovi ´c, Jo ˜ao Car- reira, and Andrew Zisserman. Input-level inductive biases for 3D reconstruction. In Proceedings of the IEEE Interna- tional Conference on Computer Vision (CVPR), 2022. 8
2022
-
[119]
Met- ric3d: Towards zero-shot metric 3d prediction from a single image
Wei Yin, Chi Zhang, Hao Chen, Zhipeng Cai, Gang Yu, Kaixuan Wang, Xiaozhi Chen, and Chunhua Shen. Met- ric3d: Towards zero-shot metric 3d prediction from a single image. In ICCV, 2023. 3, 4
2023
-
[120]
Dreamsparse: Escaping from plato’s cave with 2d frozen diffusion model given sparse views
Paul Yoo, Jiaxian Guo, Yutaka Matsuo, and Shixiang Shane Gu. Dreamsparse: Escaping from plato’s cave with 2d frozen diffusion model given sparse views. arXiv preprint arXiv:2306.03414, 2023. 2
2023 arXiv
-
[121]
pixelnerf: Neural radiance fields from one or few images
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In CVPR, 2021. 1, 2, 3, 4, 6
2021
-
[122]
Mvimgnet: A large-scale dataset of multi-view images
Xianggang Yu, Mutian Xu, Yidan Zhang, Haolin Liu, Chongjie Ye, Yushuang Wu, Zizheng Yan, Tianyou Liang, Guanying Chen, Shuguang Cui, and Xiaoguang Han. Mvimgnet: A large-scale dataset of multi-view images. In CVPR, 2023. 2
2023
-
[123]
Taskonomy: Disentangling task transfer learning
Amir R Zamir, Alexander Sax, , William B Shen, Leonidas Guibas, Jitendra Malik, and Silvio Savarese. Taskonomy: Disentangling task transfer learning. In 2018 IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR). IEEE, 2018. 2
2018
-
[124]
Monst3r: A simple approach for esti- mating geometry in the presence of motion
Junyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jam- pani, Trevor Darrell, Forrester Cole, Deqing Sun, and Ming-Hsuan Yang. Monst3r: A simple approach for esti- mating geometry in the presence of motion. arXiv preprint arxiv:2410.03825, 2024. 9
-
[125]
Stereo magnification: Learning view synthesis using multiplane images
Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: Learning view synthesis using multiplane images. In SIGGRAPH, 2018. 2, 6, 7 20
2018
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.