REVIEW 4 minor 2 cited by
LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis
T0 review · 0 major / 4 minor · reviewed 2026-07-13 · grok-4.5
Pith's one-line read Even without building an explicit 3D model, novel-view synthesis still gains from 3D-aware features taken from a geometry-pretrained encoder, yielding real-time state-of-the-art feed-forward rendering.
desk verdict Solid systems paper: VGGT-initialized highway encoder-decoder delivers real SOTA real-time feed-forward NVS with clean ablations; no load-bearing flaw. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
LagerNVS: a highway encoder-decoder whose encoder is initialized from a geometry-pretrained multi-view transformer backbone, reading out concatenated last local- and global-attention tokens as 3D-aware latents; a lightweight transformer decoder then conditions on Plücker ray maps of the target camera and renders the novel view. The highway path avoids a fixed token bottleneck so capacity scales with the number of source images while encoding cost is amortized.
What would settle it
Train the same highway architecture from scratch or from generic 2D features at full scale and check whether the reported multi-decibel PSNR gap from 3D pre-training disappears on RealEstate10k and DL3DV; or freeze the encoder entirely and verify that reflections and textures remain unusable as the ablations claim.
Extended reading notes
Core claim
Reconstruction-free novel view synthesis still benefits strongly from 3D-aware latent features obtained by initializing the encoder from a network pre-trained with explicit 3D supervision. Combined with a highway encoder-decoder that preserves full source-image information flow and end-to-end photometric fine-tuning, this produces state-of-the-art deterministic feed-forward NVS (including 31.4 PSNR on RealEstate10k), real-time rendering, and generalization with or without source cameras.
Load-bearing premise
The intermediate tokens taken from a geometry-pretrained backbone still carry enough color, texture, and reflectance after fine-tuning to support high-fidelity image rendering; if they mostly encode shape and discard appearance, the photometric gains collapse.
Editorial extensions
If this is right
- Reconstruction-free feed-forward NVS can outperform methods that still emit explicit 3D Gaussians or radiance fields.
- Real-time 512×512 rendering on a single GPU becomes practical for roughly up to nine source views with a pure neural decoder.
- One model trained on a large multi-dataset mix can handle posed and unposed, square and non-square, ego-centric and 360° inputs without retuning.
- The same decoder can be fine-tuned with a denoising objective to sample plausible completions of occluded or extrapolated regions instead of averaging them away.
- Geometry foundation models that never see a rendering loss during pre-training force expensive end-to-end fine-tuning if they are to be reused for appearance-critical tasks.
Reading between the lines
- Geometry backbones that jointly optimize a photometric rendering head during pre-training would likely transfer to NVS with far less fine-tuning than pure reconstruction models.
- The highway-versus-bottleneck distinction is likely to matter for other multi-view tasks where token capacity, not just geometry, limits quality (correspondence, relighting, material estimation).
- Pairing 3D-aware encoder initialization with compute-scaling recipes for encoder-decoder transformers could raise quality further at fixed training budget.
- If appearance is systematically discarded by pure geometry pre-training, multi-task pre-training that keeps both geometry and photometry may become the default recipe for multi-view foundation models.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. LagerNVS is a feed-forward encoder–decoder for novel view synthesis that avoids explicit 3D reconstruction. The encoder is initialized from VGGT (pre-trained with explicit 3D supervision) and produces per-image latent tokens; a lightweight ViT decoder, conditioned on a Plücker ray map of the target camera, renders the novel view. The authors compare highway vs. bottleneck encoder–decoder and decoder-only designs, train on a large multi-dataset mix, and report state-of-the-art deterministic NVS (31.39 PSNR on RealEstate10k under LVSM’s protocol), real-time decoding (30+ FPS at 512² with up to 9 views), operation with or without source cameras, and a preliminary diffusion decoder for generative extrapolation. Ablations isolate 3D pre-training, end-to-end fine-tuning, and architecture.
Significance. If the reported gains hold under the matched protocols, the work is a clear advance for reconstruction-free NVS. It shows that strong 3D inductive bias can be injected via pre-trained latent features rather than explicit geometry, while still delivering real-time rendering and strong generalization (including unposed and in-the-wild inputs). The highway encoder–decoder design, multi-dataset training, and public code/checkpoints make the result immediately usable and extensible. The diffusion fine-tuning experiment further indicates a practical path from deterministic to generative NVS without redesigning the encoder.
minor comments (4)
- Table 2 and the main Re10k numbers lack error bars or multi-seed statistics; a short note on run-to-run variance would strengthen confidence in the +1.7 dB margin.
- The shared-focal-length assumption (Sec. 3 and Limitations) is stated clearly but could be flagged earlier in the method section so readers know the scope of the camera model before the experiments.
- Fig. 7 and the occlusion examples (Fig. A3) are informative; adding a brief quantitative measure of failure modes (e.g., high-frequency texture or large baseline) would help readers gauge remaining limitations.
- Clarify the exact definition of the canonical FoV k0 used at test time when cameras are unknown (π/2 vs. 53.13°) in one place to avoid confusion between the main text and the v1/v2 appendix note.
Circularity Check
No significant circularity: SOTA claims rest on external photometric metrics and held-out benchmarks after end-to-end fine-tuning, not on any quantity forced by definition or self-citation chain.
full rationale
LagerNVS is an empirical encoder-decoder NVS system. The encoder is initialized from VGGT (overlapping authors) and then fine-tuned end-to-end with L2 + perceptual losses on multi-view image tuples; the decoder is a standard ViT that maps Plücker ray tokens plus latent features to RGB. All reported numbers (31.4 PSNR on Re10k, gains vs LVSM/DepthSplat/AnySplat, real-time FPS, generalization) are measured by standard image metrics on held-out target views under fixed evaluation protocols. Ablations (Table 2) explicitly contrast frozen vs fine-tuned VGGT, 3D vs 2D vs no pre-training, and highway vs bottleneck vs decoder-only architectures; none of these comparisons reduce the target PSNR/SSIM/LPIPS to a fitted free parameter or to a self-cited uniqueness theorem. The single self-citation of VGGT supplies only an initialization that is subsequently optimized and evaluated independently; it is not load-bearing for the central claim. No equation equates a reported metric to an input by construction, no uniqueness result is imported to forbid alternatives, and no known empirical pattern is merely renamed. Score 1 reflects only the ordinary (non-circular) use of an overlapping-author backbone as a starting point.
Assumptions & free parameters
free parameters (4)
- loss weights λ2, λp
- decoder depth / width (ViT-B, 12 blocks)
- canonical horizontal FoV k0 = π/2 (or 53.13°)
- scene-scale normalizations w1, w2
assumptions (4)
- domain assumption Photometric L2 + perceptual loss is a sufficient training signal for high-quality NVS
- domain assumption VGGT’s last local and global attention tokens contain transferable 3D-aware features useful for appearance rendering after fine-tuning
- ad hoc to paper Source and target cameras share identical focal length (and equal horizontal/vertical FoV)
- domain assumption Scenes are approximately static
invented entities (1)
-
highway encoder-decoder latent geometry tokens
Cite this review
Pith. "Pith review of LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis." pith.science (2026). https://pith.science/paper/M2AKLFVL
@misc{pith2026260320176,
author = {Pith},
title = {Pith review of: LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis},
year = {2026},
howpublished = {\url{https://pith.science/paper/M2AKLFVL}},
note = {Machine review of arXiv:2603.20176}
}
read the original abstract
Recent work has shown that neural networks can perform 3D tasks such as Novel View Synthesis (NVS) without explicit 3D reconstruction. Even so, we argue that strong 3D inductive biases are still helpful in the design of such networks. We show this point by introducing LagerNVS, an encoder-decoder neural network for NVS that builds on `3D-aware' latent features. The encoder is initialized from a 3D reconstruction network pre-trained using explicit 3D supervision. This is paired with a lightweight decoder, and trained end-to-end with photometric losses. LagerNVS achieves state-of-the-art deterministic feed-forward Novel View Synthesis (including 31.4 PSNR on Re10k), with and without known cameras, renders in real time, generalizes to in-the-wild data, and can be paired with a diffusion decoder for generative extrapolation.
Forward citations
Cited by 2 Pith papers
-
InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis
A single-image feed-forward Gaussian splatting method that samples supports from predicted depth and decodes Gaussian attributes implicitly, improving cross-dataset large-baseline novel view synthesis.
-
RayViT: Ray-Conditioned Visual Representations for Viewpoint-Robust Imitation Learning
Conditioning a pretrained ViT on per-pixel Plücker camera rays — via a gated-cross-attention class token and patch-level ray embeddings — makes imitation-learned manipulation policies substantially more robust to came...
Reference graph
Works this paper leans on
-
[1]
Adelson and R
E. Adelson and R. Bergen.The Plenoptic Function and the Elements of Early Vision. MIT Press, 1991. 3
1991
-
[2]
PyTorch 2: Faster machine learning through dynamic Python bytecode transformation and graph compilation
Jason Ansel, Edward Yang, Horace He, et al. PyTorch 2: Faster machine learning through dynamic Python bytecode transformation and graph compilation. InProc. ACM Inter- national Conference on Architectural Support for Program- ming Languages and Operating Systems (ASPLOS), 2024. 2
2024
-
[3]
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization.arXiv.cs, abs/1607.06450, 2016. 4, 15
arXiv 2016
-
[4]
Lindell, Zan Gojcic, Sanja Fidler, Huan Ling, Jun Gao, and Xuanchi Ren
Sherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang, Yifeng Jiang, Haithem Turki, Andrea Tagliasacchi, David B. Lindell, Zan Gojcic, Sanja Fidler, Huan Ling, Jun Gao, and Xuanchi Ren. Lyra: Generative 3D scene recon- struction via video diffusion model self-distillation.arXiv, 2509.19296, 2025. 3
arXiv 2025
-
[5]
ReCamMaster: camera-controlled gen- erative rendering from a single video
Jianhong Bai, Menghan Xia, Xiao Fu, Xintao Wang, Lianrui Mu, Jinwen Cao, Zuozhu Liu, Haoji Hu, Xiang Bai, Pengfei Wan, and Di Zhang. ReCamMaster: camera-controlled gen- erative rendering from a single video. InProc. ICCV, 2025. 3, 16
2025
-
[6]
ARK- itscenes - a diverse real-world dataset for 3d indoor scene understanding using mobile RGB-d data
Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Tal Dimry, Yuri Feigin, Peter Fu, Thomas Gebauer, Brandon Joffe, Daniel Kurz, Arik Schwartz, and Elad Shulman. ARK- itscenes - a diverse real-world dataset for 3d indoor scene understanding using mobile RGB-d data. InProc. NeurIPS,
-
[7]
Gortler, and Michael F
Chris Buehler, Michael Bosse, Leonard McMillan, Steven J. Gortler, and Michael F. Cohen. Unstructured lumigraph ren- dering. InProc. SIGGRAPH, 2001. 3
2001
-
[8]
Chan, Koki Nagano, Matthew A
Eric R. Chan, Koki Nagano, Matthew A. Chan, Alexan- der W. Bergman, Jeong Joon Park, Axel Levy, Miika Aittala, Shalini De Mello, Tero Karras, and Gordon Wetzstein. Gen- erative novel view synthesis with 3D-aware diffusion mod- els. InProc. ICCV, 2023. 3, 15
2023
Show all 97 references
-
[9]
pixelSplat: 3D Gaussian splats from image pairs for scalable generalizable 3D reconstruction
David Charatan, Sizhe Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelSplat: 3D Gaussian splats from image pairs for scalable generalizable 3D reconstruction. InProc. CVPR,
-
[10]
MVSNeRF: Fast generalizable radiance field reconstruction from multi-view stereo
Anpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang, Fanbo Xiang, Jingyi Yu, and Hao Su. MVSNeRF: Fast generalizable radiance field reconstruction from multi-view stereo. InProc. ICCV, 2021. 2
2021
-
[11]
MVSplat: efficient 3D gaussian splatting from sparse multi-view images
Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. MVSplat: efficient 3D gaussian splatting from sparse multi-view images. InProc. ECCV, 2024. 1, 3, 6
2024
-
[12]
MVS- plat360: Feed-forward 360 Scene Synthesis from Sparse Views
Yuedong Chen, Chuanxia Zheng, Haofei Xu, Bohan Zhuang, Andrea Vedaldi, Tat-Jen Cham, and Jianfei Cai. MVS- plat360: Feed-forward 360 Scene Synthesis from Sparse Views. InProc. NeurIPS, 2024. 1, 3
2024
-
[13]
FlashAttention-2: Faster attention with better par- allelism and work partitioning
Tri Dao. FlashAttention-2: Faster attention with better par- allelism and work partitioning. InProc. ICLR, 2024. 5
2024
-
[14]
Fu, Stefano Ermon, Atri Rudra, and Christopher R´e
Tri Dao, Daniel Y . Fu, Stefano Ermon, Atri Rudra, and Christopher R´e. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. InProc. NeurIPS, 2022. 5
2022
-
[15]
Vision transformers need registers
Timoth ´ee Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski. Vision transformers need registers. InProc. ICLR, 2024. 5
2024
-
[16]
Jimmy Ba Diederik P. Kingma. Adam: A method for stochastic optimization. InProc. ICLR, 2015. 18
2015
-
[17]
An image is worth 16×16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16×16 words: Transformers for image recognition ...
2021
-
[18]
MASt3R-SfM: a fully-integrated solution for unconstrained structure-from-motion
Bardienus Duisterhof, Lojze Zust, Philippe Weinzaepfel, Vincent Leroy, Yohann Cabon, and Jerome Revaud. MASt3R-SfM: a fully-integrated solution for unconstrained structure-from-motion. InProc. 3DV, 2025. 3
2025
-
[19]
Sigmoid- weighted linear units for neural network function approxima- tion in reinforcement learning.Neural Networks, 107, 2018
Stefan Elfwing, Eiji Uchibe, and Kenji Doya. Sigmoid- weighted linear units for neural network function approxima- tion in reinforcement learning.Neural Networks, 107, 2018. 15
2018
-
[20]
FlowR: flowing from sparse to dense 3d re- constructions
Tobias Fischer, Samuel Rota Bul `o, Yung-Hsu Yang, Nikhil Varma Keetha, Lorenzo Porzi, Norman M ¨uller, Katja Schwarz, Jonathon Luiten, Marc Pollefeys, and Peter Kontschieder. FlowR: flowing from sparse to dense 3d re- constructions. InProc. ICCV, 2025. 3
2025
-
[21]
Barron, and Ben Poole
Ruiqi Gao, Aleksander Holynski, Philipp Henzler, Arthur Brussee, Ricardo Martin-Brualla, Pratul Srinivasan, Jonathan T. Barron, and Ben Poole. CAT3D: create anything in 3d with multi-view diffusion models. InProc. NeurIPS,
-
[22]
Gortler, Radek Grzeszczuk, Richard Szeliski, and Michael F
Steven J. Gortler, Radek Grzeszczuk, Richard Szeliski, and Michael F. Cohen. The lumigraph. InProc. SIGGRAPH,
-
[23]
Susskind, Christian Theobalt, Lingjie Liu, and Ravi Ramamoorthi
Jiatao Gu, Alex Trevithick, Kai-En Lin, Josh M. Susskind, Christian Theobalt, Lingjie Liu, and Ravi Ramamoorthi. NerfDiff: Single-image view synthesis with nerf-guided dis- tillation from 3D-aware diffusion.arXiv.cs, abs/2302.10109,
-
[24]
Query-key normalization for transformers
Alex Henry, Prudhvi Raj Dachapally, Shubham Pawar, and Yuxuan Chen. Query-key normalization for transformers. In Findings of EMNLP, 2020. 5
2020
-
[25]
Denoising diffu- sion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. InProc. NeurIPS, 2020. 2
2020
-
[26]
Unifying corre- spondence, pose and nerf for generalized pose-free novel view synthesis.Proc
Sunghwan Hong, Jaewoo Jung, Heeseong Shin, Jiaolong Yang, Seungryong Kim, and Chong Luo. Unifying corre- spondence, pose and nerf for generalized pose-free novel view synthesis.Proc. CVPR, 2024. 3
2024
-
[27]
sim- ple diffusion: End-to-end diffusion for high resolution im- ages
Emiel Hoogeboom, Jonathan Heek, and Tim Salimans. sim- ple diffusion: End-to-end diffusion for high resolution im- ages. InProc. ICML, 2023. 14
2023
-
[28]
Generative camera dolly: Ex- treme monocular dynamic novel view synthesis
Basile Van Hoorick, Rundi Wu, Ege Ozguroglu, Kyle Sar- gent, Ruoshi Liu, Pavel Tokmakov, Achal Dave, Changxi Zheng, and Carl V ondrick. Generative camera dolly: Ex- treme monocular dynamic novel view synthesis. InProc. ECCV, 2024. 3
2024
-
[29]
DeepMVS: Learning multi-view stereopsis
Po-Han Huang, Kevin Matzen, Johannes Kopf, Narendra Ahuja, and Jia-Bin Huang. DeepMVS: Learning multi-view stereopsis. InProc. CVPR, 2018. 17
2018
-
[30]
LVT: large- scale scene reconstruction via local view transformers
Tooba Imtiaz, Lucy Chai, Kathryn Heal, Xuan Luo, Jungyeon Park, Jennifer Dy, and John Flynn. LVT: large- scale scene reconstruction via local view transformers. In Proc. SIGGRAPH Asia, 2025. 1
2025
-
[31]
Stable virtual camera: Generative view synthesis with diffusion models
Jensen, Zhou, Hang Gao, Vikram V oleti, Aaryaman Vasishta, Chun-Han Yao, Mark Boss, Philip Torr, Christian Rupprecht, and Varun Jampani. Stable virtual camera: Generative view synthesis with diffusion models. InProc. ICCV, 2025. 3
2025
-
[32]
RayZer: a self-supervised large view synthesis model
Hanwen Jiang, Hao Tan, Peng Wang, Haian Jin, Yue Zhao, Sai Bi, Kai Zhang, Fujun Luan, Kalyan Sunkavalli, Qixing Huang, and Georgios Pavlakos. RayZer: a self-supervised large view synthesis model. InProc. ICCV, 2025. 1, 2, 3, 4, 13, 14
2025
-
[33]
AnySplat: feed-forward 3D Gaussian Splatting from unconstrained views.ACM Trans
Lihan Jiang, Yucheng Mao, Linning Xu, Tao Lu, Kerui Ren, Yichen Jin, Xudong Xu, Mulin Yu, Jiangmiao Pang, Feng Zhao, Dahua Lin, and Bo Dai. AnySplat: feed-forward 3D Gaussian Splatting from unconstrained views.ACM Trans. on Graphics (TOG), 44(6):1–16, 2025. 1, 2, 3, 4, 7, 8
2025
-
[34]
LVSM: a large view synthesis model with minimal 3D inductive bias
Haian Jin, Hanwen Jiang, Hao Tan, Kai Zhang, Sai Bi, Tianyuan Zhang, Fujun Luan, Noah Snavely, and Zexiang Xu. LVSM: a large view synthesis model with minimal 3D inductive bias. InProc. ICLR, 2025. 1, 2, 3, 4, 6, 7, 15
2025
-
[35]
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In Proc. ECCV, 2016. 5, 14
2016
-
[36]
3D Gaussian Splatting for real-time radiance field rendering.Proc
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3D Gaussian Splatting for real-time radiance field rendering.Proc. SIGGRAPH, 42(4), 2023. 1, 2
2023
-
[37]
3D Gaussian Splatting for Real-Time Radiance Field Rendering.ACM Trans
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3D Gaussian Splatting for Real-Time Radiance Field Rendering.ACM Trans. on Graphics (TOG),
-
[38]
Mitchel, and Vincent Sitzmann
Evan Kim, Hyunwoo Ryu, Thomas W. Mitchel, and Vincent Sitzmann. Scaling view synthesis transformers. InProc. CVPR, 2026. 3, 13
2026
-
[39]
EDEN: Multimodal Synthetic Dataset of Enclosed garDEN Scenes
Hoang-An Le, Partha Das, Thomas Mensink, Sezer Karaoglu, and Theo Gevers. EDEN: Multimodal Synthetic Dataset of Enclosed garDEN Scenes. InProc. WACV, 2021. 17
2021
-
[40]
Ground- ing Image Matching in 3D with MAST3R
Vincent Leroy, Yohann Cabon, and J´erˆome Revaud. Ground- ing Image Matching in 3D with MAST3R. InProc. ECCV,
-
[41]
Pla- taniotis, Sergey Tulyakov, and Jian Ren
Hanwen Liang, Junli Cao, Vidit Goel, Guocheng Qian, Sergei Korolev, Demetri Terzopoulos, Konstantinos N. Pla- taniotis, Sergey Tulyakov, and Jian Ren. Wonderland: Navi- gating 3D scenes from a single image. InProc. CVPR, 2025. 3
2025
-
[42]
Vision transformer for nerf-based view synthesis from a single input image
Kai-En Lin, Yen-Chen Lin, Wei-Sheng Lai, Tsung-Yi Lin, Yi-Chang Shih, and Ravi Ramamoorthi. Vision transformer for nerf-based view synthesis from a single input image. In Proc. WACV, 2023. 2
2023
-
[43]
Common Diffusion Noise Schedules and Sample Steps are Flawed
Shanchuan Lin, Bingchen Liu, Jiashi Li, and Xiao Yang. Common Diffusion Noise Schedules and Sample Steps are Flawed. InProc. WACV, 2024. 14
2024
-
[44]
DL3DV-10K: a large-scale scene dataset for deep learning-based 3d vision
Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, Xuanmao Li, Xingpeng Sun, Rohan Ashok, Aniruddha Mukherjee, Hao Kang, Xiangrui Kong, Gang Hua, Tianyi Zhang, Bedrich Benes, and Aniket Bera. DL3DV-10K: a large-scale sce...
2024
-
[45]
Re- conX: reconstruct any scene from sparse views with video diffusion model.arXiv, 2408.16767, 2024
Fangfu Liu, Wenqiang Sun, Hanyang Wang, Yikai Wang, Haowen Sun, Junliang Ye, Jun Zhang, and Yueqi Duan. Re- conX: reconstruct any scene from sparse views with video diffusion model.arXiv, 2408.16767, 2024. 3
2024 arXiv
-
[46]
Zero-1-to-3: Zero-shot one image to 3D object
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tok- makov, Sergey Zakharov, and Carl V ondrick. Zero-1-to-3: Zero-shot one image to 3D object. InProc. ICCV, 2023. 3
2023
-
[47]
Zhang, Natalia Neverova, Andrea Vedaldi, Roman Shapovalov, and David Novotny
Xingchen Liu, Piyush Tayal, Jianyuan Wang, Jesus Zarzar, Tom Monnier, Konstantinos Tertikas, Jiali Duan, Antoine Toisoul, Jason Y . Zhang, Natalia Neverova, Andrea Vedaldi, Roman Shapovalov, and David Novotny. uCO3D uncommon objects in 3D. InProc. CVPR, 2025. 17
2025
-
[48]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view syn- thesis. InProc. ECCV, 2020. 1, 2
2020
-
[49]
Maxime Oquab, Timoth ´ee Darcet, Th´eo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael ...
2024
-
[50]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. InProc. ICCV, 2023. 8, 14
2023
-
[51]
Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction
Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotny. Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. InProc. ICCV, 2021. 5, 8
2021
-
[52]
GEN3C: 3D-informed world-consistent video generation with precise camera con- trol
Xuanchi Ren, Tianchang Shen, Jiahui Huang, Huan Ling, Yifan Lu, Merlin Nimier-David, Thomas M ¨uller, Alexander Keller, Sanja Fidler, and Jun Gao. GEN3C: 3D-informed world-consistent video generation with precise camera con- trol. InProc. CVPR, 2025. 3
2025
-
[53]
Susskind
Mike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M. Susskind. Hypersim: A photorealistic syn- thetic dataset for holistic indoor scene understanding. In Proc. ICCV, 2021. 17
2021
-
[54]
High-resolution image syn- thesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- thesis with latent diffusion models. InProc. CVPR, 2022. 8, 16
2022
-
[55]
Mehdi S. M. Sajjadi, Henning Meyer, Etienne Pot, Urs Bergmann, Klaus Greff, Noha Radwan, Suhani V ora, Mario Lucic, Daniel Duckworth, Alexey Dosovitskiy, Jakob Uszkoreit, Thomas A. Funkhouser, and Andrea Tagliasacchi. Scene representation transformer: Geometry-free novel view ...
2021 arXiv
-
[56]
Mehdi S. M. Sajjadi, Henning Meyer, Etienne Pot, Urs Bergmann, Klaus Greff, Noha Radwan, Suhani V ora, Mario Lucic, Daniel Duckworth, Alexey Dosovitskiy, Jakob Uszkoreit, Thomas Funkhouser, and Andrea Tagliasacchi. Scene representation transformer: Geometry-free novel view syn...
2022
-
[57]
ZeroNVS: Zero-shot 360-degree view synthesis from a single real im- age.arXiv.cs, abs/2310.17994, 2023
Kyle Sargent, Zizhang Li, Tanmay Shah, Charles Herrmann, Hong-Xing Yu, Yunzhi Zhang, Eric Ryan Chan, Dmitry La- gun, Li Fei-Fei, Deqing Sun, and Jiajun Wu. ZeroNVS: Zero-shot 360-degree view synthesis from a single real im- age.arXiv.cs, abs/2310.17994, 2023. 3
2023 arXiv
-
[58]
Structure-from-motion revisited
Johannes Lutz Sch ¨onberger and Jan-Michael Frahm. Structure-from-motion revisited. InProc. CVPR, 2016. 17
2016
-
[59]
A benchmark and a baseline for robust multi- view depth estimation
Philipp Schr ¨oppel, Jan Bechtold, Artemij Amiranashvili, and Thomas Brox. A benchmark and a baseline for robust multi- view depth estimation. InProc. 3DV, 2022. 17
2022
-
[60]
FlashAttention-3: Fast and accurate attention with asynchrony and low-precision
Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao. FlashAttention-3: Fast and accurate attention with asynchrony and low-precision. In Proc. NeurIPS, 2024. 5
2024
-
[61]
Light field networks: Neu- ral scene representations with single-evaluation rendering
Vincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum, and Fr´edo Durand. Light field networks: Neu- ral scene representations with single-evaluation rendering. In Proc. NeurIPS, 2021. 3
2021
-
[62]
Splatt3R: Zero-shot gaussian splatting from uncalibrated image pairs.arXiv, 2408.13912, 2024
Brandon Smart, Chuanxia Zheng, Iro Laina, and Vic- tor Adrian Prisacariu. Splatt3R: Zero-shot gaussian splatting from uncalibrated image pairs.arXiv, 2408.13912, 2024. 1, 3
2024 arXiv
-
[63]
Denois- ing diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois- ing diffusion implicit models. InProc. ICLR, 2021. 14
2021
-
[64]
Highway networks
Rupesh Kumar Srivastava, Klaus Greff, and J ¨urgen Schmid- huber. Highway networks. InProc. ICML Workshops, 2015. 4
2015
-
[65]
Splatter Image: Ultra-fast single-view 3D recon- struction
Stanislaw Szymanowicz, Christian Rupprecht, and Andrea Vedaldi. Splatter Image: Ultra-fast single-view 3D recon- struction. InProceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), 2024. 1, 3, 8
2024
-
[66]
Henriques, Christian Rup- precht, and Andrea Vedaldi
Stanislaw Szymanowicz, Eldar Insafutdinov, Chuanxia Zheng, Dylan Campbell, Jo ˜ao F. Henriques, Christian Rup- precht, and Andrea Vedaldi. Flash3D: Feed-forward gen- eralisable 3D scene reconstruction from a single image. In Proceedings of the International Conference on 3D Vi...
2025
-
[67]
Zhang, Pratul Srinivasan, Ruiqi Gao, Arthur Brussee, Aleksander Holynski, Ricardo Martin-Brualla, Jonathan T
Stanislaw Szymanowicz, Jason Y . Zhang, Pratul Srinivasan, Ruiqi Gao, Arthur Brussee, Aleksander Holynski, Ricardo Martin-Brualla, Jonathan T. Barron, and Philipp Henzler. Bolt3D: Generating 3D scenes in seconds. InProc. ICCV,
-
[68]
Camera pose and calibration from 4 or 5 known 3D points
Bill Triggs. Camera pose and calibration from 4 or 5 known 3D points. InProc. ICCV, 1999. 18
1999
-
[69]
Suhani V ora, Noha Radwan, Klaus Greff, Henning Meyer, Kyle Genova, Mehdi S. M. Sajjadi, Etienne Pot, Andrea Tagliasacchi, and Daniel Duckworth. NeSF: Neural semantic fields for generalizable semantic segmentation of 3D scenes. Trans. on Machine Learning Research, 2022. 17
2022
-
[70]
Vggt: Vi- sual geometry grounded transformer
Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. Vggt: Vi- sual geometry grounded transformer. InProc. CVPR, 2025. 2, 3, 4, 6, 15, 17
2025
-
[71]
DUSt3R: Geometric 3D vision made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. DUSt3R: Geometric 3D vision made easy. InProc. CVPR, 2024. 3, 6
2024
-
[72]
TartanAir: a dataset to push the limits of visual SLAM
Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor, and Se- bastian Scherer. TartanAir: a dataset to push the limits of visual SLAM. InProc. IROS, 2020. 17
2020
-
[73]
Yifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang, Yang Zhou, Zizun Li, Junyi Chen, Jiangmiao Pang, Chun- hua Shen, and Tong He.π 3: Permutation-equivariant visual geometry learning. InProc. ICLR, 2026. 3
2026
-
[74]
Bovik, Hamid R
Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image quality assessment: from error visibility to structural similarity.IEEE Trans. on Image Processing, 13(4), 2004. 5
2004
-
[75]
Novel view synthesis with diffusion models
Daniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho, Andrea Tagliasacchi, and Mohammad Norouzi. Novel view synthesis with diffusion models. In Proc. ICLR, 2023. 3
2023
-
[76]
Srinivasan, Dor Verbin, Jonathan T
Rundi Wu, Ben Mildenhall, Philipp Henzler, Keunhong Park, Ruiqi Gao, Daniel Watson, Pratul P. Srinivasan, Dor Verbin, Jonathan T. Barron, Ben Poole, and Aleksander Holynski. ReconFusion: 3D Reconstruction with Diffusion Priors. InProc. CVPR, 2024. 3, 8, 13
2024
-
[77]
Barron, and Aleksander Holynski
Rundi Wu, Ruiqi Gao, Ben Poole, Alex Trevithick, Changxi Zheng, Jonathan T. Barron, and Aleksander Holynski. CAT4D: create anything in 4D with multi-view video dif- fusion models. InProc. CVPR, 2025. 3
2025
-
[78]
RGBD objects in the wild: Scaling real-world 3D object learning from RGB-D videos
Hongchi Xia, Yang Fu, Sifei Liu, and Xiaolong Wang. RGBD objects in the wild: Scaling real-world 3D object learning from RGB-D videos. InProc. CVPR, 2024. 5, 17
2024
-
[79]
GaussianRoom: improving 3D Gaussian splatting with SDF guidance and monocular cues for indoor scene reconstruc- tion
Haodong Xiang, Xinghui Li, Kai Cheng, Xiansong Lai, Wanting Zhang, Zhichao Liao, Long Zeng, and Xueping Liu. GaussianRoom: improving 3D Gaussian splatting with SDF guidance and monocular cues for indoor scene reconstruc- tion. InProc. ICRA, 2025. 1
2025
-
[80]
DepthSplat: Connecting Gaussian Splatting and Depth
Haofei Xu, Songyou Peng, Fangjinhua Wang, Hermann Blum, Daniel Barath, Andreas Geiger, and Marc Pollefeys. DepthSplat: Connecting Gaussian Splatting and Depth. In Proc. CVPR, 2025. 3, 7, 13
2025
-
[81]
Blendedmvs: A large- scale dataset for generalized multi-view stereo networks
Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan. Blendedmvs: A large- scale dataset for generalized multi-view stereo networks. In Proc. CVPR, 2020. 17
2020
-
[82]
No Pose, No Prob- lem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images
Botao Ye, Sifei Liu, Haofei Xu, Li Xueting, Marc Pollefeys, Ming-Hsuan Yang, and Peng Songyou. No Pose, No Prob- lem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images. InProc. ICLR, 2025. 1, 3, 7, 8
2025
-
[83]
Yonosplat: You only need one model for feedfor- ward 3d gaussian splatting.Proc
Botao Ye, Boqi Chen, Haofei Xu, Daniel Barath, and Marc Pollefeys. Yonosplat: You only need one model for feedfor- ward 3d gaussian splatting.Proc. ICLR, 2026. 3
2026
-
[84]
PixelNeRF: Neural radiance fields from one or few images
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. PixelNeRF: Neural radiance fields from one or few images. InProc. CVPR, 2021. 2
2021
-
[85]
ViewCrafter: taming video diffusion models for high-fidelity novel view synthesis
Wangbo Yu, Jinbo Xing, Li Yuan, Wenbo Hu, Xiaoyu Li, Zhipeng Huang, Xiangjun Gao, Tien-Tsin Wong, Ying Shan, and Yonghong Tian. ViewCrafter: taming video diffusion models for high-fidelity novel view synthesis. pages 1–18,
-
[86]
Shen, Leonidas J
Amir Roshan Zamir, Alexander Sax, William B. Shen, Leonidas J. Guibas, Jitendra Malik, and Silvio Savarese. Taskonomy: Disentangling task transfer learning. InProc. CVPR, 2018. 17
2018
-
[87]
Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani
Jason Y . Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani. Cameras as rays: Pose estimation via ray diffusion. InProc. ICLR, 2024. 4, 16
2024
-
[88]
GS-LRM: Large Re- construction Model for 3D Gaussian Splatting
Kai Zhang, Sai Bi, Hao Tan, Yuanbo Xiangli, Nanxuan Zhao, Kalyan Sunkavalli, and Zexiang Xu. GS-LRM: Large Re- construction Model for 3D Gaussian Splatting. InProc. ECCV, 2024. 3, 6
2024
-
[89]
Efros, Eli Shecht- man, and Oliver Wang
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. InProc. CVPR, 2018. 5
2018
-
[90]
Flare: Feed-forward geometry, appearance and camera estimation from uncalibrated sparse views
Shangzhan Zhang, Jianyuan Wang, Yinghao Xu, Nan Xue, Christian Rupprecht, Xiaowei Zhou, Yujun Shen, and Gor- don Wetzstein. Flare: Feed-forward geometry, appearance and camera estimation from uncalibrated sparse views. In Proc. CVPR, 2025. 3, 7, 8, 13
2025
-
[91]
Stereo magnification: Learning view syn- thesis using multiplane images.ACM Trans
Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: Learning view syn- thesis using multiplane images.ACM Trans. on Graphics (TOG), 37(4):1–12, 2018. 2, 5, 6, 17
2018
-
[92]
camera token,
Chen Ziwen, Hao Tan, Kai Zhang, Sai Bi, Fujun Luan, Yi- cong Hong, Li Fuxin, and Zexiang Xu. Long-LRM: Long- sequence Large Reconstruction Model for Wide-coverage Gaussian Splats.Proc. ICCV, 2025. 3 LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis Supp...
2025
-
[93]
Upsize for VGGT Optional cameras V 1024× : V 11gi ×Linear SiLU Linear
-
[94]
full atten- tion
Project cameras Embedder Aggregator Add VGGT ‘camera token’ Local attention Global attention Repeat N times Linear Norm Concat. Output 3D repr. Figure A4.Encoder.The encoder takes as sourceVimagesand, optionally,Vcameras. The images are upsized (1) to the dimension expected by...
-
[95]
If only camera poses are available for a particular train- ing scene, thenw 2 = 0and we chooseλso that that λw1 := 1/1.35
-
[96]
• We choose aλ
If both camera poses and points are available, the we proceed as follows. • We choose aλ. To do so, with 50% probability, we chooseλso thatλw 1 = 1/1.35(following LVSM) and with 50% probability so thatλw 2 = 1. • We choose which, if any, scaling factor to drop out. With 1/3 pr...
-
[97]
the null token, keeping only scale parameterλw 1, so that the source cameras are not available to the model, but scale is unambiguous
Unposed” can also accept camera poses as conditioning, and we use it as our final model. the null token, keeping only scale parameterλw 1, so that the source cameras are not available to the model, but scale is unambiguous. The network is able to implicitly learn the meaning o...
Reviewed July 13, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.