Pith. sign in

REVIEW 4 minor 2 cited by

LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis

T0 review · 0 major / 4 minor · reviewed 2026-07-13 · grok-4.5

Pith's one-line read Even without building an explicit 3D model, novel-view synthesis still gains from 3D-aware features taken from a geometry-pretrained encoder, yielding real-time state-of-the-art feed-forward rendering.

desk verdict Solid systems paper: VGGT-initialized highway encoder-decoder delivers real SOTA real-time feed-forward NVS with clean ablations; no load-bearing flaw. read the letter →

arxiv 2603.20176 v3 pith:M2AKLFVL submitted 2026-03-20 cs.CV

classification cs.CV
keywords novelviewsynthesisfeed-forwardNVSlatentgeometry3D-awarefeaturesencoder-decodertransformersreal-timerenderingcamera-freediffusiondecoder
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that skipping explicit 3D reconstruction does not mean discarding 3D inductive bias. It introduces LagerNVS, an encoder-decoder that initializes its encoder from a multi-view reconstruction network trained with geometric supervision, extracts intermediate latent tokens that carry that bias, and pairs them with a lightweight decoder trained end-to-end on photometric losses. A “highway” design keeps per-image tokens instead of forcing a fixed-size bottleneck, so information from the source views reaches the renderer intact. The result is deterministic feed-forward novel view synthesis that beats prior reconstruction-free and feed-forward 3D methods on standard benchmarks, runs in real time at 512 resolution, works with or without known cameras, generalizes to in-the-wild inputs, and can be lightly adapted into a diffusion decoder for generative fill-in. A reader who wants high-quality new views without slow per-scene optimization or fragile explicit geometry cares because the same network does the job fast, robustly, and with a single set of weights.

What carries the argument

LagerNVS: a highway encoder-decoder whose encoder is initialized from a geometry-pretrained multi-view transformer backbone, reading out concatenated last local- and global-attention tokens as 3D-aware latents; a lightweight transformer decoder then conditions on Plücker ray maps of the target camera and renders the novel view. The highway path avoids a fixed token bottleneck so capacity scales with the number of source images while encoding cost is amortized.

What would settle it

Train the same highway architecture from scratch or from generic 2D features at full scale and check whether the reported multi-decibel PSNR gap from 3D pre-training disappears on RealEstate10k and DL3DV; or freeze the encoder entirely and verify that reflections and textures remain unusable as the ablations claim.

Watch

Extended reading notes

Core claim

Reconstruction-free novel view synthesis still benefits strongly from 3D-aware latent features obtained by initializing the encoder from a network pre-trained with explicit 3D supervision. Combined with a highway encoder-decoder that preserves full source-image information flow and end-to-end photometric fine-tuning, this produces state-of-the-art deterministic feed-forward NVS (including 31.4 PSNR on RealEstate10k), real-time rendering, and generalization with or without source cameras.

Load-bearing premise

The intermediate tokens taken from a geometry-pretrained backbone still carry enough color, texture, and reflectance after fine-tuning to support high-fidelity image rendering; if they mostly encode shape and discard appearance, the photometric gains collapse.

Editorial extensions

If this is right

  • Reconstruction-free feed-forward NVS can outperform methods that still emit explicit 3D Gaussians or radiance fields.
  • Real-time 512×512 rendering on a single GPU becomes practical for roughly up to nine source views with a pure neural decoder.
  • One model trained on a large multi-dataset mix can handle posed and unposed, square and non-square, ego-centric and 360° inputs without retuning.
  • The same decoder can be fine-tuned with a denoising objective to sample plausible completions of occluded or extrapolated regions instead of averaging them away.
  • Geometry foundation models that never see a rendering loss during pre-training force expensive end-to-end fine-tuning if they are to be reused for appearance-critical tasks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Geometry backbones that jointly optimize a photometric rendering head during pre-training would likely transfer to NVS with far less fine-tuning than pure reconstruction models.
  • The highway-versus-bottleneck distinction is likely to matter for other multi-view tasks where token capacity, not just geometry, limits quality (correspondence, relighting, material estimation).
  • Pairing 3D-aware encoder initialization with compute-scaling recipes for encoder-decoder transformers could raise quality further at fixed training budget.
  • If appearance is systematically discarded by pure geometry pre-training, multi-task pre-training that keeps both geometry and photometry may become the default recipe for multi-view foundation models.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 4 minor

Summary. LagerNVS is a feed-forward encoder–decoder for novel view synthesis that avoids explicit 3D reconstruction. The encoder is initialized from VGGT (pre-trained with explicit 3D supervision) and produces per-image latent tokens; a lightweight ViT decoder, conditioned on a Plücker ray map of the target camera, renders the novel view. The authors compare highway vs. bottleneck encoder–decoder and decoder-only designs, train on a large multi-dataset mix, and report state-of-the-art deterministic NVS (31.39 PSNR on RealEstate10k under LVSM’s protocol), real-time decoding (30+ FPS at 512² with up to 9 views), operation with or without source cameras, and a preliminary diffusion decoder for generative extrapolation. Ablations isolate 3D pre-training, end-to-end fine-tuning, and architecture.

Significance. If the reported gains hold under the matched protocols, the work is a clear advance for reconstruction-free NVS. It shows that strong 3D inductive bias can be injected via pre-trained latent features rather than explicit geometry, while still delivering real-time rendering and strong generalization (including unposed and in-the-wild inputs). The highway encoder–decoder design, multi-dataset training, and public code/checkpoints make the result immediately usable and extensible. The diffusion fine-tuning experiment further indicates a practical path from deterministic to generative NVS without redesigning the encoder.

minor comments (4)
  1. Table 2 and the main Re10k numbers lack error bars or multi-seed statistics; a short note on run-to-run variance would strengthen confidence in the +1.7 dB margin.
  2. The shared-focal-length assumption (Sec. 3 and Limitations) is stated clearly but could be flagged earlier in the method section so readers know the scope of the camera model before the experiments.
  3. Fig. 7 and the occlusion examples (Fig. A3) are informative; adding a brief quantitative measure of failure modes (e.g., high-frequency texture or large baseline) would help readers gauge remaining limitations.
  4. Clarify the exact definition of the canonical FoV k0 used at test time when cameras are unknown (π/2 vs. 53.13°) in one place to avoid confusion between the main text and the v1/v2 appendix note.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: SOTA claims rest on external photometric metrics and held-out benchmarks after end-to-end fine-tuning, not on any quantity forced by definition or self-citation chain.

full rationale

LagerNVS is an empirical encoder-decoder NVS system. The encoder is initialized from VGGT (overlapping authors) and then fine-tuned end-to-end with L2 + perceptual losses on multi-view image tuples; the decoder is a standard ViT that maps Plücker ray tokens plus latent features to RGB. All reported numbers (31.4 PSNR on Re10k, gains vs LVSM/DepthSplat/AnySplat, real-time FPS, generalization) are measured by standard image metrics on held-out target views under fixed evaluation protocols. Ablations (Table 2) explicitly contrast frozen vs fine-tuned VGGT, 3D vs 2D vs no pre-training, and highway vs bottleneck vs decoder-only architectures; none of these comparisons reduce the target PSNR/SSIM/LPIPS to a fitted free parameter or to a self-cited uniqueness theorem. The single self-citation of VGGT supplies only an initialization that is subsequently optimized and evaluated independently; it is not load-bearing for the central claim. No equation equates a reported metric to an input by construction, no uniqueness result is imported to forbid alternatives, and no known empirical pattern is merely renamed. Score 1 reflects only the ordinary (non-circular) use of an overlapping-author backbone as a starting point.

Assumptions & free parameters 4 free parameters · 4 assumptions · 1 invented entities

The paper is an empirical systems contribution. It inherits standard photometric losses, transformer blocks and the VGGT architecture; the only free parameters are ordinary training hyper-parameters and the architectural choice of which VGGT layers to read. No new physical entities or ad-hoc conservation laws are introduced.

free parameters (4)
  • loss weights λ2, λp
    Relative weighting of L2 and perceptual losses; standard but chosen by the authors and not derived.
  • decoder depth / width (ViT-B, 12 blocks)
    Control variable for real-time budget; selected to keep decoding ≥30 FPS.
  • canonical horizontal FoV k0 = π/2 (or 53.13°)
    Nominal value used when source cameras are unknown to avoid focal-length leakage; chosen by design.
  • scene-scale normalizations w1, w2
    Two alternative scale factors (camera-baseline vs point-cloud) randomly dropped during training; necessary for unposed/single-view operation.
assumptions (4)
  • domain assumption Photometric L2 + perceptual loss is a sufficient training signal for high-quality NVS
    Standard in the NVS literature; invoked throughout Sec. 3.3.
  • domain assumption VGGT’s last local and global attention tokens contain transferable 3D-aware features useful for appearance rendering after fine-tuning
    Core inductive-bias claim of Sec. 3.1; supported by ablation but not proven a priori.
  • ad hoc to paper Source and target cameras share identical focal length (and equal horizontal/vertical FoV)
    Stated limitation in Sec. 3 and Limitations; training data never randomizes intrinsics.
  • domain assumption Scenes are approximately static
    Inherited from all multi-view training sets used.
invented entities (1)
  • highway encoder-decoder latent geometry tokens
    purpose: Camera-independent intermediate representation that preserves per-image features without a fixed-size bottleneck
    Architectural construct; no claim of a new physical quantity, only a design pattern.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis." pith.science (2026). https://pith.science/paper/M2AKLFVL

@misc{pith2026260320176,
  author       = {Pith},
  title        = {Pith review of: LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/M2AKLFVL}},
  note         = {Machine review of arXiv:2603.20176}
}
read the original abstract

Recent work has shown that neural networks can perform 3D tasks such as Novel View Synthesis (NVS) without explicit 3D reconstruction. Even so, we argue that strong 3D inductive biases are still helpful in the design of such networks. We show this point by introducing LagerNVS, an encoder-decoder neural network for NVS that builds on `3D-aware' latent features. The encoder is initialized from a 3D reconstruction network pre-trained using explicit 3D supervision. This is paired with a lightweight decoder, and trained end-to-end with photometric losses. LagerNVS achieves state-of-the-art deterministic feed-forward Novel View Synthesis (including 31.4 PSNR on Re10k), with and without known cameras, renders in real time, generalizes to in-the-wild data, and can be paired with a diffusion decoder for generative extrapolation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A single-image feed-forward Gaussian splatting method that samples supports from predicted depth and decodes Gaussian attributes implicitly, improving cross-dataset large-baseline novel view synthesis.

  2. RayViT: Ray-Conditioned Visual Representations for Viewpoint-Robust Imitation Learning

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Conditioning a pretrained ViT on per-pixel Plücker camera rays — via a gated-cross-attention class token and patch-level ray embeddings — makes imitation-learned manipulation policies substantially more robust to came...

Reference graph

Works this paper leans on

97 extracted references · 6 linked inside Pith · cited by 2 Pith papers

  1. [1]

    Adelson and R

    E. Adelson and R. Bergen.The Plenoptic Function and the Elements of Early Vision. MIT Press, 1991. 3

  2. [2]

    PyTorch 2: Faster machine learning through dynamic Python bytecode transformation and graph compilation

    Jason Ansel, Edward Yang, Horace He, et al. PyTorch 2: Faster machine learning through dynamic Python bytecode transformation and graph compilation. InProc. ACM Inter- national Conference on Architectural Support for Program- ming Languages and Operating Systems (ASPLOS), 2024. 2

  3. [3]

    Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization.arXiv.cs, abs/1607.06450, 2016. 4, 15

  4. [4]

    Lindell, Zan Gojcic, Sanja Fidler, Huan Ling, Jun Gao, and Xuanchi Ren

    Sherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang, Yifeng Jiang, Haithem Turki, Andrea Tagliasacchi, David B. Lindell, Zan Gojcic, Sanja Fidler, Huan Ling, Jun Gao, and Xuanchi Ren. Lyra: Generative 3D scene recon- struction via video diffusion model self-distillation.arXiv, 2509.19296, 2025. 3

  5. [5]

    ReCamMaster: camera-controlled gen- erative rendering from a single video

    Jianhong Bai, Menghan Xia, Xiao Fu, Xintao Wang, Lianrui Mu, Jinwen Cao, Zuozhu Liu, Haoji Hu, Xiang Bai, Pengfei Wan, and Di Zhang. ReCamMaster: camera-controlled gen- erative rendering from a single video. InProc. ICCV, 2025. 3, 16

  6. [6]

    ARK- itscenes - a diverse real-world dataset for 3d indoor scene understanding using mobile RGB-d data

    Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Tal Dimry, Yuri Feigin, Peter Fu, Thomas Gebauer, Brandon Joffe, Daniel Kurz, Arik Schwartz, and Elad Shulman. ARK- itscenes - a diverse real-world dataset for 3d indoor scene understanding using mobile RGB-d data. InProc. NeurIPS,

  7. [7]

    Gortler, and Michael F

    Chris Buehler, Michael Bosse, Leonard McMillan, Steven J. Gortler, and Michael F. Cohen. Unstructured lumigraph ren- dering. InProc. SIGGRAPH, 2001. 3

  8. [8]

    Chan, Koki Nagano, Matthew A

    Eric R. Chan, Koki Nagano, Matthew A. Chan, Alexan- der W. Bergman, Jeong Joon Park, Axel Levy, Miika Aittala, Shalini De Mello, Tero Karras, and Gordon Wetzstein. Gen- erative novel view synthesis with 3D-aware diffusion mod- els. InProc. ICCV, 2023. 3, 15

Show all 97 references
  1. [9]

    pixelSplat: 3D Gaussian splats from image pairs for scalable generalizable 3D reconstruction

    David Charatan, Sizhe Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelSplat: 3D Gaussian splats from image pairs for scalable generalizable 3D reconstruction. InProc. CVPR,

  2. [10]

    MVSNeRF: Fast generalizable radiance field reconstruction from multi-view stereo

    Anpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang, Fanbo Xiang, Jingyi Yu, and Hao Su. MVSNeRF: Fast generalizable radiance field reconstruction from multi-view stereo. InProc. ICCV, 2021. 2

  3. [11]

    MVSplat: efficient 3D gaussian splatting from sparse multi-view images

    Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. MVSplat: efficient 3D gaussian splatting from sparse multi-view images. InProc. ECCV, 2024. 1, 3, 6

  4. [12]

    MVS- plat360: Feed-forward 360 Scene Synthesis from Sparse Views

    Yuedong Chen, Chuanxia Zheng, Haofei Xu, Bohan Zhuang, Andrea Vedaldi, Tat-Jen Cham, and Jianfei Cai. MVS- plat360: Feed-forward 360 Scene Synthesis from Sparse Views. InProc. NeurIPS, 2024. 1, 3

  5. [13]

    FlashAttention-2: Faster attention with better par- allelism and work partitioning

    Tri Dao. FlashAttention-2: Faster attention with better par- allelism and work partitioning. InProc. ICLR, 2024. 5

  6. [14]

    Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

    Tri Dao, Daniel Y . Fu, Stefano Ermon, Atri Rudra, and Christopher R´e. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. InProc. NeurIPS, 2022. 5

  7. [15]

    Vision transformers need registers

    Timoth ´ee Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski. Vision transformers need registers. InProc. ICLR, 2024. 5

  8. [16]

    Jimmy Ba Diederik P. Kingma. Adam: A method for stochastic optimization. InProc. ICLR, 2015. 18

  9. [17]

    An image is worth 16×16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16×16 words: Transformers for image recognition ...

  10. [18]

    MASt3R-SfM: a fully-integrated solution for unconstrained structure-from-motion

    Bardienus Duisterhof, Lojze Zust, Philippe Weinzaepfel, Vincent Leroy, Yohann Cabon, and Jerome Revaud. MASt3R-SfM: a fully-integrated solution for unconstrained structure-from-motion. InProc. 3DV, 2025. 3

  11. [19]

    Sigmoid- weighted linear units for neural network function approxima- tion in reinforcement learning.Neural Networks, 107, 2018

    Stefan Elfwing, Eiji Uchibe, and Kenji Doya. Sigmoid- weighted linear units for neural network function approxima- tion in reinforcement learning.Neural Networks, 107, 2018. 15

  12. [20]

    FlowR: flowing from sparse to dense 3d re- constructions

    Tobias Fischer, Samuel Rota Bul `o, Yung-Hsu Yang, Nikhil Varma Keetha, Lorenzo Porzi, Norman M ¨uller, Katja Schwarz, Jonathon Luiten, Marc Pollefeys, and Peter Kontschieder. FlowR: flowing from sparse to dense 3d re- constructions. InProc. ICCV, 2025. 3

  13. [21]

    Barron, and Ben Poole

    Ruiqi Gao, Aleksander Holynski, Philipp Henzler, Arthur Brussee, Ricardo Martin-Brualla, Pratul Srinivasan, Jonathan T. Barron, and Ben Poole. CAT3D: create anything in 3d with multi-view diffusion models. InProc. NeurIPS,

  14. [22]

    Gortler, Radek Grzeszczuk, Richard Szeliski, and Michael F

    Steven J. Gortler, Radek Grzeszczuk, Richard Szeliski, and Michael F. Cohen. The lumigraph. InProc. SIGGRAPH,

  15. [23]

    Susskind, Christian Theobalt, Lingjie Liu, and Ravi Ramamoorthi

    Jiatao Gu, Alex Trevithick, Kai-En Lin, Josh M. Susskind, Christian Theobalt, Lingjie Liu, and Ravi Ramamoorthi. NerfDiff: Single-image view synthesis with nerf-guided dis- tillation from 3D-aware diffusion.arXiv.cs, abs/2302.10109,

  16. [24]

    Query-key normalization for transformers

    Alex Henry, Prudhvi Raj Dachapally, Shubham Pawar, and Yuxuan Chen. Query-key normalization for transformers. In Findings of EMNLP, 2020. 5

  17. [25]

    Denoising diffu- sion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. InProc. NeurIPS, 2020. 2

  18. [26]

    Unifying corre- spondence, pose and nerf for generalized pose-free novel view synthesis.Proc

    Sunghwan Hong, Jaewoo Jung, Heeseong Shin, Jiaolong Yang, Seungryong Kim, and Chong Luo. Unifying corre- spondence, pose and nerf for generalized pose-free novel view synthesis.Proc. CVPR, 2024. 3

  19. [27]

    sim- ple diffusion: End-to-end diffusion for high resolution im- ages

    Emiel Hoogeboom, Jonathan Heek, and Tim Salimans. sim- ple diffusion: End-to-end diffusion for high resolution im- ages. InProc. ICML, 2023. 14

  20. [28]

    Generative camera dolly: Ex- treme monocular dynamic novel view synthesis

    Basile Van Hoorick, Rundi Wu, Ege Ozguroglu, Kyle Sar- gent, Ruoshi Liu, Pavel Tokmakov, Achal Dave, Changxi Zheng, and Carl V ondrick. Generative camera dolly: Ex- treme monocular dynamic novel view synthesis. InProc. ECCV, 2024. 3

  21. [29]

    DeepMVS: Learning multi-view stereopsis

    Po-Han Huang, Kevin Matzen, Johannes Kopf, Narendra Ahuja, and Jia-Bin Huang. DeepMVS: Learning multi-view stereopsis. InProc. CVPR, 2018. 17

  22. [30]

    LVT: large- scale scene reconstruction via local view transformers

    Tooba Imtiaz, Lucy Chai, Kathryn Heal, Xuan Luo, Jungyeon Park, Jennifer Dy, and John Flynn. LVT: large- scale scene reconstruction via local view transformers. In Proc. SIGGRAPH Asia, 2025. 1

  23. [31]

    Stable virtual camera: Generative view synthesis with diffusion models

    Jensen, Zhou, Hang Gao, Vikram V oleti, Aaryaman Vasishta, Chun-Han Yao, Mark Boss, Philip Torr, Christian Rupprecht, and Varun Jampani. Stable virtual camera: Generative view synthesis with diffusion models. InProc. ICCV, 2025. 3

  24. [32]

    RayZer: a self-supervised large view synthesis model

    Hanwen Jiang, Hao Tan, Peng Wang, Haian Jin, Yue Zhao, Sai Bi, Kai Zhang, Fujun Luan, Kalyan Sunkavalli, Qixing Huang, and Georgios Pavlakos. RayZer: a self-supervised large view synthesis model. InProc. ICCV, 2025. 1, 2, 3, 4, 13, 14

  25. [33]

    AnySplat: feed-forward 3D Gaussian Splatting from unconstrained views.ACM Trans

    Lihan Jiang, Yucheng Mao, Linning Xu, Tao Lu, Kerui Ren, Yichen Jin, Xudong Xu, Mulin Yu, Jiangmiao Pang, Feng Zhao, Dahua Lin, and Bo Dai. AnySplat: feed-forward 3D Gaussian Splatting from unconstrained views.ACM Trans. on Graphics (TOG), 44(6):1–16, 2025. 1, 2, 3, 4, 7, 8

  26. [34]

    LVSM: a large view synthesis model with minimal 3D inductive bias

    Haian Jin, Hanwen Jiang, Hao Tan, Kai Zhang, Sai Bi, Tianyuan Zhang, Fujun Luan, Noah Snavely, and Zexiang Xu. LVSM: a large view synthesis model with minimal 3D inductive bias. InProc. ICLR, 2025. 1, 2, 3, 4, 6, 7, 15

  27. [35]

    Perceptual losses for real-time style transfer and super-resolution

    Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In Proc. ECCV, 2016. 5, 14

  28. [36]

    3D Gaussian Splatting for real-time radiance field rendering.Proc

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3D Gaussian Splatting for real-time radiance field rendering.Proc. SIGGRAPH, 42(4), 2023. 1, 2

  29. [37]

    3D Gaussian Splatting for Real-Time Radiance Field Rendering.ACM Trans

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3D Gaussian Splatting for Real-Time Radiance Field Rendering.ACM Trans. on Graphics (TOG),

  30. [38]

    Mitchel, and Vincent Sitzmann

    Evan Kim, Hyunwoo Ryu, Thomas W. Mitchel, and Vincent Sitzmann. Scaling view synthesis transformers. InProc. CVPR, 2026. 3, 13

  31. [39]

    EDEN: Multimodal Synthetic Dataset of Enclosed garDEN Scenes

    Hoang-An Le, Partha Das, Thomas Mensink, Sezer Karaoglu, and Theo Gevers. EDEN: Multimodal Synthetic Dataset of Enclosed garDEN Scenes. InProc. WACV, 2021. 17

  32. [40]

    Ground- ing Image Matching in 3D with MAST3R

    Vincent Leroy, Yohann Cabon, and J´erˆome Revaud. Ground- ing Image Matching in 3D with MAST3R. InProc. ECCV,

  33. [41]

    Pla- taniotis, Sergey Tulyakov, and Jian Ren

    Hanwen Liang, Junli Cao, Vidit Goel, Guocheng Qian, Sergei Korolev, Demetri Terzopoulos, Konstantinos N. Pla- taniotis, Sergey Tulyakov, and Jian Ren. Wonderland: Navi- gating 3D scenes from a single image. InProc. CVPR, 2025. 3

  34. [42]

    Vision transformer for nerf-based view synthesis from a single input image

    Kai-En Lin, Yen-Chen Lin, Wei-Sheng Lai, Tsung-Yi Lin, Yi-Chang Shih, and Ravi Ramamoorthi. Vision transformer for nerf-based view synthesis from a single input image. In Proc. WACV, 2023. 2

  35. [43]

    Common Diffusion Noise Schedules and Sample Steps are Flawed

    Shanchuan Lin, Bingchen Liu, Jiashi Li, and Xiao Yang. Common Diffusion Noise Schedules and Sample Steps are Flawed. InProc. WACV, 2024. 14

  36. [44]

    DL3DV-10K: a large-scale scene dataset for deep learning-based 3d vision

    Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, Xuanmao Li, Xingpeng Sun, Rohan Ashok, Aniruddha Mukherjee, Hao Kang, Xiangrui Kong, Gang Hua, Tianyi Zhang, Bedrich Benes, and Aniket Bera. DL3DV-10K: a large-scale sce...

  37. [45]

    Re- conX: reconstruct any scene from sparse views with video diffusion model.arXiv, 2408.16767, 2024

    Fangfu Liu, Wenqiang Sun, Hanyang Wang, Yikai Wang, Haowen Sun, Junliang Ye, Jun Zhang, and Yueqi Duan. Re- conX: reconstruct any scene from sparse views with video diffusion model.arXiv, 2408.16767, 2024. 3

  38. [46]

    Zero-1-to-3: Zero-shot one image to 3D object

    Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tok- makov, Sergey Zakharov, and Carl V ondrick. Zero-1-to-3: Zero-shot one image to 3D object. InProc. ICCV, 2023. 3

  39. [47]

    Zhang, Natalia Neverova, Andrea Vedaldi, Roman Shapovalov, and David Novotny

    Xingchen Liu, Piyush Tayal, Jianyuan Wang, Jesus Zarzar, Tom Monnier, Konstantinos Tertikas, Jiali Duan, Antoine Toisoul, Jason Y . Zhang, Natalia Neverova, Andrea Vedaldi, Roman Shapovalov, and David Novotny. uCO3D uncommon objects in 3D. InProc. CVPR, 2025. 17

  40. [48]

    Srinivasan, Matthew Tancik, Jonathan T

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view syn- thesis. InProc. ECCV, 2020. 1, 2

  41. [49]

    Maxime Oquab, Timoth ´ee Darcet, Th´eo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael ...

  42. [50]

    Scalable diffusion models with transformers

    William Peebles and Saining Xie. Scalable diffusion models with transformers. InProc. ICCV, 2023. 8, 14

  43. [51]

    Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction

    Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotny. Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. InProc. ICCV, 2021. 5, 8

  44. [52]

    GEN3C: 3D-informed world-consistent video generation with precise camera con- trol

    Xuanchi Ren, Tianchang Shen, Jiahui Huang, Huan Ling, Yifan Lu, Merlin Nimier-David, Thomas M ¨uller, Alexander Keller, Sanja Fidler, and Jun Gao. GEN3C: 3D-informed world-consistent video generation with precise camera con- trol. InProc. CVPR, 2025. 3

  45. [53]

    Susskind

    Mike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M. Susskind. Hypersim: A photorealistic syn- thetic dataset for holistic indoor scene understanding. In Proc. ICCV, 2021. 17

  46. [54]

    High-resolution image syn- thesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- thesis with latent diffusion models. InProc. CVPR, 2022. 8, 16

  47. [55]

    Mehdi S. M. Sajjadi, Henning Meyer, Etienne Pot, Urs Bergmann, Klaus Greff, Noha Radwan, Suhani V ora, Mario Lucic, Daniel Duckworth, Alexey Dosovitskiy, Jakob Uszkoreit, Thomas A. Funkhouser, and Andrea Tagliasacchi. Scene representation transformer: Geometry-free novel view ...

  48. [56]

    Mehdi S. M. Sajjadi, Henning Meyer, Etienne Pot, Urs Bergmann, Klaus Greff, Noha Radwan, Suhani V ora, Mario Lucic, Daniel Duckworth, Alexey Dosovitskiy, Jakob Uszkoreit, Thomas Funkhouser, and Andrea Tagliasacchi. Scene representation transformer: Geometry-free novel view syn...

  49. [57]

    ZeroNVS: Zero-shot 360-degree view synthesis from a single real im- age.arXiv.cs, abs/2310.17994, 2023

    Kyle Sargent, Zizhang Li, Tanmay Shah, Charles Herrmann, Hong-Xing Yu, Yunzhi Zhang, Eric Ryan Chan, Dmitry La- gun, Li Fei-Fei, Deqing Sun, and Jiajun Wu. ZeroNVS: Zero-shot 360-degree view synthesis from a single real im- age.arXiv.cs, abs/2310.17994, 2023. 3

  50. [58]

    Structure-from-motion revisited

    Johannes Lutz Sch ¨onberger and Jan-Michael Frahm. Structure-from-motion revisited. InProc. CVPR, 2016. 17

  51. [59]

    A benchmark and a baseline for robust multi- view depth estimation

    Philipp Schr ¨oppel, Jan Bechtold, Artemij Amiranashvili, and Thomas Brox. A benchmark and a baseline for robust multi- view depth estimation. InProc. 3DV, 2022. 17

  52. [60]

    FlashAttention-3: Fast and accurate attention with asynchrony and low-precision

    Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao. FlashAttention-3: Fast and accurate attention with asynchrony and low-precision. In Proc. NeurIPS, 2024. 5

  53. [61]

    Light field networks: Neu- ral scene representations with single-evaluation rendering

    Vincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum, and Fr´edo Durand. Light field networks: Neu- ral scene representations with single-evaluation rendering. In Proc. NeurIPS, 2021. 3

  54. [62]

    Splatt3R: Zero-shot gaussian splatting from uncalibrated image pairs.arXiv, 2408.13912, 2024

    Brandon Smart, Chuanxia Zheng, Iro Laina, and Vic- tor Adrian Prisacariu. Splatt3R: Zero-shot gaussian splatting from uncalibrated image pairs.arXiv, 2408.13912, 2024. 1, 3

  55. [63]

    Denois- ing diffusion implicit models

    Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois- ing diffusion implicit models. InProc. ICLR, 2021. 14

  56. [64]

    Highway networks

    Rupesh Kumar Srivastava, Klaus Greff, and J ¨urgen Schmid- huber. Highway networks. InProc. ICML Workshops, 2015. 4

  57. [65]

    Splatter Image: Ultra-fast single-view 3D recon- struction

    Stanislaw Szymanowicz, Christian Rupprecht, and Andrea Vedaldi. Splatter Image: Ultra-fast single-view 3D recon- struction. InProceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), 2024. 1, 3, 8

  58. [66]

    Henriques, Christian Rup- precht, and Andrea Vedaldi

    Stanislaw Szymanowicz, Eldar Insafutdinov, Chuanxia Zheng, Dylan Campbell, Jo ˜ao F. Henriques, Christian Rup- precht, and Andrea Vedaldi. Flash3D: Feed-forward gen- eralisable 3D scene reconstruction from a single image. In Proceedings of the International Conference on 3D Vi...

  59. [67]

    Zhang, Pratul Srinivasan, Ruiqi Gao, Arthur Brussee, Aleksander Holynski, Ricardo Martin-Brualla, Jonathan T

    Stanislaw Szymanowicz, Jason Y . Zhang, Pratul Srinivasan, Ruiqi Gao, Arthur Brussee, Aleksander Holynski, Ricardo Martin-Brualla, Jonathan T. Barron, and Philipp Henzler. Bolt3D: Generating 3D scenes in seconds. InProc. ICCV,

  60. [68]

    Camera pose and calibration from 4 or 5 known 3D points

    Bill Triggs. Camera pose and calibration from 4 or 5 known 3D points. InProc. ICCV, 1999. 18

  61. [69]

    Suhani V ora, Noha Radwan, Klaus Greff, Henning Meyer, Kyle Genova, Mehdi S. M. Sajjadi, Etienne Pot, Andrea Tagliasacchi, and Daniel Duckworth. NeSF: Neural semantic fields for generalizable semantic segmentation of 3D scenes. Trans. on Machine Learning Research, 2022. 17

  62. [70]

    Vggt: Vi- sual geometry grounded transformer

    Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. Vggt: Vi- sual geometry grounded transformer. InProc. CVPR, 2025. 2, 3, 4, 6, 15, 17

  63. [71]

    DUSt3R: Geometric 3D vision made easy

    Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. DUSt3R: Geometric 3D vision made easy. InProc. CVPR, 2024. 3, 6

  64. [72]

    TartanAir: a dataset to push the limits of visual SLAM

    Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor, and Se- bastian Scherer. TartanAir: a dataset to push the limits of visual SLAM. InProc. IROS, 2020. 17

  65. [73]

    Yifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang, Yang Zhou, Zizun Li, Junyi Chen, Jiangmiao Pang, Chun- hua Shen, and Tong He.π 3: Permutation-equivariant visual geometry learning. InProc. ICLR, 2026. 3

  66. [74]

    Bovik, Hamid R

    Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image quality assessment: from error visibility to structural similarity.IEEE Trans. on Image Processing, 13(4), 2004. 5

  67. [75]

    Novel view synthesis with diffusion models

    Daniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho, Andrea Tagliasacchi, and Mohammad Norouzi. Novel view synthesis with diffusion models. In Proc. ICLR, 2023. 3

  68. [76]

    Srinivasan, Dor Verbin, Jonathan T

    Rundi Wu, Ben Mildenhall, Philipp Henzler, Keunhong Park, Ruiqi Gao, Daniel Watson, Pratul P. Srinivasan, Dor Verbin, Jonathan T. Barron, Ben Poole, and Aleksander Holynski. ReconFusion: 3D Reconstruction with Diffusion Priors. InProc. CVPR, 2024. 3, 8, 13

  69. [77]

    Barron, and Aleksander Holynski

    Rundi Wu, Ruiqi Gao, Ben Poole, Alex Trevithick, Changxi Zheng, Jonathan T. Barron, and Aleksander Holynski. CAT4D: create anything in 4D with multi-view video dif- fusion models. InProc. CVPR, 2025. 3

  70. [78]

    RGBD objects in the wild: Scaling real-world 3D object learning from RGB-D videos

    Hongchi Xia, Yang Fu, Sifei Liu, and Xiaolong Wang. RGBD objects in the wild: Scaling real-world 3D object learning from RGB-D videos. InProc. CVPR, 2024. 5, 17

  71. [79]

    GaussianRoom: improving 3D Gaussian splatting with SDF guidance and monocular cues for indoor scene reconstruc- tion

    Haodong Xiang, Xinghui Li, Kai Cheng, Xiansong Lai, Wanting Zhang, Zhichao Liao, Long Zeng, and Xueping Liu. GaussianRoom: improving 3D Gaussian splatting with SDF guidance and monocular cues for indoor scene reconstruc- tion. InProc. ICRA, 2025. 1

  72. [80]

    DepthSplat: Connecting Gaussian Splatting and Depth

    Haofei Xu, Songyou Peng, Fangjinhua Wang, Hermann Blum, Daniel Barath, Andreas Geiger, and Marc Pollefeys. DepthSplat: Connecting Gaussian Splatting and Depth. In Proc. CVPR, 2025. 3, 7, 13

  73. [81]

    Blendedmvs: A large- scale dataset for generalized multi-view stereo networks

    Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan. Blendedmvs: A large- scale dataset for generalized multi-view stereo networks. In Proc. CVPR, 2020. 17

  74. [82]

    No Pose, No Prob- lem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images

    Botao Ye, Sifei Liu, Haofei Xu, Li Xueting, Marc Pollefeys, Ming-Hsuan Yang, and Peng Songyou. No Pose, No Prob- lem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images. InProc. ICLR, 2025. 1, 3, 7, 8

  75. [83]

    Yonosplat: You only need one model for feedfor- ward 3d gaussian splatting.Proc

    Botao Ye, Boqi Chen, Haofei Xu, Daniel Barath, and Marc Pollefeys. Yonosplat: You only need one model for feedfor- ward 3d gaussian splatting.Proc. ICLR, 2026. 3

  76. [84]

    PixelNeRF: Neural radiance fields from one or few images

    Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. PixelNeRF: Neural radiance fields from one or few images. InProc. CVPR, 2021. 2

  77. [85]

    ViewCrafter: taming video diffusion models for high-fidelity novel view synthesis

    Wangbo Yu, Jinbo Xing, Li Yuan, Wenbo Hu, Xiaoyu Li, Zhipeng Huang, Xiangjun Gao, Tien-Tsin Wong, Ying Shan, and Yonghong Tian. ViewCrafter: taming video diffusion models for high-fidelity novel view synthesis. pages 1–18,

  78. [86]

    Shen, Leonidas J

    Amir Roshan Zamir, Alexander Sax, William B. Shen, Leonidas J. Guibas, Jitendra Malik, and Silvio Savarese. Taskonomy: Disentangling task transfer learning. InProc. CVPR, 2018. 17

  79. [87]

    Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani

    Jason Y . Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani. Cameras as rays: Pose estimation via ray diffusion. InProc. ICLR, 2024. 4, 16

  80. [88]

    GS-LRM: Large Re- construction Model for 3D Gaussian Splatting

    Kai Zhang, Sai Bi, Hao Tan, Yuanbo Xiangli, Nanxuan Zhao, Kalyan Sunkavalli, and Zexiang Xu. GS-LRM: Large Re- construction Model for 3D Gaussian Splatting. InProc. ECCV, 2024. 3, 6

  81. [89]

    Efros, Eli Shecht- man, and Oliver Wang

    Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. InProc. CVPR, 2018. 5

  82. [90]

    Flare: Feed-forward geometry, appearance and camera estimation from uncalibrated sparse views

    Shangzhan Zhang, Jianyuan Wang, Yinghao Xu, Nan Xue, Christian Rupprecht, Xiaowei Zhou, Yujun Shen, and Gor- don Wetzstein. Flare: Feed-forward geometry, appearance and camera estimation from uncalibrated sparse views. In Proc. CVPR, 2025. 3, 7, 8, 13

  83. [91]

    Stereo magnification: Learning view syn- thesis using multiplane images.ACM Trans

    Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: Learning view syn- thesis using multiplane images.ACM Trans. on Graphics (TOG), 37(4):1–12, 2018. 2, 5, 6, 17

  84. [92]

    camera token,

    Chen Ziwen, Hao Tan, Kai Zhang, Sai Bi, Fujun Luan, Yi- cong Hong, Li Fuxin, and Zexiang Xu. Long-LRM: Long- sequence Large Reconstruction Model for Wide-coverage Gaussian Splats.Proc. ICCV, 2025. 3 LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis Supp...

  85. [93]

    Upsize for VGGT Optional cameras V 1024× : V 11gi ×Linear SiLU Linear

  86. [94]

    full atten- tion

    Project cameras Embedder Aggregator Add VGGT ‘camera token’ Local attention Global attention Repeat N times Linear Norm Concat. Output 3D repr. Figure A4.Encoder.The encoder takes as sourceVimagesand, optionally,Vcameras. The images are upsized (1) to the dimension expected by...

  87. [95]

    If only camera poses are available for a particular train- ing scene, thenw 2 = 0and we chooseλso that that λw1 := 1/1.35

  88. [96]

    • We choose aλ

    If both camera poses and points are available, the we proceed as follows. • We choose aλ. To do so, with 50% probability, we chooseλso thatλw 1 = 1/1.35(following LVSM) and with 50% probability so thatλw 2 = 1. • We choose which, if any, scaling factor to drop out. With 1/3 pr...

  89. [97]

    the null token, keeping only scale parameterλw 1, so that the source cameras are not available to the model, but scale is unambiguous

    Unposed” can also accept camera poses as conditioning, and we use it as our final model. the null token, keeping only scale parameterλw 1, so that the source cameras are not available to the model, but scale is unambiguous. The network is able to implicitly learn the meaning o...

Pith tools

Reviewed July 13, 2026 · model on record in the stance chip above.