REVIEW 3 major objections 6 minor 97 references
Decoupling 3D geometry from rendering turns ordinary images into 20K interactive navigation worlds that train agents better than traditional simulators.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 02:16 UTC pith:QEO7QG2X
load-bearing objection Solid end-to-end neural data engine for VLN; SOTA zero-shot Habitat numbers and a real Stretch study are real, but free-space fidelity of the feed-forward graph is still under-measured. the 3 major comments →
Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Scalable neural simulation that explicitly decouples metric 3D spatial anchoring (feed-forward feature Gaussians) from photorealistic observation synthesis (geometry-aware one-step pixel flow) can produce interactive environments whose generated vision-language-action data trains navigation policies that outperform models trained inside conventional simulators and transfer zero-shot to the real world.
What carries the argument
Geometry-Aware One-Step Pixel Flow: an alpha-gated MeanFlow renderer that keeps reliable Gaussian projections intact in high-opacity regions while generating missing structure in low-opacity regions in a single forward pass, enabling real-time panoramic RGB-D at ~40 FPS.
Load-bearing premise
The reconstructed feature Gaussians and opacity-gated completions must produce free-space connectivity accurate enough that collision-aware trajectories transfer to both mesh simulators and real homes.
What would settle it
Train the same navigation architecture only on Image2Sim data, then measure success rate and path efficiency on held-out Habitat validation scenes and on a physical robot in rooms never seen in training; if performance collapses relative to in-domain Habitat baselines or real-world trials fail systematically near obstacles, the geometry-transfer claim fails.
If this is right
- Navigation progress becomes primarily limited by available image and video collections rather than by expensive 3D scanning or hand-authored assets.
- Policies can be trained at multi-million-sample scale with continuous log-linear gains that have not yet saturated at 10 million samples.
- Zero-shot cross-simulator and real-robot transfer becomes achievable without any Habitat or real-world fine-tuning.
- The same pipeline can automatically expand both scene diversity and instruction diversity (path-following, object-centric, human-demand styles) in one automated loop.
Where Pith is reading between the lines
- If free-space accuracy holds, the same decoupling recipe should extend to other closed-loop embodied tasks that need metric geometry plus photorealism, such as mobile manipulation.
- The unsaturated scaling curve implies that further growth in consumer video corpora could continue to lift navigation performance without new simulation engines.
- Opacity-gated one-step flow may be a general pattern for any setting where partial 3D evidence must be completed without overwriting reliable measurements.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Image2Sim constructs interactive neural environments from posed RGB-D sequences by decoupling feed-forward 3D feature-Gaussian scene construction from a Geometry-Aware One-Step Pixel Flow renderer that completes sparse Gaussian projections into panoramic RGB-D. A collision-aware motion engine and VLM instruction pipeline convert ~20K scenes into >10M vision-language-action samples. Image2Nav, trained exclusively in these environments, reports new zero-shot SOTA on R2R-CE, RxR-CE, and REVERIE-CE inside Habitat and improved success rates on a Stretch 3 robot (path-follow 11/20, goal-oriented 9/20). A log-linear scaling curve from human R2R/RxR data through 1M–10M generated samples is presented, together with rendering ablations and novel-view metrics.
Significance. If the transfer results hold under stronger geometry checks, the work offers a practical route past the fidelity–scale tradeoff that has limited embodied navigation data: real scans do not scale, synthetic assets leave large sim-to-real gaps, and pure video generators lack persistent free-space structure. The explicit decoupling of metric Gaussian anchoring from one-step generative completion, the automated trajectory/instruction engine, and the unsaturated scaling curve from 35k to 10M samples are concrete contributions. Zero-shot Habitat SOTA after exclusive Image2Sim training and a real Stretch 3 study are valuable evidence that neural simulation can serve as a training substrate rather than only an offline renderer. The manuscript also ships a clear three-stage training curriculum, MeanFlow/JVP one-step formulation, and component ablations that make the rendering design inspectable.
major comments (3)
- [§3.5, App. D, Table 1, Table 2–3] The central transfer claim (exclusive Image2Sim training → zero-shot Habitat SOTA and real Stretch 3 gains) requires that the traversable voxel graph M (Sec. 3.5, App. D) have free-space connectivity and collision statistics close enough to Habitat meshes and real homes. Table 1 and Fig. 4 report only RGB-D novel-view metrics (PSNR/SSIM/LPIPS/FPS). There is no free-space IoU, clearance-error, or collision false-positive/negative evaluation of M against Habitat ground-truth meshes on shared Matterport/HM3D scenes. Without this, data volume alone (Table 3) remains a plausible alternative explanation for the Habitat gains. A geometry-fidelity ablation—or at least mesh-aligned free-space metrics on overlapping scenes—is load-bearing for the physically grounded simulation claim.
- [§4.4, Table 3, Figure 5] Table 3 and Fig. 5 show large gains when scaling from human R2R/RxR to Image2Sim-1M/5M/10M, but do not isolate environment geometry from instruction/trajectory volume. A controlled comparison that holds the instruction and path distribution fixed while replacing Image2Sim free-space with high-fidelity Habitat meshes (or the reverse) would test whether the neural free-space graph itself contributes, or whether the gains are primarily from more diverse language-action supervision over a larger scene pool. The current design cannot separate these factors.
- [§4.6, Table 5] Real-world evaluation (Table 5) uses N=20 trials per condition with no confidence intervals, variance, or failure-mode breakdown. The absolute gains (path-follow SR 8/20→11/20; goal-oriented 5/20→9/20) are directionally supportive but under-powered for a strong sim-to-real claim. Expanding trials, reporting binomial CIs or bootstrap intervals, and cataloguing failure modes (geometry errors vs. instruction mismatch vs. control) would make the real-world evidence proportionate to the paper’s main claim.
minor comments (6)
- [§5 Limitations] Limitations correctly flag compact renderer capacity, missing contact dynamics, and VLM linguistic bias; these should be cross-referenced more explicitly when interpreting Table 2 SOTA claims so readers do not over-read physical fidelity.
- [§4.3, App. B.2] App. B.2 usefully clarifies that Image2Nav on human R2R/RxR alone underperforms baselines and that gains come from the 10M samples; a short pointer to this clarification in the main §4.3 text would prevent misattribution to architecture.
- [§3.3, Eq. (4)–(5)] Eq. (4)–(5) introduce σ_small / σ_large without numerical values in the main text; listing them (or pointing to App. B.3) would aid reproducibility of the alpha-gated source state.
- [Figure 4] Figure 4 is informative but would benefit from a quantitative callout (e.g., local PSNR in hole regions) so the qualitative completion claim is easier to compare with Table 1.
- [§3.2–3.5] Minor notation: M is used both for the number of Gaussians (Eq. 1) and the traversable graph (Sec. 3.5); disambiguating would avoid confusion.
- [Abstract, Table 6] Abstract says “near 20K” / “nearly 20K”; Table 6 reports 19,936—align wording for precision.
Circularity Check
No circular derivation: simulator construction, data synthesis, and zero-shot transfer claims are empirical and evaluated on external domains.
full rationale
The paper's load-bearing chain is engineering plus empirical evaluation, not a mathematical derivation that reduces outputs to inputs by construction. Feed-forward feature Gaussians (Eq. 1) and the alpha-gated MeanFlow renderer (Eqs. 4–8) are trained with standard reconstruction, alignment, flow, distillation, and LPIPS losses against held-out panoramic RGB-D; novel-view metrics (Table 1) and ablations (Table 4) are measured against external ground truth, not forced by the loss definitions. The motion engine builds a traversable voxel graph M from the reconstructed Gaussians and plans collision-aware trajectories; these are then annotated by an off-the-shelf VLM. Navigation models are trained exclusively inside the resulting neural environments (including re-rendered R2R/RxR plus 10 M generated samples) and evaluated zero-shot inside Habitat (R2R-CE, RxR-CE, REVERIE-CE) and on a physical Stretch 3 never seen in training. Success rates, scaling curves (Table 3), and real-world trials (Table 5) are therefore external measurements, not algebraic identities or fitted parameters renamed as predictions. Self-citations (Dynam3D, D3D-VLP) appear only as competing baselines in Table 2 and do not supply uniqueness theorems or load-bearing premises. No step matches the enumerated circularity patterns; the evaluation domains remain independent of the training substrate.
Axiom & Free-Parameter Ledger
free parameters (5)
- loss weights (λ_rec, λ_align, λ_flow, λ_distill, λ_perc)
- EMA decay γ=0.999 for self-distillation teacher
- agent radius 0.15 m, eye height 1.25 m, step 0.25 m, turn 15°
- safety margin r_safe and weight w_safe in collision-aware path cost
- noise scales σ_small / σ_large in alpha-gated source state
axioms (5)
- domain assumption Posed RGB-D (or RGB recovered by off-the-shelf 3D foundation models) supplies metric geometry accurate enough for navigable free-space extraction.
- domain assumption DINOv3 patch features are a reliable semantic prior for both Gaussian lifting and SPADE conditioning.
- ad hoc to paper MeanFlow average-velocity parameterization plus JVP yields a stable one-step map from noisy Gaussian projections to photorealistic RGB-D.
- domain assumption VLM (Qwen3-VL-32B) annotations of macro-step image-action sequences produce instructions whose distribution is useful for policy learning.
- domain assumption Navigation-level collision (voxel clearance + sliding) is a sufficient physical model for the claimed transfer gains.
invented entities (3)
-
Geometry-Aware One-Step Pixel Flow renderer
no independent evidence
-
Feed-forward feature-Gaussian scene representation (Image2Sim G)
no independent evidence
-
Image2Sim automated embodied data engine
no independent evidence
read the original abstract
Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains constrained by the lack of scalable, high-fidelity, and physically grounded interactive environments. Although real-world scanned datasets offer visual realism, they are limited by scale. In contrast, synthetic simulators scale more easily but often exhibit large sim-to-real gaps. We introduce Image2Sim, a real-time neural simulation framework that constructs high-quality interactive environments from posed RGB-D image sequences. The central idea is to decouple 3D spatial anchoring from photorealistic observation synthesis. For scene construction, Image2Sim uses a feed-forward feature Gaussian model that lifts posed RGB-D observations into a 3D feature-Gaussian representation in a single pass. For rendering, we propose a Geometry-Aware One-Step Pixel Flow model that transforms sparse and noisy Gaussian projections into high-quality panoramic RGB-D observations. Image2Sim also serves as a fully automated embodied data engine that generates high-fidelity observations, executable actions, and diverse navigation instructions at scale. It converts large collections of videos and images into nearly 20K interactive scenes and synthesizes more than 10 million navigation training samples. Navigation models trained entirely in these neural environments achieve strong improvements on major benchmarks and transfer effectively to real-world zero-shot settings. These results suggest that scalable neural simulation can serve as a practical training substrate for embodied navigation at scale.
Figures
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[2]
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. Gemini: a family of highly capable multimodal models.arXiv preprint arXiv:2312.11805, 2023
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[3]
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timo- thée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971, 2023
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[4]
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4015–4026, 2023
work page 2023
-
[5]
Oriane Siméoni, Huy V V o, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, et al. Dinov3. arXiv preprint arXiv:2508.10104, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[6]
Vggt: Visual geometry grounded transformer
Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. Vggt: Visual geometry grounded transformer. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 5294–5306, 2025
work page 2025
-
[7]
High- resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High- resolution image synthesis with latent diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022
work page 2022
-
[8]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. InProceedings of the IEEE/CVF international conference on computer vision, pages 4195–4205, 2023
work page 2023
-
[9]
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, et al. V-jepa 2: Self-supervised video models enable understanding, prediction and planning.arXiv preprint arXiv:2506.09985, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[10]
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 3674–3683, 2018
work page 2018
-
[11]
Beyond the nav- graph: Vision-and-language navigation in continuous environments
Jacob Krantz, Erik Wijmans, Arjun Majumdar, Dhruv Batra, and Stefan Lee. Beyond the nav- graph: Vision-and-language navigation in continuous environments. InComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXVIII 16, pages 104–120. Springer, 2020
work page 2020
-
[12]
Room-Across- Room: Multilingual vision-and-language navigation with dense spatiotemporal grounding
Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge. Room-Across- Room: Multilingual vision-and-language navigation with dense spatiotemporal grounding. In Conference on Empirical Methods for Natural Language Processing (EMNLP), 2020. 10
work page 2020
-
[13]
Reverie: Remote embodied visual referring expression in real indoor environments
Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. Reverie: Remote embodied visual referring expression in real indoor environments. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9982–9991, 2020
work page 2020
-
[14]
Devendra Singh Chaplot, Dhiraj Prakashchand Gandhi, Abhinav Gupta, and Russ R Salakhut- dinov. Object goal navigation using goal-oriented semantic exploration.Advances in Neural Information Processing Systems, 33:4247–4258, 2020
work page 2020
-
[15]
Matterport3d: Learning from rgb-d data in indoor environments
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niebner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. InInternational Conference on 3D Vision (3DV), 2017
work page 2017
-
[16]
Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai
Santhosh Kumar Ramakrishnan, Aaron Gokaslan, Erik Wijmans, Oleksandr Maksymets, Alexan- der Clegg, John M Turner, Eric Undersander, Wojciech Galuba, Andrew Westbury, Angel X Chang, et al. Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai. InThirty-fifth Conference on Neural Information Processing Systems Datasets and B...
-
[17]
Gibson env: Real-world perception for embodied agents
Fei Xia, Amir R Zamir, Zhiyang He, Alexander Sax, Jitendra Malik, and Silvio Savarese. Gibson env: Real-world perception for embodied agents. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 9068–9079, 2018
work page 2018
-
[18]
Habitat: A platform for embodied ai research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, et al. Habitat: A platform for embodied ai research. InProceedings of the IEEE/CVF international conference on computer vision, pages 9339–9347, 2019
work page 2019
-
[19]
Matt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs, Kiana Ehsani, Jordi Salvador, Winson Han, Eric Kolve, Aniruddha Kembhavi, and Roozbeh Mottaghi. Procthor: Large-scale embodied ai using procedural generation.Advances in Neural Information Processing Systems, 35:5982–5994, 2022
work page 2022
-
[20]
Mukul Khanna, Yongsen Mao, Hanxiao Jiang, Sanjay Haresh, Brennan Shacklett, Dhruv Batra, Alexander Clegg, Eric Undersander, Angel X Chang, and Manolis Savva. Habitat synthetic scenes dataset (hssd-200): An analysis of 3d scene scale and realism tradeoffs for objectgoal navigation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...
work page 2024
-
[21]
AI2-THOR: An Interactive 3D Environment for Visual AI
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Matt Deitke, Kiana Ehsani, Daniel Gordon, Yuke Zhu, et al. Ai2-thor: An interactive 3d environment for visual ai.arXiv preprint arXiv:1712.05474, 2017
work page internal anchor Pith review Pith/arXiv arXiv 2017
-
[22]
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Do- minik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127, 2023
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[23]
Genie: Generative interactive environments
Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al. Genie: Generative interactive environments. InForty-first International Conference on Machine Learning, 2024
work page 2024
-
[24]
Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, and Yann LeCun. Navigation world models. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 15791–15801, 2025
work page 2025
-
[25]
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[26]
The Replica Dataset: A Digital Replica of Indoor Spaces
Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, et al. The replica dataset: A digital replica of indoor spaces.arXiv preprint arXiv:1906.05797, 2019. 11
work page internal anchor Pith review Pith/arXiv arXiv 1906
-
[27]
Liuyi Wang, Xinyuan Xia, Hui Zhao, Hanqing Wang, Tai Wang, Yilun Chen, Chengju Liu, Qijun Chen, and Jiangmiao Pang. Rethinking the embodied gap in vision-and-language navigation: A holistic study of physical and visual disparities. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 9455–9465, 2025
work page 2025
-
[28]
Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning
Mayank Mittal, Pascal Roth, James Tigue, Antoine Richard, Octi Zhang, Peter Du, Antonio Serrano-Munoz, Xinjie Yao, René Zurbrügg, Nikita Rudin, et al. Isaac lab: A gpu-accelerated simulation framework for multi-modal robot learning.arXiv preprint arXiv:2511.04831, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[29]
Sihao Lin, Zerui Li, Xunyi Zhao, Gengze Zhou, Liuyi Wang, Rong Wei, Rui Tang, Juncheng Li, Hanqing Wang, Jiangmiao Pang, et al. Vlnverse: A benchmark for vision-language navigation with versatile, embodied, realistic simulation and evaluation.arXiv preprint arXiv:2512.19021, 2025
-
[30]
Kiana Ehsani, Tanmay Gupta, Rose Hendrix, Jordi Salvador, Luca Weihs, Kuo-Hao Zeng, Ku- nal Pratap Singh, Yejin Kim, Winson Han, Alvaro Herrasti, et al. Spoc: Imitating shortest paths in simulation enables effective navigation and manipulation in the real world. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 162...
work page 2024
-
[31]
Yejin Kim, Wilbert Pumacay, Omar Rayyan, Max Argus, Winson Han, Eli VanderBilt, Jordi Salvador, Abhay Deshpande, Rose Hendrix, Snehal Jauhri, et al. Molmospaces: A large-scale open ecosystem for robot navigation and manipulation.arXiv preprint arXiv:2602.11337, 2026
-
[32]
Holodeck: Language guided generation of 3d embodied ai environments
Yue Yang, Fan-Yun Sun, Luca Weihs, Eli VanderBilt, Alvaro Herrasti, Winson Han, Jiajun Wu, Nick Haber, Ranjay Krishna, Lingjie Liu, et al. Holodeck: Language guided generation of 3d embodied ai environments. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16227–16237, 2024
work page 2024
-
[33]
RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
Soroush Nasiriany, Abhiram Maddukuri, Lance Zhang, Adeet Parikh, Aaron Lo, Abhishek Joshi, Ajay Mandlekar, and Yuke Zhu. Robocasa: Large-scale simulation of everyday tasks for generalist robots.arXiv preprint arXiv:2406.02523, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[34]
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoor- thi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021
work page 2021
-
[35]
3d gaussian splatting for real-time radiance field rendering.ACM Trans
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George Drettakis, et al. 3d gaussian splatting for real-time radiance field rendering.ACM Trans. Graph., 42(4):139–1, 2023
work page 2023
-
[36]
Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields
Shijie Zhou, Haoran Chang, Sicheng Jiang, Zhiwen Fan, Zehao Zhu, Dejia Xu, Pradyumna Chari, Suya You, Zhangyang Wang, and Achuta Kadambi. Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21676–21685, 2024
work page 2024
-
[37]
Vid2sim: Realistic and interactive simulation from video for urban navigation
Ziyang Xie, Zhizheng Liu, Zhenghao Peng, Wayne Wu, and Bolei Zhou. Vid2sim: Realistic and interactive simulation from video for urban navigation. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 1581–1591, 2025
work page 2025
-
[38]
Jiahang Liu, Yuanxing Duan, Jiazhao Zhang, Minghan Li, Shaoan Wang, Zhizheng Zhang, and He Wang. Navgsim: High-fidelity gaussian splatting simulator for large-scale navigation.arXiv preprint arXiv:2603.15186, 2026
-
[39]
Bingchen Miao, Rong Wei, Zhiqi Ge, Shiqi Gao, Jingzhe Zhu, Renhan Wang, Siliang Tang, Jun Xiao, Rui Tang, Juncheng Li, et al. Towards physically executable 3d gaussian for embodied navigation.arXiv preprint arXiv:2510.21307, 2025
-
[40]
Embodiedsplat: Personalized real-to-sim-to-real navigation with gaussian splats from a mobile device
Gunjan Chhablani, Xiaomeng Ye, Muhammad Zubair Irshad, and Zsolt Kira. Embodiedsplat: Personalized real-to-sim-to-real navigation with gaussian splats from a mobile device. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 25431– 25441, 2025. 12
work page 2025
-
[41]
Xinhao Liu, Jiaqi Li, Youming Deng, Ruxin Chen, Yingjia Zhang, Yifei Ma, Li Guo, Yiming Li, Jing Zhang, and Chen Feng. Wanderland: Geometrically grounded simulation for open-world embodied ai.arXiv preprint arXiv:2511.20620, 2025
-
[42]
Seungyeon Yoo, Youngseok Jang, Dabin Kim, Youngsoo Han, Seungwoo Jung, and H Jin Kim. Ready-go: Real-to-sim dynamic 3d gaussian splatting simulation for environment-specific visual navigation with moving obstacles.arXiv preprint arXiv:2602.11575, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[43]
Alejandro Escontrela, Justin Kerr, Arthur Allshire, Jonas Frey, Rocky Duan, Carmelo Sferrazza, and Pieter Abbeel. Gaussgym: An open-source real-to-sim framework for learning locomotion from pixels.arXiv preprint arXiv:2510.15352, 2025
-
[44]
Floor flattening of image-based 3d recon- struction for mobile robots
Seonghwan Sim, Yeji Kim, and Sung Soo Hwang. Floor flattening of image-based 3d recon- struction for mobile robots. In2026 IEEE International Conference on Artificial Intelligence and eXtended and Virtual Reality (AIxVR), pages 441–446. IEEE, 2026
work page 2026
-
[45]
pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction
David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 19457–19467, 2024
work page 2024
-
[46]
Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images
Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images. InEuropean conference on computer vision, pages 370–386. Springer, 2024
work page 2024
-
[47]
Yunsong Wang, Tianxin Huang, Hanlin Chen, and Gim Hee Lee. Freesplat: Generalizable 3d gaussian splatting towards free view synthesis of indoor scenes.Advances in Neural Information Processing Systems, 37:107326–107349, 2024
work page 2024
-
[48]
Botao Ye, Boqi Chen, Haofei Xu, Daniel Barath, and Marc Pollefeys. Yonosplat: You only need one model for feedforward 3d gaussian splatting.arXiv preprint arXiv:2511.07321, 2025
-
[49]
Suyoung Lee, Jaeyoung Chung, Kihoon Kim, Jaeyoo Huh, Gunhee Lee, Minsoo Lee, and Kyoung Mu Lee. Omnisplat: Taming feed-forward 3d gaussian splatting for omnidirectional im- ages with editable capabilities. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 16356–16365, 2025
work page 2025
-
[50]
Splatter-360: Generalizable 360 gaussian splatting for wide-baseline panoramic images
Zheng Chen, Chenming Wu, Zhelun Shen, Chen Zhao, Weicai Ye, Haocheng Feng, Errui Ding, and Song-Hai Zhang. Splatter-360: Generalizable 360 gaussian splatting for wide-baseline panoramic images. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 21590–21599, 2025
work page 2025
-
[51]
Long-LRM++: Preserving Fine Details in Feed-Forward Wide-Coverage Reconstruction
Chen Ziwen, Hao Tan, Peng Wang, Zexiang Xu, and Li Fuxin. Long-lrm++: Preserving fine details in feed-forward wide-coverage reconstruction.arXiv preprint arXiv:2512.10267, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[52]
Lihan Jiang, Yucheng Mao, Linning Xu, Tao Lu, Kerui Ren, Yichen Jin, Xudong Xu, Mulin Yu, Jiangmiao Pang, Feng Zhao, et al. Anysplat: Feed-forward 3d gaussian splatting from unconstrained views.ACM Transactions on Graphics (TOG), 44(6):1–16, 2025
work page 2025
-
[53]
Jie Chen, Yuxin Cai, Yizhuo Wang, Ruofei Bai, Yuhong Cao, Jun Li, Yau Wei Yun, and Guillaume Sartoretti. Imaginav: Scalable embodied navigation via generative visual prediction and inverse dynamics.arXiv preprint arXiv:2603.13833, 2026
-
[54]
Soon: Scenario oriented object navigation with graph-based exploration
Fengda Zhu, Xiwen Liang, Yi Zhu, Qizhi Yu, Xiaojun Chang, and Xiaodan Liang. Soon: Scenario oriented object navigation with graph-based exploration. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12689–12699, 2021
work page 2021
-
[55]
Daniel Fried, Ronghang Hu, V olkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. Speaker- follower models for vision-and-language navigation.Advances in neural information processing systems, 31, 2018. 13
work page 2018
-
[56]
Towards learning a generic agent for vision-and-language navigation via pre-training
Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin, and Jianfeng Gao. Towards learning a generic agent for vision-and-language navigation via pre-training. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 13137–13146, 2020
work page 2020
-
[57]
Aishwarya Kamath, Peter Anderson, Su Wang, Jing Yu Koh, Alexander Ku, Austin Waters, Yinfei Yang, Jason Baldridge, and Zarana Parekh. A new path: Scaling vision-and-language navigation with synthetic instructions and imitation learning. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10813–10823, 2023
work page 2023
-
[58]
Learning from unlabeled 3d environments for vision-and-language navigation
Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid, and Ivan Laptev. Learning from unlabeled 3d environments for vision-and-language navigation. InEuropean Conference on Computer Vision, pages 638–655. Springer, 2022
work page 2022
-
[59]
Scaling data generation in vision-and-language navigation
Zun Wang, Jialu Li, Yicong Hong, Yi Wang, Qi Wu, Mohit Bansal, Stephen Gould, Hao Tan, and Yu Qiao. Scaling data generation in vision-and-language navigation. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 12009–12020, 2023
work page 2023
-
[60]
Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
Zun Wang, Jialu Li, Yicong Hong, Songze Li, Kunchang Li, Shoubin Yu, Yi Wang, Yu Qiao, Yali Wang, Mohit Bansal, et al. Bootstrapping language-guided navigation learning with self-refining data flywheel.arXiv preprint arXiv:2412.08467, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[61]
NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM
Zihan Wang, Yaohui Zhu, Gim Hee Lee, and Yachun Fan. Navrag: Generating user de- mand instructions for embodied navigation through retrieval-augmented llm.arXiv preprint arXiv:2502.11142, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[62]
Songze Li, Zun Wang, Gengze Zhou, Jialu Li, Xiangyu Zeng, Limin Wang, Yu Qiao, Qi Wu, Mohit Bansal, and Yi Wang. Learning goal-oriented language-guided navigation with self- improving demonstrations at scale.arXiv preprint arXiv:2509.24910, 2025
-
[63]
Embodied navigation foundation model.arXiv preprint arXiv:2509.12129, 2025
Jiazhao Zhang, Anqi Li, Yunpeng Qi, Minghan Li, Jiahang Liu, Shaoan Wang, Haoran Liu, Gengze Zhou, Yuze Wu, Xingxing Li, et al. Embodied navigation foundation model.arXiv preprint arXiv:2509.12129, 2025
-
[64]
Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation
Shuo Wang, Yucheng Wang, Guoxin Lian, Yongcai Wang, Maiyue Chen, Kaihui Wang, Bo Zhang, Zhizhong Su, Yutian Zhou, Wanting Li, et al. Progress-think: Semantic progress reasoning for vision-language navigation.arXiv preprint arXiv:2511.17097, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[65]
NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
Gengze Zhou, Yicong Hong, Zun Wang, Xin Eric Wang, and Qi Wu. Navgpt-2: Un- leashing navigational reasoning capability for large vision-language models.arXiv preprint arXiv:2407.12366, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[66]
$\pi^3$: Permutation-Equivariant Visual Geometry Learning
Yifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang, Yang Zhou, Zizun Li, Junyi Chen, Jiangmiao Pang, Chunhua Shen, and Tong He. Pi3: Permutation-equivariant visual geometry learning.arXiv preprint arXiv:2507.13347, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[67]
Depth Anything 3: Recovering the Visual Space from Any Views
Haotong Lin, Sili Chen, Junhao Liew, Donny Y Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[68]
Semantic image synthesis with spatially-adaptive normalization
Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. Semantic image synthesis with spatially-adaptive normalization. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2337–2346, 2019
work page 2019
-
[69]
Flow Matching for Generative Modeling
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling.arXiv preprint arXiv:2210.02747, 2022
work page internal anchor Pith review Pith/arXiv arXiv 2022
-
[70]
Mean Flows for One-step Generative Modeling
Zhengyang Geng, Mingyang Deng, Xingjian Bai, J Zico Kolter, and Kaiming He. Mean flows for one-step generative modeling.arXiv preprint arXiv:2505.13447, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[71]
One-step Latent-free Image Generation with Pixel Mean Flows
Yiyang Lu, Susie Lu, Qiao Sun, Hanhong Zhao, Zhicheng Jiang, Xianbang Wang, Tianhong Li, Zhengyang Geng, and Kaiming He. One-step latent-free image generation with pixel mean flows.arXiv preprint arXiv:2601.22158, 2026. 14
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[72]
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. InProceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021
work page 2021
-
[73]
The unrea- sonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unrea- sonable effectiveness of deep features as a perceptual metric. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 586–595, 2018
work page 2018
-
[74]
The marathon 2: A navigation system
Steve Macenski, Francisco Martín, Ruffin White, and Jonatan Ginés Clavero. The marathon 2: A navigation system. In2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 2718–2725. IEEE, 2020
work page 2020
-
[75]
Steve Macenski, Tom Moore, David V Lu, Alexey Merzlyakov, and Michael Ferguson. From the desks of ros maintainers: A survey of modern & capable mobile robotics algorithms in the robot operating system 2.Robotics and Autonomous Systems, 168:104493, 2023
work page 2023
-
[76]
Regulated pure pursuit for robot path tracking.Autonomous Robots, 47(6):685–694, 2023
Steve Macenski, Shrijit Singh, Francisco Martín, and Jonatan Ginés. Regulated pure pursuit for robot path tracking.Autonomous Robots, 47(6):685–694, 2023
work page 2023
-
[77]
Realsee3d: A large-scale multi-view rgb-d dataset of indoor scenes (version 1.0), 2025
Linyuan Li, Yan Wu, Xi Li, Lingli Wang, Tong Rao, Jie Zhou, Cihui Pan, and Xinchen Hui. Realsee3d: A large-scale multi-view rgb-d dataset of indoor scenes (version 1.0), 2025. URL https://doi.org/10.5281/zenodo.17826243
-
[78]
Structured3d: A large photo-realistic dataset for structured 3d modeling
Jia Zheng, Junfei Zhang, Jing Li, Rui Tang, Shenghua Gao, and Zihan Zhou. Structured3d: A large photo-realistic dataset for structured 3d modeling. InProceedings of The European Conference on Computer Vision (ECCV), 2020
work page 2020
-
[79]
Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data
Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Yuri Feigin, Peter Fu, Thomas Gebauer, Daniel Kurz, Tal Dimry, Brandon Joffe, Arik Schwartz, et al. Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data. InThirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1)
-
[80]
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017
work page 2017
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.