Pith. sign in

REVIEW 3 major objections 6 minor 97 references

Decoupling 3D geometry from rendering turns ordinary images into 20K interactive navigation worlds that train agents better than traditional simulators.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 02:16 UTC pith:QEO7QG2X

load-bearing objection Solid end-to-end neural data engine for VLN; SOTA zero-shot Habitat numbers and a real Stretch study are real, but free-space fidelity of the feed-forward graph is still under-measured. the 3 major comments →

arxiv 2607.05765 v1 pith:QEO7QG2X submitted 2026-07-07 cs.CV cs.RO

Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator

classification cs.CV cs.RO
keywords embodied navigationneural simulation3D Gaussian splattingpixel flowvision-language navigationsim-to-realfeed-forward reconstructiondata engine
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Embodied navigation has been held back by a scarcity of large, realistic, interactive 3D environments: real scans are accurate but few, while synthetic simulators are plentiful but leave a big sim-to-real gap. Image2Sim solves this by turning posed RGB-D videos and images into real-time neural simulators. It first builds an explicit 3D feature-Gaussian scene in one feed-forward pass, then uses a geometry-aware one-step pixel-flow model to fill holes and produce clean panoramic RGB-D views. A collision-aware motion engine plus an automated language annotator then generate executable trajectories and instructions, converting nearly 20K scenes into more than 10 million training samples. Agents trained only inside these neural worlds set new zero-shot records on standard Habitat benchmarks and improve success rates when transferred to a real home robot.

Core claim

Scalable neural simulation that explicitly decouples metric 3D spatial anchoring (feed-forward feature Gaussians) from photorealistic observation synthesis (geometry-aware one-step pixel flow) can produce interactive environments whose generated vision-language-action data trains navigation policies that outperform models trained inside conventional simulators and transfer zero-shot to the real world.

What carries the argument

Geometry-Aware One-Step Pixel Flow: an alpha-gated MeanFlow renderer that keeps reliable Gaussian projections intact in high-opacity regions while generating missing structure in low-opacity regions in a single forward pass, enabling real-time panoramic RGB-D at ~40 FPS.

Load-bearing premise

The reconstructed feature Gaussians and opacity-gated completions must produce free-space connectivity accurate enough that collision-aware trajectories transfer to both mesh simulators and real homes.

What would settle it

Train the same navigation architecture only on Image2Sim data, then measure success rate and path efficiency on held-out Habitat validation scenes and on a physical robot in rooms never seen in training; if performance collapses relative to in-domain Habitat baselines or real-world trials fail systematically near obstacles, the geometry-transfer claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Navigation progress becomes primarily limited by available image and video collections rather than by expensive 3D scanning or hand-authored assets.
  • Policies can be trained at multi-million-sample scale with continuous log-linear gains that have not yet saturated at 10 million samples.
  • Zero-shot cross-simulator and real-robot transfer becomes achievable without any Habitat or real-world fine-tuning.
  • The same pipeline can automatically expand both scene diversity and instruction diversity (path-following, object-centric, human-demand styles) in one automated loop.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If free-space accuracy holds, the same decoupling recipe should extend to other closed-loop embodied tasks that need metric geometry plus photorealism, such as mobile manipulation.
  • The unsaturated scaling curve implies that further growth in consumer video corpora could continue to lift navigation performance without new simulation engines.
  • Opacity-gated one-step flow may be a general pattern for any setting where partial 3D evidence must be completed without overwriting reliable measurements.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. Image2Sim constructs interactive neural environments from posed RGB-D sequences by decoupling feed-forward 3D feature-Gaussian scene construction from a Geometry-Aware One-Step Pixel Flow renderer that completes sparse Gaussian projections into panoramic RGB-D. A collision-aware motion engine and VLM instruction pipeline convert ~20K scenes into >10M vision-language-action samples. Image2Nav, trained exclusively in these environments, reports new zero-shot SOTA on R2R-CE, RxR-CE, and REVERIE-CE inside Habitat and improved success rates on a Stretch 3 robot (path-follow 11/20, goal-oriented 9/20). A log-linear scaling curve from human R2R/RxR data through 1M–10M generated samples is presented, together with rendering ablations and novel-view metrics.

Significance. If the transfer results hold under stronger geometry checks, the work offers a practical route past the fidelity–scale tradeoff that has limited embodied navigation data: real scans do not scale, synthetic assets leave large sim-to-real gaps, and pure video generators lack persistent free-space structure. The explicit decoupling of metric Gaussian anchoring from one-step generative completion, the automated trajectory/instruction engine, and the unsaturated scaling curve from 35k to 10M samples are concrete contributions. Zero-shot Habitat SOTA after exclusive Image2Sim training and a real Stretch 3 study are valuable evidence that neural simulation can serve as a training substrate rather than only an offline renderer. The manuscript also ships a clear three-stage training curriculum, MeanFlow/JVP one-step formulation, and component ablations that make the rendering design inspectable.

major comments (3)
  1. [§3.5, App. D, Table 1, Table 2–3] The central transfer claim (exclusive Image2Sim training → zero-shot Habitat SOTA and real Stretch 3 gains) requires that the traversable voxel graph M (Sec. 3.5, App. D) have free-space connectivity and collision statistics close enough to Habitat meshes and real homes. Table 1 and Fig. 4 report only RGB-D novel-view metrics (PSNR/SSIM/LPIPS/FPS). There is no free-space IoU, clearance-error, or collision false-positive/negative evaluation of M against Habitat ground-truth meshes on shared Matterport/HM3D scenes. Without this, data volume alone (Table 3) remains a plausible alternative explanation for the Habitat gains. A geometry-fidelity ablation—or at least mesh-aligned free-space metrics on overlapping scenes—is load-bearing for the physically grounded simulation claim.
  2. [§4.4, Table 3, Figure 5] Table 3 and Fig. 5 show large gains when scaling from human R2R/RxR to Image2Sim-1M/5M/10M, but do not isolate environment geometry from instruction/trajectory volume. A controlled comparison that holds the instruction and path distribution fixed while replacing Image2Sim free-space with high-fidelity Habitat meshes (or the reverse) would test whether the neural free-space graph itself contributes, or whether the gains are primarily from more diverse language-action supervision over a larger scene pool. The current design cannot separate these factors.
  3. [§4.6, Table 5] Real-world evaluation (Table 5) uses N=20 trials per condition with no confidence intervals, variance, or failure-mode breakdown. The absolute gains (path-follow SR 8/20→11/20; goal-oriented 5/20→9/20) are directionally supportive but under-powered for a strong sim-to-real claim. Expanding trials, reporting binomial CIs or bootstrap intervals, and cataloguing failure modes (geometry errors vs. instruction mismatch vs. control) would make the real-world evidence proportionate to the paper’s main claim.
minor comments (6)
  1. [§5 Limitations] Limitations correctly flag compact renderer capacity, missing contact dynamics, and VLM linguistic bias; these should be cross-referenced more explicitly when interpreting Table 2 SOTA claims so readers do not over-read physical fidelity.
  2. [§4.3, App. B.2] App. B.2 usefully clarifies that Image2Nav on human R2R/RxR alone underperforms baselines and that gains come from the 10M samples; a short pointer to this clarification in the main §4.3 text would prevent misattribution to architecture.
  3. [§3.3, Eq. (4)–(5)] Eq. (4)–(5) introduce σ_small / σ_large without numerical values in the main text; listing them (or pointing to App. B.3) would aid reproducibility of the alpha-gated source state.
  4. [Figure 4] Figure 4 is informative but would benefit from a quantitative callout (e.g., local PSNR in hole regions) so the qualitative completion claim is easier to compare with Table 1.
  5. [§3.2–3.5] Minor notation: M is used both for the number of Gaussians (Eq. 1) and the traversable graph (Sec. 3.5); disambiguating would avoid confusion.
  6. [Abstract, Table 6] Abstract says “near 20K” / “nearly 20K”; Table 6 reports 19,936—align wording for precision.

Circularity Check

0 steps flagged

No circular derivation: simulator construction, data synthesis, and zero-shot transfer claims are empirical and evaluated on external domains.

full rationale

The paper's load-bearing chain is engineering plus empirical evaluation, not a mathematical derivation that reduces outputs to inputs by construction. Feed-forward feature Gaussians (Eq. 1) and the alpha-gated MeanFlow renderer (Eqs. 4–8) are trained with standard reconstruction, alignment, flow, distillation, and LPIPS losses against held-out panoramic RGB-D; novel-view metrics (Table 1) and ablations (Table 4) are measured against external ground truth, not forced by the loss definitions. The motion engine builds a traversable voxel graph M from the reconstructed Gaussians and plans collision-aware trajectories; these are then annotated by an off-the-shelf VLM. Navigation models are trained exclusively inside the resulting neural environments (including re-rendered R2R/RxR plus 10 M generated samples) and evaluated zero-shot inside Habitat (R2R-CE, RxR-CE, REVERIE-CE) and on a physical Stretch 3 never seen in training. Success rates, scaling curves (Table 3), and real-world trials (Table 5) are therefore external measurements, not algebraic identities or fitted parameters renamed as predictions. Self-citations (Dynam3D, D3D-VLP) appear only as competing baselines in Table 2 and do not supply uniqueness theorems or load-bearing premises. No step matches the enumerated circularity patterns; the evaluation domains remain independent of the training substrate.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 3 invented entities

The central claim rests on standard neural-rendering and imitation-learning machinery plus a handful of paper-specific design choices (alpha-gated source state, MeanFlow one-step objective, EMA teacher, collision cost weights). No new physical constants are fitted; free parameters are ordinary training hyper-parameters and kinematic constants. The main invented entities are the composite Image2Sim pipeline and the Geometry-Aware One-Step Pixel Flow renderer.

free parameters (5)
  • loss weights (λ_rec, λ_align, λ_flow, λ_distill, λ_perc)
    Stage-dependent coefficients that balance reconstruction, alignment, flow, distillation and LPIPS; set by curriculum rather than derived.
  • EMA decay γ=0.999 for self-distillation teacher
    Hand-chosen momentum for the teacher network; affects single-step stability.
  • agent radius 0.15 m, eye height 1.25 m, step 0.25 m, turn 15°
    Kinematic constants matching Habitat defaults; free design choices that define the action space and collision footprint.
  • safety margin r_safe and weight w_safe in collision-aware path cost
    Planner hyper-parameters that bias trajectories away from obstacles; not derived from first principles.
  • noise scales σ_small / σ_large in alpha-gated source state
    Control how much the pixel-flow model trusts versus regenerates projected evidence.
axioms (5)
  • domain assumption Posed RGB-D (or RGB recovered by off-the-shelf 3D foundation models) supplies metric geometry accurate enough for navigable free-space extraction.
    Stated in §3.1; all subsequent collision graphs and trajectories inherit any depth/pose error.
  • domain assumption DINOv3 patch features are a reliable semantic prior for both Gaussian lifting and SPADE conditioning.
    Frozen backbone used throughout §3.2–3.3 without independent verification of feature-to-layout fidelity.
  • ad hoc to paper MeanFlow average-velocity parameterization plus JVP yields a stable one-step map from noisy Gaussian projections to photorealistic RGB-D.
    Adopted from recent MeanFlow papers and specialized with alpha gating; correctness is empirical.
  • domain assumption VLM (Qwen3-VL-32B) annotations of macro-step image-action sequences produce instructions whose distribution is useful for policy learning.
    §3.6 and Limitations acknowledge possible linguistic bias and semantic mismatch.
  • domain assumption Navigation-level collision (voxel clearance + sliding) is a sufficient physical model for the claimed transfer gains.
    Explicitly limited in the Limitations section; richer contact dynamics are omitted.
invented entities (3)
  • Geometry-Aware One-Step Pixel Flow renderer no independent evidence
    purpose: Map sparse/noisy Gaussian projections to high-fidelity panoramic RGB-D in a single forward pass under alpha and semantic conditioning.
    Core novel module; no independent external measurement of the flow field outside the paper's own ablations.
  • Feed-forward feature-Gaussian scene representation (Image2Sim G) no independent evidence
    purpose: Persistent 3D spatial and semantic anchor built in one pass from posed RGB-D.
    Combines existing 3DGS and feature ideas into a specific dual-stream encoder used as the simulator backbone.
  • Image2Sim automated embodied data engine no independent evidence
    purpose: Jointly produce observations, collision-aware actions and multi-style language instructions at 10 M scale.
    System-level invention whose value is demonstrated only by the downstream navigation numbers.

pith-pipeline@v1.1.0-grok45 · 29829 in / 3411 out tokens · 63054 ms · 2026-07-11T02:16:43.479090+00:00 · methodology

0 comments
read the original abstract

Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains constrained by the lack of scalable, high-fidelity, and physically grounded interactive environments. Although real-world scanned datasets offer visual realism, they are limited by scale. In contrast, synthetic simulators scale more easily but often exhibit large sim-to-real gaps. We introduce Image2Sim, a real-time neural simulation framework that constructs high-quality interactive environments from posed RGB-D image sequences. The central idea is to decouple 3D spatial anchoring from photorealistic observation synthesis. For scene construction, Image2Sim uses a feed-forward feature Gaussian model that lifts posed RGB-D observations into a 3D feature-Gaussian representation in a single pass. For rendering, we propose a Geometry-Aware One-Step Pixel Flow model that transforms sparse and noisy Gaussian projections into high-quality panoramic RGB-D observations. Image2Sim also serves as a fully automated embodied data engine that generates high-fidelity observations, executable actions, and diverse navigation instructions at scale. It converts large collections of videos and images into nearly 20K interactive scenes and synthesizes more than 10 million navigation training samples. Navigation models trained entirely in these neural environments achieve strong improvements on major benchmarks and transfer effectively to real-world zero-shot settings. These results suggest that scalable neural simulation can serve as a practical training substrate for embodied navigation at scale.

Figures

Figures reproduced from arXiv: 2607.05765 by Gim Hee Lee, Seungjun Lee, Yinghao Xu, Zihan Wang.

Figure 1
Figure 1. Figure 1: Comparison of (a) traditional navigation data pipeline and (b) our Image2Sim framework. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Framework of the feed-forward GS encoder and one-step pixel flow model. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Framework of the motion simulation engine and automated instruction generation. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Visualization of Gaussian Splatting and Pixel Flow renderings in high-noise sparse scenes. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Scaling law curve of training data. 4.5 Ablation Study We analyze the physical and geometric princi￾ples underlying our pixel-flow architecture in [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Distribution of instruction lengths (number of words) in the generated dataset. [PITH_FULL_IMAGE:figures/full_fig_p017_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Distribution of trajectory lengths (in meters). [PITH_FULL_IMAGE:figures/full_fig_p018_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Distribution of trajectory steps (number of discrete actions). [PITH_FULL_IMAGE:figures/full_fig_p018_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Model Architecture of Image2Sim Renderer (Please zoom in for better viewing). [PITH_FULL_IMAGE:figures/full_fig_p020_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Visualization of navigable voxels and path planning. [PITH_FULL_IMAGE:figures/full_fig_p022_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

97 extracted references · 97 canonical work pages · 25 internal anchors

  1. [1]

    GPT-4 Technical Report

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023

  2. [2]

    Gemini: A Family of Highly Capable Multimodal Models

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. Gemini: a family of highly capable multimodal models.arXiv preprint arXiv:2312.11805, 2023

  3. [3]

    LLaMA: Open and Efficient Foundation Language Models

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timo- thée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971, 2023

  4. [4]

    Segment anything

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4015–4026, 2023

  5. [5]

    Oriane Siméoni, Huy V V o, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, et al. Dinov3. arXiv preprint arXiv:2508.10104, 2025

  6. [6]

    Vggt: Visual geometry grounded transformer

    Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. Vggt: Visual geometry grounded transformer. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 5294–5306, 2025

  7. [7]

    High- resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High- resolution image synthesis with latent diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022

  8. [8]

    Scalable diffusion models with transformers

    William Peebles and Saining Xie. Scalable diffusion models with transformers. InProceedings of the IEEE/CVF international conference on computer vision, pages 4195–4205, 2023

  9. [9]

    V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

    Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, et al. V-jepa 2: Self-supervised video models enable understanding, prediction and planning.arXiv preprint arXiv:2506.09985, 2025

  10. [10]

    Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

    Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 3674–3683, 2018

  11. [11]

    Beyond the nav- graph: Vision-and-language navigation in continuous environments

    Jacob Krantz, Erik Wijmans, Arjun Majumdar, Dhruv Batra, and Stefan Lee. Beyond the nav- graph: Vision-and-language navigation in continuous environments. InComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXVIII 16, pages 104–120. Springer, 2020

  12. [12]

    Room-Across- Room: Multilingual vision-and-language navigation with dense spatiotemporal grounding

    Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge. Room-Across- Room: Multilingual vision-and-language navigation with dense spatiotemporal grounding. In Conference on Empirical Methods for Natural Language Processing (EMNLP), 2020. 10

  13. [13]

    Reverie: Remote embodied visual referring expression in real indoor environments

    Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. Reverie: Remote embodied visual referring expression in real indoor environments. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9982–9991, 2020

  14. [14]

    Object goal navigation using goal-oriented semantic exploration.Advances in Neural Information Processing Systems, 33:4247–4258, 2020

    Devendra Singh Chaplot, Dhiraj Prakashchand Gandhi, Abhinav Gupta, and Russ R Salakhut- dinov. Object goal navigation using goal-oriented semantic exploration.Advances in Neural Information Processing Systems, 33:4247–4258, 2020

  15. [15]

    Matterport3d: Learning from rgb-d data in indoor environments

    Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niebner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. InInternational Conference on 3D Vision (3DV), 2017

  16. [16]

    Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai

    Santhosh Kumar Ramakrishnan, Aaron Gokaslan, Erik Wijmans, Oleksandr Maksymets, Alexan- der Clegg, John M Turner, Eric Undersander, Wojciech Galuba, Andrew Westbury, Angel X Chang, et al. Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai. InThirty-fifth Conference on Neural Information Processing Systems Datasets and B...

  17. [17]

    Gibson env: Real-world perception for embodied agents

    Fei Xia, Amir R Zamir, Zhiyang He, Alexander Sax, Jitendra Malik, and Silvio Savarese. Gibson env: Real-world perception for embodied agents. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 9068–9079, 2018

  18. [18]

    Habitat: A platform for embodied ai research

    Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, et al. Habitat: A platform for embodied ai research. InProceedings of the IEEE/CVF international conference on computer vision, pages 9339–9347, 2019

  19. [19]

    Procthor: Large-scale embodied ai using procedural generation.Advances in Neural Information Processing Systems, 35:5982–5994, 2022

    Matt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs, Kiana Ehsani, Jordi Salvador, Winson Han, Eric Kolve, Aniruddha Kembhavi, and Roozbeh Mottaghi. Procthor: Large-scale embodied ai using procedural generation.Advances in Neural Information Processing Systems, 35:5982–5994, 2022

  20. [20]

    Habitat synthetic scenes dataset (hssd-200): An analysis of 3d scene scale and realism tradeoffs for objectgoal navigation

    Mukul Khanna, Yongsen Mao, Hanxiao Jiang, Sanjay Haresh, Brennan Shacklett, Dhruv Batra, Alexander Clegg, Eric Undersander, Angel X Chang, and Manolis Savva. Habitat synthetic scenes dataset (hssd-200): An analysis of 3d scene scale and realism tradeoffs for objectgoal navigation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...

  21. [21]

    AI2-THOR: An Interactive 3D Environment for Visual AI

    Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Matt Deitke, Kiana Ehsani, Daniel Gordon, Yuke Zhu, et al. Ai2-thor: An interactive 3d environment for visual ai.arXiv preprint arXiv:1712.05474, 2017

  22. [22]

    Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

    Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Do- minik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127, 2023

  23. [23]

    Genie: Generative interactive environments

    Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al. Genie: Generative interactive environments. InForty-first International Conference on Machine Learning, 2024

  24. [24]

    Navigation world models

    Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, and Yann LeCun. Navigation world models. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 15791–15801, 2025

  25. [25]

    Qwen3-VL Technical Report

    Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025

  26. [26]

    The Replica Dataset: A Digital Replica of Indoor Spaces

    Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, et al. The replica dataset: A digital replica of indoor spaces.arXiv preprint arXiv:1906.05797, 2019. 11

  27. [27]

    Rethinking the embodied gap in vision-and-language navigation: A holistic study of physical and visual disparities

    Liuyi Wang, Xinyuan Xia, Hui Zhao, Hanqing Wang, Tai Wang, Yilun Chen, Chengju Liu, Qijun Chen, and Jiangmiao Pang. Rethinking the embodied gap in vision-and-language navigation: A holistic study of physical and visual disparities. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 9455–9465, 2025

  28. [28]

    Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning

    Mayank Mittal, Pascal Roth, James Tigue, Antoine Richard, Octi Zhang, Peter Du, Antonio Serrano-Munoz, Xinjie Yao, René Zurbrügg, Nikita Rudin, et al. Isaac lab: A gpu-accelerated simulation framework for multi-modal robot learning.arXiv preprint arXiv:2511.04831, 2025

  29. [29]

    Vlnverse: A benchmark for vision-language navigation with versatile, embodied, realistic simulation and evaluation.arXiv preprint arXiv:2512.19021, 2025

    Sihao Lin, Zerui Li, Xunyi Zhao, Gengze Zhou, Liuyi Wang, Rong Wei, Rui Tang, Juncheng Li, Hanqing Wang, Jiangmiao Pang, et al. Vlnverse: A benchmark for vision-language navigation with versatile, embodied, realistic simulation and evaluation.arXiv preprint arXiv:2512.19021, 2025

  30. [30]

    Spoc: Imitating shortest paths in simulation enables effective navigation and manipulation in the real world

    Kiana Ehsani, Tanmay Gupta, Rose Hendrix, Jordi Salvador, Luca Weihs, Kuo-Hao Zeng, Ku- nal Pratap Singh, Yejin Kim, Winson Han, Alvaro Herrasti, et al. Spoc: Imitating shortest paths in simulation enables effective navigation and manipulation in the real world. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 162...

  31. [31]

    Molmospaces: A large-scale open ecosystem for robot navigation and manipulation.arXiv preprint arXiv:2602.11337, 2026

    Yejin Kim, Wilbert Pumacay, Omar Rayyan, Max Argus, Winson Han, Eli VanderBilt, Jordi Salvador, Abhay Deshpande, Rose Hendrix, Snehal Jauhri, et al. Molmospaces: A large-scale open ecosystem for robot navigation and manipulation.arXiv preprint arXiv:2602.11337, 2026

  32. [32]

    Holodeck: Language guided generation of 3d embodied ai environments

    Yue Yang, Fan-Yun Sun, Luca Weihs, Eli VanderBilt, Alvaro Herrasti, Winson Han, Jiajun Wu, Nick Haber, Ranjay Krishna, Lingjie Liu, et al. Holodeck: Language guided generation of 3d embodied ai environments. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16227–16237, 2024

  33. [33]

    RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

    Soroush Nasiriany, Abhiram Maddukuri, Lance Zhang, Adeet Parikh, Aaron Lo, Abhishek Joshi, Ajay Mandlekar, and Yuke Zhu. Robocasa: Large-scale simulation of everyday tasks for generalist robots.arXiv preprint arXiv:2406.02523, 2024

  34. [34]

    Nerf: Representing scenes as neural radiance fields for view synthesis

    Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoor- thi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021

  35. [35]

    3d gaussian splatting for real-time radiance field rendering.ACM Trans

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George Drettakis, et al. 3d gaussian splatting for real-time radiance field rendering.ACM Trans. Graph., 42(4):139–1, 2023

  36. [36]

    Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields

    Shijie Zhou, Haoran Chang, Sicheng Jiang, Zhiwen Fan, Zehao Zhu, Dejia Xu, Pradyumna Chari, Suya You, Zhangyang Wang, and Achuta Kadambi. Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21676–21685, 2024

  37. [37]

    Vid2sim: Realistic and interactive simulation from video for urban navigation

    Ziyang Xie, Zhizheng Liu, Zhenghao Peng, Wayne Wu, and Bolei Zhou. Vid2sim: Realistic and interactive simulation from video for urban navigation. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 1581–1591, 2025

  38. [38]

    Navgsim: High-fidelity gaussian splatting simulator for large-scale navigation.arXiv preprint arXiv:2603.15186, 2026

    Jiahang Liu, Yuanxing Duan, Jiazhao Zhang, Minghan Li, Shaoan Wang, Zhizheng Zhang, and He Wang. Navgsim: High-fidelity gaussian splatting simulator for large-scale navigation.arXiv preprint arXiv:2603.15186, 2026

  39. [39]

    Towards physically executable 3d gaussian for embodied navigation.arXiv preprint arXiv:2510.21307, 2025

    Bingchen Miao, Rong Wei, Zhiqi Ge, Shiqi Gao, Jingzhe Zhu, Renhan Wang, Siliang Tang, Jun Xiao, Rui Tang, Juncheng Li, et al. Towards physically executable 3d gaussian for embodied navigation.arXiv preprint arXiv:2510.21307, 2025

  40. [40]

    Embodiedsplat: Personalized real-to-sim-to-real navigation with gaussian splats from a mobile device

    Gunjan Chhablani, Xiaomeng Ye, Muhammad Zubair Irshad, and Zsolt Kira. Embodiedsplat: Personalized real-to-sim-to-real navigation with gaussian splats from a mobile device. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 25431– 25441, 2025. 12

  41. [41]

    Wanderland: Geometrically grounded simulation for open-world embodied ai.arXiv preprint arXiv:2511.20620, 2025

    Xinhao Liu, Jiaqi Li, Youming Deng, Ruxin Chen, Yingjia Zhang, Yifei Ma, Li Guo, Yiming Li, Jing Zhang, and Chen Feng. Wanderland: Geometrically grounded simulation for open-world embodied ai.arXiv preprint arXiv:2511.20620, 2025

  42. [42]

    ReaDy-Go: Real-to-Sim Dynamic 3D Gaussian Splatting Simulation for Environment-Specific Visual Navigation with Moving Obstacles

    Seungyeon Yoo, Youngseok Jang, Dabin Kim, Youngsoo Han, Seungwoo Jung, and H Jin Kim. Ready-go: Real-to-sim dynamic 3d gaussian splatting simulation for environment-specific visual navigation with moving obstacles.arXiv preprint arXiv:2602.11575, 2026

  43. [43]

    Gaussgym: An open-source real-to-sim framework for learning locomotion from pixels.arXiv preprint arXiv:2510.15352, 2025

    Alejandro Escontrela, Justin Kerr, Arthur Allshire, Jonas Frey, Rocky Duan, Carmelo Sferrazza, and Pieter Abbeel. Gaussgym: An open-source real-to-sim framework for learning locomotion from pixels.arXiv preprint arXiv:2510.15352, 2025

  44. [44]

    Floor flattening of image-based 3d recon- struction for mobile robots

    Seonghwan Sim, Yeji Kim, and Sung Soo Hwang. Floor flattening of image-based 3d recon- struction for mobile robots. In2026 IEEE International Conference on Artificial Intelligence and eXtended and Virtual Reality (AIxVR), pages 441–446. IEEE, 2026

  45. [45]

    pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction

    David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 19457–19467, 2024

  46. [46]

    Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images

    Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images. InEuropean conference on computer vision, pages 370–386. Springer, 2024

  47. [47]

    Freesplat: Generalizable 3d gaussian splatting towards free view synthesis of indoor scenes.Advances in Neural Information Processing Systems, 37:107326–107349, 2024

    Yunsong Wang, Tianxin Huang, Hanlin Chen, and Gim Hee Lee. Freesplat: Generalizable 3d gaussian splatting towards free view synthesis of indoor scenes.Advances in Neural Information Processing Systems, 37:107326–107349, 2024

  48. [48]

    Yonosplat: You only need one model for feedforward 3d gaussian splatting.arXiv preprint arXiv:2511.07321, 2025

    Botao Ye, Boqi Chen, Haofei Xu, Daniel Barath, and Marc Pollefeys. Yonosplat: You only need one model for feedforward 3d gaussian splatting.arXiv preprint arXiv:2511.07321, 2025

  49. [49]

    Omnisplat: Taming feed-forward 3d gaussian splatting for omnidirectional im- ages with editable capabilities

    Suyoung Lee, Jaeyoung Chung, Kihoon Kim, Jaeyoo Huh, Gunhee Lee, Minsoo Lee, and Kyoung Mu Lee. Omnisplat: Taming feed-forward 3d gaussian splatting for omnidirectional im- ages with editable capabilities. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 16356–16365, 2025

  50. [50]

    Splatter-360: Generalizable 360 gaussian splatting for wide-baseline panoramic images

    Zheng Chen, Chenming Wu, Zhelun Shen, Chen Zhao, Weicai Ye, Haocheng Feng, Errui Ding, and Song-Hai Zhang. Splatter-360: Generalizable 360 gaussian splatting for wide-baseline panoramic images. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 21590–21599, 2025

  51. [51]

    Long-LRM++: Preserving Fine Details in Feed-Forward Wide-Coverage Reconstruction

    Chen Ziwen, Hao Tan, Peng Wang, Zexiang Xu, and Li Fuxin. Long-lrm++: Preserving fine details in feed-forward wide-coverage reconstruction.arXiv preprint arXiv:2512.10267, 2025

  52. [52]

    Anysplat: Feed-forward 3d gaussian splatting from unconstrained views.ACM Transactions on Graphics (TOG), 44(6):1–16, 2025

    Lihan Jiang, Yucheng Mao, Linning Xu, Tao Lu, Kerui Ren, Yichen Jin, Xudong Xu, Mulin Yu, Jiangmiao Pang, Feng Zhao, et al. Anysplat: Feed-forward 3d gaussian splatting from unconstrained views.ACM Transactions on Graphics (TOG), 44(6):1–16, 2025

  53. [53]

    Imaginav: Scalable embodied navigation via generative visual prediction and inverse dynamics.arXiv preprint arXiv:2603.13833, 2026

    Jie Chen, Yuxin Cai, Yizhuo Wang, Ruofei Bai, Yuhong Cao, Jun Li, Yau Wei Yun, and Guillaume Sartoretti. Imaginav: Scalable embodied navigation via generative visual prediction and inverse dynamics.arXiv preprint arXiv:2603.13833, 2026

  54. [54]

    Soon: Scenario oriented object navigation with graph-based exploration

    Fengda Zhu, Xiwen Liang, Yi Zhu, Qizhi Yu, Xiaojun Chang, and Xiaodan Liang. Soon: Scenario oriented object navigation with graph-based exploration. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12689–12699, 2021

  55. [55]

    Speaker- follower models for vision-and-language navigation.Advances in neural information processing systems, 31, 2018

    Daniel Fried, Ronghang Hu, V olkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. Speaker- follower models for vision-and-language navigation.Advances in neural information processing systems, 31, 2018. 13

  56. [56]

    Towards learning a generic agent for vision-and-language navigation via pre-training

    Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin, and Jianfeng Gao. Towards learning a generic agent for vision-and-language navigation via pre-training. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 13137–13146, 2020

  57. [57]

    A new path: Scaling vision-and-language navigation with synthetic instructions and imitation learning

    Aishwarya Kamath, Peter Anderson, Su Wang, Jing Yu Koh, Alexander Ku, Austin Waters, Yinfei Yang, Jason Baldridge, and Zarana Parekh. A new path: Scaling vision-and-language navigation with synthetic instructions and imitation learning. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10813–10823, 2023

  58. [58]

    Learning from unlabeled 3d environments for vision-and-language navigation

    Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid, and Ivan Laptev. Learning from unlabeled 3d environments for vision-and-language navigation. InEuropean Conference on Computer Vision, pages 638–655. Springer, 2022

  59. [59]

    Scaling data generation in vision-and-language navigation

    Zun Wang, Jialu Li, Yicong Hong, Yi Wang, Qi Wu, Mohit Bansal, Stephen Gould, Hao Tan, and Yu Qiao. Scaling data generation in vision-and-language navigation. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 12009–12020, 2023

  60. [60]

    Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel

    Zun Wang, Jialu Li, Yicong Hong, Songze Li, Kunchang Li, Shoubin Yu, Yi Wang, Yu Qiao, Yali Wang, Mohit Bansal, et al. Bootstrapping language-guided navigation learning with self-refining data flywheel.arXiv preprint arXiv:2412.08467, 2024

  61. [61]

    NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM

    Zihan Wang, Yaohui Zhu, Gim Hee Lee, and Yachun Fan. Navrag: Generating user de- mand instructions for embodied navigation through retrieval-augmented llm.arXiv preprint arXiv:2502.11142, 2025

  62. [62]

    Learning goal-oriented language-guided navigation with self- improving demonstrations at scale.arXiv preprint arXiv:2509.24910, 2025

    Songze Li, Zun Wang, Gengze Zhou, Jialu Li, Xiangyu Zeng, Limin Wang, Yu Qiao, Qi Wu, Mohit Bansal, and Yi Wang. Learning goal-oriented language-guided navigation with self- improving demonstrations at scale.arXiv preprint arXiv:2509.24910, 2025

  63. [63]

    Embodied navigation foundation model.arXiv preprint arXiv:2509.12129, 2025

    Jiazhao Zhang, Anqi Li, Yunpeng Qi, Minghan Li, Jiahang Liu, Shaoan Wang, Haoran Liu, Gengze Zhou, Yuze Wu, Xingxing Li, et al. Embodied navigation foundation model.arXiv preprint arXiv:2509.12129, 2025

  64. [64]

    Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation

    Shuo Wang, Yucheng Wang, Guoxin Lian, Yongcai Wang, Maiyue Chen, Kaihui Wang, Bo Zhang, Zhizhong Su, Yutian Zhou, Wanting Li, et al. Progress-think: Semantic progress reasoning for vision-language navigation.arXiv preprint arXiv:2511.17097, 2025

  65. [65]

    NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models

    Gengze Zhou, Yicong Hong, Zun Wang, Xin Eric Wang, and Qi Wu. Navgpt-2: Un- leashing navigational reasoning capability for large vision-language models.arXiv preprint arXiv:2407.12366, 2024

  66. [66]

    $\pi^3$: Permutation-Equivariant Visual Geometry Learning

    Yifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang, Yang Zhou, Zizun Li, Junyi Chen, Jiangmiao Pang, Chunhua Shen, and Tong He. Pi3: Permutation-equivariant visual geometry learning.arXiv preprint arXiv:2507.13347, 2025

  67. [67]

    Depth Anything 3: Recovering the Visual Space from Any Views

    Haotong Lin, Sili Chen, Junhao Liew, Donny Y Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth anything 3: Recovering the visual space from any views.arXiv preprint arXiv:2511.10647, 2025

  68. [68]

    Semantic image synthesis with spatially-adaptive normalization

    Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. Semantic image synthesis with spatially-adaptive normalization. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2337–2346, 2019

  69. [69]

    Flow Matching for Generative Modeling

    Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling.arXiv preprint arXiv:2210.02747, 2022

  70. [70]

    Mean Flows for One-step Generative Modeling

    Zhengyang Geng, Mingyang Deng, Xingjian Bai, J Zico Kolter, and Kaiming He. Mean flows for one-step generative modeling.arXiv preprint arXiv:2505.13447, 2025

  71. [71]

    One-step Latent-free Image Generation with Pixel Mean Flows

    Yiyang Lu, Susie Lu, Qiao Sun, Hanhong Zhao, Zhicheng Jiang, Xianbang Wang, Tianhong Li, Zhengyang Geng, and Kaiming He. One-step latent-free image generation with pixel mean flows.arXiv preprint arXiv:2601.22158, 2026. 14

  72. [72]

    Emerging properties in self-supervised vision transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. InProceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021

  73. [73]

    The unrea- sonable effectiveness of deep features as a perceptual metric

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unrea- sonable effectiveness of deep features as a perceptual metric. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 586–595, 2018

  74. [74]

    The marathon 2: A navigation system

    Steve Macenski, Francisco Martín, Ruffin White, and Jonatan Ginés Clavero. The marathon 2: A navigation system. In2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 2718–2725. IEEE, 2020

  75. [75]

    From the desks of ros maintainers: A survey of modern & capable mobile robotics algorithms in the robot operating system 2.Robotics and Autonomous Systems, 168:104493, 2023

    Steve Macenski, Tom Moore, David V Lu, Alexey Merzlyakov, and Michael Ferguson. From the desks of ros maintainers: A survey of modern & capable mobile robotics algorithms in the robot operating system 2.Robotics and Autonomous Systems, 168:104493, 2023

  76. [76]

    Regulated pure pursuit for robot path tracking.Autonomous Robots, 47(6):685–694, 2023

    Steve Macenski, Shrijit Singh, Francisco Martín, and Jonatan Ginés. Regulated pure pursuit for robot path tracking.Autonomous Robots, 47(6):685–694, 2023

  77. [77]

    Realsee3d: A large-scale multi-view rgb-d dataset of indoor scenes (version 1.0), 2025

    Linyuan Li, Yan Wu, Xi Li, Lingli Wang, Tong Rao, Jie Zhou, Cihui Pan, and Xinchen Hui. Realsee3d: A large-scale multi-view rgb-d dataset of indoor scenes (version 1.0), 2025. URL https://doi.org/10.5281/zenodo.17826243

  78. [78]

    Structured3d: A large photo-realistic dataset for structured 3d modeling

    Jia Zheng, Junfei Zhang, Jing Li, Rui Tang, Shenghua Gao, and Zihan Zhou. Structured3d: A large photo-realistic dataset for structured 3d modeling. InProceedings of The European Conference on Computer Vision (ECCV), 2020

  79. [79]

    Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data

    Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Yuri Feigin, Peter Fu, Thomas Gebauer, Daniel Kurz, Tal Dimry, Brandon Joffe, Arik Schwartz, et al. Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data. InThirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1)

  80. [80]

    Scannet: Richly-annotated 3d reconstructions of indoor scenes

    Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017

Showing first 80 references.