Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Multimodal Perception for Goal-oriented Navigation: A Survey

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This survey tries to establish that goal-oriented navigation methods across four task paradigms can be organized by six inference domains — the main mechanism each agent uses to reason about its environment — and that this organization…

desk verdict Useful organizing lens, but the taxonomy is applied inconsistently and the cross-task insights inherit that unreliability. read the letter →

arxiv 2504.15643 v1 pith:4TTSWAC5 submitted 2025-04-22 cs.RO

classification cs.RO
keywords goal-orientednavigationinferencedomainsmultimodalperceptionembodiedAIobjectgoalaudio-visualimagesurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey tries to establish that the many methods for goal-oriented embodied navigation, despite differing goals and sensor setups, can be understood through one lens: the inference domain, or the main mechanism by which an agent reasons about its environment. It sorts roughly 200 methods into six such domains — latent maps, implicit representations, graph-based reasoning, linguistic reasoning, embedding-based matching, and diffusion-based generation — and then uses that sorting to show that the same computational foundations reappear across point-goal, object-goal, image-goal, and audio-goal navigation. If the taxonomy holds, differences between navigation tasks are less fundamental than usually presented: tasks differ in goal specification and sensory input, but the underlying reasoning machinery is shared. The payoff would be a principled basis for transferring ideas like map uncertainty, visual-language embeddings, or generative map completion from one navigation paradigm to another.

What carries the argument

The central object is the inference domain taxonomy: a classification of an agent's primary environmental reasoning mechanism, defined as the fundamental way the agent processes and uses information to navigate. The six domains — latent map-based, implicit representation, graph-based, linguistic, embedding-based, and diffusion model-based — are the categories into which the survey places each method. This taxonomy carries the argument by converting roughly 200 method descriptions into comparable categories, which then permits the cross-task patterns, historical progression, and strengths-and-limitations comparisons that the survey reports.

What would settle it

Take the methods listed in Tables 3 through 6 and have independent annotators assign each to exactly one inference domain using the Section 1.1 definitions; if a substantial fraction cannot be uniquely assigned — as already happens with NOMAD [37], which appears as an ObjectNav diffusion method in Section 7.1.6 yet as an ImageNav diffusion method in Table 5 and Section 5.3.1 — then the cross-task patterns derived from those assignments do not follow.

Watch

Extended reading notes

Core claim

The central claim is that goal-oriented navigation research can be organized by inference domains, six primary environmental reasoning mechanisms. Latent map-based methods construct explicit spatial or semantic maps; implicit representation methods learn end-to-end policies without maps; graph-based methods reason over relational structures; linguistic methods use large language models and common-sense knowledge; embedding-based methods use pretrained vision-language models for zero-shot goal matching; and diffusion-based methods generate maps or trajectories. Applying this lens across PointNav, ObjectNav, ImageNav, and AudioGoalNav, the paper argues, reveals recurring patterns: maps grow from geometric to semantic across tasks, implicit methods specialize through odometry versus visual correspondence, graphs shift from object-scene structures to topological ones, language becomes more valuable as semantic complexity increases, embeddings trade off pretrained knowledge against the sensory gap, and diffusion is concentrated in tasks with high semantic complexity and partial observability. The survey's contribution is this cross-task synthesis, not new experiments.

Load-bearing premise

The framework assumes every method has exactly one main environmental reasoning mechanism that places it in exactly one of the six inference domains, and that assumption is already strained by the paper's own classification of NOMAD in two different domains.

Editorial extensions

If this is right

  • If the taxonomy is right, techniques validated in one navigation task can be transferred to another when the methods share an inference domain, such as frontier-based exploration moving from PointNav to ObjectNav or ImageNav.
  • Diffusion-model-based navigation is predicted to be most valuable in settings with high semantic complexity and partial observability, which should concentrate future generative-model work in ObjectNav and semantic audio-visual navigation rather than in purely geometric PointNav.
  • Embedding-based zero-shot methods are expected to work best when the sensory gap between pretraining and navigation is small, so visual-language embeddings like CLIP should transfer well to ObjectNav but need task-specific adaptation for audio-goal navigation.
  • The historical progression from explicit maps to implicit representations appears as a genuine trend across all four tasks, not an artifact of one benchmark or one method family.
  • Language-based reasoning should grow in importance as the semantic complexity of a navigation task rises, keeping PointNav largely non-linguistic while making LLM-based reasoning central to ObjectNav and AudioGoalNav.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy could be turned into a predictive design rule: given a new navigation task, estimate its semantic complexity and degree of partial observability, then choose an inference domain accordingly; the survey documents the pattern but does not state this rule.
  • A testable extension follows from the shared-foundation claim: a component such as a latent map module or a CLIP-based goal scorer trained on one task should transfer to another task with little fine-tuning, which the survey does not directly test.
  • The paper's own placement of NOMAD in both the ObjectNav diffusion discussion and the ImageNav diffusion table suggests that single-domain assignment may be too rigid; a multi-label or probabilistic assignment could make the taxonomy more faithful.
  • A systematic re-annotation of all surveyed methods by independent readers, measuring agreement on domain assignment, would convert the taxonomy from an editorial claim into a measurable classification scheme.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This survey reviews approximately 200 papers on goal-oriented embodied navigation, covering PointNav, ObjectNav, ImageNav, and AudioGoalNav, and organizes the reviewed methods according to six "inference domains": latent map, implicit representation, graph, linguistic, embedding, and diffusion. It provides formal task definitions, dataset and simulator comparisons, evaluation metrics, per-task method tables, and a discussion section that derives cross-task insights from the inference-domain assignments. The paper's central claim is that the inference-domain lens reveals recurring computational patterns across navigation paradigms and explains how shared foundations support diverse tasks.

Significance. If the taxonomy is applied consistently, this survey would be a useful organizing resource for the embodied-AI navigation community. It assembles a broad and current bibliography, gives a compact comparison of datasets and simulators, and presents a readable narrative from explicit map-based methods to diffusion-based generative approaches. The cross-task observations in Section 7.1 are interesting and potentially valuable. However, the paper's main contribution is precisely the taxonomy, so the consistency of the domain assignments is load-bearing. The internal inconsistencies described below currently prevent the central claim from being fully supported. The paper does not provide machine-checked proofs or code, but for a survey this is not required.

major comments (3)
  1. [Section 7.1.6, Table 5, Section 5.3.1] NOMAD [37] is presented as an ImageNav diffusion method in Table 5 and Section 5.3.1, but Section 7.1.6 cites it as an ObjectNav diffusion approach, stating that the diffusion domain is "most developed in ObjectNav with approaches like NOMAD [37] and DAR [109]." This is a direct internal contradiction in the assignments on which the cross-task analysis rests, and it demonstrates that the single-domain assignment rule is not being applied reproducibly. Please correct the classification or state a rule that makes both statements true.
  2. [Section 4, Table 4] Section 4.1 defines three ObjectNav categories (Modular-Based, End-To-End, Zero-Shot) and refers to Table 4 as showing them, but Table 4 organizes ObjectNav methods by inference domain, with rows for Latent Map, Graph, Implicit Representation, Linguistic, CLIP/BLIP Embeddings, and Diffusion Learning. Sections 4.2 through 4.4 follow the three architecture categories rather than the inference-domain organization promised in Section 1.1. This structural mismatch makes it unclear whether ObjectNav methods are primarily organized by architecture type or by inference domain, and it weakens the survey's unifying-lens claim.
  3. [Section 1.1] The survey does not provide an operational rule for determining a method's "main environmental reasoning mechanism." The cross-task insights in Section 7.1 are derived from assigning each method to exactly one domain, so the absence of a decision rule is load-bearing. The NOMAD contradiction noted above shows that assignments are not currently reproducible. Please provide a concrete decision rule, for example based on which component produces the long-horizon planning decision, and re-verify all table entries against it.
minor comments (5)
  1. [Section 2.3.5, Table 2] The text states that "AI2-THOR offers photorealistic environments based on 3D scans of real-world spaces," but Table 2 classifies AI2-THOR as "Near photorealistic (synthetic)." This appears to be a factual error; AI2-THOR is a synthetic environment, not a 3D-scan-based one.
  2. [Section 2.4.4, Table 6] The metric SWS is defined in Section 2.4.4 as "Success rate when silent," but Table 6's footnote defines SWS as "Success weighted by Time Steps." Please align the definition and the table notation.
  3. [Section 4.4.2] The heading "Linguistic Inference Dominance" appears to be a typo for "Linguistic Inference Domain."
  4. [References] References [124] and [125] contain misspelled author names, "Waswani" and "Alexey," which should be corrected to "Vaswani" and "Dosovitskiy," respectively.
  5. [Table 4] The Architecture Type column of Table 4 mixes the category names Modular/End-To-End/Zero-Shot with inference-domain rows, which adds to the Section 4/Table 4 ambiguity; a note explaining the relationship between the two axes would help.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's inference-domain taxonomy is an organizing framework, not a derivation from fitted data or self-citations.

full rationale

This survey makes no quantitative predictions and fits no parameters; its claimed contribution is an organizing taxonomy (inference domains) applied to roughly 200 cited papers. For a survey, circularity would require the taxonomy's findings to be logically equivalent to the taxonomy's definitions, or the citations to be self-referential and load-bearing. Neither holds. The six domains in Section 1.1 are defined by architectural families (latent maps, implicit representations, graphs, linguistic, embeddings, diffusion), and the cross-task observations in Section 7.1 are descriptive summaries of which papers the authors placed in each cell. Those observations are contingent: they could be wrong if the assignments were wrong, and indeed the paper contains an internal labeling inconsistency (NOMAD appears as an ImageNav diffusion method in Table 5 and Section 5.3.1 but is cited as an ObjectNav diffusion approach in Section 7.1.6). That inconsistency is a correctness and consistency concern, not a circular reduction: the claimed pattern is not forced by the taxonomy's definitions in the way a fitted parameter is forced by the data it was fit to. No self-citation chain or imported uniqueness theorem is used to justify the framework; the citations at Section 1.1 are to external method papers. The survey is therefore self-contained as a literature review, and the circularity score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 1 invented entities

The survey's conclusions rest on the completeness and correctness of its taxonomy and on the accuracy of its reading of roughly 200 cited papers. No quantitative model, fitted parameter, or formal derivation is involved.

assumptions (2)
  • domain assumption The six inference domains (latent map, implicit, graph, linguistic, embedding, diffusion) are an exhaustive and faithful partition of multimodal navigation methods.
    Introduced in Section 1.1; the survey assigns every reviewed method to one domain, but provides no independent criterion or validation for the assignment.
  • domain assumption The approximately 200 cited articles are representative and sufficient for the survey's conclusions.
    Stated in the abstract; no search protocol, inclusion criteria, or comparison with other surveys is given.
invented entities (1)
  • Inference domains
    purpose: Conceptual taxonomy used to classify navigation methods and derive cross-task insights.
    Introduced by the authors; no external falsifiable prediction; its value is heuristic and depends on whether the classification helps readers.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multimodal Perception for Goal-oriented Navigation: A Survey." pith.science (2026). https://pith.science/paper/4TTSWAC5

@misc{pith2026250415643,
  author       = {Pith},
  title        = {Pith review of: Multimodal Perception for Goal-oriented Navigation: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4TTSWAC5}},
  note         = {Machine review of arXiv:2504.15643}
}
read the original abstract

Goal-oriented navigation presents a fundamental challenge for autonomous systems, requiring agents to navigate complex environments to reach designated targets. This survey offers a comprehensive analysis of multimodal navigation approaches through the unifying perspective of inference domains, exploring how agents perceive, reason about, and navigate environments using visual, linguistic, and acoustic information. Our key contributions include organizing navigation methods based on their primary environmental reasoning mechanisms across inference domains; systematically analyzing how shared computational foundations support seemingly disparate approaches across different navigation tasks; identifying recurring patterns and distinctive strengths across various navigation paradigms; and examining the integration challenges and opportunities of multimodal perception to enhance navigation capabilities. In addition, we review approximately 200 relevant articles to provide an in-depth understanding of the current landscape.

Figures

Figures reproduced from arXiv: 2504.15643 by the authors.

Figure 1
Figure 1. Timeline of the historical development of navigation tasks and their representative approaches. Different colors [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Implicit Representation Learning Inference Domain: [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Latent Map Based Inference Domain: This domain [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Graph Based Inference Domain: This domain con [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 6
Figure 6. Figure 6: Diffusion Model Based Inference Domain: This do [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs

    cs.RO 2025-12 conditional novelty 6.0 of 10

    VL-LN Bench turns instance-goal navigation into an interactive dialog task, contributes a 41k-trajectory house-scale benchmark with a GPT-4o oracle, and shows active questioning improves embodied agents' success.

Reference graph

Works this paper leans on

198 extracted references · 34 canonical work pages · cited by 1 Pith paper

  1. [37]

    Nomad: Goal masked diffusion policies for navigation and exploration,

    A. Sridhar, D. Shah, C. Glossop, and S. Levine, “Nomad: Goal masked diffusion policies for navigation and exploration,” in 2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2024, pp. 63–70. 1, 12, 14

  2. [109]

    Diffusion as reasoning: Enhancing object goal navigation with llm-biased diffusion model,

    Y. Ji, Y. Liu, Z. Wang, B. Ma, Z. Xie, and H. Liu, “Diffusion as reasoning: Enhancing object goal navigation with llm-biased diffusion model,” arXiv preprint arXiv:2410.21842, 2024. 7, 10, 14

  3. [1]

    A frontier-based approach for autonomous explo- ration,

    B. Yamauchi, “A frontier-based approach for autonomous explo- ration,” in Proceedings 1997 IEEE International Symposium on Com- putational Intelligence in Robotics and Automation CIRA’97.’Towards New Computational Principles for Robotics and Automation’ . IEEE, 1997, pp. 146–151. 1, 8, 9, 10, 11, 13

  4. [2]

    A fast marching level set method for monoton- ically advancing fronts

    J. A. Sethian, “A fast marching level set method for monoton- ically advancing fronts.” proceedings of the National Academy of Sciences, vol. 93, no. 4, pp. 1591–1595, 1996. 1, 8

  5. [3]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778. 1, 6, 8, 10, 11, 12, 13

  6. [4]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017. 1, 6, 11

  7. [5]

    Language models are few-shot learners,

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P . Dhari- wal, A. Neelakantan, P . Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020. 1, 9, 10, 11

  8. [6]

    Gpt-4 technical report,

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 ,

Show all 198 references
  1. [7]

    Soundspaces: Audio- visual navigation in 3d environments,

    C. Chen, U. Jain, C. Schissler, S. V . A. Gari, Z. Al-Halah, V . K. Ithapu, P . Robinson, and K. Grauman, “Soundspaces: Audio- visual navigation in 3d environments,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 1...

  2. [8]

    Learning representa- tions from audio-visual spatial alignment,

    P . Morgado, Y. Li, and N. Nvasconcelos, “Learning representa- tions from audio-visual spatial alignment,” Advances in Neural Information Processing Systems, vol. 33, pp. 4733–4744, 2020. 1, 13 SUBMITTED IEEE TRANSACTIONS ON P A TTERN ANAL YSIS AND MACHINE INTELLIGENCE 16

  3. [9]

    On evaluation of embodied navigation agents,

    P . Anderson, A. Chang, D. S. Chaplot, A. Dosovitskiy, S. Gupta, V . Koltun, J. Kosecka, J. Malik, R. Mottaghi, M. Savva et al. , “On evaluation of embodied navigation agents,” arXiv preprint arXiv:1807.06757, 2018. 1, 4, 6

  4. [10]

    Objectnav revisited: On evaluation of embodied agents navigating to objects,

    D. Batra, A. Gokaslan, A. Kembhavi, O. Maksymets, R. Mottaghi, M. Savva, A. Toshev, and E. Wijmans, “Objectnav revisited: On evaluation of embodied agents navigating to objects,” arXiv preprint arXiv:2006.13171, 2020. 1, 6, 11

  5. [11]

    Learning to explore using active neural slam,

    D. S. Chaplot, D. Gandhi, S. Gupta, A. Gupta, and R. Salakhut- dinov, “Learning to explore using active neural slam,” arXiv preprint arXiv:2004.05155, 2020. 1, 5, 6, 8, 11, 14

  6. [12]

    Hm3d-ovon: A dataset and benchmark for open-vocabulary object goal navigation,

    N. Yokoyama, R. Ramrakhya, A. Das, D. Batra, and S. Ha, “Hm3d-ovon: A dataset and benchmark for open-vocabulary object goal navigation,” arXiv preprint arXiv:2409.14296 , 2024. 1, 2, 7, 8

  7. [13]

    Pirlnav: Pretraining with imitation and rl finetuning for objectnav,

    R. Ramrakhya, D. Batra, E. Wijmans, and A. Das, “Pirlnav: Pretraining with imitation and rl finetuning for objectnav,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 17 896–17 906. 1, 7, 8

  8. [14]

    Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames,

    E. Wijmans, A. Kadian, A. Morcos, S. Lee, I. Essa, D. Parikh, M. Savva, and D. Batra, “Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames,” arXiv preprint arXiv:1911.00357, 2019. 1, 2, 5, 6

  9. [15]

    The sur- prising effectiveness of visual odometry techniques for embodied pointgoal navigation,

    X. Zhao, H. Agrawal, D. Batra, and A. G. Schwing, “The sur- prising effectiveness of visual odometry techniques for embodied pointgoal navigation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 16 127–16 136. 2, 5, 14

  10. [16]

    Is mapping necessary for realistic pointgoal navigation?

    R. Partsey, E. Wijmans, N. Yokoyama, O. Dobosevych, D. Batra, and O. Maksymets, “Is mapping necessary for realistic pointgoal navigation?” in Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , 2022, pp. 17 232–17 241. 2, 5

  11. [17]

    Entl: Embodied navi- gation trajectory learner,

    K. Kotar, A. Walsman, and R. Mottaghi, “Entl: Embodied navi- gation trajectory learner,” in Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, 2023, pp. 10 863–10 872. 2, 5, 6

  12. [18]

    Mpvo: Motion-prior based visual odometry for pointgoal navigation,

    S. Paul, B. Bhowmick et al. , “Mpvo: Motion-prior based visual odometry for pointgoal navigation,” arXiv preprint arXiv:2411.04796, 2024. 2, 5, 14

  13. [19]

    Object goal navigation using goal-oriented semantic explo- ration,

    D. S. Chaplot, D. P . Gandhi, A. Gupta, and R. R. Salakhutdi- nov, “Object goal navigation using goal-oriented semantic explo- ration,” Advances in Neural Information Processing Systems, vol. 33, pp. 4247–4258, 2020. 2, 7, 8

  14. [20]

    Ion: Instance-level object navigation,

    W. Li, X. Song, Y. Bai, S. Zhang, and S. Jiang, “Ion: Instance-level object navigation,” in Proceedings of the 29th ACM International Conference on Multimedia , ser. MM ’21. New York, NY, USA: Association for Computing Machinery, 2021, p. 4343–4352. [Online]. Available: https:...

  15. [21]

    Simple but effective: Clip embeddings for embodied ai,

    A. Khandelwal, L. Weihs, R. Mottaghi, and A. Kembhavi, “Simple but effective: Clip embeddings for embodied ai,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition, 2022, pp. 14 829–14 838. 2, 7, 10, 14

  16. [22]

    Habitat- web: Learning embodied object-search strategies from human demonstrations at scale,

    R. Ramrakhya, E. Undersander, D. Batra, and A. Das, “Habitat- web: Learning embodied object-search strategies from human demonstrations at scale,” in Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition , 2022, pp. 5173–

  17. [23]

    Zson: Zero-shot object-goal navigation using multimodal goal embeddings,

    A. Majumdar, G. Aggarwal, B. Devnani, J. Hoffman, and D. Batra, “Zson: Zero-shot object-goal navigation using multimodal goal embeddings,” Advances in Neural Information Processing Systems , vol. 35, pp. 32 340–32 352, 2022. 2, 7, 11, 14

  18. [24]

    L3mvn: Leveraging large language models for visual target navigation,

    B. Yu, H. Kasaei, and M. Cao, “L3mvn: Leveraging large language models for visual target navigation,” in 2023 IEEE/RSJ Interna- tional Conference on Intelligent Robots and Systems (IROS) . IEEE, 2023, pp. 3554–3560. 2, 7, 9, 14

  19. [25]

    Object-goal visual navigation via effective exploration of relations among historical navigation states,

    H. Du, L. Li, Z. Huang, and X. Yu, “Object-goal visual navigation via effective exploration of relations among historical navigation states,” 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 2563–2573, 2023. [Online]. Available: https://api.sema...

  20. [26]

    Sg-nav: Online 3d scene graph prompting for llm-based zero-shot object navigation,

    H. Yin, X. Xu, Z. Wu, J. Zhou, and J. Lu, “Sg-nav: Online 3d scene graph prompting for llm-based zero-shot object navigation,” arXiv preprint arXiv:2410.08189, 2024. 2, 7, 11

  21. [27]

    Trajectory dif- fusion for objectgoal navigation,

    X. Yu, S. Zhang, X. Song, X. Qin, and S. Jiang, “Trajectory dif- fusion for objectgoal navigation,” Advances in Neural Information Processing Systems, vol. 37, pp. 110 388–110 411, 2024. 2, 7, 10

  22. [28]

    Topo- logical semantic graph memory for image-goal navigation,

    N. Kim, O. Kwon, H. Yoo, Y. Choi, J. Park, and S. Oh, “Topo- logical semantic graph memory for image-goal navigation,” in Conference on Robot Learning. PMLR, 2023, pp. 393–402. 1, 2, 12, 14

  23. [29]

    Fgprompt: Fine-grained goal prompting for image-goal navigation,

    X. Sun, P . Chen, J. Fan, J. Chen, T. Li, and M. Tan, “Fgprompt: Fine-grained goal prompting for image-goal navigation,” Ad- vances in Neural Information Processing Systems, vol. 36, pp. 12 054– 12 073, 2023. 2, 11, 12, 14

  24. [30]

    Memonav: Working memory model for visual navigation,

    H. Li, Z. Wang, X. Yang, Y. Yang, S. Mei, and Z. Zhang, “Memonav: Working memory model for visual navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 17 913–17 922. 2, 11, 12, 14

  25. [31]

    Object instance retrieval in assistive robotics: Leveraging fine-tuned simsiam with multi-view images based on 3d semantic map,

    T. Sakaguchi, A. Taniguchi, Y. Hagiwara, L. El Hafi, S. Hasegawa, and T. Taniguchi, “Object instance retrieval in assistive robotics: Leveraging fine-tuned simsiam with multi-view images based on 3d semantic map,” in 2024 IEEE/RSJ International Conference on Intelligent Robots...

  26. [32]

    Look, listen, and act: Towards audio-visual embodied navigation,

    C. Gan, Y. Zhang, J. Wu, B. Gong, and J. B. Tenenbaum, “Look, listen, and act: Towards audio-visual embodied navigation,” in 2020 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2020, pp. 9701–9707. 2, 12, 13, 14

  27. [33]

    Semantic audio-visual navigation,

    C. Chen, Z. Al-Halah, and K. Grauman, “Semantic audio-visual navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 15 516–15 525. 2, 13, 14

  28. [34]

    Avlen: Audio- visual-language embodied navigation in 3d environments,

    S. Paul, A. Roy-Chowdhury, and A. Cherian, “Avlen: Audio- visual-language embodied navigation in 3d environments,” Ad- vances in Neural Information Processing Systems , vol. 35, pp. 6236– 6249, 2022. 2, 13

  29. [35]

    Catch me if you hear me: Audio-visual navigation in complex unmapped environments with moving sounds,

    A. Younes, D. Honerkamp, T. Welschehold, and A. Valada, “Catch me if you hear me: Audio-visual navigation in complex unmapped environments with moving sounds,” IEEE Robotics and Automation Letters, vol. 8, no. 2, pp. 928–935, 2023. 2, 4, 13

  30. [36]

    Rila: Reflective and imaginative language agent for zero-shot semantic audio-visual navigation,

    Z. Yang, J. Liu, P . Chen, A. Cherian, T. K. Marks, J. Le Roux, and C. Gan, “Rila: Reflective and imaginative language agent for zero-shot semantic audio-visual navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 16 251...

  31. [38]

    Hartley and A

    R. Hartley and A. Zisserman, Multiple view geometry in computer vision. Cambridge university press, 2003. 1, 9

  32. [39]

    Cognitive mapping and planning for visual navigation,

    S. Gupta, J. Davidson, S. Levine, R. Sukthankar, and J. Malik, “Cognitive mapping and planning for visual navigation,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2616–2625. 1, 6, 13

  33. [40]

    Poni: Potential functions for objectgoal navigation with interaction-free learning,

    S. K. Ramakrishnan, D. S. Chaplot, Z. Al-Halah, J. Malik, and K. Grauman, “Poni: Potential functions for objectgoal navigation with interaction-free learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 18 890–18 900. 1, 4,...

  34. [41]

    Vtnet: Visual transformer network for object goal navigation,

    H. Du, X. Yu, and L. Zheng, “Vtnet: Visual transformer network for object goal navigation,” arXiv preprint arXiv:2105.09447, 2021. 1, 7, 8, 14

  35. [42]

    Learning phrase rep- resentations using rnn encoder-decoder for statistical machine translation,

    K. Cho, B. Van Merri ¨enboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase rep- resentations using rnn encoder-decoder for statistical machine translation,” arXiv preprint arXiv:1406.1078, 2014. 1, 10, 12

  36. [43]

    Long short-term memory,

    J. Schmidhuber, S. Hochreiter et al., “Long short-term memory,” Neural Comput, vol. 9, no. 8, pp. 1735–1780, 1997. 1, 8

  37. [45]

    Image retrieval using scene graphs,

    J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. Shamma, M. Bern- stein, and L. Fei-Fei, “Image retrieval using scene graphs,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3668–3678. 1

  38. [46]

    Graph r-cnn for scene graph generation,

    J. Yang, J. Lu, S. Lee, D. Batra, and D. Parikh, “Graph r-cnn for scene graph generation,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 670–685. 1

  39. [47]

    Semi-supervised classification with graph convolutional networks,

    T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” arXiv preprint arXiv:1609.02907 ,

  40. [48]

    The graph neural network model,

    F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Mon- SUBMITTED IEEE TRANSACTIONS ON P A TTERN ANAL YSIS AND MACHINE INTELLIGENCE 17 fardini, “The graph neural network model,” IEEE transactions on neural networks, vol. 20, no. 1, pp. 61–80, 2008. 1

  41. [49]

    Learning hierarchical re- lationships for object-goal navigation,

    A. Pal, Y. Qiu, and H. Christensen, “Learning hierarchical re- lationships for object-goal navigation,” in Conference on Robot Learning. PMLR, 2021, pp. 517–528. 1, 7, 8, 9

  42. [50]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agar- wal, G. Sastry, A. Askell, P . Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763. 1, 9, ...

  43. [51]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P . Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems, vol. 33, pp. 6840–6851, 2020. 1, 10

  44. [52]

    Score-based generative modeling through stochastic differential equations,

    Y. Song, J. Sohl-Dickstein, D. P . Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” arXiv preprint arXiv:2011.13456 , 2020. 1, 10

  45. [53]

    Deep unsupervised learning using nonequilibrium thermody- namics,

    J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermody- namics,” in International conference on machine learning . pmlr, 2015, pp. 2256–2265. 1, 10

  46. [54]

    Photorealistic text-to-image diffusion models with deep language understanding,

    C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans et al. , “Photorealistic text-to-image diffusion models with deep language understanding,” Advances in neural information process- ing systems, vol. 35...

  47. [55]

    Dreamfusion: Text-to-3d using 2d diffusion,

    B. Poole, A. Jain, J. T. Barron, and B. Mildenhall, “Dreamfusion: Text-to-3d using 2d diffusion,” arXiv preprint arXiv:2209.14988 ,

  48. [56]

    Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai,

    S. K. Ramakrishnan, A. Gokaslan, E. Wijmans, O. Maksymets, A. Clegg, J. Turner, E. Undersander, W. Galuba, A. Westbury, A. X. Chang et al., “Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai,” arXiv preprint arXiv:2109.08238, 2021. 3, 4

  49. [57]

    Gibson env: Real-world perception for embodied agents,

    F. Xia, A. R. Zamir, Z. He, A. Sax, J. Malik, and S. Savarese, “Gibson env: Real-world perception for embodied agents,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 9068–9079. 3

  50. [58]

    Matterport3d: Learning from rgb-d data in indoor environments,

    A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y. Zhang, “Matterport3d: Learning from rgb-d data in indoor environments,”arXiv preprint arXiv:1709.06158, 2017. 3

  51. [59]

    Ai2- thor: An interactive 3d environment for visual ai,

    E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Her- rasti, M. Deitke, K. Ehsani, D. Gordon, Y. Zhu et al. , “Ai2- thor: An interactive 3d environment for visual ai,” arXiv preprint arXiv:1712.05474, 2017. 3, 4

  52. [60]

    Robothor: An open simulation-to-real embodied ai plat- form,

    M. Deitke, W. Han, A. Herrasti, A. Kembhavi, E. Kolve, R. Mot- taghi, J. Salvador, D. Schwenk, E. VanderBilt, M. Wallingford et al. , “Robothor: An open simulation-to-real embodied ai plat- form,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recogni...

  53. [61]

    Procthor: Large-scale embodied ai using procedural genera- tion,

    M. Deitke, E. VanderBilt, A. Herrasti, L. Weihs, K. Ehsani, J. Salvador, W. Han, E. Kolve, A. Kembhavi, and R. Mottaghi, “Procthor: Large-scale embodied ai using procedural genera- tion,” Advances in Neural Information Processing Systems , vol. 35, pp. 5982–5994, 2022. 3, 15

  54. [62]

    Scannet: Richly-annotated 3d reconstructions of indoor scenes,

    A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 5828–5839. 3, 8

  55. [63]

    The replica dataset: A digital replica of indoor spaces,

    J. Straub, T. Whelan, L. Ma, Y. Chen, E. Wijmans, S. Green, J. J. En- gel, R. Mur-Artal, C. Ren, S. Verma et al., “The replica dataset: A digital replica of indoor spaces,” arXiv preprint arXiv:1906.05797,

  56. [64]

    Soundspaces 2.0: A simulation platform for visual-acoustic learning,

    C. Chen, C. Schissler, S. Garg, P . Kobernik, A. Clegg, P . Calamia, D. Batra, P . Robinson, and K. Grauman, “Soundspaces 2.0: A simulation platform for visual-acoustic learning,” Advances in Neural Information Processing Systems, vol. 35, pp. 8896–8911, 2022. 4, 13

  57. [65]

    Habitat: A platform for embodied ai research,

    M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V . Koltun, J. Maliket al., “Habitat: A platform for embodied ai research,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 9339–9347. 4, 6, 9

  58. [66]

    3d scene graph: A structure for unified semantics, 3d space, and camera,

    I. Armeni, Z.-Y. He, J. Gwak, A. R. Zamir, M. Fischer, J. Malik, and S. Savarese, “3d scene graph: A structure for unified semantics, 3d space, and camera,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 5664–5673. 4

  59. [67]

    Comparison of model-free and model-based learning-informed planning for pointgoal navigation,

    Y. Li, A. Debnath, G. J. Stein, and J. Kosecka, “Comparison of model-free and model-based learning-informed planning for pointgoal navigation,” arXiv preprint arXiv:2212.08801, 2022. 5

  60. [68]

    Uncertainty-driven planner for exploration and navigation,

    G. Georgakis, B. Bucher, A. Arapin, K. Schmeckpeper, N. Matni, and K. Daniilidis, “Uncertainty-driven planner for exploration and navigation,” in 2022 International Conference on Robotics and Automation (ICRA). IEEE, 2022, pp. 11 295–11 302. 5

  61. [69]

    Moda: Map style trans- fer for self-supervised domain adaptation of embodied agents,

    E. S. Lee, J. Kim, S. Park, and Y. M. Kim, “Moda: Map style trans- fer for self-supervised domain adaptation of embodied agents,” in European Conference on Computer Vision . Springer, 2022, pp. 338–354. 5

  62. [70]

    Splitnet: Sim2sim and task2task transfer for embodied visual navigation,

    D. Gordon, A. Kadian, D. Parikh, J. Hoffman, and D. Batra, “Splitnet: Sim2sim and task2task transfer for embodied visual navigation,” in Proceedings of the IEEE/CVF International Confer- ence on Computer Vision, 2019, pp. 1022–1031. 5

  63. [71]

    Allenact: A framework for embodied ai research,

    L. Weihs, J. Salvador, K. Kotar, U. Jain, K.-H. Zeng, R. Mottaghi, and A. Kembhavi, “Allenact: A framework for embodied ai research,” arXiv preprint arXiv:2008.12760, 2020. 5, 6

  64. [72]

    Auxiliary tasks speed up learning point goal navigation,

    J. Ye, D. Batra, E. Wijmans, and A. Das, “Auxiliary tasks speed up learning point goal navigation,” in Conference on Robot Learning . PMLR, 2021, pp. 498–516. 5, 6

  65. [73]

    Auxiliary tasks for efficient learning of point-goal navigation,

    S. S. Desai and S. Lee, “Auxiliary tasks for efficient learning of point-goal navigation,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021, pp. 717–725. 5, 6

  66. [74]

    Integrating egocentric localization for more realistic point-goal navigation agents,

    S. Datta, O. Maksymets, J. Hoffman, S. Lee, D. Batra, and D. Parikh, “Integrating egocentric localization for more realistic point-goal navigation agents,” in Conference on Robot Learning . PMLR, 2021, pp. 313–328. 5, 6

  67. [75]

    Robustnav: Towards benchmarking robustness in embodied navigation,

    P . Chattopadhyay, J. Hoffman, R. Mottaghi, and A. Kembhavi, “Robustnav: Towards benchmarking robustness in embodied navigation,” in Proceedings of the IEEE/CVF International Confer- ence on Computer Vision, 2021, pp. 15 691–15 700. 5, 6

  68. [76]

    What do navigation agents learn about their environment?

    K. Dwivedi, G. Roig, A. Kembhavi, and R. Mottaghi, “What do navigation agents learn about their environment?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition, 2022, pp. 10 276–10 285. 5, 6

  69. [77]

    Vision-and-language navigation: Interpreting visually- grounded navigation instructions in real environments,

    P . Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. S ¨underhauf, I. Reid, S. Gould, and A. Van Den Hen- gel, “Vision-and-language navigation: Interpreting visually- grounded navigation instructions in real environments,” in Pro- ceedings of the IEEE conference on computer...

  70. [78]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P . Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, pro- ceedings, part III 18...

  71. [79]

    Simple and scalable predictive uncertainty estimation using deep ensem- bles,

    B. Lakshminarayanan, A. Pritzel, and C. Blundell, “Simple and scalable predictive uncertainty estimation using deep ensem- bles,” Advances in neural information processing systems , vol. 30,

  72. [80]

    Unpaired image-to- image translation using cycle-consistent adversarial networks,

    J.-Y. Zhu, T. Park, P . Isola, and A. A. Efros, “Unpaired image-to- image translation using cycle-consistent adversarial networks,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 2223–2232. 5

  73. [81]

    Real-time object navigation with deep neural networks and hierarchical reinforcement learning,

    A. Staroverov, D. A. Yudin, I. Belkin, V . Adeshkin, Y. K. Solo- mentsev, and A. I. Panov, “Real-time object navigation with deep neural networks and hierarchical reinforcement learning,” IEEE Access, vol. 8, pp. 195 608–195 621, 2020. 7, 9

  74. [82]

    Learning to map for active semantic goal navigation,

    G. Georgakis, B. Bucher, K. Schmeckpeper, S. Singh, and K. Dani- ilidis, “Learning to map for active semantic goal navigation,” arXiv preprint arXiv:2106.15648, 2021. 7, 8

  75. [83]

    Navigating to objects in unseen environments by distance prediction,

    M. Zhu, B. Zhao, and T. Kong, “Navigating to objects in unseen environments by distance prediction,” in 2022 IEEE/RSJ Interna- tional Conference on Intelligent Robots and Systems (IROS) . IEEE, 2022, pp. 10 571–10 578. 7, 8

  76. [84]

    Peanut: Predicting and navigating to unseen targets,

    A. J. Zhai and S. Wang, “Peanut: Predicting and navigating to unseen targets,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2023, pp. 10 926– 10 935. 7, 8, 14

  77. [85]

    Renderable neural radiance map for visual navigation,

    O. Kwon, J. Park, and S. Oh, “Renderable neural radiance map for visual navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 9099–9108. 7, 8

  78. [86]

    3d-aware object goal navigation via simultaneous exploration SUBMITTED IEEE TRANSACTIONS ON P A TTERN ANAL YSIS AND MACHINE INTELLIGENCE 18 and identification,

    J. Zhang, L. Dai, F. Meng, Q. Fan, X. Chen, K. Xu, and H. Wang, “3d-aware object goal navigation via simultaneous exploration SUBMITTED IEEE TRANSACTIONS ON P A TTERN ANAL YSIS AND MACHINE INTELLIGENCE 18 and identification,” in Proceedings of the IEEE/CVF Conference on Comput...

  79. [87]

    Self-supervised object goal navigation with in-situ finetuning,

    S. Y. Min, Y.-H. H. Tsai, W. Ding, A. Farhadi, R. Salakhutdinov, Y. Bisk, and J. Zhang, “Self-supervised object goal navigation with in-situ finetuning,” in 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2023, pp. 7119–7126. 7, 9

  80. [88]

    Skill fusion in hybrid robotic framework for visual object goal naviga- tion,

    A. Staroverov, K. Muravyev, K. Yakovlev, and A. I. Panov, “Skill fusion in hybrid robotic framework for visual object goal naviga- tion,” Robotics, vol. 12, no. 4, p. 104, 2023. 7

  81. [89]

    How to not train your dragon: Training-free embodied object goal navigation with semantic frontiers,

    J. Chen, G. Li, S. Kumar, B. Ghanem, and F. Yu, “How to not train your dragon: Training-free embodied object goal navigation with semantic frontiers,” arXiv preprint arXiv:2305.16925, 2023. 7, 9

  82. [90]

    Goat: Go to any thing,

    M. Chang, T. Gervet, M. Khanna, S. Yenamandra, D. Shah, S. Y. Min, K. Shah, C. Paxton, S. Gupta, D. Batra et al., “Goat: Go to any thing,” arXiv preprint arXiv:2311.06430, 2023. 7

  83. [91]

    Goat-bench: A benchmark for multi-modal lifelong navigation,

    M. Khanna, R. Ramrakhya, G. Chhablani, S. Yenamandra, T. Gervet, M. Chang, Z. Kira, D. S. Chaplot, D. Batra, and R. Mottaghi, “Goat-bench: A benchmark for multi-modal lifelong navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 20...

  84. [92]

    Hierarchical object-to-zone graph for object navigation,

    S. Zhang, X. Song, Y. Bai, W. Li, Y. Chu, and S. Jiang, “Hierarchical object-to-zone graph for object navigation,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 15 130–15 140. 7, 9, 14

  85. [93]

    Visual navigation in real-world indoor environments using end-to-end deep rein- forcement learning,

    J. Kulh ´anek, E. Derner, and R. Babu ˇska, “Visual navigation in real-world indoor environments using end-to-end deep rein- forcement learning,” IEEE Robotics and Automation Letters , vol. 6, no. 3, pp. 4345–4352, 2021. 7, 8

  86. [94]

    Object memory transformer for object goal navigation,

    R. Fukushima, K. Ota, A. Kanezaki, Y. Sasaki, and Y. Yoshiyasu, “Object memory transformer for object goal navigation,” in 2022 International conference on robotics and automation (ICRA) . IEEE, 2022, pp. 11 288–11 294. 7, 8

  87. [95]

    Learning to terminate in object navigation,

    Y. Song, A. Nguyen, and C.-Y. Lee, “Learning to terminate in object navigation,” arXiv preprint arXiv:2309.16164, 2023. 7, 8, 14

  88. [96]

    Offline visual rep- resentation learning for embodied navigation,

    K. Yadav, R. Ramrakhya, A. Majumdar, V .-P . Berges, S. Kuhar, D. Batra, A. Baevski, and O. Maksymets, “Offline visual rep- resentation learning for embodied navigation,” in Workshop on Reincarnating Reinforcement Learning at ICLR 2023 , 2023. 7

  89. [97]

    Ovrl-v2: A simple state-of-art baseline for imagenav and objectnav,

    K. Yadav, A. Majumdar, R. Ramrakhya, N. Yokoyama, A. Baevski, Z. Kira, O. Maksymets, and D. Batra, “Ovrl-v2: A simple state-of-art baseline for imagenav and objectnav,” arXiv preprint arXiv:2303.07798, 2023. 6, 7, 8

  90. [98]

    Object goal naviga- tion with recursive implicit maps,

    S. Chen, T. Chabal, I. Laptev, and C. Schmid, “Object goal naviga- tion with recursive implicit maps,” in 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2023, pp. 7089–

  91. [99]

    Find what you want: learning demand-conditioned object attribute space for demand-driven navigation,

    H. Wang, A. G. H. Chen, X. Li, M. Wu, and H. Dong, “Find what you want: learning demand-conditioned object attribute space for demand-driven navigation,” Advances in Neural Information Processing Systems, vol. 36, 2024. 7, 9

  92. [100]

    Esc: Exploration with soft commonsense constraints for zero-shot object navigation,

    K. Zhou, K. Zheng, C. Pryor, Y. Shen, H. Jin, L. Getoor, and X. E. Wang, “Esc: Exploration with soft commonsense constraints for zero-shot object navigation,” in International Conference on Machine Learning. PMLR, 2023, pp. 42 829–42 842. 7, 11

  93. [101]

    Imagine before go: Self-supervised generative map for object goal navigation,

    S. Zhang, X. Yu, X. Song, X. Wang, and S. Jiang, “Imagine before go: Self-supervised generative map for object goal navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 16 414–16 425. 7, 10, 14

  94. [102]

    Openfmnav: Towards open- set zero-shot object navigation via vision-language foundation models,

    Y. Kuang, H. Lin, and M. Jiang, “Openfmnav: Towards open- set zero-shot object navigation via vision-language foundation models,” arXiv preprint arXiv:2402.10670, 2024. 7, 11

  95. [103]

    Voronav: Voronoi-based zero-shot object navigation with large language model,

    P . Wu, Y. Mu, B. Wu, Y. Hou, J. Ma, S. Zhang, and C. Liu, “Voronav: Voronoi-based zero-shot object navigation with large language model,” arXiv preprint arXiv:2401.02695, 2024. 7, 11

  96. [104]

    Bridging zero-shot object navigation and foundation models through pixel-guided navigation skill,

    W. Cai, S. Huang, G. Cheng, Y. Long, P . Gao, C. Sun, and H. Dong, “Bridging zero-shot object navigation and foundation models through pixel-guided navigation skill,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 5228–5234. 7, 11

  97. [105]

    Zero experi- ence required: Plug & play modular transfer learning for seman- tic visual navigation,

    Z. Al-Halah, S. K. Ramakrishnan, and K. Grauman, “Zero experi- ence required: Plug & play modular transfer learning for seman- tic visual navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 17 031–17 041. 7, 10

  98. [106]

    Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation,

    S. Y. Gadre, M. Wortsman, G. Ilharco, L. Schmidt, and S. Song, “Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 23 171–23 181. 7, 10

  99. [107]

    Aligning knowledge graph with visual percep- tion for object-goal navigation,

    N. Xu, W. Wang, R. Yang, M. Qin, Z. Lin, W. Song, C. Zhang, J. Gu, and C. Li, “Aligning knowledge graph with visual percep- tion for object-goal navigation,” arXiv preprint arXiv:2402.18892 ,

  100. [108]

    Vlfm: Vision-language frontier maps for zero-shot semantic naviga- tion,

    N. Yokoyama, S. Ha, D. Batra, J. Wang, and B. Bucher, “Vlfm: Vision-language frontier maps for zero-shot semantic naviga- tion,” in 2024 IEEE International Conference on Robotics and Au- tomation (ICRA). IEEE, 2024, pp. 42–48. 7, 11

  101. [110]

    Target-driven visual navigation in indoor scenes using deep reinforcement learning,

    Y. Zhu, R. Mottaghi, E. Kolve, J. J. Lim, A. Gupta, L. Fei-Fei, and A. Farhadi, “Target-driven visual navigation in indoor scenes using deep reinforcement learning,” in 2017 IEEE international conference on robotics and automation (ICRA). IEEE, 2017, pp. 3357–

  102. [111]

    Learning to navigate in cities without a map,

    P . Mirowski, M. Grimes, M. Malinowski, K. M. Hermann, K. An- derson, D. Teplyashin, K. Simonyan, A. Zisserman, R. Hadsell et al., “Learning to navigate in cities without a map,” Advances in neural information processing systems, vol. 31, 2018. 6

  103. [112]

    Deep learning,

    Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature, vol. 521, no. 7553, pp. 436–444, 2015. 6

  104. [113]

    Asynchronous methods for deep reinforcement learn- ing,

    V . Mnih, “Asynchronous methods for deep reinforcement learn- ing,” arXiv preprint arXiv:1602.01783, 2016. 6, 9

  105. [114]

    Proximal policy optimization algorithms,

    J. Schulman, F. Wolski, P . Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017. 6, 8

  106. [115]

    A reduction of imitation learning and structured prediction to no-regret online learning,

    S. Ross, G. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in Proceedings of the fourteenth international conference on artificial intelligence and statistics . JMLR Workshop and Conference Pro- ceedings, 201...

  107. [116]

    Hybrid computing using a neural network with dynamic external memory,

    A. Graves, G. Wayne, M. Reynolds, T. Harley, I. Danihelka, A. Grabska-Barwi ´nska, S. G. Colmenarejo, E. Grefenstette, T. Ra- malho, J. Agapiou et al. , “Hybrid computing using a neural network with dynamic external memory,” Nature, vol. 538, no. 7626, pp. 471–476, 2016. 6

  108. [117]

    Unsupervised predictive memory in a goal-directed agent,

    G. Wayne, C.-C. Hung, D. Amos, M. Mirza, A. Ahuja, A. Grabska- Barwinska, J. Rae, P . Mirowski, J. Z. Leibo, A. Santoro et al. , “Unsupervised predictive memory in a goal-directed agent,” arXiv preprint arXiv:1803.10760, 2018. 6

  109. [118]

    End-to-end training of deep visuomotor policies,

    S. Levine, C. Finn, T. Darrell, and P . Abbeel, “End-to-end training of deep visuomotor policies,”Journal of Machine Learning Research, vol. 17, no. 39, pp. 1–40, 2016. 6

  110. [119]

    Masked autoencoders are scalable vision learners,

    K. He, X. Chen, S. Xie, Y. Li, P . Doll ´ar, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 16 000–16 009. 6, 8, 10

  111. [120]

    End-to-end object detection with transformers,

    N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in European conference on computer vision . Springer, 2020, pp. 213–229. 8

  112. [121]

    Vision- based navigation using deep reinforcement learning,

    J. Kulh ´anek, E. Derner, T. De Bruin, and R. Babu ˇska, “Vision- based navigation using deep reinforcement learning,” in 2019 european conference on mobile robots (ECMR). IEEE, 2019, pp. 1–8. 8

  113. [122]

    Efficient parallel methods for deep reinforcement learning,

    A. V . Clemente, H. N. Castej ´on, and A. Chandra, “Efficient parallel methods for deep reinforcement learning,” arXiv preprint arXiv:1705.04862, 2017. 8

  114. [123]

    Visual genome: Connecting language and vision using crowdsourced dense image annotations,

    R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” International journal of computer vision , vol. 123, pp. 32–73, 2017. 8, 9

  115. [124]

    Attention is all you need,

    A. Waswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS, 2017. 8

  116. [125]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    D. Alexey, “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv: 2010.11929, 2020. 8

  117. [126]

    Sigmoid loss for language image pre-training,

    X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 11 975– 11 986. 8 SUBMITTED IEEE TRANSACTIONS ON P A TTERN ANAL YSIS AND MACHINE INTELLIGENCE 19

  118. [127]

    Scaling open- vocabulary object detection,

    M. Minderer, A. Gritsenko, and N. Houlsby, “Scaling open- vocabulary object detection,” Advances in Neural Information Pro- cessing Systems, vol. 36, 2024. 8

  119. [128]

    Mask r-cnn,

    K. He, G. Gkioxari, P . Doll ´ar, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2961–2969. 8

  120. [129]

    Pyramid scene parsing network,

    H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2881–2890. 8

  121. [130]

    Efficient object search with belief road map using mobile robot,

    C. Wang, J. Cheng, J. Wang, X. Li, and M. Q.-H. Meng, “Efficient object search with belief road map using mobile robot,” IEEE Robotics and Automation Letters , vol. 3, no. 4, pp. 3081–3088, 2018. 8

  122. [131]

    Fusion-aware point convolution for online semantic 3d scene segmentation,

    J. Zhang, C. Zhu, L. Zheng, and K. Xu, “Fusion-aware point convolution for online semantic 3d scene segmentation,” in Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 4534–4543. 8

  123. [132]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P . P . Srinivasan, M. Tancik, J. T. Barron, R. Ra- mamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM , vol. 65, no. 1, pp. 99–106, 2021. 9, 11

  124. [133]

    A simple framework for contrastive learning of visual representations,

    T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in International conference on machine learning . PMLR, 2020, pp. 1597–1607. 9, 10

  125. [134]

    Deeplab: Semantic image segmentation with deep convo- lutional nets, atrous convolution, and fully connected crfs,

    L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convo- lutional nets, atrous convolution, and fully connected crfs,” IEEE transactions on pattern analysis and machine intelligence , vol. 40, no. 4, pp. 834–84...

  126. [135]

    Orb-slam: a versatile and accurate monocular slam system,

    R. Mur-Artal, J. M. M. Montiel, and J. D. Tardos, “Orb-slam: a versatile and accurate monocular slam system,” IEEE transactions on robotics, vol. 31, no. 5, pp. 1147–1163, 2015. 9

  127. [136]

    Bert: Pre-training of deep bidirectional transformers for language understanding,

    J. Devlin, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805 ,

  128. [137]

    Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras,

    R. Mur-Artal and J. D. Tard ´os, “Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras,” IEEE transactions on robotics, vol. 33, no. 5, pp. 1255–1262, 2017. 9

  129. [138]

    Faster r-cnn: Towards real-time object detection with region proposal networks,

    S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” IEEE transactions on pattern analysis and machine intelligence , vol. 39, no. 6, pp. 1137–1149, 2016. 9

  130. [139]

    Planning and acting in partially observable stochastic domains,

    L. P . Kaelbling, M. L. Littman, and A. R. Cassandra, “Planning and acting in partially observable stochastic domains,” Artificial intelligence, vol. 101, no. 1-2, pp. 99–134, 1998. 9

  131. [140]

    Learning contextual dependence with convolutional hierarchical recurrent neural networks,

    Z. Zuo, B. Shuai, G. Wang, X. Liu, X. Wang, B. Wang, and Y. Chen, “Learning contextual dependence with convolutional hierarchical recurrent neural networks,” IEEE Transactions on Image Processing, vol. 25, no. 7, pp. 2983–2996, 2016. 9

  132. [141]

    Minigpt-4: Enhancing vision-language understanding with advanced large language models,

    D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt-4: Enhancing vision-language understanding with advanced large language models,” arXiv preprint arXiv:2304.10592, 2023. 9

  133. [142]

    Palm: Scaling language modeling with pathways,

    A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P . Barham, H. W. Chung, C. Sutton, S. Gehrmann et al., “Palm: Scaling language modeling with pathways,” Journal of Machine Learning Research, vol. 24, no. 240, pp. 1–113, 2023. 9

  134. [143]

    Lamda: Language models for dialog applications,

    R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kul- shreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du et al. , “Lamda: Language models for dialog applications,” arXiv preprint arXiv:2201.08239, 2022. 9

  135. [144]

    Glm: General language model pretraining with autoregressive blank infilling,

    Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang, “Glm: General language model pretraining with autoregressive blank infilling,” arXiv preprint arXiv:2103.10360, 2021. 10, 11

  136. [145]

    Diffusion models beat gans on im- age synthesis,

    P . Dhariwal and A. Nichol, “Diffusion models beat gans on im- age synthesis,” Advances in neural information processing systems , vol. 34, pp. 8780–8794, 2021. 10

  137. [146]

    Improved denoising diffusion probabilistic models,

    A. Q. Nichol and P . Dhariwal, “Improved denoising diffusion probabilistic models,” in International conference on machine learn- ing. PMLR, 2021, pp. 8162–8171. 10

  138. [147]

    Diffusion policy: Visuomotor policy learning via action diffusion,

    C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” The International Journal of Robotics Research, p. 02783649241273668, 2023. 10

  139. [148]

    Elucidating the design space of diffusion-based generative models,

    T. Karras, M. Aittala, T. Aila, and S. Laine, “Elucidating the design space of diffusion-based generative models,” Advances in neural information processing systems , vol. 35, pp. 26 565–26 577,

  140. [149]

    Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

    J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in International conference on machine learning. PMLR, 2023, pp. 19 730–19 742. 10, 11, 13

  141. [150]

    Alfred: A benchmark for interpreting grounded instructions for everyday tasks,

    M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mot- taghi, L. Zettlemoyer, and D. Fox, “Alfred: A benchmark for interpreting grounded instructions for everyday tasks,” in Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 1...

  142. [151]

    Learning to navigate in complex environments,

    P . Mirowski, R. Pascanu, F. Viola, H. Soyer, A. J. Ballard, A. Ban- ino, M. Denil, R. Goroshin, L. Sifre, K. Kavukcuoglu et al. , “Learning to navigate in complex environments,” arXiv preprint arXiv:1611.03673, 2016. 10

  143. [152]

    Domain randomization for transferring deep neural networks from simulation to the real world,

    J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P . Abbeel, “Domain randomization for transferring deep neural networks from simulation to the real world,” in 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS). IEEE, 2017, pp. 23–30. 10, 15

  144. [153]

    Yolov7: Train- able bag-of-freebies sets new state-of-the-art for real-time object detectors,

    C.-Y. Wang, A. Bochkovskiy, and H.-Y. M. Liao, “Yolov7: Train- able bag-of-freebies sets new state-of-the-art for real-time object detectors,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 7464–7475. 11

  145. [154]

    Grounding dino: Marrying dino with grounded pre-training for open-set object detection,

    S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su et al. , “Grounding dino: Marrying dino with grounded pre-training for open-set object detection,” in European Conference on Computer Vision. Springer, 2024, pp. 38–55. 11

  146. [155]

    Faster segment anything: Towards lightweight sam for mobile applications,

    C. Zhang, D. Han, Y. Qiao, J. U. Kim, S.-H. Bae, S. Lee, and C. S. Hong, “Faster segment anything: Towards lightweight sam for mobile applications,” arXiv preprint arXiv:2306.14289, 2023. 11

  147. [156]

    Grounded language- image pre-training,

    L. H. Li, P . Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang et al. , “Grounded language- image pre-training,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 10 965–10 975. 11

  148. [157]

    Chain-of-thought prompting elicits reasoning in large language models,

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhou et al., “Chain-of-thought prompting elicits reasoning in large language models,” Advances in neural information processing systems, vol. 35, pp. 24 824–24 837, 2022. 11

  149. [158]

    Hinge-loss markov random fields and probabilistic soft logic,

    S. H. Bach, M. Broecheler, B. Huang, and L. Getoor, “Hinge-loss markov random fields and probabilistic soft logic,” Journal of Machine Learning Research, vol. 18, no. 109, pp. 1–67, 2017. 11

  150. [159]

    Voronoi diagrams—a survey of a fundamen- tal geometric data structure,

    F. Aurenhammer, “Voronoi diagrams—a survey of a fundamen- tal geometric data structure,” ACM Computing Surveys (CSUR) , vol. 23, no. 3, pp. 345–405, 1991. 11

  151. [160]

    Llama-adapter: Efficient fine-tuning of language models with zero-init attention,

    R. Zhang, J. Han, C. Liu, P . Gao, A. Zhou, X. Hu, S. Yan, P . Lu, H. Li, and Y. Qiao, “Llama-adapter: Efficient fine-tuning of language models with zero-init attention,” arXiv preprint arXiv:2303.16199, 2023. 11

  152. [161]

    Image- goal navigation in complex environments via modular learning,

    Q. Wu, J. Wang, J. Liang, X. Gong, and D. Manocha, “Image- goal navigation in complex environments via modular learning,” IEEE Robotics and Automation Letters , vol. 7, no. 3, pp. 6902–6909,

  153. [162]

    Memory-augmented reinforce- ment learning for image-goal navigation,

    L. Mezghan, S. Sukhbaatar, T. Lavril, O. Maksymets, D. Batra, P . Bojanowski, and K. Alahari, “Memory-augmented reinforce- ment learning for image-goal navigation,” in 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2022, pp. 3316–3323. ...

  154. [163]

    Navigating to objects specified by images,

    J. Krantz, T. Gervet, K. Yadav, A. Wang, C. Paxton, R. Mottaghi, D. Batra, J. Malik, S. Lee, and D. S. Chaplot, “Navigating to objects specified by images,” in Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision , 2023, pp. 10 916–10 925. 11, 12

  155. [164]

    Feudal networks for visual navigation,

    F. Johnson, B. B. Cao, A. Ashok, S. Jain, and K. Dana, “Feudal networks for visual navigation,” arXiv preprint arXiv:2402.12498 ,

  156. [165]

    Litevloc: Map-lite visual localization for image goal navigation,

    J. Jiao, J. He, C. Liu, S. Aegidius, X. Hu, T. Braud, and D. Kanoulas, “Litevloc: Map-lite visual localization for image goal navigation,” arXiv preprint arXiv:2410.04419, 2024. 11, 12

  157. [166]

    End-to-end (instance)-image goal navigation through correspondence as an emergent phenomenon,

    G. Bono, L. Antsfeld, B. Chidlovskii, P . Weinzaepfel, and C. Wolf, “End-to-end (instance)-image goal navigation through correspondence as an emergent phenomenon,” arXiv preprint arXiv:2309.16634, 2023. 11, 12, 14

  158. [167]

    Last-mile embodied visual navigation,

    J. Wasserman, K. Yadav, G. Chowdhary, A. Gupta, and U. Jain, “Last-mile embodied visual navigation,” in Conference on Robot Learning. PMLR, 2023, pp. 666–678. 11, 12 SUBMITTED IEEE TRANSACTIONS ON P A TTERN ANAL YSIS AND MACHINE INTELLIGENCE 20

  159. [168]

    Enhancing exploratory capability of visual navigation using uncertainty of implicit scene representation,

    Y. Wang, Q. Liu, Z. Liu, and H. Wang, “Enhancing exploratory capability of visual navigation using uncertainty of implicit scene representation,” in 2024 IEEE/RSJ International Conference on Intel- ligent Robots and Systems (IROS) . IEEE, 2024, pp. 13 824–13 829. 11, 12

  160. [169]

    Soon: Scenario oriented object navigation with graph-based explo- ration,

    F. Zhu, X. Liang, Y. Zhu, Q. Yu, X. Chang, and X. Liang, “Soon: Scenario oriented object navigation with graph-based explo- ration,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 12 689–12 699. 11

  161. [170]

    Representation learning with contrastive predictive coding,

    A. v. d. Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” arXiv preprint arXiv:1807.03748 ,

  162. [171]

    Neural episodic control,

    A. Pritzel, B. Uria, S. Srinivasan, A. P . Badia, O. Vinyals, D. Has- sabis, D. Wierstra, and C. Blundell, “Neural episodic control,” in International conference on machine learning . PMLR, 2017, pp. 2827–2836. 11

  163. [172]

    Superpoint: Self- supervised interest point detection and description,

    D. DeTone, T. Malisiewicz, and A. Rabinovich, “Superpoint: Self- supervised interest point detection and description,” in Proceed- ings of the IEEE conference on computer vision and pattern recognition workshops, 2018, pp. 224–236. 11, 13

  164. [173]

    Superglue: Learning feature matching with graph neural net- works,

    P .-E. Sarlin, D. DeTone, T. Malisiewicz, and A. Rabinovich, “Superglue: Learning feature matching with graph neural net- works,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 4938–4947. 11

  165. [174]

    Unsupervised visual representation learning by synchronous momentum grouping,

    B. Pang, Y. Zhang, Y. Li, J. Cai, and C. Lu, “Unsupervised visual representation learning by synchronous momentum grouping,” in European Conference on Computer Vision . Springer, 2022, pp. 265–282. 11

  166. [175]

    Orb: An efficient alternative to sift or surf,

    E. Rublee, V . Rabaud, K. Konolige, and G. Bradski, “Orb: An efficient alternative to sift or surf,” in 2011 International conference on computer vision. Ieee, 2011, pp. 2564–2571. 11

  167. [176]

    Information-based compact pose slam,

    V . Ila, J. M. Porta, and J. Andrade-Cetto, “Information-based compact pose slam,” IEEE Transactions on Robotics, vol. 26, no. 1, pp. 78–93, 2009. 11

  168. [177]

    Croco: Self-supervised pre-training for 3d vision tasks by cross-view completion,

    P . Weinzaepfel, V . Leroy, T. Lucas, R. Br´egier, Y. Cabon, V . Arora, L. Antsfeld, B. Chidlovskii, G. Csurka, and J. Revaud, “Croco: Self-supervised pre-training for 3d vision tasks by cross-view completion,” Advances in Neural Information Processing Systems , vol. 35, pp. 3...

  169. [178]

    Film: Visual reasoning with a general conditioning layer,

    E. Perez, F. Strub, H. De Vries, V . Dumoulin, and A. Courville, “Film: Visual reasoning with a general conditioning layer,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, no. 1, 2018. 11

  170. [179]

    Epnp: Efficient perspective-n-point camera pose estimation,

    V . Lepetit, F. Moreno-Noguer, and P . Fua, “Epnp: Efficient perspective-n-point camera pose estimation,” Int. J. Comput. Vis , vol. 81, no. 2, pp. 155–166, 2009. 11

  171. [180]

    Exploring simple siamese representation learning,

    X. Chen and K. He, “Exploring simple siamese representation learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 15 750–15 758. 11

  172. [181]

    Semi- parametric topological memory for navigation,

    N. Savinov, A. Dosovitskiy, and V . Koltun, “Semi- parametric topological memory for navigation,” arXiv preprint arXiv:1803.00653, 2018. 12

  173. [182]

    A behavioral approach to vi- sual navigation with graph localization networks,

    K. Chen, J. P . De Vicente, G. Sepulveda, F. Xia, A. Soto, M. V ´azquez, and S. Savarese, “A behavioral approach to vi- sual navigation with graph localization networks,” arXiv preprint arXiv:1903.00445, 2019. 12

  174. [183]

    Learning to learn how to learn: Self-adaptive visual navigation using meta-learning,

    M. Wortsman, K. Ehsani, M. Rastegari, A. Farhadi, and R. Mot- taghi, “Learning to learn how to learn: Self-adaptive visual navigation using meta-learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 6750–6759. 12

  175. [184]

    Learning to set waypoints for audio-visual navigation,

    C. Chen, S. Majumder, Z. Al-Halah, R. Gao, S. K. Ramakrishnan, and K. Grauman, “Learning to set waypoints for audio-visual navigation,” arXiv preprint arXiv:2008.09622, 2020. 12, 13, 14

  176. [185]

    Sim2real transfer for audio-visual navigation with frequency-adaptive acoustic field prediction,

    C. Chen, J. Ramos, A. Tomar, and K. Grauman, “Sim2real transfer for audio-visual navigation with frequency-adaptive acoustic field prediction,” in 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2024, pp. 8595–8602. 12, 13

  177. [186]

    Sound adversarial audio-visual navigation,

    Y. Yu, W. Huang, F. Sun, C. Chen, Y. Wang, and X. Liu, “Sound adversarial audio-visual navigation,” arXiv preprint arXiv:2202.10910, 2022. 13

  178. [187]

    Omnidirectional information gathering for knowledge transfer-based audio-visual navigation,

    J. Chen, W. Wang, S. Liu, H. Li, and Y. Yang, “Omnidirectional information gathering for knowledge transfer-based audio-visual navigation,” in Proceedings of the IEEE/CVF International Confer- ence on Computer Vision, 2023, pp. 10 993–11 003. 13, 14

  179. [188]

    Multi-goal audio-visual naviga- tion using sound direction map,

    H. Kondoh and A. Kanezaki, “Multi-goal audio-visual naviga- tion using sound direction map,” in 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2023, pp. 5219–5226. 13

  180. [189]

    Caven: an em- bodied conversational agent for efficient audio-visual navigation in noisy environments,

    X. Liu, S. Paul, M. Chatterjee, and A. Cherian, “Caven: an em- bodied conversational agent for efficient audio-visual navigation in noisy environments,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 4, 2024, pp. 3765–3773. 13, 15

  181. [190]

    Audio visual language maps for robot navigation,

    C. Huang, O. Mees, A. Zeng, and W. Burgard, “Audio visual language maps for robot navigation,” in International Symposium on Experimental Robotics. Springer, 2023, pp. 105–117. 13, 14, 15

  182. [191]

    Visualechoes: Spatial image representation learning through echolocation,

    R. Gao, C. Chen, Z. Al-Halah, C. Schissler, and K. Grauman, “Visualechoes: Spatial image representation learning through echolocation,” in Computer Vision–ECCV 2020: 16th European Con- ference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IX 16 . Springer, 2020, pp. 658–676. 12

  183. [192]

    Mapnet: An allocentric spatial memory for mapping environments,

    J. F. Henriques and A. Vedaldi, “Mapnet: An allocentric spatial memory for mapping environments,” in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 8476–8484. 12

  184. [193]

    Visual representations for semantic target driven navigation,

    A. Mousavian, A. Toshev, M. Fi ˇser, J. Ko ˇseck´a, A. Wahid, and J. Davidson, “Visual representations for semantic target driven navigation,” in 2019 International Conference on Robotics and Au- tomation (ICRA). IEEE, 2019, pp. 8846–8852. 13

  185. [194]

    Sta- bilizing transformers for reinforcement learning,

    E. Parisotto, F. Song, J. Rae, R. Pascanu, C. Gulcehre, S. Jayaku- mar, M. Jaderberg, R. L. Kaufman, A. Clark, S. Noury et al., “Sta- bilizing transformers for reinforcement learning,” in International conference on machine learning. PMLR, 2020, pp. 7487–7498. 13

  186. [195]

    Krishnamurthy, Partially observed Markov decision processes

    V . Krishnamurthy, Partially observed Markov decision processes . Cambridge university press, 2016. 13

  187. [196]

    Netvlad: Cnn architecture for weakly supervised place recog- nition,

    R. Arandjelovic, P . Gronat, A. Torii, T. Pajdla, and J. Sivic, “Netvlad: Cnn architecture for weakly supervised place recog- nition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 5297–5307. 13

  188. [197]

    Scaling open-vocabulary image segmentation with image-level labels,

    G. Ghiasi, X. Gu, Y. Cui, and T.-Y. Lin, “Scaling open-vocabulary image segmentation with image-level labels,” in European confer- ence on computer vision. Springer, 2022, pp. 540–557. 13

  189. [198]

    Audioclip: Extend- ing clip to image, text and audio,

    A. Guzhov, F. Raue, J. Hees, and A. Dengel, “Audioclip: Extend- ing clip to image, text and audio,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 976–980. 13

  190. [199]

    Conceptfusion: Open-set multimodal 3d mapping,

    K. M. Jatavallabhula, A. Kuwajerwala, Q. Gu, M. Omama, T. Chen, A. Maalouf, S. Li, G. Iyer, S. Saryazdi, N. Keetha et al., “Conceptfusion: Open-set multimodal 3d mapping,” arXiv preprint arXiv:2302.07241, 2023. 13

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.