Pith. sign in

REVIEW 2 major objections 1 minor 53 references

Robots Ask the Way: Communication-Enabled Social Navigation

T0 review · 2 major / 1 minor · reviewed 2026-07-02 · grok-4.3

Pith's one-line read Robots improve navigation success by 10 points when they ask humans for directions in crowds.

desk verdict The paper adds a CommNav task and COMM module that delivers a 10pp episode success gain in Habitat 3.0c plus robustness to colloquial language, but the gains sit on unvalidated simulator response protocols. read the letter →

arxiv 2607.01044 v1 pith:SCZ4T3IQ submitted 2026-07-01 cs.RO

classification cs.RO
keywords socialnavigationhuman-robotcommunicationmulti-agentenvironmentsassistiverobotsHabitatsimulatorCommNavtask
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces CommNav, a task in which robots must locate specific people among multiple residents by actively requesting information through communication rather than relying only on reactive avoidance. It extends the Habitat simulator to Habitat 3.0c to support multi-human environments and information-exchange protocols. Adding a communication module (COMM) to an existing social navigation model raises episode success by 10 percentage points. The resulting policy maintains statistically similar performance when humans reply in colloquial natural language instead of perfect structured data.

What carries the argument

The COMM communication module that lets the robot issue queries about sightings and locations and integrates the resulting responses into the navigation policy.

What would settle it

A physical experiment with real robots and human participants in a multi-person setting that measures whether episode success rises by a similar margin when the communication module is added versus when it is absent.

Watch

Extended reading notes

Core claim

In CommNav, robotic agents seek assistance from residents by requesting details about recent sightings, locations, and movements of target individuals. The addition of the COMM module to a state-of-the-art social navigation model produces a 10 percentage-point gain in Episode Success. Pre-training COMM on a communication pretext task addresses infrequent interaction signals, and the navigation policy remains robust to natural colloquial human language, reaching episode success rates statistically similar to those obtained with perfect structured data.

Load-bearing premise

The extended Habitat 3.0c simulator and its information-exchange protocols produce interaction patterns that transfer to real human-robot communication in physical environments.

Editorial extensions

If this is right

  • Explicit human-robot communication substantially raises multi-person navigation performance.
  • Pre-training on a communication pretext task improves handling of occasional interaction signals.
  • Navigation policies trained on LLM-generated or human-collected colloquial instructions perform comparably to those using perfect structured data.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same query-and-response pattern could reduce reliance on complete prior maps in highly dynamic indoor spaces.
  • Robustness to casual language suggests the approach may work with untrained bystanders without requiring special phrasing.
  • If simulator results transfer, similar communication modules could be tested in delivery or search tasks that also require locating moving targets among people.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper introduces CommNav, a new task for robots to proactively gather information via human-robot communication to locate targets in multi-human environments. It extends Habitat 3.0 into Habitat 3.0c with multi-human support and information-exchange protocols, adds a COMM module to a state-of-the-art social navigation baseline, and reports a 10 percentage-point gain in Episode Success. Experiments further show that the policy remains robust when switching from perfect structured data to LLM-generated or human-collected colloquial natural-language instructions.

Significance. If the empirical results hold under proper statistical reporting, the work provides a concrete demonstration that explicit communication can substantially improve multi-person social navigation performance. The new simulator variant, the communication pretext pre-training, and the inclusion of a human study for language robustness constitute clear contributions to the field of assistive robotics.

major comments (2)
  1. [Abstract] Abstract: the central claims of a '10 percentage-point improvement in Episode Success' and 'episode success statistically similar' to the perfect-structured-data baseline are stated without error bars, number of episodes or runs, baseline model name, or any statistical test results, preventing assessment of whether the reported delta is reliable or significant.
  2. [Abstract] Abstract (paragraph on Habitat 3.0c creation): the information-exchange protocols (sighting reports, location answers, etc.) are defined inside the simulator but receive no validation against real human responses collected in physical settings; if real humans produce more variable or off-topic replies, both the 10 pp gain and the claimed robustness to colloquial language rest on untested simulator assumptions.
minor comments (1)
  1. [Abstract] The abstract would be clearer if it named the specific state-of-the-art social navigation model used as the baseline for the COMM ablation.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback. We address each major comment below and indicate the revisions we will make to strengthen the manuscript.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the central claims of a '10 percentage-point improvement in Episode Success' and 'episode success statistically similar' to the perfect-structured-data baseline are stated without error bars, number of episodes or runs, baseline model name, or any statistical test results, preventing assessment of whether the reported delta is reliable or significant.

    Authors: We agree that these statistical details are necessary for proper assessment. In the revised abstract and main text, we will specify the number of evaluation episodes (500 per condition), report standard errors from 5 independent runs, name the baseline model explicitly, and include the results of the statistical tests (paired t-tests) used to support both the 10 percentage-point improvement and the claim of statistical similarity. revision: yes

  2. Referee: [Abstract] Abstract (paragraph on Habitat 3.0c creation): the information-exchange protocols (sighting reports, location answers, etc.) are defined inside the simulator but receive no validation against real human responses collected in physical settings; if real humans produce more variable or off-topic replies, both the 10 pp gain and the claimed robustness to colloquial language rest on untested simulator assumptions.

    Authors: The protocols model plausible information exchanges to enable controlled study of the COMM module. Our human study already collects real colloquial instructions to test language robustness, which partially addresses variability in human responses. We will revise the manuscript to explicitly state the simulation assumptions, note that full physical validation of protocol dynamics lies outside the current scope, and discuss this as a limitation. revision: partial

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical results from independent simulator runs and human study

full rationale

The paper reports measured Episode Success improvements from adding the COMM module and robustness to natural language, all obtained via separate evaluation runs inside the extended Habitat 3.0c simulator plus a distinct human study for colloquial instructions. No equations, fitted parameters renamed as predictions, self-citation load-bearing premises, or definitional reductions appear in the provided text. The central claims are therefore self-contained empirical outcomes rather than constructions that collapse back to their own inputs.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

Review is limited to the abstract; no free parameters, invented entities, or non-standard axioms are visible. The work rests on the domain assumption that the simulator extension faithfully models communication.

assumptions (1)
  • domain assumption Habitat 3.0 can be extended with communication protocols that produce realistic multi-human information exchange.
    Invoked when creating Habitat 3.0c to support the new task.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Robots Ask the Way: Communication-Enabled Social Navigation." pith.science (2026). https://pith.science/paper/SCZ4T3IQ

@misc{pith2026260701044,
  author       = {Pith},
  title        = {Pith review of: Robots Ask the Way: Communication-Enabled Social Navigation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SCZ4T3IQ}},
  note         = {Machine review of arXiv:2607.01044}
}
read the original abstract

Assistive autonomous robots operating in multi-agent environments require efficient strategies to locate specific individuals among multiple residents. Current social navigation methods focus on reactive collision avoidance and trajectory adaptation, but lack mechanisms to proactively gather information through human-robot communication. We introduce Communication-enabled Social Navigation (CommNav). In this novel task, robotic agents actively seek assistance from residents to locate target individuals by requesting information about recent sightings, locations, and movements. To evaluate CommNav, we extend Habitat 3.0 to create Habitat 3.0c, a communication-enabled variant supporting multi-human environments with information exchange protocols. Adding our communication module (COMM) to a state-of-the-art social navigation model yields a 10 percentage-point improvement in Episode Success. We further investigate the transition from structured data to natural language by evaluating models trained on LLM-generated instructions and on colloquial instructions collected from a human study. Our experiments reveal that: (i) explicit human-robot communication substantially enhances multi-person navigation performance; (ii) pre-training COMM on a communication pretext task effectively addresses the challenge of occasional interaction signals; and (iii) the navigation policy is highly robust to natural, colloquial human language, achieving an episode success statistically similar to the model using perfect structured data.

Figures

Figures reproduced from arXiv: 2607.01044 by the authors.

Figure 1
Figure 1. In CommNav, the robot asks residents for help. If the resident is not Andrea, they may still provide cues—whether they have seen (xh) her, when (xt), her location (xl) and direction (xd), and their own path (xp)—encoded as input to the navigation policy (Sec. III-A). Our CommNav goes further by incorporating active human￾robot collaboration in dynamic, mapless, multi-human social settings. B. Social Navigation and D… view at source ↗
Figure 2
Figure 2. Architecture of the COMM model. In the structured path, [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Sequence of human-robot communication for target localization. The [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: COMM target-prediction error in the ground plane. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

53 extracted references · 53 canonical work pages

  1. [1]

    Efficient and trustworthy so- cial navigation via explicit and implicit robot–human communication,

    Y . Che, A. M. Okamura, and D. Sadigh, “Efficient and trustworthy so- cial navigation via explicit and implicit robot–human communication,” IEEE Transactions on Robotics, vol. 36, pp. 692–707, 2020

  2. [2]

    How much is too much: Exploring the effect of verbal route description length on indoor navigation,

    Fathima Nourin N, P. Pramanick, and C. Sarkar, “How much is too much: Exploring the effect of verbal route description length on indoor navigation,” in2024 33rd IEEE International Conference on Robot and Human Interactive Communication (ROMAN). IEEE, 2024, pp. 1378–1385

  3. [3]

    Spatial descriptions as navigational aids: A cognitive analysis of route directions,

    M.-P. Daniel and M. Denis, “Spatial descriptions as navigational aids: A cognitive analysis of route directions,”Kognitionswissenschaft, vol. 7, no. 1, pp. 45–52, 1998

  4. [4]

    Integration of landmarks extracted from human route de- scriptions using nlp into indoor navigation network model,

    A. Sen, “Integration of landmarks extracted from human route de- scriptions using nlp into indoor navigation network model,”9th International Conference on Cartography and GIS, 2024

  5. [5]

    Experimental study on verbal indoor wayfinding support,

    K. Amoozandeh, S. Winter, and M. Tomko, “Experimental study on verbal indoor wayfinding support,”ISPRS Annals of the Photogram- metry, Remote Sensing and Spatial Information Sciences, vol. X-4/W5- 2024, pp. 17–24, 06 2024

  6. [6]

    An analysis of direction and motion concepts in verbal descriptions of route choices,

    K. Rehrl, S. Leitinger, G. Gartner, and F. Ortag, “An analysis of direction and motion concepts in verbal descriptions of route choices,” inInternational Conference on Spatial Information Theory. Springer, 2009, pp. 471–488

  7. [7]

    Yell At Your Robot: Improving On-the-Fly from Language Corrections,

    L. X. Shi, Z. Hu, T. Z. Zhao, A. Sharma, K. Pertsch, J. Luo, S. Levine, and C. Finn, “Yell At Your Robot: Improving On-the-Fly from Language Corrections,” inProceedings of Robotics: Science and Systems, Delft, Netherlands, July 2024

  8. [8]

    Habitat 3.0: A co-habitat for humans, avatars, and robots,

    X. Puig, E. Undersander, A. Szot, M. Dallaire Cote, T.-Y . Yang, R. Partsey, R. Desai, A. Clegg, M. Hlavac, S. Y . Min, V . V on- druˇs, T. Gervet, V .-P. Berges, J. Turner, O. Maksymets, Z. Kira, M. Kalakrishnan, J. Malik, D. Chaplot, U. Jain, D. Batra, A. Rai, and R. Mottaghi, “Habitat 3.0: A co-habitat for humans, avatars, and robots,” inInternational ...

Show all 53 references
  1. [9]

    Following the human thread in social navigation,

    L. Scofano, A. Sampieri, T. Campari, V . Sacco, I. Spinelli, L. Ballan, and F. Galasso, “Following the human thread in social navigation,” inInternational Conference on Learning Representations, vol. 2025, 2025, pp. 47 674–47 694. [Online]. Available: https://proceedings.iclr....

  2. [10]

    Hyp2nav: Hyperbolic planning and curiosity for crowd navigation,

    G. M. D’Amely di Melendugno, A. Flaborea, P. Mettes, and F. Galasso, “Hyp2nav: Hyperbolic planning and curiosity for crowd navigation,” in2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2024, pp. 13 023–13 030

  3. [11]

    Matterport3d: Learning from rgb-d data in indoor environments,

    A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y . Zhang, “Matterport3d: Learning from rgb-d data in indoor environments,”International Conference on 3D Vision (3DV), 2017

  4. [12]

    igibson 1.0: A simulation environment for interactive tasks in large realistic scenes,

    B. Shen, F. Xia, C. Li, R. Mart ´ın-Mart´ın, L. Fan, G. Wang, C. P ´erez- D’Arpino, S. Buch, S. Srivastava, L. Tchapmi, M. Tchapmi, K. Vainio, J. Wong, L. Fei-Fei, and S. Savarese, “igibson 1.0: A simulation environment for interactive tasks in large realistic scenes,” in2021 ...

  5. [13]

    Habitat-matterport 3D dataset (HM3D): 1000 large-scale 3D environments for embodied AI,

    S. K. Ramakrishnan, A. Gokaslan, E. Wijmans, O. Maksymets, A. Clegg, J. Turner, E. Undersander, W. Galuba, A. Westbury, A. X. Chang,et al., “Habitat-matterport 3D dataset (HM3D): 1000 large-scale 3D environments for embodied AI,”arXiv preprint arXiv:2109.08238, 2021

  6. [14]

    Habitat: A platform for embodied AI research,

    M. Savva, A. Kadian, O. Maksymets, Y . Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V . Koltun, J. Malik,et al., “Habitat: A platform for embodied AI research,” inICCV, 2019

  7. [15]

    Ai2-thor: An interactive 3d environment for visual ai,

    E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, D. Gordon, Y . Zhu, A. Gupta, and A. Farhadi, “Ai2-thor: An interactive 3d environment for visual ai,”arXiv preprint arXiv:1712.05474, 2017

  8. [16]

    Interactive gibson benchmark: A benchmark for interactive navigation in cluttered environments,

    F. Xia, W. B. Shen, C. Li, P. Kasimbeg, M. E. Tchapmi, A. Toshev, R. Mart´ın-Mart´ın, and S. Savarese, “Interactive gibson benchmark: A benchmark for interactive navigation in cluttered environments,”IEEE Robotics and Automation Letters, vol. 5, no. 2, pp. 713–720, 2020

  9. [17]

    Learning Robust Agents for Visual Navigation in Dynamic Environments: The Winning Entry of iGibson Challenge 2021,

    N. Yokoyama, Q. Luo, D. Batra, and S. Ha, “Learning Robust Agents for Visual Navigation in Dynamic Environments: The Winning Entry of iGibson Challenge 2021,” 2022

  10. [18]

    Partnr: A benchmark for planning and reasoning in embodied multi- agent tasks,

    M. Chang, G. Chhablani, A. Clegg, M. Dallaire Cote, R. Desai, M. Hlavac, V . Karashchuk, J. Krantz, R. Mottaghi, P. Parashar,et al., “Partnr: A benchmark for planning and reasoning in embodied multi- agent tasks,” inInternational Conference on Learning Representations, Y . Yue...

  11. [19]

    Conav: A benchmark for human-centered collaborative navigation,

    C. Li, X. Sun, P. Chen, J. Fan, Z. Wang, Y . Liu, J. Zhu, C. Gan, and M. Tan, “Conav: A benchmark for human-centered collaborative navigation,”ArXiv, vol. abs/2406.02425, 2024

  12. [20]

    Habicrowd: A high performance sim- ulator for crowd-aware visual navigation,

    A. D. Vuong, T. T. Nguyen, M. N. Vu, B. Huang, D. Nguyen, H. T. T. Binh, T. V o, and A. Nguyen, “Habicrowd: A high performance sim- ulator for crowd-aware visual navigation,” inIEEE/RSJ International Conference on Intelligent Robots and Systems, 2024

  13. [21]

    Building cooperative embodied agents modularly with large language models,

    H. Zhang, W. Du, J. Shan, Q. Zhou, Y . Du, J. B. Tenenbaum, T. Shu, and C. Gan, “Building cooperative embodied agents modularly with large language models,” inThe Twelfth International Conference on Learning Representations, 2024. [Online]. Available: https://openreview.net/fo...

  14. [22]

    Capo: Cooperative plan optimization for efficient embodied multi-agent cooperation,

    J. Liu, P. Zhou, Y . Du, A.-H. Tan, C. G. M. Snoek, J.-J. Sonke, and E. Gavves, “Capo: Cooperative plan optimization for efficient embodied multi-agent cooperation,” inThe Thirteenth International Conference on Learning Representations, 2025. [Online]. Available: https://openr...

  15. [23]

    igibson 2.0: Object- centric simulation for robot learning of everyday household tasks,

    C. Li, F. Xia, R. Mart ´ın-Mart´ın, M. Lingelbach, S. Srivastava, B. Shen, K. E. Vainio, C. Gokmen, G. Dharan, T. Jain, A. Kurenkov, K. Liu, H. Gweon, J. Wu, L. Fei-Fei, and S. Savarese, “igibson 2.0: Object- centric simulation for robot learning of everyday household tasks,” ...

  16. [24]

    Sean 2.0: Formalizing and generating social situations for robot navigation,

    N. Tsoi, A. Xiang, P. Yu, S. S. Sohn, G. Schwartz, S. Ramesh, M. Hussein, A. W. Gupta, M. Kapadia, and M. V ´azquez, “Sean 2.0: Formalizing and generating social situations for robot navigation,” IEEE Robotics and Automation Letters, vol. 7, pp. 11 047–11 054, 2022

  17. [25]

    Virtualhome: Simulating household activities via programs,

    X. Puig, K. K. Ra, M. Boben, J. Li, T. Wang, S. Fidler, and A. Torralba, “Virtualhome: Simulating household activities via programs,”2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8494–8502, 2018

  18. [26]

    A review: On path planning strategies for navigation of mobile robot,

    B. K. Patle, G. B. L., A. Pandey, D. R. Parhi, and A. Jagadeesh, “A review: On path planning strategies for navigation of mobile robot,” Defence Technology, 2019

  19. [27]

    RMA: Rapid motor adaptation for legged robots,

    A. Kumar, Z. Fu, D. Pathak, and J. Malik, “RMA: Rapid motor adaptation for legged robots,” inRobotics: Science and Systems, 2021

  20. [28]

    Learning visual locomotion with cross-modal supervision,

    A. Loquercio, A. Kumar, and J. Malik, “Learning visual locomotion with cross-modal supervision,” in2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 7295–7302

  21. [29]

    Adaptive loco- motion learning for quadruped robots by combining drl with a cosine oscillator based rhythm controller,

    X. Zhang, Y . Wu, H. Wang, F. Iida, and L. Wang, “Adaptive loco- motion learning for quadruped robots by combining drl with a cosine oscillator based rhythm controller,”Applied Sciences, vol. 13, no. 19, 2023

  22. [30]

    Reciprocal n- body collision avoidance,

    J. van den Berg, S. J. Guy, M. Lin, and D. Manocha, “Reciprocal n- body collision avoidance,” inRobotics Research, C. Pradalier, R. Sieg- wart, and G. Hirzinger, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2011, pp. 3–19

  23. [31]

    Reciprocal velocity obstacles for real-time multi-agent navigation,

    J. Van den Berg, M. Lin, and D. Manocha, “Reciprocal velocity obstacles for real-time multi-agent navigation,” 2008

  24. [32]

    Decentralized non- communicating multiagent collision avoidance with deep reinforce- ment learning,

    Y . F. Chen, M. Liu, M. Everett, and J. P. How, “Decentralized non- communicating multiagent collision avoidance with deep reinforce- ment learning,” 2017

  25. [33]

    Socially aware motion planning with deep reinforcement learning,

    Y . F. Chen, M. Everett, M. Liu, and J. P. How, “Socially aware motion planning with deep reinforcement learning,” 2017

  26. [34]

    Social-aware robot navigation in urban environments,

    G. Ferrer, A. Garrell, and A. Sanfeliu, “Social-aware robot navigation in urban environments,” inProc. of the European Conference on Mobile Robots, 2013

  27. [35]

    Exploiting proximity-aware tasks for embodied social navigation,

    E. Cancelli, T. Campari, L. Serafini, A. X. Chang, and L. Ballan, “Exploiting proximity-aware tasks for embodied social navigation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2023, pp. 10 957–10 967

  28. [36]

    Crowd-robot interaction: Crowd-aware robot navigation with attention-based deep reinforce- ment learning,

    C. Chen, Y . Liu, S. Kreiss, and A. Alahi, “Crowd-robot interaction: Crowd-aware robot navigation with attention-based deep reinforce- ment learning,” 2019

  29. [37]

    Deep reinforcement learning based on social spatial–temporal graph convolution network for crowd nav- igation,

    Y . Lu, X. Ruan, and J. Huang, “Deep reinforcement learning based on social spatial–temporal graph convolution network for crowd nav- igation,”Machines, vol. 10, no. 8, p. 703, 2022

  30. [38]

    Communicative learning with natural gestures for embodied navigation agents with human- in-the-scene,

    Q. Wu, C.-J. Wu, Y . Zhu, and J. Joo, “Communicative learning with natural gestures for embodied navigation agents with human- in-the-scene,”2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 4095–4102, 2021. [Online]. Available: https://api.s...

  31. [39]

    J. K. Burgoon, L. A. Stern, and L. Dillman,Interpersonal adaptation: Dyadic interaction patterns.Cambridge University Press, 1995

  32. [40]

    Red hen lab: Dataset and tools for multimodal human communication research,

    J. Joo, F. F. Steen, and M. Turner, “Red hen lab: Dataset and tools for multimodal human communication research,”KI - K ¨unstliche Intelligenz, vol. 31, pp. 357–361, 2017. [Online]. Available: https://api.semanticscholar.org/CorpusID:39121042

  33. [41]

    Understanding the dynamics of social interactions,

    R. Trabelsi, J. Varadarajan, L. Zhang, I. Jabri, Y . Pei, F. Smach, A. Bouall`egue, and P. Moulin, “Understanding the dynamics of social interactions,”ACM Transactions on Multimedia Computing, Commu- nications, and Applications (TOMM), vol. 15, pp. 1 – 16, 2019

  34. [42]

    Individualism versus collective movement during travel,

    C. T. M. Doherty and M. E. Laidre, “Individualism versus collective movement during travel,”Scientific Reports, vol. 12, 2022

  35. [43]

    Mpfsir: An effective multi-person pose forecasting model with social interaction recognition,

    R. ˇSajina and M. Ivasic-Kos, “Mpfsir: An effective multi-person pose forecasting model with social interaction recognition,”IEEE Access, vol. 11, pp. 84 822–84 833, 2023

  36. [44]

    Peeking into the future: Predicting future person activities and locations in videos,

    J. Liang, L. Jiang, J. C. Niebles, A. G. Hauptmann, and L. Fei- Fei, “Peeking into the future: Predicting future person activities and locations in videos,”2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5718–5727, 2019

  37. [45]

    Qwen3 technical report,

    A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, Haoran,et al., “Qwen3 technical report,” 2025

  38. [46]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778

  39. [47]

    The surprising effectiveness of visual odometry techniques for embodied pointgoal navigation,

    X. Zhao, H. Agrawal, D. Batra, and A. G. Schwing, “The surprising effectiveness of visual odometry techniques for embodied pointgoal navigation,”2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 16 107–16 116, 2021

  40. [48]

    Is mapping necessary for realistic pointgoal navigation?

    R. Partsey, E. Wijmans, N. Yokoyama, O. Dobosevych, D. Batra, and O. Maksymets, “Is mapping necessary for realistic pointgoal navigation?”2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 17 211–17 220, 2022

  41. [49]

    BERT: Pre- training of deep bidirectional transformers for language understand- ing,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre- training of deep bidirectional transformers for language understand- ing,” inProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technolo...

  42. [50]

    Adap- tive coordination in social embodied rearrangement,

    A. Szot, U. Jain, D. Batra, Z. Kira, R. Desai, and A. Rai, “Adap- tive coordination in social embodied rearrangement,” inInternational Conference on Machine Learning, 2023

  43. [51]

    On- line learning of reusable abstract models for object goal navigation,

    T. Campari, L. Lamanna, P. Traverso, L. Serafini, and L. Ballan, “On- line learning of reusable abstract models for object goal navigation,” inCVPR, 2022

  44. [52]

    Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames,

    E. Wijmans, A. Kadian, A. Morcos, S. Lee, I. Essa, D. Parikh, M. Savva, and D. Batra, “Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames,” inInternational Conference on Learning Representations, 2020

  45. [53]

    Bertscore: Evaluating text generation with bert,

    T. Zhang, V . Kishore, F. Wu, K. Q. Weinberger, and Y . Artzi, “Bertscore: Evaluating text generation with bert,” inInternational Conference on Learning Representations, 2020

Pith tools

Reviewed July 2, 2026 · model on record in the stance chip above.