REVIEW 4 major objections 4 minor 2 cited by
Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues
T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A training-free method that asks the user only when uncertain beats trained instance-navigation policies.
desk verdict New interactive instance-nav task and a sensible training-free baseline, but the headline numbers rest on a simulated user validated on only 40 easy episodes. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carriage of the argument is a two-module loop wrapped around any navigation policy. The Self-Questioner obtains a coarse VLM description of each detected candidate, asks an LLM to generate attribute-specific yes/no questions, and restricts VLM answers to "Yes", "No", or "I don't know" so that the Shannon entropy of the three-token output distribution, normalized by its maximum, gives a per-attribute uncertainty in [0,1]. Attributes above a threshold τ are dropped when the LLM recomposes the refined description S_refined. The Interaction Trigger then asks the LLM for a single alignment score s between S_refined and the accumulated facts F_t about the target, along with one question to the user; s ≥ τ_stop halts, s < τ_skip continues silently, and the middle band triggers the user question, whose answer updates F_t. The Normalized-Entropy estimator is the component that lets the agent trust its own perception, and the paper evaluates it separately on IDKVQA, a 502-question dataset with triple human annotations, where it scores 21.12 on the Effective Reliability metric versus 20.45 for the best baseline.
What would settle it
Run the full 1,649-episode CoIN-Bench with human participants who are given only the target image and compare their answers with the VLM-simulated user on the exact same episodes; if success rate or questions-asked diverges beyond the variation seen in the 40-episode pilot, the simulated-user results that drive the headline comparison would not transfer to real users.
Extended reading notes
Core claim
The central claim is that a training-free agent can solve instance object navigation by actively managing its own perceptual uncertainty and asking the human user only when a computed alignment score says a question will help. On CoIN-Bench, AIUTA achieves 6.67% success on the Val Unseen split, above the best trained baseline (PSL, 4.58%), while the zero-shot navigator it wraps, VLFM, scores near zero because it cannot tell the target apart from the roughly five same-category distractors per episode. On the seen splits AIUTA is competitive with trained baselines, and across all splits it succeeds with fewer than two user questions per successful episode (NQ between 1.13 and 1.67). The authors also report a 40-episode human study in which real humans and the VLM-simulated user give the same success rate (42.5%), which they take as evidence that the simulated protocol is a reliable stand-in for large-scale evaluation.
Load-bearing premise
The load-bearing assumption is that a vision-language model shown a high-resolution image of the target answers the agent's questions like a real human, a check the paper runs on only 40 episodes with detectable targets.
Editorial extensions
If this is right
- A user can start an instance-navigation task with a bare category phrase, and the agent, not the user, supplies the descriptive effort through targeted questions.
- On unseen object categories, a training-free pipeline can beat policies trained for the task, suggesting instance disambiguation is more about perception-and-query reasoning than about navigation-policy training.
- The NQ metric plus the simulated-user protocol make large dialogue-based navigation benchmarks reproducible without recruiting hundreds of human annotators.
- Restricting VLM answers to a three-way vocabulary turns open-ended hallucination into a measurable entropy signal, giving a reusable abstention rule for visual question answering.
- Because AIUTA operates independently of the navigation policy, the same interaction reasoning can be plugged into any detector-equipped navigator.
Reading between the lines
- If the simulated-user result transfers, the evaluation design generalizes: any dialogic embodied task with a visual referent could be benchmarked by giving a VLM a high-resolution image of the goal, making human studies a verification step rather than the main data source.
- The normalized-entropy abstention rule is not specific to navigation; it could serve as a generic hallucination filter for any constrained VQA pipeline, including caption-refinement systems.
- A natural extension the paper does not test is making the two thresholds τ_skip and τ_stop adaptive, since a fixed threshold that works on GOAT-Bench scenes may need recalibration in more cluttered or occluded environments.
- The benchmark filters out episodes where the target is hard to see, so the reported question counts likely understate what real deployment would require; testing on unfiltered, occluded episodes is an open empirical question.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Collaborative Instance object Navigation (CoIN), a task in which an embodied agent receives only a minimal category-level instruction and must resolve instance-level ambiguity through open-ended, template-free dialogue with a human user. The authors propose AIUTA, a training-free method that combines a VLFM navigation policy with two LLM/VLM-based modules: a Self-Questioner that refines observations through self-dialogue and normalized-entropy uncertainty estimation, and an Interaction Trigger that decides whether to ask the user, continue exploring, or stop. They also introduce CoIN-Bench, a benchmark of 1,649 episodes derived from GOAT-Bench with multiple distractors, and an IDKVQA dataset for evaluating VLM uncertainty. Experiments with a VLM-simulated user report that AIUTA outperforms trained InstanceObjectNav methods on Val Unseen and surpasses the zero-shot VLFM baseline, while a 40-episode human study is claimed to validate the simulation.
Significance. If the central claims are substantiated, the paper makes a useful contribution: it defines a practical interactive navigation task, shows that a training-free pipeline of pretrained VLMs and LLMs can meaningfully disambiguate object instances with a small number of questions, and provides a reproducible benchmark with a simulated user. The IDKVQA dataset and the proposed normalized-entropy uncertainty estimator are also potentially valuable. However, the current evidence is not yet sufficient to support the quantitative claims, because the human validation is small and biased, the main results rely entirely on a simulated user whose fidelity is not established on the actual evaluation distribution, and no significance testing or confidence intervals are reported.
major comments (4)
- [Sec. 6 'Validation with real human' and Supp. E] The human validation that underpins the simulated-user setup uses only 40 episodes, all explicitly selected to have 'detectable target instances.' This is an easy subset, as shown by the SR of 42.5% on this set versus 6.67% on Val Unseen in Table 2. The paper states 'we observe no significant differences' but reports no statistical test, no confidence intervals, and no effect size. With 40 episodes, the study cannot detect all but extremely large differences in SR, SPL, or NQ. Since every headline result in Table 2 and every ablation in Table 4 uses the VLM-simulated user, the quantitative validity of the central claim is not established. Please provide a proper statistical comparison (e.g., bootstrap confidence intervals or a hypothesis test) and validate on a more representative sample, or explicitly frame the simulated-user results as preliminary.
- [Sec. 6 'Implementation Details'] The thresholds tau = 0.75, tau_stop = 7, and tau_skip = 5 are described only as set 'as they yield the best result.' If these values were selected on the CoIN-Bench evaluation splits, all reported SR, SPL, and NQ numbers are optimistically biased. The paper does not state that tuning was performed on a separate validation split (such as the Ablation split used in Ablation I). Please clarify the tuning procedure, and either report results across a range of thresholds or use a nested validation to assess the sensitivity of the headline comparisons.
- [Ablation II: 'VLM uncertainty estimation on IDKVQA'] The principal evidence for the claimed superiority of Normalized-Entropy over Energy Score is a difference of 0.67 points (21.12 vs. 20.45) on the Effective Reliability metric in Table 5. No variance, confidence intervals, or significance tests are provided, and the gap is small relative to the scale of the metric. The sensitivity analysis in Fig. 3 uses normalized values and does not by itself demonstrate that the difference in the main metric is meaningful. Please report standard errors or statistical tests, and consider additional datasets or question distributions to support the claim of 'more reliable uncertainty measure.'
- [Table 2 and Sec. 6 'Results with simulated user-agent interaction'] All comparisons in Table 2 are reported as point estimates without any measure of variability. Given the stochastic components (VLFM, Grounding-DINO, LLaVA, GPT-4o), the differences between AIUTA and trained baselines on Val Unseen (SR 6.67% vs. 2.61%) and on other splits may be within run-to-run noise. The paper's central claim that AIUTA 'outperforms training-based methods on Val Unseen' would be strengthened by reporting standard errors across multiple runs or episodes, or by a statistical test for the SR differences.
minor comments (4)
- [Abstract] The sentence 'the agent actively resolve uncertainties' should be 'the agent actively resolves uncertainties.'
- [Sec. 6 'Validation with real human'] The paper reports NQ = 1.29 for real human versus 1.10 for simulated in Table 3, but no standard deviations or per-episode distributions are provided, making it impossible to judge the variability of this key metric.
- [Supp. F] The 'Distractor Success' metric is a useful diagnostic, but the text would benefit from explaining how it relates to the main SR metric, particularly whether the agent stops at a distractor because of the interaction policy or because of the navigation policy.
- [Sec. 5 'Metrics'] The definition of NQ as 'average Number of Questions asked in successful episodes' should be clarified: does it count questions across all interactions in an episode, and is it averaged only over successful episodes or over all episodes?
Circularity Check
No significant circularity: AIUTA's headline comparisons rest on an external benchmark and independent human-annotated IDKVQA evaluation, not on a self-citation or fitted-parameter chain.
full rationale
AIUTA's processing chain is not circular. The self-questioner and interaction trigger produce descriptions and alignment scores from pretrained VLM/LLM calls, and the CoIN-Bench episodes are drawn from the external GOAT-Bench evaluation splits; the trained baselines (Monolithic, PSL, OVON) and zero-shot VLFM are evaluated under the same protocol, so the reported success-rate advantage is an empirical comparison rather than a quantity reconstructed from its inputs. The Normalized-Entropy uncertainty estimator is defined by Shannon entropy normalized over a three-token vocabulary (Eq. 2) and is validated on IDKVQA with human annotations, so the comparison against MaxProb, LP, and Energy Score does not reduce to the method's own definition. The single self-citation ([46]) appears only in a related-work list and is not load-bearing. Thresholds τ, τ_stop, and τ_skip are tuned on validation data, but they are not relabeled as predictions and do not by construction determine success. The main weakness is that the simulated user is validated on only 40 episodes selected for detectable targets, and the claim of no statistical difference is not backed by a reported test; this is an external-validity concern about whether real human answers match the VLM, not a circularity in the derivation, because no equation or fitted quantity is reused as evidence for itself.
Assumptions & free parameters
free parameters (4)
- tau (uncertainty threshold) =
0.75
- tau_stop (alignment stop threshold) =
7
- tau_skip (alignment skip threshold) =
5
- max_interaction_rounds =
4
assumptions (4)
- domain assumption The user is aware of full details about the target instance and is collaborative, providing truthful responses.
- domain assumption The VLM-simulated user with access to the target image produces responses similar to real human responses.
- domain assumption Shannon entropy of the constrained three-token answer distribution is a valid measure of VLM perception uncertainty.
- domain assumption GOAT-Bench and HM3DSem provide a suitable base for evaluating interactive navigation.
Cite this review
Pith. "Pith review of Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues." pith.science (2026). https://pith.science/paper/F4KVVQOQ
@misc{pith2026241201250,
author = {Pith},
title = {Pith review of: Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues},
year = {2026},
howpublished = {\url{https://pith.science/paper/F4KVVQOQ}},
note = {Machine review of arXiv:2412.01250}
}
read the original abstract
Language-driven instance object navigation assumes that human users initiate the task by providing a detailed description of the target instance to the embodied agent. While this description is crucial for distinguishing the target from visually similar instances in a scene, providing it prior to navigation can be demanding for human. To bridge this gap, we introduce Collaborative Instance object Navigation (CoIN), a new task setting where the agent actively resolve uncertainties about the target instance during navigation in natural, template-free, open-ended dialogues with human. We propose a novel training-free method, Agent-user Interaction with UncerTainty Awareness (AIUTA), which operates independently from the navigation policy, and focuses on the human-agent interaction reasoning with Vision-Language Models (VLMs) and Large Language Models (LLMs). First, upon object detection, a Self-Questioner model initiates a self-dialogue within the agent to obtain a complete and accurate observation description with a novel uncertainty estimation technique. Then, an Interaction Trigger module determines whether to ask a question to the human, continue or halt navigation, minimizing user input. For evaluation, we introduce CoIN-Bench, with a curated dataset designed for challenging multi-instance scenarios. CoIN-Bench supports both online evaluation with humans and reproducible experiments with simulated user-agent interactions. On CoIN-Bench, we show that AIUTA serves as a competitive baseline, while existing language-driven instance navigation methods struggle in complex multi-instance scenes. Code and benchmark will be available upon acceptance at https://intelligolabs.github.io/CoIN/
Figures
Figures from the paper (7 more)
Forward citations
Cited by 2 Pith papers
-
VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs
VL-LN Bench turns instance-goal navigation into an interactive dialog task, contributes a 41k-trajectory house-scale benchmark with a GPT-4o oracle, and shows active questioning improves embodied agents' success.
-
ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People
ObjectFinder combines YOLO-World and GPT-4 in smart glasses to let blind users search for any spoken object, and an eight-person study found most users preferred it to BeMyAI and Google Lookout.
Reference graph
Works this paper leans on
-
[1]
Etpnav: Evolving topo- logical planning for vision-language navigation in continuous environments
Dong An, Hanqing Wang, Wenguan Wang, Zun Wang, Yan Huang, Keji He, and Liang Wang. Etpnav: Evolving topo- logical planning for vision-language navigation in continuous environments. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024. 2
work page 2024
-
[2]
On Evaluation of Embodied Navigation Agents
Peter Anderson, Angel Chang, Devendra Singh Chaplot, Alexey Dosovitskiy, Saurabh Gupta, Vladlen Koltun, Jana Kosecka, Jitendra Malik, Roozbeh Mottaghi, Manolis Savva, et al. On Evaluation of Embodied Navigation Agents. arXiv preprint arXiv:1807.06757, 2018. 3, 6, 4
arXiv 2018
-
[3]
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko S ¨underhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3674–3683,
-
[4]
Personalized Instance-based Navigation Toward User-Specific Objects in Realistic Envi- ronments
Luca Barsellotti, Roberto Bigazzi, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. Personalized Instance-based Navigation Toward User-Specific Objects in Realistic Envi- ronments. In The Thirty-eight Conference on Neural Infor- mation Processing Systems Datasets and Benchmarks Track,
-
[5]
ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects
Dhruv Batra, Aaron Gokaslan, Aniruddha Kembhavi, Olek- sandr Maksymets, Roozbeh Mottaghi, Manolis Savva, Alexander Toshev, and Erik Wijmans. ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects. arXiv preprint arXiv:2006.13171, 2020. 3, 6
arXiv 2006
-
[6]
Object Goal Naviga- tion using Goal-Oriented Semantic Exploration
Devendra Singh Chaplot, Dhiraj Prakashchand Gandhi, Abhi- nav Gupta, and Russ R Salakhutdinov. Object Goal Naviga- tion using Goal-Oriented Semantic Exploration. In Advances in Neural Information Processing Systems, pages 4247–4258. Curran Associates, Inc., 2020. 2
work page 2020
-
[7]
Ta-Chung Chi, Minmin Shen, Mihail Eric, Seokhwan Kim, and Dilek Hakkani-tur. Just Ask: An Interactive Learning Framework for Vision and Language Navigation.Proceedings of the AAAI Conference on Artificial Intelligence , 34(03): 2459–2466, 2020. 3
work page 2020
-
[8]
Mobility vla: Multimodal instruction navigation with long-context vlms and topological graphs
Hao-Tien Lewis Chiang, Zhuo Xu, Zipeng Fu, Mithun George Jacob, Tingnan Zhang, Tsang-Wei Edward Lee, Wenhao Yu, Connor Schenck, David Rendleman, Dhruv Shah, et al. Mobility vla: Multimodal instruction navigation with long-context vlms and topological graphs. arXiv preprint arXiv:2407.07775, 2024. 4, 5
arXiv 2024
Show all 63 references
-
[9]
Think, Act, and Ask: Open-World Interactive Personalized Robot Navi- gation
Yinpei Dai, Run Peng, Sikai Li, and Joyce Chai. Think, Act, and Ask: Open-World Interactive Personalized Robot Navi- gation. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 3296–3303, 2024. 3, 6, 1
2024
-
[10]
Spoc: Imitating shortest paths in simulation enables effective navi- gation and manipulation in the real world
Kiana Ehsani, Tanmay Gupta, Rose Hendrix, Jordi Salvador, Luca Weihs, Kuo-Hao Zeng, Kunal Pratap Singh, Yejin Kim, Winson Han, Alvaro Herrasti, Ranjay Krishna, Dustin Schwenk, Eli VanderBilt, and Aniruddha Kembhavi. Spoc: Imitating shortest paths in simulation enables effectiv...
2024
-
[11]
CoWs on Pasture: Base- lines and Benchmarks for Language-Driven Zero-Shot Object Navigation
Samir Yitzhak Gadre, Mitchell Wortsman, Gabriel Ilharco, Ludwig Schmidt, and Shuran Song. CoWs on Pasture: Base- lines and Benchmarks for Language-Driven Zero-Shot Object Navigation. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2023. 3, 7
2023
-
[12]
Alexa Arena: A User- Centric Interactive Platform for Embodied AI
Qiaozi Gao, Govind Thattai, Suhaila Shakiah, Xiaofeng Gao, Shreyas Pansare, Vasu Sharma, Gaurav Sukhatme, Hangjie Shi, Bofei Yang, Desheng Zhang, Lucy Hu, Karthika Aru- mugam, Shui Hu, Matthew Wen, Dinakar Guthy, Shunan Chung, Rohan Khanna, Osman Ipek, Leslie Ball, Kate Bland,...
2023
-
[13]
Groq - Accelerated AI Inference
Groq. Groq - Accelerated AI Inference. https://groq. com/, 2024. Accessed: Mar. 7, 2025. 5
2024
-
[14]
Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P
Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P. Bigham. VizWiz Grand Challenge: Answering Visual Questions from Blind People. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 2018. 6
2018
-
[15]
Gpt-4o system card
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024. 7
2024 arXiv
-
[16]
Survey of Hallucination in Natural Language Generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of Hallucination in Natural Language Generation. ACM Comput. Surv., 55(12), 2023. 3
2023
-
[17]
GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation
Mukul Khanna, Ram Ramrakhya, Gunjan Chhablani, Sriram Yenamandra, Theophile Gervet, Matthew Chang, Zsolt Kira, Devendra Singh Chaplot, Dhruv Batra, and Roozbeh Mot- taghi. GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation. In CVPR, page 16373–16383. IEEE, 2024. 2, 3,...
2024
-
[18]
Instance-Specific Image Goal Navi- gation: Training Embodied Agents to Find Object Instances
Jacob Krantz, Stefan Lee, Jitendra Malik, Dhruv Batra, and Devendra Singh Chaplot. Instance-Specific Image Goal Navi- gation: Training Embodied Agents to Find Object Instances. arXiv preprint arXiv:2211.15876, 2022. 3, 6
2022 arXiv
-
[19]
OpenFMNav: Towards Open-Set Zero-Shot Object Navigation via Vision- Language Foundation Models
Yuxuan Kuang, Hai Lin, and Meng Jiang. OpenFMNav: Towards Open-Set Zero-Shot Object Navigation via Vision- Language Foundation Models. In Findings of the Association for Computational Linguistics: NAACL 2024, pages 338–351, Mexico City, Mexico, 2024. Association for Computatio...
2024
-
[20]
BLIP- 2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP- 2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In Proceedings 9 of the 40th International Conference on Machine Learning. JMLR.org, 2023. 4, 1
2023
-
[21]
ION: Instance-level Object Navigation
Weijie Li, Xinhang Song, Yubing Bai, Sixian Zhang, and Shuqiang Jiang. ION: Instance-level Object Navigation. In ACM MM, pages 4343–4352. ACM, 2021. 2
2021
-
[22]
Visual Instruction Tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual Instruction Tuning. In Advances in Neural Information Processing Systems, pages 34892–34916. Curran Associates, Inc., 2023. 5
2023
-
[23]
LLaV A-NeXT: Improved reasoning, OCR, and world knowledge, 2024
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. LLaV A-NeXT: Improved reasoning, OCR, and world knowledge, 2024. 7
2024
-
[24]
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection. arXiv preprint arXiv:2303.05499, 2023. 4
2023 arXiv
-
[25]
Paying More At- tention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs
Shi Liu, Kecheng Zheng, and Wei Chen. Paying More At- tention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs. In Computer Vision - ECCV 2024. Springer Nature Switzerland, 2025. 2, 3, 4, 5
2024
-
[26]
Energy-based Out-of-distribution Detection
Weitang Liu, Xiaoyun Wang, John Owens, and Yixuan Li. Energy-based Out-of-distribution Detection. In Advances in Neural Information Processing Systems, pages 21464–21475. Curran Associates, Inc., 2020. 2, 7, 8, 6
2020
-
[27]
CA VEN: An Embodied Conversational Agent for Efficient Audio-Visual Navigation in Noisy Environments
Xiulong Liu, Sudipta Paul, Moitreya Chatterjee, and Anoop Cherian. CA VEN: An Embodied Conversational Agent for Efficient Audio-Visual Navigation in Noisy Environments. Proceedings of the AAAI Conference on Artificial Intelligence, 38(4):3765–3773, 2024. 3
2024
-
[28]
ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings
Arjun Majumdar, Gunjan Aggarwal, Bhavika Devnani, Judy Hoffman, and Dhruv Batra. ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings. In Advances in Neural Information Processing Systems , pages 32340– 32352. Curran Associates, Inc., 2022. 3, 4
2022
-
[29]
FindThis: Language-Driven Object Dis- ambiguation in Indoor Environments
Arjun Majumdar, Fei Xia, Brian Ichter, Dhruv Batra, and Leonidas Guibas. FindThis: Language-Driven Object Dis- ambiguation in Indoor Environments. In Proceedings of The 7th Conference on Robot Learning, pages 1335–1347. PMLR,
-
[30]
UMAP: Uniform Manifold Approximation and Projection
Leland McInnes, John Healy, Nathaniel Saul, and Lukas Großberger. UMAP: Uniform Manifold Approximation and Projection. Journal of Open Source Software , 3(29):861,
-
[31]
Help, Anna! Visual Navi- gation with Natural Multimodal Assistance via Retrospective Curiosity-Encouraging Imitation Learning
Khanh Nguyen and Hal Daum´e III. Help, Anna! Visual Navi- gation with Natural Multimodal Assistance via Retrospective Curiosity-Encouraging Imitation Learning. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International J...
2019
-
[32]
Vision-Based Navigation With Language- Based Assistance via Imitation Learning With Indirect In- tervention
Khanh Nguyen, Debadeepta Dey, and Bill Brockett, Chrnd Dolan. Vision-Based Navigation With Language- Based Assistance via Imitation Learning With Indirect In- tervention. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2019. 3
2019
-
[33]
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
Hadas Orgad, Michael Toker, Zorik Gekhman, Roi Re- ichart, Idan Szpektor, Hadas Kotek, and Yonatan Belinkov. LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations. arXiv preprint arXiv:2410.02707, 2024. 3
-
[34]
A VLEN: Audio-Visual-Language Embodied Navigation in 3D Environments
Sudipta Paul, Amit Roy-Chowdhury, and Anoop Cherian. A VLEN: Audio-Visual-Language Embodied Navigation in 3D Environments. In Advances in Neural Information Pro- cessing Systems, pages 6236–6249. Curran Associates, Inc.,
-
[35]
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
Yusu Qian, Haotian Zhang, Yinfei Yang, and Zhe Gan. How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts. In Neurips Safe Generative AI Workshop 2024, 2024. 2, 3, 4, 5
2024
-
[36]
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of th...
2021
-
[37]
Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI
Santhosh Kumar Ramakrishnan, Aaron Gokaslan, Erik Wij- mans, Oleksandr Maksymets, Alexander Clegg, John Turner, Eric Undersander, Wojciech Galuba, Andrew Westbury, An- gel Chang, Manolis Savva, Yili Zhao, and Dhruv Batra. Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale ...
2021
-
[38]
PIRLNav: Pretraining with Imitation and RL Finetuning for OBJECTNA V
Ram Ramrakhya, Dhruv Batra, Erik Wijmans, and Abhishek Das. PIRLNav: Pretraining with Imitation and RL Finetuning for OBJECTNA V . In2023 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR). IEEE, 2023. 3
2023
-
[39]
UNMuTe: Unifying Navigation and Mul- timodal Dialogue-like Text Generation
Niyati Rawal, Roberto Bigazzi, Lorenzo Baraldi, and Rita Cucchiara. UNMuTe: Unifying Navigation and Mul- timodal Dialogue-like Text Generation. arXiv preprint arXiv:2408.04423, 2024. 3
2024 arXiv
-
[40]
”Sentence-BERT: Sen- tence Embeddings using Siamese BERT-Networks”
Nils Reimers and Iryna Gurevych. ”Sentence-BERT: Sen- tence Embeddings using Siamese BERT-Networks”. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP- IJCN...
2019
-
[41]
Allen Z. Ren, Anushri Dixit, Alexandra Bodrova, Sumeet Singh, Stephen Tu, Noah Brown, Peng Xu, Leila Takayama, Fei Xia, Jake Varley, Zhenjia Xu, Dorsa Sadigh, Andy Zeng, and Anirudha Majumdar. Robots That Ask For Help: Uncer- tainty Alignment for Large Language Model Planners....
2023
-
[42]
Habitat: A Platform for Embodied AI Research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra. Habitat: A Platform for Embodied AI Research. In Proceedings of the IEEE/CVF International Conferen...
2019
-
[43]
A mathematical theory of commu- nication
Claude Elwood Shannon. A mathematical theory of commu- nication. The Bell system technical journal, 27(3):379–423,
-
[44]
Ask4Help: Learning to Leverage an Expert for Embodied Tasks
Kunal Pratap Singh, Luca Weihs, Alvaro Herrasti, Jonghyun Choi, Aniruddha Kembhavi, and Roozbeh Mottaghi. Ask4Help: Learning to Leverage an Expert for Embodied Tasks. In Advances in Neural Information Processing Sys- tems, pages 16221–16232. Curran Associates, Inc., 2022. 3
2022
-
[45]
Prioritized Semantic Learning for Zero-shot Instance Navigation
Xander Sun, Louis Lau, Hoyard Zhi, Ronghe Qiu, and Junwei Liang. Prioritized Semantic Learning for Zero-shot Instance Navigation. In Computer Vision - ECCV 2024 . Springer Nature Switzerland, 2025. 3, 7, 4
2024
-
[46]
Mind the Error! Detection and Localiza- tion of Instruction Errors in Vision-and-Language Navigation
Francesco Taioli, Stefano Rosa, Alberto Castellini, Lorenzo Natale, Alessio Del Bue, Alessandro Farinelli, Marco Cristani, and Yiming Wang. Mind the Error! Detection and Localiza- tion of Instruction Errors in Vision-and-Language Navigation. In 2024 IEEE/RSJ International Conf...
2024
-
[47]
Vision-and-Dialog Navigation
Jesse Thomason, Michael Murray, Maya Cakmak, and Luke Zettlemoyer. Vision-and-Dialog Navigation. In Proceedings of the Conference on Robot Learning, pages 394–406. PMLR,
-
[48]
Eyes Wide Shut? Exploring the Vi- sual Shortcomings of Multimodal LLMs
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie. Eyes Wide Shut? Exploring the Vi- sual Shortcomings of Multimodal LLMs. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), page 9568–9578. IEEE, 2024. 2, 3, 4, 5
2024
-
[49]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkor- eit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is All you Need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, ...
2017
-
[50]
Reliable Visual Question Answering: Abstain Rather Than Answer Incorrectly
Spencer Whitehead, Suzanne Petryk, Vedaad Shakib, Joseph Gonzalez, Trevor Darrell, Anna Rohrbach, and Marcus Rohrbach. Reliable Visual Question Answering: Abstain Rather Than Answer Incorrectly. In Computer Vision – ECCV 2022, page 148–166. Springer Nature Switzerland, 2022. 8, 5
2022
-
[51]
Habitat Challenge 2023, 2023
Karmesh Yadav, Jacob Krantz, Ram Ramrakhya, San- thosh Kumar Ramakrishnan, Jimmy Yang, Austin Wang, John Turner, Aaron Gokaslan, Vincent-Pierre Berges, Roozbeh Mootaghi, Oleksandr Maksymets, Angel X Chang, Manolis Savva, Alexander Clegg, Devendra Singh Chaplot, and Dhruv Batra...
2023
-
[52]
Yamauchi
B. Yamauchi. A frontier-based approach for autonomous exploration. In Proceedings 1997 IEEE International Sympo- sium on Computational Intelligence in Robotics and Automa- tion CIRA’97. ’Towards New Computational Principles for Robotics and Automation’, pages 146–151, 1997. 3
1997
-
[53]
VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation
Naoki Yokoyama, Sehoon Ha, Dhruv Batra, Jiuguang Wang, and Bernadette Bucher. VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation. In 2024 IEEE In- ternational Conference on Robotics and Automation (ICRA), page 42–48. IEEE, 2024. 3, 4, 7, 8
2024
-
[54]
HM3D-OVON: A Dataset and Bench- mark for Open-V ocabulary Object Goal Navigation
Naoki Yokoyama, Ram Ramrakhya, Abhishek Das, Dhruv Batra, and Sehoon Ha. HM3D-OVON: A Dataset and Bench- mark for Open-V ocabulary Object Goal Navigation. In2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 5543–5550, 2024. 3, 7, 1, 4
2024
-
[55]
L3MVN: Leveraging Large Language Models for Visual Target Navi- gation
Bangguo Yu, Hamidreza Kasaei, and Ming Cao. L3MVN: Leveraging Large Language Models for Visual Target Navi- gation. In 2023 IEEE/RSJ International Conference on Intel- ligent Robots and Systems (IROS). IEEE, 2023. 3
2023
-
[56]
Sigmoid Loss for Language Image Pre-Training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid Loss for Language Image Pre-Training. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023, pages 11941– 11952. IEEE, 2023. 4
2023
-
[57]
Faster Segment Anything: Towards Lightweight SAM for Mobile Applications
Chaoning Zhang, Dongshen Han, Yu Qiao, Jung Uk Kim, Sung-Ho Bae, Seungkyu Lee, and Choong Seon Hong. Faster Segment Anything: Towards Lightweight SAM for Mobile Applications. arXiv preprint arXiv:2306.14289, 2023. 4
2023 arXiv
-
[58]
TriHelper: Zero-Shot Object Navigation with Dynamic Assistance
Lingfeng Zhang, Qiang Zhang, Hao Wang, Erjia Xiao, Zixuan Jiang, Honglei Chen, and Renjing Xu. TriHelper: Zero-Shot Object Navigation with Dynamic Assistance. arXiv preprint arXiv:2403.15223, 2024. 3
2024 arXiv
-
[59]
The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision- Language Models? In Computer Vision - ECCV 2024
Qinyu Zhao, Ming Xu, Kartik Gupta, Akshay Asthana, Liang Zheng, and Stephen Gould. The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision- Language Models? In Computer Vision - ECCV 2024 . Springer Nature Switzerland, 2025. 3, 7, 8, 6
2024
-
[60]
ESC: Explo- ration with Soft Commonsense Constraints for Zero-shot Ob- ject Navigation
Kaiwen Zhou, Kaizhi Zheng, Connor Pryor, Yilin Shen, Hongxia Jin, Lise Getoor, and Xin Eric Wang. ESC: Explo- ration with Soft Commonsense Constraints for Zero-shot Ob- ject Navigation. In Proceedings of the 40th International Con- ference on Machine Learning, pages 42829–42842. PMLR,
-
[61]
ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions
Deyao Zhu, Jun Chen, Kilichbek Haydarov, Xiaoqian Shen, Wenxuan Zhang, and Mohamed Elhoseiny. ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions. Transactions on Machine Learning Re- search, 2024. 3
2024
-
[62]
Lim, Abhinav Gupta, Li Fei-Fei, and Ali Farhadi
Yuke Zhu, Roozbeh Mottaghi, Eric Kolve, Joseph J. Lim, Abhinav Gupta, Li Fei-Fei, and Ali Farhadi. Target-driven Visual Navigation in Indoor Scenes using Deep Reinforce- ment Learning. In 2017 IEEE International Conference on Robotics and Automation (ICRA) , page 3357–3364. IEEE,
2017
-
[2017]
cabinet”, “bed
2, 3 11 Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues Supplementary Material In this supplementary material, we first provide additional details regarding the CoIN-Bench dataset (Sec. A), including an overview of t...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.