A zero-shot navigation pipeline using an LLM to extract landmarks, SigLIP to find goal panoramas, and GPT-4o grounding plus dynamic programming to rank paths achieves 88.9% nDTW on R2R-Habitat and 70% Precision@10 for landmark retrieval.
Speaker-follower models for vision-and- language navigation.Advances in neural information processing systems, 31, 2018
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation
A zero-shot navigation pipeline using an LLM to extract landmarks, SigLIP to find goal panoramas, and GPT-4o grounding plus dynamic programming to rank paths achieves 88.9% nDTW on R2R-Habitat and 70% Precision@10 for landmark retrieval.