A zero-shot navigation pipeline using an LLM to extract landmarks, SigLIP to find goal panoramas, and GPT-4o grounding plus dynamic programming to rank paths achieves 88.9% nDTW on R2R-Habitat and 70% Precision@10 for landmark retrieval.
Language models are few-shot learners.Ad- vances in neural information processing systems, 33: 1877–1901, 2020
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation
A zero-shot navigation pipeline using an LLM to extract landmarks, SigLIP to find goal panoramas, and GPT-4o grounding plus dynamic programming to rank paths achieves 88.9% nDTW on R2R-Habitat and 70% Precision@10 for landmark retrieval.