REVIEW 4 cited by
Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
A long-term goal of AI research is to build intelligent agents that can communicate with humans in natural language, perceive the environment, and perform real-world tasks. Vision-and-Language Navigation (VLN) is a fundamental and interdisciplinary research topic towards this goal, and receives increasing attention from natural language processing, computer vision, robotics, and machine learning communities. In this paper, we review contemporary studies in the emerging field of VLN, covering tasks, evaluation metrics, methods, etc. Through structured analysis of current progress and challenges, we highlight the limitations of current VLN and opportunities for future work. This paper serves as a thorough reference for the VLN research community.
Forward citations
Cited by 4 Pith papers
-
When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents
Spatial-memory staleness is a measurable safety failure for VLM agents: stale memory increases deaths, and visual auditing of stale entries is highly model-dependent.
-
ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension
ReMeREC introduces a relation-aware multi-entity referring expression comprehension framework and the ReMeX dataset, reporting state-of-the-art grounding and relation prediction, with some evaluation caveats.
-
Active Test-time Vision-Language Navigation
ATENA uses episodic success/failure labels and a mixture entropy objective to adapt vision-language navigation policies at test time, improving REVERIE, R2R, and R2R-CE benchmarks.
-
Text-guided Generation of Efficient Personalized Inspection Plans
A training-free pipeline uses a vision-language model and segmentation to convert text instructions into smooth, order-respecting drone inspection trajectories in known 3D maps.
Discussion (0). Sign in to comment.