REVIEW 10 cited by
General Evaluation for Instruction Conditioned Navigation using Dynamic Time Warping
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In instruction conditioned navigation, agents interpret natural language and their surroundings to navigate through an environment. Datasets for studying this task typically contain pairs of these instructions and reference trajectories. Yet, most evaluation metrics used thus far fail to properly account for the latter, relying instead on insufficient similarity comparisons. We address fundamental flaws in previously used metrics and show how Dynamic Time Warping (DTW), a long known method of measuring similarity between two time series, can be used for evaluation of navigation agents. For such, we define the normalized Dynamic Time Warping (nDTW) metric, that softly penalizes deviations from the reference path, is naturally sensitive to the order of the nodes composing each path, is suited for both continuous and graph-based evaluations, and can be efficiently calculated. Further, we define SDTW, which constrains nDTW to only successful paths. We collect human similarity judgments for simulated paths and find nDTW correlates better with human rankings than all other metrics. We also demonstrate that using nDTW as a reward signal for Reinforcement Learning navigation agents improves their performance on both the Room-to-Room (R2R) and Room-for-Room (R4R) datasets. The R4R results in particular highlight the superiority of SDTW over previous success-constrained metrics.
Forward citations
Cited by 10 Pith papers
-
Goal-oriented Navigation Instruction Generation with Tour Video Priors
VideoNIG tests whether multimodal models can turn tour videos into executable navigation instructions, and a two-stage curriculum with preference optimization improves their spatial reasoning.
-
GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation
GC-VLN decomposes a navigation instruction into a graph of spatial constraints, solves the constraints with an optimizer, and beats prior zero-shot methods on VLN-CE benchmarks without any training.
-
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
By iteratively retraining on automatically generated corrective examples derived from its own wrong paths, CorrectNav reports new state-of-the-art success rates of 65.1% (R2R-CE) and 69.3% (RxR-CE).
-
NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments
NavMorph combines an RSSM-based latent world model with an online-updated contextual memory, reporting consistent VLN-CE gains on R2R-CE and RxR-CE.
-
Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
Iterative navigator-generator collaboration, where the navigator filters generated instructions and the rebuilt generator rewrites low-quality ones, raises R2R navigation SPL to 78% and instruction SPICE to 26.2.
-
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
One navigation model with state-adaptive mixture-of-experts routing matches or exceeds task-specific agents on several of seven navigation benchmarks.
-
Hijacking Vision-and-Language Navigation Agents with Adversarial Environmental Attacks
A whitebox adversarial attack that repaints a single 3D object can redirect or stop a pretrained Vision-and-Language Navigation agent on unseen instructions.
-
Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation
A CP-decomposed, instruction-conditioned latent bottleneck (CompactNav) improves VLN-CE success rate by about 2% over prior state of the art on two benchmarks.
-
Nav-R1: Reasoning and Navigation in Embodied Scenes
Nav-R1 uses a 110K synthetic CoT dataset, GRPO with three rewards, and a fast-in-slow system to set new SOTA on R2R-CE, RxR-CE, and HM3D-OVON.
-
A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI
A review of navigation and manipulation simulators, datasets, and methods, framed around the sim-to-real gap.
Discussion (0). Continue with ORCID to comment.