Pith. sign in

REVIEW 10 cited by

General Evaluation for Instruction Conditioned Navigation using Dynamic Time Warping

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.05446 v2 pith:WKZ2KIJK submitted 2019-07-11 cs.RO cs.AIcs.CL

classification cs.ROcs.AIcs.CL
keywords metricsnavigationndtwtimeagentsdynamicevaluationsimilarity
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In instruction conditioned navigation, agents interpret natural language and their surroundings to navigate through an environment. Datasets for studying this task typically contain pairs of these instructions and reference trajectories. Yet, most evaluation metrics used thus far fail to properly account for the latter, relying instead on insufficient similarity comparisons. We address fundamental flaws in previously used metrics and show how Dynamic Time Warping (DTW), a long known method of measuring similarity between two time series, can be used for evaluation of navigation agents. For such, we define the normalized Dynamic Time Warping (nDTW) metric, that softly penalizes deviations from the reference path, is naturally sensitive to the order of the nodes composing each path, is suited for both continuous and graph-based evaluations, and can be efficiently calculated. Further, we define SDTW, which constrains nDTW to only successful paths. We collect human similarity judgments for simulated paths and find nDTW correlates better with human rankings than all other metrics. We also demonstrate that using nDTW as a reward signal for Reinforcement Learning navigation agents improves their performance on both the Room-to-Room (R2R) and Room-for-Room (R4R) datasets. The R4R results in particular highlight the superiority of SDTW over previous success-constrained metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Goal-oriented Navigation Instruction Generation with Tour Video Priors

    cs.CV 2026-08 conditional novelty 6.0 of 10

    VideoNIG tests whether multimodal models can turn tour videos into executable navigation instructions, and a two-stage curriculum with preference optimization improves their spatial reasoning.

  2. GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation

    cs.RO 2025-09 conditional novelty 6.0 of 10

    GC-VLN decomposes a navigation instruction into a graph of spatial constraints, solves the constraints with an optimizer, and beats prior zero-shot methods on VLN-CE benchmarks without any training.

  3. CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model

    cs.RO 2025-08 unverdicted novelty 6.0 of 10

    By iteratively retraining on automatically generated corrective examples derived from its own wrong paths, CorrectNav reports new state-of-the-art success rates of 65.1% (R2R-CE) and 69.3% (RxR-CE).

  4. NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments

    cs.CV 2025-06 conditional novelty 6.0 of 10

    NavMorph combines an RSSM-based latent world model with an online-updated contextual memory, reporting consistent VLN-CE gains on R2R-CE and RxR-CE.

  5. Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Iterative navigator-generator collaboration, where the navigator filters generated instructions and the rebuilt generator rewrites low-quality ones, raises R2R navigation SPL to 78% and instruction SPICE to 26.2.

  6. SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts

    cs.CV 2024-12 conditional novelty 6.0 of 10

    One navigation model with state-adaptive mixture-of-experts routing matches or exceeds task-specific agents on several of seven navigation benchmarks.

  7. Hijacking Vision-and-Language Navigation Agents with Adversarial Environmental Attacks

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A whitebox adversarial attack that repaints a single 3D object can redirect or stop a pretrained Vision-and-Language Navigation agent on unseen instructions.

  8. Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation

    cs.CV 2026-07 reject novelty 4.0 of 10

    A CP-decomposed, instruction-conditioned latent bottleneck (CompactNav) improves VLN-CE success rate by about 2% over prior state of the art on two benchmarks.

  9. Nav-R1: Reasoning and Navigation in Embodied Scenes

    cs.RO 2025-09 reject novelty 4.0 of 10

    Nav-R1 uses a 110K synthetic CoT dataset, GRPO with three rewards, and a fast-in-slow system to set new SOTA on R2R-CE, RxR-CE, and HM3D-OVON.

  10. A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI

    cs.RO 2025-05 conditional novelty 2.0 of 10

    A review of navigation and manipulation simulators, datasets, and methods, framed around the sim-to-real gap.

Pith tools