Pith. sign in

REVIEW 3 cited by

Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1811.10092 v2 pith:CB2FYJEM submitted 2018-11-25 cs.CV cs.AIcs.CLcs.RO

classification cs.CVcs.AIcs.CLcs.RO
keywords cross-modalmatchingenvironmentsgroundinglearningimitationinstructionsnavigation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Vision-language navigation (VLN) is the task of navigating an embodied agent to carry out natural language instructions inside real 3D environments. In this paper, we study how to address three critical challenges for this task: the cross-modal grounding, the ill-posed feedback, and the generalization problems. First, we propose a novel Reinforced Cross-Modal Matching (RCM) approach that enforces cross-modal grounding both locally and globally via reinforcement learning (RL). Particularly, a matching critic is used to provide an intrinsic reward to encourage global matching between instructions and trajectories, and a reasoning navigator is employed to perform cross-modal grounding in the local visual scene. Evaluation on a VLN benchmark dataset shows that our RCM model significantly outperforms previous methods by 10% on SPL and achieves the new state-of-the-art performance. To improve the generalizability of the learned policy, we further introduce a Self-Supervised Imitation Learning (SIL) method to explore unseen environments by imitating its own past, good decisions. We demonstrate that SIL can approximate a better and more efficient policy, which tremendously minimizes the success rate performance gap between seen and unseen environments (from 30.7% to 11.7%).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. World-Consistent Data Generation for Vision-and-Language Navigation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A 3D-guided data augmentation method generates world-consistent panoramic training data that improves a VLN agent's performance on unseen environments over prior augmentation baselines.

  2. Transferable Representation Learning in Vision-and-Language Navigation

    cs.CV 2019-08 reject novelty 6.0 of 10

    Auxiliary cross-modal alignment and future-scene prediction pretraining is claimed to improve VLN agents, but the paper's own ablations show no benefit over no-pretraining baselines.

  3. Walking with MIND: Mental Imagery eNhanceD Embodied QA

    cs.CV 2019-08 conditional novelty 6.0 of 10

    A mental imagery module that predicts future views and treats them as short-term subgoals improves an embodied agent's navigation and question-answering accuracy in simulation.

Pith tools