Pith. sign in

REVIEW 1 cited by

Robust Navigation with Language Pretraining and Stochastic Sampling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.02244 v1 pith:DHHTIB3E submitted 2019-09-05 cs.CL cs.CVcs.LG

classification cs.CLcs.CVcs.LG
keywords actionactionsdecodinggeneralizeinstructionslanguagelearnnavigation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Core to the vision-and-language navigation (VLN) challenge is building robust instruction representations and action decoding schemes, which can generalize well to previously unseen instructions and environments. In this paper, we report two simple but highly effective methods to address these challenges and lead to a new state-of-the-art performance. First, we adapt large-scale pretrained language models to learn text representations that generalize better to previously unseen instructions. Second, we propose a stochastic sampling scheme to reduce the considerable gap between the expert actions in training and sampled actions in test, so that the agent can learn to correct its own mistakes during long sequential action decoding. Combining the two techniques, we achieve a new state of the art on the Room-to-Room benchmark with 6% absolute gain over the previous best result (47% -> 53%) on the Success Rate weighted by Path Length metric.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Active Test-time Vision-Language Navigation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    ATENA uses episodic success/failure labels and a mixture entropy objective to adapt vision-language navigation policies at test time, improving REVERIE, R2R, and R2R-CE benchmarks.

Pith tools