Pith. sign in

REVIEW 10 cited by

Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.07954 v1 pith:5W4G55OH submitted 2020-10-15 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords multilingualinstructionlanguagenavigationpathsroom-across-roomvision-and-languageaddressing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Room-Across-Room (RxR), a new Vision-and-Language Navigation (VLN) dataset. RxR is multilingual (English, Hindi, and Telugu) and larger (more paths and instructions) than other VLN datasets. It emphasizes the role of language in VLN by addressing known biases in paths and eliciting more references to visible entities. Furthermore, each word in an instruction is time-aligned to the virtual poses of instruction creators and validators. We establish baseline scores for monolingual and multilingual settings and multitask learning when including Room-to-Room annotations. We also provide results for a model that learns from synchronized pose traces by focusing only on portions of the panorama attended to in human demonstrations. The size, scope and detail of RxR dramatically expands the frontier for research on embodied language agents in simulated, photo-realistic environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation

    cs.CV 2026-02 conditional novelty 7.0 of 10

    LangMap is a human-verified navigation benchmark with 18K tasks spanning scene-, room-, region-, and instance-level goals in real-world 3D scans, covering 414 object categories.

  2. Joint On-and-Off Policy Learning for Vision-and-Language Navigation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    JOP-VLN combines DAgger imitation learning with GRPO reinforcement learning, using high-entropy trajectory filtering and error-correction prioritization, achieving 69.9% SR on R2R Val-Unseen.

  3. Structured Observation Language for Efficient and Generalizable Vision-Language Navigation

    cs.CV 2026-03 reject novelty 6.0 of 10

    SOL-Nav encodes RGB-D observations as grid-organized text and uses a 0.6B text-embedding model with four classification heads to predict navigation action blocks, reporting SOTA/comparable R2R-CE/RxR-CE results with 1...

  4. Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models

    cs.RO 2026-02 conditional novelty 6.0 of 10

    Adding a real-world supervised loss to simulation reinforcement learning improves real-robot success and data efficiency for VLA co-training.

  5. SkeNa: Learning to Navigate Unseen Environments Based on Abstract Hand-Drawn Maps

    cs.RO 2025-08 unverdicted novelty 6.0 of 10

    A navigation agent can follow abstract hand-drawn sketch maps to reach goals in unseen indoor environments, backed by a new 54k-pair dataset and a model with a 105 percent relative SPL gain.

  6. Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A VLN agent that recursively imagines future views and layouts in a fixed-size neural grid, and adaptively aligns instruction parts to grid cells, achieves state-of-the-art success rates on R2R-CE and ObjectNav.

  7. StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A streaming navigation framework combining a sliding-window KV cache with depth-based token pruning achieves state-of-the-art results on VLN-CE benchmarks with bounded context and low latency.

  8. SkyVLN: Vision-and-Language Navigation and NMPC Control for UAVs in Urban Environments

    cs.RO 2025-07 conditional novelty 5.0 of 10

    An LLM-and-NMPC drone navigation framework that reports 42.4% success on unseen AVDN test data, versus 16.6% for NavGPT, using spatial verbalization and a path memory graph.

  9. Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A dual-branch text-imagination system, with one LLM branch for state estimation and one for candidate-direction description, improves R2R navigation success over prior LLM-based VLN methods.

  10. Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey

    cs.CV 2025-07 conditional novelty 2.0 of 10

    A survey organizing feature matching research by modality, from SIFT to transformer-based dense matchers and vision-language models.

Pith tools