REVIEW 10 cited by
Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce Room-Across-Room (RxR), a new Vision-and-Language Navigation (VLN) dataset. RxR is multilingual (English, Hindi, and Telugu) and larger (more paths and instructions) than other VLN datasets. It emphasizes the role of language in VLN by addressing known biases in paths and eliciting more references to visible entities. Furthermore, each word in an instruction is time-aligned to the virtual poses of instruction creators and validators. We establish baseline scores for monolingual and multilingual settings and multitask learning when including Room-to-Room annotations. We also provide results for a model that learns from synchronized pose traces by focusing only on portions of the panorama attended to in human demonstrations. The size, scope and detail of RxR dramatically expands the frontier for research on embodied language agents in simulated, photo-realistic environments.
Forward citations
Cited by 10 Pith papers
-
LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation
LangMap is a human-verified navigation benchmark with 18K tasks spanning scene-, room-, region-, and instance-level goals in real-world 3D scans, covering 414 object categories.
-
Joint On-and-Off Policy Learning for Vision-and-Language Navigation
JOP-VLN combines DAgger imitation learning with GRPO reinforcement learning, using high-entropy trajectory filtering and error-correction prioritization, achieving 69.9% SR on R2R Val-Unseen.
-
Structured Observation Language for Efficient and Generalizable Vision-Language Navigation
SOL-Nav encodes RGB-D observations as grid-organized text and uses a 0.6B text-embedding model with four classification heads to predict navigation action blocks, reporting SOTA/comparable R2R-CE/RxR-CE results with 1...
-
Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models
Adding a real-world supervised loss to simulation reinforcement learning improves real-robot success and data efficiency for VLA co-training.
-
SkeNa: Learning to Navigate Unseen Environments Based on Abstract Hand-Drawn Maps
A navigation agent can follow abstract hand-drawn sketch maps to reach goals in unseen indoor environments, backed by a new 54k-pair dataset and a model with a 105 percent relative SPL gain.
-
Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation
A VLN agent that recursively imagines future views and layouts in a fixed-size neural grid, and adaptively aligns instruction parts to grid cells, achieves state-of-the-art success rates on R2R-CE and ObjectNav.
-
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
A streaming navigation framework combining a sliding-window KV cache with depth-based token pruning achieves state-of-the-art results on VLN-CE benchmarks with bounded context and low latency.
-
SkyVLN: Vision-and-Language Navigation and NMPC Control for UAVs in Urban Environments
An LLM-and-NMPC drone navigation framework that reports 42.4% success on unseen AVDN test data, versus 16.6% for NavGPT, using spatial verbalization and a path memory graph.
-
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
A dual-branch text-imagination system, with one LLM branch for state estimation and one for candidate-direction description, improves R2R navigation success over prior LLM-based VLN methods.
-
Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey
A survey organizing feature matching research by modality, from SIFT to transformer-based dense matchers and vision-language models.
Discussion (0). Continue with ORCID to comment.