REVIEW 7 cited by
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Vision-and-Language Navigation (VLN) has gained increasing attention over recent years and many approaches have emerged to advance their development. The remarkable achievements of foundation models have shaped the challenges and proposed methods for VLN research. In this survey, we provide a top-down review that adopts a principled framework for embodied planning and reasoning, and emphasizes the current methods and future opportunities leveraging foundation models to address VLN challenges. We hope our in-depth discussions could provide valuable resources and insights: on one hand, to milestone the progress and explore opportunities and potential roles for foundation models in this field, and on the other, to organize different challenges and solutions in VLN to foundation model researchers.
Forward citations
Cited by 7 Pith papers
-
From Region Arrival to Instance-Level Grounding in Vision-and-Language Navigation
REALM, a visibility-aware plug-and-play last-meters module trained on the new REVERIE-AIM dataset, consistently raises instance proximity and grounding success on four VLN backbones.
-
3D-Aware VLMs with Implicit and Explicit Geometries
Fusing a VLM's normal 2D tokens with implicit geometry tokens from a video-geometry encoder plus tokens from its own reconstructed depth maps improves 3D detection, grounding, captioning, and spatial reasoning.
-
AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation
AdvNav is a gradient-free attack that overlays Perlin noise on a VLN agent's camera and uses behavior feedback plus genetic search, breaking 49.70-87.30% of successful R2R navigations.
-
OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding
OpenGround grounds open-world 3D targets by planning a task chain and dynamically expanding the object lookup table through online 2D segmentation and 3D lifting, achieving SOTA zero-shot ScanRefer accuracy and 46.2% ...
-
TrackVLA: Embodied Visual Tracking in the Wild
A single vision-language-action model jointly trained on recognition and tracking data follows described targets at the best reported levels on a public benchmark and transfers zero-shot from simulation to a real quad...
-
CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process Thinking
CogDDN uses a fast heuristic VLM paired with a slow analytic reflection process and a growing knowledge base to navigate to objects that implicitly satisfy a user's demand, with large reported gains on AI2Thor DDN benchmarks.
-
When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs
A survey that classifies VLM attacks by goal and data manipulation strategy, and reviews defenses and metrics.
Discussion (0). Continue with ORCID to comment.