Pith. sign in

REVIEW 7 cited by

Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.07035 v2 pith:FPIMSWZB submitted 2024-07-09 cs.CL cs.CV

classification cs.CLcs.CV
keywords foundationmodelschallengesmethodsnavigationopportunitiessurveyvision-and-language
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-and-Language Navigation (VLN) has gained increasing attention over recent years and many approaches have emerged to advance their development. The remarkable achievements of foundation models have shaped the challenges and proposed methods for VLN research. In this survey, we provide a top-down review that adopts a principled framework for embodied planning and reasoning, and emphasizes the current methods and future opportunities leveraging foundation models to address VLN challenges. We hope our in-depth discussions could provide valuable resources and insights: on one hand, to milestone the progress and explore opportunities and potential roles for foundation models in this field, and on the other, to organize different challenges and solutions in VLN to foundation model researchers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Region Arrival to Instance-Level Grounding in Vision-and-Language Navigation

    cs.RO 2026-07 conditional novelty 6.5 of 10

    REALM, a visibility-aware plug-and-play last-meters module trained on the new REVERIE-AIM dataset, consistently raises instance proximity and grounding success on four VLN backbones.

  2. 3D-Aware VLMs with Implicit and Explicit Geometries

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Fusing a VLM's normal 2D tokens with implicit geometry tokens from a video-geometry encoder plus tokens from its own reconstructed depth maps improves 3D detection, grounding, captioning, and spatial reasoning.

  3. AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation

    cs.AI 2026-07 conditional novelty 6.0 of 10

    AdvNav is a gradient-free attack that overlays Perlin noise on a VLN agent's camera and uses behavior feedback plus genetic search, breaking 49.70-87.30% of successful R2R navigations.

  4. OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding

    cs.CV 2025-12 conditional novelty 6.0 of 10

    OpenGround grounds open-world 3D targets by planning a task chain and dynamically expanding the object lookup table through online 2D segmentation and 3D lifting, achieving SOTA zero-shot ScanRefer accuracy and 46.2% ...

  5. TrackVLA: Embodied Visual Tracking in the Wild

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A single vision-language-action model jointly trained on recognition and tracking data follows described targets at the best reported levels on a public benchmark and transfers zero-shot from simulation to a real quad...

  6. CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process Thinking

    cs.AI 2025-07 conditional novelty 5.0 of 10

    CogDDN uses a fast heuristic VLM paired with a slow analytic reflection process and a growing knowledge base to navigate to objects that implicitly satisfy a user's demand, with large reported gains on AI2Thor DDN benchmarks.

  7. When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs

    cs.CV 2025-02 conditional novelty 3.0 of 10

    A survey that classifies VLM attacks by goal and data manipulation strategy, and reviews defenses and metrics.

Pith tools