Pith. sign in

REVIEW 10 cited by

Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.07775 v2 pith:Y36KFHQ6 submitted 2024-07-10 cs.RO cs.AI

classification cs.ROcs.AI
keywords navigationmultimodalmobilitygoalpolicyvideovlmsdemonstration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

An elusive goal in navigation research is to build an intelligent agent that can understand multimodal instructions including natural language and image, and perform useful navigation. To achieve this, we study a widely useful category of navigation tasks we call Multimodal Instruction Navigation with demonstration Tours (MINT), in which the environment prior is provided through a previously recorded demonstration video. Recent advances in Vision Language Models (VLMs) have shown a promising path in achieving this goal as it demonstrates capabilities in perceiving and reasoning about multimodal inputs. However, VLMs are typically trained to predict textual output and it is an open research question about how to best utilize them in navigation. To solve MINT, we present Mobility VLA, a hierarchical Vision-Language-Action (VLA) navigation policy that combines the environment understanding and common sense reasoning power of long-context VLMs and a robust low-level navigation policy based on topological graphs. The high-level policy consists of a long-context VLM that takes the demonstration tour video and the multimodal user instruction as input to find the goal frame in the tour video. Next, a low-level policy uses the goal frame and an offline constructed topological graph to generate robot actions at every timestep. We evaluated Mobility VLA in a 836m^2 real world environment and show that Mobility VLA has a high end-to-end success rates on previously unsolved multimodal instructions such as "Where should I return this?" while holding a plastic bin. A video demonstrating Mobility VLA can be found here: https://youtu.be/-Tof__Q8_5s

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ABot-N1: Toward a General Visual Language Navigation Foundation Model

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    A slow–fast VLN model that routes five navigation tasks through CoT plus image-space pixel goals reaches SOTA on established and new urban benchmarks, including 77.3% POI arrival.

  2. PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation

    cs.RO 2026-02 conditional novelty 6.0 of 10

    AgenticLab's closed-loop planning-language pipeline lets different vision-language models drive a real robot, and benchmark tests show action-verification quality, not planning, determines long-horizon success.

  3. TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A navigation pipeline that bridges object-level global planning with traversability-aware local control, using only RGB images and pretrained models, improves success over prior zero-shot and learned baselines in simulation.

  4. CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models

    cs.RO 2025-08 conditional novelty 6.0 of 10

    Counterfactual language-action relabeling raises instruction-following success from about 26% to 53% in real-world navigation tests.

  5. Adversarial Attacks on Robotic Vision Language Action Models

    cs.RO 2025-06 conditional novelty 6.0 of 10

    Text-based adversarial suffixes can make OpenVLA robot policies elicit chosen target actions with over 90% success on one-hot targets and persist across rollout steps.

  6. GraphPad: Inference-Time 3D Scene Graph Updates for Embodied Question Answering

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Allowing a vision-language model to edit its own 3D scene graph during inference improves embodied question answering from 52.3% to 55.3% on OpenEQA.

  7. Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Real-time SegFormer green/red overlays reduce OmniVLA far-waypoint error 27-44% on Grand Tour language goals mainly by shortening trajectories, with little help for image goals.

  8. PixelNav: Towards Model-based Vision-Only Navigation with Topological Graphs

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A camera-only navigation system combining visual place recognition, traversability segmentation, and model predictive control over a topological graph, evaluated on a real robot.

  9. ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning

    cs.RO 2025-05 conditional novelty 5.0 of 10

    Fine-tuning a vision-language-action robot model on teacher-generated reasoning rationales raises average simulated manipulation success by up to 8.6 percentage points over the SpatialVLA baseline.

  10. Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning

    cs.RO 2025-08 reject novelty 4.0 of 10

    A review that categorizes large-model-empowered embodied AI into hierarchical and end-to-end decision-making, imitation and reinforcement learning, and world models.

Pith tools