Pith. sign in

REVIEW 32 cited by

NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.04453 v2 pith:FU5ZF6UD submitted 2024-12-05 cs.RO cs.CV

classification cs.ROcs.CV
keywords navilaactionslow-levelrobotbenchmarkslanguageleggedlocomotion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper proposes to solve the problem of Vision-and-Language Navigation with legged robots, which not only provides a flexible way for humans to command but also allows the robot to navigate through more challenging and cluttered scenes. However, it is non-trivial to translate human language instructions all the way to low-level leg joint actions. We propose NaVILA, a 2-level framework that unifies a Vision-Language-Action model (VLA) with locomotion skills. Instead of directly predicting low-level actions from VLA, NaVILA first generates mid-level actions with spatial information in the form of language, (e.g., "moving forward 75cm"), which serves as an input for a visual locomotion RL policy for execution. NaVILA substantially improves previous approaches on existing benchmarks. The same advantages are demonstrated in our newly developed benchmarks with IsaacLab, featuring more realistic scenes, low-level controls, and real-world robot experiments. We show more results at https://navila-bot.github.io/

Discussion (0). Sign in to comment.

Forward citations

Cited by 32 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Unmodified software-agent harnesses, given only a monocular RGB camera and four discrete actions, achieve 68–78% success on zero-shot R2R-CE navigation, rivaling trained and workflow-based systems.

  2. Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A new 120-mission drone benchmark finds the best off-the-shelf multimodal AI completes 34.8% of missions versus 84.4% for humans, with scaling helping but not closing the gap.

  3. ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset

    cs.RO 2026-07 conditional novelty 6.0 of 10

    ACME is a socially navigated robot and pedestrian trajectory dataset covering 8 sites in 5 countries with 7 robot embodiments, including human-verified BEV tracks and robot speech annotations.

  4. EA-Nav: Learning Safe Visual Navigation Policies with Embodiment Awareness

    cs.RO 2026-07 conditional novelty 6.0 of 10

    An imitation-learning navigation model conditioned on the robot's body dimensions reduces collisions and improves success across embodiments, using pseudo-labeled internet video pretraining and risk-augmented fine-tuning.

  5. ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A single 8B backbone unifies spatial perception, decision making, navigation/manipulation, and progress estimation with SSR+ merging, reporting gains on most spatial benchmarks and competitive action/progress results.

  6. VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    VEGA reconstructs local geometry from monocular egocentric video to create supervised trajectories that train a flow-matching VLA policy, yielding lower collision rates on a new benchmark and in real-world tests.

  7. CART: Context-Aware Terrain Adaptation using Temporal Sequence Selection for Legged Robots

    cs.RO 2026-04 unverdicted novelty 6.0 of 10

    CART learns a vision–proprioception terrain context and uses Temporal Sequence Selection to cut base oscillation by up to 41% in simulation and 22% on Spot outdoors, with a 5% higher sim success rate.

  8. Structured Observation Language for Efficient and Generalizable Vision-Language Navigation

    cs.CV 2026-03 reject novelty 6.0 of 10

    SOL-Nav encodes RGB-D observations as grid-organized text and uses a 0.6B text-embedding model with four classification heads to predict navigation action blocks, reporting SOTA/comparable R2R-CE/RxR-CE results with 1...

  9. Robix: A Unified Model for Robot Interaction, Reasoning and Planning

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A three-stage-trained VLM unifies robot planning and dialogue, and beats commercial VLMs on the authors' interactive-task benchmarks.

  10. CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models

    cs.RO 2025-08 conditional novelty 6.0 of 10

    Counterfactual language-action relabeling raises instruction-following success from about 26% to 53% in real-world navigation tests.

  11. CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model

    cs.RO 2025-08 unverdicted novelty 6.0 of 10

    By iteratively retraining on automatically generated corrective examples derived from its own wrong paths, CorrectNav reports new state-of-the-art success rates of 65.1% (R2R-CE) and 69.3% (RxR-CE).

  12. Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation

    cs.RO 2025-08 conditional novelty 6.0 of 10

    Low within-subdataset diversity and large between-subdataset differences cause shortcut learning in generalist robot policies, and targeted augmentation can mitigate it.

  13. EmbRACE-3K: Embodied Reasoning and Action in Complex Environments

    cs.CV 2025-07 conditional novelty 6.0 of 10

    EmbRACE-3K provides a photorealistic, closed-loop embodied benchmark with stepwise reasoning annotations, and fine-tuning Qwen2.5-VL on it improves embodied task success in simulation.

  14. Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation?

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A behavior-cloning policy trained only on frozen SigLIP embeddings reaches 74% of language-specified targets in a simple simulator, versus 100% for a state-aware expert, and takes 3.2x more steps.

  15. OctoNav: Towards Generalist Embodied Navigation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    OctoNav-R1, trained with SFT, GRPO, and online RL on the new OctoNav-Bench, achieves 19.4% overall success on mixed-instruction navigation, more than double the best baseline.

  16. Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

    cs.CV 2025-05 reject novelty 6.0 of 10

    VeBrain unifies perception, spatial reasoning, and robot control in one MLLM by representing control as keypoint detection plus skill selection, with a robotic adapter for deployment.

  17. TrackVLA: Embodied Visual Tracking in the Wild

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A single vision-language-action model jointly trained on recognition and tracking data follows described targets at the best reported levels on a public benchmark and transfers zero-shot from simulation to a real quad...

  18. Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Edge-cloud speculative decoding runs faster when early exits in the server model let the client pre-draft the next candidate tokens before final verification is complete.

  19. Omni-Perception: Omnidirectional Collision Avoidance for Legged Locomotion in Dynamic Environments

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Omni-Perception is an end-to-end RL policy for legged robots that processes raw LiDAR point clouds with PD-RiskNet to achieve omnidirectional collision avoidance, validated in simulation and on a Unitree G1.

  20. Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Real-time SegFormer green/red overlays reduce OmniVLA far-waypoint error 27-44% on Grand Tour language goals mainly by shortening trajectories, with little help for image goals.

  21. MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Pyramidal multi-resolution visual memory plus single-token mid-level actions let a 4B–8B VLM navigate continuous indoor environments at 14 FPS with SOTA R2R/RxR success rates.

  22. SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    A GAN framework that translates unlabeled medical images between classes and fuses ensemble, time-averaged pseudo-labels outperforms six prior GAN semi-supervised methods on MedMNIST at 5-50 labels per class.

  23. Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A self-critique, revision, and verification loop makes small vision-language models produce more detailed and more executable robot plans, beating their own baselines and, on the paper's judge-based evaluation, plans ...

  24. Hierarchical Vision-Language Planning for Multi-Step Humanoid Manipulation

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A three-layer hierarchical system using a VLM planner and VLM skill monitor with imitation-learned skills and an RL tracking policy achieved 73% success on a real humanoid pick-and-place task.

  25. Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A dual-branch text-imagination system, with one LLM branch for state estimation and one for candidate-direction description, improves R2R navigation success over prior LLM-based VLN methods.

  26. Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation

    cs.CV 2026-07 conditional novelty 4.0 of 10

    FSC-VLN: a future-prediction training signal improves long-horizon vision-language navigation on R2R, with no future frames needed at test time.

  27. Exp2VLA: Enabling Vision-Language-Action for Drone Navigation from Expert Demonstrations

    cs.RO 2026-07 conditional novelty 4.0 of 10

    Expert demonstrations distilled into LeRobot-format data let fine-tuned mid-scale VLAs navigate multi-object drone scenes from language, with partial zero-shot color-shape generalization in simulation and SITL.

  28. Nav-R1: Reasoning and Navigation in Embodied Scenes

    cs.RO 2025-09 reject novelty 4.0 of 10

    Nav-R1 uses a 110K synthetic CoT dataset, GRPO with three rewards, and a fast-in-slow system to set new SOTA on R2R-CE, RxR-CE, and HM3D-OVON.

  29. Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment

    cs.CV 2025-08 reject novelty 4.0 of 10

    VEME, a dual-memory cross-modal alignment framework built on Qwen-2.5-VL, reports modest gains on VLN-CE and VSI-Bench that are contradicted by its own internal numbers.

  30. SPG: Style-Prompting Guidance for Style-Specific Content Creation

    cs.GR 2025-08 unverdicted novelty 4.0 of 10

    SPG is not described anywhere in the supplied text; the body is a different paper (OVSegDT) about robot navigation.

  31. LOVON: Legged Open-Vocabulary Object Navigator

    cs.RO 2025-07 reject novelty 4.0 of 10

    LOVON integrates an LLM planner, a blur-filtered object detector, and a small learned motion model to navigate legged robots to user-specified objects over long horizons, claiming near-perfect simulation success and r...

  32. LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A unified vision-language-action model that emits a sub-task description followed by a discrete action token outperforms modular and action-only baselines on simulated long-horizon tabletop tasks.

Pith tools