REVIEW 32 cited by
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper proposes to solve the problem of Vision-and-Language Navigation with legged robots, which not only provides a flexible way for humans to command but also allows the robot to navigate through more challenging and cluttered scenes. However, it is non-trivial to translate human language instructions all the way to low-level leg joint actions. We propose NaVILA, a 2-level framework that unifies a Vision-Language-Action model (VLA) with locomotion skills. Instead of directly predicting low-level actions from VLA, NaVILA first generates mid-level actions with spatial information in the form of language, (e.g., "moving forward 75cm"), which serves as an input for a visual locomotion RL policy for execution. NaVILA substantially improves previous approaches on existing benchmarks. The same advantages are demonstrated in our newly developed benchmarks with IsaacLab, featuring more realistic scenes, low-level controls, and real-world robot experiments. We show more results at https://navila-bot.github.io/
Forward citations
Cited by 32 Pith papers
-
Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation
Unmodified software-agent harnesses, given only a monocular RGB camera and four discrete actions, achieve 68–78% success on zero-shot R2R-CE navigation, rivaling trained and workflow-based systems.
-
Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents
A new 120-mission drone benchmark finds the best off-the-shelf multimodal AI completes 34.8% of missions versus 84.4% for humans, with scaling helping but not closing the gap.
-
ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset
ACME is a socially navigated robot and pedestrian trajectory dataset covering 8 sites in 5 countries with 7 robot embodiments, including human-verified BEV tracks and robot speech annotations.
-
EA-Nav: Learning Safe Visual Navigation Policies with Embodiment Awareness
An imitation-learning navigation model conditioned on the robot's body dimensions reduces collisions and improves success across embodiments, using pseudo-labeled internet video pretraining and risk-augmented fine-tuning.
-
ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI
A single 8B backbone unifies spatial perception, decision making, navigation/manipulation, and progress estimation with SSR+ merging, reporting gains on most spatial benchmarks and competitive action/progress results.
-
VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision
VEGA reconstructs local geometry from monocular egocentric video to create supervised trajectories that train a flow-matching VLA policy, yielding lower collision rates on a new benchmark and in real-world tests.
-
CART: Context-Aware Terrain Adaptation using Temporal Sequence Selection for Legged Robots
CART learns a vision–proprioception terrain context and uses Temporal Sequence Selection to cut base oscillation by up to 41% in simulation and 22% on Spot outdoors, with a 5% higher sim success rate.
-
Structured Observation Language for Efficient and Generalizable Vision-Language Navigation
SOL-Nav encodes RGB-D observations as grid-organized text and uses a 0.6B text-embedding model with four classification heads to predict navigation action blocks, reporting SOTA/comparable R2R-CE/RxR-CE results with 1...
-
Robix: A Unified Model for Robot Interaction, Reasoning and Planning
A three-stage-trained VLM unifies robot planning and dialogue, and beats commercial VLMs on the authors' interactive-task benchmarks.
-
CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
Counterfactual language-action relabeling raises instruction-following success from about 26% to 53% in real-world navigation tests.
-
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
By iteratively retraining on automatically generated corrective examples derived from its own wrong paths, CorrectNav reports new state-of-the-art success rates of 65.1% (R2R-CE) and 69.3% (RxR-CE).
-
Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
Low within-subdataset diversity and large between-subdataset differences cause shortcut learning in generalist robot policies, and targeted augmentation can mitigate it.
-
EmbRACE-3K: Embodied Reasoning and Action in Complex Environments
EmbRACE-3K provides a photorealistic, closed-loop embodied benchmark with stepwise reasoning annotations, and fine-tuning Qwen2.5-VL on it improves embodied task success in simulation.
-
Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation?
A behavior-cloning policy trained only on frozen SigLIP embeddings reaches 74% of language-specified targets in a simple simulator, versus 100% for a state-aware expert, and takes 3.2x more steps.
-
OctoNav: Towards Generalist Embodied Navigation
OctoNav-R1, trained with SFT, GRPO, and online RL on the new OctoNav-Bench, achieves 19.4% overall success on mixed-instruction navigation, more than double the best baseline.
-
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
VeBrain unifies perception, spatial reasoning, and robot control in one MLLM by representing control as keypoint detection plus skill selection, with a robotic adapter for deployment.
-
TrackVLA: Embodied Visual Tracking in the Wild
A single vision-language-action model jointly trained on recognition and tracking data follows described targets at the best reported levels on a public benchmark and transfers zero-shot from simulation to a real quad...
-
Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits
Edge-cloud speculative decoding runs faster when early exits in the server model let the client pre-draft the next candidate tokens before final verification is complete.
-
Omni-Perception: Omnidirectional Collision Avoidance for Legged Locomotion in Dynamic Environments
Omni-Perception is an end-to-end RL policy for legged robots that processes raw LiDAR point clouds with PD-RiskNet to achieve omnidirectional collision avoidance, validated in simulation and on a Unitree G1.
-
Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies
Real-time SegFormer green/red overlays reduce OmniVLA far-waypoint error 27-44% on Grand Tour language goals mainly by shortening trajectories, with little help for image goals.
-
MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation
Pyramidal multi-resolution visual memory plus single-token mid-level actions let a 4B–8B VLM navigate continuous indoor environments at 14 FPS with SOTA R2R/RxR success rates.
-
SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation
A GAN framework that translates unlabeled medical images between classes and fuses ensemble, time-averaged pseudo-labels outperforms six prior GAN semi-supervised methods on MedMNIST at 5-50 labels per class.
-
Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning
A self-critique, revision, and verification loop makes small vision-language models produce more detailed and more executable robot plans, beating their own baselines and, on the paper's judge-based evaluation, plans ...
-
Hierarchical Vision-Language Planning for Multi-Step Humanoid Manipulation
A three-layer hierarchical system using a VLM planner and VLM skill monitor with imitation-learned skills and an RL tracking policy achieved 73% success on a real humanoid pick-and-place task.
-
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
A dual-branch text-imagination system, with one LLM branch for state estimation and one for candidate-direction description, improves R2R navigation success over prior LLM-based VLN methods.
-
Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation
FSC-VLN: a future-prediction training signal improves long-horizon vision-language navigation on R2R, with no future frames needed at test time.
-
Exp2VLA: Enabling Vision-Language-Action for Drone Navigation from Expert Demonstrations
Expert demonstrations distilled into LeRobot-format data let fine-tuned mid-scale VLAs navigate multi-object drone scenes from language, with partial zero-shot color-shape generalization in simulation and SITL.
-
Nav-R1: Reasoning and Navigation in Embodied Scenes
Nav-R1 uses a 110K synthetic CoT dataset, GRPO with three rewards, and a fast-in-slow system to set new SOTA on R2R-CE, RxR-CE, and HM3D-OVON.
-
Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment
VEME, a dual-memory cross-modal alignment framework built on Qwen-2.5-VL, reports modest gains on VLN-CE and VSI-Bench that are contradicted by its own internal numbers.
-
SPG: Style-Prompting Guidance for Style-Specific Content Creation
SPG is not described anywhere in the supplied text; the body is a different paper (OVSegDT) about robot navigation.
-
LOVON: Legged Open-Vocabulary Object Navigator
LOVON integrates an LLM planner, a blur-filtered object detector, and a small learned motion model to navigate legged robots to user-specified objects over long horizons, claiming near-perfect simulation success and r...
-
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
A unified vision-language-action model that emits a sub-task description followed by a discrete action token outperforms modular and action-only baselines on simulated long-horizon tabletop tasks.
Discussion (0). Sign in to comment.