Pith. sign in

Paper Citation Record · LEDGER

NaVILA: Legged Robot Vision-Language-Action Model for Navigation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 76 inbound Pith citation observations for arXiv:2412.04453.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04453 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 76 of 76 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:22:21.019683Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:09:43.511252Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fac3d417-f91f-4a1e-a41c-8d39f05ae5cc · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.167538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:792c0673f2fc888088e86f9f88be83f2f61c60cd7b3b3cbec70409e6822f8f18

Observation 836504b2-6e5d-48a7-9dc3-fa941533a835 · inbound

FAST: Efficient Action Tokenization for Vision-Language-Action Models cites this paper.

FAST: Efficient Action Tokenization for Vision-Language-Action Models NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:31.934448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:faafd207c0cb864c4a6f4b6de1e8196703b6e4924ff0f1c7efa8ce5d6083f613

Observation 64d3b1c8-7ccc-47a6-9849-95ad8e8a46b4 · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.989303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:2250286ce8d91f102e88d0e1f0bc2059c596d7e0bf5aff7e2166e071be317a69

Observation 6e5caf3e-bffc-4b63-a42d-14a7215d9f47 · inbound

Omni-Perception: Omnidirectional Collision Avoidance for Legged Locomotion in Dynamic Environments cites this paper.

Omni-Perception: Omnidirectional Collision Avoidance for Legged Locomotion in Dynamic Environments NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:21.019683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:21.019683Z digest=sha256:2961f91f20934ac742d3bb359e3d712dc76e4bedf1c442fed0ad8c4db12ce6dc

Observation c73ab88c-39f1-44a1-b111-8abc20fdbeba · inbound

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation cites this paper.

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:05.254189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:05.254189Z digest=sha256:d7b4811506340a3ba4293b3e39379c9e452c356a971c68e8fd01567c09101148

Observation 59f4dc6f-7814-4d5a-8ac6-845a407fbd16 · inbound

Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits cites this paper.

Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:16.513264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:16.513264Z digest=sha256:8459d7e150a2bacc879c185890e0004689028e33c8956ec8a41af95abdb5af6d

Observation 2eef2a0e-9ef3-4ab0-ba8c-0c7f41be2c87 · inbound

TrackVLA: Embodied Visual Tracking in the Wild cites this paper.

TrackVLA: Embodied Visual Tracking in the Wild NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:39.930337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:39.930337Z digest=sha256:94b4c1c94844411ecf9adce92329f15f05fcd7d46000474a0de9a069614b7685

Observation 38816cc3-5d58-4c66-9782-0b54dcd0d5d0 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.666040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.666040Z digest=sha256:fc242a2a54ebb88f79514ee05eb8879267617c507663b8e6b9a652f7ea6d45a0

Observation 396dad35-108d-4f25-beea-2953d1be9b1d · inbound

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks cites this paper.

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:51.175126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:51.175126Z digest=sha256:713ca9b5f61f8abe9634572db6914ff43d2f4391af96ca3bdf63e2aaf22f26d6

Observation c65785dc-dc54-48cc-8a95-cbdfc092f495 · inbound

Real-Time Execution of Action Chunking Flow Policies cites this paper.

Real-Time Execution of Action Chunking Flow Policies NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:18:51.663506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T14:18:51.613045Z digest=sha256:6689560976b20938be63f7204f0a7e1ba0a3a467982e612b14454ec9ca27bf68

Observation d427094e-e49b-4e57-8b62-ef3e1532db70 · inbound

OctoNav: Towards Generalist Embodied Navigation cites this paper.

OctoNav: Towards Generalist Embodied Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:40.740461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:44:40.740461Z digest=sha256:9d738c0f3a99ed30c61e2f497be529fed6022cac3d1e318149d9cfbcf3d0da67

Observation ed17b4aa-87dc-430f-a22e-7824d1b5cc3f · inbound

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? cites this paper.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.630568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.630568Z digest=sha256:245289c1adc325732c4288aeaf4be2f4b254fb2a56942313e1aa142e44eb601b

Observation c7e8f828-9987-4e2f-887e-4330e5f0a59c · inbound

Hierarchical Vision-Language Planning for Multi-Step Humanoid Manipulation cites this paper.

Hierarchical Vision-Language Planning for Multi-Step Humanoid Manipulation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:52.812767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:52.812767Z digest=sha256:dd0859596629a72f5b9e69c350e3dbf230d1549e1aeefa7625b48b6de47daafb

Observation 80347b0f-17f0-42af-ba60-2eea9961e4d0 · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 153

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:08:35.152944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:4dd8d28c68249c5ad219f53de788275a4a778bd937f67130b656cc38ff6b455e

Observation ce5d2b74-c720-4015-ae13-4a7af8fd3cc0 · inbound

LOVON: Legged Open-Vocabulary Object Navigator cites this paper.

LOVON: Legged Open-Vocabulary Object Navigator NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:01:24.035518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:01:24.035518Z digest=sha256:bd55db72e0eddd2ec4c93caf215fa6c13d78e06eb070a55826bf5749d66e06d3

Observation 74ba35c4-3fda-4de6-8faf-1e5eeae78498 · inbound

Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning cites this paper.

Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:47.046180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:30:47.046180Z digest=sha256:e37dd5313940e4cd64d056989b5c0756391a2d7520182b7d8727887c752be90e

Observation 1a6cdeb1-03f9-4203-8b30-d80f04486d72 · inbound

EmbRACE-3K: Embodied Reasoning and Action in Complex Environments cites this paper.

EmbRACE-3K: Embodied Reasoning and Action in Complex Environments NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:33:08.649485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:33:08.649485Z digest=sha256:3c28db0b95eab966fd54d3b8f6a390355ed16ef40e2102e6ec6b5f2a04600fd9

Observation 0ee95326-0488-4216-a12b-b8dff9a04446 · inbound

Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation cites this paper.

Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T22:50:43.715480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:50:43.715480Z digest=sha256:16a14e1a78408e08b27666006952c3e7944aed156c0bf268bf2bbf64e4638166

Observation a42816a0-a620-41d5-9747-aa27b76a829a · inbound

SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation cites this paper.

SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T22:46:42.049089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:46:42.049089Z digest=sha256:b73dbe7c2cdfec3cc9c5376e018b88aaedfc98191b15be320eca8bb1e0d05008

Observation ded1f365-b984-4087-a96d-a245e3a528f9 · inbound

CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model cites this paper.

CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:07.196850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:31:07.196850Z digest=sha256:218a6963e87929af3904248e6932381c2e46d18d3b0204faf86213a716079c4a

Observation f291dc62-9fa7-4f61-b385-8632ae5e69f7 · inbound

SPG: Style-Prompting Guidance for Style-Specific Content Creation cites this paper.

SPG: Style-Prompting Guidance for Style-Specific Content Creation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:56.869301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:56.869301Z digest=sha256:579f88dde4f0553ff7bc020cdba73dbb6e756df7d23d083793dc6800f37e8356

Observation c7b4a824-620b-40e8-989f-60541a878e8f · inbound

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models cites this paper.

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:28.436191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:28.436191Z digest=sha256:6c16028647c180731abfabe0c3c75dc7dc63ea0d3558863bf2fe49211368f34b

Observation dc10f7e0-86da-404d-8597-82dfcfea904a · inbound

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment cites this paper.

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T13:54:49.258914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:54:49.258914Z digest=sha256:4d81d538d916f2e5b63c92a9db16779b690d80dcb480edd6370d7760451e9aea

Observation 347d0f17-a116-4e37-9d7f-8a777b4109dd · inbound

Robix: A Unified Model for Robot Interaction, Reasoning and Planning cites this paper.

Robix: A Unified Model for Robot Interaction, Reasoning and Planning NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T12:59:02.133626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:59:02.133626Z digest=sha256:7333544d682276ae395fc8ece4ccfa9da7066bfd39b2b5b94716ed440a7d6695

Observation 6eb69b24-e8c1-4332-bb3e-2b5f293abe54 · inbound

Nav-R1: Reasoning and Navigation in Embodied Scenes cites this paper.

Nav-R1: Reasoning and Navigation in Embodied Scenes NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T17:31:38.800522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:31:38.800522Z digest=sha256:04689d5c8936981da771a1731c624703222604c5a843b957e650f6e7857dceb9

Observation 64ed98ce-ebd6-4121-af1d-d60f1761d8ef · inbound

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation cites this paper.

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:51:09.111792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T08:46:51.278024Z digest=sha256:909308173968bf9c1aa879958dbb910e56044d27bd3b760129da9391acf66045

Observation 3fe1c212-e7fd-4c3c-8912-a7930b086f03 · inbound

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation cites this paper.

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:05:15.979141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T21:02:38.013115Z digest=sha256:bfd31d125a7f169ba6a3cce08a3b26bf2a88139a6d1afaa42b5aa8f9309df4f4

Observation 22394142-5144-4124-9272-e14f4f880d3f · inbound

AstraNav-World: World Model for Foresight Control and Consistency cites this paper.

AstraNav-World: World Model for Foresight Control and Consistency NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:28:20.825261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T19:23:58.769472Z digest=sha256:2d8e99fd4365d5de32fd99c2bf6bb0234b99a4ab466901d4d587b58954db5041

Observation 9a86e709-e12f-4f98-b08b-8ebdb98f2a4a · inbound

VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness cites this paper.

VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:08.250364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T14:50:58.269817Z digest=sha256:4a24a114453b82378c9e6cf45a5a2b499531d89bf7421d10d66540d05dd28caa

Observation 5db8e094-517c-4759-b35a-d9a0eb429c3d · inbound

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation cites this paper.

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T17:15:42.360530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:15:42.360530Z digest=sha256:39238c53cf44df96d07b3973b5748b38d33407115018939e843ca556427cfc7f

Observation b22d2ba4-de02-411a-9a1c-08d2a5beb5fc · inbound

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation cites this paper.

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T05:43:08.245773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:43:08.245773Z digest=sha256:9afd15ab5a1ebb9fdf067005946557fa1e15a912d95a5e0bfac2b3c23fbcdf2a

Observation 68842129-da80-4f90-be5e-25b2877a6721 · inbound

Learning Task-Invariant Properties via Dreamer: Enabling Efficient Policy Transfer for Quadruped Robots cites this paper.

Learning Task-Invariant Properties via Dreamer: Enabling Efficient Policy Transfer for Quadruped Robots NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:53:11.811580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T19:49:23.086498Z digest=sha256:6cba604630b0d646da3f6ffd2335461fc9029deb59309b03532369a89221f68e

Observation 36db350c-d7f4-4364-9154-3106a2bf5bec · inbound

HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation cites this paper.

HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:41:01.469951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:01:13.350219Z digest=sha256:e6f85116bc1b3a01357334973b2f6897c464c2eba11df572143fd1c985f5d4ec

Observation a483cbf6-61b7-4079-acc5-4094ee500782 · inbound

Visually-grounded Humanoid Agents cites this paper.

Visually-grounded Humanoid Agents NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:21:04.468428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:10:32.710853Z digest=sha256:2fe33028c0b3bc1715d460d7ad235bd2d50ca4c51932b26547acc1234e291ab6

Observation f7431370-e85e-42e1-8961-779024c897df · inbound

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap cites this paper.

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:29.674502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:48:08.135538Z digest=sha256:267b833ec5ab3aa02b92aa4f626c9aadb11a4305e45707532255ad488c535da5

Observation 7b9d5fbe-234e-48c0-9b58-099b5820b2bb · inbound

CART: Context-Aware Terrain Adaptation using Temporal Sequence Selection for Legged Robots cites this paper.

CART: Context-Aware Terrain Adaptation using Temporal Sequence Selection for Legged Robots NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:55:25.378958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T12:51:40.003948Z digest=sha256:ccbb0cf34caaa06d7a313147644fb5783f512381be04cf6c06e4fb4bd0fdd614

Observation 6266283e-8f1e-4979-9cb4-72232016f271 · inbound

CART: Context-Aware Terrain Adaptation using Temporal Sequence Selection for Legged Robots cites this paper.

CART: Context-Aware Terrain Adaptation using Temporal Sequence Selection for Legged Robots NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T20:22:05.957224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:22:05.957224Z digest=sha256:cd1725f25144f4a42040366148a9f8b06080658b143c64a9083257888c6e9d70

Observation 84d45e28-7c51-45f1-8658-aac6ce8e678d · inbound

Think before Go: Hierarchical Reasoning for Image-goal Navigation cites this paper.

Think before Go: Hierarchical Reasoning for Image-goal Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:51:10.522123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T05:43:27.972164Z digest=sha256:d92fbd4b7c5b8571d8e2134035b3f08dafd7f01d70ce6f2dded4ae8d8289bd5f

Observation fe194dbe-af9d-46ea-968b-9dbfa45f548c · inbound

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation cites this paper.

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:45.831019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T06:50:34.310831Z digest=sha256:dec5b3cd0aaee3e0e4df675eebba42d01d96442a265af16eef0fc98e5b1c044a

Observation 163b26c4-4daa-4f7b-950c-d0651a2e70c7 · inbound

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation cites this paper.

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:31.261652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T05:05:33.606975Z digest=sha256:46d3676fac26f761826c163028a0447023d2b2ce59414e4626854241e5353eed

Observation 388dea21-e8b9-4e71-8a59-2c428a5b97e3 · inbound

Beyond Isolation: A Unified Benchmark for General-Purpose Navigation cites this paper.

Beyond Isolation: A Unified Benchmark for General-Purpose Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:27.549019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:33:53.557357Z digest=sha256:64c858669cd9ae16b42fdbec345600a1201262cb8638dd931b4b7d0d0ed55d0f

Observation b01388b8-e8c8-4d07-b0bd-accfe5ee30ea · inbound

Terrain Consistent Reference-Guided RL for Humanoid Navigation Autonomy cites this paper.

Terrain Consistent Reference-Guided RL for Humanoid Navigation Autonomy NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T15:23:07.747403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T15:22:38.188537Z digest=sha256:afa90f9438c858ae40f1054e16d99edda1ab00af708ad0d2308c3022759d0837

Observation 55735769-cfc8-49e1-91ae-76d4454c88e3 · inbound

SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation cites this paper.

SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:33:19.148446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T13:31:16.012419Z digest=sha256:63d4e232075c1b2be362a54c7632d998bbc38292699a5dc3a8323991b008f5d5

Observation 84154134-46ff-4693-a428-37c25129d403 · inbound

SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation cites this paper.

SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:45:00.664705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T19:42:41.238072Z digest=sha256:e3cc697b7988ecb25be57d24471cf29f37d9ebb4b767c68b618e7dde8368b11a

Observation 20405ac0-091c-4448-8806-d39c7ef3aef3 · inbound

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation cites this paper.

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:46:10.584337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T06:45:35.920991Z digest=sha256:8945b4f9cc93a0195be8db84f5c67b9485bbedd149191dbd86e0e37a1f6d4fcf

Observation 2120b067-32a6-4a06-8585-b872eaeb6435 · inbound

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation cites this paper.

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T04:46:04.562531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T04:46:03.020800Z digest=sha256:2cf32c28cea285b31feff985513c9dbac3e463ff1c6d13928b22ab845e5ef391

Observation deadbe25-e701-4ce6-8412-8e2732944916 · inbound

G-DRAGON: Geospatial Reasoning and Dynamic Planning for Retrieval-Augmented Outdoor Navigation cites this paper.

G-DRAGON: Geospatial Reasoning and Dynamic Planning for Retrieval-Augmented Outdoor Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:53:59.368576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T21:49:54.950611Z digest=sha256:f19e3f606044262c028ca55bf039edf02f9fcf66e657cc6f54e04769a0aa9d94

Observation 0ad83405-314b-48d1-861d-db91b4970d60 · inbound

Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation cites this paper.

Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:03:40.848584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T17:01:20.073666Z digest=sha256:7b341a32dee4a84d4e8766f34c69202378a14c1223e57ad329268a335b84ca43

Observation 7d8c41c4-e31e-4da3-bf8b-14aaa25f3361 · inbound

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation cites this paper.

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T11:53:23.916822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T11:48:14.888295Z digest=sha256:1289e0adafe93e78beb3c49cb21eba9cac4e2d7ce9c8734577420d237fbcebde

Observation 1e1a8380-b7ab-4042-b09e-9aeb5418ec3e · inbound

PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs cites this paper.

PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:53:40.418154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T16:53:17.504094Z digest=sha256:6e8c798a8b47f052a1706da7dcf4a643c4b9619de5ea2564fd77fbf8823560cb

Observation 2bd71fac-4c4b-4551-9a56-6e6f7fdb41bb · inbound

Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation cites this paper.

Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.934190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:49:30.585724Z digest=sha256:c43642987ef338e3aa1a33014f7981873f1fc120ae435b4ec0d0319794403477

Observation b79fac24-9072-4f62-8a59-08199f454518 · inbound

Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation cites this paper.

Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:15.795177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T15:39:13.803571Z digest=sha256:24328a1d762753483ca0d62a67676dcfd67b7b762fb6855f1f2e9e242b3f76cf

Observation c09ffcf2-259f-4859-bf6e-bc2d6c216df2 · inbound

Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation cites this paper.

Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:26.123860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T11:01:08.668612Z digest=sha256:1b86a79427de8b1f6a4b3c3876cdaa1ed2b109f504394a24604ca5fde965c493

Observation d7115002-5b6a-4bdb-83f5-cbf262b684e5 · inbound

Beyond Waypoints: A Trajectory-Centric Waypointing Paradigm for Vision-Language Navigation cites this paper.

Beyond Waypoints: A Trajectory-Centric Waypointing Paradigm for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.750835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T21:40:58.329546Z digest=sha256:e022d70a794a188affb0f067364c612fd55aca874ae9bad2b84713ab1f8a9e64

Observation 17ad59cb-7c8d-4d41-9b1f-2347ee9a9ca5 · inbound

Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models cites this paper.

Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.671095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T12:52:38.764427Z digest=sha256:d9500621d79636dbc391e4b141272a8d7269317844d60fcaae22cc205350f446

Observation 8eabb674-170a-4a53-8da8-8543c03f3941 · inbound

From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation cites this paper.

From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:38:04.652491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T09:29:43.030058Z digest=sha256:5a65f0b53cc04ce4380e9f76321c9e5d1b5290a6672edd71bb688b1b5d08c149

Observation 899df63f-9ade-4e5b-85f9-6121d359961f · inbound

VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision cites this paper.

VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:38:59.469725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T00:15:26.619238Z digest=sha256:6232c7048d117eb26ee607e0ba7bfdfedcbe67642ebbbee66570fa006b21d36a

Observation cbada972-9842-43c1-b840-9bb5f6f2b4bc · inbound

VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision cites this paper.

VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T11:03:51.561853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:03:51.561853Z digest=sha256:6018e5c0c7192d2f03c954cabb90e5e41cb664129688aec43c6968db3b7514b4

Observation 8ccd85cd-0fa4-4d8d-9ba3-6250c061aa6c · inbound

Slow Brain, Fast Planner: Latency-Resilient VLM-Augmented Urban Navigation cites this paper.

Slow Brain, Fast Planner: Latency-Resilient VLM-Augmented Urban Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:09:33.984552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T17:16:52.173943Z digest=sha256:93e0207d1943a88362f000e79983994b86a270182381fe7288fed8aae279d777

Observation 091a2de6-e56a-4a62-a619-e2b06ba794e9 · inbound

Vesta: A Generalist Embodied Reasoning Model cites this paper.

Vesta: A Generalist Embodied Reasoning Model NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:29:35.757784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T16:55:12.518255Z digest=sha256:5a50e601b8e2a12e2c4aed54688a55b18075126dcf8318cc1b24368c54d06f9c

Observation 38150bff-1b9d-4291-993b-8050aec8de85 · inbound

BIT-Nav: Brain-Inspired Trajectory Memory for Embodied Navigation cites this paper.

BIT-Nav: Brain-Inspired Trajectory Memory for Embodied Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:09:37.370731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T13:52:31.453516Z digest=sha256:eeddfda266a5bcf67ea92ea183a88d9be0ff0ea59583659f4fa515c28b3b21fc

Observation 6f8c30ea-56fe-49c7-8600-cc3687652cac · inbound

FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation cites this paper.

FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:09:43.512947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T10:24:56.607921Z digest=sha256:e5a3a65ebb264cf150312f0770a867c764c2bfcac1cca6ec50065a651e0cb8f6

Observation c560a56e-eacc-4fe8-a460-cefc703ba6df · inbound

SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks cites this paper.

SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:33:54.613197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T04:42:23.040915Z digest=sha256:558c8921ce54fcfacaca519db03cbca96b3211393d0d6652500f86a03b105987

Observation 613b6a72-ca78-421f-b77f-9dec93d39f73 · inbound

Vision-Language Models for Deployable Social Robot Navigation: Bridging Semantic Reasoning and Low-Level Control cites this paper.

Vision-Language Models for Deployable Social Robot Navigation: Bridging Semantic Reasoning and Low-Level Control NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:54:34.519791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T09:52:51.549135Z digest=sha256:8d90159ee1456f03511656e8ea75594c19990dff8aef433b771c7669ecccefe8

Observation 73fef3cb-f37c-4a72-9c36-98d11587368b · inbound

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation cites this paper.

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:54:49.928777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T05:05:56.065261Z digest=sha256:425bd47a8415c9795e2651f91cc5ce288aeb3af17854094fb6920884dd216ad4

Observation c132e310-40d5-4fa2-92b5-e7a178b3f5ab · inbound

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation cites this paper.

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:08:21.406594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T14:04:46.400336Z digest=sha256:232dcbe089f4b8fb8d3dfa0ffc90296e8e54a776e5480c7a709b2ea7d1c6fa7c

Observation a3f816f0-a771-4f4e-9e6b-66e40e2249b1 · inbound

Exp2VLA: Enabling Vision-Language-Action for Drone Navigation from Expert Demonstrations cites this paper.

Exp2VLA: Enabling Vision-Language-Action for Drone Navigation from Expert Demonstrations NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T04:33:48.376442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:33:48.376442Z digest=sha256:28ff18b34806e915d21fb8b89a7504541ef43d11579a09f99883e91af685f28e

Observation 60e3694c-c7e2-4522-8943-3ba9f774cccf · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:792c9020ee6a18c10e7607f65948173b015480a4d5473aa85689b414b2ee3dae

Observation b6e50527-e702-4417-aa23-4552559563ad · inbound

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies cites this paper.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:56b165b16f46474035e8698d69f7d146bcbfb8233cff2e554e60788837b3b138

Observation f4729f4d-613e-4a48-bbbf-f24dfdff1b23 · inbound

Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation cites this paper.

Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-01T16:20:01.293615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:20:01.293615Z digest=sha256:5d451da533394bb8fda6f38fed6b186390911213d9aac580a9157a1f5edc8331

Observation 4dcd7d93-6b4e-406e-a652-5bedd565ddcb · inbound

EA-Nav: Learning Safe Visual Navigation Policies with Embodiment Awareness cites this paper.

EA-Nav: Learning Safe Visual Navigation Policies with Embodiment Awareness NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:30:22.146055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:30:22.146055Z digest=sha256:054dd95f948fc085343ed86cd365de91cbfda0699526fb9aada916286698f0c7

Observation 7e252989-0e19-47ea-a185-1296bbc615d6 · inbound

ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset cites this paper.

ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 292

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:39.791761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:39.791761Z digest=sha256:e468ed1fbcc6c0f8614d8090e32ad470e935af5f74e1c423390bcfcfb68b7ae8

Observation 79160368-0112-47d0-b764-c75123e98fe3 · inbound

Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents cites this paper.

Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T06:06:22.870692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:06:22.870692Z digest=sha256:21e31783cc43abe10ad8e8fa22600b1ba3b7ced259e055e1eb81dc4091302697

Observation 51a0af7a-2ead-4729-b06b-9357ef1cda00 · inbound

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation cites this paper.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.068111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.068111Z digest=sha256:9551a3b0b28e873927ee78f7f81a2a360f58d9284b68c1680af11bf0d87c611d

Observation fda33824-fdbd-4099-8d6b-0184621082ef · inbound

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation cites this paper.

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T00:43:30.833308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:43:30.833308Z digest=sha256:808d76d99584ef8ec8ab6a9e679901960dd63e237e84fd7955f6d0faa6f729a5

Observation b0ffe209-0598-4085-a7b1-ca7ea417fa2f · inbound

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation cites this paper.

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T03:24:27.168988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:24:27.168988Z digest=sha256:41fb7a3b0913f616629c47ac7a8a2488a03f612bb190db4dfac6becb89ecdf5e