Pith. sign in

Paper Citation Record · LEDGER

NaVILA: Legged Robot Vision-Language-Action Model for Navigation

As of 19 August 2026, this Paper Citation Record lists 100 of 122 outbound references and 85 inbound Pith citation observations for arXiv:2412.04453.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04453 v2

Coverage vector

measured 100 of 122 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:30:18.197698Z

measured 185 of 185 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 85 of 85 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:50:21.311984Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:09:43.511252Z

Reference resolution

100 of 122 outbound references displayed

  • verified exact0
  • verified fuzzy58
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a363256c-9a04-4119-ab42-3c5db7d6dd56 · outbound

This paper cites Vision-and- language navigation: Interpreting visually-grounded navigation instructions in real environments.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Vision-and- language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.831508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.831508Z digest=sha256:9737c5ad6a9526d7408b4a3e72eb08268dd19ad73c135b74b62faf1d034e289a

Observation d84794ca-ccd1-47a0-a619-badf00fe81af · outbound

This paper cites Reinforced cross-modal match- ing and self-supervised imitation learning for vision- language navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Reinforced cross-modal match- ing and self-supervised imitation learning for vision- language navigation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.836242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.836242Z digest=sha256:6e9542668559963ce947c988c5d8ec60228207d88ebe1274fd85bbeb24324ebc

Observation 49ac0c45-089f-4424-9575-5f80791f90bd · outbound

This paper cites Learning to explore using active neural slam.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Learning to explore using active neural slam

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.840331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.840331Z digest=sha256:e52a8a17e57a694079f94b0f461452c964c2c30e7f3ee5bf167d935888560031

Observation d1ff295d-4ef9-48da-9f66-c083f0067529 · outbound

This paper cites Object goal navigation using goal-oriented semantic exploration.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Object goal navigation using goal-oriented semantic exploration

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.845598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.845598Z digest=sha256:8104a4b2dcfe9eb115e81c4c40e3fd8c9438f9cf30bd29d293296968d1daf74d

Observation 59b44ac5-0c8f-426c-94ee-b53caf157428 · outbound

This paper cites Neural topological slam for visual navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Neural topological slam for visual navigation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.851532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.851532Z digest=sha256:4948774e4d150b55a5ba17959610f06ed02d68e36d7f972050cb3abb10c657e7

Observation 91f2abfc-1957-4d3f-b9b3-9d19712f5f05 · outbound

This paper cites Habitat-web: Learning embodied object- search strategies from human demonstrations at scale.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Habitat-web: Learning embodied object- search strategies from human demonstrations at scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.855931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.855931Z digest=sha256:79148d55475266c871497cb4dc94179d684ae3de78a6686e5729aae3caba7142

Observation 570268ab-f4bf-47d1-9f36-89bccd297f4b · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.859854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.859854Z digest=sha256:603074b5df9b0bc2256af4ca685b10a3d61498f177baeeb87219c57e280f41c7

Observation 0115472b-3e5d-4cc9-a46a-132900ce24fa · outbound

This paper cites Openvla: An open-source vision-language-action model.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Openvla: An open-source vision-language-action model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.864098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.864098Z digest=sha256:989d2ec82fde7692649e4d4f98bf1d031312858287dfe471f55c8dfce596f756

Observation a595661a-31d2-4e3a-9459-4973ba399154 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x mod- els.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Open x-embodiment: Robotic learning datasets and rt-x mod- els

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.867800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.867800Z digest=sha256:8068d4a32fdb64e859a549eb51128c391187460fcfb98d2530fe21fbe93829bc

Observation 911cf3e0-e316-4d1c-975d-f41f01f0b0b1 · outbound

This paper cites Spa- tialvlm: Endowing vision-language models with spatial reasoning capabilities.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Spa- tialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.871382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.871382Z digest=sha256:44073d2be00794f24f3c21e4258f79f39a3ae1f5ac739284686fc7f4e2b1a24a

Observation 58830bb4-a099-408d-839b-57f9c27ccb2c · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision- language models.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Spatialrgpt: Grounded spatial reasoning in vision- language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.874729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.874729Z digest=sha256:93605047446ed2a881a121e2d934590925f2c51667303e37b008cc39be1bc953

Observation 05dc268e-3be0-4dc9-bf9c-9691b9a38488 · outbound

This paper cites Navid: Video-based vlm plans the next step for vision-and-language navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Navid: Video-based vlm plans the next step for vision-and-language navigation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.878371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.878371Z digest=sha256:98b9ef461a0703b7ca20131cdd7fdfd1a75af8f4c3bd1dc8bc4a6334bf6d2613

Observation 5b737d52-472c-4fc1-b137-2d64785dd41a · outbound

This paper cites Vila: On pre-training for visual language models.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Vila: On pre-training for visual language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.882456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.882456Z digest=sha256:5ab7eb692310cf6117c9e25f7a49cc078095c52ad1eb55420cbe0b55276fe4cb

Observation 65b3d0cb-92c5-4fb2-aafd-e039216f4add · outbound

This paper cites Vila-u: a unified foundation model integrating visual understanding and generation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Vila-u: a unified foundation model integrating visual understanding and generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.886040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.886040Z digest=sha256:9c2a73d642cba2d1e13eec6d0249321b76e4161604a9d3ab3e29aadf1b814111

Observation b722a745-26d8-4db6-8b7a-b9932ccb9d8a · outbound

This paper cites Vila 2: Vila augmented vila.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Vila 2: Vila augmented vila

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.889655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.889655Z digest=sha256:dfb7529179667722890c875188ec3191942a4b15386664bcdd14e60503a7c30d

Observation 378b2d1a-8f0d-4086-a726-c7c1049068d5 · outbound

This paper cites Longvila: Scaling long- context visual language models for long videos.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Longvila: Scaling long- context visual language models for long videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.893308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.893308Z digest=sha256:571ed1eb25e208354624988410e6892576f171356c3b3c842dece1a5196c57f8

Observation a369e17e-c6ee-410e-954d-16bb275b8738 · outbound

This paper cites X-vila: Cross-modality alignment for large language model.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation X-vila: Cross-modality alignment for large language model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.897116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.897116Z digest=sha256:5eaebc4355c3615eb57f373634e7df92fc0a79f2312c259cf1450b90c519de39

Observation 559cbdd2-7d62-421e-b43d-bd57471b3cae · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Lita: Language instructed temporal-localization assistant

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.900546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.900546Z digest=sha256:cc697e69ba7b8dc53445557e5bbd4a0acaf3dbeea6979b67114ab61ec8a93b78

Observation f4dba096-e6a8-4d49-840a-16ff475015d7 · outbound

This paper cites Nvila: Efficient frontier visual language models.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Nvila: Efficient frontier visual language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.904040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.904040Z digest=sha256:92e929a9d02ea16eca59a9fc0b1276f668fc25d10a98a689a7c629da715df842

Observation 4881a756-aa64-4b9f-a570-54e1e1b9b507 · outbound

This paper cites Visual instruction tuning.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Visual instruction tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.907773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.907773Z digest=sha256:30844117d05b26ee9810ae1650bb0df7b04a1389fe5f47fa56ac1bb80bbd31ce

Observation 40257328-7073-4b89-9ad0-2750db60f4db · outbound

This paper cites Coyo-700m: Image-text pair dataset.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Coyo-700m: Image-text pair dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.911633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.911633Z digest=sha256:288bfae440b1153de4b25acfbdb91970b610f97b4e79cede3d4bed5c8352a430

Observation 9aab1108-f263-497c-9d0b-7e78489217d6 · outbound

This paper cites Multimodal c4: An open, billion-scale corpus of images interleaved with text.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Multimodal c4: An open, billion-scale corpus of images interleaved with text

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.915170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.915170Z digest=sha256:a24669acad2456fa8e5517ef45fc306f99e9a961b451fac5ccf13f81f828d10a

Observation b52bf830-7b78-4975-9d3e-597987ec8b1a · outbound

This paper cites Improved baselines with visual instruction tuning.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Improved baselines with visual instruction tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.918942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.918942Z digest=sha256:e47775b77d0d6c67e273938114279b78a744ae028170021546babefb64ac8a6e

Observation 76cb3944-cfb3-4372-a4f8-8f07b623a703 · outbound

This paper cites Airbert: In-domain pretraining for vision-and-language navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Airbert: In-domain pretraining for vision-and-language navigation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.922686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.922686Z digest=sha256:16670c308376d3d316d9316968420c3ef1210f32172a34562e643d219188a2f6

Observation 3c99551f-8731-44eb-930b-e39e9e20490d · outbound

This paper cites Improving vision-and-language navigation with image-text pairs from the web.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Improving vision-and-language navigation with image-text pairs from the web

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.926611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.926611Z digest=sha256:7a9ed575ffea961d305f8afa8d3c7dfea2bad319df9e4ec782fefaccec8ebffc

Observation 788944ab-8b84-4a78-b01a-50120c87446b · outbound

This paper cites Learning vision- and-language navigation from youtube videos.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Learning vision- and-language navigation from youtube videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.930553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.930553Z digest=sha256:02084a66e235215463893282dd6e391544956faceb665855c0ffa5e0fa12af0e

Observation dffa3edc-7555-4171-b21c-8f4f09ccac35 · outbound

This paper cites Grounding image matching in 3d with mast3r.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Grounding image matching in 3d with mast3r

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.933948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.933948Z digest=sha256:458d78238af93adfd86ba90c729318847cf7845e716500b21c0863df62f3ee60

Observation d3c6e179-c46f-49cf-93f9-51b690655d83 · outbound

This paper cites Hello gpt-4o, 2024.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Hello gpt-4o, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.937623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.937623Z digest=sha256:c2bfef4be887043a574e537d3e970e19db18dd911a4af0335848d96736ebfb39

Observation da16f356-3b44-4ff3-b4c3-c572eb74ddeb · outbound

This paper cites Beyond the nav-graph: Vision and language navigation in continuous environments.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Beyond the nav-graph: Vision and language navigation in continuous environments

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.942500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.942500Z digest=sha256:331888dcd8210f9708cff4b135dc05dd68a4300d2570964ba67564a066c8ba0e

Observation 4e1f93db-0c9b-4786-91b9-f6412e9ae0d0 · outbound

This paper cites Room-across-room: Multilingual vision-and-language navigation with dense spatiotem- poral grounding.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Room-across-room: Multilingual vision-and-language navigation with dense spatiotem- poral grounding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.945879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.945879Z digest=sha256:bf9a3c078af27241031d3cf09dc83e0eed75c2c308aad623777b1a9419c9b9e1

Observation 34249108-43b1-4f4b-9f12-db518486fb2b · outbound

This paper cites Learning to navigate unseen environments: Back translation with environmental dropout.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Learning to navigate unseen environments: Back translation with environmental dropout

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.949478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.949478Z digest=sha256:8d6302024f53296d26619f90e0f5c5871deaec9b1b3278dbf8572557a8c3be5d

Observation 38a4a32d-5441-48b0-9719-2dfd89473c76 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Scanqa: 3d question answering for spatial scene understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.952702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.952702Z digest=sha256:bc1d744d54cf27899503fa4edc056200118a4a69ce6815062c194f0c3424ab60

Observation bd9c79c8-a349-4d41-9a2b-9157de80c1be · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Sharegpt4v: Improving large multi-modal models with better captions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.956028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.956028Z digest=sha256:5e3bb6340b0344f99d32360d640e83920902f3265d70294ac998a9db7881a218

Observation 1fd49c3f-bff9-4ea4-8392-56ee24df0c16 · outbound

This paper cites Video-chatgpt: Towards de- tailed video understanding via large vision and language models.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Video-chatgpt: Towards de- tailed video understanding via large vision and language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.959008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.959008Z digest=sha256:8ce015cfab13512cd7ac8e10b4323ca46ef7df32682346dc837b6803020f8a0a

Observation 42287389-5a65-46e5-9e33-26d0b8e3f60a · outbound

This paper cites Extending regular expressions with context operators and parse extraction.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Extending regular expressions with context operators and parse extraction

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.961816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.961816Z digest=sha256:8be3ceb92cab124ddd611877266800a0a67fe3cf22da81de22c4a0ce6fd1b678

Observation 27a06b9c-6ffa-4631-a160-cd3fd2058a21 · outbound

This paper cites Orbit: A unified simulation framework for interactive robot learning environments.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Orbit: A unified simulation framework for interactive robot learning environments

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.964732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.964732Z digest=sha256:1dace90d74b4279dc2567dadff5b4089676a4b8a8412d842868a05b079d87d28

Observation c00eb566-2ce2-43e5-bde4-a7809a56bc3c · outbound

This paper cites Proximal policy optimiza- tion algorithms.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Proximal policy optimiza- tion algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.968342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.968342Z digest=sha256:512a7a49be3a4a286f3ae32e7a47541c0a0264f505358ba91de66d70bdd4d9e5

Observation 658ab226-eebe-4a8d-8709-0c0b27a14540 · outbound

This paper cites Learning quadrupedal locomotion over challenging terrain.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Learning quadrupedal locomotion over challenging terrain

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.971318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.971318Z digest=sha256:ddcdf798dc0640904cee9b4f4047fb4ec4326b7e25e7059df081a4eafe0130d3

Observation 0a7520c2-85db-440e-8c16-b9913bcd0fe9 · outbound

This paper cites Learn- ing robust perceptive locomotion for quadrupedal robots in the wild.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Learn- ing robust perceptive locomotion for quadrupedal robots in the wild

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.974198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.974198Z digest=sha256:111e86291a0c1b425b9fbd96527cd60a348fea9458593f206b57baa122af10e9

Observation a52950d2-2359-4958-b3d9-4674ea5b8aa2 · outbound

This paper cites Rapid locomotion via reinforcement learning.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Rapid locomotion via reinforcement learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.978565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.978565Z digest=sha256:6b2703fbde99612dbb228fdbb337aba5ed069688ad79f43822768ed106a56e6c

Observation 4df7d7b8-2ae4-4d74-9ffe-e5fdf0dc7e83 · outbound

This paper cites Learn- ing robust autonomous navigation and locomotion for wheeled-legged robots.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Learn- ing robust autonomous navigation and locomotion for wheeled-legged robots

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.982299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.982299Z digest=sha256:0c70445dc1521b930175a471421e94e1add1caf0937e3e7d3900e7769cbeb0ca

Observation 6ed90a7e-ec21-4c97-9ff6-2e696f3c1e31 · outbound

This paper cites Bridging the gap between learning in discrete and continuous environments for vision-and-language navi- gation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Bridging the gap between learning in discrete and continuous environments for vision-and-language navi- gation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:17.985828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:17.985828Z digest=sha256:cc2a16b4e8677f741a3647e72fbed40c495b625daad6cbfb51de44d7e115fa5c

Observation 47ea631a-dded-44f8-b46e-29e787c50265 · outbound

This paper cites Waypoint models for instruction-guided navigation in continuous environ- ments.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Waypoint models for instruction-guided navigation in continuous environ- ments

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.175613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:17.989293Z digest=sha256:417337aea8f899e8d0bdcb83c1738f9afb0a8838ffd83215641c7f1e803b5eda

Observation b09310a8-0cba-4858-a773-4d435c63cebe · outbound

This paper cites Sim-2-sim transfer for vision-and-language navigation in continuous environ- ments.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Sim-2-sim transfer for vision-and-language navigation in continuous environ- ments

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.166517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:17.992882Z digest=sha256:9683ea19bc18a6ddcbcfc7e8163e622c7647f3b863c06d8500bafa4c2f49bbdc

Observation febb700a-a77f-4679-8494-70247f3dca8d · outbound

This paper cites Gridmm: Grid memory map for vision- and-language navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Gridmm: Grid memory map for vision- and-language navigation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.156675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:17.996441Z digest=sha256:fb5449bed55020890b2eeb8d860ccfddf0c5aa8c35cf5d5d4995b1e979dce044

Observation 6fb11541-972e-441b-a928-69c4b5e022a6 · outbound

This paper cites Learning navigational visual representations with se- mantic map supervision.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Learning navigational visual representations with se- mantic map supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.146249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.000086Z digest=sha256:74c00357413236fdd8882ea7e5527b27fe40ea2fc6e7719f370f7261cb033233

Observation 43421e5f-c1da-4d46-a0be-5a344640446d · outbound

This paper cites Dreamwalker: Mental planning for contin- uous vision-language navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Dreamwalker: Mental planning for contin- uous vision-language navigation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.134725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.004455Z digest=sha256:ef99e9052d9a7878b0ebd990d1c7454e976c5c53a2bac1ff92fdcb6a037ad890

Observation cf8ec2cd-05bb-41d7-b2ce-91fbc1171ff4 · outbound

This paper cites 1st place solutions for rxr-habitat vision-and-language nav- igation competition.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation 1st place solutions for rxr-habitat vision-and-language nav- igation competition

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.123027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.008151Z digest=sha256:d39f9f6f5f4014e10d4b83e15a3e59e09e8ae0cf2fb26386744abd8886461c5c

Observation 5d1bef50-e143-4d6c-a505-b2388fd4d18c · outbound

This paper cites Etpnav: Evolv- ing topological planning for vision-language navigation in continuous environments.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Etpnav: Evolv- ing topological planning for vision-language navigation in continuous environments

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.111020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.011594Z digest=sha256:66823bae034b671750e917ef41e64622ebc095afa909754ea2a3a699fc2cd8b7

Observation 4b08ae4b-f300-4a04-a335-cc5e44873171 · outbound

This paper cites Lookahead exploration with neural radiance representation for con- tinuous vision-language navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Lookahead exploration with neural radiance representation for con- tinuous vision-language navigation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.100909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.015078Z digest=sha256:fa4f9ee492b5667556c0241ffc377e1ec65a0eb87c89d0642d004aa9a67d7652

Observation dc1d74ff-aded-4e15-9399-86acae984a74 · outbound

This paper cites Bevbert: Multimodal map pre-training for language-guided navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Bevbert: Multimodal map pre-training for language-guided navigation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.090412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.018877Z digest=sha256:5cfcbcb06dea060e8108b631587b6d32d960a6803d988cc9e17435fe0ef45b44

Observation 52b2357d-cdce-48a5-a0f3-e42bebcb84e0 · outbound

This paper cites Scaling data generation in vision-and-language naviga- tion.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Scaling data generation in vision-and-language naviga- tion

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.079230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.022680Z digest=sha256:dd0ea5c62196a9be49b867aa7b313057a49a89f617af52df575729cc278eed0f

Observation 52f18842-3c32-4575-9805-d35b1afe0491 · outbound

This paper cites Topological planning with transformers for vision-and-language navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Topological planning with transformers for vision-and-language navigation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.067347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.026744Z digest=sha256:0182797366baedabc147401b59867377bf4b3a554b3f7b48fa0900dbf864ef20

Observation a9d66aaa-8382-4724-9e11-8ad16f82b235 · outbound

This paper cites Language-aligned waypoint (law) supervision for vision-and-language navigation in continuous environments.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Language-aligned waypoint (law) supervision for vision-and-language navigation in continuous environments

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.055941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.030545Z digest=sha256:fa378b62c77796a31cb4b6d0f2a5bad6d1008558e14c7deff5c8c76549049206

Observation 10a49391-f1e2-452b-b672-65b3f8b8f1f1 · outbound

This paper cites Cross-modal map learning for vision and language navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Cross-modal map learning for vision and language navigation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.045962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.034016Z digest=sha256:64c0b3b05fa7039848cd48742bfd20f45741d5df14841026cc630df28bfe3d25

Observation d1fc72cf-8589-4b90-969e-a40fc9d0ffe0 · outbound

This paper cites Weakly- supervised multi-granularity map learning for vision- and-language navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Weakly- supervised multi-granularity map learning for vision- and-language navigation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.034336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.037585Z digest=sha256:1d99e3bb0f7500295082aa6d55e1ccf7b2ba145d8187f1219025de5573a17bc4

Observation a9ad9588-d3b4-4ac1-b561-84df3c3886a4 · outbound

This paper cites Affordances-oriented planning using foundation models for continuous vision-language navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Affordances-oriented planning using foundation models for continuous vision-language navigation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.022390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.041249Z digest=sha256:913aef03b15b5c2291b700f5863a75a2bdffb0b38be7d09408acecfea589e605

Observation aed9c626-93fa-4af7-a5b5-0efefccf9468 · outbound

This paper cites Beyond the nav-graph: Vision- and-language navigation in continuous environments.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Beyond the nav-graph: Vision- and-language navigation in continuous environments

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.011844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.044935Z digest=sha256:e069e5ac0e3d0f14b0b5bb88bac7690951a6335138d27f544eb86c32110827ca

Observation a35b5962-e278-41b2-8d07-08a71924346f · outbound

This paper cites Li, Gaowen Liu, Mingkui Tan, and Chuang Gan.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Li, Gaowen Liu, Mingkui Tan, and Chuang Gan

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:19.000309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.048729Z digest=sha256:2aafac28e6cf9265e38d56c42d697f0a8a976f63493642d6d05dcbca5c3a5094

Observation ec12e8a0-0f45-4e9b-a76f-938433fa0a4a · outbound

This paper cites Towards learning a generalist model for embodied navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Towards learning a generalist model for embodied navigation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.989035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.053107Z digest=sha256:bc951b9cfba84670dd869fb2c4ea4d36bbcf11fcc3f63fa993b6eecfc4218f9b

Observation fdaf41ca-a68b-48a6-93bc-1742da40723f · outbound

This paper cites Scene-llm: Extending language model for 3d visual understanding and reasoning.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Scene-llm: Extending language model for 3d visual understanding and reasoning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.976172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.057131Z digest=sha256:12207d2e743cddefe5fd7623f17366486d55dbb0543756bb1b5ed8f8137b2314

Observation c9ab142d-3c1f-4452-8626-45a16f0c55e2 · outbound

This paper cites An embodied generalist agent in 3d world.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation An embodied generalist agent in 3d world

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.963231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.061045Z digest=sha256:b087eb76d002115d808e6492eb8bde9780c557b3c45645f6109691fa4dc8d0c9

Observation 8cda76cb-cfa7-4054-a680-331f0a34d6a6 · outbound

This paper cites Deep modular co-attention networks for visual question answering.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Deep modular co-attention networks for visual question answering

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.951415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.064592Z digest=sha256:b155c258cf0c6b9702ee6cbec1a46b544050ec9a67dcb00d96aa24963b813e42

Observation a839c7b8-00c0-491e-827d-5c52b8cb752e · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.939559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.068196Z digest=sha256:9a6466779e3021fb177a39ed30c484cceee0334e6e9ad34b001fd27d5cb48b98

Observation c8c102a5-fd58-40aa-8a5b-ba07e1aef51e · outbound

This paper cites 3d- llm: Injecting the 3d world into large language models.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation 3d- llm: Injecting the 3d world into large language models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.928179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.072384Z digest=sha256:8f833ee170f8f8e915a445dacefade69440cc677661de1eecccd6fb221c9547d

Observation bb6eca6d-821b-4875-9daf-0046ba95b65a · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.916970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.075896Z digest=sha256:ef9af7b0acfae3f282eb918dd70f5bb39f62d4f2a0921597d6fe3ccfa6bd6127

Observation 9772d1cf-c948-4466-a4c5-fd713197f4cb · outbound

This paper cites Chat-scene: Bridging 3d scene and large language mod- els with object identifiers.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Chat-scene: Bridging 3d scene and large language mod- els with object identifiers

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.905676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.079595Z digest=sha256:f1ea8145eb99188f9d35a4047ea7c1d3c7f713ae2cde3200c819c99ef2e68753

Observation b5bff175-0b86-4386-9dc5-725f84714d4b · outbound

This paper cites Deep whole-body control: Learning a unified policy for ma- nipulation and locomotion.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Deep whole-body control: Learning a unified policy for ma- nipulation and locomotion

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.893437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.083279Z digest=sha256:86e3dc835fc9262231f1ba8479d3dca46f8725adfa178045df23f6e582749611

Observation 79208f9c-74ad-4304-9292-b0b51c84219c · outbound

This paper cites Habitat: A platform for embodied ai research.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Habitat: A platform for embodied ai research

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.882359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.087220Z digest=sha256:d0b52a196bf541d0db274960c5fe1da5e12c399921ca5618e169b24bc44163a7

Observation a2436d9b-a4ef-4d8e-b8b1-7a933ff0c338 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Robust speech recognition via large-scale weak supervision

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.870761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.089979Z digest=sha256:07d887344628a051abe6b950448af538ea0eea882203891e3eb8348b5c2a5ffc

Observation 817ab644-2f4d-40f9-b335-5cbaa4cb875f · outbound

This paper cites Obstacle avoidance and naviga- tion in the real world by a seeing robot rover.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Obstacle avoidance and naviga- tion in the real world by a seeing robot rover

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.859776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.092970Z digest=sha256:653d5ef75c5f6b71051e72ab0c16ceeef9b3b2697db72188d8287d9e4d04d9ef

Observation 768bacb8-77ba-465d-bbd4-d2e010554ccc · outbound

This paper cites Sonar-based real-world mapping and navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Sonar-based real-world mapping and navigation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.847566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.096102Z digest=sha256:8a7947ac831232f39d34386305fb640b08ad5cf0823c4230ce635115f45643b7

Observation c1177870-473c-4f9c-ad67-bb4c23cb61ac · outbound

This paper cites Robust monte carlo localization for mobile robots.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Robust monte carlo localization for mobile robots

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.836025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.099092Z digest=sha256:b4030f16263a13ca7be225b3d4adf2151f6fc85acccad3940f53e467db31ff0b

Observation 7e2754ef-6ace-4c45-8185-bf5c0fa83369 · outbound

This paper cites Navigating to objects in the real world.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Navigating to objects in the real world

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.825550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.102227Z digest=sha256:2815178418a04b12ce311789fee1a918dd3b168b4a21ec1d72a0f8bf863ac818

Observation 4ebb8a00-e351-40d0-8e52-bbc16aaa01d6 · outbound

This paper cites Minerva: A second-generation museum tour-guide robot.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Minerva: A second-generation museum tour-guide robot

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.814922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.105109Z digest=sha256:2df2f88fa9cf2f966385ec031e3a01dc52401d5fe516bed434a1750835b2fb85

Observation ed81421a-b222-4ec5-8499-d088410b203d · outbound

This paper cites Kinectfusion: Real-time dense surface mapping and tracking.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Kinectfusion: Real-time dense surface mapping and tracking

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.806121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.108169Z digest=sha256:535a9d7f4f348f90082ac61e3c8954d92c83888a4ca949f962e6969a15da6c99

Observation e555f857-e0ad-417c-a19e-4d76eaf8dd69 · outbound

This paper cites Monoslam: Real-time single camera slam.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Monoslam: Real-time single camera slam

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.797079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.111431Z digest=sha256:5ea754432a825941d79807356ddf4ee3f83e744275008a03be648ec413694351

Observation fee5065f-9262-4295-9d39-f15335605d15 · outbound

This paper cites Visual-inertial navigation, mapping and localization: A scalable real- time causal approach.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Visual-inertial navigation, mapping and localization: A scalable real- time causal approach

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.787344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.115159Z digest=sha256:4670090f7483f1e158a339ee2182f9ac897930832c7940a0e48d02d410e70b65

Observation 9fad0211-9fdc-4b78-99cf-13ab9e3ef97a · outbound

This paper cites Gated-attention architectures for task-oriented language grounding.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Gated-attention architectures for task-oriented language grounding

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.776546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.118952Z digest=sha256:c6b297ffaeae8852875d53ec14374a34060214fb0ae2a95af3773702af2dba38

Observation 074665ac-1205-4171-9e32-ecb935306093 · outbound

This paper cites End-to-end driving via conditional imitation learning.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation End-to-end driving via conditional imitation learning

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.765935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.122588Z digest=sha256:fd19f1320e32bcc0823c4f241aabaf51c29ce5dd2140ac3767316b3c1660f0dd

Observation 03f69ab1-da58-4f8c-a927-36739973faae · outbound

This paper cites Human-level control through deep reinforcement learning.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Human-level control through deep reinforcement learning

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.754810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.126445Z digest=sha256:b7e0012f45bf9e1c2d23864c13ab8cd6570f42b6a460434d11a1a09538cc9ff0

Observation 694ac538-c33d-4070-b5e4-863dc42e8d3a · outbound

This paper cites Continuous control with deep reinforce- ment learning.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Continuous control with deep reinforce- ment learning

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.744149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.130311Z digest=sha256:5d58dd98b7a95421fb59849c58ef832fd27d7393ba9eaea28d4ca55596822ca5

Observation 88252727-b064-4087-b32e-9a5fb7e24972 · outbound

This paper cites Reverie: Remote embodied visual referring expression in real indoor environments.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Reverie: Remote embodied visual referring expression in real indoor environments

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.732798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.134064Z digest=sha256:ff9ec2c097f689e5a68630d6eda5660c69a66346d35641e0ca2c5f3e18819e08

Observation 810bbb99-44e4-4bcd-8c52-c29cfc621296 · outbound

This paper cites Matterport3d: Learning from rgb-d data in indoor environments.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Matterport3d: Learning from rgb-d data in indoor environments

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.722233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.137801Z digest=sha256:d4b43a247a9e04fdaccfa65813641b085ccbae0dc6fbc3cbdcd7e014231c6998

Observation 28b42a25-66aa-4504-8c65-cb168d0fa6b9 · outbound

This paper cites Speaker-follower models for vision-and- language navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Speaker-follower models for vision-and- language navigation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.711465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.141768Z digest=sha256:d3629cd66493b8856b2128eba8ae96d9462fbe8328cb5f2b575e67f299e081b8

Observation 8d86520f-65b8-444b-a154-fc8dd8d0b522 · outbound

This paper cites Self-monitoring navigation agent via auxiliary progress estimation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Self-monitoring navigation agent via auxiliary progress estimation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.701082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.145628Z digest=sha256:42cf49cb33392591965e25b1b3b7629b9a47133f1a04685478ac27df39a79982

Observation 24cac0a8-a7df-4165-a822-731150688f90 · outbound

This paper cites Tactical rewind: Self-correction via backtracking in vision-and-language navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Tactical rewind: Self-correction via backtracking in vision-and-language navigation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.690082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.149708Z digest=sha256:ae1a52c5024535051bd00f9a5e1fa18ab8e104d3b985bc3477a22b1240e61551

Observation c594b6ea-60dc-4d2c-8d78-9d24bf5a1478 · outbound

This paper cites Language and visual entity rela- tionship graph for agent navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Language and visual entity rela- tionship graph for agent navigation

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.680363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.153608Z digest=sha256:746a2a7fc40aa98d76905bb0a5d6869c35c7ce05b6ffdb16ebcd9b146bdd3dba

Observation 7b1a9998-751d-47df-801c-00a4cc314965 · outbound

This paper cites History aware multimodal transformer for vision-and-language navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation History aware multimodal transformer for vision-and-language navigation

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.670495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.157221Z digest=sha256:b37d2a44281c09a1b66283dcdc25c667dc4ec2d81e5c6858bb0a13932fc9df55

Observation f3aaaf58-9ce7-4078-9c22-983ce7e00219 · outbound

This paper cites Mapgpt: Map-guided prompting with adaptive path planning for vision-and- language navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Mapgpt: Map-guided prompting with adaptive path planning for vision-and- language navigation

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.660046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.161044Z digest=sha256:054ee4ffe8be884a0db4bb737261bc5a69fec5c0baff0cd2ef23ed5a8d959739

Observation 1fba79c5-22a9-4a44-91b9-ccc40f8a6136 · outbound

This paper cites Navgpt-2: Unleashing navigational reasoning capability for large vision-language models.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Navgpt-2: Unleashing navigational reasoning capability for large vision-language models

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.648783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.164687Z digest=sha256:e0db0c29fbde69d19bac130e90a0856b08f9a7ebb5f72fb3b006a81ebe7ee042

Observation eb2f9478-1a99-44ef-956a-d6ff04ee56a0 · outbound

This paper cites Robust navigation with language pretraining and stochastic sampling.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Robust navigation with language pretraining and stochastic sampling

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.637499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.168252Z digest=sha256:3ec63e2b25d0c25dd008ec6558091f3b06ef18f32a86cebfb5909a4ac6b1f0fc

Observation ca0fbb2e-7ade-4a31-a7fc-cdca3270c246 · outbound

This paper cites A new path: Scaling vision-and-language navigation with synthetic instruc- tions and imitation learning.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation A new path: Scaling vision-and-language navigation with synthetic instruc- tions and imitation learning

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.626468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.172132Z digest=sha256:d04e16aa8d0d7188645a20ee36ea02bede59804af55ab9c1d497148cfe02d13b

Observation 590399c6-d6f6-404a-8f8a-70744b063820 · outbound

This paper cites Hierarchical cross-modal agent for robotics vision-and-language navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Hierarchical cross-modal agent for robotics vision-and-language navigation

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.614924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.175874Z digest=sha256:230dba68767dc6f23bd9f2fcf40df89540d61f8b6206b7b4507f34e8fa8b516a

Observation f96e3868-d9cf-4090-89e3-24f4062d0c67 · outbound

This paper cites Octo: An open-source generalist robot policy.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Octo: An open-source generalist robot policy

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.602929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.179458Z digest=sha256:47a077142e03fa2da6cdf062dde75e7299437145f71b3c739c60289be45718a0

Observation ead327bc-0b54-4785-84a6-aa6763529d22 · outbound

This paper cites Scaling cross-embodied learning: One policy for manipulation, navigation, lo- comotion and aviation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Scaling cross-embodied learning: One policy for manipulation, navigation, lo- comotion and aviation

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.592189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.182975Z digest=sha256:1e4b9eef97e1c305adb8061b5f8b69a6c67df283976f370e76dcef555b674605

Observation 8213583d-38df-4f8b-93f6-56352140919d · outbound

This paper cites Pushing the limits of cross- embodiment learning for manipulation and navigation.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Pushing the limits of cross- embodiment learning for manipulation and navigation

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.580919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.186581Z digest=sha256:73bffc28c474c47f5c72e8d7ff945f930b9a188856c75e37a5e71e2c22be1cc9

Observation 61d3d485-36e4-4b27-a63d-8259bf2cd161 · outbound

This paper cites Poliformer: Scaling on-policy rl with transformers results in mas- terful navigators.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Poliformer: Scaling on-policy rl with transformers results in mas- terful navigators

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.570311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.190271Z digest=sha256:27f995c2ccc39cf22dea045efd46db3acea4df73cc6b21c36365098c50c7ada8

Observation 467acaa3-6b37-43ef-a20c-c5456742bbfc · outbound

This paper cites Gnm: A general navigation model to drive any robot.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Gnm: A general navigation model to drive any robot

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.559163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.193977Z digest=sha256:1e228e9794e21e250164aec06b2ba29af62830cfb0c254f98197c3b7b8b36565

Observation f0ab59dc-4a01-4c91-8f96-9b5d706f4b89 · outbound

This paper cites Nomad: Goal masked diffusion policies for navigation and exploration.

NaVILA: Legged Robot Vision-Language-Action Model for Navigation Nomad: Goal masked diffusion policies for navigation and exploration

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:18.548094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T21:30:18.197698Z digest=sha256:600243c10faf08bb42bc7f5afbb14c4e6bcbddb151f5e27ae38abbe46ecce436

Pith citing papers

Observation fac3d417-f91f-4a1e-a41c-8d39f05ae5cc · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.167538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:6955059268378458f5706d347c65cf6dd76651584561664f31bd870c19618f6c

Observation 87032ec4-0de6-42c4-8414-021e309234ab · inbound

doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation cites this paper.

doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:17:47.689534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:17:47.689534Z digest=sha256:6728cbf0fd4993a87f7812d41f325b1b7a37153bd605986ea148748a5728b8cc

Observation 836504b2-6e5d-48a7-9dc3-fa941533a835 · inbound

FAST: Efficient Action Tokenization for Vision-Language-Action Models cites this paper.

FAST: Efficient Action Tokenization for Vision-Language-Action Models NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:52:31.934448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:7d0319d9bb77226fdfe0488397a859ec6503c4fe71b3ee47c9ff8ebfaee24112

Observation 64d3b1c8-7ccc-47a6-9849-95ad8e8a46b4 · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.989303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:0b8716d1af81eee4c46b8611b7d031750f683ec94b8e08d33c6bb3e57feccae7

Observation 5bef1725-593f-413c-94d9-7ccf30136d2b · inbound

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding cites this paper.

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:21.311984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:21.311984Z digest=sha256:29c1c7a3e1a5c6720e6e7b8e4aad1ea664eac136716a62771916eb895fb38e55

Observation 257800ec-aff5-4796-b802-35990f8bda71 · inbound

RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning cites this paper.

RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:12.624697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:11:12.624697Z digest=sha256:60b3c7b53d74bae8ececc29225fba0fc92150691bb48e1add73df893172f9599

Observation 4c7dc942-9294-4115-b044-85ea5dfe130a · inbound

SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning cites this paper.

SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:41:30.849601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:41:30.849601Z digest=sha256:31d4bb27edf8409a95434d41efe7613c6c40cc04b6c765a9bf3f4662f8791382

Observation 176cfe1f-fe6c-4c1b-8191-a8d663bae688 · inbound

A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI cites this paper.

A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:46:40.167610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:46:40.167610Z digest=sha256:1ea1b806fe8ba397634d3fc5b1207122fad5bdc114688bfadfff21635be0cb81

Observation 6e5caf3e-bffc-4b63-a42d-14a7215d9f47 · inbound

Omni-Perception: Omnidirectional Collision Avoidance for Legged Locomotion in Dynamic Environments cites this paper.

Omni-Perception: Omnidirectional Collision Avoidance for Legged Locomotion in Dynamic Environments NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:21.019683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:21.019683Z digest=sha256:003177df84094254d44e42c8bf39cea49a436a652a554f6197a1d1b3f8e4ba59

Observation c73ab88c-39f1-44a1-b111-8abc20fdbeba · inbound

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation cites this paper.

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:05.254189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:05.254189Z digest=sha256:997e02dbe7834181cee8c9d41bf179ab25a9518917d9e4261a9c69648f3a0fc4

Observation 59f4dc6f-7814-4d5a-8ac6-845a407fbd16 · inbound

Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits cites this paper.

Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:16.513264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:16.513264Z digest=sha256:ccd00382923a879e391d9a2e3bf61d324586b1b7b83c1bb641e9c24a6f8d93fe

Observation 2eef2a0e-9ef3-4ab0-ba8c-0c7f41be2c87 · inbound

TrackVLA: Embodied Visual Tracking in the Wild cites this paper.

TrackVLA: Embodied Visual Tracking in the Wild NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:58:39.930337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:58:39.930337Z digest=sha256:bd0cbf390d786338a953ab85c267c2bac0a251781e5edced57ce3e419b35b89f

Observation 38816cc3-5d58-4c66-9782-0b54dcd0d5d0 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:09.666040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:09.666040Z digest=sha256:05a9117f2f8bfae65c4f9d70771b59bb5f125d4010172f96c3de3956d321a0b0

Observation 396dad35-108d-4f25-beea-2953d1be9b1d · inbound

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks cites this paper.

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:51.175126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:51.175126Z digest=sha256:1011e2027f8bd8f7ab6bcfe297981c1ce075bdbb4b901a93fedce913a93157be

Observation c65785dc-dc54-48cc-8a95-cbdfc092f495 · inbound

Real-Time Execution of Action Chunking Flow Policies cites this paper.

Real-Time Execution of Action Chunking Flow Policies NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:18:51.663506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T14:18:51.613045Z digest=sha256:23fbef322f332d5015c6b7b654ec4d8db5870284cc6517f0fe9a6ab83975dedf

Observation d427094e-e49b-4e57-8b62-ef3e1532db70 · inbound

OctoNav: Towards Generalist Embodied Navigation cites this paper.

OctoNav: Towards Generalist Embodied Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:40.740461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:44:40.740461Z digest=sha256:011be2c7d3135cf96914311fad7bb9b84f374f573763ec9887a406dd1b381742

Observation ed17b4aa-87dc-430f-a22e-7824d1b5cc3f · inbound

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? cites this paper.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.630568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.630568Z digest=sha256:d1fe0015b9f4f3daaf5abaab4be4cc260cbc84471fdaf15a733bb67e9c842275

Observation c7e8f828-9987-4e2f-887e-4330e5f0a59c · inbound

Hierarchical Vision-Language Planning for Multi-Step Humanoid Manipulation cites this paper.

Hierarchical Vision-Language Planning for Multi-Step Humanoid Manipulation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:52.812767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:52.812767Z digest=sha256:b6f82ed81ec2f42ba3e87dbd90dbf49cde1b0c57942c50b19e13e46aade0b153

Observation 80347b0f-17f0-42af-ba60-2eea9961e4d0 · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 153

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:08:35.152944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:83bed647827d95877482fee6abc08029a9dfcd564b84bf70b30e72bf19be6804

Observation ce5d2b74-c720-4015-ae13-4a7af8fd3cc0 · inbound

LOVON: Legged Open-Vocabulary Object Navigator cites this paper.

LOVON: Legged Open-Vocabulary Object Navigator NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:01:24.035518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:01:24.035518Z digest=sha256:29afe95ad1d2b5fcc42899e078194705e70bb80ee527bc8f2093b7ecaee2cb30

Observation 74ba35c4-3fda-4de6-8faf-1e5eeae78498 · inbound

Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning cites this paper.

Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:30:47.046180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:30:47.046180Z digest=sha256:027126ee806db6f19c87dd696f635660f879e95250507d376061716dd329eca4

Observation 1a6cdeb1-03f9-4203-8b30-d80f04486d72 · inbound

EmbRACE-3K: Embodied Reasoning and Action in Complex Environments cites this paper.

EmbRACE-3K: Embodied Reasoning and Action in Complex Environments NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:33:08.649485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:33:08.649485Z digest=sha256:7806df6599394cc389f5bb90ef7beaf25362e7fc781a1c876ccfdb997dd8bf9b

Observation 0ee95326-0488-4216-a12b-b8dff9a04446 · inbound

Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation cites this paper.

Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T22:50:43.715480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:50:43.715480Z digest=sha256:26e3bb9b7ed87214cf013cb447a0f2ebd24c4ebd30cff9843ebdf997341fb799

Observation a42816a0-a620-41d5-9747-aa27b76a829a · inbound

SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation cites this paper.

SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T22:46:42.049089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:46:42.049089Z digest=sha256:3b90dc57d285b12d0c1fee9acb819321ed3db7f5f96ef1a5c3681d101b285d92

Observation ded1f365-b984-4087-a96d-a245e3a528f9 · inbound

CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model cites this paper.

CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:07.196850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:31:07.196850Z digest=sha256:c0e2b1002d0169cd1662a730e882d407fb0a960b4ed58675180a5eddcdbd764a

Observation f291dc62-9fa7-4f61-b385-8632ae5e69f7 · inbound

SPG: Style-Prompting Guidance for Style-Specific Content Creation cites this paper.

SPG: Style-Prompting Guidance for Style-Specific Content Creation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:56.869301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:56.869301Z digest=sha256:bbc033f2ffcd6d11004d5c3c1aa0227692207b4df9d80ffc2213754181f02716

Observation c7b4a824-620b-40e8-989f-60541a878e8f · inbound

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models cites this paper.

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:28.436191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:28.436191Z digest=sha256:fb0827dd0c95a643221a29b7fdc9e145acf1f7f94214006400418390fd2aad00

Observation 958c5fc4-594e-44d3-905f-748fd98e4e2c · inbound

From reactive to cognitive: brain-inspired spatial intelligence for embodied agents cites this paper.

From reactive to cognitive: brain-inspired spatial intelligence for embodied agents NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:20.527129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:20.527129Z digest=sha256:7c17f4d4e71e609021c73ef3ff58cbb1aa6131dc7d03592012d07f681f679ea7

Observation dc10f7e0-86da-404d-8597-82dfcfea904a · inbound

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment cites this paper.

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T13:54:49.258914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:54:49.258914Z digest=sha256:2456a52cf3f2edd4a194a1e2d0c8c8112d196fe9904f8df9d1b4d35bc1769196

Observation 347d0f17-a116-4e37-9d7f-8a777b4109dd · inbound

Robix: A Unified Model for Robot Interaction, Reasoning and Planning cites this paper.

Robix: A Unified Model for Robot Interaction, Reasoning and Planning NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T12:59:02.133626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:59:02.133626Z digest=sha256:0465a4dc752b3aefd480450a42d8bf5670a3d37770360b0257921a47bfdce29c

Observation 6eb69b24-e8c1-4332-bb3e-2b5f293abe54 · inbound

Nav-R1: Reasoning and Navigation in Embodied Scenes cites this paper.

Nav-R1: Reasoning and Navigation in Embodied Scenes NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T17:31:38.800522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:31:38.800522Z digest=sha256:a027ca6b421a81f63f08dcf4812fe7f499f970ef733c6ddcfaf2e5f0834af407

Observation 64ed98ce-ebd6-4121-af1d-d60f1761d8ef · inbound

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation cites this paper.

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:51:09.111792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T08:46:51.278024Z digest=sha256:6f3958a80cc9e2548788db55b571080b52ea94035922e85bc9a3ba0cdabcb7f3

Observation 3fe1c212-e7fd-4c3c-8912-a7930b086f03 · inbound

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation cites this paper.

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:05:15.979141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T21:02:38.013115Z digest=sha256:8ab645842cc6b3b25e39a28e9fb1b1c6b334a0b4fa769a7180606263f94348ad

Observation 22394142-5144-4124-9272-e14f4f880d3f · inbound

AstraNav-World: World Model for Foresight Control and Consistency cites this paper.

AstraNav-World: World Model for Foresight Control and Consistency NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:28:20.825261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T19:23:58.769472Z digest=sha256:e1fc6102bbba518bd36c9dacdd89f85ebba42440fd173be925e1fc0414178e90

Observation 9a86e709-e12f-4f98-b08b-8ebdb98f2a4a · inbound

VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness cites this paper.

VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:51:08.250364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T14:50:58.269817Z digest=sha256:9b26a1067b49bb26369763c993b870042e9d91d3aee56cb125d315fc68b6b7b4

Observation 5db8e094-517c-4759-b35a-d9a0eb429c3d · inbound

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation cites this paper.

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T17:15:42.360530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:15:42.360530Z digest=sha256:987aff732c47d9b86eab4853e22a0a0bbcf24cda0c5fce6214857ee849968c7e

Observation b22d2ba4-de02-411a-9a1c-08d2a5beb5fc · inbound

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation cites this paper.

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T05:43:08.245773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:43:08.245773Z digest=sha256:6f601f8d74c01b210a6ad057c9571ded93cf372d71ca339a7ae31588b4f77092

Observation 68842129-da80-4f90-be5e-25b2877a6721 · inbound

Learning Task-Invariant Properties via Dreamer: Enabling Efficient Policy Transfer for Quadruped Robots cites this paper.

Learning Task-Invariant Properties via Dreamer: Enabling Efficient Policy Transfer for Quadruped Robots NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:53:11.811580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T19:49:23.086498Z digest=sha256:ab362fa320f50741abd5602b475368ca3e47923e5b32bd1e1e0ec1a2be968cae

Observation 36db350c-d7f4-4364-9154-3106a2bf5bec · inbound

HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation cites this paper.

HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:41:01.469951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:01:13.350219Z digest=sha256:d78b61c64178d79c87d9ae6223c5aec0f8e7fcfcc477f6cadc6ded6ce76fd3ad

Observation a483cbf6-61b7-4079-acc5-4094ee500782 · inbound

Visually-grounded Humanoid Agents cites this paper.

Visually-grounded Humanoid Agents NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:21:04.468428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:10:32.710853Z digest=sha256:97788eaa12c2f1a08ec6de964037db6fe4531416f1c0bfbbbc82c8d7558e1817

Observation f7431370-e85e-42e1-8961-779024c897df · inbound

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap cites this paper.

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:29.674502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T13:48:08.135538Z digest=sha256:20d67474491a4b30cb314a30f11e9017f46bdb9145a87b44ef34723bdc32832b

Observation 7b9d5fbe-234e-48c0-9b58-099b5820b2bb · inbound

CART: Context-Aware Terrain Adaptation using Temporal Sequence Selection for Legged Robots cites this paper.

CART: Context-Aware Terrain Adaptation using Temporal Sequence Selection for Legged Robots NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:55:25.378958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T12:51:40.003948Z digest=sha256:2dabac8f3b2a1e7b80be63e9eb55eaca19aa4fbfb5a0b230fcc4560c803bf7e6

Observation 6266283e-8f1e-4979-9cb4-72232016f271 · inbound

CART: Context-Aware Terrain Adaptation using Temporal Sequence Selection for Legged Robots cites this paper.

CART: Context-Aware Terrain Adaptation using Temporal Sequence Selection for Legged Robots NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T20:22:05.957224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:22:05.957224Z digest=sha256:f21d6aeedccfbae42ab4fd09f730a73951b1ede564f2bdd4d1705c37195e66c2

Observation 84d45e28-7c51-45f1-8658-aac6ce8e678d · inbound

Think before Go: Hierarchical Reasoning for Image-goal Navigation cites this paper.

Think before Go: Hierarchical Reasoning for Image-goal Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:51:10.522123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T05:43:27.972164Z digest=sha256:b789beed9b5861970b2290814af98990f3e7a75e66b5ccccabbf373c535aac35

Observation fe194dbe-af9d-46ea-968b-9dbfa45f548c · inbound

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation cites this paper.

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:45.831019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T06:50:34.310831Z digest=sha256:e119dd6c4104d9a9e4b0745e3ed9673c0457b921f57c4e9ae65823c910563750

Observation 163b26c4-4daa-4f7b-950c-d0651a2e70c7 · inbound

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation cites this paper.

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:31.261652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T05:05:33.606975Z digest=sha256:ce228d77dedc99b0d8d84ae41fa7bcd97b6e54d504c114402e1219591abed7cf

Observation 388dea21-e8b9-4e71-8a59-2c428a5b97e3 · inbound

Beyond Isolation: A Unified Benchmark for General-Purpose Navigation cites this paper.

Beyond Isolation: A Unified Benchmark for General-Purpose Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:27.549019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T04:33:53.557357Z digest=sha256:366aa6fe79649c7d0cf97fb24a0111095c4f60ee4b2b65ee9a3382beaa7016d1

Observation b01388b8-e8c8-4d07-b0bd-accfe5ee30ea · inbound

Terrain Consistent Reference-Guided RL for Humanoid Navigation Autonomy cites this paper.

Terrain Consistent Reference-Guided RL for Humanoid Navigation Autonomy NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T15:23:07.747403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T15:22:38.188537Z digest=sha256:ed090c2a32ea9a92b60175bcd6626037b4c167ad127b36c2d227ebfd581896da

Observation 55735769-cfc8-49e1-91ae-76d4454c88e3 · inbound

SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation cites this paper.

SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:33:19.148446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T13:31:16.012419Z digest=sha256:e3e768e336899cac471cdb42f477d804d0a1e5a0d6cb741a799cb5b2f739fc86

Observation 84154134-46ff-4693-a428-37c25129d403 · inbound

SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation cites this paper.

SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:45:00.664705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T19:42:41.238072Z digest=sha256:385c4c5b8364daefd0f363154f581e9dd395ae2ec2a41e12aa15d64bf9117b15

Observation 20405ac0-091c-4448-8806-d39c7ef3aef3 · inbound

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation cites this paper.

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:46:10.584337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T06:45:35.920991Z digest=sha256:1521c2c3746dea4ddab851d3114583c455b7baa9fb267ae0bca66cca995b64ae

Observation 2120b067-32a6-4a06-8585-b872eaeb6435 · inbound

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation cites this paper.

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T04:46:04.562531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T04:46:03.020800Z digest=sha256:da1cc991a19079768db908717575fb5039e38b278ca3fba37cff6a6dd804c3e5

Observation deadbe25-e701-4ce6-8412-8e2732944916 · inbound

G-DRAGON: Geospatial Reasoning and Dynamic Planning for Retrieval-Augmented Outdoor Navigation cites this paper.

G-DRAGON: Geospatial Reasoning and Dynamic Planning for Retrieval-Augmented Outdoor Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:53:59.368576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T21:49:54.950611Z digest=sha256:55f7a5625d247e3dffe466a853a66b68ab4cc77f65b4aa318484617c2e7f17c1

Observation 0ad83405-314b-48d1-861d-db91b4970d60 · inbound

Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation cites this paper.

Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:03:40.848584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T17:01:20.073666Z digest=sha256:49110e2228974dd56222f91c196b542eeb807dc2f91df411a3505cb68393afd2

Observation 7d8c41c4-e31e-4da3-bf8b-14aaa25f3361 · inbound

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation cites this paper.

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T11:53:23.916822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T11:48:14.888295Z digest=sha256:aec9da63cf88398759adaf2b9f84dafc7d828d833cc922d25473c13715f1ca33

Observation 1e1a8380-b7ab-4042-b09e-9aeb5418ec3e · inbound

PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs cites this paper.

PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:53:40.418154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T16:53:17.504094Z digest=sha256:b2e7b7a63a9c83a56fd152fb38416af8826395e4e6fe0a0dc36d53f7bbb37073

Observation 2bd71fac-4c4b-4551-9a56-6e6f7fdb41bb · inbound

Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation cites this paper.

Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.934190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T14:49:30.585724Z digest=sha256:966e5628b961046e84ed9053721e7c1d7ce2a2eb66964cd118a5614e92ff4d94

Observation b79fac24-9072-4f62-8a59-08199f454518 · inbound

Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation cites this paper.

Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:15.795177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T15:39:13.803571Z digest=sha256:22bbebf8b9b464579ccba633f62362d32963de6c5b698d166db844b629cb8197

Observation c09ffcf2-259f-4859-bf6e-bc2d6c216df2 · inbound

Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation cites this paper.

Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:26.123860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T11:01:08.668612Z digest=sha256:6a21dda69f0876f054267670ca7ead0ef3414c44745b795e6c3a1913cba05c81

Observation d7115002-5b6a-4bdb-83f5-cbf262b684e5 · inbound

Beyond Waypoints: A Trajectory-Centric Waypointing Paradigm for Vision-Language Navigation cites this paper.

Beyond Waypoints: A Trajectory-Centric Waypointing Paradigm for Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.750835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T21:40:58.329546Z digest=sha256:5294dd4734cb8a214e505ef7117000bfd3837ced39785a74356361121d728274

Observation 17ad59cb-7c8d-4d41-9b1f-2347ee9a9ca5 · inbound

Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models cites this paper.

Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.671095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T12:52:38.764427Z digest=sha256:bd18d10f7b1ebc5eb2e9ab223830c769688181e175c30073c5cfd435cdcdd368

Observation 8eabb674-170a-4a53-8da8-8543c03f3941 · inbound

From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation cites this paper.

From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:38:04.652491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T09:29:43.030058Z digest=sha256:0012553f14bf755e55d9adb0e3e278918891b34d9f6c514eb1fa3abb309f68e0

Observation 899df63f-9ade-4e5b-85f9-6121d359961f · inbound

VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision cites this paper.

VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:38:59.469725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T00:15:26.619238Z digest=sha256:031c7b6964a37cdf42a2a44884030bc023a0ca124be9b9422013e4f392b4c05d

Observation cbada972-9842-43c1-b840-9bb5f6f2b4bc · inbound

VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision cites this paper.

VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T11:03:51.561853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:03:51.561853Z digest=sha256:03e11393501817f461f2d804e83c4604f686ef0708c7f5829afa654ac72056e7

Observation 8ccd85cd-0fa4-4d8d-9ba3-6250c061aa6c · inbound

Slow Brain, Fast Planner: Latency-Resilient VLM-Augmented Urban Navigation cites this paper.

Slow Brain, Fast Planner: Latency-Resilient VLM-Augmented Urban Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:09:33.984552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T17:16:52.173943Z digest=sha256:9ede71fc155b3ff345c815f6d797950f7421a24100b1c09de3931598ef99b1fc

Observation 091a2de6-e56a-4a62-a619-e2b06ba794e9 · inbound

Vesta: A Generalist Embodied Reasoning Model cites this paper.

Vesta: A Generalist Embodied Reasoning Model NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:29:35.757784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T16:55:12.518255Z digest=sha256:e4d32b29ee9e89a254dc2e42bb751d79e25364d007109e684bc33f1b0256bf7a

Observation 38150bff-1b9d-4291-993b-8050aec8de85 · inbound

BIT-Nav: Brain-Inspired Trajectory Memory for Embodied Navigation cites this paper.

BIT-Nav: Brain-Inspired Trajectory Memory for Embodied Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:09:37.370731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T13:52:31.453516Z digest=sha256:9bb80dba4f22c1bf6299c407a927ee9e679cac36f668a8cab33b975f204656ef

Observation 6f8c30ea-56fe-49c7-8600-cc3687652cac · inbound

FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation cites this paper.

FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:09:43.512947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T10:24:56.607921Z digest=sha256:9d6547ae25a45f44172747cfb0370d966a34473a5ce569fa61882697ad37e34c

Observation c560a56e-eacc-4fe8-a460-cefc703ba6df · inbound

SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks cites this paper.

SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:33:54.613197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T04:42:23.040915Z digest=sha256:b4770d503f94e03873241188e4c9f53455a0bba846a129a8652967b6d3f4be9e

Observation 613b6a72-ca78-421f-b77f-9dec93d39f73 · inbound

Vision-Language Models for Deployable Social Robot Navigation: Bridging Semantic Reasoning and Low-Level Control cites this paper.

Vision-Language Models for Deployable Social Robot Navigation: Bridging Semantic Reasoning and Low-Level Control NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:54:34.519791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T09:52:51.549135Z digest=sha256:db1a93c0da794c4f639296c74ab2872e052f81fb9cfd7788d0ef0f1e794ba91e

Observation 73fef3cb-f37c-4a72-9c36-98d11587368b · inbound

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation cites this paper.

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:54:49.928777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T05:05:56.065261Z digest=sha256:c8b404575e290344953a0d2797b1495bc23f019893ab39e35509a19b27ecf730

Observation c132e310-40d5-4fa2-92b5-e7a178b3f5ab · inbound

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation cites this paper.

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:08:21.406594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-03T14:04:46.400336Z digest=sha256:87914c012e89c1406fabcf93aed9c1adadaa400998f7e3b60b8bfa01b89818d9

Observation a3f816f0-a771-4f4e-9e6b-66e40e2249b1 · inbound

Exp2VLA: Enabling Vision-Language-Action for Drone Navigation from Expert Demonstrations cites this paper.

Exp2VLA: Enabling Vision-Language-Action for Drone Navigation from Expert Demonstrations NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T04:33:48.376442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:33:48.376442Z digest=sha256:eeb39000e588bf11b01a68f7ca2fb38772bcfa9de73e43550253941fc22fbcca

Observation 60e3694c-c7e2-4522-8943-3ba9f774cccf · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:5c63e7414a13a4bebfa6dcfc6c000711f78ff5702586b437406c7910967161a2

Observation b6e50527-e702-4417-aa23-4552559563ad · inbound

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies cites this paper.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:d873853c293d8709dc7cc19391fcf237ed7d99767758a1c83a530b6dd9fb1b17

Observation f4729f4d-613e-4a48-bbbf-f24dfdff1b23 · inbound

Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation cites this paper.

Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-01T16:20:01.293615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:20:01.293615Z digest=sha256:4fbf66519d6ac21941e4dc6fcd98dafc01ef83ecca0123a5b117b60effe82bc2

Observation 4dcd7d93-6b4e-406e-a652-5bedd565ddcb · inbound

EA-Nav: Learning Safe Visual Navigation Policies with Embodiment Awareness cites this paper.

EA-Nav: Learning Safe Visual Navigation Policies with Embodiment Awareness NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:30:22.146055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:30:22.146055Z digest=sha256:d5f8af6df376afda817a46fc6b1bc7aa9ec93e241081ac60559c2bb3f69a6716

Observation 7e252989-0e19-47ea-a185-1296bbc615d6 · inbound

ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset cites this paper.

ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 292

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:39.791761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:39.791761Z digest=sha256:47017ee9674658f17a8988ac11c8fa6dc600ed76da6c4c2f29b989cd5789a73c

Observation 79160368-0112-47d0-b764-c75123e98fe3 · inbound

Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents cites this paper.

Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T06:06:22.870692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:06:22.870692Z digest=sha256:c0bfaca452d2ea8019daea0db9761243adfc2518859f6486379920ad8c54e2d1

Observation 51a0af7a-2ead-4729-b06b-9357ef1cda00 · inbound

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation cites this paper.

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-30T20:42:03.068111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:42:03.068111Z digest=sha256:cb7c98948645ffd094803ef5e6d1639db3b85f594202a0207b69701396e2d805

Observation fda33824-fdbd-4099-8d6b-0184621082ef · inbound

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation cites this paper.

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T00:43:30.833308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:43:30.833308Z digest=sha256:9db94ee4b02406397b64787eca222258dcec286c961df0e986362830dda0cbfe

Observation b0ffe209-0598-4085-a7b1-ca7ea417fa2f · inbound

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation cites this paper.

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T03:24:27.168988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:24:27.168988Z digest=sha256:a5089fe09460b16033c6bf926ec07c7dacf0e0480155e8e34520851dc01686a2

Observation 73f86b6f-c688-413c-996f-c89124c25906 · inbound

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN cites this paper.

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T11:29:20.418365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:29:20.418365Z digest=sha256:a104bdbfd6207b2377740de6e8675614ef28124e41de455d7250d797695997dd

Observation 11c936a4-78c7-458d-821f-0029b6b698df · inbound

Goal-oriented Navigation Instruction Generation with Tour Video Priors cites this paper.

Goal-oriented Navigation Instruction Generation with Tour Video Priors NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:35.040088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:34:35.040088Z digest=sha256:d50082aa8ddf08cc8fdb51154cdb2669ce85e13868c46b58332eca07f5b155f6

Observation a860a33c-fd01-4431-890e-a63af14b9b69 · inbound

HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments cites this paper.

HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:41.907818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:41.907818Z digest=sha256:bb75e33ebc13e30917fcd7ff82abe4cba6b64483d4b7708fabf95e5984dece5a