Pith. sign in

Paper Citation Record · LEDGER

ABot-N1: Toward a General Visual Language Navigation Foundation Model

As of 20 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 2 inbound Pith citation observations for arXiv:2607.10383.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.10383 v3

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:19:46.694802Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T03:24:25.215191Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved80
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation adb8d8a6-5568-4bf9-b57d-d7983547f11e · outbound

This paper cites 1st Place Solutions for RxR-Habitat Vision-and-Language Navigation Competition (CVPR 2022).

ABot-N1: Toward a General Visual Language Navigation Foundation Model 1st Place Solutions for RxR-Habitat Vision-and-Language Navigation Competition (CVPR 2022)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:36.808445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:36.808445Z digest=sha256:c04db9ed5ed2248897c59e4b00e452de0babf831a4c15ea9d7e101559a45a7af

Observation b6304653-9a51-4724-b28a-20db22fe21ce · outbound

This paper cites Etpnav: Evolving topological planning for vision-language navigation in continuous environments.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Etpnav: Evolving topological planning for vision-language navigation in continuous environments.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:36.880741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:36.880741Z digest=sha256:10706e0d97e9cf683de745aeec09a497b5227d7b723d6d9e338a19d07a031146

Observation ec2a35d4-3b96-4de9-9bfe-a61f76517456 · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:36.938706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:36.938706Z digest=sha256:1b0e6671a935265fabc68e69a067686ffc24565eae04d82217766bc99197bbe9

Observation 53affc0d-cfba-4b9c-9336-f00b74a31862 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:36.990464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:36.990464Z digest=sha256:002d20d81f154693853b356d75ca6989abb72d2bba1a134b8461d2d8f5e9891f

Observation 0f34bf93-085e-4bd1-863e-09d364244fbe · outbound

This paper cites ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects.

ABot-N1: Toward a General Visual Language Navigation Foundation Model ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.070208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.070208Z digest=sha256:64876b95bb5f472dfa607132e1fbd1ebb5fed65d1da1d80a7b1b20846a05a0b3

Observation 826ef829-ab27-4aef-b0e6-544bf1774399 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.117575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.117575Z digest=sha256:cd2774bee480c4e2e5c9221f80f343e6efce5a32339c25ddc95623e66ee413b6

Observation 122a4698-8e1a-4fce-917b-11a3a3fad552 · outbound

This paper cites Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.181342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.181342Z digest=sha256:8a01ee0804c04a8ca0bf1801d0a239546df9ef39c88ef8eb9ef4fe2d42ef1a88

Observation 86df4f93-2068-4319-ac8f-a1a774594377 · outbound

This paper cites Affordances-oriented planning using foundation models for continuous vision-language navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Affordances-oriented planning using foundation models for continuous vision-language navigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.263414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.263414Z digest=sha256:e13a81da1c71a399a8d4519f853e95ce024f6ba28bc34b4badbd0e20d721a5f4

Observation 5125dade-d346-4bec-bbc3-7e6a973086a1 · outbound

This paper cites AstraNav-World: World Model for Foresight Control and Consistency.

ABot-N1: Toward a General Visual Language Navigation Foundation Model AstraNav-World: World Model for Foresight Control and Consistency

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.335878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.335878Z digest=sha256:27c126cb370a3774cd6f0f3d018ad3a1050ac335164df141eb1ad0edc95e1d08

Observation c1e8090a-5e9b-4c25-bf3f-76865e3d3d62 · outbound

This paper cites Topological planning with transformers for vision-and-language navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Topological planning with transformers for vision-and-language navigation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.400923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.400923Z digest=sha256:996eb80c16f9b3954ac9a00c9d47e5e583810db86c352357263568d805c9251f

Observation ebb3fbdb-5ac6-4751-a227-0b18a806e801 · outbound

This paper cites Weakly- supervised multi-granularity map learning for vision-and-language navigation.Advances in Neural Information Processing Systems, 35:38149–38161, 2022.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Weakly- supervised multi-granularity map learning for vision-and-language navigation.Advances in Neural Information Processing Systems, 35:38149–38161, 2022

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.487932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.487932Z digest=sha256:afc39389a07b342bb9b3ff0ad0545c9a3ba2647ccac5b70f2611df8494aaa762

Observation 9525adab-4ff2-4b9a-9be7-4ae88c9da44b · outbound

This paper cites Explore like humans: Autonomous exploration with online sg-memo construction for embodied agents,.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Explore like humans: Autonomous exploration with online sg-memo construction for embodied agents,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.589776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.589776Z digest=sha256:c6ff6a6a7b2440b0ed078b6d2fd60f5d58eca1886ae2ba53084f3ae1cbf716cb

Observation ad62a71f-2a0d-447f-b742-d70398530c9a · outbound

This paper cites Socialnav: Training human-inspired foundation model for socially-aware embodied navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Socialnav: Training human-inspired foundation model for socially-aware embodied navigation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.754962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.754962Z digest=sha256:8a0492dbc8187dacc0c8fce08d5e477a13281d938355264ac372c64b047aa2f7

Observation d4cd4a0c-1303-4220-bb7c-e8301ad32e1f · outbound

This paper cites Socialnav: Training human-inspired foundation model for socially-aware embodied navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Socialnav: Training human-inspired foundation model for socially-aware embodied navigation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.827587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.827587Z digest=sha256:644380cb48e01ebd6b76ec25028d88202302c97ed115f019f4f57113dc7b60b3

Observation 111dc39e-bfdb-495e-839a-0f743d99968b · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision-language models.Advances in Neural Information Processing Systems, 37:135062–135093, 2024.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Spatialrgpt: Grounded spatial reasoning in vision-language models.Advances in Neural Information Processing Systems, 37:135062–135093, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.918214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.918214Z digest=sha256:4cdc15b455ef51358ff7d4794484ed8e1b2ca6febe5d4ec4d692f6246310c51c

Observation 2816f1cb-6124-4612-887f-33fc2ba2146c · outbound

This paper cites Navila: Legged robot vision-language-action model for navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Navila: Legged robot vision-language-action model for navigation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.960793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.960793Z digest=sha256:64d58a6105d76420f16c843b4adec0c68313d7433d03476a53c6e8a7017731da

Observation b23e3ce2-1839-463e-808a-8e25af104d0d · outbound

This paper cites Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.021123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.021123Z digest=sha256:5e92c0f4ddd3e1e42e174e1cf68cb217aa35fe6dc2689a4f3102fb77d26ba125

Observation 4601b2dd-ff75-40ca-bb6f-ae604f0de3c0 · outbound

This paper cites Abot-n0: Technical report on the vla foundation model for versatile embodied navigation, 2026.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Abot-n0: Technical report on the vla foundation model for versatile embodied navigation, 2026

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.081301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.081301Z digest=sha256:e703ca41f5782a89e040ee79f4e742b7c1c337eaa95d9a286b359374940837f9

Observation b3dd50d3-e7b1-43d0-88b9-d4473b4f9f1d · outbound

This paper cites Vpn: Visual prompt navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Vpn: Visual prompt navigation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.136227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.136227Z digest=sha256:1a3fcde99325472458acd1e2ac20f7489a749eb90e3898e68961a9b0f80eb837

Observation 891b3320-bfeb-4dff-ad9b-346ee9c97ece · outbound

This paper cites Helix: A vision-language-action model for generalist humanoid control.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Helix: A vision-language-action model for generalist humanoid control

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.216170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.216170Z digest=sha256:86fff343bc514edfd0e4663dd1f21cb070d577e3f47975126e90fd002a143248

Observation 3bbe455c-b2ac-4aee-81e2-8f0e24068f23 · outbound

This paper cites POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.321058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.321058Z digest=sha256:33256540a588d38e0b2bc437d81ee6a94ca6fcd6e9082ed2858d793eab23f447

Observation 84dc7f60-b8c4-40fa-9c95-7575481f4859 · outbound

This paper cites Vision-and-language navigation: A survey of tasks, methods, and future directions.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Vision-and-language navigation: A survey of tasks, methods, and future directions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.370727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.370727Z digest=sha256:ec2447d87cd121e2312b59a3a6538364ef557175b15447230ab18122a8a922ca

Observation 83394366-1244-42e5-bc9a-e593ebae063d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ABot-N1: Toward a General Visual Language Navigation Foundation Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.433524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.433524Z digest=sha256:4e49b4ecdee01aed6a13f082fb6cc10c0766c67c84d6469bfdd9399926bcd08a

Observation faec5ddf-a971-4758-bd72-5d3f25e535ab · outbound

This paper cites A novel vision-based tracking algorithm for a human-following mobile robot.IEEE Transactions on Systems, Man, and Cybernetics: Systems, 47(7):1415–1427, 2016.

ABot-N1: Toward a General Visual Language Navigation Foundation Model A novel vision-based tracking algorithm for a human-following mobile robot.IEEE Transactions on Systems, Man, and Cybernetics: Systems, 47(7):1415–1427, 2016

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.441120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.441120Z digest=sha256:0a53d91c5e6574b71dc0f1d22d93276f9eec814873ccec0de0891361e2a407db

Observation 1bfa2779-a80d-4126-a669-60a79fa00688 · outbound

This paper cites Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Bridging the gap between learning in discrete and continuous environments for vision-and-language navigation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.449475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.449475Z digest=sha256:46348ccb90d88582ea8bbbce96d91d261f5148449aa3d3f0e3cf1c234b66b490

Observation 74343176-14e1-443d-bf2a-69a4e3d1a491 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

ABot-N1: Toward a General Visual Language Navigation Foundation Model $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.498013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.498013Z digest=sha256:ae293ebb0c2aa8f057b79ea33288f2c69fe80917017a1c95ddcbfc30c124865e

Observation eb18c568-df6a-4dcd-b467-c79204643797 · outbound

This paper cites A comprehensive review of recent advancements in vision-and-language navigation.Discover Computing, 29(1):167, 2026.

ABot-N1: Toward a General Visual Language Navigation Foundation Model A comprehensive review of recent advancements in vision-and-language navigation.Discover Computing, 29(1):167, 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.592299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.592299Z digest=sha256:336ce9ba71b6dcf535aa00fb58a0715cc25af61826d1897b0f6bab1068757c9e

Observation a10459ff-3da6-4574-b3a3-2d57746b6d7e · outbound

This paper cites Sim-2-sim transfer for vision-and-language navigation in continuous environments.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Sim-2-sim transfer for vision-and-language navigation in continuous environments

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.749263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.749263Z digest=sha256:71dc12afe9ba86954642b190cbf7c4bfe51a9ad77fcb86acf2d009b24141a0a2

Observation 92565456-8fbb-4f8b-ba79-4abbee239075 · outbound

This paper cites Beyond the nav-graph: Vision- and-language navigation in continuous environments.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Beyond the nav-graph: Vision- and-language navigation in continuous environments

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:38.893334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:38.893334Z digest=sha256:e8d987624164ba7aea872ab1dcc2d256f650886eeea04c82d1bed08a7f66f64a

Observation 2e19769c-7526-40be-986e-9b7f4eaa3641 · outbound

This paper cites Waypoint models for instruction-guided navigation in continuous environments.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Waypoint models for instruction-guided navigation in continuous environments

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.022007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.022007Z digest=sha256:0bf27fc495201b9ac74d42b722e90f1028e061c1968f804d4fde13ebdb39d603

Observation 302cfabe-c7cb-4ee1-a8da-240a88e2ad90 · outbound

This paper cites Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.190648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.190648Z digest=sha256:62eac19554dafc385d9faddeba6f95541aa9d427fe76a5e4cac2a245d194dcc9

Observation 394ab79b-b505-4b3b-8678-4aa890450fdc · outbound

This paper cites Large-scale model-enhanced vision-language navigation: Recent advances, practical applications, and future challenges.Sensors, 26(7):2022, 2026.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Large-scale model-enhanced vision-language navigation: Recent advances, practical applications, and future challenges.Sensors, 26(7):2022, 2026

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.270648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.270648Z digest=sha256:78ae2a658059d206406698d4e2645e61d47e7c0b33b9f3f3169f4b54a3768f0f

Observation ad0bed1f-a59d-45a7-ae5d-5c9f82e49d7e · outbound

This paper cites Navcot: Boosting llm-based vision-and-language navigation via learning disentangled reasoning.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Navcot: Boosting llm-based vision-and-language navigation via learning disentangled reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.467980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.467980Z digest=sha256:7cfe9688ebce8594c7f5c87adf2a001c19abcd747efda3b86e9c000b80c73d0f

Observation 156c63b7-23f7-4127-b6c8-34bb149e5a5d · outbound

This paper cites Conflict-averse gradient descent for multi-task learning.Advances in neural information processing systems, 34:18878–18890, 2021.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Conflict-averse gradient descent for multi-task learning.Advances in neural information processing systems, 34:18878–18890, 2021

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.574382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.574382Z digest=sha256:81c2b593d406d0760fa9bd608eb8e0594ab3c052d675cf60b6645e799697bf26

Observation 50477096-bb03-49b9-8cb8-4eaab3decd3a · outbound

This paper cites Navforesee: A unified vision-language world model for hierarchical planning and dual-horizon navigation prediction.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Navforesee: A unified vision-language world model for hierarchical planning and dual-horizon navigation prediction

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.691873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.691873Z digest=sha256:036304dda90a6670d112e45dafcee22549aeb467cc516139f3ceb5fc05f71e7d

Observation 90a00645-7917-432c-8dd8-0bb8d514a38d · outbound

This paper cites Navforesee: A unified vision-language world model for hierarchical planning and dual-horizon navigation prediction.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Navforesee: A unified vision-language world model for hierarchical planning and dual-horizon navigation prediction

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.799646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.799646Z digest=sha256:ebce7614fb810d36e18fcf0d7c154f1a397dca8bf5bbf2e7cd50ae6aae4755fc

Observation ad2bda31-0b80-4a07-b4ce-0494b10127f7 · outbound

This paper cites Trackvla++: Unleashing reasoning and memory capabilities in vla models for embodied visual tracking.arXiv preprint arXiv:2510.07134, 2025.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Trackvla++: Unleashing reasoning and memory capabilities in vla models for embodied visual tracking.arXiv preprint arXiv:2510.07134, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.883480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.883480Z digest=sha256:2676b0989c04bc94b8cdbf787639f6d44809ed779b0026188942d6279607f05a

Observation 2bd60e77-0852-4b8a-8f26-986ca9ea610c · outbound

This paper cites Citywalker: Learning embodied urban navigation from web-scale videos.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Citywalker: Learning embodied urban navigation from web-scale videos

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:39.946903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:39.946903Z digest=sha256:1c4e6fed6f5874982521c916d2ceaa1a38ae58d1d5360b5bf7de9ccc117eab59

Observation 5b425577-fa8e-4f88-904d-a268d96ded91 · outbound

This paper cites InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment.

ABot-N1: Toward a General Visual Language Navigation Foundation Model InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:40.032137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:40.032137Z digest=sha256:a215ba15f0cdf7cb801db66a58d5cb4a1284d4733cf53dd822a81db5c213f121

Observation f1c82efe-76e2-4865-85bc-b0e94dae1bf4 · outbound

This paper cites Pivot: Iterative visual prompting elicits actionable knowledge for vlms.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Pivot: Iterative visual prompting elicits actionable knowledge for vlms

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:40.091709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:40.091709Z digest=sha256:71ad78c2a5a804df8df1b0f4441c579885c35cd0742baba66acd2bc44b2c9e64

Observation 187b3f55-7677-4926-9bf6-ff4120c9f179 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

ABot-N1: Toward a General Visual Language Navigation Foundation Model GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:40.319034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:40.319034Z digest=sha256:a5c4d8eb412e3cd7fa2353ddd0b97a1e8fca0ac0a2e225ee7fa7ca8102a3cc4d

Observation 1e0dc0c5-6860-45b1-8596-bf19425790df · outbound

This paper cites OpenFrontier: General Navigation with Visual-Language Grounded Frontiers.

ABot-N1: Toward a General Visual Language Navigation Foundation Model OpenFrontier: General Navigation with Visual-Language Grounded Frontiers

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:40.471855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:40.471855Z digest=sha256:6f2b4f0a0d71d264c0d7a0f85961b64e81ef97a5738d2b37d2329d27b36854d8

Observation d087e6c2-8e6a-45eb-8ee1-879b33f5ea76 · outbound

This paper cites Habitat: A platform for embodied ai research.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Habitat: A platform for embodied ai research

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:40.601537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:40.601537Z digest=sha256:21fc34e0c45ec5a7c3e166eb23295ea998fa06171a75e8021dc44e0bca0ceb50

Observation 59ddb180-bbda-45af-b020-75ea7f9aff43 · outbound

This paper cites Gnm: A general navigation model to drive any robot.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Gnm: A general navigation model to drive any robot

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:40.794162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:40.794162Z digest=sha256:a3772dc5c6c40b02a97305c15478ec39387f5ca08951f9111d01ebd7a766601c

Observation 597f69b9-bd76-43ca-b68d-50b2c5ad3eac · outbound

This paper cites Vint: A foundation model for visual navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Vint: A foundation model for visual navigation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.001578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.001578Z digest=sha256:05f02556e2e02de58ef9174f7507f471ada5fb546fa5dc83624edafa95d15da2

Observation 7cadf22a-03b2-4143-8bf3-e1a1e5b5c8d7 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.156501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.156501Z digest=sha256:f5ad3a27e1eecce5484431dbe6128792602ce8257e9e03b35eb8a62aa1c8729e

Observation 5efd68fd-27c8-462e-ab83-5827f5d3a342 · outbound

This paper cites Hume: Introducing System-2 Thinking in Visual-Language-Action Model.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.336447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.336447Z digest=sha256:00ca2b02ba25fc0cd5249dd62868cb8f261cd508a7ca224ff0320486fc2994ad

Observation 91b3af87-24f6-42c2-9a9f-42ec0f4198b5 · outbound

This paper cites Nomad: Goal masked diffusion policies for navigation and exploration.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Nomad: Goal masked diffusion policies for navigation and exploration

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.436082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.436082Z digest=sha256:d8c09a2e7433728eabd50a14f73556a796f2c528a32205f9dead964d19d087e8

Observation 8a3f2dc4-a7f6-4331-a51e-a42b46d0a661 · outbound

This paper cites Emma-x: An embodied multimodal action model with grounded chain of thought and look-ahead spatial reasoning.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Emma-x: An embodied multimodal action model with grounded chain of thought and look-ahead spatial reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.559440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.559440Z digest=sha256:68780975e3b9f9a89acb9c072c0144af3855d6e425d9148c3f61d6089cf5332d

Observation 0ddff957-a98c-47dc-a4ff-74feac8f942d · outbound

This paper cites Robobrain 2.0 technical report.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Robobrain 2.0 technical report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.747678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.747678Z digest=sha256:c3327336982198c3c326ee9f8792d3d2217be4bfc0cf3a08302a39e5e6fb7d12

Observation 146d8fdc-2eca-462f-98c9-4d0503f45d03 · outbound

This paper cites Robobrain 2.5: Depth in sight, time in mind.arXiv preprint arXiv:2601.14352, 2026.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Robobrain 2.5: Depth in sight, time in mind.arXiv preprint arXiv:2601.14352, 2026

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.949051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.949051Z digest=sha256:7b5ec89dc81ef333bc17dd2f63dc79acaf9648f9d95477912328b8c7e547aaaf

Observation c17b63ee-6b9e-4f45-9987-cd14d64924f9 · outbound

This paper cites Drivevlm: The convergence of autonomous driving and large vision-language models.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Drivevlm: The convergence of autonomous driving and large vision-language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:42.085991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:42.085991Z digest=sha256:5fbab6e3b2bfc76fcd1623881cafeeba6f538abcfa709843d35c8a61e64dc9a6

Observation 66d353e3-7077-4572-b75d-540e94739956 · outbound

This paper cites Dreamwalker: Mental planning for continuous vision-language navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Dreamwalker: Mental planning for continuous vision-language navigation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:42.224826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:42.224826Z digest=sha256:c54e9f6a148053550897f613aba3daa3bcd0f4cce3b302a136c7fdf728b82405

Observation 441e92bc-1557-4a7d-a989-e32f160bb4ef · outbound

This paper cites Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:42.400030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:42.400030Z digest=sha256:9526fcb8c97bf4cd6e29209cdec9afda9511b9b1c71eebde21da4e923b2938d4

Observation 7066ddf5-e0c3-4f18-b96d-5b2649a436d7 · outbound

This paper cites Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:42.537690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:42.537690Z digest=sha256:718354b89bce1a4095f5bd04df64cc00ab3baa7d601c4ed504570394bdfe1711

Observation d2a5e682-9c0f-4b86-b398-958e6ccf2a58 · outbound

This paper cites TrackVLA: Embodied Visual Tracking in the Wild.

ABot-N1: Toward a General Visual Language Navigation Foundation Model TrackVLA: Embodied Visual Tracking in the Wild

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:42.675991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:42.675991Z digest=sha256:ba8643c596cfa913f019bc54c344607a837897e52c9f1e3c180d58ca69e0d4ac

Observation 1d498e91-58a5-4e3f-abee-5514a3f38ab6 · outbound

This paper cites Gridmm: Grid memory map for vision-and-language navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Gridmm: Grid memory map for vision-and-language navigation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:42.887647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:42.887647Z digest=sha256:987e5f9ec66b1fbc41bbe04037252f685e26818e64b0544d52c5e50af2711738

Observation f5bd57f7-9021-4d13-9c35-84914e1681f7 · outbound

This paper cites Lookahead exploration with neural radiance representation for continuous vision-language navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Lookahead exploration with neural radiance representation for continuous vision-language navigation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:43.042916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:43.042916Z digest=sha256:ed7c94dcea3ae0088c0dcc84772fe4ad37c6f13c1b0d5410251c3924f88cd382

Observation 343b73d3-f947-49f5-af83-c92c35158c7b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:43.223398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:43.223398Z digest=sha256:415763d50b3205f91459034a0a47e38ea6e08f490b55c0f61b118bb6d80ed517

Observation 13e2aa2c-7a73-4bdb-8fe8-13be17d0adbe · outbound

This paper cites Ground slow, move fast: A dual-system foundation model for generalizable vision-and-language navigation.arXiv preprint arXiv:2512.08186, 2025.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Ground slow, move fast: A dual-system foundation model for generalizable vision-and-language navigation.arXiv preprint arXiv:2512.08186, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:43.356041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:43.356041Z digest=sha256:0169cc8c4dee847036de142cf856ece906496ad886075522fc8180149dc264fb

Observation 35ace54b-82f0-409c-bcfd-8e99ef9ab06b · outbound

This paper cites StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling.

ABot-N1: Toward a General Visual Language Navigation Foundation Model StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:43.503753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:43.503753Z digest=sha256:cf14e1b64180a2d4088c48b8d9cda790ee0cbd3785005ded3cb07cb67731f845

Observation 4bc63c77-f5ad-48e3-b10e-f20deda7c58c · outbound

This paper cites Nav-r2 dual-relation reasoning for generalizable open-vocabulary object-goal navigation, 2025.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Nav-r2 dual-relation reasoning for generalizable open-vocabulary object-goal navigation, 2025

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:43.638832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:43.638832Z digest=sha256:ac2d79d35741913e3df1ed65bfe9aa0993c95a0c0b4652088ec893a57e7cb6b3

Observation f240ba48-7982-4785-a6eb-3933a3dfd00e · outbound

This paper cites Omninav: A unified framework for prospective exploration and visual-language navigation,.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Omninav: A unified framework for prospective exploration and visual-language navigation,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:43.887148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:43.887148Z digest=sha256:1138c5f6cd205ce53e56c48e7f3a0a23d70602e0a5da07b825005dcb7eb18d22

Observation b3c11b26-4533-4378-870a-ed559ff17874 · outbound

This paper cites an unresolved cited work.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:44.163112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:44.163112Z digest=sha256:5ec2ddd981807856d117ac255105e610b075668e32f3279e40f01a30d4bdd5e9

Observation f87437b6-e352-4f3e-88cb-186f17837e35 · outbound

This paper cites AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:44.848057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:44.848057Z digest=sha256:14cc90e0c760fbabbb94fa6c03ee9ee35cbeb6f522b440324caf30a30aee9e88

Observation fb7e05a8-f520-459b-ae16-e2e899d97805 · outbound

This paper cites Ce-nav: Flow-guided reinforcement refinement for cross-embodiment local navigation.arXiv preprint arXiv:2509.23203, 2025.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Ce-nav: Flow-guided reinforcement refinement for cross-embodiment local navigation.arXiv preprint arXiv:2509.23203, 2025

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:44.595755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:44.595755Z digest=sha256:e35aa753168119b413f8badc9b2de0849a58e2bc57d919ace06c7c2e52908b01

Observation c5740e3e-0e1e-4ceb-84d7-b7bf2ee97dee · outbound

This paper cites Gradient surgery for multi-task learning.Advances in neural information processing systems, 33:5824–5836, 2020.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Gradient surgery for multi-task learning.Advances in neural information processing systems, 33:5824–5836, 2020

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.185975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.185975Z digest=sha256:66a63a7fbd62dbdd23eeb4f9436f324d6af4c1022683d1ea586cf7372e2bf06f

Observation dd6cb1f0-d359-4913-aeb3-9fb9a95af901 · outbound

This paper cites Hm3d-ovon: A dataset and benchmark for open-vocabulary object goal navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Hm3d-ovon: A dataset and benchmark for open-vocabulary object goal navigation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.033252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.033252Z digest=sha256:36ade389724d219bd72c04c3605ca0516c9d084854a23a04da92b496c415ae5a

Observation 0032cf60-d307-4307-a2a7-0583942824a2 · outbound

This paper cites Robotic control via embodied chain-of-thought reasoning.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Robotic control via embodied chain-of-thought reasoning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.353049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.353049Z digest=sha256:27e82b0c3d0bd7098937cbb504856c229ea4ac3a80a4fb97566dfb955acd7048

Observation c1ffa229-320e-430d-820c-a04ef96401db · outbound

This paper cites Correctnav: Self-correction flywheel empowers vision-language-action navigation model.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Correctnav: Self-correction flywheel empowers vision-language-action navigation model

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.272916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.272916Z digest=sha256:4e0adda90940306aea5a46db603f88a7eabfa85754162bc0f8124f6cebfe5ecb

Observation bfdbe320-9b50-442c-89dc-f11389031e3d · outbound

This paper cites Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.554365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.554365Z digest=sha256:fbae5ab5e7634cc00da0d11cff56e7073383115b46f6e506021468925a8b1ec5

Observation f45fc376-bbdd-4e1b-bf69-de16055f0d5f · outbound

This paper cites PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators.

ABot-N1: Toward a General Visual Language Navigation Foundation Model PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.435794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.435794Z digest=sha256:91b53fc5b8e39a3bad1028e1dc1ab311d2f2c308f8b564c1d7fc4439375f311a

Observation 510c865f-b4da-407a-b17f-b9ddbceb467d · outbound

This paper cites Embodied navigation foundation model, 2025.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Embodied navigation foundation model, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.893363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.893363Z digest=sha256:b8254f5d0296d0fec6cfc5ca110e24a2a759e790be42e333ae5b264be46fd3bb

Observation 84b45b46-bdc3-4a0e-bfa2-c00f9278b93b · outbound

This paper cites NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.694354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.694354Z digest=sha256:32675913b57cfab2bf94e5af13d389f9fbe89c9ec4930bc2ba679e9863f53965

Observation c4a90c1c-ddee-426a-bd76-fac3191b2434 · outbound

This paper cites Cot-vla: Visual chain-of-thought reasoning for vision-language-action models.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Cot-vla: Visual chain-of-thought reasoning for vision-language-action models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:46.111431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:46.111431Z digest=sha256:04a6a4d65cb21530ffb7ad502f161e9d340ada3b7183dbd201a280ee0a7dca92

Observation 59865f02-98db-450f-8038-67d1c7f55cd0 · outbound

This paper cites Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:45.977650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:45.977650Z digest=sha256:6cfbd6d2a8ec4f956c125388c0b31e267ea39b9566144db819e25492d30ad2d5

Observation 1b2f08d2-583e-427c-bc98-06e325f44c00 · outbound

This paper cites Navgpt-2: Unleashing navigational reasoning capability for large vision-language models.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Navgpt-2: Unleashing navigational reasoning capability for large vision-language models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:46.494248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:46.494248Z digest=sha256:22f3decc9cadb387d746a9e3642acd439c4eb2fbceddce633cd0def817c3d0d9

Observation 3c065a18-e89e-4221-b477-1683adbf7e21 · outbound

This paper cites Empowering embodied visual tracking with visual foundation models and offline rl.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Empowering embodied visual tracking with visual foundation models and offline rl

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:46.276765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:46.276765Z digest=sha256:6a93e955ae6935e05bd77c017a571ee609f203087193f19cb7be4b94cfa684aa

Observation e57645ce-5f1f-40f3-8174-76c56f8fb83f · outbound

This paper cites Fantasyvln: Unified multimodal chain-of-thought reasoning for vision-and-language navigation.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Fantasyvln: Unified multimodal chain-of-thought reasoning for vision-and-language navigation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:46.694802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:46.694802Z digest=sha256:dd1921451f5066eb9b3ae84b36c1532eb7801f8a62473dbeaa71e8c6596ff7b7

Observation d5937e12-c0c7-4f15-b192-086bebb1c23e · outbound

This paper cites Explore Like Humans: Autonomous Exploration with Online SG-Memo Construction for Embodied Agents.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Explore Like Humans: Autonomous Exploration with Online SG-Memo Construction for Embodied Agents

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:37.659711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:37.659711Z digest=sha256:224dbc516b375d0ce0247b0d0b27ce32731e5a2085919a05265c3bc7f2511a15

Pith citing papers

Observation 25261757-7e38-47d7-8624-7e27c07703b7 · inbound

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation cites this paper.

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation ABot-N1: Toward a General Visual Language Navigation Foundation Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T00:43:27.294478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:43:27.294478Z digest=sha256:80ff3d9f0e693509965f16acd4787f6169af86e9f320d852555ba8b39582593a

Observation 3a8e0699-478c-4ebc-ae1a-20becba05d77 · inbound

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation cites this paper.

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation ABot-N1: Toward a General Visual Language Navigation Foundation Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T03:24:25.215191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:24:25.215191Z digest=sha256:67d79373e944188be283f0521a6bdffc5682ccebd5bba41537b279e9d33b3da4