Pith. sign in

Paper Citation Record · LEDGER

ViNT: A Foundation Model for Visual Navigation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 53 inbound Pith citation observations for arXiv:2306.14846.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.14846 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 53 of 53 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:04:10.941000Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:08.552638Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b36a4887-8aca-4a60-b288-7bcdea25da7f · inbound

Learning Interactive Real-World Simulators cites this paper.

Learning Interactive Real-World Simulators ViNT: A Foundation Model for Visual Navigation

Reference 166

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T02:15:18.607033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T02:15:18.265190Z digest=sha256:dca06190dddbf3588f02c851feb7711fe9d3b146c3da772eafffbaa6ce62d66e

Observation 5d2783ab-23b6-40cb-9504-886ea939271e · inbound

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset cites this paper.

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset ViNT: A Foundation Model for Visual Navigation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:51:18.813272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T05:51:18.508352Z digest=sha256:aad20a1f2efb1b9c65f670a3d8a57e6e55f2905c658e364afdf79ef0984407b5

Observation 39e852db-8e93-4270-83bb-5c330482b405 · inbound

Octo: An Open-Source Generalist Robot Policy cites this paper.

Octo: An Open-Source Generalist Robot Policy ViNT: A Foundation Model for Visual Navigation

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:26:15.361975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T00:26:15.163358Z digest=sha256:2f46db5f8c65060c0f6808c41965af5dd61138abd2284c0e82ce16566e137923

Observation e8386bbb-770c-46cf-8bb5-82a311ab90b0 · inbound

OpenVLA: An Open-Source Vision-Language-Action Model cites this paper.

OpenVLA: An Open-Source Vision-Language-Action Model ViNT: A Foundation Model for Visual Navigation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:46:36.428955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:46:35.942338Z digest=sha256:86be4a5c356473775c3a315f65818e6b6f89e32feeeef6a089f51fa23efd284a

Observation 66941f01-0ffd-4c08-a963-5caaf9c79c81 · inbound

Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight cites this paper.

Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight ViNT: A Foundation Model for Visual Navigation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:55:25.130017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T04:53:33.503100Z digest=sha256:46d19814c6c095d87621b88288d6fbdcb5117054a153121058549aa2222bc6f6

Observation f825b08a-225c-4e22-a183-1b4b4ed3de42 · inbound

Predictive Red Teaming: Breaking Policies Without Breaking Robots cites this paper.

Predictive Red Teaming: Breaking Policies Without Breaking Robots ViNT: A Foundation Model for Visual Navigation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:10.941000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:10.941000Z digest=sha256:771951233ffac515004145305ca3e50bf63d15a39256ad1fa1ed0873cea610ad

Observation e8dedf86-1f6a-465f-977c-2584ea469877 · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization ViNT: A Foundation Model for Visual Navigation

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.947021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:b077256f137deafaa311a45a1818b67a8bde61d2cfbfc7362aca2d005a507b51

Observation 45463947-70e1-4228-b576-c606f23ff2f8 · inbound

Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments cites this paper.

Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments ViNT: A Foundation Model for Visual Navigation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:25:04.971498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:25:04.971498Z digest=sha256:9fdfc64f52a87ae1e1950acc48988b39263a6b68131e59195650fa52ac1c8111

Observation a3b9629b-edf7-4985-9bd4-aee7a3b1c63a · inbound

ThermalLoc: A Vision Transformer-Based Approach for Robust Thermal Camera Relocalization in Large-Scale Environments cites this paper.

ThermalLoc: A Vision Transformer-Based Approach for Robust Thermal Camera Relocalization in Large-Scale Environments ViNT: A Foundation Model for Visual Navigation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:39.314742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:39.314742Z digest=sha256:2bb29a393607d94f3f3db4edc4a32e38a7a2dac1a790148132108c4d77395185

Observation d53be860-d51d-44f6-871f-14939fd49404 · inbound

Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface cites this paper.

Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface ViNT: A Foundation Model for Visual Navigation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:12.361685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:12.361685Z digest=sha256:0aae5f48ad1ae46f61796f3cca351f1b30c7e5c9e6d683f228298bbc2aed1ab9

Observation 3f906f59-76aa-43b6-89d6-5a93af40085e · inbound

OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference cites this paper.

OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference ViNT: A Foundation Model for Visual Navigation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:25.667794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:45:25.667794Z digest=sha256:324eadb44387d7e4299e7f811bfbe6338e370d4ff5990251f5ace4408c1f0035

Observation b5c086d3-636b-4df9-96a8-81861f380df6 · inbound

MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding cites this paper.

MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding ViNT: A Foundation Model for Visual Navigation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:16.219197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:16.219197Z digest=sha256:7b9767104afad3c2064da89a07e01872b1a2a50d4d0b8a3c594615e5db35455f

Observation 86a39f3a-db25-45be-881c-7c321934740d · inbound

PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models cites this paper.

PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models ViNT: A Foundation Model for Visual Navigation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T14:59:03.607088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:59:03.607088Z digest=sha256:545f30c19af12749f969ec050f38fd261de301f72f094ba9612b8a240894030e

Observation 0abed561-5e51-4d04-a671-9199c4446750 · inbound

From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning cites this paper.

From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning ViNT: A Foundation Model for Visual Navigation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:11:28.510190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:11:28.510190Z digest=sha256:c464ecded76957a65da970ac572236115d7bc8bc5b16b7ead4b3fd0311c71929

Observation c4176d68-2e40-4bfe-a753-303db470432b · inbound

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation cites this paper.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation ViNT: A Foundation Model for Visual Navigation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:41.942248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:41.942248Z digest=sha256:5e1bb47b3a2da9e5865829982203ad0d0a376be3cd61820c5f4e799939dc7808

Observation 97f304a3-34f0-48be-b5d6-64447c219eee · inbound

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models cites this paper.

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models ViNT: A Foundation Model for Visual Navigation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:27.032727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:27.032727Z digest=sha256:5c6f7c5c0590ed96a9d31e97de3ce963757b8cca1438d925f27e012e4517ea44

Observation 296f0711-4772-4115-b7eb-e037b94a6f41 · inbound

DUViN: Diffusion-Based Underwater Visual Navigation via Knowledge-Transferred Depth Features cites this paper.

DUViN: Diffusion-Based Underwater Visual Navigation via Knowledge-Transferred Depth Features ViNT: A Foundation Model for Visual Navigation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T11:21:05.054748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:21:05.054748Z digest=sha256:f9c4d154a6f23f49d211adf0a69d56924975a21576034ab9bde2232dff4b6959

Observation dec2dce5-537b-40d7-a202-44debc262802 · inbound

MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning cites this paper.

MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning ViNT: A Foundation Model for Visual Navigation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:54.645609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:37:54.645609Z digest=sha256:7cdcfb37a864d05f669a255e60a99d877ff4b0b24b6537d703cb4b6c57116dc9

Observation 47d1f422-b9d8-42dd-a79c-d98119105243 · inbound

MATT-Diff: Multimodal Active Target Tracking by Diffusion Policy cites this paper.

MATT-Diff: Multimodal Active Target Tracking by Diffusion Policy ViNT: A Foundation Model for Visual Navigation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:42:07.427047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T21:41:44.772556Z digest=sha256:49c5a36cfbbd2b97c7a4993ea85b3ff0deda61e0fa9f571c85334be73b4cab2a

Observation c8d781ae-5875-47a1-9c8f-d70f6e7285ea · inbound

Splatblox: Traversability-Aware Gaussian Splatting for Outdoor Robot Navigation cites this paper.

Splatblox: Traversability-Aware Gaussian Splatting for Outdoor Robot Navigation ViNT: A Foundation Model for Visual Navigation

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:04:09.042373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T06:03:45.608023Z digest=sha256:b0aa6fd5fd513a2c221a82eda8bc907fc518bb12681d298b470c931c9a7cc66c

Observation 90e2115f-e72e-4b12-a315-776551a94066 · inbound

CostNav: A Navigation Benchmark for Real-World Economic-Cost Evaluation of Physical AI Agents cites this paper.

CostNav: A Navigation Benchmark for Real-World Economic-Cost Evaluation of Physical AI Agents ViNT: A Foundation Model for Visual Navigation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T20:24:36.804719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:24:36.804719Z digest=sha256:ec3832bd4e6dd3d6672ff881e4ccc946c5a522f10eb6a8e4b43e686ec7022f45

Observation 02abab78-1ada-49a4-9146-9cf470127c0d · inbound

AID: Agent Intent from Diffusion for Multi-Agent Informative Path Planning cites this paper.

AID: Agent Intent from Diffusion for Multi-Agent Informative Path Planning ViNT: A Foundation Model for Visual Navigation

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T02:58:55.122084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T02:55:11.396778Z digest=sha256:fd2e905625b775a224ff4663667bed1c380c37f9f3e3832fd33e537397ed9a47

Observation 6f3df3c4-ee9c-4659-895a-90074b0e39e1 · inbound

Learning to Localize Reference Trajectories in Image-Space for Visual Navigation cites this paper.

Learning to Localize Reference Trajectories in Image-Space for Visual Navigation ViNT: A Foundation Model for Visual Navigation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T21:56:35.817252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:56:35.817252Z digest=sha256:30735f3eff8d6b2967c59e82446a9fc522711770cd927ebbdfb2e7314320f3f4

Observation 4e8c844d-bbf4-41cf-ad0b-672e941a9e7c · inbound

Approximate Imitation Learning for Event-based Quadrotor Flight in Cluttered Environments cites this paper.

Approximate Imitation Learning for Event-based Quadrotor Flight in Cluttered Environments ViNT: A Foundation Model for Visual Navigation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-15T13:10:42.531155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:10:42.531155Z digest=sha256:4fd5f6d48409bfe97359f49e7190c5273e9cef02a1a63aaa35505b4d421c294e

Observation b8a1c452-39d6-4731-8b11-86af9a830586 · inbound

RAE-NWM: Navigation World Model in Dense Visual Representation Space cites this paper.

RAE-NWM: Navigation World Model in Dense Visual Representation Space ViNT: A Foundation Model for Visual Navigation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-15T12:07:05.640150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:07:05.640150Z digest=sha256:c70ee75541aca1eab3841492ca6b34340e54bdd13cea6f8eaade61d3295ca328

Observation 8d205b90-6e6b-4ee9-aa56-d8a6bce10e5d · inbound

STRNet: Visual Navigation with Spatio-Temporal Representation through Dynamic Graph Aggregation cites this paper.

STRNet: Visual Navigation with Spatio-Temporal Representation through Dynamic Graph Aggregation ViNT: A Foundation Model for Visual Navigation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:13:13.493766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T20:11:36.717810Z digest=sha256:08eb76d7c8ca3b7552217d859d9fcccb759e9bd8c845a111b1a72febd61de56b

Observation 89b88992-5782-4582-8245-f78c936a7f4f · inbound

Rectified Schr\"odinger Bridge Matching for Few-Step Visual Navigation cites this paper.

Rectified Schr\"odinger Bridge Matching for Few-Step Visual Navigation ViNT: A Foundation Model for Visual Navigation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:00:48.611135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:25:36.704959Z digest=sha256:69228e734f78673d362e41046b21ecec8ac24f83a9670687a56759c73d25fd83

Observation 9cce0e6f-dd7e-4808-a0ca-ede5a09f008e · inbound

How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace cites this paper.

How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace ViNT: A Foundation Model for Visual Navigation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:25:57.050349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:09:39.510770Z digest=sha256:177a826753956935bbb5a2e967856c35ec3185a7caafccd40ce30dcf97cc554b

Observation 061c57a9-588c-454f-a3ff-076cadee399a · inbound

Human Cognition in Machines: A Unified Perspective of World Models cites this paper.

Human Cognition in Machines: A Unified Perspective of World Models ViNT: A Foundation Model for Visual Navigation

Reference 147

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:12:26.201103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:12:15.663761Z digest=sha256:7d01c98d4c509248e029fb57752d68009cde5c844f13857c7b853a8b2e16419c

Observation 6ed31ce3-54cc-4393-9118-efadcf7054ab · inbound

NavOL: Navigation Policy with Online Imitation Learning cites this paper.

NavOL: Navigation Policy with Online Imitation Learning ViNT: A Foundation Model for Visual Navigation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:21.907111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:51:53.848800Z digest=sha256:21ba8f388d767c184f999c786fb601f56464179533282bf49eb3eeb85db7d557

Observation b7040a10-19db-4a64-9b42-018e365017f4 · inbound

NavRL++: A System-Level Framework for Improving Sim-to-Real Transfer in Reinforcement Learning-Based Robot Navigation cites this paper.

NavRL++: A System-Level Framework for Improving Sim-to-Real Transfer in Reinforcement Learning-Based Robot Navigation ViNT: A Foundation Model for Visual Navigation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:18:54.542544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T19:16:06.363772Z digest=sha256:4f97d38b2d2bff3a9b93692ae0cf963995ebd1bc349db19c985b26f66a6e2e4b

Observation f1f397ed-7472-448b-a023-6c533351503d · inbound

Improved Baselines with Representation Autoencoders cites this paper.

Improved Baselines with Representation Autoencoders ViNT: A Foundation Model for Visual Navigation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:43:15.342413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:40:14.358108Z digest=sha256:c3429ba641fd693e9bfc725a4d2d76444e8ad7308de245be39e3f1c2142d3153

Observation c3d78722-6edd-4e6f-a662-480de3c89b15 · inbound

Autonomous Frontier-Based Exploration with VLM Guidance cites this paper.

Autonomous Frontier-Based Exploration with VLM Guidance ViNT: A Foundation Model for Visual Navigation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:24.310430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T04:37:35.682778Z digest=sha256:569a97ed7ebd001aae0b6a44002cc3380bab70ddef8687820c0245c4022d375f

Observation 605043c3-8d2a-4628-8f39-577c7ef2a6e2 · inbound

World Models as Group Actions cites this paper.

World Models as Group Actions ViNT: A Foundation Model for Visual Navigation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.360069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T13:17:42.720825Z digest=sha256:b718a7a35c41e325dbe5ccd97cc323812e76ae50705befb85e31af7818599935

Observation 95e3068d-ae8a-4563-b635-90a93fa1ad63 · inbound

Drift-Resistant Navigation World Model with Anchored Epipolar Guidance cites this paper.

Drift-Resistant Navigation World Model with Anchored Epipolar Guidance ViNT: A Foundation Model for Visual Navigation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:14:40.998750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T13:07:20.197513Z digest=sha256:96e25ab32cac3cb32a562105a0cbd8a0baf49913a357e795f8bd6d3afc21ff51

Observation 95c14893-4724-4966-a61c-bbadeaceccd5 · inbound

Sentinel: Embodied Cooperative Spatial Reasoning and Planning cites this paper.

Sentinel: Embodied Cooperative Spatial Reasoning and Planning ViNT: A Foundation Model for Visual Navigation

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:04:01.451884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:58:41.050531Z digest=sha256:1cb1df10e280107b4507b7c246553155d15e0a5fe23e6dcc6919a4412a3b7db5

Observation 77733835-f772-40fe-b4c0-247fc1edec81 · inbound

Look Further: Socially-Compliant Navigation System in Residential Buildings cites this paper.

Look Further: Socially-Compliant Navigation System in Residential Buildings ViNT: A Foundation Model for Visual Navigation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:03:40.865324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T17:01:18.162750Z digest=sha256:44e27468e2032aa22683af1fd5b6815501df736913e7309d3c4e6d0bb86136dd

Observation 657cc18f-a988-48ed-b808-8972ceb6515b · inbound

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation cites this paper.

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation ViNT: A Foundation Model for Visual Navigation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T11:53:23.879641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T11:48:14.888295Z digest=sha256:854902f74a8d16749e24c0fc7b83e9f9e1463e5e29519d9083278deef1cbc3f2

Observation b9f10a16-8489-4590-a084-9ff074af788f · inbound

Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation cites this paper.

Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation ViNT: A Foundation Model for Visual Navigation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:15.766063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T15:39:13.803571Z digest=sha256:01e8081a6fa13651583888e0801fc9a1f5cb7520ec2fcdcf382f4df1c1e3807f

Observation 99530b7f-5da1-4f2c-9db9-d9f6f8902c60 · inbound

Test-time Adversarial Takeover: A Real-time Hijacking Interface against Robotic Diffusion Policies cites this paper.

Test-time Adversarial Takeover: A Real-time Hijacking Interface against Robotic Diffusion Policies ViNT: A Foundation Model for Visual Navigation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.602340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T13:04:32.423616Z digest=sha256:f47303e3e4c7237b435c1f9413940f08d9ed3cf21ebd46df0a9343c63a76f18e

Observation 0a9ddb97-e826-4dbd-a98c-ce4a8a45a628 · inbound

Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models cites this paper.

Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models ViNT: A Foundation Model for Visual Navigation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.691295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T12:52:38.764427Z digest=sha256:1a288e0e46a3d7cba1b0b3d1ddf709f0997fa07210b2e4054bc45a50c3a3bc3e

Observation a922ce2f-ecd0-4232-8a05-eb0f98af1726 · inbound

From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation cites this paper.

From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation ViNT: A Foundation Model for Visual Navigation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:38:04.666292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T09:29:43.030058Z digest=sha256:2afcf0aaff8607270c984b91bd60dacae5ca02ee2e92512365ebdaf8677b8f51

Observation 661ef9d2-4339-4654-9882-fbc7c1398a43 · inbound

NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation cites this paper.

NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation ViNT: A Foundation Model for Visual Navigation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:34.030334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T06:30:17.719105Z digest=sha256:fd1d0494e9b0333b78967592e4d77ec88d1763afcc5b5ed29f38c705ab4a002b

Observation ec350992-313e-4f4f-aab2-a37f10debbd7 · inbound

Slow Brain, Fast Planner: Latency-Resilient VLM-Augmented Urban Navigation cites this paper.

Slow Brain, Fast Planner: Latency-Resilient VLM-Augmented Urban Navigation ViNT: A Foundation Model for Visual Navigation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:09:33.966990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T17:16:52.173943Z digest=sha256:d447ac3c16616a5bd768fca646e1e9170111586125dd506a05d2bc6ec7fb40bc

Observation afbe0154-3e88-4f32-8e4c-187e97abf4f6 · inbound

NavWM: A Unified Navigation World Model for Foresight-Driven Planning cites this paper.

NavWM: A Unified Navigation World Model for Foresight-Driven Planning ViNT: A Foundation Model for Visual Navigation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:29:56.754455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T00:42:30.237545Z digest=sha256:e37dff36961530c7aee03560ed890fd193583dc987716b7c5e9d080300036274

Observation 0678cc9e-6dc4-469d-9dda-6fb8346f5348 · inbound

Learning Robot Visual Navigation in Crowds via Intention-Aware Scene Representations cites this paper.

Learning Robot Visual Navigation in Crowds via Intention-Aware Scene Representations ViNT: A Foundation Model for Visual Navigation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:08.555240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T19:06:41.116675Z digest=sha256:4081543f53e7419684680ce89543b1d24db0397203dc85e1a51184ac8b07e27e

Observation 334c8c08-cd41-4ab2-8cef-b583b016dec2 · inbound

Learning to Navigate Efficiently with Only 0.58M Trainable Parameters cites this paper.

Learning to Navigate Efficiently with Only 0.58M Trainable Parameters ViNT: A Foundation Model for Visual Navigation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T07:33:00.386358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:33:00.386358Z digest=sha256:4f3ec9023497ee7ccbe53e0e7deb6f5f3703ea60adaf534343e8451f92eadd52

Observation 6db71e94-3207-4520-889b-7cf54c88d980 · inbound

Learning to Navigate Efficiently with Only 0.58M Trainable Parameters cites this paper.

Learning to Navigate Efficiently with Only 0.58M Trainable Parameters ViNT: A Foundation Model for Visual Navigation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T07:09:14.936674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:09:14.936674Z digest=sha256:254c29dd50992e13126414a2c033b2cec6d27c4b8f51c8efc6f55f3760aa071e

Observation 5a79a019-96fe-4527-a696-b9995511514e · inbound

SeeSE3: Emergence of 3D Space in Vision Features cites this paper.

SeeSE3: Emergence of 3D Space in Vision Features ViNT: A Foundation Model for Visual Navigation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T02:49:56.083126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:49:56.083126Z digest=sha256:39d32832f124a1ea81ce22e941ee4d74611578ee800fe646de11468b6ea39f7f

Observation 7b5c56b6-b556-4e0b-841c-3128b91a13ed · inbound

G2-Nav: Grounded and Guarded Vision-Language Costmaps for Robot Social Navigation cites this paper.

G2-Nav: Grounded and Guarded Vision-Language Costmaps for Robot Social Navigation ViNT: A Foundation Model for Visual Navigation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T19:29:21.514289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:29:21.514289Z digest=sha256:ba9e2202df77b10785090484c4c61b987bf19ed5b0a755b4397f624f7f08731c

Observation a655afa5-7faf-4430-99c7-a88d41869dba · inbound

ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset cites this paper.

ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset ViNT: A Foundation Model for Visual Navigation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:39.811308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:39.811308Z digest=sha256:d70c387d864c567400c26eef6c69b31706c70c8ebb9e5a181ae1603324b6c763

Observation 8bf941ef-1f3b-47b7-a8bc-f503debb84b3 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills ViNT: A Foundation Model for Visual Navigation

Reference 226

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.206962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.206962Z digest=sha256:f49313d24abdf3d23a3fc9ef3e9b44ce5f9ba201febc16667dc7df719ea7470c

Observation d4e6f2ed-93b4-412d-815c-fe561c056ec6 · inbound

UniNav: A Unified World-Action Diffusion Model for Visual Navigation cites this paper.

UniNav: A Unified World-Action Diffusion Model for Visual Navigation ViNT: A Foundation Model for Visual Navigation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T22:45:35.867887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:45:35.867887Z digest=sha256:b398d5e0792125152901877fd6c42d60a8bcdecd1c8e86d66dda5f49b638b641