Pith. sign in

Paper Citation Record · LEDGER

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation?

As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2506.14507.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14507 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:01.919344Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T15:39:59.878169Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T23:13:15.774007Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0499a38-1ba4-43d3-b5d9-aeb475c79970 · outbound

This paper cites Getting vit in shape: Scaling laws for compute-optimal model design,.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Getting vit in shape: Scaling laws for compute-optimal model design,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.555563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:23:00.467925Z digest=sha256:d98d1bec95f158fb2f6e35e87b5160b11dc63a588b6941017c0c1de4cf9cdbf2

Observation 40e9d7d8-c1a9-4405-acf3-7240dd96a381 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabilities, 2024.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Spatialvlm: Endow- ing vision-language models with spatial reasoning capabilities, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.531508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:23:00.566662Z digest=sha256:3900532f89885d104779de31b29dc0ca437bb9754492eeedf77c746f606ef27a

Observation ed17b4aa-87dc-430f-a22e-7824d1b5cc3f · outbound

This paper cites NaVILA: Legged Robot Vision-Language-Action Model for Navigation.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.630568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.630568Z digest=sha256:ad04d50965615884d2b69d1dd95bcdfb6b40b05c2fda7bfaff8558410a99610d

Observation 3c22e683-59c9-4af2-9a43-b6de6d942bcc · outbound

This paper cites CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.698604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.698604Z digest=sha256:0d0b914b2fa68ea68bd827156e6b3b89dabbe0d31422419e4980b0b97ab98c98

Observation e5e2af88-9c60-4abf-99d4-6749b39fe012 · outbound

This paper cites Learning latent dynamics for planning from pixels, 2019.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Learning latent dynamics for planning from pixels, 2019

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.510339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:23:00.783826Z digest=sha256:d2f3ae27902e7818c54b193c8ad06c03dfcef71c065631e7bc53ca385bccb154

Observation 481ba7a2-3ddb-421e-8953-a4f958356b26 · outbound

This paper cites Mastering Diverse Domains through World Models.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Mastering Diverse Domains through World Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.842906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.842906Z digest=sha256:cf8070d118d4df28b901aeba449857fd38c66f445f25323976f165a576b58507

Observation 0040f366-669a-4909-8d08-e03002d5417c · outbound

This paper cites Visual language maps for robot navigation, 2023.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Visual language maps for robot navigation, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.486795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:23:00.924004Z digest=sha256:ad3e92d56896705ee5f5d2bfdda7fb18472c9486e8c2336dd9db787fdb7aade8

Observation 987b5551-2994-42c3-9cad-e7c8825761ad · outbound

This paper cites BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.963437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.963437Z digest=sha256:abdfbb4f64af72b5deaf7f04d7417fbf5613579e62372b77736d4ec20cc0d6c6

Observation a2d883d2-6d9b-4aea-8c14-1e0f207b4e51 · outbound

This paper cites ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.968433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.968433Z digest=sha256:448021921f434a1e12ebd8adce3e508585d7d6ed193d466762182e384c71ea8b

Observation c9a54b4f-259b-4157-ab69-694ee5098b09 · outbound

This paper cites Distilling Realizable Students from Unrealizable Teachers.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Distilling Realizable Students from Unrealizable Teachers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.023823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.023823Z digest=sha256:388997df0b20960fd97afdb17fdc87ca4edf9076dda200c2cd335907c8928b8e

Observation 1049ce38-98fd-44b0-a912-d685f452942f · outbound

This paper cites Constrained Behavior Cloning for Robotic Learning.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Constrained Behavior Cloning for Robotic Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.104551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.104551Z digest=sha256:6cb09656e94ac6127b819073e3c73a3417d95854a7a6003645f6c1b5fd460000

Observation 8389b828-de7b-4414-97ee-71e84bdaa3ed · outbound

This paper cites ZSON: Zero-shot object-goal navigation using multimodal goal embeddings.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? ZSON: Zero-shot object-goal navigation using multimodal goal embeddings

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.465031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:23:01.190903Z digest=sha256:91a76d937859d03fab8302f81c18118a01b081097e8ae5a16136e95dcd14d9e5

Observation d45a9c07-9fde-4449-b71f-79f99530b241 · outbound

This paper cites Orbit: A unified simulation framework for interactive robot learning environ- ments.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Orbit: A unified simulation framework for interactive robot learning environ- ments

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.237600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.237600Z digest=sha256:78ecc505333c9b0bd0b8c99086802042903900817fdd95cdaf686ba15d1ab9a8

Observation c28bb70d-a4d7-4b7d-93d9-cfd44f9e2958 · outbound

This paper cites Language-Conditioned Offline RL for Multi-Robot Navigation.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Language-Conditioned Offline RL for Multi-Robot Navigation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.301764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.301764Z digest=sha256:52fda9b707e0722c408f105ba5484e13ed6b1b369fbfdce745c58ff2afa549f8

Observation 173f7ab1-8664-42c0-a57b-22d6abed694e · outbound

This paper cites R3m: A uni- versal visual representation for robot manipulation,.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? R3m: A uni- versal visual representation for robot manipulation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.446687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:23:01.359403Z digest=sha256:a7d4415180f6e9bba804d4baae161d2fd6708c064298c356c4297dfa70e7184c

Observation 9032cfc7-51d1-4bf7-b85a-1c46bfcda798 · outbound

This paper cites Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.486633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.486633Z digest=sha256:984882ae28f823d165fb1a5f7f2950f2c492e99a1fe5b0c991f3a1afb60301b3

Observation 4dc5ed25-58df-4883-b710-8c2e4851e655 · outbound

This paper cites Isaac sim.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Isaac sim

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.425411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:23:01.494307Z digest=sha256:2f7e2c046635aba8549ed4629d8cda3b7c1020937b44ef000359a3db81056351

Observation 30a9500b-81ca-4169-880a-acde40926f46 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Learning Transferable Visual Models From Natural Language Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.525716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.525716Z digest=sha256:66fb9d96de3b3f25199b2478f1d04b5970d171d11189e5a24734bfee7cbadef8

Observation c66757b3-508a-4717-aa66-2dc5e06524df · outbound

This paper cites Latent Plans for Task-Agnostic Offline Reinforcement Learning.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Latent Plans for Task-Agnostic Offline Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:23:02.127836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:23:01.592832Z digest=sha256:f1c60e3408083389b88b8305ad2701c0022fde39f5cb133a46db74a0686c0d4a

Observation 3f7cb8c2-75c1-4d4f-921f-d2a1ed275cf4 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.655232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.655232Z digest=sha256:fb92832c713bd0aad1d3973041ed7c0f3a3db729676ffabcb242c84a1821430e

Observation 3fa81a0a-5c6a-4212-8c6e-6968a2ccc98d · outbound

This paper cites Vlm-social-nav: Socially aware robot navigation through scoring using vision-language models, 2024.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Vlm-social-nav: Socially aware robot navigation through scoring using vision-language models, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.404975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:23:01.734952Z digest=sha256:c79e10698e6e34af22b82fd05f4cadafc3aa59b0dbda0f05e57342a013147f9b

Observation 775a35c4-692f-4f85-976b-691f70366aa1 · outbound

This paper cites Open-World Object Manipulation using Pre-trained Vision-Language Models.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.797762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.797762Z digest=sha256:9a4b19efd0655d1211cbeb2475a3f7b33cad7554e91cbb9dd7b112d1f2ca27ad

Observation c50d8b92-015b-4c1a-8568-5378e6677997 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Sigmoid Loss for Language Image Pre-Training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.862440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.862440Z digest=sha256:24e4fbd3d7c2f29d569aa2815b6cca229e88579f2c0eeda10a6e946fd89d84b6

Observation 8e4c8fb4-92e4-4e61-b02a-a8b3e0cf5ee3 · outbound

This paper cites RT-2: Vision- language-action models transfer web knowledge to robotic control.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? RT-2: Vision- language-action models transfer web knowledge to robotic control

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:02.385497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:23:01.919344Z digest=sha256:64f39f16fa9fde94a4483b55f60ceb14252b12cd92e4e291aa562e3a2c8b3945

Observation 41596763-d6ca-4f0a-ad20-f00e05c1f39b · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? R3M: A Universal Visual Representation for Robot Manipulation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:01.438793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:01.438793Z digest=sha256:1f88129e28b4bb7ecdc5887a82f2cc2dddb725cc51a4102dee2f660d6d04791d

Observation b7bdfaf7-9860-469c-845b-9238aa34cd6f · outbound

This paper cites Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.522629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.522629Z digest=sha256:54c55fda2dc69ae0c063e00caf58278975d511504052736d1b4ba8440bb10e62

Pith citing papers

Observation bcff18d3-129c-4ce3-a3e5-0edbe42ff4a2 · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation?

Reference 248

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:13:15.777167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:373c3c6201d9105acbc358cc3ac705952ac5f207bd924e29998c45ff95d0f50f

Observation b634dbf6-4d76-4ab7-98b0-f537bcf8a414 · inbound

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation cites this paper.

A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation?

Reference 238

Resolution
unresolved
no resolver link, observed 2026-07-14T15:39:59.878169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:39:59.878169Z digest=sha256:1f5d697c6e684c97fae8836651c9901bf18da9868e7cfc0413c5e3bbb774d8a6