Pith. sign in

Paper Citation Record · LEDGER

CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2211.16649.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.16649 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:31:23.364399Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:09:59.557766Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 02e1bf55-4766-44e9-b8df-eae642b669fc · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:25:59.157170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:0ae1e54ca25926aab6ac3405bc2f866b3c8885a1630fc0f722c5c8a6eae310bd

Observation 28af2476-28df-4681-af85-cadd6162b821 · inbound

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks cites this paper.

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:51:36.288054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T19:51:36.137985Z digest=sha256:e36ed013b5dedf8d89389b1e987fff4eff85daf45df268c2a026a473613ba91d

Observation d12c190c-8d2f-45ff-83d0-4fc9f528a92b · inbound

LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models cites this paper.

LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T13:31:23.364399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:31:23.364399Z digest=sha256:0dc03fb1279add8f8f6582d94600913df96743c8a55c33e157399462e0d236b9

Observation 9be30479-98b8-44f2-a96e-87b6c3783f3a · inbound

Demonstrating CavePI: Autonomous Exploration of Underwater Caves by Semantic Guidance cites this paper.

Demonstrating CavePI: Autonomous Exploration of Underwater Caves by Semantic Guidance CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T19:38:00.276632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:38:00.276632Z digest=sha256:f0adee4786ddf9d8defdf676f45b65ebe6df1d5a0a0cbd47aaf03845012fea56

Observation 3d7e7b90-3d0c-463d-a931-6e44fc8ac08d · inbound

TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation cites this paper.

TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T13:14:30.081511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:14:30.081511Z digest=sha256:e338d8c766a6c25e8cdaefe72c11832fd3a6e76aa49e21776a4348c17ac737ae

Observation 3c22e683-59c9-4af2-9a43-b6de6d942bcc · inbound

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? cites this paper.

Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation? CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:00.698604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:00.698604Z digest=sha256:e821eac8c42fc3dc5c8a0bb67896cf5ba409e9afbb927ddeeb1918c5935278e9

Observation d9b511d7-f51c-44ef-bfff-bcde7103d964 · inbound

OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference cites this paper.

OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:23.231918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:45:23.231918Z digest=sha256:6d45149f8461f071c02b8c728b9cd5feba099406b32e04147e3e39a1ef22fda1

Observation 2647ea3d-4f34-44d7-9f22-98653abaa3d0 · inbound

DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes cites this paper.

DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:43:16.523905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T18:42:01.266249Z digest=sha256:abe7876fb0a20b16e6f475e1b2b358ac1e9851d600acc3b323b20cda6152515e

Observation ae29507b-a075-4f3f-818d-cc19d812ecfb · inbound

How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace cites this paper.

How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:25:57.266116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:09:39.510770Z digest=sha256:0c3d1fc4b61d5abb73663e84d3be9ce41e554f747b76dfd6dd527fcaa57a85c6

Observation 7ae10c10-5436-48a9-b4e9-38ce84c6f502 · inbound

FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation cites this paper.

FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:17:37.783102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:13:12.029402Z digest=sha256:80e80d9dece4cda37dc6a8ed91ef881572789139663b3d79b0e952a5dec7daf6

Observation 4f670a4a-86c0-4cd4-9a19-9c12edc3ff19 · inbound

NORM-Nav: Zero-Shot Mobile Robot Navigation with Natural Language Behavioral Constraints cites this paper.

NORM-Nav: Zero-Shot Mobile Robot Navigation with Natural Language Behavioral Constraints CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:12:44.705181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T20:11:52.658283Z digest=sha256:0d075d66307bf6c355df59c7f66e3f4f165a1d71200cf7dad226af4fe97d36c5

Observation bf3fa07c-7c51-4b30-a3e2-ec3b920608f4 · inbound

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation cites this paper.

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T11:53:23.904587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T11:48:14.888295Z digest=sha256:0bb688cbdeb8ebd1588f592bfef57e10b5b2fc3a6c738c194b27f61dd431f142

Observation 3ba348a1-0196-4ec2-90ea-b11d30a25461 · inbound

Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control cites this paper.

Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:09:59.559277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T23:51:39.882029Z digest=sha256:aa450eaf9928a3576a0373f7608107b47a3ac21bcc02ee035dac5f1559ea029e

Observation 8d2c8d4b-e06d-4a01-93ba-54e93aa04528 · inbound

SocietyBench: Forecasting Counterfactual Social-World Evolution cites this paper.

SocietyBench: Forecasting Counterfactual Social-World Evolution CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-05T04:17:18.712661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:17:18.712661Z digest=sha256:2f6bd56b3e5b199f767eafccad6abf9c64ed12dc25691932fbd3a9b576acf109