Pith. sign in

REVIEW 6 cited by

CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.16649 v1 pith:KSSBYJVO submitted 2022-11-30 cs.CV cs.AIcs.CLcs.RO

classification cs.CVcs.AIcs.CLcs.RO
keywords clipzero-shotlanguagesuccesscapabilityenvironmentsfollowingmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Household environments are visually diverse. Embodied agents performing Vision-and-Language Navigation (VLN) in the wild must be able to handle this diversity, while also following arbitrary language instructions. Recently, Vision-Language models like CLIP have shown great performance on the task of zero-shot object recognition. In this work, we ask if these models are also capable of zero-shot language grounding. In particular, we utilize CLIP to tackle the novel problem of zero-shot VLN using natural language referring expressions that describe target objects, in contrast to past work that used simple language templates describing object classes. We examine CLIP's capability in making sequential navigational decisions without any dataset-specific finetuning, and study how it influences the path that an agent takes. Our results on the coarse-grained instruction following task of REVERIE demonstrate the navigational capability of CLIP, surpassing the supervised baseline in terms of both success rate (SR) and success weighted by path length (SPL). More importantly, we quantitatively show that our CLIP-based zero-shot approach generalizes better to show consistent performance across environments when compared to SOTA, fully supervised learning approaches when evaluated via Relative Change in Success (RCS).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SocietyBench: Forecasting Counterfactual Social-World Evolution

    cs.CL 2026-08 conditional novelty 7.0 of 10

    A new benchmark measures LLM social-world forecasting on anonymized real events, finding the best model reaches 75/100 and agent scaffolding does not help.

  2. Can Pretrained Vision-Language Embeddings Alone Guide Robot Navigation?

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A behavior-cloning policy trained only on frozen SigLIP embeddings reaches 74% of language-specified targets in a simple simulator, versus 100% for a state-aware expert, and takes 3.2x more steps.

  3. TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A zero-shot navigation pipeline using an LLM to extract landmarks, SigLIP to find goal panoramas, and GPT-4o grounding plus dynamic programming to rank paths achieves 88.9% nDTW on R2R-Habitat and 70% Precision@10 for...

  4. LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    LoRA-TTT improves CLIP's zero-shot accuracy under distribution shift by test-time training only low-rank adapters in the image encoder, using entropy and masked-class-token consistency losses.

  5. OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Object-based environment inference via kernel density estimates on learned object features achieves zero-shot room retrieval and beats scene-based CLIP.

  6. Demonstrating CavePI: Autonomous Exploration of Underwater Caves by Semantic Guidance

    cs.RO 2025-02 conditional novelty 5.0 of 10

    A low-cost AUV tracks a cave guide line using lightweight semantic segmentation and a PID controller, with credible tank and open-water results but only limited, partially unsuccessful cave tests.

Pith tools