Pith. sign in

REVIEW 4 cited by

VLN-Game: Vision-Language Equilibrium Search for Zero-Shot Semantic Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.11609 v1 pith:MWHNAHMK submitted 2024-11-18 cs.RO

classification cs.RO
keywords navigationtargetvln-gamelanguageframeworkobjectsearchenvironment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Following human instructions to explore and search for a specified target in an unfamiliar environment is a crucial skill for mobile service robots. Most of the previous works on object goal navigation have typically focused on a single input modality as the target, which may lead to limited consideration of language descriptions containing detailed attributes and spatial relationships. To address this limitation, we propose VLN-Game, a novel zero-shot framework for visual target navigation that can process object names and descriptive language targets effectively. To be more precise, our approach constructs a 3D object-centric spatial map by integrating pre-trained visual-language features with a 3D reconstruction of the physical environment. Then, the framework identifies the most promising areas to explore in search of potential target candidates. A game-theoretic vision language model is employed to determine which target best matches the given language description. Experiments conducted on the Habitat-Matterport 3D (HM3D) dataset demonstrate that the proposed framework achieves state-of-the-art performance in both object goal navigation and language-based navigation tasks. Moreover, we show that VLN-Game can be easily deployed on real-world robots. The success of VLN-Game highlights the promising potential of using game-theoretic methods with compact vision-language models to advance decision-making capabilities in robotic systems. The supplementary video and code can be accessed via the following link: https://sites.google.com/view/vln-game.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Instance-Enriched Semantic Maps for Visual Language Navigation

    cs.RO 2026-07 unverdicted novelty 6.0 of 10

    Instance-enriched 2.5D semantic maps plus LLM expert fusion improve VLN object retrieval by >17% and success by >23% while cutting storage ~96% versus 3D baselines.

  2. MAG-Nav: Language-Driven Object Navigation Leveraging Memory-Reserved Active Grounding

    cs.RO 2025-08 conditional novelty 6.0 of 10

    MAG-Nav uses active viewpoint selection and memory replay with GPT-4o to achieve state-of-the-art 40.8% success in zero-shot language-driven object navigation on GOAT-Bench/HM3D.

  3. CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process Thinking

    cs.AI 2025-07 conditional novelty 5.0 of 10

    CogDDN uses a fast heuristic VLM paired with a slow analytic reflection process and a growing knowledge base to navigate to objects that implicitly satisfy a user's demand, with large reported gains on AI2Thor DDN benchmarks.

  4. Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers

    cs.AI 2025-02 conditional novelty 5.0 of 10

    A taxonomy-based survey of bidirectional game theory and LLM research, spanning evaluation, alignment, economic competition, and LLM-driven game solving.

Pith tools