Pith. sign in

REVIEW 22 cited by

TopV-Nav: Unlocking the Top-View Spatial Reasoning Potential of MLLM for Zero-shot Object Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.16425 v2 pith:KQOV7ZTK submitted 2024-11-25 cs.CV cs.AIcs.RO

classification cs.CVcs.AIcs.RO
keywords spatialtop-viewinformationreasoningobjectpotentialtopv-navdirectly
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The Zero-Shot Object Navigation (ZSON) task requires embodied agents to find a previously unseen object by navigating in unfamiliar environments. Such a goal-oriented exploration heavily relies on the ability to perceive, understand, and reason based on the spatial information of the environment. However, current LLM-based approaches convert visual observations to language descriptions and reason in the linguistic space, leading to the loss of spatial information. In this paper, we introduce TopV-Nav, an MLLM-based method that directly reasons on the top-view map with sufficient spatial information. To fully unlock the MLLM's spatial reasoning potential in top-view perspective, we propose the Adaptive Visual Prompt Generation (AVPG) method to adaptively construct semantically-rich top-view map. It enables the agent to directly utilize spatial information contained in the top-view map to conduct thorough reasoning. Besides, we design a Dynamic Map Scaling (DMS) mechanism to dynamically zoom top-view map at preferred scales, enhancing local fine-grained reasoning. Additionally, we devise a Potential Target Driven (PTD) mechanism to predict and to utilize target locations, facilitating global and human-like exploration. Experiments on MP3D and HM3D datasets demonstrate the superiority of our TopV-Nav.

Discussion (0). Sign in to comment.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SkillNav: Score-Level Skill Intervention for Zero-Shot Object Goal Navigation

    cs.RO 2026-07 conditional novelty 7.0 of 10

    A training-free skill layer that modifies the VLM's value map improves zero-shot object-goal navigation SPL by up to 6.0 points on MP3D and HM3D.

  2. MCNav: Memory-Aware Dynamic Cognitive Map for Zero-shot Goal-oriented Navigation

    cs.RO 2026-05 unverdicted novelty 7.0 of 10

    MCNav builds a dynamic cognitive map with goal re-validation and missed-goal re-exploration to reach state-of-the-art results on instance-level zero-shot navigation in HM3D environments.

  3. NavOne: One-Step Global Planning for Vision-Language Navigation on Top-Down Maps

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    NavOne enables one-step global navigation planning on top-down maps using a unified multi-modal framework, achieving state-of-the-art results and up to 80x speedup on the new R2R-TopDown dataset.

  4. NavOne: One-Step Global Planning for Vision-Language Navigation on Top-Down Maps

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    NavOne reformulates vision-language navigation as single-step global path planning on top-down maps, delivering state-of-the-art results and 8x-80x speedups over prior map-based and egocentric baselines.

  5. Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A three-stage parse-search-confirm MLLM pipeline plus structured spatial memory sets training-free SOTA on AVDN, matching or beating several supervised methods on ANDH and ANDH-Full.

  6. VTM-Nav: Harnessing Cross-Episode Experience for Object-Goal Navigation with Hierarchical Visual-Topological Memory

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A hierarchical room-and-object memory that persists across independent ObjectNav episodes yields small success-rate gains, but most of the gain comes from within-episode memory rather than the cross-episode component.

  7. NavOne: One-Step Global Planning for Vision-Language Navigation on Top-Down Maps

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    NavOne performs one-step global path planning for vision-language navigation by predicting dense path probabilities directly on fused multi-modal top-down maps.

  8. NavOne: One-Step Global Planning for Vision-Language Navigation on Top-Down Maps

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    NavOne enables one-step global path planning for vision-language navigation on top-down maps via a unified neural framework, achieving SOTA among map-based methods with 8x and 80x speedups on the new R2R-TopDown dataset.

  9. Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    GLMap combines explicit 3D Gaussians with multi-scale language semantics in a dual-modality structure and uses an analytical Gaussian Estimator for incremental map building, improving zero-shot performance on navigati...

  10. FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    FineCog-Nav uses fine-grained cognitive modules driven by foundation models to outperform zero-shot baselines in UAV navigation and introduces the AerialVLN-Fine benchmark with refined instructions.

  11. OVAL: Open-Vocabulary Augmented Memory Model for Lifelong Object Goal Navigation

    cs.RO 2026-04 unverdicted novelty 6.0 of 10

    OVAL introduces an open-vocabulary memory model with structured descriptors and multi-value frontier scoring to enable efficient lifelong object goal navigation in unseen settings.

  12. HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    HiRO-Nav adaptively triggers reasoning only on high-entropy actions via a hybrid training pipeline and shows better success-token trade-offs than always-reason or never-reason baselines on the CHORES-S benchmark.

  13. ReMemNav: A Rethinking and Memory-Augmented Framework for Zero-Shot Object Navigation

    cs.RO 2026-03 conditional novelty 6.0 of 10

    ReMemNav improves zero-shot object navigation success and efficiency by integrating episodic memory and rethinking with VLMs, achieving SR/SPL gains of 1.7%/7.0% on HM3D v0.1, 18.2%/11.1% on HM3D v0.2, and 8.7%/7.9% on MP3D.

  14. SignScene: Visual Sign Grounding for Mapless Navigation

    cs.RO 2026-02 conditional novelty 6.0 of 10

    A sign-centric abstract map lets a vision-language model turn navigational sign instructions into correct paths 88.6% of the time across nine environment types.

  15. MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation

    cs.CV 2026-02 unverdicted novelty 6.0 of 10

    MerNav's Memory-Execute-Review framework improves success rates in zero-shot object goal navigation by 5-8% over baselines on four datasets while outperforming both training-free and supervised methods on key benchmarks.

  16. Visual-Language-Guided Task Planning for Horticultural Robots

    cs.RO 2026-01 conditional novelty 6.0 of 10

    A vision-language model drives a simulated greenhouse robot through simple crop-inspection tasks with ~87% success, but long multi-target tasks collapse to under 10% success.

  17. SkeNa: Learning to Navigate Unseen Environments Based on Abstract Hand-Drawn Maps

    cs.RO 2025-08 unverdicted novelty 6.0 of 10

    A navigation agent can follow abstract hand-drawn sketch maps to reach goals in unseen indoor environments, backed by a new 54k-pair dataset and a model with a 105 percent relative SPL gain.

  18. IntentNav: Learning Spatial-Visual Object Navigation from Human Demonstrations

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    IntentNav is a spatial-visual imitation framework that infers human search intent via frontier labeling to train VLM policies for object navigation, reporting SOTA on MP3D and HM3D benchmarks with zero-shot transfer t...

  19. CLUE: Adaptively Prioritized Contextual Cues by Leveraging a Unified Semantic Map for Effective Zero-Shot Object-Goal Navigation

    cs.RO 2026-05 unverdicted novelty 5.0 of 10

    CLUE adaptively weights room-type and object-co-location cues from an LLM to construct a unified semantic value map that improves success rate and efficiency in zero-shot object-goal navigation.

  20. Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards

    cs.LG 2026-03 unverdicted novelty 5.0 of 10

    Discounted Beta-Bernoulli reward estimation reduces variance and variance collapse in group RLVR, improving GRPO Acc@8 on reasoning benchmarks at no extra cost.

  21. A Deployable Embodied Vision-Language Navigation System with Hierarchical Cognition and Context-Aware Exploration

    cs.RO 2026-04 unverdicted novelty 4.0 of 10

    Introduces a hierarchical VLN architecture with asynchronous layers, incremental memory graph, and WTRP-based exploration that improves success and efficiency on resource-constrained robots.

  22. A Deployable Embodied Vision-Language Navigation System with Hierarchical Cognition and Context-Aware Exploration

    cs.RO 2026-04 unverdicted novelty 4.0 of 10

    A modular VLN architecture builds a cognitive memory graph, decomposes it for VLM reasoning, and solves a weighted traveling repairman problem for context-aware exploration to achieve real-time performance and higher ...

Pith tools