Pith. sign in

REVIEW 3 cited by

Open-vocabulary Mobile Manipulation in Unseen Dynamic Environments with 3D Semantic Maps

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.18115 v1 pith:VTIF4C5B submitted 2024-06-26 cs.RO cs.AIcs.CV

classification cs.ROcs.AIcs.CV
keywords semanticmanipulationspatialdynamicframeworkinstructionslanguagemobile
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Open-Vocabulary Mobile Manipulation (OVMM) is a crucial capability for autonomous robots, especially when faced with the challenges posed by unknown and dynamic environments. This task requires robots to explore and build a semantic understanding of their surroundings, generate feasible plans to achieve manipulation goals, adapt to environmental changes, and comprehend natural language instructions from humans. To address these challenges, we propose a novel framework that leverages the zero-shot detection and grounded recognition capabilities of pretraining visual-language models (VLMs) combined with dense 3D entity reconstruction to build 3D semantic maps. Additionally, we utilize large language models (LLMs) for spatial region abstraction and online planning, incorporating human instructions and spatial semantic context. We have built a 10-DoF mobile manipulation robotic platform JSR-1 and demonstrated in real-world robot experiments that our proposed framework can effectively capture spatial semantics and process natural language user instructions for zero-shot OVMM tasks under dynamic environment settings, with an overall navigation and task success rate of 80.95% and 73.33% over 105 episodes, and better SFT and SPL by 157.18% and 19.53% respectively compared to the baseline. Furthermore, the framework is capable of replanning towards the next most probable candidate location based on the spatial semantic context derived from the 3D semantic map when initial plans fail, keeping an average success rate of 76.67%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SD-OVON: A Semantics-aware Dataset and Benchmark Generation Pipeline for Open-Vocabulary Object Navigation in Dynamic Scenes

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A dataset and benchmark pipeline that generates dynamic indoor scenes and Open-Vocabulary Object Navigation episodes by placing movable objects on receptacles according to LLM-derived commonsense rules.

  2. Language-Conditioned Open-Vocabulary Mobile Manipulation with Pretrained Models

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A robot system that combines GPT-4, vision-language maps, and a CLIPort-style network follows free-form household commands across rooms in simulation, reaching 10.2% average success on unseen tasks and beating two bas...

  3. OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A vision-language model fine-tuned on 572K synthetic simulation examples improves open-world mobile manipulation action decisions and object grounding over GPT-4o, with 21.9% full-task success in simulation and 90% ac...

Pith tools