REVIEW 10 cited by
GOAT: GO to Any Thing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In deployment scenarios such as homes and warehouses, mobile robots are expected to autonomously navigate for extended periods, seamlessly executing tasks articulated in terms that are intuitively understandable by human operators. We present GO To Any Thing (GOAT), a universal navigation system capable of tackling these requirements with three key features: a) Multimodal: it can tackle goals specified via category labels, target images, and language descriptions, b) Lifelong: it benefits from its past experience in the same environment, and c) Platform Agnostic: it can be quickly deployed on robots with different embodiments. GOAT is made possible through a modular system design and a continually augmented instance-aware semantic memory that keeps track of the appearance of objects from different viewpoints in addition to category-level semantics. This enables GOAT to distinguish between different instances of the same category to enable navigation to targets specified by images and language descriptions. In experimental comparisons spanning over 90 hours in 9 different homes consisting of 675 goals selected across 200+ different object instances, we find GOAT achieves an overall success rate of 83%, surpassing previous methods and ablations by 32% (absolute improvement). GOAT improves with experience in the environment, from a 60% success rate at the first goal to a 90% success after exploration. In addition, we demonstrate that GOAT can readily be applied to downstream tasks such as pick and place and social navigation.
Forward citations
Cited by 10 Pith papers
-
SplatSearch: Instance Image Goal Navigation for Mobile Robots using 3D Gaussian Splatting and Diffusion Models
SplatSearch combines sparse-view 3D Gaussian Splatting, multi-view diffusion inpainting, and semantic/visual frontier scoring to achieve viewpoint-invariant instance image-goal navigation in unknown environments.
-
TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
A navigation pipeline that bridges object-level global planning with traversability-aware local control, using only RGB images and pretrained models, improves success over prior zero-shot and learned baselines in simulation.
-
WoMAP: World Models For Embodied Open-Vocabulary Object Localization
WoMAP generates training data from Gaussian Splatting scenes, distills detector confidence into a latent world model, and uses that model to refine vision-language action proposals for open-vocabulary object localization.
-
GET: Goal-directed Exploration and Targeting for Large-Scale Unknown Environments
A framework coupling LLM direction proposals with an external evaluator and a Gaussian-mixture experience memory reduces object-search path length and time versus heuristic baselines in real-world large-scale robot trials.
-
VL-Explore: Zero-shot Vision-Language Exploration and Target Discovery by Mobile Robots
A monocular, map-free navigation pipeline uses CLIP scores on six image tiles to explore rooms and discover a target in real time.
-
Learning Category-level Last-meter Navigation from RGB Demonstrations of a Single-instance
An RGB-only imitation-learning policy trained on one chair generalizes to unseen chairs and environments for last-meter base positioning, though reported success depends on an added heuristic stopping rule and a 0.3 m...
-
N2M: Bridging Navigation and Manipulation by Learning Pose Preference from Rollout
N2M predicts preferable base poses for manipulation policies from ego-centric point clouds, learned from rollouts, lifting success from 3% to 54% in the PnPCounterToCab task.
-
IRS: Instance-Level 3D Scene Graphs via Room Prior Guided LiDAR-Camera Fusion
IRS builds instance-level 3D scene graphs faster by using LiDAR room priors to constrain and parallelize semantic fusion from vision-language models.
-
DELIVER: A System for LLM-Guided Coordinated Multi-Robot Pickup and Delivery using Voronoi-Based Relay Planning
DELIVER combines an LLM parser, Voronoi territory division, and relay handoffs so multiple robots can cooperatively deliver an item from a single spoken instruction.
-
SPG: Style-Prompting Guidance for Style-Specific Content Creation
SPG is not described anywhere in the supplied text; the body is a different paper (OVSegDT) about robot navigation.
Discussion (0). Continue with ORCID to comment.