REVIEW 4 cited by
ESC: Exploration with Soft Commonsense Constraints for Zero-shot Object Navigation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The ability to accurately locate and navigate to a specific object is a crucial capability for embodied agents that operate in the real world and interact with objects to complete tasks. Such object navigation tasks usually require large-scale training in visual environments with labeled objects, which generalizes poorly to novel objects in unknown environments. In this work, we present a novel zero-shot object navigation method, Exploration with Soft Commonsense constraints (ESC), that transfers commonsense knowledge in pre-trained models to open-world object navigation without any navigation experience nor any other training on the visual environments. First, ESC leverages a pre-trained vision and language model for open-world prompt-based grounding and a pre-trained commonsense language model for room and object reasoning. Then ESC converts commonsense knowledge into navigation actions by modeling it as soft logic predicates for efficient exploration. Extensive experiments on MP3D, HM3D, and RoboTHOR benchmarks show that our ESC method improves significantly over baselines, and achieves new state-of-the-art results for zero-shot object navigation (e.g., 288% relative Success Rate improvement than CoW on MP3D).
Forward citations
Cited by 4 Pith papers
-
VTM-Nav: Harnessing Cross-Episode Experience for Object-Goal Navigation with Hierarchical Visual-Topological Memory
A hierarchical room-and-object memory that persists across independent ObjectNav episodes yields small success-rate gains, but most of the gain comes from within-episode memory rather than the cross-episode component.
-
OpenGuide: Assistive Object Retrieval in Indoor Spaces for Individuals with Visual Impairments
OpenGuide combines vision-language value maps, frontier exploration, and POMDP planning to locate multiple objects in unfamiliar indoor spaces, reaching about 55% success in simulation and 54% in real-world trials.
-
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
A streaming navigation framework combining a sliding-window KV cache with depth-based token pruning achieves state-of-the-art results on VLN-CE benchmarks with bounded context and low latency.
-
Imaging the disk-halo interface of NGC 891: a 2.7 kpc-thick molecular gas disk
NGC 891’s molecular gas has a thin (~360 pc) plus thick (~1.1 kpc FWHM) disk, with CO detected to 1.3–1.4 kpc and up to ~27% of H2 mass in the thick component.
Discussion (0). Sign in to comment.