REVIEW 10 cited by
Instance-Specific Image Goal Navigation: Training Embodied Agents to Find Object Instances
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We consider the problem of embodied visual navigation given an image-goal (ImageNav) where an agent is initialized in an unfamiliar environment and tasked with navigating to a location 'described' by an image. Unlike related navigation tasks, ImageNav does not have a standardized task definition which makes comparison across methods difficult. Further, existing formulations have two problematic properties; (1) image-goals are sampled from random locations which can lead to ambiguity (e.g., looking at walls), and (2) image-goals match the camera specification and embodiment of the agent; this rigidity is limiting when considering user-driven downstream applications. We present the Instance-specific ImageNav task (InstanceImageNav) to address these limitations. Specifically, the goal image is 'focused' on some particular object instance in the scene and is taken with camera parameters independent of the agent. We instantiate InstanceImageNav in the Habitat Simulator using scenes from the Habitat-Matterport3D dataset (HM3D) and release a standardized benchmark to measure community progress.
Forward citations
Cited by 10 Pith papers
-
LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation
LangMap is a human-verified navigation benchmark with 18K tasks spanning scene-, room-, region-, and instance-level goals in real-world 3D scans, covering 414 object categories.
-
VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method
A new vision-only long-horizon aerial navigation benchmark (VoLN-UAV) and a baseline method (VoLN-MLLM) that achieve 7.4%/4.5%/1.8% success on unseen Easy/Normal/Hard episodes.
-
Learning to Localize Reference Trajectories in Image-Space for Visual Navigation
LoTIS localizes a reference RGB trajectory in the robot's current view, predicting image-space coordinates, visibility, and distance to provide robot-agnostic guidance for navigation.
-
SplatSearch: Instance Image Goal Navigation for Mobile Robots using 3D Gaussian Splatting and Diffusion Models
SplatSearch combines sparse-view 3D Gaussian Splatting, multi-view diffusion inpainting, and semantic/visual frontier scoring to achieve viewpoint-invariant instance image-goal navigation in unknown environments.
-
ObjectReact: Learning Object-Relative Control for Visual Navigation
A controller trained purely on object-relative costmaps beats image-goal navigation on alt-goal, shortcut, and reverse-route tasks, with nearly height-invariant performance and qualitative sim-to-real transfer.
-
TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
A navigation pipeline that bridges object-level global planning with traversability-aware local control, using only RGB images and pretrained models, improves success over prior zero-shot and learned baselines in simulation.
-
OctoNav: Towards Generalist Embodied Navigation
OctoNav-R1, trained with SFT, GRPO, and online RL on the new OctoNav-Bench, achieves 19.4% overall success on mixed-instruction navigation, more than double the best baseline.
-
Hierarchical Scoring with 3D Gaussian Splatting for Instance Image-Goal Navigation
A two-stage scorer, CLIP semantic prefiltering plus DINOv2 geometric matching over a 3D Gaussian map, achieves a 0.784 success rate on HM3D instance image-goal navigation.
-
A Large Catalog of DA White Dwarf Characteristics Using SDSS and Gaia Observations
The paper publishes the largest catalog of DA white dwarf measurements to date, using SDSS DR19 plus earlier SDSS data and Gaia, and reports a systematic offset between SDSS-V and older SDSS measurements.
-
NavBench: Probing Multimodal Large Language Models for Embodied Navigation
NavBench introduces a two-part benchmark for zero-shot embodied navigation evaluation, showing that MLLMs' navigation comprehension correlates with execution and that temporal progress tracking is a major bottleneck.
Discussion (0). Continue with ORCID to comment.