Pith. sign in

REVIEW 10 cited by

Instance-Specific Image Goal Navigation: Training Embodied Agents to Find Object Instances

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.15876 v1 pith:P3TWSSDA submitted 2022-11-29 cs.CV

classification cs.CV
keywords agentimageimagenavnavigationcameraembodiedgoalimage-goals
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We consider the problem of embodied visual navigation given an image-goal (ImageNav) where an agent is initialized in an unfamiliar environment and tasked with navigating to a location 'described' by an image. Unlike related navigation tasks, ImageNav does not have a standardized task definition which makes comparison across methods difficult. Further, existing formulations have two problematic properties; (1) image-goals are sampled from random locations which can lead to ambiguity (e.g., looking at walls), and (2) image-goals match the camera specification and embodiment of the agent; this rigidity is limiting when considering user-driven downstream applications. We present the Instance-specific ImageNav task (InstanceImageNav) to address these limitations. Specifically, the goal image is 'focused' on some particular object instance in the scene and is taken with camera parameters independent of the agent. We instantiate InstanceImageNav in the Habitat Simulator using scenes from the Habitat-Matterport3D dataset (HM3D) and release a standardized benchmark to measure community progress.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation

    cs.CV 2026-02 conditional novelty 7.0 of 10

    LangMap is a human-verified navigation benchmark with 18K tasks spanning scene-, room-, region-, and instance-level goals in real-world 3D scans, covering 414 object categories.

  2. VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A new vision-only long-horizon aerial navigation benchmark (VoLN-UAV) and a baseline method (VoLN-MLLM) that achieve 7.4%/4.5%/1.8% success on unseen Easy/Normal/Hard episodes.

  3. Learning to Localize Reference Trajectories in Image-Space for Visual Navigation

    cs.RO 2026-02 conditional novelty 6.0 of 10

    LoTIS localizes a reference RGB trajectory in the robot's current view, predicting image-space coordinates, visibility, and distance to provide robot-agnostic guidance for navigation.

  4. SplatSearch: Instance Image Goal Navigation for Mobile Robots using 3D Gaussian Splatting and Diffusion Models

    cs.RO 2025-11 conditional novelty 6.0 of 10

    SplatSearch combines sparse-view 3D Gaussian Splatting, multi-view diffusion inpainting, and semantic/visual frontier scoring to achieve viewpoint-invariant instance image-goal navigation in unknown environments.

  5. ObjectReact: Learning Object-Relative Control for Visual Navigation

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A controller trained purely on object-relative costmaps beats image-goal navigation on alt-goal, shortcut, and reverse-route tasks, with nearly height-invariant performance and qualitative sim-to-real transfer.

  6. TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A navigation pipeline that bridges object-level global planning with traversability-aware local control, using only RGB images and pretrained models, improves success over prior zero-shot and learned baselines in simulation.

  7. OctoNav: Towards Generalist Embodied Navigation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    OctoNav-R1, trained with SFT, GRPO, and online RL on the new OctoNav-Bench, achieves 19.4% overall success on mixed-instruction navigation, more than double the best baseline.

  8. Hierarchical Scoring with 3D Gaussian Splatting for Instance Image-Goal Navigation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A two-stage scorer, CLIP semantic prefiltering plus DINOv2 geometric matching over a 3D Gaussian map, achieves a 0.784 success rate on HM3D instance image-goal navigation.

  9. A Large Catalog of DA White Dwarf Characteristics Using SDSS and Gaia Observations

    astro-ph.SR 2025-08 unverdicted novelty 5.0 of 10

    The paper publishes the largest catalog of DA white dwarf measurements to date, using SDSS DR19 plus earlier SDSS data and Gaia, and reports a systematic offset between SDSS-V and older SDSS measurements.

  10. NavBench: Probing Multimodal Large Language Models for Embodied Navigation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    NavBench introduces a two-part benchmark for zero-shot embodied navigation evaluation, showing that MLLMs' navigation comprehension correlates with execution and that temporal progress tracking is a major bottleneck.

Pith tools