Pith. sign in

REVIEW 2 cited by

End-to-End (Instance)-Image Goal Navigation through Correspondence as an Emergent Phenomenon

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.16634 v1 pith:E4E5RJYL submitted 2023-09-28 cs.CV

classification cs.CV
keywords goalcorrespondencelearningperceptionproblemvisualdifficultenvironments
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Most recent work in goal oriented visual navigation resorts to large-scale machine learning in simulated environments. The main challenge lies in learning compact representations generalizable to unseen environments and in learning high-capacity perception modules capable of reasoning on high-dimensional input. The latter is particularly difficult when the goal is not given as a category ("ObjectNav") but as an exemplar image ("ImageNav"), as the perception module needs to learn a comparison strategy requiring to solve an underlying visual correspondence problem. This has been shown to be difficult from reward alone or with standard auxiliary tasks. We address this problem through a sequence of two pretext tasks, which serve as a prior for what we argue is one of the main bottleneck in perception, extremely wide-baseline relative pose estimation and visibility prediction in complex scenes. The first pretext task, cross-view completion is a proxy for the underlying visual correspondence problem, while the second task addresses goal detection and finding directly. We propose a new dual encoder with a large-capacity binocular ViT model and show that correspondence solutions naturally emerge from the training signals. Experiments show significant improvements and SOTA performance on the two benchmarks, ImageNav and the Instance-ImageNav variant, where camera intrinsics and height differ between observation and goal.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Memory Proxy Maps for Visual Navigation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A three-tier feudal navigation agent with a self-supervised memory proxy map, a human-imitation waypoint network, and a low-level action classifier achieves state-of-the-art image-goal navigation in unseen Gibson envi...

  2. Multimodal Perception for Goal-oriented Navigation: A Survey

    cs.RO 2025-04 conditional novelty 2.0 of 10

    A literature survey that categorizes multimodal goal-oriented navigation methods into six inference domains and claims this taxonomy reveals cross-task computational patterns.

Pith tools