Pith. sign in

REVIEW 3 cited by

Learning Cross-view Visual Geo-localization without Ground Truth

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.12702 v1 pith:PVDIAOCC submitted 2024-03-19 cs.CV

classification cs.CV
keywords unlabeledcross-viewdatatrainingadaptermodelsadaptationchallenges
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cross-View Geo-Localization (CVGL) involves determining the geographical location of a query image by matching it with a corresponding GPS-tagged reference image. Current state-of-the-art methods predominantly rely on training models with labeled paired images, incurring substantial annotation costs and training burdens. In this study, we investigate the adaptation of frozen models for CVGL without requiring ground truth pair labels. We observe that training on unlabeled cross-view images presents significant challenges, including the need to establish relationships within unlabeled data and reconcile view discrepancies between uncertain queries and references. To address these challenges, we propose a self-supervised learning framework to train a learnable adapter for a frozen Foundation Model (FM). This adapter is designed to map feature distributions from diverse views into a uniform space using unlabeled data exclusively. To establish relationships within unlabeled data, we introduce an Expectation-Maximization-based Pseudo-labeling module, which iteratively estimates associations between cross-view features and optimizes the adapter. To maintain the robustness of the FM's representation, we incorporate an information consistency module with a reconstruction loss, ensuring that adapted features retain strong discriminative ability across views. Experimental results demonstrate that our proposed method achieves significant improvements over vanilla FMs and competitive accuracy compared to supervised methods, while necessitating fewer training parameters and relying solely on unlabeled data. Evaluation of our adaptation for task-specific models further highlights its broad applicability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cross-View Image Set Geo-Localization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Using several unordered ground-view photos as a query set improves cross-view geo-localization accuracy, and the proposed FlexGeo model achieves state-of-the-art results on a new six-city benchmark.

  2. Cross-View Geo-Localization with Street-View and VHR Satellite Imagery in Decentrality Settings

    cs.CV 2024-12 conditional novelty 6.0 of 10

    The authors introduce the decentrality problem in cross-view geo-localization and present DReSS, a worldwide dataset with strong query-reference offsets, plus AuxGeo, an auxiliary-task method that achieves state-of-th...

  3. Scale-adaptive UAV Geo-localization via Height-aware Partition Learning

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A height-aware partition network (SaLPN) adjusts the size of feature partitions based on relative drone/satellite height, improving UAV geo-localization accuracy under scale mismatches.

Pith tools