Pith. sign in

REVIEW 2 cited by

UniLoc: Towards Universal Place Recognition Using Any Single Modality

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.12079 v1 pith:NK7BLAV2 submitted 2024-12-16 cs.CV

classification cs.CV
keywords placerecognitionunilocmatchingachievingcross-modalgreatermethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To date, most place recognition methods focus on single-modality retrieval. While they perform well in specific environments, cross-modal methods offer greater flexibility by allowing seamless switching between map and query sources. It also promises to reduce computation requirements by having a unified model, and achieving greater sample efficiency by sharing parameters. In this work, we develop a universal solution to place recognition, UniLoc, that works with any single query modality (natural language, image, or point cloud). UniLoc leverages recent advances in large-scale contrastive learning, and learns by matching hierarchically at two levels: instance-level matching and scene-level matching. Specifically, we propose a novel Self-Attention based Pooling (SAP) module to evaluate the importance of instance descriptors when aggregated into a place-level descriptor. Experiments on the KITTI-360 dataset demonstrate the benefits of cross-modality for place recognition, achieving superior performance in cross-modal settings and competitive results also for uni-modal scenarios. Our project page is publicly available at https://yan-xia.github.io/projects/UniLoc/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SGLoc: Semantic Localization System for Camera Pose Estimation from 3D Gaussian Splatting Representation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A semantic retrieval and rendering-refinement pipeline estimates camera poses from 3D Gaussian Splatting maps without an initial pose prior, reporting state-of-the-art median errors on 7Scenes and 12Scenes.

  2. G2IA: Geometry-Guided Instance-Aware Retrieval and Refinement for Cross-Modal Place Recognition

    cs.CV 2026-06 conditional novelty 5.0 of 10

    G2IA improves image-to-LiDAR place recognition by combining visual-geometry and instance-aware descriptors with a shape-and-layout candidate re-ranking stage.

Pith tools