Pith. sign in

REVIEW 7 cited by

Efficient and Robust 2D-to-BEV Representation Learning via Geometry-guided Kernel Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.04584 v1 pith:VFH6ZTNG submitted 2022-06-09 cs.CV

classification cs.CV
keywords representationkernellearningtransformercamerad-to-bevgeometry-guidedgreat
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Learning Bird's Eye View (BEV) representation from surrounding-view cameras is of great importance for autonomous driving. In this work, we propose a Geometry-guided Kernel Transformer (GKT), a novel 2D-to-BEV representation learning mechanism. GKT leverages the geometric priors to guide the transformer to focus on discriminative regions and unfolds kernel features to generate BEV representation. For fast inference, we further introduce a look-up table (LUT) indexing method to get rid of the camera's calibrated parameters at runtime. GKT can run at $72.3$ FPS on 3090 GPU / $45.6$ FPS on 2080ti GPU and is robust to the camera deviation and the predefined BEV height. And GKT achieves the state-of-the-art real-time segmentation results, i.e., 38.0 mIoU (100m$\times$100m perception range at a 0.5m resolution) on the nuScenes val set. Given the efficiency, effectiveness, and robustness, GKT has great practical values in autopilot scenarios, especially for real-time running systems. Code and models will be available at \url{https://github.com/hustvl/GKT}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mapping like a Skeptic: Probabilistic BEV Projection for Online HD Mapping

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A probabilistic BEV projection with learned offsets and confidence-based temporal fusion improves online HD map accuracy and generalization.

  2. Coherent Online Road Topology Estimation and Reasoning with Standard-Definition Maps

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Score jointly detects lane segments, road boundaries, and traffic elements, estimates lane topology, and associates traffic elements with lanes, using SD map priors and temporal fusion to reach state-of-the-art on Ope...

  3. RTMap: Real-Time Recursive Mapping with Change Detection and Localization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    RTMap combines online HD mapping, map-based localization, and change detection in one model, improving the map over repeated traversals via probabilistic crowdsourced fusion.

  4. SafeMap: Robust HD Map Construction from Incomplete Observations

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SafeMap improves HD map construction accuracy under missing camera views by reconstructing the missing perspective features with Gaussian-sampled attention and correcting the BEV features through distillation.

  5. Learning to Generate Vectorized Maps at Intersections with Multiple Roadside Cameras

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MRC-VMap learns to generate vectorized intersection maps directly from four roadside camera images without explicit camera calibration.

  6. What Really Matters for Robust Multi-Sensor HD Map Construction?

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Combining data augmentation, cross-modal attention fusion, and modality dropout training improves robustness of camera-LiDAR HD map construction under 13 synthetic sensor corruptions and raises clean nuScenes mAP to 77.0.

  7. MapFusion: A Novel BEV Feature Fusion Network for Multi-modal Map Construction

    cs.CV 2025-02 conditional novelty 4.0 of 10

    A plug-in attention-and-gating fusion method improves multi-modal HD map and BEV map construction by a few points on nuScenes and Argoverse2.

Pith tools