REVIEW 7 cited by
Efficient and Robust 2D-to-BEV Representation Learning via Geometry-guided Kernel Transformer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Learning Bird's Eye View (BEV) representation from surrounding-view cameras is of great importance for autonomous driving. In this work, we propose a Geometry-guided Kernel Transformer (GKT), a novel 2D-to-BEV representation learning mechanism. GKT leverages the geometric priors to guide the transformer to focus on discriminative regions and unfolds kernel features to generate BEV representation. For fast inference, we further introduce a look-up table (LUT) indexing method to get rid of the camera's calibrated parameters at runtime. GKT can run at $72.3$ FPS on 3090 GPU / $45.6$ FPS on 2080ti GPU and is robust to the camera deviation and the predefined BEV height. And GKT achieves the state-of-the-art real-time segmentation results, i.e., 38.0 mIoU (100m$\times$100m perception range at a 0.5m resolution) on the nuScenes val set. Given the efficiency, effectiveness, and robustness, GKT has great practical values in autopilot scenarios, especially for real-time running systems. Code and models will be available at \url{https://github.com/hustvl/GKT}.
Forward citations
Cited by 7 Pith papers
-
Mapping like a Skeptic: Probabilistic BEV Projection for Online HD Mapping
A probabilistic BEV projection with learned offsets and confidence-based temporal fusion improves online HD map accuracy and generalization.
-
Coherent Online Road Topology Estimation and Reasoning with Standard-Definition Maps
Score jointly detects lane segments, road boundaries, and traffic elements, estimates lane topology, and associates traffic elements with lanes, using SD map priors and temporal fusion to reach state-of-the-art on Ope...
-
RTMap: Real-Time Recursive Mapping with Change Detection and Localization
RTMap combines online HD mapping, map-based localization, and change detection in one model, improving the map over repeated traversals via probabilistic crowdsourced fusion.
-
SafeMap: Robust HD Map Construction from Incomplete Observations
SafeMap improves HD map construction accuracy under missing camera views by reconstructing the missing perspective features with Gaussian-sampled attention and correcting the BEV features through distillation.
-
Learning to Generate Vectorized Maps at Intersections with Multiple Roadside Cameras
MRC-VMap learns to generate vectorized intersection maps directly from four roadside camera images without explicit camera calibration.
-
What Really Matters for Robust Multi-Sensor HD Map Construction?
Combining data augmentation, cross-modal attention fusion, and modality dropout training improves robustness of camera-LiDAR HD map construction under 13 synthetic sensor corruptions and raises clean nuScenes mAP to 77.0.
-
MapFusion: A Novel BEV Feature Fusion Network for Multi-modal Map Construction
A plug-in attention-and-gating fusion method improves multi-modal HD map and BEV map construction by a few points on nuScenes and Argoverse2.
Discussion (0). Continue with ORCID to comment.