Pith. sign in

REVIEW 10 cited by

MapTR: Structured Modeling and Learning for Online Vectorized HD Map Construction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.14437 v2 pith:WYQOKJA6 submitted 2022-08-30 cs.CV cs.RO

classification cs.CVcs.RO
keywords maptrconstructiondrivingachieveselementexistinghigherlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

High-definition (HD) map provides abundant and precise environmental information of the driving scene, serving as a fundamental and indispensable component for planning in autonomous driving system. We present MapTR, a structured end-to-end Transformer for efficient online vectorized HD map construction. We propose a unified permutation-equivalent modeling approach, i.e., modeling map element as a point set with a group of equivalent permutations, which accurately describes the shape of map element and stabilizes the learning process. We design a hierarchical query embedding scheme to flexibly encode structured map information and perform hierarchical bipartite matching for map element learning. MapTR achieves the best performance and efficiency with only camera input among existing vectorized map construction approaches on nuScenes dataset. In particular, MapTR-nano runs at real-time inference speed ($25.1$ FPS) on RTX 3090, $8\times$ faster than the existing state-of-the-art camera-based method while achieving $5.0$ higher mAP. Even compared with the existing state-of-the-art multi-modality method, MapTR-nano achieves $0.7$ higher mAP, and MapTR-tiny achieves $13.5$ higher mAP and $3\times$ faster inference speed. Abundant qualitative results show that MapTR maintains stable and robust map construction quality in complex and various driving scenes. MapTR is of great application value in autonomous driving. Code and more demos are available at \url{https://github.com/hustvl/MapTR}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TRIG: Trajectory-Rig Decoupled Metric Geometry Learning

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    TRIG factorizes multi-camera poses into ego-trajectory and static rig geometry, with decoupled supervision and sparse temporal-spatial attention, claiming SOTA metric depth, pose, and 3D reconstruction on five driving...

  2. End-to-End Generation of City-Scale Vectorized Maps by Crowdsourced Vehicles

    cs.RO 2025-07 conditional novelty 6.0 of 10

    Fusing vectorized map elements from multiple crowdsourced vehicles with a trip-aware transformer improves online HD map accuracy over single-vehicle methods on the Navinfo Dataset.

  3. S2GO: Streaming Sparse Gaussian Occupancy Prediction

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A sparse-query, streaming Gaussian occupancy predictor achieves state-of-the-art 3D semantic occupancy on nuScenes and KITTI with real-time inference.

  4. MapKD: Unlocking Prior Knowledge with Cross-Modal Distillation for Efficient Online HD Map Construction

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A teacher-coach-student distillation framework transfers lidar and map-prior knowledge into a camera-only HD map model, improving nuScenes mIoU by 6.68 and mAP by 10.94 over the student baseline.

  5. PriorFusion: Unified Integration of Priors for Robust Road Perception in Autonomous Driving

    cs.CV 2025-07 conditional novelty 5.0 of 10

    PriorFusion integrates semantic segmentation, SVD-based shape templates, and a truncated diffusion decoder to improve vectorized road element perception, reporting state-of-the-art mAP on nuScenes.

  6. GTAD: Global Temporal Aggregation Denoising Learning for 3D Semantic Occupancy Prediction

    cs.CV 2025-07 conditional novelty 5.0 of 10

    GTAD combines an in-model latent denoising network with global temporal interaction to improve camera-based 3D semantic occupancy prediction, reporting 40.76 mIoU on Occ3D-nuScenes at 12 epochs.

  7. Reusing Attention for One-stage Lane Topology Understanding

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A one-stage transformer with attention reuse predicts lane and traffic-element topology directly, improving accuracy and speed on OpenLane-V2.

  8. World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model

    cs.CV 2025-07 conditional novelty 5.0 of 10

    World4Drive couples multiple driving intentions with a latent world model to generate, score, and select trajectories, reporting state-of-the-art perception-free planning on nuScenes and NavSim.

  9. DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A synthetic dataset and visual prompting framework improve VLM-based driving risk prediction, with the main evidence coming from an unreleased private test set.

  10. Learning to Generate Vectorized Maps at Intersections with Multiple Roadside Cameras

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MRC-VMap learns to generate vectorized intersection maps directly from four roadside camera images without explicit camera calibration.

Pith tools