Pith. sign in

REVIEW 2 cited by

MapNeXt: Revisiting Training and Scaling Practices for Online Vectorized HD Map Construction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.07323 v1 pith:ACOPT4HA submitted 2024-01-14 cs.CV

classification cs.CV
keywords maptrconstructionmodelperformancescalingtrainingarchitecturemapnext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

High-Definition (HD) maps are pivotal to autopilot navigation. Integrating the capability of lightweight HD map construction at runtime into a self-driving system recently emerges as a promising direction. In this surge, vision-only perception stands out, as a camera rig can still perceive the stereo information, let alone its appealing signature of portability and economy. The latest MapTR architecture solves the online HD map construction task in an end-to-end fashion but its potential is yet to be explored. In this work, we present a full-scale upgrade of MapTR and propose MapNeXt, the next generation of HD map learning architecture, delivering major contributions from the model training and scaling perspectives. After shedding light on the training dynamics of MapTR and exploiting the supervision from map elements thoroughly, MapNeXt-Tiny raises the mAP of MapTR-Tiny from 49.0% to 54.8%, without any architectural modifications. Enjoying the fruit of map segmentation pre-training, MapNeXt-Base further lifts the mAP up to 63.9% that has already outperformed the prior art, a multi-modality MapTR, by 1.4% while being $\sim1.8\times$ faster. Towards pushing the performance frontier to the next level, we draw two conclusions on practical model scaling: increased query favors a larger decoder network for adequate digestion; a large backbone steadily promotes the final accuracy without bells and whistles. Building upon these two rules of thumb, MapNeXt-Huge achieves state-of-the-art performance on the challenging nuScenes benchmark. Specifically, we push the mapless vision-only single-model performance to be over 78% for the first time, exceeding the best model from existing methods by 16%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SafeMap: Robust HD Map Construction from Incomplete Observations

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SafeMap improves HD map construction accuracy under missing camera views by reconstructing the missing perspective features with Gaussian-sampled attention and correcting the BEV features through distillation.

  2. MapExpert: Online HD Map Construction with Simple and Efficient Sparse Map Element Expert

    cs.CV 2024-12 conditional novelty 5.0 of 10

    MapExpert uses shape-specific sparse expert networks and a learnable temporal fusion module to improve online HD map construction by about 1.4-1.8 mAP over MapTracker on nuScenes and Argoverse2.

Pith tools