Pith. sign in

REVIEW 2 minor 52 references

GaussianMap: Learning Gaussian Representation for Multi-Sensor Online HD Map Construction

T0 review · 0 major / 2 minor · reviewed 2026-07-01 · grok-4.3

Pith's one-line read GaussianMap replaces fixed BEV grids with learned Gaussian primitives that adaptively represent sparse map elements for online HD map construction.

desk verdict Gaussian primitives replace fixed BEV grids for map construction and the paper claims SOTA on nuScenes and Argoverse 2, but the abstract supplies no numbers or ablations to check the gains. read the letter →

arxiv 2606.31177 v1 pith:EBGOUG4B submitted 2026-06-30 cs.CV

classification cs.CV
keywords HDmapconstructionGaussianrepresentationBEVfeaturesonlinemappingmulti-sensorfusionvectorizedpredictionautonomousdriving
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Existing online HD map methods encode scenes with uniform high-resolution BEV grids, which waste capacity on empty space while struggling with the fine localization needed for vectorized road elements. GaussianMap instead maintains a set of adaptive Gaussian primitives on the BEV plane, each carrying its own geometric parameters and feature vector, so representational effort concentrates only where map elements exist. A feed-forward encoder refines these primitives by modeling their interactions and fusing camera and LiDAR observations, after which the primitives are splatted into a BEV feature map for final vectorized decoding. Experiments on nuScenes and Argoverse 2 show this approach reaches state-of-the-art accuracy in both camera-only and camera-LiDAR settings.

What carries the argument

A collection of learnable Gaussian primitives on the BEV plane, each defined by geometric parameters and a feature vector, that are refined by a feed-forward encoder and then splatted into a BEV feature map.

What would settle it

An experiment in which GaussianMap produces higher localization error on map element vertices than a comparable fixed-grid baseline on the same nuScenes or Argoverse 2 validation splits.

Watch

Extended reading notes

Core claim

The paper introduces a Gaussian representation consisting of a modest number of learnable primitives on the BEV plane; each primitive encodes a flexible local region through explicit geometric properties together with an attached feature vector. A Gaussian encoder progressively refines the set via interaction modeling and multi-sensor aggregation, after which the primitives are splatted into a dense BEV feature map that is decoded into vectorized map elements.

Load-bearing premise

A modest number of learned Gaussian primitives can capture the fine-grained geometry and localization of map elements at least as accurately as fixed high-resolution grids.

Editorial extensions

If this is right

  • Representational capacity is allocated only to map-relevant regions instead of uniform grids.
  • The same Gaussian encoder supports both camera-only and camera-LiDAR fusion inputs.
  • The splatted BEV feature map can be decoded into vectorized predictions without additional post-processing stages.
  • Performance gains appear on standard autonomous-driving map benchmarks without requiring extra sensor modalities.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same adaptive primitive mechanism could be applied to other sparse scene-understanding tasks such as lane detection or traffic-sign localization.
  • Because the primitives carry explicit geometric parameters, they may enable direct metric evaluation of map accuracy without rasterization.
  • Reducing the number of primitives further could trade accuracy for lower memory and faster inference on embedded hardware.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 2 minor

Summary. The manuscript introduces GaussianMap, an online HD map construction method that learns an adaptive set of Gaussian primitives on the BEV plane to represent the scene, each encoding local geometry and features. A feed-forward Gaussian encoder refines the primitives via interaction modeling and multi-sensor (camera/LiDAR) feature aggregation; the primitives are then splatted to a BEV feature map and decoded into vectorized map elements. The central claim is that this representation is more efficient than fixed-resolution BEV grids for sparse map elements and yields state-of-the-art results on nuScenes and Argoverse 2 in both camera-only and camera-LiDAR fusion settings.

Significance. If the reported performance gains hold under rigorous evaluation, the shift from dense fixed grids to adaptive Gaussian primitives could improve representational efficiency for HD mapping tasks. The explicit commitment to public code release supports reproducibility, which is a strength for this line of work.

minor comments (2)
  1. The abstract states SOTA results but does not reference specific quantitative tables or ablation studies; adding a sentence pointing to the main result table would improve readability for readers scanning the abstract.
  2. Notation for the Gaussian primitives (position, covariance, feature vector) should be introduced with explicit symbols in the method section to avoid ambiguity when describing the splatting operation.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for the positive review and recommendation of minor revision. The recognition of the efficiency advantages of adaptive Gaussian primitives over fixed BEV grids for sparse map elements is appreciated.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity detected

full rationale

The paper proposes a feed-forward Gaussian encoder that produces adaptive primitives on the BEV plane, followed by splatting and vectorized decoding. All performance claims rest on empirical results from external benchmarks (nuScenes and Argoverse 2) rather than any internal derivation that reduces to fitted inputs or self-citations. No equations, uniqueness theorems, or ansatzes are shown that loop back to the method's own definitions or prior author work; the architecture is presented as a standard learned model whose validity is tested outside the training distribution.

Assumptions & free parameters 0 free parameters · 1 assumptions · 1 invented entities

Based solely on the abstract, the central claim rests on the domain assumption that map elements are spatially sparse and that adaptive Gaussians can allocate capacity more effectively than uniform grids; no free parameters or invented entities beyond the Gaussian primitives themselves are described.

assumptions (1)
  • domain assumption Map elements are spatially sparse yet require fine-grained geometric localization, rendering uniform BEV grids redundant.
    Stated in the second sentence of the abstract as the motivation for moving away from fixed-resolution grids.
invented entities (1)
  • Gaussian primitives
    purpose: Adaptive local region encoding on the BEV plane with geometric properties and feature vectors
    Introduced as the core scene representation; no independent evidence outside the paper is provided in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of GaussianMap: Learning Gaussian Representation for Multi-Sensor Online HD Map Construction." pith.science (2026). https://pith.science/paper/EBGOUG4B

@misc{pith2026260631177,
  author       = {Pith},
  title        = {Pith review of: GaussianMap: Learning Gaussian Representation for Multi-Sensor Online HD Map Construction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EBGOUG4B}},
  note         = {Machine review of arXiv:2606.31177}
}
read the original abstract

Autonomous driving systems benefit from high-definition (HD) maps that provide critical information about road infrastructure. The online construction of HD maps offers a scalable approach to generate local vectorized maps from onboard sensor observations. Existing methods commonly adopt bird's-eye-view (BEV) features as the intermediate scene representation, encoding the surrounding space with fixed-resolution dense grids. However, map elements are spatially sparse yet require fine-grained geometric localization, making uniformly allocated BEV representations redundant and less effective for vectorized map prediction. In this work, we propose GaussianMap, an online HD map construction framework that learns an adaptive Gaussian representation of the surrounding scene. This representation consists of a set of Gaussian primitives on the BEV plane, each encoding a flexible local region with geometric properties and a feature vector, allowing the model to allocate representational capacity to map-relevant regions. To generate such a representation from sensor observations, we introduce a feed-forward Gaussian encoder that progressively refines these primitives through Gaussian interaction modeling and multi-sensor feature aggregation. The refined Gaussian representation is then splatted into a BEV feature map and decoded into vectorized map predictions. Extensive experiments on nuScenes and Argoverse 2 datasets demonstrate that GaussianMap achieves state-of-the-art performance in both camera-only and camera-LiDAR fusion settings. Our code will be made publicly available.

Figures

Figures reproduced from arXiv: 2606.31177 by the authors.

Figure 1
Figure 1. Motivation of GaussianMap. Compared with grid-based BEV features that uniformly allocate representational capacity, GaussianMap represents the surrounding scene with adaptive Gaus￾sian primitives to focus on map-relevant regions and perseve fine￾grained geometry. Gaussian colors are used only for visualization. features as the intermediate scene representation. Although BEV features provide a structured top-down spa… view at source ↗
Figure 2
Figure 2. Overall framework of GaussianMap. The framework takes multi-camera images and an optional LiDAR point cloud as inputs. The Gaussian encoder generates an adaptive Gaussian representation on the BEV plane from the extracted sensor features. This representation is converted into a BEV feature map via Gaussian-to-BEV splatting, from which the map decoder produces the vectorized map prediction. For clarity, we visualize … view at source ↗
Figure 3
Figure 3. Qualitative results on the nuScenes validation set. (a) Camera-only setting. (b) Camera-LiDAR fusion setting. We compare the vectorized map predictions by GaussianMap with those from MapQR and MapQR+DAMap, using the ground truth as the reference. For GaussianMap, we additionally visualize the Gaussians alongside the predictions, retaining those with opacity higher than 0.4 for clarity. interaction modeling and multi… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

52 extracted references · 52 canonical work pages

  1. [1]

    Online high-definition map construction for autonomous vehicles: A comprehensive survey,

    H. Lyu, J. S. Berrio Perez, Y . Huang, K. Li, M. Shan, and S. Worrall, “Online high-definition map construction for autonomous vehicles: A comprehensive survey,”Journal of Sensor and Actuator Networks, vol. 14, no. 1, p. 15, 2025

  2. [2]

    High-definition maps: Comprehensive survey, challenges, and future perspectives,

    G. Elghazaly, R. Frank, S. Harvey, and S. Safko, “High-definition maps: Comprehensive survey, challenges, and future perspectives,” IEEE Open Journal of Intelligent Transportation Systems, vol. 4, pp. 527–550, 2023

  3. [3]

    Loam: Lidar odometry and mapping in real-time

    J. Zhang, S. Singhet al., “Loam: Lidar odometry and mapping in real-time.” inRobotics: Science and systems, vol. 2, no. 9. Berkeley, CA, 2014, pp. 1–9

  4. [4]

    Lego-loam: Lightweight and ground- optimized lidar odometry and mapping on variable terrain,

    T. Shan and B. Englot, “Lego-loam: Lightweight and ground- optimized lidar odometry and mapping on variable terrain,” in2018 IEEE/RSJ international conference on intelligent robots and systems (IROS). IEEE, 2018, pp. 4758–4765

  5. [5]

    Long-term map maintenance pipeline for autonomous vehicles,

    J. S. Berrio, S. Worrall, M. Shan, and E. Nebot, “Long-term map maintenance pipeline for autonomous vehicles,”IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 8, pp. 10 427–10 440, 2021

  6. [6]

    Hdmapnet: An online hd map construction and evaluation framework,

    Q. Li, Y . Wang, Y . Wang, and H. Zhao, “Hdmapnet: An online hd map construction and evaluation framework,” inInternational Conference on Robotics and Automation (ICRA). IEEE, 2022, pp. 4628–4634

  7. [7]

    Vectormapnet: End-to-end vectorized hd map learning,

    Y . Liu, T. Yuan, Y . Wang, Y . Wang, and H. Zhao, “Vectormapnet: End-to-end vectorized hd map learning,” inInternational Conference on Machine Learning. PMLR, 2023, pp. 22 352–22 369

  8. [8]

    Maptr: Structured modeling and learning for online vectorized hd map construction,

    B. Liao, S. Chen, X. Wang, T. Cheng, Q. Zhang, W. Liu, and C. Huang, “Maptr: Structured modeling and learning for online vectorized hd map construction,” inThe Eleventh International Conference on Learning Representations, 2023

Show all 52 references
  1. [9]

    Maprf: Weakly supervised online hd map construction via nerf-guided self-training,

    H. Lyu, T. Monninger, J. S. B. Perez, M. Shan, Z. Ming, and S. Worrall, “Maprf: Weakly supervised online hd map construction via nerf-guided self-training,” inIEEE International Conference on Intelligent Transportation Systems (ITSC). IEEE, 2026

  2. [10]

    Maptrv2: An end-to-end framework for online vectorized hd map construction,

    B. Liao, S. Chen, Y . Zhang, B. Jiang, Q. Zhang, W. Liu, C. Huang, and X. Wang, “Maptrv2: An end-to-end framework for online vectorized hd map construction,”International Journal of Computer Vision, vol. 133, no. 3, pp. 1352–1374, 2025

  3. [11]

    Leveraging enhanced queries of point sets for vectorized map construction,

    Z. Liu, X. Zhang, G. Liu, J. Zhao, and N. Xu, “Leveraging enhanced queries of point sets for vectorized map construction,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 461–477

  4. [12]

    Damap: Distance-aware mapnet for high quality hd map construction,

    J. Dong, C. Li, Y . Lin, J. Fu, S. Zhou, and N. Zheng, “Damap: Distance-aware mapnet for high quality hd map construction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 5285–5294

  5. [13]

    3d gaussian splatting for real-time radiance field rendering

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, G. Drettakiset al., “3d gaussian splatting for real-time radiance field rendering.”ACM Trans. Graph., vol. 42, no. 4, pp. 139–1, 2023

  6. [14]

    Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,

    J. Philion and S. Fidler, “Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,” inEuropean conference on computer vision. Springer, 2020, pp. 194–210

  7. [15]

    Bev- former: learning bird’s-eye-view representation from lidar-camera via spatiotemporal transformers,

    Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Q. Yu, and J. Dai, “Bev- former: learning bird’s-eye-view representation from lidar-camera via spatiotemporal transformers,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 3, pp. 2020–2036, 2024

  8. [16]

    Bevfusion: Multi-task multi-sensor fusion with unified bird’s- eye view representation,

    Z. Liu, H. Tang, A. Amini, X. Yang, H. Mao, D. L. Rus, and S. Han, “Bevfusion: Multi-task multi-sensor fusion with unified bird’s- eye view representation,” in2023 IEEE international conference on robotics and automation (ICRA). IEEE, 2023, pp. 2774–2781

  9. [17]

    End-to-end vectorized hd- map construction with piecewise bezier curve,

    L. Qiao, W. Ding, X. Qiu, and C. Zhang, “End-to-end vectorized hd- map construction with piecewise bezier curve,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 13 218–13 228

  10. [18]

    Pivotnet: Vectorized pivot learning for end-to-end hd map construction,

    W. Ding, L. Qiao, X. Qiu, and C. Zhang, “Pivotnet: Vectorized pivot learning for end-to-end hd map construction,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 3672–3682

  11. [19]

    Streammapnet: Streaming mapping network for vectorized online hd map construc- tion,

    T. Yuan, Y . Liu, Y . Wang, Y . Wang, and H. Zhao, “Streammapnet: Streaming mapping network for vectorized online hd map construc- tion,” inProceedings of the IEEE/CVF Winter Conference on Appli- cations of Computer Vision, 2024, pp. 7356–7365

  12. [20]

    Mgmap: Mask-guided learning for online vectorized hd map construction,

    X. Liu, S. Wang, W. Li, R. Yang, J. Chen, and J. Zhu, “Mgmap: Mask-guided learning for online vectorized hd map construction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 14 812–14 821

  13. [21]

    Refdiffmap: Diffusion-guided progressive refinement for vectorized hd map con- struction,

    W. Gao, E. Chang, J. Fu, Z. Zhu, S. Chen, and N. Zheng, “Refdiffmap: Diffusion-guided progressive refinement for vectorized hd map con- struction,”IEEE Robotics and Automation Letters, vol. 11, no. 3, pp. 2554–2561, 2026

  14. [22]

    Admap: Anti-disturbance framework for vectorized hd map construction,

    H. Hu, F. Wang, Y . Wang, L. Hu, J. Xu, and Z. Zhang, “Admap: Anti-disturbance framework for vectorized hd map construction,” in European Conference on Computer Vision. Springer, 2024, pp. 311– 326

  15. [23]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoor- thi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,”Communications of the ACM, vol. 65, no. 1, pp. 99–106, 2021

  16. [24]

    4d gaussian splatting for real-time dynamic scene rendering,

    G. Wu, T. Yi, J. Fang, L. Xie, X. Zhang, W. Wei, W. Liu, Q. Tian, and X. Wang, “4d gaussian splatting for real-time dynamic scene rendering,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 20 310–20 320

  17. [25]

    Dreamgaussian: Generative gaussian splatting for efficient 3d content creation,

    J. Tang, J. Ren, H. Zhou, Z. Liu, and G. Zeng, “Dreamgaussian: Generative gaussian splatting for efficient 3d content creation,” in International Conference on Learning Representations, vol. 2024, 2024, pp. 33 879–33 896

  18. [26]

    Langsplat: 3d language gaussian splatting,

    M. Qin, W. Li, J. Zhou, H. Wang, and H. Pfister, “Langsplat: 3d language gaussian splatting,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 051–20 060

  19. [27]

    Gaussianformer: Scene as gaussians for vision-based 3d semantic occupancy predic- tion,

    Y . Huang, W. Zheng, Y . Zhang, J. Zhou, and J. Lu, “Gaussianformer: Scene as gaussians for vision-based 3d semantic occupancy predic- tion,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 376–393

  20. [28]

    Gaussianformer-2: Probabilistic gaussian superposition for efficient 3d occupancy prediction,

    Y . Huang, A. Thammatadatrakoon, W. Zheng, Y . Zhang, D. Du, and J. Lu, “Gaussianformer-2: Probabilistic gaussian superposition for efficient 3d occupancy prediction,” inProceedings of the computer vision and pattern recognition conference, 2025, pp. 27 477–27 486

  21. [29]

    Gaussianbev: 3d gaussian representation meets perception models for bev segmentation,

    F. Chabot, N. Granger, and G. Lapouge, “Gaussianbev: 3d gaussian representation meets perception models for bev segmentation,” in2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). IEEE, 2025, pp. 2250–2259

  22. [30]

    Toward real-world bev per- ception: Depth uncertainty estimation via gaussian splatting,

    S.-W. Lu, Y .-H. Tsai, and Y .-T. Chen, “Toward real-world bev per- ception: Depth uncertainty estimation via gaussian splatting,” inPro- ceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 17 124–17 133

  23. [31]

    Gaussianad: Gaussian-centric end- to-end autonomous driving,

    W. Zheng, J. Wu, Y . Zheng, S. Zuo, Z. Xie, L. Yang, Y . Pan, Z. Hao, P. Jia, X. Langet al., “Gaussianad: Gaussian-centric end- to-end autonomous driving,”arXiv preprint arXiv:2412.10371, 2024

  24. [32]

    Pointpainting: Sequential fusion for 3d object detection,

    S. V ora, A. H. Lang, B. Helou, and O. Beijbom, “Pointpainting: Sequential fusion for 3d object detection,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 4604–4612

  25. [33]

    Pointaugmenting: Cross- modal augmentation for 3d object detection,

    C. Wang, C. Ma, M. Zhu, and X. Yang, “Pointaugmenting: Cross- modal augmentation for 3d object detection,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 11 794–11 803

  26. [34]

    Multimodal virtual point 3d detection,

    T. Yin, X. Zhou, and P. Kr ¨ahenb¨uhl, “Multimodal virtual point 3d detection,”Advances in Neural Information Processing Systems, vol. 34, pp. 16 494–16 507, 2021

  27. [35]

    Multi-view 3d object detection network for autonomous driving,

    X. Chen, H. Ma, J. Wan, B. Li, and T. Xia, “Multi-view 3d object detection network for autonomous driving,” inProceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2017, pp. 1907–1915

  28. [36]

    Centerfusion: Center-based radar and camera fusion for 3d object detection,

    R. Nabati and H. Qi, “Centerfusion: Center-based radar and camera fusion for 3d object detection,” inProceedings of the IEEE/CVF winter conference on applications of computer vision, 2021, pp. 1527–1536

  29. [37]

    Transfusion: Robust lidar-camera fusion for 3d object detection with transformers,

    X. Bai, Z. Hu, X. Zhu, Q. Huang, Y . Chen, H. Fu, and C.-L. Tai, “Transfusion: Robust lidar-camera fusion for 3d object detection with transformers,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 1090–1099

  30. [38]

    Bevfusion: A simple and robust lidar-camera fusion framework,

    T. Liang, H. Xie, K. Yu, Z. Xia, Z. Lin, Y . Wang, T. Tang, B. Wang, and Z. Tang, “Bevfusion: A simple and robust lidar-camera fusion framework,”Advances in neural information processing systems, vol. 35, pp. 10 421–10 434, 2022

  31. [39]

    Mapfusion: A novel bev feature fusion network for multi-modal map construction,

    X. Hao, Y . Diao, M. Wei, Y . Yang, P. Hao, R. Yin, H. Zhang, W. Li, S. Zhao, and Y . Liu, “Mapfusion: A novel bev feature fusion network for multi-modal map construction,”Information Fusion, vol. 119, p. 103018, 2025

  32. [40]

    What really matters for robust multi-sensor hd map construction?

    X. Hao, Y . Zhao, Y . Ji, L. Dai, P. Hao, D. Li, S. Cheng, and R. Yin, “What really matters for robust multi-sensor hd map construction?” in2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2025, pp. 1298–1304

  33. [41]

    Sef-map: Subspace-decomposed expert fusion for robust multimodal hd map prediction,

    H. Fu, L. Zhang, H. Li, R. Hu, Z. Li, G. Liu, Z. Tan, L. Chen, H. Ye, and X. Hao, “Sef-map: Subspace-decomposed expert fusion for robust multimodal hd map prediction,” inInternational Conference on Robotics and Automation (ICRA). IEEE, 2026

  34. [42]

    Deformable detr: Deformable transformers for end-to-end object detection,

    X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr: Deformable transformers for end-to-end object detection,” in International Conference on Learning Representations, 2021

  35. [43]

    Efficient and robust 2d-to-bev representation learning via geometry- guided kernel transformer,

    S. Chen, T. Cheng, X. Wang, W. Meng, Q. Zhang, and W. Liu, “Efficient and robust 2d-to-bev representation learning via geometry- guided kernel transformer,”arXiv preprint arXiv:2206.04584, 2022

  36. [44]

    nuscenes: A multimodal dataset for autonomous driving,

    H. Caesar, V . Bankiti, A. H. Lang, S. V ora, V . E. Liong, Q. Xu, A. Krishnan, Y . Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” inProceedings of the IEEE/CVF Conference on Computer vision and pattern recognition, 2020, pp. 11 621–11 631

  37. [45]

    Argoverse 2: Next generation datasets for self-driving perception and forecasting,

    B. Wilson, W. Qi, T. Agarwal, J. Lambert, J. Singh, S. Khandelwal, B. Pan, R. Kumar, A. Hartnett, J. K. Pontes, D. Ramanan, P. Carr, and J. Hays, “Argoverse 2: Next generation datasets for self-driving perception and forecasting,” inProceedings of the Neural Information Proces...

  38. [46]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” inProceedings of the IEEE Conference on Computer vision and pattern recognition, 2016, pp. 770–778

  39. [47]

    Second: Sparsely embedded convolutional detection,

    Y . Yan, Y . Mao, and B. Li, “Second: Sparsely embedded convolutional detection,”Sensors, vol. 18, no. 10, p. 3337, 2018

  40. [48]

    Online map vectorization for autonomous driving: A rasterization perspective,

    G. Zhang, J. Lin, S. Wu, Y . Song, Z. Luo, Y . Xue, S. Lu, and Z. Wang, “Online map vectorization for autonomous driving: A rasterization perspective,”Advances in Neural Information Processing Systems, 2023

  41. [49]

    Heightmapnet: Explicit height modeling for end-to-end hd map learning,

    W. Qiu, S. Pang, H. Zhang, J. Fang, and J. Xue, “Heightmapnet: Explicit height modeling for end-to-end hd map learning,” in2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). IEEE, 2025, pp. 6022–6031

  42. [50]

    Himap: Hybrid representation learning for end-to-end vectorized hd map construction,

    Y . Zhou, H. Zhang, J. Yu, Y . Yang, S. Jung, S.-I. Park, and B. Yoo, “Himap: Hybrid representation learning for end-to-end vectorized hd map construction,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 15 396–15 406

  43. [51]

    Relmap: Enhancing online map construction with class-aware spatial relation and semantic priors,

    T. Cai, Y . Zhang, Z. Zhou, Z. Huang, and J. Ma, “Relmap: Enhancing online map construction with class-aware spatial relation and semantic priors,” inInternational Conference on Robotics and Automation (ICRA). IEEE, 2026

  44. [52]

    Online vectorized hd map construction using geometry,

    Z. Zhang, Y . Zhang, X. Ding, F. Jin, and X. Yue, “Online vectorized hd map construction using geometry,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 73–90

Pith tools

Reviewed July 1, 2026 · model on record in the stance chip above.