Pith. sign in

REVIEW 3 cited by

M$^2$BEV: Multi-Camera Joint 3D Detection and Segmentation with Unified Birds-Eye View Representation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.05088 v2 pith:ICWJG52P submitted 2022-04-11 cs.CV

classification cs.CV
keywords detectionsegmentationtasksunifiedbenefitdesignsefficiencyefficient
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

In this paper, we propose M$^2$BEV, a unified framework that jointly performs 3D object detection and map segmentation in the Birds Eye View~(BEV) space with multi-camera image inputs. Unlike the majority of previous works which separately process detection and segmentation, M$^2$BEV infers both tasks with a unified model and improves efficiency. M$^2$BEV efficiently transforms multi-view 2D image features into the 3D BEV feature in ego-car coordinates. Such BEV representation is important as it enables different tasks to share a single encoder. Our framework further contains four important designs that benefit both accuracy and efficiency: (1) An efficient BEV encoder design that reduces the spatial dimension of a voxel feature map. (2) A dynamic box assignment strategy that uses learning-to-match to assign ground-truth 3D boxes with anchors. (3) A BEV centerness re-weighting that reinforces with larger weights for more distant predictions, and (4) Large-scale 2D detection pre-training and auxiliary supervision. We show that these designs significantly benefit the ill-posed camera-based 3D perception tasks where depth information is missing. M$^2$BEV is memory efficient, allowing significantly higher resolution images as input, with faster inference speed. Experiments on nuScenes show that M$^2$BEV achieves state-of-the-art results in both 3D object detection and BEV segmentation, with the best single model achieving 42.5 mAP and 57.0 mIoU in these two tasks, respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    OVBEVSeg produces open-vocabulary bird's-eye-view semantic maps on nuScenes by projecting CLIP labels through 3D detections, constraining Gaussian splats with BEV occupancy, and distilling the geometry into a real-tim...

  2. Collaborative Perceiver: Elevating Vision-based 3D Object Detection via Local Density-Aware Spatial Occupancy

    cs.CV 2025-07 reject novelty 5.0 of 10

    A multi-task camera model that adds local-density-aware occupancy prediction to 3D object detection reports strong nuScenes scores, but internal inconsistencies and missing code prevent confirmation.

  3. MATS: A novel multi-modality multi-task learning framework for 3D perception in autonomous driving

    cs.CV 2026-07 conditional novelty 4.0 of 10

    Multiple simple adaptive BEV fusion branches plus task-gated Mixture-of-Experts cut multi-task negative transfer on nuScenes camera–LiDAR detection and map segmentation.

Pith tools