REVIEW 3 cited by
M$^2$BEV: Multi-Camera Joint 3D Detection and Segmentation with Unified Birds-Eye View Representation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
In this paper, we propose M$^2$BEV, a unified framework that jointly performs 3D object detection and map segmentation in the Birds Eye View~(BEV) space with multi-camera image inputs. Unlike the majority of previous works which separately process detection and segmentation, M$^2$BEV infers both tasks with a unified model and improves efficiency. M$^2$BEV efficiently transforms multi-view 2D image features into the 3D BEV feature in ego-car coordinates. Such BEV representation is important as it enables different tasks to share a single encoder. Our framework further contains four important designs that benefit both accuracy and efficiency: (1) An efficient BEV encoder design that reduces the spatial dimension of a voxel feature map. (2) A dynamic box assignment strategy that uses learning-to-match to assign ground-truth 3D boxes with anchors. (3) A BEV centerness re-weighting that reinforces with larger weights for more distant predictions, and (4) Large-scale 2D detection pre-training and auxiliary supervision. We show that these designs significantly benefit the ill-posed camera-based 3D perception tasks where depth information is missing. M$^2$BEV is memory efficient, allowing significantly higher resolution images as input, with faster inference speed. Experiments on nuScenes show that M$^2$BEV achieves state-of-the-art results in both 3D object detection and BEV segmentation, with the best single model achieving 42.5 mAP and 57.0 mIoU in these two tasks, respectively.
Forward citations
Cited by 3 Pith papers
-
Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints
OVBEVSeg produces open-vocabulary bird's-eye-view semantic maps on nuScenes by projecting CLIP labels through 3D detections, constraining Gaussian splats with BEV occupancy, and distilling the geometry into a real-tim...
-
Collaborative Perceiver: Elevating Vision-based 3D Object Detection via Local Density-Aware Spatial Occupancy
A multi-task camera model that adds local-density-aware occupancy prediction to 3D object detection reports strong nuScenes scores, but internal inconsistencies and missing code prevent confirmation.
-
MATS: A novel multi-modality multi-task learning framework for 3D perception in autonomous driving
Multiple simple adaptive BEV fusion branches plus task-gated Mixture-of-Experts cut multi-task negative transfer on nuScenes camera–LiDAR detection and map segmentation.
Discussion (0). Continue with ORCID to comment.