Pith. sign in

REVIEW 6 cited by

GraphBEV: Towards Robust BEV Feature Alignment for Multi-Modal 3D Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.11848 v4 pith:GVK6SYHB submitted 2024-03-18 cs.CV

classification cs.CV
keywords cameragraphlidarfeaturesfusionmisalignmentaligndepth
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Integrating LiDAR and camera information into Bird's-Eye-View (BEV) representation has emerged as a crucial aspect of 3D object detection in autonomous driving. However, existing methods are susceptible to the inaccurate calibration relationship between LiDAR and the camera sensor. Such inaccuracies result in errors in depth estimation for the camera branch, ultimately causing misalignment between LiDAR and camera BEV features. In this work, we propose a robust fusion framework called Graph BEV. Addressing errors caused by inaccurate point cloud projection, we introduce a Local Align module that employs neighbor-aware depth features via Graph matching. Additionally, we propose a Global Align module to rectify the misalignment between LiDAR and camera BEV features. Our Graph BEV framework achieves state-of-the-art performance, with an mAP of 70.1\%, surpassing BEV Fusion by 1.6\% on the nuscenes validation set. Importantly, our Graph BEV outperforms BEV Fusion by 8.3\% under conditions with misalignment noise.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy Prediction

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SDGOCC improves multimodal 3D occupancy prediction by using LiDAR depth and semantic masks to guide camera-to-BEV transformation, achieving state-of-the-art mIoU on Occ3D-nuScenes.

  2. RCTrans: Radar-Camera Transformer via Radar Densifier and Sequential Decoder for 3D Object Detection

    cs.CV 2024-12 conditional novelty 6.0 of 10

    RCTrans achieves new state-of-the-art radar-camera 3D detection on nuScenes by densifying radar BEV features and using a pruning sequential decoder for query-based fusion.

  3. TiGDistill-BEV: Multi-view BEV 3D Object Detection via Target Inner-Geometry Learning Distillation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A LiDAR-to-camera distillation method that supervises relative depth inside object foregrounds and distills BEV feature relationships to boost camera-only 3D object detection.

  4. Timealign: A multi-modal object detection method for time misalignment fusing in autonomous driving

    cs.CV 2024-12 conditional novelty 4.0 of 10

    TimeAlign uses Swin-LSTM prediction and camera-guided combination of past and observed LiDAR BEV features to partially recover 3D detection accuracy under LiDAR time lag.

  5. Reflective Teacher: Semi-Supervised Multimodal 3D Object Detection in Bird's-Eye-View via Uncertainty Measure

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A semi-supervised teacher-student detector that applies a memory-aware regularizer and uncertainty weighting achieves near-full-supervised 3D detection accuracy with 22 to 25 percent of labels.

  6. CoreNet: Conflict Resolution Network for Point-Pixel Misalignment and Sub-Task Suppression of 3D LiDAR-Camera Object Detection

    cs.CV 2025-01

Pith tools