Pith. sign in

REVIEW 3 cited by

FusionFormer: A Multi-sensory Fusion in Bird's-Eye-View and Temporal Consistent Transformer for 3D Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.05257 v3 pith:OBXZSEUM submitted 2023-09-11 cs.CV

classification cs.CV
keywords fusionbirddetectionfeaturesobjectfeaturefusionformermethod
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Multi-sensor modal fusion has demonstrated strong advantages in 3D object detection tasks. However, existing methods that fuse multi-modal features require transforming features into the bird's eye view space and may lose certain information on Z-axis, thus leading to inferior performance. To this end, we propose a novel end-to-end multi-modal fusion transformer-based framework, dubbed FusionFormer, that incorporates deformable attention and residual structures within the fusion encoding module. Specifically, by developing a uniform sampling strategy, our method can easily sample from 2D image and 3D voxel features spontaneously, thus exploiting flexible adaptability and avoiding explicit transformation to the bird's eye view space during the feature concatenation process. We further implement a residual structure in our feature encoder to ensure the model's robustness in case of missing an input modality. Through extensive experiments on a popular autonomous driving benchmark dataset, nuScenes, our method achieves state-of-the-art single model performance of 72.6% mAP and 75.1% NDS in the 3D object detection task without test time augmentation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LiDAR-based End-to-end Temporal Perception for Vehicle-Infrastructure Cooperation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    LET-VIC is an end-to-end lidar framework for vehicle-infrastructure cooperative detection and tracking that fuses temporal and multi-view features and learns to compensate calibration errors, outperforming the tested ...

  2. EVT: Efficient View Transformation for Multi-Modal 3D Object Detection

    cs.CV 2024-11 conditional novelty 6.0 of 10

    EVT achieves state-of-the-art 75.3% NDS on the nuScenes test set by using LiDAR-guided adaptive sampling and projection for image-to-BEV transformation, along with a geometry-aware transformer decoder.

  3. Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey

    cs.RO 2025-08 conditional novelty 5.0 of 10

    A safety-critical survey that organizes BEV perception into single-modality, multimodal, and collaborative stages and consolidates robustness evidence that multimodal fusion degrades far less than single-modality perc...

Pith tools