Pith. sign in

REVIEW 16 cited by

BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.17270 v2 pith:SEM36NV6 submitted 2022-03-31 cs.CV

classification cs.CV
keywords bevformerspatialinformationtemporalautonomousdrivingimagesmulti-camera
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

3D visual perception tasks, including 3D detection and map segmentation based on multi-camera images, are essential for autonomous driving systems. In this work, we present a new framework termed BEVFormer, which learns unified BEV representations with spatiotemporal transformers to support multiple autonomous driving perception tasks. In a nutshell, BEVFormer exploits both spatial and temporal information by interacting with spatial and temporal space through predefined grid-shaped BEV queries. To aggregate spatial information, we design spatial cross-attention that each BEV query extracts the spatial features from the regions of interest across camera views. For temporal information, we propose temporal self-attention to recurrently fuse the history BEV information. Our approach achieves the new state-of-the-art 56.9\% in terms of NDS metric on the nuScenes \texttt{test} set, which is 9.0 points higher than previous best arts and on par with the performance of LiDAR-based baselines. We further show that BEVFormer remarkably improves the accuracy of velocity estimation and recall of objects under low visibility conditions. The code is available at \url{https://github.com/zhiqi-li/BEVFormer}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Factorized Dense Routing approximates unconstrained 2D-to-3D feature mixing by hierarchical tensor contractions, yielding global-context occupancy prediction that remains robust without camera extrinsics.

  2. VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    VISA improves closed-set 3D occupancy mIoU on nuScenes by using VLM instance audits as reliability-weighted semantic supervisors during training of existing world models.

  3. 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection

    cs.CV 2025-07 conditional novelty 7.0 of 10

    3D-MOOD is the first end-to-end monocular 3D object detector for open-set classes and novel scenes, achieving SOTA on Omni3D and on new open-set benchmarks.

  4. RayOcc: Occlusion-Aware Ray Occupancy Estimation via Gaussian Mixture Intensity

    cs.CV 2026-07 conditional novelty 6.0 of 10

    RayOcc models each camera ray as a non-normalized Gaussian mixture with Poisson-based occupancy probabilities, allowing multiple depth hypotheses per ray and improving Gaussian-initialized 3D occupancy prediction on nuScenes.

  5. Semantic Causality-Aware Vision-Based 3D Occupancy Prediction

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A class-conditional gradient loss (Causal Loss) plus channel-grouped lifting, learnable camera offsets, and normalized convolution raises Occ3D mIoU by 1.2/0.8 points and cuts the camera-noise mIoU drop from 32% to 7%.

  6. S2GO: Streaming Sparse Gaussian Occupancy Prediction

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A sparse-query, streaming Gaussian occupancy predictor achieves state-of-the-art 3D semantic occupancy on nuScenes and KITTI with real-time inference.

  7. Panoramic Scene Understanding: A Survey from Distortion-Aware Engineering to Sphere-Native Modeling

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    Survey organizing panoramic scene analysis literature by architectural design and training paradigm, identifying the absence of methods achieving both strict spherical equivariance and full reuse of perspective-pretra...

  8. NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    OmniDreams turns the Cosmos video-diffusion model into a real-time, action-conditioned driving simulator and shows that its internal representations can be fine-tuned into a competitive driving policy.

  9. UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

    cs.CV 2026-01 conditional novelty 5.0 of 10

    A unified VLM for autonomous driving that couples trajectory planning with future-frame image generation improves open- and closed-loop planning metrics on Bench2Drive and nuScenes.

  10. Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models

    cs.LG 2025-09 conditional novelty 5.0 of 10

    Multi-car bird's-eye view tokens injected into a frozen vision-language model improve simulated V2I link prediction accuracy by up to 13.9 points on average.

  11. Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey

    cs.RO 2025-08 conditional novelty 5.0 of 10

    A safety-critical survey that organizes BEV perception into single-modality, multimodal, and collaborative stages and consolidates robustness evidence that multimodal fusion degrades far less than single-modality perc...

  12. Learning to Generate Vectorized Maps at Intersections with Multiple Roadside Cameras

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MRC-VMap learns to generate vectorized intersection maps directly from four roadside camera images without explicit camera calibration.

  13. SHTOcc: Effective 3D Occupancy Prediction with Sparse Head and Tail Voxels

    cs.CV 2025-05 reject novelty 5.0 of 10

    SHTOcc combines attention-based sparse voxel selection with decoupled classifier retraining for 3D occupancy prediction, reporting efficiency gains and small, partly inconsistent accuracy improvements.

  14. CogAD: Cognitive-Hierarchy Guided End-to-End Autonomous Driving

    cs.RO 2025-05 conditional novelty 5.0 of 10

    CogAD reports state-of-the-art open-loop and closed-loop planning results by combining hierarchical scene-to-instance perception with intent-to-trajectory planning and dual-level uncertainty.

  15. DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion

    cs.CV 2025-10 conditional novelty 4.0 of 10

    DriveGen3D makes long driving-video synthesis and 3D scene reconstruction practical by caching only the conditional diffusion branch, quantizing cross-view attention, and fusing temporal context into a feed-forward Ga...

  16. SKGE-SWIN: End-To-End Autonomous Vehicle Waypoint Prediction and Navigation Using Skip Stage Swin Transformer

    cs.CV 2025-08 reject novelty 4.0 of 10

    Adding skip connections between Swin Transformer stages raises the CARLA Driving Score of the authors' end-to-end driving model from 29.7 (x13 CNN baseline) to 37.1 on Town05, in a single reported evaluation run.

Pith tools