Pith. sign in

REVIEW 20 cited by

FB-OCC: 3D Occupancy Prediction based on Forward-Backward View Transformation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.01492 v1 pith:65FNWAVO submitted 2023-07-04 cs.CV cs.RO

classification cs.CVcs.RO
keywords fb-bevoccupancypredictionworkshopautonomouschallengecvprdesigns
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This technical report summarizes the winning solution for the 3D Occupancy Prediction Challenge, which is held in conjunction with the CVPR 2023 Workshop on End-to-End Autonomous Driving and CVPR 23 Workshop on Vision-Centric Autonomous Driving Workshop. Our proposed solution FB-OCC builds upon FB-BEV, a cutting-edge camera-based bird's-eye view perception design using forward-backward projection. On top of FB-BEV, we further study novel designs and optimization tailored to the 3D occupancy prediction task, including joint depth-semantic pre-training, joint voxel-BEV representation, model scaling up, and effective post-processing strategies. These designs and optimization result in a state-of-the-art mIoU score of 54.19% on the nuScenes dataset, ranking the 1st place in the challenge track. Code and models will be released at: https://github.com/NVlabs/FB-BEV.

Discussion (0). Sign in to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Factorized Dense Routing approximates unconstrained 2D-to-3D feature mixing by hierarchical tensor contractions, yielding global-context occupancy prediction that remains robust without camera extrinsics.

  2. VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    VISA improves closed-set 3D occupancy mIoU on nuScenes by using VLM instance audits as reliability-weighted semantic supervisors during training of existing world models.

  3. SparseOcc++: Geometry-Aware Sparse Latent Representation for Semantic Occupancy Prediction

    cs.CV 2026-07 accept novelty 6.5 of 10

    SparseOcc++ decouples geometry completion (via orthogonal SCF regression on sparse anchors) from semantics, improving IoU 2.3 points and running 3.9 imes faster than SparseOcc on nuScenes while 5.9 imes faster than Oc...

  4. RayOcc: Occlusion-Aware Ray Occupancy Estimation via Gaussian Mixture Intensity

    cs.CV 2026-07 conditional novelty 6.0 of 10

    RayOcc models each camera ray as a non-normalized Gaussian mixture with Poisson-based occupancy probabilities, allowing multiple depth hypotheses per ray and improving Gaussian-initialized 3D occupancy prediction on nuScenes.

  5. OccTrack360: 4D Panoptic Occupancy Tracking from Surround-View Fisheye Cameras

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A new benchmark and baseline method for 4D panoptic occupancy tracking with surround-view fisheye cameras, built from KITTI-360.

  6. DVGT: Driving Visual Geometry Transformer

    cs.CV 2025-12 conditional novelty 6.0 of 10

    DVGT predicts metric-scaled global 3D point maps and ego poses from unposed multi-view driving video, beating prior geometry models on several driving benchmarks.

  7. Semantic Causality-Aware Vision-Based 3D Occupancy Prediction

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A class-conditional gradient loss (Causal Loss) plus channel-grouped lifting, learnable camera offsets, and normalized convolution raises Occ3D mIoU by 1.2/0.8 points and cuts the camera-noise mIoU drop from 32% to 7%.

  8. GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A camera-only Gaussian-surfel pipeline reconstructs full Waymo scenes, converts them to binary occupancy labels, and trains CVT-Occ to generalize on Occ3D-Waymo and Occ3D-nuScenes at a level close to or above LiDAR-la...

  9. SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy Prediction

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SDGOCC improves multimodal 3D occupancy prediction by using LiDAR depth and semantic masks to guide camera-to-BEV transformation, achieving state-of-the-art mIoU on Occ3D-nuScenes.

  10. Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SceneDINO performs semantic scene completion from a single image in a fully unsupervised way by lifting self-supervised DINO features into a 3D feature field trained with multi-view consistency.

  11. VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow Prediction

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A training-only Gaussian splatting loss, which renders predicted 3D semantics and motion into 2D camera views, improves semantic occupancy and scene flow prediction across several camera-based models.

  12. S2GO: Streaming Sparse Gaussian Occupancy Prediction

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A sparse-query, streaming Gaussian occupancy predictor achieves state-of-the-art 3D semantic occupancy on nuScenes and KITTI with real-time inference.

  13. GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A unified framework converts surface geometry priors into sparse Gaussian occupancy predictions and extends it to multi-view and temporal inputs.

  14. Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method

    cs.CV 2025-10 conditional novelty 5.0 of 10

    UniScenev2 scales occupancy-centric driving-scene generation to NuPlan scale, releasing a 3.6M-frame semantic-occupancy dataset and jointly generating occupancy, video, and LiDAR that beats published baselines on its ...

  15. Occupancy Learning with Spatiotemporal Memory

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    ST-Occ improves 3D occupancy prediction for self-driving by storing a compact scene-level memory of past frames and conditioning current predictions on it with uncertainty-aware attention, gaining 3 mIoU over prior st...

  16. Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A humanoid-specific multimodal occupancy perception system with a new dataset, sensor layout, and a fusion network that claims state-of-the-art results on its own benchmark.

  17. GaussianFusionOcc: A Seamless Sensor Fusion Approach for 3D Occupancy Prediction Using 3D Gaussians

    cs.CV 2025-07 conditional novelty 5.0 of 10

    GaussianFusionOcc fuses camera, LiDAR, and radar features through deformable attention to refine semantic 3D Gaussians, improving 3D occupancy prediction on nuScenes while lowering memory use and latency.

  18. QuadricFormer: Scene as Superquadrics for 3D Semantic Occupancy Prediction

    cs.CV 2025-06 conditional novelty 5.0 of 10

    QuadricFormer represents 3D scenes as a probabilistic mixture of superquadrics, improving accuracy and efficiency over Gaussian-based occupancy prediction on nuScenes.

  19. Diffusion-Based Generative Models for 3D Occupancy Prediction in Autonomous Driving

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Diffusion-based generative models, using discrete categorical diffusion conditioned on BEV features, improve 3D occupancy prediction and downstream planning for autonomous driving.

  20. CaR1: A Multi-Modal Baseline for BEV Vehicle Segmentation via Camera-Radar Fusion

    cs.RO 2025-09 conditional novelty 4.0 of 10

    CaR1 achieves 57.6 IoU on nuScenes BEV vehicle segmentation by fusing camera features with grid-scattered radar features via adaptive weighting.

Pith tools