REVIEW 6 cited by
Vision-based 3D occupancy prediction in autonomous driving: a review and outlook
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In recent years, autonomous driving has garnered escalating attention for its potential to relieve drivers' burdens and improve driving safety. Vision-based 3D occupancy prediction, which predicts the spatial occupancy status and semantics of 3D voxel grids around the autonomous vehicle from image inputs, is an emerging perception task suitable for cost-effective perception system of autonomous driving. Although numerous studies have demonstrated the greater advantages of 3D occupancy prediction over object-centric perception tasks, there is still a lack of a dedicated review focusing on this rapidly developing field. In this paper, we first introduce the background of vision-based 3D occupancy prediction and discuss the challenges in this task. Secondly, we conduct a comprehensive survey of the progress in vision-based 3D occupancy prediction from three aspects: feature enhancement, deployment friendliness and label efficiency, and provide an in-depth analysis of the potentials and challenges of each category of methods. Finally, we present a summary of prevailing research trends and propose some inspiring future outlooks. To provide a valuable reference for researchers, a regularly updated collection of related papers, datasets, and codes is organized at https://github.com/zya3d/Awesome-3D-Occupancy-Prediction.
Forward citations
Cited by 6 Pith papers
-
Event-aided Semantic Scene Completion
Fusing event-camera data during the 2D-to-3D lifting step improves semantic scene completion accuracy and robustness on a new real-world benchmark and on corrupted SemanticKITTI.
-
S2GO: Streaming Sparse Gaussian Occupancy Prediction
A sparse-query, streaming Gaussian occupancy predictor achieves state-of-the-art 3D semantic occupancy on nuScenes and KITTI with real-time inference.
-
An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training
An end-to-end, non-autoregressive 3D occupancy world model warps dynamic voxels via predicted flow, moves static voxels by pose, and uses image-based rendering supervision, achieving state-of-the-art results on three ...
-
LightOcc: Lightweight Spatial Embedding for Efficient Vision-based 3D Occupancy Prediction
LightOcc uses a one-channel occupancy volume, rearranged into tri-perspective views, to add height information to BEV features, reaching 47.24 mIoU on Occ3D-nuScenes with 8 history frames.
-
ViPOcc: Leveraging Visual Priors from Vision Foundation Models for Single-View 3D Occupancy Prediction
ViPOcc combines Depth Anything V2 metric depth priors and Grounded-SAM instance masks with a NeRF-based occupancy network to improve single-view 3D occupancy and depth prediction.
-
A Survey of World Models for Autonomous Driving
A survey presenting a three-branch taxonomy of world models for autonomous driving, plus benchmark tables comparing representative generation and planning methods on nuScenes, Waymo, Occ3D, and CarlaSC.
Discussion (0). Continue with ORCID to comment.