REVIEW 3 cited by
OccTransformer: Improving BEVFormer for 3D camera-only occupancy prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This technical report presents our solution, "occTransformer" for the 3D occupancy prediction track in the autonomous driving challenge at CVPR 2023. Our method builds upon the strong baseline BEVFormer and improves its performance through several simple yet effective techniques. Firstly, we employed data augmentation to increase the diversity of the training data and improve the model's generalization ability. Secondly, we used a strong image backbone to extract more informative features from the input data. Thirdly, we incorporated a 3D unet head to better capture the spatial information of the scene. Fourthly, we added more loss functions to better optimize the model. Additionally, we used an ensemble approach with the occ model BevDet and SurroundOcc to further improve the performance. Most importantly, we integrated 3D detection model StreamPETR to enhance the model's ability to detect objects in the scene. Using these methods, our solution achieved 49.23 miou on the 3D occupancy prediction track in the autonomous driving challenge.
Forward citations
Cited by 3 Pith papers
-
OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
OccVLA trains a vision-language-action model to predict 3D occupancy as an auxiliary output, improving nuScenes trajectory planning and 3D VQA from camera images only, with the occupancy branch disabled at inference.
-
VisHall3D: Monocular Semantic Scene Completion from Reconstructing the Visible Regions to Hallucinating the Invisible Regions
VisHall3D splits monocular 3D scene completion into visible-region reconstruction (VisFrontierNet) and invisible-region hallucination (OcclusionMAE), reporting SOTA mIoU of 17.46 and 20.95 on SemanticKITTI and SSCBenc...
-
Spatiotemporal Decoupling for Efficient Vision-Based Occupancy Forecasting
EfficientOCF forecasts 3D occupancy by decoupling it into 2D BEV occupancy, height, and instance flow, achieving state-of-the-art accuracy and 82.33 ms inference on autonomous driving datasets.
Discussion (0). Continue with ORCID to comment.