Pith. sign in

REVIEW 2 cited by

CVT-Occ: Cost Volume Temporal Fusion for 3D Occupancy Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.13430 v3 pith:NZYXZJIU submitted 2024-09-20 cs.CV cs.AI

classification cs.CVcs.AI
keywords costcvt-occoccupancypredictionvolumeapproachfeaturesfusion
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Vision-based 3D occupancy prediction is significantly challenged by the inherent limitations of monocular vision in depth estimation. This paper introduces CVT-Occ, a novel approach that leverages temporal fusion through the geometric correspondence of voxels over time to improve the accuracy of 3D occupancy predictions. By sampling points along the line of sight of each voxel and integrating the features of these points from historical frames, we construct a cost volume feature map that refines current volume features for improved prediction outcomes. Our method takes advantage of parallax cues from historical observations and employs a data-driven approach to learn the cost volume. We validate the effectiveness of CVT-Occ through rigorous experiments on the Occ3D-Waymo dataset, where it outperforms state-of-the-art methods in 3D occupancy prediction with minimal additional computational cost. The code is released at \url{https://github.com/Tsinghua-MARS-Lab/CVT-Occ}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Language Driven Occupancy Prediction

    cs.CV 2024-11 conditional novelty 6.0 of 10

    LOcc transfers text labels from images through LiDAR points to voxels to create dense pseudo-labeled 3D language ground truth, and uses it to train occupancy models that outperform prior zero-shot open-vocabulary methods.

  2. GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A world model operating on 3D Gaussians forecasts the current occupancy from the previous frame and current RGB, improving mIoU by about 2 points on nuScenes without meaningful added latency.

Pith tools