Pith. sign in

REVIEW 5 cited by

OVO: Open-Vocabulary Occupancy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.16133 v2 pith:TIK27IKJ submitted 2023-05-25 cs.CV cs.AIcs.LGcs.RO

classification cs.CVcs.AIcs.LGcs.RO
keywords occupancypredictionsemantictrainingannotationsapproachdataframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Semantic occupancy prediction aims to infer dense geometry and semantics of surroundings for an autonomous agent to operate safely in the 3D environment. Existing occupancy prediction methods are almost entirely trained on human-annotated volumetric data. Although of high quality, the generation of such 3D annotations is laborious and costly, restricting them to a few specific object categories in the training dataset. To address this limitation, this paper proposes Open Vocabulary Occupancy (OVO), a novel approach that allows semantic occupancy prediction of arbitrary classes but without the need for 3D annotations during training. Keys to our approach are (1) knowledge distillation from a pre-trained 2D open-vocabulary segmentation model to the 3D occupancy network, and (2) pixel-voxel filtering for high-quality training data generation. The resulting framework is simple, compact, and compatible with most state-of-the-art semantic occupancy prediction models. On NYUv2 and SemanticKITTI datasets, OVO achieves competitive performance compared to supervised semantic occupancy prediction approaches. Furthermore, we conduct extensive analyses and ablation studies to offer insights into the design of the proposed framework. Our code is publicly available at https://github.com/dzcgaara/OVO.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    VISA improves closed-set 3D occupancy mIoU on nuScenes by using VLM instance audits as reliability-weighted semantic supervisors during training of existing world models.

  2. O3N: Omnidirectional Open-Vocabulary Occupancy Prediction for Urban Autonomous Agents

    cs.CV 2026-03 conditional novelty 6.5 of 10

    O3N is the first open-vocabulary occupancy prediction method that takes a single omnidirectional RGB image and labels 3D voxels with both seen and unseen semantic classes.

  3. SAM4D: Segment Anything in Camera and LiDAR Streams

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SAM4D is a promptable model that segments and tracks objects across camera and LiDAR streams with cross-modal prompts, trained on pseudo-labels generated by an automated data engine.

  4. TGRIP: A Text-Guided Approach to Vehicle Instance Prediction in Autonomous Driving

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Auxiliary CLIP-derived BEV semantic supervision during training improves nuScenes end-to-end vehicle instance prediction over a geometric-only baseline, with the semantic head removed at inference.

  5. Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities

    cs.RO 2025-09 conditional novelty 4.0 of 10

    Foundation-model perception for autonomous driving is surveyed through four capability lenses: generalized knowledge, spatial understanding, multi-sensor robustness, and temporal understanding.

Pith tools