Pith. sign in

REVIEW 2 cited by

OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.03650 v1 pith:5K73IYME submitted 2024-04-04 cs.CV

classification cs.CV
keywords featuresopennerfsegmentationclipimageimagesmodelsopen-set
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large visual-language models (VLMs), like CLIP, enable open-set image segmentation to segment arbitrary concepts from an image in a zero-shot manner. This goes beyond the traditional closed-set assumption, i.e., where models can only segment classes from a pre-defined training set. More recently, first works on open-set segmentation in 3D scenes have appeared in the literature. These methods are heavily influenced by closed-set 3D convolutional approaches that process point clouds or polygon meshes. However, these 3D scene representations do not align well with the image-based nature of the visual-language models. Indeed, point cloud and 3D meshes typically have a lower resolution than images and the reconstructed 3D scene geometry might not project well to the underlying 2D image sequences used to compute pixel-aligned CLIP features. To address these challenges, we propose OpenNeRF which naturally operates on posed images and directly encodes the VLM features within the NeRF. This is similar in spirit to LERF, however our work shows that using pixel-wise VLM features (instead of global CLIP features) results in an overall less complex architecture without the need for additional DINO regularization. Our OpenNeRF further leverages NeRF's ability to render novel views and extract open-set VLM features from areas that are not well observed in the initial posed images. For 3D point cloud segmentation on the Replica dataset, OpenNeRF outperforms recent open-vocabulary methods such as LERF and OpenScene by at least +4.9 mIoU.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DiSCO-3D : Discovering and segmenting Sub-Concepts from Open-vocabulary queries in NeRF

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A NeRF-based method jointly performs unsupervised semantic clustering and CLIP-guided relevancy to discover and segment query-relevant sub-concepts in 3D scenes, with a new Replica benchmark.

  2. UnMix-NeRF: Spectral Unmixing Meets Neural Radiance Fields

    eess.IV 2025-06 conditional novelty 6.0 of 10

    A NeRF-based framework jointly performs hyperspectral novel view synthesis and unsupervised material segmentation by learning per-point spectral abundances over a global endmember dictionary.

Pith tools