Pith. sign in

REVIEW 2 cited by

Model2Scene: Learning 3D Scene Representation via Contrastive Language-CAD Models Pre-training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.16956 v1 pith:IUHM3HD7 submitted 2023-09-29 cs.CV

classification cs.CV
keywords scenemodelsmodel2sceneobjectpointrepresentationchallengescontrastive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current successful methods of 3D scene perception rely on the large-scale annotated point cloud, which is tedious and expensive to acquire. In this paper, we propose Model2Scene, a novel paradigm that learns free 3D scene representation from Computer-Aided Design (CAD) models and languages. The main challenges are the domain gaps between the CAD models and the real scene's objects, including model-to-scene (from a single model to the scene) and synthetic-to-real (from synthetic model to real scene's object). To handle the above challenges, Model2Scene first simulates a crowded scene by mixing data-augmented CAD models. Next, we propose a novel feature regularization operation, termed Deep Convex-hull Regularization (DCR), to project point features into a unified convex hull space, reducing the domain gap. Ultimately, we impose contrastive loss on language embedding and the point features of CAD models to pre-train the 3D network. Extensive experiments verify the learned 3D scene representation is beneficial for various downstream tasks, including label-free 3D object salient detection, label-efficient 3D scene perception and zero-shot 3D semantic segmentation. Notably, Model2Scene yields impressive label-free 3D object salient detection with an average mAP of 46.08\% and 55.49\% on the ScanNet and S3DIS datasets, respectively. The code will be publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PanoSLAM: Panoptic 3D Scene Reconstruction via Gaussian SLAM

    cs.CV 2024-12 conditional novelty 6.0 of 10

    PanoSLAM is a Gaussian Splatting SLAM system that produces label-free 3D panoptic maps from RGB-D video by lifting and refining 2D panoptic predictions in 3D.

  2. OVGaussian: Generalizable 3D Gaussian Segmentation with Open Vocabularies

    cs.CV 2024-12 conditional novelty 6.0 of 10

    This paper introduces a dataset and a cross-modal network that predicts renderable open-vocabulary semantic attributes for 3D Gaussian scenes, enabling segmentation of new scenes without scene-specific fine-tuning.

Pith tools