Pith. sign in

REVIEW 5 cited by

CLIP-GS: CLIP-Informed Gaussian Splatting for View-Consistent 3D Indoor Semantic Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.14249 v2 pith:7NEGTMP7 submitted 2024-04-22 cs.CV

classification cs.CV
keywords semanticgaussianclipcoherentindoorunderstandingachievingapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Exploiting 3D Gaussian Splatting (3DGS) with Contrastive Language-Image Pre-Training (CLIP) models for open-vocabulary 3D semantic understanding of indoor scenes has emerged as an attractive research focus. Existing methods typically attach high-dimensional CLIP semantic embeddings to 3D Gaussians and leverage view-inconsistent 2D CLIP semantics as Gaussian supervision, resulting in efficiency bottlenecks and deficient 3D semantic consistency. To address these challenges, we present CLIP-GS, efficiently achieving a coherent semantic understanding of 3D indoor scenes via the proposed Semantic Attribute Compactness (SAC) and 3D Coherent Regularization (3DCR). SAC approach exploits the naturally unified semantics within objects to learn compact, yet effective, semantic Gaussian representations, enabling highly efficient rendering (>100 FPS). 3DCR enforces semantic consistency in 2D and 3D domains: In 2D, 3DCR utilizes refined view-consistent semantic outcomes derived from 3DGS to establish cross-view coherence constraints; in 3D, 3DCR encourages features similar among 3D Gaussian primitives associated with the same object, leading to more precise and coherent segmentation results. Extensive experimental results demonstrate that our method remarkably suppresses existing state-of-the-art approaches, achieving mIoU improvements of 21.20% and 13.05% on ScanNet and Replica datasets, respectively, while maintaining real-time rendering speed. Furthermore, our approach exhibits superior performance even with sparse input data, substantiating its robustness.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Representing bounding-box coordinates as axis-specific hundreds, tens, and ones tokens, plus a geometry-aware GRPO reward, improves generative visual grounding accuracy.

  2. ArtChart: Faithful Artistic Chart Generation with Integrated Text Rendering

    cs.CV 2026-07 conditional novelty 6.0 of 10

    ArtChart, a ControlNet + GRPO + multi-expert distillation system, achieves about 9.1/10 math, 9.5/10 text, and 7.7/10 layout on a new 2K bilingual artistic-chart benchmark, well above open baselines.

  3. IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Training a coordinate-generating VLM with an IoU-aware distillation loss from a teacher that sees the ground-truth box marked on the image improves referring-expression grounding by ~3-4 accuracy points.

  4. Gaussian Splatting Feature Fields for Privacy-Preserving Visual Localization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A self-supervised 3D Gaussian feature field, with cluster-derived segmentations, is used for accurate camera pose refinement and privacy-preserving visual localization.

  5. Hi-LSplat: Hierarchical 3D Language Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Hi-LSplat trains language-augmented 3D Gaussians with a three-level semantic tree and instance/part contrastive losses, improving open-vocabulary 3D segmentation and localization on eight datasets.

Pith tools