Pith. sign in

REVIEW 9 cited by

Semantic Gaussians: Open-Vocabulary Scene Understanding with 3D Gaussian Splatting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.15624 v2 pith:MJC45WJR submitted 2024-03-22 cs.CV

classification cs.CV
keywords semanticgaussiansscenesegmentationunderstandingmethodsopen-vocabularyapplications
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Open-vocabulary 3D scene understanding presents a significant challenge in computer vision, with wide-ranging applications in embodied agents and augmented reality systems. Existing methods adopt neurel rendering methods as 3D representations and jointly optimize color and semantic features to achieve rendering and scene understanding simultaneously. In this paper, we introduce Semantic Gaussians, a novel open-vocabulary scene understanding approach based on 3D Gaussian Splatting. Our key idea is to distill knowledge from 2D pre-trained models to 3D Gaussians. Unlike existing methods, we design a versatile projection approach that maps various 2D semantic features from pre-trained image encoders into a novel semantic component of 3D Gaussians, which is based on spatial relationship and need no additional training. We further build a 3D semantic network that directly predicts the semantic component from raw 3D Gaussians for fast inference. The quantitative results on ScanNet segmentation and LERF object localization demonstates the superior performance of our method. Additionally, we explore several applications of Semantic Gaussians including object part segmentation, instance segmentation, scene editing, and spatiotemporal segmentation with better qualitative results over 2D and 3D baselines, highlighting its versatility and effectiveness on supporting diverse downstream tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. E3DGS: Unified Geometric-Photometric Equivariance for 3D Gaussian Splatting via Color-as-Geometry Embedding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    3D Gaussian view-dependent colors are repacked as 3×3 matrices so geometry and color rotate together, giving exact rotation-equivariant recognition and world modeling in 3DGS.

  2. NLI4VolVis: Natural Language Interaction for Volume Visualization via LLM Multi-Agents and Editable 3D Gaussian Splatting

    cs.HC 2025-07 conditional novelty 6.0 of 10

    NLI4VolVis integrates multi-agent large language models, editable 3D Gaussian splatting, and CLIP-based querying so users can explore, query, and edit volume visualizations through natural language.

  3. PointGS: Point Attention-Aware Sparse View Synthesis with Gaussian Splatting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PointGS improves few-shot 3D Gaussian splatting by fusing multi-view image features per 3D point and refining them with a neighbor-attention network before decoding Gaussian colors.

  4. PIG: Physically-based Multi-Material Interaction with 3D Gaussians

    cs.GR 2025-06 conditional novelty 6.0 of 10

    PIG couples depth-based 3D object segmentation with MLS-MPM physics and adaptive eigen-clamping of Gaussian deformations to create multi-material interactions inside 3D Gaussian scenes.

  5. Enhancing LLM Training via Spectral Clipping

    cs.LG 2026-03 unverdicted novelty 5.0 of 10

    SPECTRA improves LLM pretraining via post-clipping of update spectral norms and optional pre-clipping of gradient spikes, framed as Composite Frank-Wolfe regularization.

  6. PGOV3D: Open-Vocabulary 3D Semantic Segmentation with Partial-to-Global Curriculum

    cs.CV 2025-06 conditional novelty 5.0 of 10

    PGOV3D reports 59.5 mIoU on ScanNet for open-vocabulary 3D segmentation by pretraining on partial RGB-D views and then fine-tuning on full scenes with self-generated pseudo labels.

  7. Hi-LSplat: Hierarchical 3D Language Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Hi-LSplat trains language-augmented 3D Gaussians with a three-level semantic tree and instance/part contrastive losses, improving open-vocabulary 3D segmentation and localization on eight datasets.

  8. LEG-SLAM: Real-Time Language-Enhanced Gaussian Splatting for SLAM

    cs.CV 2025-06 conditional novelty 5.0 of 10

    LEG-SLAM is a real-time RGB-D SLAM that jointly renders photorealistic images and open-vocabulary semantic masks by distilling PCA-compressed DINOv2 features into 3D Gaussians.

  9. The ALMA-QUARKS Survey: III. Clump-to-core fragmentation and search for high-mass starless cores

    astro-ph.GA 2025-08 unverdicted novelty 4.0 of 10

    In 139 infrared-bright massive protoclusters, ALMA resolves 1562 cores whose separations are much smaller than the Jeans length, and finds only two candidate high-mass starless cores.

Pith tools