Pith. sign in

REVIEW 6 cited by

Gaussian Grouping: Segment and Edit Anything in 3D Scenes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.00732 v2 pith:WFNZRBSZ submitted 2023-12-01 cs.CV cs.AI

classification cs.CVcs.AI
keywords gaussiananythingscenesegmentgroupingsceneseditediting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent Gaussian Splatting achieves high-quality and real-time novel-view synthesis of the 3D scenes. However, it is solely concentrated on the appearance and geometry modeling, while lacking in fine-grained object-level scene understanding. To address this issue, we propose Gaussian Grouping, which extends Gaussian Splatting to jointly reconstruct and segment anything in open-world 3D scenes. We augment each Gaussian with a compact Identity Encoding, allowing the Gaussians to be grouped according to their object instance or stuff membership in the 3D scene. Instead of resorting to expensive 3D labels, we supervise the Identity Encodings during the differentiable rendering by leveraging the 2D mask predictions by Segment Anything Model (SAM), along with introduced 3D spatial consistency regularization. Compared to the implicit NeRF representation, we show that the discrete and grouped 3D Gaussians can reconstruct, segment and edit anything in 3D with high visual quality, fine granularity and efficiency. Based on Gaussian Grouping, we further propose a local Gaussian Editing scheme, which shows efficacy in versatile scene editing applications, including 3D object removal, inpainting, colorization, style transfer and scene recomposition. Our code and models are at https://github.com/lkeab/gaussian-grouping.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  2. LabelGS: Label-Aware 3D Gaussian Splatting for 3D Scene Segmentation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    LabelGS assigns 2D video-tracking labels to the most-contributing 3D Gaussians, with depth-based occlusion masking and a projection filter, reporting better mIoU/PSNR than Feature-3DGS with roughly 22x faster training.

  3. DCHM: Depth-Consistent Human Modeling for Multiview Detection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DCHM uses superpixel-based Gaussian Splatting to make monocular depth estimates multiview-consistent, producing point clouds that yield state-of-the-art label-free pedestrian detection on Wildtrack, Terrace, and MultiviewX.

  4. MUVOD: A Novel Multi-view Video Object Segmentation Dataset and A Benchmark for 3D Segmentation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MUVOD provides multi-view video panoptic masks for 17 dynamic scenes (459 instances, 73 categories) and benchmarks for multi-view video and 3D object segmentation.

  5. DSG-World: Learning a 3D Gaussian World Model from Dual State Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    DSG-World builds two segmented 3D Gaussian fields from two scene states and trains them with mutual consistency, enabling novel-state simulation without inpainting or dense capture.

  6. Hi-LSplat: Hierarchical 3D Language Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Hi-LSplat trains language-augmented 3D Gaussians with a three-level semantic tree and instance/part contrastive losses, improving open-vocabulary 3D segmentation and localization on eight datasets.

Pith tools