Pith. sign in

REVIEW 9 cited by

Segment Any 3D Gaussians

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.00860 v3 pith:EBBL22YA submitted 2023-12-01 cs.CV

classification cs.CV
keywords segmentationsagasegmentaffinityfeaturegaussiansmulti-granularityd-gs
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents SAGA (Segment Any 3D GAussians), a highly efficient 3D promptable segmentation method based on 3D Gaussian Splatting (3D-GS). Given 2D visual prompts as input, SAGA can segment the corresponding 3D target represented by 3D Gaussians within 4 ms. This is achieved by attaching an scale-gated affinity feature to each 3D Gaussian to endow it a new property towards multi-granularity segmentation. Specifically, a scale-aware contrastive training strategy is proposed for the scale-gated affinity feature learning. It 1) distills the segmentation capability of the Segment Anything Model (SAM) from 2D masks into the affinity features and 2) employs a soft scale gate mechanism to deal with multi-granularity ambiguity in 3D segmentation through adjusting the magnitude of each feature channel according to a specified 3D physical scale. Evaluations demonstrate that SAGA achieves real-time multi-granularity segmentation with quality comparable to state-of-the-art methods. As one of the first methods addressing promptable segmentation in 3D-GS, the simplicity and effectiveness of SAGA pave the way for future advancements in this field. Our code will be released.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SaaF: Scene-Specific Ambiguity-Aware 3D Language Fields towards Interactive Real-World Object Retrieval

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SaaF embeds metric-learned CLIP features into 3D Gaussians and uses the L2 norm of a compressed text query to decide whether to ask the user for clarification before retrieving an object.

  2. EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An end-to-end 3D editing framework achieves high-fidelity local edits from coarse bounding boxes and 2D image prompts using region-aware loss reweighting and a large-scale parts-derived training dataset.

  3. TexGS-VolVis: Expressive Scene Editing for Volume Visualization via Textured Gaussian Splatting

    cs.GR 2025-07 conditional novelty 6.0 of 10

    A textured Gaussian splatting framework enables flexible image- and text-driven style editing of volume visualizations with real-time rendering.

  4. VolSegGS: Segmentation and Tracking in Dynamic Volumetric Scenes via Deformable 3D Gaussians

    cs.GR 2025-07 conditional novelty 6.0 of 10

    VolSegGS reconstructs dynamic volumetric scenes from rendered images with deformable 3D Gaussians and enables real-time interactive segmentation and tracking of regions over time.

  5. MUVOD: A Novel Multi-view Video Object Segmentation Dataset and A Benchmark for 3D Segmentation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MUVOD provides multi-view video panoptic masks for 17 dynamic scenes (459 instances, 73 categories) and benchmarks for multi-view video and 3D object segmentation.

  6. PIG: Physically-based Multi-Material Interaction with 3D Gaussians

    cs.GR 2025-06 conditional novelty 6.0 of 10

    PIG couples depth-based 3D object segmentation with MLS-MPM physics and adaptive eigen-clamping of Gaussian deformations to create multi-material interactions inside 3D Gaussian scenes.

  7. Latent Radiance Fields with 3D-aware 2D Representations

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A three-stage pipeline makes VAE latent codes 3D-consistent and builds a latent radiance field, improving photorealistic novel-view synthesis in latent space.

  8. Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering

    cs.CV 2025-10 conditional novelty 5.0 of 10

    UGSDF achieves state-of-the-art novel-view rendering of dynamic urban objects without LiDAR or 3D motion annotations by jointly optimizing SDFs and 3D Gaussians under 2D depth and point-tracking priors.

  9. Hi-LSplat: Hierarchical 3D Language Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Hi-LSplat trains language-augmented 3D Gaussians with a three-level semantic tree and instance/part contrastive losses, improving open-vocabulary 3D segmentation and localization on eight datasets.

Pith tools