REVIEW 9 cited by
Segment Any 3D Gaussians
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper presents SAGA (Segment Any 3D GAussians), a highly efficient 3D promptable segmentation method based on 3D Gaussian Splatting (3D-GS). Given 2D visual prompts as input, SAGA can segment the corresponding 3D target represented by 3D Gaussians within 4 ms. This is achieved by attaching an scale-gated affinity feature to each 3D Gaussian to endow it a new property towards multi-granularity segmentation. Specifically, a scale-aware contrastive training strategy is proposed for the scale-gated affinity feature learning. It 1) distills the segmentation capability of the Segment Anything Model (SAM) from 2D masks into the affinity features and 2) employs a soft scale gate mechanism to deal with multi-granularity ambiguity in 3D segmentation through adjusting the magnitude of each feature channel according to a specified 3D physical scale. Evaluations demonstrate that SAGA achieves real-time multi-granularity segmentation with quality comparable to state-of-the-art methods. As one of the first methods addressing promptable segmentation in 3D-GS, the simplicity and effectiveness of SAGA pave the way for future advancements in this field. Our code will be released.
Forward citations
Cited by 9 Pith papers
-
SaaF: Scene-Specific Ambiguity-Aware 3D Language Fields towards Interactive Real-World Object Retrieval
SaaF embeds metric-learned CLIP features into 3D Gaussians and uses the L2 norm of a compressed text query to decide whether to ask the user for clarification before retrieving an object.
-
EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning
An end-to-end 3D editing framework achieves high-fidelity local edits from coarse bounding boxes and 2D image prompts using region-aware loss reweighting and a large-scale parts-derived training dataset.
-
TexGS-VolVis: Expressive Scene Editing for Volume Visualization via Textured Gaussian Splatting
A textured Gaussian splatting framework enables flexible image- and text-driven style editing of volume visualizations with real-time rendering.
-
VolSegGS: Segmentation and Tracking in Dynamic Volumetric Scenes via Deformable 3D Gaussians
VolSegGS reconstructs dynamic volumetric scenes from rendered images with deformable 3D Gaussians and enables real-time interactive segmentation and tracking of regions over time.
-
MUVOD: A Novel Multi-view Video Object Segmentation Dataset and A Benchmark for 3D Segmentation
MUVOD provides multi-view video panoptic masks for 17 dynamic scenes (459 instances, 73 categories) and benchmarks for multi-view video and 3D object segmentation.
-
PIG: Physically-based Multi-Material Interaction with 3D Gaussians
PIG couples depth-based 3D object segmentation with MLS-MPM physics and adaptive eigen-clamping of Gaussian deformations to create multi-material interactions inside 3D Gaussian scenes.
-
Latent Radiance Fields with 3D-aware 2D Representations
A three-stage pipeline makes VAE latent codes 3D-consistent and builds a latent radiance field, improving photorealistic novel-view synthesis in latent space.
-
Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering
UGSDF achieves state-of-the-art novel-view rendering of dynamic urban objects without LiDAR or 3D motion annotations by jointly optimizing SDFs and 3D Gaussians under 2D depth and point-tracking priors.
-
Hi-LSplat: Hierarchical 3D Language Gaussian Splatting
Hi-LSplat trains language-augmented 3D Gaussians with a three-level semantic tree and instance/part contrastive losses, improving open-vocabulary 3D segmentation and localization on eight datasets.
Discussion (0). Continue with ORCID to comment.