REVIEW 20 cited by
SAM3D: Segment Anything in 3D Scenes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this work, we propose SAM3D, a novel framework that is able to predict masks in 3D point clouds by leveraging the Segment-Anything Model (SAM) in RGB images without further training or finetuning. For a point cloud of a 3D scene with posed RGB images, we first predict segmentation masks of RGB images with SAM, and then project the 2D masks into the 3D points. Later, we merge the 3D masks iteratively with a bottom-up merging approach. At each step, we merge the point cloud masks of two adjacent frames with the bidirectional merging approach. In this way, the 3D masks predicted from different frames are gradually merged into the 3D masks of the whole 3D scene. Finally, we can optionally ensemble the result from our SAM3D with the over-segmentation results based on the geometric information of the 3D scenes. Our approach is experimented with ScanNet dataset and qualitative results demonstrate that our SAM3D achieves reasonable and fine-grained 3D segmentation results without any training or finetuning of SAM.
Forward citations
Cited by 20 Pith papers
-
GReFEM: Multimodal LLMs as Zero-Shot Semantic Assistants for Physics-Guided 3D Mesh Refinement
GReFEM shows MLLMs zero-shot isolate load-activated geometric features for volumetric mesh refinement with higher precision than matched-budget geometric heuristics.
-
JOPP-3D: Joint Open Vocabulary Semantic Segmentation on Point Clouds and Panoramas
A training-free pipeline jointly segments panoramic images and reconstructed point clouds with open-vocabulary language queries via tangential decomposition, instance proposals, CLIP alignment, and depth-based label t...
-
CDIS: Cross-Dimensional Class-Agnostic 3D Instance Segmentation via 2D Mask Tracking and 3D-2D Projection Merging
A training-free pipeline that tracks 2D masks frame-to-frame and associates them with 3D superpoints achieves 33.2 AP on ScanNet200 and 28.2 AP on ScanNet++, beating or matching prior zero-shot 3D instance segmentatio...
-
OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations
A single-stage image-based detector that, trained with pseudo boxes from SAM segments and CLIP features, detects and classifies arbitrary indoor objects in 3D at 0.3 seconds per scene.
-
ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting
ObjectGS unifies 3D Gaussian scene reconstruction with object-level segmentation by binding each object to local anchors with fixed one-hot ID encodings, improving open-vocabulary and panoptic segmentation.
-
InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered Scenes
InstaScene combines Gaussian-based instance decomposition with generative completion to produce complete, scene-aligned 3D object models from cluttered scenes.
-
SAM4D: Segment Anything in Camera and LiDAR Streams
SAM4D is a promptable model that segments and tracks objects across camera and LiDAR streams with cross-modal prompts, trained on pseudo-labels generated by an automated data engine.
-
BoxFusion: Reconstruction-Free Open-Vocabulary 3D Object Detection via Real-Time Multi-View Box Fusion
BoxFusion fuses per-frame 3D bounding box proposals from Cubify Anything and CLIP semantics into open-vocabulary 3D detections, reporting state-of-the-art AP among online methods without dense reconstruction.
-
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
Training-time alignment of π0 visual features with object-centric SAM3D 3D features improves VLA manipulation performance while keeping RGB-language-only inference.
-
Distill, Diffuse, Segment: Unsupervised 3D Semantic Segmentation for Autonomous Driving Based on Multi-Level Distillation and Graph Diffusion
DDS performs annotation-free 3D semantic scene understanding by combining multi-granularity distillation from 2D visuals with graph-diffusion segmentation and cluster-to-category association, yielding up to 5.9% oAcc,...
-
Unified Semantic Transformer for 3D Scene Understanding
UNITE is a feed-forward transformer that predicts geometry plus semantic, instance, open-vocabulary, and articulation features for indoor 3D scenes from RGB images in one pass.
-
Integrating SAM Supervision for 3D Weakly Supervised Point Cloud Segmentation
A weakly supervised 3D point cloud segmentation method that back-projects Semantic-SAM 2D masks into 3D, propagates sparse labels inside masks, and uses reliability-filtered pseudo labels, reporting state-of-the-art m...
-
SERES: Semantic-aware neural reconstruction from sparse views
A semantic-aware implicit reconstruction method claims 44% and 20% lower Chamfer distance than SparseNeuS and VolRecon, and 69%/68% error reductions as a NeuS/Neuralangelo plugin.
-
scI2CL: Effectively Integrating Single-cell Multi-omics by Intra- and Inter-omics Contrastive Learning
The abstract claims a state-of-the-art single-cell multi-omics integration method with new cell-subtype and trajectory findings, but the supplied full text is a different paper.
-
GeoSAM2: Unleashing the Power of SAM2 for 3D Part Segmentation
A prompt-controllable 3D part segmentation method that adapts SAM2 with LoRA and geometry fusion on rendered normal and point maps, then back-projects multi-view masks to the mesh.
-
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding
MVOV3D corrects noise in multi-view vision-language features via region-level CLIP encoding, caption-based text features, and geometric pooling, achieving 14.7% mIoU on ScanNet200 and 16.2% on Matterport160 without tr...
-
OV-MAP: Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
OV-MAP projects 2D masks into 3D and uses mesh-area voting to create zero-shot, open-vocabulary 3D instance segmentation maps.
-
Hi-LSplat: Hierarchical 3D Language Gaussian Splatting
Hi-LSplat trains language-augmented 3D Gaussians with a three-level semantic tree and instance/part contrastive losses, improving open-vocabulary 3D segmentation and localization on eight datasets.
-
IRS: Instance-Level 3D Scene Graphs via Room Prior Guided LiDAR-Camera Fusion
IRS builds instance-level 3D scene graphs faster by using LiDAR room priors to constrain and parallelize semantic fusion from vision-language models.
-
Details Matter for Indoor Open-vocabulary 3D Instance Segmentation
A carefully engineered pipeline of 2D grounding, 3D tracking, proposal merging, and Alpha-CLIP classification with a standardized similarity filter achieves state-of-the-art open-vocabulary 3D instance segmentation on...
Discussion (0). Sign in to comment.