Pith. sign in

REVIEW 20 cited by

SAM3D: Segment Anything in 3D Scenes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.03908 v1 pith:AVQPUUBJ submitted 2023-06-06 cs.CV

classification cs.CV
keywords maskssam3dapproachimagespointresultscloudfinetuning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we propose SAM3D, a novel framework that is able to predict masks in 3D point clouds by leveraging the Segment-Anything Model (SAM) in RGB images without further training or finetuning. For a point cloud of a 3D scene with posed RGB images, we first predict segmentation masks of RGB images with SAM, and then project the 2D masks into the 3D points. Later, we merge the 3D masks iteratively with a bottom-up merging approach. At each step, we merge the point cloud masks of two adjacent frames with the bidirectional merging approach. In this way, the 3D masks predicted from different frames are gradually merged into the 3D masks of the whole 3D scene. Finally, we can optionally ensemble the result from our SAM3D with the over-segmentation results based on the geometric information of the 3D scenes. Our approach is experimented with ScanNet dataset and qualitative results demonstrate that our SAM3D achieves reasonable and fine-grained 3D segmentation results without any training or finetuning of SAM.

Discussion (0). Sign in to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GReFEM: Multimodal LLMs as Zero-Shot Semantic Assistants for Physics-Guided 3D Mesh Refinement

    cs.GR 2026-07 conditional novelty 6.5 of 10

    GReFEM shows MLLMs zero-shot isolate load-activated geometric features for volumetric mesh refinement with higher precision than matched-budget geometric heuristics.

  2. JOPP-3D: Joint Open Vocabulary Semantic Segmentation on Point Clouds and Panoramas

    cs.CV 2026-03 conditional novelty 6.5 of 10

    A training-free pipeline jointly segments panoramic images and reconstructed point clouds with open-vocabulary language queries via tangential decomposition, instance proposals, CLIP alignment, and depth-based label t...

  3. CDIS: Cross-Dimensional Class-Agnostic 3D Instance Segmentation via 2D Mask Tracking and 3D-2D Projection Merging

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A training-free pipeline that tracks 2D masks frame-to-frame and associates them with 3D superpoints achieves 33.2 AP on ScanNet200 and 28.2 AP on ScanNet++, beating or matching prior zero-shot 3D instance segmentatio...

  4. OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A single-stage image-based detector that, trained with pseudo boxes from SAM segments and CLIP features, detects and classifies arbitrary indoor objects in 3D at 0.3 seconds per scene.

  5. ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting

    cs.GR 2025-07 conditional novelty 6.0 of 10

    ObjectGS unifies 3D Gaussian scene reconstruction with object-level segmentation by binding each object to local anchors with fixed one-hot ID encodings, improving open-vocabulary and panoptic segmentation.

  6. InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered Scenes

    cs.CV 2025-07 conditional novelty 6.0 of 10

    InstaScene combines Gaussian-based instance decomposition with generative completion to produce complete, scene-aligned 3D object models from cluttered scenes.

  7. SAM4D: Segment Anything in Camera and LiDAR Streams

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SAM4D is a promptable model that segments and tracks objects across camera and LiDAR streams with cross-modal prompts, trained on pseudo-labels generated by an automated data engine.

  8. BoxFusion: Reconstruction-Free Open-Vocabulary 3D Object Detection via Real-Time Multi-View Box Fusion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    BoxFusion fuses per-frame 3D bounding box proposals from Cubify Anything and CLIP semantics into open-vocabulary 3D detections, reporting state-of-the-art AP among online methods without dense reconstruction.

  9. SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Training-time alignment of π0 visual features with object-centric SAM3D 3D features improves VLA manipulation performance while keeping RGB-language-only inference.

  10. Distill, Diffuse, Segment: Unsupervised 3D Semantic Segmentation for Autonomous Driving Based on Multi-Level Distillation and Graph Diffusion

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    DDS performs annotation-free 3D semantic scene understanding by combining multi-granularity distillation from 2D visuals with graph-diffusion segmentation and cluster-to-category association, yielding up to 5.9% oAcc,...

  11. Unified Semantic Transformer for 3D Scene Understanding

    cs.CV 2025-12 reject novelty 5.0 of 10

    UNITE is a feed-forward transformer that predicts geometry plus semantic, instance, open-vocabulary, and articulation features for indoor 3D scenes from RGB images in one pass.

  12. Integrating SAM Supervision for 3D Weakly Supervised Point Cloud Segmentation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A weakly supervised 3D point cloud segmentation method that back-projects Semantic-SAM 2D masks into 3D, propagates sparse labels inside masks, and uses reliability-filtered pseudo labels, reporting state-of-the-art m...

  13. SERES: Semantic-aware neural reconstruction from sparse views

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    A semantic-aware implicit reconstruction method claims 44% and 20% lower Chamfer distance than SparseNeuS and VolRecon, and 69%/68% error reductions as a NeuS/Neuralangelo plugin.

  14. scI2CL: Effectively Integrating Single-cell Multi-omics by Intra- and Inter-omics Contrastive Learning

    q-bio.GN 2025-08 unverdicted novelty 5.0 of 10

    The abstract claims a state-of-the-art single-cell multi-omics integration method with new cell-subtype and trajectory findings, but the supplied full text is a different paper.

  15. GeoSAM2: Unleashing the Power of SAM2 for 3D Part Segmentation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A prompt-controllable 3D part segmentation method that adapts SAM2 with LoRA and geometry fusion on rendered normal and point maps, then back-projects multi-view masks to the mesh.

  16. Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MVOV3D corrects noise in multi-view vision-language features via region-level CLIP encoding, caption-based text features, and geometric pooling, achieving 14.7% mIoU on ScanNet200 and 16.2% on Matterport160 without tr...

  17. OV-MAP: Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots

    cs.CV 2025-06 conditional novelty 5.0 of 10

    OV-MAP projects 2D masks into 3D and uses mesh-area voting to create zero-shot, open-vocabulary 3D instance segmentation maps.

  18. Hi-LSplat: Hierarchical 3D Language Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Hi-LSplat trains language-augmented 3D Gaussians with a three-level semantic tree and instance/part contrastive losses, improving open-vocabulary 3D segmentation and localization on eight datasets.

  19. IRS: Instance-Level 3D Scene Graphs via Room Prior Guided LiDAR-Camera Fusion

    cs.RO 2025-06 conditional novelty 5.0 of 10

    IRS builds instance-level 3D scene graphs faster by using LiDAR room priors to constrain and parallelize semantic fusion from vision-language models.

  20. Details Matter for Indoor Open-vocabulary 3D Instance Segmentation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A carefully engineered pipeline of 2D grounding, 3D tracking, proposal merging, and Alpha-CLIP classification with a standardized similarity filter achieves state-of-the-art open-vocabulary 3D instance segmentation on...

Pith tools