Pith. sign in

REVIEW 18 cited by

Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.00615 v1 pith:QW7KGQUK submitted 2023-09-01 cs.CV cs.AIcs.CLcs.LGcs.MM

classification cs.CVcs.AIcs.CLcs.LGcs.MM
keywords point-bindpoint-llmmulti-modalitypointaligningapplicationscloudsembedding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Point-Bind, a 3D multi-modality model aligning point clouds with 2D image, language, audio, and video. Guided by ImageBind, we construct a joint embedding space between 3D and multi-modalities, enabling many promising applications, e.g., any-to-3D generation, 3D embedding arithmetic, and 3D open-world understanding. On top of this, we further present Point-LLM, the first 3D large language model (LLM) following 3D multi-modal instructions. By parameter-efficient fine-tuning techniques, Point-LLM injects the semantics of Point-Bind into pre-trained LLMs, e.g., LLaMA, which requires no 3D instruction data, but exhibits superior 3D and multi-modal question-answering capacity. We hope our work may cast a light on the community for extending 3D point clouds to multi-modality applications. Code is available at https://github.com/ZiyuGuo99/Point-Bind_Point-LLM.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A multimodal language model that dynamically routes queries to the most relevant scene modalities and modality-specialized experts, achieving state-of-the-art results on five 3D benchmarks.

  2. Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    ViPS fuses five complementary visual priors into an MLLM via lightweight distillation proxies and query-conditioned dynamic weighting, reporting state-of-the-art results on VSI-Bench and ScanNet-series benchmarks.

  3. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  4. MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    MV-GEL uses a learned view selector and a fine-tuned vision-language segmentation model to localize text-described faces and edges on 3D meshes.

  5. Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

    cs.CV 2026-03 unverdicted novelty 6.0 of 10

    VEGA-3D extracts intermediate spatiotemporal features from a pretrained video diffusion model and fuses them into MLLMs to improve geometric and embodied reasoning without explicit 3D supervision.

  6. Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the Wild

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    CAV-SAM reformulates reference segmentation as pseudo-video object segmentation using diffusion-based semantic transitions and test-time geometric alignment, claiming over 5% improvement over state-of-the-art.

  7. MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh

    cs.GR 2025-08 unverdicted novelty 6.0 of 10

    MeshLLM improves LLM-based 3D mesh understanding and generation through primitive decomposition, a 1500k+ sample dataset, and topology-focused training strategies.

  8. Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A two-stage reasoning-segmentation method plus a new LLM-generated 3D dataset improves spatial reasoning in 3D multimodal large language models on several benchmarks.

  9. Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Video-grounded RL with two-stage 2D grounding and SAM2 lifting enables 3D object localization and QA without dense 3D instance supervision.

  10. Rethinking Multimodal Few-Shot 3D Point Cloud Segmentation: From Fused Refinement to Decoupled Arbitration

    cs.CV 2026-01 conditional novelty 5.0 of 10

    Decoupling geometric and semantic feature pathways in a few-shot 3D segmentation model yields modest but consistent mIoU improvements over the fused-refinement baseline on S3DIS and ScanNet.

  11. Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Argus fuses multi-view images and camera poses with 3D point cloud features in a frozen-LLM Q-Former architecture, improving 3D question answering, grounding, and scene description over prior 3D-LMMs.

  12. Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Fast3D prunes up to 90% of object-centric visual tokens in 3D MLLMs while preserving about 96.8% of original benchmark performance, using a trained attention predictor and adaptive layer-wise pruning.

  13. MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MINT-CoT-7B interleaves fine-grained visual tokens into each math reasoning step and reports 73.70 on MathVista-Math, 64.72 on GeoQA, and 69.6 on MMStar-Math.

  14. Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs

    cs.CV 2025-06 conditional novelty 5.0 of 10

    3D scene-centric VLMs underutilize the 3D encoder, overfit to linguistic and answer-frequency shortcuts, and the proposed 3D-RDQA dataset helps expose and mitigate this problem.

  15. Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A sparse mixture-of-experts 3D multimodal LLM adaptively fuses RGB, RGBD, BEV, point cloud, and voxel tokens, achieving SOTA on several ScanNet-based 3D scene understanding benchmarks.

  16. Aligning Proteins and Language: A Foundation Model for Protein Retrieval

    q-bio.BM 2025-05 conditional novelty 5.0 of 10

    A CLIP-style model aligns protein surface point clouds with GO-derived captions and achieves roughly 60% Top-5 zero-shot retrieval on PDB and 36% on EMDB.

  17. RelMap: Reliable Spatiotemporal Sensor Data Visualization via Imputative Spatial Interpolation

    cs.LG 2025-08 unverdicted novelty 4.0 of 10

    RelMap combines GNN-based imputation with spatial interpolation and uncertainty-aware heatmaps for spatiotemporal sensor data.

  18. NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding

    cs.CV 2025-07 conditional novelty 4.0 of 10

    NeuroVoxel-LM combines dynamic multi-resolution voxelization with attention-based pooling of NeRF weights, reporting faster 3D feature extraction and modestly better NeRF captioning than fixed-resolution and max-pooli...

Pith tools