Pith. sign in

REVIEW 3 cited by

ShapeLLM: Universal 3D Object Understanding for Embodied Interaction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.17766 v3 pith:OCSXGWPO submitted 2024-02-27 cs.CV

classification cs.CV
keywords shapellmreconunderstandingembodiedinteractionencodergeometryobject
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents ShapeLLM, the first 3D Multimodal Large Language Model (LLM) designed for embodied interaction, exploring a universal 3D object understanding with 3D point clouds and languages. ShapeLLM is built upon an improved 3D encoder by extending ReCon to ReCon++ that benefits from multi-view image distillation for enhanced geometry understanding. By utilizing ReCon++ as the 3D point cloud input encoder for LLMs, ShapeLLM is trained on constructed instruction-following data and tested on our newly human-curated benchmark, 3D MM-Vet. ReCon++ and ShapeLLM achieve state-of-the-art performance in 3D geometry understanding and language-unified 3D interaction tasks, such as embodied visual grounding. Project page: https://qizekun.github.io/shapellm/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Revisiting 3D LLM Benchmarks: Are We Really Testing 3D Capabilities?

    cs.AI 2025-02 conditional novelty 6.0 of 10

    Vision-language models given rendered 2D images of point clouds can outperform specialized 3D LLMs on object-level benchmarks, showing these benchmarks do not isolate 3D understanding.

  2. 3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding

    cs.CV 2025-01 conditional novelty 6.0 of 10

    3UR-LLM encodes 3D point clouds with a frozen detector, compresses the features into 32 tokens, and beats 3D-LLM by 7.1 CIDEr on ScanQA with less training time.

  3. 4DPC$^2$hat: Towards Dynamic Point Cloud Understanding with Failure-Aware Bootstrapping

    cs.CV 2026-02 conditional novelty 5.0 of 10

    A failure-aware bootstrapping MLLM with a new 200K-QA dataset becomes the first system to caption and answer questions about dynamic 4D point clouds, beating static-3D baselines by large margins.

Pith tools