Pith. sign in

REVIEW 5 cited by

SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.09340 v3 pith:3KVOJIOH submitted 2024-01-17 cs.CV cs.AIcs.CLcs.LGcs.RO

classification cs.CVcs.AIcs.CLcs.LGcs.RO
keywords vision-languagelearninggroundedscenesgroundingsceneversechallengesdata
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

3D vision-language grounding, which focuses on aligning language with the 3D physical environment, stands as a cornerstone in the development of embodied agents. In comparison to recent advancements in the 2D domain, grounding language in 3D scenes faces several significant challenges: (i) the inherent complexity of 3D scenes due to the diverse object configurations, their rich attributes, and intricate relationships; (ii) the scarcity of paired 3D vision-language data to support grounded learning; and (iii) the absence of a unified learning framework to distill knowledge from grounded 3D data. In this work, we aim to address these three major challenges in 3D vision-language by examining the potential of systematically upscaling 3D vision-language learning in indoor environments. We introduce the first million-scale 3D vision-language dataset, SceneVerse, encompassing about 68K 3D indoor scenes and comprising 2.5M vision-language pairs derived from both human annotations and our scalable scene-graph-based generation approach. We demonstrate that this scaling allows for a unified pre-training framework, Grounded Pre-training for Scenes (GPS), for 3D vision-language learning. Through extensive experiments, we showcase the effectiveness of GPS by achieving state-of-the-art performance on all existing 3D visual grounding benchmarks. The vast potential of SceneVerse and GPS is unveiled through zero-shot transfer experiments in the challenging 3D vision-language tasks. Project website: https://scene-verse.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A two-stage reasoning-segmentation method plus a new LLM-generated 3D dataset improves spatial reasoning in 3D multimodal large language models on several benchmarks.

  2. Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A contrastive pre-training method aligns 3D point features with language at the entity level and achieves state-of-the-art open-vocabulary semantic segmentation on ScanNet.

  3. SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models

    cs.CV 2025-04 conditional novelty 6.0 of 10

    SORT3D is a zero-shot 3D grounding system where an LLM calls hand-built spatial functions and uses 2D captions, matching or beating prior zero-shot methods on several view-dependent subsets while running on real robots.

  4. LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LSceneLLM chooses task-relevant 3D regions via LLM attention, magnifies their details, and improves large-scene 3D question answering, planning, and captioning.

  5. Foundational Models for 3D Point Clouds: A Survey and Outlook

    cs.CV 2025-01 conditional novelty 4.0 of 10

    A structured review of methods that build or adapt 2D foundation models and LLMs for 3D point cloud tasks, with a proposed taxonomy and curated paper list.

Pith tools