Pith. sign in

REVIEW 3 cited by

Context-Aware Entity Grounding with Open-Vocabulary 3D Scene Graphs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.15940 v1 pith:USGXBGTP submitted 2023-09-27 cs.RO cs.CV

classification cs.ROcs.CV
keywords localizationopen-vocabularyovsgscenecontext-awaredatasetentityexperiments
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present an Open-Vocabulary 3D Scene Graph (OVSG), a formal framework for grounding a variety of entities, such as object instances, agents, and regions, with free-form text-based queries. Unlike conventional semantic-based object localization approaches, our system facilitates context-aware entity localization, allowing for queries such as ``pick up a cup on a kitchen table" or ``navigate to a sofa on which someone is sitting". In contrast to existing research on 3D scene graphs, OVSG supports free-form text input and open-vocabulary querying. Through a series of comparative experiments using the ScanNet dataset and a self-collected dataset, we demonstrate that our proposed approach significantly surpasses the performance of previous semantic-based localization techniques. Moreover, we highlight the practical application of OVSG in real-world robot navigation and manipulation experiments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A training-free pipeline that disambiguates text queries and infers viewpoints improves zero-shot 3D visual grounding, reaching 64.06% Acc@0.5 on ScanRefer.

  2. 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A learnable 3D scene graph representation that injects semantic relation embeddings between objects improves LLM performance on 3D grounding, captioning, and question answering.

  3. Time is on my sight: scene graph filtering for dynamic environment perception in an LLM-driven robot

    cs.RO 2024-11 conditional novelty 4.0 of 10

    A particle filter improves object localization in a dynamically updated semantic scene graph used by an LLM-driven robot planner.

Pith tools