Pith. sign in

REVIEW 5 cited by

ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.16650 v1 pith:FUUMMBBH submitted 2023-09-28 cs.RO cs.CV

classification cs.ROcs.CV
keywords planningconceptgraphsmodelsrepresentationsemanticapproachesdownstreamhttps
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

For robots to perform a wide variety of tasks, they require a 3D representation of the world that is semantically rich, yet compact and efficient for task-driven perception and planning. Recent approaches have attempted to leverage features from large vision-language models to encode semantics in 3D representations. However, these approaches tend to produce maps with per-point feature vectors, which do not scale well in larger environments, nor do they contain semantic spatial relationships between entities in the environment, which are useful for downstream planning. In this work, we propose ConceptGraphs, an open-vocabulary graph-structured representation for 3D scenes. ConceptGraphs is built by leveraging 2D foundation models and fusing their output to 3D by multi-view association. The resulting representations generalize to novel semantic classes, without the need to collect large 3D datasets or finetune models. We demonstrate the utility of this representation through a number of downstream planning tasks that are specified through abstract (language) prompts and require complex reasoning over spatial and semantic concepts. (Project page: https://concept-graphs.github.io/ Explainer video: https://youtu.be/mRhNkQwRYnc )

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Memory for Attention: Language-Conditioned Re-Perception with a Vision--Language--Motion Map

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Persistent map memory schedules budgeted re-perception better than memoryless VLM priors, with gain equal to Var(√λ), and language-conditioned VLMM needs both open-vocabulary relevance and per-instance dynamics.

  2. Vision-Language-Motion Maps: An Open-Vocabulary, Uncertainty-Aware, Queryable Motion Attribute for 3D Scene Maps

    cs.RO 2026-07 conditional novelty 6.0 of 10

    VLMM is a 3D map representation where each object carries a fused, uncertainty-aware motion attribute (language-based movability prior + observed geometric motion) that makes motion queries such as 'what is moving' an...

  3. What to Distinguish and How? Opportunities and Challenges of Augmenting Multiple, Cluttered Objects in Complex Scenes for People with Low Vision

    cs.HC 2026-07 conditional novelty 6.0 of 10

    For people with low vision, AR overlays that rank objects by importance redirect attention toward high-priority objects, but multi-object augmentation lowers overall scene recall and creates new visual-confusion problems.

  4. Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes

    cs.CV 2025-08 conditional novelty 6.0 of 10

    OVODA combines a 3DETR-style detector with a frozen foundation model to detect novel objects and attributes in 3D scenes, and the OVAD dataset adds spatial and motion attribute labels to nuScenes.

  5. SPARK: Graph-Based Online Semantic Integration System for Robot Task Planning

    cs.RO 2025-06 conditional novelty 4.0 of 10

    SPARK updates a 3D scene graph online from verbal, text, and gesture cues, and replans robot fetch tasks based on the updated graph.

Pith tools