Pith. sign in

REVIEW 4 cited by

Clio: Real-time Task-Driven Open-Set 3D Scene Graphs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.13696 v4 pith:FCS7NQOB submitted 2024-04-21 cs.RO

classification cs.RO
keywords scenesemanticcontributionobjectsrobottask-drivencliogranularity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern tools for class-agnostic image segmentation (e.g., SegmentAnything) and open-set semantic understanding (e.g., CLIP) provide unprecedented opportunities for robot perception and mapping. While traditional closed-set metric-semantic maps were restricted to tens or hundreds of semantic classes, we can now build maps with a plethora of objects and countless semantic variations. This leaves us with a fundamental question: what is the right granularity for the objects (and, more generally, for the semantic concepts) the robot has to include in its map representation? While related work implicitly chooses a level of granularity by tuning thresholds for object detection, we argue that such a choice is intrinsically task-dependent. The first contribution of this paper is to propose a task-driven 3D scene understanding problem, where the robot is given a list of tasks in natural language and has to select the granularity and the subset of objects and scene structure to retain in its map that is sufficient to complete the tasks. We show that this problem can be naturally formulated using the Information Bottleneck (IB), an established information-theoretic framework. The second contribution is an algorithm for task-driven 3D scene understanding based on an Agglomerative IB approach, that is able to cluster 3D primitives in the environment into task-relevant objects and regions and executes incrementally. The third contribution is to integrate our task-driven clustering algorithm into a real-time pipeline, named Clio, that constructs a hierarchical 3D scene graph of the environment online using only onboard compute, as the robot explores it. Our final contribution is an extensive experimental campaign showing that Clio not only allows real-time construction of compact open-set 3D scene graphs, but also improves the accuracy of task execution by limiting the map to relevant semantic concepts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Ella: Embodied Social Agents with Lifelong Memory

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Ella, an embodied social agent with a name-centric semantic memory and a spatiotemporal episodic memory, outperformed two re-implemented baselines in social influence and leadership tasks in a 3D simulation.

  2. Semantics-aware Predictive Inspection Path Planning

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A semantics-aware inspection planner that predicts repeated structures in unseen space reduces mission time by 12 to 19 percent in real ballast tanks while maintaining coverage.

  3. GraphPad: Inference-Time 3D Scene Graph Updates for Embodied Question Answering

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Allowing a vision-language model to edit its own 3D scene graph during inference improves embodied question answering from 52.3% to 55.3% on OpenEQA.

  4. CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity

    cs.RO 2025-06

Pith tools