Pith. sign in

REVIEW 10 cited by

3D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and Humans

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.06289 v2 pith:UACMJHKZ submitted 2020-02-15 cs.RO cs.AIcs.CV

classification cs.ROcs.AIcs.CV
keywords scenegraphsdynamicperceptionspatialactionablehumannodes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a unified representation for actionable spatial perception: 3D Dynamic Scene Graphs. Scene graphs are directed graphs where nodes represent entities in the scene (e.g. objects, walls, rooms), and edges represent relations (e.g. inclusion, adjacency) among nodes. Dynamic scene graphs (DSGs) extend this notion to represent dynamic scenes with moving agents (e.g. humans, robots), and to include actionable information that supports planning and decision-making (e.g. spatio-temporal relations, topology at different levels of abstraction). Our second contribution is to provide the first fully automatic Spatial PerceptIon eNgine(SPIN) to build a DSG from visual-inertial data. We integrate state-of-the-art techniques for object and human detection and pose estimation, and we describe how to robustly infer object, robot, and human nodes in crowded scenes. To the best of our knowledge, this is the first paper that reconciles visual-inertial SLAM and dense human mesh tracking. Moreover, we provide algorithms to obtain hierarchical representations of indoor environments (e.g. places, structures, rooms) and their relations. Our third contribution is to demonstrate the proposed spatial perception engine in a photo-realistic Unity-based simulator, where we assess its robustness and expressiveness. Finally, we discuss the implications of our proposal on modern robotics applications. 3D Dynamic Scene Graphs can have a profound impact on planning and decision-making, human-robot interaction, long-term autonomy, and scene prediction. A video abstract is available at https://youtu.be/SWbofjhyPzI

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A training-free pipeline that disambiguates text queries and infers viewpoints improves zero-shot 3D visual grounding, reaching 64.06% Acc@0.5 on ScanRefer.

  2. Robotic Contextual Awareness for Human-Robot Collaboration and Environmental Understanding

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Novel person re-identification with continual adaptation plus submap LiDAR SLAM, ground-aware filtering, Gaussian Scan Context, and multi-modal semantic mapping improve robotic contextual awareness for HRC and navigation.

  3. From Scene-Centric to Observer-Centric: Modeling Observer-Aware Relations for 3D Scene Graph Generation

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    TAD decouples 3DSGG relation reasoning by predicate transformation behavior to achieve yaw-robust predictions without rotation augmentation.

  4. Rheos: Modelling Continuous Motion Dynamics in Hierarchical 3D Scene Graphs

    cs.RO 2026-03 conditional novelty 6.0 of 10

    Rheos embeds online semi-wrapped Gaussian mixture models of directional motion into 3D scene graph navigational nodes and outperforms discrete histogram baselines on continuous and discrete metrics.

  5. Interleaved LLM and Motion Planning for Generalized Multi-Object Collection in Large Scene Graphs

    cs.RO 2025-07 conditional novelty 6.0 of 10

    Inter-LLM interleaves LLM task selection with motion-planning cost feedback through a multimodal similarity estimator, reporting 30% lower mission cost than SayPlan and MoMa-LLM in a simulated household setting.

  6. Pixels-to-Graph: Real-time Integration of Building Information Models and Scene Graphs for Semantic-Geometric Human-Robot Understanding

    cs.RO 2025-06 conditional novelty 6.0 of 10

    Pix2G generates hierarchical scene graphs with object, scene, room, and building layers, on CPU only, by combining 2D object detection, GAN-based map denoising, and BEV room segmentation.

  7. SLAM in Low-Light Environments: Project Report

    cs.RO 2026-07 conditional novelty 5.0 of 10

    On five LaMARia sequences, only Kimera-VIO tracked all to completion; DPVO/DPV-SLAM never lost tracking but had roughly 100 m absolute error.

  8. Towards Terrain-Aware Task-Driven 3D Scene Graph Generation in Outdoor Environments

    cs.RO 2025-06 conditional novelty 5.0 of 10

    An outdoor 3D scene graph pipeline using LiDAR-camera fusion, CLIP embeddings, and per-terrain Voronoi graphs is demonstrated on a campus dataset with qualitative results.

  9. Advances in 4D Representation: Geometry, Motion, and Interaction

    cs.CV 2025-10 conditional novelty 4.0 of 10

    A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.

  10. Open-Vocabulary Indoor Object Grounding with 3D Hierarchical Scene Graph

    cs.CV 2025-07 conditional novelty 4.0 of 10

    OVIGo-3DHSG builds a five-level scene graph (building, floor, room, location, object) and uses LLM reasoning over relevant subgraphs to ground open-vocabulary objects in multi-floor indoor scenes.

Pith tools