REVIEW 10 cited by
3D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and Humans
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present a unified representation for actionable spatial perception: 3D Dynamic Scene Graphs. Scene graphs are directed graphs where nodes represent entities in the scene (e.g. objects, walls, rooms), and edges represent relations (e.g. inclusion, adjacency) among nodes. Dynamic scene graphs (DSGs) extend this notion to represent dynamic scenes with moving agents (e.g. humans, robots), and to include actionable information that supports planning and decision-making (e.g. spatio-temporal relations, topology at different levels of abstraction). Our second contribution is to provide the first fully automatic Spatial PerceptIon eNgine(SPIN) to build a DSG from visual-inertial data. We integrate state-of-the-art techniques for object and human detection and pose estimation, and we describe how to robustly infer object, robot, and human nodes in crowded scenes. To the best of our knowledge, this is the first paper that reconciles visual-inertial SLAM and dense human mesh tracking. Moreover, we provide algorithms to obtain hierarchical representations of indoor environments (e.g. places, structures, rooms) and their relations. Our third contribution is to demonstrate the proposed spatial perception engine in a photo-realistic Unity-based simulator, where we assess its robustness and expressiveness. Finally, we discuss the implications of our proposal on modern robotics applications. 3D Dynamic Scene Graphs can have a profound impact on planning and decision-making, human-robot interaction, long-term autonomy, and scene prediction. A video abstract is available at https://youtu.be/SWbofjhyPzI
Forward citations
Cited by 10 Pith papers
-
TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding
A training-free pipeline that disambiguates text queries and infers viewpoints improves zero-shot 3D visual grounding, reaching 64.06% Acc@0.5 on ScanRefer.
-
Robotic Contextual Awareness for Human-Robot Collaboration and Environmental Understanding
Novel person re-identification with continual adaptation plus submap LiDAR SLAM, ground-aware filtering, Gaussian Scan Context, and multi-modal semantic mapping improve robotic contextual awareness for HRC and navigation.
-
From Scene-Centric to Observer-Centric: Modeling Observer-Aware Relations for 3D Scene Graph Generation
TAD decouples 3DSGG relation reasoning by predicate transformation behavior to achieve yaw-robust predictions without rotation augmentation.
-
Rheos: Modelling Continuous Motion Dynamics in Hierarchical 3D Scene Graphs
Rheos embeds online semi-wrapped Gaussian mixture models of directional motion into 3D scene graph navigational nodes and outperforms discrete histogram baselines on continuous and discrete metrics.
-
Interleaved LLM and Motion Planning for Generalized Multi-Object Collection in Large Scene Graphs
Inter-LLM interleaves LLM task selection with motion-planning cost feedback through a multimodal similarity estimator, reporting 30% lower mission cost than SayPlan and MoMa-LLM in a simulated household setting.
-
Pixels-to-Graph: Real-time Integration of Building Information Models and Scene Graphs for Semantic-Geometric Human-Robot Understanding
Pix2G generates hierarchical scene graphs with object, scene, room, and building layers, on CPU only, by combining 2D object detection, GAN-based map denoising, and BEV room segmentation.
-
SLAM in Low-Light Environments: Project Report
On five LaMARia sequences, only Kimera-VIO tracked all to completion; DPVO/DPV-SLAM never lost tracking but had roughly 100 m absolute error.
-
Towards Terrain-Aware Task-Driven 3D Scene Graph Generation in Outdoor Environments
An outdoor 3D scene graph pipeline using LiDAR-camera fusion, CLIP embeddings, and per-terrain Voronoi graphs is demonstrated on a campus dataset with qualitative results.
-
Advances in 4D Representation: Geometry, Motion, and Interaction
A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.
-
Open-Vocabulary Indoor Object Grounding with 3D Hierarchical Scene Graph
OVIGo-3DHSG builds a five-level scene graph (building, floor, room, location, object) and uses LLM reasoning over relevant subgraphs to ground open-vocabulary objects in multi-floor indoor scenes.
Discussion (0). Continue with ORCID to comment.