Pith. sign in

REVIEW 13 cited by

Hydra: A Real-time Spatial Perception System for 3D Scene Graph Construction and Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.13360 v2 pith:DGKDUXYU submitted 2022-01-31 cs.RO

classification cs.RO
keywords scenegraphgraphsperceptionreal-timespatialalgorithmsbuild
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

3D scene graphs have recently emerged as a powerful high-level representation of 3D environments. A 3D scene graph describes the environment as a layered graph where nodes represent spatial concepts at multiple levels of abstraction and edges represent relations between concepts. While 3D scene graphs can serve as an advanced "mental model" for robots, how to build such a rich representation in real-time is still uncharted territory. This paper describes a real-time Spatial Perception System, a suite of algorithms to build a 3D scene graph from sensor data in real-time. Our first contribution is to develop real-time algorithms to incrementally construct the layers of a scene graph as the robot explores the environment; these algorithms build a local Euclidean Signed Distance Function (ESDF) around the current robot location, extract a topological map of places from the ESDF, and then segment the places into rooms using an approach inspired by community-detection techniques. Our second contribution is to investigate loop closure detection and optimization in 3D scene graphs. We show that 3D scene graphs allow defining hierarchical descriptors for loop closure detection; our descriptors capture statistics across layers in the scene graph, ranging from low-level visual appearance to summary statistics about objects and places. We then propose the first algorithm to optimize a 3D scene graph in response to loop closures; our approach relies on embedded deformation graphs to simultaneously correct all layers of the scene graph. We implement the proposed Spatial Perception System into a architecture named Hydra, that combines fast early and mid-level perception processes with slower high-level perception. We evaluate Hydra on simulated and real data and show it is able to reconstruct 3D scene graphs with an accuracy comparable with batch offline methods despite running online.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility

    cs.CV 2026-07 accept novelty 6.5 of 10

    Frozen VGGT layers contain hierarchical co-visibility signals that a <7.5M MoE head extracts to raise Co-VisiON pairwise IoU* by >25% and multiview by ~10% over prior work.

  2. Just-In-Time Scene Graph Growth: Combating Perceptual Saturation in Long-Horizon Robotics

    cs.CV 2026-07 conditional novelty 6.0 of 10

    JITOMA keeps 3D scene graph nodes dormant until a natural-language query activates only task-relevant objects, bounding active graph size and captioning latency while preserving grounding accuracy.

  3. Relational Semantic Reasoning on 3D Scene Graphs for Open World Interactive Object Search

    cs.RO 2026-03 accept novelty 6.0 of 10

    SCOUT matches LLM planners on open-world interactive object search by scoring 3D scene-graph nodes with lightweight models distilled from LLM relational priors, at far lower compute cost.

  4. Ella: Embodied Social Agents with Lifelong Memory

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Ella, an embodied social agent with a name-centric semantic memory and a spatiotemporal episodic memory, outperformed two re-implemented baselines in social influence and leadership tasks in a 3D simulation.

  5. Pixels-to-Graph: Real-time Integration of Building Information Models and Scene Graphs for Semantic-Geometric Human-Robot Understanding

    cs.RO 2025-06 conditional novelty 6.0 of 10

    Pix2G generates hierarchical scene graphs with object, scene, room, and building layers, on CPU only, by combining 2D object detection, GAN-based map denoising, and BEV room segmentation.

  6. FloAff-Kitchen: Bridging Navigation and Manipulation via Canonical and Progressive Floor Affordance Learning

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Canonical local floor geometry plus progressive skill adaptation predicts robot base placements that raise simulated kitchen mobile-manipulation success over prior FloAff methods.

  7. M2H-MX: Multi-Task Semantic and Geometric Perception for Real-Time Monocular 3D Scene Graph Construction

    cs.CV 2026-03 conditional novelty 5.0 of 10

    A DINOv3-based multi-task depth/semantics front end cuts monocular ScanNet ATE by 60.7% and improves NYUDv2 dense prediction when plugged into an unmodified Mono-Hydra mapping pipeline.

  8. OpenNavMap: Multi-Session Appearance-Based Topometric Mapping for Scalable Visual Navigation

    cs.RO 2026-01 conditional novelty 5.0 of 10

    OpenNavMap shows that an image graph plus on-demand 3D reconstruction can match structure-based maps for visual localization and navigation.

  9. Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering

    cs.RO 2025-10 conditional novelty 5.0 of 10

    A training-free pipeline constructs a spatiotemporal knowledge graph from egocentric video, enabling low-latency, explainable embodied question answering.

  10. N2M: Bridging Navigation and Manipulation by Learning Pose Preference from Rollout

    cs.RO 2025-09 conditional novelty 5.0 of 10

    N2M predicts preferable base poses for manipulation policies from ego-centric point clouds, learned from rollouts, lifting success from 3% to 54% in the PnPCounterToCab task.

  11. IRS: Instance-Level 3D Scene Graphs via Room Prior Guided LiDAR-Camera Fusion

    cs.RO 2025-06 conditional novelty 5.0 of 10

    IRS builds instance-level 3D scene graphs faster by using LiDAR room priors to constrain and parallelize semantic fusion from vision-language models.

  12. Open-Vocabulary Indoor Object Grounding with 3D Hierarchical Scene Graph

    cs.CV 2025-07 conditional novelty 4.0 of 10

    OVIGo-3DHSG builds a five-level scene graph (building, floor, room, location, object) and uses LLM reasoning over relevant subgraphs to ground open-vocabulary objects in multi-floor indoor scenes.

  13. CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity

    cs.RO 2025-06

Pith tools