REVIEW 3 cited by
The Bare Necessities: Designing Simple, Effective Open-Vocabulary Scene Graphs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
3D open-vocabulary scene graph methods are a promising map representation for embodied agents, however many current approaches are computationally expensive. In this paper, we reexamine the critical design choices established in previous works to optimize both efficiency and performance. We propose a general scene graph framework and conduct three studies that focus on image pre-processing, feature fusion, and feature selection. Our findings reveal that commonly used image pre-processing techniques provide minimal performance improvement while tripling computation (on a per object view basis). We also show that averaging feature labels across different views significantly degrades performance. We study alternative feature selection strategies that enhance performance without adding unnecessary computational costs. Based on our findings, we introduce a computationally balanced approach for 3D point cloud segmentation with per-object features. The approach matches state-of-the-art classification accuracy while achieving a threefold reduction in computation.
Forward citations
Cited by 3 Pith papers
-
Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured Environments
Prior-SG uses an LLM-generated probabilistic prior graph and graph-cut inference to segment robot maps into task-relevant functional regions.
-
GraphPilot: Grounded Scene Graph Conditioning for Language-Based Autonomous Driving
Training language-based driving agents with serialized traffic scene graphs improves closed-loop driving scores by up to roughly 20%, and the benefit persists even when graphs are omitted at inference.
-
Towards Terrain-Aware Task-Driven 3D Scene Graph Generation in Outdoor Environments
An outdoor 3D scene graph pipeline using LiDAR-camera fusion, CLIP embeddings, and per-terrain Voronoi graphs is demonstrated on a campus dataset with qualitative results.
Discussion (0). Sign in to comment.