REVIEW 14 cited by
ConceptFusion: Open-set Multimodal 3D Mapping
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Building 3D maps of the environment is central to robot navigation, planning, and interaction with objects in a scene. Most existing approaches that integrate semantic concepts with 3D maps largely remain confined to the closed-set setting: they can only reason about a finite set of concepts, pre-defined at training time. Further, these maps can only be queried using class labels, or in recent work, using text prompts. We address both these issues with ConceptFusion, a scene representation that is (1) fundamentally open-set, enabling reasoning beyond a closed set of concepts and (ii) inherently multimodal, enabling a diverse range of possible queries to the 3D map, from language, to images, to audio, to 3D geometry, all working in concert. ConceptFusion leverages the open-set capabilities of today's foundation models pre-trained on internet-scale data to reason about concepts across modalities such as natural language, images, and audio. We demonstrate that pixel-aligned open-set features can be fused into 3D maps via traditional SLAM and multi-view fusion approaches. This enables effective zero-shot spatial reasoning, not needing any additional training or finetuning, and retains long-tailed concepts better than supervised approaches, outperforming them by more than 40% margin on 3D IoU. We extensively evaluate ConceptFusion on a number of real-world datasets, simulated home environments, a real-world tabletop manipulation task, and an autonomous driving platform. We showcase new avenues for blending foundation models with 3D open-set multimodal mapping. For more information, visit our project page https://concept-fusion.github.io or watch our 5-minute explainer video https://www.youtube.com/watch?v=rkXgws8fiDs
Forward citations
Cited by 14 Pith papers
-
3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
A feature-diversity-guided token compression pipeline for 3D VLMs keeps 94.7% of original QA accuracy at 128 tokens and runs 1.92x faster than uncompressed LLaVA-3D.
-
Memory for Attention: Language-Conditioned Re-Perception with a Vision--Language--Motion Map
Persistent map memory schedules budgeted re-perception better than memoryless VLM priors, with gain equal to Var(√λ), and language-conditioned VLMM needs both open-vocabulary relevance and per-instance dynamics.
-
Vision-Language-Motion Maps: An Open-Vocabulary, Uncertainty-Aware, Queryable Motion Attribute for 3D Scene Maps
VLMM is a 3D map representation where each object carries a fused, uncertainty-aware motion attribute (language-based movability prior + observed geometric motion) that makes motion queries such as 'what is moving' an...
-
Just-In-Time Scene Graph Growth: Combating Perceptual Saturation in Long-Horizon Robotics
JITOMA keeps 3D scene graph nodes dormant until a natural-language query activates only task-relevant objects, bounding active graph size and captioning latency while preserving grounding accuracy.
-
CoSAG: Compact Semantic Anchor Gaussians via Training-Free Rate-Distortion Coding
Training-free closed-form lift plus spatially predictive entropy coding of Gaussian-to-anchor bindings yields sub-megabyte open-vocabulary 3D fields that match or beat prior accuracy.
-
MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes
MV-GEL uses a learned view selector and a fine-tuned vision-language segmentation model to localize text-described faces and edges on 3D meshes.
-
FunFact: Building Probabilistic Functional 3D Scene Graphs via Factor-Graph Reasoning
Joint factor-graph inference over LLM- and geometry-constrained functional edges yields higher recall and substantially lower calibration error than independent pairwise functional scene-graph methods.
-
VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation
A vision-language model writes a simulation program—grounded by segmentation and 3D tools—that predicts physically plausible futures from an image and caption, outperforming video generators on modified benchmarks.
-
Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes
OVODA combines a 3DETR-style detector with a frozen foundation model to detect novel objects and attributes in 3D scenes, and the OVAD dataset adds spatial and motion attribute labels to nuScenes.
-
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
A benchmark that scores vision-language models by reconstructing the 3D scene behind an image as executable Blender code finds the models fail mainly on spatial precision, not tool usage.
-
Satellites Reveal Mobility: A Commuting Origin-destination Flow Generator for Global Cities
Satellite imagery plus population is enough to generate commuting origin-destination flows that closely match models using detailed sociodemographic and point-of-interest data.
-
OV-MAP: Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
OV-MAP projects 2D masks into 3D and uses mesh-area voting to create zero-shot, open-vocabulary 3D instance segmentation maps.
-
Details Matter for Indoor Open-vocabulary 3D Instance Segmentation
A carefully engineered pipeline of 2D grounding, 3D tracking, proposal merging, and Alpha-CLIP classification with a standardized similarity filter achieves state-of-the-art open-vocabulary 3D instance segmentation on...
- CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity
Discussion (0). Continue with ORCID to comment.