Pith. sign in

REVIEW 14 cited by

ConceptFusion: Open-set Multimodal 3D Mapping

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.07241 v3 pith:5SRC3FYE submitted 2023-02-14 cs.CV cs.AIcs.RO

classification cs.CVcs.AIcs.RO
keywords conceptsopen-setconceptfusionmapsapproachesmultimodalaudioenabling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Building 3D maps of the environment is central to robot navigation, planning, and interaction with objects in a scene. Most existing approaches that integrate semantic concepts with 3D maps largely remain confined to the closed-set setting: they can only reason about a finite set of concepts, pre-defined at training time. Further, these maps can only be queried using class labels, or in recent work, using text prompts. We address both these issues with ConceptFusion, a scene representation that is (1) fundamentally open-set, enabling reasoning beyond a closed set of concepts and (ii) inherently multimodal, enabling a diverse range of possible queries to the 3D map, from language, to images, to audio, to 3D geometry, all working in concert. ConceptFusion leverages the open-set capabilities of today's foundation models pre-trained on internet-scale data to reason about concepts across modalities such as natural language, images, and audio. We demonstrate that pixel-aligned open-set features can be fused into 3D maps via traditional SLAM and multi-view fusion approaches. This enables effective zero-shot spatial reasoning, not needing any additional training or finetuning, and retains long-tailed concepts better than supervised approaches, outperforming them by more than 40% margin on 3D IoU. We extensively evaluate ConceptFusion on a number of real-world datasets, simulated home environments, a real-world tabletop manipulation task, and an autonomous driving platform. We showcase new avenues for blending foundation models with 3D open-set multimodal mapping. For more information, visit our project page https://concept-fusion.github.io or watch our 5-minute explainer video https://www.youtube.com/watch?v=rkXgws8fiDs

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. 3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A feature-diversity-guided token compression pipeline for 3D VLMs keeps 94.7% of original QA accuracy at 128 tokens and runs 1.92x faster than uncompressed LLaVA-3D.

  2. Memory for Attention: Language-Conditioned Re-Perception with a Vision--Language--Motion Map

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Persistent map memory schedules budgeted re-perception better than memoryless VLM priors, with gain equal to Var(√λ), and language-conditioned VLMM needs both open-vocabulary relevance and per-instance dynamics.

  3. Vision-Language-Motion Maps: An Open-Vocabulary, Uncertainty-Aware, Queryable Motion Attribute for 3D Scene Maps

    cs.RO 2026-07 conditional novelty 6.0 of 10

    VLMM is a 3D map representation where each object carries a fused, uncertainty-aware motion attribute (language-based movability prior + observed geometric motion) that makes motion queries such as 'what is moving' an...

  4. Just-In-Time Scene Graph Growth: Combating Perceptual Saturation in Long-Horizon Robotics

    cs.CV 2026-07 conditional novelty 6.0 of 10

    JITOMA keeps 3D scene graph nodes dormant until a natural-language query activates only task-relevant objects, bounding active graph size and captioning latency while preserving grounding accuracy.

  5. CoSAG: Compact Semantic Anchor Gaussians via Training-Free Rate-Distortion Coding

    cs.CV 2026-07 accept novelty 6.0 of 10

    Training-free closed-form lift plus spatially predictive entropy coding of Gaussian-to-anchor bindings yields sub-megabyte open-vocabulary 3D fields that match or beat prior accuracy.

  6. MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    MV-GEL uses a learned view selector and a fine-tuned vision-language segmentation model to localize text-described faces and edges on 3D meshes.

  7. FunFact: Building Probabilistic Functional 3D Scene Graphs via Factor-Graph Reasoning

    cs.CV 2026-04 accept novelty 6.0 of 10

    Joint factor-graph inference over LLM- and geometry-constrained functional edges yields higher recall and substantially lower calibration error than independent pairwise functional scene-graph methods.

  8. VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A vision-language model writes a simulation program—grounded by segmentation and 3D tools—that predicts physically plausible futures from an image and caption, outperforming video generators on modified benchmarks.

  9. Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes

    cs.CV 2025-08 conditional novelty 6.0 of 10

    OVODA combines a 3DETR-style detector with a frozen foundation model to detect novel objects and attributes in 3D scenes, and the OVAD dataset adds spatial and motion attribute labels to nuScenes.

  10. IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A benchmark that scores vision-language models by reconstructing the 3D scene behind an image as executable Blender code finds the models fail mainly on spatial precision, not tool usage.

  11. Satellites Reveal Mobility: A Commuting Origin-destination Flow Generator for Global Cities

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Satellite imagery plus population is enough to generate commuting origin-destination flows that closely match models using detailed sociodemographic and point-of-interest data.

  12. OV-MAP: Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots

    cs.CV 2025-06 conditional novelty 5.0 of 10

    OV-MAP projects 2D masks into 3D and uses mesh-area voting to create zero-shot, open-vocabulary 3D instance segmentation maps.

  13. Details Matter for Indoor Open-vocabulary 3D Instance Segmentation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A carefully engineered pipeline of 2D grounding, 3D tracking, proposal merging, and Alpha-CLIP classification with a standardized similarity filter achieves state-of-the-art open-vocabulary 3D instance segmentation on...

  14. CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity

    cs.RO 2025-06

Pith tools