Pith. sign in

REVIEW 5 cited by

ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.11417 v1 pith:CKVSR433 submitted 2023-08-22 cs.CV

classification cs.CV
keywords scannetimagesscenescenessemanticannotatedbenchmarkcapture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present ScanNet++, a large-scale dataset that couples together capture of high-quality and commodity-level geometry and color of indoor scenes. Each scene is captured with a high-end laser scanner at sub-millimeter resolution, along with registered 33-megapixel images from a DSLR camera, and RGB-D streams from an iPhone. Scene reconstructions are further annotated with an open vocabulary of semantics, with label-ambiguous scenarios explicitly annotated for comprehensive semantic understanding. ScanNet++ enables a new real-world benchmark for novel view synthesis, both from high-quality RGB capture, and importantly also from commodity-level images, in addition to a new benchmark for 3D semantic scene understanding that comprehensively encapsulates diverse and ambiguous semantic labeling scenarios. Currently, ScanNet++ contains 460 scenes, 280,000 captured DSLR images, and over 3.7M iPhone RGBD frames.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vision as Unified Multimodal Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A single unified multimodal model matches leading task-specialized vision systems across detection, segmentation, dense geometry, and multi-view 3D by casting all outputs as native text or image generation.

  2. Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Multi-view relational distillation transfers geometric knowledge to vision-language models by matching cross-view patch similarity matrices, improving spatial reasoning with minimal overhead.

  3. Discovering and using Spelke segments

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SpelkeNet, a self-supervised video world model, discovers Spelke segments in static images by aggregating motion correlations across imagined pokes.

  4. ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding

    cs.CV 2025-01 conditional novelty 6.0 of 10

    ViGiL3D is a 350-prompt diagnostic dataset showing that existing 3D visual grounding models lose 20 or more points on linguistically diverse prompts compared to ScanRefer.

  5. THUD++: Large-Scale Dynamic Indoor Scene Dataset and Benchmark for Mobile Robots

    cs.RO 2024-12 conditional novelty 5.0 of 10

    THUD++ is a 13-scene dynamic indoor RGB-D and trajectory dataset with benchmarks showing existing algorithms struggle in crowded mobile-robot environments.

Pith tools