Pith. sign in

REVIEW 4 cited by

Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.11290 v3 pith:5FEEX3KG submitted 2023-06-20 cs.CV

classification cs.CV
keywords datasetagentsscenesscenesynthetictrainedgeneralizationnavigation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We contribute the Habitat Synthetic Scene Dataset, a dataset of 211 high-quality 3D scenes, and use it to test navigation agent generalization to realistic 3D environments. Our dataset represents real interiors and contains a diverse set of 18,656 models of real-world objects. We investigate the impact of synthetic 3D scene dataset scale and realism on the task of training embodied agents to find and navigate to objects (ObjectGoal navigation). By comparing to synthetic 3D scene datasets from prior work, we find that scale helps in generalization, but the benefits quickly saturate, making visual fidelity and correlation to real-world scenes more important. Our experiments show that agents trained on our smaller-scale dataset can match or outperform agents trained on much larger datasets. Surprisingly, we observe that agents trained on just 122 scenes from our dataset outperform agents trained on 10,000 scenes from the ProcTHOR-10K dataset in terms of zero-shot generalization in real-world scanned environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VideoMat: Extracting PBR Materials from Video Diffusion Models

    cs.GR 2025-06 conditional novelty 7.0 of 10

    VideoMat uses a finetuned video diffusion model, intrinsic decomposition, and differentiable path tracing to extract PBR material maps for known 3D geometry from text or image prompts.

  2. WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    WildShadowRemover fine-tunes a pretrained video diffusion model with LoRA plus detail-injection and depth conditioning to produce temporally consistent shadow-free videos, trained on a new synthetic dataset.

  3. TRELLIS-Enhanced Surface Features for Comprehensive Intracranial Aneurysm Analysis

    cs.CV 2025-09 conditional novelty 6.0 of 10

    TRELLIS-derived surface features improve aneurysm classification, segmentation, and hemodynamic simulation, including a 15% lower blood-flow prediction error.

  4. Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 23-dimension benchmark finds that even the best VLMs score near random on motion trajectory, temporal extension, and several prediction tasks, far below humans, suggesting weak internal world models.

Pith tools