Pith. sign in

REVIEW 2 cited by

A Recipe for Generating 3D Worlds From a Single Image

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.16611 v1 pith:HPCLQW52 submitted 2025-03-20 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords generatingimageworldsapproachinpaintingminimalmodelmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce a recipe for generating immersive 3D worlds from a single image by framing the task as an in-context learning problem for 2D inpainting models. This approach requires minimal training and uses existing generative models. Our process involves two steps: generating coherent panoramas using a pre-trained diffusion model and lifting these into 3D with a metric depth estimator. We then fill unobserved regions by conditioning the inpainting model on rendered point clouds, requiring minimal fine-tuning. Tested on both synthetic and real images, our method produces high-quality 3D environments suitable for VR display. By explicitly modeling the 3D structure of the generated environment from the start, our approach consistently outperforms state-of-the-art, video synthesis-based methods along multiple quantitative image quality metrics. Project Page: https://katjaschwarz.github.io/worlds/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Voyager is a video diffusion model that jointly generates RGB and depth from one image, enabling direct 3D scene reconstruction and long-range camera exploration.

  2. HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A staged pipeline generates layered, mesh-based 3D worlds from text or images by combining panoramic diffusion, semantic layer decomposition, and video-based expansion.

Pith tools