Pith. sign in

REVIEW 2 cited by

Sampling 3D Gaussian Scenes in Seconds with Latent Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.13099 v1 pith:FSJLZGOU submitted 2024-06-18 cs.CV cs.LG

classification cs.CVcs.LG
keywords scenesdiffusionlatentmodelmodelscomplexgaussiangenerative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a latent diffusion model over 3D scenes, that can be trained using only 2D image data. To achieve this, we first design an autoencoder that maps multi-view images to 3D Gaussian splats, and simultaneously builds a compressed latent representation of these splats. Then, we train a multi-view diffusion model over the latent space to learn an efficient generative model. This pipeline does not require object masks nor depths, and is suitable for complex scenes with arbitrary camera positions. We conduct careful experiments on two large-scale datasets of complex real-world scenes -- MVImgNet and RealEstate10K. We show that our approach enables generating 3D scenes in as little as 0.2 seconds, either from scratch, from a single input view, or from sparse input views. It produces diverse and high-quality results while running an order of magnitude faster than non-latent diffusion models and earlier NeRF-based generative models

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation

    cs.CV 2025-01 conditional novelty 7.0 of 10

    DiffSplat repurposes image diffusion models to generate multi-view Gaussian splat grids, using a rendering loss for 3D consistency and achieving state-of-the-art text- and image-conditioned 3D generation.

  2. Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A feed-forward system that generates object-level and scene-level 3D Gaussian scenes from text in about eight seconds by diffusing multi-view RGB-D latent codes and decoding them into pixel-aligned 3D Gaussians.

Pith tools