Pith. sign in

REVIEW 3 cited by

3D-LDM: Neural Implicit 3D Shape Generation with Latent Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.00842 v2 pith:6SGNMO67 submitted 2022-12-01 cs.CV

classification cs.CV
keywords generationdiffusionlatentshapesallowsimageimplicitmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models have shown great promise for image generation, beating GANs in terms of generation diversity, with comparable image quality. However, their application to 3D shapes has been limited to point or voxel representations that can in practice not accurately represent a 3D surface. We propose a diffusion model for neural implicit representations of 3D shapes that operates in the latent space of an auto-decoder. This allows us to generate diverse and high quality 3D surfaces. We additionally show that we can condition our model on images or text to enable image-to-3D generation and text-to-3D generation using CLIP embeddings. Furthermore, adding noise to the latent codes of existing shapes allows us to explore shape variations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unifi3D: A Study on 3D Representations for Generation and Reconstruction in a Common Framework

    cs.GR 2025-09 conditional novelty 6.0 of 10

    SDF grids reconstruct best, Dual Octrees score best on automatic generation metrics, but users prefer SDF output, and reconstruction plus compression errors make up a large share of generation error.

  2. Collaborative Multi-Modal Coding for High-Quality 3D Generation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    TriMM fuses RGB, RGB-D, and point-cloud encoding into a shared triplane latent space and generates 3D assets from a single image with a latent diffusion model.

  3. GANFusion: Feed-Forward Text-to-3D with Diffusion in GAN Space

    cs.CV 2024-12 conditional novelty 4.0 of 10

    Text-conditioned 3D human generation is achieved by distilling a 2D-supervised GAN's triplane space into a text-conditioned diffusion model, avoiding 3D supervision and test-time optimization.

Pith tools