Pith. sign in

REVIEW 6 cited by

3DShape2VecSet: A 3D Shape Representation for Neural Fields and Generative Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.11445 v3 pith:JZDAWJ3S submitted 2023-01-26 cs.CV cs.GR

classification cs.CVcs.GR
keywords representationfieldsneuralshapegenerationgenerativelatentmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce 3DShape2VecSet, a novel shape representation for neural fields designed for generative diffusion models. Our shape representation can encode 3D shapes given as surface models or point clouds, and represents them as neural fields. The concept of neural fields has previously been combined with a global latent vector, a regular grid of latent vectors, or an irregular grid of latent vectors. Our new representation encodes neural fields on top of a set of vectors. We draw from multiple concepts, such as the radial basis function representation and the cross attention and self-attention function, to design a learnable representation that is especially suitable for processing with transformers. Our results show improved performance in 3D shape encoding and 3D shape generative modeling tasks. We demonstrate a wide variety of generative applications: unconditioned generation, category-conditioned generation, text-conditioned generation, point-cloud completion, and image-conditioned generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Part-level 3D Object Generation via Dual Volume Packing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    From a single image, a 3D latent diffusion model generates all parts of an object at once by packing the part structure into two non-overlapping volumes.

  2. Zero-1-to-G: Taming Pretrained 2D Diffusion Model for Direct 3D Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A single image can be turned into a 3D Gaussian splat model by fine-tuning a pretrained 2D diffusion model to output decomposed multi-view splatter attribute images.

  3. Light Transport-aware Diffusion Posterior Sampling for Single-View Reconstruction of 3D Volumes

    cs.CV 2025-01 reject novelty 6.0 of 10

    A diffusion-prior-guided differentiable volume renderer (PDPS) reconstructs 3D clouds from a single image, using a new monoplanar latent representation and a synthetic cloud dataset.

  4. Coherent 3D Scene Diffusion From a Single RGB Image

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single RGB image is converted into a coherent 3D scene by denoising all object poses and shapes simultaneously with a diffusion model.

  5. Wavelet Latent Diffusion (Wala): Billion-Parameter 3D Generative Model with Compact Wavelet Encodings

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Wavelet Latent Diffusion (WaLa) shrinks 3D shapes to 6,912-variable latent codes and trains billion-parameter diffusion models that generate 256^3 geometry in 2-4 seconds, claiming state-of-the-art results.

  6. Visualizing the Invisible: A Generative AR System for Intuitive Multi-Modal Sensor Data Presentation

    cs.HC 2024-12 conditional novelty 5.0 of 10

    Vivar uses barycentric interpolation in a pre-trained CLIP embedding space to generate AR visualizations of multi-modal sensor data, with caching that speeds generation 11x.

Pith tools