REVIEW 6 cited by
3DShape2VecSet: A 3D Shape Representation for Neural Fields and Generative Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce 3DShape2VecSet, a novel shape representation for neural fields designed for generative diffusion models. Our shape representation can encode 3D shapes given as surface models or point clouds, and represents them as neural fields. The concept of neural fields has previously been combined with a global latent vector, a regular grid of latent vectors, or an irregular grid of latent vectors. Our new representation encodes neural fields on top of a set of vectors. We draw from multiple concepts, such as the radial basis function representation and the cross attention and self-attention function, to design a learnable representation that is especially suitable for processing with transformers. Our results show improved performance in 3D shape encoding and 3D shape generative modeling tasks. We demonstrate a wide variety of generative applications: unconditioned generation, category-conditioned generation, text-conditioned generation, point-cloud completion, and image-conditioned generation.
Forward citations
Cited by 6 Pith papers
-
Efficient Part-level 3D Object Generation via Dual Volume Packing
From a single image, a 3D latent diffusion model generates all parts of an object at once by packing the part structure into two non-overlapping volumes.
-
Zero-1-to-G: Taming Pretrained 2D Diffusion Model for Direct 3D Generation
A single image can be turned into a 3D Gaussian splat model by fine-tuning a pretrained 2D diffusion model to output decomposed multi-view splatter attribute images.
-
Light Transport-aware Diffusion Posterior Sampling for Single-View Reconstruction of 3D Volumes
A diffusion-prior-guided differentiable volume renderer (PDPS) reconstructs 3D clouds from a single image, using a new monoplanar latent representation and a synthetic cloud dataset.
-
Coherent 3D Scene Diffusion From a Single RGB Image
A single RGB image is converted into a coherent 3D scene by denoising all object poses and shapes simultaneously with a diffusion model.
-
Wavelet Latent Diffusion (Wala): Billion-Parameter 3D Generative Model with Compact Wavelet Encodings
Wavelet Latent Diffusion (WaLa) shrinks 3D shapes to 6,912-variable latent codes and trains billion-parameter diffusion models that generate 256^3 geometry in 2-4 seconds, claiming state-of-the-art results.
-
Visualizing the Invisible: A Generative AR System for Intuitive Multi-Modal Sensor Data Presentation
Vivar uses barycentric interpolation in a pre-trained CLIP embedding space to generate AR visualizations of multi-modal sensor data, with caching that speeds generation 11x.
Discussion (0). Continue with ORCID to comment.