REVIEW 4 cited by
An Object is Worth 64x64 Pixels: Generating 3D Object via Image Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce a new approach for generating realistic 3D models with UV maps through a representation termed "Object Images." This approach encapsulates surface geometry, appearance, and patch structures within a 64x64 pixel image, effectively converting complex 3D shapes into a more manageable 2D format. By doing so, we address the challenges of both geometric and semantic irregularity inherent in polygonal meshes. This method allows us to use image generation models, such as Diffusion Transformers, directly for 3D shape generation. Evaluated on the ABO dataset, our generated shapes with patch structures achieve point cloud FID comparable to recent 3D generative models, while naturally supporting PBR material generation.
Forward citations
Cited by 4 Pith papers
-
LL3M: Large Language 3D Modelers
A multi-agent LLM system generates editable 3D assets as Blender Python code, using documentation retrieval and visual self-critique to refine results.
-
Efficient Part-level 3D Object Generation via Dual Volume Packing
From a single image, a 3D latent diffusion model generates all parts of an object at once by packing the part structure into two non-overlapping volumes.
-
Zero-1-to-G: Taming Pretrained 2D Diffusion Model for Direct 3D Generation
A single image can be turned into a 3D Gaussian splat model by fine-tuning a pretrained 2D diffusion model to output decomposed multi-view splatter attribute images.
-
Wavelet Latent Diffusion (Wala): Billion-Parameter 3D Generative Model with Compact Wavelet Encodings
Wavelet Latent Diffusion (WaLa) shrinks 3D shapes to 6,912-variable latent codes and trains billion-parameter diffusion models that generate 256^3 geometry in 2-4 seconds, claiming state-of-the-art results.
Discussion (0). Continue with ORCID to comment.