Pith. sign in

REVIEW 4 cited by

An Object is Worth 64x64 Pixels: Generating 3D Object via Image Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.03178 v1 pith:JZTA4UY3 submitted 2024-08-06 cs.CV cs.GRcs.LG

classification cs.CVcs.GRcs.LG
keywords generationimagemodelsobjectapproachdiffusiongeneratingpatch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce a new approach for generating realistic 3D models with UV maps through a representation termed "Object Images." This approach encapsulates surface geometry, appearance, and patch structures within a 64x64 pixel image, effectively converting complex 3D shapes into a more manageable 2D format. By doing so, we address the challenges of both geometric and semantic irregularity inherent in polygonal meshes. This method allows us to use image generation models, such as Diffusion Transformers, directly for 3D shape generation. Evaluated on the ABO dataset, our generated shapes with patch structures achieve point cloud FID comparable to recent 3D generative models, while naturally supporting PBR material generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LL3M: Large Language 3D Modelers

    cs.GR 2025-08 conditional novelty 6.0 of 10

    A multi-agent LLM system generates editable 3D assets as Blender Python code, using documentation retrieval and visual self-critique to refine results.

  2. Efficient Part-level 3D Object Generation via Dual Volume Packing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    From a single image, a 3D latent diffusion model generates all parts of an object at once by packing the part structure into two non-overlapping volumes.

  3. Zero-1-to-G: Taming Pretrained 2D Diffusion Model for Direct 3D Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A single image can be turned into a 3D Gaussian splat model by fine-tuning a pretrained 2D diffusion model to output decomposed multi-view splatter attribute images.

  4. Wavelet Latent Diffusion (Wala): Billion-Parameter 3D Generative Model with Compact Wavelet Encodings

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Wavelet Latent Diffusion (WaLa) shrinks 3D shapes to 6,912-variable latent codes and trains billion-parameter diffusion models that generate 256^3 geometry in 2-4 seconds, claiming state-of-the-art results.

Pith tools