Pith. sign in

REVIEW 8 cited by

Dora: Sampling and Benchmarking for 3D Shape Variational Auto-Encoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.17808 v3 pith:WBFRFICA submitted 2024-12-23 cs.CV

classification cs.CV
keywords reconstructionsamplingshapegeometricgenerationqualitystrategycomplexity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Recent 3D content generation pipelines commonly employ Variational Autoencoders (VAEs) to encode shapes into compact latent representations for diffusion-based generation. However, the widely adopted uniform point sampling strategy in Shape VAE training often leads to a significant loss of geometric details, limiting the quality of shape reconstruction and downstream generation tasks. We present Dora-VAE, a novel approach that enhances VAE reconstruction through our proposed sharp edge sampling strategy and a dual cross-attention mechanism. By identifying and prioritizing regions with high geometric complexity during training, our method significantly improves the preservation of fine-grained shape features. Such sampling strategy and the dual attention mechanism enable the VAE to focus on crucial geometric details that are typically missed by uniform sampling approaches. To systematically evaluate VAE reconstruction quality, we additionally propose Dora-bench, a benchmark that quantifies shape complexity through the density of sharp edges, introducing a new metric focused on reconstruction accuracy at these salient geometric features. Extensive experiments on the Dora-bench demonstrate that Dora-VAE achieves comparable reconstruction quality to the state-of-the-art dense XCube-VAE while requiring a latent space at least 8$\times$ smaller (1,280 vs. > 10,000 codes).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unifi3D: A Study on 3D Representations for Generation and Reconstruction in a Common Framework

    cs.GR 2025-09 conditional novelty 6.0 of 10

    SDF grids reconstruct best, Dual Octrees score best on automatic generation metrics, but users prefer SDF output, and reconstruction plus compression errors make up a large share of generation error.

  2. AutoPartGen: Autogressive 3D Part Generation and Discovery

    cs.CV 2025-07 conditional novelty 6.0 of 10

    AutoPartGen generates 3D objects as a sequence of latent-space parts, conditioning each new part on previously generated parts, and reports state-of-the-art part completion on PartObjaverse-Tiny.

  3. Efficient Part-level 3D Object Generation via Dual Volume Packing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    From a single image, a 3D latent diffusion model generates all parts of an object at once by packing the part structure into two non-overlapping volumes.

  4. PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PartCrafter generates several separable 3D part meshes at once from a single image by fine-tuning a pretrained 3D diffusion transformer with part identity tokens and local-global attention.

  5. ViewCraft3D: High-Fidelity and View-Consistent 3D Vector Graphics Synthesis

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A two-stage method that fits 3D Bézier curves to a reconstructed mesh and refines them with a 3D diffusion prior, producing view-consistent 3D vector graphics from a single image in about 30 minutes.

  6. Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Direct3D-S2 uses a new Spatial Sparse Attention mechanism to train a sparse-volume diffusion transformer at 1024^3 resolution on 8 GPUs.

  7. Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes Modeling

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A sparse, deformable marching-cubes representation plus a sparse-convolution VAE convert rough meshes into watertight 1024^3 surfaces with less detail loss, faster training, and better downstream 3D generation than pr...

  8. UniTEX: Universal High Fidelity Generative Texturing for 3D Shapes

    cs.CV 2025-05 conditional novelty 5.0 of 10

    UniTEX generates textures for 3D shapes by predicting continuous volumetric texture functions, bypassing UV maps.

Pith tools