Pith. sign in

REVIEW 5 cited by

VolumeDiffusion: Flexible Text-to-3D Generation with Efficient Volumetric Encoder

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.11459 v3 pith:SHYA434D submitted 2023-12-18 cs.CV

classification cs.CV
keywords generationmodelobjecttext-to-3dvolumesdiffusionefficientencoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces a pioneering 3D volumetric encoder designed for text-to-3D generation. To scale up the training data for the diffusion model, a lightweight network is developed to efficiently acquire feature volumes from multi-view images. The 3D volumes are then trained on a diffusion model for text-to-3D generation using a 3D U-Net. This research further addresses the challenges of inaccurate object captions and high-dimensional feature volumes. The proposed model, trained on the public Objaverse dataset, demonstrates promising outcomes in producing diverse and recognizable samples from text prompts. Notably, it empowers finer control over object part characteristics through textual cues, fostering model creativity by seamlessly combining multiple concepts within a single object. This research significantly contributes to the progress of 3D generation by introducing an efficient, flexible, and scalable representation methodology.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A causal transformer with 3D RoPE generates vector-quantized 3D Gaussian latent grids autoregressively, enabling unconditional synthesis, completion, and open-ended outpainting of indoor scenes.

  2. Incorporating Pre-trained Diffusion Models in Solving the Schr\"odinger Bridge Problem

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Schrödinger Bridge models can be trained with diffusion-style mean, terminus, and flow-matching losses and initialized from pretrained diffusion models, improving image generation and unpaired translation.

  3. Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A video-to-4D model that encodes mesh animations into compact Gaussian variation latents and diffuses them conditioned on the video and a canonical Gaussian splat.

  4. Efficient Part-level 3D Object Generation via Dual Volume Packing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    From a single image, a 3D latent diffusion model generates all parts of an object at once by packing the part structure into two non-overlapping volumes.

  5. Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation

    cs.GR 2025-08 unverdicted novelty 4.0 of 10

    A visual prompt engineering system for text-to-3D generation uses multi-view MLLM scoring and interactive visualizations to help designers create models faster, with 70.5% time reduction and higher quality ratings (4....

Pith tools