Pith. sign in

REVIEW 13 cited by

Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.14832 v2 pith:3QAD2RBF submitted 2024-05-23 cs.CV

classification cs.CV
keywords direct3dmodelscalablediffusiongenerationimage-to-3dimageslatent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generating high-quality 3D assets from text and images has long been challenging, primarily due to the absence of scalable 3D representations capable of capturing intricate geometry distributions. In this work, we introduce Direct3D, a native 3D generative model scalable to in-the-wild input images, without requiring a multiview diffusion model or SDS optimization. Our approach comprises two primary components: a Direct 3D Variational Auto-Encoder (D3D-VAE) and a Direct 3D Diffusion Transformer (D3D-DiT). D3D-VAE efficiently encodes high-resolution 3D shapes into a compact and continuous latent triplane space. Notably, our method directly supervises the decoded geometry using a semi-continuous surface sampling strategy, diverging from previous methods relying on rendered images as supervision signals. D3D-DiT models the distribution of encoded 3D latents and is specifically designed to fuse positional information from the three feature maps of the triplane latent, enabling a native 3D generative model scalable to large-scale 3D datasets. Additionally, we introduce an innovative image-to-3D generation pipeline incorporating semantic and pixel-level image conditions, allowing the model to produce 3D shapes consistent with the provided conditional image input. Extensive experiments demonstrate the superiority of our large-scale pre-trained Direct3D over previous image-to-3D approaches, achieving significantly better generation quality and generalization ability, thus establishing a new state-of-the-art for 3D content creation. Project page: https://nju-3dv.github.io/projects/Direct3D/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space

    cs.CV 2025-08 conditional novelty 7.0 of 10

    A training-free 3D editing method that inverts a source asset into TRELLIS latent space and replaces latents plus attention K/V tokens in unedited regions during re-denosing.

  2. EditFlow3D: Automated Local Editing of 3D Assets with Trajectory Preservation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Mask-guided differential flow with a soft preservation loss enables training-free local 3D editing that keeps unedited regions close to the source asset.

  3. UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routing

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Routing each 3D voxel to its most informative conditioning image via a model-intrinsic Voxel Reference Score unlocks robust unconstrained multi-image 3D generation without retraining.

  4. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  5. TInR: Exploring Tool-Internalized Reasoning in Large Language Models

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    TInR-U internalizes tool knowledge into LLMs via bidirectional alignment, supervised fine-tuning, and reinforcement learning, outperforming standard tool-integrated reasoning in both in-domain and out-of-domain evaluations.

  6. Unifi3D: A Study on 3D Representations for Generation and Reconstruction in a Common Framework

    cs.GR 2025-09 conditional novelty 6.0 of 10

    SDF grids reconstruct best, Dual Octrees score best on automatic generation metrics, but users prefer SDF output, and reconstruction plus compression errors make up a large share of generation error.

  7. Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A video-to-4D model that encodes mesh animations into compact Gaussian variation latents and diffuses them conditioned on the video and a canonical Gaussian splat.

  8. PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PoseMaster produces a 3D character mesh from one image and a target 3D skeleton, preserving identity and pose in a single unified model, and it outperforms two-stage 2D-to-3D baselines on the VRoid pose canonicalizati...

  9. Efficient Part-level 3D Object Generation via Dual Volume Packing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    From a single image, a 3D latent diffusion model generates all parts of an object at once by packing the part structure into two non-overlapping volumes.

  10. Squeeze3D: Your 3D Generation Model is Secretly an Extreme Neural Compressor

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Two small mapping networks connect a frozen 3D encoder to a frozen 3D generator, so the generator decompresses objects from latent codes as small as 3 KB, achieving up to 2187x compression on meshes.

  11. Few-step Flow for 3D Generation via Marginal-Data Transport Distillation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    MDT-dist distills a pretrained 3D flow model into a 1-2 step generator using velocity matching plus velocity distillation, cutting TRELLIS inference from 6.1s to 0.68s while approximately preserving generation quality.

  12. XSpecMesh: Quality-Preserving Auto-Regressive Mesh Generation Acceleration via Multi-Head Speculative Decoding

    cs.GR 2025-07 conditional novelty 5.0 of 10

    XSpecMesh speeds up auto-regressive mesh generation by about 1.7x using multi-head speculative decoding with cross-attention heads and a probability threshold verification, while keeping output quality close to the ba...

  13. ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ShapeLLM-Omni unifies text, image, and 3D generation and understanding in one autoregressive LLM using discrete 3D tokens and a new 3D-Alpaca training dataset.

Pith tools