Pith. sign in

REVIEW 8 cited by

3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.12957 v2 pith:6WVE7WPP submitted 2024-09-19 cs.CV cs.GR

classification cs.CVcs.GR
keywords assetsdtopia-xlgenerativehigh-qualitydiffusionprimitiveexistingmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The increasing demand for high-quality 3D assets across various industries necessitates efficient and automated 3D content creation. Despite recent advancements in 3D generative models, existing methods still face challenges with optimization speed, geometric fidelity, and the lack of assets for physically based rendering (PBR). In this paper, we introduce 3DTopia-XL, a scalable native 3D generative model designed to overcome these limitations. 3DTopia-XL leverages a novel primitive-based 3D representation, PrimX, which encodes detailed shape, albedo, and material field into a compact tensorial format, facilitating the modeling of high-resolution geometry with PBR assets. On top of the novel representation, we propose a generative framework based on Diffusion Transformer (DiT), which comprises 1) Primitive Patch Compression, 2) and Latent Primitive Diffusion. 3DTopia-XL learns to generate high-quality 3D assets from textual or visual inputs. We conduct extensive qualitative and quantitative experiments to demonstrate that 3DTopia-XL significantly outperforms existing methods in generating high-quality 3D assets with fine-grained textures and materials, efficiently bridging the quality gap between generative models and real-world applications.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VideoMat: Extracting PBR Materials from Video Diffusion Models

    cs.GR 2025-06 conditional novelty 7.0 of 10

    VideoMat uses a finetuned video diffusion model, intrinsic decomposition, and differentiable path tracing to extract PBR material maps for known 3D geometry from text or image prompts.

  2. DualMat: PBR Material Estimation via Coherent Dual-Path Diffusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    DualMat is a dual-path diffusion model combining an albedo-optimized pretrained latent path with a material-specialized compact latent path, using feature distillation and rectified flow to estimate PBR materials from...

  3. Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A video-to-4D model that encodes mesh animations into compact Gaussian variation latents and diffuses them conditioned on the video and a canonical Gaussian splat.

  4. 4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    4D-Animal fits SMAL animal models to video using silhouette, part, pixel, and tracking losses from off-the-shelf 2D models, removing the need for sparse keypoint annotations.

  5. PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PoseMaster produces a 3D character mesh from one image and a target 3D skeleton, preserving identity and pose in a single unified model, and it outperforms two-stage 2D-to-3D baselines on the VRoid pose canonicalizati...

  6. Efficient Part-level 3D Object Generation via Dual Volume Packing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    From a single image, a 3D latent diffusion model generates all parts of an object at once by packing the part structure into two non-overlapping volumes.

  7. Squeeze3D: Your 3D Generation Model is Secretly an Extreme Neural Compressor

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Two small mapping networks connect a frozen 3D encoder to a frozen 3D generator, so the generator decompresses objects from latent codes as small as 3 KB, achieving up to 2187x compression on meshes.

  8. ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ShapeLLM-Omni unifies text, image, and 3D generation and understanding in one autoregressive LLM using discrete 3D tokens and a new 3D-Alpaca training dataset.

Pith tools