Pith. sign in

REVIEW 5 cited by

RichDreamer: A Generalizable Normal-Depth Diffusion Model for Detail Richness in Text-to-3D

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.16918 v2 pith:HFHQ4VIY submitted 2023-11-28 cs.CV cs.AI

classification cs.CVcs.AI
keywords diffusionmodelgeneralizableimagesnormalsalbedodetailexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Lifting 2D diffusion for 3D generation is a challenging problem due to the lack of geometric prior and the complex entanglement of materials and lighting in natural images. Existing methods have shown promise by first creating the geometry through score-distillation sampling (SDS) applied to rendered surface normals, followed by appearance modeling. However, relying on a 2D RGB diffusion model to optimize surface normals is suboptimal due to the distribution discrepancy between natural images and normals maps, leading to instability in optimization. In this paper, recognizing that the normal and depth information effectively describe scene geometry and be automatically estimated from images, we propose to learn a generalizable Normal-Depth diffusion model for 3D generation. We achieve this by training on the large-scale LAION dataset together with the generalizable image-to-depth and normal prior models. In an attempt to alleviate the mixed illumination effects in the generated materials, we introduce an albedo diffusion model to impose data-driven constraints on the albedo component. Our experiments show that when integrated into existing text-to-3D pipelines, our models significantly enhance the detail richness, achieving state-of-the-art results. Our project page is https://aigc3d.github.io/richdreamer/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PartGen: Part-level 3D Generation and Reconstruction with Multi-View Diffusion Models

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A multi-view diffusion pipeline that segments 3D objects into parts, completes occluded or invisible parts, and reconstructs them into a compositional 3D asset.

  2. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  3. PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    PixGS is a single-stage pixel-space diffusion model that directly produces high-quality 3D Gaussian Splats from text or images in ~1s, outperforming multi-stage latent methods on standard benchmarks.

  4. MCMat: Multiview-Consistent and Physically Accurate PBR Material Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A two-stage diffusion transformer pipeline generates multi-view-consistent, relightable PBR material maps for 3D meshes, reporting state-of-the-art FID/KID scores on 70 Objaverse models and higher user-study ratings t...

  5. DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DrivingRecon predicts 4D Gaussians of street scenes from surround-view video in one forward pass, using a novel Prune and Dilate Block to reduce redundant overlapping points.

Pith tools