Pith. sign in

REVIEW 8 cited by

DMV3D: Denoising Multi-View Diffusion using 3D Large Reconstruction Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.09217 v1 pith:4OMAOM4M submitted 2023-11-15 cs.CV

classification cs.CV
keywords reconstructiondmv3dmulti-viewdiffusiongenerationmodeldenoisediverse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We propose \textbf{DMV3D}, a novel 3D generation approach that uses a transformer-based 3D large reconstruction model to denoise multi-view diffusion. Our reconstruction model incorporates a triplane NeRF representation and can denoise noisy multi-view images via NeRF reconstruction and rendering, achieving single-stage 3D generation in $\sim$30s on single A100 GPU. We train \textbf{DMV3D} on large-scale multi-view image datasets of highly diverse objects using only image reconstruction losses, without accessing 3D assets. We demonstrate state-of-the-art results for the single-image reconstruction problem where probabilistic modeling of unseen object parts is required for generating diverse reconstructions with sharp textures. We also show high-quality text-to-3D generation results outperforming previous 3D diffusion models. Our project website is at: https://justimyhxu.github.io/projects/dmv3d/ .

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos

    cs.GR 2025-06 conditional novelty 7.0 of 10

    A single feed-forward transformer predicts per-pixel deformable 3D Gaussians with dense scene flow from a posed monocular video, enabling real-time dynamic view synthesis and 3D tracking.

  2. Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A video diffusion backbone fine-tuned on 4M densely captioned 360-degree renderings generates spatially consistent multi-view images for 3D assets from image plus detailed text input.

  3. DualMat: PBR Material Estimation via Coherent Dual-Path Diffusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    DualMat is a dual-path diffusion model combining an albedo-optimized pretrained latent path with a material-specialized compact latent path, using feature distillation and rectified flow to estimate PBR materials from...

  4. Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Can3Tok tokenizes scene-level 3D Gaussian splats into canonical latent tokens with normalization and saliency filtering, enabling reconstruction and text/image-to-3D generation.

  5. EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    EarthCrafter generates 600-meter-scale 3D Earth scenes using separate latent diffusion models for structure and texture, conditioned on semantics, images, or nothing.

  6. FlexPainter: Flexible and Multi-View Consistent Texture Generation

    cs.GR 2025-06 conditional novelty 6.0 of 10

    FlexPainter combines multi-view grid generation, UV-space view synchronization with a learned weighting network, and multi-modal embedding control to generate consistent, high-resolution textures from text and image prompts.

  7. SMPL Normal Map Is All You Need for Single-view Textured Human Reconstruction

    cs.CV 2025-06 conditional novelty 5.0 of 10

    By fine-tuning the LGM object reconstruction model with SMPL normal maps and a normal-map constraint, SEHR reconstructs full clothed 3D humans from one image in a single forward pass.

  8. ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ShapeLLM-Omni unifies text, image, and 3D generation and understanding in one autoregressive LLM using discrete 3D tokens and a new 3D-Alpaca training dataset.

Pith tools