Pith. sign in

REVIEW 6 cited by

IM-3D: Iterative Multiview Diffusion and Reconstruction for High-Quality 3D Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.08682 v1 pith:7DSGQMBI submitted 2024-02-13 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords reconstructioncombineddirectlydistillationgenerationgeneratorgeneratorshigh-quality
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Most text-to-3D generators build upon off-the-shelf text-to-image models trained on billions of images. They use variants of Score Distillation Sampling (SDS), which is slow, somewhat unstable, and prone to artifacts. A mitigation is to fine-tune the 2D generator to be multi-view aware, which can help distillation or can be combined with reconstruction networks to output 3D objects directly. In this paper, we further explore the design space of text-to-3D models. We significantly improve multi-view generation by considering video instead of image generators. Combined with a 3D reconstruction algorithm which, by using Gaussian splatting, can optimize a robust image-based loss, we directly produce high-quality 3D outputs from the generated views. Our new method, IM-3D, reduces the number of evaluations of the 2D generator network 10-100x, resulting in a much more efficient pipeline, better quality, fewer geometric inconsistencies, and higher yield of usable 3D assets.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Single-image novel view synthesis is decomposed into panorama outpainting plus keyframe-conditioned video diffusion, producing loop-consistent scene tours.

  2. Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A video diffusion backbone fine-tuned on 4M densely captioned 360-degree renderings generates spatially consistent multi-view images for 3D assets from image plus detailed text input.

  3. Efficient multi-view training for 3D Gaussian Splatting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Training 3D Gaussian Splatting with multiple images per iteration, using partial rendering and a 3D-aware SSIM loss, improves novel-view synthesis quality over single-view training.

  4. FlexPainter: Flexible and Multi-View Consistent Texture Generation

    cs.GR 2025-06 conditional novelty 6.0 of 10

    FlexPainter combines multi-view grid generation, UV-space view synchronization with a learned weighting network, and multi-modal embedding control to generate consistent, high-resolution textures from text and image prompts.

  5. S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A three-stage pipeline generates animatable 3D Gaussian head avatars from one image by diffusion-based splat synthesis, FLAME fitting, and inverse-distance binding with scale adaptation.

  6. DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A multi-view conditioning framework that improves controllable novel view synthesis and 3D reconstruction by injecting fused 3D latents into frozen image and video diffusion models.

Pith tools