Pith. sign in

REVIEW 11 cited by

Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.17843 v2 pith:4MSJCO2V submitted 2023-06-30 cs.CV

classification cs.CV
keywords magic123priorsdiffusiongeneratedgenerationgeometryhigh-qualityimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Magic123, a two-stage coarse-to-fine approach for high-quality, textured 3D meshes generation from a single unposed image in the wild using both2D and 3D priors. In the first stage, we optimize a neural radiance field to produce a coarse geometry. In the second stage, we adopt a memory-efficient differentiable mesh representation to yield a high-resolution mesh with a visually appealing texture. In both stages, the 3D content is learned through reference view supervision and novel views guided by a combination of 2D and 3D diffusion priors. We introduce a single trade-off parameter between the 2D and 3D priors to control exploration (more imaginative) and exploitation (more precise) of the generated geometry. Additionally, we employ textual inversion and monocular depth regularization to encourage consistent appearances across views and to prevent degenerate solutions, respectively. Magic123 demonstrates a significant improvement over previous image-to-3D techniques, as validated through extensive experiments on synthetic benchmarks and diverse real-world images. Our code, models, and generated 3D assets are available at https://github.com/guochengqian/Magic123.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 75 citations worldwide. Full citation record

  1. 3D PixBrush: Image-Guided Local Texture Synthesis

    cs.GR 2025-07 conditional novelty 7.0 of 10

    A method that uses a reference image to automatically predict a localization mask and synthesize a matching local texture on a 3D mesh.

  2. ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    ELSA3D introduces elastic semantic anchoring via sparse anchor tokens and a scale-aware octree tokenizer to unify 3D generation and captioning at reduced computational cost.

  3. MoGA: 3D Generative Avatar Prior for Monocular Gaussian Avatar Reconstruction

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Fitting a learned 3D Gaussian avatar prior to six diffusion-hallucinated views reconstructs an animatable, high-fidelity avatar from a single image.

  4. DreamGrasp: Zero-Shot 3D Multi-Object Reconstruction from Partial-View Images for Robotic Manipulation

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A zero-shot pipeline reconstructs per-instance 3D geometry in cluttered scenes from two partial RGB views by combining diffusion-based score distillation with learned instance features and text-guided refinement.

  5. LoomNet: Enhancing Multi-View Image Generation via Latent Space Weaving

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LoomNet generates consistent multi-view images from a single input by fusing per-view latent predictions onto shared triplane features before rendering.

  6. Controllable 3D Placement of Objects with Scene-Aware Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Projecting a color-coded 3D bounding box into a ControlNet conditioning map gives diffusion inpainting models precise control over vehicle orientation and placement in driving scenes.

  7. Zero-P-to-3: Zero-Shot Partial-View Images to 3D Object

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Zero-P-to-3 fuses multi-view diffusion, a restoration prior, and a coarse 3D Gaussian rendering in DDIM sampling, then refines with rotated views, and reports improved invisible-region reconstruction from partial-view...

  8. SynthDrive: Scalable Real2Sim2Real Sensor Simulation Pipeline for High-Fidelity Asset Generation and Driving Data Synthesis

    cs.CV 2025-09 conditional novelty 5.0 of 10

    SynthDrive automatically mines images of rare objects, reconstructs them as 3D assets from a single view, and synthesizes driving footage that modestly improves detection of those objects.

  9. Splat4D: Diffusion-Enhanced 4D Gaussian Splatting for Temporally and Spatially Consistent Content Creation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Splat4D generates temporally and spatially consistent 4D Gaussian scenes from monocular video by coupling multi-view diffusion, image enhancement, and uncertainty-guided video diffusion refinement.

  10. TextMesh4D: Zero-shot Text-to-4D Mesh Generation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    TextMesh4D generates text-conditioned dynamic meshes by combining a Jacobian Deformation Field, video score distillation, and a local-global semantic regularizer in a zero-shot pipeline.

  11. DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A multi-view conditioning framework that improves controllable novel view synthesis and 3D reconstruction by injecting fused 3D latents into frozen image and video diffusion models.

Pith tools