REVIEW 11 cited by
Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present Magic123, a two-stage coarse-to-fine approach for high-quality, textured 3D meshes generation from a single unposed image in the wild using both2D and 3D priors. In the first stage, we optimize a neural radiance field to produce a coarse geometry. In the second stage, we adopt a memory-efficient differentiable mesh representation to yield a high-resolution mesh with a visually appealing texture. In both stages, the 3D content is learned through reference view supervision and novel views guided by a combination of 2D and 3D diffusion priors. We introduce a single trade-off parameter between the 2D and 3D priors to control exploration (more imaginative) and exploitation (more precise) of the generated geometry. Additionally, we employ textual inversion and monocular depth regularization to encourage consistent appearances across views and to prevent degenerate solutions, respectively. Magic123 demonstrates a significant improvement over previous image-to-3D techniques, as validated through extensive experiments on synthetic benchmarks and diverse real-world images. Our code, models, and generated 3D assets are available at https://github.com/guochengqian/Magic123.
Forward citations
Cited by 11 Pith papers
-
3D PixBrush: Image-Guided Local Texture Synthesis
A method that uses a reference image to automatically predict a localization mask and synthesize a matching local texture on a 3D mesh.
-
ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation
ELSA3D introduces elastic semantic anchoring via sparse anchor tokens and a scale-aware octree tokenizer to unify 3D generation and captioning at reduced computational cost.
-
MoGA: 3D Generative Avatar Prior for Monocular Gaussian Avatar Reconstruction
Fitting a learned 3D Gaussian avatar prior to six diffusion-hallucinated views reconstructs an animatable, high-fidelity avatar from a single image.
-
DreamGrasp: Zero-Shot 3D Multi-Object Reconstruction from Partial-View Images for Robotic Manipulation
A zero-shot pipeline reconstructs per-instance 3D geometry in cluttered scenes from two partial RGB views by combining diffusion-based score distillation with learned instance features and text-guided refinement.
-
LoomNet: Enhancing Multi-View Image Generation via Latent Space Weaving
LoomNet generates consistent multi-view images from a single input by fusing per-view latent predictions onto shared triplane features before rendering.
-
Controllable 3D Placement of Objects with Scene-Aware Diffusion Models
Projecting a color-coded 3D bounding box into a ControlNet conditioning map gives diffusion inpainting models precise control over vehicle orientation and placement in driving scenes.
-
Zero-P-to-3: Zero-Shot Partial-View Images to 3D Object
Zero-P-to-3 fuses multi-view diffusion, a restoration prior, and a coarse 3D Gaussian rendering in DDIM sampling, then refines with rotated views, and reports improved invisible-region reconstruction from partial-view...
-
SynthDrive: Scalable Real2Sim2Real Sensor Simulation Pipeline for High-Fidelity Asset Generation and Driving Data Synthesis
SynthDrive automatically mines images of rare objects, reconstructs them as 3D assets from a single view, and synthesizes driving footage that modestly improves detection of those objects.
-
Splat4D: Diffusion-Enhanced 4D Gaussian Splatting for Temporally and Spatially Consistent Content Creation
Splat4D generates temporally and spatially consistent 4D Gaussian scenes from monocular video by coupling multi-view diffusion, image enhancement, and uncertainty-guided video diffusion refinement.
-
TextMesh4D: Zero-shot Text-to-4D Mesh Generation
TextMesh4D generates text-conditioned dynamic meshes by combining a Jacobian Deformation Field, video score distillation, and a local-global semantic regularizer in a zero-shot pipeline.
-
DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation
A multi-view conditioning framework that improves controllable novel view synthesis and 3D reconstruction by injecting fused 3D latents into frozen image and video diffusion models.
Discussion (0). Sign in to comment.