Pith. sign in

REVIEW 5 cited by

One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.16928 v1 pith:SCFDKRVK submitted 2023-06-29 cs.CV cs.AIcs.RO

classification cs.CVcs.AIcs.RO
keywords imagereconstructionsinglemethoddiffusioninputmeshmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Single image 3D reconstruction is an important but challenging task that requires extensive knowledge of our natural world. Many existing methods solve this problem by optimizing a neural radiance field under the guidance of 2D diffusion models but suffer from lengthy optimization time, 3D inconsistency results, and poor geometry. In this work, we propose a novel method that takes a single image of any object as input and generates a full 360-degree 3D textured mesh in a single feed-forward pass. Given a single image, we first use a view-conditioned 2D diffusion model, Zero123, to generate multi-view images for the input view, and then aim to lift them up to 3D space. Since traditional reconstruction methods struggle with inconsistent multi-view predictions, we build our 3D reconstruction module upon an SDF-based generalizable neural surface reconstruction method and propose several critical training strategies to enable the reconstruction of 360-degree meshes. Without costly optimizations, our method reconstructs 3D shapes in significantly less time than existing methods. Moreover, our method favors better geometry, generates more 3D consistent results, and adheres more closely to the input image. We evaluate our approach on both synthetic data and in-the-wild images and demonstrate its superiority in terms of both mesh quality and runtime. In addition, our approach can seamlessly support the text-to-3D task by integrating with off-the-shelf text-to-image diffusion models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. 3D PixBrush: Image-Guided Local Texture Synthesis

    cs.GR 2025-07 conditional novelty 7.0 of 10

    A method that uses a reference image to automatically predict a localization mask and synthesize a matching local texture on a 3D mesh.

  2. CLONE: Continuous Latent Optimization for Normal Estimation via 3D Gaussian Splatting

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    CLONE estimates surface normals from images via a differentiable 3D Gaussian splatting loop, using photometric loss instead of normal labels.

  3. Efficient Part-level 3D Object Generation via Dual Volume Packing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    From a single image, a 3D latent diffusion model generates all parts of an object at once by packing the part structure into two non-overlapping volumes.

  4. EgoAnimate: Generating Human Animations from Egocentric top-down Views

    cs.CV 2025-07 conditional novelty 4.0 of 10

    EgoAnimate synthesizes a frontal T-pose image from an egocentric top-down photo using a fine-tuned Stable Diffusion model, then animates it with off-the-shelf image-to-motion methods to produce an animatable avatar.

  5. RGBTrack: Fast, Robust Depth-Free 6D Pose Estimation and Tracking

    cs.CV 2025-06

Pith tools