Pith. sign in

REVIEW 6 cited by

Zero-1-to-3: Zero-shot One Image to 3D Object

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.11328 v1 pith:I5NA7QDT submitted 2023-03-20 cs.CV cs.GRcs.RO

classification cs.CVcs.GRcs.RO
keywords cameradiffusionimageimagesobjectdatasetlearnmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Zero-1-to-3, a framework for changing the camera viewpoint of an object given just a single RGB image. To perform novel view synthesis in this under-constrained setting, we capitalize on the geometric priors that large-scale diffusion models learn about natural images. Our conditional diffusion model uses a synthetic dataset to learn controls of the relative camera viewpoint, which allow new images to be generated of the same object under a specified camera transformation. Even though it is trained on a synthetic dataset, our model retains a strong zero-shot generalization ability to out-of-distribution datasets as well as in-the-wild images, including impressionist paintings. Our viewpoint-conditioned diffusion approach can further be used for the task of 3D reconstruction from a single image. Qualitative and quantitative experiments show that our method significantly outperforms state-of-the-art single-view 3D reconstruction and novel view synthesis models by leveraging Internet-scale pre-training.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation

    cs.GR 2026-07 conditional novelty 6.0 of 10

    A feed-forward model reconstructs a layered, simulation-ready 3D Gaussian world from multi-view driving video in ~1.5 s, with quality approaching per-scene optimized reconstruction.

  2. 3D-Telepathy: Reconstructing 3D Objects from EEG Signals

    cs.CV 2025-06 conditional novelty 6.0 of 10

    3D-Telepathy reconstructs 3D objects from EEG signals by combining a dual self-attention EEG encoder with stable diffusion and variational score distillation into a NeRF, and reports best 2D-frame metrics among compar...

  3. A Definition and Roadmap for World Models

    cs.AI 2026-07 conditional novelty 5.0 of 10

    A perspective article defining world models as finite-resource compression of physical state transitions and outlining a roadmap toward physical AGI via unified representations and interactive simulators.

  4. Few-step Flow for 3D Generation via Marginal-Data Transport Distillation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    MDT-dist distills a pretrained 3D flow model into a 1-2 step generator using velocity matching plus velocity distillation, cutting TRELLIS inference from 6.1s to 0.68s while approximately preserving generation quality.

  5. VEIGAR: View-consistent Explicit Inpainting and Geometry Alignment for 3D object Removal

    cs.GR 2025-06 conditional novelty 5.0 of 10

    VEIGAR is a pipeline for 3D object removal in Gaussian Splatting that uses deep stereo depth projection and a scale-invariant depth loss to achieve faster training and comparable quality to prior state-of-the-art.

  6. 2D Instance Editing in 3D Space

    cs.CV 2025-07 reject novelty 4.0 of 10

    A 2D-to-3D-to-2D editing system that segments an object, reconstructs it as 3D Gaussians, deforms it under a rigidity constraint, and inpaints it back into the original image.

Pith tools