REVIEW 6 cited by
Zero-1-to-3: Zero-shot One Image to 3D Object
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce Zero-1-to-3, a framework for changing the camera viewpoint of an object given just a single RGB image. To perform novel view synthesis in this under-constrained setting, we capitalize on the geometric priors that large-scale diffusion models learn about natural images. Our conditional diffusion model uses a synthetic dataset to learn controls of the relative camera viewpoint, which allow new images to be generated of the same object under a specified camera transformation. Even though it is trained on a synthetic dataset, our model retains a strong zero-shot generalization ability to out-of-distribution datasets as well as in-the-wild images, including impressionist paintings. Our viewpoint-conditioned diffusion approach can further be used for the task of 3D reconstruction from a single image. Qualitative and quantitative experiments show that our method significantly outperforms state-of-the-art single-view 3D reconstruction and novel view synthesis models by leveraging Internet-scale pre-training.
Forward citations
Cited by 6 Pith papers
-
Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation
A feed-forward model reconstructs a layered, simulation-ready 3D Gaussian world from multi-view driving video in ~1.5 s, with quality approaching per-scene optimized reconstruction.
-
3D-Telepathy: Reconstructing 3D Objects from EEG Signals
3D-Telepathy reconstructs 3D objects from EEG signals by combining a dual self-attention EEG encoder with stable diffusion and variational score distillation into a NeRF, and reports best 2D-frame metrics among compar...
-
A Definition and Roadmap for World Models
A perspective article defining world models as finite-resource compression of physical state transitions and outlining a roadmap toward physical AGI via unified representations and interactive simulators.
-
Few-step Flow for 3D Generation via Marginal-Data Transport Distillation
MDT-dist distills a pretrained 3D flow model into a 1-2 step generator using velocity matching plus velocity distillation, cutting TRELLIS inference from 6.1s to 0.68s while approximately preserving generation quality.
-
VEIGAR: View-consistent Explicit Inpainting and Geometry Alignment for 3D object Removal
VEIGAR is a pipeline for 3D object removal in Gaussian Splatting that uses deep stereo depth projection and a scale-invariant depth loss to achieve faster training and comparable quality to prior state-of-the-art.
-
2D Instance Editing in 3D Space
A 2D-to-3D-to-2D editing system that segments an object, reconstructs it as 3D Gaussians, deforms it under a rigidity constraint, and inpaints it back into the original image.
Discussion (0). Sign in to comment.