REVIEW 20 cited by
ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce a 3D-aware diffusion model, ZeroNVS, for single-image novel view synthesis for in-the-wild scenes. While existing methods are designed for single objects with masked backgrounds, we propose new techniques to address challenges introduced by in-the-wild multi-object scenes with complex backgrounds. Specifically, we train a generative prior on a mixture of data sources that capture object-centric, indoor, and outdoor scenes. To address issues from data mixture such as depth-scale ambiguity, we propose a novel camera conditioning parameterization and normalization scheme. Further, we observe that Score Distillation Sampling (SDS) tends to truncate the distribution of complex backgrounds during distillation of 360-degree scenes, and propose "SDS anchoring" to improve the diversity of synthesized novel views. Our model sets a new state-of-the-art result in LPIPS on the DTU dataset in the zero-shot setting, even outperforming methods specifically trained on DTU. We further adapt the challenging Mip-NeRF 360 dataset as a new benchmark for single-image novel view synthesis, and demonstrate strong performance in this setting. Our code and data are at http://kylesargent.github.io/zeronvs/
Forward citations
Cited by 20 Pith papers
-
LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding
LSD-3D generates explicit, 3D-consistent driving scenes by combining a generated proxy mesh with geometry-grounded distillation from a 2D diffusion model.
-
GPS as a Control Signal for Image Generation
A diffusion model conditioned on GPS tags and text can generate location-specific images and reconstruct 3D landmarks via score distillation sampling, without explicit pose estimation.
-
Wonderland: Navigating 3D Scenes from a Single Image
A feed-forward pipeline reconstructs 3D Gaussian scenes from single images by regressing 3DGS directly from camera-conditioned video diffusion latents.
-
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
A diffusion model trained on 1 million 360-degree videos synthesizes novel views with camera translation and enables 3D reconstruction from a single image.
-
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space
A single pixel-space diffusion model jointly performs 3D scene reconstruction and generation by supervising flow matching on rendered multi-view images, matching SOTA reconstruction and outperforming latent-space generation.
-
LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis
Initializing a highway encoder-decoder NVS network from VGGT 3D-aware features yields 31.4 PSNR on RealEstate10k with real-time decoding and optional unposed inputs.
-
Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP
Diffusing in a unified 3D representation (geometry plus appearance and semantics) produces more cross-view-consistent 3D Gaussian scenes than 2D latent pipelines.
-
CharacterShot: Controllable and Consistent 4D Character Animation
A new pipeline generates pose-controlled, view-consistent 4D character animations from one reference image and a 2D pose sequence, backed by a new 13,115-character dataset and benchmark.
-
Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
Vid-CamEdit re-synthesizes monocular videos along user-defined camera paths by conditioning a video diffusion model on 2D flows derived from estimated 3D geometry, without training on multi-view video data.
-
Towards In-the-wild 3D Plane Reconstruction from a Single Image
ZeroPlane trains a Transformer plane reconstructor on 560K images spanning 10 indoor and outdoor datasets and outperforms prior methods in zero-shot evaluations on NYUv2, 7-Scenes, ParallelDomain, and ApolloScape.
-
LIRM: Large Inverse Rendering Model for Progressive Reconstruction of Shape, Materials and View-dependent Radiance Fields
A feed-forward transformer reconstructs shape, PBR materials, and view-dependent radiance from 3 to 6 posed images in under a second, rivaling slower optimization-based inverse rendering.
-
Pippo: High-Resolution Multi-View Humans from a Single Image
A single-image multi-view diffusion transformer generates 1K-resolution turnaround views of humans, with attention biasing for many views and a new reprojection-error metric.
-
Matrix3D: Large Photogrammetry Model All-in-One
A single multi-modal diffusion transformer trained with masked learning performs pose estimation, depth prediction, and novel view synthesis in one model, reporting SOTA pose and NVS numbers.
-
UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting Control
UniAvatar integrates FLAME-based 3D motion rendering and SH-based illumination rendering into a diffusion talking-head model, enabling separate or combined control of motion and lighting in generated videos.
-
LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors
LiftImage3D generates small-motion video clips from one image, registers them with MASt3R, and fits a distortion-aware 3D Gaussian field whose canonical scene renders new views.
-
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
See3D proposes a pose-free visual condition for multi-view diffusion trained on web videos, claiming SOTA single- and sparse-view 3D generation, but the evaluation protocol leaks ground-truth information and mixes ben...
-
Novel View Extrapolation with Video Diffusion Priors
A training-free video-diffusion refinement stage makes radiance field and point cloud renderings look more realistic when the camera moves far outside the training views.
-
Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation
A feed-forward system that generates object-level and scene-level 3D Gaussian scenes from text in about eight seconds by diffusing multi-view RGB-D latent codes and decoding them into pixel-aligned 3D Gaussians.
-
DSplats: 3D Generation by Denoising Splats-Based Multiview Diffusion Models
DSplats is a single-stage multiview diffusion model that denoises latents through a 3D Gaussian reconstructor, achieving state-of-the-art single-image-to-3D reconstruction on the Google Scanned Objects benchmark.
-
Dynamic View Synthesis as an Inverse Problem
Dynamic view synthesis from a monocular video is achieved by redesigning the noise initialization of a pretrained video diffusion model using a recursive interpolation and a stochastic latent modulation.
Discussion (0). Continue with ORCID to comment.