REVIEW 12 cited by
NVS-Solver: Video Diffusion Model as Zero-Shot Novel View Synthesizer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
By harnessing the potent generative capabilities of pre-trained large video diffusion models, we propose NVS-Solver, a new novel view synthesis (NVS) paradigm that operates \textit{without} the need for training. NVS-Solver adaptively modulates the diffusion sampling process with the given views to enable the creation of remarkable visual experiences from single or multiple views of static scenes or monocular videos of dynamic scenes. Specifically, built upon our theoretical modeling, we iteratively modulate the score function with the given scene priors represented with warped input views to control the video diffusion process. Moreover, by theoretically exploring the boundary of the estimation error, we achieve the modulation in an adaptive fashion according to the view pose and the number of diffusion steps. Extensive evaluations on both static and dynamic scenes substantiate the significant superiority of our NVS-Solver over state-of-the-art methods both quantitatively and qualitatively. \textit{ Source code in } \href{https://github.com/ZHU-Zhiyu/NVS_Solver}{https://github.com/ZHU-Zhiyu/NVS$\_$Solver}.
Forward citations
Cited by 12 Pith papers
-
UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models
UniGeo unifies geometric guidance across three levels in video models to reduce geometric drift and improve consistency in camera-controllable image editing.
-
FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control
A single video-diffusion framework composites both static images and dynamic footage along user-defined trajectories by transporting canonical foreground latents directly into the background latent sequence.
-
UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models
A video-generation world model that warps positional encodings of memory frames to target viewpoints achieves state-of-the-art long-term consistency and camera control.
-
InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem
A training-free, near-zero-overhead inverse solver for novel-view video generation and inpainting that projects masks into continuous multi-channel latent masks and applies DDS with conjugate gradient in latent space.
-
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
EPiC trains a 30M-parameter visibility-aware ControlNet on mask-based anchor videos from 5,000 in-the-wild videos and 500 steps, reaching SOTA camera accuracy on RealEstate10K and MiraData.
-
F3D-Gaus: Feed-forward 3D-aware Generation on ImageNet with Cycle-Aggregative Gaussian Splatting
F3D-Gaus predicts a pixel-aligned 3D Gaussian representation from a single RGB-D image and uses cycle-aggregative self-supervision plus video-prior refinement to render consistent novel views from monocular training d...
-
StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
StreetCrafter conditions a video diffusion model on LiDAR point cloud renderings to synthesize controllable street views, and distills it into a real-time 3D Gaussian representation.
-
LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors
LiftImage3D generates small-motion video clips from one image, registers them with MASt3R, and fits a distortion-aware 3D Gaussian field whose canonical scene renders new views.
-
Trajectory Attention for Fine-grained Video Motion Control
An auxiliary trajectory attention branch, added to temporal attention in video diffusion models, improves camera motion control precision while preserving generation quality.
-
PE-Field 4D: Video Generation Models as Canvas
Warping reference tokens' positional encodings into the target view, with depth offsets and frame-level compression fixes, improves geometry-aware camera control in video diffusion transformers.
-
Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation
A feed-forward system that generates object-level and scene-level 3D Gaussian scenes from text in about eight seconds by diffusing multi-view RGB-D latent codes and decoding them into pixel-aligned 3D Gaussians.
-
DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation
A multi-view conditioning framework that improves controllable novel view synthesis and 3D reconstruction by injecting fused 3D latents into frozen image and video diffusion models.
Discussion (0). Continue with ORCID to comment.