Pith. sign in

REVIEW 3 cited by

DreamSparse: Escaping from Plato's Cave with 2D Frozen Diffusion Model Given Sparse Views

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.03414 v4 pith:SGXMT7TG submitted 2023-06-06 cs.CV cs.AIcs.GR

classification cs.CVcs.AIcs.GR
keywords imagesdiffusionnovelviewviewsdreamsparsemodelpre-trained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Synthesizing novel view images from a few views is a challenging but practical problem. Existing methods often struggle with producing high-quality results or necessitate per-object optimization in such few-view settings due to the insufficient information provided. In this work, we explore leveraging the strong 2D priors in pre-trained diffusion models for synthesizing novel view images. 2D diffusion models, nevertheless, lack 3D awareness, leading to distorted image synthesis and compromising the identity. To address these problems, we propose DreamSparse, a framework that enables the frozen pre-trained diffusion model to generate geometry and identity-consistent novel view image. Specifically, DreamSparse incorporates a geometry module designed to capture 3D features from sparse views as a 3D prior. Subsequently, a spatial guidance model is introduced to convert these 3D feature maps into spatial information for the generative process. This information is then used to guide the pre-trained diffusion model, enabling it to generate geometrically consistent images without tuning it. Leveraging the strong image priors in the pre-trained diffusion models, DreamSparse is capable of synthesizing high-quality novel views for both object and scene-level images and generalising to open-set images. Experimental results demonstrate that our framework can effectively synthesize novel view images from sparse views and outperforms baselines in both trained and open-set category images. More results can be found on our project page: https://sites.google.com/view/dreamsparse-webpage.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion

    cs.CV 2025-01 conditional novelty 6.0 of 10

    MVGD jointly generates novel-view images and scale-consistent depth maps with a pixel-level diffusion model, reporting state-of-the-art scores on several view synthesis and depth benchmarks.

  2. AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers

    cs.CV 2024-11 conditional novelty 6.0 of 10

    AC3D improves camera control in video diffusion transformers by conditioning only early denoising steps and the first 8 of 32 blocks, and by adding 20K static-camera dynamic videos to training.

  3. PaintScene4D: Consistent 4D Scene Generation from Text Prompts

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A training-free pipeline that turns one text-to-video clip into a multi-view 4D scene renderable along user-chosen camera paths.

Pith tools