Pith. sign in

REVIEW 3 cited by

Consistent-1-to-3: Consistent Image to 3D View Synthesis via Geometry-aware Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.03020 v2 pith:GQPYZJPZ submitted 2023-10-04 cs.CV

classification cs.CV
keywords imagemodelsnovelviewviewsapproachesattentionconsistency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Zero-shot novel view synthesis (NVS) from a single image is an essential problem in 3D object understanding. While recent approaches that leverage pre-trained generative models can synthesize high-quality novel views from in-the-wild inputs, they still struggle to maintain 3D consistency across different views. In this paper, we present Consistent-1-to-3, which is a generative framework that significantly mitigates this issue. Specifically, we decompose the NVS task into two stages: (i) transforming observed regions to a novel view, and (ii) hallucinating unseen regions. We design a scene representation transformer and view-conditioned diffusion model for performing these two stages respectively. Inside the models, to enforce 3D consistency, we propose to employ epipolor-guided attention to incorporate geometry constraints, and multi-view attention to better aggregate multi-view information. Finally, we design a hierarchy generation paradigm to generate long sequences of consistent views, allowing a full 360-degree observation of the provided object image. Qualitative and quantitative evaluation over multiple datasets demonstrates the effectiveness of the proposed mechanisms against state-of-the-art approaches. Our project page is at https://jianglongye.com/consistent123/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Consistent Flow Distillation for Text-to-3D Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Consistent Flow Distillation (CFD) guides 3D generation by denoising rendered views with a noise field that is consistent across camera views on the object surface.

  2. LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LiftImage3D generates small-motion video clips from one image, registers them with MASt3R, and fits a distortion-aware 3D Gaussian field whose canonical scene renders new views.

  3. GraphicsDreamer: Image to 3D Generation with Physical Consistency

    cs.GR 2024-12 conditional novelty 5.0 of 10

    From one image, GraphicsDreamer generates multi-view color, geometry, and PBR material maps, then reconstructs a clean, UV-unwrapped 3D mesh usable in graphics engines.

Pith tools