Pith. sign in

REVIEW 5 cited by

iFusion: Inverting Diffusion for Pose-Free Reconstruction from Sparse Views

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.17250 v1 pith:TGJMON2L submitted 2023-12-28 cs.CV

classification cs.CV
keywords viewsdiffusionnovelreconstructionmodelobjectposeview
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present iFusion, a novel 3D object reconstruction framework that requires only two views with unknown camera poses. While single-view reconstruction yields visually appealing results, it can deviate significantly from the actual object, especially on unseen sides. Additional views improve reconstruction fidelity but necessitate known camera poses. However, assuming the availability of pose may be unrealistic, and existing pose estimators fail in sparse view scenarios. To address this, we harness a pre-trained novel view synthesis diffusion model, which embeds implicit knowledge about the geometry and appearance of diverse objects. Our strategy unfolds in three steps: (1) We invert the diffusion model for camera pose estimation instead of synthesizing novel views. (2) The diffusion model is fine-tuned using provided views and estimated poses, turned into a novel view synthesizer tailored for the target object. (3) Leveraging registered views and the fine-tuned diffusion model, we reconstruct the 3D object. Experiments demonstrate strong performance in both pose estimation and novel view synthesis. Moreover, iFusion seamlessly integrates with various reconstruction methods and enhances them.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed Images

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A dual-stream diffusion model that generates novel views and condition-view camera poses together, removing the need for external pose estimation in multi-view novel view synthesis.

  2. Sparse-view Pose Estimation and Reconstruction via Analysis by Generative Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SparseAGS jointly refines initial camera poses and reconstructs 3D from sparse views using multi-view SDS diffusion priors and explicit outlier removal.

  3. UNOPose: Unseen Object Pose Estimation with an Unposed RGB-D Reference Image

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A single unposed RGB-D reference image is enough to estimate the 6D pose of an unseen object, outperforming prior reference-based methods on BOP datasets.

  4. Pragmatist: Multiview Conditional Diffusion Models for High-Fidelity 3D Reconstruction from Unposed Sparse Views

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Pragmatist turns sparse unposed photos of an object into a high-fidelity 3D mesh by generating consistent canonical views with a diffusion model, reconstructing a triplane mesh, then refining camera poses and texture ...

  5. Sparse-View 3D Reconstruction: Recent Advances and Open Challenges

    cs.CV 2025-07 conditional novelty 3.0 of 10

    A comprehensive survey that organizes sparse-view 3D reconstruction methods into geometry-based, NeRF, 3DGS, and diffusion-based categories, with benchmarks and open challenges.

Pith tools