Pith. sign in

REVIEW 3 cited by

Hybrid Fourier Score Distillation for Efficient One Image to 3D Object Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.20669 v2 pith:PCU2PYNV submitted 2024-05-31 cs.CV

classification cs.CV
keywords generationpriorsappearancefourierimagemodelsdifferentdiffusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Single image-to-3D generation is pivotal for crafting controllable 3D assets. Given its under-constrained nature, we attempt to leverage 3D geometric priors from a novel view diffusion model and 2D appearance priors from an image generation model to guide the optimization process. We note that there is a disparity between the generation priors of these two diffusion models, leading to their different appearance outputs. Specifically, image generation models tend to deliver more detailed visuals, whereas novel view models produce consistent yet over-smooth results across different views. Directly combining them leads to suboptimal effects due to their appearance conflicts. Hence, we propose a 2D-3D hybrid Fourier Score Distillation objective function, hy-FSD. It optimizes 3D Gaussians using 3D priors in spatial domain to ensure geometric consistency, while exploiting 2D priors in the frequency domain through Fourier transform for better visual quality. hy-FSD can be integrated into existing 3D generation methods and produce significant performance gains. With this technique, we further develop an image-to-3D generation pipeline to create high-quality 3D objects within one minute, named Fourier123. Extensive experiments demonstrate that Fourier123 excels in efficient generation with rapid convergence speed and visually-friendly generation results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. 4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A two-stage cascaded video diffusion model generates 16-view consistent videos from a monocular video, enabling higher-quality 4D content reconstruction.

  2. OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A drag-style motion control method for 360 degree image-to-video generation, built on spherical trajectory estimation and joint fine-tuning of a pretrained video diffusion model.

  3. RIGI: Rectifying Image-to-3D Generation Inconsistency via Uncertainty-aware Learning

    cs.CV 2024-11 conditional novelty 4.0 of 10

    RIGI improves image-to-3D generation by estimating pixel-wise uncertainty from the difference between two 3D Gaussian models and using it to reweight the reconstruction loss, reducing artifacts from inconsistent multi...

Pith tools