Pith. sign in

REVIEW 2 cited by

Customizing Text-to-Image Diffusion with Object Viewpoint Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.12333 v2 pith:WYDICFBI submitted 2024-04-18 cs.CV

classification cs.CV
keywords objectcontrolviewpointcustomizationdiffusionmodeltext-to-imagewhile
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Model customization introduces new concepts to existing text-to-image models, enabling the generation of these new concepts/objects in novel contexts. However, such methods lack accurate camera view control with respect to the new object, and users must resort to prompt engineering (e.g., adding ``top-view'') to achieve coarse view control. In this work, we introduce a new task -- enabling explicit control of the object viewpoint in the customization of text-to-image diffusion models. This allows us to modify the custom object's properties and generate it in various background scenes via text prompts, all while incorporating the object viewpoint as an additional control. This new task presents significant challenges, as one must harmoniously merge a 3D representation from the multi-view images with the 2D pre-trained model. To bridge this gap, we propose to condition the diffusion process on the 3D object features rendered from the target viewpoint. During training, we fine-tune the 3D feature prediction modules to reconstruct the object's appearance and geometry, while reducing overfitting to the input multi-view images. Our method outperforms existing image editing and model customization baselines in preserving the custom object's identity while following the target object viewpoint and the text prompt.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GPS as a Control Signal for Image Generation

    cs.CV 2025-01 conditional novelty 7.0 of 10

    A diffusion model conditioned on GPS tags and text can generate location-specific images and reconstruct 3D landmarks via score distillation sampling, without explicit pose estimation.

  2. PreciseCam: Precise Camera Control for Text-to-Image Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    PreciseCam enables precise camera control (roll, pitch, vFoV, distortion) in text-to-image generation by conditioning SDXL with Perspective Field maps and a new dataset of 57,380 images.

Pith tools