REVIEW 3 cited by
DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Subject-driven text-to-image generation aims to generate customized images of the given subject based on the text descriptions, which has drawn increasing attention. Existing methods mainly resort to finetuning a pretrained generative model, where the identity-relevant information (e.g., the boy) and the identity-irrelevant information (e.g., the background or the pose of the boy) are entangled in the latent embedding space. However, the highly entangled latent embedding may lead to the failure of subject-driven text-to-image generation as follows: (i) the identity-irrelevant information hidden in the entangled embedding may dominate the generation process, resulting in the generated images heavily dependent on the irrelevant information while ignoring the given text descriptions; (ii) the identity-relevant information carried in the entangled embedding can not be appropriately preserved, resulting in identity change of the subject in the generated images. To tackle the problems, we propose DisenBooth, an identity-preserving disentangled tuning framework for subject-driven text-to-image generation. Specifically, DisenBooth finetunes the pretrained diffusion model in the denoising process. Different from previous works that utilize an entangled embedding to denoise each image, DisenBooth instead utilizes disentangled embeddings to respectively preserve the subject identity and capture the identity-irrelevant information. We further design the novel weak denoising and contrastive embedding auxiliary tuning objectives to achieve the disentanglement. Extensive experiments show that our proposed DisenBooth framework outperforms baseline models for subject-driven text-to-image generation with the identity-preserved embedding. Additionally, by combining the identity-preserved embedding and identity-irrelevant embedding, DisenBooth demonstrates more generation flexibility and controllability
Forward citations
Cited by 3 Pith papers
-
Interact-Custom: Customized Human Object Interaction Image Generation
Interact-Custom generates customized human-object interaction images by first generating a foreground mask from the prompt and then using that mask to guide identity-preserving diffusion generation.
-
Comparison Reveals Commonality: Customized Image Generation through Contrastive Inversion
Contrastive Inversion disentangles the common concept from per-image auxiliary tokens via InfoNCE loss, then fine-tunes only the target cross-attention pathway, matching DisenBooth's numbers while claiming better qual...
-
Training Free Stylized Abstraction
A training-free framework coupling VLLM-based identity distillation with cross-domain rectified flow inversion generates identity-preserving stylized abstractions from a single reference image, evaluated by a new GPT-...
Discussion (0). Continue with ORCID to comment.