REVIEW 12 cited by
ViCo: Plug-and-play Visual Condition for Personalized Text-to-image Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Personalized text-to-image generation using diffusion models has recently emerged and garnered significant interest. This task learns a novel concept (e.g., a unique toy), illustrated in a handful of images, into a generative model that captures fine visual details and generates photorealistic images based on textual embeddings. In this paper, we present ViCo, a novel lightweight plug-and-play method that seamlessly integrates visual condition into personalized text-to-image generation. ViCo stands out for its unique feature of not requiring any fine-tuning of the original diffusion model parameters, thereby facilitating more flexible and scalable model deployment. This key advantage distinguishes ViCo from most existing models that necessitate partial or full diffusion fine-tuning. ViCo incorporates an image attention module that conditions the diffusion process on patch-wise visual semantics, and an attention-based object mask that comes at no extra cost from the attention module. Despite only requiring light parameter training (~6% compared to the diffusion U-Net), ViCo delivers performance that is on par with, or even surpasses, all state-of-the-art models, both qualitatively and quantitatively. This underscores the efficacy of ViCo, making it a highly promising solution for personalized text-to-image generation without the need for diffusion model fine-tuning. Code: https://github.com/haoosz/ViCo
Forward citations
Cited by 12 Pith papers
-
DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
DSH-Bench supplies a hierarchical 58-category subject set, difficulty/scenario labels, and a human-aligned SICS metric that exposes systematic failures of 19 subject-driven T2I models.
-
DreamCache: Finetuning-Free Lightweight Personalized Image Generation via Feature Caching
DreamCache achieves zero-shot personalized image generation by caching reference features from one denoising step and injecting them through 25M-parameter adapters.
-
APT: Adaptive Personalized Training for Diffusion Models with Limited Data
APT detects overfitting during diffusion fine-tuning and uses adaptive augmentation, loss weighting, feature-statistics regularization, and attention alignment to preserve prior knowledge while learning new concepts.
-
BridgeIV: Bridging Customized Image and Video Generation through Test-Time Autoregressive Identity Propagation
BridgeIV improves subject consistency in customized text-to-video generation by warping attention maps and self-attention values across frames, then refining latents with a CLIP-based reward.
-
StyleBlend: Enhancing Style-Specific Content Creation in Text-to-Image Diffusion Models
StyleBlend learns few-shot artistic style as separate layout and texture components and blends them during diffusion sampling to improve text-aligned, style-specific image generation.
-
LoRA.rar: Learning to Merge LoRAs via Hypernetworks for Subject-Style Conditioned Image Generation
A hypernetwork pretrained on pairs of subject and style LoRAs predicts column-wise merging coefficients, enabling real-time, high-quality joint subject-style image personalization.
-
DreamBlend: Advancing Personalized Fine-tuning of Text-to-Image Diffusion Models
DreamBlend guides an overfit fine-tuned checkpoint with cross-attention maps from an underfit checkpoint, improving subject fidelity, prompt fidelity, and diversity in personalized text-to-image generation.
-
PersonaCraft: Personalized and Controllable Full-Body Multi-Human Scene Generation Using Occlusion-Aware 3D-Conditioned Diffusion
PersonaCraft adds SMPLx depth and normal conditioning, occlusion boundary enhancement, and occlusion-aware classifier-free guidance to diffusion models, enabling controllable multi-person images that preserve both fac...
-
Steering Guidance for Personalized Text-to-Image Diffusion Models
Weight-interpolated null-text weak model in classifier-free guidance improves subject fidelity with minimal text-fidelity loss in personalized text-to-image diffusion.
-
Hybrid Scandium Aluminum Nitride/Silicon Nitride Integrated Photonic Circuits
The abstract reports a low-loss ScAlN/Si3N4 hybrid waveguide, but the full text is an unrelated diffusion-model paper, leaving the photonics claim without supporting evidence.
-
PIDiff: Image Customization for Personalized Identities with Diffusion Models
The paper reports a per-identity fine-tuning method using StyleGAN's layered face codes as visual prompts, and it claims improved identity preservation and text alignment over previous baselines.
-
Text-to-Image Synthesis: A Decade Survey
A decade-spanning survey categorizes over 440 text-to-image papers by architecture, research problem, dataset, and evaluation metric.
Discussion (0). Continue with ORCID to comment.