Pith. sign in

REVIEW 5 cited by

ViCo: Plug-and-play Visual Condition for Personalized Text-to-image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.00971 v2 pith:DGLEX3AK submitted 2023-06-01 cs.CV cs.AI

classification cs.CVcs.AI
keywords vicodiffusiongenerationmodelpersonalizedtext-to-imagevisualfine-tuning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Personalized text-to-image generation using diffusion models has recently emerged and garnered significant interest. This task learns a novel concept (e.g., a unique toy), illustrated in a handful of images, into a generative model that captures fine visual details and generates photorealistic images based on textual embeddings. In this paper, we present ViCo, a novel lightweight plug-and-play method that seamlessly integrates visual condition into personalized text-to-image generation. ViCo stands out for its unique feature of not requiring any fine-tuning of the original diffusion model parameters, thereby facilitating more flexible and scalable model deployment. This key advantage distinguishes ViCo from most existing models that necessitate partial or full diffusion fine-tuning. ViCo incorporates an image attention module that conditions the diffusion process on patch-wise visual semantics, and an attention-based object mask that comes at no extra cost from the attention module. Despite only requiring light parameter training (~6% compared to the diffusion U-Net), ViCo delivers performance that is on par with, or even surpasses, all state-of-the-art models, both qualitatively and quantitatively. This underscores the efficacy of ViCo, making it a highly promising solution for personalized text-to-image generation without the need for diffusion model fine-tuning. Code: https://github.com/haoosz/ViCo

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation

    cs.CV 2026-03 unverdicted novelty 7.0 of 10

    DSH-Bench supplies a hierarchical 58-category subject set, difficulty/scenario labels, and a human-aligned SICS metric that exposes systematic failures of 19 subject-driven T2I models.

  2. APT: Adaptive Personalized Training for Diffusion Models with Limited Data

    cs.CV 2025-07 conditional novelty 6.0 of 10

    APT detects overfitting during diffusion fine-tuning and uses adaptive augmentation, loss weighting, feature-statistics regularization, and attention alignment to preserve prior knowledge while learning new concepts.

  3. StyleBlend: Enhancing Style-Specific Content Creation in Text-to-Image Diffusion Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    StyleBlend learns few-shot artistic style as separate layout and texture components and blends them during diffusion sampling to improve text-aligned, style-specific image generation.

  4. Steering Guidance for Personalized Text-to-Image Diffusion Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Weight-interpolated null-text weak model in classifier-free guidance improves subject fidelity with minimal text-fidelity loss in personalized text-to-image diffusion.

  5. Hybrid Scandium Aluminum Nitride/Silicon Nitride Integrated Photonic Circuits

    physics.optics 2025-08 reject novelty 5.0 of 10

    The abstract reports a low-loss ScAlN/Si3N4 hybrid waveguide, but the full text is an unrelated diffusion-model paper, leaving the photonics claim without supporting evidence.

Pith tools