Pith. sign in

REVIEW 2 cited by

InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.03411 v1 pith:JFKSGOQ2 submitted 2023-04-06 cs.CV

classification cs.CV
keywords conceptimagetest-timefinetuningimageslearnmodelpre-trained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in personalized image generation allow a pre-trained text-to-image model to learn a new concept from a set of images. However, existing personalization approaches usually require heavy test-time finetuning for each concept, which is time-consuming and difficult to scale. We propose InstantBooth, a novel approach built upon pre-trained text-to-image models that enables instant text-guided image personalization without any test-time finetuning. We achieve this with several major components. First, we learn the general concept of the input images by converting them to a textual token with a learnable image encoder. Second, to keep the fine details of the identity, we learn rich visual feature representation by introducing a few adapter layers to the pre-trained model. We train our components only on text-image pairs without using paired images of the same concept. Compared to test-time finetuning-based methods like DreamBooth and Textual-Inversion, our model can generate competitive results on unseen concepts concerning language-image alignment, image fidelity, and identity preservation while being 100 times faster.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. R^2MoE: Redundancy-Removal Mixture of Experts for Lifelong Concept Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    R2MoE adds per-concept LoRA experts with routing distillation and expert pruning, reporting 0.19% forgetting and 15.2M added parameters on CustomConcept101.

  2. Multitwine: Multi-Object Compositing with Text and Layout Control

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A single diffusion model simultaneously composites multiple objects into a scene with text and layout control, outperforming sequential insertion on interacting cases.

Pith tools