Pith. sign in

REVIEW 11 cited by

StyleDrop: Text-to-Image Generation in Any Style

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.00983 v1 pith:MUO44IYS submitted 2023-06-01 cs.CV cs.AI

classification cs.CVcs.AI
keywords stylestyledroptext-to-imagedesigneffectsimageimagesimpressive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Pre-trained large text-to-image models synthesize impressive images with an appropriate use of text prompts. However, ambiguities inherent in natural language and out-of-distribution effects make it hard to synthesize image styles, that leverage a specific design pattern, texture or material. In this paper, we introduce StyleDrop, a method that enables the synthesis of images that faithfully follow a specific style using a text-to-image model. The proposed method is extremely versatile and captures nuances and details of a user-provided style, such as color schemes, shading, design patterns, and local and global effects. It efficiently learns a new style by fine-tuning very few trainable parameters (less than $1\%$ of total model parameters) and improving the quality via iterative training with either human or automated feedback. Better yet, StyleDrop is able to deliver impressive results even when the user supplies only a single image that specifies the desired style. An extensive study shows that, for the task of style tuning text-to-image models, StyleDrop implemented on Muse convincingly outperforms other methods, including DreamBooth and textual inversion on Imagen or Stable Diffusion. More results are available at our project website: https://styledrop.github.io

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Imagine for Me: Creative Conceptual Blending of Real Images and Text via Blended Attention

    cs.CV 2025-06 conditional novelty 7.0 of 10

    IT-Blender blends a real image and a text prompt by injecting clean reference features into a trained attention module, improving disentangled concept blending in SD and FLUX.

  2. LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow Transformers

    cs.CV 2025-05 conditional novelty 7.0 of 10

    LoRAShop localizes each LoRA's effect to attention-derived spatial masks inside a Flux transformer, enabling training-free multi-concept image generation and editing.

  3. ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Test-time tuning of video diffusion models collapses generation toward the source video; ElasticTTT counters this with noisy targets, contrastive source-prompt guidance, and asynchronous region-wise noise scheduling, ...

  4. Calligrapher: Freestyle Text Image Customization

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Calligrapher trains a style encoder and in-context inference on self-distilled FLUX outputs to redraw arbitrary text in the visual style of a reference image.

  5. FairyGen: Storied Cartoon Video from a Single Child-Drawn Character

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A pipeline that generates story-driven cartoon videos from one child-drawn character by separating foreground style, background synthesis, and motion learning.

  6. BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A dual-stream diffusion model trained with Blender-render conditioning, source masking, and object jittering performs 3D-grounded multi-object editing and compositing better than existing baselines on three video datasets.

  7. Noise Consistency Regularization for Improved Subject-Driven Image Synthesis

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Adding consistency-to-pretrained and multiplicative-noise consistency losses to fine-tuning improves subject identity and background diversity over DreamBooth on a 30-subject benchmark.

  8. CDST: Color Disentangled Style Transfer for Universal Style Reference Customization

    cs.CV 2025-05 conditional novelty 6.0 of 10

    CDST disentangles color from style via greyscale style input and a color histogram stream, enabling zero-shot style transfer with separate color control and a new characteristics-preserved mode.

  9. OmniStyle: Filtering High Quality Style Transfer Data at Scale

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new million-triplet dataset and a diffusion transformer model that performs text-guided and image-guided style transfer, with a filtering pipeline used to curate high-quality training examples.

  10. DICT: Data Injection and Contrastive Trajectory Refinement for Conditional Image Generation with Diffusion Models

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Noise-perturbed condition injection plus contrastive trajectory refinement improves training-free conditional diffusion sampling across style transfer, super-resolution and deblurring.

  11. StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation

    cs.CV 2025-05 reject novelty 4.0 of 10

    StyleAR enables autoregressive image generation models to do style-aligned text-to-image generation using only binary text-image data, via self-reconstruction training and style-enhanced tokens.

Pith tools