Pith. sign in

REVIEW 5 cited by

Visual Style Prompting with Swapping Self-Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12974 v2 pith:H3CFHTFY submitted 2024-02-20 cs.CV

classification cs.CV
keywords styleimagesvisualapproachchallengescontentelementsensuring
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the evolving domain of text-to-image generation, diffusion models have emerged as powerful tools in content creation. Despite their remarkable capability, existing models still face challenges in achieving controlled generation with a consistent style, requiring costly fine-tuning or often inadequately transferring the visual elements due to content leakage. To address these challenges, we propose a novel approach, \ours, to produce a diverse range of images while maintaining specific style elements and nuances. During the denoising process, we keep the query from original features while swapping the key and value with those from reference features in the late self-attention layers. This approach allows for the visual style prompting without any fine-tuning, ensuring that generated images maintain a faithful style. Through extensive evaluation across various styles and text prompts, our method demonstrates superiority over existing approaches, best reflecting the style of the references and ensuring that resulting images match the text prompts most accurately. Our project page is available https://curryjung.github.io/VisualStylePrompt/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    USO trains one DiT model for subject-driven, style-driven, and joint generation by disentangling content and style from triplet data and adding a style-reward objective, claiming SOTA on USO-Bench.

  2. Category-Aware 3D Object Composition with Disentangled Texture and Shape Multi-view Diffusion

    cs.CV 2025-09 conditional novelty 5.0 of 10

    C33D blends a 3D model with an object category by generating a fused front view, then using texture and shape multi-view diffusion plus adaptive inversion to reconstruct a novel, consistent 3D model.

  3. Towards Efficient Exemplar Based Image Editing with Multimodal VLMs

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ReEdit transfers exemplar-based edits to new images by conditioning Stable Diffusion on a LLaVA-written caption plus a CLIP edit-direction vector, with no per-example optimization.

  4. QR-LoRA: Efficient and Disentangled Fine-tuning via QR Decomposition for Customized Generation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    QR-LoRA freezes the QR-decomposed basis of pretrained weights, trains only a residual matrix, and reports halved trainable parameters with improved content-style disentanglement in diffusion models.

  5. StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation

    cs.CV 2025-05 reject novelty 4.0 of 10

    StyleAR enables autoregressive image generation models to do style-aligned text-to-image generation using only binary text-image data, via self-reconstruction training and style-enhanced tokens.

Pith tools