Pith. sign in

REVIEW 14 cited by

Measuring Style Similarity in Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.01292 v1 pith:N7U5TSVO submitted 2024-04-01 cs.CV cs.LG

classification cs.CVcs.LG
keywords styleimagemodelsusedimagestrainingartistscontent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative models are now widely used by graphic designers and artists. Prior works have shown that these models remember and often replicate content from their training data during generation. Hence as their proliferation increases, it has become important to perform a database search to determine whether the properties of the image are attributable to specific training data, every time before a generated image is used for professional purposes. Existing tools for this purpose focus on retrieving images of similar semantic content. Meanwhile, many artists are concerned with style replication in text-to-image models. We present a framework for understanding and extracting style descriptors from images. Our framework comprises a new dataset curated using the insight that style is a subjective property of an image that captures complex yet meaningful interactions of factors including but not limited to colors, textures, shapes, etc. We also propose a method to extract style descriptors that can be used to attribute style of a generated image to the images used in the training dataset of a text-to-image model. We showcase promising results in various style retrieval tasks. We also quantitatively and qualitatively analyze style attribution and matching in the Stable Diffusion model. Code and artifacts are available at https://github.com/learn2phoenix/CSD.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Style Similarity Scores Fail: Diagnosing Raw CSD Cosine in Artist-Style Evaluation

    cs.CV 2026-05 conditional novelty 7.0 of 10

    Raw CSD cosine similarity produces negative discrimination gaps for many artists and does not support absolute style-fidelity interpretation, but CSLS readout on frozen backbones reduces failures and improves AUC.

  2. Quantifying Cross-Modality Memorization in Vision-Language Models

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Fine-tuning VLMs on image-only or text-only personas yields a significant, asymmetric cross-modal memorization gap that persists with model scale, unlearning, and multi-hop reasoning.

  3. FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Video editing can be learned from image-edit pairs that are synthetically warped into videos, plus self-distillation losses that align image and video outputs.

  4. Evaluating Intellectual Property Guardrails of Generative Image Models: A Technical Report

    cs.CV 2026-07 conditional novelty 6.0 of 10

    All 14 tested text-to-image models readily generate recognizable IP; private models refuse at highly uneven rates, with commercial logos refused least and generated most.

  5. GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization

    cs.CV 2026-01 conditional novelty 6.0 of 10

    GimmBO uses preference-based Bayesian optimization with a sparse, sum-bounded search space to help users interactively discover adapter merges in 20-30 dimensional model-merging spaces.

  6. Neural Scene Designer: Self-Styled Semantic Image Manipulation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    NSD uses a contrastively learned style embedding from the input image itself, fed through a second cross-attention branch, to make diffusion-based inpainting results match the surrounding scene's style.

  7. USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    USO trains one DiT model for subject-driven, style-driven, and joint generation by disentangling content and style from triplet data and adding a style-reward objective, claiming SOTA on USO-Bench.

  8. AIComposer: Any Style and Content Image Composition via Feature Integration

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A nearly training-free SDXL pipeline composes foreground content with background style using a small MLP that merges CLIP image features, removing the need for text prompts.

  9. From Imitation to Innovation: The Emergence of AI Unique Artistic Styles and the Challenge of Copyright Protection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ArtBulb uses style-description-guided multimodal clustering combined with MLLMs to judge whether AI-generated artworks have a unique, consistent, prompt-accurate style eligible for copyright protection.

  10. OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    OneIG-Bench introduces a 2,440-prompt, six-dimension benchmark with automated metrics for text-to-image models, covering alignment, text, reasoning, style, and diversity in English and Chinese.

  11. UNIC: Unified In-Context Video Editing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    One diffusion transformer handles ID insert, swap, delete, stylization, propagation, and re-camera control in a single model using in-context token concatenation with task-aware positional encoding and bias.

  12. MMD Guidance: Training-Free Distribution Adaptation for Diffusion Models via Maximum Mean Discrepancy Guidance

    cs.LG 2026-01 conditional novelty 5.0 of 10

    MMD Guidance adapts pretrained diffusion models at inference time by adding gradients of Maximum Mean Discrepancy between generated and reference latents, aligning samples with a target distribution without retraining.

  13. A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A survey and framework that categorizes generative model unlearning by point-wise versus concept-wise objectives, parameter-based versus non-parametric methods, and completeness/utility/efficiency evaluation.

  14. MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting

    cs.CV 2025-06 conditional novelty 4.0 of 10

    MTADiffusion improves text-guided object inpainting by training on a new 5M-image mask-text dataset with edge prediction and style-consistency losses.

Pith tools