Pith. sign in

Scaling Concept With Text-Guided Diffusion Models

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Text-guided diffusion models have revolutionized generative tasks by producing high-fidelity content from text descriptions. They have also enabled an editing paradigm where concepts can be replaced through text conditioning (e.g., a dog to a tiger). In this work, we explore a novel approach: instead of replacing a concept, can we enhance or suppress the concept itself? Through an empirical study, we identify a trend where concepts can be decomposed in text-guided diffusion models. Leveraging this insight, we introduce ScalingConcept, a simple yet effective method to scale decomposed concepts up or down in real input without introducing new elements. To systematically evaluate our approach, we present the WeakConcept-10 dataset, where concepts are imperfect and need to be enhanced. More importantly, ScalingConcept enables a variety of novel zero-shot applications across image and audio domains, including tasks such as canonical pose generation and generative sound highlighting or removal.

citation-role summary

background 1

citation-polarity summary

fields

cs.SD 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

ZeroSep: Separate Anything in Audio with Zero Training

cs.SD · 2025-05-29 · conditional · novelty 6.0

Latent inversion of a mixed audio into a pretrained text-guided diffusion model, followed by denoising with classifier-free guidance weight 1, performs zero-training source separation.

citing papers explorer

Showing 1 of 1 citing paper.

  • ZeroSep: Separate Anything in Audio with Zero Training cs.SD · 2025-05-29 · conditional · none · ref 2025 · internal anchor

    Latent inversion of a mixed audio into a pretrained text-guided diffusion model, followed by denoising with classifier-free guidance weight 1, performs zero-training source separation.