Pith. sign in

REVIEW 2 cited by

Scaling Concept With Text-Guided Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.24151 v1 pith:7KWTRPCR submitted 2024-10-31 cs.CV cs.CL

classification cs.CVcs.CL
keywords conceptsconceptdiffusionmodelstext-guidedapproachdecomposedgenerative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-guided diffusion models have revolutionized generative tasks by producing high-fidelity content from text descriptions. They have also enabled an editing paradigm where concepts can be replaced through text conditioning (e.g., a dog to a tiger). In this work, we explore a novel approach: instead of replacing a concept, can we enhance or suppress the concept itself? Through an empirical study, we identify a trend where concepts can be decomposed in text-guided diffusion models. Leveraging this insight, we introduce ScalingConcept, a simple yet effective method to scale decomposed concepts up or down in real input without introducing new elements. To systematically evaluate our approach, we present the WeakConcept-10 dataset, where concepts are imperfect and need to be enhanced. More importantly, ScalingConcept enables a variety of novel zero-shot applications across image and audio domains, including tasks such as canonical pose generation and generative sound highlighting or removal.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ZeroSep: Separate Anything in Audio with Zero Training

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Latent inversion of a mixed audio into a pretrained text-guided diffusion model, followed by denoising with classifier-free guidance weight 1, performs zero-training source separation.

  2. KinMo: Kinematic-aware Human Motion Understanding and Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    KinMo adds hierarchical body-part-level text annotations to HumanML3D and shows that progressive text-motion alignment improves retrieval, generation, editing, and trajectory control.

Pith tools