REVIEW 6 cited by
Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present a method to create interpretable concept sliders that enable precise control over attributes in image generations from diffusion models. Our approach identifies a low-rank parameter direction corresponding to one concept while minimizing interference with other attributes. A slider is created using a small set of prompts or sample images; thus slider directions can be created for either textual or visual concepts. Concept Sliders are plug-and-play: they can be composed efficiently and continuously modulated, enabling precise control over image generation. In quantitative experiments comparing to previous editing techniques, our sliders exhibit stronger targeted edits with lower interference. We showcase sliders for weather, age, styles, and expressions, as well as slider compositions. We show how sliders can transfer latents from StyleGAN for intuitive editing of visual concepts for which textual description is difficult. We also find that our method can help address persistent quality issues in Stable Diffusion XL including repair of object deformations and fixing distorted hands. Our code, data, and trained sliders are available at https://sliders.baulab.info/
Forward citations
Cited by 6 Pith papers
-
CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
CompSlider learns to synthesize image-conditioning latents from multiple attribute sliders at once, aiming for more disentangled and structure-preserving multi-attribute control in text-to-image generation.
-
Learning Distribution-Wise Control in Representation Space for Language Models
Stochastic distribution-wise interventions that learn a mean and variance in representation space improve ReFT-based language model reasoning, with the largest gains from restricting randomness to the first quarter of layers.
-
MARBLE: Material Recomposition and Blending in CLIP-Space
MARBLE performs material blending and parametric material-attribute control by manipulating CLIP image embeddings and injecting them into a specific U-Net block of a pre-trained diffusion model.
-
Dual-Process Image Generation
Backpropagating a VLM's yes/no answer probabilities into an image generator's LoRA weights distills new visual controls and commonsense rules into a fast feed-forward pass.
-
Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment
A reparameterization recipe that lets pre-trained Stable Diffusion checkpoints be finetuned as flow matching models, giving faster convergence and better performance under parameter-efficient constraints.
-
Beyond Sliders: Mastering the Art of Diffusion-based Image Manipulation
Beyond Sliders augments Concept Sliders with perceptual, adversarial, and an undefined triplet loss, claiming better real-world edits, but the evidence is weak and the derivation is not valid.
Discussion (0). Sign in to comment.