Pith. sign in

REVIEW 8 cited by

Interpreting the Weight Space of Customized Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.09413 v3 pith:A2UC6QHZ submitted 2024-06-13 cs.CV cs.GRcs.LG

classification cs.CVcs.GRcs.LG
keywords spacemodelmodelsdiffusionidentityweightweightscustomized
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We investigate the space of weights spanned by a large collection of customized diffusion models. We populate this space by creating a dataset of over 60,000 models, each of which is a base model fine-tuned to insert a different person's visual identity. We model the underlying manifold of these weights as a subspace, which we term weights2weights. We demonstrate three immediate applications of this space that result in new diffusion models -- sampling, editing, and inversion. First, sampling a set of weights from this space results in a new model encoding a novel identity. Next, we find linear directions in this space corresponding to semantic edits of the identity (e.g., adding a beard), resulting in a new model with the original identity edited. Finally, we show that inverting a single image into this space encodes a realistic identity into a model, even if the input image is out of distribution (e.g., a painting). We further find that these linear properties of the diffusion model weight space extend to other visual concepts. Our results indicate that the weight space of fine-tuned diffusion models can behave as an interpretable meta-latent space producing new models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can this Model Also Recognize Dogs? Zero-Shot Model Search from Weights

    cs.LG 2025-02 conditional novelty 7.0 of 10

    ProbeLog represents each classifier output by its responses to fixed probe images and uses CLIP to answer text queries, achieving 43.8% top-1 accuracy when searching 1,500 ImageNet-trained models for a concept.

  2. HyperNet Fields: Efficiently Training Hypernetworks without Ground Truth by Learning Weight Trajectories

    cs.LG 2024-12 conditional novelty 7.0 of 10

    A hypernetwork is trained to predict the whole weight trajectory of a task network by matching the task loss gradient at each step, removing the need for per-sample ground truth weights.

  3. SliderSpace: Decomposing the Visual Capabilities of Diffusion Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    SliderSpace uses PCA on CLIP embeddings of a diffusion model's own samples, then trains low-rank adapters for each principal component, turning them into composable image control sliders.

  4. A LoRA is Worth a Thousand Pictures

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LoRA weight vectors, projected with PCA and a per-PC calibration, cluster and retrieve artistic styles more accurately than CLIP, DINO, and style-specialized image features.

  5. FluxSpace: Disentangled Semantic Editing in Rectified Flow Transformers

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FluxSpace performs training-free, disentangled semantic editing in rectified flow transformers by combining attention outputs with prompt-derived linear directions.

  6. Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models

    cs.AI 2024-11 conditional novelty 6.0 of 10

    Fine-tuning text-to-image diffusion models on benign data can reactivate suppressed unsafe concepts, and training the task adapter separately from a frozen safety LoRA prevents this.

  7. MyTimeMachine: Personalized Facial Age Transformation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A personalized facial age transformation method that uses an adapter network on top of the SAM global aging model, trained with 10 to 50 photos of one person, to produce re-aged images that resemble that person's actu...

  8. LoRA Diffusion: Zero-Shot LoRA Synthesis for Diffusion Model Personalization

    cs.LG 2024-12 reject novelty 3.0 of 10

    A VAE plus diffusion hypernetwork synthesizes Stable Diffusion LoRAs for faces from ArcFace embeddings, aiming for zero-shot personalization without per-user fine-tuning.

Pith tools