Pith. sign in

REVIEW 10 cited by

An analytic theory of creativity in convolutional diffusion models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.20292 v2 pith:PLHS2GO2 submitted 2024-12-28 cs.LG cond-mat.dis-nncs.AIq-bio.NCstat.ML

classification cs.LGcond-mat.dis-nncs.AIq-bio.NCstat.ML
keywords modelsdiffusioncreativitylocaltheoryanalyticscore-matchingtraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We obtain an analytic, interpretable and predictive theory of creativity in convolutional diffusion models. Indeed, score-matching diffusion models can generate highly original images that lie far from their training data. However, optimal score-matching theory suggests that these models should only be able to produce memorized training examples. To reconcile this theory-experiment gap, we identify two simple inductive biases, locality and equivariance, that: (1) induce a form of combinatorial creativity by preventing optimal score-matching; (2) result in fully analytic, completely mechanistically interpretable, local score (LS) and equivariant local score (ELS) machines that, (3) after calibrating a single time-dependent hyperparameter can quantitatively predict the outputs of trained convolution only diffusion models (like ResNets and UNets) with high accuracy (median $r^2$ of $0.95, 0.94, 0.94, 0.96$ for our top model on CIFAR10, FashionMNIST, MNIST, and CelebA). Our model reveals a locally consistent patch mosaic mechanism of creativity, in which diffusion models create exponentially many novel images by mixing and matching different local training set patches at different scales and image locations. Our theory also partially predicts the outputs of pre-trained self-attention enabled UNets (median $r^2 \sim 0.77$ on CIFAR10), revealing an intriguing role for attention in carving out semantic coherence from local patch mosaics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Toward a mechanistic understanding of inference in visual cortex and diffusion models

    q-bio.NC 2026-07 reject novelty 7.0 of 10

    A sparse-coding circuit with learned pairwise interactions, trained by score matching, reproduces diffusion-model-like contour completion and claims to expose the mechanism behind it.

  2. A Random Matrix Theory Perspective on the Consistency of Diffusion Models

    cs.LG 2026-02 conditional novelty 7.0 of 10

    Finite-sample randomness in linear diffusion models is equivalent to a renormalized noise scale, and cross-split disagreement follows a factorized law scaling as 1/n.

  3. Emergence of Nonequilibrium Latent Cycles in Unsupervised Generative Modeling

    cond-mat.stat-mech 2025-12 unverdicted novelty 7.0 of 10

    In an asymmetric visible-hidden Markov-chain generative model, likelihood-type training spontaneously produces nonequilibrium latent cycles, and models with stronger cycles reproduce digit-class frequencies more faithfully.

  4. Spooky Action at a Distance: Normalization Layers Enable Side-Channel Spatial Communication

    cs.LG 2025-07 conditional novelty 7.0 of 10

    InstanceNorm, GroupNorm, and BatchNorm can act as spatial communication channels that let CNNs aggregate information from well beyond their local receptive field.

  5. GuessBench: Sensemaking Multimodal Creativity in the Wild

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A Minecraft-based benchmark shows vision-language models often fail to decode player-built creations, with accuracy falling sharply for rare concepts and low-resource languages.

  6. Momentum Guidance: Plug-and-Play Guidance for Flow Models

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Momentum Guidance improves flow-model sample quality by extrapolating the current velocity away from an exponential moving average of past velocities, with no extra model evaluations.

  7. Ambient Diffusion Omni: Training Good Models with Bad Data

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Ambient Diffusion Omni trains diffusion models on mixed-quality data by learning when corrupted images can be treated as clean, improving generation quality and diversity.

  8. Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    In overparameterized diffusion models, generalization happens first and memorization starts later, with the memorization time growing linearly with dataset size.

  9. Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Direct Ascent Synthesis generates recognizable images from CLIP embeddings by optimizing a sum of multi-resolution image components, requiring no generative training.

  10. Density Ratio Estimation with Conditional Probability Paths

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Conditional Time Score Matching estimates density ratios by regressing closed-form conditional time scores, giving faster learning and theoretical error bounds.

Pith tools