REVIEW 10 cited by
An analytic theory of creativity in convolutional diffusion models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We obtain an analytic, interpretable and predictive theory of creativity in convolutional diffusion models. Indeed, score-matching diffusion models can generate highly original images that lie far from their training data. However, optimal score-matching theory suggests that these models should only be able to produce memorized training examples. To reconcile this theory-experiment gap, we identify two simple inductive biases, locality and equivariance, that: (1) induce a form of combinatorial creativity by preventing optimal score-matching; (2) result in fully analytic, completely mechanistically interpretable, local score (LS) and equivariant local score (ELS) machines that, (3) after calibrating a single time-dependent hyperparameter can quantitatively predict the outputs of trained convolution only diffusion models (like ResNets and UNets) with high accuracy (median $r^2$ of $0.95, 0.94, 0.94, 0.96$ for our top model on CIFAR10, FashionMNIST, MNIST, and CelebA). Our model reveals a locally consistent patch mosaic mechanism of creativity, in which diffusion models create exponentially many novel images by mixing and matching different local training set patches at different scales and image locations. Our theory also partially predicts the outputs of pre-trained self-attention enabled UNets (median $r^2 \sim 0.77$ on CIFAR10), revealing an intriguing role for attention in carving out semantic coherence from local patch mosaics.
Forward citations
Cited by 10 Pith papers
-
Toward a mechanistic understanding of inference in visual cortex and diffusion models
A sparse-coding circuit with learned pairwise interactions, trained by score matching, reproduces diffusion-model-like contour completion and claims to expose the mechanism behind it.
-
A Random Matrix Theory Perspective on the Consistency of Diffusion Models
Finite-sample randomness in linear diffusion models is equivalent to a renormalized noise scale, and cross-split disagreement follows a factorized law scaling as 1/n.
-
Emergence of Nonequilibrium Latent Cycles in Unsupervised Generative Modeling
In an asymmetric visible-hidden Markov-chain generative model, likelihood-type training spontaneously produces nonequilibrium latent cycles, and models with stronger cycles reproduce digit-class frequencies more faithfully.
-
Spooky Action at a Distance: Normalization Layers Enable Side-Channel Spatial Communication
InstanceNorm, GroupNorm, and BatchNorm can act as spatial communication channels that let CNNs aggregate information from well beyond their local receptive field.
-
GuessBench: Sensemaking Multimodal Creativity in the Wild
A Minecraft-based benchmark shows vision-language models often fail to decode player-built creations, with accuracy falling sharply for rare concepts and low-resource languages.
-
Momentum Guidance: Plug-and-Play Guidance for Flow Models
Momentum Guidance improves flow-model sample quality by extrapolating the current velocity away from an exponential moving average of past velocities, with no extra model evaluations.
-
Ambient Diffusion Omni: Training Good Models with Bad Data
Ambient Diffusion Omni trains diffusion models on mixed-quality data by learning when corrupted images can be treated as clean, improving generation quality and diversity.
-
Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models
In overparameterized diffusion models, generalization happens first and memorization starts later, with the memorization time growing linearly with dataset size.
-
Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models
Direct Ascent Synthesis generates recognizable images from CLIP embeddings by optimizing a sum of multi-resolution image components, requiring no generative training.
-
Density Ratio Estimation with Conditional Probability Paths
Conditional Time Score Matching estimates density ratios by regressing closed-form conditional time scores, giving faster learning and theoretical error bounds.
Discussion (0). Continue with ORCID to comment.