Pith. sign in

REVIEW 11 cited by

Generalization in diffusion models arises from geometry-adaptive harmonic representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.02557 v3 pith:OSRK4UFQ submitted 2023-10-04 cs.CV cs.LG

Generalization in diffusion models arises from geometry-adaptive harmonic representations

classification cs.CV cs.LG
keywords trainedharmonicimagewhenbasisdenoisingdensitydnns
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Deep neural networks (DNNs) trained for image denoising are able to generate high-quality samples with score-based reverse diffusion algorithms. These impressive capabilities seem to imply an escape from the curse of dimensionality, but recent reports of memorization of the training set raise the question of whether these networks are learning the "true" continuous density of the data. Here, we show that two DNNs trained on non-overlapping subsets of a dataset learn nearly the same score function, and thus the same density, when the number of training images is large enough. In this regime of strong generalization, diffusion-generated images are distinct from the training set, and are of high visual quality, suggesting that the inductive biases of the DNNs are well-aligned with the data density. We analyze the learned denoising functions and show that the inductive biases give rise to a shrinkage operation in a basis adapted to the underlying image. Examination of these bases reveals oscillating harmonic structures along contours and in homogeneous regions. We demonstrate that trained denoisers are inductively biased towards these geometry-adaptive harmonic bases since they arise not only when the network is trained on photographic images, but also when it is trained on image classes supported on low-dimensional manifolds for which the harmonic basis is suboptimal. Finally, we show that when trained on regular image classes for which the optimal basis is known to be geometry-adaptive and harmonic, the denoising performance of the networks is near-optimal.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. An exact information theory of generalization phase transitions in Bayesian diffusion models

    cs.LG 2026-07 conditional novelty 8.0

    Bayesian diffusion models memorize training data when mutual information between restricted observations and training data exceeds log dataset size, and generalize otherwise.

  2. Generating quantum ensembles via reverse-time quantum diffusions

    quant-ph 2026-06 unverdicted novelty 8.0

    The paper establishes a reverse-time quantum diffusion framework that generates complex quantum ensembles from simple distributions by deriving and learning a feedback Hamiltonian from forward trajectory data.

  3. Diffusion Model Attribution via Spectral Coupling of Denoiser Responses

    cs.CV 2026-06 unverdicted novelty 7.0

    SDS extracts stable spectral signatures from diffusion model denoisers via frequency-controlled perturbations, achieving 99.9% attribution accuracy across eight models and 96.2% under prompt shift.

  4. Grokking of Diffusion Models: Case Study on Modular Addition

    cs.LG 2026-04 unverdicted novelty 7.0

    Diffusion models show grokking on modular addition by composing periodic operand representations in simple data regimes or by separating arithmetic computation from visual denoising across timesteps in varied regimes.

  5. Generalization and Memorization in Rectified Flow

    cs.LG 2026-03 accept novelty 7.0

    Rectified Flow models peak in membership-inference vulnerability at the flow midpoint under uniform training; U-shaped timestep sampling suppresses memorization without harming FID.

  6. From Score Matching to Diffusion: A Fine-Grained Error Analysis in the Gaussian Setting

    cs.LG 2025-03 unverdicted novelty 7.0

    In the Gaussian setting the Wasserstein error of score-matching-plus-diffusion sampling equals a kernel norm of the data power spectrum whose kernel is determined by the four error sources and the algorithm parameters.

  7. Mechanisms of Misgeneralization in Physical Sequence Modeling

    cs.LG 2026-05 unverdicted novelty 6.0

    Generative sequence models for physical tasks exhibit physical misgeneralization where local prediction errors propagate through physical measurements to distort aggregate distributions over quantities like distance o...

  8. Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data

    cs.LG 2026-04 unverdicted novelty 6.0

    Uniform-based discrete diffusion models behave as associative memories that retrieve unseen data, with a dataset-size-driven memorization-to-generalization transition detectable via conditional entropy of token predictions.

  9. Diffusion Models Memorize in Training -- and Generalize in Inference

    cs.LG 2026-03 unverdicted novelty 6.0

    Diffusion models overfit denoising loss at intermediate noise but generalize in inference as model error smooths the flow field and sampling paths avoid memorized noisy training data.

  10. A Theoretical Analysis of Memory and Overfitting Phenomena in Stochastic Interpolation Models

    cs.LG 2026-06 unverdicted novelty 5.0

    In the oracle continuous-time setting, stochastic interpolation models recover training samples exactly, with deviations controlled by discretization and estimation errors, leading to theoretical definitions of overfi...

  11. BADiff: Bandwidth Adaptive Diffusion Model

    cs.CV 2025-10 unverdicted novelty 5.0

    BADiff introduces joint training of diffusion models with quality conditioning derived from bandwidth to enable adaptive early-stop sampling that preserves appropriate perceptual quality.