Pith. sign in

REVIEW 2 cited by

Diffusion Models are Minimax Optimal Distribution Estimators

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.01861 v1 pith:R6RNXLY6 submitted 2023-03-03 stat.ML cs.LG

classification stat.MLcs.LG
keywords diffusiondistributionmodelingdatadistancefunctionminimaxmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While efficient distribution learning is no doubt behind the groundbreaking success of diffusion modeling, its theoretical guarantees are quite limited. In this paper, we provide the first rigorous analysis on approximation and generalization abilities of diffusion modeling for well-known function spaces. The highlight of this paper is that when the true density function belongs to the Besov space and the empirical score matching loss is properly minimized, the generated data distribution achieves the nearly minimax optimal estimation rates in the total variation distance and in the Wasserstein distance of order one. Furthermore, we extend our theory to demonstrate how diffusion models adapt to low-dimensional data distributions. We expect these results advance theoretical understandings of diffusion modeling and its ability to generate verisimilar outputs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 8 citations worldwide. Full citation record

  1. Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    In overparameterized diffusion models, generalization happens first and memorization starts later, with the memorization time growing linearly with dataset size.

  2. Adaptivity and Convergence of Probability Flow ODEs in Diffusion Generative Models

    stat.ML 2025-01 conditional novelty 5.0 of 10

    With accurate score estimates, the probability flow ODE sampler reaches O(k/T) total-variation error, where k is the intrinsic dimension of the target distribution.

Pith tools