Pith. sign in

REVIEW 2 cited by

Maximum Likelihood Training of Score-Based Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.09258 v4 pith:V7V2G2JR submitted 2021-01-22 stat.ML cs.LG

classification stat.MLcs.LG
keywords modelsdiffusionscore-basedlikelihoodlog-likelihoodmaximumtrainingcombination
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Score-based diffusion models synthesize samples by reversing a stochastic process that diffuses data to noise, and are trained by minimizing a weighted combination of score matching losses. The log-likelihood of score-based diffusion models can be tractably computed through a connection to continuous normalizing flows, but log-likelihood is not directly optimized by the weighted combination of score matching losses. We show that for a specific weighting scheme, the objective upper bounds the negative log-likelihood, thus enabling approximate maximum likelihood training of score-based diffusion models. We empirically observe that maximum likelihood training consistently improves the likelihood of score-based diffusion models across multiple datasets, stochastic processes, and model architectures. Our best models achieve negative log-likelihoods of 2.83 and 3.76 bits/dim on CIFAR-10 and ImageNet 32x32 without any data augmentation, on a par with state-of-the-art autoregressive models on these tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 46 citations worldwide. Full citation record

  1. Combining complex Langevin dynamics with score-based and energy-based diffusion models

    hep-lat 2025-10 conditional novelty 5.0 of 10

    Energy-based diffusion models trained on complex Langevin data produce an explicit energy function for the sampled distribution, enabling MCMC without re-simulation.

  2. Fokker-Planck to Callan-Symanzik: evolution of weight matrices under training

    cs.LG 2025-01 conditional novelty 4.0 of 10

    Weight-matrix probability densities in a toy autoencoder are evolved with the Fokker-Planck equation driven by the ADAM update, and the resulting output distributions roughly match training at epoch 5.

Pith tools