Pith. sign in

REVIEW 14 cited by

How to Train Your Energy-Based Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.03288 v2 pith:4IHUFDEB submitted 2021-01-09 cs.LG stat.ML

classification cs.LGstat.ML
keywords modelsebmstrainingapproachesconstantnormalizingenergy-basedprobabilistic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Energy-Based Models (EBMs), also known as non-normalized probabilistic models, specify probability density or mass functions up to an unknown normalizing constant. Unlike most other probabilistic models, EBMs do not place a restriction on the tractability of the normalizing constant, thus are more flexible to parameterize and can model a more expressive family of probability distributions. However, the unknown normalizing constant of EBMs makes training particularly difficult. Our goal is to provide a friendly introduction to modern approaches for EBM training. We start by explaining maximum likelihood training with Markov chain Monte Carlo (MCMC), and proceed to elaborate on MCMC-free approaches, including Score Matching (SM) and Noise Constrastive Estimation (NCE). We highlight theoretical connections among these three approaches, and end with a brief survey on alternative training methods, which are still under active research. Our tutorial is targeted at an audience with basic understanding of generative models who want to apply EBMs or start a research project in this direction.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 81 citations worldwide. Full citation record

  1. Bayesian Experimental Design via Score Matching

    stat.ML 2026-07 conditional novelty 7.0 of 10

    SCOREBED isolates EIG double intractability in a policy-independent score-matching stage, then trains design policies with a singly intractable gradient estimator, enabling cheap multi-policy selection.

  2. Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Motion Generation

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    STREAM decouples text (via AdaLN) from music (via energy-based BEAM attention) to generate editable, musically aligned dance motions with a new annotated dataset and editability metric.

  3. Energy-based Tissue Manifolds for Longitudinal Multiparametric MRI Analysis

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Patient-specific energy manifolds from baseline mpMRI scans act as fixed geometric references to monitor longitudinal evolution of voxel distributions in sequence space for neuro-oncology proof-of-concept cases.

  4. Transferable Direct Prompt Injection via Activation-Guided MCMC Sampling

    cs.AI 2025-09 conditional novelty 7.0 of 10

    An activation-guided energy model plus MCMC sampling creates transferable direct prompt injection attacks in a black-box setting, reaching 49.6% average attack success across five LLMs.

  5. Projected Energy Matching for Generative 3D Priors

    eess.IV 2026-07 conditional novelty 6.0 of 10

    Helmholtz Distillation plus Negative Caching amortizes Energy Matching to 3D CT volumes, producing a conservative prior that improves FID and sparse-view reconstruction over pure flow baselines.

  6. Diffusion-based Annealed Boltzmann Generators : benefits, pitfalls and hopes

    stat.ML 2026-01 conditional novelty 6.0 of 10

    Even a perfect diffusion model yields poor annealed Boltzmann generators when coupled through first-order stochastic denoising kernels, while deterministic transport maps and second-order kernels improve; with learned...

  7. Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction

    cs.LG 2025-12 unverdicted novelty 6.0 of 10

    Under the chain rule of probability, autoregressive models and energy-based models are in exact bijection in function space, making the global optimum of teacher forcing equivalent to an energy-based model with implic...

  8. Learning Latent Energy-Based Models via Interacting Particle Langevin Dynamics

    stat.ML 2025-10 conditional novelty 6.0 of 10

    A particle Langevin algorithm (EBIPLA) trains latent energy-based models via maximum marginal likelihood, with convergence bounds and competitive image generation.

  9. Estimating Rate-Distortion Functions Using the Energy-Based Model

    cs.IT 2025-07 conditional novelty 6.0 of 10

    An energy-based training objective derived from the rate-distortion dual gives a single-network algorithm that estimates rate-distortion curves and reconstructs optimal conditional distributions.

  10. ToxBench: A Binding Affinity Prediction Benchmark with AB-FEP-Calculated Labels for Human Estrogen Receptor Alpha

    cs.LG 2025-07 conditional novelty 6.0 of 10

    ToxBench provides 8,770 AB-FEP-computed binding affinities for ERα-ligand complexes, and the proposed DualBind model learns to predict them orders of magnitude faster.

  11. Navigating the Latent Space Dynamics of Neural Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Autoencoders implicitly define a latent vector field whose attractors encode the model's memorized and generalized knowledge, enabling data-free probing and out-of-distribution detection.

  12. Joint Flow Matching for Generator-Consistent Classification

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Assigning images and labels opposite endpoints in a flow gives one model that both generates and classifies, with forward and backward passes sampling from the same joint.

  13. A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Tunable energy landscapes whose thermal averages equal sigmoid, softmax, and matrix-vector products can, in principle, form the basis of a low-energy analog computer, with a superconducting double-well device as a fir...

  14. Jarzynski Reweighting and Sampling Dynamics for Training Energy-Based Models: Theoretical Analysis of Different Transition Kernels

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Jarzynski reweighting, which estimates normalization constants from out-of-equilibrium paths, is shown to apply to a broad class of sampling kernels, including drift-based stochastic interpolants and RBM Gibbs samplin...

Pith tools