REVIEW 14 cited by
How to Train Your Energy-Based Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Energy-Based Models (EBMs), also known as non-normalized probabilistic models, specify probability density or mass functions up to an unknown normalizing constant. Unlike most other probabilistic models, EBMs do not place a restriction on the tractability of the normalizing constant, thus are more flexible to parameterize and can model a more expressive family of probability distributions. However, the unknown normalizing constant of EBMs makes training particularly difficult. Our goal is to provide a friendly introduction to modern approaches for EBM training. We start by explaining maximum likelihood training with Markov chain Monte Carlo (MCMC), and proceed to elaborate on MCMC-free approaches, including Score Matching (SM) and Noise Constrastive Estimation (NCE). We highlight theoretical connections among these three approaches, and end with a brief survey on alternative training methods, which are still under active research. Our tutorial is targeted at an audience with basic understanding of generative models who want to apply EBMs or start a research project in this direction.
Forward citations
Cited by 14 Pith papers
-
Bayesian Experimental Design via Score Matching
SCOREBED isolates EIG double intractability in a policy-independent score-matching stage, then trains design policies with a singly intractable gradient estimator, enabling cheap multi-policy selection.
-
Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Motion Generation
STREAM decouples text (via AdaLN) from music (via energy-based BEAM attention) to generate editable, musically aligned dance motions with a new annotated dataset and editability metric.
-
Energy-based Tissue Manifolds for Longitudinal Multiparametric MRI Analysis
Patient-specific energy manifolds from baseline mpMRI scans act as fixed geometric references to monitor longitudinal evolution of voxel distributions in sequence space for neuro-oncology proof-of-concept cases.
-
Transferable Direct Prompt Injection via Activation-Guided MCMC Sampling
An activation-guided energy model plus MCMC sampling creates transferable direct prompt injection attacks in a black-box setting, reaching 49.6% average attack success across five LLMs.
-
Projected Energy Matching for Generative 3D Priors
Helmholtz Distillation plus Negative Caching amortizes Energy Matching to 3D CT volumes, producing a conservative prior that improves FID and sparse-view reconstruction over pure flow baselines.
-
Diffusion-based Annealed Boltzmann Generators : benefits, pitfalls and hopes
Even a perfect diffusion model yields poor annealed Boltzmann generators when coupled through first-order stochastic denoising kernels, while deterministic transport maps and second-order kernels improve; with learned...
-
Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
Under the chain rule of probability, autoregressive models and energy-based models are in exact bijection in function space, making the global optimum of teacher forcing equivalent to an energy-based model with implic...
-
Learning Latent Energy-Based Models via Interacting Particle Langevin Dynamics
A particle Langevin algorithm (EBIPLA) trains latent energy-based models via maximum marginal likelihood, with convergence bounds and competitive image generation.
-
Estimating Rate-Distortion Functions Using the Energy-Based Model
An energy-based training objective derived from the rate-distortion dual gives a single-network algorithm that estimates rate-distortion curves and reconstructs optimal conditional distributions.
-
ToxBench: A Binding Affinity Prediction Benchmark with AB-FEP-Calculated Labels for Human Estrogen Receptor Alpha
ToxBench provides 8,770 AB-FEP-computed binding affinities for ERα-ligand complexes, and the proposed DualBind model learns to predict them orders of magnitude faster.
-
Navigating the Latent Space Dynamics of Neural Models
Autoencoders implicitly define a latent vector field whose attractors encode the model's memorized and generalized knowledge, enabling data-free probing and out-of-distribution detection.
-
Joint Flow Matching for Generator-Consistent Classification
Assigning images and labels opposite endpoints in a flow gives one model that both generates and classifies, with forward and backward passes sampling from the same joint.
-
A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing
Tunable energy landscapes whose thermal averages equal sigmoid, softmax, and matrix-vector products can, in principle, form the basis of a low-energy analog computer, with a superconducting double-well device as a fir...
-
Jarzynski Reweighting and Sampling Dynamics for Training Energy-Based Models: Theoretical Analysis of Different Transition Kernels
Jarzynski reweighting, which estimates normalization constants from out-of-equilibrium paths, is shown to apply to a broad class of sampling kernels, including drift-based stochastic interpolants and RBM Gibbs samplin...
Discussion (0). Sign in to comment.