REVIEW 12 cited by
Energy-Based Diffusion Language Models for Text Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Despite remarkable progress in autoregressive language models, alternative generative paradigms beyond left-to-right generation are still being actively explored. Discrete diffusion models, with the capacity for parallel generation, have recently emerged as a promising alternative. Unfortunately, these models still underperform the autoregressive counterparts, with the performance gap increasing when reducing the number of sampling steps. Our analysis reveals that this degradation is a consequence of an imperfect approximation used by diffusion models. In this work, we propose Energy-based Diffusion Language Model (EDLM), an energy-based model operating at the full sequence level for each diffusion step, introduced to improve the underlying approximation used by diffusion models. More specifically, we introduce an EBM in a residual form, and show that its parameters can be obtained by leveraging a pretrained autoregressive model or by finetuning a bidirectional transformer via noise contrastive estimation. We also propose an efficient generation algorithm via parallel important sampling. Comprehensive experiments on language modeling benchmarks show that our model can consistently outperform state-of-the-art diffusion models by a significant margin, and approaches autoregressive models' perplexity. We further show that, without any generation performance drop, our framework offers a 1.3$\times$ sampling speedup over existing diffusion models. Reproduced code is available at https://github.com/MinkaiXu/Energy-Diffusion-LLM.
Forward citations
Cited by 12 Pith papers
-
Gumbel Distillation for Parallel Text Generation
Conditioning parallel decoders on Gumbel noise sampled from an autoregressive teacher's Gumbel-Max process improves generation quality on LM1B and OpenWebText.
-
IDLM: Inverse-distilled Diffusion Language Models
IDLM distills pretrained discrete diffusion language models into few-step generators, cutting inference steps by 4–64× with roughly matched GenPPL and entropy.
-
Fine-Tuning Masked Diffusion for Provable Self-Correction
PRISM fine-tunes any masked diffusion model with a binary-cross-entropy loss so its new head provably estimates per-token quality p(x_i=y_i|y⊕m_i) and can remask low-quality tokens at inference.
-
Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
TarFlowLM models language in a continuous latent space with transformer-based autoregressive normalizing flows, using mixture-CDF and Rosenblatt couplings, and reports competitive NELBO on TEXT8 and OpenWebText.
-
Theoretical Benefit and Limitation of Diffusion Language Model
Masked diffusion language models have a metric-dependent efficiency tradeoff: near-optimal perplexity in constant steps, but sequence-level correctness needs linearly many steps in the worst case.
-
Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions
Masked diffusion models trained order-agnostically can solve puzzles better than autoregressive models when inference unmasking order is chosen adaptively by confidence.
-
Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images?
Diffusion models trained with denoising score matching can follow coarse image rules but cannot reliably reproduce fine-grained inter-feature rules, and a two-layer network analysis shows a constant error for such rules.
-
Categorical Schr\"odinger Bridge Matching
The paper proves that discrete-time iterative Markovian fitting converges to the Schrödinger Bridge on finite discrete spaces and introduces CSBM, a practical matching algorithm for categorical data.
-
Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers
SNLP reduces symbolic FHE bootstraps from 53 to 20 on a 0.5B model with +1.2% PPL degradation and lower polynomial-error amplification than sequential inference.
-
On the Separability of Information in Diffusion Models
Diffusion models devote most of their information budget to class-agnostic texture, and the small class-relevant slice is what classifier-free guidance amplifies.
-
A Survey on Diffusion Language Models
A comprehensive survey of diffusion language models covering taxonomy, training and inference techniques, and comparisons with autoregressive models.
-
Generative AI for Autonomous Driving: A Review
A review of generative models (VAEs, GANs, diffusion, transformers, LLMs) applied to map generation, scenario generation, trajectory prediction, and motion planning for autonomous driving.
Discussion (0). Continue with ORCID to comment.