REVIEW 8 cited by
Self-conditioned Embedding Diffusion for Text Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Can continuous diffusion models bring the same performance breakthrough on natural language they did for image generation? To circumvent the discrete nature of text data, we can simply project tokens in a continuous space of embeddings, as is standard in language modeling. We propose Self-conditioned Embedding Diffusion, a continuous diffusion mechanism that operates on token embeddings and allows to learn flexible and scalable diffusion models for both conditional and unconditional text generation. Through qualitative and quantitative evaluation, we show that our text diffusion models generate samples comparable with those produced by standard autoregressive language models - while being in theory more efficient on accelerator hardware at inference time. Our work paves the way for scaling up diffusion models for text, similarly to autoregressive models, and for improving performance with recent refinements to continuous diffusion.
Forward citations
Cited by 8 Pith papers
-
Hacking Generative Perplexity: Why Unconditional Text Evaluation Needs Distributional Metrics
Zero-parameter naive samplers achieve state-of-the-art generative perplexity while producing incoherent text, proving the metric is unsound; distributional divergences like MAUVE and energy distance correctly rank the...
-
Subliminal Clocks: Latent Time Modelling in Diffusion Language Models
DLMs encode a decodable latent timestep signal in residual activations that can be steered to predictably change model confidence and entropy.
-
CANDI: Hybrid Discrete-Continuous Diffusion Models
CANDI combines masked and Gaussian corruption in one noising process, letting discrete diffusion models use continuous gradients for joint updates and guidance.
-
Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving
On Waymo Sim Agents, LLM-style tokenization, positional embeddings, pretraining, RL post-training, and test-time search can be adapted to improve motion generation, but not all transfer without domain-specific changes.
-
DiffusionGemma Technical Report
DiffusionGemma turns a 25B-parameter open-weight autoregressive MoE into a discrete diffusion model that produces ~20 tokens per forward pass and ~1,500 tokens/sec on one H100, establishing a speed-quality tradeoff po...
-
DLM-One: Diffusion Language Models for One-Step Sequence Generation
DLM-One distills a continuous diffusion language model into a one-step student, achieving roughly 500x inference speedup while staying within a few percent of the teacher on BLEU, ROUGE, and BERTScore, with substantia...
-
The Philosophy and Physics of Duality
A philosophical monograph that surveys dualities across physics and proposes a 'geometric view of theories' for theoretical equivalence, realism, and explanation.
-
A Survey on Diffusion Language Models
A comprehensive survey of diffusion language models covering taxonomy, training and inference techniques, and comparisons with autoregressive models.
Discussion (0). Continue with ORCID to comment.