Pith. sign in

REVIEW 10 cited by

Simplified and Generalized Masked Diffusion for Discrete Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.04329 v4 pith:HBOV2F7K submitted 2024-06-06 cs.LG stat.ML

classification cs.LGstat.ML
keywords modelsdiffusionmaskeddiscretemodelingautoregressivedataframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Masked (or absorbing) diffusion is actively explored as an alternative to autoregressive models for generative modeling of discrete data. However, existing work in this area has been hindered by unnecessarily complex model formulations and unclear relationships between different perspectives, leading to suboptimal parameterization, training objectives, and ad hoc adjustments to counteract these issues. In this work, we aim to provide a simple and general framework that unlocks the full potential of masked diffusion models. We show that the continuous-time variational objective of masked diffusion models is a simple weighted integral of cross-entropy losses. Our framework also enables training generalized masked diffusion models with state-dependent masking schedules. When evaluated by perplexity, our models trained on OpenWebText surpass prior diffusion language models at GPT-2 scale and demonstrate superior performance on 4 out of 5 zero-shot language modeling tasks. Furthermore, our models vastly outperform previous discrete diffusion models on pixel-level image modeling, achieving 2.75 (CIFAR-10) and 3.40 (ImageNet 64x64) bits per dimension that are better than autoregressive models of similar sizes. Our code is available at https://github.com/google-deepmind/md4.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hacking Generative Perplexity: Why Unconditional Text Evaluation Needs Distributional Metrics

    cs.CL 2026-06 conditional novelty 7.0 of 10

    Zero-parameter naive samplers achieve state-of-the-art generative perplexity while producing incoherent text, proving the metric is unsound; distributional divergences like MAUVE and energy distance correctly rank the...

  2. CANDI: Hybrid Discrete-Continuous Diffusion Models

    cs.LG 2025-10 conditional novelty 6.0 of 10

    CANDI combines masked and Gaussian corruption in one noising process, letting discrete diffusion models use continuous gradients for joint updates and guidance.

  3. Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows

    cs.LG 2025-07 conditional novelty 6.0 of 10

    TarFlowLM models language in a continuous latent space with transformer-based autoregressive normalizing flows, using mixture-CDF and Rosenblatt couplings, and reports competitive NELBO on TEXT8 and OpenWebText.

  4. Neuro-Symbolic Generative Diffusion Models for Physically Grounded, Robust, and Safe Generation

    cs.LG 2025-06 conditional novelty 6.0 of 10

    NSD adds iterative constraint projection to both continuous and discrete diffusion sampling, achieving near-zero constraint violations across molecular, robotic, material, and language generation tasks.

  5. Learning Distributions over Permutations and Rankings with Factorized Representations

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Factorized codes for permutations let standard transformers learn arbitrary permutation distributions with guaranteed validity, and the paper adds a fast insertion-vector decoding theorem plus two new benchmarks.

  6. Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Adaptive Classifier-Free Guidance (A-CFG) re-masks low-confidence tokens in the unconditional input at each generation step, improving reasoning and planning accuracy for masked diffusion language models.

  7. Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    In overparameterized diffusion models, generalization happens first and memorization starts later, with the memorization time growing linearly with dataset size.

  8. Theoretical Benefit and Limitation of Diffusion Language Model

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Masked diffusion language models have a metric-dependent efficiency tradeoff: near-optimal perplexity in constant steps, but sequence-level correctness needs linearly many steps in the worst case.

  9. Masked Diffusion Language Models with Frequency-Informed Training

    cs.CL 2025-09 conditional novelty 4.0 of 10

    Masked diffusion language models trained on 100M words match a hybrid GPT-BERT baseline on BabyLM tests, with a rare-word-focused masking variant.

  10. Adaptive Mask-guided K-space Diffusion for Accelerated MRI Reconstruction

    eess.IV 2025-06 reject novelty 4.0 of 10

    AMDM reconstructs undersampled MRI by masking k-space frequency components with adaptive masks inside a diffusion model, and reports large PSNR gains over baseline methods.

Pith tools