Pith. sign in

REVIEW 4 cited by

Argmax Flows and Multinomial Diffusion: Learning Categorical Distributions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.05379 v3 pith:RBEIBHN4 submitted 2021-02-10 stat.ML cs.CLcs.LG

classification stat.MLcs.CLcs.LG
keywords diffusionargmaxflowscategoricaldatamultinomialcontinuousgenerative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative flows and diffusion models have been predominantly trained on ordinal data, for example natural images. This paper introduces two extensions of flows and diffusion for categorical data such as language or image segmentation: Argmax Flows and Multinomial Diffusion. Argmax Flows are defined by a composition of a continuous distribution (such as a normalizing flow), and an argmax function. To optimize this model, we learn a probabilistic inverse for the argmax that lifts the categorical data to a continuous space. Multinomial Diffusion gradually adds categorical noise in a diffusion process, for which the generative denoising process is learned. We demonstrate that our method outperforms existing dequantization approaches on text modelling and modelling on image segmentation maps in log-likelihood.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 36 citations worldwide. Full citation record

  1. D$e^+e^-$ffusion: Capturing the Beam-Beam Physics of $e^+e^-$ Collisions with Diffusion Models

    hep-ph 2026-07 conditional novelty 6.0 of 10

    A diffusion model trained on GuineaPig++ reproduces FCC-ee beam-induced pair-production distributions at particle and detector level, about 10^4 times faster.

  2. Scaling Probabilistic Circuits via Monarch Matrices

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Structured Monarch matrices, derived from circuit multiplication, let probabilistic circuits scale to larger hidden sizes and beat prior tractable models at lower FLOP cost.

  3. Masked Diffusion Language Models with Frequency-Informed Training

    cs.CL 2025-09 conditional novelty 4.0 of 10

    Masked diffusion language models trained on 100M words match a hybrid GPT-BERT baseline on BabyLM tests, with a rare-word-focused masking variant.

  4. The Good, The Efficient and the Inductive Biases: Exploring Efficiency in Deep Learning Through the Use of Inductive Biases

    cs.LG 2024-11 conditional novelty 3.0 of 10

    A dissertation synthesizing the author's papers on continuous kernel convolutions and symmetry-preserving architectures, claiming these inductive biases improve deep learning efficiency.

Pith tools