REVIEW 4 major objections 6 minor 1 cited by
[MASK] is All You Need
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read One masked-token recipe unifies diffusion, masked generation, and segmentation in a single model.
desk verdict A useful large-scale empirical study of discrete flow matching with an overclaimed SOTA headline and a real train/inference gap in the segmentation-as-unmasking claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Discrete interpolants driven by a masking schedule $\kappa_t$ are the central object: the interpolated state is $p_{t|0,1}(x|x_0,x_1) = (1-\kappa_t)\delta_{[M]}(x) + \kappa_t\delta_{x_1}(x)$, and the vector field to learn is $u_t(x_t) = \frac{\dot\kappa_t}{1-\kappa_t}[p_{1|t}(x_1|x_t,t;\theta) - \delta_{x_t}(x)]$. The schedule ($\kappa_t = t$ linear, cosine, quadratic, and others) controls how gradually real tokens replace the [MASK] token. The training signal is masked cross-entropy with a weighting $w(t)$, where only positions that are [MASK] in $x_t$ contribute; the paper finds this masking necessary to avoid overfitting in vision and finds $w(t)=1$ better than the ELBO-derived weight $\dot\kappa_t/(1-\kappa_t)$. The other load-bearing design choice is the implicit-timestep network $p(x_1|x_t;\theta)$, which omits $t$ and therefore behaves like a masked generative model during sampling; a final argmax over logits at the last step, called churning, removes leftover [MASK] tokens and fixes scheduler misalignment.
What would settle it
Probe an implicit-timestep model's ability to read corruption level from the mask alone: train it on masked inputs generated by two different schedulers or timesteps that produce the same visible [MASK] pattern for the same clean sequence, then see whether supplying the true timestep changes the predicted clean-token distribution. If the timestep helps, the mask pattern alone does not encode the corruption level; equivalently, if a probe classifier cannot recover the remaining mask ratio from $x_t$ accurately, the implicit model has no channel for that information and the central premise fails.
Extended reading notes
Core claim
The central claim is that unmasking is a common language for generation and dense prediction. Concretely, the paper defines a discrete interpolant under a masking schedule $\kappa_t$, with $p_{t|0,1}(x|x_0,x_1) = (1-\kappa_t)\delta_{x_0}(x) + \kappa_t\delta_{x_1}(x)$, where $x_0$ is the all-[MASK] state and $x_1$ is the clean token sequence. A network trained with masked cross-entropy to predict $x_1$ from $x_t$ acts as the vector field of a discrete flow, and the same network can be sampled as an explicit-timestep diffusion model, as an implicit-timestep model, or with masked-generative greedy unmasking. The paper's distinctive move is dropping the timestep: because the scheduler is monotone, the pattern of [MASK] tokens in $x_t$ already encodes how corrupted the input is. That is the step that connects discrete diffusion to masked generative models and lets segmentation be framed as unmasking. With image and segmentation-mask tokens concatenated and masked under a shared schedule, one training run serves image-conditioned segmentation, mask-conditioned image generation, and joint modeling of the two modalities.
Load-bearing premise
The load-bearing premise is that a masked token sequence itself reveals how much corruption has been applied, so the network can safely ignore the timestep; if two different timesteps can produce indistinguishable masked inputs, the implicit model conflates corruption levels and the claimed bridge to masked generative models weakens.
Editorial extensions
If this is right
- A single trained discrete model can be sampled in explicit-timestep diffusion style, implicit-timestep diffusion style, or masked-generative greedy style, using the same weights.
- Segmentation becomes an unmasking task: after joint training on image and mask token pairs, one checkpoint can return segmentation masks conditioned on images or images conditioned on masks.
- Removing the explicit timestep yields simpler, order-flexible sampling that is upper-bounded by the token length and can handle row-by-row or editing-style schedules where a global timestep is hard to define.
- Masked cross-entropy with a constant weight plus a final argmax step improves low-step sampling and corrects the mismatch when sampling with a scheduler different from training.
- The same discrete recipe scales from images to video; adapting the continuous-state Latte model to discrete tokens gives a better FVD on FaceForensics.
Reading between the lines
- If the implicit-timestep premise holds for discrete tokens, the same time-agnostic idea is worth testing in continuous-state diffusion for editing and arbitrary-order sampling; the paper notes this possibility but does not demonstrate it.
- Segmentation-as-unmasking should extend to any tokenizable dense prediction target, such as depth, surface normals, or object detection, since the paper's argument only relies on a shared discrete vocabulary and a masked joint distribution.
- A practical testable extension is to use the residual [MASK]-token rate after sampling as a proxy for scheduler misalignment: if churning helps, the failure mode is mostly leftover masked tokens, and monitoring that rate could predict when a new scheduler is safe to use.
- Because the paper identifies irreversible unmasking as the main limitation, adding a smoothing or corrector term to allow remasking could convert the framework into a fully reversible stochastic interpolant; this is a future direction the paper only sketches.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces "Discrete Interpolants," a discrete-state flow-matching framework for vision. It defines masking schedules κt that interpolate between fully masked tokens and data tokens, trains a model with masked cross-entropy to predict original tokens, and supports explicit-timestep, implicit-timestep, and MaskGIT-style sampling. It also proposes training once on image–segmentation-mask pairs to model a joint distribution and then sampling conditionally in either direction, framing semantic segmentation as an unmasking process. Experiments are reported on ImageNet 256, MS-COCO, FaceForensics, and Cityscapes, with the central claims being a unified design-space analysis of masked generative and discrete diffusion models and a demonstration that discriminative tasks can be recast as unmasking.
Significance. If the claims are supported, the paper would provide a useful single discrete-token recipe for generation, conditional generation, and dense prediction, and its explicit-versus-implicit timestep analysis is a practical contribution. The paper ships extensive ablations of sampling steps, softmax temperature, CFG scale, schedulers, Gumbel noise, and argmax churning, and it is transparent that several of these are empirical tuning findings rather than theoretical predictions. The theoretical scaffolding is largely imported from discrete flow matching and masked diffusion literature, so the novelty lies in the vision-domain unification and empirical study rather than in new theory. The results are potentially valuable but currently contain an overstatement about MS-COCO state-of-the-art and a training–inference mismatch in the conditional segmentation protocol that undermines the strongest new claim.
major comments (4)
- [§4.2.1, Table 2] The sentence "Our method achieves state-of-the-art performance compared to both continuous-state and discrete-state models" is contradicted by Table 2: U-ViT attains FID 5.48 while the Implicit Timestep Model attains 5.65, so the continuous-state comparison is not state-of-the-art, and the Explicit Timestep Model is 6.03. The defensible claim is state-of-the-art among the discrete-state baselines listed (VQ-Diffusion at 19.75). Because the gap to U-ViT is 0.17 FID and no error bars or multiple seeds are reported, the present wording overstates the result; please restrict the claim to discrete-state models or reword it as competitive, and report variance or at least note single-run metrics.
- [§3.4–§3.5, Eq. (5), Table 5] Eq. (5) trains on z_t = x_t ⊕ y_t with both modalities corrupted by the same masking schedule, but the conditional sampling used for Table 5 and Figs. 15–16 conditions on one clean modality while unmasking the other. For typical token counts (Lx, Ly on the order of 256–1024), a training example in which all condition tokens are clean and all target tokens are masked occurs with probability (1−κt)^{Lx} κt^{Ly}, which is negligible except at t near 0 or 1. The model is therefore evaluated on inputs far outside the training distribution, so the mIOU/FID numbers in Table 5 do not establish the claimed "train once, flexible conditional sampling" recipe. The paper should train with an unmasked-condition protocol, mask the condition at inference according to κt, or provide an explicit ablation justifying the mismatch. In addition, §4.1 mentions a 0.1 conditional dropout for classifier-free guidance, but Eq. (5) contains no condition term, so it is unclear how the unconditional branch in Eq. (6) was trained.
- [§4.2.1, Table 4] The abstract and contribution list claim "competitive" results on FaceForensics compared to counterpart discrete-state models, but Table 4 compares only against Latte, a continuous-state model. No discrete-state video baseline is included, so the comparison class stated in the claims is not actually evaluated. Please add discrete-state video baselines (for example, MAGVIT or a discrete adaptation of Latte) or explicitly restrict the claim to the continuous-state comparison shown.
- [§4.2.2, Table 5] Table 5 reports mIOU of 89.1 and 90.1 and FID of 34.4 and 33.8 on Cityscapes without any baseline segmentation method, so the discriminative value of the segmentation-as-unmasking framework is not quantified relative to existing approaches. Even if the training–inference mismatch above is resolved, the reader cannot tell whether these mIOU numbers are strong, and the statement in §4.2.2 that the [MASK] token can be leveraged to reframe discriminative tasks should be supported by a comparison to at least one standard segmentation model or to prior discrete diffusion discriminative work.
minor comments (6)
- [Section 3.5] The heading "Classifier-free Gudiance" contains a typo and should read "Classifier-free Guidance."
- [Section 5] The opening sentence "Our work Stochastic Interpolant extends discrete flow matching theory to vision tasks" is imprecise; the theory is from Gat et al. [22] and the paper builds on and applies it, so the sentence should say "builds on" or "applies" rather than "extends," or it should clearly specify what is newly extended.
- [Algorithm 1] The notation in Algorithm 1 is informal: the line `p(x1|xt+Δt, t+Δt; θ) ← Cat[δxt(t + Δt) + ut(xt)Δt]` conflates a distribution with a sampling rule, uses `δxt(t + Δt)` ambiguously, and the pseudocode should explicitly state that only masked positions are resampled and unmasked positions are carried over, matching the prose description.
- [Section 4.1] The paper should state whether code and checkpoints are released; the FID, FVD, and mIOU tables cannot otherwise be independently reproduced, especially given the sensitivity of FID to the evaluation protocol.
- [Section 3.4] The sentence "we share the mask schedule between two different modalities, we find it empirically works well" is a run-on sentence, and the sharing assumption is presented only as an empirical finding; a brief explanation of why a shared schedule is expected to work would help readers assess the design choice.
- [Table 1 and Figures 2–3] Scheduler names are not fully harmonized: Table 1 lists "Arccos," Figure 2 uses "Arcsine" and "Arccos," and Figure 3 uses "ArcCos"; please unify the naming convention across tables, figures, and text.
Circularity Check
No load-bearing circularity: the central framework is imported from externally cited discrete flow matching, and self-citations are not load-bearing.
full rationale
The paper's central derivation is not circular. The Discrete Interpolants construction (Eq. 1), the cross-entropy training loss (Eqs. 2-4), and the masking/weighting modifications are explicitly built on Discrete Flow Matching [22,63] and related external discrete-diffusion work, not on the authors' own prior results. The Implicit Timestep Model is justified by an explicitly stated assumption: for a reversible scheduler, the mask ratio encodes the timestep, so the timestep dependence can be dropped. That is a stated modeling premise, not a conclusion defined in terms of its own output. The segmentation-as-unmasking extension uses the joint loss in Eq. 5 and standard classifier-free guidance in Eq. 6; the skeptic's observation that conditional sampling feeds a clean condition while training corrupts both modalities is a train/inference distribution mismatch and a validity concern, not a circular reduction. The ablation findings (temperature sweet spot, CFG strength, argmax churning) are empirical tuning results, not fitted parameters renamed as predictions. Several self-citations appear in the related work and appendix (e.g., [19,20,33-37,68,76]), but none carries a load-bearing argument for the central claims; the theoretical backbone is externally cited. The paper's COCO 'state-of-the-art' wording is also imprecise relative to Table 2, where U-ViT achieves 5.48 FID against the method's 5.65, but that is a correctness/accuracy issue, not circularity. No step was found where a claimed prediction or unification reduces by construction to its own inputs.
Assumptions & free parameters
free parameters (8)
- Masking schedule κt =
linear by default; root, cosine, arccos, quadratic, cubic compared
- Weighting w(t) =
w(t) = 1
- Softmax temperature =
~0.8 for ETM/ITM, ~1.2 for MGM-style sampling
- Classifier-free guidance scale ω =
~3 optimal in ablations; final runs use values like 3.0 or 4.5
- Gumbel noise =
linear annealing with Gumbel temperature 4.5 (for MGM-style sampling)
- Top-p =
0.9
- argmax churning =
applied to logits at the last sampling step
- Number of sampling steps =
1000 default, fixed step size
assumptions (6)
- standard math Kolmogorov equation for discrete-state probability paths
- domain assumption Interpolants operation can be factorized token-wise
- domain assumption Masked data xt inherently contains timestep information
- ad hoc to paper Sharing one masking schedule across modalities works
- domain assumption Classifier-free guidance extends to discrete tokens
- domain assumption Pretrained SD-VQ-F8 tokenizer gives a suitable discrete state space
Cite this review
Pith. "Pith review of [MASK] is All You Need." pith.science (2026). https://pith.science/paper/J6OJDROQ
@misc{pith2026241206787,
author = {Pith},
title = {Pith review of: [MASK] is All You Need},
year = {2026},
howpublished = {\url{https://pith.science/paper/J6OJDROQ}},
note = {Machine review of arXiv:2412.06787}
}
read the original abstract
In generative models, two paradigms have gained attraction in various applications: next-set prediction-based Masked Generative Models and next-noise prediction-based Non-Autoregressive Models, e.g., Diffusion Models. In this work, we propose using discrete-state models to connect them and explore their scalability in the vision domain. First, we conduct a step-by-step analysis in a unified design space across two types of models including timestep-independence, noise schedule, temperature, guidance strength, etc in a scalable manner. Second, we re-cast typical discriminative tasks, e.g., image segmentation, as an unmasking process from [MASK] tokens on a discrete-state model. This enables us to perform various sampling processes, including flexible conditional sampling by only training once to model the joint distribution. All aforementioned explorations lead to our framework named Discrete Interpolants, which enables us to achieve state-of-the-art or competitive performance compared to previous discrete-state based methods in various benchmarks, like ImageNet256, MS COCO, and video dataset FaceForensics. In summary, by leveraging [MASK] in discrete-state models, we can bridge Masked Generative and Non-autoregressive Diffusion models, as well as generative and discriminative tasks.
Figures
Figures from the paper (16 more)
Forward citations
Cited by 1 Pith paper
-
MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing
A masking-augmented diffusion objective plus pause-token inference scaling modestly improves instruction adherence and source preservation for OmniGen-based image editing.
Reference graph
Works this paper leans on
-
[1]
Building nor- malizing flows with stochastic interpolants
Michael S Albergo and Eric Vanden-Eijnden. Building nor- malizing flows with stochastic interpolants. In ICLR, 2023. 2
2023
-
[2]
Stochastic interpolants: A unifying framework for flows and diffusions
Michael S Albergo, Nicholas M Boffi, and Eric Vanden- Eijnden. Stochastic interpolants: A unifying framework for flows and diffusions. arXiv preprint arXiv:2303.08797, 2023. 2, 14
arXiv 2023
-
[3]
Self-supervised learning from images with a joint-embedding predictive architecture
Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bo- janowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas. Self-supervised learning from images with a joint-embedding predictive architecture. In CVPR, pages 15619–15629, 2023. 5
2023
-
[4]
Structured denoising diffusion models in discrete state-spaces
Jacob Austin, Daniel D Johnson, Jonathan Ho, Daniel Tarlow, and Rianne Van Den Berg. Structured denoising diffusion models in discrete state-spaces. NeurIPS, 34:17981–17993,
-
[5]
Cold diffusion: Inverting arbitrary image transforms without noise
Arpit Bansal, Eitan Borgnia, Hong-Min Chu, Jie Li, Hamid Kazemi, Furong Huang, Micah Goldblum, Jonas Geiping, and Tom Goldstein. Cold diffusion: Inverting arbitrary image transforms without noise. NeurIPS, 36, 2024. 5
2024
-
[6]
All are worth words: a vit backbone for score-based diffusion models
Fan Bao, Chongxuan Li, Yue Cao, and Jun Zhu. All are worth words: a vit backbone for score-based diffusion models. CVPR, 2023. 5, 6
2023
-
[7]
A pytorch reproduction of masked generative image transformer, 2023
Victor Besnier and Mickael Chen. A pytorch reproduction of masked generative image transformer, 2023. 7
2023
-
[8]
Generative flows on discrete state-spaces: Enabling multimodal flows with applications to protein co-design
Andrew Campbell, Jason Yim, Regina Barzilay, Tom Rain- forth, and Tommi Jaakkola. Generative flows on discrete state-spaces: Enabling multimodal flows with applications to protein co-design. ICML, 2024. 2, 3, 14, 15
2024
Show all 94 references
-
[9]
Maskgit: Masked generative image transformer
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T Freeman. Maskgit: Masked generative image transformer. In CVPR, pages 11315–11325, 2022. 1, 2, 4, 6, 7, 8, 14, 15
2022
-
[10]
Murphy, William T
Huiwen Chang, Han Zhang, Jarred Barber, AJ Maschinot, Jos´e Lezama, Lu Jiang, Ming-Hsuan Yang, Kevin P. Murphy, William T. Freeman, Michael Rubinstein, Yuanzhen Li, and Dilip Krishnan. Muse: Text-to-image generation via masked generative transformers. ICML, 2023. 2
2023
-
[11]
Denoising with a joint-embedding predictive architecture,
Dengsheng Chen, Jie Hu, Xiaoming Wei, and Enhua Wu. Denoising with a joint-embedding predictive architecture,
-
[12]
Re-imagen: Retrieval-augmented text-to-image gen- erator
Wenhu Chen, Hexiang Hu, Chitwan Saharia, and William W Cohen. Re-imagen: Retrieval-augmented text-to-image gen- erator. arXiv preprint arXiv:2209.14491, 2022. 6
2022 arXiv
-
[13]
Yunlu Chen, Vincent Tao Hu, Efstratios Gavves, Thomas Mensink, Pascal Mettes, Pengwan Yang, and Cees G. M. Snoek. Pointmixup: Augmentation for point clouds. In ECCV, 2020. 14
2020
-
[14]
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In CVPR, 2009. 16
2009
-
[15]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. In NeurIPS, 2021. 6
2021
-
[16]
Continuous diffusion for categorical data
Sander Dieleman, Laurent Sartran, Arman Roshannai, Niko- lay Savinov, Yaroslav Ganin, Pierre H Richemond, Arnaud Doucet, Robin Strudel, Chris Dyer, Conor Durkan, et al. Continuous diffusion for categorical data. arXiv preprint arXiv:2211.15089, 2022. 9, 14
2022 arXiv
-
[17]
Imagebart: Bidirectional context with multinomial diffusion for autoregressive image synthesis
Patrick Esser, Robin Rombach, Andreas Blattmann, and Bjorn Ommer. Imagebart: Bidirectional context with multinomial diffusion for autoregressive image synthesis. NeurIPS, 34: 3518–3532, 2021. 15
2021
-
[18]
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In CVPR,
-
[19]
Diffusion mod- els and representation learning: A survey
Michael Fuest, Pingchuan Ma, Ming Gui, Johannes S Fis- cher, Vincent Tao Hu, and Bjorn Ommer. Diffusion mod- els and representation learning: A survey. arXiv preprint arXiv:2407.00783, 2024. 1, 16
2024 arXiv
-
[20]
Distillation of diffusion features for semantic correspondence
Frank Fundel, Johannes Schusterbauer, Vincent Tao Hu, and Bj¨orn Ommer. Distillation of diffusion features for semantic correspondence. WACV, 2025. 16
2025
-
[21]
Make-a-scene: Scene-based text-to-image generation with human priors
Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman. Make-a-scene: Scene-based text-to-image generation with human priors. In ECCV, pages 89–106. Springer, 2022. 6
2022
-
[22]
Discrete flow matching
Itai Gat, Tal Remez, Neta Shaul, Felix Kreuk, Ricky TQ Chen, Gabriel Synnaeve, Yossi Adi, and Yaron Lipman. Discrete flow matching. NeurIPS, 2024. 1, 2, 3, 5, 14, 15
2024
-
[23]
Making llama see and draw with seed tokenizer
Yuying Ge, Sijie Zhao, Ziyun Zeng, Yixiao Ge, Chen Li, Xintao Wang, and Ying Shan. Making llama see and draw with seed tokenizer. ICLR, 2024. 15
2024
-
[24]
Scaling diffusion language models via adaptation from autoregressive models
Shansan Gong, Shivam Agarwal, Yizhe Zhang, Jiacheng Ye, Lin Zheng, Mukai Li, Chenxin An, Peilin Zhao, Wei Bi, Jiawei Han, et al. Scaling diffusion language models via adaptation from autoregressive models. arXiv preprint arXiv:2410.17891, 2024. 2, 4
-
[25]
Your classifier is secretly an energy based model and you should treat it like one
Will Grathwohl, Kuan-Chieh Wang, J¨orn-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky. Your classifier is secretly an energy based model and you should treat it like one. arXiv preprint arXiv:1912.03263,
1912 arXiv
-
[26]
Vector quan- tized diffusion model for text-to-image synthesis
Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. Vector quan- tized diffusion model for text-to-image synthesis. In CVPR, pages 10696–10706, 2022. 2, 6
2022
-
[27]
Fischer, Ulrich Prestel, Pingchuan Ma, Olga Grebenkova Dmytro Kotovenko, Stefan Andreas Baumann, Vincent Tao Hu, and Bj¨orn Ommer
Ming Gui, Johannes S. Fischer, Ulrich Prestel, Pingchuan Ma, Olga Grebenkova Dmytro Kotovenko, Stefan Andreas Baumann, Vincent Tao Hu, and Bj¨orn Ommer. Depthfm: Fast monocular depth estimation with flow matching. In AAAI,
-
[28]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In CVPR, 2022. 15
2022
-
[29]
Dice: Discrete inversion enabling controllable editing for multinomial diffusion and masked generative models
Xiaoxiao He, Ligong Han, Quan Dao, Song Wen, Minhao Bai, Di Liu, Han Zhang, Martin Renqiang Min, Felix Juefei- Xu, Chaowei Tan, et al. Dice: Discrete inversion enabling controllable editing for multinomial diffusion and masked generative models. arXiv preprint arXiv:2410.08207...
-
[30]
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. In NeurIPS Workshop, 2021. 5
2021
-
[31]
Denoising diffu- sion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. In NeurIPS, 2020. 1, 14
2020
-
[32]
Cascaded diffusion models for high fidelity image generation.Journal of Machine Learning Research, 23(47):1–33, 2022
Jonathan Ho, Chitwan Saharia, William Chan, David J Fleet, Mohammad Norouzi, and Tim Salimans. Cascaded diffusion models for high fidelity image generation.Journal of Machine Learning Research, 23(47):1–33, 2022. 6
2022
-
[33]
Guided dif- fusion from self-supervised diffusion features
Vincent Tao Hu, Yunlu Chen, Mathilde Caron, Yuki M Asano, Cees GM Snoek, and Bjorn Ommer. Guided dif- fusion from self-supervised diffusion features. arXiv preprint arXiv:2312.08825, 2023. 16
2023 arXiv
-
[34]
Self-guided diffusion mod- els
Vincent Tao Hu, David W Zhang, Yuki M Asano, Gertjan J Burghouts, and Cees GM Snoek. Self-guided diffusion mod- els. In CVPR, pages 18413–18422, 2023. 1
2023
-
[35]
Zigma: A dit-style zigzag mamba diffusion model
Vincent Tao Hu, Stefan Andreas Baumann, Ming Gui, Olga Grebenkova, Pingchuan Ma, Johannes Schusterbauer, and Bj¨orn Ommer. Zigma: A dit-style zigzag mamba diffusion model. In ECCV, 2024. 1
2024
-
[36]
Flow matching for conditional text generation in a few sampling steps
Vincent Tao Hu, Di Wu, Yuki Asano, Pascal Mettes, Basura Fernando, Bj¨orn Ommer, and Cees Snoek. Flow matching for conditional text generation in a few sampling steps. In EACL, pages 380–392, 2024. 15
2024
-
[37]
Latent space editing in transformer- based flow matching
Vincent Tao Hu, Wei Zhang, Meng Tang, Pascal Mettes, Deli Zhao, and Cees Snoek. Latent space editing in transformer- based flow matching. In AAAI, pages 2247–2255, 2024. 1
2024
-
[38]
Layoutdm: Discrete diffusion model for controllable layout generation
Naoto Inoue, Kotaro Kikuchi, Edgar Simo-Serra, Mayu Otani, and Kota Yamaguchi. Layoutdm: Discrete diffusion model for controllable layout generation. In CVPR, pages 10167–10176,
-
[39]
Bert: Pre-training of deep bidirectional transform- ers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. Bert: Pre-training of deep bidirectional transform- ers for language understanding. InProceedings of naacL-HLT, page 2. Minneapolis, Minnesota, 2019. 3, 15
2019
-
[40]
Computa- tional tradeoffs in image synthesis: Diffusion, masked-token, and next-token prediction
Maciej Kilian, Varun Japan, and Luke Zettlemoyer. Computa- tional tradeoffs in image synthesis: Diffusion, masked-token, and next-token prediction. arXiv preprint arXiv:2405.13218,
-
[41]
Understanding diffusion ob- jectives as the elbo with simple data augmentation
Diederik Kingma and Ruiqi Gao. Understanding diffusion ob- jectives as the elbo with simple data augmentation. NeurIPS, 36, 2024. 3
2024
-
[42]
Variational diffusion models
Diederik Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. Variational diffusion models. In NeurIPS, 2021. 3
2021
-
[43]
The open im- ages dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al. The open im- ages dataset v4: Unified image classification, object detection, and visual relationship detection...
1956
-
[44]
Autoregressive image generation using resid- ual quantization
Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han. Autoregressive image generation using resid- ual quantization. In CVPR, pages 11523–11532, 2022. 6
2022
-
[45]
Your diffusion model is secretly a zero-shot classifier
Alexander C Li, Mihir Prabhudesai, Shivam Duggal, Ellis Brown, and Deepak Pathak. Your diffusion model is secretly a zero-shot classifier. In ICCV, pages 2206–2217, 2023. 1
2023
-
[46]
Mage: Masked generative en- coder to unify representation learning and image synthesis
Tianhong Li, Huiwen Chang, Shlok Mishra, Han Zhang, Dina Katabi, and Dilip Krishnan. Mage: Masked generative en- coder to unify representation learning and image synthesis. In CVPR, pages 2142–2152, 2023. 1, 2, 5
2023
-
[47]
Autoregressive image generation without vector quantization
Tianhong Li, Yonglong Tian, He Li, Mingyang Deng, and Kaiming He. Autoregressive image generation without vector quantization. arXiv preprint arXiv:2406.11838, 2024. 15
2024 arXiv
-
[48]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 16
2014
-
[49]
Think while you generate: Discrete diffusion with planned denois- ing
Sulin Liu, Juno Nam, Andrew Campbell, Hannes St¨ark, Yilun Xu, Tommi Jaakkola, and Rafael G´omez-Bombarelli. Think while you generate: Discrete diffusion with planned denois- ing. arXiv preprint arXiv:2410.06264, 2024. 2, 9, 14
2024 arXiv
-
[50]
Pyramid diffusion for fine 3d large scene generation
Yuheng Liu, Xinke Li, Xueting Li, Lu Qi, Chongshou Li, and Ming-Hsuan Yang. Pyramid diffusion for fine 3d large scene generation. arXiv preprint arXiv:2311.12085, 2023. 2
2023 arXiv
-
[51]
Unified-io: A unified model for vision, language, and multi-modal tasks
Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mot- taghi, and Aniruddha Kembhavi. Unified-io: A unified model for vision, language, and multi-modal tasks. In ICLR, 2022. 15
2022
-
[52]
Open-magvit2: An open-source project toward democratizing auto-regressive visual generation.arXiv preprint arXiv:2409.04410, 2024
Zhuoyan Luo, Fengyuan Shi, Yixiao Ge, Yujiu Yang, Limin Wang, and Ying Shan. Open-magvit2: An open-source project toward democratizing auto-regressive visual generation.arXiv preprint arXiv:2409.04410, 2024. 6
2024 arXiv
-
[53]
Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers
Nanye Ma, Mark Goldstein, Michael S Albergo, Nicholas M Boffi, Eric Vanden-Eijnden, and Saining Xie. Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers. ECCV, 2024. 2
2024
-
[54]
Latte: Latent diffusion transformer for video generation
Xin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen, Ziwei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. Latte: Latent diffusion transformer for video generation. arXiv preprint arXiv:2401.03048, 2024. 7, 8
2024 arXiv
-
[55]
Discrete representations strengthen vision transformer robustness
Chengzhi Mao, Lu Jiang, Mostafa Dehghani, Carl V ondrick, Rahul Sukthankar, and Irfan Essa. Discrete representations strengthen vision transformer robustness. arXiv preprint arXiv:2111.10493, 2021. 15
2021 arXiv
-
[56]
Scal- ing up masked diffusion models on text
Shen Nie, Fengqi Zhu, Chao Du, Tianyu Pang, Qian Liu, Guangtao Zeng, Min Lin, and Chongxuan Li. Scal- ing up masked diffusion models on text. arXiv preprint arXiv:2410.18514, 2024. 2
2024 arXiv
-
[57]
Your absorbing discrete diffusion secretly models the conditional distributions of clean data, 2024
Jingyang Ou, Shen Nie, Kaiwen Xue, Fengqi Zhu, Jiacheng Sun, Zhenguo Li, and Chongxuan Li. Your absorbing discrete diffusion secretly models the conditional distributions of clean data, 2024. 2, 4, 15
2024
-
[58]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. ICCV, 2023. 6
2023
-
[59]
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M¨uller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis. ICLR, 2024. 15
2024
-
[60]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 1, 5, 6, 15, 16
2022
-
[61]
Simple and effective masked diffu- sion language models
Subham Sekhar Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin T Chiu, Alexander Rush, and V olodymyr Kuleshov. Simple and effective masked diffu- sion language models. NeurIPS, 2024. 2, 4
2024
-
[62]
Baumann, Vincent Tao Hu, and Bj ¨orn Ommer
Johannes Schusterbauer, Ming Gui, Pingchuan Ma, Nick Stracke, Stefan A. Baumann, Vincent Tao Hu, and Bj ¨orn Ommer. Boosting latent diffusion with flow matching. In ECCV, 2024. 1
2024
-
[63]
Simplified and generalized masked dif- fusion for discrete data
Jiaxin Shi, Kehang Han, Zhe Wang, Arnaud Doucet, and Michalis K Titsias. Simplified and generalized masked dif- fusion for discrete data. arXiv preprint arXiv:2406.04329,
-
[64]
Training and inference on any-order autoregressive models the right way
Andy Shih, Dorsa Sadigh, and Stefano Ermon. Training and inference on any-order autoregressive models the right way. NeurIPS, 35:2762–2775, 2022. 2
2022
-
[65]
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In ICML, 2015. 1
2015
-
[66]
Generative modeling by estimating gradients of the data distribution
Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. NeurIPS, 32,
-
[67]
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In ICLR, 2021. 1, 3
2021
-
[68]
Cleandift: Diffusion features without noise
Nick Stracke, Stefan Andreas Baumann, Kolja Bauer, Frank Fundel, and Bj ¨orn Ommer. Cleandift: Diffusion features without noise. arXiv preprint arXiv:2412.03439, 2024. 16
2024
-
[69]
Autoregressive model beats diffusion: Llama for scalable image generation
Peize Sun, Yi Jiang, Shoufa Chen, Shilong Zhang, Bingyue Peng, Ping Luo, and Zehuan Yuan. Autoregressive model beats diffusion: Llama for scalable image generation. arXiv preprint arXiv:2406.06525, 2024. 6, 15
2024 arXiv
-
[70]
Improved vector quantized diffusion models
Zhicong Tang, Shuyang Gu, Jianmin Bao, Dong Chen, and Fang Wen. Improved vector quantized diffusion models. arXiv preprint arXiv:2205.16007, 2022. 2, 4, 14, 15
2022 arXiv
-
[71]
Chameleon: Mixed-modal early-fusion foundation models
Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818, 2024. 16
2024 arXiv
-
[72]
Simulation-free schr\” odinger bridges via score and flow matching
Alexander Tong, Nikolay Malkin, Kilian Fatras, Lazar Atanackovic, Yanlei Zhang, Guillaume Huguet, Guy Wolf, and Yoshua Bengio. Simulation-free schr\” odinger bridges via score and flow matching. Proceedings of The 27th Inter- national Conference on Artificial Intelligence and ...
-
[73]
Givt: Generative infinite-vocabulary transformers
Michael Tschannen, Cian Eastwood, and Fabian Mentzer. Givt: Generative infinite-vocabulary transformers. ECCV,
-
[74]
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. NeurIPS, 30, 2017. 15
2017
-
[75]
Phenaki: Variable length video generation from open domain textual descriptions
Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kinder- mans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. Phenaki: Variable length video generation from open domain textual descriptions. In ICLR, 2022. 2
2022
-
[76]
Scaling image tokenizers with grouped spherical quantization
Jiangtao Wang, Zhen Qin, Yifan Zhang, Vincent Tao Hu, Bj¨orn Ommer, Rania Briq, and Stefan Kesselheim. Scaling image tokenizers with grouped spherical quantization. arXiv preprint arXiv:2412.02632, 2024. 15
2024 arXiv
-
[77]
Segrefiner: Towards model- agnostic segmentation refinement with discrete diffusion pro- cess
Mengyu Wang, Henghui Ding, Jun Hao Liew, Jiajun Liu, Yao Zhao, and Yunchao Wei. Segrefiner: Towards model- agnostic segmentation refinement with discrete diffusion pro- cess. arXiv preprint arXiv:2312.12425, 2023. 2
2023 arXiv
-
[78]
Maskbit: Embedding-free image generation via bit tokens
Mark Weber, Lijun Yu, Qihang Yu, Xueqing Deng, Xi- aohui Shen, Daniel Cremers, and Liang-Chieh Chen. Maskbit: Embedding-free image generation via bit tokens. arXiv:2409.16211, 2024. 1, 2
2024 arXiv
-
[79]
Show-o: One single transformer to unify multimodal understanding and genera- tion
Jinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang, Weihao Wang, Kevin Qinghong Lin, Yuchao Gu, Zhijie Chen, Zhenheng Yang, and Mike Zheng Shou. Show-o: One single transformer to unify multimodal understanding and genera- tion. arXiv preprint arXiv:2408.12528, 2024. 1, 15
2024 arXiv
-
[80]
Autoregressive models in vision: A survey
Jing Xiong, Gongye Liu, Lun Huang, Chengyue Wu, Taiqiang Wu, Yao Mu, Yuan Yao, Hui Shen, Zhongwei Wan, Jinfa Huang, et al. Autoregressive models in vision: A survey. arXiv preprint arXiv:2411.05902, 2024. 16
2024 arXiv
-
[81]
Vector-quantized image modeling with improved vqgan
Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. Vector-quantized image modeling with improved vqgan. ICLR, 2022. 6
2022
-
[82]
Scaling autoregressive mod- els for content-rich text-to-image generation
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gun- jan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. Scaling autoregressive mod- els for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789, 2(3):5, 2022. 6
2022 arXiv
-
[83]
Magvit: Masked generative video transformer
Lijun Yu, Yong Cheng, Kihyuk Sohn, Jos ´e Lezama, Han Zhang, Huiwen Chang, Alexander G Hauptmann, Ming- Hsuan Yang, Yuan Hao, Irfan Essa, et al. Magvit: Masked generative video transformer. In CVPR, pages 10459–10469,
-
[84]
Language model beats diffusion–tokenizer is key to visual generation
Lijun Yu, Jos´e Lezama, Nitesh B Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Agrim Gupta, Xiuye Gu, Alexander G Hauptmann, et al. Language model beats diffusion–tokenizer is key to visual generation. ICLR,
-
[85]
An image is worth 32 tokens for reconstruction and generation
Qihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen, Daniel Cremers, and Liang-Chieh Chen. An image is worth 32 tokens for reconstruction and generation. NeurIPS, 2024. 1
2024
-
[86]
Representa- tion alignment for generation: Training diffusion transformers is easier than you think
Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie. Representa- tion alignment for generation: Training diffusion transformers is easier than you think. arXiv preprint arXiv:2410.06940,
-
[87]
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In ICLR, 2017. 14
2017
-
[88]
Masked diffusion models are secretly time-agnostic masked models and exploit inaccurate categorical sampling
Kaiwen Zheng, Yongxin Chen, Hanzi Mao, Ming-Yu Liu, Jun Zhu, and Qinsheng Zhang. Masked diffusion models are secretly time-agnostic masked models and exploit inaccurate categorical sampling. arXiv preprint arXiv:2409.02908, 2024. 2, 4
2024 arXiv
-
[89]
Transfusion: Predict the next token and diffuse images with one multi-modal model
Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michi- hiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, and Omer Levy. Transfusion: Predict the next token and diffuse images with one multi-modal model. arXiv preprint arXiv:2408.11039, 2024. 1, 15, 16
2024 arXiv
-
[90]
Lafite: Towards language-free training for text-to-image generation
Yufan Zhou, Ruiyi Zhang, Changyou Chen, Chunyuan Li, Chris Tensmeyer, Tong Yu, Jiuxiang Gu, Jinhui Xu, and Tong Sun. Lafite: Towards language-free training for text-to-image generation. In CVPR, 2022. 6 A. Appendix Contents
2022
-
[91]
Discrete Interpolants
Method 2 3.1. Discrete Interpolants . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 3.2. Training . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 3.3. Sampling . ....
-
[92]
Experimental Detail
Experiment 5 4.1. Experimental Detail . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 4.2. Experimental Result . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 4.2.1 Main Result ...
-
[93]
Conclusion & Future Works 8
-
[94]
Appendix 13 B
Acknowledgment 9 A . Appendix 13 B . Potential Impact 14 C . Extra Discussions 14 C.1. Scheduler . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 C.2. Connections with Other Methods . . . . . . . . . . . . . . . . ....
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.