Pith. sign in

REVIEW 12 cited by

Fast Training of Diffusion Models with Masked Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.09305 v2 pith:A7BULS5I submitted 2023-06-15 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords trainingmaskedpatchesdiffusionmodelsgenerativetransformertransformers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose an efficient approach to train large diffusion models with masked transformers. While masked transformers have been extensively explored for representation learning, their application to generative learning is less explored in the vision domain. Our work is the first to exploit masked training to reduce the training cost of diffusion models significantly. Specifically, we randomly mask out a high proportion (e.g., 50%) of patches in diffused input images during training. For masked training, we introduce an asymmetric encoder-decoder architecture consisting of a transformer encoder that operates only on unmasked patches and a lightweight transformer decoder on full patches. To promote a long-range understanding of full patches, we add an auxiliary task of reconstructing masked patches to the denoising score matching objective that learns the score of unmasked patches. Experiments on ImageNet-256x256 and ImageNet-512x512 show that our approach achieves competitive and even better generative performance than the state-of-the-art Diffusion Transformer (DiT) model, using only around 30% of its original training time. Thus, our method shows a promising way of efficiently training large transformer-based diffusion models without sacrificing the generative performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    SRC-Flow compresses RAE features via a Semantic Representation Compressor into a low-dimensional space, enabling normalizing flows to reach gFID 1.65 on ImageNet 256x256 and 2.07 on 512x512 while retaining exact likelihoods.

  2. ELT: Elastic Looped Transformers for Visual Generation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Weight-shared looped transformers trained with intra-loop self-distillation match MaskGIT-class FID/FVD at roughly 4x fewer parameters and support any-time inference across loop counts.

  3. Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation

    cs.CV 2026-02 unverdicted novelty 6.0 of 10

    Re-injecting shallow text features into deeper MMDiT blocks counteracts measured 'prompt forgetting' and improves instruction following in SD3, SD3.5, FLUX, and Qwen-Image without retraining.

  4. MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Training diffusion models on a mixture of higher-noise interpolations (MixFlow) improves generation FID across SiT, REPA, RAE and SD3.5, reaching ImageNet 256 gFID 1.43 after post-training.

  5. REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Nonlinear multi-layer compression of frozen VFM patch semantics, jointly denoised with VAE latents, improves ImageNet 256x256 FID (12.9 vs 15.2 for REG at SiT-B/2, 400K) and accelerates convergence over REPA/ReDi/REG.

  6. Missing Fine Details in Images: Last Seen in High Frequencies

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A wavelet-based VAE that trains low- and high-frequency branches separately improves image reconstruction and diffusion generation.

  7. MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A masking-augmented diffusion objective plus pause-token inference scaling modestly improves instruction adherence and source preservation for OmniGen-based image editing.

  8. Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DisCon treats discrete image tokens as conditioning signals rather than targets, letting a continuous autoregressive model refine details and reach gFID 1.38 on ImageNet-256.

  9. EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Enforcing scale and rotation equivariance in pretrained image autoencoders via a reconstruction loss on transformed latents speeds up and improves latent generative models.

  10. Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders

    cs.LG 2026-02 conditional novelty 5.0 of 10

    Training flow matching along sphere geodesics with a curvature-aware loss weight lets standard DiT-B converge on DINOv2 features (FID 3.37 with guidance), contradicting the need for width scaling.

  11. Improving Joint Embedding Predictive Architecture with Diffusion Noise

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Injecting EDM-style noise into masked-token position embeddings and adding two auxiliary losses improves I-JEPA's linear-probing accuracy by about 1.5 points on ImageNet-1K.

  12. DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers

    cs.CV 2026-08

Pith tools