Pith. sign in

REVIEW 4 cited by

Masked Autoencoders Are Effective Tokenizers for Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.03444 v2 pith:7RZYR5LD submitted 2025-02-05 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords latentmodelsdiffusiongenerationspaceautoencodersbetterdiscriminative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in latent diffusion models have demonstrated their effectiveness for high-resolution image synthesis. However, the properties of the latent space from tokenizer for better learning and generation of diffusion models remain under-explored. Theoretically and empirically, we find that improved generation quality is closely tied to the latent distributions with better structure, such as the ones with fewer Gaussian Mixture modes and more discriminative features. Motivated by these insights, we propose MAETok, an autoencoder (AE) leveraging mask modeling to learn semantically rich latent space while maintaining reconstruction fidelity. Extensive experiments validate our analysis, demonstrating that the variational form of autoencoders is not necessary, and a discriminative latent space from AE alone enables state-of-the-art performance on ImageNet generation using only 128 tokens. MAETok achieves significant practical improvements, enabling a gFID of 1.69 with 76x faster training and 31x higher inference throughput for 512x512 generation. Our findings show that the structure of the latent space, rather than variational constraints, is crucial for effective diffusion models. Code and trained models are released.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Physics-Informed Distillation of Diffusion Models for PDE-Constrained Generation

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Post-hoc distillation with a PDE-residual loss on final samples avoids the Jensen gap and yields one-step physics-constrained generation.

  2. BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    A volumetric MAE tokenizer decouples clinical embedding from reconstruction to support both 23-task linear probing and conditional 3D brain MRI generation via DiT.

  3. DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer

    cs.CV 2025-07 conditional novelty 5.0 of 10

    DC-AR generates 512x512 images in 12 masked autoregressive steps plus 20 diffusion refinement steps, using a 32x compressed 2D tokenizer, and reports gFID 5.49 on MJHQ-30K.

  4. Adaptive Mask-guided K-space Diffusion for Accelerated MRI Reconstruction

    eess.IV 2025-06 reject novelty 4.0 of 10

    AMDM reconstructs undersampled MRI by masking k-space frequency components with adaptive masks inside a diffusion model, and reports large PSNR gains over baseline methods.

Pith tools