Pith. sign in

REVIEW 2 cited by

MAGE: MAsked Generative Encoder to Unify Representation Learning and Image Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.09117 v2 pith:R6JZZFQV submitted 2022-11-16 cs.CV

classification cs.CV
keywords representationlearningmagegenerativeimagegenerationmaskingencoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative modeling and representation learning are two key tasks in computer vision. However, these models are typically trained independently, which ignores the potential for each task to help the other, and leads to training and model maintenance overheads. In this work, we propose MAsked Generative Encoder (MAGE), the first framework to unify SOTA image generation and self-supervised representation learning. Our key insight is that using variable masking ratios in masked image modeling pre-training can allow generative training (very high masking ratio) and representation learning (lower masking ratio) under the same training framework. Inspired by previous generative models, MAGE uses semantic tokens learned by a vector-quantized GAN at inputs and outputs, combining this with masking. We can further improve the representation by adding a contrastive loss to the encoder output. We extensively evaluate the generation and representation learning capabilities of MAGE. On ImageNet-1K, a single MAGE ViT-L model obtains 9.10 FID in the task of class-unconditional image generation and 78.9% top-1 accuracy for linear probing, achieving state-of-the-art performance in both image generation and representation learning. Code is available at https://github.com/LTH14/mage.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    1D binary image latents reduce a 1024x1024 image to 128 discrete tokens and support text-to-image generation with diffusion and autoregressive models.

  2. MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval

    cs.CV 2025-10 conditional novelty 5.0 of 10

    MSAM introduces two drone-video/text datasets and a CLIP-based multi-semantic pooling model that reports 0.6–3.8 point R@1 gains over earlier video-text retrieval methods.

Pith tools