Pith. sign in

REVIEW 11 cited by

Generating Images with Sparse Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.03841 v1 pith:OFYD3BQB submitted 2021-03-05 cs.CV stat.ML

classification cs.CVstat.ML
keywords imageshighimagemodelsapproacharchitecturelikelihood-basedmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The high dimensionality of images presents architecture and sampling-efficiency challenges for likelihood-based generative models. Previous approaches such as VQ-VAE use deep autoencoders to obtain compact representations, which are more practical as inputs for likelihood-based models. We present an alternative approach, inspired by common image compression methods like JPEG, and convert images to quantized discrete cosine transform (DCT) blocks, which are represented sparsely as a sequence of DCT channel, spatial location, and DCT coefficient triples. We propose a Transformer-based autoregressive architecture, which is trained to sequentially predict the conditional distribution of the next element in such sequences, and which scales effectively to high resolution images. On a range of image datasets, we demonstrate that our approach can generate high quality, diverse images, with sample metric scores competitive with state of the art methods. We additionally show that simple modifications to our method yield effective image colorization and super-resolution models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SPARE: Structural Parameter-Free Affinity Regularization for Flow Matching

    cs.CV 2026-08 conditional novelty 7.0 of 10

    A parameter-free regularizer that aligns intermediate token affinities to clean VAE latent affinities, including cross-image pairs, lowers FID on ImageNet with SiT backbones at matched training budgets.

  2. MentalThink: Shaping Thoughts in Mental SVG World

    cs.AI 2026-07 conditional novelty 7.0 of 10

    MLLMs that generate and render SVG sketches as multi-turn intermediate reasoning steps reach 55.1% on VSIBench and 76.0% on MindCube, far above the Qwen2.5-VL-7B backbone.

  3. Post-Training Pruning for Diffusion Transformers

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    DiT-Pruning keeps CLIP and FID nearly unchanged at 50% sparsity on FLUX and PixArt by an energy-motivated squared-weight metric plus clustering-aware granularity, beating Wanda and magnitude baselines.

  4. Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training?

    cs.CV 2026-01 conditional novelty 6.0 of 10

    A self-supervised two-stage training method—VAE-latent feature alignment then feature-level classifier-free guidance—lets DiT models match or beat DINO-guided REPA training without any external feature extractor.

  5. REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Nonlinear multi-layer compression of frozen VFM patch semantics, jointly denoised with VAE latents, improves ImageNet 256x256 FID (12.9 vs 15.2 for REG at SiT-B/2, 400K) and accelerates convergence over REPA/ReDi/REG.

  6. Transition Models: Rethinking the Generative Learning Objective

    cs.LG 2025-09 conditional novelty 6.0 of 10

    TiM trains a single diffusion-type model on arbitrary time-interval transitions, achieving strong one-step and multi-step text-to-image generation with 865M parameters.

  7. PixNerd: Pixel Neural Field Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PixNerd is a single-stage pixel-space diffusion transformer that uses predicted neural field weights to decode large patches, reaching 2.15 FID on ImageNet 256 without a VAE.

  8. DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A post-training quantization method that combines learned channel scaling and power-of-two scaling keeps diffusion image quality high at 4-bit weight, 6-bit activation precision.

  9. MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MPQ-DMv2 adds binary residual quantization, temporal relation distillation, and SVD-initialized LoRA to mixed-precision quantization, improving low-bit diffusion model generation quality.

  10. Native-Resolution Image Synthesis

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A single diffusion transformer trained on native-resolution ImageNet achieves state-of-the-art FID at 256 and 512, and extrapolates to 1024 and 1536 with moderate degradation.

  11. Contrastive Flow Matching

    cs.CV 2025-06 reject novelty 2.0 of 10

    Contrastive Flow Matching adds a negative flow-target term to the standard flow-matching loss, reporting large empirical gains, but the closed-form solution shows the term only applies a global rescaling and shift, no...

Pith tools