Pith. sign in

REVIEW 3 cited by

StraIT: Non-autoregressive Generation with Stratified Image Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.00750 v1 pith:OX3CISWE submitted 2023-03-01 cs.CV

classification cs.CV
keywords imagestraitexistinggenerationstratifiedgenerativeguidancemodeling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose Stratified Image Transformer(StraIT), a pure non-autoregressive(NAR) generative model that demonstrates superiority in high-quality image synthesis over existing autoregressive(AR) and diffusion models(DMs). In contrast to the under-exploitation of visual characteristics in existing vision tokenizer, we leverage the hierarchical nature of images to encode visual tokens into stratified levels with emergent properties. Through the proposed image stratification that obtains an interlinked token pair, we alleviate the modeling difficulty and lift the generative power of NAR models. Our experiments demonstrate that StraIT significantly improves NAR generation and out-performs existing DMs and AR methods while being order-of-magnitude faster, achieving FID scores of 3.96 at 256*256 resolution on ImageNet without leveraging any guidance in sampling or auxiliary image classifiers. When equipped with classifier-free guidance, our method achieves an FID of 3.36 and IS of 259.3. In addition, we illustrate the decoupled modeling process of StraIT generation, showing its compelling properties on applications including domain transfer.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. STORM: Benchmarking Visual Rating of MLLMs with a Comprehensive Ordinal Regression Dataset

    cs.CV 2025-06 conditional novelty 6.0 of 10

    STORM is a new multi-domain ordinal-regression benchmark with coarse-to-fine Chain-of-Thought prompts that improves MLLM zero-shot visual rating, though the 'universal' claim is bounded by its five curated domains.

  2. Plug-and-Play Context Feature Reuse for Efficient Masked Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ReCAP interleaves full model evaluations with lightweight steps that reuse cached context features, delivering up to 2.4x faster masked generation with minimal FID loss.

  3. Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    MaskGIL, a bidirectional-LLaMA masked autoregressive model, generates ImageNet 256x256 images with FID 3.71 in 8 steps, and also supports text-driven and speech-driven generation.

Pith tools