Pith. sign in

REVIEW 8 cited by

Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.02591 v1 pith:57AIJRR2 submitted 2023-09-05 cs.LG cs.CLcs.CV

classification cs.LGcs.CLcs.CV
keywords multi-modalcm3leongenerationmodelmodelsdemonstratelanguagemethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present CM3Leon (pronounced "Chameleon"), a retrieval-augmented, token-based, decoder-only multi-modal language model capable of generating and infilling both text and images. CM3Leon uses the CM3 multi-modal architecture but additionally shows the extreme benefits of scaling up and tuning on more diverse instruction-style data. It is the first multi-modal model trained with a recipe adapted from text-only language models, including a large-scale retrieval-augmented pre-training stage and a second multi-task supervised fine-tuning (SFT) stage. It is also a general-purpose model that can do both text-to-image and image-to-text generation, allowing us to introduce self-contained contrastive decoding methods that produce high-quality outputs. Extensive experiments demonstrate that this recipe is highly effective for multi-modal models. CM3Leon achieves state-of-the-art performance in text-to-image generation with 5x less training compute than comparable methods (zero-shot MS-COCO FID of 4.88). After SFT, CM3Leon can also demonstrate unprecedented levels of controllability in tasks ranging from language-guided image editing to image-controlled generation and segmentation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 27 citations worldwide. Full citation record

  1. AR-RAG: Autoregressive Retrieval Augmentation for Image Generation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Autoregressive patch-level retrieval augmentation improves text-to-image generation on GenEval, DPG-Bench, and Midjourney-30K, with a training-free decoding variant and a fine-tuned variant.

  2. HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    VLAC-Cut-guided multi-robot HITL post-training reaches 80–95% success and 1.7–4.2× throughput over the base VLA, outperforming HITL-only under the same human budget.

  3. Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought

    cs.CV 2026-01 conditional novelty 6.0 of 10

    A visual chain-of-thought model that interleaves text plans and rendered 2D states improves compositional object assembly, hitting 88.4% component numeracy and 84.8% structural topology on a new benchmark.

  4. Transition Matching: Scalable and Flexible Generative Modeling

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Transition Matching unifies flow matching and continuous autoregressive generation as discrete-time Markov processes, with three variants that improve text-to-image quality and speed.

  5. CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CuRe scores text-to-image systems by how much their output changes as prompts add cultural details, and reports better agreement with human ratings than existing proxies.

  6. Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications

    cs.CV 2026-07 conditional novelty 4.0 of 10

    A survey of instruction-based image editing plus a new 21-task benchmark, CDD-IIE, on which ten open models are scored by human experts.

  7. Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A unified image understanding and generation model with decoupled visual encoders achieves competitive benchmark scores on both tasks.

  8. Next Block Prediction: Video Generation via Semi-Autoregressive Modeling

    cs.CV 2025-02 conditional novelty 3.0 of 10

    NBP generates video rows in parallel with block-wise attention, achieving 11x faster inference and better FVD than next-token autoregressive models.

Pith tools