Pith. sign in

REVIEW 2 cited by

Jointly Training Large Autoregressive Multimodal Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.15564 v2 pith:JCOZG6DA submitted 2023-09-27 cs.LG cs.CLcs.CV

classification cs.LGcs.CLcs.CV
keywords modelmodelsmultimodalautoregressivegeneratinggenerationoutputsaddress
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In recent years, advances in the large-scale pretraining of language and text-to-image models have revolutionized the field of machine learning. Yet, integrating these two modalities into a single, robust model capable of generating seamless multimodal outputs remains a significant challenge. To address this gap, we present the Joint Autoregressive Mixture (JAM) framework, a modular approach that systematically fuses existing text and image generation models. We also introduce a specialized, data-efficient instruction-tuning strategy, tailored for mixed-modal generation tasks. Our final instruct-tuned model demonstrates unparalleled performance in generating high-quality multimodal outputs and represents the first model explicitly designed for this purpose.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Harmonizing and Merging Source Models for CLIP-based Domain Generalization

    cs.CV 2025-06 conditional novelty 6.0 of 10

    HAM trains per-domain CLIP encoders, enriches them with confident cross-domain samples, aligns their update directions, and merges them with redundancy trimming, reaching 79.0% average accuracy on five DG benchmarks w...

  2. LMFusion: Adapting Pretrained Language Models for Multimodal Generation

    cs.CL 2024-12 conditional novelty 6.0 of 10

    LMFusion freezes Llama-3's language modules and trains parallel image modules, yielding better image captioning and comparable image generation than Transfusion at half the training FLOPs.

Pith tools