REVIEW 2 cited by
Jointly Training Large Autoregressive Multimodal Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In recent years, advances in the large-scale pretraining of language and text-to-image models have revolutionized the field of machine learning. Yet, integrating these two modalities into a single, robust model capable of generating seamless multimodal outputs remains a significant challenge. To address this gap, we present the Joint Autoregressive Mixture (JAM) framework, a modular approach that systematically fuses existing text and image generation models. We also introduce a specialized, data-efficient instruction-tuning strategy, tailored for mixed-modal generation tasks. Our final instruct-tuned model demonstrates unparalleled performance in generating high-quality multimodal outputs and represents the first model explicitly designed for this purpose.
Forward citations
Cited by 2 Pith papers
-
Harmonizing and Merging Source Models for CLIP-based Domain Generalization
HAM trains per-domain CLIP encoders, enriches them with confident cross-domain samples, aligns their update directions, and merges them with redundancy trimming, reaching 79.0% average accuracy on five DG benchmarks w...
-
LMFusion: Adapting Pretrained Language Models for Multimodal Generation
LMFusion freezes Llama-3's language modules and trains parallel image modules, yielding better image captioning and comparable image generation than Transfusion at half the training FLOPs.
Discussion (0). Continue with ORCID to comment.