Pith. sign in

REVIEW 11 cited by

Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.17172 v1 pith:R3SLMFZP submitted 2023-12-28 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords modelmultimodalaudiounderstandingactionunified-ioautoregressivediverse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Unified-IO 2, the first autoregressive multimodal model that is capable of understanding and generating image, text, audio, and action. To unify different modalities, we tokenize inputs and outputs -- images, text, audio, action, bounding boxes, etc., into a shared semantic space and then process them with a single encoder-decoder transformer model. Since training with such diverse modalities is challenging, we propose various architectural improvements to stabilize model training. We train our model from scratch on a large multimodal pre-training corpus from diverse sources with a multimodal mixture of denoisers objective. To learn an expansive set of skills, such as following multimodal instructions, we construct and finetune on an ensemble of 120 datasets with prompts and augmentations. With a single unified model, Unified-IO 2 achieves state-of-the-art performance on the GRIT benchmark and strong results in more than 35 benchmarks, including image generation and understanding, natural language understanding, video and audio understanding, and robotic manipulation. We release all our models to the research community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MentalThink: Shaping Thoughts in Mental SVG World

    cs.AI 2026-07 conditional novelty 7.0 of 10

    MLLMs that generate and render SVG sketches as multi-turn intermediate reasoning steps reach 55.1% on VSIBench and 76.0% on MindCube, far above the Qwen2.5-VL-7B backbone.

  2. Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Multi-TW is the first Traditional Chinese benchmark to evaluate multimodal models on both image-text and audio-text questions while also measuring inference latency.

  3. UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    UniGen shows a 1.5B model trained on open data can beat larger systems on image understanding and generation once it verifies its own outputs with chain-of-thought and Best-of-N selection.

  4. Next Patch Prediction for Autoregressive Visual Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Averaging neighboring image tokens into patches during training lets autoregressive image models train faster and generate higher-quality images, with inference unchanged.

  5. MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A new 1.47M-instance biomedical instruction-tuning dataset improves a mixed-modal 7B model's medical VQA accuracy by 18 to 26 percentage points over GPT-4o and Chameleon.

  6. Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Diffusion-conditioned denoising trains a video tokenizer whose features support both video question answering and text-to-video generation in a single LLM.

  7. One Diffusion to Generate Them All

    cs.CV 2024-11 conditional novelty 6.0 of 10

    OneDiffusion shows that a single 2.8B-parameter diffusion model, trained by treating all tasks as frame sequences with varying noise scales, can handle image generation and image understanding tasks bidirectionally.

  8. BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching

    cs.LG 2024-11 conditional novelty 6.0 of 10

    BlendServe combines resource-aware batching with prefix sharing using a resource-aware prefix tree and dual scanner, achieving up to 1.44x throughput vs vLLM/SGLang in offline LLM inference.

  9. GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A weakly supervised latent-variable agent improves multimodal instruction following by combining VAE self-imitating on unlabeled data with a likelihood-based alignment of labeled and video latents.

  10. Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

    cs.CV 2025-01 conditional novelty 4.0 of 10

    Valley2, a 7B-scale open-source multimodal model, reports second-best OpenCompass average (67.4) among sub-10B models and the highest score (79.66) on its own in-house Ecom-VQA benchmark.

  11. Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A comprehensive survey that frames multimodal understanding and generation as next token prediction and proposes a five-part taxonomy.

Pith tools