Pith. sign in

REVIEW 2 cited by

4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.09406 v2 pith:Q557VXYI submitted 2024-06-13 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords modalitiesmodelmodelstrainingdiverselikemultimodaltasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform diverse tasks are limited by the (usually rather small) number of modalities and tasks they are trained on. In this paper, we expand upon the capabilities of them by training a single model on tens of highly diverse modalities and by performing co-training on large-scale multimodal datasets and text corpora. This includes training on several semantic and geometric modalities, feature maps from recent state of the art models like DINOv2 and ImageBind, pseudo labels of specialist models like SAM and 4DHumans, and a range of new modalities that allow for novel ways to interact with the model and steer the generation, for example image metadata or color palettes. A crucial step in this process is performing discrete tokenization on various modalities, whether they are image-like, neural network feature maps, vectors, structured data like instance segmentation or human poses, or data that can be represented as text. Through this, we expand on the out-of-the-box capabilities of multimodal models and specifically show the possibility of training one model to solve at least 3x more tasks/modalities than existing ones and doing so without a loss in performance. This enables more fine-grained and controllable multimodal generation capabilities and allows us to study the distillation of models trained on diverse data and objectives into a unified model. We successfully scale the training to a three billion parameter model using tens of modalities and different datasets. The resulting models and training code are open sourced at 4m.epfl.ch.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Stable Diffusion features, especially when conditioned on the question, improve vision-centric multimodal question answering when fused with CLIP.

  2. BiFold: Bimanual Cloth Folding with Language Guidance

    cs.RO 2025-01 conditional novelty 6.0 of 10

    A language-conditioned model predicts where two robot arms should pick and place cloth to fold it, trained on a newly auto-annotated bimanual dataset.

Pith tools