Pith. sign in

REVIEW 4 cited by

Multimodal Masked Autoencoders Learn Transferable Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.14204 v3 pith:NOMJVEX5 submitted 2022-05-27 cs.CV

Multimodal Masked Autoencoders Learn Transferable Representations

classification cs.CV
keywords datam3aelearnmaskedmultimodalcontrastivedownstreamimage-text
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Building scalable models to learn from diverse, multimodal data remains an open challenge. For vision-language data, the dominant approaches are based on contrastive learning objectives that train a separate encoder for each modality. While effective, contrastive learning approaches introduce sampling bias depending on the data augmentations used, which can degrade performance on downstream tasks. Moreover, these methods are limited to paired image-text data, and cannot leverage widely-available unpaired data. In this paper, we investigate whether a large multimodal model trained purely via masked token prediction, without using modality-specific encoders or contrastive learning, can learn transferable representations for downstream tasks. We propose a simple and scalable network architecture, the Multimodal Masked Autoencoder (M3AE), which learns a unified encoder for both vision and language data via masked token prediction. We provide an empirical study of M3AE trained on a large-scale image-text dataset, and find that M3AE is able to learn generalizable representations that transfer well to downstream tasks. Surprisingly, we find that M3AE benefits from a higher text mask ratio (50-90%), in contrast to BERT whose standard masking ratio is 15%, due to the joint training of two data modalities. We also provide qualitative analysis showing that the learned representation incorporates meaningful information from both image and language. Lastly, we demonstrate the scalability of M3AE with larger model size and training time, and its flexibility to train on both paired image-text data as well as unpaired data.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Biosignal Fingerprinting: A Cross-Modal PPG-ECG Foundation Model

    cs.LG 2026-05 unverdicted novelty 6.0

    A cross-modal masked autoencoder creates reusable biosignal fingerprints that match or exceed specialist models on seven cardiovascular tasks using only single-modality input.

  2. Self-Supervised Learning with a Multi-Task Latent Space Objective

    cs.CV 2026-02 conditional novelty 6.0

    Assigning a dedicated predictor to each view type stabilizes multi-crop Siamese SSL and, combined with asymmetric cutout views, yields consistent ImageNet gains over BYOL, SimSiam, and MoCo v3.

  3. Evidential Fusion Network for Multimodal Survival Prediction under Missing Modalities

    cs.LG 2026-06 unverdicted novelty 5.0

    EMMS uses evidential fusion based on Dempster-Shafer theory to handle missing modalities in multimodal survival prediction without generative imputation, reporting SOTA results and calibrated uncertainty on four cance...

  4. Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis

    cs.LG 2026-06 unverdicted novelty 4.0

    Foundation model representations from images and transcriptomics carry complementary signals for cancer classification; multimodal fusion improves results mainly when no modality dominates, and conformal prediction re...