Pith. sign in

REVIEW 2 cited by

M6: A Chinese Multimodal Pretrainer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.00823 v4 pith:WAP7ZMD6 submitted 2021-03-01 cs.CL

classification cs.CL
keywords chinesemodelpretrainingbilliondownstreamimageslargestmulti-modality
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we construct the largest dataset for multimodal pretraining in Chinese, which consists of over 1.9TB images and 292GB texts that cover a wide range of domains. We propose a cross-modal pretraining method called M6, referring to Multi-Modality to Multi-Modality Multitask Mega-transformer, for unified pretraining on the data of single modality and multiple modalities. We scale the model size up to 10 billion and 100 billion parameters, and build the largest pretrained model in Chinese. We apply the model to a series of downstream applications, and demonstrate its outstanding performance in comparison with strong baselines. Furthermore, we specifically design a downstream task of text-guided image generation, and show that the finetuned M6 can create high-quality images with high resolution and abundant details.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Safe and Certifiable AI Systems: Concepts, Challenges, and Lessons Learned

    cs.CY 2025-09 conditional novelty 5.0 of 10

    The paper presents the TÜV AUSTRIA Trusted AI audit catalog, a statistical framework based on the Stochastic Application Domain Definition, minimum performance requirements, and independent-sample testing for certifyi...

  2. Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval

    cs.CV 2025-05 conditional novelty 4.0 of 10

    RDB improves remote sensing image-text retrieval mean recall by 1.15 to 2 percent over fully fine-tuned GeoRSCLIP using an asymmetric adapter and a dual-task consistency loss.

Pith tools