REVIEW 9 cited by
Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Despite recent advances in multimodal large language models (MLLMs), their development has predominantly focused on English- and western-centric datasets and tasks, leaving most of the world's languages and diverse cultural contexts underrepresented. This paper introduces Pangea, a multilingual multimodal LLM trained on PangeaIns, a diverse 6M instruction dataset spanning 39 languages. PangeaIns features: 1) high-quality English instructions, 2) carefully machine-translated instructions, and 3) culturally relevant multimodal tasks to ensure cross-cultural coverage. To rigorously assess models' capabilities, we introduce PangeaBench, a holistic evaluation suite encompassing 14 datasets covering 47 languages. Results show that Pangea significantly outperforms existing open-source models in multilingual settings and diverse cultural contexts. Ablation studies further reveal the importance of English data proportions, language popularity, and the number of multimodal training samples on overall performance. We fully open-source our data, code, and trained checkpoints, to facilitate the development of inclusive and robust multilingual MLLMs, promoting equity and accessibility across a broader linguistic and cultural spectrum.
Forward citations
Cited by 9 Pith papers
-
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
A Wikidata-derived multilingual multimodal dataset improves cultural understanding of a vision-language model, yielding state-of-the-art results on cultural benchmarks among open models.
-
Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times
Video-language models perform far below humans on a new quadrilingual benchmark that tests understanding of action completion and duration through grammatical aspect.
-
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
This paper quantifies Western-centric bias in MMLU, releases Global-MMLU across 42 languages with human-verified translations, and shows model rankings shift on culturally sensitive versus agnostic subsets.
-
VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models
VARCO-VISION-14B is a Korean-English vision-language model that reports strong results among similar-size open models and introduces five Korean multimodal benchmarks.
-
mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
mSTEB is a new 200+ language speech and text benchmark showing that LLMs perform substantially worse on low-resource African and Americas/Oceania languages, especially in speech tasks.
-
AIN: The Arabic INclusive Large Multimodal Model
AIN, a 7B-parameter Arabic English multimodal model fine-tuned from Qwen2-VL on 3.6M samples, reports state-of-the-art Arabic scores including a 3.4-point average gain over GPT-4o on CAMEL-Bench.
-
Bridging Language Barriers in Healthcare: A Study on Arabic LLMs
The optimal Arabic-English training-data ratio for a medical LLM varies by task, and fine-tuning alone does not reliably improve Arabic clinical performance.
-
Maya: An Instruction Finetuned Multilingual Multimodal Model
Maya, an 8B multilingual multimodal model built on Aya-23 and SigLIP, shows small gains over PALO-7B on a PALO-based benchmark after finetuning on PALO instruction data.
-
The Multilingual Divide and Its Impact on Global AI Safety
The language gap in AI models creates safety disparities across languages, and closing it requires funding multilingual datasets, transparency, and research.
Discussion (0). Continue with ORCID to comment.