Pith. sign in

REVIEW 9 cited by

Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.16153 v3 pith:F5ZCGSY4 submitted 2024-10-21 cs.CL cs.CV

classification cs.CLcs.CV
keywords multimodallanguagesmultilingualculturaldiversemodelspangeacontexts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite recent advances in multimodal large language models (MLLMs), their development has predominantly focused on English- and western-centric datasets and tasks, leaving most of the world's languages and diverse cultural contexts underrepresented. This paper introduces Pangea, a multilingual multimodal LLM trained on PangeaIns, a diverse 6M instruction dataset spanning 39 languages. PangeaIns features: 1) high-quality English instructions, 2) carefully machine-translated instructions, and 3) culturally relevant multimodal tasks to ensure cross-cultural coverage. To rigorously assess models' capabilities, we introduce PangeaBench, a holistic evaluation suite encompassing 14 datasets covering 47 languages. Results show that Pangea significantly outperforms existing open-source models in multilingual settings and diverse cultural contexts. Ablation studies further reveal the importance of English data proportions, language popularity, and the number of multimodal training samples on overall performance. We fully open-source our data, code, and trained checkpoints, to facilitate the development of inclusive and robust multilingual MLLMs, promoting equity and accessibility across a broader linguistic and cultural spectrum.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Grounding Multilingual Multimodal LLMs With Cultural Knowledge

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A Wikidata-derived multilingual multimodal dataset improves cultural understanding of a vision-language model, yielding state-of-the-art results on cultural benchmarks among open models.

  2. Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Video-language models perform far below humans on a new quadrilingual benchmark that tests understanding of action completion and duration through grammatical aspect.

  3. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

    cs.CL 2024-12 conditional novelty 6.0 of 10

    This paper quantifies Western-centric bias in MMLU, releases Global-MMLU across 42 languages with human-verified translations, and shows model rankings shift on culturally sensitive versus agnostic subsets.

  4. VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    VARCO-VISION-14B is a Korean-English vision-language model that reports strong results among similar-size open models and introduces five Korean multimodal benchmarks.

  5. mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks

    cs.CL 2025-06 conditional novelty 5.0 of 10

    mSTEB is a new 200+ language speech and text benchmark showing that LLMs perform substantially worse on low-resource African and Americas/Oceania languages, especially in speech tasks.

  6. AIN: The Arabic INclusive Large Multimodal Model

    cs.CV 2025-01 conditional novelty 5.0 of 10

    AIN, a 7B-parameter Arabic English multimodal model fine-tuned from Qwen2-VL on 3.6M samples, reports state-of-the-art Arabic scores including a 3.4-point average gain over GPT-4o on CAMEL-Bench.

  7. Bridging Language Barriers in Healthcare: A Study on Arabic LLMs

    cs.CL 2025-01 conditional novelty 5.0 of 10

    The optimal Arabic-English training-data ratio for a medical LLM varies by task, and fine-tuning alone does not reliably improve Arabic clinical performance.

  8. Maya: An Instruction Finetuned Multilingual Multimodal Model

    cs.CV 2024-12 conditional novelty 4.0 of 10

    Maya, an 8B multilingual multimodal model built on Aya-23 and SigLIP, shows small gains over PALO-7B on a PALO-based benchmark after finetuning on PALO instruction data.

  9. The Multilingual Divide and Its Impact on Global AI Safety

    cs.AI 2025-05 conditional novelty 3.0 of 10

    The language gap in AI models creates safety disparities across languages, and closing it requires funding multilingual datasets, transparency, and research.

Pith tools