Pith. sign in

REVIEW 5 cited by

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.16508 v4 pith:OUSH23QF submitted 2024-11-25 cs.CV cs.CL

classification cs.CVcs.CL
keywords languagesalm-benchlmmsdiversemodelsbenchmarkculturalculturally
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivities, and support low-resource languages, all while effectively integrating corresponding visual cues. In pursuit of culturally diverse global multimodal models, our proposed All Languages Matter Benchmark (ALM-bench) represents the largest and most comprehensive effort to date for evaluating LMMs across 100 languages. ALM-bench challenges existing models by testing their ability to understand and reason about culturally diverse images paired with text in various languages, including many low-resource languages traditionally underrepresented in LMM research. The benchmark offers a robust and nuanced evaluation framework featuring various question formats, including true/false, multiple choice, and open-ended questions, which are further divided into short and long-answer categories. ALM-bench design ensures a comprehensive assessment of a model's ability to handle varied levels of difficulty in visual and linguistic reasoning. To capture the rich tapestry of global cultures, ALM-bench carefully curates content from 13 distinct cultural aspects, ranging from traditions and rituals to famous personalities and celebrations. Through this, ALM-bench not only provides a rigorous testing ground for state-of-the-art open and closed-source LMMs but also highlights the importance of cultural and linguistic inclusivity, encouraging the development of models that can serve diverse global populations effectively. Our benchmark is publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA

    cs.CL 2025-05 conditional novelty 7.0 of 10

    EverGreenQA and EG-E5 provide a multilingual, human-labeled evergreen question classifier that improves self-knowledge estimation and QA dataset curation.

  2. STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery

    eess.IV 2025-09 conditional novelty 6.0 of 10

    A new 1,000-video benchmark of stroke patients performing box-and-block sub-actions, with raw frames and 2D skeletons, establishes baseline action classification accuracy for seven models.

  3. Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Multi-TW is the first Traditional Chinese benchmark to evaluate multimodal models on both image-text and audio-text questions while also measuring inference latency.

  4. Mitigating Response Delays in Free-Form Conversations with LLM-powered Intelligent Virtual Agents

    cs.HC 2025-07 conditional novelty 5.0 of 10

    Natural conversational fillers improve perceived response time for VR agents when LLM responses are delayed above four seconds, while artificial wait indicators do not.

  5. LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A multilingual visual question-answering benchmark across 11 languages and 5 social attributes, evaluated on 7 large multimodal models.

Pith tools