Pith. sign in

REVIEW 3 cited by

JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.17250 v2 pith:MXTSBNXP submitted 2024-10-22 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords japanesesubsetculturaljmmmulmmsunderstandingbenchmarkcontext
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Accelerating research on Large Multimodal Models (LMMs) in non-English languages is crucial for enhancing user experiences across broader populations. In this paper, we introduce JMMMU (Japanese MMMU), the first large-scale Japanese benchmark designed to evaluate LMMs on expert-level tasks based on the Japanese cultural context. To facilitate comprehensive culture-aware evaluation, JMMMU features two complementary subsets: (i) culture-agnostic (CA) subset, where the culture-independent subjects (e.g., Math) are selected and translated into Japanese, enabling one-to-one comparison with its English counterpart MMMU; and (ii) culture-specific (CS) subset, comprising newly crafted subjects that reflect Japanese cultural context. Using the CA subset, we observe performance drop in many LMMs when evaluated in Japanese, which is purely attributable to language variation. Using the CS subset, we reveal their inadequate Japanese cultural understanding. Further, by combining both subsets, we identify that some LMMs perform well on the CA subset but not on the CS subset, exposing a shallow understanding of the Japanese language that lacks depth in cultural understanding. We hope this work will not only help advance LMM performance in Japanese but also serve as a guideline to create high-standard, culturally diverse benchmarks for multilingual LMM development. The project page is https://mmmu-japanese-benchmark.github.io/JMMMU/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MyCulture: Exploring Malaysia's Diverse Culture under Low-Resource Language Constraints

    cs.CL 2025-08 reject novelty 6.0 of 10

    MyCulture, a new Malay-language cultural benchmark, shows LLM accuracy drops by at least 17% when multiple-choice questions are converted to an open-ended format.

  2. J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM

    cs.CV 2024-12 conditional novelty 6.0 of 10

    J-EDI QA is a new 100-image Japanese multiple-choice benchmark for deep-sea organism identification; OpenAI o1 scored 50%, GPT-4o 39%, and non-expert humans about 40%.

  3. Development of a Large-scale Dataset of Chest Computed Tomography Reports in Japanese and a High-performance Finding Classification Model

    cs.CL 2024-12 conditional novelty 4.0 of 10

    The paper creates CT-RATE-JPN, a Japanese version of the CT-RATE CT report dataset, and CT-BERT-JPN, a Japanese BERT model that classifies 18 chest CT findings in Japanese reports.

Pith tools