Pith. sign in

REVIEW 9 cited by

CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.05967 v2 pith:AFY3QMJ5 submitted 2024-06-10 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords languagesmodelsbenchmarkculturalcvqavisualansweringdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual Question Answering (VQA) is an important task in multimodal AI, and it is often used to test the ability of vision-language models to understand and reason on knowledge present in both visual and textual data. However, most of the current VQA models use datasets that are primarily focused on English and a few major world languages, with images that are typically Western-centric. While recent efforts have tried to increase the number of languages covered on VQA datasets, they still lack diversity in low-resource languages. More importantly, although these datasets often extend their linguistic range via translation or some other approaches, they usually keep images the same, resulting in narrow cultural representation. To address these limitations, we construct CVQA, a new Culturally-diverse multilingual Visual Question Answering benchmark, designed to cover a rich set of languages and cultures, where we engage native speakers and cultural experts in the data collection process. As a result, CVQA includes culturally-driven images and questions from across 30 countries on four continents, covering 31 languages with 13 scripts, providing a total of 10k questions. We then benchmark several Multimodal Large Language Models (MLLMs) on CVQA, and show that the dataset is challenging for the current state-of-the-art models. This benchmark can serve as a probing evaluation suite for assessing the cultural capability and bias of multimodal models and hopefully encourage more research efforts toward increasing cultural awareness and linguistic diversity in this field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Across nine Asian languages, multilingual LLMs favor Western cultural entities in 30-40% of culturally grounded contexts, with model-specific sentiment biases and extraction accuracy gaps.

  2. VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A multilingual, multi-page document retrieval benchmark with 35K+ QA pairs shows MLLM retrievers lead but still fail on tables and low-resource languages.

  3. Grounding Multilingual Multimodal LLMs With Cultural Knowledge

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A Wikidata-derived multilingual multimodal dataset improves cultural understanding of a vision-language model, yielding state-of-the-art results on cultural benchmarks among open models.

  4. Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Marco-Bench-MIF localizes IFEval into 30 languages with cultural adaptation and finds large resource gaps and scale effects in multilingual instruction following.

  5. CliniDial: A Naturally Occurring Multimodal Dialogue Dataset for Team Reflection in Action During Clinical Operation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new multimodal dataset of simulated operating-room team dialogues shows that existing LLMs and fine-tuned models reach only about 51% macro F1 on team reflection behavior classification, indicating substantial room ...

  6. CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention

    cs.CL 2025-06 conditional novelty 6.0 of 10

    An inference-time attention-shift intervention aligns non-English queries' cross-modal attention with English, cutting multilingual object hallucination in LVLMs on POPE and MME.

  7. RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    RusCode is a new 1,250-prompt Russian/English benchmark for cultural awareness in text-to-image models, with human evaluation showing Russian-trained models outperform general models.

  8. Towards General Continuous Memory for Vision-Language Models

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A vision-language model can act as its own continuous memory encoder, compressing external multimodal knowledge into eight embeddings that improve reasoning when prepended to the frozen model.

  9. The Human Labour of Data Work: Capturing Cultural Diversity through World Wide Dishes

    cs.CY 2025-02 conditional novelty 4.0 of 10

    A design retrospective of World Wide Dishes identifies three dimensions of community ambassador labor, trust building, accessibility, and cultural contextualization, as essential to participatory dataset creation.

Pith tools