REVIEW 9 cited by
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Visual Question Answering (VQA) is an important task in multimodal AI, and it is often used to test the ability of vision-language models to understand and reason on knowledge present in both visual and textual data. However, most of the current VQA models use datasets that are primarily focused on English and a few major world languages, with images that are typically Western-centric. While recent efforts have tried to increase the number of languages covered on VQA datasets, they still lack diversity in low-resource languages. More importantly, although these datasets often extend their linguistic range via translation or some other approaches, they usually keep images the same, resulting in narrow cultural representation. To address these limitations, we construct CVQA, a new Culturally-diverse multilingual Visual Question Answering benchmark, designed to cover a rich set of languages and cultures, where we engage native speakers and cultural experts in the data collection process. As a result, CVQA includes culturally-driven images and questions from across 30 countries on four continents, covering 31 languages with 13 scripts, providing a total of 10k questions. We then benchmark several Multimodal Large Language Models (MLLMs) on CVQA, and show that the dataset is challenging for the current state-of-the-art models. This benchmark can serve as a probing evaluation suite for assessing the cultural capability and bias of multimodal models and hopefully encourage more research efforts toward increasing cultural awareness and linguistic diversity in this field.
Forward citations
Cited by 9 Pith papers
-
Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages
Across nine Asian languages, multilingual LLMs favor Western cultural entities in 30-40% of culturally grounded contexts, with model-specific sentiment biases and extraction accuracy gaps.
-
VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
A multilingual, multi-page document retrieval benchmark with 35K+ QA pairs shows MLLM retrievers lead but still fail on tables and low-resource languages.
-
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
A Wikidata-derived multilingual multimodal dataset improves cultural understanding of a vision-language model, yielding state-of-the-art results on cultural benchmarks among open models.
-
Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models
Marco-Bench-MIF localizes IFEval into 30 languages with cultural adaptation and finds large resource gaps and scale effects in multilingual instruction following.
-
CliniDial: A Naturally Occurring Multimodal Dialogue Dataset for Team Reflection in Action During Clinical Operation
A new multimodal dataset of simulated operating-room team dialogues shows that existing LLMs and fine-tuned models reach only about 51% macro F1 on team reflection behavior classification, indicating substantial room ...
-
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
An inference-time attention-shift intervention aligns non-English queries' cross-modal attention with English, cutting multilingual object hallucination in LVLMs on POPE and MME.
-
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation
RusCode is a new 1,250-prompt Russian/English benchmark for cultural awareness in text-to-image models, with human evaluation showing Russian-trained models outperform general models.
-
Towards General Continuous Memory for Vision-Language Models
A vision-language model can act as its own continuous memory encoder, compressing external multimodal knowledge into eight embeddings that improve reasoning when prepended to the frozen model.
-
The Human Labour of Data Work: Capturing Cultural Diversity through World Wide Dishes
A design retrospective of World Wide Dishes identifies three dimensions of community ambassador labor, trust building, accessibility, and cultural contextualization, as essential to participatory dataset creation.
Discussion (0). Continue with ORCID to comment.