Pith. sign in

REVIEW 5 cited by

CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.01282 v1 pith:I6AHPNSR submitted 2025-01-02 cs.AI cs.CLcs.CV

classification cs.AIcs.CLcs.CV
keywords culturalmodelsunderstandingconceptsperformancevlmscharacterizingcountries
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-language models (VLMs) have advanced human-AI interaction but struggle with cultural understanding, often misinterpreting symbols, gestures, and artifacts due to biases in predominantly Western-centric training data. In this paper, we construct CultureVerse, a large-scale multimodal benchmark covering 19, 682 cultural concepts, 188 countries/regions, 15 cultural concepts, and 3 question types, with the aim of characterizing and improving VLMs' multicultural understanding capabilities. Then, we propose CultureVLM, a series of VLMs fine-tuned on our dataset to achieve significant performance improvement in cultural understanding. Our evaluation of 16 models reveals significant disparities, with a stronger performance in Western concepts and weaker results in African and Asian contexts. Fine-tuning on our CultureVerse enhances cultural perception, demonstrating cross-cultural, cross-continent, and cross-dataset generalization without sacrificing performance on models' general VLM benchmarks. We further present insights on cultural generalization and forgetting. We hope that this work could lay the foundation for more equitable and culturally aware multimodal AI systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Failing to See or Failing to Know? Attributing Errors in Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.5 of 10

    VLM wrong answers in knowledge-intensive visual QA can be attributed to four decision points—recognition, visual evidence, answer success, factual access—with different pre-generation representations best predicting e...

  2. When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Parallel, role-specialized prompt agents improve cultural relevance in text-to-video generation, with a new cross-cultural benchmark showing the largest gains for location cues.

  3. Evaluation of Cultural Competence of Vision-Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    The paper proposes five theory-informed frameworks from visual cultural studies for evaluating cultural competence in vision-language models.

  4. SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression

    cs.CL 2025-07 conditional novelty 5.0 of 10

    SARA combines short natural-language snippets with vector-compressed summaries of the remaining retrieved documents, improving RAG answer quality under 512/1024-token context budgets.

  5. Domain Specific Benchmarks for Evaluating Multimodal Large Language Models

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.

Pith tools