Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Evaluation of Cultural Competence of Vision-Language Models

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Five cultural dimensions AI vision benchmarks skip

desk verdict A genuine position paper that maps visual cultural studies onto VLM evaluation; the five-framework synthesis is new and useful, but the 'must be considered' claim is unvalidated without a pilot. read the letter →

arxiv 2505.22793 v2 pith:E3FFWJYD submitted 2025-05-28 cs.CV cs.CL

classification cs.CVcs.CL
keywords culturalcompetencevision-languagemodelsvisualstudiessemioticsevaluationbenchmarksbiasmultimodalAIpositionpaper
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that current cultural benchmarks for vision-language models (VLMs) mostly test whether a model can name a culturally tagged object, which treats culture as a list of visible things. It proposes a paradigm shift toward theory-informed evaluation grounded in visual cultural studies, and surveys 35 recent works to show that cultural categories are chosen inconsistently and rarely connected to cultural theory. The five frameworks it puts forward are processual grounding (who defines and judges culture), material and embodied culture (what objects are and how they are made, used, and socially situated), symbolic and semiotic encoding (how meaning is layered beyond the literal), contextual interpretation (who evaluates and how perception is framed), and temporality (how meaning changes over time). For each dimension, the paper suggests annotation schemas and evaluation tasks, such as contrastive captions that distinguish literal from symbolic readings. If the paper is right, high scores on today's cultural recognition benchmarks are not evidence of cultural competence, and future evaluations should be insider-driven, context-aware, and time-sensitive.

What carries the argument

The central object is an inventory of five cultural dimensions, each anchored to a named theory or method: the emic/etic distinction and participatory visual methods (photo elicitation, Photovoice); Tilley et al.'s material-culture taxonomy; semiotic frameworks of Barthes, de Saussure, Peirce, and Panofsky together with Hall's high-context/low-context distinction; Hall's encoding/decoding model of contextual interpretation; and Appadurai's notion of the social life of things for temporality. This inventory does the argument's work by functioning as a checklist: for each dimension, the paper proposes concrete annotation categories or evaluation tasks that current object-recognition benchmarks omit, which turns abstract cultural theory into a template for future benchmark design.

What would settle it

A concrete test would be to build an evaluation around the five proposed dimensions with community insiders doing the labeling, then measure inter-annotator agreement and compare model rankings on those tasks with rankings on existing object-recognition benchmarks; if agreement is near chance, or if the rankings coincide perfectly, the added dimensions carry no measurable signal. A second, simpler falsifier is that if the proposed annotation schemas cannot be implemented at all because annotators cannot consistently identify the categories, the framework fails at the operationalization step the paper delegates to the community.

Watch

Extended reading notes

Core claim

The paper's central claim is that cultural competence in VLMs cannot be measured by recognizing entities such as foods, clothing, or rituals, because culture operates through dimensions that object naming does not capture. Drawing on visual cultural studies, it proposes a set of five frameworks that should guide annotation and evaluation: processual grounding, based on emic versus etic perspectives and participatory methods such as photo elicitation and Photovoice; material culture, based on a taxonomy of object types, visual properties, spatial relations, use contexts, symbolic elements, social associations, and production markers; symbolic and semiotic encoding, based on the distinction between denotation and connotation, Saussurean signifier and signified, Peirce's firstness, secondness, and thirdness, and Panofsky's three iconological levels; contextual interpretation, based on Hall's encoding and decoding model, high-context versus low-context communication, and insider-led evaluation; and temporality, based on the social life of things and the historicity of meaning. For each framework, the paper gives example annotation schemas and benchmark tasks, such as contrastive captions that distinguish literal from symbolic readings, questions about use and production, and diachronic archives. The paper is a position paper: it does not run experiments with these frameworks, and explicitly leaves operationalization to the research community.

Load-bearing premise

The proposal stands or falls on the assumption that the theoretical dimensions from visual cultural studies can be turned into reliable annotation categories and evaluation tasks that annotators and models can actually use.

Editorial extensions

If this is right

  • Current cultural VQA and recognition benchmarks measure surface naming, so high scores on them should not be taken as evidence of cultural competence.
  • Future cultural datasets should annotate use contexts, production markers, social associations, symbolic layers, and time period, not just entity categories.
  • Evaluation protocols should shift judgment of cultural fidelity to insiders, using participatory methods, rather than external annotators applying fixed templates.
  • New tasks would include contrastive captions pitting literal against symbolic meanings, gesture interpretation, advertisement-context analysis, and diachronic questions over historical image collections.
  • Text-to-image and advertisement generation should be assessed on how well they handle high-context versus low-context communication and whose frame defines the prompt.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next step, not in the paper, is to turn Panofsky's three levels into a task hierarchy and test whether model accuracy degrades monotonically from pre-iconographic description to iconological interpretation; such a gradient would give the framework a quantitative signature.
  • If the five dimensions are adopted, the field's current practice of reusing generic taxonomies would need justification from cultural theory, which could change how datasets are compared and make cross-dataset generalization a meaningful evaluation target.
  • The proposal likely extends beyond evaluation into training: if culture is layered, situated, and time-dependent, then datasets built from static, crowd-sourced images may be structurally insufficient, and participatory or archival collection methods may be needed to supply the missing signal.
  • A testable extension is a codebook study where community insiders annotate the same images on all five dimensions; low inter-annotator agreement would strengthen the paper's claim that culture is contested and context-dependent, while high agreement would show the dimensions are operationalizable.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This position paper argues that current benchmarks and evaluations of cultural competence in vision-language models (VLMs) operate at a surface level, largely reducing culture to object recognition or naming. Drawing on visual cultural studies, semiotics, cultural studies, and visual studies, the authors survey 35 recent papers and propose five frameworks that they argue "must be considered" for a more complete cultural analysis of images: processual grounding, material and embodied culture, symbolic and semiotic encoding, contextual interpretation, and temporality. Each framework is tied to existing humanities theory (e.g., Peirce, Panofsky, Hall, Tilley) and to suggested evaluation tasks, such as contrastive caption selection, insider-led text-to-image evaluation, and a context score for advertisements. The paper is explicitly a position piece: it states in the introduction that it "does not aim to experiment" and leaves operationalization to the community. The conclusion reiterates that the goal is to revisit foundational theories of culture to inform future development and evaluation of multimodal systems.

Significance. If the proposed frameworks can be operationalized, this paper could substantially shift how the multimodal community thinks about cultural evaluation: from a recognition-centric view to a meaning-oriented analysis grounded in decades of humanities scholarship. The synthesis of visual cultural studies with VLM evaluation is timely and valuable, and the paper provides a useful catalog of theoretical tools and concrete task proposals, including the well-developed contrastive-caption idea in Table 2 and the advertisement context-score questions in Appendix A.2. The survey of 35 papers, despite methodological gaps, is a useful contribution in identifying fragmentation in existing work. However, the central claim of the paper—that these frameworks "must be considered"—is a hypothesis, not a demonstrated result. Its significance therefore depends on future validation of the reliability and validity of the proposed measurements, which the manuscript does not provide.

major comments (3)
  1. [Section 4 and Tables 2-3, Appendix A.2] The central claim that the five frameworks "must be considered" rests on an untested assumption: that high-inference constructs such as Peircean thirdness, Panofsky's iconological levels, Hall's high-/low-context distinction, and the material-culture taxonomy can be reliably annotated by humans and used to produce valid model evaluations. The paper itself states in Section 1 that it "does not aim to experiment" and leaves operationalization to the community, and it offers no inter-annotator agreement data, construct-validity evidence, or comparison with existing surface-level benchmarks. Without at least a proof-of-concept that one of the proposed tasks (e.g., the contrastive-caption task in Table 2) yields consistent labels and non-redundant signal beyond current benchmarks, the necessity claim is unsupported. I recommend either adding a pilot study on one dimension or explicitly reframing the contribution as a research agenda that requires validation rather than as an established requirement.
  2. [Section 3.1 and Section 4.3 / Appendix A.2] The paragraph on "High-Context and Low-Context Cultural Messages" contains an inversion of Hall's definitions. The text says: "In high-context cultural messages, people tend to respond to make meaning constructed through straightforward and unambiguous language, often in the form of text." This is the opposite of Hall's definition quoted immediately before it, where high-context communication has most information already in the person and little in the explicitly coded message. Since the context-score proposal in Appendix A.2 and the discussion in Section 4.3 depend on this distinction, the error is not merely typographical; it could mislead readers applying the proposed metrics. Please correct the definitions and ensure the subsequent discussion is consistent with Hall's original account.
  3. [Section 2.1 and Appendix A.5 (Tables 5-6)] The survey of 35 papers is presented as the empirical basis for the claimed gap in the literature, but no systematic review methodology is reported. The appendix does not specify the search strategy, inclusion/exclusion criteria, or a coding protocol for the categories "material culture" and "semiotically layered". Several rows in Table 5 assign categories to papers that are listed as not clearly mentioning cultural concepts (e.g., ViTextVQA, HaVQA, SEA-VQA), suggesting the categorizations may be applied inconsistently. Without a transparent protocol, the reader cannot verify the claim that "a significant number of papers do not elaborate visual categories... or their methodologies grounded in visual cultural studies." Please add the review protocol and, if possible, inter-coder agreement for the taxonomy.
minor comments (4)
  1. [Throughout] There are several typos and spelling inconsistencies: "Pansofsky" should be "Panofsky" (Sections 3.1 and elsewhere), "These datasets cme from different sources" should be "come" (Section 2.1), "we observed" in Section 2.2 should be followed by a period and capital letter, "acount" should be "account" (Section 3.4), and a missing space appears in "such asdataset curation" (Section 4.2) and "such ascontrastive caption evaluation" (Section 4.3).
  2. [Figure 1] The text below Figure 1 states that the gap shown "exists in various datasets with the stated goal of studying cultural knowledge in VLMs," but the figure only illustrates CVQA. Please name or cite at least two or three additional datasets to support this generalization, or adjust the wording to indicate that this is an illustrative example.
  3. [Appendix A.2] The table caption is redundant: "Proposed questions to test the visual complexity of advertisements, taken from Hornikx and le Pair (2017)" already appears in the main text, and the caption then repeats "Proposed questions to test ad's visual complexity taken from Hornikx and le Pair (2017)…". Please simplify the caption and avoid duplication.
  4. [Section 4.1] The term "enregisterment-aware evaluation" is introduced without a definition or reference in the main text; consider providing a brief explanation or a pointer to the source (Nakassis, 2023) so that readers outside linguistics can follow the argument.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the five frameworks are grounded in external humanities scholarship, and the paper's self-citations are contextual survey entries rather than load-bearing premises.

full rationale

This is a position paper that makes no quantitative predictions and fits no parameters. Its central proposal—that evaluation of VLM cultural competence should be grounded in visual cultural studies—is developed from external sources (Peirce, Panofsky, Hall, Tilley, Barthes, Appadurai, etc.), not from the authors' prior results. The self-citations that appear (Yadav et al. 2025, Cao et al. 2024, Li et al. 2025, Karamolegkou et al. 2024, Hershcovich et al. 2022) are used as examples of existing work in the survey or as background motivation; none of the five frameworks in Section 4 depends on these citations for its content. The paper explicitly states that it 'does not aim to experiment with these frameworks' and leaves operationalization to the community, so there are no fitted inputs renamed as predictions and no derivation chain whose conclusion equals its assumptions. The acknowledged open question—whether the theoretical dimensions can be made into reliable annotation tasks—concerns empirical validation, not circularity. The derivation is thus self-contained in the relevant sense.

Assumptions & free parameters 0 free parameters · 4 assumptions · 1 invented entities

The central argument rests on domain assumptions about what culture is and how it can be measured. No free parameters are fitted. The five proposed frameworks are new conceptual constructs with no independent empirical evidence. The paper acknowledges its value assumption in the Limitations section.

assumptions (4)
  • domain assumption Cultural meaning in images can be systematically identified and annotated through the proposed dimensions.
    The entire paper assumes that the five frameworks are the right lens for cultural analysis and can be made operational, but no annotation study or validation is provided.
  • domain assumption Current VLM cultural evaluations are inadequate because they focus on surface-level recognition.
    Section 2 argues this based on a survey of 35 papers, but no controlled comparison demonstrates that deeper frameworks yield different or better evaluations.
  • domain assumption Improving cultural competence of VLMs is beneficial.
    The Limitations section explicitly states this is an underlying assumption, acknowledging it is not always true.
  • domain assumption The three disciplines of cultural studies, semiotics, and visual studies are representative and sufficient to ground cultural analysis.
    The Limitations section notes these fields are 'just one path' and that other works may be missing from the review.
invented entities (1)
  • Five proposed frameworks: processual grounding, material and embodied culture, symbolic and semiotic encoding, contextual interpretation, and temporality
    purpose: To organize cultural dimensions that VLM evaluation should consider, beyond object naming.
    These are conceptual constructs introduced in Section 4. They are grounded in prior humanities literature but are not empirically validated and do not make falsifiable predictions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluation of Cultural Competence of Vision-Language Models." pith.science (2026). https://pith.science/paper/E3FFWJYD

@misc{pith2026250522793,
  author       = {Pith},
  title        = {Pith review of: Evaluation of Cultural Competence of Vision-Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E3FFWJYD}},
  note         = {Machine review of arXiv:2505.22793}
}
read the original abstract

Modern vision-language models (VLMs) often fail at cultural competency evaluations and benchmarks. Given the diversity of applications built upon VLMs, there is renewed interest in understanding how they encode cultural nuances. While individual aspects of this problem have been studied, we still lack a comprehensive framework for systematically identifying and annotating the nuanced cultural dimensions present in images for VLMs. This position paper argues that foundational methodologies from visual culture studies (cultural studies, semiotics, and visual studies) are necessary for cultural analysis of images. Building upon this review, we propose a set of five frameworks, corresponding to cultural dimensions, that must be considered for a more complete analysis of the cultural competencies of VLMs.

Figures

Figures reproduced from arXiv: 2505.22793 by the authors.

Figure 1
Figure 1. An image from a cultural evaluation dataset (CVQA (Romero et al., 2025)), compared [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Plot showing the frequency of use of models for evaluation of cultural knowledge. All the [PITH_FULL_IMAGE:figures/full_fig_p018_2.png] view at source ↗
Figure 3
Figure 3. An AI-generated advertisement (with prompt ‘draw a liquid detergent advertisement for [PITH_FULL_IMAGE:figures/full_fig_p020_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Overview of the distribution of concepts across the literature studied (Appendix A.6). Each [PITH_FULL_IMAGE:figures/full_fig_p021_4.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Across 26 LVLMs on MemeBench, all models show a 14.6-29.0% gap between visual coverage and cultural-knowledge coverage, and retrieval raises knowledge while lowering visual coverage.

Reference graph

Works this paper leans on

117 extracted references · 55 canonical work pages · cited by 1 Pith paper

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Utkarsh Agarwal, Kumar Tanmay, Aditi Khandelwal, and Monojit Choudhury. 2024. Ethical reasoning and moral value alignment of llms depend on the language we prompt them in. arXiv preprint arXiv:2404.18460

  4. [4]

    Aysan Aghazadeh and Adriana Kovashka. 2024. Cap: Evaluation of persuasive and creative image generation. arXiv preprint arXiv:2412.10426

  5. [5]

    Badr Alkhamissi, Muhammad ElNokrashy, Mai Alkhamissi, and Mona Diab. 2024. Investigating cultural alignment of large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12404--12422

  6. [6]

    Arjun Appadurai. 1988. The social life of things: Commodities in cultural perspective. Cambridge University Press

  7. [7]

    Michael Argyle, Mark Cook, and Duncan Cramer. 1994. Gaze and mutual gaze. The British Journal of Psychiatry, 165(6):848--850

  8. [8]

    Taylor Arnold and Lauren Tilton. 2023. Distant Viewing: Computational Exploration of Digital Images . MIT Press

Show all 117 references
  1. [9]

    Sarah H Awad. 2020. The social life of images. Visual Studies, 35(1):28--39

  2. [10]

    Yujin Baek, ChaeHun Park, Jaeseok Kim, Yu-Jung Heo, Du-Seong Chang, and Jaegul Choo. 2024. Evaluating visual and cultural interpretation: The k-viscuit benchmark with human-vlm collaboration. CoRR

  3. [11]

    Roland Barthes. 1977. Image, Music, Text. Hill and Wang

  4. [12]

    Venkatesh Babu, and Danish Pruthi

    Abhipsa Basu, R. Venkatesh Babu, and Danish Pruthi. 2023. https://doi.org/10.1109/ICCV51070.2023.00474 Inspecting the geographical representativeness of images from text-to-image models . In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 5113--5124

  5. [13]

    Aaron Olaf Batty and Ruslan Suvorov. 2024. Visual cues and listening. In The Routledge Handbook of Second Language Acquisition and Listening, pages 307--318. Routledge

  6. [14]

    John Berger. 1973. Ways of Seeing. Penguin

  7. [15]

    Shaily Bhatt and Fernando Diaz. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.942 Extrinsic evaluation of cultural competence in large language models . In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 16055--16074, Miami, Florida, USA. A...

  8. [16]

    Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, and Aylin Caliskan. 2023. https://doi.org/10.1145/3593013.3594095 Easily accessible text-to-image generation amplifies demographic stereotyp...

  9. [17]

    Yi Bin, Wenhao Shi, Yujuan Ding, Zhiqiang Hu, Zheng Wang, Yang Yang, See-Kiong Ng, and Heng Tao Shen. 2024. Gallerygpt: Analyzing paintings with large multimodal models. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 7734--7743

  10. [18]

    Steven Bird. 2020. https://doi.org/10.18653/v1/2020.coling-main.313 Decolonising speech and language technology . In Proceedings of the 28th International Conference on Computational Linguistics, pages 3504--3519, Barcelona, Spain (Online). International Committee on Computati...

  11. [19]

    Steven Bird. 2024. https://doi.org/10.18653/v1/2024.acl-long.797 Must NLP be extractive? In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14915--14929, Bangkok, Thailand. Association for Computational Linguistics

  12. [20]

    Ray L Birdwhistell. 1970. Some meta-communicational thoughts about communicational studies. Language behavior: A book of readings in communication, pages 265--270

  13. [21]

    Marie-Louise Brunner and Stefan Diemer. 2021. Multimodal meaning making: The annotation of nonverbal elements in multimodal corpus transcription. Research in Corpus Linguistics, 9(1):63--88

  14. [22]

    Yong Cao, Wenyan Li, Jiaang Li, Yifei Yuan, Antonia Karamolegkou, and Daniel Hershcovich. 2024. Exploring visual culture awareness in gpt-4v: A comprehensive probing. arXiv preprint arXiv:2402.06015

  15. [23]

    Ji r \'i C en e k and S a s inka C en e k. 2015. https://doi.org/10.15503/jecs20151.187.206 Cross-cultural differences in visual perception . Journal of Education Culture and Society, 6(1):187--206

  16. [24]

    Daniel Chandler. 2002. The Basics. Routledge London, UK

  17. [25]

    Daniel Chandler. 2022. Semiotics: The Basics. Routledge

  18. [26]

    Rochelle Choenni and Ekaterina Shutova. 2024. Self-alignment: Improving alignment of cultural values in llms via in-context learning. arXiv preprint arXiv:2408.16482

  19. [27]

    Ferdinand de Saussure. 1916. Cours de linguistique générale. Payot

  20. [28]

    Paul Ekman and Wallace V Friesen. 1969. The repertoire of nonverbal behavior: Categories, origins, usage, and coding. semiotica, 1(1):49--98

  21. [29]

    Fengping Gao et al. 2005. Japanese: A heavily culture-laden language. Journal of Intercultural Communication, 5(3):1--09

  22. [30]

    Christian Haerpfer, Ronald Inglehart, Alejandro Moreno, Christian Welzel, Kseniya Kizilova, Jaime Diez-Medrano, Marta Lagos, Pippa Norris, Eduard Ponarin, and Bi Puranen. 2022. https://doi.org/10.14281/18241.18 World values survey wave 7 (2017–2022) cross-national data-set

  23. [31]

    Melissa Hall, Candace Ross, Adina Williams, Nicolas Carion, Michal Drozdzal, and Adriana Romero Soriano. 2023. https://arxiv.org/abs/2308.06198 Dig in: Evaluating disparities in image generations with indicators for geographic diversity . Preprint, arXiv:2308.06198

  24. [32]

    Stuart Hall. 1997. Representation: Cultural Representations and Signifying Practices. SAGE Publications

  25. [33]

    Stuart Hall. 2007. Encoding and decoding in the television discourse. In CCCS selected working papers, pages 402--414. Routledge

  26. [34]

    Ben Halpern. 1955. The dynamic elements of culture. Ethics, 65(4):235--249

  27. [35]

    Douglas Harper. 1988. Visual sociology: Expanding sociological vision. The american sociologist, 19:54--70

  28. [36]

    Marvin Harris. 1968. Current Anthropology, http://www.jstor.org/stable/2740497 9(5):519--533

  29. [37]

    Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, et al. 2022. Challenges and strategies in cross-cultural nlp. arXiv preprint arXiv:2203.10020

  30. [38]

    Geert Hofstede. 1983. https://doi.org/10.1080/00208825.1983.11656358 National cultures in four dimensions: A research-based theory of cultural differences among nations . International Studies of Management and Organization, 13(1-2):46--74

  31. [39]

    Jos Hornikx and Rob le Pair. 2017. The influence of high-/low-context culture on perceived ad complexity and liking. Journal of Global Marketing, 30(4):228--237

  32. [40]

    Akshita Jha, Aida Mostafazadeh Davani, Chandan K Reddy, Shachi Dave, Vinodkumar Prabhakaran, and Sunipa Dev. 2023. https://doi.org/10.18653/v1/2023.acl-long.548 S ee GULL : A stereotype benchmark with broad geo-cultural coverage leveraging generative models . In Proceedings of...

  33. [41]

    Akshita Jha, Vinodkumar Prabhakaran, Remi Denton, Sarah Laszlo, Shachi Dave, Rida Qadri, Chandan Reddy, and Sunipa Dev. 2024. Visage: A global-scale analysis of visual stereotypes in text-to-image generation. In Proceedings of the 62nd Annual Meeting of the Association for Com...

  34. [42]

    Antonia Karamolegkou, Phillip Rust, Ruixiang Cui, Yong Cao, Anders S gaard, and Daniel Hershcovich. 2024. https://doi.org/10.18653/v1/2024.hucllm-1.5 Vision-language models under cultural and inclusive considerations . In Proceedings of the 1st Human-Centered Large Language Mo...

  35. [43]

    Adam Kendon. 2004. Gesture: Visible action as utterance. Cambridge University Press

  36. [44]

    Mary Ritchie Key and Bernard Comrie. 2015. https://www.eva.mpg.de/linguistics/past-research-resources/language-history/intercontinental-dictionary-series-ids/ Intercontinental Dictionary Series (IDS) . Leipzig. Accessed May 22, 2025

  37. [45]

    Simran Khanuja, Sathyanarayanan Ramamoorthy, Yueqi Song, and Graham Neubig. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.573 An image speaks a thousand words, but can everyone listen? on image transcreation for cultural relevance . In Proceedings of the 2024 Conference on...

  38. [46]

    Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of com...

  39. [47]

    Tatiana Larina. 2015. Culture-specific communicative styles as a framework for interpreting linguistic and cultural idiosyncrasies. International Review of Pragmatics, 7(2):195--215

  40. [48]

    Cheng Li, Mengzhuo Chen, Jindong Wang, Sunayana Sitaram, and Xing Xie. 2024 a . Culturellm: Incorporating cultural differences into large language models. Advances in Neural Information Processing Systems, 37:84799--84838

  41. [49]

    Jiaang Li, Yifei Yuan, Wenyan Li, Mohammad Aliannejadi, Daniel Hershcovich, Anders S gaard, Ivan Vuli \'c , Wenxuan Zhang, Paul Pu Liang, Yang Deng, et al. 2025. Ravenea: A benchmark for multimodal retrieval-augmented visual culture understanding. arXiv preprint arXiv:2505.14462

  42. [50]

    Wenyan Li, Crystina Zhang, Jiaang Li, Qiwei Peng, Raphael Tang, Li Zhou, Weijia Zhang, Guimin Hu, Yifei Yuan, Anders S gaard, et al. 2024 b . Foodieqa: A multimodal dataset for fine-grained understanding of chinese food culture. In Proceedings of the 2024 Conference on Empiric...

  43. [51]

    Zhi Li and Yin Zhang. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.18 Cultural concept adaptation on multimodal reasoning . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 262--276, Singapore. Association for Computational ...

  44. [52]

    Meaning through motion: Duet--a multimodal dataset for kinesics analysis in dyadic activities

    Cheyu Lin, Katherine A Flanigan, and Sirajum Munir. Meaning through motion: Duet--a multimodal dataset for kinesics analysis in dyadic activities. In NeurIPS 2024 Workshop on Behavioral Machine Learning

  45. [53]

    Kenneth C. Lindsay. 1966. http://www.jstor.org/stable/30199202 Art, art history, and the computer . Computers and the Humanities, 1(2):27--30

  46. [54]

    Bingshuai Liu, Longyue Wang, Chenyang Lyu, Yong Zhang, Jinsong Su, Shuming Shi, and Zhaopeng Tu. 2024 a . On the cultural gap in text-to-image generation. In ECAI 2024, pages 930--937. IOS Press

  47. [55]

    Chen Liu, Fajri Koto, Timothy Baldwin, and Iryna Gurevych. 2024 b . Are multilingual llms culturally-diverse reasoners? an investigation into multicultural proverbs and sayings. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computat...

  48. [56]

    Chen Cecilia Liu, Iryna Gurevych, and Anna Korhonen. 2024 c . Culturally aware and adapted nlp: A taxonomy and a survey of the state of the art. CoRR

  49. [57]

    Fangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy, Nigel Collier, and Desmond Elliott. 2021. Visually grounded reasoning across languages and cultures. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10467--10485

  50. [58]

    Shudong Liu, Yiqiao Jin, Cheng Li, Derek F Wong, Qingsong Wen, Lichao Sun, Haipeng Chen, Xing Xie, and Jindong Wang. 2025. Culturevlm: Characterizing and improving cultural understanding of vision-language models for over 100 countries. arXiv preprint arXiv:2501.01282

  51. [59]

    Zhixuan Liu, Youeun Shin, Beverley-Claire Okogwu, Youngsik Yun, Lia Coleman, Peter Schaldenbrand, Jihie Kim, and Jean Oh. 2023. Towards equitable representation in text-to-image synthesis models with the cross-cultural understanding benchmark (ccub) dataset. arXiv preprint arX...

  52. [60]

    Jinshun Long and Jun He. 2021. Cultural semiotics and the related interpretation. In 2021 International Conference on Public Relations and Social Sciences (ICPRSS 2021), pages 1268--1272. Atlantis Press

  53. [61]

    Oscar Ma\ n as, Benno Krojer, and Aishwarya Agrawal. 2024. https://doi.org/10.1609/aaai.v38i5.28212 Improving automatic vqa evaluation using large language models . In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on In...

  54. [62]

    Jabez Magomere, Shu Ishida, Tejumade Afonja, Aya Salama, Daniel Kochin, Foutse Yuehgoh, Imane Hamzaoui, Raesetje Sefala, Aisha Alaagib, Elizaveta Semenova, et al. 2024. You are what you eat? feeding foundation models a regionally diverse food dataset of world wide dishes. CoRR

  55. [63]

    Lev Manovich. 2020. Cultural analytics. MIT Press

  56. [64]

    Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition, pages 3195--3204

  57. [65]

    David Matsumoto and Hyisung C Hwang. 2013. Cultural similarities and differences in emblematic gestures. Journal of Nonverbal Behavior, 37:1--27

  58. [66]

    David Matsumoto, Seung Hee Yoo, and Sanae Nakagawa. 2008. Culture, emotion regulation, and adjustment. Journal of personality and social psychology, 94(6):925

  59. [67]

    Nicholas Mirzoeff. 1999. An introduction to visual culture, volume 274. Routledge London

  60. [68]

    WJ Thomas Mitchell. 1995. Picture theory: Essays on verbal and visual representation. University of Chicago Press

  61. [69]

    Surbhi Mittal, Arnav Sudan, Mayank Vatsa, Richa Singh, Tamar Glaser, and Tal Hassner. 2024. Navigating text-to-image generative bias across indic languages. In European Conference on Computer Vision, pages 53--67. Springer

  62. [70]

    Youssef Mohamed, Runjia Li, Ibrahim Said Ahmad, Kilichbek Haydarov, Philip Torr, Kenneth Ward Church, and Mohamed Elhoseiny. 2024. No culture left behind: Artelingo-28, a benchmark of wikiart with captions in 28 languages. arXiv preprint arXiv:2411.03769

  63. [71]

    Altyn N Muratova, BZ SharaMazhitayeva, A Sarybayeva, and ZK Kelmaganbetova. 2021. Non-verbal signs and secret communication as universal signs of intercultural communication. Rupkatha journal on interdisciplinary studies in humanities, 13(1):1--9

  64. [72]

    Gregory Murphy. 2004. The big book of concepts. MIT press

  65. [73]

    Constantine V Nakassis. 2023. A linguistic anthropology of images. Annual Review of Anthropology, 52(1):73--91

  66. [74]

    Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy, Sjoerd Steenkiste, Lisa Hendricks, Karolina Stanczak, and Aishwarya Agrawal. 2024 a . Benchmarking vision language models for cultural understanding. In Proceedings of the 2024 Conference on Empirical Methods in Natural Lan...

  67. [75]

    Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy, Sjoerd Van Steenkiste, Lisa Anne Hendricks, Karolina Stanczak, and Aishwarya Agrawal. 2024 b . https://doi.org/10.18653/v1/2024.emnlp-main.329 Benchmarking vision language models for cultural understanding . In Proceedings ...

  68. [76]

    Tuan-Phong Nguyen, Simon Razniewski, Aparna Varde, and Gerhard Weikum. 2023. Extracting cultural commonsense knowledge at scale. In Proceedings of the ACM Web Conference 2023, pages 1907--1917

  69. [77]

    Malvina Nikandrou, Georgios Pantazopoulos, Nikolas Vitsakis, Ioannis Konstas, and Alessandro Suglia. 2025. https://aclanthology.org/2025.naacl-long.402/ CROPE : Evaluating in-context adaptation of vision and language models to culture-specific concepts . In Proceedings of the ...

  70. [78]

    Nisbett and Takahiko Masuda

    Richard E. Nisbett and Takahiko Masuda. 2003. https://doi.org/10.1073/pnas.1934527100 Culture and point of view . Proceedings of the National Academy of Sciences of the United States of America, 100(19):11163--11170

  71. [79]

    Richard E Nisbett and Yuri Miyamoto. 2005. The influence of culture: holistic versus analytic perception. Trends in cognitive sciences, 9(10):467--473

  72. [80]

    Erwin Panofsky. 1939. Iconology

  73. [81]

    Shantipriya Parida, Idris Abdulmumin, Shamsuddeen Hassan Muhammad, Aneesh Bose, Guneet Singh Kohli, Ibrahim Sa’id Ahmad, Ketan Kotwal, Sayan Deb Sarkar, Ond r ej Bojar, and Habeebah Kakudi. 2023. Havqa: A dataset for visual question answering and multimodal research in hausa l...

  74. [82]

    Bolette Sandford Pedersen, Nathalie S rensen, Sanni Nimb, Dorte Haltrup Hansen, Sussi Olsen, and Ali Al-Laith. 2025. Evaluating llm-generated explanations of metaphors--a culture-sensitive study of danish. In Proceedings of the Joint 25th Nordic Conference on Computational Lin...

  75. [83]

    Charles S. Peirce. 1868. On a new list of categories. Proceedings of the American Academy of Arts and Sciences, 7:287--298

  76. [84]

    Kenneth L. Pike. 1967. https://doi.org/doi:10.1515/9783111657158 Language in Relation to a Unified Theory of the Structure of Human Behavior . De Gruyter Mouton, Berlin, Boston

  77. [85]

    Ramaswamy, Sing Yu Lin, Dora Zhao, Aaron Adcock, Laurens van der Maaten, Deepti Ghadiyaram, and Olga Russakovsky

    Vikram V. Ramaswamy, Sing Yu Lin, Dora Zhao, Aaron Adcock, Laurens van der Maaten, Deepti Ghadiyaram, and Olga Russakovsky. 2023. https://proceedings.neurips.cc/paper_files/paper/2023/file/d08b6801f24dda81199079a3371d77f9-Paper-Datasets_and_Benchmarks.pdf Geode: a geographical...

  78. [86]

    Rieko Maruta Richardson and Sandi W Smith. 2007. The influence of high/low-context culture and power distance on choice of communication media: Students’ media choice to communicate with professors in japan and america. International Journal of Intercultural Relations, 31(4):479--501

  79. [87]

    Milton Rokeach. 2006. https://doi.org/10.4135/9781412952675.n244 Rokeach values survey . In Jeffrey H. Greenhaus and Gerard A. Callanan, editors, Encyclopedia of Career Development, pages 701--701. SAGE Publications, Inc

  80. [88]

    David Romero, Chenyang Lyu, Haryo Wibowo, Santiago G \'o ngora, Aishik Mandal, Sukannya Purkayastha, Jesus-German Ortiz-Barajas, Emilio Cueva, Jinheon Baek, Soyeong Jeong, et al. 2025. Cvqa: Culturally-diverse multilingual visual question answering benchmark. Advances in Neura...

  81. [89]

    Muna Numan Said, Aarib Zaidi, Rabia Usman, Sonia Okon, Praneeth Medepalli, Kevin Zhu, Vasu Sharma, and Sean O'Brien. 2025. Deconstructing bias: A multifaceted framework for diagnosing cultural and compositional inequities in text-to-image generative models. arXiv preprint arXi...

  82. [90]

    Barry Salt. 1974. Statistical style analysis of motion pictures. Film quarterly, 28(1):13--22

  83. [91]

    Michael Saxon and William Yang Wang. 2023. https://doi.org/10.18653/v1/2023.acl-long.266 Multilingual conceptual coverage in text-to-image models . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4831--4...

  84. [92]

    Florian Schneider, Carolin Holtermann, Chris Biemann, and Anne Lauscher. 2025. https://arxiv.org/abs/2502.13766 Gimmick -- globally inclusive multimodal multitask cultural knowledge benchmarking . Preprint, arXiv:2502.13766

  85. [93]

    Nithish Kannen Senthilkumar, Arif Ahmad, Marco Andreetto, Vinodkumar Prabhakaran, Utsav Prabhu, Adji Bousso Dieng, Pushpak Bhattacharyya, and Shachi Dave. 2024. Beyond aesthetics: Cultural competence in text-to-image models. Advances in Neural Information Processing Systems, 3...

  86. [94]

    Susan Sontag. 1977. On Photography. Farrar, Straus and Giroux

  87. [95]

    Lukas Struppek, Dom Hintersdorf, Felix Friedrich, Manuel br, Patrick Schramowski, and Kristian Kersting. 2024. https://doi.org/10.1613/jair.1.15388 Exploiting cultural biases via homoglyphs in text-to-image synthesis . J. Artif. Int. Res., 78

  88. [96]

    Marita Sturken and Lisa Cartwright. 2001. Practices of looking, volume 2009. Oxford University Press Oxford

  89. [97]

    Christopher Tilley, Susanne Kuechler-Fogden, and Webb Keane. 2005. Handbook of material culture

  90. [98]

    Yuri Tsivian. 2009. Cinemetrics, part of the humanities' cyberstructure. In B. Freisleben, J. Garncarz, and M. Grauer, editors, Digital Tools in Media Studies: Analysis and Research: an Overview, pages 93--100. Transcript Verlag, Bielefeld

  91. [99]

    Norawit Urailertprasert, Peerat Limkonchotiwat, Supasorn Suwajanakorn, and Sarana Nutanong. 2024. https://doi.org/10.18653/v1/2024.alvr-1.15 SEA - VQA : S outheast A sian cultural context dataset for visual question answering . In Proceedings of the 3rd Workshop on Advances in...

  92. [100]

    Quan Van Nguyen, Dan Quang Tran, Huy Quang Pham, Thang Kien-Bao Nguyen, Nghia Hieu Nguyen, Kiet Van Nguyen, and Ngan Luu-Thuy Nguyen. 2024. Vitextvqa: A large-scale visual question answering dataset for evaluating vietnamese text comprehension in images. arXiv preprint arXiv:2...

  93. [101]

    Mor Ventura, Eyal Ben-David, Anna Korhonen, and Roi Reichart. 2025. https://doi.org/10.1162/tacl_a_00732 Navigating cultural chasms: Exploring and unlocking the cultural pov of text-to-image models . Transactions of the Association for Computational Linguistics, 13:142--166

  94. [102]

    Jean-Paul Vinay and Jean Darbelnet. 1995. Comparative stylistics of french and english

  95. [103]

    Caroline Wang and Mary Ann Burris. 1997. Photovoice: Concept, methodology, and use for participatory needs assessment. Health education & behavior, 24(3):369--387

  96. [104]

    Yuxuan Wang, Yijun Liu, Fei Yu, Chen Huang, Kexin Li, Zhiguo Wan, Wanxiang Che, and Hongyang Chen. 2025. Cvlue: A new benchmark dataset for chinese vision-language understanding evaluation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 8196--8204

  97. [105]

    Raymond Williams. 1983. Culture and Society, 1780-1950. Columbia University Press

  98. [106]

    Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Wang Yutong, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Anirban Das, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Chen...

  99. [107]

    Bo Xu, Junzhe Zheng, Jiayuan He, Yuxuan Sun, Hongfei Lin, Liang Zhao, and Feng Xia. 2024. Generating multimodal metaphorical features for meme understanding. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 447--455

  100. [108]

    Srishti Yadav, Zhi Zhang, Daniel Hershcovich, and Ekaterina Shutova. 2025. Beyond words: Exploring cultural value sensitivity in multimodal models. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 7592--7608

  101. [109]

    Andre Ye, Sebastin Santy, Jena D Hwang, Amy X Zhang, and Ranjay Krishna. 2023. Computer vision datasets and models exhibit cultural and linguistic diversity in perception. arXiv preprint arXiv:2310.14356

  102. [110]

    Fulong Ye, Guang Liu, Xinya Wu, and Ledell Wu. 2024. Altdiffusion: A multilingual text-to-image diffusion model. In Proceedings of the AAAI conference on artificial intelligence, volume 38, pages 6648--6656

  103. [111]

    Akhila Yerukola, Saadia Gabriel, Nanyun Peng, and Maarten Sap. 2025. Mind the gesture: Evaluating ai sensitivity to culturally offensive non-verbal gestures. arXiv preprint arXiv:2502.17710

  104. [112]

    Da Yin, Liunian Harold Li, Ziniu Hu, Nanyun Peng, and Kai-Wei Chang. 2021. Broaden the vision: Geo-diverse visual commonsense reasoning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2115--2129

  105. [113]

    Youngsik Yun and Jihie Kim. 2024. https://doi.org/10.24963/ijcai.2024/180 Cic: a framework for culturally-aware image captioning . In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI '24

  106. [114]

    Lili Zhang, Xi Liao, Zaijia Yang, Baihang Gao, Chunjie Wang, Qiuling Yang, and Deshun Li. 2024. https://doi.org/10.1145/3613904.3642877 Partiality and misconception: Investigating cultural representativeness in text-to-image models . In Proceedings of the 2024 CHI Conference o...

  107. [115]

    Na Zhang and Guansheng Ma. 2020. https://doi.org/10.1186/s42779-020-0045-z Nutritional characteristics and health effects of regional cuisines in china . Journal of Ethnic Foods, 7(1):7

  108. [116]

    Kang Zhao, Xinyu Zhao, Zhipeng Jin, Yi Yang, Wen Tao, Cong Han, Shuanglong Li, and Lin Liu. 2024. https://api.semanticscholar.org/CorpusID:271114473 Enhancing baidu multimodal advertisement with chinese text-to-image generation via bilingual alignment and caption synthesis . P...

  109. [117]

    Naitian Zhou, David Bamman, and Isaac L Bleaman. 2025. Culture is not trivia: Sociocultural theory for cultural nlp. arXiv preprint arXiv:2502.12057

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.