REVIEW 4 major objections 6 minor 54 references
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation
T0 review · 4 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Russian culture benchmark exposes text-to-image gap.
desk verdict Useful new Russian T2I benchmark, but the headline model ranking has an unstated prompt-language confound and no statistical support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the RusCode dataset itself: 1,250 bilingual prompts built from 19 categories and 125 subcategories of the Russian cultural code, with a reference image per prompt. The cultural code is defined as the set of concepts a native member of a community regularly encounters and uses for communication, which outsiders may not understand. The prompts are deliberately complex, embedding the target concept in a context (for example, 'art photography, aerial view of the Bolshoi Theatre in Moscow, evening, sunset'), so that a model must recognize the entity and render it with culturally correct detail. The evaluation mechanism is a blind side-by-side human choice task: raters see one image from each of two unnamed models and pick the one that best matches the text prompt. That mechanism converts fuzzy cultural knowledge into a measurable preference ranking.
What would settle it
Re-run the side-by-side evaluation with two groups of raters -- one screened for native-level Russian cultural background and one not -- and compute agreement within and across groups; if the rankings diverge or within-group agreement is low, the reported model ordering would not be a stable measure of cultural awareness. Alternatively, an automatic analysis of the generated images' error types could check whether 'international' substitutes appear at the rates the paper describes.
Extended reading notes
Core claim
The central claim is that current text-to-image models have a measurable cultural-awareness gap for Russian culture, and that the gap can be quantified with a purpose-built benchmark. The paper argues that the RusCode dataset -- 1,250 deliberately complex prompts created by 13 native-speaker prompt engineers under expert guidance, covering art, literature, folklore, food, holidays, science, sites, and everyday inscriptions -- is a valid instrument for this measurement. On a side-by-side human evaluation, generations from Kandinsky 3.1 and YandexART 2 were chosen as more accurate than those from Stable Diffusion 3 and DALL-E 3, which the authors attribute to Russian-language or Russian-user-centered training data. The authors also show that automatic scoring with CLIP cannot substitute for this human judgment, because its scores are high for all models and do not correlate with the human rankings. The intended upshot is that cultural awareness is a distinct quality axis for text-to-image systems, separable from photorealism and general prompt fidelity.
Load-bearing premise
The evaluation's validity rests on the assumption that the 48 human raters are culturally competent judges of Russian concepts; the paper does not report screening for Russian background or native-level Russian, nor does it give inter-annotator agreement.
Editorial extensions
If this is right
- The RusCode dataset can be reused as a bilingual test suite, in Russian and English, for any text-to-image model claiming cultural competence.
- Because the side-by-side results favor models with Russian-oriented training, the benchmark supports the paper's explanation that cultural awareness comes largely from culturally relevant training data.
- Automatic metrics such as CLIP Score should not be trusted for cultural-awareness evaluation; the gap between CLIP and human rankings is itself a finding.
- The included reference images allow evaluations by people unfamiliar with Russian culture, widening the benchmark's use beyond native speakers.
- Treating model censorship or refusal to generate as a failure case means future audits of cultural awareness must report blocked prompts separately.
Reading between the lines
- A direct extension would be to fine-tune a general model on Russian cultural data and re-run the same side-by-side test; the paper's attribution of the gap to training data predicts a measurable improvement.
- The same category-and-reference-image recipe could be applied to other under-represented cultures, yielding comparable benchmarks for a multilingual cultural map of text-to-image models.
- Because the benchmark includes reference images, it can support a scalable screening test in which non-experts judge whether a generated image matches the reference, which could be used to build automatic cultural-awareness metrics.
- The paper's admission that its 19 categories narrow a complex cultural code suggests that results should be read as coverage of common concepts, not as a complete statement about Russian culture.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces RusCode, a benchmark dataset for evaluating the cultural awareness of text-to-image (T2I) generation models with respect to Russian visual culture. The dataset consists of 1250 prompts (each with a Russian version and an English translation) organized into 19 categories and 125 subcategories, with a reference image per prompt. Using this dataset, the authors evaluate four T2I models—Stable Diffusion 3, DALL-E 3, Kandinsky 3.1, and YandexART 2—through a 48-participant human side-by-side study and a CLIP-score analysis. They report that Kandinsky 3.1 and YandexART 2 significantly outperform Stable Diffusion 3 and DALL-E 3, and they argue that automatic metrics like CLIP score do not capture cultural awareness. The dataset is publicly released under the MIT license.
Significance. If the benchmark is valid, it fills a clear gap in multicultural T2I evaluation by providing a Russian-language and Russian-culture resource, complementing existing benchmarks that are mostly English-centric. The dataset construction is described in reasonable detail: prompt engineers from 13 professional backgrounds, expert filtering, and a transparent taxonomy. The human evaluation is a genuine effort with 48 participants and side-by-side comparisons. However, the central comparative claim—that Russian-focused models significantly outperform general models on Russian cultural awareness—is not yet statistically supported and is potentially confounded by the unspecified prompt language. The benchmark itself is a useful contribution, but the paper's headline result requires more rigorous evidence before it can be accepted.
major comments (4)
- [Section 5 (Human evaluation)] The language of the prompts used for each model is never specified. The dataset is bilingual (Section 4.1: 'Each prompt is presented in Russian and has an English translation variant'), and the CLIP-score analysis explicitly uses English embeddings (Table 2). If Kandinsky 3.1 and YandexART 2 were prompted in Russian while Stable Diffusion 3 and DALL-E 3 received English translations, the observed differences could reflect language comprehension rather than cultural awareness. Please state the prompt language for each model in the human evaluation, and either run the comparison in both languages or explicitly analyze the language factor.
- [Section 5, Figure 5] The claim that Kandinsky 3.1 and YandexART 2 'significantly outperform' Stable Diffusion 3 and DALL-E 3 is not supported by any statistical analysis. No confidence intervals, no significance tests (e.g., Wilcoxon signed-rank test), no inter-annotator agreement (e.g., Cohen's kappa), and no information about the number of judgments per comparison pair are reported. Without these, the word 'significantly' is unjustified. Please report per-pair counts, variance, and appropriate significance tests.
- [Section 5 and Appendix A] The exclusion of Midjourney v6 from the main comparison is asymmetric. The paper states that Midjourney could generate only 974 of 1250 prompts due to censorship, so it is not included in the main text; however, Appendix A reports that on those 974 prompts, Kandinsky 3.1 and Midjourney show 'competitive quality.' This suggests that the conclusion about Russian-focused models outperforming general models depends heavily on how censorship is handled. Please either include a full comparison with Midjourney (treating censorship as a separate outcome) or provide a sensitivity analysis showing how the ranking changes when censored prompts are excluded.
- [Section 5, 'Human evaluation' paragraph] The 48 evaluators are described by age and professional fields, but the paper does not report whether they were screened for Russian cultural background or native-level proficiency. Since the task is to judge the correctness of Russian cultural concepts, evaluators who are not culturally competent cannot provide valid labels. Please describe the screening process and consider reporting inter-annotator agreement separately for culturally knowledgeable and non-knowledgeable raters.
minor comments (6)
- [Section 3] In the sentence 'Their collective efforts have resulted identifying of 19 main categories,' 'resulted identifying' should be 'resulted in identifying.'
- [Section 6] In the first sentence, 'proof' appears where 'prove' is intended: 'The generation results of popular T2I models proof the existence.'
- [Table 2] The CLIP scores are reported without any measure of variance across prompts or across seeds. Since the values differ by less than one point, a confidence interval or standard deviation is needed to judge whether these differences are meaningful.
- [Section 4.2] The description of using 'popular queries in search engines related to Russian culture' is vague; please specify the search engines, the query collection method, and whether this introduced a bias toward particular types of cultural concepts.
- [Appendix B] The per-category results are shown only in the appendix and are not analyzed in the main text. A short discussion of which categories show the largest and smallest gaps between models would strengthen the paper's conclusions about where cultural knowledge is lacking.
- [Section 8 (Limitations)] The limitations section does not mention the potential confound between prompt language and cultural knowledge, despite the dataset being bilingual; this should be acknowledged explicitly.
Circularity Check
No circular derivation: benchmark construction and human evaluation are externally grounded; self-citations are explanatory only.
full rationale
The paper's derivation chain is a dataset-construction plus human-evaluation pipeline, not a mathematical derivation. The benchmark prompts (Section 4.1) were created by 13 prompt-engineers and filtered by two professional prompt-engineers, while the human evaluation (Section 5) used 48 raters 'who were not involved in the creation of the dataset.' No parameter is fitted to the evaluation outcomes, and no quantity used as a 'prediction' is defined in terms of the measured outcome. The only self-citations are explanatory: the Discussion attributes Kandinsky 3.1's advantage to Russian-language training data by citing the authors' own technical report (Arkhipkin et al., 2024), but the evaluation result itself comes from independent side-by-side human judgments, and that citation is not used to derive the ranking. The taxonomy of 19 categories is self-defined, but that is a benchmark-construction choice, not a circular reduction; the central comparative claim is empirically grounded in rater choices. The skeptical concern about unspecified prompt language in Section 5 is a potential experimental confound, not a circularity, because it does not make the measured outcome equal to an input by construction. Overall, no circular step can be exhibited from the paper's text, so the circularity score is low; the slight elevation reflects the presence of minor, non-load-bearing self-citations.
Assumptions & free parameters
free parameters (3)
- Prompts per subcategory =
10
- Number of categories and subcategories =
19 categories, 125 subcategories
- Number of human evaluators and pairs per evaluator =
48 raters, about 125 pairs each
assumptions (4)
- ad hoc to paper The Russian cultural code is adequately represented by the 19 categories and 125 subcategories identified by the authors' expert team.
- domain assumption Human side-by-side preference is a valid measure of cultural awareness in T2I generation.
- domain assumption The 48 evaluators are culturally competent judges of Russian cultural concepts.
- domain assumption Reference images from open sources correctly represent the prompt concepts.
Cite this review
Pith. "Pith review of RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation." pith.science (2026). https://pith.science/paper/WUQS2GBQ
@misc{pith2026250207455,
author = {Pith},
title = {Pith review of: RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/WUQS2GBQ}},
note = {Machine review of arXiv:2502.07455}
}
read the original abstract
Text-to-image generation models have gained popularity among users around the world. However, many of these models exhibit a strong bias toward English-speaking cultures, ignoring or misrepresenting the unique characteristics of other language groups, countries, and nationalities. The lack of cultural awareness can reduce the generation quality and lead to undesirable consequences such as unintentional insult, and the spread of prejudice. In contrast to the field of natural language processing, cultural awareness in computer vision has not been explored as extensively. In this paper, we strive to reduce this gap. We propose a RusCode benchmark for evaluating the quality of text-to-image generation containing elements of the Russian cultural code. To do this, we form a list of 19 categories that best represent the features of Russian visual culture. Our final dataset consists of 1250 text prompts in Russian and their translations into English. The prompts cover a wide range of topics, including complex concepts from art, popular culture, folk traditions, famous people's names, natural objects, scientific achievements, etc. We present the results of a human evaluation of the side-by-side comparison of Russian visual concepts representations using popular generative models.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Vladimir Arkhipkin, Andrei Filatov, Viacheslav Vasilev, Anastasia Maltseva, Said Azizov, Igor Pavlov, Julia Agafonova, Andrey Kuznetsov, and Denis Dimitrov. 2024. https://arxiv.org/abs/2312.03511 Kandinsky 3.0 technical report . Preprint, arXiv:2312.03511
arXiv 2024
-
[4]
Vladimir Arkhipkin, Zein Shaheen, Viacheslav Vasilev, Elizaveta Dakhova, Andrey Kuznetsov, and Denis Dimitrov. 2023. https://arxiv.org/abs/2311.13073 Fusionframes: Efficient architectural aspects for text-to-video generation pipeline . Preprint, arXiv:2311.13073
work page Pith review arXiv 2023
-
[5]
Vladimir Arkhipkin, Zein Shaheen, Viacheslav Vasilev, Elizaveta Dakhova, Konstantin Sobolev, Andrey Kuznetsov, and Denis Dimitrov. 2025. https://doi.org/10.1109/ACCESS.2024.3522510 Improveyourvideos: Architectural improvements for text-to-video generation pipeline . IEEE Access, 13:1986--2003
- [6]
-
[7]
Venkatesh Babu, and Danish Pruthi
Abhipsa Basu, R. Venkatesh Babu, and Danish Pruthi. 2023. https://doi.org/10.1109/ICCV51070.2023.00474 Inspecting the geographical representativeness of images from text-to-image models . In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 5113--5124
-
[8]
Federico Becattini, Pietro Bongini, Luana Bulla, Alberto Del Bimbo, Ludovica Marinucci, Misael Mongiov\` , and Valentina Presutti. 2023. https://doi.org/10.1145/3590773 Viscounth: A large-scale multilingual visual question answering dataset for cultural heritage . ACM Trans. Multimedia Comput. Commun. Appl., 19(6)
doi:10.1145/3590773 2023
Show all 54 references
-
[9]
Hugo Berg, Siobhan Hall, Yash Bhalgat, Hannah Kirk, Aleksandar Shtedritski, and Max Bain. 2022. https://aclanthology.org/2022.aacl-main.61 A prompt array keeps the bias away: Debiasing vision-language models with adversarial learning . In Proceedings of the 2nd Conference of t...
2022
-
[10]
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, Wesam Manassra, Prafulla Dhariwa, Casey Chu, Yunxin Jiao, and Aditya Ramesh. 2023. Improving image generation with better captions
2023
-
[11]
Mehar Bhatia, Sahithya Ravi, Aditya Chinchure, Eunjeong Hwang, and Vered Shwartz. 2024. https://arxiv.org/abs/2407.00263 From local concepts to universals: Evaluating the multicultural understanding of vision-language models . Preprint, arXiv:2407.00263
2024 arXiv
-
[12]
James Billington. 2010. The icon and axe: An interpretative history of Russian culture. Vintage
2010
-
[14]
Abeba Birhane, Sepehr Dehdashtian, Vinay Prabhu, and Vishnu Boddeti. 2024. https://doi.org/10.1145/3630106.3658968 The dark side of dataset scaling: Evaluating racial classification in multimodal models . In Proceedings of the 2024 ACM Conference on Fairness, Accountability, a...
2024
-
[15]
Emanuele Bugliarello, Fangyu Liu, Jonas Pfeiffer, Siva Reddy, Desmond Elliott, Edoardo Maria Ponti, and Ivan Vuli \'c . 2022. https://proceedings.mlr.press/v162/bugliarello22a.html IGLUE : A benchmark for transfer learning across modalities, tasks, and languages . In Proceedin...
2022
-
[16]
Samuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio, Xiaohong Li, Adhiguna Kuncoro, Sebastian Ruder, Zhi Yuan Lim, Syafri Bahar, Masayu Khodra, Ayu Purwarianti, and Pascale Fung. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.699 I ndo NLG : Benchmark and...
2021 doi
-
[17]
Yong Cao, Min Chen, and Daniel Hershcovich. 2024 a . https://aclanthology.org/2024.findings-eacl.63 Bridging cultural nuances in dialogue agents through cultural value surveys . In Findings of the Association for Computational Linguistics: EACL 2024, pages 929--945, St. Julian...
2024
-
[18]
Yong Cao, Yova Kementchedjhieva, Ruixiang Cui, Antonia Karamolegkou, Li Zhou, Megan Dare, Lucia Donatelli, and Daniel Hershcovich. 2024 b . https://doi.org/10.1162/tacl_a_00634 Cultural Adaptation of Recipes . Transactions of the Association for Computational Linguistics, 12:80--99
2024 doi
-
[19]
Jaemin Cho, Abhay Zala, and Mohit Bansal. 2023. Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models. In ICCV
2023
-
[20]
Colton Clemmer, Junhua Ding, and Yunhe Feng. 2024. Precisedebias: An automatic prompt engineering approach for generative ai to mitigate image demographic biases. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 8596--8605
2024
-
[21]
John Corner. 1980. https://doi.org/10.1177/016344378000200107 Codes and cultural analysis . Media, Culture & Society, 2(1)
1980 doi
-
[22]
Nassim Dehouche. 2021. https://doi.org/10.1109/ACCESS.2021.3136898 Implicit stereotypes in pre-trained classifiers . IEEE Access, 9:167936--167947
2021
-
[23]
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach. 2024. https://arxiv.org/abs/2...
2024 arXiv
-
[24]
Alena Fenogenova, Artem Chervyakov, Nikita Martynov, Anastasia Kozlova, Maria Tikhonova, Albina Akhmetgareeva, Anton Emelyanov, Denis Shevelev, Pavel Lebedev, Leonid Sinev, Ulyana Isaeva, Katerina Kolomeytseva, Daniil Moskovskiy, Elizaveta Goncharova, Nikita Savushkin, Polina ...
2024 doi
-
[25]
Orlando Figes. 2002. Natasha's dance: A cultural history of Russia. Macmillan
2002
-
[26]
Sourojit Ghosh, Pranav Narayanan Venkit, Sanjana Gautam, Shomir Wilson, and Aylin Caliskan. 2024. Do generative ai models output harm while representing non-western cultures: Evidence from a community-centered approach. arXiv preprint arXiv:2407.14779
2024 arXiv
-
[27]
Mikhail Goloubkov. 2013. https://doi.org/10.1016/j.euras.2012.02.001 Literature and the russian cultural code at the beginning of the 21st century . Journal of Eurasian Studies, 4(1):107--113. 20 Years of the Collapse of the Fomer Soviet Union
2013 doi
-
[28]
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders S gaard. 2022. https://doi.org/10.186...
2022 doi
-
[29]
Yuichi Inoue, Kento Sasaki, Yuma Ochi, Kazuki Fujii, Kotaro Tanahashi, and Yu Yamaguchi. 2024. https://arxiv.org/abs/2404.07824 Heron-bench: A benchmark for evaluating vision language models in japanese . Preprint, arXiv:2404.07824
2024 arXiv
-
[30]
Nithish Kannen, Arif Ahmad, Marco Andreetto, Vinodkumar Prabhakaran, Utsav Prabhu, Adji Bousso Dieng, Pushpak Bhattacharyya, and Shachi Dave. 2024. https://arxiv.org/abs/2407.06863 Beyond aesthetics: Cultural competence in text-to-image models . Preprint, arXiv:2407.06863
2024 arXiv
-
[31]
Sergey Kastryulin, Artem Konev, Alexander Shishenya, Eugene Lyapustin, Artem Khurshudov, Alexander Tselousov, Nikita Vinokurov, Denis Kuznedelev, Alexander Markovich, Grigoriy Livshits, Alexey Kirillov, Anastasiia Tabisheva, Liubov Chubarova, Marina Kaminskaia, Alexander Ustyu...
2024 arXiv
-
[32]
Eunsu Kim, Juyoung Suk, Philhoon Oh, Haneul Yoo, James Thorne, and Alice Oh. 2024. https://aclanthology.org/2024.lrec-main.296 CLI c K : A benchmark dataset of cultural and linguistic intelligence in K orean . In Proceedings of the 2024 Joint International Conference on Comput...
2024
-
[33]
u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt\
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K\" u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt\" a schel, Sebastian Riedel, and Douwe Kiela. 2020. https://proceedings.neurips.cc/paper_files/paper/2020/file/6b493230205f78...
2020
-
[34]
Yulong Liu, Guibo Zhu, Bin Zhu, Qi Song, Guojing Ge, Haoran Chen, GuanHui Qiao, Ru Peng, Lingxiang Wu, and Jinqiao Wang. 2022. https://proceedings.neurips.cc/paper_files/paper/2022/file/6a386d703b50f1cf1f61ab02a15967bb-Paper-Datasets_and_Benchmarks.pdf Taisu: A 166m large-scal...
2022
-
[35]
Alexandra Sasha Luccioni, Christopher Akiki, Margaret Mitchell, and Yacine Jernite. 2024. Stable bias: evaluating societal representations in diffusion models. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS '23, Red Hook, NY,...
2024
-
[36]
Midjourney. 2022. Midjourney. https://www.midjourney.com/
2022
-
[37]
Nusrat Jahan Mim, Dipannita Nandi, Sadaf Sumyia Khan, Arundhuti Dey, and Syed Ishtiaque Ahmed. 2024. https://doi.org/10.1145/3613904.3641951 In-between visuals and visible: The impacts of text-to-image generative ai tools on digital image-making practices in the global south ....
2024
-
[38]
Ojha, Akanksha Bansal, Deepak Alok, John P
Sourabrata Mukherjee, Atul Kr. Ojha, Akanksha Bansal, Deepak Alok, John P. McCrae, and Ondrej Dusek. 2024. https://aclanthology.org/2024.inlg-main.41 Multilingual text style transfer: Datasets & models for I ndian languages . In Proceedings of the 17th International Natural La...
2024
-
[39]
Ranjita Naik and Besmira Nushi. 2023. https://doi.org/10.1145/3600211.3604711 Social biases through the text-to-image generation lens . In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, AIES '23, page 786–808, New York, NY, USA. Association for Computi...
2023
-
[40]
Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy, Sjoerd van Steenkiste, Lisa Anne Hendricks, Karolina Stańczak, and Aishwarya Agrawal. 2024. https://arxiv.org/abs/2407.10920 Benchmarking vision language models for cultural understanding . Preprint, arXiv:2407.10920
2024 arXiv
-
[41]
Denis Peskov, Viktor Hangya, Jordan Boyd-Graber, and Alexander Fraser. 2021. https://doi.org/10.18653/v1/2021.findings-emnlp.315 Adapting entities across languages and cultures . In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3725--3750, Punta ...
2021 doi
-
[42]
Bennett, and Emily Denton
Rida Qadri, Renee Shelby, Cynthia L. Bennett, and Emily Denton. 2023. https://doi.org/10.1145/3593013.3594016 Ai’s regimes of representation: A community-centered study of text-to-image models in south asia . In Proceedings of the 2023 ACM Conference on Fairness, Accountabilit...
2023
-
[43]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. https://arxiv.org/abs/2103.00020 Learning transferable visual models from natural lan...
2021 arXiv
-
[44]
David Romero, Chenyang Lyu, Haryo Akbarianto Wibowo, Teresa Lynn, Injy Hamed, Aditya Nanda Kishore, Aishik Mandal, Alina Dragonetti, Artem Abzaliev, Atnafu Lambebo Tonja, Bontu Fufa Balcha, Chenxi Whitehouse, Christian Salamea, Dan John Velasco, David Ifeoluwa Adelani, David L...
2024 arXiv
-
[45]
Richard Stites. 1992. Russian popular culture: Entertainment and society since 1900. Cambridge University Press
1992
-
[46]
Lukas Struppek, Dom Hintersdorf, Felix Friedrich, Manuel br, Patrick Schramowski, and Kristian Kersting. 2024. https://doi.org/10.1613/jair.1.15388 Exploiting cultural biases via homoglyphs in text-to-image synthesis . J. Artif. Int. Res., 78
2024 doi
-
[47]
Ekaterina Taktasheva, Maxim Bazhukov, Kirill Koncha, Alena Fenogenova, Ekaterina Artemova, and Vladislav Mikhailov. 2024. https://arxiv.org/abs/2406.19232 Rublimp: Russian benchmark of linguistic minimal pairs . Preprint, arXiv:2406.19232
2024 arXiv
-
[48]
Mor Ventura, Eyal Ben-David, Anna Korhonen, and Roi Reichart. 2023. Navigating cultural chasms: Exploring and unlocking the cultural pov of text-to-image models. arXiv preprint arXiv:2310.01929
2023 arXiv
-
[49]
Arkhipkin Vladimir, Viacheslav Vasilev, Andrei Filatov, Igor Pavlov, Julia Agafonova, Nikolai Gerasimenko, Anna Averchenkova, Evelina Mironova, Bukashkin Anton, Konstantin Kulikov, Andrey Kuznetsov, and Denis Dimitrov. 2024. https://doi.org/10.18653/v1/2024.emnlp-demo.48 Kandi...
2024 doi
-
[50]
Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, and William Isaac. 2023. https://arxiv.org/abs/2310.11986 Sociotechnical safety evalu...
2023 arXiv
-
[51]
Haryo Akbarianto Wibowo, Erland Hilman Fuadi, Made Nindyatama Nityasya, Radityo Eko Prasojo, and Alham Fikri Aji. 2023. Copal-id: Indonesian language reasoning with local culture and nuances. arXiv preprint arXiv:2311.01012
2023 arXiv
-
[52]
Anna Wierzbicka. 2002. https://doi.org/10.1525/eth.2002.30.4.401 Russian cultural scripts: The theory of cultural scripts and its applications . Ethos, 30(4):401--432
2002 doi
-
[53]
Yue Yu, Yuchen Zhuang, Jieyu Zhang, Yu Meng, Alexander J Ratner, Ranjay Krishna, Jiaming Shen, and Chao Zhang. 2023. https://proceedings.neurips.cc/paper_files/paper/2023/file/ae9500c4f5607caf2eff033c67daa9d7-Paper-Datasets_and_Benchmarks.pdf Large language model as attributed...
2023
- [54]
-
[55]
Li Zhou, Antonia Karamolegkou, Wenyu Chen, and Daniel Hershcovich. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.845 Cultural compass: Predicting transfer learning success in offensive language detection with cultural features . In Findings of the Association for Compu...
2023 doi
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.