REVIEW 4 major objections 4 minor 1 cited by
SESGO: Spanish Evaluation of Stereotypical Generative Outputs
T0 review · 4 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A 4,156-prompt Spanish bias evaluation finds that culturally specific Latin American stereotypes expose bias that translated English benchmarks miss.
desk verdict A useful, reusable Spanish bias benchmark with a sensible metric; the dataset labels carry the weight, and the cross-lingual 'mitigation' claim outruns the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the SESGO dataset together with a signed bias score. SESGO builds scenarios from documented Latin American sayings, such as “A Black man without a master is like a child without a father,” across gender, race, class, and xenophobia, and each prompt appears in ambiguous and disambiguated forms with positive and negative questions and Target, Other, and Unknown answer options. The bias score is the Euclidean distance between the ideal point, where accuracy is 1 and $F(\mathrm{Target}) = F(\mathrm{Other})$, and the observed point, with a sign indicating which group bears the detected bias: bias score $= \sigma \sqrt{(1-\mathrm{acc})^{2} + (F(\mathrm{Target})-F(\mathrm{Other}))^{2}}$. This metric is what lets the paper separate raw accuracy from the direction of model errors.
What would settle it
Take a random subset of SESGO prompts and have independent raters from several Latin American countries assign Target, Other, Unknown, and the correct disambiguated answer without seeing model results, then recompute all bias scores with those labels. If the direction or magnitude of model rankings changes substantially under alternative ground truth, the reported bias patterns depend on the authors' stereotype readings rather than on stable model behavior.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that culturally situated Spanish prompts reveal bias patterns that translated benchmarks obscure. Using ambiguous and disambiguated questions adapted from the BBQ methodology but grounded in Latin American popular sayings, the authors find that several leading models answer incorrectly more often in Spanish and, when they err, target historically marginalized groups more strongly. They also find that bias scores rise on full culturally grounded prompts compared with direct translations, and that temperature changes do not materially shift bias. The paper presents this as the first systematic evaluation of commercial LLMs on culturally specific bias in Spanish.
Load-bearing premise
The load-bearing premise is that every SESGO ground-truth label—who is the target of a saying, what the correct disambiguated answer is, and when Unknown is right—accurately encodes a genuinely held Latin American stereotype; the authors note these labels were not validated through user studies or expert review.
Editorial extensions
If this is right
- A model that appears balanced on translated English prompts can still show systematic bias on culturally grounded Spanish prompts.
- English-language safety training cannot be assumed to protect Spanish-speaking users; mitigation must be validated on native-language, culturally specific data.
- Because temperature does not change bias scores, decoding-time adjustments are not a practical lever; the levers are upstream training and fine-tuning.
- The modular prompt structure can be extended to new stereotypes, bias categories, and languages, giving other regions a ready-made template for culturally aware evaluation.
- Benchmarks that claim multilingual coverage by translation should be read as measuring linguistic equivalence, not cultural safety.
Reading between the lines
- Beyond the paper: if the dataset's labels encode Colombian-centric readings of sayings, bias scores could shift when the same items are rated by people in Mexico, Argentina, or Spain; a multi-country label validation would test how much of the reported effect is regional.
- Beyond the paper: the findings predict that harmful chatbot outputs in Spanish will concentrate in narratives tied to local out-groups, such as Venezuelan migrants in Colombia, rather than in generic translated stereotypes.
- Beyond the paper: the same accuracy-plus-error-direction metric could be applied to culturally grounded prompt sets in other regional or Indigenous languages to map where English-centric mitigation fails most.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SESGO, a Spanish-language benchmark and evaluation framework for measuring stereotypical bias in instruction-tuned LLMs. The dataset contains 4,156 prompts built from Latin American sayings and cultural expressions, organized into four bias categories (racism, gender, classism, xenophobia), each with ambiguous and disambiguated versions and Target/Other/Unknown answer options. The authors define a bias score (Eq. 1) that combines accuracy with the difference between two error-direction metrics, evaluate six commercial or openly available models, and report (i) that bias is higher on ambiguous Spanish prompts, (ii) that matched Spanish/English comparisons reveal cross-lingual differences including direction reversals, and (iii) that bias scores are largely stable across sampling temperatures. The code and dataset are publicly released.
Significance. If the gold labels are culturally valid, SESGO is a valuable resource: it moves beyond translated English benchmarks, documents region-specific stereotypes, and offers a simple interpretable metric that separates accuracy from error direction. The matched English/Spanish comparison is a useful design for isolating language effects, and the public code and data release support reproducibility. However, the headline claims are not yet established: the correctness of every score depends on unvalidated labels, the cross-lingual and temperature comparisons lack uncertainty quantification, and the claimed failure of English-optimized mitigation is not directly tested. With label validation and more careful statistical reporting, the contribution would be solid.
major comments (4)
- [SESGO Dataset and Discussion (limitations)] All accuracy and bias scores in Tables 2-3 and Figures 3-7 inherit the gold labels Target/Other/Unknown, but no validation of these labels is reported: there is no inter-annotator agreement, expert review, or pilot study, and the annotation protocol is not described. The authors themselves state in the Discussion that the prompts 'could be refined through user studies or expert reviews ... to validate the representativeness of the dataset and bias categories.' This is load-bearing rather than a routine limitation: if a nontrivial fraction of ambiguous prompts actually contains a culturally recognizable stereotype that makes a demographic answer inferable, then treating Unknown as the gold answer would penalize culturally competent responses and inflate bias scores for models that choose that answer. Please add validation of the labels, or at minimum report agreement and a sensitivity analysis showing that the main cross-lingual and temperature conclusions are robust to plausible label noise.
- [Cross-Linguistic Transferability of Bias Mitigation] The conclusion and abstract claim that 'bias mitigation techniques optimized for English fail to transfer effectively to Spanish contexts,' but the paper does not apply or ablate any mitigation technique. Figures 5-6 and Tables A1-A2 compare model behavior on English versus Spanish matched prompts; this is an observational language comparison, not a test of mitigation transfer. The claim should be reframed as cross-lingual bias differences, or the experiment should actually include an English-optimized mitigation intervention and measure its effect on the Spanish prompts.
- [Bias Presence Across Sampling Temperatures, Table A3] The claim that bias scores 'remain remarkably stable across temperature values from 0.1 to 1' is not supported by the reported point estimates. In Table A3, DeepSeek R1's disambiguated bias score moves from -0.243 (T=0.1) to +0.261 (T=0.5) and 0.000 (T=1.0), including sign reversals; other models also show small but non-negligible variation. Moreover, no confidence intervals, standard errors, or significance tests are reported for any temperature comparison, and the models are stochastic, so repeated sampling is needed. Please add uncertainty quantification and soften or revise the stability claim accordingly.
- [Metrics for Bias Quantification, Eq. (1)] Equation (1) is not fully specified. The text describes the ideal model as the point (acc=1, F(Target)=F(Other)), but that condition is a line, not a point; the formula sqrt((1-acc)^2 + (F(Target)-F(Other))^2) is not the Euclidean distance to (1,0,0) in the (F(Target), F(Other)) plane, and the sign sigma is introduced but never defined in terms of the response data. Since every reported bias score in Tables 2-3 and Figures 3-7 uses this metric, please provide an explicit definition of sigma (e.g., sign(F(Target)-F(Other)) with a stated convention) and explain the intended geometry, including how the equality F(Target)+F(Other)=1-acc is used.
minor comments (4)
- [Xenophobia and Racism subsections] There are typos in the text: 'an LLMs risk perpetuating these biases' should be 'LLMs risk perpetuating these biases,' and 'whit indigenous and black communities' should be 'with indigenous and black communities.'
- [Tables 2 and 3] The caption 'Best value for each bias category (in absolute value) in bold' is ambiguous because lower bias score is better but the tables also report direction; please state precisely what is bolded and why a nonzero signed score can be 'best.'
- [Appendix A3] The appendix table reports a column labeled 'Ft-Fo' without defining it in the caption; please define F(Target), F(Other), and Ft-Fo in the appendix or refer explicitly to the main-text definitions.
- [Discussion] The Discussion states that 'bias scores consistently increase when models operate in ambiguous contexts in Spanish compared to English,' but the Cross-Linguistic section reports that three out of six models show this pattern in ambiguous contexts; please correct this inconsistency so the Discussion matches the evidence.
Circularity Check
No circularity: SESGO's bias metric is an explicit operational definition and the model conclusions are empirical measurements, not derived from the benchmark's labels by construction.
full rationale
The paper's derivation chain is self-contained and empirical. The SESGO dataset is constructed from documented Latin American sayings and expressions, with labels (Target/Other/Unknown) assigned by the authors; this is the standard benchmark-construction step, and the authors explicitly acknowledge the need for external validation in the Discussion ('these could be refined through user studies or expert reviews... to validate the representativeness of the dataset and bias categories'). The bias metric in Equation (1) is an explicit definition combining accuracy with the difference between error directions; it is not fitted to any outcome and no parameter is tuned to produce the reported conclusions. The cross-lingual and temperature results are direct measurements relative to that metric, not consequences of the metric's definition. The paper builds on the external BBQ methodology (Parrish et al. 2022) and external El Barómetro narratives; there is no load-bearing self-citation or imported uniqueness theorem. The central claims about Spanish bias patterns, English-to-Spanish transfer failure, and temperature stability are empirical findings that could in principle have come out differently, so they are not forced by construction. The unvalidated gold labels are a validity/representativeness limitation, not a circularity, and the paper itself flags this limitation.
Assumptions & free parameters
assumptions (3)
- domain assumption In ambiguous prompts, the Unknown option is the correct answer because information is insufficient.
- domain assumption The Target and Other labels and the stereotype direction encoded in each prompt correctly reflect documented Latin American stereotypes.
- domain assumption Forcing models to choose among three fixed options via a system prompt is a valid proxy for bias in generative outputs.
Cite this review
Pith. "Pith review of SESGO: Spanish Evaluation of Stereotypical Generative Outputs." pith.science (2026). https://pith.science/paper/BJRMOFF2
@misc{pith2026250903329,
author = {Pith},
title = {Pith review of: SESGO: Spanish Evaluation of Stereotypical Generative Outputs},
year = {2026},
howpublished = {\url{https://pith.science/paper/BJRMOFF2}},
note = {Machine review of arXiv:2509.03329}
}
read the original abstract
This paper addresses the critical gap in evaluating bias in multilingual Large Language Models (LLMs), with a specific focus on Spanish language within culturally-aware Latin American contexts. Despite widespread global deployment, current evaluations remain predominantly US-English-centric, leaving potential harms in other linguistic and cultural contexts largely underexamined. We introduce a novel, culturally-grounded framework for detecting social biases in instruction-tuned LLMs. Our approach adapts the underspecified question methodology from the BBQ dataset by incorporating culturally-specific expressions and sayings that encode regional stereotypes across four social categories: gender, race, socioeconomic class, and national origin. Using more than 4,000 prompts, we propose a new metric that combines accuracy with the direction of error to effectively balance model performance and bias alignment in both ambiguous and disambiguated contexts. To our knowledge, our work presents the first systematic evaluation examining how leading commercial LLMs respond to culturally specific bias in the Spanish language, revealing varying patterns of bias manifestation across state-of-the-art models. We also contribute evidence that bias mitigation techniques optimized for English do not effectively transfer to Spanish tasks, and that bias patterns remain largely consistent across different sampling temperatures. Our modular framework offers a natural extension to new stereotypes, bias categories, or languages and cultural contexts, representing a significant step toward more equitable and culturally-aware evaluation of AI systems in the diverse linguistic environments where they operate.
Figures
Forward citations
Cited by 1 Pith paper
-
Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders
A pilot probe finds weak, non-robust evidence that Qwen2.5-7B internally represents Colombian identity from a single implicit cue; the only nominally significant effect is driven by unrestricted, confabulated national...
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Abid, A.; Farooqi, M.; and Zou, J. 2021. Persistent anti-muslim bias in large language models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, 298--306
work page 2021
-
[4]
L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774
arXiv 2023
-
[5]
Acosta, D.; Blouin, C.; and Freier, L. F. 2019. La emigraci \'o n venezolana: respuestas latinoamericanas. Documento de Trabajo, no. 3 (2a época), Madrid, Fundación Carolina
work page 2019
-
[6]
Adelman, J., ed. 1999. Colonial Legacies: The Problem of Persistence in Latin American History. Routledge, 1st edition
work page 1999
-
[7]
Amnesty International . 2022. ' F ear and X enophobia': V enezuelan people in P eru facing discrimination and exclusion. Original Spanish document. Report number: AMR 46/5238/2022
work page 2022
-
[8]
Anthropic . 2025. Supported Countries and Regions . https://www.anthropic.com/claude-ai-locations. Accessed: 2025-04-23
work page 2025
Show all 62 references
-
[9]
Anthropic, A. 2024. The claude 3 model family: Opus, sonnet, haiku. https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf. Accessed: 2025-04-23
2024
-
[10]
Aponte de Torres, S. 1995. Las Guajibiadas / Silvia Aponte. Corpes-Orinoquía-Libros. Villavicencio: Imprenta Departamental, 2a. ed. edition
1995
-
[11]
Aristizabal, A.; and Garcias, H. 2020. Refranes, dichos, ag \"u eros y creencias populares. Fondo de publicaciones del Valle del Cauca
2020
-
[12]
M.; Gebru, T.; McMillan-Major, A.; and Shmitchell, S
Bender, E. M.; Gebru, T.; McMillan-Major, A.; and Shmitchell, S. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, 610--623
2021
-
[13]
L.; Barocas, S.; Daum \'e III, H.; and Wallach, H
Blodgett, S. L.; Barocas, S.; Daum \'e III, H.; and Wallach, H. 2020. Language (Technology) is Power: A Critical Survey of “Bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5454--5476
2020
-
[14]
Y.; Saligrama, V.; and Kalai, A
Bolukbasi, T.; Chang, K.-W.; Zou, J. Y.; Saligrama, V.; and Kalai, A. T. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. Advances in neural information processing systems, 29
2016
-
[15]
J.; and Narayanan, A
Caliskan, A.; Bryson, J. J.; and Narayanan, A. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334): 183--186
2017
-
[16]
Castellanos Guerrero, A.; and Land \'a zury Ben \' tez, G. 2012. Racismos y otras formas de intolerancia de Norte a Sur en Am \'e rica Latina . Universidad Aut \'o noma Metropolitana
2012
-
[17]
Chavez, L. 2020. The Latino threat: Constructing immigrants, citizens, and the nation. Stanford University Press
2020
-
[18]
Corrales, J. 2015. The politics of LGBT rights in Latin America and the Caribbean: Research agendas. European Review of Latin American and Caribbean Studies/Revista Europea de Estudios Latinoamericanos y del Caribe, 53--62
2015
-
[19]
de Souza, N. M. F.; and Rodrigues, L. 2022. Gender violence and feminist resistance in Latin America. International Feminist Journal of Politics, 24(1): 5--15
2022
-
[20]
El Barómetro . 2024. Plataforma de monitoreo de narrativas xenófobas en redes sociales en América Latina. https://barometro.org/. Accessed: 2024-06-15
2024
-
[21]
Evstafev, E. 2025. The Paradox of Stochasticity: Limited Creativity and Computational Decoupling in Temperature-Varied LLM Outputs of Structured Fictional Data. arXiv preprint arXiv:2502.08515
2025 arXiv
-
[22]
Fairclough, N. 2013. Language and power. Routledge
2013
-
[23]
H.; Lipton, Z
Feffer, M.; Sinha, A.; Deng, W. H.; Lipton, Z. C.; and Heidari, H. 2024. Red-Teaming for generative AI: Silver bullet or security theater? In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, 421--437
2024
-
[24]
O.; Rossi, R
Gallegos, I. O.; Rossi, R. A.; Barrow, J.; Tanjim, M. M.; Kim, S.; Dernoncourt, F.; Yu, T.; Zhang, R.; and Ahmed, N. K. 2024. Bias and fairness in large language models: A survey. Computational Linguistics, 50(3): 1097--1179
2024
-
[25]
Garrido-Muñoz, I.; Martínez-Santiago, F.; and Montejo-Ráez, A. 2024. MarIA and BETO are sexist: evaluating gender bias in large language models for Spanish. Language Resources and Evaluation, 58: 1387--1417
2024
-
[26]
Gasparini, L.; and Lustig, N. 2011. The Rise and Fall of Income Inequality in Latin America. Technical report, CEDLAS, Universidad Nacional de La Plata
2011
-
[27]
Gasparini, L.; and Marchionni, M. 2015. Bridging Gender Gaps? The Rise and Deceleration of Female Labor Force Participation in Latin America. CEDLAS Working Paper No. 185
2015
-
[28]
Gaviria, A.; Medina, C.; and Palau, M. d. M. 2010. Las consecuencias econ \'o micas de un nombre at \' pico. El caso colombiano. El trimestre econ \'o mico , 77(307): 535--556
2010
-
[29]
Gemini Team Google . 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805
2023 arXiv
-
[30]
Google. 2025 a . Gemini 2.0 Flash. https://cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/2-0-flash3. Accessed: 2025-05-10
2025
-
[31]
Google. 2025 b . Gemini Supported Countries & Territories. https://support.google.com/gemini/answer/13575153. Accessed: 2025-04-23
2025
-
[32]
Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P.; Bi, X.; et al. 2025. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv preprint arXiv:2501.12948
2025 arXiv
-
[33]
Hoffman, K.; and Centeno, M. A. 2003. The lopsided continent: inequality in Latin America. Annual Review of Sociology, 29(1): 363--390
2003
-
[34]
Huang, Y.; and Xiong, D. 2024. CBBQ: A Chinese Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), 2917--2929
2024
-
[35]
Inter-Agency Coordination Platform for Refugees and Migrants from Venezuela . 2025. Refugees and migrants from Venezuela. Inter-Agency Coordination Platform
2025
-
[36]
Jaramillo-Echeverri, J.; and \'A lvarez, A. 2023. The Persistence of Segregation in Education: Evidence from Historical Elites and Ethnic Surnames in Colombia. Technical Report 58, Banco de la Rep \'u blica de Colombia
2023
-
[37]
Jin, J.; Kim, J.; Lee, N.; Yoo, H.; Oh, A.; and Lee, H. 2024. KoBBQ: Korean bias benchmark for question answering. Transactions of the Association for Computational Linguistics, 12: 507--524
2024
-
[38]
Kessler, G.; and Dimarco, S. 2013. J \'o venes, polic \' a y estigmatizaci \'o n territorial en la periferia de Buenos Aires. Espacio abierto, 22(2): 221--243
2013
-
[39]
Lucy, J. A. 1996. The Scope of Linguistic Relativity: an Analysis and Review of Empirical Research. In Gumperz, J. J.; and Levinson, S. C., eds., Rethinking Linguistic Relativity, 37--69. Cambridge: Cambridge University Press
1996
-
[40]
Martinkov \'a , S.; Sta \'n czak, K.; and Augenstein, I. 2023. Measuring Gender Bias in West Slavic Language Models. In Proceedings of the 9th Workshop on Slavic Natural Language Processing 2023 (SlavicNLP 2023), 146--154
2023
-
[41]
J.; and Barrera-Mellado, I
Medina-Hern \'a ndez, E.; Fern \'a ndez-G \'o mez, M. J.; and Barrera-Mellado, I. 2021. Gender inequality in Latin America: A multidimensional analysis based on ECLAC indicators. Sustainability, 13(23): 13140
2021
-
[42]
Meta. 2024. The Llama 3 Herd of Models . arXiv e-prints, arXiv:2407.21783
2024 arXiv
-
[43]
Neplenbroek, V.; Bisazza, A.; and Fern \'a ndez, R. 2024. MBBQ: A Dataset for Cross-Lingual Comparison of Stereotypes in Generative LLMs. In First Conference on Language Modeling
2024
-
[44]
N \'e v \'e ol, A.; Dupont, Y.; Bezan c on, J.; and Fort, K. 2022. French CrowS-pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English. In Proceedings of the 60th Annual Meeting of the Association for Computati...
2022
-
[45]
OpenAI. 2024 a . GPT-4o System Card. arXiv:2410.21276
2024 arXiv
-
[46]
OpenAI. 2024 b . Multilingual Massive Multitask Language Understanding (MMMLU). https://huggingface.co/datasets/openai/MMMLU. Accessed: 2025-04-23
2024
-
[47]
OpenAI. 2025. ChatGPT Supported Countries. https://help.openai.com/en/articles/7947663-chatgpt-supported-countries. Accessed: 2025-04-23
2025
-
[48]
Orenguteng. 2024. Llama-3.1-8B-Lexi-Uncensored-V2. https://huggingface.co/Orenguteng/Llama-3.1-8B-Lexi-Uncensored-V2. Accessed: 2025-04-14
2024
-
[49]
M.; and Bowman, S
Parrish, A.; Chen, A.; Nangia, N.; Padmakumar, V.; Phang, J.; Thompson, J.; Htut, P. M.; and Bowman, S. 2022. BBQ : A hand-built bias benchmark for question answering. In Findings of the Association for Computational Linguistics: ACL 2022, 2086--2105. Dublin, Ireland: Associat...
2022
-
[50]
Portes, A.; and Hoffman, K. 2003. Latin American class structures: Their composition and change during the neoliberal era. Latin American research review, 38(1): 41--82
2003
-
[51]
Reimers, F.; et al. 2000. Unequal schools, unequal chances: The challenges to equal opportunity in the Americas, volume 5. Harvard University Press Cambridge, MA
2000
-
[52]
Renze, M. 2024. The effect of sampling temperature on problem solving in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, 7346--7356
2024
-
[53]
Salgado, M.; and Castillo, J. 2023. Inequality and Stratification in Latin America. In The Oxford Handbook of Social Stratification. Oxford University Press. ISBN 9780197539484
2023
-
[54]
Saralegi, X.; and Zulaika, M. 2025. BasqBBQ: A QA Benchmark for Assessing Social Biases in LLMs for Basque, a Low-Resource Language. In Proceedings of the 31st International Conference on Computational Linguistics, 4753--4767
2025
-
[55]
W.; Tay, Y.; Ruder, S.; Zhou, D.; Das, D.; and Wei, J
Shi, F.; Suzgun, M.; Freitag, M.; Wang, X.; Srivats, S.; Vosoughi, S.; Chung, H. W.; Tay, Y.; Ruder, S.; Zhou, D.; Das, D.; and Wei, J. 2023. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Learning Representations
2023
-
[56]
Stanford University . 2025. The 2025 AI Index Report
2025
-
[57]
Telles, E. 2014. Pigmentocracies: Ethnicity, Race, and Color in Latin America. University of North Carolina Press. ISBN 9781469617831
2014
-
[58]
E.; Bailey, S
Telles, E. E.; Bailey, S. R.; Davoudpour, S.; and Freeman, N. C. 2025. Racial inequality in Latin America. Oxford Open Economics, 4(Supplement\_1): i200--i218
2025
-
[59]
B.; Solberg, V
Torres, J. B.; Solberg, V. S. H.; and Carlstrom, A. H. 2002. The myth of sameness among Latino men and their machismo. American Journal of Orthopsychiatry, 72(2): 163--181
2002
-
[60]
E.; and Osueke, A
Tsounta, M. E.; and Osueke, A. 2014. What is behind Latin America’s declining income inequality? International Monetary Fund
2014
-
[61]
Yanaka, H.; Han, N.; Kumon, R.; Lu, J.; Takeshita, M.; Sekizawa, R.; Kato, T.; and Arai, H. 2024. Analyzing social biases in japanese large language models. arXiv preprint arXiv:2406.02050
2024 arXiv
-
[62]
Zhao, Y.; Wang, B.; Wang, Y.; Zhao, D.; Jin, X.; Zhang, J.; He, R.; and Hou, Y. 2024. A Comparative Study of Explicit and Implicit Gender Biases in Large Language Models via Self-evaluation. In Proceedings of the 2024 Joint International Conference on Computational Linguistics...
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.