REVIEW 2 major objections 6 minor 1 cited by
EsBBQ and CaBBQ: The Spanish and Catalan Bias Benchmarks for Question Answering
T0 review · 2 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Spanish and Catalan bias benchmarks for question answering show that higher QA accuracy often comes with greater reliance on stereotypes.
desk verdict A careful and genuinely adapted Spanish/Catalan bias QA resource, but its headline accuracy and bias numbers need re-scoring before you trust them. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the template-instantiation scheme inherited from BBQ and adapted here: for each stereotype, a template yields paired contexts (ambiguous, stereotypical disambiguated, anti-stereotypical disambiguated), two question polarities (negative and non-negative), and three answer options (target group, non-target group, unknown), with all ordering permutations and placeholder values systematically generated. The bias score is the load-bearing measurement: in ambiguous items it is the difference in prediction frequencies between stereotypical and anti-stereotypical answers divided by the total, and in disambiguated items it is the accuracy on stereotypical minus accuracy on anti-stereotypical instances. The survey with its 63% retention threshold, the manual translation by a Spanish translation graduate, and the validation of newly created templates by 17 annotators are what make the benchmark specific to Spain rather than a machine translation of the U.S.-centred original.
What would settle it
Run a replication survey with a demographically representative sample of Spanish residents that asks the same stereotype-elicitation questions; if a large fraction of the benchmark items (say, more than a third) are not recognized as prevalent by this new sample, the ground-truth labels would not stand, and the reported bias scores would be measuring the benchmark's own assumptions rather than model bias.
Extended reading notes
Core claim
On the paper's own terms, the central contribution is the construction of two parallel, culturally adapted benchmarks that measure stereotyping in a multiple-choice QA setting. Each template instantiates an ambiguous context where the correct answer is always 'unknown' and a disambiguated context whose correct answer is either stereotypical or anti-stereotypical; by computing accuracy and a bias score (the difference between stereotypical and anti-stereotypical answer rates, or between accuracies on the two disambiguated variants), the dataset separates task competence from stereotype reliance. The evaluated models show positive bias scores across the board, meaning all of them favour stereotypical answers in at least some measure, and the disambiguated results show larger models and instruction-tuned models often combining the highest QA accuracy with the highest bias scores, notably in the Physical Appearance and SES categories. The paper interprets this as evidence that, in Spanish and Catalan, high QA capability and stereotype reliance travel together, and it raises the open question whether social biases are inseparable from the linguistic and world knowledge the models acquire.
Load-bearing premise
The load-bearing premise is that the stereotypes validated through the majority-skewed survey actually are the prevalent harmful stereotypes of Spanish society, so the benchmark's ground-truth 'stereotypical' answers really measure social bias.
Editorial extensions
If this is right
- Researchers and practitioners can now audit stereotype reliance in Spanish and Catalan QA with open, culturally adapted benchmarks instead of relying on English or U.S.-centric resources.
- Model developers can use the ambiguous-context accuracy and the bias score as separate axes, so a model that merely answers stereotypically is not mistaken for a good QA system.
- The reported positive bias scores across all evaluated models imply that none of the tested LLMs is free of stereotype reliance in these languages, even when its overall answers look correct.
- The correlation between disambiguated accuracy and bias score implies that improving a model's QA performance on Spanish and Catalan will not automatically reduce social bias; the two may need to be optimised separately.
- The category breakdown identifies Physical Appearance and SES as the dimensions driving the average bias, while Religion shows anti-stereotypical answers, pointing to where bias mitigation should be targeted.
Reading between the lines
- The benchmark's 705-respondent survey skews toward majority groups (the paper acknowledges this), so the ground-truth stereotype labels may overrepresent majority perspectives; a replication with stratified sampling could confirm or revise specific items, especially for minority-targeted categories.
- Because the same template set underlies both ESBBQ and CaBBQ, the two benchmarks could be used to test whether models that are strong in Spanish transfer their stereotype reliance to Catalan or whether the Catalan version catches different confounds, a comparison the paper does not draw explicitly.
- The template-instantiation scheme, with all permutations and no random sampling, makes the benchmark deterministic; this means a future model can be scored at different sizes or training stages and any change in bias score is attributable to the model rather than to random dataset draws, an extension that would sharpen the reported size/bias trend.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces EsBBQ and CaBBQ, two parallel Spanish and Catalan bias benchmarks for multiple-choice question answering, adapted from the English BBQ dataset. The construction relies on a public survey of 705 respondents in Spain, manual cultural and linguistic translation, category restructuring (e.g., splitting Gender Identity into Gender and LGBTQIA, adding a Spanish Region category), and 110 newly created templates validated by 17 annotators. The authors evaluate 17 model variants from several families using a zero-shot multiple-choice setup and report accuracy and bias scores for ambiguous and disambiguated contexts. The headline empirical claims are that models tend to fail to select the correct unknown answer in ambiguous scenarios and that higher QA accuracy often correlates with greater reliance on social biases.
Significance. If validated, the resource itself is a valuable contribution: it provides open, culturally adapted bias-evaluation data for Spanish and Catalan, fills a clear gap in non-U.S. and non-English bias benchmarks, and is documented with unusual care including survey instruments, annotator demographics, template statistics, and a detailed limitation section. The authors are explicit about the survey's skew toward majority groups and about the lack of intersectionality, which is a strength. The code and data are released, and the template-instantiation process eliminates the randomization confounds present in the original BBQ. However, the empirical evaluation has a load-bearing methodological flaw in the scoring protocol (comparing raw log-likelihoods across answers of different lengths), which undermines the abstract's ambiguous-context accuracy claim and the reported accuracy–bias trend. The dataset itself can be re-evaluated with a corrected protocol, so the contribution remains significant after revision.
major comments (2)
- [§4.2 and Appendix E.2] The evaluation protocol scores each candidate answer by the raw sequence log-likelihood of the full prompt-plus-answer, but the answer strings are of very different lengths: target and non-target answers are short noun phrases or names, while the unknown expressions range from 'No sé' to 'No se puede determinar' (Table 6). Raw log-likelihoods are not length-normalized, so longer sequences accumulate more negative scores; in ambiguous contexts the correct answer is always one of these unknown expressions, which are systematically longer than the other options. This mechanically deflates Accambig (Eq. 1) and inflates Biasambig (Eq. 3). The authors should re-score with per-token log-likelihood (average log-likelihood per token) or an equivalent length-controlled comparison, and re-report all ambiguous-context results, including the abstract's claim that 'models tend to fail to choose the correct answer in ambiguous scenarios.'
- [§4.4 and Figure 3b] The conclusion that 'high QA accuracy often correlates with greater reliance on social biases' is supported only by visual inspection of aggregate scores over a small set of models, without statistical tests or uncertainty quantification. After correcting the scoring protocol, the correlation should be quantified (e.g., rank correlation with a confidence interval) and checked against the possibility that it is driven by the ambiguous-context scoring artifact or by a small number of outliers such as the instructed Salamandra 7B model, which has the highest Biasambig and also very high disambiguated accuracy.
minor comments (6)
- [Throughout] The acronym is inconsistent: the title and GitHub use 'CaBBQ', but the abstract and body text use 'CABBQ' (e.g., Section 1, Section 3, Table 2). Please standardize to one spelling.
- [§2.2] The sentence 'Previous literature have already adressed the measurement...' contains a typo: 'adressed' should be 'addressed'.
- [Appendix E, last paragraph] The Catalan placeholder example list includes 'la mujer' (Spanish) in the CA section; presumably this should be 'la dona' to match the Catalan pattern.
- [§3.2.5] The criterion for dropping NEWLY-CREATED templates is stated as 'failed to reach agreement among four annotators', but the exact agreement rule (e.g., majority of 3/4, unanimous, or a fixed agreement threshold) is not specified; please clarify and report the measured agreement level for the retained templates.
- [Figure 2] The legend terms 'Gray-squared numbers', 'circled numbers', 'blue squares', and 'orange-squared numbers' are difficult to map to the diagram; a cleaner schematic or more explicit labels would improve readability.
- [§4.1 and Table 3] The paper alternates between 'instructed' and 'instruction-tuned' to describe the same model variants; please define the term and use it consistently throughout.
Circularity Check
No circularity: benchmark labels come from external survey/references and model scores are empirical outputs; self-citations are not load-bearing.
full rationale
The paper's core contribution is a Spanish/Catalan adaptation of BBQ, with stereotype labels established by an external survey, secondary references, and annotator validation, not by the evaluated models. The bias and accuracy metrics (Eqs. 1-4) are computed from model outputs against this external gold standard, so the reported trends are empirical measurements rather than consequences of the benchmark's construction. Self-citations (Salamandra technical report, Mina et al. 2025) appear only for model description and for a standard caution about answer-order effects; neither supplies the benchmark's validity. The acknowledged skew in survey demographics affects stereotype coverage but does not make the evaluation circular, since the labels remain external to the models. The possible confound that 'unknown' answer strings are longer than target/non-target options is a methodological validity concern about log-likelihood scoring, not a circularity: no equation or fitted parameter makes the headline result true by definition, and the sign and magnitude of the bias scores still depend on model behavior. The central claim is therefore self-contained with respect to external evidence.
Assumptions & free parameters
free parameters (2)
- Survey stereotype inclusion threshold =
retention rule: frequency >= 3, or frequency >= 2 with at least one affected report; resulting 63% of 289 unique…
- Minimum annotator agreement for NEWLY-CREATED templates =
agreement among 4 annotators; 6.45% templates dropped
assumptions (4)
- domain assumption Survey responses and cited references accurately capture the prevalence of harmful stereotypes in Spanish society.
- domain assumption The original BBQ templates and stereotype annotations are a valid starting point for measuring social bias.
- domain assumption Comparing raw whole-string log-likelihoods across answer choices of different lengths is a valid way to select a model's answer.
- domain assumption Social bias can be measured in a multiple-choice QA task where the correct answer is unknown in ambiguous contexts.
Cite this review
Pith. "Pith review of EsBBQ and CaBBQ: The Spanish and Catalan Bias Benchmarks for Question Answering." pith.science (2026). https://pith.science/paper/BV4MQ5U2
@misc{pith2026250711216,
author = {Pith},
title = {Pith review of: EsBBQ and CaBBQ: The Spanish and Catalan Bias Benchmarks for Question Answering},
year = {2026},
howpublished = {\url{https://pith.science/paper/BV4MQ5U2}},
note = {Machine review of arXiv:2507.11216}
}
read the original abstract
Previous literature has largely shown that Large Language Models (LLMs) perpetuate social biases learnt from their pre-training data. Given the notable lack of resources for social bias evaluation in languages other than English, and for social contexts outside of the United States, this paper introduces the Spanish and the Catalan Bias Benchmarks for Question Answering (EsBBQ and CaBBQ). Based on the original BBQ, these two parallel datasets are designed to assess social bias across 10 categories using a multiple-choice QA setting, now adapted to the Spanish and Catalan languages and to the social context of Spain. We report evaluation results on different LLMs, factoring in model family, size and variant. Our results show that models tend to fail to choose the correct answer in ambiguous scenarios, and that high QA accuracy often correlates with greater reliance on social biases.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 1 Pith paper
-
Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
Unsigned differential activations locate a few GLU-MLP neurons whose zeroing surgically destabilizes demographic bias while retaining ~99.5% of measured capabilities.
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Afra Feyza Akyürek, Sejin Paik, Muhammed Yusuf Kocyigit, Seda Akbiyik, Şerife Leman Runyun, and Derry Wijaya. 2022. https://arxiv.org/abs/2205.11605 On measuring social biases in prompt-based multi-task learning . Preprint, arXiv:2205.11605
work page Pith review arXiv 2022
-
[4]
Duarte M. Alves, José Pombal, Nuno M. Guerreiro, Pedro H. Martins, João Alves, Amin Farajian, Ben Peters, Ricardo Rei, Patrick Fernandes, Sweta Agrawal, Pierre Colombo, José G. C. de Souza, and André F. T. Martins. 2024. https://arxiv.org/abs/2402.17733 Tower: An open multilingual large language model for translation-related tasks . Preprint, arXiv:2402.17733
arXiv 2024
-
[5]
Eleftherios Avramidis, Annika Gr \"u tzner-Zahn, Manuel Brack, Patrick Schramowski, Pedro Ortiz Suarez, Malte Ostendorff, Fabio Barth, Shushen Manakhimova, Vivien Macketanz, Georg Rehm, and Kristian Kersting. 2024. https://doi.org/10.18653/v1/2024.wmt-1.23 Occiglot at WMT 24: E uropean open-source large language models evaluated on translation . In Procee...
-
[6]
Solon Barocas, Kate Crawford, Aaron Shapiro, and Hanna Wallach. 2017. The problem with bias: From allocative to representational harms in machine learning. In SIGCIS conference paper
work page 2017
-
[8]
Hilary B Bergsieker, Lisa M Leslie, Vanessa S Constantine, and Susan T Fiske. 2012. Stereotyping by omission: eliminate the negative, accentuate the positive. Journal of personality and social psychology, 102(6):1214
work page 2012
-
[9]
Steven Bird. 2020. https://doi.org/10.18653/v1/2020.coling-main.313 Decolonising speech and language technology . In Proceedings of the 28th International Conference on Computational Linguistics, pages 3504--3519, Barcelona, Spain (Online). International Committee on Computational Linguistics
Show all 60 references
-
[10]
Su Lin Blodgett, Solon Barocas, Hal Daum \'e III, and Hanna Wallach. 2020. https://doi.org/10.18653/v1/2020.acl-main.485 Language (technology) is power: A critical survey of bias in NLP . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistic...
2020 doi
-
[11]
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021. https://doi.org/10.18653/v1/2021.acl-long.81 Stereotyping N orwegian salmon: An inventory of pitfalls in fairness benchmark datasets . In Proceedings of the 59th Annual Meeting of the Asso...
2021 doi
-
[12]
Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016. https://arxiv.org/abs/1607.06520 Man is to computer programmer as woman is to homemaker? debiasing word embeddings . Preprint, arXiv:1607.06520
2016 arXiv
-
[13]
Hudson, Ehsan Adeli, Russ B
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ B. Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri S. Chatterji, Annie S. Chen, Kathleen Cr...
2021 arXiv
-
[14]
Bryson, and Arvind Narayanan
Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017. https://doi.org/10.1126/science.aal4230 Semantics derived automatically from language corpora contain human-like biases . Science, 356(6334):183–186
2017 doi
-
[15]
Centro de Investigaciones Sociológicas (CIS) . 2024. https://www.cis.es/documents/d/cis/es3489marMT_a Barómetro de diciembre 2024 . Technical Report 3489, Centro de Investigaciones Sociológicas (CIS)
2024
-
[16]
Kate Crawford. 2017. The trouble with bias. https://www.youtube.com/watch?v=fMym_BKWQzk. Keynote at NeurIPS
2017
-
[17]
Severino Da Dalt, Joan Llop, Irene Baucells, Marc Pamies, Yishi Xu, Aitor Gonzalez-Agirre, and Marta Villegas. 2024. https://aclanthology.org/2024.lrec-main.650/ FLOR : On the effectiveness of language adaptation . In Proceedings of the 2024 Joint International Conference on C...
2024
-
[18]
Pieter Delobelle, Ewoenam Tokpo, Toon Calders, and Bettina Berendt. 2022. https://doi.org/10.18653/v1/2022.naacl-main.122 Measuring fairness with biased rulers: A comparative study on bias metrics for pre-trained language models . In Proceedings of the 2022 Conference of the N...
2022 doi
-
[19]
Sunipa Dev, Emily Sheng, Jieyu Zhao, Aubrie Amstutz, Jiao Sun, Yu Hou, Mattie Sanseverino, Jiin Kim, Akihiro Nishi, Nanyun Peng, and Kai-Wei Chang. 2022. https://doi.org/10.18653/v1/2022.findings-aacl.24 On measures of biases and harms in NLP . In Findings of the Association f...
2022 doi
-
[20]
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021. https://doi.org/10.1145/3442188.3445924 Bold: Dataset and metrics for measuring biases in open-ended language generation . In Proceedings of the 2021 ACM Confere...
2021
-
[21]
Fanny Ducel, Aur \'e lie N \'e v \'e ol, and Kar \"e n Fort. 2023. https://inria.hal.science/hal-04171198 Bias Identification in Language Models is Biased . In Workshop on Algorithmic Injustice 2023 , Amsterdam, Netherlands
2023
-
[23]
Virginia Felkner, Ho-Chun Herbert Chang, Eugene Jang, and Jonathan May. 2023 b . https://doi.org/10.18653/v1/2023.acl-long.507 W ino Q ueer: A community-in-the-loop benchmark for anti- LGBTQ + bias in large language models . In Proceedings of the 61st Annual Meeting of the Ass...
2023 doi
-
[24]
Gallegos, Ryan A
Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. 2024. https://doi.org/10.1162/coli_a_00524 Bias and fairness in large language models: A survey . Computational Linguistics, 50(3):1097--1179
2024 doi
-
[25]
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac'h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang S...
2023 doi
-
[26]
Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Mu \ n oz S \'a nchez, Mugdha Pandya, and Adam Lopez. 2021. https://doi.org/10.18653/v1/2021.acl-long.150 Intrinsic bias metrics do not correlate with application bias . In Proceedings of the 59th Annual Meeting of the Asso...
2021 doi
-
[27]
Seraphina Goldfarb-Tarrant, Eddie Ungless, Esma Balkir, and Su Lin Blodgett. 2023. https://2023.aclweb.org/ This prompt is measuring < mask > : Evaluating bias evaluation in language models . In Findings of the Association for Computational Linguistics: ACL 2023, pages 2209--2...
2023
-
[28]
Aitor Gonzalez-Agirre, Marc Pàmies, Joan Llop, Irene Baucells, Severino Da Dalt, Daniel Tamayo, José Javier Saiz, Ferran Espuña, Jaume Prats, Javier Aula-Blasco, Mario Mina, Iñigo Pikabea, Adrián Rubio, Alexander Shvets, Anna Sallés, Iñaki Lacunza, Jorge Palomar, Júlia Falcão,...
2025 arXiv
-
[29]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...
2024 arXiv
-
[30]
Wei Guo and Aylin Caliskan. 2021. https://doi.org/10.1145/3461702.3462536 Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases . In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, AIES '21,...
2021
-
[31]
Hilton and William von Hippel
James L. Hilton and William von Hippel. 1996. https://doi.org/10.1146/annurev.psych.47.1.237 Stereotypes . Annual Review of Psychology, 47(1):237--271
1996 doi
-
[32]
Dirk Hovy and Shrimai Prabhumoye. 2021. Five sources of bias in natural language processing. Language and linguistics compass, 15(8):e12432
2021
-
[33]
Yufei Huang and Deyi Xiong. 2024. https://aclanthology.org/2024.lrec-main.260/ CBBQ : A C hinese bias benchmark dataset curated with human- AI collaboration for large language models . In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Lang...
2024
-
[34]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...
2023 arXiv
-
[35]
Jiho Jin, Woosung Kang, Junho Myung, and Alice Oh. 2025. https://arxiv.org/abs/2503.06987 Social bias benchmark for generation: A comparison of generation and qa-based evaluations . Preprint, arXiv:2503.06987
2025 arXiv
-
[36]
Jiho Jin, Jiseon Kim, Nayeon Lee, Haneul Yoo, Alice Oh, and Hwaran Lee. 2024. https://doi.org/10.1162/tacl_a_00661 K o BBQ : K orean bias benchmark for question answering . Transactions of the Association for Computational Linguistics, 12:507--524
2024 doi
-
[37]
Masahiro Kaneko, Danushka Bollegala, and Naoaki Okazaki. 2022. https://aclanthology.org/2022.coling-1.111/ Debiasing isn`t enough! -- on the effectiveness of debiasing MLM s and their social biases in downstream tasks . In Proceedings of the 29th International Conference on Co...
2022
-
[38]
Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov. 2019. https://doi.org/10.18653/v1/W19-3823 Measuring bias in contextualized word representations . In Proceedings of the First Workshop on Gender Bias in Natural Language Processing, pages 166--172, Flor...
2019 doi
-
[39]
Tao Li, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Vivek Srikumar. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.311 UNQOVER ing stereotyping biases via underspecified questions . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages ...
2020 doi
-
[40]
Guerreiro, Ricardo Rei, Duarte M
Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro, Ricardo Rei, Duarte M. Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, Pierre Colombo, Barry Haddow, José G. C. de Souza, Alexandra Birch, and André F. T. Martins. 2024. https://a...
2024 arXiv
-
[41]
Bowman, and Rachel Rudinger
Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019. https://doi.org/10.18653/v1/N19-1063 On measuring social biases in sentence encoders . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational...
2019 doi
-
[42]
Mario Mina, Valle Ruiz-Fern \'a ndez, J \'u lia Falc \ a o, Luis Vasquez-Reina, and Aitor Gonzalez-Agirre. 2025. https://aclanthology.org/2025.coling-main.120/ Cognitive biases, task complexity, and result intepretability in large language models . In Proceedings of the 31st I...
2025
-
[43]
Eleazar Ros Moreno, Antonio Daniel Garc \' a Rojas, and Mar \' a Jos \'e Carrasco Mac \' as. 2024. Realidad actual sobre la inclusi \'o n del colectivo trans* en espa \ n a. G \'e neros , 13(3):175--196
2024
-
[44]
Miguel Moya and Alba Moya-Garófano. 2021. https://reunido.uniovi.es/index.php/PST/article/view/17070 Evolución de los estereotipos de género en españa: de 1985 a 2018 . Psicothema, 33(Número 1):53–59
2021
-
[45]
Daniel Naber. 2003. https://www.danielnaber.de/languagetool/download/style_and_grammar_checker.pdf A rule-based style and grammar checker
2003
-
[46]
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. https://doi.org/10.18653/v1/2021.acl-long.416 S tereo S et: Measuring stereotypical bias in pretrained language models . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Inte...
2021 doi
-
[47]
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.154 C row S -pairs: A challenge dataset for measuring social biases in masked language models . In Proceedings of the 2020 Conference on Empirical Methods in Na...
2020 doi
-
[48]
Vera Neplenbroek, Arianna Bisazza, and Raquel Fernández. 2024. https://arxiv.org/abs/2406.07243 Mbbq: A dataset for cross-lingual comparison of stereotypes in generative llms . Preprint, arXiv:2406.07243
2024 arXiv
-
[49]
Debora Nozza, Federico Bianchi, and Dirk Hovy. 2021. https://doi.org/10.18653/v1/2021.naacl-main.191 HONEST : Measuring hurtful sentence completion in language models . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Lin...
2021 doi
-
[50]
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R. Bowman. 2021. https://arxiv.org/abs/2110.08193 BBQ: A hand-built bias benchmark for question answering . CoRR, abs/2110.08193
2021 arXiv
-
[51]
Elizabeth Peel, Sonja J Ellis, and Damien W Riggs. 2021. Lesbian, gay, bisexual and transgender people: Prejudice, stereotyping, discrimination and social change. In The Routledge international handbook of discrimination, prejudice and stereotyping, pages 104--117. Routledge
2021
-
[52]
Mat \'u s Pikuliak, Ivana Be n ov \'a , and Viktor Bachrat \'y . 2023. https://doi.org/10.18653/v1/2023.eacl-main.265 In-depth look at word filling societal bias measures . In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Lingu...
2023 doi
-
[53]
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, ...
2025 arXiv
-
[54]
Jaime Rossell. 2020. https://www.inclusion.gob.es/oberaxe/ficheros/documentos/LaNoDiscriminacionMotivosReligiosos.pdf La no discriminación por motivos religiosos . Ministerio de Inclusión, Seguridad Social y Migraciones
2020
-
[55]
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018. https://doi.org/10.18653/v1/N18-2002 Gender bias in coreference resolution . In Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computational Linguistics: H...
2018 doi
-
[56]
Xabier Saralegi and Muitze Zulaika. 2025. https://aclanthology.org/2025.coling-main.318/ B asq BBQ : A QA benchmark for assessing social biases in LLM s for B asque, a low-resource language . In Proceedings of the 31st International Conference on Computational Linguistics, pag...
2025
-
[57]
Nikil Selvam, Sunipa Dev, Daniel Khashabi, Tushar Khot, and Kai-Wei Chang. 2023. https://doi.org/10.18653/v1/2023.acl-short.118 The tail wagging the dog: Dataset construction biases of social bias benchmarks . In Proceedings of the 61st Annual Meeting of the Association for Co...
2023 doi
-
[58]
Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.625 I `m sorry to hear that : Finding new biases in language models with a holistic descriptor dataset . In Proceedings of the 2022 Confe...
2022 doi
-
[59]
Elisa Celis
Yi Chern Tan and L. Elisa Celis. 2019. Assessing social and intersectional biases in contextualized word representations. Curran Associates Inc., Red Hook, NY, USA
2019
-
[60]
Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon, ...
2025 arXiv
-
[61]
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. https://arxiv.org/abs/1804.06876 Gender bias in coreference resolution: Evaluation and debiasing methods
2018 arXiv
-
[62]
Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh
Tony Z. Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. https://arxiv.org/abs/2102.09690 Calibrate before use: Improving few-shot performance of language models . Preprint, arXiv:2102.09690
2021 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.