REVIEW 50 references
Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory
T0 review · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Across 16 open-weight language models, gender is encoded as a sex-linked binary, transgender and nonbinary terms are less probable than random nouns, and mental illness is more probable after those terms; these effects grow with model size.
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Extended reading notes
Core claim
The paper's central claim, stated in the abstract, is that 'language models tend to encode gender as a binary category tied to biological sex, and that gendered terms that do not neatly fall into one of these binary categories are erased and pathologized', and additionally that 'larger models... learn stronger associations between gender and sex.' If true, this means gender bias in LMs is not only a stereotype problem but a construction problem: the models define what genders are 'imaginable', reproducing a folk understanding of gender.
Load-bearing premise
The scaling result in Figure 2 depends on the Folk-Subversive LPR (Eq. 1), which sums raw conditional probabilities without normalizing by the marginal probability of each gender term. If 'a man' and 'a woman' are simply more probable completions in any neutral context due to token frequency, a positive LPR and its increase with model size could partly reflect base rates rather than learned sex-gender associations. The matched-guise Sex-Gender LPR (Eq. 2) controls for this, but the headline size trend is built on Eq. 1.
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
assumptions (4)
- domain assumption Gender performativity theory provides the normative standard for labeling gender terms as 'folk' or 'subversive'.
- domain assumption Conditional probabilities from LMs are a valid operationalization of how models 'construct' gender.
- ad hoc to paper The 47 random nouns from wonderwords constitute a valid non-human baseline for 'meaningful embedding' of personhood.
- domain assumption The sex characteristic sets S_F and S_M align with the folk binary (testosterone/estrogen, XY/XX, etc.).
Cite this review
Pith. "Pith review of Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory." pith.science (2026). https://pith.science/paper/3TPW5JAO
@misc{pith2026250514080,
author = {Pith},
title = {Pith review of: Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory},
year = {2026},
howpublished = {\url{https://pith.science/paper/3TPW5JAO}},
note = {Machine review of arXiv:2505.14080}
}
read the original abstract
Language models encode and subsequently perpetuate harmful gendered stereotypes. Research has succeeded in mitigating some of these harms, e.g. by dissociating non-gendered terms such as occupations from gendered terms such as 'woman' and 'man'. This approach, however, remains superficial given that associations are only one form of prejudice through which gendered harms arise. Critical scholarship on gender, such as gender performativity theory, emphasizes how harms often arise from the construction of gender itself, such as conflating gender with biological sex. In language models, these issues could lead to the erasure of transgender and gender diverse identities and cause harms in downstream applications, from misgendering users to misdiagnosing patients based on wrong assumptions about their anatomy. For FAccT research on gendered harms to go beyond superficial linguistic associations, we advocate for a broader definition of 'gender bias' in language models. We operationalize insights on the construction of gender through language from gender studies literature and then empirically test how 16 language models of different architectures, training datasets, and model sizes encode gender. We find that language models tend to encode gender as a binary category tied to biological sex, and that gendered terms that do not neatly fall into one of these binary categories are erased and pathologized. Finally, we show that larger models, which achieve better results on performance benchmarks, learn stronger associations between gender and sex, further reinforcing a narrow understanding of gender. Our findings lead us to call for a re-evaluation of how gendered harms in language models are defined and addressed.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Paul Baker. 2008. Sexed Texts. Cambridge University Press, Cambridge, UK
work page 2008
-
[2]
Abeba Birhane, Sepehr Dehdashtian, Vinay Prabhu, and Vishnu Boddeti. 2024. The Dark Side of Dataset Scaling: Evaluating Racial Classification in Multimodal Models. In The 2024 ACM Conference on Fairness, Accountability, and Transparency. ACM, Rio de Janeiro Brazil, 1229–1244
work page 2024
-
[3]
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. Lan- guage (Technology) Is Power: A Critical Survey of “Bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, ...
work page 2020
-
[4]
Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Proceedings of the 30th International Conference on Neural Information Processing Systems (Barcelona, Spain) (NIPS’16). Curran Associates Inc., Red Hook, NY, USA, 4356–4364
work page 2016
-
[5]
Judith Butler. 2006. Gender Trouble: Feminism and the Subversion of Identity . Routledge, New York
work page 2006
-
[6]
Judith Butler. 2011. Bodies That Matter: On the Discursive Limits of Sex . Routledge, London
work page 2011
-
[7]
Judith Butler. 2019. Gender in Translation: Beyond Monolingualism.philoSOPHIA: A Journal of Continental Feminism 9, 1 (2019), 1–25. doi:10.1353/phi.2019.0011
arXiv 2019
-
[8]
Gabrielle R Chiaramonte and Ronald Friend. 2006. Medical students’ and resi- dents’ gender bias in the diagnosis, treatment, and interpretation of coronary heart disease symptoms. Health Psychology 25, 3 (2006), 255
work page 2006
Show all 50 references
-
[9]
Zowie Davy. 2015. The DSM-5 and the Politics of Diagnosing Transpeople. Archives of Sexual Behavior 44, 5 (July 2015), 1165–1176. doi:10.1007/s10508-015- 0573-6
2015 doi
-
[10]
Zowie Davy, Anniken Sørlie, and Amets Suess Schwend. 2018. Democratising di- agnoses? The role of the depathologisation perspective in constructing corporeal trans citizenship. Critical Social Policy 38, 1 (Feb. 2018), 13–34
2018
-
[11]
Simone De Beauvoir. 1953. The Second Sex. Knopf, New York
1953
-
[12]
Hannah Devinney, Jenny Björklund, and Henrik Björklund. 2022. Theories of “Gender” in NLP Bias Research. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’22). Association for Computing Machinery, New York, NY, USA, 2083–2102
2022
-
[13]
Dewey and Melissa M
Jodie M. Dewey and Melissa M. Gesbeck. 2017. (Dys) Functional Diagnosing: Mental Health Diagnosis, Medicalization, and the Making of Transgender Patients. Humanity & Society 41, 1 (Feb. 2017), 37–72
2017
-
[14]
Harnoor Dhingra, Preetiha Jayashanker, Sayali Moghe, and Emma Strubell. 2023. Queer people are people first: Deconstructing sexual identity stereotypes in large language models
2023
-
[15]
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. In Proceedings of the 2021 Conference on Empirical Methods...
2021
-
[16]
Jack Drescher. 2015. Queer diagnoses revisited: The past and future of homosex- uality and gender diagnoses in DSM and ICD. International Review of Psychiatry 27, 5 (Sept. 2015), 386–395
2015
-
[17]
Jack Drescher, Peggy Cohen-Kettenis, and Sam Winter. 2012. Minding the body: Situating gender identity diagnoses in the ICD-11. International Review of Psy- chiatry 24, 6 (Dec. 2012), 568–577
2012
-
[18]
Virginia Felkner, Ho-Chun Herbert Chang, Eugene Jang, and Jonathan May
-
[19]
Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023. From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models. InProceedings of the 61st Annual Meeting of the Association for Computationa...
2023
-
[20]
Michel Foucault. 1980. Herculine Barbin: Being the Recently Discovered Mem- oirs of a Nineteenth-century French Hermaphrodite . Pantheon Books, New York. Translated by Richard McDougall
1980
-
[21]
Gallegos, Ryan A
Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. 2024. Bias and Fairness in Large Language Models: A Survey. Computational Linguistics 50, 3 (Sept. 2024), 1097–1179
2024
-
[22]
Ian Hacking. 2007. Kinds of people: Moving targets. In Proceedings of the British Academy, Vol. 151. Oxford University Press, Oxford, UK, 285–318g
2007
-
[23]
Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024. AI Generates Covertly Racist Decisions about People Based on Their Dialect. Nature 633, 8028 (Sept. 2024), 147–154
2024
-
[24]
Emma Inch. 2016. Changing Minds: The Psycho-Pathologization of Trans People. International Journal of Mental Health 45, 3 (July 2016), 193–204
2016
-
[25]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thoma...
2023 arXiv
-
[26]
Os Keyes. 2018. The Misgendering Machines: Trans/HCI Implications of Auto- matic Gender Recognition. Proceedings of the ACM on Human-Computer Interac- tion 2, CSCW (Nov. 2018), 88:1–88:22
2018
-
[27]
Kilavuz, Anna Korhonen, and Hinrich Schuetze
Abdullatif Köksal, Omer Yalcin, Ahmet Akbiyik, M. Kilavuz, Anna Korhonen, and Hinrich Schuetze. 2023. Language-Agnostic Bias Detection in Language Models with Bias Probing. In Findings of the Association for Computational Linguistics: EMNLP 2023. Association for Computational ...
2023
-
[28]
Smith, and Yejin Choi
Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi. 2021. DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts. In Proceedings of the 59th Annual Meeting of the Association for Computational Lin...
2021
-
[29]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 [cs.CL]
2019 arXiv
-
[30]
Renee Molloy, Gabrielle Brand, Ian Munro, and Nicole Pope. 2023. Seeing the complete picture: A systematic review of mental health consumer and health professional experiences of diagnostic overshadowing.Journal of Clinical Nursing 32, 9-10 (2023), 1662–1673
2023
-
[31]
Thekla Morgenroth and Michelle K Ryan. 2018. Gender trouble in social psychol- ogy: How can Butler’s work inform experimental social psychologists’ conceptu- alization of gender? Frontiers in Psychology 9 (2018), 1320
2018
-
[32]
I’m fully who I am
Anaelia Ovalle, Palash Goyal, Jwala Dhamala, Zachary Jaggers, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2023. “I’m fully who I am”: Towards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation. InProceedings of the 2023 ...
2023
-
[33]
Anaelia Ovalle, Krunoslav Lehman Pavasovic, Louis Martin, Luke Zettlemoyer, Eric Michael Smith, Adina Williams, and Levent Sagun. 2024. The Root Shapes the Fruit: On the Persistence of Gender-Exclusive Harms in Aligned Language Models. arXiv:2411.03700 [cs]
2024 arXiv
-
[34]
Matúš Pikuliak, Stefan Oresko, Andrea Hrckova, and Marian Simko. 2024. Women Are Beautiful, Men Are Leaders: Gender Stereotypes in Machine Translation and Language Modeling. In Findings of the Association for Computational Linguis- tics: EMNLP 2024, Yaser Al-Onaizan, Mohit Ban...
2024
-
[35]
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023. Fine-Tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
2023
-
[36]
Organizers Of Queerinai, Anaelia Ovalle, Arjun Subramonian, Ashwin Singh, Claas Voelcker, Danica J Sutherland, Davide Locatelli, Eva Breznik, Filip Klubicka, Hang Yuan, et al. 2023. Queer in AI: A case study in community-led participatory AI. In Proceedings of the 2023 ACM Con...
2023
-
[37]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners
2019
-
[38]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2023. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv:1910.10683 [cs.LG]
2023 arXiv
-
[39]
Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang
-
[40]
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019. The Woman Worked as a Babysitter: On Biases in Language Generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Pro- cessing and the 9th International Joint Conference on...
2019
-
[41]
In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd- Graber, and Naoaki Okazaki (Eds.)
On Second Thought, Let’s Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd- Graber, and Naoaki Okazaki (Eds.). Associat...
-
[42]
Ryan Steed, Swetasudha Panda, Ari Kobren, and Michael Wick. 2022. Upstream Mitigation Is Not All You Need: Testing the Bias Transfer Hypothesis in Pre- Trained Language Models. In Proceedings of the 60th Annual Meeting of the As- sociation for Computational Linguistics (Volume...
2022
-
[43]
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, Miles McCain, Alex Newhouse, Jason Blazakis, Kris McGuffie, and Jasmine Wang. 2019. Release Strategies and the Social Impacts of ...
2019 arXiv
-
[44]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guil- laume Lample. 2023. LLaMA: Open and Efficient Foundation ...
2023 arXiv
-
[45]
Yolande Strengers, Lizhen Qu, Qiongkai Xu, and Jarrod Knibbe. 2020. Adhering, Steering, and Queering: Treatment of Gender in Natural Language Generation. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (CHI ’20). Association for Computing Machin...
2020
-
[46]
Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li. 2023. DecodingTrust: A Com...
2023
-
[47]
Ekaterina Vylomova, Laura Rimell, Trevor Cohn, and Timothy Baldwin. 2015. Take and took, gaggle and goose, book and read: Evaluating the utility of vector differences for lexical relation learning
2015
-
[48]
Amy (Azure) Zhou. 2024. Queer Bias in Natural Language Processing: Towards More Expansive Frameworks of Gender and Sexuality in NLP Bias Research. GRACE: Global Review of AI Community Ethics 2, 1 (Jan. 2024), 1–9. Gender Trouble in Language Models FAccT ’25, June 23–26, 2025, ...
2024
-
[49]
Monique Wittig. 1985. The mark of gender. Feminist issues 5, 2 (1985), 3–12
1985
-
[2023]
In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.)
WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Associat...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.