Pith. sign in

REVIEW 50 references

Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory

T0 review · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Across 16 open-weight language models, gender is encoded as a sex-linked binary, transgender and nonbinary terms are less probable than random nouns, and mental illness is more probable after those terms; these effects grow with model size.

arxiv 2505.14080 v1 pith:3TPW5JAO submitted 2025-05-20 cs.CL cs.AIcs.CYcs.HC

classification cs.CLcs.AIcs.CYcs.HC
keywords genderlanguagemodelsgenderedharmsassociationsencodeterms
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Language models are trained on huge amounts of text, and they learn the patterns in that text. One pattern is a simple story about gender: men have male bodies and women have female bodies, everyone fits into one of these two groups, and that is all. The researchers behind this paper wanted to test whether this 'folk understanding' of gender shows up in what models predict. They used short sentence templates. For example, they fed models sentences like 'The person who has testosterone is ___' and asked how likely the model thought each ending was: 'a man', 'a woman', 'nonbinary', 'genderqueer', and so on. They also used templates like 'The person who is transgender has ___' and measured how likely different illnesses were. They ran this on 16 open-weight models, from small 60-million-parameter T5 to 70-billion-parameter Llama. The results are consistent. In almost every model, 'a man' is the most likely completion for male sex characteristics, and 'a woman' for female sex characteristics, while transgender and nonbinary terms get very low probabilities, often lower than random nouns like 'windscreen'. The models also tend to make mental illness words more likely than physical illness words after transgender and nonbinary contexts: words like 'body dysmorphia' and 'post-traumatic stress' appear at the top. And the larger the model, the stronger these patterns are: the correlation between model size and a folk-understanding score is 0.89.
Extended reading notes

Core claim

The paper's central claim, stated in the abstract, is that 'language models tend to encode gender as a binary category tied to biological sex, and that gendered terms that do not neatly fall into one of these binary categories are erased and pathologized', and additionally that 'larger models... learn stronger associations between gender and sex.' If true, this means gender bias in LMs is not only a stereotype problem but a construction problem: the models define what genders are 'imaginable', reproducing a folk understanding of gender.

Load-bearing premise

The scaling result in Figure 2 depends on the Folk-Subversive LPR (Eq. 1), which sums raw conditional probabilities without normalizing by the marginal probability of each gender term. If 'a man' and 'a woman' are simply more probable completions in any neutral context due to token frequency, a positive LPR and its increase with model size could partly reflect base rates rather than learned sex-gender associations. The matched-guise Sex-Gender LPR (Eq. 2) controls for this, but the headline size trend is built on Eq. 1.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters are fitted; the LPR metrics are defined directly from theory and probabilities. The main extra assumptions are the operationalizations: Butler's theory as the normative frame, conditional probability as a probe of gender construction, the particular sex characteristic sets, and the random-noun baseline. These are assumptions rather than fitted constants.

assumptions (4)
  • domain assumption Gender performativity theory provides the normative standard for labeling gender terms as 'folk' or 'subversive'.
    Invoked in Section 2.1 to define the folk understanding and in Eq. 1's delta_FOLK assignments.
  • domain assumption Conditional probabilities from LMs are a valid operationalization of how models 'construct' gender.
    Used throughout Section 3 to justify interpreting P(g|Context(s)) as encoding gender concepts rather than mere co-occurrence.
  • ad hoc to paper The 47 random nouns from wonderwords constitute a valid non-human baseline for 'meaningful embedding' of personhood.
    Section 3.1.1; the list includes abstract nouns like 'condition' and 'spirituality' that are not clearly non-human, and all completions are ungrammatical without an article.
  • domain assumption The sex characteristic sets S_F and S_M align with the folk binary (testosterone/estrogen, XY/XX, etc.).
    Section 3.1.1; the matched pairs assume these characteristics are the canonical folk markers and that all other characteristics are irrelevant.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory." pith.science (2026). https://pith.science/paper/3TPW5JAO

@misc{pith2026250514080,
  author       = {Pith},
  title        = {Pith review of: Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3TPW5JAO}},
  note         = {Machine review of arXiv:2505.14080}
}
read the original abstract

Language models encode and subsequently perpetuate harmful gendered stereotypes. Research has succeeded in mitigating some of these harms, e.g. by dissociating non-gendered terms such as occupations from gendered terms such as 'woman' and 'man'. This approach, however, remains superficial given that associations are only one form of prejudice through which gendered harms arise. Critical scholarship on gender, such as gender performativity theory, emphasizes how harms often arise from the construction of gender itself, such as conflating gender with biological sex. In language models, these issues could lead to the erasure of transgender and gender diverse identities and cause harms in downstream applications, from misgendering users to misdiagnosing patients based on wrong assumptions about their anatomy. For FAccT research on gendered harms to go beyond superficial linguistic associations, we advocate for a broader definition of 'gender bias' in language models. We operationalize insights on the construction of gender through language from gender studies literature and then empirically test how 16 language models of different architectures, training datasets, and model sizes encode gender. We find that language models tend to encode gender as a binary category tied to biological sex, and that gendered terms that do not neatly fall into one of these binary categories are erased and pathologized. Finally, we show that larger models, which achieve better results on performance benchmarks, learn stronger associations between gender and sex, further reinforcing a narrow understanding of gender. Our findings lead us to call for a re-evaluation of how gendered harms in language models are defined and addressed.

Figures

Figures reproduced from arXiv: 2505.14080 by the authors.

Figure 1
Figure 1. Probability of gender-related predictions by sex-related context [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Alignment of Models with Folk Understanding of [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Alignment of gendered terms with male vs. female sex characteristics [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Distribution of Gender–Illness Log Probability Ratio per Gender Context [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Illnesses Most and Least Associated with Gender Contexts [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Probability of gender-related predictions by sex-related context [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: Alignment of gendered terms with male versus female sex characteristics [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 8
Figure 8. Figure 8: Probability of Mental versus Physical Illness Predictions by Gender Context [PITH_FULL_IMAGE:figures/full_fig_p015_8.png]
Figure 9
Figure 9. Figure 9: Illnesses Most and Least Associated with Gender Contexts [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

50 extracted references · 43 canonical work pages

  1. [1]

    Paul Baker. 2008. Sexed Texts. Cambridge University Press, Cambridge, UK

  2. [2]

    Abeba Birhane, Sepehr Dehdashtian, Vinay Prabhu, and Vishnu Boddeti. 2024. The Dark Side of Dataset Scaling: Evaluating Racial Classification in Multimodal Models. In The 2024 ACM Conference on Fairness, Accountability, and Transparency. ACM, Rio de Janeiro Brazil, 1229–1244

  3. [3]

    Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. Lan- guage (Technology) Is Power: A Critical Survey of “Bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, ...

  4. [4]

    Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. In Proceedings of the 30th International Conference on Neural Information Processing Systems (Barcelona, Spain) (NIPS’16). Curran Associates Inc., Red Hook, NY, USA, 4356–4364

  5. [5]

    Judith Butler. 2006. Gender Trouble: Feminism and the Subversion of Identity . Routledge, New York

  6. [6]

    Judith Butler. 2011. Bodies That Matter: On the Discursive Limits of Sex . Routledge, London

  7. [7]

    Judith Butler. 2019. Gender in Translation: Beyond Monolingualism.philoSOPHIA: A Journal of Continental Feminism 9, 1 (2019), 1–25. doi:10.1353/phi.2019.0011

  8. [8]

    Gabrielle R Chiaramonte and Ronald Friend. 2006. Medical students’ and resi- dents’ gender bias in the diagnosis, treatment, and interpretation of coronary heart disease symptoms. Health Psychology 25, 3 (2006), 255

Show all 50 references
  1. [9]

    Zowie Davy. 2015. The DSM-5 and the Politics of Diagnosing Transpeople. Archives of Sexual Behavior 44, 5 (July 2015), 1165–1176. doi:10.1007/s10508-015- 0573-6

  2. [10]

    Zowie Davy, Anniken Sørlie, and Amets Suess Schwend. 2018. Democratising di- agnoses? The role of the depathologisation perspective in constructing corporeal trans citizenship. Critical Social Policy 38, 1 (Feb. 2018), 13–34

  3. [11]

    Simone De Beauvoir. 1953. The Second Sex. Knopf, New York

  4. [12]

    Hannah Devinney, Jenny Björklund, and Henrik Björklund. 2022. Theories of “Gender” in NLP Bias Research. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’22). Association for Computing Machinery, New York, NY, USA, 2083–2102

  5. [13]

    Dewey and Melissa M

    Jodie M. Dewey and Melissa M. Gesbeck. 2017. (Dys) Functional Diagnosing: Mental Health Diagnosis, Medicalization, and the Making of Transgender Patients. Humanity & Society 41, 1 (Feb. 2017), 37–72

  6. [14]

    Harnoor Dhingra, Preetiha Jayashanker, Sayali Moghe, and Emma Strubell. 2023. Queer people are people first: Deconstructing sexual identity stereotypes in large language models

  7. [15]

    Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. In Proceedings of the 2021 Conference on Empirical Methods...

  8. [16]

    Jack Drescher. 2015. Queer diagnoses revisited: The past and future of homosex- uality and gender diagnoses in DSM and ICD. International Review of Psychiatry 27, 5 (Sept. 2015), 386–395

  9. [17]

    Jack Drescher, Peggy Cohen-Kettenis, and Sam Winter. 2012. Minding the body: Situating gender identity diagnoses in the ICD-11. International Review of Psy- chiatry 24, 6 (Dec. 2012), 568–577

  10. [18]

    Virginia Felkner, Ho-Chun Herbert Chang, Eugene Jang, and Jonathan May

  11. [19]

    Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023. From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models. InProceedings of the 61st Annual Meeting of the Association for Computationa...

  12. [20]

    Michel Foucault. 1980. Herculine Barbin: Being the Recently Discovered Mem- oirs of a Nineteenth-century French Hermaphrodite . Pantheon Books, New York. Translated by Richard McDougall

  13. [21]

    Gallegos, Ryan A

    Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. 2024. Bias and Fairness in Large Language Models: A Survey. Computational Linguistics 50, 3 (Sept. 2024), 1097–1179

  14. [22]

    Ian Hacking. 2007. Kinds of people: Moving targets. In Proceedings of the British Academy, Vol. 151. Oxford University Press, Oxford, UK, 285–318g

  15. [23]

    Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024. AI Generates Covertly Racist Decisions about People Based on Their Dialect. Nature 633, 8028 (Sept. 2024), 147–154

  16. [24]

    Emma Inch. 2016. Changing Minds: The Psycho-Pathologization of Trans People. International Journal of Mental Health 45, 3 (July 2016), 193–204

  17. [25]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thoma...

  18. [26]

    Os Keyes. 2018. The Misgendering Machines: Trans/HCI Implications of Auto- matic Gender Recognition. Proceedings of the ACM on Human-Computer Interac- tion 2, CSCW (Nov. 2018), 88:1–88:22

  19. [27]

    Kilavuz, Anna Korhonen, and Hinrich Schuetze

    Abdullatif Köksal, Omer Yalcin, Ahmet Akbiyik, M. Kilavuz, Anna Korhonen, and Hinrich Schuetze. 2023. Language-Agnostic Bias Detection in Language Models with Bias Probing. In Findings of the Association for Computational Linguistics: EMNLP 2023. Association for Computational ...

  20. [28]

    Smith, and Yejin Choi

    Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi. 2021. DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts. In Proceedings of the 59th Annual Meeting of the Association for Computational Lin...

  21. [29]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 [cs.CL]

  22. [30]

    Renee Molloy, Gabrielle Brand, Ian Munro, and Nicole Pope. 2023. Seeing the complete picture: A systematic review of mental health consumer and health professional experiences of diagnostic overshadowing.Journal of Clinical Nursing 32, 9-10 (2023), 1662–1673

  23. [31]

    Thekla Morgenroth and Michelle K Ryan. 2018. Gender trouble in social psychol- ogy: How can Butler’s work inform experimental social psychologists’ conceptu- alization of gender? Frontiers in Psychology 9 (2018), 1320

  24. [32]

    I’m fully who I am

    Anaelia Ovalle, Palash Goyal, Jwala Dhamala, Zachary Jaggers, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2023. “I’m fully who I am”: Towards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation. InProceedings of the 2023 ...

  25. [33]

    Anaelia Ovalle, Krunoslav Lehman Pavasovic, Louis Martin, Luke Zettlemoyer, Eric Michael Smith, Adina Williams, and Levent Sagun. 2024. The Root Shapes the Fruit: On the Persistence of Gender-Exclusive Harms in Aligned Language Models. arXiv:2411.03700 [cs]

  26. [34]

    Matúš Pikuliak, Stefan Oresko, Andrea Hrckova, and Marian Simko. 2024. Women Are Beautiful, Men Are Leaders: Gender Stereotypes in Machine Translation and Language Modeling. In Findings of the Association for Computational Linguis- tics: EMNLP 2024, Yaser Al-Onaizan, Mohit Ban...

  27. [35]

    Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023. Fine-Tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

  28. [36]

    Organizers Of Queerinai, Anaelia Ovalle, Arjun Subramonian, Ashwin Singh, Claas Voelcker, Danica J Sutherland, Davide Locatelli, Eva Breznik, Filip Klubicka, Hang Yuan, et al. 2023. Queer in AI: A case study in community-led participatory AI. In Proceedings of the 2023 ACM Con...

  29. [37]

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners

  30. [38]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2023. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv:1910.10683 [cs.LG]

  31. [39]

    Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang

  32. [40]

    Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019. The Woman Worked as a Babysitter: On Biases in Language Generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Pro- cessing and the 9th International Joint Conference on...

  33. [41]

    In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd- Graber, and Naoaki Okazaki (Eds.)

    On Second Thought, Let’s Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd- Graber, and Naoaki Okazaki (Eds.). Associat...

  34. [42]

    Ryan Steed, Swetasudha Panda, Ari Kobren, and Michael Wick. 2022. Upstream Mitigation Is Not All You Need: Testing the Bias Transfer Hypothesis in Pre- Trained Language Models. In Proceedings of the 60th Annual Meeting of the As- sociation for Computational Linguistics (Volume...

  35. [43]

    Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, Miles McCain, Alex Newhouse, Jason Blazakis, Kris McGuffie, and Jasmine Wang. 2019. Release Strategies and the Social Impacts of ...

  36. [44]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guil- laume Lample. 2023. LLaMA: Open and Efficient Foundation ...

  37. [45]

    Yolande Strengers, Lizhen Qu, Qiongkai Xu, and Jarrod Knibbe. 2020. Adhering, Steering, and Queering: Treatment of Gender in Natural Language Generation. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (CHI ’20). Association for Computing Machin...

  38. [46]

    Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li

    Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li. 2023. DecodingTrust: A Com...

  39. [47]

    Ekaterina Vylomova, Laura Rimell, Trevor Cohn, and Timothy Baldwin. 2015. Take and took, gaggle and goose, book and read: Evaluating the utility of vector differences for lexical relation learning

  40. [48]

    Amy (Azure) Zhou. 2024. Queer Bias in Natural Language Processing: Towards More Expansive Frameworks of Gender and Sexuality in NLP Bias Research. GRACE: Global Review of AI Community Ethics 2, 1 (Jan. 2024), 1–9. Gender Trouble in Language Models FAccT ’25, June 23–26, 2025, ...

  41. [49]

    Monique Wittig. 1985. The mark of gender. Feminist issues 5, 2 (1985), 3–12

  42. [2023]

    In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.)

    WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Associat...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.