Pith. sign in

REVIEW 4 major objections 5 minor 3 cited by

Queer NLP: A Critical Survey on Literature Gaps, Biases and Trends

T0 review · 4 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read Queer NLP research is growing but reactive, English-centric, and rarely community-led, a systematic survey finds.

desk verdict First systematic map of queer NLP in the ACL Anthology; the corpus-building is the weak spot, not the core findings. read the letter →

arxiv 2602.16151 v4 pith:U6JAEXWB submitted 2026-02-18 cs.CY

classification cs.CY
keywords queerNLPLGBTQIA+literaturesurveybiasinnaturallanguageprocessingintersectionalitystakeholderinvolvementdiversityethics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey aims to establish what natural language processing (NLP) research about LGBTQIA+ communities actually looks like, by systematically collecting and annotating papers from the field's main publication archive through early 2026. The authors argue that queer NLP is growing but lopsided: most papers are reactive—they expose bias in existing models—rather than proactive, and they lean heavily on template-based evaluations and data augmentation. The paper also finds that the research is dominated by English and Western European languages, rarely takes an intersectional view, and seldom involves queer stakeholders despite citing harm to queer communities as motivation. The survey matters because it turns scattered findings into a map of gaps and a call for future work that is multilingual, community-centered, and not limited to post-hoc fixes.

What carries the argument

The load-bearing apparatus is the survey's annotation framework: each paper is coded for NLP task, method, explicit queer groups, harms addressed, language and geographic region of the data, whether intersectionality is addressed, and whether stakeholders are involved. The central distinction is 'reactive' versus 'proactive' research—detecting bias in systems that exist versus building new systems or solutions—and the paper uses this contrast to organize its findings and future agenda.

What would settle it

A concrete check is to rerun the annotation on an expanded corpus: search beyond the main archive with additional identity keywords (including community terms in languages such as Arabic, Hindi, Portuguese, and Japanese) and include regional conference proceedings and preprint servers. If the expanded sample shows proactive mitigation papers are as common as reactive bias papers, or that non-English papers form a large share of the work, then the survey's central claims about a predominantly reactive, English-centric field would not survive.

Watch

Extended reading notes

Core claim

The central claim is that queer NLP, as represented in the main NLP archive, has three structural features: a reactive orientation in which identifying bias far outpaces mitigating it; a language distribution heavily weighted toward English; and a systematic absence of stakeholder involvement and intersectionality, even though most papers name queer communities as their motivation. The authors reach this by annotating papers along task, method, named queer groups, harms, language and region of data, treatment of intersectionality, and stakeholder engagement. They conclude that these gaps are not random but reflect methods—template prompting, augmentation, post-hoc debiasing—that are convenie

Load-bearing premise

The load-bearing premise is that the keyword-based search of the main publication archive, expanded by the authors' familiarity with the field, yields a representative and complete set of queer NLP papers; if many relevant papers use different terminology, appear in non-English venues, or exist only as preprints, then the quantitative claims about English dominance and the reactive/proactive split would be skewed.

Editorial extensions

If this is right

  • If the field is as reactive as the survey shows, existing bias benchmarks are an incomplete evidence base for understanding queer harm in NLP.
  • Community-in-the-loop evaluation, such as building benchmarks with queer annotators, would have to shift from a rare exception to a standard practice.
  • The survey's examples of non-English shared tasks—like hope-speech detection and slur-reclamation—provide concrete templates for multilingual queer NLP that could be adapted beyond those languages.
  • A shift toward proactive mitigation would reallocate research effort from reporting model failures to designing systems that support queer users from the start.
  • The paper's structural suggestions—allowing users to opt out of classification and building dynamic evaluations—imply that future systems should accommodate fluid identities and shifting context over time.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Our inference: the reactive pattern may be reinforced by publishing incentives, since bias findings are comparatively easy to benchmark, and the survey does not examine this incentive structure.
  • Our inference: because the survey is tied to one publication archive, it may undercount proactive and non-English work published in regional venues or as preprints, so the true gap could be smaller than reported.
  • Our inference: the paper's discussion of 'critical refusal' points to a limit of inclusion—some users may want systems that decline to classify them at all, a design principle that goes beyond better benchmarks.
  • Our inference: a direct extension would apply the same annotation protocol to non-archive venues and preprint repositories; a finding that proactive work is common there would qualify the survey's conclusions as archive-specific.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper presents a systematic, PRISMA-style survey of NLP research addressing LGBTQIA+ topics in the ACL Anthology. The authors describe a keyword search, manual filtering, and community-based expansion resulting in a corpus of 86 ACL papers (the abstract states n=122). They annotate papers along axes including task, method, queer groups, harms, language, intersectionality, and stakeholder involvement; group them into seven categories; and find that most work is reactive (bias discovery), English-centric, and lacking stakeholder/intersectional engagement. The paper also offers qualitative discussion and future directions from queer theory.

Significance. If the corpus is accepted as representative, the survey is a valuable and timely map: it is, to my knowledge, the first systematic survey of queer NLP across the ACL Anthology, and its descriptive statistics give a concrete basis for claims about Anglocentrism and reactive research. Strengths include the community-led author team, explicit positionality, the open-source living repository, and the careful distinction between ACL and non-ACL works. The qualitative sections on refusal and non-English venues are thought-provoking. However, the quantitative findings inherit the reproducibility limitations described below.

major comments (4)
  1. [Abstract/§3.1] The sample size is inconsistent: the Abstract states n=122 and 'all such papers published in the ACL Anthology', while §3.1 says the final survey covers 86 ACL papers. Figure 2/Table 1 report 66 English papers, and 66/86 = 76.7%, so the headline percentages are computed on 86, not 122. The discrepancy is not explained (e.g., the 122 could include non-ACL papers, duplicates, or pre-2025 search results). Since the paper's central quantitative claims depend on the denominator, the authors must reconcile these numbers and state which n underlies each statistic.
  2. [§3.1 Paper Selection] The paper selection is not fully reproducible despite the PRISMA label. The keyword list (queer, aromantic, gender non*, lgbt*, agender, glbt, lesbian, gay, bisexual, transgender) omits common alternative terms such as nonbinary, homophobia, transphobia, gender-neutral, and sexual orientation. From 3,864 search hits, manual filtering yielded 55 papers; the authors then 'expanded' the list based on personal knowledge and Semantic Scholar and added 19 ACL 2025 papers. No PRISMA flow diagram, inclusion/exclusion log, or complete list of the 86 papers is provided in the manuscript (the repository is cited, but the manuscript should be self-contained). Because the headline claims about English-centrism, reactive research, and stakeholder gaps are proportions over this corpus, a non-reproducible or biased selection could change the conclusions. The Limitations section acknowledges this risk, b
  3. [§3.2 Annotation Process] Inter-rater reliability is computed on only 11 of 86 papers, and the reported agreement is 90.09% with Cohen's kappa = .79 for stakeholder involvement and .62 for intersectionality. Kappa = .62 is moderate, not high. Intersectionality is one of the three headline gap claims (Figure 1 and §5.3). With single annotation for the other 75 papers and inductive coding used to form the seven categories, the reliability evidence is too thin to support the strength of the intersectionality and stakeholder-gap claims. Please report per-axis kappas with confidence intervals, the full contingency tables, and a larger reliability sample, or soften the corresponding quantitative claims.
  4. [§5.1/§5.2] The classification of papers as 'reactive' versus 'proactive' appears to be interpreted by the annotators, but the coding scheme for this distinction is not specified in §3.2 or §5.2. Since the Abstract's central claim is that 'most papers take a reactive rather than a proactive approach,' the operational definition of this binary (e.g., which categories count as proactive, how mitigation-only vs. evaluation-only papers are coded) should be stated explicitly and supported by reliability evidence.
minor comments (5)
  1. [§5.1] 'queer ALC papers' should be 'queer ACL papers'.
  2. [References] 'Retrivied' is misspelled (Human Rights Campaign, Nonbinary Wiki entries).
  3. [Figure 2] The caption should state explicitly that percentages refer to the share of papers, not languages, and that the total exceeds 100% because a paper can cover multiple languages.
  4. [Table 1] The header 'Language # Singletons (1 each)' is confusing; consider separating singleton count from the language list.
  5. [§3.2] Please clarify whether 'language diversity' reliability is perfect because all 11 papers were coded identically, and report the number of categories considered.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's central claims are descriptive aggregates of its own annotated corpus, not fitted predictions; self-citations are part of the surveyed literature and not load-bearing.

full rationale

The paper is a systematic literature survey. Its headline findings (reactive vs. proactive research, English-centric skew, stakeholder gaps) are computed by annotating the collected papers (Section 3.2) and are not derived from any equation, fitted parameter, uniqueness theorem, or empirical prediction. The corpus is constructed via keyword search and author expansion (Section 3.1), but this is a selection method, not a circular derivation; the paper explicitly labels it as a limitation: "our initial paper search utilized a list of queer identity keywords... might lead to a selection that misses publications from other venues, other non-English venues and queer aspects not captured by our chosen keywords." The abstract's n=122 versus the 86 ACL papers in Section 3.1 is an accounting inconsistency and a transparency issue, not a reduction of a conclusion to an input. The survey cites many works by its own authors (Devinney, Gautam, Lauscher, Nozza, Ovalle, Subramonian, et al.), but those citations are used as examples within the surveyed literature and as supporting discussion; the central gap findings come from annotation of the whole set, not from those papers' conclusions alone. No load-bearing step reduces to self-citation or to the definition of a variable. Therefore no circularity is present.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

A survey has no free parameters or invented entities. Its load-bearing assumptions are about scope and annotation reliability, which affect all reported percentages and gap claims.

assumptions (3)
  • domain assumption The ACL Anthology is a sufficiently representative corpus of NLP research on queer topics.
    The survey's scope is limited to ACL, and the authors acknowledge in §6.2 and Limitations that non-ACL venues may host relevant work. This assumption directly shapes all counts.
  • domain assumption The keyword list (e.g., queer, aromantic, gender non*, lgbt*, etc.) captures the relevant queer NLP literature.
    In §3.1, the authors use a keyword list from Taylor et al. (2024). They admit in Limitations that queer aspects not captured by the keywords may be missed, which could bias gap estimates.
  • domain assumption Annotator judgments are reliable enough for aggregate statistics.
    Most papers are single-annotated; only 11 papers are double-annotated for inter-rater reliability (§3.2). Cohen's Kappa is .79 and .62 for stakeholder involvement and intersectionality, but other axes are not validated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Queer NLP: A Critical Survey on Literature Gaps, Biases and Trends." pith.science (2026). https://pith.science/paper/U6JAEXWB

@misc{pith2026260216151,
  author       = {Pith},
  title        = {Pith review of: Queer NLP: A Critical Survey on Literature Gaps, Biases and Trends},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/U6JAEXWB}},
  note         = {Machine review of arXiv:2602.16151}
}
read the original abstract

Natural language processing (NLP) technologies are rapidly reshaping how language is created, processed, and interpreted by humans. With current and potential applications in hiring, law, healthcare, and other areas that impact people's lives, understanding and mitigating harms towards marginalized groups is critical. In this survey, we examine NLP research papers that explicitly address the relationship between LGBTQIA+ communities and NLP technologies. We systematically review all such papers published in the ACL Anthology up until February 2026 (n=122), to answer the following research questions: (1) What are current research trends? (2) What gaps exist in terms of topics and methods? (3) What areas are open for future work? We find that while the number of papers on queer NLP has grown within the last few years, most papers take a reactive rather than a proactive approach, focusing on shortcomings of existing systems rather than creating new solutions. Our survey uncovers many opportunities for future work, especially regarding stakeholder involvement, intersectionality, interdisciplinarity, and languages other than English. We also offer an outlook from a queer studies perspective, highlighting understudied topics and blind spots in the harms addressed in NLP papers. Beyond being a roadmap of what has been done, this survey is a call to action for work towards more just and inclusive NLP technologies.

Figures

Figures reproduced from arXiv: 2602.16151 by the authors.

Figure 1
Figure 1. The majority of queer NLP papers pub￾lished in the ACL Anthology are focused on En￾glish, disregard intersectionality and omit stake￾holders. 1 Introduction Natural language processing (NLP) is a field that focuses on the interaction between computers and human language, aiming for machines to analyze, understand, and generate natural language. As NLP technologies become increasingly integrated into everyday life, u… view at source ↗
Figure 2
Figure 2. ), geographical region, and domain of the data used, whether intersectionality is explicitly addressed,2 the involvement of any stakeholders and limitations named in the paper. We then group all papers into the 7 categories used in Section 4 using inductive coding. Papers can belong to mul￾tiple categories. While each paper is only read and annotated by a single person, to calculate inter￾rater reliability we random… view at source ↗
Figure 3
Figure 3. Comparing paper categories by (a) intersectionality, (b) language diversity, and (c) stakeholder [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    GRUFF dataset shows LLMs agree well with masculine and feminine German pronouns but fail on neopronouns and distractors, with occupational stereotypes poorly correlated across cases.

  2. Do Language Models Know Their Slang? Queer Slang Understanding in User-Generated Content

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A new dataset and evaluation show that LLMs define queer slang better when given domain framing and sentential context, but still below a paraphrase-based reference bound.

  3. The queer Hero versus the Fool bias of the queer trait: An archetypometric analysis of the collective portrayal of queerness in fictional stories

    cs.CY 2026-07 conditional novelty 6.0 of 10

    Characters with the highest queer scores present as Heroes/Adventurers, but the straight-queer trait itself loads strongly toward Fool across stories.

Reference graph

Works this paper leans on

17 extracted references · 4 linked inside Pith · cited by 3 Pith papers

  1. [7]

    Rosemary A

    Overview of HOPE at IberLEF 2023: Multilingual hope speech detection.Proce- samiento del Lenguaje Natural, 71:371–381. Rosemary A. Joyce. 2015. Heteronormativity. In The International Encyclopedia of Human Sexu- ality, pages 501–581. John Wiley & Sons. Os Keyes. 2019. Counting the countless. Pub- lished in the online magazineReal Life. Retriv- ied October...

  2. [8]

    InProceedings of the 29th International Conference on Computational Linguistics, pages 1221–1232, Gyeongju, Re- public of Korea

    Welcome to the modern world of pro- nouns: Identity-inclusive natural language pro- cessing beyond gender. InProceedings of the 29th International Conference on Computational Linguistics, pages 1221–1232, Gyeongju, Re- public of Korea. International Committee on Computational Linguistics. Anne Lauscher, Debora Nozza, Ehm Miltersen, Archie Crowley, and Dir...

  3. [9]

    InProceedings of the First Work- shop on Cross-Cultural Considerations in NLP (C3NLP), pages 16–24, Dubrovnik, Croatia

    A cross-lingual study of homotransphobia on Twitter. InProceedings of the First Work- shop on Cross-Cultural Considerations in NLP (C3NLP), pages 16–24, Dubrovnik, Croatia. As- sociation for Computational Linguistics. Christina Lu and David Jurgens. 2022. The sub- tle language of exclusion: Identifying the toxic speech of trans-exclusionary radical femini...

  4. [11]

    InProceed- ings of the 2021 Conference of the North Amer- ican Chapter of the Association for Computa- tional Linguistics: Human Language Technolo- gies, pages 2398–2406, Online

    HONEST: Measuring hurtful sentence completion in language models. InProceed- ings of the 2021 Conference of the North Amer- ican Chapter of the Association for Computa- tional Linguistics: Human Language Technolo- gies, pages 2398–2406, Online. Association for Computational Linguistics. Debora Nozza, Federico Bianchi, Anne Lauscher, and Dirk Hovy. 2022. M...

  5. [12]

    I’m fully who I am

    HODI at EV ALITA 2023: Overview of the first shared task on homotransphobia detec- tion in Italian. InProceedings of the Eighth Evaluation Campaign of Natural Language Pro- cessing and Speech Tools for Italian. Final Work- shop (EVALITA 2023), volume 3473 ofCEUR Workshop Proceedings. CEUR-WS.org. Anaelia Ovalle, Palash Goyal, Jwala Dhamala, Zachary Jagger...

  6. [13]

    In2024 IEEE Spoken Language Tech- nology Workshop (SLT), pages 526–532

    Beyond the binary: Limitations and pos- sibilities of gender-related speech technology research. In2024 IEEE Spoken Language Tech- nology Workshop (SLT), pages 526–532. Vinícius Santos, Felipe Henriques, and Gustavo Guedes. 2022. O discurso de ódio homofóbico no Twitter a partir da análise de dados. In Anais do XI Brazilian Workshop on Social Net- work An...

  7. [14]

    democratization

    Queer Waves: a German speech dataset capturing gender and sexual diversity from pod- casts and YouTube. InProceedings of Inter- speech 2025, pages 679–683, Rotterdam, The Netherlands. Atli Sigurgeirsson and Eddie L. Ungless. 2024. Just because we camp, doesn’t mean we should: The ethics of modelling queer voices. InPro- ceedings of Interspeech 2024, pages...

  8. [15]

    InProceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, CCS ’24, pages 1196–1210, Salt Lake City, Utah, USA

    GenderCARE: A comprehensive frame- work for assessing and reducing gender bias in large language models. InProceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, CCS ’24, pages 1196–1210, Salt Lake City, Utah, USA. Associa- tion for Computing Machinery. Jordan Taylor, Ellen Simpson, Anh-Ton Tran, Jed R. Brubaker, Sarah E...

Show all 17 references
  1. [16]

    InPro- ceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ’24), Hon- olulu, Hawaii, USA

    Cruising queer HCI on the DL: A litera- ture review of LGBTQ+ people in HCI. InPro- ceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ’24), Hon- olulu, Hawaii, USA. Association for Computing Machinery. Joshua Tint. 2025. Guardrails, not guidance: U...

  2. [17]

    unnatural or dangerous

    A rewriting approach for gender inclusiv- ity in Portuguese. InFindings of the Association for Computational Linguistics: EMNLP 2023, pages 8747–8759, Singapore. Association for Computational Linguistics. Andreas Waldis, Joel Birrer, Anne Lauscher, and Iryna Gurevych. 2024. Th...

  3. [2016]

    John Benjamins Publishing Com- pany

    What is pejoration, and how can it be ex- pressed in language? In Rita Finkbeiner, Jörg Meibauer, and Heike Wiese, editors,Pejoration, pages 1–18. John Benjamins Publishing Com- pany. Michel Foucault. 1978.The History of Sexuality, Volume 1: An Introduction. Pantheon Books, Ne...

  4. [2018]

    stereotype

    Hurtlex: A multilingual lexicon of words to hurt. InProceedings of the Fifth Italian Con- ference on Computational Linguistics (CLiC-it 2018), volume 2253 ofCEUR Workshop Pro- ceedings, Torino, Italy. CEUR-WS.org. Gemma Bel-Enguix, Helena Gómez-Adorno, Ger- ardo Sierra, Juan V...

  5. [2021]

    In Proceedings of the 3rd Conference on Conversa- tional User Interfaces, CUI ’21, Bilbao (online), Spain

    LGBTQ-AI? Exploring expressions of gender and sexual orientation in chatbots. In Proceedings of the 3rd Conference on Conversa- tional User Interfaces, CUI ’21, Bilbao (online), Spain. Association for Computing Machinery. Paul Engelmann, Peter Trolle, and Christian Hard- meier...

  6. [2022]

    sexuality

    Gender spectrum speech corpus. OR- TOLANG (Open Resources and TOols for LAN- Guage). Sourojit Ghosh and Aylin Caliskan. 2023. Chat- GPT perpetuates gender bias in machine trans- lation and ignores non-gendered pronouns: Findings across Bengali and five other low- resource lang...

  7. [2023]

    In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5352–5367, Toronto, Canada

    MISGENDERED: Limits of large lan- guage models in understanding pronouns. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5352–5367, Toronto, Canada. Association for Computational Linguistics. Dirk Hovy ...

  8. [2024]

    In Proceedings of the 2024 Conference on Empir- ical Methods in Natural Language Processing, pages 3078–3105, Miami, Florida, USA

    From insights to actions: The impact of interpretability and analysis research on NLP. In Proceedings of the 2024 Conference on Empir- ical Methods in Natural Language Processing, pages 3078–3105, Miami, Florida, USA. Asso- ciation for Computational Linguistics. Daniel Najafal...

  9. [2025]

    Razvan Amironesei and Mark Diaz

    Charting the landscape of African NLP: Mapping progress and shaping the road ahead.Computing Research Repository, arXiv:2505.21315. Razvan Amironesei and Mark Diaz. 2023. Re- lationality and offensive speech: A research agenda. InThe 7th Workshop on Online Abuse and Harms (WOA...

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.