{"id":"2291b2ee-7134-4e03-8645-cfdd50632b4d","arxiv_id":"2505.14080","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across 16 open-weight language models, gender is encoded as a sex-linked binary, transgender and nonbinary terms are less probable than random nouns, and mental illness is more probable after those terms; these effects grow with model size.","lead":"This paper tests how 16 language models associate gender with biological sex and illness, and shows they mostly tie gender to a male/female binary while giving transgender and nonbinary terms low probability and linking them to mental illness. Larger models match these narrow gender stereotypes more strongly, which matters for how AI systems treat gender-diverse people in real applications.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unnormalized Folk-Subversive LPR (Eq. 1) confounds token-frequency base rates with learned sex-gender associations; the Figure 2 scaling trend may reflect LM fluency, not stronger conflation of sex and gender.","rationale":"The paper's most novel and policy-relevant quantitative claim is that sex-gender conflation increases with model size, used to argue that scaling alone will not fix (and may worsen) essentialist gender constructs. That claim rests entirely on Folk-Subversive LPR, which is not normalized by term frequency or a neutral baseline. Since larger models are better at predicting frequent words like 'a man' and assigning tiny probabilities to rare identity terms, a positive and increasing LPR can be produced without any learned association between testosterone and 'man' or estrogen and 'woman'. The paper already contains the right control, the matched-guise Sex-Gender LPR in Eq. 2, but does not use it for the scaling analysis. Testing the trend with an aggregated matched-guise metric is straightforward and decisive. The pathologization findings use matched-guise comparisons for each illness and are less exposed to this confound, and the open-weight reproducible setup is a strength; however, the scaling conclusion is central to the abstract, so the conditional verdict is appropriate until the base-rate control is run.","tokens_in":24186,"tokens_out":16347,"duration_ms":157596,"concrete_test":"Recompute Figure 2 using a base-rate-controlled aggregate of the matched-guise Sex-Gender LPR (Eq. 2): for each model, average over the five male-female context pairs the folk-aligned quantities log P(a man|male s) - log P(a man|female s') and log P(a woman|female s) - log P(a woman|male s'), then correlate with model size (Spearman). If the correlation drops below significance or within-family monotonicity breaks, the Figure 2 trend is a token-frequency artefact rather than evidence of stronger sex-gender conflation; if the trend persists, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline scaling result (Figure 2, rho = 0.89) is computed from Folk-Subversive LPR (Eq. 1), which averages raw log conditional probabilities log P(g|Context(s)) with +1 for folk-aligned pairs and -1 for subversive pairs. This metric does not subtract the marginal log probability of g or compare against a neutral context. Because 'a man' and 'a woman' are high-frequency completions after any 'The person...' prefix, while 'genderqueer' and 'two-spirit' are rare tokens, the LPR is positive even for a model with no sex-gender association; larger models produce sharper next-token distributions, so the gap can widen purely from better language modeling. The matched-guise Sex-Gender LPR (Eq. 2) controls for base rates by comparing the same gender term across matched male/female contexts, but Figure 2 is not built from an aggregated version of it. In addition, Eq. 1 as printed divides by |G| while summing 10 positive and 60 negative terms, which would overweight the rare subversive terms by a factor of six; the reported range 0.42-3.22 suggests the implemented metric differs from the displayed formula, further weakening the reproducibility of the size trend.","agreement_with_reader":"agree"},"referee_report":null,"author_rebuttal":null,"desk_editor":null,"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-07T15:40:57.094512+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}