{"id":"589bbae0-5489-4c48-a2ca-f39aabf9d18a","arxiv_id":"2505.12054","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A 14-probe benchmark on 12 LLMs finds consistent gender stereotype reasoning and unbalanced character representation across models.","lead":"GenderBench is a new open-source suite of 14 tests that measures gender bias in large language models. It evaluated 12 current chatbots and found that they repeat gender stereotypes in creative writing and show signs of favoritism in hiring decisions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Free-text gender probes silently drop gender-neutral outputs; if coverage rates vary across models, the reported convergence in 'character generation' metrics may be an artifact of pronoun-detection selection.","rationale":"The reader's weakest_assumption about probe validity and simple rules is correct in spirit, but I identified a more specific and, I think, more damaging mechanism: the free-text character-generation probes compute gender representation and stereotypical alignment only on outputs that contain a detectable gendered pronoun. The paper never discloses what fraction of outputs is discarded. If that fraction is large or varies across models, the headline convergence could be a shared artifact of how modern LLMs handle pronoun use under these particular prompts, rather than evidence of common bias in the full distribution of generated characters. This concern is load-bearing because the 'creative writing is the most affected use case' result and the 'disproportionate number of female characters' observation come directly from these probes. A conditional acceptance is appropriate: the authors should report coverage rates and either include gender-neutral outputs or explicitly rescope the claim to 'gendered outputs'. The paper is otherwise transparent, open-source, and thoughtfully limited, which supports a conditional rather than a rejection verdict. My recommendation keeps the reader's CONDITIONAL verdict unchanged, while adding a concrete verification step that would settle whether the concern is real.","tokens_in":14178,"tokens_out":7568,"duration_ms":84238,"concrete_test":"Run JobsLum, Inventories, and GestCreative on a representative subset of at least five models and, for each model-probe pair, compute the fraction of generated profiles that contain at least one gendered pronoun. Then recompute masculine_rate and stereotype_rate by (a) treating gender-neutral outputs as a separate third category and (b) imputing a neutrality-adjusted denominator. If any model-probe pair has coverage below 80%, or if the cross-model ranking or the sign of the gender imbalance changes materially under either recomputation, the convergence claim in Section 3.2 is an artifact of pronoun-detection selection. Report the coverage fractions alongside the metrics in the paper.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—striking cross-model convergence with consistent weaknesses in character generation—rests substantially on probes GestCreative, Inventories, and JobsLum (Section 2.3), which detect character gender 'by observing pronouns'. Outputs without a gendered pronoun are excluded from the gendered classification, yet the paper does not report the fraction of generated profiles that receive a gender classification per model or probe. Because newer instruction-tuned models are often explicitly trained to avoid gendered pronouns or to use singular 'they', the excluded subset can be large and model-dependent. The reported masculine_rate and stereotype_rate are therefore conditional statistics computed only over the gendered subset, and the denominators shrink or become biased in ways that can make models appear more similar than they are. If, for example, one model produces mostly neutral profiles and only uses a female pronoun for 'nurturing', while another uses gendered pronouns widely but with a different balance, the two can receive similar stereotype_rate values even though their full output distributions differ. This is a concrete, load-bearing instantiation of the paper's own concession that evaluations rely on 'simple, high-precision rules and heuristics' (Section 2.2); the Limitations section on Prompts does not address this selection-bias mechanism. Without per-model coverage reporting or a handling of gender-neutral outputs, the 'consistent weaknesses' in representational harms may be an artifact of the pronoun-detection pipeline rather than a property of the LLMs' actual generated texts.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces GenderBench, an open-source evaluation suite for gender bias in LLMs, consisting of 14 probes that produce 19 harm metrics across three categories: outcome disparity, stereotypical reasoning, and representational harms. The authors evaluate 12 LLMs from different providers and sizes, using rule-based scoring rather than LLM-as-a-judge, and report bootstrapped confidence intervals along with a four-tier severity scale. The main empirical claim is a striking convergence across LLMs: models consistently struggle with stereotypical reasoning and with equitable gender representation in free-text character generation, while performing better in some decision-making and affective-computing tasks. The paper also reports a tendency toward female preference in creative writing and some decision scenarios.","tokens_in":1417,"tokens_out":1350,"duration_ms":46714,"significance":"If the measurements are valid, GenderBench is a useful community resource: it is open-source, includes a large prompt set (60,469 prompts), spans multiple previously separate bias-evaluation methodologies, and avoids the reproducibility problems of LLM-as-a-judge by using explicit rules. The decomposition of gender bias into separately measured harms is a valuable framing, and the publication of the library and raw evaluation infrastructure supports reproducibility. The empirical claims about cross-model convergence and the 'jagged frontier' of gender-bias severity are interesting and falsifiable. However, the validity of the central convergence claim depends on the measurement assumptions in the free-text probes, and those assumptions are not currently documented with enough detail.","major_comments":[{"comment":"The free-text probes detect character gender 'by observing pronouns,' but the paper never reports the fraction of generated profiles that contain a gendered pronoun, nor the per-model coverage rate. The metrics masculine_rate and stereotype_rate are therefore computed only over the gendered subset. Because instruction-tuned models are often explicitly trained to avoid gendered pronouns or to use singular 'they,' the excluded subset can be large and model-dependent. This makes cross-model comparisons of representational-harm metrics conditional on an unmeasured selection variable. For example, a model that writes mostly gender-neutral profiles and only uses a female pronoun for 'nurturing' could receive a similar stereotype_rate to a model that uses gendered pronouns widely but with a different balance. I request per-model coverage statistics for each free-text probe, an explicit treatment of gender-neutral outputs (e.g., a separate 'neutral' category or a defined handling rule), and a sensitivity analysis of masculine_rate and stereotype_rate to the coverage definition. Without this, the claim that creative writing is 'the most affected use case' and the associated convergence evidence are not fully supported.","section":"Section 2.3 (GestCreative, Inventories, JobsLum) and Section 3.2, Figure 1"},{"comment":"The paper acknowledges that most probes use only one prompt template, and this is a real threat to the central generalization claim. Since the same template is used for every model, the observed cross-model convergence could reflect a common sensitivity to that particular wording rather than a stable property of the models. The limitation is stated in the Limitations section, but it is not quantified or bounded. I ask for at least a small multi-template sensitivity analysis on a subset of probes (e.g., varying the wording of GestCreative, Inventories, and one decision-making probe) or for the conclusions to be explicitly restricted to the exact prompts used, with correspondingly weaker generalizations about LLM behavior in general.","section":"Limitations (Prompts); Section 2.3"},{"comment":"The paper states that most probes report how many prompts failed to elicit a valid response, but none of these counts appear in the paper. This is especially important for the free-text probes discussed above and for multiple-choice probes where a model might answer outside the allowed options. I request that the per-probe and per-model valid-response rates be reported, and that the metrics be examined for sensitivity to the inclusion or exclusion of invalid responses.","section":"Section 3.1 and Section 2.2"}],"minor_comments":[{"comment":"The text says 'e.g., gpt-4 model with HiringBloomberg probe,' but the evaluated models are gpt-4o and gpt-4o-mini; please specify which model is meant.","section":"Section 3.2, Figure 1"},{"comment":"The normalization procedure used to project metrics to [0,1] is not defined. Please state whether the normalization is per-probe, per-model, or global, and describe the computation of Pearson correlations in Figure 2 (e.g., number of metrics, whether they are averaged across models).","section":"Table 2 and Figure 2"},{"comment":"There is a typo: 'traits associted with masculinity' should read 'traits associated with masculinity.'","section":"Section 2.3, Inventories"},{"comment":"The probe name 'BusinessVocabulary' is split across lines as 'BusinessV ocabulary' in several places; this should be fixed.","section":"Table 1"},{"comment":"Generation parameters use temperature 1 and top-p 1, which are high-variance settings. It would be helpful to state whether the reported bootstrap intervals account for the sampling randomness from these settings, and whether any deterministic decoding was used for comparison.","section":"Section 3.1"},{"comment":"The four-tier severity thresholds are described as subjective and based on expert judgment, but the threshold values themselves are not reported in the paper. Please include the exact thresholds for all metrics, preferably in an appendix, so that the color-coded figures are interpretable and the severity labels can be audited.","section":"Section 2.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of the journal and the self-citation to GEST is reasonable. The main concern is that the central convergence claim depends on coverage and denominator behavior in the free-text probes, which is fixable with additional reporting and analysis. I would be willing to look at a revised version that includes per-model coverage statistics, a handling of gender-neutral outputs, and at least a minimal prompt-sensitivity check."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"GenderBench is worth your time. It bundles 14 gender-bias probes into one open-source library, applies them to 12 models under a single protocol, and reports bootstrapped confidence intervals. The integrated comparison is genuinely new, and the choice to score outputs with simple rules rather than an LLM judge is a real plus for reproducibility. The headline finding — that models converge in their bias patterns — is plausible and interesting.\n\nThe soft spot is real, though. Three free-text probes (GestCreative, Inventories, JobsLum) classify the gender of generated characters by looking for pronouns. Outputs without a gendered pronoun are silently dropped from the gendered classification, but the paper never reports how many outputs fall into that drop per model or probe. Newer instruction-tuned models are often trained to avoid gendered pronouns or to use singular “they,” so the excluded subset can be large and model-dependent. That makes masculine_rate and stereotype_rate conditional statistics over an unknown, possibly shrinking denominator. Two models with very different full output distributions could end up with similar scores just because one excludes most outputs and the other does not. That does not sink the whole paper, but it does mean the “striking convergence” in representational harms may be partly an artifact of the pronoun-detection pipeline. The Limitations section mentions prompt sensitivity but not this selection mechanism, so it needs an explicit fix: per-model coverage reporting, and ideally a handling of gender-neutral outputs.\n\nAlso worth flagging is the claim to be “the most detailed and complete assessment of gender biases in LLMs to date.” That is too strong for a suite that uses one prompt template per probe and a convenience sample of 12 models, even though the limitations section is otherwise candid. A versioned code release and baseline comparisons to the source studies for each probe would help too.\n\nWho this is for: people building fairness benchmarks, model evaluators, and governance folks who want a single scorecard. It deserves serious peer review, not a desk reject. The benchmark is useful and the convergence claim is valuable if it survives the coverage analysis. Recommendation: send to review with a request for coverage reporting and a softened completeness claim.","headline":"A useful, open-source gender-bias benchmark with a plausible convergence finding, but the pronoun-based gender detection silently drops a potentially model-dependent subset of outputs, which could inflate the apparent convergence.","tokens_in":14992,"tokens_out":2666,"would_cite":true,"duration_ms":25972,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Twelve contemporary LLMs converge on the same gender-bias profile: stereotypical reasoning and skewed character generation recur across providers and sizes, while decision and emotion tasks look largely healthy.","keywords":["GenderBench","gender bias","LLM evaluation","benchmark suite","stereotypical reasoning","representational harms","outcome disparity","gender stereotypes"],"falsifier":"Re-run all 14 probes on the same 12 models with several paraphrased templates per probe, and with updated occupation and trait inventories. If severity tiers swing substantially across templates — for example, a model that is catastrophic on occupation-based stereotypical reasoning under one phrasing becomes healthy under another — the claimed convergence would be shown to be an artifact of shared prompt sensitivity rather than a stable behavioral trait. If the same probes stay in the same severity tiers across templates and models, the convergence claim is confirmed; a supporting check is recomputing the cross-model correlation matrix per harm category and seeing whether the smaller models keep showing weaker alignment on the same probes.","tokens_in":13939,"feed_emoji":"⚖️","tokens_out":14680,"duration_ms":119688,"temperature":0.7,"pith_summary":"GenderBench is an open-source evaluation suite that measures gender biases in LLMs through 14 pre-packaged probes covering 19 harmful behaviors, from hiring and medical decisions to creative writing. The paper's central empirical claim is that twelve LLMs — spanning different providers and sizes — display what it calls \"a striking convergence\": they consistently rely on stereotypical reasoning (picking stereotype-consistent answers even when the facts do not support them) and produce skewed gender representation in generated characters, while decision-making and emotion-attribution tasks come out largely healthy. Alongside the measurements, the paper contributes a decomposable methodology: each harmful behavior gets its own metric and a four-tier severity label (healthy, cautionary, critical, catastrophic), so bias is reported dimension by dimension rather than as a single score. If the convergence claim is right, gender-bias behavior is not idiosyncratic to individual models but a shared property of current training practices, and a compact probe suite can monitor it as new models are released.","feed_headline":"12 LLMs share one gender-bias profile","feed_subtitle":"A 14-probe suite shows the same weak spots in every model — and the same healthy spots.","key_machinery":"The carrying mechanism is the probe, defined as a self-contained, pre-packaged experiment: a fixed set of prompts plus an evaluation methodology that scores outputs with simple, high-precision rules — multiple-choice, yes/no, and constrained natural-language formats — deliberately avoiding any machine model as judge. Each probe yields one or more metrics that quantify a specific harmful behavior, and each metric maps to a four-tier severity scale (healthy, cautionary, critical, catastrophic) whose thresholds encode an egalitarian standard under which any unfair gender difference counts as harm. The GenderBench harness bundles 14 probes totaling 60,469 prompts, repeats prompts with minor variations such as shuffled answer order, and computes bootstrapped confidence intervals to stabilize measurements. The decomposition itself does the conceptual work: because each behavior is measured independently, the suite can separate areas of ubiquitous weakness from areas of relative strength within the same model, which is what makes the cross-model convergence visible.","core_discovery":"The paper's central claim is that gender bias in LLMs is best understood as a decomposable collection of measurable behaviors, and that measured this way, twelve current LLMs converge on the same profile: the same weaknesses recur across providers and model sizes, while the same areas look healthy. In particular, all evaluated models exhibit stereotypical reasoning and skewed gender representation in character generation, with creative-writing probes (character profiles built from traits, mottoes, or occupations) showing the largest bias; decision-making probes such as hiring and medical diagnosis come out mostly healthy, with isolated exceptions. The paper also reports a directional pattern of preferential treatment for women — female characters are generated more often, women are favored in relationship-conflict judgments, and they receive a slight advantage in some decision scenarios — and suggests this convergence reflects standardization in training methodology. A companion claim is that publication bias toward positive findings has obscured areas of relative strength, so the suite deliberately includes probes where models perform well.","pith_inferences":["A direct testable extension the paper leaves implicit: rerunning the 14 probes with several paraphrased templates per probe would show whether the cross-model convergence survives wording changes or is partly an artifact of shared prompt sensitivity — the paper itself flags the single-template design as a stated limitation.","The reported preferential treatment for women has an unstated consequence for fairness practice: if simple parity is the goal, some models would need adjustment in the pro-woman direction, which is why direction-sensitive metrics rather than absolute disparities are the more informative quantities.","Because the stereotype, occupation, and trait inventories are anchored in contemporary Western norms, a neighbouring study could re-run the same harness with culturally different inventories; finding the same weak spots would strengthen the convergence claim, while finding different ones would bound it.","The paper's deliberate avoidance of LLM-as-a-judge and reliance on constrained output formats leaves open whether the convergence reflects latent associations or forced-choice behavior; letting models answer the same probes in free text would separate those two possibilities."],"forward_implications":["If the convergence holds, a compact suite like GenderBench can serve as a monitoring instrument: rerunning the probes on newly released models should predict where bias will appear, letting developers and auditors check specific behaviors instead of designing evaluations from scratch.","The jagged-frontier point the paper makes implies that a healthy result on any covered probe cannot certify a model as unbiased; the paper states this explicitly as 'non-existence of proof is not a proof of non-existence,' so certifications must be framed as coverage-limited.","The observed preferential treatment for women implies that mitigations aimed at restoring parity — and debates about what neutral behavior means — must account for a bias direction opposite to the historically assumed male-centric one.","The finding that creative-writing and occupation-based character generation carry the strongest stereotypical reasoning implies that content-generation and business-communication applications are the most likely deployment contexts for gender-biased outputs to surface, not high-stakes classifiers.","The paper's argument that per-prompt alignment tuning does not address global behavioral properties such as corpus-level gender representation implies that correcting these biases will require different interventions than current alignment pipelines provide."],"supporting_citations":[{"why":"Supplies the BBQ dataset of ambiguous two-character scenarios that the Bbq probe adapts to measure how often stereotypical reasoning wins over logic.","marker":"Parrish et al., 2022"},{"why":"Provides the reference-letter generation setup and the gendered vocabulary inventories behind the BusinessVocabulary probe.","marker":"Wan et al., 2023"},{"why":"Provides the discrimeval high-stakes decision scenarios (loans, rentals) used by the DiscriminationTamkin probe.","marker":"Tamkin et al., 2023"},{"why":"Provides the DiversityMedQA medical questions that the DiversityMedQa probe re-genders to measure accuracy disparity across patient genders.","marker":"Rawat et al., 2024"},{"why":"Supplies the hiring-decision experiment that the HiringAn probe extends across roles and candidate genders.","marker":"An et al., 2024"},{"why":"Provides the CV-ranking protocol that the HiringBloomberg probe uses to detect candidate-selection disparity.","marker":"Yin et al., 2024"},{"why":"Gives the occupation-to-character generation task behind the JobsLum probe.","marker":"Lum et al., 2025"},{"why":"Provides the relationship-conflict scenarios and the gender-swap protocol behind the RelationshipLevy probe.","marker":"Levy et al., 2024"},{"why":"Contributes the GEST dataset of stereotypical statements and mottoes used by the Gest and GestCreative probes.","marker":"Pikuliak et al., 2024"}],"fun_headline_variants":["12 LLMs, one gender-bias fingerprint","LLMs pass high-stakes, fail creative gender tests","Same gender bias weak spots across all 12 LLMs","Gender bias in LLMs: consistent weak and strong areas","All 12 LLMs share identical gender-bias profile"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that each probe's single prompt template, rule-based scoring, and stereotype and occupation lists faithfully capture the gender harm it claims to measure; if the wording of one template triggers or suppresses biased behavior, or if the lists are outdated or culturally narrow, the severity labels and the cross-model convergence could change.","fun_headline_variants_meta":{"raw":{"variants":["12 LLMs, one gender-bias fingerprint","LLMs pass high-stakes, fail creative gender tests","Same gender bias weak spots across all 12 LLMs","Gender bias in LLMs: consistent weak and strong areas","All 12 LLMs share identical gender-bias profile"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000273,"raw_usage":{"total_tokens":1571,"prompt_tokens":814,"completion_tokens":757,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":430,"completion_tokens_details":{"reasoning_tokens":678}},"tokens_in":430,"tokens_out":757,"duration_ms":7625,"temperature":1.0,"reasoning_tokens":678,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:41:29.621880+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run all 14 probes on the same 12 models with several paraphrased templates per probe, and with updated occupation and trait inventories. If severity tiers swing substantially across templates — for example, a model that is catastrophic on occupation-based stereotypical reasoning under one phrasing becomes healthy under another — the claimed convergence would be shown to be an artifact of shared prompt sensitivity rather than a stable behavioral trait. If the same probes stay in the same severity tiers across templates and models, the convergence claim is confirmed; a supporting check is recomputing the cross-model correlation matrix per harm category and seeing whether the smaller models keep showing weaker alignment on the same probes.","supporting_citations":[],"review_version":1}