{"id":"97c97864-5091-4c83-a5cd-315244ca0267","arxiv_id":"2505.11795","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across five LLMs, sexism classification agrees more with female than male annotators, and demographic persona prompting does not reliably shift agreement toward the target group.","lead":"This paper tests whether large language models can be steered by demographic instructions to mimic different human perspectives on sexist tweets. It finds that the models already agree more with female annotators than male ones, and that adding persona instructions changes agreement only inconsistently.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The female-bias result in Table 2 may be an artifact of higher within-group agreement among female annotators rather than of LLM alignment with female judgments.","rationale":"The reader's CONDITIONAL verdict is the right overall level: the raw α values are internally consistent, but the interpretive claims outrun the methods. I agree with the reader that α between an LLM and a small demographic subgroup is a fragile operationalization, and that the paper's 'incapable of adopting personas' phrasing in Section 2 is contradicted by its own Table 3, which shows several sizable improvements (e.g., GPT-4 with age personas). My concern is a different, more fundamental mechanism: the female-vs-male α comparison may be driven by differing within-group agreement among human annotators, not by LLM alignment with female judgments. The paper's check on gender label proportions does not address this. The bootstrap confidence intervals reported only as 'smaller than 0.001' are implausibly tight and should be reported numerically; this is a reporting flaw but secondary given the size of the gender gap. The contamination footnote in Section 1 is internally inconsistent about the public availability of labels, but it does not affect the computations. No code or supplementary material is provided, so reanalysis currently cannot be independently run. For these reasons, I would keep the reader's CONDITIONAL verdict: accept only after a reanalysis that separates LLM-to-individual agreement from within-group human agreement.","tokens_in":8767,"tokens_out":13872,"duration_ms":154221,"concrete_test":"Recompute Table 2 using only LLM–human pairs: for each tweet, compute pairwise agreement (or per-annotator Cohen's κ) between the LLM label and each individual female and male annotator, then average within gender. In the same run, report within-gender Krippendorff's α for the female-only and male-only annotator sets. If the female advantage persists in the pairwise LLM-to-individual comparison while within-gender α is comparable, the bias claim is supported; if it disappears or reverses, Table 2's pattern is a group-variance artifact. This reanalysis requires only the EXIST labels already used and the stored LLM outputs, so it is a direct check rather than a new experiment.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central RQ1 claim is that all five LLMs agree more with female than male annotators, interpreted in Section 5 as a female-skewed perspective. The supporting measure is Krippendorff's α between the LLM label sequence and the demographic subgroup's label multiset. Section 3 does not specify how multiple human annotations per tweet enter α; if the standard form is used, α's observed-disagreement term includes pairwise disagreements among the female annotators themselves and among the male annotators themselves, because each tweet has three female and three male labels. α_F and α_M are therefore functions of LLM–human agreement and of within-group human agreement. The paper's t-test on gender means (p = 0.237) only rules out a marginal difference in sexist-label proportions; it does not rule out higher internal consistency among female annotators. If female raters are more homogeneous than male raters, α_F can exceed α_M even when the LLM's labels are no closer (or are closer) to male judgments. The headline 'LLMs exhibit bias toward certain demographics' is thus not yet separated from a label-variance artifact. This is more fundamental than the persona-adoption validity issue: it affects RQ1 itself, and the motivation for RQ2 depends on RQ1.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether LLMs exhibit demographic bias when classifying tweets as sexist, and whether demographic persona prompting can mitigate such bias. Using the EXIST 2023 dataset, which contains multiple annotations per tweet stratified by gender and age, the authors compute Krippendorff's alpha between each of five LLMs (GPT-3.5, GPT-4, GPT-4o, Mistral, Qwen) and human annotator subgroups. They report that all five models agree more with female annotators than with male annotators, and that age patterns are inconsistent. They then add gender- and age-specific persona instructions to the prompts and find that the effects on agreement are mixed and unpredictable, concluding that demographic-based persona prompting cannot be relied upon to mitigate bias in sexism classification.","tokens_in":8886,"tokens_out":5780,"duration_ms":54935,"significance":"If the results are correct, the paper provides a valuable empirical result: LLM-based sexism moderation inherits a female-skewed judgment perspective, and simple demographic persona prompting is not a dependable debiasing intervention. The study's strengths include the use of a multi-annotator perspectivist dataset, evaluation across five models spanning both API and open-source families, and explicit bootstrap resampling. However, the central measure of agreement is confounded by within-group annotator agreement, so the headline RQ1 claim requires re-analysis. If the pattern persists after controlling for this confound, the paper would be an important contribution to perspectivist IR and fairness evaluation.","major_comments":[{"comment":"The Krippendorff's alpha values comparing an LLM to a demographic subgroup are computed on a reliability matrix that contains multiple human annotations per tweet (three female and three male; two per age group). In the standard formulation, the observed disagreement includes pairwise disagreements among the human annotators within the subgroup. If female annotators are more internally consistent than male annotators, alpha_F can exceed alpha_M even when the LLM's labels are no closer to female judgments. The reported cross-group human alpha of 0.477 does not control for this. The authors must recompute alignment using an LLM–human pairwise measure (e.g., average alpha between the LLM and each individual annotator) or a model that separates within-group human disagreement from LLM–human agreement. Without this, the RQ1 claim that LLMs 'exhibit bias toward female annotators' is not supported.","section":"Section 3, 'Bias Analysis'; Tables 2 and 3"},{"comment":"The claimed consistency across all five models rests on the confounded alpha values. For example, GPT-3.5 alpha_F = 0.415 versus alpha_M = 0.371, a difference of 0.044. The paper states that all bootstrap confidence intervals are 'smaller than 0.001' but provides no interval values or details on the bootstrap procedure for the differences. The authors should report the confidence intervals for the alpha differences and, more importantly, for the recomputed pairwise measure suggested above, so that the significance of the gender pattern can be assessed.","section":"Section 4, first paragraph; Table 2"},{"comment":"The statement that the LLMs 'seemed to be incapable of doing so' (adopting personas) is too strong relative to the evidence. The experiments use one specific prompt template and a single set of demographic labels; failure to move agreement with the EXIST annotators could result from a weak persona intervention or from the metric's insensitivity, rather than from a fundamental inability to adopt personas. The paper's more cautious formulation in Section 5 -- 'inconsistent and unpredictable outcomes... cannot be relied upon' -- is defensible and should replace the 'incapable' framing throughout.","section":"Section 1 and Section 5"}],"minor_comments":[{"comment":"The sentence 'All of the measured Krippendorff's alpha coefficients had confidence intervals smaller than 0.001' is ambiguous; please clarify whether this refers to interval width and report the bootstrap method and at least one concrete interval.","section":"Section 3, 'Bias Analysis'"},{"comment":"The text says the training and development subsets were merged, but Table 1 reports 7,958 total tweets; please clarify whether this is the train+dev size and how the test set was handled.","section":"Section 3, 'Data'"},{"comment":"The prompt development process is unclear: three candidate prompts are first tested on 20 tweets, then the prompt is further optimized with o1-preview; please state explicitly which version was used as the baseline prompt in all experiments.","section":"Section 3, 'Prompt Creation'"},{"comment":"Please describe how the cross-group human alpha of 0.477 was computed (e.g., treating all six annotators as coders, or aggregating per group); this will help readers interpret the between-group consistency and the subsequent LLM comparisons.","section":"Table 2"},{"comment":"The age-group subscripts in the model names, such as GPT-3.518−22, are visually ambiguous; consider using a clearer notation, for example GPT-3.5[18-22], in the table and text.","section":"Table 3 caption"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of SIGIR and the topic is timely, but the central methodological issue -- the confounding of Krippendorff's alpha with within-group annotator agreement -- must be addressed before the RQ1 conclusion can be accepted. Please require the authors to provide a re-analysis using a pairwise LLM–human agreement measure and to report whether the female-alignment pattern survives."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nShort version: the paper is a clean, small empirical study that deserves a referee's time, but its headline claim — that five LLMs are biased toward female annotators — is not actually settled by the data as analyzed. The stress-test concern lands. Table 2 is consistent with a female-bias story, but also with the more boring possibility that female annotators in EXIST simply agree with each other more than male annotators do. Krippendorff's alpha between an LLM and a demographic group is computed over the group's multiple labels; if the group is internally homogeneous, alpha goes up even when the LLM's labels are no more similar to that group's judgments. The paper's t-test on mean sexist-label proportions (p = 0.237) doesn't address within-group agreement. So the central RQ1 conclusion, and the motivation for RQ2, rest on an unexamined confound. That's the soft spot, and it's not minor.\n\nWhat's genuinely new: nobody else has run this five-model persona comparison on EXIST 2023, and the specific observation that GPT-3.5, GPT-4, GPT-4o, Mistral, and Qwen all show higher alpha with female annotators is a useful data point — if the reanalysis separates alignment from label variance. The persona-prompting result is also fairly stated most of the time: inconsistent, unpredictable, not reliable. Only the word \"incapable\" in the introduction overreaches; the tables show mixed effects, not no effect. The paper also engages the right prior literature on persona prompting.\n\nMinor issues: all bootstrap CIs are reported as smaller than 0.001 without actual values, which is implausibly tight and unverifiable. The prompt was selected on twenty tweets from the same training/development data used for evaluation, which is a mild selection concern. And the footnote asserting the dataset could not have been used as LLM training is at odds with the public availability of EXIST 2023 train/dev splits. None of those break the raw measurements.\n\nBottom line: someone building or evaluating sexism detection systems, and anyone working on perspectivist evaluation, will want to know this exists and then see the follow-up. I wouldn't cite the female-bias claim as established. But the paper deserves serious peer review — a knowledgeable reviewer can ask for the subgroup-agreement decomposition, and that is exactly what the process should do.","headline":"The female-bias headline is not settled: the alpha computation mixes LLM alignment with demographic-group internal agreement, so the central claim may be a variance artifact; still, this is a useful small study that deserves peer review.","tokens_in":9548,"tokens_out":2832,"would_cite":false,"duration_ms":30806,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM sexism labels skew female; persona prompts don't fix it","keywords":["sexism detection","perspectivism","large language models","demographic bias","persona prompting","annotation disagreement","content moderation","Krippendorff alpha"],"falsifier":"Run the same experiment on a held-out panel of annotators per demographic group, using several phrasings of the same persona instruction; if instructed-persona models improve agreement with their target group more often than they lower it across models and phrasings, the paper's conclusion that demographic personas cannot be relied on is overturned.","tokens_in":8483,"feed_emoji":"♀️","tokens_out":9897,"duration_ms":87751,"temperature":0.7,"pith_summary":"The paper asks whether large language models can be steered into judging tweets from the perspective of a particular demographic group, and whether those models already carry a demographic slant. Using the EXIST 2023 collection, where each tweet carries labels from six annotators stratified by gender and age, the authors find that five LLMs (GPT-3.5, GPT-4, GPT-4o, Mistral, and Qwen) all agree more with female annotators than with male annotators when deciding whether a tweet is sexist. Age patterns are less consistent, with different models aligning with different age groups. The paper then tests demographic 'persona' instructions that tell the model to adopt a gender and age group, and finds the effects inconsistent and unpredictable: some models move toward the instructed group, others move away. The authors conclude that demographic-based persona prompting cannot be relied on as a debiasing method for sexism classification.","feed_headline":"LLM sexism labels skew female; persona prompts don't fix it","feed_subtitle":"Across five LLMs, female-annotator agreement is higher, and persona instructions shift results unpredictably.","key_machinery":"The load-bearing object is the EXIST 2023 collection, which preserves annotation disagreement: each of 7,958 tweets is labeled by six annotators stratified by gender and age, so a demographic group's perspective can be read directly from its label distribution. The paper's measure is Krippendorff's $\\alpha$, an inter-annotator agreement coefficient, computed between each model's labels and each demographic subgroup's labels, with bootstrap confidence intervals below 0.001. The intervention is persona prompting: a prompt that states 'your demographic information is: sex X, age group Y' before asking the model to classify tweets as sexist or not. That setup lets the paper ask whether the instruction moves the model's labels toward the instructed group's labels.","core_discovery":"The central discovery is that LLM-based sexism classification is not perspective-neutral: every one of the five tested models agrees more closely, in Krippendorff's $\\alpha$, with the labels of female annotators than with male annotators on the EXIST 2023 sexism-detection task. For example, GPT-3.5 reaches $\\alpha=0.415$ against female annotators and $0.371$ against male annotators, while GPT-4o reaches $0.228$ against female and $0.191$ against male annotators. Age-group alignment is model-specific: GPT-3.5 and Mistral align most with the 46+ group, GPT-4 and Qwen with 23–45, and GPT-4o with 18–22. Instructing the model to adopt a demographic persona changes agreement in both directions, improving it for some model–group pairs and worsening it for others, so the paper concludes that demographic-based persona prompting cannot be relied on to mitigate bias.","pith_inferences":["A testable extension of the gender finding is that the same female-leaning gap should appear in other subjective classification tasks, such as hate-speech or sentiment labeling; if it does not, the effect may be specific to sexism content rather than a general demographic default.","The persona results are tied to one optimized prompt template, so it remains open whether richer persona descriptions, in-context examples, or instruction-tuned variants would produce reliable persona adoption; that would narrow, rather than overturn, the paper's negative conclusion.","The authors' framing implies a practical diagnostic the paper does not develop: moderation systems could report which demographic group a model's judgments align with, and treat large deviations from the female-aligned default as a trigger for human review.","If the female-aligned default reflects pretraining-data composition rather than something about sexism, then models trained with different data mixes should show different gap sizes; comparing models of the same family trained on different data would test that."],"forward_implications":["An LLM-based moderation pipeline that labels sexist tweets will, under these results, reproduce a perspective closer to the female annotators in EXIST, meaning the set of flagged tweets will differ from what a male-leaning system would flag.","Adding a demographic instruction is not a reliable way to steer that perspective: outcome changes are model- and group-specific, so persona prompting should not be treated as a debiasing control.","Because the age group with highest agreement differs by model, claims about LLM demographic bias on this task need to be model-specific rather than general.","When LLMs are used as stand-ins for human annotators, their demographic alignment is part of the measurement and should be reported; otherwise reported quality scores may actually measure closeness to one group's perspective."],"supporting_citations":[{"why":"It supplies the EXIST 2023 collection with tweets labeled by annotators stratified by gender and age, preserving disagreement.","marker":"[18]"},{"why":"It defines the EXIST 2023 task and the soft-labeling framework the paper uses for perspectivist evaluation.","marker":"[19]"},{"why":"It provides the perspectivist approach that motivates keeping diverse annotations instead of gold-standard labels.","marker":"[10]"},{"why":"It demonstrates that personas in system prompts do not consistently improve LLM performance, which the paper extends to demographic personas in sexism detection.","marker":"[28]"},{"why":"It shows that LLM-generated personas carry demographic biases, motivating the persona experiments.","marker":"[22]"},{"why":"It positions LLMs as human-machine annotation collaborators, the setting in which the paper's bias findings matter.","marker":"[9]"}],"fun_headline_variants":["LLM sexism detectors skew female, personas don't fix","Persona prompts fail to debias LLM sexism labels","All five LLMs side with female annotators on sexism","Demographic prompts shift LLM bias unpredictably","Sexism AI ignores persona, favors female views"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that agreement between the model's labels and one demographic subgroup's labels, measured by a single alpha value, is a valid sign both of alignment with that demographic and of whether the model actually adopted the instructed persona; a model could take on a persona and still disagree with the specific, small group of annotators in EXIST.","fun_headline_variants_meta":{"raw":{"variants":["LLM sexism detectors skew female, personas don't fix","Persona prompts fail to debias LLM sexism labels","All five LLMs side with female annotators on sexism","Demographic prompts shift LLM bias unpredictably","Sexism AI ignores persona, favors female views"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000113,"raw_usage":{"total_tokens":1006,"prompt_tokens":828,"completion_tokens":178,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":444,"completion_tokens_details":{"reasoning_tokens":98}},"tokens_in":444,"tokens_out":178,"duration_ms":2364,"temperature":1.0,"reasoning_tokens":98,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:47:47.613704+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same experiment on a held-out panel of annotators per demographic group, using several phrasings of the same persona instruction; if instructed-persona models improve agreement with their target group more often than they lower it across models and phrasings, the paper's conclusion that demographic personas cannot be relied on is overturned.","supporting_citations":[{"cited_title":"Kurita, N","cited_arxiv_id":null,"evidence_quote":"It supplies the EXIST 2023 collection with tweets labeled by annotators stratified by gender and age, preserving disagreement."},{"cited_title":"Plaza, J","cited_arxiv_id":null,"evidence_quote":"It defines the EXIST 2023 task and the soft-labeling framework the paper uses for perspectivist evaluation."},{"cited_title":"Frenda, G","cited_arxiv_id":null,"evidence_quote":"It provides the perspectivist approach that motivates keeping diverse annotations instead of gold-standard labels."}],"review_version":1}