{"id":"7c0ac383-0eaf-4484-83b1-c84dfa640698","arxiv_id":"2411.17338","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM bias scores change depending on whether 'unbiased' means equal treatment across groups or close alignment with US workforce statistics.","lead":"This paper proposes a new way to measure bias in large language models: instead of only checking whether a model treats demographic groups equally, it checks whether the model's answers match real-world workforce statistics. The authors show that models rank differently under these two standards, and argue that bias evaluation should use several criteria at once.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MS cannot certify statistical alignment: the perfect-alignment line is Score=(1-r)(2*Stat-1), so both slope and intercept depend on the refusal rate; the paper's fixed beta=2 target is misspecified and confounds refusal with alignment.","rationale":"The paper's descriptive point that bias rankings depend on the evaluation criterion is supported by Table 2, and the authors provide code and acknowledge limitations. My concern targets the core metric: the perfect alignment condition is a line with both slope and intercept, and the slope target itself depends on the refusal rate, which the paper measures but does not incorporate. This is a concrete, fixable issue rather than a rejection of the multi-criteria framing. The reader noted the missing intercept as a secondary assumption; I elevate it and extend it to the refusal-rate dependence because it affects all model evaluations and the human-survey interpretation, not just the survey wording. I recommend no change to the reader's CONDITIONAL verdict: the authors can resolve the concern by reporting fitted intercepts, using a distance-to-line alignment error, and adjusting the target slope by the refusal rate. Doing so would also clarify whether the survey result actually supports the 'statistically aligned' interpretation.","tokens_in":13534,"tokens_out":12585,"duration_ms":119327,"concrete_test":"For each model in Table 2 and for the human survey, re-fit Score(x)=beta*Statistics(x)+beta0 on the same occupations and report the fitted intercept beta0. Then compare against the refusal-adjusted perfect line beta=2(1-MR), beta0=-(1-MR), using each model's own refusal rate, and re-rank models by RMSE to that line. If the model with MS closest to 2 (e.g., GPT-4o mini on the gender persona task) has beta0 far from -(1-MR) or a large RMSE, slope-only MS does not measure statistical alignment and the paper's ranking by MS is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 defines MS as the slope beta of Score(x)=beta*Statistics(x)+beta0 and claims that 'the perfectly statistically aligned line ... has beta of 2.' This target ignores two features of the paper's own setup. First, the regression line has an intercept: with Statistics(x)=female ratio f and Score(x)=P(female)-P(male), perfect alignment gives Score=(1-r)(2f-1), i.e., slope 2(1-r) and intercept -(1-r), where r is the refusal rate on that item. Only when r=0 is the slope 2. Second, because MR is estimated per model (0.255 for humans, up to 0.489 for Llama2 7B Chat), models with different refusal rates should not be compared against the same slope target. A model that refuses often will have a lower MS even if its non-refusal choices perfectly track demographics. The paper's comparison of MS to the constant 2, and its interpretation of GPT-4o mini's MS about 2 as 'strong alignment,' is therefore not justified by the metric's definition. The human-survey result (MS=1.414) is also ambiguous: with MR=0.255 the refusal-adjusted target is 2(1-0.255)=1.49, so the conclusion depends on which target is used. Because MS is the proposed fact-based metric, this is a load-bearing correctness issue for the central claim that statistical alignment is a distinct and valid evaluation axis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes three criteria for evaluating social bias in LLM responses: balance (MB), refusal (MR), and a new 'statistical alignment' score (MS), defined as the slope of the regression of a gendered or age-related stereotype score on real-world demographic ratios from U.S. Bureau of Labor Statistics. The authors report a 58-participant survey in which respondents chose the 'least objectionable' occupation for a male or female persona, and they use the resulting MB, MR, and MS to argue that people prefer non-refusing, statistically aligned LLM outputs. They then evaluate 14 open-weight models and three GPT models on a WinoBias coreference task and a persona-based occupation selection task, finding that model bias rankings differ across the three metrics. The paper concludes that bias assessment should be pluralistic and that MS is a valid complementary metric.","tokens_in":13840,"tokens_out":18958,"duration_ms":178065,"significance":"If the proposed metric and survey evidence were sound, this paper would make a useful contribution by offering an external, fact-based anchor for LLM bias evaluation that is complementary to equality-based measures. The core descriptive finding that rankings change across MB, MR, and MS is supported by Table 2, and the use of external BLS statistics and released code are strengths. However, the interpretation of MS as 'statistical alignment' is compromised by the paper's treatment of refusal, and the human-preference claim is not supported by the survey design. Both are load-bearing for the paper's central claims, so the manuscript requires substantive revision rather than minor polishing.","major_comments":[{"comment":"The claim that the perfect alignment line has slope beta = 2 is not derived from the task's choice structure and is not correct under a natural accounting of UNKNOWN responses. In the two-occupation-plus-UNKNOWN setup, the selection probabilities in each context must sum to 1 - r, where r is the refusal rate. Under a simple complementary-pair model, the aligned score is Score(x) = (1 - r)(2f_x - 1), giving slope 2(1 - r) and intercept -(1 - r); only r = 0 gives slope 2. For the actual occupation pairs in Table 9, the female ratios are not complementary (e.g., Carpenter 0.03 paired with Designer 0.55), so the target is not even a line with slope 2 in general. As a result, Table 2's 'MS (≈ 2)' column compares models with very different refusal rates against a single constant: GPT-4o mini (MR = 0.001) is scored against a target near 2, while Llama2 7B Chat (MR = 0.489) would need a target near 1.02 under the complementary-pair model. The human result is also ambiguous: with MR = 0.255 the adjusted target is about 1.49, so MS = 1.414 is close to that target, whereas the paper's fixed target of 2 makes it look far. The authors should specify a generative model for choices, adjust for refusal, report beta0 and fit diagnostics, and revisit conclusions that depend on comparing MS to 2, including the 'over-debiasing' interpretation in Section 5.2.","section":"Section 3 (Scoring Metrics) and Table 2"},{"comment":"The survey does not measure what the abstract claims. The prompt in Table 10 asks participants to 'choose the response that feels the least objectionable' to the question 'What is a more suitable job for a male/female living in the United States?' with options such as Carpenter, Designer, and Not Sure. This is a forced choice among occupations; it does not present actual LLM outputs, refusal messages, or different generation styles, and it does not ask participants to rate alignment or preference for model behavior. The resulting low MR and high MS therefore describe respondents' own choices and stereotypes (or social desirability), not their perception of LLM outputs. The sentence in Section 4 that 'humans tend to prefer LLM responses that avoid refusal yet are statistically aligned' is an overclaim. In addition, Table 1 reports only point estimates for 58 participants, with no confidence intervals or significance tests, so the distance from MS = 1.414 to any target is unquantified. A survey that presents participants with actual model outputs and asks for preference or acceptability ratings would be needed to support the claim.","section":"Section 4 and Appendix A.3"},{"comment":"The MR metric conflates refusal with genuine ambiguity in the coreference task. The WinoBias split used is explicitly the ambiguous subset, where the pronoun has no unique referent; the paper's own prompt includes UNKNOWN as a possible answer to 'determine who the pronoun refers to.' Selecting UNKNOWN in such cases is an accurate response to ambiguity, not necessarily an act of refusing to answer. Yet MR is labeled 'refusal' throughout, and coreference MR values such as Llama2 7B Chat's 0.489 are interpreted as approximating a 'refusing state.' The same index is more plausibly interpreted as refusal in the persona task, where the model is asked to choose an occupation for a persona. To keep the equality-based criterion meaningful, the paper should separate ambiguity-driven UNKNOWN from refusal-driven UNKNOWN, or at least justify why the coreference UNKNOWN count is treated as refusal.","section":"Section 5.1 and Table 2"}],"minor_comments":[{"comment":"The notation P(x|x, g1) is confusing because x appears both as the event and as the conditioning variable; use a clearer notation such as P(x | g1, item) or define the conditioning on the occupation pair explicitly.","section":"Section 3"},{"comment":"The instruction-tuned model list names 'Mistral-Plus-7B', but Table 2 reports 'Mistral 7B Instruct'; these names should be reconciled, and the model's exact checkpoint should be clarified.","section":"Appendix A.2"},{"comment":"The paper says 'we processed the option with the highest logit value as the model's choice,' but for GPT-3.5, GPT-4, and GPT-4o mini accessed through the OpenAI API, raw logits are not generally available; the paper should state how these choices were obtained (e.g., logprobs, repeated sampling, or text parsing).","section":"Appendix A.2"},{"comment":"No standard errors, confidence intervals, or per-prompt counts are reported for MB, MR, or MS, which makes it difficult to interpret small differences such as MS = 0.161 versus 0.185 for Llama3 8B and Llama2 13B.","section":"Table 2"},{"comment":"The expert/non-expert differences (MR 0.320 versus 0.206, MS 1.345 versus 1.452) are reported without statistical tests; with 25 and 33 participants these differences may not be meaningful.","section":"Section 4 and Table 1"},{"comment":"The sentence 'responses previously considered biased may actually be statistically aligned and, therefore, may not be biased' is a normative conclusion that goes beyond the metric; it should be rephrased as 'less biased under the statistical-alignment criterion' to avoid implying that alignment with current occupational segregation is an absence of bias.","section":"Section 5.2"}],"recommendation":"major_revision","confidential_remarks":"The paper has an interesting descriptive result: Table 2 does show that model bias rankings differ across MB, MR, and MS. However, the two headline contributions need substantive work. The MS target must be corrected or redefined to account for refusal, and the human-preference claim needs either a redesigned survey or a substantial reinterpretation. The coreference MR conflation is also a correctable but nontrivial issue. I would not reject outright, but this is more than a minor revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi [Colleague],\n\nYou should know this paper before using its metric. The core idea—measure LLM bias under multiple criteria (balance, refusal, statistical alignment) and show rankings shift—is genuinely useful. The specific finding that a model can look unbiased on MB and MR but anti-stereotypical on MS is worth being aware of. The code is public and the evaluation across many models is thorough.\n\nThe catch is the statistical alignment metric, MS. The paper defines MS as the slope of Score(x) = β·Statistics(x) + β0 and claims the perfect-alignment line has β = 2. That is only true when the refusal rate MR is zero. If a respondent refuses with probability r, the observed selection probability is (1-r) times the fully observed probability, so Score(x) = (1-r)(2f-1). The slope is 2(1-r) and the intercept -(1-r), not slope 2. Since MR varies across models (0.001 for GPT-4 to 0.489 for Llama2 7B Chat), comparing all MS scores against the same constant 2 systematically rewards low-refusal models and punishes high-refusal ones. The human result (MS=1.414 with MR=0.255) is actually close to the refusal-adjusted target of 1.49, but the paper interprets it against 2. This is a load-bearing issue, not a quibble.\n\nSecond soft spot: the human survey does not actually ask participants to evaluate LLM outputs. It asks them to choose the least objectionable occupation for a male/female persona. The claim that 'humans prefer statistically aligned LLM outputs' is an overreach. The survey tells you something about people's stated preferences for occupations, not about their perception of LLM responses.\n\nThird, there are no error bars on any of the metrics. With 58 participants and 20 occupations, the regression slope has substantial uncertainty. That is easily fixed but should not be ignored.\n\nWhere does this leave the paper? The descriptive multi-criteria finding holds: model rankings do depend on the criterion. The paper is worth a serious referee because this framing is important and the metric is fixable. The authors should either redefine MS to be refusal-adjusted (e.g., estimate r and compare slopes against 2(1-r)), report confidence intervals, and soften the survey claim to match their design. I would not cite the MS metric as published, but I would follow revisions.\n\nRecommendation: send to peer review; flag the beta=2 problem early.","headline":"A useful multi-criteria bias framing undermined by a confounded regression target: MS conflates refusal with alignment.","tokens_in":14338,"tokens_out":4414,"would_cite":false,"duration_ms":54615,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM bias is criterion-dependent: the paper proposes a fact-based statistical alignment metric (MS) that can rank a model as unbiased under one criterion and biased under another.","keywords":["LLM bias","statistical alignment","bias metrics","fact-based criteria","human survey","gender bias","age bias","real-world statistics"],"falsifier":"Re-run the preference survey with actual LLM responses at different MS values while holding refusal rate fixed: if respondents do not prefer outputs near MS = 2 over outputs near MS = 1.4, the human-preference justification fails. Separately, recompute model rankings using distance to the full line Score = 2·Statistics − 1 (slope and intercept) instead of slope alone; a large reordering would show that the metric's target is under-specified.","tokens_in":13312,"feed_emoji":"⚖️","tokens_out":5296,"duration_ms":48965,"temperature":0.7,"pith_summary":"This paper argues that there is no single correct answer to whether a language model is biased. It proposes measuring bias against three distinct criteria: balanced output across demographic groups, refusal to answer stereotype-laden questions, and statistical alignment with real-world demographic ratios. The central new claim is that statistical alignment, quantified by a regression slope MS, captures a bias dimension that the other two criteria miss. A 58-person survey suggests people prefer generated answers that avoid refusal yet stay close to real-world distributions, and evaluating open models plus commercial APIs shows that the same model can rank best under one criterion and worst under another. If accepted, the paper establishes that bias assessments must report multiple criteria rather than a single score.","feed_headline":"Same AI looks biased or fair depending on the yardstick","feed_subtitle":"New metric measures whether model choices match real-world demographics; bias rankings flip across criteria.","key_machinery":"The central machinery is the trio of scores MB, MR, and MS, computed from a two-option-plus-UNKNOWN choice design. The occupation-specific score Score(x) separates association by group; MB averages |Score(x)| to measure balance; MR measures refusal frequency; and MS, the slope of Score(x) regressed on real-world gender or age ratios from US Bureau of Labor Statistics, measures statistical alignment, with β = 2 as the target. The key identity is the regression equation Score(x) = β·Statistics(x) + β0, which converts raw choice ratios into a number comparable across tasks and models.","core_discovery":"The paper's core claim is that bias in LLM outputs is criterion-dependent, and that a fact-based criterion called statistical alignment should join equality-based criteria in bias evaluation. For each occupation x, the paper computes Score(x) = P(x|x,g1) − P(x|x,g2), the difference between how often a male/female (or youth/elderly) context selects that occupation. MB is the mean absolute score, with a balanced target of 0; MR is the rate of UNKNOWN or refusal choices; and MS is the slope β of the regression Score(x) = β·Statistics(x) + β0, with β = 2 defined as the perfectly statistically aligned line. On their survey, human respondents scored MB = 0.479, MR = 0.255, and MS = 1.414, which the authors read as evidence that people prefer outputs that avoid refusal while largely tracking real-world ratios. Across models, the same system can appear balanced or biased depending on the metric; for instance, GPT-series models score high on MS while scoring poorly on equality-based balance, and RLHF can push models below zero MS into anti-stereotypical territory.","pith_inferences":["Editorial: if MS gains traction, 'debiasing' would be reframed from equalizing outputs to calibrating them to demographic baselines, which could legitimate some currently stereotype-aligned outputs as unbiased.","Editorial: the slope-only definition of alignment could be tested by fitting the same data with the intercept fixed at -1, matching the perfect line Score = 2·Statistics − 1; if rankings shift, the target is under-specified.","Editorial: the human survey did not present actual alternative LLM responses, so a direct preference test between a model at MS ≈ 2 and one at MS ≈ 1.4, matched for refusal rate, would isolate whether people truly prefer statistical alignment or merely dislike refusal.","Editorial: extending MS to non-binary gender, multi-ethnic, or non-US demographic categories is a natural next step, but it requires re-deriving the perfect-alignment target, since slope 2 is specific to the two-group choice design."],"forward_implications":["If the central claim holds, a model cannot be called 'the least biased' without specifying the criterion, since rankings under MB, MR, and MS diverge for the same model.","Instruction-tuning and RLHF can improve balance or refusal while simultaneously moving a model away from statistical alignment, sometimes into anti-stereotypical responses with negative MS.","Human preference data imply that safety-style refusal is not the default ideal: with a low refusal rate and MS closer to 2 than to 0, respondents accepted biased-sounding but statistically plausible answers.","Fact-based bias assessment is bounded by the statistics it uses: the paper's gender and age categories are binary and US-based, so MS values are not directly portable to other populations.","The trade-off between MB and MS is structural, because aligning with real-world skew conflicts with treating groups equally, so both metrics should be reported together."],"supporting_citations":[{"why":"Supplies the real-world gender and age ratios from US Bureau of Labor Statistics that define the statistical alignment target for MS.","marker":"[39]"},{"why":"Provides the ambiguous coreference sentences, occupation pairs, and occupation list used in both the model tasks and the human survey.","marker":"[45]"},{"why":"Supplies the social-science definition of stereotype accuracy that motivates the fact-based criteria and the MS metric.","marker":"[16]"},{"why":"Provides the persona-assignment instruction templates used in the occupation selection task.","marker":"[13]"},{"why":"Represents the equality/refusal-oriented bias benchmark that the paper contrasts with its statistical alignment approach.","marker":"[29]"}],"fun_headline_variants":["Bias is in the eye of the criterion","LLM bias flips with fact-based vs equality metrics","Statistical alignment uncovers hidden bias trade-offs","Same model, many bias verdicts depending on the metric","Fact-based lens reveals bias is criterion-dependent"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 58-person survey measures a preference for statistical alignment rather than participants' own stereotypes or a desire to avoid awkward refusals, and that the slope-2 line fully encodes the statistically aligned state.","fun_headline_variants_meta":{"raw":{"variants":["Bias is in the eye of the criterion","LLM bias flips with fact-based vs equality metrics","Statistical alignment uncovers hidden bias trade-offs","Same model, many bias verdicts depending on the metric","Fact-based lens reveals bias is criterion-dependent"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000196,"raw_usage":{"total_tokens":1377,"prompt_tokens":975,"completion_tokens":402,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":329}},"tokens_in":591,"tokens_out":402,"duration_ms":4062,"temperature":1.0,"reasoning_tokens":329,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:13:32.357804+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the preference survey with actual LLM responses at different MS values while holding refusal rate fixed: if respondents do not prefer outputs near MS = 2 over outputs near MS = 1.4, the human-preference justification fails. Separately, recompute model rankings using distance to the full line Score = 2·Statistics − 1 (slope and intercept) instead of slope alone; a large reordering would show that the metric's target is under-specified.","supporting_citations":[{"cited_title":"Labor Force Statistics from the Current Population Survey","cited_arxiv_id":null,"evidence_quote":"Supplies the real-world gender and age ratios from US Bureau of Labor Statistics that define the statistical alignment target for MS."},{"cited_title":"Gender bias in coreference resolution: Evaluation and debiasing methods","cited_arxiv_id":null,"evidence_quote":"Provides the ambiguous coreference sentences, occupation pairs, and occupation list used in both the model tasks and the human survey."},{"cited_title":"Definition and assessment of accuracy in social stereotypes","cited_arxiv_id":null,"evidence_quote":"Supplies the social-science definition of stereotype accuracy that motivates the fact-based criteria and the MS metric."},{"cited_title":"Bias Runs Deep: Implicit reasoning biases in persona-assigned LLMs","cited_arxiv_id":null,"evidence_quote":"Provides the persona-assignment instruction templates used in the occupation selection task."},{"cited_title":"BBQ: A hand-built bias benchmark for question answering","cited_arxiv_id":null,"evidence_quote":"Represents the equality/refusal-oriented bias benchmark that the paper contrasts with its statistical alignment approach."}],"review_version":1}