REVIEW 5 major objections 5 minor 34 references
Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models
T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Persona prompts barely move MMLU scores, but the same models rate female answers more leniently and advise lower salaries for women, minorities, and refugees; the paper argues the pay gap is the more salient bias probe.
desk verdict Salary-advice bias is real and worth reporting; the 'surface fairness' null is under-powered, so the paper's headline contrast is only half established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a three-rung experimental ladder that escalates the social stakes of the same underlying question set: (1) persona-prompted MMLU accuracy, (2) grading a persona-typed user's answer as correct or not, and (3) soliciting a specific dollar salary-negotiation figure. The ladder is what carries the argument: it isolates knowledge-expertise effects from social-evaluation effects and from socio-economic advice effects using the same models, personae, and MMLU-derived fields. The statistical machinery is McNemar tests for paired accuracy and grading comparisons, Mann-Whitney and Kruskal-Wallis tests for salary distributions, and Bonferroni correction for multiple comparisons; the compound-persona construction (taking the highest- and lowest-salary personae and merging them) is the mechanism that shows biases combine rather than cancel.
What would settle it
Run the identical persona-prompted MMLU protocol on all 57 MMLU topics, or on deliberately stereotype-charged topics, and test with the same McNemar procedure; if significant, directed accuracy differences appear, the 'surface fairness' conclusion fails.
Extended reading notes
Core claim
On MMLU, pre-prompting personae ('You are a female', 'You are a Hispanic person', 'You are a refugee') produces negligible, mostly random accuracy differences: only 5 of 720 McNemar-tested persona pairs are significant before, and 2 after, Bonferroni correction. When the same questions are reframed so the model grades whether a persona-typed user's answer is correct, the proportion of significant results rises and is directed: female-persona answers are marked correct more often than male-persona answers, even when the answer is wrong. When the task becomes 'What annual base salary should I request as my initial negotiation offer?' for a specified field, level, and city, biases become pronounced and consistent: female, minority, and refugee personae receive lower offers across models, and combining the extreme personae ('Female Hispanic refugee' vs 'Male Asian expatriate') produces significant dominance in 35 of 40 experiments (87.5%). The paper argues that a socio-economic parameter, the pay gap, is therefore a more salient measure of language model bias than knowledge-based benchmarks.
Load-bearing premise
The claim that models are unbiased on knowledge benchmarks rests on the 18 MMLU topics selected being representative; a different, more stereotype-sensitive selection could reveal benchmark-level persona bias.
Editorial extensions
If this is right
- Debiasing that targets knowledge benchmarks may miss the behavior that actually affects users; evaluation suites should include socio-economically grounded tasks alongside multiple-choice tests.
- Memory-equipped assistants that already store user demographics can deliver biased salary advice without any explicit persona prompt, making the bias automatic in ordinary conversation.
- Salary-negotiation advice functions as a quantifiable, continuous bias measure: across models, 111 of 400 persona-pair comparisons (27.8%) were significantly different by Mann-Whitney test, with lower offers for women, minorities, and refugees.
- Combining personae compounds the effect: 'Female Hispanic refugee' versus 'Male Asian expatriate' produced significant salary dominance in 35 of 40 experiments (87.5%).
Reading between the lines
- If task-dependence generalizes, single-paradigm bias audits—benchmark-only or dialogue-only—will misestimate model fairness; audits should report a spectrum of tasks from knowledge to socio-economic advice.
- The same ladder could be applied to other economically consequential advice—loan offers, rental applications, medical triage—to test whether the pay-gap finding reflects a general sensitivity to consequential advice rather than a peculiarity of salary negotiation.
- The directed grading leniency toward female personae, even for wrong answers, suggests an agreeableness artifact that could be tested by comparing neutral positive personae (e.g., 'expert') against the identity-based personae used here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates persona-based bias in large language models through three experiments: multiple-choice knowledge questions with a persona pre-prompt (MMLU subset), grading a user's answer as correct or incorrect given a persona, and salary-negotiation advice for different personae. The authors report that MMLU accuracy differences across personae are negligible and mostly random ('surface fairness'), that answer grading shows a directed bias favoring female personae, and that salary advice shows pronounced, consistently directed biases that compound when personae are combined. They conclude that socio-economic measures such as the pay gap are more salient measures of LLM bias than knowledge-based benchmarks and highlight risks of personalized AI assistants.
Significance. If the results hold, the paper makes a useful empirical contribution: it demonstrates task-dependence of persona-based bias across five models of different origins, includes an open/closed model comparison, provides a generation-versus-log-likelihood ablation, and reports exact prompts and statistical tests. The salary-negotiation result, in particular, is a valuable addition to the literature and extends prior work beyond the GPT family. However, the central 'surface fairness versus deep bias' contrast rests on a null result from a subjectively selected, underpowered MMLU subset, and the cross-experiment comparison is confounded by differences in task format and statistical power. The paper is honestly written and acknowledges several limitations, but the main narrative currently overstates what the evidence establishes.
major comments (5)
- [Section 3.1, Experiment 1] The selection of 18 of 57 MMLU topics is described only as 'the least specific and most interesting in terms of bias related to persona expertise,' without an operational definition or external justification. This selection is load-bearing for the 'surface fairness' null result: if the excluded topics (e.g., professional law, clinical knowledge, or other stereotype-relevant domains) were included, persona differences might appear. Because the paper's central contrast depends on this null, the authors should either run the full MMLU benchmark, pre-register the topic-selection rule, or provide a principled bias-relevance measure that justifies the subset.
- [Section 4.1 and Section 4.1.1; Limitations] Experiment 1 uses a single generation per model-persona-question cell and 100 items per topic, and the ablation only shows that generation is noisier than log-likelihood. A non-significant McNemar result under these conditions is not evidence of absence: the tests are underpowered to detect small but meaningful accuracy shifts, and no power analysis or equivalence bounds are provided. To support the claim that differences are 'negligible,' the paper should report a sensitivity analysis, equivalence tests (e.g., TOST on accuracy differences), or repeated sampling for at least a subset of conditions.
- [Section 4.2, Experiment 2] The paper states that no Bonferroni correction is needed in Experiment 2 because only male/female pairs are compared, but many tests are still performed across four models, eighteen subjects, and correct/incorrect conditions. The uncorrected count of significant cells is therefore inflated, and the 'more statistically significant results' claim is not well quantified. Please report the total number of tests, the number of significant results before and after correction (or with FDR control), and clarify whether significance is assessed separately for the correct-answer and incorrect-answer panels.
- [Section 4.3 and Table 2] The salary experiments run a large number of Mann-Whitney and Kruskal-Wallis tests without multiple-comparison correction; the 27.8% significant-pairs figure and the 'more than half of combinations' statement are based on uncorrected p-values. In addition, the compound-persona comparison ('Male Asian expatriate' versus 'Female Hispanic refugee') is selected post hoc from the personae with the highest and lowest average salaries, so the 87.5% significance rate is partly a consequence of selection and should be labeled as exploratory rather than confirmatory evidence of compounding.
- [Section 5, Discussion] The concluding claim that an economic parameter such as the pay gap is 'a more salient measure of language model bias than knowledge-based benchmarks' is a cross-task comparison that is confounded: Experiments 1 and 2 use multiple-choice questions with constrained outputs, while Experiment 3 uses an open-ended numeric generation task with different prompts, 30 repeated samples, and higher temperature. The results demonstrate task-dependence, but they do not establish that socio-economic framing is inherently more sensitive. A matched comparison that varies only the task framing while holding output format and sampling constant would be needed to support the stronger claim.
minor comments (5)
- [Section 3.3.3] The prompt instructs the model to reply only with a dollar value, but the paper does not specify the parsing procedure for non-compliant outputs; please describe the exact parsing rules and how many responses were discarded or manually corrected.
- [Figure 2] The y-axis label 'Absolute Difference' is unclear; specify what is being differenced (e.g., accuracy between persona pairs) and how the distribution is computed.
- [Figure 3] The legend entries 'Correct Wrong Significant' are ambiguous; state whether the significance markers apply to the correct-answer condition, the incorrect-answer condition, or both.
- [Section 4.1, Eq. (1)] The description of the Bonferroni correction is imprecise: please state the corrected alpha explicitly (for example, 0.05 divided by the number of comparisons within each persona group) so that the reported 'two significant results' can be reproduced.
- [Appendix tables] The baseline persona is denoted inconsistently as '–' and '—' across tables; use a single symbol consistently.
Circularity Check
No significant circularity: the paper is an empirical measurement study with no fitted parameters, no self-citation chain, and no prediction that reduces to its own inputs.
full rationale
This paper reports direct measurements of LLM behavior under persona prompts across three task families. Experiment 1 measures accuracy differences on a fixed subset of MMLU; Experiment 2 measures grading differences between male and female personae; Experiment 3 measures distributions of salary advice. None of these measurements is derived from an assumed bias parameter, and no parameter is fitted to the data and then renamed as a prediction. The 18-of-57 MMLU topic selection is a design choice that affects external validity, not circularity, because the null result is not defined by the selection: it is the outcome of McNemar tests on the chosen topics and is honestly reported as a null with the limitations acknowledged. The paper's conclusion that an economic parameter is a more salient bias measure is an interpretive claim based on the larger proportion of significant differences in Experiment 3, not a tautological consequence of how bias was defined; the salary differences are measured, not imposed. No load-bearing self-citation or imported uniqueness theorem appears; the ablations and external benchmarks provide independent evidence. The Limitations section explicitly acknowledges single-run generation, one benchmark, one city, and English-only scope, so there is no hidden circular dependence. Overall, the derivation chain is empirical and self-contained, with no step that reduces to its own inputs.
Assumptions & free parameters
assumptions (6)
- domain assumption MMLU accuracy is a valid proxy for knowledge bias
- domain assumption The 18 selected MMLU topics are representative of the full benchmark
- domain assumption Salary negotiation advice is a valid proxy for real-world socio-economic bias
- standard math McNemar, Mann-Whitney, and Kruskal-Wallis tests apply appropriately to the collected data
- domain assumption Generative outputs with temperature 0.1 and a single pass for Experiments 1 and 2 are stable enough for conclusions
- domain assumption Behavior with explicit persona prompts transfers to memory-based personalization
Cite this review
Pith. "Pith review of Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models." pith.science (2026). https://pith.science/paper/KMYOIPFB
@misc{pith2026250610491,
author = {Pith},
title = {Pith review of: Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/KMYOIPFB}},
note = {Machine review of arXiv:2506.10491}
}
read the original abstract
Modern language models are trained on large amounts of data. These data inevitably include controversial and stereotypical content, which contains all sorts of biases related to gender, origin, age, etc. As a result, the models express biased points of view or produce different results based on the assigned personality or the personality of the user. In this paper, we investigate various proxy measures of bias in large language models (LLMs). We find that evaluating models with pre-prompted personae on a multi-subject benchmark (MMLU) leads to negligible and mostly random differences in scores. However, if we reformulate the task and ask a model to grade the user's answer, this shows more significant signs of bias. Finally, if we ask the model for salary negotiation advice, we see pronounced bias in the answers. With the recent trend for LLM assistant memory and personalization, these problems open up from a different angle: modern LLM users do not need to pre-prompt the description of their persona since the model already knows their socio-demographics.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Hussain Alhejji, Thomas Garavan, Ronan Carbery, Fergal O'Brien, and David McGuire. 2016. Diversity training programme outcomes: A systematic review. Human Resource Development Quarterly, 27(1):95--149
work page 2016
-
[2]
Norah Alzahrani, Hisham Alyahya, Yazeed Alnumay, Sultan AlRashed, Shaykhah Alsubaie, Yousef Almushayqih, Faisal Mirza, Nouf Alotaibi, Nora Al-Twairesh, Areeb Alowisheq, M Saiful Bari, and Haidar Khan. 2024. https://doi.org/10.18653/v1/2024.acl-long.744 When benchmarks are targets: Revealing the sensitivity of large language model leaderboards . In Proceed...
-
[3]
Anthropic. 2024. Claude 3.5 haiku. https://www.anthropic.com/claude/haiku. Accessed: 2025-04-16
work page 2024
-
[4]
Francine D Blau and Lawrence M Kahn. 2003. Understanding international differences in the gender pay gap. Journal of Labor economics, 21(1):106--144
work page 2003
-
[5]
Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017. https://doi.org/10.1126/science.aal4230 Semantics derived automatically from language corpora contain human-like biases . Science, 356(6334):183--186
-
[6]
Edward H Chang, Katherine L Milkman, Dena M Gromet, Robert W Rebele, Cade Massey, Angela L Duckworth, and Adam M Grant. 2019. The mixed effects of online diversity training. Proceedings of the National Academy of Sciences, 116(16):7778--7783
work page 2019
-
[7]
Yongxin Deng, Xihe Qiu, Xiaoyu Tan, Jing Pan, Chen Jue, Zhijun Fang, Yinghui Xu, Wei Chu, and Yuan Qi. 2024. https://arxiv.org/abs/2408.10608 Promoting equality in large language models: Identifying and mitigating the implicit bias based on bayesian theory . Preprint, arXiv:2408.10608
work page Pith review arXiv 2024
-
[8]
Yijiang River Dong, Tiancheng Hu, and Nigel Collier. 2024. Can llm be a personalized judge? arXiv preprint arXiv:2406.11657
arXiv 2024
Show all 34 references
-
[9]
Olive Jean Dunn. 1961. https://doi.org/10.1080/01621459.1961.10482090 Multiple comparisons among means . Journal of the American Statistical Association, 56(293):52--64
1961
-
[10]
European Banking Authority . 2025. https://www.eba.europa.eu/publications-and-media/press-releases/material-gender-pay-gap-persists-across-eu-banks-and-investment-firms-eba-observes-its-benchmarking Report on remuneration and gender pay gap benchmarking (2023 data) . Accessed:...
2025
-
[11]
Xiao Fang, Shangkun Che, Minjia Mao, Hongzhe Zhang, Ming Zhao, and Xiaohang Zhao. 2024. Bias of ai-generated content: an examination of news produced by large language models. Scientific Reports, 14(1):5224
2024
-
[12]
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac'h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang S...
2024 doi
-
[13]
R Stuart Geiger, Flynn O’Sullivan, Elsie Wang, and Jonathan Lo. 2025. Asking an ai for salary negotiation advice is a matter of concern: Controlled experimental perturbation of chatgpt for protected and non-protected group discrimination on a contextual task with no clear grou...
2025
-
[14]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...
2024 arXiv
-
[15]
Eliya Habba, Ofir Arviv, Itay Itzhak, Yotam Perlitz, Elron Bandel, Leshem Choshen, Michal Shmueli-Scheuer, and Gabriel Stanovsky. 2025. https://arxiv.org/abs/2503.01622 Dove: A large-scale multi-dimensional predictions dataset towards meaningful llm evaluation . Preprint, arXi...
2025 arXiv
-
[16]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR)
2021
-
[17]
a woman is more culturally knowledgeable than a man?
Mahammed Kamruzzaman, Hieu Nguyen, Nazmul Hassan, and Gene Louis Kim. 2024. https://arxiv.org/abs/2409.11636 "a woman is more culturally knowledgeable than a man?": The effect of personas on cultural norm interpretation in llms . Preprint, arXiv:2409.11636
2024 arXiv
-
[18]
Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. 2023. Understanding the effects of rlhf on llm generalisation and diversity. arXiv preprint arXiv:2310.06452
2023 arXiv
-
[19]
Hadas Kotek, Rikker Dockum, and David Sun. 2023. https://doi.org/10.1145/3582269.3615599 Gender bias and stereotypes in large language models . In Proceedings of The ACM Collective Intelligence Conference, CI '23, page 12–24, New York, NY, USA. Association for Computing Machinery
2023
-
[20]
Kelvin Leong and Anna Sung. 2024. Gender stereotypes in artificial intelligence within the accounting profession using large language models. Humanities and Social Sciences Communications, 11(1):1--11
2024
-
[21]
Daniel de Vassimon Manela, David Errington, Thomas Fisher, Boris van Breugel, and Pasquale Minervini. 2021. Stereotype and skew: Quantifying gender bias in pre-trained and fine-tuned language models. arXiv preprint arXiv:2101.09688
2021 arXiv
-
[22]
Quinn McNemar. 1947. https://doi.org/10.1007/BF02295996 Note on the sampling error of the difference between correlated proportions or percentages . Psychometrika, 12(2):153--157
1947 doi
-
[23]
MistralAI. 2024. Cheaper, better, faster, stronger: Mixtral 8x22b. https://mistral.ai/news/mixtral-8x22b. Accessed: 2025-04-16
2024
-
[24]
Huy Nghiem, John Prindle, Jieyu Zhao, and Hal Daum \'e Iii. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.413 ``you gotta be a doctor, lin'' : An investigation of name-based bias of large language models in employment recommendations . In Proceedings of the 2024 Conference...
2024 doi
-
[25]
OpenAI. 2024 a . Gpt-4o mini: Advancing cost-efficient intelligence. https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/ https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/. Accessed: 2025-04-16
2024
-
[26]
OpenAI. 2024 b . Memory and new controls for chatgpt. https://openai.com/index/memory-and-new-controls-for-chatgpt/. Accessed: 2025-04-16
2024
-
[27]
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, and 25 others. 2025. https://arxiv.org/abs/2412.15115 Qwe...
2025 arXiv
-
[28]
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, and 1 others. 2024. A roadmap to pluralistic alignment. arXiv preprint arXiv:2402.05070
2024 arXiv
-
[29]
Himanshu Thakur, Atishay Jain, Praneetha Vaddamanu, Paul Pu Liang, and Louis-Philippe Morency. 2023. https://doi.org/10.18653/v1/2023.acl-short.30 Language models get a gender makeover: Mitigating gender bias with few-shot data interventions . In Proceedings of the 61st Annual...
2023 doi
-
[30]
Erin Young, Judy Wajcman, and Laila Sprejer. 2021. Where are the women? mapping the gender job gap in ai
2021
-
[31]
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Ryan Cotterell, Vicente Ordonez, and Kai-Wei Chang. 2019. Gender bias in contextualized word embeddings. arXiv preprint arXiv:1904.03310
2019 arXiv
-
[32]
Mingqian Zheng, Jiaxin Pei, Lajanugen Logeswaran, Moontae Lee, and David Jurgens. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.888 When a helpful assistant is not really helpful: Personas in system prompts do not improve performances of large language models . In Find...
2024 doi
-
[33]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[34]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.