Pith. sign in

REVIEW 1 major objections 6 minor 96 references

Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances

T0 review · 1 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper claims that generative AI unfairness metrics lose validity because they skip systematization, and that decomposing unfairness into harms, morally arbitrary factors, and morally decisive factors—the Fair Equality of Chances…

desk verdict A clear, honest conceptual paper that usefully brings Fair Equality of Chances to the systematization stage of GenAI fairness measurement, with a genuinely diagnostic case study; the main soft spot is real but explicitly acknowledged. read the letter →

arxiv 2507.04641 v1 pith:APII6TTO submitted 2025-07-07 cs.CY

classification cs.CY
keywords generativeAIfairnessmeasurementvalidityFairEqualityofChancessystematizationmorallyarbitraryfactorsdecisivebenchmarkdesign
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Generative AI fairness metrics often promise to measure unfairness but end up measuring something else, because their designers jump straight from a vague goal to a concrete formula. This paper claims that the missing step is systematization—an explicit decomposition of what unfairness means in a given context—and that this omission has compromised the validity of existing benchmarks. To fill that gap, the authors adapt the Fair Equality of Chances principle from political philosophy, breaking outcome unfairness into three constituents: the harm or benefit produced by the system, the morally arbitrary factors that should not drive inequalities, and the morally decisive factors that justify different treatment. The central claim is that working through these three components before choosing a metric exposes where validity is lost and yields measurements that actually capture the intended unfairness concept. If the framework works, fairness measurement for generative AI becomes a structured design task rather than an arbitrary selection from a crowded benchmark zoo.

What carries the argument

The Fair Equality of Chances (FEC) framework, imported from predictive AI fairness literature, is the engine of the argument. It states that the distribution of harm/benefit should be equal across groups differing only in morally arbitrary factors, conditional on morally decisive factors; formally, $F^h(.|s,d) = F^h(.|s',d)$ for arbitrary factors $s,s'$ and deservingness level $d$. The paper's move is to treat this not as a metric but as a systematization checklist: every generative AI fairness measurement must answer three questions—what counts as harm/benefit, which factors are morally arbitrary, and which are morally decisive—before operationalization. The framework also supplies a prioritization principle (prevalence, severity, distribution) to allocate measurement effort.

What would settle it

A concrete test: have two independent groups of stakeholders apply the FEC decomposition to the same generative AI application, such as a mental health support chatbot. If their classifications of factors as morally decisive versus morally arbitrary agree at chance level, even after facilitated deliberation, then the central premise that this distinction can be reliably systematized fails, and with it the claimed improvement in measurement validity.

Watch

Extended reading notes

Core claim

The paper's central discovery is that the validity of generative AI unfairness measurements can be assessed and improved by systematizing the unfairness construct through a three-part decomposition grounded in the Fair Equality of Chances principle. This decomposition requires specifying the harm/benefit at stake, the morally arbitrary factors that must not affect the distribution of that harm/benefit, and the morally decisive factors that justify differential treatment. The authors argue that many existing metrics, including Marked Persons, Counterfactual Sentiment Bias, and Psycholinguistic Norms, fail on at least one of these dimensions—for example by assuming all groups warrant equal treatment even when factors like occupation legitimately alter outcomes, or by leaving the selection of morally arbitrary attributes undocumented. Bypassing systematization, they claim, is the source of the misalignment between what metrics report and what unfairness actually means in context.

Load-bearing premise

The framework's power rests on the assumption that evaluators can reliably and non-arbitrarily distinguish factors that are morally arbitrary from those that are morally decisive in a given generative AI context; if stakeholders cannot agree on that line, the systematization step cannot deliver the validity gains it promises.

Editorial extensions

If this is right

  • Existing generative AI fairness benchmarks that skip systematization likely misreport unfairness; their scores should be treated as provisional pending decomposition under the three-part framework.
  • New benchmark design can follow a structured checklist—specify harms/benefits, morally arbitrary factors, and morally decisive factors—to preempt validity threats during the design phase.
  • Difference-aware metrics that acknowledge morally decisive factors can contradict equality-based benchmarks; with FEC systematization such contradictions become interpretable rather than puzzling.
  • Benchmark documentation should disclose assumptions about morally arbitrary and morally decisive factors, enabling meta-evaluation by stakeholders and affected communities.
  • Prioritization based on prevalence, severity, and distribution helps allocate limited evaluation resources toward the most consequential fairness concerns.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same three-part decomposition could serve as a validation checklist for other sociotechnical constructs beyond unfairness, such as safety or truthfulness, since the underlying issue—jumping from vague concept to operational formula—is generic.
  • A quantitative consequence one could test: benchmarks revised under this framework should show divergent unfairness scores precisely in cases where morally decisive factors correlate with protected attributes; this would turn the paper's philosophical point into an empirical one.
  • The framework implies that benchmark documentation should become a normative artifact: disclosing morally decisive and morally arbitrary assumptions makes evaluation choices contestable by affected communities, shifting part of the burden of fairness from model builders to deliberative processes.
  • The paper leaves open how to weight morally decisive factors; a natural extension would be a sensitivity analysis showing how unfairness rankings change as weights vary, helping stakeholders understand the stakes of their normative choices.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 6 minor

Summary. This paper argues that the validity of (un)fairness measurements for generative AI is compromised when measurement designers move directly from contextualization to operationalization, bypassing systematization. The authors propose a systematization framework based on the Fair Equality of Chances (FEC) principle, decomposing a contextualized unfairness construct into three constituents: the harm/benefit produced by the system, morally arbitrary factors, and morally decisive factors. They apply the framework in a case study of three widely used stereotyping metrics (Marked Persons, Counterfactual Sentiment Bias, Psycholinguistic Norms), claiming to expose validity threats such as unspecified harms, unexamined assumptions about morally arbitrary factors, and failure to recognize justified differential treatment. The paper concludes with recommendations for improving existing benchmarks and a discussion of limitations.

Significance. The contribution is timely and relevant: there is a recognized gap between fairness benchmarks and the constructs they purport to measure, and the FEC lens provides a principled way to make normative assumptions explicit. The paper is careful to acknowledge that it does not resolve normative disagreement over which factors are morally arbitrary or decisive, and it explicitly disclaims any guarantee of validity. Its case study is illustrative rather than empirical, which is appropriate for a systematization proposal. The framework's usefulness as a diagnostic tool is demonstrated convincingly: the decomposition usefully isolates, for example, the hidden assumption in Psycholinguistic Norms that occupation is morally arbitrary. The honest limitations, clear structure, and grounding in prior measurement theory are strengths. The main weakness is that the central mechanism--the classification of morally arbitrary versus decisive factors--is not specified tightly enough to fully support the stronger claim that following the framework reduces threats to validity; this point is addressed in the major comments.

major comments (1)
  1. [§5.3 and §7] The paper's central claim is that systematizing unfairness through the FEC decomposition improves measurement validity. This claim depends on the evaluator's ability to reliably classify factors as morally arbitrary or morally decisive, but the manuscript offers no procedure for this classification: §5.3 only says that identifying the distinction 'requires careful analysis of indirect relationships and their correlations' and suggests establishing 'justifiable thresholds,' while §7 recommends stakeholder input and documentation without specifying how disagreements are to be adjudicated. Because the classification is underdetermined, two well-intentioned teams could decompose the same contextualized construct differently and produce different 'validated' metrics, in which case the framework relabels rather than reduces ad hoc choices. The Ethical Considerations candidly state that the framework cannot guarantee validity, but the stronger assertion that following the framework 'will have identified and reduced the threats to measurement validity' remains unsupported. I recommend either adding a concrete, replicable process (e.g., a participatory deliberation protocol with sensitivity analysis and pre-registered classification rules) or explicitly and consistently framing the framework as a diagnostic lens rather than a generative method.
minor comments (6)
  1. [§5.1] The phrase 'course-grained' should be 'coarse-grained'; it appears twice in the text.
  2. [Figure 1] The figure contains typos: 'discrimnation' should be 'discrimination' and 'examplified' should be 'exemplified'.
  3. [§6 and Table 1] The metric is referred to as 'Marked Persons' in the body and Table 1 but as 'Marked Personas' in the reference list; the naming should be made consistent.
  4. [Definition 2.1] The formal definition F_h(.|s,d) is never used after its introduction; the paper should connect it to the three constituents (harm/benefit b, morally arbitrary s, morally decisive d) or omit it to avoid a gap between the formal and informal expositions.
  5. [§5.3] The statement that morally decisive and morally arbitrary factors 'must be mutually exclusive' should be clarified as 'mutually exclusive classifications of the same factor in a given context,' since a factor such as occupation can be morally arbitrary in one context and decisive in another.
  6. [Ethical Considerations] The candid statement that the framework cannot guarantee validity is useful; consider moving a version of it earlier in the paper (e.g., §1 or §5) to temper the stronger wording elsewhere.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the FEC-based framework is transparently imported from prior published work and applied analytically to GenAI measurement validity, with explicit qualifications and no prediction that is equivalent to its input.

full rationale

The paper makes no quantitative predictions and fits no parameters, so the patterns of fitted-input-called-prediction or by-construction equation identity do not apply. The FEC decomposition (harm/benefit, morally arbitrary factors, morally decisive factors) is explicitly imported from prior work by the same research group: the paper says, 'Building on a well-studied view in political philosophy (Heidari et al. 2019; Loi, Herlitz, and Heidari 2024), we define outcome unfairness as the unequal treatment of individuals on the grounds that they possess attributes belonging or ascribed to socially salient groups, but that are morally irrelevant to the task at hand,' and later, 'We propose to define context-aware outcome unfairness measurements for GenAI systems by extending Heidari et al. (2019)'s extension of the FEC principle.' This is an acknowledged intellectual inheritance rather than a disguised assumption presented as a derivation. The case study in Section 6 decomposes three existing metrics into the FEC components and interprets their validity threats; that is an analytical application of a stated normative lens, not a prediction that reduces to its own input. The paper also qualifies its central claim: the Ethical Considerations section states, 'we cannot guarantee that any measurement is valid. Rather, by following our proposed framework, one will have identified and reduced the threats to measurement validity which very commonly surface due to improper or a lack of systematization during measurement design,' and the Introduction acknowledges that the work 'does not resolve normative disagreements regarding the appropriate choice for each of these three pillars of fairness.' The framework's conclusions are therefore conditional on accepting FEC as a normative standard, which is a transparent premise rather than a circular step. External anchors such as Chouldechova et al. (2024) and Wang et al. (2025) are also cited to ground the validity framework and the morally arbitrary/decisive distinction. No specific reduction of a claimed result to a fitted parameter or to a self-citation chain can be exhibited.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no entities beyond its conceptual framework. It relies on normative assumptions from the authors' prior FEC work and from social science measurement theory. These assumptions are reasonable for a position paper but are not proven or empirically grounded.

assumptions (4)
  • domain assumption The Fair Equality of Chances principle, as defined in Definition 2.1, is the appropriate normative benchmark for evaluating unfairness in GenAI outcomes.
    Borrowed from Heidari et al. 2019 and Loi, Herlitz, and Heidari 2024, Section 2.5. The paper does not independently justify why FEC is the right lens rather than alternative fairness frameworks.
  • domain assumption The four-component measurement framework (contextualize, systematize, operationalize, and apply) from Chouldechova et al. 2024 is the correct characterization of measurement design.
    Invoked in Section 3 as the structural backbone of the paper without critical examination.
  • domain assumption Harms can be meaningfully taxonomized into allocative, representational, social systems, and interpersonal categories.
    Uses taxonomies from Shelby et al. 2023 and Bird et al. 2020 in Section 5.1, assuming these categories cover the relevant space of GenAI harms.
  • domain assumption Morally arbitrary and morally decisive factors are mutually exclusive and identifiable within a given context.
    Assumed in Section 5.3. The paper acknowledges the difficulty and discusses correlations between the two, but the framework relies on the ability to separate them in practice.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances." pith.science (2026). https://pith.science/paper/APII6TTO

@misc{pith2026250704641,
  author       = {Pith},
  title        = {Pith review of: Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/APII6TTO}},
  note         = {Machine review of arXiv:2507.04641}
}
read the original abstract

Disparities in the societal harms and impacts of Generative AI (GenAI) systems highlight the critical need for effective unfairness measurement approaches. While numerous benchmarks exist, designing valid measurements requires proper systematization of the unfairness construct. Yet this process is often neglected, resulting in metrics that may mischaracterize unfairness by overlooking contextual nuances, thereby compromising the validity of the resulting measurements. Building on established (un)fairness measurement frameworks for predictive AI, this paper focuses on assessing and improving the validity of the measurement task. By extending existing conceptual work in political philosophy, we propose a novel framework for evaluating GenAI unfairness measurement through the lens of the Fair Equality of Chances framework. Our framework decomposes unfairness into three core constituents: the harm/benefit resulting from the system outcomes, morally arbitrary factors that should not lead to inequality in the distribution of harm/benefit, and the morally decisive factors, which distinguish subsets that can justifiably receive different treatments. By examining fairness through this structured lens, we integrate diverse notions of (un)fairness while accounting for the contextual dynamics that shape GenAI outcomes. We analyze factors contributing to each component and the appropriate processes to systematize and measure each in turn. This work establishes a foundation for developing more valid (un)fairness measurements for GenAI systems.

Figures

Figures reproduced from arXiv: 2507.04641 by the authors.

Figure 1
Figure 1. Overview of the systematization process of the unfairness construct using our proposed framework for valid measure [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

96 extracted references · 62 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    A.; Scheidegger, C.; and Venkatasubramanian, S

    Abbasi, M.; Friedler, S. A.; Scheidegger, C.; and Venkatasubramanian, S. 2019. Fairness in representation: quantifying stereotyping as a representational harm. In Proceedings of the 2019 SIAM International Conference on Data Mining, 801--809. SIAM

  4. [4]

    Adcock, R.; and Collier, D. 2001. Measurement Validity: A Shared Standard for Qualitative and Quantitative Research. American Political Science Review, 95(3): 529–546

  5. [5]

    Challenges in Measuring Bias via Open-Ended Language Generation

    Akyürek, A. F.; Kocyigit, M. Y.; Paik, S.; and Wijaya, D. 2022. Challenges in Measuring Bias via Open-Ended Language Generation. arXiv:2205.11601

  6. [6]

    Al-kfairy, M.; Mustafa, D.; Kshetri, N.; Insiew, M.; and Alfandi, O. 2024. Ethical challenges and solutions of generative AI: An interdisciplinary perspective. In Informatics, volume 11, 58. MDPI

  7. [7]

    Angwin, J.; Larson, J.; Mattu, S.; and Kirchner, L. 2022. Machine bias. In Ethics of data and analytics, 254--264. Auerbach Publications

  8. [8]

    A.; Brubach, B.; Desmarais, S.; Horowitz, A

    Bao, M.; Zhou, A.; Zottola, S. A.; Brubach, B.; Desmarais, S.; Horowitz, A. S.; Lum, K.; and Venkatasubramanian, S. 2021. It's COMPAS licated: The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1)

Show all 96 references
  1. [9]

    Bartlett, R.; Morse, A.; Stanton, R.; and Wallace, N. 2022. Consumer-lending discrimination in the FinTech era. Journal of Financial Economics, 143(1): 30--56

  2. [10]

    Bell, A.; Bynum, L.; Drushchak, N.; Zakharchenko, T.; Rosenblatt, L.; and Stoyanovich, J. 2023. The possibility of fairness: Revisiting the impossibility theorem in practice. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, 400--422

  3. [11]

    K.; Dey, K.; Hind, M.; Hoffman, S

    Bellamy, R. K.; Dey, K.; Hind, M.; Hoffman, S. C.; Houde, S.; Kannan, K.; Lohia, P.; Martino, J.; Mehta, S.; Mojsilovi \'c , A.; et al. 2019. AI Fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias. IBM Journal of Research and Development, 63(4/5): 4--1

  4. [12]

    H.; and Hutchinson, B

    Berman, G.; Cooper, N.; Deng, W. H.; and Hutchinson, B. 2024. Troubling Taxonomies in GenAI Evaluation. arXiv:2410.22985

  5. [13]

    Binns, R.; Van Kleek, M.; Veale, M.; Lyngs, U.; Zhao, J.; and Shadbolt, N. 2018. 'It's Reducing a Human Being to a Percentage': Perceptions of Justice in Algorithmic Decisions. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, CHI '18, 1–14. New ...

  6. [14]

    Bird, S.; Dud \' k, M.; Edgar, R.; Horn, B.; Lutz, R.; Milan, V.; Sameki, M.; Wallach, H.; and Walker, K. 2020. Fairlearn: A toolkit for assessing and improving fairness in AI. Microsoft, Tech. Rep. MSR-TR-2020-32

  7. [15]

    Bird, S.; Hutchinson, B.; Kenthapadi, K.; K c man, E.; and Mitchell, M. 2019. Fairness-Aware Machine Learning: Practical Challenges and Lessons Learned. In Companion Proceedings of The 2019 World Wide Web Conference, WWW '19, 1297–1298. New York, NY, USA: Association for Compu...

  8. [16]

    Blili-Hamelin, B.; and Hancox-Li, L. 2023. Making Intelligence: Ethical Values in IQ and ML Benchmarks. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, FAccT '23, 271–284. New York, NY, USA: Association for Computing Machinery. ISBN 979...

  9. [17]

    L.; Lopez, G.; Olteanu, A.; Sim, R.; and Wallach, H

    Blodgett, S. L.; Lopez, G.; Olteanu, A.; Sim, R.; and Wallach, H. 2021. Stereotyping N orwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets. In Zong, C.; Xia, F.; Li, W.; and Navigli, R., eds., Proceedings of the 59th Annual Meeting of the Association for C...

  10. [18]

    A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M

    Bommasani, R.; Hudson, D. A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M. S.; Bohg, J.; Bosselut, A.; Brunskill, E.; Brynjolfsson, E.; Buch, S.; Card, D.; Castellon, R.; Chatterji, N.; Chen, A.; Creel, K.; Davis, J. Q.; Demszky, D.; Donahue, C.; Doumbouya, M.;...

  11. [19]

    R.; and Dahl, G

    Bowman, S. R.; and Dahl, G. 2021. What Will it Take to Fix Benchmarking in Natural Language Understanding? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4843--4855. Online: Ass...

  12. [20]

    Caton, S.; and Haas, C. 2024. Fairness in machine learning: A survey. ACM Computing Surveys, 56(7): 1--38

  13. [21]

    Cheng, H.-F.; Stapleton, L.; Wang, R.; Bullock, P.; Chouldechova, A.; Wu, Z. S. S.; and Zhu, H. 2021. Soliciting stakeholders’ fairness notions in child maltreatment predictive systems. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 1--17

  14. [22]

    Cheng, M.; Durmus, E.; and Jurafsky, D. 2023. Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models. In Rogers, A.; Boyd-Graber, J.; and Okazaki, N., eds., Proceedings of the 61st Annual Meeting of the Association for Computational Linguisti...

  15. [23]

    Chouldechova, A. 2017. Fair prediction with disparate impact: A study of bias in recidivism prediction instruments. Big data, 5(2): 153--163

  16. [24]

    F.; Corvi, E.; Dow, P

    Chouldechova, A.; Atalla, C.; Barocas, S.; Cooper, A. F.; Corvi, E.; Dow, P. A.; Garcia-Gathright, J.; Pangakis, N.; Reed, S.; Sheng, E.; Vann, D.; Vogel, M.; Washington, H.; and Wallach, H. 2024. A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, ...

  17. [25]

    Chu, Z.; Wang, Z.; and Zhang, W. 2024. Fairness in Large Language Models: A Taxonomic Survey. SIGKDD Exploration Newsletter, 26(1): 34–48

  18. [26]

    D.; Nilforoshan, H.; Shroff, R.; and Goel, S

    Corbett-Davies, S.; Gaebler, J. D.; Nilforoshan, H.; Shroff, R.; and Goel, S. 2024. The measure and mismeasure of fairness. J. Mach. Learn. Res., 24(1)

  19. [27]

    Coston, A.; Kawakami, A.; Zhu, H.; Holstein, K.; and Heidari, H. 2023. A validity perspective on evaluating the justified use of data-driven decision-making algorithms. In 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 690--704. IEEE

  20. [28]

    Cotter, A.-M. M. 2016. Race matters: An international legal analysis of race discrimination. Routledge

  21. [29]

    L.; Madaio, M.; Daum \'e Iii, H.; Harrington, C.; and Wallach, H

    Cunningham, J.; Blodgett, S. L.; Madaio, M.; Daum \'e Iii, H.; Harrington, C.; and Wallach, H. 2024. Understanding the Impacts of Language Technologies' Performance Disparities on A frican A merican Language Speakers. In Findings of the Association for Computational Linguistic...

  22. [30]

    Dastin, J. 2018. Amazon scraps secret AI recruiting tool that showed bias against women. https://www.reuters.com/article/world/insight-amazon-scraps-secret-ai-recruiting-tool-that-showed-bias-against-women-idUSKCN1MK0AG/&ved=2ahUKEwi2lpCG7v-KAxVfEFkFHTfACJkQFnoECBAQAQ\. Access...

  23. [31]

    L.; and Talat, Z

    Delobelle, P.; Attanasio, G.; Nozza, D.; Blodgett, S. L.; and Talat, Z. 2024. Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLP . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 21669--21691...

  24. [32]

    K.; Calders, T.; and Berendt, B

    Delobelle, P.; Tokpo, E. K.; Calders, T.; and Berendt, B. 2022. Measuring fairness with biased rulers: A comparative study on bias metrics for pre-trained language models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational ...

  25. [33]

    H.; Yildirim, N.; Chang, M.; Eslami, M.; Holstein, K.; and Madaio, M

    Deng, W. H.; Yildirim, N.; Chang, M.; Eslami, M.; Holstein, K.; and Madaio, M. 2023. Investigating Practices and Opportunities for Cross-functional Collaboration around AI Fairness in Industry Practice. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and...

  26. [34]

    Dhamala, J.; Sun, T.; Kumar, V.; Krishna, S.; Pruksachatkun, Y.; Chang, K.-W.; and Gupta, R. 2021. Bold: Dataset and metrics for measuring biases in open-ended language generation. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, 862--872

  27. [35]

    Drost, E. A. 2011. Validity and reliability in social science research. Education Research and perspectives, 38(1): 105--123

  28. [36]

    Dworkin, R. 1985. A matter of principle. Oxford University Press

  29. [37]

    Feffer, M.; Martelaro, N.; and Heidari, H. 2023. The AI Incident Database as an Educational Tool to Raise Awareness of AI Harms: A Classroom Exploration of Efficacy, Limitations, & Future Improvements. In Proceedings of the 3rd ACM Conference on Equity and Access in Algorithms...

  30. [38]

    L.; Klein, D.; and Talat, Z

    Fleisig, E.; Blodgett, S. L.; Klein, D.; and Talat, Z. 2024. The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lang...

  31. [39]

    C.; and Kiritchenko, S

    Fraser, K. C.; and Kiritchenko, S. 2024. Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images. arXiv:2402.05779

  32. [40]

    A.; Scheidegger, C.; and Venkatasubramanian, S

    Friedler, S. A.; Scheidegger, C.; and Venkatasubramanian, S. 2021. The (im) possibility of fairness: Different value systems require different mechanisms for fair decision making. Communications of the ACM, 64(4): 136--143

  33. [41]

    Fulton, R.; Fulton, D.; Hayes, N.; and Kaplan, S. 2024. The Transformation Risk-Benefit Model of Artificial Intelligence: Balancing Risks and Benefits Through Practical Solutions and Use Cases. arXiv:2406.11863

  34. [42]

    Fuster, A.; Goldsmith-Pinkham, P.; Ramadorai, T.; and Walther, A. 2022. Predictably unequal? The effects of machine learning on credit markets. The Journal of Finance, 77(1): 5--47

  35. [43]

    O.; Rossi, R

    Gallegos, I. O.; Rossi, R. A.; Barrow, J.; Tanjim, M. M.; Kim, S.; Dernoncourt, F.; Yu, T.; Zhang, R.; and Ahmed, N. K. 2024. Bias and fairness in large language models: A survey. Computational Linguistics, 50: 1097--1179

  36. [44]

    S.; and Chouldechova, A

    Guerdan, L.; Barocas, S.; Holstein, K.; Wallach, H.; Wu, Z. S.; and Chouldechova, A. 2025. Validating LLM-as-a-Judge Systems in the Absence of Gold Labels. arXiv:2503.05965

  37. [45]

    Gupta, S.; Shrivastava, V.; Deshpande, A.; Kalyan, A.; Clark, P.; Sabharwal, A.; and Khot, T. 2024. Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs. arXiv:2311.04892

  38. [46]

    Hardt, M. 2025. The emerging science of machine learning benchmarks. Online at https://mlbenchmarks.org. Manuscript. Accessed May 2025

  39. [47]

    L.; Chouldechova, A.; Garcia-Gathright, J.; Olteanu, A.; and Wallach, H

    Harvey, E.; Sheng, E.; Blodgett, S. L.; Chouldechova, A.; Garcia-Gathright, J.; Olteanu, A.; and Wallach, H. 2024. Gaps Between Research and Practice When Measuring Representational Harms Caused by LLM-Based Systems. arXiv:2411.15662

  40. [48]

    P.; and Krause, A

    Heidari, H.; Loi, M.; Gummadi, K. P.; and Krause, A. 2019. A Moral Framework for Understanding Fair ML through Economic Models of Equality of Opportunity. In Proceedings of the Conference on Fairness, Accountability, and Transparency, 181–190. Association for Computing Machine...

  41. [49]

    Holstein, K.; Wortman Vaughan, J.; Daum\' e , H.; Dudik, M.; and Wallach, H. 2019. Improving Fairness in Machine Learning Systems: What Do Industry Practitioners Need? In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, CHI '19, 1–16. New York, NY,...

  42. [50]

    Hsu, B.; Mazumder, R.; Nandy, P.; and Basu, K. 2022. Pushing the limits of fairness impossibility: Who's the fairest of them all? Advances in Neural Information Processing Systems, 35: 32749--32761

  43. [51]

    Huang, P.-S.; Zhang, H.; Jiang, R.; Stanforth, R.; Welbl, J.; Rae, J.; Maini, V.; Yogatama, D.; and Kohli, P. 2020. Reducing Sentiment Bias in Language Models via Counterfactual Evaluation. In Findings of the Association for Computational Linguistics: EMNLP 2020, 65--83. Assoc...

  44. [52]

    L.; and Luetge, C

    Hunkenschroer, A. L.; and Luetge, C. 2022. Ethics of AI-enabled recruiting and selection: A review and research agenda. Journal of Business Ethics, 178(4): 977--1007

  45. [53]

    Jiang, Z.; Han, X.; Fan, C.; Yang, F.; Mostafavi, A.; and Hu, X. 2022. Generalized demographic parity for group fairness. In International Conference on Learning Representations

  46. [54]

    Khaitan, T. 2015. A theory of discrimination law. Oxford University Press

  47. [55]

    Kleinberg, J.; Mullainathan, S.; and Raghavan, M. 2017. Inherent Trade-Offs in the Fair Determination of Risk Scores . In Papadimitriou, C. H., ed., 8th Innovations in Theoretical Computer Science Conference (ITCS 2017), volume 67 of Leibniz International Proceedings in Inform...

  48. [56]

    Lee, M. K. 2018. Understanding perception of algorithmic decisions: Fairness, trust, and emotion in response to algorithmic management. Big Data & Society, 5(1): 2053951718756684

  49. [57]

    Lefranc, A.; Pistolesi, N.; and Trannoy, A. 2009. Equality of opportunity and luck: Definitions and testable conditions, with an application to income in France. Journal of public economics, 93(11-12): 1189--1207

  50. [58]

    Li, Y.; Du, M.; Song, R.; Wang, X.; and Wang, Y. 2024. A Survey on Fairness in Large Language Models. arXiv:2308.10149

  51. [59]

    D.; Ré, C.; Acosta-Navas, D.; Hudson, D

    Liang, P.; Bommasani, R.; Lee, T.; Tsipras, D.; Soylu, D.; Yasunaga, M.; Zhang, Y.; Narayanan, D.; Wu, Y.; Kumar, A.; Newman, B.; Yuan, B.; Yan, B.; Zhang, C.; Cosgrove, C.; Manning, C. D.; Ré, C.; Acosta-Navas, D.; Hudson, D. A.; Zelikman, E.; Durmus, E.; Ladhak, F.; Rong, F....

  52. [60]

    Lippert-Rasmussen, K. 2013. Born free and equal?: A philosophical inquiry into the nature of discrimination. Oxford University Press

  53. [61]

    L.; Blodgett, S

    Liu, Y. L.; Blodgett, S. L.; Cheung, J. C. K.; Liao, Q. V.; Olteanu, A.; and Xiao, Z. 2024. ECBD: Evidence-Centered Benchmark Design for NLP. arXiv:2406.08723

  54. [62]

    Loi, M.; Herlitz, A.; and Heidari, H. 2024. Fair equality of chances for prediction-based decisions. Economics and Philosophy, 40(3): 557--580

  55. [63]

    Madaio, M.; Egede, L.; Subramonyam, H.; Wortman Vaughan, J.; and Wallach, H. 2022. Assessing the Fairness of AI Systems: AI Practitioners' Processes, Challenges, and Needs for Support. Proc. ACM Hum.-Comput. Interact., 6(CSCW1)

  56. [64]

    A.; Chen, J.; Wallach, H.; and Wortman Vaughan, J

    Madaio, M. A.; Chen, J.; Wallach, H.; and Wortman Vaughan, J. 2024. Tinker, Tailor, Configure, Customize: The Articulation Work of Contextualizing an AI Fairness Checklist. Proc. ACM Hum.-Comput. Interact., 8(CSCW1)

  57. [65]

    A.; Stark, L.; Wortman Vaughan, J.; and Wallach, H

    Madaio, M. A.; Stark, L.; Wortman Vaughan, J.; and Wallach, H. 2020. Co-Designing Checklists to Understand Organizational Challenges and Opportunities around Fairness in AI. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, CHI '20, 1–14. New Yor...

  58. [66]

    McGregor, S. 2021. Preventing Repeated Real World AI Failures by Cataloging Incidents: The AI Incident Database. Proceedings of the AAAI Conference on Artificial Intelligence, 35(17): 15458--15463

  59. [67]

    Mitchell, S.; Potash, E.; Barocas, S.; D'Amour, A.; and Lum, K. 2021. Algorithmic fairness: Choices, assumptions, and definitions. Annual review of statistics and its application, 8(1): 141--163

  60. [68]

    Moreau, S. 2010. What Is Discrimination? Philosophy & Public Affairs, 38(2): 143--179

  61. [69]

    Mun, J.; Jiang, L.; Liang, J.; Cheong, I.; DeCairo, N.; Choi, Y.; Kohno, T.; and Sap, M. 2024. Particip-ai: A democratic surveying framework for anticipating future ai use cases, harms and benefits. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7...

  62. [70]

    Obermeyer, Z.; Powers, B.; Vogeli, C.; and Mullainathan, S. 2019. Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464): 447--453

  63. [71]

    Perez, E.; Ringer, S.; Lukosiute, K.; Nguyen, K.; Chen, E.; Heiner, S.; Pettit, C.; Olsson, C.; Kundu, S.; Kadavath, S.; et al. 2023. Discovering Language Model Behaviors with Model-Written Evaluations. In Findings of the Association for Computational Linguistics: ACL 2023, 13...

  64. [72]

    Plank, B. 2022. The “Problem” of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 10671--10682

  65. [73]

    Poole-Dayan, E.; Roy, D.; and Kabbara, J. 2024. LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users. In Neurips Safe Generative AI Workshop

  66. [74]

    D.; Denton, E.; Bender, E

    Raji, I. D.; Denton, E.; Bender, E. M.; Hanna, A.; and Paullada, A. 2021. AI and the Everything in the Whole Wide World Benchmark. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track

  67. [75]

    Rawte, V.; Priya, P.; Tonmoy, S. M. T. I.; Zaman, S. M. M.; Sheth, A.; and Das, A. 2023. Exploring the Relationship between LLM Hallucinations and Prompt Linguistic Nuances: Readability, Formality, and Concreteness. arXiv:2309.11064

  68. [76]

    Roemer, J. E. 2002. Equality of opportunity: A progress report. Social Choice and Welfare, 19(2): 455--471

  69. [77]

    E.; and Trannoy, A

    Roemer, J. E.; and Trannoy, A. 2015. Equality of opportunity. In Handbook of income distribution, volume 2, 217--300. Elsevier

  70. [78]

    Ruf, B.; and Detyniecki, M. 2021. Towards the Right Kind of Fairness in AI. arXiv:2102.08453

  71. [79]

    Saha, D.; Schumann, C.; Mcelfresh, D.; Dickerson, J.; Mazurek, M.; and Tschantz, M. 2020. Measuring non-expert comprehension of machine learning fairness metrics. In International Conference on Machine Learning, 8377--8387. PMLR

  72. [80]

    T.; and Ghani, R

    Saleiro, P.; Kuester, B.; Hinkson, L.; London, J.; Stevens, A.; Anisfeld, A.; Rodolfa, K. T.; and Ghani, R. 2019. Aequitas: A Bias and Fairness Audit Toolkit. arXiv:1811.05577

  73. [81]

    Sap, M.; Card, D.; Gabriel, S.; Choi, Y.; and Smith, N. A. 2019. The Risk of Racial Bias in Hate Speech Detection. In Korhonen, A.; Traum, D.; and M \`a rquez, L., eds., Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 1668--1678. Floren...

  74. [82]

    Sap, M.; Swayamdipta, S.; Vianna, L.; Zhou, X.; Choi, Y.; and Smith, N. A. 2022. Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computatio...

  75. [83]

    A.; Huang, K.; DeFilippis, E.; Radanovic, G.; Parkes, D

    Saxena, N. A.; Huang, K.; DeFilippis, E.; Radanovic, G.; Parkes, D. C.; and Liu, Y. 2019. How Do Fairness Definitions Fare? Examining Public Attitudes Towards Algorithmic Definitions of Fairness. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, AIES '...

  76. [84]

    Sclar, M.; Choi, Y.; Tsvetkov, Y.; and Suhr, A. 2023. Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting. In The Twelfth International Conference on Learning Representations

  77. [85]

    Scurich, N.; and Monahan, J. 2016. Evidence-based sentencing: Public openness and opposition to using gender, age, and race as risk factors for recidivism. Law and Human Behavior, 40(1): 36

  78. [86]

    Sharma, S. 2024. Benefits or concerns of AI: A multistakeholder responsibility. Futures, 103328

  79. [87]

    Shelby, R.; Rismani, S.; Henne, K.; Moon, A.; Rostamzadeh, N.; Nicholas, P.; Yilla-Akbari, N.; Gallegos, J.; Smart, A.; Garcia, E.; et al. 2023. Sociotechnical harms of algorithmic systems: Scoping a taxonomy for harm reduction. In Proceedings of the 2023 AAAI/ACM Conference o...

  80. [88]

    K.; Grundy, E

    Slattery, P.; Saeri, A. K.; Grundy, E. A. C.; Graham, J.; Noetel, M.; Uuk, R.; Dao, J.; Pour, S.; Casper, S.; and Thompson, N. 2024. The AI Risk Repository: A Comprehensive Meta-Review, Database, and Taxonomy of Risks From Artificial Intelligence. arXiv:2408.12622

  81. [89]

    L.; Chen, C.; III, H

    Solaiman, I.; Talat, Z.; Agnew, W.; Ahmad, L.; Baker, D.; Blodgett, S. L.; Chen, C.; III, H. D.; Dodge, J.; Duan, I.; Evans, E.; Friedrich, F.; Ghosh, A.; Gohar, U.; Hooker, S.; Jernite, Y.; Kalluri, R.; Lusoli, A.; Leidinger, A.; Lin, M.; Lin, X.; Luccioni, S.; Mickel, J.; Mi...

  82. [90]

    Trewin, S.; Basson, S.; Muller, M.; Branham, S.; Treviranus, J.; Gruen, D.; Hebert, D.; Lyckowski, N.; and Manser, E. 2019. Considerations for AI fairness for people with disabilities. AI Matters, 5(3): 40–63

  83. [91]

    F.; Wang, A.; Atalla, C.; Barocas, S.; Blodgett, S

    Wallach, H.; Desai, M.; Cooper, A. F.; Wang, A.; Atalla, C.; Barocas, S.; Blodgett, S. L.; Chouldechova, A.; Corvi, E.; Dow, P. A.; Garcia-Gathright, J.; Olteanu, A.; Pangakis, N.; Reed, S.; Sheng, E.; Vann, D.; Vaughan, J. W.; Vogel, M.; Washington, H.; and Jacobs, A. Z. 2025...

  84. [92]

    E.; and Koyejo, S

    Wang, A.; Phan, M.; Ho, D. E.; and Koyejo, S. 2025. Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs. arXiv:2502.01926

  85. [93]

    A.; Comanescu, R.; Chang, O.; Rodriguez, M.; Beroshi, J.; Bloxwich, D.; Proleev, L.; Chen, J.; Farquhar, S.; Ho, L.; Gabriel, I.; Dafoe, A.; and Isaac, W

    Weidinger, L.; Barnhart, J.; Brennan, J.; Butterfield, C.; Young, S.; Hawkins, W.; Hendricks, L. A.; Comanescu, R.; Chang, O.; Rodriguez, M.; Beroshi, J.; Bloxwich, D.; Proleev, L.; Chen, J.; Farquhar, S.; Ho, L.; Gabriel, I.; Dafoe, A.; and Isaac, W. 2024. Holistic Safety and...

  86. [94]

    Zhao, D.; Andrews, J. T. A.; Papakyriakopoulos, O.; and Xiang, A. 2024. Position: Measure Dataset Diversity, Don't Just Claim It. arXiv:2407.08188

  87. [95]

    X.; Chen, X.; Lin, Y.; Wen, J.-R.; and Han, J

    Zhou, K.; Zhu, Y.; Chen, Z.; Chen, W.; Zhao, W. X.; Chen, X.; Lin, Y.; Wen, J.-R.; and Han, J. 2023. Don't Make Your LLM an Evaluation Benchmark Cheater. arXiv:2311.01964

  88. [96]

    Zimmermann, A.; and Lee-Stronach, C. 2022. Proceed with Caution. Canadian Journal of Philosophy, 52(1): 6–25

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.