REVIEW 3 major objections 5 minor 132 references
The Value of Disagreement in AI Design, Evaluation, and Alignment
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper argues that AI pipelines that suppress disagreement risk both accuracy and fairness, and it builds a framework to decide when and how to preserve dissent.
desk verdict A clearly argued conceptual framework that names a real risk—perspectival homogenization—but whose practical prescriptions outrun the evidence on transferring face-to-face group benefits to asynchronous, incentivized AI pipelines. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the three-stage task model together with the four epistemic rationales. The three stages—antecedent (selecting perspectives), process (structuring interaction), and outcomes (documenting and communicating)—provide a scaffold for locating where homogenization enters and where interventions belong. Each stage is linked to epistemic mechanisms: cognitive diversity and information elaboration, standpoint-based epistemic advantage, the justificatory and division-of-labor effects of disagreement, and disagreement as higher-order evidence. These mechanisms do the argumentative work of explaining when disagreement is genuinely valuable and why suppressing it is costly.
What would settle it
Run a preregistered comparison in a hate-speech annotation task: one arm uses majority-vote aggregation with isolated annotators; the other uses structured deliberation among a standpoint-diverse panel with recorded justifications. If the deliberation arm does not improve detection of target harms or does not yield better-calibrated labels than the majority-vote arm, the central claims about epistemic benefits in AI tasks would be weakened.
Extended reading notes
Core claim
The core claim is that 'perspectival homogenization'—an aspect of an AI system's design, evaluation, or alignment that excludes or attenuates relevant disagreement—constitutes a coupled ethical-epistemic risk that should be managed as a procedural risk throughout the AI lifecycle. The paper develops a normative framework that ties three stages of AI development tasks (antecedent, process, outcomes) to four epistemic rationales: diversity of perspectives expands cognitive resources and improves information exchange; marginalized standpoints offer situated knowledge and epistemic advantage; active disagreement motivates justification and divides cognitive labor; and the fact and content of disagreement provide higher-order evidence that should calibrate confidence. The framework yields practical recommendations: disagreement matters in complex objective tasks, not only subjective ones; inclusion should target achieved standpoints rather than demographic membership; task design should use network structures and justification-seeking communication rather than isolated judgment; and documentation should preserve disagreement to support coherent decisions across stages.
Load-bearing premise
The framework assumes that the epistemic mechanisms shown in face-to-face or philosophical settings—diversity's cognitive benefits, standpoint advantage, justificatory disagreement, and higher-order evidence—actually operate in real AI development contexts such as crowdsourced annotation, red-teaming, and preference elicitation, which are often online, asynchronous, and structured to avoid interaction.
Editorial extensions
If this is right
- Majority-vote aggregation and isolated annotation are called into question for tasks with relevant diversity, since they preclude the mechanisms that generate epistemic benefits.
- Disagreement should be taken seriously in tasks like clinical labeling and red-teaming—ones usually treated as objective—because they involve complexity, uncertainty, or situated knowledge.
- Including marginalized people is not enough; teams should operationalize participation around achieved standpoints, such as community leaders and advocacy experts.
- Two levers—communication topology and communicating justifications rather than bare judgments—allow realizing disagreement's benefits while managing friction.
- Documentation of disagreement should be coherent across stages, so downstream users can use it as higher-order evidence to calibrate confidence in AI outputs.
Reading between the lines
- A testable consequence is that in crowdsourcing platforms, adding structured deliberation or justification exchange should improve label quality and reliability over majority-vote baselines in tasks with high contextual complexity, mirroring lab findings.
- The same framework predicts that AI systems trained with disagreement-preserving annotations will be better calibrated—expressing higher uncertainty on contested inputs—than systems trained on majority labels; this can be measured at deployment.
- The higher-order evidence rationale generalizes: presenting users with preserved disagreement rather than synthetic consensus could reduce sycophantic echoing of user views, a direction the paper gestures at but does not fully develop.
- Even if the epistemic transfer fails in online, asynchronous, AI-mediated settings, the framework still offers procedural fairness justification for preserving disagreement; the two rationales are separable.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that standard AI development practice—majority-vote aggregation, homogeneous participant pools, and documentation that reports only aggregate labels—systematically obscures disagreement, and it introduces the notion of 'perspectival homogenization' to name the resulting coupled ethical-epistemic risk. It develops a normative framework based on four epistemic rationales for valuing disagreement: cognitive diversity and information elaboration, standpoint epistemology, productive argumentative discourse, and higher-order evidence. These rationales are mapped onto three stages of AI development tasks (antecedent, process, outcomes), yielding practical recommendations: expand the scope of tasks in which disagreement is treated as epistemically relevant; distinguish achieved standpoints from demographic attributes; replace isolated independent annotation with networked collective structures; modify communication structure and content; and document disagreement coherently. The paper is primarily a conceptual and normative contribution, aimed at grounding emerging perspectivist, participatory, and pluralistic approaches to AI development.
Significance. If the framework holds, it provides a valuable unifying vocabulary and normative justification for a growing body of work on annotator disagreement, participatory AI, and pluralistic alignment. The paper is careful in several places: it restricts the argument to 'relevant' disagreement, emphasizes that diversity is not always beneficial, and distinguishes demographic representation from standpoint-based expertise. It also draws on four independent research traditions and does not rely circularly on the authors' own prior work. The main significance is conceptual: it reframes disagreement handling as a procedural ethical-epistemic issue and identifies concrete design levers. The practical recommendations, however, depend on empirical premises about the transfer of small-group epistemic mechanisms to AI pipelines; this is the main weakness and the reason I am not recommending acceptance without revision.
major comments (3)
- [Section 5.3] The claim that 'the epistemic benefits of diversity and disagreement do not stem from aggregating isolated judgments' and that independent task designs 'preclude the benefits of both diversity and epistemically productive disagreement' is stronger than the cited evidence supports. The mechanisms reviewed in Sections 4.1–4.3 are largely studied in face-to-face or organizationally embedded groups, where social presence, accountability, and shared justification norms are present. The target settings—crowdsourced annotation, red-teaming, and preference elicitation—are often asynchronous, anonymous, and economically incentivized, and these conditions can weaken or reverse those mechanisms (e.g., through anchoring, social desirability, or strategic responding). The paper should engage the wisdom-of-crowds literature showing that independence, not interaction, often drives aggregate accuracy, and it should either present direct evidence for interaction benefits in AI pipeline settings or reframe the recommendations as conditional design hypotheses with explicit scope conditions. This is load-bearing because Section 5.4's network-topology and justification-exchange prescriptions inherit the same unvalidated premise.
- [Sections 3.1 and 5.2] The definition of perspectival homogenization turns on 'relevant' disagreement, but the paper's characterization of relevance is partly circular: relevant disagreement is said to be disagreement that contributes to 'epistemically and ethically better outcomes,' and the framework is then offered as the way to determine that. To make the concept operational for practitioners, the paper should provide more explicit criteria—for example, linking relevance to task complexity, situated knowledge, evidential diversity, and the presence of achieved standpoints—and should illustrate how to apply those criteria to a concrete annotation or red-teaming case. Without such criteria, the charge of unjustified homogenization risks being applied only retrospectively.
- [Sections 4.4 and 5.5] The higher-order evidence rationale treats the existence of disagreement as a reason to reduce confidence, but in adversarial or incentive-distorted settings (e.g., red-teaming or paid crowdwork) disagreement may reflect strategic behavior rather than independent epistemic signals. The paper notes 'all else equal' but does not discuss how practitioners can distinguish epistemically meaningful disagreement from noise or strategic responding. A short discussion of this distinction would strengthen the outcome-stage recommendations and prevent a misapplication of the framework.
minor comments (5)
- [Section 3 heading] The heading 'PERSPECTIV AL HOMOGENIZATION' contains an unintended space and should read 'PERSPECTIVAL HOMOGENIZATION.'
- [Section 5.5] The phrase 'annotator selection, compostion, and condition' contains a typo; 'compostion' should be 'composition.'
- [Figure 1] The figure caption does not explain the three-stage boxes beyond naming them; one sentence defining the antecedent, process, and outcome stages would help readers navigate Sections 4 and 5.
- [Section 4.1] The discussion of 'social identities, broadly construed' would benefit from an early caveat that demographic diversity is a proxy, not a guarantee, of cognitive diversity; the point appears later in the paper but earlier placement would prevent misreading.
- [Section 4.4] The term 'higher-order evidence' is used in a technical epistemological sense; a one-sentence gloss aimed at a computer science audience would improve accessibility.
Circularity Check
No significant circularity: the framework's normative conclusions are grounded in independent research traditions, with self-citations only in a supporting role.
full rationale
The paper is a normative and conceptual framework rather than an empirical derivation: it contains no fitted parameters, no equations, and no prediction that is statistically forced by construction. The central notion of perspectival homogenization is defined as the exclusion or attenuation of 'relevant' disagreement, but the paper explicitly does not treat its normative conclusion as following from that definition alone; it says that determining 'what constitutes relevant diversity and disagreement... requires a clearer understanding of what makes disagreement productive' and then supplies that content through four independent research lines: cognitive diversity and information elaboration (Steel et al., Phillips, Hong and Page), standpoint epistemology (Harding, Wylie, Toole), argumentative discourse (Mercier and Sperber, Longino), and higher-order evidence (Christensen). The authors' own prior work (Fazelpour and De-Arteaga 2022; Steel et al. 2019, co-authored with Fazelpour) is cited as supporting literature, and Section 4.1 states that the discussion 'closely follows' [34], but the mechanisms are also anchored in external sources, so the self-citation is not load-bearing for the paper's main claim. Section 5.5's specific entropy-versus-proportionality guidance is imported from [34], but it is an illustrative application, not the load-bearing step. The skeptic's concern that face-to-face group mechanisms may not transfer to asynchronous AI pipelines is an external-validity or empirical-support objection, not a circularity objection; under the paper's stated assumptions, the derivation is self-contained.
Assumptions & free parameters
assumptions (7)
- domain assumption Situated Knowledge Thesis: what individuals know depends on their social situation.
- domain assumption Achievement Thesis: standpoints require collective consciousness raising, not mere demographic membership.
- domain assumption Epistemic Advantage Thesis: marginalized standpoints are epistemically advantaged in certain domains.
- domain assumption Empirical generalization that socioculturally diverse groups outperform homogeneous ones on complex tasks through cognitive and information elaboration pathways.
- domain assumption Disagreement provides higher-order evidence that should reduce confidence.
- domain assumption Coupled ethical-epistemic issues are best addressed by making value decisions transparent.
- domain assumption Procedural risk should be assessed throughout a process rather than only by outcomes.
invented entities (1)
-
perspectival homogenization
Cite this review
Pith. "Pith review of The Value of Disagreement in AI Design, Evaluation, and Alignment." pith.science (2026). https://pith.science/paper/BLUITCRW
@misc{pith2026250507772,
author = {Pith},
title = {Pith review of: The Value of Disagreement in AI Design, Evaluation, and Alignment},
year = {2026},
howpublished = {\url{https://pith.science/paper/BLUITCRW}},
note = {Machine review of arXiv:2505.07772}
}
read the original abstract
Disagreements are widespread across the design, evaluation, and alignment pipelines of artificial intelligence (AI) systems. Yet, standard practices in AI development often obscure or eliminate disagreement, resulting in an engineered homogenization that can be epistemically and ethically harmful, particularly for marginalized groups. In this paper, we characterize this risk, and develop a normative framework to guide practical reasoning about disagreement in the AI lifecycle. Our contributions are two-fold. First, we introduce the notion of perspectival homogenization, characterizing it as a coupled ethical-epistemic risk that arises when an aspect of an AI system's development unjustifiably suppresses disagreement and diversity of perspectives. We argue that perspectival homogenization is best understood as a procedural risk, which calls for targeted interventions throughout the AI development pipeline. Second, we propose a normative framework to guide such interventions, grounded in lines of research that explain why disagreement can be epistemically beneficial, and how its benefits can be realized in practice. We apply this framework to key design questions across three stages of AI development tasks: when disagreement is epistemically valuable; whose perspectives should be included and preserved; how to structure tasks and navigate trade-offs; and how disagreement should be documented and communicated. In doing so, we challenge common assumptions in AI practice, offer a principled foundation for emerging participatory and pluralistic approaches, and identify actionable pathways for future work in AI design and governance.
Figures
Reference graph
Works this paper leans on
-
[1]
Lama Ahmad, Sandhini Agarwal, Michael Lampe, and Pamela Mishkin. 2025. OpenAI’s Approach to External Red Teaming for AI Models and Systems.arXiv preprint arXiv:2503.16431(2025)
arXiv 2025
-
[2]
Elizabeth Anderson. 2006. The epistemology of democracy.Episteme3, 1-2 (2006), 8–22
2006
-
[3]
Anthropic. 2024. Challenges in red teaming AI systems. https://www.anthropic. com/news/challenges-in-red-teaming-ai-systems Accessed: 2025-01-22
2024
-
[4]
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al
-
[5]
Moya Bailey. 2021. Misogynoir transformed: Black women’s digital resistance. InMisogynoir transformed. New York University Press
2021
-
[6]
2023.Fairness and machine learning: Limitations and opportunities
Solon Barocas, Moritz Hardt, and Arvind Narayanan. 2023.Fairness and machine learning: Limitations and opportunities. MIT press
2023
-
[7]
Valerio Basile, Michael Fell, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, Massimo Poesio, and Alexandra Uma. 2021. We Need to Consider Disagreement in Evaluation. InProceedings of the 1st Workshop on Benchmarking: Past, Present and Future, Kenneth Church, Mark Liberman, and Valia Kordoni (Eds.). Association for Computational Linguistics, On...
-
[8]
Julia B Bear and Anita Williams Woolley. 2011. The role of gender in team collaboration and performance.Interdisciplinary science reviews36, 2 (2011), 146–153
2011
Show all 132 references
-
[9]
Stevie Bergman, Nahema Marchal, John Mellor, Shakir Mohamed, Iason Gabriel, and William Isaac. 2024. STELA: a community-centred approach to norm elicitation for AI alignment.Scientific Reports14, 1 (2024), 6616
2024
-
[10]
Emily Black, Hadi Elzayn, Alexandra Chouldechova, Jacob Goldin, and Daniel Ho. 2022. Algorithmic fairness and vertical equity: Income fairness with irs tax audit models. InProceedings of the 2022 ACM Conference on Fairness, Account- ability, and Transparency. 1479–1503
2022
-
[11]
James Bohman. 2007. Political communication and the epistemic value of diversity: Deliberation and legitimation in media societies.Communication Theory17, 4 (2007), 348–355
2007
-
[12]
Rishi Bommasani, Kathleen A Creel, Ananya Kumar, Dan Jurafsky, and Percy S Liang. 2022. Picking on the same person: Does algorithmic monoculture lead to outcome homogenization?Advances in Neural Information Processing Systems 35 (2022), 3663–3678
2022
-
[13]
Daniel Braun. 2024. I beg to differ: how disagreement is handled in the anno- tation of legal machine learning data sets.Artificial intelligence and law32, 3 (2024), 839–862
2024
-
[14]
Liam Kofi Bright. 2024. Duboisian leadership through standpoint epistemology. The Monist107, 1 (2024), 82–97
2024
-
[15]
Federico Cabitza, Andrea Campagner, and Valerio Basile. 2023. Toward a per- spectivist turn in ground truthing for predictive computing. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 6860–6868
2023
-
[16]
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al. 2023. Open problems and fundamental limitations of reinforcement learning from human feedback.arXiv preprint ar...
2023 arXiv
-
[17]
I’ma strong independent Black woman
Stephanie Castelin and Grace White. 2022. “I’ma strong independent Black woman”: The strong Black woman schema and mental health in college-aged Black women.Psychology of Women Quarterly46, 2 (2022), 196–208
2022
-
[18]
David Christensen. 2009. Disagreement as Evidence: The Epistemology of Controversy.Philosophy Compass4, 5 (2009), 756–767. https://doi.org/10.1111/ j.1747-9991.2009.00237.x
2009
-
[19]
Robert B Cialdini and Noah J Goldstein. 2004. Social influence: Compliance and conformity.Annu. Rev. Psychol.55, 1 (2004), 591–621
2004
-
[20]
Patricia Hill Collins. 1986. Learning from the outsider within: The sociological significance of Black feminist thought.Social problems33, 6 (1986), s14–s32
1986
-
[21]
Matteo Colombo and Stephan Hartmann. 2017. Bayesian cognitive science, unification, and explanation.The British Journal for the Philosophy of Science (2017)
2017
-
[22]
Nancy J Cooke, Jamie C Gorman, Christopher W Myers, and Jasmine L Duran
-
[23]
Amanda Coston, Anna Kawakami, Haiyi Zhu, Ken Holstein, and Hoda Heidari
-
[24]
Helen De Cruz and Johan De Smedt. 2013. The Value of Epistemic Disagreement in Scientific Practice. The Case of Homo Floresiensis.Studies in History and Philosophy of Science Part A44, 2 (2013), 169–177. https://doi.org/10.1016/j. shpsa.2013.02.002
2013 doi
-
[25]
Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022. Dealing with disagreements: Looking beyond the majority vote in subjective annotations.Transactions of the Association for Computational Linguistics10 (2022), 92–110
2022
-
[26]
Fernando Delgado, Stephen Yang, Michael Madaio, and Qian Yang. 2023. The Participatory Turn in AI Design: Theoretical Foundations and the Current State of Practice. InProceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization(Boston,...
2023
-
[27]
Morton Deutsch and Harold B Gerard. 1955. A study of normative and informa- tional social influences upon individual judgment.The journal of abnormal and social psychology51, 3 (1955), 629
1955
-
[28]
Mark Díaz, Ian Kivlichan, Rachel Rosen, Dylan Baker, Razvan Amironesei, Vin- odkumar Prabhakaran, and Emily Denton. 2022. Crowdworksheets: Accounting for individual and collective identities underlying crowdsourced dataset anno- tation. InProceedings of the 2022 ACM Conference...
2022
-
[29]
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018. Measuring and mitigating unintended bias in text classification. InProceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society. 67–73
2018
-
[30]
Xiaoni Duan, Chien-Ju Ho, and Ming Yin. 2020. Does exposure to diverse perspectives mitigate biases in crowdwork? an explorative study. InProceedings of the aaai conference on human computation and crowdsourcing, Vol. 8. 155–158
2020
-
[31]
Joann G Elmore, Gary M Longton, Patricia A Carney, Berta M Geller, Tracy Onega, Anna NA Tosteson, Heidi D Nelson, Margaret S Pepe, Kimberly H Allison, Stuart J Schnitt, et al. 2015. Diagnostic concordance among pathologists interpreting breast biopsy specimens.Jama313, 11 (201...
2015
-
[32]
Tyna Eloundou and Teddy Lee. 2024. Democratic inputs to AI grant program: lessons learned and implementation plans. Available at: https://openai.com/blog/democraticinputs-to-ai-grant-program-update (ac- cessed 11 February 2025)
2024
-
[33]
Sina Fazelpour and David Danks. 2021. Algorithmic bias: Senses, sources, solutions.Philosophy Compass16, 8 (2021), e12760
2021
-
[34]
Sina Fazelpour and Maria De-Arteaga. 2022. Diversity in sociotechnical machine learning systems.Big Data & Society9, 1 (2022), 20539517221082027
2022
-
[35]
Sina Fazelpour, Zachary C Lipton, and David Danks. 2022. Algorithmic fairness and the situated dynamics of justice.Canadian Journal of Philosophy52, 1 (2022), 44–60
2022
-
[36]
Sina Fazelpour and Daniel Steel. 2022. Diversity, trust, and conformity: A simulation study.Philosophy of Science89, 2 (2022), 209–231
2022
-
[37]
Will Fleisher. 2018. Rational Endorsement.Philosophical Studies175, 10 (2018), 2649–2675. https://doi.org/10.1007/s11098-017-0976-4
2018 doi
-
[38]
Will Fleisher. 2019. Endorsement and Assertion.Noûs55, 2 (2019), 363–384. https://doi.org/10.1111/nous.12315
2019 doi
-
[39]
Will Fleisher. 2020. How to Endorse Conciliationism.Synthese198, 10 (2020), 9913–9939. https://doi.org/10.1007/s11229-020-02695-z
2020 doi
-
[40]
Will Fleisher. 2022. Understanding, idealization, and explainable AI.Episteme 19, 4 (2022), 534–560
2022
-
[41]
Eve Fleisig, Rediet Abebe, and Dan Klein. 2023. When the majority is wrong: Modeling annotator disagreement for subjective tasks. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 6715–6726
2023
-
[42]
Eve Fleisig, Su Lin Blodgett, Dan Klein, and Zeerak Talat. 2024. The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels.arXiv preprint arXiv:2405.05860(2024)
2024 arXiv
-
[43]
Bryan Frances and Jonathan Matheson. 2024. Disagreement. InThe Stanford Encyclopedia of Philosophy(Winter 2024 ed.), Edward N. Zalta and Uri Nodelman (Eds.). Metaphysics Research Lab, Stanford University
2024
-
[44]
Phoebe Friesen and Jordan Goldstein. 2022. Standpoint theory and the psy sciences: Can marginalization and critical engagement lead to an epistemic advantage?Hypatia37, 4 (2022), 659–687
2022
-
[45]
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al. 2022. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned.arXiv preprint arXiv:2209...
2022 arXiv
-
[46]
Mitchell L Gordon, Michelle S Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S Bernstein. 2022. Jury learning: Integrating dissenting voices into machine learning models. InProceedings of the 2022 CHI Conference on Human Factors in Computing S...
2022
-
[47]
Sandra Harding. 2004. A Socially Relevant Philosophy of Science? Resources From Standpoint Theory’s Controversiality.Hypatia19, 1 (2004), 25–47. https: //doi.org/10.1111/j.1527-2001.2004.tb01267.x FAccT ’25, June 23–26, 2025, Athens, Greece Sina Fazelpour and Will Fleisher
2004
-
[48]
Sandra Harding. 2009. Standpoint theories: Productively controversial.Hypatia 24, 4 (2009), 192–200
2009
-
[49]
2019.Objectivity and diversity: Another logic of scientific research
Sandra Harding. 2019.Objectivity and diversity: Another logic of scientific research. University of Chicago Press
2019
-
[50]
2004.The feminist standpoint theory reader: Intellectual and political controversies
Sandra G Harding. 2004.The feminist standpoint theory reader: Intellectual and political controversies. Psychology Press
2004
-
[51]
Remco Heesen, Liam Kofi Bright, and Andrew Zucker. 2019. Vindicating method- ological triangulation.Synthese196 (2019), 3067–3081
2019
-
[52]
Babak Heydari, Pedram Heydari, and Mohsen Mosleh. 2020. Not all bridges connect: integration in multi-community networks.The Journal of Mathematical Sociology44, 4 (2020), 199–220
2020
-
[53]
Babak Heydari, Mohsen Mosleh, and Kia Dalili. 2015. Efficient network struc- tures with separable heterogeneous connection costs.Economics Letters134 (2015), 82–85
2015
-
[54]
Lu Hong and Scott E Page. 2004. Groups of diverse problem solvers can out- perform groups of high-ability problem solvers.Proceedings of the National Academy of Sciences101, 46 (2004), 16385–16389
2004
-
[55]
Chris Jay Hoofnagle, Ashkan Soltani, Nathaniel Good, and Dietrich J Wambach
-
[56]
Sophie Horowitz. 2025. Higher-Order Evidence. InThe Stanford Encyclopedia of Philosophy(Spring 2025 ed.), Edward N. Zalta and Uri Nodelman (Eds.). Metaphysics Research Lab, Stanford University
2025
-
[57]
Saffron Huang, Divya Siddarth, Liane Lovitt, Thomas I Liao, Esin Durmus, Alex Tamkin, and Deep Ganguli. 2024. Collective constitutional ai: Aligning a language model with public input. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency. 1395–1417
2024
-
[58]
Kristen Intemann. 2010. 25 years of feminist empiricism and standpoint theory: Where are we now?Hypatia25, 4 (2010), 778–796
2010
-
[59]
Abigail Z Jacobs and Hanna Wallach. 2021. Measurement and fairness. InPro- ceedings of the 2021 ACM conference on fairness, accountability, and transparency. 375–385
2021
-
[60]
Shomik Jain, Vinith Suriyakumar, Kathleen Creel, and Ashia Wilson. 2024. Algorithmic Pluralism: A Structural Approach To Equal Opportunity. InThe 2024 ACM Conference on Fairness, Accountability, and Transparency. 197–206
2024
-
[61]
Akshita Jha, Aida Davani, Chandan K Reddy, Shachi Dave, Vinodkumar Prab- hakaran, and Sunipa Dev. 2023. SeeGULL: A stereotype benchmark with broad geo-cultural coverage leveraging generative models.arXiv preprint arXiv:2305.11840(2023)
2023 arXiv
-
[62]
Anna Jobin, Marcello Ienca, and Effy Vayena. 2019. The global landscape of AI ethics guidelines.Nature Machine Intelligence1, 9 (2019), 389–399
2019
-
[63]
Shivani Kapania, Alex S Taylor, and Ding Wang. 2023. A hunt for the snark: Annotator diversity in data practices. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–15
2023
-
[64]
S Vittal Katikireddi and Sean A Valles. 2015. Coupled ethical–epistemic anal- ysis of public health research and practice: Categorizing variables to improve population health and equity.American journal of public health105, 1 (2015), e36–e42
2015
-
[65]
Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, et al. 2024. The PRISM Alignment Project: What Participatory, Representative and Individualised Human Feedback Reveals About t...
2024 arXiv
-
[66]
Philip Kitcher. 1990. The Division of Cognitive Labor.Journal of Philosophy87, 1 (1990), 5–22. https://doi.org/10.2307/2026796
1990 doi
-
[67]
Jon Kleinberg, Jens Ludwig, Sendhil Mullainathan, and Cass R Sunstein. 2018. Discrimination in the Age of Algorithms.Journal of Legal Analysis10 (2018), 113–174
2018
-
[68]
Jon Kleinberg and Manish Raghavan. 2021. Algorithmic monoculture and social welfare.Proceedings of the National Academy of Sciences118, 22 (2021), e2018340118
2021
-
[69]
Tzu-Sheng Kuo, Quan Ze Chen, Amy X Zhang, Jane Hsieh, Haiyi Zhu, and Kenneth Holstein. 2024. PolicyCraft: Supporting Collaborative and Partici- patory Policy Design through Case-Grounded Deliberation.arXiv preprint arXiv:2409.15644(2024)
2024 arXiv
-
[70]
Tzu-Sheng Kuo, Aaron Lee Halfaker, Zirui Cheng, Jiwoo Kim, Meng-Hsin Wu, Tongshuang Wu, Kenneth Holstein, and Haiyi Zhu. 2024. Wikibench: Community-driven data curation for ai evaluation on wikipedia. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1–24
2024
-
[71]
Jaakko Kuorikoski and Caterina Marchionni. 2016. Evidential diversity and the triangulation of phenomena.Philosophy of Science83, 2 (2016), 227–247
2016
-
[72]
Hélène Landemore. 2017. Beyond the fact of disagreement? The epistemic turn in deliberative democracy.Social Epistemology31, 3 (2017), 277–295
2017
-
[73]
Jürgen Landes. 2021. The variety of evidence thesis and its independence of degrees of independence.Synthese198, 11 (2021), 10611–10641
2021
-
[74]
Seth Lazar. 2024. Legitimacy, Authority, and Democratic.Oxford Studies in Political Philosophy Volume 10(2024), 28
2024
-
[75]
Min Kyung Lee, Daniel Kusbit, Anson Kahng, Ji Tae Kim, Xinran Yuan, Allissa Chan, Daniel See, Ritesh Noothigattu, Siheon Lee, Alexandros Psomas, et al. 2019. WeBuildAI: Participatory framework for algorithmic governance.Proceedings of the ACM on Human-Computer Interaction3, CS...
2019
-
[76]
Adam Dahlgren Lindström, Leila Methnani, Lea Krause, Petter Ericson, Íñigo Martínez de Rituerto de Troya, Dimitri Coelho Mollo, and Roel Dobbe. 2024. AI Alignment through Reinforcement Learning from Human Feedback? Contradic- tions and Limitations.arXiv preprint arXiv:2406.18346(2024)
2024 arXiv
-
[77]
Zachary Lipton, Julian McAuley, and Alexandra Chouldechova. 2018. Does mitigating ML’s impact disparity require treatment disparity?Advances in neural information processing systems31 (2018)
2018
-
[78]
1990.Science as social knowledge: Values and objectivity in scientific inquiry
Helen E Longino. 1990.Science as social knowledge: Values and objectivity in scientific inquiry. Princeton University Press
1990
-
[79]
2015.The Epistemology of Disagreement
Jonathan Matheson. 2015.The Epistemology of Disagreement. Palgrave, New York
2015
-
[80]
Hugo Mercier and Dan Sperber. 2011. Argumentation: Its Adaptiveness and Efficacy.Behavioral and Brain Sciences34, 2 (2011), 94–111. https://doi.org/10. 1017/s0140525x10003031
2011
-
[81]
2000.Deliberative democracy or agonistic pluralism
Chantal Mouffe. 2000.Deliberative democracy or agonistic pluralism. Reihe Poli- tikwissenschaft / Institut für Höhere Studien, Abt. Politikwissenschaft, Vol. 72. Institut für Höhere Studien (IHS), Wien, Wien. 17 pages
2000
-
[82]
Ryan Muldoon. 2013. Diversity and the division of cognitive labor.Philosophy Compass8, 2 (2013), 117–125
2013
-
[83]
2006.Hearing the other side: Deliberative versus participatory democracy
Diana C Mutz. 2006.Hearing the other side: Deliberative versus participatory democracy. Cambridge University Press
2006
-
[84]
Hasti Narimanzadeh, Arash Badie-Modiri, Iuliia G Smirnova, and Ted Hsuan Yun Chen. 2023. Crowdsourcing subjective annotations using pairwise comparisons reduces bias and error compared to the majority-vote method.Proceedings of the ACM on Human-Computer Interaction7, CSCW2 (20...
2023
-
[85]
Helen Nissenbaum. 2019. Contextual integrity up and down the data food chain. Theoretical inquiries in law20, 1 (2019), 221–256
2019
-
[86]
2017.The diversity bonus
Scott E Page. 2017.The diversity bonus. Princeton University Press
2017
-
[87]
Samir Passi and Solon Barocas. 2019. Problem formulation and fairness. In Proceedings of the conference on fairness, accountability, and transparency. 39–48
2019
-
[88]
Ethan Perez, Sam Ringer, Kamil˙e Lukoši¯ut˙e, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al
-
[89]
Joshua C Peterson, Ruairidh M Battleday, Thomas L Griffiths, and Olga Rus- sakovsky. 2019. Human uncertainty makes classification more robust. InPro- ceedings of the IEEE/CVF international conference on computer vision. 9617–9626
2019
-
[90]
Katherine W Phillips. 2017. Commentary. What is the real value of diversity in organizations? Questioning our assumptions. InThe diversity bonus. Princeton University Press, 223–246
2017
-
[91]
Katherine W Phillips, Katie A Liljenquist, and Margaret A Neale. 2009. Is the pain worth the gain? The advantages and liabilities of agreeing with socially distinct newcomers.Personality and Social Psychology Bulletin35, 3 (2009), 336–350
2009
-
[92]
Katherine W Phillips and Denise Lewin Loyd. 2006. When surface and deep- level diversity collide: The effects on dissenting group members.Organizational behavior and human decision processes99, 2 (2006), 143–160
2006
-
[93]
arXiv preprint arXiv:2212.09251(2022)
Discovering language model behaviors with model-written evaluations. arXiv preprint arXiv:2212.09251(2022)
2022 arXiv
-
[94]
Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Díaz. 2021. On releasing Annotator-Level labels and information in datasets.arXiv preprint arXiv:2110.05699(2021)
2021 arXiv
-
[95]
Inioluwa Deborah Raji, Andrew Smart, Rebecca N White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes. 2020. Closing the AI accountability gap: Defining an end-to-end frame- work for internal algorithmic auditing. InProceedi...
2020
-
[96]
Quinetta M Roberson. 2019. Diversity in the workplace: A review, synthesis, and future research agenda.Annual Review of Organizational Psychology and Organizational Behavior6 (2019), 69–88
2019
-
[97]
Pratik S Sachdeva, Renata Barreto, Claudia von Vacano, and Chris J Kennedy
-
[98]
Gaile Pohlhaus. 2002. Knowing communities: An investigation of Harding’s standpoint epistemology.Social epistemology16, 3 (2002), 283–293
2002
-
[99]
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect?. In International Conference on Machine Learning. PMLR, 29971–30004
2023
-
[100]
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A Smith. 2019. The risk of racial bias in hate speech detection. InProceedings of the 57th annual meeting of the association for computational linguistics. 1668–1678. The Value of Disagreement in AI Design, Evaluat...
2019
-
[101]
Mike Schaekermann, Graeme Beaton, Minahz Habib, Andrew Lim, Kate Larson, and Edith Law. 2019. Understanding expert disagreement in medical data analysis through structured adjudication.Proceedings of the ACM on Human- Computer Interaction3, CSCW (2019), 1–23
2019
-
[102]
2000.Cognitive task analysis
Jan Maarten Schraagen, Susan F Chipman, and Valerie L Shalin. 2000.Cognitive task analysis. Psychology Press
2000
-
[103]
InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency
Assessing annotator identity sensitivity via item response theory: A case study in a hate speech corpus. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency. 1585–1603
2022
-
[104]
Catharine Saint-Croix. 2020. Privilege and Position: Formal Tools for Standpoint Epistemology.Res Philosophica97, 4 (2020), 489–524. https://doi.org/10.11612/ resphil.1953
2020
-
[105]
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, et al. 2024. A roadmap to pluralistic alignment.arXiv preprint arXiv:2402.05070(2024)
2024 arXiv
-
[106]
Dominik Stammbach, Philine Widmer, Eunjung Cho, Caglar Gulcehre, and Elliott Ash. 2024. Aligning Large Language Models with Diverse Political Viewpoints.arXiv preprint arXiv:2406.14155(2024)
2024 arXiv
-
[107]
Daniel Steel, Sina Fazelpour, Bianca Crewe, and Kinley Gillette. 2019. Informa- tion elaboration and epistemic effects of diversity.Synthese(2019), 1–21
2019
-
[108]
Daniel Steel, Sina Fazelpour, Kinley Gillette, Bianca Crewe, and Michael Burgess
-
[109]
Christopher Small, Michael Bjorkegren, Timo Erkkilä, Lynette Shaw, and Colin Megill. 2021. Polis: Scaling deliberation by mapping high dimensional opinion spaces.Recerca: revista de pensament i anàlisi26, 2 (2021)
2021
-
[110]
2007.Social empiricism
Miriam Solomon. 2007.Social empiricism. MIT press
2007
-
[111]
Elham Tabassi. 2023. Artificial Intelligence Risk Management Framework (AI RMF 1.0). (2023)
2023
-
[112]
Johanna Thoma. 2015. The Epistemic Division of Labor Revisited.Philosophy of Science82, 3 (2015), 454–472. https://doi.org/10.1086/681768
2015 doi
-
[113]
Briana Toole. 2021. Recent Work in Standpoint Epistemology.Analysis81, 2 (2021), 338–350. https://doi.org/10.1093/analys/anab026
2021 doi
-
[114]
Nancy Tuana. 2010. Leading with ethics, aiming for policy: New opportunities for philosophy of science.Synthese177 (2010), 471–492
2010
-
[115]
Nancy Tuana. 2013. Embedding philosophers in the practices of science: bringing humanities to the sciences.Synthese190, 11 (2013), 1955–1973
2013
-
[116]
Michael Strevens. 2003. The Role of the Priority Rule in Science.Journal of Philosophy100, 2 (2003), 55–79. https://doi.org/10.5840/jphil2003100224
2003 doi
-
[117]
Harini Suresh, Emily Tseng, Meg Young, Mary Gray, Emma Pierson, and Karen Levy. 2024. Participation in the age of foundation models. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency. 1609–1621
2024
-
[118]
Hamed Valizadegan, Quang Nguyen, and Milos Hauskrecht. 2012. Learning medical diagnosis models from multiple experts. InAMIA annual symposium proceedings, Vol. 2012. 921
2012
-
[119]
Kate Vredenburgh. 2022. The right to explanation.Journal of Political Philosophy 30, 2 (2022), 209–229
2022
-
[120]
Zeerak Waseem. 2016. Are you a racist or am i seeing things? annotator influence on hate speech detection on twitter. InProceedings of the first workshop on NLP and computational social science. 138–142
2016
-
[121]
Amy A Winecoff and Miranda Bogen. 2024. Improving governance outcomes through AI documentation: Bridging theory and practice.arXiv preprint arXiv:2409.08960(2024)
2024 arXiv
-
[122]
Alison Wylie. 2003. Why Standpoint Matters. InScience and other cultures: issues in philosophies of science and technology, Robert Figueroa and Sandra G. Harding (Eds.). Routledge, 26–48
2003
-
[123]
Nancy Tuana. 2017. Understanding coupled ethical-epistemic issues relevant to climate modeling and decision support science.Scientific integrity and ethics in the geosciences(2017), 155–173
2017
-
[124]
Alexandra N Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2021. Learning from disagreement: A survey.Journal of Artificial Intelligence Research72 (2021), 1385–1470
2021
-
[125]
Weiyu Zhang. 2015. Perceived procedural fairness in deliberation: Predictors and effects.Communication Research42, 3 (2015), 345–364
2015
-
[130]
Alison Wylie. 2017. What knowers know well: standpoint theory and gender archeology.Scientiae Studia15, 1 (2017), 13–38
2017
-
[131]
annotator ratio- nales
Omar Zaidan, Jason Eisner, and Christine Piatko. 2007. Using “annotator ratio- nales” to improve machine learning for text categorization. InHuman language technologies 2007: The conference of the North American chapter of the association for computational linguistics; proceed...
2007
-
[2012]
Behavioral advertising: The offer you can’t refuse.Harv. L. & Pol’y Rev.6 (2012), 273
2012
-
[2013]
Interactive team cognition.Cognitive science37, 2 (2013), 255–285
2013
-
[2018]
European Journal for Philosophy of Science8, 3 (2018), 761–780
Multiple diversity concepts and their ethical-epistemic implications. European Journal for Philosophy of Science8, 3 (2018), 761–780
2018
-
[2022]
Training a helpful and harmless assistant with reinforcement learning from human feedback.arXiv preprint arXiv:2204.05862(2022)
2022 arXiv
-
[2023]
In2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML)
A validity perspective on evaluating the justified use of data-driven decision-making algorithms. In2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML). IEEE, 690–704
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.