REVIEW 2 major objections 4 minor 52 references
Practitioner Insights on Fairness Requirements in the AI Development Life Cycle: An Interview Study
T0 review · 2 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read Fairness in AI is widely recognized but inconsistently applied: interviews with 26 practitioners across 23 countries find fairness requirements are often undocumented, validated ad hoc, and deprioritized against performance and deadlines.
desk verdict Competent, honestly reported interview study confirming that fairness is inconsistently practiced and deprioritized; the novelty claim is overstated and the evidence is self-report, but it deserves a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The study advances its argument with a semi-structured interview protocol of 15 open-ended questions mapped to four research questions: awareness and definition, translation and documentation of fairness requirements, emergence and challenges early in the SDLC, and implementation, validation, and trade-offs. Responses were analyzed with thematic analysis using open coding and card-sorting sessions to merge codes into themes, with per-theme saturation reported. This protocol is what lets the authors trace fairness from practitioners' conceptual understanding through requirements engineering and validation, and it is the backbone of the claim that practices are inconsistent.
What would settle it
An observational study that audits the actual artifacts of AI/ML teams—issue trackers, model cards, test reports, and requirement documents—and finds formal fairness documentation and validation in most projects would weaken the claim that fairness practice is inconsistent and deprioritized.
Extended reading notes
Core claim
Through a thematic analysis of 26 interviews with AI/ML practitioners in 23 countries, the study finds a consistent split between recognition and practice. Participants demonstrate an implicit understanding of fairness across three lenses—data quality and outcome fairness, model fairness, and ethical/operational fairness—and they readily identify early-lifecycle concerns such as dataset imbalance, under-representation, and sensitive attributes. Yet the translation of those concerns into requirements is uneven: some teams use metrics, model cards, and issue trackers, while others keep fairness only in meetings or not at all. Validation ranges from explicit group and individual fairness metric
Load-bearing premise
The conclusions rest on treating what 26 practitioners said in interviews as an accurate picture of what their teams actually do; if participants overstated or forgot their fairness work, the reported inconsistency and deprioritization could be mischaracterized.
Editorial extensions
If this is right
- If the claim holds, organizations cannot close the fairness gap with training alone; they need enforced documentation and metric requirements, because awareness already exists.
- Fairness that is not written into product requirements, acceptance criteria, or model cards will predictably lose to deadlines and feature work, as practitioners described.
- Using accuracy, F1-score, or confusion matrices as fairness proxies provides false assurance, since several participants equated high accuracy with fairness.
- Because fairness concerns typically enter at the data stage, early auditing of data collection, labeling, and representation is the highest-leverage intervention point.
- Context-specific fairness definitions must be chosen and agreed with stakeholders before metric selection, since conflicting definitions undermine validation and trade-off decisions.
Reading between the lines
- Implicit in the findings: a direct audit of team artifacts might reveal that documented practice is even thinner than the interviews suggest, because self-reports can overstate formal process.
- Applying the same interview protocol to teams in regulated sectors or teams with dedicated responsible-AI roles would likely show a different trade-off pattern, making the generalizability boundary testable.
- The themes suggest a concrete intervention—a lightweight fairness checklist covering data balance, sensitive-attribute handling, metric selection, and documentation—that could be evaluated for traceability improvements in real projects.
- The fact that some participants only recognized fairness after concrete scenarios suggests that terminology itself is a barrier; scenario-based elicitation may uncover more fairness awareness than direct questions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a qualitative interview study of 26 AI/ML practitioners from 23 countries, aimed at understanding how fairness is perceived, translated into requirements, handled in the early SDLC, and implemented, validated, and traded off against other project goals. Using thematic analysis with two independent coders and consensus reconciliation, the authors identify themes such as data quality and outcome fairness, inconsistent documentation, absence of formal fairness metrics, and the frequent deprioritization of fairness in favor of performance, deadlines, and features. The paper concludes with recommendations for standardized fairness definitions, metrics, and processes across the AI development lifecycle.
Significance. If the findings are accepted as accurate, the study provides a useful, holistic account of fairness requirements in real-world AI/ML development, complementing prior interview and survey studies by focusing on the entire SDLC rather than isolated stages. The paper's strengths include a systematic thematic analysis with two coders, per-theme saturation reporting, representative participant quotes, a supplementary coding spreadsheet for traceability, and a diverse participant sample. These methodological features make the study's descriptive findings credible and reproducible. However, the significance of the contribution depends on whether the authors' practice-level claims can be supported by self-report data alone.
major comments (2)
- [Abstract; Sections 4.4, 6.3.3, 6.4] The central claim—'practices are inconsistent, and fairness is often deprioritized'—is stated as a fact about organizational practice, but the evidence is entirely self-report. The paper concedes no formal member checks (Section 4.4), no triangulation with artifacts (Section 6.3.3), and no formal validation of the interview guide against the RQs (Section 6.4). Real-time prompts such as 'Am I correct?' can also steer responses. Because the contribution is presented as a practice-oriented account, this gap between 'practitioners reported X' and 'practices are X' is load-bearing. Please temper the abstract and all RQ summaries to 'participants reported/reported inconsistent practices,' or add an explicit justification for why self-report is treated as sufficient evidence for the practice-level claim.
- [Section 4.6 (Data Saturation)] The saturation account is internally inconsistent: it reports that validation-challenge saturation occurred at participant P25, but then states 'the overall saturation for the main aspects of fairness was reached by participant P24.' If one theme saturated only at P25, the overall claim cannot be P24. Please reconcile or rephrase the per-theme vs overall saturation claims, since the sufficiency of the sample is a methodological point that reviewers and readers will check. Also clarify whether saturation is being used as an ex-post description or as a stopping rule, as the current wording is ambiguous.
minor comments (4)
- [Table 6] The caption reads 'Final themes for answering R3'; this should be 'RQ3' for consistency with the other tables.
- [Section 4.7] In the 'Report production' paragraph, 'experiences in my context' should be 'experiences in our context' (or 'their context'), as the first-person singular appears to be an editing artifact.
- [Table 1] There are small typographical issues: 'Finance & Baking' should be 'Finance & Banking'; 'Geo-spatial Vision .5' is clearer as '0.5 years'; and P14's country entry 'PK' appears without a space. Please proofread the demographic table.
- [Section 6.2] The phrase 'our study provides the first holistic view' is a strong novelty claim. Given the acknowledged limitations regarding member checks and triangulation, consider softening to 'a holistic view' or substantiating the 'first' more precisely against the related-work comparison.
Circularity Check
No significant circularity: the interview findings are self-contained; minor same-author citations are contextual, not load-bearing.
full rationale
The paper's primary chain is empirical: 26 semi-structured interviews were coded through Braun & Clarke thematic analysis (Section 4.7), with representative quotes and initial codes in Tables 2-7 and a supplementary traceability spreadsheet [2]. The central finding - that practitioners recognize fairness dimensions but apply them inconsistently, often deprioritizing fairness - is a summary of these coded interview themes (Sections 5.1-5.4), not a mathematical or model-derived prediction. No parameter is fitted to a subset of the data and then used to 'predict' a closely related quantity; the saturation statements in Section 4.6 are descriptive stopping points in data collection, not out-of-sample predictions. The authors cite their own gray-literature study [32] in the Background and Discussion, and [12] shares a co-author, but those citations are contextual support for related work; the interview findings do not depend on the truth of [32]. The explicit limitations - no formal member checks (Section 4.4), no triangulation with project artifacts (Section 6.3.3), and no formal validation of the interview guide (Section 6.4) - weaken the strength of the evidence about actual organizational practice, but the gap is between self-reports and reality, not between a derived result and its own inputs. Under the required standard, no circular step can be exhibited with specific textual evidence of a definitional or constructional reduction.
Assumptions & free parameters
assumptions (4)
- domain assumption Interviewee self-reports accurately reflect real fairness practices.
- domain assumption Saturation claimed at participant P24 justifies the breadth of conclusions.
- domain assumption The 26-participant convenience sample supports cross-role and cross-domain thematic insight.
- domain assumption Braun and Clarke thematic analysis is a valid framework for this data.
Cite this review
Pith. "Pith review of Practitioner Insights on Fairness Requirements in the AI Development Life Cycle: An Interview Study." pith.science (2026). https://pith.science/paper/MZKCUO3E
@misc{pith2026251213830,
author = {Pith},
title = {Pith review of: Practitioner Insights on Fairness Requirements in the AI Development Life Cycle: An Interview Study},
year = {2026},
howpublished = {\url{https://pith.science/paper/MZKCUO3E}},
note = {Machine review of arXiv:2512.13830}
}
read the original abstract
Nowadays, Artificial Intelligence (AI), particularly Machine Learning (ML) and Large Language Models (LLMs), is widely applied across various contexts. However, the corresponding models often operate as black boxes, leading them to unintentionally act unfairly towards different demographic groups. This has led to a growing focus on fairness in AI software recently, alongside the traditional focus on the effectiveness of AI models. Through 26 semi-structured interviews with practitioners from different application domains and with varied backgrounds across 23 countries, we conducted research on fairness requirements in AI from software engineering perspective. Our study assesses the participants' awareness of fairness in AI / ML software and its application within the Software Development Life Cycle (SDLC), from translating fairness concerns into requirements to assessing their arising early in the SDLC. It also examines fairness through the key assessment dimensions of implementation, validation, evaluation, and how it is balanced with trade-offs involving other priorities, such as addressing all the software functionalities and meeting critical delivery deadlines. Findings of our thematic qualitative analysis show that while our participants recognize the aforementioned AI fairness dimensions, practices are inconsistent, and fairness is often deprioritized with noticeable knowledge gaps. This highlights the need for agreement with relevant stakeholders on well-defined, contextually appropriate fairness definitions, the corresponding evaluation metrics, and formalized processes to better integrate fairness into AI/ML projects.
Reference graph
Works this paper leans on
-
[1]
Interview Guide: Fairness Requirements in AI/ML Development
2025. Interview Guide: Fairness Requirements in AI/ML Development. https://docs.google.com/document/d/ 1m1N4Vz3pUWjiNSdbxj9LbpzIn20oBzwTzoC_nidWn4I
2025
-
[2]
Representative quotes and Thematic Analysis Application
2025. Representative quotes and Thematic Analysis Application. https://docs.google.com/spreadsheets/d/13dwcnEVtzbca_jd1nbRWwib73beXj-- YwwpiO3cn3uA/edit?gid=0#gid=0
2025
-
[3]
Julien Kiesse Bahangulu and Louis Owusu-Berko. 2025. Algorithmic bias, data ethics, and governance: Ensuring fairness, transparency and compliance in AI-powered business analytics applications.World J Adv Res Rev25, 2 (2025), 1746–63
2025
-
[4]
Sebastian Baltes and Paul Ralph. 2022. Sampling in software engineering research: A critical review and guidelines.Empirical Software Engineering 27, 4 (2022), 94
2022
-
[5]
Pierre W Banks, John C Hagedorn II, Alexandria Soybel, Delayne Michelle Coleman, Gabriel Rivera, and Namita Bhardwaj. 2025. Multiple mini interviews vs traditional interviews: investigating racial and socioeconomic differences in interview processes.Advances in Medical Education and Practice(2025), 157–163
2025
-
[6]
Luciano Baresi, Chiara Criscuolo, and Carlo Ghezzi. 2023. Understanding fairness requirements for ml-based software. In2023 IEEE 31st International Requirements Engineering Conference (RE). IEEE, 341–346
2023
-
[7]
Hi. I’m Molly, Your Virtual Interviewer!
Shreyan Biswas, Ji-Youn Jung, Abhishek Unnam, Kuldeep Yadav, Shreyansh Gupta, and Ujwal Gadiraju. 2024. “Hi. I’m Molly, Your Virtual Interviewer!” Exploring the impact of race and gender in AI-powered virtual interview experiences. InProceedings of the AAAI Conference on Human Computation and Crowdsourcing, Vol. 12. 12–22
2024
-
[8]
Zhenpeng Chen, Jie M Zhang, Max Hort, Mark Harman, and Federica Sarro. 2024. Fairness testing: A comprehensive survey and analysis of trends. ACM Transactions on Software Engineering and Methodology33, 5 (2024), 1–59
2024
Show all 52 references
-
[9]
Lu Cheng, Kush R Varshney, and Huan Liu. 2021. Socially responsible ai algorithms: Issues, purposes, and challenges.Journal of Artificial Intelligence Research71 (2021), 1137–1181
2021
-
[10]
Sapna Cheryan, Sianna A Ziegler, Amanda K Montoya, and Lily Jiang. 2017. Why are some STEM fields more gender balanced than others? Psychological bulletin143, 1 (2017), 1
2017
-
[11]
Victoria Clarke and Virginia Braun. 2017. Thematic analysis.The journal of positive psychology12, 3 (2017), 297–298
2017
-
[12]
Ronnie de Souza Santos, Matheus de Morais Leça, Reydne Santos, and Cleyton Magalhaes. 2025. Software Fairness Testing in Practice.arXiv e-prints (2025), arXiv–2506
2025
-
[13]
Greg Demirchyan. 2025. Algorithmic fairness: challenges to building an effective regulatory regime.Frontiers in Artificial Intelligence8 (2025), 1637134
2025
-
[14]
Wesley Hanwen Deng, Manish Nagireddy, Michelle Seng Ah Lee, Jatinder Singh, Zhiwei Steven Wu, Kenneth Holstein, and Haiyi Zhu. 2022. Exploring how machine learning practitioners (try to) use fairness toolkits. InProceedings of the 2022 ACM Conference on Fairness, Accountabilit...
2022
-
[15]
Wesley Hanwen Deng, Nur Yildirim, Monica Chang, Motahhare Eslami, Kenneth Holstein, and Michael Madaio. 2023. Investigating practices and opportunities for cross-functional collaboration around AI fairness in industry practice. InProceedings of the 2023 ACM Conference on Fairn...
2023
-
[16]
Alessandro Fabris, Nina Baranowska, Matthew J Dennis, David Graus, Philipp Hacker, Jorge Saldivar, Frederik Zuiderveen Borgesius, and Asia J Biega. 2025. Fairness and bias in algorithmic hiring: A multidisciplinary survey.ACM Transactions on Intelligent Systems and Technology1...
2025
-
[17]
Carmine Ferrara, Francesco Casillo, Carmine Gravino, Andrea De Lucia, and Fabio Palomba. 2024. Refair: Toward a context-aware recommender for fairness requirements engineering. InProceedings of the IEEE/ACM 46th International Conference on Software Engineering. 1–12
2024
-
[18]
Carmine Ferrara, Giulia Sellitto, Filomena Ferrucci, Fabio Palomba, and Andrea De Lucia. 2024. Fairness-aware machine learning engineering: how far are we?Empirical software engineering29, 1 (2024), 9
2024
-
[19]
Emilio Ferrara. 2024. Fairness and bias in artificial intelligence: A brief survey of sources, impacts, and mitigation strategies.Sci6, 1 (2024), 3
2024
-
[20]
Patricia I Fusch Ph D and Lawrence R Ness. 2015. Are we there yet? Data saturation in qualitative research. (2015)
2015
-
[21]
Greg Guest, Arwen Bunce, and Laura Johnson. 2006. How many interviews are enough? An experiment with data saturation and variability.Field methods18, 1 (2006), 59–82
2006
-
[22]
Lakshitha Gunasekara, Nicole El-Haber, Swati Nagpal, Harsha Moraliyage, Zafar Issadeen, Milos Manic, and Daswin De Silva. 2025. A Systematic Review of Responsible Artificial Intelligence Principles and Practice.Applied System Innovation8, 4 (2025), 97
2025
-
[23]
Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miro Dudik, and Hanna Wallach. 2019. Improving fairness in machine learning systems: What do industry practitioners need?. InProceedings of the 2019 CHI conference on human factors in computing systems. 1–16
2019
-
[24]
Oskar Jakobsson and Zuzana Rohacova. 2024. Requirement representation for safety-critical and fairness aware automotive perception systems. (2024). Paper Under Review Practitioner Insights on Fairness Requirements in the AI Development Life Cycle: An Interview Study 27
2024
-
[25]
Neil Joshi and Phil Burlina. 2021. AI fairness via domain adaptation.arXiv preprint arXiv:2104.01109(2021)
2021 arXiv
-
[26]
Qinghua Lu, Liming Zhu, Xiwei Xu, Jon Whittle, and Zhenchang Xing. 2022. Towards a roadmap on software engineering for responsible AI. In Proceedings of the 1st International Conference on AI Engineering: software Engineering for AI. 101–112
2022
-
[27]
2024.Advancing Ethical and Responsible AI: Exploring Fairness, Privacy, and Explainability through Causal Perspectives
Karima Makhlouf. 2024.Advancing Ethical and Responsible AI: Exploring Fairness, Privacy, and Explainability through Causal Perspectives. Ph. D. Dissertation. École polytechnique, France. https://tel.archives-ouvertes.fr/tel-04775522 PhD thesis, Artificial Intelligence [cs.AI]
2024
-
[28]
Louise McCormack and Malika Bendechache. 2024. Ethical ai governance: Methods for evaluating trustworthy ai.arXiv preprint arXiv:2409.07473 (2024)
2024 arXiv
-
[29]
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A survey on bias and fairness in machine learning. ACM computing surveys (CSUR)54, 6 (2021), 1–35
2021
-
[30]
Dena F Mujtaba and Nihar R Mahapatra. 2025. Behind the Screens: Uncovering Bias in AI-Driven Video Interview Assessments Using Counterfactuals. arXiv preprint arXiv:2505.12114(2025)
2025
-
[31]
Kelvin Mwita. 2022. Factors influencing data saturation in qualitative studies.A vailable at SSRN 4889752(2022)
2022
-
[32]
Thanh Nguyen, Chaima Boufaied, and Ronnie de Souza Santos. 2025. A Gray Literature Study on Fairness Requirements in AI-enabled Software Engineering.arXiv preprint arXiv:2512.07990(2025)
2025
-
[33]
Aastha Pant, Rashina Hoda, Chakkrit Tantithamthavorn, and Burak Turhan. 2025. Navigating fairness: practitioners’ understanding, challenges, and strategies in AI/ML development.Empirical Software Engineering30, 3 (2025), 1–38
2025
-
[34]
Urja Pawar, Donna O’shea, Susan Rea, and Ruairi O’reilly. 2020. Explainable AI in healthcare. In2020 international conference on cyber situational awareness, data analytics and assessment (CyberSA). IEEE, 1–2
2020
-
[35]
Nga Pham, Hung Pham Ngoc, and Anh Nguyen-Duc. 2025. Fairness for machine learning software in education: A systematic mapping study. Journal of Systems and Software219 (2025), 112244
2025
-
[36]
Mahima Pushkarna, Andrew Zaldivar, and Oddur Kjartansson. 2022. Data cards: Purposeful and transparent dataset documentation for responsible ai. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency. 1776–1826
2022
-
[37]
Bogdana Rakova, Jingying Yang, Henriette Cramer, and Rumman Chowdhury. 2021. Where responsible AI meets reality: Practitioner perspectives on enablers for shifting organizational practices.Proceedings of the ACM on Human-Computer Interaction5, CSCW1 (2021), 1–23
2021
-
[38]
Qusai Ramadan, Jukka Ruohonen, Abhishek Tiwari, Adam Alami, and Zeyd Boukhers. 2025. Towards Systematic Specification and Verification of Fairness Requirements: A Position Paper.arXiv preprint arXiv:2509.20387(2025)
2025
-
[39]
Seamus Ryan, Wanling Cai, Robert Bowman, and Gavin Doherty. 2025. Fairness Challenges in the Design of Machine Learning Applications for Healthcare.ACM Transactions on Computing for Healthcare6, 4 (2025), 1–26
2025
-
[40]
Seamus Ryan, Camille Nadal, and Gavin Doherty. 2023. Integrating fairness in the software design process: An interview study with hci and ml experts.IEEE Access11 (2023), 29296–29313
2023
-
[41]
Nripsuta Ani Saxena, Karen Huang, Evan DeFilippis, Goran Radanovic, David C Parkes, and Yang Liu. 2020. How do fairness definitions fare? Testing public attitudes towards three algorithmic definitions of fairness in loan allocations.Artificial Intelligence283 (2020), 103238
2020
-
[42]
Donghee Shin and Yong Jin Park. 2019. Role of fairness, accountability, and transparency in algorithmic affordance.Computers in Human Behavior 98 (2019), 277–284
2019
-
[43]
Vivek Singh, Anshuman Singh, and Kailash Joshi. 2022. Fair CRISP-DM: Embedding fairness in machine learning (ML) development life cycle. (2022)
2022
-
[44]
Jessie J Smith, Michael Madaio, Robin Burke, and Casey Fiesler. 2025. Pragmatic Fairness: Evaluating ML Fairness Within the Constraints of Industry. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. 628–638
2025
-
[45]
Aakash Sorathiya and Gouri Ginde. 2024. Ethical software requirements from user reviews: A systematic literature review.arXiv preprint arXiv:2410.01833(2024)
2024 arXiv
-
[46]
2009.Card sorting: Designing usable categories
Donna Spencer. 2009.Card sorting: Designing usable categories. Rosenfeld Media, Brooklyn, NY
2009
-
[47]
Julia Stoyanovich, Meike Zehlike, and Ke Yang. 2023. Fairness in Ranking: From Values to Technical Choices and Back. InCompanion of the 2023 International Conference on Management of Data. 7–12
2023
-
[48]
Xiaoli Tang and Han Yu. 2025. Towards trustworthy AI-empowered real-time bidding for online advertisement auctioning.Comput. Surveys57, 6 (2025), 1–36
2025
-
[49]
Nguyen Van Tuan, Tran Minh Quang, et al . 2020. Ethical Implications of Artificial Intelligence: A Systematic Review of Bias, Fairness, and Accountability.Artificial Intelligence and Machine Learning Review1, 1 (2020), 1–7
2020
-
[50]
Gianmario Voria, Giulia Sellitto, Carmine Ferrara, Francesco Abate, Andrea De Lucia, Filomena Ferrucci, Gemma Catolino, and Fabio Palomba. 2024. A catalog of fairness-aware practices in machine learning engineering.arXiv preprint arXiv:2408.16683(2024)
2024 arXiv
-
[51]
Yi Xu, Zhiyun Chen, and Mengyuan Dong. 2025. Shaping the Fairness Journey: The Roles of AI Literacy, Explanation, and Interpersonal Interaction in AI Interviews.International Journal of Human-Computer Studies(2025), 103629
2025
-
[52]
Yifan Yang, Mingquan Lin, Han Zhao, Yifan Peng, Furong Huang, and Zhiyong Lu. 2024. A survey of recent methods for addressing AI fairness and bias in biomedicine.Journal of Biomedical Informatics154 (2024), 104646
2024
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.