REVIEW 3 major objections 5 minor 43 references
A Question Bank to Assess AI Inclusivity: Mapping out the Journey from Diversity Errors to Inclusion Excellence
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper proposes a 253-question bank, organized into five pillars, for assessing AI inclusivity.
desk verdict Genuinely useful question bank for AI inclusivity, but the same-model simulated validation can't support the paper's validation claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the question bank itself: 253 questions organized under the five pillars of Humans, Data, Process, System, and Governance. The construction history is the argument — eight sequential versions that merge manual drafting from D&I guidelines, large-language-model prompts based on those guidelines, questions derived from a systematic review of D&I challenges, and fifteen questions from an existing responsible-AI question bank related to bias and fairness. The validation machinery is a simulated user study in which a large language model generated 70 personas over 14 AI-related roles, answered five research questions about the bank's relevance and usefulness, and produced feedback that the authors manually analyzed to refine the bank.
What would settle it
Recruit real AI practitioners from the same 14 roles and several of the same domains, ask them the five research questions about the 253-question bank, and compare their judgments of relevance, clarity, and usefulness with the simulated personas' responses; a large divergence would show the simulated study does not by itself validate the bank.
Extended reading notes
Core claim
The central claim is that AI inclusivity can be assessed with a dedicated question bank rather than left to general ethical principles or fairness metrics. The paper proposes that its 253 questions, mapped to the five pillars of humans, data, process, system, and governance, capture the relevant D&I considerations across the AI lifecycle, and that the simulated user study with 70 personas across 14 AI roles supports the bank's relevance, usefulness, educational value, and domain applicability. Feedback from the simulated personas led to refinements in seven questions, and the authors present this as validation that the bank is ready for adoption while acknowledging it has not yet been tested in real projects.
Load-bearing premise
The load-bearing premise is that responses from 70 AI-generated personas accurately stand in for how real AI practitioners in those roles and domains would judge the question bank; if the simulated voices are not representative, the study's validation claim loses its support.
Editorial extensions
If this is right
- Organizations can use the 253 questions as a pre-deployment checklist to surface exclusion risks in people, data, development processes, system behavior, and governance.
- Regulators and internal auditors get a common reference point for asking D&I questions that existing risk and explainability assessments do not cover.
- The five-pillar organization lets different roles, such as data scientists, product managers, UX designers, and policy advisors, see which inclusivity concerns fall in their lane.
- The bank can serve as an awareness and training tool, particularly for entry-level practitioners, by turning D&I principles into concrete yes/no questions.
Reading between the lines
- My inference: the real test of the bank is empirical; a deployment study with actual AI teams would reveal whether the questions are actionable or only sensible on paper.
- My inference: because much of the content derives from existing guidelines and a large language model's expansion of them, the bank may be blind to exclusion patterns that are not yet documented in those sources; mining AI-incident reports could feed new questions.
- My inference: the five-pillar structure invites aggregation into a maturity score or dashboard, which the paper lists only as future work but could be built directly from the current questions and would make the tool more useful to executives and regulators.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a 253-question bank for assessing the inclusivity of AI systems, organized into five pillars (Humans, Data, Process, System, Governance). The authors describe an iterative construction process that draws on D&I guidelines, a systematic literature review of D&I challenges, an existing Responsible AI question bank, and GPT-4o-generated question prompts. They then report a "simulated user study" in which 70 GPT-4o-generated personas answer five research questions about the relevance and usefulness of the question bank. The paper presents descriptive insights from this simulated evaluation, a comparison with prior AI question banks, and a discussion of threats to validity. The central claim is that the question bank, validated through this process, is an actionable tool for researchers, practitioners, and policymakers.
Significance. If the validation were sound, the question bank could be a genuinely useful practical resource for integrating D&I considerations into AI development and governance. The authors are transparent about their construction process, make a dataset publicly available ([37]), and address an underexplored niche relative to existing XAI and RAI question banks. The literature synthesis and the articulation of five assessment pillars have value. However, the paper's main evidential claim rests on a simulated evaluation in which the same large language model generated both many of the questions and all of the evaluator personas and responses. That is a self-referential loop that cannot support the claimed validity or effectiveness of the instrument. The paper's own limitation statement (Section VII.A) concedes that the question bank has not been tested in real-world projects, which directly undercuts the abstract's and conclusion's assertions of affirmation. The contribution is therefore contingent on a substantial reframing of the validation claim and on making the full instrument available for inspection.
major comments (3)
- [Section IV, Steps 1-4; Section VII.A] The simulated user study is a same-model self-assessment. GPT-4o generated many of the questions (as described in Section III.B for V2, V4, and V6) and also generated the 70 personas and their answers to the research questions. The authors' manual review checks coherence and alignment with study objectives only; it cannot establish that the responses represent real AI practitioners' judgments. Section VII.A concedes that the question bank "has not yet been tested in real-world AI development projects" and that "simulated feedback may not fully capture the complexities." These admissions directly contradict the abstract's characterization of the validation as rigorous and the Conclusion's statement (Section VIII) that the simulated user study findings "affirm its relevance and effectiveness." The simulation is not a valid proxy for human evaluation, so the central validation claim is unsupported.
- [Section VI.B, Table I and Conclusion] All reported findings—role-relevance frequencies in Table I, domain applicability themes, usefulness categories, and educational value—are derived from GPT-4o-generated persona responses. These are not empirical observations about AI practitioners; they are outputs of the same model that helped generate the question bank. The "insights from Q1-Q4" therefore describe the behavior of a language model, not the professional reasoning of data scientists, policy advisors, or UX designers. Consequently, the Conclusion's claim that the question bank is "relevant and effective" is not supported by the evidence presented in Section VI.B.
- [Section V; Section III.B] The full 253-question instrument is not included in the manuscript. Section V provides only aggregate counts per pillar and a handful of illustrative examples, while the cited dataset ([37]) is described as the simulated user study data rather than the complete question bank. Since the central contribution is an actionable assessment tool, the complete question bank should be included in the paper or in a clearly labeled supplementary appendix. Without it, readers cannot evaluate the content, assess pillar coverage, or use the tool in practice, which undermines the paper's stated purpose.
minor comments (5)
- [Section III.B V8 and Section IV Step 5] The manuscript is internally inconsistent about the final version number: Section III.B says the simulated study led to "the development of V8," while Section IV Step 5 states "we arrived at the final version (V9) of our question bank." Please reconcile the version numbering and clarify whether the 253-question count refers to V8, V9, or both.
- [Section III.A] The statement that using GPT "ensur[ed] ... minimizing human bias in the question formation process" is not substantiated; LLM-generated questions can readily encode biases from training data, and the authors' human-in-the-loop review is the actual bias-mitigation mechanism and should be credited as such.
- [Section I] The claims that "our research is unique" and that no existing study has proposed a structured question bank for inclusive AI are stronger than the later comparison in Section VI.D supports; please soften the introduction to acknowledge the adjacent XAI and RAI question banks while noting the distinct D&I focus.
- [Section VI.B, Table I] Please add a note explaining how the "frequency of relevant questions" was computed from the persona responses, and include a total row or column so readers can interpret the pillar-wise counts in context.
- [Throughout] There are typographical and phrasing issues, such as "that allows for a structured" in Section III.B V1 and the repeated "why do you think so?" in the research questions; a careful proofreading pass is needed.
Circularity Check
The simulated user study is a same-model self-assessment: GPT-4o generated many of the questions and then, as 70 personas, rated those questions; the paper itself concedes no real-world testing.
-
fitted input called prediction
[Section III.B (V2, V4, V6) and Section IV Steps 2 and 4; cf. Section VII.A]
"In V2 of the question bank, we leveraged GPT-4o to independently generate questions based on the D&I guidelines [13]. ... we once again employed GPT-4o to generate a set of questions based on the identified challenges ... Once the roles were finalized, we proceeded to persona creation using GPT-4o. ... For each of the 70 personas, we used GPT-4o to generate responses to the five research questions using our question bank as the reference [37]."
The validation loop uses the same model on both sides: GPT-4o generated a large portion of the question bank (V2, V4, V6), and then GPT-4o created the 70 personas and generated their answers to research questions about that same question bank. The findings that the question bank is 'relevant' and 'effective' are therefore the model's assessment of its own output, not independent evidence from human practitioners. The manual review only checks coherence and alignment with study objectives; it cannot establish that the simulated personas represent real AI professionals' judgments.
full rationale
The construction of the 253-question bank is not circular: it is grounded in external D&I guidelines, a Responsible AI question bank, and a systematic literature review, with manual author review at each version. The self-citations to the authors' prior guidelines [13] and systematic review [1] are legitimate prior work used as inputs, not as validation. The circularity lies in the validation claim. The simulated user study uses GPT-4o to generate both the questions (V2, V4, V6) and the personas' evaluations, so the reported 'relevance and effectiveness' findings are a same-model self-assessment rather than evidence of real-world utility. The paper's own threat-to-validity section concedes that the question bank has not been tested in real projects and that simulated feedback may not capture real deployment complexity, which confirms that this validation cannot carry the central claim. Because the artifact itself is independently assembled but its key validation step reduces by construction to the same model's output, a score of 6 is appropriate.
Assumptions & free parameters
assumptions (5)
- domain assumption The five pillars (Humans, Data, Process, System, Governance) are the correct and sufficient organizing framework for AI inclusivity.
- domain assumption The 46 D&I guidelines from Zowghi and da Rimini [13] are a valid grounding for generating assessment questions.
- domain assumption The 55 challenges for D&I in AI and 24 challenges for AI for D&I from Shams et al. [1] are comprehensive and correctly mapped to questions.
- domain assumption The RAI question bank [17] is a reliable external benchmark for identifying bias and fairness questions.
- ad hoc to paper GPT-4o-generated personas and their responses can serve as a valid proxy for real user feedback.
invented entities (1)
-
70 AI-generated personas
Cite this review
Pith. "Pith review of A Question Bank to Assess AI Inclusivity: Mapping out the Journey from Diversity Errors to Inclusion Excellence." pith.science (2026). https://pith.science/paper/AIIAJIG2
@misc{pith2026250618538,
author = {Pith},
title = {Pith review of: A Question Bank to Assess AI Inclusivity: Mapping out the Journey from Diversity Errors to Inclusion Excellence},
year = {2026},
howpublished = {\url{https://pith.science/paper/AIIAJIG2}},
note = {Machine review of arXiv:2506.18538}
}
read the original abstract
Ensuring diversity and inclusion (D&I) in artificial intelligence (AI) is crucial for mitigating biases and promoting equitable decision-making. However, existing AI risk assessment frameworks often overlook inclusivity, lacking standardized tools to measure an AI system's alignment with D&I principles. This paper introduces a structured AI inclusivity question bank, a comprehensive set of 253 questions designed to evaluate AI inclusivity across five pillars: Humans, Data, Process, System, and Governance. The development of the question bank involved an iterative, multi-source approach, incorporating insights from literature reviews, D&I guidelines, Responsible AI frameworks, and a simulated user study. The simulated evaluation, conducted with 70 AI-generated personas related to different AI jobs, assessed the question bank's relevance and effectiveness for AI inclusivity across diverse roles and application domains. The findings highlight the importance of integrating D&I principles into AI development workflows and governance structures. The question bank provides an actionable tool for researchers, practitioners, and policymakers to systematically assess and enhance the inclusivity of AI systems, paving the way for more equitable and responsible AI technologies.
Figures
Reference graph
Works this paper leans on
-
[37]
The Simulates User Study Dataset of the Paper Titled “A Question Bank to Assess Inclusive AI
R. A. Shams, D. Zowghi, and M. Bano, “The Simulates User Study Dataset of the Paper Titled “A Question Bank to Assess Inclusive AI”,”
-
[1]
Ai and the quest for diversity and inclusion: A systematic literature review,
R. A. Shams, D. Zowghi, and M. Bano, “Ai and the quest for diversity and inclusion: A systematic literature review,”AI and Ethics, pp. 1–28, 2023
work page 2023
-
[2]
Ai for all: defining the what, why, and how of inclusive ai,
T. Avellan, S. Sharma, and M. Turunen, “Ai for all: defining the what, why, and how of inclusive ai,” inProceedings of the 23rd International Conference on Academic Mindtrek, 2020, pp. 142–144
work page 2020
-
[3]
What does a software engi- neer look like? exploring societal stereotypes in llms,
M. Bano, H. Gunatilake, and R. Hoda, “What does a software engi- neer look like? exploring societal stereotypes in llms,”arXiv preprint arXiv:2501.03569, 2025
arXiv 2025
-
[4]
Artificial intelligence in healthcare: past, present and future,
F. Jiang, Y . Jiang, H. Zhi, Y . Dong, H. Li, S. Ma, Y . Wang, Q. Dong, H. Shen, and Y . Wang, “Artificial intelligence in healthcare: past, present and future,”Stroke and vascular neurology, vol. 2, no. 4, 2017
work page 2017
-
[5]
Diversity and Inclusion in AI for Recruitment: Lessons from Industry Workshop
M. Bano, D. Zowghi, F. Mourao, S. Kaur, and T. Zhang, “Diversity and inclusion in ai for recruitment: Lessons from industry workshop,”arXiv preprint arXiv:2411.06066, 2024
work page Pith review arXiv 2024
-
[6]
Ai-driven healthcare: A survey on ensuring fairness and mitigating bias,
S. V . Chinta, Z. Wang, X. Zhang, T. D. Viet, A. Kashif, M. A. Smith, and W. Zhang, “Ai-driven healthcare: A survey on ensuring fairness and mitigating bias,”arXiv preprint arXiv:2407.19655, 2024
arXiv 2024
-
[7]
Ethical concerns while using artificial intelligence in recruitment of employees,
A. Gupta and M. Mishra, “Ethical concerns while using artificial intelligence in recruitment of employees,” 2022
work page 2022
Show all 43 references
-
[8]
Diversity, equity, and inclusion, and the deployment of artificial intelligence within the department of defense,
S. Darwish, A. Bragaw-Butler, P. Marcelli, and K. Gassner, “Diversity, equity, and inclusion, and the deployment of artificial intelligence within the department of defense,” inProceedings of the AAAI Symposium Series, vol. 3, no. 1, 2024, pp. 348–353
2024
-
[9]
Ai for all: Opera- tionalising diversity and inclusion requirements for ai systems,
M. Bano, D. Zowghi, V . Gervasi, and R. Shams, “Ai for all: Opera- tionalising diversity and inclusion requirements for ai systems,”arXiv preprint arXiv:2311.14695, 2023
2023 arXiv
-
[10]
Exploring ai’s role in supporting diversity and inclusion initiatives in multicultural marketplaces,
D. Mariyono and A. N. A. Akmal, “Exploring ai’s role in supporting diversity and inclusion initiatives in multicultural marketplaces,”Inter- national Journal of Religion, vol. 5, no. 10, pp. 10–61 707, 2024
2024
-
[11]
Ai fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias,
R. K. Bellamy, K. Dey, M. Hind, S. C. Hoffman, S. Houde, K. Kannan, P. Lohia, J. Martino, S. Mehta, A. Mojsilovi ´cet al., “Ai fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias,”IBM Journal of Research and Development, vol. 63, no. 4/5, pp. 4–1, 2019
2019
-
[12]
Diversity, equity, and inclusion in artificial intelligence: an evaluation of guidelines,
G. Cachat-Rosset and A. Klarsfeld, “Diversity, equity, and inclusion in artificial intelligence: an evaluation of guidelines,”Applied Artificial Intelligence, vol. 37, no. 1, p. 2176618, 2023
2023
-
[13]
Diversity and inclusion in artificial intelligence,
D. Zowghi and F. da Rimini, “Diversity and inclusion in artificial intelligence,”arXiv preprint arXiv:2305.12728, 2023
2023 arXiv
-
[14]
Questioning the ai: informing design practices for explainable ai user experiences,
Q. V . Liao, D. Gruen, and S. Miller, “Questioning the ai: informing design practices for explainable ai user experiences,” inProceedings of the 2020 CHI conference on human factors in computing systems, 2020, pp. 1–15
2020
-
[15]
Question- driven design process for explainable ai user experiences,
Q. V . Liao, M. Pribi ´c, J. Han, S. Miller, and D. Sow, “Question- driven design process for explainable ai user experiences,”arXiv preprint arXiv:2104.03483, 2021
2021 arXiv
-
[16]
Qb4aira: A question bank for ai risk assessment,
S. U. Lee, H. Perera, B. Xia, Y . Liu, Q. Lu, L. Zhu, O. Salvado, and J. Whittle, “Qb4aira: A question bank for ai risk assessment,”arXiv preprint arXiv:2305.09300, 2023
2023 arXiv
-
[17]
Responsible ai question bank: A comprehensive tool for ai risk assessment,
S. U. Lee, H. Perera, Y . Liu, B. Xia, Q. Lu, and L. Zhu, “Responsible ai question bank: A comprehensive tool for ai risk assessment,”arXiv preprint arXiv:2408.11820, 2024
2024 arXiv
-
[18]
Measuring responsible artificial intelligence (rai) in banking: a valid and reliable instrument,
J. Ratzan and N. Rahman, “Measuring responsible artificial intelligence (rai) in banking: a valid and reliable instrument,”AI and Ethics, pp. 1–19, 2023
2023
-
[19]
Role of artificial intelligence (ai) in meeting diversity, equality and inclusion (dei) goals,
R. B. Jora, K. K. Sodhi, P. Mittal, and P. Saxena, “Role of artificial intelligence (ai) in meeting diversity, equality and inclusion (dei) goals,” in2022 8th international conference on advanced computing and communication systems (ICACCS), vol. 1. IEEE, 2022, pp. 1687–1690
2022
-
[20]
An exploratory study on role of artificial intelligence in overcoming biases to promote diversity and inclusion practices,
B. Rathore, M. Mathur, and S. Solanki, “An exploratory study on role of artificial intelligence in overcoming biases to promote diversity and inclusion practices,”Impact of artificial intelligence on organizational transformation, pp. 147–164, 2022
2022
-
[21]
Ai for all: Identify- ing ai incidents related to diversity and inclusion,
R. A. Shams, D. Zowghi, and M. Bano, “Ai for all: Identify- ing ai incidents related to diversity and inclusion,”arXiv preprint arXiv:2408.01438, 2024
2024 arXiv
-
[22]
Accountability in human and artificial intelligence decision-making as the basis for diversity and educational inclusion,
K. Porayska-Pomsta and G. Rajendran, “Accountability in human and artificial intelligence decision-making as the basis for diversity and educational inclusion,”Artificial intelligence and inclusive education: Speculative futures and emerging practices, pp. 39–59, 2019
2019
-
[23]
Identifying explanation needs of end-users: Applying and extending the xai question bank,
L. Sipos, U. Sch ¨afer, K. Glinka, and C. M ¨uller-Birn, “Identifying explanation needs of end-users: Applying and extending the xai question bank,” inProceedings of Mensch und Computer 2023, 2023, pp. 492– 497
2023
-
[24]
Qb4aira: A question bank for responsible ai risk assessment,
S. U. Lee, H. Perera, B. Xia, Y . Liu, Q. Lu, L. Zhu, O. Salvado, and J. Whittle, “Qb4aira: A question bank for responsible ai risk assessment,” IEEE Software, 2024
2024
-
[25]
Co- designing checklists to understand organizational challenges and op- portunities around fairness in ai,
M. A. Madaio, L. Stark, J. Wortman Vaughan, and H. Wallach, “Co- designing checklists to understand organizational challenges and op- portunities around fairness in ai,” inProceedings of the 2020 CHI conference on human factors in computing systems, 2020, pp. 1–14
2020
-
[26]
Bias in artificial intelligence algorithms and recommendations for mitigation,
L. H. Nazer, R. Zatarah, S. Waldrip, J. X. C. Ke, M. Moukheiber, A. K. Khanna, R. S. Hicklen, L. Moukheiber, D. Moukheiber, H. Ma et al., “Bias in artificial intelligence algorithms and recommendations for mitigation,”PLOS Digital Health, vol. 2, no. 6, p. e0000278, 2023
2023
-
[27]
Ai risk management framework,
N. I. of Standards and Technology, “Ai risk management framework,” July 2024. [Online]. Available: https://www.nist.gov/itl/ai- risk-management-framework
2024
-
[28]
Microsoft responsible ai impact as- sessment template,
Microsoft, “Microsoft responsible ai impact as- sessment template,” June 2022. [Online]. Avail- able: https://blogs.microsoft.com/wp-content/uploads/prod/sites/5/2022/ 06/Microsoft-RAI-Impact-Assessment-Template.pdf
2022
-
[29]
Microsoft responsible ai impact as- sessment guide,
T. Microsoft, “Microsoft responsible ai impact as- sessment guide,” June 2022. [Online]. Avail- able: https://blogs.microsoft.com/wp-content/uploads/prod/sites/5/2022/ 06/Microsoft-RAI-Impact-Assessment-Guide.pdf
2022
-
[30]
The assessment list for trustworthy artificial intelligence,
E. Commission, “The assessment list for trustworthy artificial intelligence,” 2000. [Online]. Available: https://futurium.ec.europa.eu/ en/european-ai-alliance/pages/welcome-altai-portal
2000
-
[31]
Algorithmic impact assessment tool,
G. of Canada, “Algorithmic impact assessment tool,” 2023. [Online]. Available: https://www.canada.ca/en/government/system/digital- government/digital-government-innovations/responsible-use- ai/algorithmic-impact-assessment.html
2023
-
[32]
Artificial intelligence assurance framework,
N. Australia, “Artificial intelligence assurance framework,” 2022. [Online]. Available: https://www.digital.nsw.gov.au/sites/default/files/ 2022-09/nsw-government-assurance-framework.pdf
2022
-
[33]
Chatgpt vs. bard: a comparative study,
I. Ahmed, A. Roy, M. Kajol, U. Hasan, P. P. Datta, and M. R. Reza, “Chatgpt vs. bard: a comparative study,”Authorea Preprints, 2023
2023
-
[34]
User simulations for evaluating answers to question series,
J. Lin, “User simulations for evaluating answers to question series,” Information processing & management, vol. 43, no. 3, pp. 717–729, 2007
2007
-
[35]
Validating simulations of user query variants,
T. Breuer, N. Fuhr, and P. Schaer, “Validating simulations of user query variants,” inEuropean Conference on Information Retrieval. Springer, 2022, pp. 80–94
2022
-
[36]
Design models and deployment strategy to effectively manage question bank for assessment in large enterprises,
S. K. Ghosh, A. Tiwari, A. Tiwary, G. Ajeesh, and N. Raghavendra, “Design models and deployment strategy to effectively manage question bank for assessment in large enterprises,” in2012 IEEE international conference on technology enhanced education (ICTEE). IEEE, 2012, pp. 1–6
2012
-
[38]
Developing gdpr compliant user data policies for internet of things,
M. Barati, I. Petri, and O. F. Rana, “Developing gdpr compliant user data policies for internet of things,” inProceedings of the 12th IEEE/ACM International conference on utility and cloud computing, 2019, pp. 133– 141
2019
-
[39]
Data sovereignty: A review,
P. Hummel, M. Braun, M. Tretter, and P. Dabrock, “Data sovereignty: A review,”Big Data & Society, vol. 8, no. 1, p. 2053951720982012, 2021
2021
-
[40]
Is your policy compliant? a deep learning-based empirical study of privacy policies’ compliance with gdpr,
T. A. Rahat, M. Long, and Y . Tian, “Is your policy compliant? a deep learning-based empirical study of privacy policies’ compliance with gdpr,” inProceedings of the 21st Workshop on Privacy in the Electronic Society, 2022, pp. 89–102
2022
-
[41]
Language, national origin, and employment discrimi- nation: The importance of the eeoc guidelines,
A. J. Robinson, “Language, national origin, and employment discrimi- nation: The importance of the eeoc guidelines,”U. Pa. L. Rev., vol. 157, p. 1513, 2008
2008
-
[42]
Metrics, explain- ability and the european ai act proposal,
F. Sovrano, S. Sapienza, M. Palmirani, and F. Vitali, “Metrics, explain- ability and the european ai act proposal,”J, vol. 5, no. 1, pp. 126–138, 2022
2022
-
[2025]
Available: https://doi.org/10.5281/zenodo.15086407
[Online]. Available: https://doi.org/10.5281/zenodo.15086407
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.