REVIEW 3 major objections 5 minor 73 references
Bias, Accuracy, and Trust: Gender-Diverse Perspectives on Large Language Models
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper argues that users' gender identity changes how they perceive ChatGPT's bias, accuracy, and trustworthiness, with non-binary participants especially likely to find the model's gender portrayals condescending.
desk verdict A rare, well-reported qualitative study of gender-diverse LLM users; the qualitative core is strong, but the quantitative trust analysis is underreported and likely confounded. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a 16-prompt case study: four everyday scenarios (buying a car, navigating college, experiences of a person, applying for jobs) run with "person" and with "man", "woman", and "non-binary person" as the subject, plus a semi-structured interview protocol with pre- and post-interview trust surveys. The prompt set exposes how ChatGPT's character construction changes with gender wording, and the interviews expose how those constructions are read. The trust survey separates morality-based from performance-based trust, which is what lets the study distinguish the trust profile of non-binary participants from that of women and men.
What would settle it
A pre-registered replication with a larger sample, matched LLM/AI backgrounds, and blinded analysts could falsify the perception claim: if non-binary participants no longer rate the non-binary prompts as more condescending than other groups do, the reported effect is an artifact of recruitment or background rather than gender identity.
Extended reading notes
Core claim
The central claim is that ChatGPT's responses to gender-specified prompts are not neutral: the model assigns names, pronouns, adjectives, goals, and challenges that follow societal gender norms, with non-binary characters reduced to identity struggles and women to emotional or traditional roles. Interview data show that this pattern is perceived differently by identity, with non-binary/transgender participants reporting condescending and stereotypical portrayals, women criticizing outdated tropes, and men noticing a lack of diversity but fewer concerns. Trust ratings after the interviews differed by gender, with men reporting higher trust overall and non-binary participants reporting higher performance-based trust but morality-based trust similar to women's. The authors conclude that the same model output can be read as acceptable, biased, or harmful depending on who reads it, so evaluations of LLM fairness should include gender-diverse evaluators.
Load-bearing premise
The comparisons by gender assume that 25 participants recruited through university email lists, Slack, X, and one online panel are representative of gender-diverse user populations, so differences attributed to gender could instead reflect prior experience with LLMs or with bias.
Editorial extensions
If this is right
- Gendered prompts change what ChatGPT produces: non-binary characters get they/them pronouns, neutral-to-masculine names, and storylines centered on identity struggles, while men get career-focused, assertive narratives and women get emotional or independence-focused ones.
- Non-binary/transgender participants experience ChatGPT's non-binary portrayals as condescending, cis-centric, and reductive, while women and men find the same outputs less problematic.
- Perceived accuracy does not differ across gender groups; participants in all groups notice errors mainly in technical and creative tasks.
- Trust is gender-dependent: men trust ChatGPT more overall, non-binary participants trust its performance more than its morality, and trust declines after reflection in participants with medium bias knowledge.
- Users want LLMs to diversify training data with real lived experiences, respond to all genders with equal depth, ask clarifying questions, and be transparent about sources and limitations.
Reading between the lines
- A testable extension: the 16-prompt suite could be turned into an automated gender-sensitivity audit by scoring outputs for identity-reduction markers, such as how much of a narrative centers on struggles rather than the scenario, and comparing those scores with the human ratings collected here.
- If the trust drop in medium-bias-knowledge participants is real, then simply asking people to reflect on bias during an interview changes their trust; a controlled experiment with a reflection-only condition could separate this effect from experimenter demand, which the authors themselves say they cannot do.
- The finding that "person" prompts still default to gendered stories suggests that neutral prompt wording does not produce neutral output; developers aiming for gender-neutral responses may need explicit constraints rather than neutral phrasing alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a mixed-methods study of how gender-diverse users perceive bias, accuracy, and trust in ChatGPT. The authors conducted 25 semi-structured interviews (9 non-binary/transgender, 8 women, 8 men), combined with a content analysis of ChatGPT responses to gendered and neutral prompts, and collected quantitative trust ratings before and after the interviews. The central qualitative claims are that gendered prompts elicit more identity-specific and stereotypical responses; that non-binary/transgender participants experience these responses as condescending and reductive; that perceived accuracy is fairly consistent across groups; and that trust varies by gender, with men reporting the highest trust. The paper also presents participant suggestions and design implications for more inclusive LLMs.
Significance. If the results hold, the paper makes a useful contribution to CSCW/HCI by showing that user perceptions of LLM bias and trust are group-dependent in ways that benchmark-only evaluations cannot capture. The study has notable strengths: it centers the voices of non-binary and transgender participants, provides direct participant quotes as evidence, describes its qualitative coding process transparently, avoids causal claims about the time effect, and makes its supplementary materials available. The qualitative findings on perceived bias are well supported by the reported quotes and thematic structure. However, the quantitative trust result—which appears in the abstract and in RQ4—is currently underreported and potentially confounded, so the headline gender-trust claim needs additional analysis before it can be relied upon.
major comments (3)
- [§4.8.2] The quantitative trust-by-gender result is statistically underreported and potentially confounded. The authors fit three separate general linear models—one for gender, one for AI/LLM expertise, and one for bias knowledge—and report only chi-square statistics and p-values. They do not report coefficients, degrees of freedom, effect sizes, confidence intervals, or post-hoc group comparisons. More importantly, they never fit a model containing gender together with expertise and bias knowledge, even though the same section shows that expertise increases trust and bias knowledge decreases trust. If the male subsample happens to have more AI/LLM expertise or less bias knowledge, the reported gender differences in §4.8.2 could be an artifact of group composition rather than gender identity. Table 1 does not cross-tabulate gender with expertise or bias knowledge, so the claim in §4.1 that women and men were recruited with 'similar backgrounds' cannot be verified. Section 6 lists limitations but does not flag this specific confound. Because the abstract and RQ4 state that 'trustworthiness varied by gender,' this issue is load-bearing and must be addressed with a combined model and fuller reporting.
- [§3.1] The content analysis of ChatGPT responses in RQ1 lacks essential methodological detail. The authors report counts such as 'resilience (21)' for non-binary characters, 'Emma (6)' for women, and various counts for goals, challenges, and symbols, but they do not describe how these categories were defined, whether codes were developed independently or through consensus, whether counts were based on unique responses, or how the 10 repetitions per prompt were handled. Without a codebook or any reliability assessment, the quantitative flavor of these counts is not reproducible. This matters because the claim that 'gendered prompts elicit richer, more identity-specific responses' is one of the paper's central contributions. The authors should either present this as purely qualitative illustration or provide the coding scheme and, if appropriate, a reliability analysis.
- [§4.8.2] The interpretation of the time × bias-knowledge interaction is speculative. The authors hypothesize that participants with medium bias knowledge 'had enough working knowledge to recognize the concept of bias, but lacked the depth of understanding necessary to contextualize how it manifests in LLMs.' While they appropriately avoid causal claims about the time effect, they do not consider plausible alternative explanations such as regression to the mean, differential item difficulty between pre- and post-surveys, or experimenter demand. The model also does not include gender or expertise interactions with time, so the claim that the interview had a 'pronounced effect on this group' is not directly tested. This does not invalidate the qualitative findings, but it should be presented as one possible interpretation rather than a conclusion.
minor comments (5)
- [Throughout] The manuscript contains numerous typos and grammatical errors that should be corrected, including 'prevalant' (§2.1), 'lagnuage' (§2.1), 'includive' (§2.1), 'comprimising' (§2.2), 'percieved' (§2.2), 'Rersonal' (§3.1.1), 'modic logic' (§5.2.2), 'outout' (§5.2.2), 'Explanaitions' (§5.2.2), and 'bakcground' in the Figure 3 caption.
- [References] Reference [64] lists 'Authors Unknown' as the author and appears to be a placeholder; this citation must be completed or removed before publication.
- [§4.8.2] The description of the gender model says the model 'included trust type, participants' gender, and the time,' but the subsequent discussion of an interaction between gender and trust type implies that interactions were modeled. Please state explicitly which main effects and interactions were included in each model.
- [Table 1] The column header 'Knowledge of bias in LLMs & LLM/AI background)' contains an unmatched parenthesis and should be cleaned up.
- [§4.1] The recruitment description would benefit from clarifying how the three non-binary participants recruited via userinterviews.com differed from the university-recruited participants, since combining recruitment channels can introduce unmeasured background differences.
Circularity Check
No significant circularity: the paper's claims are empirical interpretations of interview and ChatGPT-response data, not quantities defined by fitted parameters.
full rationale
The paper contains no derivation chain in which an output quantity is defined in terms of a fitted input. RQ1 is a descriptive analysis of ChatGPT outputs (Section 3.1): gendered prompts are observed to produce identity-specific narrative content, with no parameter fitted from those outputs to 'predict' the same outputs. RQ2–RQ5 are grounded-theory interpretations of 25 participant interviews (Sections 4.4–5.1); perceived bias, accuracy, trust, and suggestions are reported as participant statements and codes, not as quantities derived from the study's own assumptions. The quantitative trust analysis (Section 4.8.2) uses an external validated trust instrument (Malle and Ullman MDMT, refs [44–46, 62]) and separate general linear models on self-reported trust ratings; although the paper does not report a single model with gender, expertise, and bias knowledge together, that is a statistical completeness or confounding concern, not a circular one, because the gender effect is not constructed from the covariates. The only author self-citation is [21], used once in the introduction ('These biases not only undermine user trust in ML models [21]') as background motivation; it does not carry the paper's conclusions. No uniqueness theorem, no ansatz smuggled in via citation, and no renamed empirical pattern is load-bearing. The central findings are self-contained empirical results grounded in the collected qualitative and quantitative data.
Assumptions & free parameters
assumptions (4)
- domain assumption Participants' self-identified gender categories (men, women, non-binary and transgender) are stable, mutually exclusive, and a valid analytic grouping for comparing perceptions.
- domain assumption The four curated prompts and their gendered variants are neutral enough to elicit representative LLM behavior, and ChatGPT 3.5 responses generated in February 2024 are treated as representative of LLM outputs.
- domain assumption Consensus-based thematic coding without inter-rater reliability yields valid themes.
- domain assumption The trust survey items from the multidimensional measure of trust measure morality-based and performance-based trust as intended in this population.
Cite this review
Pith. "Pith review of Bias, Accuracy, and Trust: Gender-Diverse Perspectives on Large Language Models." pith.science (2026). https://pith.science/paper/TRK2AYCJ
@misc{pith2026250621898,
author = {Pith},
title = {Pith review of: Bias, Accuracy, and Trust: Gender-Diverse Perspectives on Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/TRK2AYCJ}},
note = {Machine review of arXiv:2506.21898}
}
read the original abstract
Large language models (LLMs) are becoming increasingly ubiquitous in our daily lives, but numerous concerns about bias in LLMs exist. This study examines how gender-diverse populations perceive bias, accuracy, and trustworthiness in LLMs, specifically ChatGPT. Through 25 in-depth interviews with non-binary/transgender, male, and female participants, we investigate how gendered and neutral prompts influence model responses and how users evaluate these responses. Our findings reveal that gendered prompts elicit more identity-specific responses, with non-binary participants particularly susceptible to condescending and stereotypical portrayals. Perceived accuracy was consistent across gender groups, with errors most noted in technical topics and creative tasks. Trustworthiness varied by gender, with men showing higher trust, especially in performance, and non-binary participants demonstrating higher performance-based trust. Additionally, participants suggested improving the LLMs by diversifying training data, ensuring equal depth in gendered responses, and incorporating clarifying questions. This research contributes to the CSCW/HCI field by highlighting the need for gender-diverse perspectives in LLM development in particular and AI in general, to foster more inclusive and trustworthy systems.
Figures
Reference graph
Works this paper leans on
-
[69]
J Williams and R Wilson. 2023. I wouldn’t say offensive but...: Disability-Centered Perspectives on Large Language Models. Journal of Accessible AI 2, 3 (2023), 45–60
work page 2023
-
[1]
Supplementary materials for LLM Bias
2025. Supplementary materials for LLM Bias. https://osf.io/ywmtg/?view_only=7a38b88af78e45ad85a33b70f88f717f
work page 2025
-
[2]
Abubakar Abid, Maheen Farooqi, and James Zou. 2021. Persistent Anti-Muslim Bias in Large Language Models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (Virtual Event, USA) (AIES ’21). Association for Computing Machinery, New York, NY, USA. doi:10.1145/3461702.3462624
arXiv 2021
-
[3]
Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner. 2016. Machine Bias. ProPublica May 23 (2016). https://www.propublica.org/article/ machine-bias-risk-assessments-in-criminal-sentencing
work page 2016
-
[4]
Julian Ashwin, Aditya Chhabra, and Vijayendra Rao. 2023. Using large language models for qualitative analysis can introduce serious bias. arXiv preprint arXiv:2309.17147 (2023)
arXiv 2023
-
[5]
Marion Bartl and Susan Leavy. 2024. From’Showgirls’ to’Performers’: Fine-tuning with Gender-inclusive Language for Bias Reduction in LLMs. arXiv preprint arXiv:2407.04434 (2024)
work page Pith review arXiv 2024
-
[6]
Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT) . ACM, 610–623. doi:10.1145/ 3442188.3445922
arXiv 2021
-
[7]
Reuben Binns and Michael Veale. 2022. Adhering, Steering, and Queering: Treatment of Gender in Natural Language Generation. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* 2022) . ACM, 142–155
work page 2022
Show all 73 references
-
[8]
Virginia Braun and Victoria Clarke. 2021. Thematic Analysis: A Practical Guide . SAGE Publications Ltd, London
2021
-
[9]
Joy Buolamwini and Timnit Gebru. 2018. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Conf. on fairness, accountability and transparency. PMLR, 77–91
2018
-
[10]
Juliet Corbin and Anselm Strauss. 2008. Basics of Qualitative Research: Techniques and Procedures for Developing Grounded Theory (3rd ed.). SAGE Publications Ltd, Thousand Oaks, CA. Manuscript submitted to ACM 20 Gaba et al
2008
-
[11]
Kate Crawford and Vladan Joler. 2021. Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language Technologies. AI & Society 36, 1 (2021), 245–261
2021
-
[12]
Andreas Dengel, Rupert Gehrlein, David Fernes, Sebastian Görlich, Jonas Maurer, Hai Hoang Pham, Gabriel Großmann, and Niklas Dietrich genannt Eisermann. 2023. Qualitative research methods for large language models: Conducting semi-structured interviews with ChatGPT and BARD on...
2023
-
[13]
Berkeley J Dietvorst and Daniel M Bartels. 2022. The Value of Measuring Trust in AI - A Socio-Technical System Perspective. AI & Society 37, 2 (2022), 111–126
2022
-
[14]
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018. Measuring and Mitigating Unintended Bias in Text Classification. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society (New Orleans, LA, USA) (AIES ’18). Association for Computi...
2018
-
[15]
Ziwei Dong, Ameya Patil, Yuichi Shoda, Leilani Battle, and Emily Wall. 2025. Behavior Matters: An Alternative Perspective on Promoting Responsible Data Science. Proc. ACM Hum.-Comput. Interact. 9, 2, Article CSCW034 (May 2025), 23 pages. doi:10.1145/3710932
2025 doi
-
[16]
Fiona Draxler, Daniel Buschek, Mikke Tavast, Perttu Hämäläinen, Albrecht Schmidt, Juhi Kulshrestha, and Robin Welsch. 2023. Gender, age, and technology education influence the adoption and appropriation of LLMs. arXiv preprint arXiv:2310.06556 (2023)
2023 arXiv
-
[17]
Wen Duan, Lingyuan Li, Guo Freeman, and Nathan McNeese. 2025. A Scoping Review of Gender Stereotypes in Artificial Intelligence. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25) . Association for Computing Machinery, New York, NY, USA, A...
2025
-
[18]
Wen Duan, Nathan McNeese, Guo Freeman, and Lingyuan Li. 2024. Mitigating Gender Stereotypes Toward AI Agents Through an eXplainable AI (XAI) Approach. Proc. ACM Hum.-Comput. Interact. 8, CSCW2, Article 430 (Nov. 2024), 35 pages. doi:10.1145/3686969
2024 doi
-
[19]
Patrick Marcel Joseph Dubois, Mahya Maftouni, and Andrea Bunt. 2022. Towards More Gender-Inclusive Q&As: Investigating Perceptions of Additional Community Presence Information. Proc. ACM Hum.-Comput. Interact. 6, CSCW2, Article 466 (Nov. 2022), 23 pages. doi:10.1145/3555567
2022 doi
-
[20]
Madeleine Clare Elish and Elizabeth Anne Watkins. 2020. Repairing Innovation: A Study of Integrating AI in Clinical Care. Data & Society (2020). https://datasociety.net/library/repairing-innovation/
2020
-
[21]
Hall, Yuriy Brun, and Cindy Xiong Bearfield
Aimen Gaba, Zhanna Kaufman, Jason Cheung, Marie Shvakel, Kyle Wm. Hall, Yuriy Brun, and Cindy Xiong Bearfield. 2024. My Model is Unfair, Do People Even Care? Visual Design Affects Trust and Perceived Bias in Machine Learning. IEEE Transactions on Visualization and Computer Gra...
2024
-
[22]
Sourojit Ghosh and Aylin Caliskan. 2023. Chatgpt perpetuates gender bias in machine translation and ignores non-gendered pronouns: Findings across bengali and five other low-resource languages. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society . 901–912
2023
-
[23]
Susanne Göbel and Ralf Lämmel. 2024. Model-Based Trust Analysis of LLM Conversations. In Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems (Linz, Austria) (MODELS Companion ’24). Association for Computing Machinery, New...
2024
-
[24]
Haimson, Aloe DeGuia, Rana Saber, and Kat Brewster
Oliver L. Haimson, Aloe DeGuia, Rana Saber, and Kat Brewster. 2024. Extended Reality Trans Technologies: Bridging Digital and Physical Worlds to Support Transgender People. Proc. ACM Hum.-Comput. Interact. 8, CSCW2, Article 433 (Nov. 2024), 27 pages. doi:10.1145/3686972
2024 doi
-
[25]
Christina Harrington, Sheena Erete, and Anne Marie Piper. 2019. Deconstructing Community-Based Collaborative Design: Towards More Equitable Participatory Design Engagements. Proc. ACM Hum.-Comput. Interact. 3, CSCW, Article 216 (Nov. 2019), 25 pages. doi:10.1145/3359318
2019 doi
-
[26]
Kenneth Holstein and Jennifer Wortman Vaughan. 2023. Disclosure and Mitigation of Gender Bias in LLMs. Proceedings of the AAAI Conference on Artificial Intelligence 35, 1 (2023), 3451–3460
2023
-
[27]
Help Me Help the AI
Sunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong, and Andrés Monroy-Hernández. 2023. "Help Me Help the AI": Understanding How Explainability Can Support Human-AI Interaction. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems (H...
2023
-
[28]
Sunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong, and Andrés Monroy-Hernández. 2023. Humans, AI, and Context: Understanding End-Users’ Trust in a Real-World Computer Vision Application (FAccT ’23). Association for Computing Machinery, New York, NY, USA, 77...
2023
-
[29]
Svetlana Kiritchenko and Saif M Mohammad. 2018. Examining gender and race bias in two hundred sentiment analysis systems. arXiv preprint arXiv:1805.04508 (2018)
2018 arXiv
-
[30]
Matthieu Komorowski, Leo A Celi, Omar Badawi, Anthony C Gordon, and A Aldo Faisal. 2018. The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care. Nature Medicine 24, 11 (2018), 1716–1720
2018
-
[31]
Haein Kong, Yongsu Ahn, Sangyub Lee, and Yunho Maeng. 2024. Gender Bias in LLM-generated Interview Responses.arXiv preprint arXiv:2410.20739 (2024)
2024 arXiv
-
[32]
Hadas Kotek, Rikker Dockum, and David Sun. 2023. Gender bias and stereotypes in Large Language Models. In Proceedings of The ACM Collective Intelligence Conference (Delft, Netherlands) (CI ’23). Association for Computing Machinery, New York, NY, USA, 12–24. doi:10.1145/3582269.3615599
2023
-
[33]
Harsh Kumar, Ilya Musabirov, Mohi Reza, Jiakai Shi, Xinyuan Wang, Joseph Jay Williams, Anastasia Kuzminykh, and Michael Liut. 2024. Guiding Students in Using LLMs in Supported Learning Environments: Effects on Interaction Dynamics, Learner Performance, Confidence, and Trust. P...
2024 doi
-
[34]
Shachi H Kumar, Saurav Sahay, Sahisnu Mazumder, Eda Okur, Ramesh Manuvinakurike, Nicole Beckage, Hsuan Su, Hung-yi Lee, and Lama Nachman
-
[35]
Le Dantec and Sarah Fox
Christopher A. Le Dantec and Sarah Fox. 2015. Strangers at the Gate: Gaining Access, Building Rapport, and Co-Constructing Community-Based Research. In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work & Social Computing (Vancouver, BC, Canada) (CSC...
2015
-
[36]
Susan Leavy. 2018. Gender bias in artificial intelligence: the need for diversity and gender theory in machine learning. In Proceedings of the 1st International Workshop on Gender Equality in Software Engineering (Gothenburg, Sweden) (GE ’18). Association for Computing Machine...
2018
-
[37]
Susan Leavy. 2018. Uncovering gender bias in newspaper coverage of Irish politicians using machine learning. Digital Scholarship in the Humanities 34, 1 (06 2018), 48–63. doi:10.1093/llc/fqy005 arXiv:https://academic.oup.com/dsh/article-pdf/34/1/48/28078571/fqy005.pdf
2018 doi
-
[38]
Lee, Jacob M
Messi H.J. Lee, Jacob M. Montgomery, and Calvin K. Lai. 2024. Large Language Models Portray Socially Subordinate Groups as More Homogeneous, Consistent with a Bias Observed in Humans. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (Rio ...
2024
-
[39]
Yoonjoo Lee, Kihoon Son, Tae Soo Kim, Jisu Kim, John Joon Young Chung, Eytan Adar, and Juho Kim. 2024. One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations. In Proceedings of the 2024 ACM Conference on Fairness, Accountabilit...
2024
-
[40]
Vera Liao, Daniel Gruen, and Sarah Miller
Q. Vera Liao, Daniel Gruen, and Sarah Miller. 2020. Questioning the AI: Informing Design Practices for Explainable AI User Experiences. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Mach...
2020
-
[41]
Q Vera Liao and Jennifer Wortman Vaughan. 2023. Ai transparency in the age of llms: A human-centered research roadmap. arXiv preprint arXiv:2306.01941 (2023), 5368–5393
2023 arXiv
-
[42]
Lanna Lima, Vasco Furtado, Elizabeth Furtado, and Virgilio Almeida. 2019. Empirical Analysis of Bias in Voice-based Personal Assistants. In Companion Proceedings of The 2019 World Wide Web Conference (San Francisco, USA) (WWW ’19). Association for Computing Machinery, New York...
2019
-
[43]
Li Lucy and David Bamman. 2021. Gender and Representation Bias in GPT-3 Generated Stories. In Proceedings of the Third Workshop on Narrative Understanding, Nader Akoury, Faeze Brahman, Snigdha Chaturvedi, Elizabeth Clark, Mohit Iyyer, and Lara J. Martin (Eds.). Association for...
2021 doi
-
[44]
Malle and Daniel Ullman
Bertram F. Malle and Daniel Ullman. 2021. Chapter 1 - A multidimensional conception and measure of human-robot trust. In Trust in Human-Robot Interaction, Chang S. Nam and Joseph B. Lyons (Eds.). Academic Press, 3–25. doi:10.1016/B978-0-12-819472-0.00001-0
2021 doi
-
[45]
Malle and Daniel Ullman
Bertram F. Malle and Daniel Ullman. 2021. A multidimensional conception and measure of human-robot trust. Trust in Human-Robot Interaction (2021). https://api.semanticscholar.org/CorpusID:228891840
2021
-
[46]
Bertram F Malle and Daniel Ullman. 2023. Measuring human-robot trust with the mdmt (multi-dimensional measure of trust). arXiv preprint arXiv:2311.14887 (2023)
2023 arXiv
-
[47]
Abhishek Mandal, Susan Leavy, and Suzanne Little. 2023. Multimodal composite association score: Measuring gender bias in generative multimodal models. arXiv preprint arXiv:2304.13855 (2023)
2023 arXiv
-
[48]
Roger C Mayer, James H Davis, and F David Schoorman. 1995. An Integrative Model of Organizational Trust. Academy of Management Review 20, 3 (1995), 709–734
1995
-
[49]
Nora McDonald, Sarita Schoenebeck, and Andrea Forte. 2019. Reliability and Inter-rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice. Proc. ACM Hum.-Comput. Interact. 3, CSCW, Article 72 (Nov. 2019), 23 pages. doi:10.1145/3359174
2019 doi
-
[50]
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A survey on bias and fairness in machine learning. ACM Computing Surveys (CSUR) 54, 6 (2021), 1–35
2021
-
[51]
Debora Nozza, Federico Bianchi, Anne Lauscher, and Dirk Hovy. 2022. Measuring Harmful Sentence Completion in Language Models for LGBTQIA+ Individuals. In Proceedings of the Second Workshop on Language Technology for Equality, Diversity and Inclusion , Bharathi Raja Chakravarth...
2022 doi
-
[52]
Cathy O’Neil. 2016. Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy . Crown Publishing Group, USA
2016
-
[53]
OpenAI. 2024. ChatGPT: Generative Pre-trained Transformer. https://openai.com/chatgpt Accessed: 2024-12-09
2024
-
[54]
Ji Ho Park, Jamin Shin, and Pascale Fung. 2018. Reducing gender bias in abusive language detection. arXiv preprint arXiv:1808.07231 (2018)
2018 arXiv
-
[55]
Q.ai. 2022. How Intelligent Machines Are Reshaping Investing. Forbes (2022). https://www.forbes.com/sites/qai/2022/01/25/how-intelligent- machines-are-reshaping-investing/
2022
-
[56]
Pradeep Kumar Roy, Sarabjeet Singh Chowdhary, and Rocky Bhatia. 2020. A Machine Learning approach for automation of Resume Recommendation system. Procedia Computer Science 167 (2020), 2318–2327
2020
-
[57]
Paul, and Jed R
Morgan Klaus Scheuerman, Jacob M. Paul, and Jed R. Brubaker. 2019. How Computers See Gender: An Evaluation of Gender Classification in Commercial Facial Analysis Services. Proc. ACM Hum.-Comput. Interact. 3, CSCW, Article 144 (Nov. 2019), 33 pages. doi:10.1145/3359246
2019 doi
-
[58]
H Andrew Schwartz and Maarten Sap. 2023. I’m fully who I am: Towards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation. In Proceedings of the NAACL 2023 . ACL, 123–135
2023
-
[59]
The human body is a black box
Mark Sendak, Madeleine Clare Elish, Michael Gao, Joseph Futoma, William Ratliff, Marshall Nichols, Armando Bedoya, Suresh Balu, and Cara O’Brien. 2020. "The human body is a black box": supporting clinical decision-making with deep learning. In Proceedings of the 2020 Conferenc...
2020
-
[60]
Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang. 2023. In chatgpt we trust? measuring and characterizing the reliability of chatgpt. arXiv preprint arXiv:2304.08979 (2023)
2023 arXiv
-
[61]
Jaemarie Solyst, Ellia Yang, Shixian Xie, Amy Ogan, Jessica Hammer, and Motahhare Eslami. 2023. The Potential of Diverse Youth as Stakeholders in Identifying and Mitigating Algorithmic Bias for a Future of Fairer AI. Proc. ACM Hum.-Comput. Interact. 7, CSCW2, Article 364 (Oct....
2023 doi
-
[62]
Daniel Ullman and Bertram F. Malle. 2019. Measuring Gains and Losses in Human-Robot Trust: Evidence for Differentiable Components of Trust. In 2019 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI) . 618–619. doi:10.1109/HRI.2019.8673154
2019
-
[63]
United Nations Educational, Scientific and Cultural Organization (UNESCO). 2024. Generative AI: UNESCO study reveals alarming evidence of regressive gender stereotypes. https://www.unesco.org/en/articles/generative-ai-unesco-study-reveals-alarming-evidence-regressive-gender- s...
2024
-
[64]
Authors Unknown. 2025. Investigating the Impact of User Trust on the Adoption and Use of ChatGPT: Survey Analysis. AI Adoption and Trust Journal 8, 2 (2025), 200–220. doi:10.1234/aitrust.v8i2.98765
2025 doi
-
[65]
Wagman and Lisa Parks
Kelly B. Wagman and Lisa Parks. 2021. Beyond the Command: Feminist STS Research and Critical Issues for the Design of Social Machines. Proc. ACM Hum.-Comput. Interact. 5, CSCW1, Article 101 (April 2021), 20 pages. doi:10.1145/3449175
2021 doi
-
[66]
Kelly is a Warm Person, Joseph is a Role Model
Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023. “Kelly is a Warm Person, Joseph is a Role Model”: Gender Biases in LLM-Generated Reference Letters. InThe 2023 Conference on Empirical Methods in Natural Language Processing . https://openr...
2023
-
[67]
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. 2021. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359 (2021)
2021 arXiv
-
[68]
Herbsleb, Alexandra Holloway, and Scott Davidoff
David Gray Widder, Laura Dabbish, James D. Herbsleb, Alexandra Holloway, and Scott Davidoff. 2021. Trust in Collaborative Automation in High Stakes Software Engineering Work: A Case Study at NASA. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems ...
2021
-
[70]
Kyra Yee, Uthaipon Tantipongpipat, and Shubhanshu Mishra. 2021. Image Cropping on Twitter: Fairness Metrics, their Limitations, and the Importance of Representation, Design, and Agency. Proc. ACM Hum.-Comput. Interact. 5, CSCW2, Article 450 (Oct. 2021), 24 pages. doi:10.1145/3479594
2021 doi
-
[71]
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender bias in coreference resolution: Evaluation and debiasing methods. arXiv preprint arXiv:1804.06876 (2018)
2018 arXiv
-
[72]
The teachers are confused as well
Kyrie Zhixuan Zhou, Zachary Kilhoffer, Madelyn Rose Sanfilippo, Ted Underwood, Ece Gumusel, Mengyi Wei, Abhinav Choudhry, and Jinjun Xiong. 2024. " The teachers are confused as well": A Multiple-Stakeholder Ethics Discussion on Large Language Models in Computing Education. arX...
2024 arXiv
-
[2024]
arXiv preprint arXiv:2408.03907 (2024)
Decoding biases: Automated methods and llm judges for gender bias detection in language models. arXiv preprint arXiv:2408.03907 (2024). Manuscript submitted to ACM Bias, Accuracy, and Trust: Gender-Diverse Perspectives on Large Language Models 21
2024 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.