Pith. sign in

REVIEW 3 major objections 5 minor 73 references

Bias, Accuracy, and Trust: Gender-Diverse Perspectives on Large Language Models

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper argues that users' gender identity changes how they perceive ChatGPT's bias, accuracy, and trustworthiness, with non-binary participants especially likely to find the model's gender portrayals condescending.

desk verdict A rare, well-reported qualitative study of gender-diverse LLM users; the qualitative core is strong, but the quantitative trust analysis is underreported and likely confounded. read the letter →

arxiv 2506.21898 v2 pith:TRK2AYCJ submitted 2025-06-27 cs.HC

classification cs.HC
keywords largelanguagemodelsalgorithmicharmsgenderrepresentationtrustinAIfairresponsibleChatGPTnon-binary
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether users of different gender identities perceive bias, accuracy, and trust in ChatGPT differently, and whether the gender wording of a prompt changes what the model produces. Through 25 interviews with non-binary/transgender, women, and men participants, the authors show that gendered prompts produce more identity-specific, stereotype-aligned stories than neutral prompts, and that non-binary participants find the model's non-binary portrayals condescending, cis-centric, and reductive. Perceived accuracy was similar across groups, while trust split along gender lines: men trusted the model most, and non-binary participants trusted its performance more than its morality. The study's point is that user perception of LLM bias and trust is group-dependent in ways that benchmark-only evaluations miss, and that inclusive design should draw on gender-diverse perspectives.

What carries the argument

The machinery is a 16-prompt case study: four everyday scenarios (buying a car, navigating college, experiences of a person, applying for jobs) run with "person" and with "man", "woman", and "non-binary person" as the subject, plus a semi-structured interview protocol with pre- and post-interview trust surveys. The prompt set exposes how ChatGPT's character construction changes with gender wording, and the interviews expose how those constructions are read. The trust survey separates morality-based from performance-based trust, which is what lets the study distinguish the trust profile of non-binary participants from that of women and men.

What would settle it

A pre-registered replication with a larger sample, matched LLM/AI backgrounds, and blinded analysts could falsify the perception claim: if non-binary participants no longer rate the non-binary prompts as more condescending than other groups do, the reported effect is an artifact of recruitment or background rather than gender identity.

Watch

Extended reading notes

Core claim

The central claim is that ChatGPT's responses to gender-specified prompts are not neutral: the model assigns names, pronouns, adjectives, goals, and challenges that follow societal gender norms, with non-binary characters reduced to identity struggles and women to emotional or traditional roles. Interview data show that this pattern is perceived differently by identity, with non-binary/transgender participants reporting condescending and stereotypical portrayals, women criticizing outdated tropes, and men noticing a lack of diversity but fewer concerns. Trust ratings after the interviews differed by gender, with men reporting higher trust overall and non-binary participants reporting higher performance-based trust but morality-based trust similar to women's. The authors conclude that the same model output can be read as acceptable, biased, or harmful depending on who reads it, so evaluations of LLM fairness should include gender-diverse evaluators.

Load-bearing premise

The comparisons by gender assume that 25 participants recruited through university email lists, Slack, X, and one online panel are representative of gender-diverse user populations, so differences attributed to gender could instead reflect prior experience with LLMs or with bias.

Editorial extensions

If this is right

  • Gendered prompts change what ChatGPT produces: non-binary characters get they/them pronouns, neutral-to-masculine names, and storylines centered on identity struggles, while men get career-focused, assertive narratives and women get emotional or independence-focused ones.
  • Non-binary/transgender participants experience ChatGPT's non-binary portrayals as condescending, cis-centric, and reductive, while women and men find the same outputs less problematic.
  • Perceived accuracy does not differ across gender groups; participants in all groups notice errors mainly in technical and creative tasks.
  • Trust is gender-dependent: men trust ChatGPT more overall, non-binary participants trust its performance more than its morality, and trust declines after reflection in participants with medium bias knowledge.
  • Users want LLMs to diversify training data with real lived experiences, respond to all genders with equal depth, ask clarifying questions, and be transparent about sources and limitations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension: the 16-prompt suite could be turned into an automated gender-sensitivity audit by scoring outputs for identity-reduction markers, such as how much of a narrative centers on struggles rather than the scenario, and comparing those scores with the human ratings collected here.
  • If the trust drop in medium-bias-knowledge participants is real, then simply asking people to reflect on bias during an interview changes their trust; a controlled experiment with a reflection-only condition could separate this effect from experimenter demand, which the authors themselves say they cannot do.
  • The finding that "person" prompts still default to gendered stories suggests that neutral prompt wording does not produce neutral output; developers aiming for gender-neutral responses may need explicit constraints rather than neutral phrasing alone.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports a mixed-methods study of how gender-diverse users perceive bias, accuracy, and trust in ChatGPT. The authors conducted 25 semi-structured interviews (9 non-binary/transgender, 8 women, 8 men), combined with a content analysis of ChatGPT responses to gendered and neutral prompts, and collected quantitative trust ratings before and after the interviews. The central qualitative claims are that gendered prompts elicit more identity-specific and stereotypical responses; that non-binary/transgender participants experience these responses as condescending and reductive; that perceived accuracy is fairly consistent across groups; and that trust varies by gender, with men reporting the highest trust. The paper also presents participant suggestions and design implications for more inclusive LLMs.

Significance. If the results hold, the paper makes a useful contribution to CSCW/HCI by showing that user perceptions of LLM bias and trust are group-dependent in ways that benchmark-only evaluations cannot capture. The study has notable strengths: it centers the voices of non-binary and transgender participants, provides direct participant quotes as evidence, describes its qualitative coding process transparently, avoids causal claims about the time effect, and makes its supplementary materials available. The qualitative findings on perceived bias are well supported by the reported quotes and thematic structure. However, the quantitative trust result—which appears in the abstract and in RQ4—is currently underreported and potentially confounded, so the headline gender-trust claim needs additional analysis before it can be relied upon.

major comments (3)
  1. [§4.8.2] The quantitative trust-by-gender result is statistically underreported and potentially confounded. The authors fit three separate general linear models—one for gender, one for AI/LLM expertise, and one for bias knowledge—and report only chi-square statistics and p-values. They do not report coefficients, degrees of freedom, effect sizes, confidence intervals, or post-hoc group comparisons. More importantly, they never fit a model containing gender together with expertise and bias knowledge, even though the same section shows that expertise increases trust and bias knowledge decreases trust. If the male subsample happens to have more AI/LLM expertise or less bias knowledge, the reported gender differences in §4.8.2 could be an artifact of group composition rather than gender identity. Table 1 does not cross-tabulate gender with expertise or bias knowledge, so the claim in §4.1 that women and men were recruited with 'similar backgrounds' cannot be verified. Section 6 lists limitations but does not flag this specific confound. Because the abstract and RQ4 state that 'trustworthiness varied by gender,' this issue is load-bearing and must be addressed with a combined model and fuller reporting.
  2. [§3.1] The content analysis of ChatGPT responses in RQ1 lacks essential methodological detail. The authors report counts such as 'resilience (21)' for non-binary characters, 'Emma (6)' for women, and various counts for goals, challenges, and symbols, but they do not describe how these categories were defined, whether codes were developed independently or through consensus, whether counts were based on unique responses, or how the 10 repetitions per prompt were handled. Without a codebook or any reliability assessment, the quantitative flavor of these counts is not reproducible. This matters because the claim that 'gendered prompts elicit richer, more identity-specific responses' is one of the paper's central contributions. The authors should either present this as purely qualitative illustration or provide the coding scheme and, if appropriate, a reliability analysis.
  3. [§4.8.2] The interpretation of the time × bias-knowledge interaction is speculative. The authors hypothesize that participants with medium bias knowledge 'had enough working knowledge to recognize the concept of bias, but lacked the depth of understanding necessary to contextualize how it manifests in LLMs.' While they appropriately avoid causal claims about the time effect, they do not consider plausible alternative explanations such as regression to the mean, differential item difficulty between pre- and post-surveys, or experimenter demand. The model also does not include gender or expertise interactions with time, so the claim that the interview had a 'pronounced effect on this group' is not directly tested. This does not invalidate the qualitative findings, but it should be presented as one possible interpretation rather than a conclusion.
minor comments (5)
  1. [Throughout] The manuscript contains numerous typos and grammatical errors that should be corrected, including 'prevalant' (§2.1), 'lagnuage' (§2.1), 'includive' (§2.1), 'comprimising' (§2.2), 'percieved' (§2.2), 'Rersonal' (§3.1.1), 'modic logic' (§5.2.2), 'outout' (§5.2.2), 'Explanaitions' (§5.2.2), and 'bakcground' in the Figure 3 caption.
  2. [References] Reference [64] lists 'Authors Unknown' as the author and appears to be a placeholder; this citation must be completed or removed before publication.
  3. [§4.8.2] The description of the gender model says the model 'included trust type, participants' gender, and the time,' but the subsequent discussion of an interaction between gender and trust type implies that interactions were modeled. Please state explicitly which main effects and interactions were included in each model.
  4. [Table 1] The column header 'Knowledge of bias in LLMs & LLM/AI background)' contains an unmatched parenthesis and should be cleaned up.
  5. [§4.1] The recruitment description would benefit from clarifying how the three non-binary participants recruited via userinterviews.com differed from the university-recruited participants, since combining recruitment channels can introduce unmeasured background differences.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's claims are empirical interpretations of interview and ChatGPT-response data, not quantities defined by fitted parameters.

full rationale

The paper contains no derivation chain in which an output quantity is defined in terms of a fitted input. RQ1 is a descriptive analysis of ChatGPT outputs (Section 3.1): gendered prompts are observed to produce identity-specific narrative content, with no parameter fitted from those outputs to 'predict' the same outputs. RQ2–RQ5 are grounded-theory interpretations of 25 participant interviews (Sections 4.4–5.1); perceived bias, accuracy, trust, and suggestions are reported as participant statements and codes, not as quantities derived from the study's own assumptions. The quantitative trust analysis (Section 4.8.2) uses an external validated trust instrument (Malle and Ullman MDMT, refs [44–46, 62]) and separate general linear models on self-reported trust ratings; although the paper does not report a single model with gender, expertise, and bias knowledge together, that is a statistical completeness or confounding concern, not a circular one, because the gender effect is not constructed from the covariates. The only author self-citation is [21], used once in the introduction ('These biases not only undermine user trust in ML models [21]') as background motivation; it does not carry the paper's conclusions. No uniqueness theorem, no ansatz smuggled in via citation, and no renamed empirical pattern is load-bearing. The central findings are self-contained empirical results grounded in the collected qualitative and quantitative data.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No mathematical parameters are fitted; the assumptions concern participant categorization, prompt representativeness, and qualitative validity. No new theoretical entities are introduced beyond trust categories borrowed from prior research.

assumptions (4)
  • domain assumption Participants' self-identified gender categories (men, women, non-binary and transgender) are stable, mutually exclusive, and a valid analytic grouping for comparing perceptions.
    Invoked throughout recruitment (Section 4.1) and all group comparisons (Sections 4.6 to 4.8); no independent verification of identity categories.
  • domain assumption The four curated prompts and their gendered variants are neutral enough to elicit representative LLM behavior, and ChatGPT 3.5 responses generated in February 2024 are treated as representative of LLM outputs.
    Section 3 states prompts were crafted through extensive probing to avoid leading content, but the selection is not externally validated against other prompts or models.
  • domain assumption Consensus-based thematic coding without inter-rater reliability yields valid themes.
    Section 4.4 explicitly declines inter-rater reliability based on an interpretive framework; the replicability of themes depends on this choice.
  • domain assumption The trust survey items from the multidimensional measure of trust measure morality-based and performance-based trust as intended in this population.
    The trust scores in Section 4.8.2 rely on the validity of the measurement instrument for gender-diverse participants, which is not separately established.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bias, Accuracy, and Trust: Gender-Diverse Perspectives on Large Language Models." pith.science (2026). https://pith.science/paper/TRK2AYCJ

@misc{pith2026250621898,
  author       = {Pith},
  title        = {Pith review of: Bias, Accuracy, and Trust: Gender-Diverse Perspectives on Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TRK2AYCJ}},
  note         = {Machine review of arXiv:2506.21898}
}
read the original abstract

Large language models (LLMs) are becoming increasingly ubiquitous in our daily lives, but numerous concerns about bias in LLMs exist. This study examines how gender-diverse populations perceive bias, accuracy, and trustworthiness in LLMs, specifically ChatGPT. Through 25 in-depth interviews with non-binary/transgender, male, and female participants, we investigate how gendered and neutral prompts influence model responses and how users evaluate these responses. Our findings reveal that gendered prompts elicit more identity-specific responses, with non-binary participants particularly susceptible to condescending and stereotypical portrayals. Perceived accuracy was consistent across gender groups, with errors most noted in technical topics and creative tasks. Trustworthiness varied by gender, with men showing higher trust, especially in performance, and non-binary participants demonstrating higher performance-based trust. Additionally, participants suggested improving the LLMs by diversifying training data, ensuring equal depth in gendered responses, and incorporating clarifying questions. This research contributes to the CSCW/HCI field by highlighting the need for gender-diverse perspectives in LLM development in particular and AI in general, to foster more inclusive and trustworthy systems.

Figures

Figures reproduced from arXiv: 2506.21898 by the authors.

Figure 1
Figure 1. Research questions to understand people’s perception of bias in a real-world application of large language models. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. An overview of our study protocol was making assumptions about gender, if the statement could be perceived as offensive or harmful, why the participants thought ChatGPT responded the way it did, if they found the response surprising, and what follow-up questions they would ask ChatGPT if they had the option. Each interview ended with a summary discussion of how the participant felt about ChatGPT or LLMs in general a… view at source ↗
Figure 3
Figure 3. Before and after the interview trust ratings by self-reported gender, by the participant’s expertise/bakcground in LLM and AI [PITH_FULL_IMAGE:figures/full_fig_p015_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

73 extracted references · 36 canonical work pages

  1. [69]

    J Williams and R Wilson. 2023. I wouldn’t say offensive but...: Disability-Centered Perspectives on Large Language Models. Journal of Accessible AI 2, 3 (2023), 45–60

  2. [1]

    Supplementary materials for LLM Bias

    2025. Supplementary materials for LLM Bias. https://osf.io/ywmtg/?view_only=7a38b88af78e45ad85a33b70f88f717f

  3. [2]

    Abubakar Abid, Maheen Farooqi, and James Zou. 2021. Persistent Anti-Muslim Bias in Large Language Models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (Virtual Event, USA) (AIES ’21). Association for Computing Machinery, New York, NY, USA. doi:10.1145/3461702.3462624

  4. [3]

    Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner. 2016. Machine Bias. ProPublica May 23 (2016). https://www.propublica.org/article/ machine-bias-risk-assessments-in-criminal-sentencing

  5. [4]

    Julian Ashwin, Aditya Chhabra, and Vijayendra Rao. 2023. Using large language models for qualitative analysis can introduce serious bias. arXiv preprint arXiv:2309.17147 (2023)

  6. [5]

    Marion Bartl and Susan Leavy. 2024. From’Showgirls’ to’Performers’: Fine-tuning with Gender-inclusive Language for Bias Reduction in LLMs. arXiv preprint arXiv:2407.04434 (2024)

  7. [6]

    Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell

    Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT) . ACM, 610–623. doi:10.1145/ 3442188.3445922

  8. [7]

    Reuben Binns and Michael Veale. 2022. Adhering, Steering, and Queering: Treatment of Gender in Natural Language Generation. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* 2022) . ACM, 142–155

Show all 73 references
  1. [8]

    Virginia Braun and Victoria Clarke. 2021. Thematic Analysis: A Practical Guide . SAGE Publications Ltd, London

  2. [9]

    Joy Buolamwini and Timnit Gebru. 2018. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Conf. on fairness, accountability and transparency. PMLR, 77–91

  3. [10]

    Juliet Corbin and Anselm Strauss. 2008. Basics of Qualitative Research: Techniques and Procedures for Developing Grounded Theory (3rd ed.). SAGE Publications Ltd, Thousand Oaks, CA. Manuscript submitted to ACM 20 Gaba et al

  4. [11]

    Kate Crawford and Vladan Joler. 2021. Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language Technologies. AI & Society 36, 1 (2021), 245–261

  5. [12]

    Andreas Dengel, Rupert Gehrlein, David Fernes, Sebastian Görlich, Jonas Maurer, Hai Hoang Pham, Gabriel Großmann, and Niklas Dietrich genannt Eisermann. 2023. Qualitative research methods for large language models: Conducting semi-structured interviews with ChatGPT and BARD on...

  6. [13]

    Berkeley J Dietvorst and Daniel M Bartels. 2022. The Value of Measuring Trust in AI - A Socio-Technical System Perspective. AI & Society 37, 2 (2022), 111–126

  7. [14]

    Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018. Measuring and Mitigating Unintended Bias in Text Classification. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society (New Orleans, LA, USA) (AIES ’18). Association for Computi...

  8. [15]

    Ziwei Dong, Ameya Patil, Yuichi Shoda, Leilani Battle, and Emily Wall. 2025. Behavior Matters: An Alternative Perspective on Promoting Responsible Data Science. Proc. ACM Hum.-Comput. Interact. 9, 2, Article CSCW034 (May 2025), 23 pages. doi:10.1145/3710932

  9. [16]

    Fiona Draxler, Daniel Buschek, Mikke Tavast, Perttu Hämäläinen, Albrecht Schmidt, Juhi Kulshrestha, and Robin Welsch. 2023. Gender, age, and technology education influence the adoption and appropriation of LLMs. arXiv preprint arXiv:2310.06556 (2023)

  10. [17]

    Wen Duan, Lingyuan Li, Guo Freeman, and Nathan McNeese. 2025. A Scoping Review of Gender Stereotypes in Artificial Intelligence. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25) . Association for Computing Machinery, New York, NY, USA, A...

  11. [18]

    Wen Duan, Nathan McNeese, Guo Freeman, and Lingyuan Li. 2024. Mitigating Gender Stereotypes Toward AI Agents Through an eXplainable AI (XAI) Approach. Proc. ACM Hum.-Comput. Interact. 8, CSCW2, Article 430 (Nov. 2024), 35 pages. doi:10.1145/3686969

  12. [19]

    Patrick Marcel Joseph Dubois, Mahya Maftouni, and Andrea Bunt. 2022. Towards More Gender-Inclusive Q&As: Investigating Perceptions of Additional Community Presence Information. Proc. ACM Hum.-Comput. Interact. 6, CSCW2, Article 466 (Nov. 2022), 23 pages. doi:10.1145/3555567

  13. [20]

    Madeleine Clare Elish and Elizabeth Anne Watkins. 2020. Repairing Innovation: A Study of Integrating AI in Clinical Care. Data & Society (2020). https://datasociety.net/library/repairing-innovation/

  14. [21]

    Hall, Yuriy Brun, and Cindy Xiong Bearfield

    Aimen Gaba, Zhanna Kaufman, Jason Cheung, Marie Shvakel, Kyle Wm. Hall, Yuriy Brun, and Cindy Xiong Bearfield. 2024. My Model is Unfair, Do People Even Care? Visual Design Affects Trust and Perceived Bias in Machine Learning. IEEE Transactions on Visualization and Computer Gra...

  15. [22]

    Sourojit Ghosh and Aylin Caliskan. 2023. Chatgpt perpetuates gender bias in machine translation and ignores non-gendered pronouns: Findings across bengali and five other low-resource languages. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society . 901–912

  16. [23]

    Susanne Göbel and Ralf Lämmel. 2024. Model-Based Trust Analysis of LLM Conversations. In Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems (Linz, Austria) (MODELS Companion ’24). Association for Computing Machinery, New...

  17. [24]

    Haimson, Aloe DeGuia, Rana Saber, and Kat Brewster

    Oliver L. Haimson, Aloe DeGuia, Rana Saber, and Kat Brewster. 2024. Extended Reality Trans Technologies: Bridging Digital and Physical Worlds to Support Transgender People. Proc. ACM Hum.-Comput. Interact. 8, CSCW2, Article 433 (Nov. 2024), 27 pages. doi:10.1145/3686972

  18. [25]

    Christina Harrington, Sheena Erete, and Anne Marie Piper. 2019. Deconstructing Community-Based Collaborative Design: Towards More Equitable Participatory Design Engagements. Proc. ACM Hum.-Comput. Interact. 3, CSCW, Article 216 (Nov. 2019), 25 pages. doi:10.1145/3359318

  19. [26]

    Kenneth Holstein and Jennifer Wortman Vaughan. 2023. Disclosure and Mitigation of Gender Bias in LLMs. Proceedings of the AAAI Conference on Artificial Intelligence 35, 1 (2023), 3451–3460

  20. [27]

    Help Me Help the AI

    Sunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong, and Andrés Monroy-Hernández. 2023. "Help Me Help the AI": Understanding How Explainability Can Support Human-AI Interaction. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems (H...

  21. [28]

    Sunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong, and Andrés Monroy-Hernández. 2023. Humans, AI, and Context: Understanding End-Users’ Trust in a Real-World Computer Vision Application (FAccT ’23). Association for Computing Machinery, New York, NY, USA, 77...

  22. [29]

    Svetlana Kiritchenko and Saif M Mohammad. 2018. Examining gender and race bias in two hundred sentiment analysis systems. arXiv preprint arXiv:1805.04508 (2018)

  23. [30]

    Matthieu Komorowski, Leo A Celi, Omar Badawi, Anthony C Gordon, and A Aldo Faisal. 2018. The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care. Nature Medicine 24, 11 (2018), 1716–1720

  24. [31]

    Haein Kong, Yongsu Ahn, Sangyub Lee, and Yunho Maeng. 2024. Gender Bias in LLM-generated Interview Responses.arXiv preprint arXiv:2410.20739 (2024)

  25. [32]

    Hadas Kotek, Rikker Dockum, and David Sun. 2023. Gender bias and stereotypes in Large Language Models. In Proceedings of The ACM Collective Intelligence Conference (Delft, Netherlands) (CI ’23). Association for Computing Machinery, New York, NY, USA, 12–24. doi:10.1145/3582269.3615599

  26. [33]

    Harsh Kumar, Ilya Musabirov, Mohi Reza, Jiakai Shi, Xinyuan Wang, Joseph Jay Williams, Anastasia Kuzminykh, and Michael Liut. 2024. Guiding Students in Using LLMs in Supported Learning Environments: Effects on Interaction Dynamics, Learner Performance, Confidence, and Trust. P...

  27. [34]

    Shachi H Kumar, Saurav Sahay, Sahisnu Mazumder, Eda Okur, Ramesh Manuvinakurike, Nicole Beckage, Hsuan Su, Hung-yi Lee, and Lama Nachman

  28. [35]

    Le Dantec and Sarah Fox

    Christopher A. Le Dantec and Sarah Fox. 2015. Strangers at the Gate: Gaining Access, Building Rapport, and Co-Constructing Community-Based Research. In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work & Social Computing (Vancouver, BC, Canada) (CSC...

  29. [36]

    Susan Leavy. 2018. Gender bias in artificial intelligence: the need for diversity and gender theory in machine learning. In Proceedings of the 1st International Workshop on Gender Equality in Software Engineering (Gothenburg, Sweden) (GE ’18). Association for Computing Machine...

  30. [37]

    Susan Leavy. 2018. Uncovering gender bias in newspaper coverage of Irish politicians using machine learning. Digital Scholarship in the Humanities 34, 1 (06 2018), 48–63. doi:10.1093/llc/fqy005 arXiv:https://academic.oup.com/dsh/article-pdf/34/1/48/28078571/fqy005.pdf

  31. [38]

    Lee, Jacob M

    Messi H.J. Lee, Jacob M. Montgomery, and Calvin K. Lai. 2024. Large Language Models Portray Socially Subordinate Groups as More Homogeneous, Consistent with a Bias Observed in Humans. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (Rio ...

  32. [39]

    Yoonjoo Lee, Kihoon Son, Tae Soo Kim, Jisu Kim, John Joon Young Chung, Eytan Adar, and Juho Kim. 2024. One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations. In Proceedings of the 2024 ACM Conference on Fairness, Accountabilit...

  33. [40]

    Vera Liao, Daniel Gruen, and Sarah Miller

    Q. Vera Liao, Daniel Gruen, and Sarah Miller. 2020. Questioning the AI: Informing Design Practices for Explainable AI User Experiences. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Mach...

  34. [41]

    Q Vera Liao and Jennifer Wortman Vaughan. 2023. Ai transparency in the age of llms: A human-centered research roadmap. arXiv preprint arXiv:2306.01941 (2023), 5368–5393

  35. [42]

    Lanna Lima, Vasco Furtado, Elizabeth Furtado, and Virgilio Almeida. 2019. Empirical Analysis of Bias in Voice-based Personal Assistants. In Companion Proceedings of The 2019 World Wide Web Conference (San Francisco, USA) (WWW ’19). Association for Computing Machinery, New York...

  36. [43]

    Li Lucy and David Bamman. 2021. Gender and Representation Bias in GPT-3 Generated Stories. In Proceedings of the Third Workshop on Narrative Understanding, Nader Akoury, Faeze Brahman, Snigdha Chaturvedi, Elizabeth Clark, Mohit Iyyer, and Lara J. Martin (Eds.). Association for...

  37. [44]

    Malle and Daniel Ullman

    Bertram F. Malle and Daniel Ullman. 2021. Chapter 1 - A multidimensional conception and measure of human-robot trust. In Trust in Human-Robot Interaction, Chang S. Nam and Joseph B. Lyons (Eds.). Academic Press, 3–25. doi:10.1016/B978-0-12-819472-0.00001-0

  38. [45]

    Malle and Daniel Ullman

    Bertram F. Malle and Daniel Ullman. 2021. A multidimensional conception and measure of human-robot trust. Trust in Human-Robot Interaction (2021). https://api.semanticscholar.org/CorpusID:228891840

  39. [46]

    Bertram F Malle and Daniel Ullman. 2023. Measuring human-robot trust with the mdmt (multi-dimensional measure of trust). arXiv preprint arXiv:2311.14887 (2023)

  40. [47]

    Abhishek Mandal, Susan Leavy, and Suzanne Little. 2023. Multimodal composite association score: Measuring gender bias in generative multimodal models. arXiv preprint arXiv:2304.13855 (2023)

  41. [48]

    Roger C Mayer, James H Davis, and F David Schoorman. 1995. An Integrative Model of Organizational Trust. Academy of Management Review 20, 3 (1995), 709–734

  42. [49]

    Nora McDonald, Sarita Schoenebeck, and Andrea Forte. 2019. Reliability and Inter-rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice. Proc. ACM Hum.-Comput. Interact. 3, CSCW, Article 72 (Nov. 2019), 23 pages. doi:10.1145/3359174

  43. [50]

    Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A survey on bias and fairness in machine learning. ACM Computing Surveys (CSUR) 54, 6 (2021), 1–35

  44. [51]

    Debora Nozza, Federico Bianchi, Anne Lauscher, and Dirk Hovy. 2022. Measuring Harmful Sentence Completion in Language Models for LGBTQIA+ Individuals. In Proceedings of the Second Workshop on Language Technology for Equality, Diversity and Inclusion , Bharathi Raja Chakravarth...

  45. [52]

    Cathy O’Neil. 2016. Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy . Crown Publishing Group, USA

  46. [53]

    OpenAI. 2024. ChatGPT: Generative Pre-trained Transformer. https://openai.com/chatgpt Accessed: 2024-12-09

  47. [54]

    Ji Ho Park, Jamin Shin, and Pascale Fung. 2018. Reducing gender bias in abusive language detection. arXiv preprint arXiv:1808.07231 (2018)

  48. [55]

    Q.ai. 2022. How Intelligent Machines Are Reshaping Investing. Forbes (2022). https://www.forbes.com/sites/qai/2022/01/25/how-intelligent- machines-are-reshaping-investing/

  49. [56]

    Pradeep Kumar Roy, Sarabjeet Singh Chowdhary, and Rocky Bhatia. 2020. A Machine Learning approach for automation of Resume Recommendation system. Procedia Computer Science 167 (2020), 2318–2327

  50. [57]

    Paul, and Jed R

    Morgan Klaus Scheuerman, Jacob M. Paul, and Jed R. Brubaker. 2019. How Computers See Gender: An Evaluation of Gender Classification in Commercial Facial Analysis Services. Proc. ACM Hum.-Comput. Interact. 3, CSCW, Article 144 (Nov. 2019), 33 pages. doi:10.1145/3359246

  51. [58]

    H Andrew Schwartz and Maarten Sap. 2023. I’m fully who I am: Towards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation. In Proceedings of the NAACL 2023 . ACL, 123–135

  52. [59]

    The human body is a black box

    Mark Sendak, Madeleine Clare Elish, Michael Gao, Joseph Futoma, William Ratliff, Marshall Nichols, Armando Bedoya, Suresh Balu, and Cara O’Brien. 2020. "The human body is a black box": supporting clinical decision-making with deep learning. In Proceedings of the 2020 Conferenc...

  53. [60]

    Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang. 2023. In chatgpt we trust? measuring and characterizing the reliability of chatgpt. arXiv preprint arXiv:2304.08979 (2023)

  54. [61]

    Jaemarie Solyst, Ellia Yang, Shixian Xie, Amy Ogan, Jessica Hammer, and Motahhare Eslami. 2023. The Potential of Diverse Youth as Stakeholders in Identifying and Mitigating Algorithmic Bias for a Future of Fairer AI. Proc. ACM Hum.-Comput. Interact. 7, CSCW2, Article 364 (Oct....

  55. [62]

    Daniel Ullman and Bertram F. Malle. 2019. Measuring Gains and Losses in Human-Robot Trust: Evidence for Differentiable Components of Trust. In 2019 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI) . 618–619. doi:10.1109/HRI.2019.8673154

  56. [63]

    United Nations Educational, Scientific and Cultural Organization (UNESCO). 2024. Generative AI: UNESCO study reveals alarming evidence of regressive gender stereotypes. https://www.unesco.org/en/articles/generative-ai-unesco-study-reveals-alarming-evidence-regressive-gender- s...

  57. [64]

    Authors Unknown. 2025. Investigating the Impact of User Trust on the Adoption and Use of ChatGPT: Survey Analysis. AI Adoption and Trust Journal 8, 2 (2025), 200–220. doi:10.1234/aitrust.v8i2.98765

  58. [65]

    Wagman and Lisa Parks

    Kelly B. Wagman and Lisa Parks. 2021. Beyond the Command: Feminist STS Research and Critical Issues for the Design of Social Machines. Proc. ACM Hum.-Comput. Interact. 5, CSCW1, Article 101 (April 2021), 20 pages. doi:10.1145/3449175

  59. [66]

    Kelly is a Warm Person, Joseph is a Role Model

    Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023. “Kelly is a Warm Person, Joseph is a Role Model”: Gender Biases in LLM-Generated Reference Letters. InThe 2023 Conference on Empirical Methods in Natural Language Processing . https://openr...

  60. [67]

    Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. 2021. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359 (2021)

  61. [68]

    Herbsleb, Alexandra Holloway, and Scott Davidoff

    David Gray Widder, Laura Dabbish, James D. Herbsleb, Alexandra Holloway, and Scott Davidoff. 2021. Trust in Collaborative Automation in High Stakes Software Engineering Work: A Case Study at NASA. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems ...

  62. [70]

    Kyra Yee, Uthaipon Tantipongpipat, and Shubhanshu Mishra. 2021. Image Cropping on Twitter: Fairness Metrics, their Limitations, and the Importance of Representation, Design, and Agency. Proc. ACM Hum.-Comput. Interact. 5, CSCW2, Article 450 (Oct. 2021), 24 pages. doi:10.1145/3479594

  63. [71]

    Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender bias in coreference resolution: Evaluation and debiasing methods. arXiv preprint arXiv:1804.06876 (2018)

  64. [72]

    The teachers are confused as well

    Kyrie Zhixuan Zhou, Zachary Kilhoffer, Madelyn Rose Sanfilippo, Ted Underwood, Ece Gumusel, Mengyi Wei, Abhinav Choudhry, and Jinjun Xiong. 2024. " The teachers are confused as well": A Multiple-Stakeholder Ethics Discussion on Large Language Models in Computing Education. arX...

  65. [2024]

    arXiv preprint arXiv:2408.03907 (2024)

    Decoding biases: Automated methods and llm judges for gender bias detection in language models. arXiv preprint arXiv:2408.03907 (2024). Manuscript submitted to ACM Bias, Accuracy, and Trust: Gender-Diverse Perspectives on Large Language Models 21

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.