REVIEW 3 major objections 6 minor 139 references
The Generative AI Ethics Playbook
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper presents a practitioner-facing playbook that claims to help AI/ML teams diagnose and mitigate ethical harms at each stage of the generative AI lifecycle, with checklists, case studies, and mitigation strategies.
desk verdict A well-organized, honest synthesis of AI ethics guidance that is over-promising in its abstract: the mitigation strategies are not validated, and the core 'reduce negative impact' claim needs softening. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing machinery is the six-stage AI lifecycle, each stage paired with a transparency and documentation checklist built to make decisions and trade-offs visible to reviewers and stakeholders. Inside each stage, 'topics of interest' pose common considerations, illustrate harms through case studies, and list mitigation strategies. The shared vocabulary is the five-category harm taxonomy—representational, allocative, quality-of-service, interpersonal, and societal harms—which lets a decision made at one stage (for example, filtering a dataset) be traced to downstream social effects (for example, a model that underperforms for dialect speakers). The checklists function as audit artifacts meant to turn ethics from a one-time review into an ongoing documentation practice.
What would settle it
A controlled field study would settle it: randomize comparable AI product teams to use the playbook or not, then measure documented harms such as privacy leaks, biased or toxic outputs, and stakeholder complaints across the lifecycle. If playbook-using teams show no measurable improvement in harm rates or documentation quality over control teams, the central claim is falsified. A simpler test would survey practitioners after adoption: if most report they lacked the authority or resources to act on the checklists, the agency premise fails.
Extended reading notes
Core claim
The paper's central claim is that ethical harm in AI is not an afterthought but a property of decisions made at every stage of a system's life, and that a structured playbook can help practitioners see and reduce those harms. The authors synthesize current research and practice into stage-by-stage guidance for text, image, and multimodal generative models: problem formulation, dataset, model design, model training, model evaluation, and model use and monitoring. Each stage carries a transparency and documentation checklist, common ethical considerations, case studies of recorded harms, and mitigation strategies, all keyed to a five-part taxonomy of harms: representational, allocative, quality-of-service, interpersonal, and societal. The authors describe the playbook as a practically useful survey assembled from a review of over 100 resources and collaborative expert input, rather than as a complete or final treatment of the field.
Load-bearing premise
The load-bearing premise is that AI practitioners have considerable agency over the social impact of their work and that following the playbook's checklists and mitigation strategies actually reduces harm in real projects; if organizational constraints override that agency, or the recommended practices are ineffective, the playbook's central utility collapses.
Editorial extensions
If this is right
- Using the playbook from stage one would push teams to ask whether a task should be built at all, before any data is collected or model trained.
- Following the documentation checklists would produce a written record of decisions and trade-offs at each stage, making impact statements and external audits more feasible.
- Evaluation guidance would move teams from global accuracy to group-specific metrics and sociotechnical measurement, surfacing performance disparities that aggregate scores hide.
- Deployment guidance would add refusals, safeguards, red teaming, and human recourse as standard parts of release planning rather than reactions to incidents.
- If adopted widely, the playbook would normalize ethics review as an ongoing practice across the lifecycle instead of a one-time approval.
Reading between the lines
- The paper does not test whether following the playbook actually reduces harms; a natural next step would be an empirical study comparing teams that use it with matched teams that do not.
- The checklists could plausibly be encoded as machine-readable audit artifacts or automated into model-development tooling, though the paper does not propose this.
- The agency premise suggests that organizational support is the real constraint: a practitioner who lacks authority to act on the checklists would not be helped by the playbook alone.
- The lifecycle structure and harm taxonomy could transfer to non-generative ML systems and to policy or procurement reviews, even though the examples are drawn from generative AI.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript is a practitioner-facing playbook organized around the six stages of the AI lifecycle: problem formulation, dataset, model design, model training, model evaluation, and model use and monitoring. For each stage it provides a transparency and documentation checklist, common considerations, case studies of harms, and mitigation strategies, drawing on a literature review of over 100 sources and feedback from a set of interdisciplinary experts. The central claim is that the playbook helps practitioners diagnose potential harms and provides concrete guidance and resources for mitigation strategies to reduce the negative impact of those harms in generative AI systems.
Significance. If taken as a survey-style resource rather than as a validated intervention, the playbook is a useful contribution: it consolidates a scattered literature into a navigable lifecycle structure, gives concrete checklists and case studies, and repeatedly signals where the field lacks settled answers (e.g., §3.4.3, §6.4.3, §7.2.3). It also engages with critical scholarship on bias, measurement, and participatory methods, which strengthens its credibility as a synthesis. Its main weakness is the gap between the strong framing of the central claim and the absence of any evidence that following the recommended strategies actually reduces harm; the paper itself concedes several places where recommended mitigations are unproven, circumventable, or capable of introducing new harms.
major comments (3)
- [Section 1 and throughout] The paper's central claim—that it provides 'concrete guidance... for mitigation strategies to reduce the negative impact' of harms—is never tested. No section reports evidence that following the playbook's checklists or using its recommended tools reduces harm, and several sections concede the opposite: §3.4.3 states there is no 'best' or 'correct' way to filter data yet, §6.4.3 calls evaluation of hallucinations and misinformation a 'nascent research discipline,' and §7.2.3 admits both refusal options are 'circumventable.' To make the claim load-bearing, either reframe the contribution as a synthesis of current best practices without a validated-effect claim, or add a section on how practitioners can evaluate whether a chosen mitigation actually reduced harm in their context.
- [Table 9 and §3.3.3] Table 9 demonstrates that filtering 'bad words' from the C4 dataset disproportionately removes African American English, Hispanic English vernacular, and LGBTQ+ identity text, causing quality-of-service and representational harms. Section 3.3.3 nevertheless recommends deciding whether to filter toxic or hateful content and acknowledges tradeoffs only in the abstract. The playbook lacks a decision framework for weighing a mitigation's benefit against its documented side effects, so a practitioner following the guidance could adopt filtering that itself creates or worsens harms. This is a load-bearing gap for the promise that the playbook 'reduces the negative impact' of harms.
- [§7.2.3] The recommended refusal options are introduced as 'currently accepted best practices,' but the same subsection concedes that both options are circumventable, Table 35 shows an example of circumvention, and Table 36 shows harms from over-refusal. The playbook would be strengthened by an explicit treatment of when refusals or safeguards should be preferred over alternative mitigations or no refusal at all, and by guidance for testing circumvention during red teaming rather than treating 'circumventable' as a binary property.
minor comments (6)
- [§5.2.3] The text contains a standalone placeholder heading 'Mitigation Strategy Header' in the middle of the environmental-impact mitigation strategies; it should be completed or removed.
- [§6.3] The definition of Domain reads 'Domain: Domain: the specific context...' with the word 'Domain' duplicated.
- [§2.2.3] The heading 'Methods of Accountability.' appears twice in quick succession within the same mitigation-strategies subsection.
- [§1.0.2] In the sentence 'other stakeholders would might be affected by your work,' the phrase 'would might' should be corrected to 'who might.'
- [Abstract] The abstract says the playbook helps 'minimize negative impacts' without the caveats and uncertainties that are acknowledged in the body; aligning the abstract with those limitations would make the contribution easier to assess.
- [Table 33] The case study shifts between 'darker skin tones' and 'dark skin color' to describe the same population; one term should be used consistently throughout the table.
Circularity Check
No circularity: the playbook makes no derivation and its recommendations are attributed to external or prior peer-reviewed sources, so none of the enumerated circularity patterns apply.
full rationale
This is a curated survey/playbook rather than a derivation chain. Its central promise — "help you diagnose potential harms... providing concrete guidance and resources for mitigation strategies" (Section 1) — is not derived from any fitted quantity or mathematically defined construct, and no equation or prediction is constructed from its own inputs. The harm taxonomy is explicitly "adapted from [113, 125]" (Section 1.2), and the checklists and mitigation pointers draw on over 100 cited resources. The few author self-citations (e.g., Sap et al. on toxicity in [54, 87, 108, 109], Deng et al. on auditing in [42, 43, 44, 49, 74], Dodge et al. on PII in [117]) are used as primary-source pointers to previously published, externally reviewed work rather than as the sole justification for a conclusion. Where the playbook makes claims that could be challenged, it explicitly flags the evidentiary limits — e.g., "there is currently not a “best” or “correct” way to do data filtering yet" (Section 3.4.3), evaluating hallucinations is "a nascent research discipline" (Section 6.4.3), and both refusal options "are circumventable" (Section 7.2.3). Those concessions reduce confidence in the effectiveness of the recommended mitigations, but that is a validity/evidence concern, not a circularity concern. No step in the playbook reduces by construction to a self-citation or to a fitted input renamed as a prediction, so the appropriate circularity finding is none.
Assumptions & free parameters
assumptions (3)
- domain assumption AI practitioners have considerable agency to influence the social impacts of their research and system design.
- domain assumption The playbook assumes value pluralism, that different people have different values and may reach different ethical conclusions.
- domain assumption The cited mitigation strategies are effective in reducing harms in the reader's context.
Cite this review
Pith. "Pith review of The Generative AI Ethics Playbook." pith.science (2026). https://pith.science/paper/DHPICCHR
@misc{pith2026250110383,
author = {Pith},
title = {Pith review of: The Generative AI Ethics Playbook},
year = {2026},
howpublished = {\url{https://pith.science/paper/DHPICCHR}},
note = {Machine review of arXiv:2501.10383}
}
read the original abstract
The Generative AI Ethics Playbook provides guidance for identifying and mitigating risks of machine learning systems across various domains, including natural language processing, computer vision, and generative AI. This playbook aims to assist practitioners in diagnosing potential harms that may arise during the design, development, and deployment of datasets and models. It offers concrete strategies and resources for mitigating these risks, to help minimize negative impacts on users and society. Drawing on current best practices in both research and ethical considerations, this playbook aims to serve as a comprehensive resource for AI/ML practitioners. The intended audience of this playbook includes machine learning researchers, engineers, and practitioners who are involved in the creation and implementation of generative and multimodal models (e.g., text-to-text, image-to-image, text-to-image, text-to-video). Specifically, we provide transparency/documentation checklists, topics of interest, common questions, examples of harms through case studies, and resources and strategies to mitigate harms throughout the Generative AI lifecycle. This playbook was made collaboratively over the course of 16 months through extensive literature review of over 100 resources and peer-reviewed articles, as well as through an initial group brainstorming session with 18 interdisciplinary AI ethics experts from industry and academia, and with additional feedback from 8 experts (5 of whom were in the initial brainstorming session). We note that while this playbook provides examples, discussion, and harm mitigation strategies, research in this area is ongoing. Our playbook aims to be a practically useful survey, taking a high-level view rather than aiming for covering the entire existing body of research.
Reference graph
Works this paper leans on
-
[1]
[n. d.]. AI & Ethics: Collaborative Activities for Designers — ideo.com. https://www.ideo.com/journal/ai-ethics-collaborative-activities- for-designers. [Accessed 16-10-2024]
2024
-
[2]
[n. d.]. GitHub - fau-masters-collected-works-cgarbin/datasheet-for-dataset-template: Template for datasheet for datasets — github.com. https://github.com/fau-masters-collected-works-cgarbin/datasheet-for-dataset-template. [Accessed 16-10-2024]
2024
-
[3]
[n. d.]. Google AI Principles – Google AI — ai.google. https://ai.google/responsibility/principles/. [Accessed 15-10-2024]
2024
-
[4]
[n. d.]. Holistic Evaluation of Language Models (HELM) — crfm.stanford.edu. https://crfm.stanford.edu/helm/lite/latest/. [Accessed 27-10-2024]
2024
-
[5]
[n. d.]. Participatory data stewardship — adalovelaceinstitute.org. https://www.adalovelaceinstitute.org/report/participatory-data- stewardship/. [Accessed 22-10-2024]
2024
-
[6]
[n. d.]. PERVADE data ethics tool. https://pervade.umd.edu/pervade-data-ethics-tool. [Accessed 22-10-2024]
2024
-
[7]
[n. d.]. The Industrialization of Terrorist Propaganda: Neural Language Models and the Threat of Fake Content Generation — middlebury.edu. https://www.middlebury.edu/institute/academics/centers-initiatives/ctec/ctec-publications/industrialization-terrorist- propaganda-neural. [Accessed 16-10-2024]
2024
-
[8]
[n. d.]. The Tarot Cards Of Tech — tarotcardsoftech.artefactgroup.com. https://tarotcardsoftech.artefactgroup.com/. [Accessed 16-10-2024]
2024
Show all 139 references
-
[9]
[n. d.]. Using AI Factsheets for AI Governance | IBM Cloud Pak for Data as a Service — dataplatform.cloud.ibm.com. https://dataplatform. cloud.ibm.com/docs/content/wsj/analyze-data/factsheets-model-inventory.html?context=cpdaas. [Accessed 16-10-2024]
2024
-
[10]
GitHub - PovertyAction/PII_detection: Application and python script to identify, remove, and/or recode personally identifiable information (PII) from field experiment datasets
2020. GitHub - PovertyAction/PII_detection: Application and python script to identify, remove, and/or recode personally identifiable information (PII) from field experiment datasets. — github.com. https://github.com/PovertyAction/PII_detection. [Accessed 20-10-2024]
2020
-
[11]
Introducing the Data Measurements Tool: an Interactive Tool for Looking at Datasets — huggingface.co
2021. Introducing the Data Measurements Tool: an Interactive Tool for Looking at Datasets — huggingface.co. https://huggingface.co/ blog/data-measurements-tool. [Accessed 20-10-2024]
2021
-
[12]
GitHub - tokern/piicatcher: Scan databases and data warehouses for PII data
2023. GitHub - tokern/piicatcher: Scan databases and data warehouses for PII data. Tag tables and columns in data catalogs like Amundsen and Datahub — github.com. https://github.com/tokern/piicatcher. [Accessed 20-10-2024]
2023
-
[13]
GitHub - microsoft/presidio: Context aware, pluggable and customizable data protection and de-identification SDK for text and images — github.com
2024. GitHub - microsoft/presidio: Context aware, pluggable and customizable data protection and de-identification SDK for text and images — github.com. https://github.com/microsoft/presidio. [Accessed 20-10-2024]
2024
-
[14]
Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. 2016. Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC conference on computer and communications security . 308–318
2016
-
[15]
Amina Adadi and Mohammed Berrada. 2018. Peeking inside the black-box: a survey on explainable artificial intelligence (XAI). IEEE access 6 (2018), 52138–52160
2018
-
[16]
Wilbert G Aguilar, Darwin Alulema, Alex Limaico, and David Sandoval. 2017. Development and verification of a verbal corpus based on natural language for Ecuadorian dialect. In 2017 IEEE 11th International Conference on Semantic Computing (ICSC) . IEEE, 515–519
2017
-
[17]
Open AI. [n. d.]. https://openai.com/index/reducing-bias-and-improving-safety-in-dall-e-2/. [Accessed 26-10-2024]
2024
-
[18]
Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner. 2022. Machine bias. InEthics of data and analytics. Auerbach Publications, 254–264
2022
-
[19]
Tuomas Aura, Thomas A Kuhn, and Michael Roe. 2006. Scanning electronic documents for personally identifiable information. In Proceedings of the 5th ACM workshop on Privacy in electronic society . 41–50
2006
-
[20]
Emily M Bender and Batya Friedman. 2018. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics 6 (2018), 587–604
2018
-
[21]
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big?. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency . 610–623
2021
-
[22]
It’s complicated
Cynthia L Bennett, Cole Gleason, Morgan Klaus Scheuerman, Jeffrey P Bigham, Anhong Guo, and Alexandra To. 2021. “It’s complicated”: Negotiating accessibility and (mis) representation in image descriptions of race, gender, and disability. In Proceedings of the 2021 CHI Conferen...
2021
-
[23]
Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, and Aylin Caliskan. 2023. Easily accessible text-to-image generation amplifies demographic stereotypes at large scale. In Proceedings of the...
2023
-
[24]
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. Language (technology) is power: A critical survey of" bias" in nlp. arXiv preprint arXiv:2005.14050 (2020)
2020 arXiv
-
[25]
Su Lin Blodgett, Lisa Green, and Brendan O’Connor. 2016. Demographic dialectal variation in social media: A case study of African- American English. arXiv preprint arXiv:1608.08868 (2016)
2016 arXiv
-
[26]
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021. Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 1...
2021
-
[27]
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2022. Improving language models by retrieving from trillions of tokens. In International conf...
2022
-
[28]
Houda Bouamor, Nizar Habash, and Kemal Oflazer. 2014. A Multidialectal Parallel Corpus of Arabic.. In LREC. 1240–1245
2014
-
[29]
Joy Buolamwini and Timnit Gebru. 2018. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Conference on fairness, accountability and transparency . PMLR, 77–91
2018
-
[30]
Yang Trista Cao and Hal Daumé III. 2019. Toward gender-inclusive coreference resolution. arXiv preprint arXiv:1910.13913 (2019)
2019 arXiv
-
[31]
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. 2021. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21)....
2021
-
[32]
Natalie Garrett Casey Fiesler. 2020. Ethical Tech Starts With Addressing Ethical Debt — wired.com. https://www.wired.com/story/ opinion-ethical-tech-starts-with-addressing-ethical-debt/. [Accessed 22-10-2024]
2020
-
[33]
Isaac Caswell, Theresa Breiner, Daan Van Esch, and Ankur Bapna. 2020. Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus. arXiv preprint arXiv:2010.14571 (2020)
2020 arXiv
-
[34]
Eshwar Chandrasekharan, Mattia Samory, Shagun Jhaver, Hunter Charvat, Amy Bruckman, Cliff Lampe, Jacob Eisenstein, and Eric Gilbert. 2018. The Internet’s hidden rules: An empirical study of Reddit norm violations at micro, meso, and macro scales. Proceedings of the ACM on Huma...
2018
-
[35]
Huajie Chen, Deng Cai, Wei Dai, Zehui Dai, and Yadong Ding. 2019. Charge-based prison term prediction with deep gating network. arXiv preprint arXiv:1908.11521 (2019)
2019 arXiv
-
[36]
Myra Cheng, Tiziano Piccardi, and Diyi Yang. 2023. CoMPosT: Characterizing and evaluating caricature in LLM simulations. arXiv preprint arXiv:2310.11501 (2023)
2023 arXiv
-
[37]
Andrew R Chow. 2023. AI-human romances are flourishing—and this is just the beginning. Time https://time. com/6257790/ai-chatbots- love/(21 February 2023) (2023)
2023
-
[38]
Donavyn Coffey. 2021. M ¯aori are trying to save their language from Big Tech. Wired UK (2021)
2021
-
[39]
I Glenn Cohen, Boris Babic, Sara Gerke, Qiong Xia, Theodoros Evgeniou, and Klaus Wertenbroch. 2023. How AI can learn from the law: putting humans in the loop only on appeal. npj Digital Medicine 6, 1 (2023), 160
2023
-
[40]
Paula Czarnowska, Yogarshi Vyas, and Kashif Shah. 2021. Quantifying social biases in NLP: A generalization and empirical comparison of extrinsic fairness metrics. Transactions of the Association for Computational Linguistics 9 (2021), 1249–1267
2021
-
[41]
Nicholas Deas, Jessi Grieser, Shana Kleiner, Desmond Patton, Elsbeth Turcan, and Kathleen McKeown. 2023. Evaluation of African American language bias in natural language generation. arXiv preprint arXiv:2305.14291 (2023)
2023 arXiv
-
[42]
Wesley Hanwen Deng, Boyuan Guo, Alicia Devrio, Hong Shen, Motahhare Eslami, and Kenneth Holstein. 2023. Understanding Practices, Challenges, and Opportunities for User-Engaged Algorithm Auditing in Industry Practice. In Proceedings of the 2023 CHI Conference on Human Factors i...
2023
-
[43]
Wesley Hanwen Deng, Manish Nagireddy, Michelle Seng Ah Lee, Jatinder Singh, Zhiwei Steven Wu, Kenneth Holstein, and Haiyi Zhu. 2022. Exploring How Machine Learning Practitioners (Try To) Use Fairness Toolkits. In 2022 ACM Conference on Fairness, Accountability, and Transparenc...
2022
-
[44]
Wesley Hanwen Deng, Nur Yildirim, Monica Chang, Motahhare Eslami, Kenneth Holstein, and Michael Madaio. 2023. Investigating Practices and Opportunities for Cross-functional Collaboration around AI Fairness in Industry Practice. In Proceedings of the 2023 ACM Conference on Fair...
2023
-
[45]
Sunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian, Jeff M Phillips, and Kai-Wei Chang. 2021. Harms of gender exclusivity and challenges in non-binary representation in language technologies. arXiv preprint arXiv:2108.12084 (2021)
2021 arXiv
-
[46]
Jacob Eisenstein. 2013. What to do about bad language on the internet. In Proceedings of the 2013 conference of the North American Chapter of the association for computational linguistics: Human language technologies . 359–369
2013
-
[47]
Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Pete Walsh, Dirk Groeneveld, Luca Soldaini, Sameer Singh, et al. 2023. What’s In My Big Data? arXiv preprint arXiv:2310.20707 (2023). 60 • Jessie J. Smith, Wesley Hanwen Deng, Willi...
2023 arXiv
-
[48]
Entrepreneur en Español. 2021. Xiaoice robot users have ended up in therapy for falling in love with their artificial intelligence. https://www.expressnews.com/business/article/XiaoIce-robot-users-have-ended-up-in-therapy-for-16414790.php
2021
-
[49]
Michael Feffer, Anusha Sinha, Wesley H Deng, Zachary C Lipton, and Hoda Heidari. 2024. Red-teaming for generative ai: Silver bullet or security theater?. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , Vol. 7. 421–437
2024
-
[50]
Eve Fleisig, Rediet Abebe, and Dan Klein. 2023. When the majority is wrong: Modeling annotator disagreement for subjective tasks. arXiv preprint arXiv:2305.06626 (2023)
2023 arXiv
-
[51]
Kathleen C Fraser, Svetlana Kiritchenko, and Isar Nejadgholi. 2023. A friendly face: Do text-to-image systems rely on stereotypes when the input is under-specified? arXiv preprint arXiv:2302.07159 (2023)
2023 arXiv
-
[52]
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al. 2022. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:220...
2022 arXiv
-
[53]
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford
-
[54]
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. EMNLP Findings (2020)
2020
-
[55]
Seraphina Goldfarb-Tarrant, Eddie Ungless, Esma Balkir, and Su Lin Blodgett. 2023. This prompt is measuring< mask>: evaluating bias evaluation in language models. arXiv preprint arXiv:2305.12757 (2023)
2023 arXiv
-
[56]
Robert Gorwa, Reuben Binns, and Christian Katzenbach. 2020. Algorithmic content moderation: Technical and political challenges in the automation of platform governance. Big Data & Society 7, 1 (2020), 2053951719897945
2020
-
[57]
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan. 2015. Deep learning with limited numerical precision. In International conference on machine learning . PMLR, 1737–1746
2015
-
[58]
Moritz Hardt, Eric Price, and Nati Srebro. 2016. Equality of opportunity in supervised learning. Advances in neural information processing systems 29 (2016)
2016
-
[59]
George H Heilmeier. 2023. The Heilmeier Catechism. Defense Advanced Research Projects Agency (DARPA), https://www. darpa. mil/work-with-us/heilmeier-catechism (2023)
2023
-
[60]
Peter Henderson, Jieru Hu, Joshua Romoff, Emma Brunskill, Dan Jurafsky, and Joelle Pineau. 2020. Towards the systematic reporting of the energy and carbon footprints of machine learning. Journal of Machine Learning Research 21, 248 (2020), 1–43
2020
-
[61]
Nora Hollenstein and Noëmi Aepli. 2015. A resource for natural language processing of Swiss German dialects. University of Zurich (2015)
2015
-
[62]
Dirk Hovy. 2016. The enemy in your own camp: How well can we detect statistically-generated fake reviews–an adversarial study. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) . 351–356
2016
-
[63]
Dirk Hovy and Shannon L Spruit. 2016. The social impact of natural language processing. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) . 591–598
2016
-
[64]
HuggingFace. [n. d.]. Red-Teaming Large Language Models — huggingface.co. https://huggingface.co/blog/red-teaming. [Accessed 27-10-2024]
2024
-
[65]
Gautier Izacard and Edouard Grave. 2020. Leveraging passage retrieval with generative models for open domain question answering. arXiv preprint arXiv:2007.01282 (2020)
2020 arXiv
-
[66]
Abigail Z Jacobs and Hanna Wallach. 2021. Measurement and fairness. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency. 375–385
2021
-
[67]
Seyyed Ahmad Javadi, Chris Norval, Richard Cloete, and Jatinder Singh. 2021. Monitoring AI services for misuse. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society . 597–607
2021
-
[68]
Yacine Jernite, Huu Nguyen, Stella Biderman, Anna Rogers, Maraim Masoud, Valentin Danchev, Samson Tan, Alexandra Sasha Luccioni, Nishant Subramani, Isaac Johnson, et al. 2022. Data governance in the age of large-scale data-driven language technology. InProceedings of the 2022 ...
2022
-
[69]
Akshita Jha, Aida Davani, Chandan K Reddy, Shachi Dave, Vinodkumar Prabhakaran, and Sunipa Dev. 2023. SeeGULL: A stereotype benchmark with broad geo-cultural coverage leveraging generative models. arXiv preprint arXiv:2305.11840 (2023)
2023 arXiv
-
[70]
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019. Tinybert: Distilling bert for natural language understanding. arXiv preprint arXiv:1909.10351 (2019)
2019 arXiv
-
[71]
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2004.04906 (2020)
2020 arXiv
-
[72]
Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. 2019. Ctrl: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858 (2019)
2019 arXiv
-
[73]
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2019. Generalization through memorization: Nearest neighbor language models. arXiv preprint arXiv:1911.00172 (2019). The Generative AI Ethics Playbook • 61
2019 arXiv
-
[74]
Sara Kingsley, Jiayin Zhi, Wesley Hanwen Deng, Jaimie Lee, Sizhe Zhang, Motahhare Eslami, Kenneth Holstein, Jason I Hong, Tianshi Li, and Hong Shen. 2024. Investigating What Factors Influence Users’ Rating of Harmful Algorithmic Bias and Discrimination. In Proceedings of the A...
2024
-
[75]
Aniket Kittur, Jeffrey V Nickerson, Michael Bernstein, Elizabeth Gerber, Aaron Shaw, John Zimmerman, Matt Lease, and John Horton
-
[76]
Allison Koenecke, Andrew Nam, Emily Lake, Joe Nudell, Minnie Quartey, Zion Mengesha, Connor Toups, John R Rickford, Dan Jurafsky, and Sharad Goel. 2020. Racial disparities in automated speech recognition. Proceedings of the national academy of sciences 117, 14 (2020), 7684–7689
2020
-
[77]
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals. 2019. Dynamic evaluation of transformer language models. arXiv preprint arXiv:1904.08378 (2019)
2019 arXiv
-
[78]
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, et al. 2022. Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association...
2022
-
[79]
Amit Kulkarni. 2021. GitHub Copilot AI is leaking functional API keys. Analytics Drift (2021)
2021
-
[80]
Mina Lee, Megha Srivastava, Amelia Hardy, John Thickstun, Esin Durmus, Ashwin Paranjape, Ines Gerard-Ursin, Xiang Lisa Li, Faisal Ladhak, Frieda Rong, et al. 2022. Evaluating human-language model interaction. arXiv preprint arXiv:2212.09746 (2022)
2022 arXiv
-
[81]
Kobi Leins, Jey Han Lau, and Timothy Baldwin. 2020. Give me convenience and give her death: Who should decide what uses of NLP are appropriate, and on what basis? arXiv preprint arXiv:2005.13213 (2020)
2020 arXiv
-
[82]
Zhuohan Li, Siyuan Zhuang, Shiyuan Guo, Danyang Zhuo, Hao Zhang, Dawn Song, and Ion Stoica. 2021. Terapipe: Token-level pipeline parallelism for training large-scale language models. In International Conference on Machine Learning . PMLR, 6543–6552
2021
-
[83]
Bin Liang, Hongcheng Li, Miaoqiang Su, Xirong Li, Wenchang Shi, and Xiaofeng Wang. 2018. Detecting adversarial image examples in deep neural networks with adaptive noise reduction. IEEE Transactions on Dependable and Secure Computing 18, 1 (2018), 72–85
2018
-
[84]
Q Vera Liao and Ziang Xiao. 2023. Rethinking model evaluation as narrowing the socio-technical gap. arXiv preprint arXiv:2306.03100 (2023)
2023 arXiv
-
[85]
Daniel J Liebling, Katherine Heller, Margaret Mitchell, Mark Díaz, Michal Lahav, Niloufar Salehi, Samantha Robertson, Samy Bengio, Timnit Gebru, and Wesley Deng. 2021. Three Directions for the Design of Human-Centered Machine Translation. (2021)
2021
-
[86]
Daniel Link, Bernd Hellingrath, and Jie Ling. 2016. A Human-is-the-Loop Approach for Semi-Automated Content Moderation.. In ISCRAM
2016
-
[87]
Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A Smith, and Yejin Choi. 2021. DExperts: Decoding-time controlled text generation with experts and anti-experts. ACL (2021)
2021
-
[88]
Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, et al. 2023. A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity. arXiv preprint...
2023 arXiv
-
[89]
Kadan Lottick, Silvia Susai, Sorelle A Friedler, and Jonathan P Wilson. 2019. Energy Usage Reports: Environmental awareness as part of algorithmic accountability. arXiv preprint arXiv:1911.08354 (2019)
2019 arXiv
-
[90]
Henrietta Lyons, Senuri Wijenayake, Tim Miller, and Eduardo Velloso. 2022. What’s the appeal? Perceptions of review processes for algorithmic decisions. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems . 1–15
2022
-
[91]
Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Evan Pete Walsh, Yanai Elazar, Kyle Lo, et al. 2023. Paloma: A benchmark for evaluating language model fit. arXiv preprint arXiv:2312.10523 (2023)
2023 arXiv
-
[92]
Microsoft. 2022. Microsoft Responsible AI Impact Assessment Template. https://blogs.microsoft.com/wp-content/uploads/prod/sites/5/ 2022/06/Microsoft-RAI-Impact-Assessment-Template.pdf. [Accessed 22-10-2024]
2022
-
[93]
Heather Murphy. 2017. Why Stanford researchers tried to create a ‘gaydar’machine. The New York Times 9 (2017)
2017
-
[94]
Moin Nadeem, Anna Bethke, and Siva Reddy. 2020. StereoSet: Measuring stereotypical bias in pretrained language models. arXiv preprint arXiv:2004.09456 (2020)
2020 arXiv
-
[95]
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R Bowman. 2020. CrowS-pairs: A challenge dataset for measuring social biases in masked language models. arXiv preprint arXiv:2010.00133 (2020)
2020 arXiv
-
[96]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . 311–318
2002
-
[97]
D Patterson, J Gonzalez, Q Le, C Liang, LM Munguia, D Rothchild, D So, M Texier, and J Dean. 2021. Carbon emissions and large neural network training. arXiv 2021. arXiv preprint arXiv:2104.10350 (2021)
2021 arXiv
-
[98]
CITI Program. 2024. The Trusted Standard in Research, Ethics, Compliance, and Safety Training. https://about.citiprogram.org/. [Accessed 22-10-2024]
2024
-
[99]
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al . 2021. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint 62 • Jessie J. Smith...
2021 arXiv
-
[100]
Inioluwa Deborah Raji, Andrew Smart, Rebecca N White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes. 2020. Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing. In Proceedin...
2020
-
[101]
Swaroop Ramaswamy, Om Thakkar, Rajiv Mathews, Galen Andrew, H Brendan McMahan, and Françoise Beaufays. 2020. Training production language models without memorizing user data. arXiv preprint arXiv:2009.10031 (2020)
2020 arXiv
-
[102]
ACL Rolling Review. 2024. ACL Rolling Review — aclrollingreview.org. https://aclrollingreview.org/responsibleNLPresearch/. [Accessed 16-10-2024]
2024
-
[103]
Pedro Reviriego and Elena Merino-Gómez. 2022. Text to image generation: Leaving no language behind.arXiv preprint arXiv:2208.09333 (2022)
2022 arXiv
-
[104]
Paul Röttger, Bertie Vidgen, Dirk Hovy, and Janet B Pierrehumbert. 2021. Two contrasting data annotation paradigms for subjective NLP tasks. arXiv preprint arXiv:2112.07475 (2021)
2021 arXiv
-
[105]
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018. Gender Bias in Coreference Resolution. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (S...
2018 doi
-
[106]
Liebling Michal Lahav Katherine Heller Mark Díaz Samy Bengio Niloufar Salehi Saman- tha Robertson, Wesley Hanwen Deng
Timnit Gebru Margaret Mitchell Daniel J. Liebling Michal Lahav Katherine Heller Mark Díaz Samy Bengio Niloufar Salehi Saman- tha Robertson, Wesley Hanwen Deng. 2021. Three Directions for the Design of Human-Centered Machine Translation. In HCI + MT Workshop at EACL ’21
2021
-
[107]
Victor Sanh, Thomas Wolf, and Alexander Rush. 2020. Movement pruning: Adaptive sparsity by fine-tuning. Advances in neural information processing systems 33 (2020), 20378–20389
2020
-
[108]
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A Smith. 2019. The risk of racial bias in hate speech detection. In Proceedings of the 57th annual meeting of the association for computational linguistics . 1668–1678
2019
-
[109]
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022. Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection. In NAACL. https://aclanthology.org/2022.naacl-main.431/
2022
-
[110]
Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021. Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp. Transactions of the Association for Computational Linguistics 9 (2021), 1408–1424
2021
-
[111]
Ari Schlesinger, Kenton P O’Hara, and Alex S Taylor. 2018. Let’s talk about race: Identity, chatbots, and AI. In Proceedings of the 2018 chi conference on human factors in computing systems . 1–14
2018
-
[112]
Shawn Shan, Wenxin Ding, Josephine Passananti, Haitao Zheng, and Ben Y Zhao. 2023. Prompt-specific poisoning attacks on text-to-image generative models. arXiv preprint arXiv:2310.13828 (2023)
2023 arXiv
-
[113]
Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Rostamzadeh, Paul Nicholas, N’Mah Yilla-Akbari, Jess Gallegos, Andrew Smart, Emilio Garcia, et al. 2023. Sociotechnical harms of algorithmic systems: Scoping a taxonomy for harm reduction. In Proceedings of the 2...
2023
-
[114]
Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Boyd-Graber, and Lijuan Wang. 2022. Prompting gpt-3 to be reliable. arXiv preprint arXiv:2210.09150 (2022)
2022 arXiv
-
[115]
I’m sorry to hear that
Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022. " I’m sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset. arXiv preprint arXiv:2205.09209 (2022)
2022 arXiv
-
[116]
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2020. Energy and policy considerations for modern deep learning research. In Proceedings of the AAAI conference on artificial intelligence , Vol. 34. 13693–13696
2020
-
[117]
Nishant Subramani, Sasha Luccioni, Jesse Dodge, and Margaret Mitchell. 2023. Detecting personal information in training corpora: an analysis. In Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023) . 208–220
2023
-
[118]
Ryo Takahashi, Takashi Matsubara, and Kuniaki Uehara. 2019. Data augmentation using random image cropping and patching for deep CNNs. IEEE Transactions on Circuits and Systems for Video Technology 30, 9 (2019), 2917–2931
2019
-
[119]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023)
2023 arXiv
-
[120]
C Tran, S Bhosale, J Cross, P Koehn, S Edunov, and A Fan. 2021. Facebook ai wmt21 news translation task submission. arXiv
2021
-
[121]
James Vincent. [n. d.]. Google ‘fixed’ its racist algorithm by removing gorillas from its image-labeling tech — theverge.com. https: //www.theverge.com/2018/1/12/16882408/google-racist-gorillas-photo-recognition-algorithm-ai. [Accessed 27-10-2024]
2018
-
[122]
Paul Voigt and Axel Von dem Bussche. 2017. The eu general data protection regulation (gdpr).A Practical Guide, 1st Ed., Cham: Springer International Publishing 10, 3152676 (2017), 10–5555
2017
-
[123]
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020. Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. Advances in Neural Information Processing Systems 33 (2020), 5776–5788
2020
-
[124]
Yilun Wang and Michal Kosinski. 2018. Deep neural networks are more accurate than humans at detecting sexual orientation from facial images. Journal of personality and social psychology 114, 2 (2018), 246. The Generative AI Ethics Playbook • 63
2018
-
[125]
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al. 2022. Taxonomy of risks posed by language models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability,...
2022
-
[126]
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang. 2021. Challenges in detoxifying language models. arXiv preprint arXiv:2109.07445 (2021)
2021 arXiv
-
[127]
AI supply chain
David Gray Widder and Dawn Nafus. 2023. Dislocated accountabilities in the “AI supply chain”: Modularity and developers’ notions of responsibility. Big Data & Society 10, 1 (2023), 20539517231177620
2023
-
[128]
Genta Indra Winata, Andrea Madotto, Zhaojiang Lin, Rosanne Liu, Jason Yosinski, and Pascale Fung. 2021. Language models are few-shot multilingual learners. arXiv preprint arXiv:2109.07684 (2021)
2021 arXiv
-
[129]
Jingjing Xu, Xuancheng Ren, Junyang Lin, and Xu Sun. 2018. Diversity-promoting GAN: A cross-entropy based generative adversarial network for diversified text generation. In Proceedings of the 2018 conference on empirical methods in natural language processing . 3940–3949
2018
-
[130]
L Xue. 2020. mt5: A massively multilingual pre-trained text-to-text transformer. arXiv preprint arXiv:2010.11934 (2020)
2020 arXiv
-
[131]
Muhammad Bilal Zafar, Isabel Valera, Manuel Gomez Rogriguez, and Krishna P Gummadi. 2017. Fairness constraints: Mechanisms for fair classification. In Artificial intelligence and statistics. PMLR, 962–970
2017
-
[132]
Rich Zemel, Yu Wu, Kevin Swersky, Toni Pitassi, and Cynthia Dwork. 2013. Learning fair representations. In International conference on machine learning. PMLR, 325–333
2013
-
[133]
Brian Hu Zhang, Blake Lemoine, and Margaret Mitchell. 2018. Mitigating unwanted biases with adversarial learning. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society . 335–340
2018
-
[134]
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2017. Men also like shopping: Reducing gender bias amplification using corpus-level constraints. arXiv preprint arXiv:1707.09457 (2017)
2017 arXiv
-
[135]
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender bias in coreference resolution: Evaluation and debiasing methods. arXiv preprint arXiv:1804.06876 (2018)
2018 arXiv
-
[136]
Caleb Ziems, Jiaao Chen, Camille Harris, Jessica Anderson, and Diyi Yang. 2022. VALUE: Understanding dialect disparity in NLU. arXiv preprint arXiv:2204.03031 (2022)
2022 arXiv
-
[137]
Caleb Ziems, William Held, Jingfeng Yang, Jwala Dhamala, Rahul Gupta, and Diyi Yang. 2022. Multi-VALUE: A framework for cross-dialectal English NLP. arXiv preprint arXiv:2212.08011 (2022)
2022 arXiv
-
[2013]
In Proceedings of the 2013 conference on Computer supported cooperative work
The future of crowd work. In Proceedings of the 2013 conference on Computer supported cooperative work . 1301–1318
2013
-
[2021]
Datasheets for datasets. Commun. ACM 64, 12 (2021), 86–92
2021
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.