Pith. sign in

REVIEW 3 major objections 5 minor 101 references

Impact Assessment Card: Communicating Risks and Benefits of AI Uses

T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read The paper claims that a one-page Impact Assessment Card lets people write faster, higher-quality AI recommendation emails than a full impact assessment report does.

desk verdict The card is a genuinely useful design artifact, but the headline quantitative claim is undercut by an unresolved contradiction between the two main statistical tables. read the letter →

arxiv 2508.18919 v1 pith:3HACTWNQ submitted 2025-08-26 cs.HC

classification cs.HC
keywords impactassessmentAIgovernanceriskcommunicationdatavisualizationaccessibledocumentationEUActuserstudymodelcards
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the standard way of communicating AI risks and benefits—long impact assessment reports—leaves most people out, and proposes a one-page Impact Assessment Card that condenses the same information. Drawing on label-like formats from food and energy and on input from three focus groups, the authors designed a card with a risk summary bar, benefits, risks with mitigations, data and model details, and reporting channels. In an online study with 235 participants—AI developers, compliance experts, and U.S.-representative ordinary individuals—people using the card wrote recommendation emails faster and of higher quality than people using a full report, across all three groups. The card also scored higher on usability and was preferred by 58% of experts and 70% of ordinary participants. The intended contribution is a simple, accessible medium that supports EU AI Act-style risk disclosure and lets non-experts take part in AI governance.

What carries the argument

The Impact Assessment Card itself is the central artifact. It is a one-page template organized into three bands: a header with the system's name, a plain-language description built on five components (purpose, deployer, subject, capability, domain), and an EU AI Act risk-classification bar; a middle band with benefits, a combined risks-and-mitigations table, and a data/model performance section; and a footer with reporting channels, registered office, and certifications. The card's working mechanism is condensation plus visual encoding: short phrases, a color-coded risk summary bar, heatmaps for data and stakeholders, and a QR code linking to the full report. It is designed so the same artif

What would settle it

Have an independent organization write the full impact assessment reports for the same two systems using a regulatory template, then rerun the identical 235-participant email task; if the card no longer yields faster, higher-quality emails, the reported advantage belonged to the specific baseline report, not the card format.

Watch

Extended reading notes

Core claim

The central discovery is that a compact, visually structured one-page card can communicate AI risk-and-benefit information more effectively than a full impact assessment report, even for expert readers. The authors built the card through iterative design: 14 design patterns from the literature, three focus groups that produced 12 speculative card designs and 8 design requirements, and four internal refinement rounds. The evaluation used two hypothetical AI systems (a biometric checkout classified as high risk and a license plate detector classified as limited risk under the EU AI Act) and asked participants to write an email recommending or rejecting deployment. Blind ratings of email qualit

Load-bearing premise

The card's advantage over the report assumes the baseline report is a fair, typical example of current impact-assessment practice; if the authors' report was unusually dense or poorly written, the speed and quality gains could be an artifact of that particular comparison.

Editorial extensions

If this is right

  • Regulators and deployers could adopt the card as a public-facing companion to mandatory impact assessment reports, with the full report remaining the legal record.
  • Because the card helped technical experts too, concise documentation formats may improve internal AI governance workflows, not just public communication.
  • The card's risk summary bar gives ordinary people a direct, glanceable link to EU AI Act risk classes, potentially supporting informed consent and public accountability.
  • Cards for digital AI systems (a recommender system and a benefits-allocation assistant) suggest the format generalizes beyond physically situated systems.
  • The result implies that documentation length is not the main driver of understanding; structure and visual encoding matter.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the format advantage persists with externally produced baseline reports, one-page 'nutrition label' style summaries could become a standard regulatory disclosure layer for AI systems, sitting between full reports and certification labels.
  • A natural next test is whether the card changes actual decisions—such as opt-in or opt-out choices and procurement preferences—rather than only email-writing quality and speed.
  • The card's structure could be extended machine-readably, so risk summaries are generated automatically from structured impact-assessment data, making comparisons across competing AI services feasible.
  • Because participants preferred the card but some experts wanted more depth, a tiered design (card, then summary report, then full report) may serve both quick scans and due diligence better than either format alone.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents the design and evaluation of an Impact Assessment Card, a one-page artifact intended to communicate AI risks and benefits to diverse audiences. The design process combines a literature review, three focus groups with 12 participants, and iterative refinement. The card is then evaluated in a within-subjects online study with 235 participants (AI developers, compliance experts, and a US Census-matched sample of ordinary individuals) who write recommendation emails using either the card or a full impact assessment report for two hypothetical AI systems. The authors report that the card yields faster task completion, higher-quality emails, higher usability ratings, and greater preference across cohorts. The paper also provides example cards for four AI systems and discusses implications for AI governance.

Significance. If the quantitative findings are correct, the card is a practical, low-cost communication tool with benefits for expert and non-expert stakeholders, and the paper makes a useful contribution to CSCW/HCI work on responsible AI artifacts. Strengths include a transparent positionality statement, a public artifact repository, a concrete email-writing task with a published scoring rubric, blind rating of email quality with 85% inter-rater agreement, and an effort to recruit a census-matched public sample. However, the paper's central empirical claim currently rests on an internal inconsistency in the regression results (Tables 4 and 5) and on an author-built baseline report whose representativeness is not established. These issues must be resolved before the headline findings can be accepted.

major comments (3)
  1. [§5.4.1, Tables 4 and 5] The two analyses of task quality contradict each other on the treatment effect, which is the load-bearing result for the abstract's claim of 'higher-quality emails.' Table 4 reports the coefficient for 'Treatment type Card vs. Report' as -0.987 (p<0.001), indicating lower quality for the card under the stated coding; Table 5 reports a positive mean difference of +1.207 (p<0.001). The same sign reversal appears for 'Type of task Reject vs. Recommend' (0.880, p=0.325 vs. -0.014, p=0.719) and for 'System Plate Detector vs. Checkout' (0.150 vs. -0.156). Because the regression and mean-difference tables cannot both be correct with the same dummy coding, the reported effect size and direction for the primary outcome are unreliable. The authors need to disclose the exact coding, report corrected regression estimates, and reconcile the narrative in §5.4.1 with the numbers.
  2. [§5.1, baseline report] The comparison condition is an impact assessment report written by the same authors to 'mirror the card's content' and is not validated against published state-of-the-art reports or readability benchmarks. The large mean difference in quality (3.327 vs. 2.12) could therefore reflect the particular report's density rather than a general advantage of the card over current practice. To support the general claim that cards outperform 'current methods such as technical reports,' the authors should either benchmark the baseline against existing reports (e.g., published examples [18, 63, 82] or a readability/comprehension measure) or temper the claim to 'this card outperforms this report.' This is a generalizability issue, not a circularity one, but it affects the external validity of the headline.
  3. [§5.3 Analysis and Table 5] Each participant completed both the card and the report condition (within-subjects design), but the mean-difference testing in Table 5 reports Mann-Whitney p-values that appear to treat the 470 observations as independent. This ignores the pairing and the two observations per participant. Table 5 also reports p=0.0*** for the treatment row. The analysis should use paired tests or a mixed model with a random intercept for participants (and possibly for system/task order), and should report the exact test and sample size. The current presentation makes it impossible to assess the uncertainty of the most important comparison.
minor comments (5)
  1. [Table 4 caption] The description of random effects is unclear: 'Random effects were included to account for variability in task quality based on participants’ self-selected decisions to reject or recommend the system' conflates a fixed factor (type of task) with the random-effect structure. Please clarify that the random effect is participant, and specify the full model formula.
  2. [Appendix A.5] Tables 6–8 use 'Legal Experts' while the body uses 'Compliance experts.' Please unify terminology.
  3. [Table 5] The treatment row reports p=0.0***; use the standard notation p<0.001.
  4. [Appendix figures] There are typographical artifacts in the card images and appendix text (e.g., 'a/f_ter', 'se/t_tings', 'li/t_tle'). Please proofread the final PDF and vector graphics.
  5. [§5.4.2 qualitative analysis] The thematic analysis of open-ended responses is summarized with illustrative quotes, but no intercoder reliability or coding scheme details are reported. Adding this information would strengthen confidence in the qualitative findings.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation; the card's advantage is an empirical result, not a construction artifact.

full rationale

The paper makes no formal derivation or predictive claim; it designs an Impact Assessment Card and evaluates it against an author-built baseline report in an online study. The central claim (faster completion, higher-quality emails) is an empirical outcome of user behavior and blind-rated email quality, not a logical consequence of the card's definition. No parameter is fitted to a subset and then renamed a prediction; no uniqueness theorem is imported from prior work to force the design; no ansatz is smuggled via citation. The authors' prior publications ([6], [7], [8], [17]) are cited for context, design-process stimuli, and related tools, but they do not generate the reported treatment effect. The author-built baseline and author-built rubric create validity/bias risks, and the sign inconsistency between Table 4 and Table 5 for the treatment effect is a correctness concern, but neither is circularity in the sense of an equation reducing to its own inputs. Under the stated rules, this is a 'no significant circularity' finding (score 0-2); I assign 1 rather than 0 only to mark the self-citations and author-built comparison as minor provenance considerations.

Assumptions & free parameters 0 free parameters · 6 assumptions · 1 invented entities

This empirical design study does not perform a derivation, so the ledger lists the assumptions that must hold for the study's comparison to support the central claim, plus the new artifact the paper introduces.

assumptions (6)
  • domain assumption The baseline impact assessment report created by the authors is representative of real-world state-of-the-art reports.
    If the baseline is a strawman, the card's apparent advantage is not a fair comparison. Entered in Section 5.1.
  • domain assumption Email quality, scored by two authors with a custom rubric, is a valid measure of how well participants understood and could communicate risks and benefits.
    The central outcome metric is the authors' own Likert rubric; no external validation against actual decision quality.
  • domain assumption Prolific participants recruited as AI developers and compliance experts accurately represent those professional groups.
    Cohorts were filtered by self-reported role and AI use frequency, not by employer verification.
  • domain assumption The EU AI Act risk classifications shown on the cards (e.g., biometric checkout is high risk) are correct.
    The cards' risk summary bar depends on the authors' reading of the Act; wrong classifications would mislead users.
  • domain assumption US Census matching on age, sex, and race makes the ordinary-individual sample representative.
    The matching is limited to three demographics and excludes education and Hispanic/Latino as a co-equal category, as the authors acknowledge in Section 6.3.
  • domain assumption Attention checks and pasting restrictions ensure the online responses are genuine.
    Standard crowdwork quality controls used as described in Section 5.3.
invented entities (1)
  • Impact Assessment Card template independent evidence
    purpose: Communicate AI system risks and benefits to technical and non-technical stakeholders.
    The template is publicly available online, so others can test and reuse it; the card's claim to improve communication is falsifiable in independent user studies.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Impact Assessment Card: Communicating Risks and Benefits of AI Uses." pith.science (2026). https://pith.science/paper/3HACTWNQ

@misc{pith2026250818919,
  author       = {Pith},
  title        = {Pith review of: Impact Assessment Card: Communicating Risks and Benefits of AI Uses},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3HACTWNQ}},
  note         = {Machine review of arXiv:2508.18919}
}
read the original abstract

Communicating the risks and benefits of AI is important for regulation and public understanding. Yet current methods such as technical reports often exclude people without technical expertise. Drawing on HCI research, we developed an Impact Assessment Card to present this information more clearly. We held three focus groups with a total of 12 participants who helped identify design requirements and create early versions of the card. We then tested a refined version in an online study with 235 participants, including AI developers, compliance experts, and members of the public selected to reflect the U.S. population by age, sex, and race. Participants used either the card or a full impact assessment report to write an email supporting or opposing a proposed AI system. The card led to faster task completion and higher-quality emails across all groups. We discuss how design choices can improve accessibility and support AI governance. Examples of cards are available at: https://social-dynamics.net/ai-risks/impact-card/.

Figures

Figures reproduced from arXiv: 2508.18919 by the authors.

Figure 1
Figure 1. Template of the Impact Assessment Card showing AI risks and benefits in plain terms. Section A includes [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Speculative impact assessment cards created by 12 participants (P1-12) during three focus groups [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Impact assessment cards: preliminary (A) and refined (B) versions for the biometric checkout system. [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: The online study involved 9 steps. Initially, participants received a brief introduction to the survey [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]
Figure 5
Figure 5. Figure 5: Card outperformed report across all quantitative metrics and cohorts. It helped produce higher quality emails in less time, while being more usable and preferred for the task. 5.4.2 Qualitative Results. Through thematic analysis of participants’ free-form answers, we i…
Figure 6
Figure 6. Figure 6: Fourteen design patterns for visual representation (D1-D9) and layout (D10-D14) to communicate the risks and benefits of AI technologies, derived from the literature review [16, 37, 40, 46, 58, 66, 71, 83, 86, 101]. Also available at: https://social-dynamics.net/ai-ris…
Figure 7
Figure 7. Figure 7: Impact assessment card for a store checkout system using biometric identification, used during the [PITH_FULL_IMAGE:figures/full_fig_p029_7.png]
Figure 8
Figure 8. Figure 8: Impact assessment card for a car park monitoring system using image recognition, used during the [PITH_FULL_IMAGE:figures/full_fig_p030_8.png]
Figure 9
Figure 9. Figure 9: Impact assessment report for a store checkout system using biometric identification, used during the [PITH_FULL_IMAGE:figures/full_fig_p031_9.png]
Figure 10
Figure 10. Figure 10: Impact assessment report for a car park monitoring system using image recognition, used during the [PITH_FULL_IMAGE:figures/full_fig_p032_10.png]
Figure 11
Figure 11. Figure 11: Final impact assessment card for a store checkout system using biometric identification, including [PITH_FULL_IMAGE:figures/full_fig_p039_11.png]
Figure 12
Figure 12. Figure 12: Final impact assessment card for a car park monitoring system using image recognition, including [PITH_FULL_IMAGE:figures/full_fig_p040_12.png]
Figure 13
Figure 13. Figure 13: Final impact assessment card for a music recommender system that suggests songs to platform users [PITH_FULL_IMAGE:figures/full_fig_p041_13.png]
Figure 14
Figure 14. Figure 14: The impact assessment card for a housing benefit allocation assistant for public servants reviews [PITH_FULL_IMAGE:figures/full_fig_p042_14.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

101 extracted references · 36 canonical work pages

  1. [1]

    Philip Achimugu, Ali Selamat, Roliana Ibrahim, and Mohd Naz’ri Mahrin. 2014. A Systematic Literature Review of Software Requirements Prioritization Research. Information and Software Technology 56, 6 (2014), 568–585. doi:10.1016/j.infsof.2014.02.001

  2. [2]

    Ada Lovelace Institute. 2022. Algorithmic Impact Assessment: AIA Template . Retrieved August 15, 2025 from https://www.adalovelaceinstitute.org/resource/aia-template

  3. [3]

    Chappell, and Kristen Kennedy

    Stephanie Ballard, Karen M. Chappell, and Kristen Kennedy. 2019. Judgment Call the Game: Using Value Sensitive Design and Design Fiction to Surface Ethical Concerns Related to Technology. In Proceedings of the ACM Conference on Designing Interactive Systems (DIS) . 421–433. doi:10.1145/3322276.3323697

  4. [4]

    Bender and Batya Friedman

    Emily M. Bender and Batya Friedman. 2018. Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science. Transactions of the Association for Computational Linguistics (ACL) 6 (2018), 587–604. doi:10.1162/tacl_a_00041

  5. [6]

    Edyta Bogucka, Marios Constantinides, Sanja Šćepanović, and Daniele Quercia. 2024. AI Design: A Responsible Artificial Intelligence Framework for Prefilling Impact Assessment Reports. IEEE Internet Computing 28, 5 (2024), 37–45. doi:10.1109/MIC.2024.3451351

  6. [7]

    Edyta Bogucka, Marios Constantinides, Sanja Šćepanović, and Daniele Quercia. 2024. Co-designing an AI Impact Assessment Report Template with AI Practitioners and AI Compliance Experts. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES) 7, 1 (Oct. 2024), 168–180. doi:10.1609/aies.v7i1.31627

  7. [8]

    Edyta Bogucka, Sanja Šćepanović, and Daniele Quercia. 2024. Atlas of AI Risks: Enhancing Public Understanding of AI Risks. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing (HCOMP) 12, 1 (2024), 33–43. doi:10.1609/hcomp.v12i1.31598

  8. [9]

    Vanessa Bracamonte, Sebastian Pape, Sascha Löbner, and Frederic Tronnier. 2023. Effectiveness and Information Quality Perception of an AI Model Card: A Study Among Non-Experts. In Annual International Conference on Privacy, Security and Trust (PST) . IEEE, 1–7

Show all 101 references
  1. [10]

    Virginia Braun and Victoria Clarke. 2006. Using Thematic Analysis in Psychology. Qualitative Research in Psychology 3, 2 (2006), 77–101. doi:10.1191/1478088706qp063oa

  2. [11]

    John Brooke. 1996. SUS: A Quick and Dirty Usability Scale. Usability Evaluation in Industry 189, 3 (1996), 189–194

  3. [12]

    Alexander Buhmann and Christian Fieseler. 2021. Towards a Deliberative Framework for Responsible Innovation in Artificial Intelligence. Technology in Society 64 (2021), 101475. doi:10.1016/j.techsoc.2020.101475

  4. [13]

    Alyxander Burns, Cindy Xiong, Steven Franconeri, Alberto Cairo, and Narges Mahyar. 2020. How To Evaluate Data Visualizations Across Different Levels of Understanding. In IEEE Workshop on Evaluation and Beyond - Methodological Approaches to Visualization (BELIV). IEEE Computer ...

  5. [14]

    Zana Buçinca, Chau Minh Pham, Maurice Jakesch, Marco Tulio Ribeiro, Alexandra Olteanu, and Saleema Amershi

  6. [15]

    Lee, Kristina Rapuano, Kate Niederhoffer, Alex Liebscher, and Jeffrey Hancock

    Myra Cheng, Angela Y. Lee, Kristina Rapuano, Kate Niederhoffer, Alex Liebscher, and Jeffrey Hancock. 2025. From Tools to Thieves: Measuring and Understanding Public Perceptions of AI Through Crowdsourced Metaphors. arXiv:2501.18045

  7. [16]

    Kleinaltenkamp, Jonas Schuett, and Seth D

    Peter Cihon, Moritz J. Kleinaltenkamp, Jonas Schuett, and Seth D. Baum. 2021. AI Certification: Advancing Ethical Practice by Reducing Information Asymmetries. IEEE Transactions on Technology and Society 2, 4 (2021), 200–209. doi:10.1109/tts.2021.3077595

  8. [17]

    Marios Constantinides, Edyta Bogucka, Daniele Quercia, Susanna Kallio, and Mohammad Tahaei. 2024. RAI Guidelines: Method for Generating Responsible AI Guidelines Grounded in Regulations and Usable by (Non-)Technical Roles. Proceedings of the ACM on Human-Computer Interaction C...

  9. [18]

    Credo AI. 2024. AI Vendor Risk Profiles. Retrieved August 15, 2025 from https://www .credo.ai/ai-vendor-directory

  10. [19]

    Andrew Gary Darwin Holmes. 2020. Researcher Positionality - A Consideration of Its Influence and Place in Qualitative Research - A New Researcher Guide. Shanlax International Journal of Education 8, 4 (2020), 1–10. doi:10.34293/education.v8i4.3232

  11. [20]

    Wesley Hanwen Deng, Solon Barocas, and Jennifer Wortman Vaughan. 2025. Supporting Industry Computing Researchers in Assessing, Articulating, and Addressing the Potential Negative Societal Impact of Their Work. Proc. ACM Hum.-Comput. Interact. 9, 2, Article CSCW178 (2025), 37 p...

  12. [21]

    Alicia DeVos, Aditi Dhabalia, Hong Shen, Kenneth Holstein, and Motahhare Eslami. 2022. Toward User-Driven Algorithm Auditing: Investigating Users’ Strategies for Uncovering Harmful Algorithmic Behavior. In Proceedings of the ACM Conference on Human Factors in Computing Systems...

  13. [22]

    David Dickinson and Suzy Gallina. 2017. Information Design: Research and Practice . Routledge, Chapter Information Design in Medicine Package Leaflets. doi:10.4324/9781315585680

  14. [23]

    Berger, Vagner Figueredo De Santana, and Juana Catalina Becerra Sandoval

    Salma Elsayed-Ali, Sara E. Berger, Vagner Figueredo De Santana, and Juana Catalina Becerra Sandoval. 2023. Re- sponsible & Inclusive Cards: An Online Card Tool to Promote Critical Reflection in Technology Industry Work Practices. In Proceedings of the ACM Conference on Human F...

  15. [24]

    Eva Erman and Markus Furendal. 2024. The Democratization of Global AI Governance and the Role of Tech Companies. Nature Machine Intelligence (2024), 1–3. doi:10.1038/s42256-024-00811-z

  16. [25]

    European Comission. 2024. Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence and amending Regulations (EC) No 300/2008, (EU) No 167/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/113...

  17. [26]

    Michael Feffer, Nikolas Martelaro, and Hoda Heidari. 2023. The AI Incident Database as an Educational Tool to Raise Awareness of AI Harms: A Classroom Exploration of Efficacy, Limitations, & Future Improvements. In Proceedings of the ACM Conference on Equity and Access in Algo...

  18. [27]

    O’Halloran, Sabine Tan, and Peter Wignell

    Almudena Fernández-Fontecha, Kay L. O’Halloran, Sabine Tan, and Peter Wignell. 2019. A Multimodal Approach to Visual Thinking: The Scientific Sketchnote. Visual Communication 18, 1 (2019), 5–29. doi:10.1177/1470357218759808

  19. [28]

    Figma. 2016. Figma: The Collaborative Interface Design Tool . Retrieved August 15, 2025 from https://www .figma.com

  20. [29]

    Luciano Floridi, Josh Cowls, Monica Beltrametti, Raja Chatila, Patrice Chazerand, Virginia Dignum, Christoph Luetge, Robert Madelin, Ugo Pagallo, Francesca Rossi, et al. 2018. AI4People—an Ethical Framework for a Good AI Society: Opportunities, Risks, Principles, and Recommend...

  21. [30]

    Riccardo Fogliato, Alexandra Chouldechova, and Zachary Lipton. 2021. The Impact of Algorithmic Risk Assessments on Human Predictions and Its Analysis via Crowdsourcing Studies. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–24. doi:10.1145/3479572

  22. [31]

    Franconeri, Lace M

    Steven L. Franconeri, Lace M. Padilla, Priti Shah, Jeffrey M. Zacks, and Jessica Hullman. 2021. The Science of Visual Data Communication: What Works. Psychological Science in the Public Interest 22, 3 (2021), 110–161. doi:10.1177/ 15291006211051956 PMID: 34907835

  23. [32]

    Hall, Yuriy Brun, and Cindy Xiong Bearfield

    Aimen Gaba, Zhanna Kaufman, Jason Cheung, Marie Shvakel, Kyle Wm. Hall, Yuriy Brun, and Cindy Xiong Bearfield

  24. [34]

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021. Datasheets for Datasets. Commun. ACM 64, 12 (2021), 86–92. doi:10.1145/3458723

  25. [35]

    Pandit, and Dave Lewis

    Delaram Golpayegani, Harshvardhan J. Pandit, and Dave Lewis. 2023. To Be High-Risk, or Not To Be–Semantic Specifications and Implications of the AI Act’s High-Risk AI Applications and Harmonised Standards. In Proceedings of the ACM Conference on Fairness, Accountability, and T...

  26. [36]

    Samantha Goodman, Lana Vanderlee, Rachel Acton, Syed Mahamad, and David Hammond. 2018. The Impact of Front-of-Package Label Design on Consumer Understanding of Nutrient Amounts. Nutrients 10, 11 (2018), 1624. doi:10.3390/nu10111624

  27. [37]

    Matthew Gorton, Barbara Tocco, Ching-Hua Yeh, and Monika Hartmann. 2021. What Determines Consumers’ Use of Eco-Labels? Taking a Close Look at Label Trust. Ecological Economics 189 (2021), 107173. doi:10 .1016/ j.ecolecon.2021.107173

  28. [38]

    Ben Green and Yiling Chen. 2021. Algorithmic Risk Assessments Can Alter Human Decision-Making Processes in High-Stakes Government Contexts. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–33. doi:10.1145/3479562

  29. [39]

    Grunert, Sophie Hieke, and Josephine Wills

    Klaus G. Grunert, Sophie Hieke, and Josephine Wills. 2014. Sustainability Labels on Food Products: Consumer Motivation, Understanding and Use. Food Policy 44 (2014), 177–189. doi:10.1016/j.foodpol.2013.12.001

  30. [40]

    Sarah Holland, Ahmed Hosny, Sarah Newman, Joshua Joseph, and Kasia Chmielinski. 2020. The Dataset Nutrition Label: A Framework to Drive Higher Data Quality Standards. In Data Protection and Privacy , Dara Hallinan, Ronald Leenes, Serge Gutwirth, and Paul De Hert (Eds.). Hart P...

  31. [41]

    Hong, Adam Fourney, Derek DeBellis, and Saleema Amershi

    Matthew K. Hong, Adam Fourney, Derek DeBellis, and Saleema Amershi. 2021. Planning for Natural Language Failures With the AI Playbook. In Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI) . 1–11. doi:10.1145/3411764.3445735

  32. [42]

    Isabelle Hupont, David Fernández-Llorca, Sandra Baldassarri, and Emilia Gómez. 2024. Use Case Cards: A Use Case Reporting Framework Inspired by the European AI Act. Ethics and Information Technology 26, 2 (March 2024). doi:10.1007/s10676-024-09757-7

  33. [43]

    ISO/IEC. 2023. Information Technology – Artificial Intelligence – Management System . Standard ISO/IEC 42001:2023. International Organization for Standardization. https://www .iso.org/standard/81230.html

  34. [44]

    ISO/IEC. 2025. Information Technology – Artificial Intelligence – AI System Impact Assessment . Standard ISO/IEC 42005:2025. International Organization for Standardization. https://www .iso.org/standard/42005

  35. [45]

    Nari Johnson and Hoda Heidari. 2023. Assessing AI Impact Assessments: A Classroom Study. arXiv:2311.11193

  36. [46]

    Alexandra Jones, Bruce Neal, Belinda Reeve, Cliona Ni Mhurchu, and Anne Marie Thow. 2019. Front-of-Pack Nutrition Labelling To Promote Healthier Diets: Current Practice and Opportunities To Strengthen Regulation Worldwide.BMJ Global Health 4, 6 (2019). doi:10.1136/bmjgh-2019-001882

  37. [47]

    Marius Kaminskas, Francesco Ricci, and Markus Schedl. 2013. Location-Aware Music Recommendation Using Auto- Tagging and Hybrid Matching. In Proceedings of the ACM Conference on Recommender Systems (RecSys) (Hong Kong, China) (RecSys ’13). ACM, 17–24. doi:10.1145/2507157.2507180

  38. [48]

    Iqbal, Q

    Anna Kawakami, Shreya Chowdhary, Shamsi T. Iqbal, Q. Vera Liao, Alexandra Olteanu, Jina Suh, and Koustuv Saha. 2023. Sensing Wellbeing in the Workplace, Why and for Whom? Envisioning Impacts With Organizational Stakeholders. Proceedings of the ACM on Human-Computer Interaction...

  39. [49]

    Anna Kawakami, Daricia Wilkinson, and Alexandra Chouldechova. 2024. Do Responsible AI Artifacts Advance Stakeholder Goals? Four Key Barriers Perceived by Legal and Civil Stakeholders. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , Vol. 7. 670–682. doi:1...

  40. [50]

    Johannes Kehrer and Helwig Hauser. 2012. Visualization and Visual Analysis of Multifaceted Scientific Data: A Survey. IEEE Transactions on Visualization and Computer Graphics (TVCG) 19, 3 (2012), 495–513. doi:10.1109/TVCG.2012.110

  41. [51]

    Nutrition Label

    Patrick Gage Kelley, Joanna Bresee, Lorrie Faith Cranor, and Robert W. Reeder. 2009. A "Nutrition Label" for Privacy. In Proceedings of the Symposium on Usable Privacy and Security (SOUPS) . ACM, Article 4, 12 pages. doi:10 .1145/ 1572532.1572538

  42. [52]

    Hancock, and Michael S

    Pranav Khadpe, Ranjay Krishna, Li Fei-Fei, Jeffrey T. Hancock, and Michael S. Bernstein. 2020. Conceptual Metaphors Impact Perceptions of Human-AI Collaboration. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2 (2020), 1–26. doi:10.1145/3415234

  43. [53]

    Sandjar Kozubaev, Chris Elsden, Noura Howell, Marie Louise Juul Søndergaard, Nick Merrill, Britta Schulte, and Richmond Y. Wong. 2020. Expanding Modes of Reflection in Design Futuring. In Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI) . 1–15. doi...

  44. [54]

    Jobst Landgrebe and Barry Smith. 2022. Why Machines Will Never Rule the World: Artificial Intelligence Without Fear . Routledge

  45. [55]

    Tianshi Li, Lorrie Faith Cranor, Yuvraj Agarwal, and Jason I. Hong. 2024. Matcha: An IDE Plugin for Creating Accurate Privacy Nutrition Labels. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. (UbiComp) 8, 1, Article 33 (2024), 38 pages. doi:10.1145/3643544

  46. [56]

    Weixin Liang, Nazneen Rajani, Xinyu Yang, Ezinwanne Ozoani, Eric Wu, Yiqun Chen, Daniel Scott Smith, and James Zou. 2024. What’s Documented in AI? Systematic Analysis of 32K AI Model Cards. arXiv:2402.05160

  47. [57]

    Vera Liao and S

    Q. Vera Liao and S. Shyam Sundar. 2022. Designing for Responsible Trust in AI Systems: A Communication Perspective. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT) . 1257–1268. doi:10.1145/3531146.3533182

  48. [58]

    Joseph Lindley, Haider Ali Akmal, Franziska Pilling, and Paul Coulton. 2020. Researching AI Legibility Through Design. In Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI). 1–13. doi:10.1145/3313831.3376792

  49. [59]

    Ewa Luger, Lachlan Urquhart, Tom Rodden, and Michael Golembewski. 2015. Playing the Legal Card: Using Ideation Cards to Raise Data Protection Issues within the Design Process. In Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI) . ACM, 457–466. doi:...

  50. [60]

    Michal, and Cindy Xiong

    Prateek Mantri, Hariharan Subramonyam, Audrey L. Michal, and Cindy Xiong. 2023. How Do Viewers Synthesize Conflicting Information from Data Visualizations? IEEE Transactions on Visualization & Computer Graphics (TVCG) 29, 01 (2023), 1005–1015. doi:10.1109/TVCG.2022.3209467

  51. [61]

    Nora McDonald, Sarita Schoenebeck, and Andrea Forte. 2019. Reliability and Inter-Rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice. Proceedings of the ACM on Human-Computer Interaction 3, CSCW (2019). doi:10.1145/3359174

  52. [62]

    Jacob Metcalf, Emanuel Moss, Elizabeth Anne Watkins, Ranjit Singh, and Madeleine Clare Elish. 2021. Algorithmic Impact Assessments and Accountability: The Co-Construction of Impacts. In Proceedings of the ACM Conference on Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Artic...

  53. [63]

    Microsoft. 2022. Microsoft Responsible AI Impact Assessment Template . Retrieved August 15, 2025 from https: //blogs.microsoft.com/wp-content/uploads/prod/sites/5/2022/06/Microsoft-RAI-Impact-Assessment-Template.pdf

  54. [64]

    Matthew Miles and Michael Huberman. 1994. Qualitative Data Analysis: A Methods Sourcebook . Sage

  55. [65]

    Raji, and Timnit Gebru

    Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa D. Raji, and Timnit Gebru. 2019. Model Cards for Model Reporting. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT). 2...

  56. [66]

    Morik, Helena Kotthaus, Raphael Fischer, Sascha Mücke, Matthias Jakobs, Nico Piatkowski, Andreas Pauly, Lukas Heppe, and Danny Heinrich

    Katharina J. Morik, Helena Kotthaus, Raphael Fischer, Sascha Mücke, Matthias Jakobs, Nico Piatkowski, Andreas Pauly, Lukas Heppe, and Danny Heinrich. 2022. Yes We Care!-Certification for Machine Learning Methods Through the Care Label Framework. Frontiers in Artificial Intelli...

  57. [67]

    Muller and Sarah Kuhn

    Michael J. Muller and Sarah Kuhn. 1993. Participatory Design. Commun. ACM 36, 6 (1993), 24–28

  58. [68]

    National Institute of Standards and Technology. 2023. The EqualAI Algorithmic Impact Assessment Tool . Retrieved August 15, 2025 from https://www.equalai.org/aia

  59. [69]

    Claudio Novelli, Federico Casolari, Antonino Rotolo, Mariarosaria Taddeo, and Luciano Floridi. 2023. Taking AI Risks Seriously: A New Assessment Model for the AI Act. AI & Society (2023), 1–5. doi:10.1007/s00146-023-01723-z

  60. [70]

    2024.Revisions to OMB’s Statistical Policy Directive No

    Office of Management and Budget of the United States Government. 2024.Revisions to OMB’s Statistical Policy Directive No. 15: Standards for Maintaining, Collecting, and Presenting Federal Data on Race and Ethnicity . Retrieved August 15, 2025 from https://www.federalregister.g...

  61. [71]

    Victor Ojewale, Ryan Steed, Briana Vecchione, Abeba Birhane, and Inioluwa Deborah Raji. 2024. Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling. arXiv:2402.17861

  62. [72]

    Open Ethics. 2023. Open Ethics Label: AI Nutrition Labels . Retrieved August 15, 2025 from https://openethics .ai/label

  63. [73]

    Vinodkumar Prabhakaran, Margaret Mitchell, Timnit Gebru, and Iason Gabriel. 2022. A Human Rights-Based Approach to Responsible AI. arXiv:2210.02667

  64. [74]

    Prolific. 2014. Prolific: Quickly Find Research Participants You Can Trust . Retrieved August 15, 2025 from https: //www.prolific.com

  65. [75]

    Ghulam Jilani Quadri, Arran Zeyu Wang, Zhehao Wang, Jennifer Adorno, Paul Rosen, and Danielle Albers Szafir

  66. [76]

    Malak Sadek, Marios Constantinides, Daniele Quercia, and Celine Mougenot. 2024. Guidelines for Integrating Value Sensitive Design in Responsible AI Toolkits. Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI), Article 472 (2024), 20 pages. doi:10.114...

  67. [77]

    Johnny Saldaña. 2015. The Coding Manual for Qualitative Researchers . Sage

  68. [79]

    Kühne, Léane Wettstein, and Florian Brühlmann

    Nicolas Scharowski, Michaela Benk, Swen J. Kühne, Léane Wettstein, and Florian Brühlmann. 2023. Certification Labels for Trustworthy AI: Insights From an Empirical Mixed-Method Study. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT). 2...

  69. [80]

    Hong Shen, Alicia DeVos, Motahhare Eslami, and Kenneth Holstein. 2021. Everyday Algorithm Auditing: Under- standing the Power of Everyday Users in Surfacing Harmful Algorithmic Behaviors. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2, Article 433 (2021), 29 pag...

  70. [81]

    Jeff Sauro. 2011. A Practical Guide to the System Usability Scale: Background, Benchmarks & Best Practices . Measuring Usability LLC

  71. [82]

    Eli Sherman and Ian Eisenberg. 2024. AI Risk Profiles: A Standards Proposal for Pre-deployment AI Risk Disclosures. 23047–23052 pages. doi:10.1609/aaai.v38i21.30348

  72. [83]

    Marcel Stadelmann and Renate Schubert. 2018. How Do Different Designs of Energy Labels Influence Purchases of Household Appliances? A Field Study in Switzerland. Ecological Economics 144 (2018), 112–123. doi:10 .1016/ j.ecolecon.2017.07.031

  73. [84]

    Hong Shen, Haojian Jin, Ángel Alexander Cabrera, Adam Perer, Haiyi Zhu, and Jason I Hong. 2020. Designing Alternative Representations of Confusion Matrices To Support Non-Expert Public Understanding of Algorithm Performance. Proceedings of the ACM on Human-Computer Interaction...

  74. [85]

    Ioannis Stavrakakis, Damian Gordon, Brendan Tierney, Anna Becevel, Emma Murphy, Gordana Dodig-Crnkovic, Radu Dobrin, Viola Schiaffonati, Cristina Pereira, Svetlana Tikhonenko, et al. 2021. The Teaching of Computer Ethics on Computer Science and Related Degree Programmes. A Eur...

  75. [86]

    Kees Stuurman and Eric Lachaud. 2021. Regulating AI. A Label To Complete the Newly Proposed Act on Artificial Intelligence. SSRN Electronic Journal (2021). doi:10.2139/ssrn.3963890

  76. [87]

    Bernd Carsten Stahl, Josephina Antoniou, Nitika Bhalla, Laurence Brooks, Philip Jansen, Blerta Lindqvist, Alexey Kirichenko, Samuel Marchal, Rowena Rodrigues, Nicole Santiago, Zuzanna Warso, and David Wright. 2023. A Sys- tematic Review of Artificial Intelligence Impact Assess...

  77. [88]

    The Future of Life Institute. 2024. EU AI Act Compliance Checker . The Future of Life Institute. Retrieved August 15, 2025 from https://artificialintelligenceact.eu/assessment/eu-ai-act-compliance-checker

  78. [89]

    Wilson, John Coveney, Trevor Webb, and Samantha B

    Emma Tonkin, Annabelle M. Wilson, John Coveney, Trevor Webb, and Samantha B. Meyer. 2015. Trust in and Through Labelling–a Systematic Review and Critique.British Food Journal 117, 1 (2015), 318–338. doi:10.1108/BFJ-07-2014-0244

  79. [90]

    Ningjing Tang, Jiayin Zhi, Tzu-Sheng Kuo, Calla Kainaroi, Jeremy J Northup, Kenneth Holstein, Haiyi Zhu, Hoda Heidari, and Hong Shen. 2024. AI Failure Cards: Understanding and Supporting Grassroots Efforts to Mitigate AI Failures in Homeless Services. In Proceedings of the ACM...

  80. [91]

    United Nations. 2023. The 17 Sustainable Development Goals. Retrieved August 15, 2025 from https://sdgs.un.org/goals

  81. [92]

    Scientific United Nations Educational and Cultural Organization. 2023. Ethical Impact Assessment. A Tool of the Recommendation on the Ethics of Artificial Intelligence . Recommendation SHS/REI/BIO/REC-AIETHICS-TOOL- EIA/2023. United Nations Educational, Scientific and Cultural...

  82. [93]

    Twilio Inc. 2023. AI Nutrition Facts. Retrieved August 15, 2025 from https://nutrition-facts .ai

  83. [94]

    Census Bureau

    U.S. Census Bureau. 2022.Population by Age and Sex. Annual Social and Economic Supplement. https://www.census.gov/ library/visualizations/interactive/race-and-ethnicity-in-the-united-state-2010-and-2020-census .html

  84. [95]

    Office of Management and Budget

    U.S. Office of Management and Budget. 1997. Revisions to the Standards for the Classification of Federal Data on Race and Ethnicity. Retrieved August 15, 2025 from https://obamawhitehouse.archives.gov/omb/fedreg_1997standards

  85. [96]

    Census Bureau

    U.S. Census Bureau. 2021. Race and Ethnicity in the United States: 2010 Census and 2020 Cen- sus. https://www .census.gov/library/visualizations/interactive/race-and-ethnicity-in-the-united-state-2010-and- 2020-census.html

  86. [97]

    Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, and William Isaac. 2023. Sociotechnical Safety Evaluation of Generative AI Systems. a...

  87. [98]

    Wogalter, Christopher B

    Michael S. Wogalter, Christopher B. Mayhorn, and Olga A. Zielinska. 2015. Use of Color in Warnings . Cambridge University Press, 377–400

  88. [99]

    Wang, Chinmay Kulkarni, Lauren Wilcox, Michael Terry, and Michael Madaio

    Zijie J. Wang, Chinmay Kulkarni, Lauren Wilcox, Michael Terry, and Michael Madaio. 2024. Farsight: Fostering Responsible AI Awareness During AI Application Prototyping. In Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI) . Article 976, 40 pages. do...

  89. [100]

    Young and Michael S

    Stephen L. Young and Michael S. Wogalter. 1990. Comprehension and Memory of Instruction Manual Warnings: Conspicuous Print and Pictorial Icons. Human Factors 32, 6 (1990), 637–649. doi:10.1177/001872089003200603

  90. [101]

    Latest security patches

    Natalie Zelenka, Nina Di Cara, and Huw Day. 2021. Data Hazard Labels . Retrieved August 15, 2025 from https: //datahazards.com/index.html Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW301. Publication date: November 2025. CSCW301:28 Edyta Bogucka, Marios Constan...

  91. [102]

    Nur Yildirim, Changhoon Oh, Deniz Sayar, Kayla Brand, Supritha Challa, Violet Turri, Nina Crosby Walton, Anna Elise Wong, Jodi Forlizzi, James McCann, and John Zimmerman. 2023. Creating Design Resources to Scaffold the Ideation of AI Concepts. In Proceedings of the ACM Confere...

  92. [2023]

    arXiv:2306.03280

    AHA!: Facilitating AI Impact Assessment by Generating Examples of Harms. arXiv:2306.03280

  93. [2024]

    IEEE Transactions on Visualization and Computer Graphics (TVCG ) 30, 01 (2024), 327–337

    My Model is Unfair, Do People Even Care? Visual Design Affects Trust and Perceived Bias in Machine Learning. IEEE Transactions on Visualization and Computer Graphics (TVCG ) 30, 01 (2024), 327–337. doi:10.1109/ TVCG.2023.3327192

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.