Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper argues that measuring AI's second-order effects—shifts in behavior, work, and society—requires a new evaluation ecosystem built on context specification, field testing, and red teaming.

desk verdict A competent, well-cited position paper that names a real gap in AI evaluation, but the load-bearing feasibility premise—sandbox field tests yielding transferable evidence about long-term societal effects—is asserted, not demonstrated. read the letter →

arxiv 2505.18893 v4 pith:B7W7W6XQ submitted 2025-05-24 cs.CY cs.AI

classification cs.CYcs.AI
keywords AIevaluationbenchmarkingsecond-ordereffectsfieldtestingredteamingcontextualawarenesssociotechnicalsystemsriskassessment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This position paper argues that conventional AI evaluation, built around static single-turn benchmarks run entirely on computers, can only capture first-order effects—whether an output is accurate, toxic, biased, or stereotypical—and cannot measure the second-order effects that matter when AI is embedded in education, finance, healthcare, and employment. Those second-order effects include shifts in user behavior, workforce transformations, societal and cultural change, and long-term downstream risks. The paper contends that credible claims about these effects require a new evaluation ecosystem that specifies context up front, collects contextually informed data through field testing and red teaming, and feeds results back into design and governance. If the paper is right, current benchmark rankings and safety scores systematically overstate what we know about how AI behaves in real deployments.

What carries the argument

The load-bearing distinction is between first-order effects (immediate system outputs) and second-order effects (long-term outcomes and consequences of use), with third-order effects as broader societal changes. The machinery that carries the proposal is the value-sensitive design framework, whose conceptual, empirical, and technical methods organize three ecosystem activities: context specification (theory of change, systematization of real-world concepts, stakeholder engagement), field testing (structured multi-session observation of people using AI in a semi-controlled environment), and red teaming (expert, public, and automated attempts to surface failures and boundary conditions). This combination is what the paper says can capture what materializes when people use AI in context.

What would settle it

Compare the second-order outcomes predicted by a large multi-session field test in a simulated sandbox with outcomes measured in a natural longitudinal deployment of the same AI system; if the sandbox results do not track the deployment results (for example, because participants behave differently when observed or the sample is unrepresentative), the paper's central claim that such field testing is necessary for understanding real-world effects would be falsified.

Watch

Extended reading notes

Core claim

The central claim is that the unit of analysis for real-world AI evaluation must shift from the model to the contextual unit—the complex, adaptive behavior that emerges as people use AI in a setting. Benchmarks answer first-order questions about immediate system output; they abstract away the interdependencies between humans and AI, so they cannot reveal feedback loops, long-tail failures, or gradual declines in performance that only appear over repeated use. The paper proposes that evaluation should be initiated by a theory of change, systematize real-world concepts, engage stakeholders, and then gather empirical and technical evidence through field testing—multi-session observation of hundreds or thousands of human subjects in a simulated sandbox—and through red teaming of adversarial and off-label use. Together these methods supply the contextual awareness needed for downstream interpretation and decision making about second-order effects.

Load-bearing premise

The load-bearing premise is that behavior observed when hundreds or thousands of people use an AI system in a controlled sandbox carries over to real deployment; if sandbox behavior does not transfer to natural contexts, the proposed ecosystem would not actually deliver contextual awareness.

Editorial extensions

If this is right

  • Adoption of the ecosystem would redirect evaluation effort from leaderboard rankings toward pre-deployment and post-deployment studies of how people actually use AI systems.
  • Claims about AI safety, fairness, or societal impact would be expected to cite evidence from context-specific testing, not only static benchmark scores.
  • Model developers would need to treat systematization of real-world concepts and stakeholder input as part of the evaluation lifecycle, not as optional extras.
  • Field testing and red teaming would become standard components of AI risk assessment, especially in high-stakes domains such as education, healthcare, and employment.
  • Evaluation results would be interpreted as hypotheses about real-world effects rather than final verdicts, with continuous feedback loops informing design and governance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: The same argument implies that procurement and regulatory decisions should require deployment-context evidence, not just benchmark compliance, which the paper gestures at but does not fully develop.
  • Inference: If field testing becomes standard, its results will themselves depend on participant sampling and observer effects; the paper's feasibility claim could be tested by comparing sandbox field-test outcomes with longitudinal data from natural deployments.
  • Inference: The ecosystem could be extended to third-order effects—forecasting societal shifts such as labor-market restructuring—though the paper states these are even harder to measure and does not specify methods for them.
  • Inference: A cost-effectiveness comparison between large-scale field testing and cheaper multi-turn or agent-based simulations would be a natural next step; the paper assumes field testing is necessary but does not quantify when it is worth the expense.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This position paper argues that conventional AI evaluation methods, particularly static single-turn benchmarks, are insufficient for understanding the second-order societal effects of AI, and that a new evaluation ecosystem combining context specification, field testing, and red teaming is necessary. The paper grounds its proposal in value-sensitive design, outlines concrete activities for establishing contextual awareness and collecting contextually informed data, and suggests governance structures such as testing hubs. It explicitly acknowledges limitations, including deferred methods for analyzing contextual information and the unresolved step from evaluation outcomes to societal insights.

Significance. If the proposed ecosystem proves feasible, it would address a significant gap in AI evaluation practice and policy. The paper is a useful interdisciplinary synthesis, drawing on metrology, social science, and red teaming literatures, and it offers a concrete, actionable roadmap rather than a vague call for more contextual evaluation. Its honesty about open problems is a strength. However, the central necessity claim is not demonstrated: the paper provides no empirical evidence that field testing in sandboxes transfers to real-world settings, no comparison with less costly alternatives, and it explicitly defers the analysis methods needed to convert collected data into societal insights. The significance is therefore conditional on feasibility premises that the paper asserts but does not support.

major comments (3)
  1. [Section 3.2.1, third paragraph] The load-bearing feasibility premise is stated without support: field testing in a simulated sandbox with hundreds or thousands of subjects is claimed to 'enable the collection of real world evidence about what materializes when certain AI features are deployed to the broader public.' The paper does not discuss external validity, sample representativeness, or controls for observer and novelty effects, all of which are essential if sandbox observations are to support claims about natural deployment. This premise is central to the necessity argument, so the paper should either provide empirical evidence from pilots (e.g., the NIST ARIA program) or explicitly frame the transferability of sandbox findings as an open hypothesis with a validation plan.
  2. [Section 3, page 5, and Section 4, page 9] The paper itself states that 'Methods for analyzing contextual information will be the focus of future directions' and that 'evaluation outcomes do not automatically ladder up to societal insights.' These admissions directly weaken the claim that the proposed ecosystem is necessary to understand second-order effects. Without at least a sketch of how annotated dialogues, surveys, and behavior logs will be aggregated and validated into societal-level conclusions, the ecosystem delivers raw data but not the promised understanding. The paper should include a provisional analytical framework or clearly narrow its claim to data collection rather than societal insight.
  3. [Section 1.1 and Section 4] The paper argues that a new ecosystem is 'necessary,' but it does not compare this proposal with less costly or already established alternatives such as longitudinal observational studies, natural experiments, or program evaluation methods. Section 4 concedes that contextual work is 'slow and resource-intensive' yet provides no comparative cost-benefit analysis. To support a necessity claim, the paper should demonstrate that existing methods are insufficient in kind, not merely that benchmarking alone is insufficient. A brief comparison with program evaluation and social science field methods would also clarify what genuinely new infrastructure is required.
minor comments (5)
  1. [Section 3.2.1] The term 'simulated sandbox and reporting environment' is introduced without a clear definition; consider adding a footnote or a more explicit description of the technical and institutional setup.
  2. [Table 1] The 'Testing & Evaluation' row lists effects as '1st, 2nd' and describes performance 'in silico in vitro and in situ,' but the column header is 'Type of Effects (order).' Clarify the distinction between first- and second-order effects in the table to avoid confusing measurement of immediate outputs with longer-term impacts.
  3. [Section 4] The statement that 'fielding qualitative research surveys and conducting ethnographies are both more expensive and time-consuming than using "found data"' would benefit from a citation or a brief example to make the cost comparison concrete.
  4. [References] Reference [3] is incomplete (ending at 'Which humans?'), and reference [102] has a duplicated 'https://' in the URL; these should be corrected.
  5. [Section 1.1 and Figure 1] Figure 1 is not referenced in the text; add a pointer to it where the interdisciplinary community is discussed so that readers understand its intended role.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the paper makes no fitted or derived predictions; the central argument is a position claim, and the self-citations are illustrative rather than load-bearing.

full rationale

This is a position paper arguing that static in silico benchmarking is insufficient for understanding AI's second-order effects and that a new ecosystem of context specification, field testing, and red teaming is needed. There is no fitted parameter, no equation, and no prediction that reduces to an input; the central claim is not definitionally forced, because 'second-order effects' are defined as real-world consequences and the paper argues, rather than assumes, that only in-context methods can observe them. The main self-citations occur in Section 3.2.1, where field testing is described as 'relatively nascent, with recent work in the field of AI risk assessment [83, 84]'—references that are the authors' own NIST ARIA reports—and in Sections 3.1.3 and 3.2.2, where co-authored works [59] and [47] support stakeholder engagement and red-teaming adaptation. These citations supply examples and prior-work context, but the necessity argument does not reduce to them: benchmarking's limitations are supported by independent literature, and the paper itself flags its two largest gaps: 'Methods for analyzing contextual information will be the focus of future directions' (Section 3, page 5) and 'evaluation outcomes do not automatically ladder up to societal insights' (Section 4). The feasibility premise—that sandbox field tests with hundreds or thousands of subjects produce transferable real-world evidence—is asserted rather than demonstrated, but that is a correctness or evidence gap, not circularity, because the claim is explicitly conditional and the analysis methods are deferred. No circular step is established; the score of 2 reflects only the minor, non-load-bearing self-citations.

Assumptions & free parameters 0 free parameters · 3 assumptions · 1 invented entities

The central claim rests on three domain assumptions: benchmarks are insufficient for second-order effects, field testing transfers to real-world settings, and value-sensitive design can be scaled to AI evaluation. No free parameters or fitted values are involved. The only proposed entity is an organizational ecosystem rather than a physical or formal construct, and it has no independent evidence pending implementation.

assumptions (3)
  • domain assumption Static, single-turn in silico benchmarks cannot adequately capture second-order effects.
    Central premise of the paper, argued in Section 2.2 with cited critiques, but not empirically demonstrated in this paper.
  • domain assumption Large-scale field testing with human subjects yields valid evidence about real-world second-order effects.
    Assumed in Section 3.2.1; supported mainly by the authors' NIST ARIA planning documents [83,84], with no independent validity study presented here.
  • domain assumption Value-sensitive design's conceptual, empirical, and technical methods can be applied to AI evaluation at ecosystem scale.
    Used to organize the recommendations in Section 3; VSD is established in design contexts, but its transfer to large-scale AI evaluation is assumed rather than shown.
invented entities (1)
  • Real-world AI evaluation ecosystem
    purpose: Proposed community and testing infrastructure to measure second-order effects through testing hubs, field testing, and red teaming.
    Described in Section 4 as a future infrastructure; no pilot, prototype, or empirical evidence is provided that it will deliver the claimed insights.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects." pith.science (2026). https://pith.science/paper/B7W7W6XQ

@misc{pith2026250518893,
  author       = {Pith},
  title        = {Pith review of: Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/B7W7W6XQ}},
  note         = {Machine review of arXiv:2505.18893}
}
read the original abstract

Conventional AI evaluation approaches concentrated within the AI stack exhibit systemic limitations for exploring, navigating and resolving the human and societal factors that play out in real world deployment such as in education, finance, healthcare, and employment sectors. AI capability evaluations can capture detail about first-order effects, such as whether immediate system outputs are accurate, or contain toxic, biased or stereotypical content, but AI's second-order effects, i.e. any long-term outcomes and consequences that may result from AI use in the real world, have become a significant area of interest as the technology becomes embedded in our daily lives. These secondary effects can include shifts in user behavior, societal, cultural and economic ramifications, workforce transformations, and long-term downstream impacts that may result from a broad and growing set of risks. This position paper argues that measuring the indirect and secondary effects of AI will require expansion beyond static, single-turn approaches conducted in silico to include testing paradigms that can capture what actually materializes when people use AI technology in context. Specifically, we describe the need for data and methods that can facilitate contextual awareness and enable downstream interpretation and decision making about AI's secondary effects, and recommend requirements for a new ecosystem.

Figures

Figures reproduced from arXiv: 2505.18893 by the authors.

Figure 1
Figure 1. Disciplines at the intersection of Real World AI [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Foreign Policy AI Evaluation Gap

    cs.CY 2026-07 conditional novelty 6.0 of 10

    Public technical AI governance almost never evaluates real foreign-policy AI workflows; the paper maps that gap and proposes task-scoped, human-recombined evaluation instead of model leaderboards.

Reference graph

Works this paper leans on

111 extracted references · 33 canonical work pages · cited by 1 Pith paper

  1. [1]

    A collaborative, human-centred taxonomy of ai, algorithmic, and automation harms

    Gavin Abercrombie, Djalel Benbouzid, Paolo Giudici, Delaram Golpayegani, Julio Hernandez, Pierre Noro, Harshvardhan Pandit, Eva Paraschou, Charlie Pownall, Jyoti Prajapati, et al. A collaborative, human-centred taxonomy of ai, algorithmic, and automation harms. arXiv preprint arXiv:2407.01294, 2024

  2. [2]

    The simple macroeconomics of ai

    Daron Acemoglu. The simple macroeconomics of ai. Working Paper 32487, National Bureau of Economic Research, May 2024. URL http://www.nber.org/papers/w32487

  3. [3]

    Xue, Peter S

    Mohammad Atari, Mona J. Xue, Peter S. Park, Damián E. Blasi, and Joseph Henrich. Which humans?

  4. [4]

    Aïvodji et al

    U. Aïvodji et al. Fairwashing: the risk of rationalization. In Proceedings of the International Conference on Machine Learning, pages 161–170, 2019. doi: 10.48550/arXiv.1901.09749. URL https://doi.org/10.48550/arXiv.1901.09749

  5. [5]

    Bai et al

    Y . Bai et al. Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022. URL https://arxiv.org/abs/2204.05862

  6. [6]

    Incentives and the replication crisis in social sciences: A critical review of open science practices

    Loukas Balafoutas, Jeremy Celse, Alexandros Karakostas, and Nicholas Umashev. Incentives and the replication crisis in social sciences: A critical review of open science practices. Journal of Behavioral and Experimental Economics, 114:102327, 2025. ISSN 2214-8043. doi: https://doi.org/10.1016/j.socec.2024.102327. URL https://www.sciencedirect.com/ science...

  7. [7]

    On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623, 2021

    Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623, 2021

  8. [8]

    Participatory ai for humani- tarian innovation: A briefing paper

    Aleks Berditchevskaia, Eirini Malliaraki, and Kathy Peach. Participatory ai for humani- tarian innovation: A briefing paper. Nesta, 2021. URL https://media.nesta.org.uk/ documents/Nesta_Participatory_AI_for_humanitarian_innovation_Final.pdf

Show all 111 references
  1. [9]

    Bertrand and S

    M. Bertrand and S. Mullainathan. Are emily and greg more employable than lakisha and jamal? a field experiment on labor market discrimination. The American Economic Review, 94 (4):991–1013, 2004. URL https://www.jstor.org/stable/3592802

  2. [10]

    A metrological framework for uncertainty evaluation in machine learning classification models, May 2025

    Samuel Bilson, Maurice Cox, Anna Pustogvar, and Andrew Thompson. A metrological framework for uncertainty evaluation in machine learning classification models, May 2025. URL http://arxiv.org/abs/2504.03359. arXiv:2504.03359 [cs]

  3. [12]

    Blaire et al

    B. Blaire et al. Legal red teaming: A systematic approach to assessing legal risk of gener- ative ai models. https://www.dlapiper.com/-/media/project/dlapiper-tenant/ dlapiper/pdf/dla-piper---white-paper---ai-legal-red-teaming.pdf , 2024

  4. [13]

    Atlas of ai risks: Enhancing public understanding of ai risks

    Edyta Bogucka, Sanja Š´cepanovi´c, and Daniele Quercia. Atlas of ai risks: Enhancing public understanding of ai risks. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, volume 12, pages 33–43, 2024

  5. [14]

    Bommasani, P

    R. Bommasani, P. Liang, and T. Lee. Holistic evaluation of language models. Annals of the New York Academy of Sciences, 1525(1):140–146, 2023. doi: 10.1111/nyas.15007. URL https://nyaspubs.onlinelibrary.wiley.com/doi/abs/10.1111/nyas.15007. 10

  6. [15]

    The impact of ai on the workforce: Tasks versus jobs? Economics Letters, 244:111971, 2024

    Kathryn Bonney, Cory Breaux, Catherine Buffington, Emin Dinlersoz, Lucia Foster, Nathan Goldschlag, John Haltiwanger, Zachary Kroff, and Keith Savage. The impact of ai on the workforce: Tasks versus jobs? Economics Letters, 244:111971, 2024. ISSN 0165-1765. doi: https://doi.or...

  7. [16]

    Overcoming failures of imagination in ai infused system development and deployment, 2020

    Margarita Boyarskaya, Alexandra Olteanu, and Kate Crawford. Overcoming failures of imagination in ai infused system development and deployment, 2020. URL https://arxiv. org/abs/2011.13416

  8. [17]

    Evaluating the replicability of social science experiments in nature and science between 2010 and 2015

    Colin Camerer, Anna Dreber, Felix Holzmeister, Teck Ho, Jürgen Huber, Magnus Johannesson, Michael Kirchler, Gideon Nave, Brian Nosek, Thomas Pfeiffer, Adam Altmejd, Nick Buttrick, Taizan Chan, Yiling Chen, Eskil Forsell, Anup Gampa, Emma Heikensten, Lily Hummer, Taisuke Imai, ...

  9. [18]

    Carlini et al

    N. Carlini et al. Extracting training data from large language models. In 30th USENIX Security Symposium, 2021. URL https://www.usenix.org/conference/usenixsecurity21/ presentation/carlini-extracting

  10. [19]

    Quantifying memorization across neural language models, 2023

    Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. Quantifying memorization across neural language models, 2023. URL https://arxiv.org/abs/2202.07646

  11. [20]

    Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Itay Yona, Eric Wallace, David Rolnick, and Florian Tramèr

    Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke, Jonathan Hayase, A. Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Itay Yona, Eric Wallace, David Rolnick, and Florian Tramèr. Stealing part of a production language model, ...

  12. [21]

    Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Bıyık, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell

    Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, Tony Wang, Samuel Marks, Charbel-Raphaël Segerie, Micah Carroll, Andi Peng, Phillip Christoffersen, Mehul Damani, Stew...

  13. [22]

    D. Centola. The network science of collective intelligence. Trends in Cognitive Sciences, 26 (11):923–941, 2022. URL https://pubmed.ncbi.nlm.nih.gov/36180361/

  14. [23]

    A survey on evaluation of large language models

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. A survey on evaluation of large language models. ACM transactions on intelligent systems and technology, 15(3):1–45, 2024

  15. [24]

    Pappas, and Eric Wong

    Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries, 2024. URL https://arxiv.org/abs/2310.08419

  16. [25]

    From representational harms to quality-of-service harms: A case study on llama 2 safety safeguards

    Khaoula Chehbouni, Megha Roshan, Emmanuel Ma, Futian Wei, Afaf Taik, Jackie Chi Kit Cheung, and Golnoosh Farnadi. From representational harms to quality-of-service harms: A case study on llama 2 safety safeguards. In Findings of the Association for Computational Linguistics AC...

  17. [26]

    Feder Cooper, Emily Corvi, P

    Alexandra Chouldechova, Chad Atalla, Solon Barocas, A. Feder Cooper, Emily Corvi, P. Alex Dow, Jean Garcia-Gathright, Nicholas Pangakis, Stefanie Reed, Emily Sheng, Dan Vann, Matthew V ogel, Hannah Washington, and Hanna Wallach. A shared standard for valid measurement of gener...

  18. [27]

    D’Amour et al

    A. D’Amour et al. Underspecification presents challenges for credibility in modern machine learning. Journal of Machine Learning Research, 23(226):1–61, 2022. doi: 10.48550/arXiv. 2011.03395. URL https://doi.org/10.48550/arXiv.2011.03395. 11

  19. [28]

    Cedric E. Dawkins. The principle of good faith: Toward substantive stakeholder engagement. J Bus Ethics, 121:283–295, 2014. doi: 10.1007/s10551-013-1697-z

  20. [29]

    Dobbe, T

    R. Dobbe, T. K. Gilbert, and Y . Mintz. Hard choices in artificial intelligence. Artificial Intelligence, 300:103555, 2021. doi: 10.48550/arXiv.2106.11022. URL https://arxiv. org/abs/2106.11022

  21. [30]

    Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. Show your work: Improved reporting of experimental results, 2019. URL https://arxiv.org/abs/ 1909.03004

  22. [31]

    Harmful speech detection by language models exhibits gender-queer dialect bias

    Rebecca Dorn, Lee Kezar, Fred Morstatter, and Kristina Lerman. Harmful speech detection by language models exhibits gender-queer dialect bias. InProceedings of the 4th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, pages 1–12, 2024

  23. [32]

    Duan et al

    M. Duan et al. Do membership inference attacks work on large language models? arXiv preprint arXiv:2402.07841, 2024. Feb 12

  24. [33]

    On the impossible safety of large ai models, 2023

    El-Mahdi El-Mhamdi, Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Lê-Nguyên Hoang, Rafael Pinot, Sébastien Rouault, and John Stephan. On the impossible safety of large ai models, 2023. URL https://arxiv.org/abs/2209.15259

  25. [34]

    Ethics Owners: A New Model of Organi- zational Responsibility in Data-Driven Technology Companies

    Moss Emanuel and Metcalf Jacob. Ethics Owners: A New Model of Organi- zational Responsibility in Data-Driven Technology Companies . Data & Society, September 2020. URL https://datasociety.net/wp-content/uploads/2020/09/ Ethics-Owners_20200923-DataSociety.pdf

  26. [35]

    Can we trust ai benchmarks? an interdisciplinary review of current issues in ai evaluation, 2025

    Maria Eriksson, Erasmo Purificato, Arman Noroozian, Joao Vinagre, Guillaume Chaslot, Emilia Gomez, and David Fernandez-Llorca. Can we trust ai benchmarks? an interdisciplinary review of current issues in ai evaluation, 2025. URL https://arxiv.org/abs/2502. 06559

  27. [36]

    Code-ifying the law: How disciplinary divides afflict the development of legal software

    Nel Escher, Jeffrey Bilik, Nikola Banovic, and Ben Green. Code-ifying the law: How disciplinary divides afflict the development of legal software. Proc. ACM Hum.-Comput. Interact., 8(CSCW2), November 2024. doi: 10.1145/3686937. URL https://doi.org/10. 1145/3686937

  28. [37]

    The replication crisis has led to positive structural, procedural, and community changes

    Max Korbmacher et al. The replication crisis has led to positive structural, procedural, and community changes. Communications Psychology, 2023. doi: 10.1038/s44271-023-00003-2

  29. [38]

    Large ai models are cultural and social technologies

    Henry Farrell, Alison Gopnik, Cosma Shalizi, and James Evans. Large ai models are cultural and social technologies. Science, 387(6739):1153–1156, 2025. doi: 10.1126/science.adt9819. URL https://www.science.org/doi/abs/10.1126/science.adt9819

  30. [39]

    Value sensitive design: Theory and methods

    Batya Friedman, Peter Kahn, and Alan Borning. Value sensitive design: Theory and methods. UW CSE Technical Report, 2003

  31. [40]

    Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning, October 2016

    Yarin Gal and Zoubin Ghahramani. Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning, October 2016. URL http://arxiv.org/abs/1506. 02142. arXiv:1506.02142 [stat]

  32. [41]

    Americans use ai in everyday prod- ucts without realizing it

    Gallup and Telescope Foundation. Americans use ai in everyday prod- ucts without realizing it. URL https://www.telescopegp.com/insights/ americans-use-ai-in-everyday-products-without-realizing-it

  33. [42]

    Gao et al

    L. Gao et al. A framework for few-shot language model evaluation. https://github.com/ EleutherAI/lm-evaluation-harness , 2021

  34. [43]

    Real- toxicityprompts: Evaluating neural toxic degeneration in language models

    Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. Real- toxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462, 2020. 12

  35. [44]

    Trust, attitudes and use of artificial intelligence: A global study 2025, 2025

    Nicole Gillespie, Steven Lockey, Tabi Ward, Alexandria Macdade, and Gerard Hassed. Trust, attitudes and use of artificial intelligence: A global study 2025, 2025. URL https://figshare.unimelb.edu.au/articles/report/Trust_attitudes_and_ use_of_artificial_intelligence_A_global_s...

  36. [45]

    Greshake et al

    K. Greshake et al. Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In 16th ACM Workshop on Artificial Intelligence and Security, pages 79–90, 2023. URL https://arxiv.org/abs/2302.12173

  37. [46]

    Troy, Dario Amodei, Jared Kaplan, Jack Clark, and Deep Ganguli

    Kunal Handa, Alex Tamkin, Miles McCain, Saffron Huang, Esin Durmus, Sarah Heck, Jared Mueller, Jerry Hong, Stuart Ritchie, Tim Belonax, Kevin K. Troy, Dario Amodei, Jared Kaplan, Jack Clark, and Deep Ganguli. Which economic tasks are performed with ai? evidence from millions o...

  38. [47]

    Hoffmann and H

    M. Hoffmann and H. Frase. Adding structure to ai harm: An introduc- tion to cset’s ai harm framework. https://cset.georgetown.edu/publication/ adding-structure-to-ai-harm/ , 2023

  39. [48]

    Anna A. Ivanova. Toward best research practices in ai psychology, 2024. URL https: //arxiv.org/abs/2312.01276

  40. [49]

    Jansma, Anne M

    Sikke R. Jansma, Anne M. Dijkstra, and Menno D.T. de Jong. Co-creation in support of responsible research and innovation: an analysis of three stakeholder workshops on nanotechnology for health. Journal of Responsible Innovation , 9(1):28–48, 2021. doi: 10.1080/23299460.2021.1994195

  41. [50]

    Survey of hallucination in natural language generation

    Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM computing surveys, 55(12):1–38, 2023

  42. [51]

    Kapoor and A

    S. Kapoor and A. Narayanan. Openai’s policies hinder reproducible research on language mod- els. https://www.aisnakeoil.com/p/openais-policies-hinder-reproducible , 2023

  43. [52]

    Leakage and the reproducibility crisis in ml-based science, 2022

    Sayash Kapoor and Arvind Narayanan. Leakage and the reproducibility crisis in ml-based science, 2022. URL https://arxiv.org/abs/2207.07048

  44. [53]

    Dynabench: Rethinking benchmarking in NLP

    Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adi...

  45. [54]

    Audiocaps: Gener- ating captions for audios in the wild

    Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. Audiocaps: Gener- ating captions for audios in the wild. In NAACL-HLT, 2019

  46. [55]

    The history and risks of reinforcement learning and human feedback, 2023

    Nathan Lambert, Thomas Krendl Gilbert, and Tom Zick. The history and risks of reinforcement learning and human feedback, 2023. URL https://arxiv.org/abs/2310.13595

  47. [56]

    Dempsey, Ece Kamar, Steven M

    Susan Landau, James X. Dempsey, Ece Kamar, Steven M. Bellovin, and Robert Pool. Challenging the machine: Contestability in government ai systems, 2024. URL https: //arxiv.org/abs/2406.10430

  48. [57]

    Leask, Marlene Sandlund, Dawn A

    Calum F. Leask, Marlene Sandlund, Dawn A. Skelton, Teatske M. Altenburg, Greet Cardon, et al. Framework, principles and recommendations for utilising participatory methodologies in the co-creation and evaluation of public health interventions. Res Involv Engagem, 5(2): 1153–11...

  49. [58]

    Lenaerts-Bergmans

    B. Lenaerts-Bergmans. Data poisoning: The exploitation of generative ai. https://www. crowdstrike.com/cybersecurity-101/cyberattacks/data-poisoning/ , 2024

  50. [59]

    Ai sustainability in practice part one: Foundations for sustainable ai projects

    David Leslie, Cami Rincon, Morgan Briggs, et al. Ai sustainability in practice part one: Foundations for sustainable ai projects. https://aiethics.turing.ac.uk/modules/sustainability-1/, 2024

  51. [60]

    Li and J

    C. Li and J. Flanigan. Task contamination: Language models may not be few-shot anymore. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 18471–18480, 2024. doi: 10.48550/arXiv.2312.16337. URL https://doi.org/10.48550/arXiv.2312.16337

  52. [61]

    LLM defenses are not robust to multi-turn human jailbreaks yet

    Nathaniel Li, Ziwen Han, Ian Steneker, Willow Primack, Riley Goodside, Hugh Zhang, Zifan Wang, Cristina Menghini, and Summer Yue. LLM defenses are not robust to multi-turn human jailbreaks yet. arXiv preprint arXiv:2408.15221, 2024

  53. [62]

    Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm- Burger, Rassin Lababidi, Lennart Justen, Andrew B

    Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D. Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm- Burger, Rassin Lababidi, Lennart Justen, Andrew B. Liu, Michael Chen, Isabelle Barrass, Oliver Zhang...

  54. [63]

    A survey of multimodel large language models

    Zijing Liang, Yanjie Xu, Yifan Hong, Penghui Shang, Qi Wang, Qiang Fu, and Ke Liu. A survey of multimodel large language models. In Proceedings of the 3rd International Conference on Computer, Artificial Intelligence and Control Engineering , pages 405–409, 2024

  55. [64]

    Are we learn- ing yet? a meta review of evaluation failures across machine learning

    Thomas Liao, Rohan Taori, Deborah Raji, and Ludwig Schmidt. Are we learn- ing yet? a meta review of evaluation failures across machine learning. In J. Vanschoren and S. Yeung, editors, Proceedings of the Neural Information Pro- cessing Systems Track on Datasets and Benchmarks ...

  56. [65]

    Jaffe, Abhishek Nagaraj, Imke Reimers, Michael D

    Brent Lutes, Joshua Gans, Shane Greenstein, Adam B. Jaffe, Abhishek Nagaraj, Imke Reimers, Michael D. Smith, Rahul Telang, Catherine E. Tucker, and Joel Waldfogel. Identifying the economic implications of artificial intelligence for copyright policy (february 12, 2025). first ...

  57. [66]

    Madaio, Jingya Chen, Hanna Wallach, and Jennifer Wortman Vaughan

    Michael A. Madaio, Jingya Chen, Hanna Wallach, and Jennifer Wortman Vaughan. Tinker, tailor, configure, customize: The articulation work of contextualizing an ai fairness checklist. Proc. ACM Hum.-Comput. Interact., 8(CSCW1), April 2024. doi: 10.1145/3653705. URL https://doi.o...

  58. [67]

    Martínez

    E. Martínez. Re-evaluating GPT-4’s bar exam performance. Artificial Intelligence and Law, 30:1–24, 2024. doi: 10.1177/20539517241290220. URL https://doi.org/10.1177/ 20539517241290220

  59. [68]

    Co-creation for policy: Participatory methodologies to structure multi-stakeholder policymaking processes

    Cristian Matti and Gabriel Rissola. Co-creation for policy: Participatory methodologies to structure multi-stakeholder policymaking processes. https://www.eit.europa.eu/ sites/default/files/jrc128771_01.pdf, 2022

  60. [69]

    How the u.s

    Colleen McClain, Brian Kennedy, Jeffrey Gottfried, Monica Anderson, and Gi- ancarlo Pasquini. How the u.s. public and ai experts view artificial intelli- gence, April 2025. URL https://www.pewresearch.org/internet/2025/04/03/ how-the-us-public-and-ai-experts-view-artificial-in...

  61. [70]

    A. Mislove. Red-teaming large language models to identify novel ai risks. https://www.whitehouse.gov/ostp/news-updates/2023/08/29/ red-teaming-large-language-models-to-identify-novel-ai-risks/ , 2023

  62. [71]

    Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland, May 2022

    Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland, May 2022. Association for Computational Linguistics. URL https: //aclanthology.o...

  63. [72]

    thick alignment

    A. Nelson. Facct’23 keynote: "thick alignment". https://www.youtube.com/watch?v= Sq_XwqVTqvQ, 2023

  64. [73]

    Global law and policy tracker

    International Association of Privacy Professionals. Global law and policy tracker. URL https://iapp.org/resources/article/global-ai-legislation-tracker/

  65. [74]

    Governing with artificial intelli- gence, June 2024

    OECD Artificial Intelligence Papers. Governing with artificial intelli- gence, June 2024. URL https://www.oecd.org/en/publications/ governing-with-artificial-intelligence_26324bc2-en.html

  66. [75]

    Luona Lin Parker and Kim. U.s. workers are more worried than hopeful about future ai use in the workplace, February 2025. URL https://www.pewresearch.org/social-trends/2025/02/25/ u-s-workers-are-more-worried-than-hopeful-about-future-ai-use-in-the-workplace/

  67. [76]

    Concrete problems in ai safety, revisited, 2023

    Inioluwa Deborah Raji and Roel Dobbe. Concrete problems in ai safety, revisited, 2023. URL https://arxiv.org/abs/2401.10899

  68. [77]

    The fallacy of ai functionality

    Inioluwa Deborah Raji, I Elizabeth Kumar, Aaron Horowitz, and Andrew Selbst. The fallacy of ai functionality. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 959–972, 2022

  69. [78]

    Do imagenet classifiers generalize to imagenet?, 2019

    Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to imagenet?, 2019. URL https://arxiv.org/abs/1902.10811

  70. [79]

    Just what do you think you’re doing, dave?’ a checklist for responsible data use in nlp, 2021

    Anna Rogers, Tim Baldwin, and Kobi Leins. Just what do you think you’re doing, dave?’ a checklist for responsible data use in nlp, 2021. URL https://arxiv.org/abs/2109. 06598

  71. [80]

    everyone wants to do the model work, not the data work

    Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. “everyone wants to do the model work, not the data work”: Data cascades in high-stakes ai. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems , C...

  72. [81]

    Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerster, Tim Rocktäschel, and Roberta Raileanu

    Mikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro, Aram H. Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerster, Tim Rocktäschel, and Roberta Raileanu. Rainbow teaming: Open-ended generation of diverse adversarial prompts, 20...

  73. [82]

    Towards a standard for identifying and managing bias in artificial intelligence

    Reva Schwartz, Apostol Vassilev, Kristen Greene, Lori Perine, Andrew Burt, and Patrick Hall. Towards a standard for identifying and managing bias in artificial intelligence. Technical Report NIST SP 1270, National Institute of Standards and Technology, Gaithersburg, MD,

  74. [83]

    The nist assessing risks and impacts of ai (aria) pilot evaluation plan

    Reva Schwartz, Jonathan Fiscus, Kristen Greene, Gabriella Waters, Rumman Chowdhury, Theodore Jensen, Craig Greenberg, Afzal Godil, Razvan Amironesei, Patrick Hall, and Shomik Jain. The nist assessing risks and impacts of ai (aria) pilot evaluation plan. Technical report, Natio...

  75. [84]

    The assessing risks and impacts of ai (aria) program evaluation design document

    Reva Schwartz, Gabriella Waters, Razvan Amironesei, Craig Greenberg, Jon Fiscus, Patrick Hall, Anya Jones, Shomik Jain, Afzal Godil, Kristen Greene, Ted Jensen, and Noah Schulman. The assessing risks and impacts of ai (aria) program evaluation design document. Technical report...

  76. [85]

    A. D. Selbst et al. Fairness and abstraction in sociotechnical systems. In Proceedings of the Conference on Fairness, Accountability, and Transparency , pages 59–68, 2019. doi: 10.1145/3287560.3287598. URL https://doi.org/10.1145/3287560.3287598

  77. [86]

    Sociotech- nical harms of algorithmic systems: Scoping a taxonomy for harm reduction

    Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Rostamzadeh, Paul Nicholas, N’Mah Yilla-Akbari, Jess Gallegos, Andrew Smart, Emilio Garcia, et al. Sociotech- nical harms of algorithmic systems: Scoping a taxonomy for harm reduction. In Proceedings of the 2023 ...

  78. [87]

    Taskbench: Benchmarking large language models for task automation

    Yongliang Shen, Kaitao Song, Xu Tan, Wenqi Zhang, Kan Ren, Siyu Yuan, Weiming Lu, Dongsheng Li, and Yueting Zhuang. Taskbench: Benchmarking large language models for task automation. Advances in Neural Information Processing Systems, 37:4540–4574, 2024

  79. [88]

    Shumailov et al

    I. Shumailov et al. Sponge examples: Energy-latency attacks on neural networks. In IEEE European Symposium on Security and Privacy (EuroS&P) , pages 212–231, 2021. URL https://arxiv.org/abs/2006.03463

  80. [89]

    The leaderboard illusion, 2025

    Shivalika Singh, Yiyang Nan, Alex Wang, Daniel D’Souza, Sayash Kapoor, Ahmet Üstün, Sanmi Koyejo, Yuntian Deng, Shayne Longpre, Noah Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker. The leaderboard illusion, 2025. URL https://arxiv.org/abs/2504. 20879

  81. [90]

    Participation is not a design fix for machine learning

    Mona Sloane, Emanuel Moss, Olaitan Awomolo, and Laura Forlano. Participation is not a design fix for machine learning. EAAMO ’22, New York, NY , USA, 2022. Association for Computing Machinery. ISBN 9781450394772. doi: 10.1145/3551624.3555285. URL https://doi.org/10.1145/355162...

  82. [91]

    Slota, Kenneth R

    Stephen C. Slota, Kenneth R. Fleischmann, Sherri Greenberg, Nitin Verma, Brenna Cummings, Lan Li, and Chris Shenefiel. Many hands make many fingers to point: Challenges in creating accountable ai. AI and Society, 38(4):1287–1299, 2023. doi: 10.1007/s00146-021-01302-0

  83. [92]

    Smith, Saleema Amershi, Solon Barocas, Hanna Wallach, and Jennifer Wort- man Vaughan

    Jessie J. Smith, Saleema Amershi, Solon Barocas, Hanna Wallach, and Jennifer Wort- man Vaughan. Real ml: Recognizing, exploring, and articulating limitations of machine learning research. In 2022 ACM Conference on Fairness Accountability and Transparency, FAccT ’22, page 587–5...

  84. [93]

    Srivastava et al

    A. Srivastava et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models, 2023. URL https://arxiv.org/abs/2206.04615

  85. [94]

    Beyond memorization: Violating privacy via inference with large language models, 2024

    Robin Staab, Mark Vero, Mislav Balunovi ´c, and Martin Vechev. Beyond memorization: Violating privacy via inference with large language models, 2024. URL https://arxiv. org/abs/2310.07298

  86. [95]

    Stadler, T

    C. Stadler, T. Rajwani, and F. Karaba. Solutions to the exploration/exploitation dilemma: networks as a new level of analysis. International Journal of Management Reviews, 16(2): 172–193, 2013. doi: 10.1111/ijmr.12015. URL https://doi.org/10.1111/ijmr.12015

  87. [96]

    The State of AI Governance Research: AI Safety and Reliability in Real World Commercial Deployment

    Ilan Strauss, Isobel Moure, Tim O’Reilly, and Sruly Rosenblat. The State of AI Governance Research: AI Safety and Reliability in Real World Commercial Deployment. AI Disclosures Project, Social Science Research Council, April 2025. doi: 10.35650/aidp.4112.d.2025. URL http://dx...

  88. [97]

    Sumers, Jared Mueller, William McEachen, Wes Mitchell, Shan Carter, Jack Clark, Jared Kaplan, and Deep Ganguli

    Alex Tamkin, Miles McCain, Kunal Handa, Esin Durmus, Liane Lovitt, Ankur Rathi, Saffron Huang, Alfred Mountfield, Jerry Hong, Stuart Ritchie, Michael Stern, Brian Clarke, Landon Goldberg, Theodore R. Sumers, Jared Mueller, William McEachen, Wes Mitchell, Shan Carter, Jack Clar...

  89. [98]

    Theofanos, Y

    M. Theofanos, Y . Choong, and T. Jensen. Ai use taxonomy: A human-centered approach. Technical Report NIST AI 200-1, National Institute of Standards and Technology, Gaithersburg, MD, 2024. URL https://doi.org/10.6028/NIST.AI.200-1

  90. [99]

    Thomas and D

    R. Thomas and D. Uminsky. Reliance on metrics is a fundamental challenge for ai. https: //arxiv.org/pdf/2002.08512, 2020

  91. [100]

    Analytical results for uncertainty propagation through trained ma- chine learning regression models, May 2024

    Andrew Thompson. Analytical results for uncertainty propagation through trained ma- chine learning regression models, May 2024. URL http://arxiv.org/abs/2404.11224. arXiv:2404.11224 [cs]

  92. [101]

    E. L. Trist and K. W. Bamforth. Some social and psychological consequences of the long- wall method of coal-getting: An examination of the psychological situation and defences of a work group in relation to the social structure and technological content of the work system. Hum...

  93. [102]

    Prime minister sets out blueprint to turbocharge ai

    UK Department for Science, Innovation and Technology. Prime minister sets out blueprint to turbocharge ai. https://https://www.gov.uk/government/news/ prime-minister-sets-out-blueprint-to-turbocharge-ai , 2025

  94. [103]

    A Deeper Look into Aleatoric and Epistemic Uncertainty Disentanglement, April 2022

    Matias Valdenegro-Toro and Daniel Saromo. A Deeper Look into Aleatoric and Epistemic Uncertainty Disentanglement, April 2022. URL http://arxiv.org/abs/2204.09308. arXiv:2204.09308 [cs]

  95. [104]

    Feder Cooper, Angelina Wang, Chad Atalla, Solon Baro- cas, Su Lin Blodgett, Alexandra Chouldechova, Emily Corvi, P

    Hanna Wallach, Meera Desai, A. Feder Cooper, Angelina Wang, Chad Atalla, Solon Baro- cas, Su Lin Blodgett, Alexandra Chouldechova, Emily Corvi, P. Alex Dow, Jean Garcia- Gathright, Alexandra Olteanu, Nicholas Pangakis, Stefanie Reed, Emily Sheng, Dan Vann, Jennifer Wortman Vau...

  96. [105]

    Decodingtrust: A comprehensive assess- ment of trustworthiness in gpt models

    Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al. Decodingtrust: A comprehensive assess- ment of trustworthiness in gpt models. In NeurIPS, 2023

  97. [106]

    Calibration in Deep Learning: A Survey of the State-of-the-Art, May 2024

    Cheng Wang. Calibration in Deep Learning: A Survey of the State-of-the-Art, May 2024. URL http://arxiv.org/abs/2308.01222. arXiv:2308.01222 [cs]

  98. [107]

    Sociotechnical safety evaluation of generative ai systems,

    Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, and William Isaac. Sociotechnical safety evaluation of generative ai systems,

  99. [108]

    Roberts, and Mary L

    Alice Qian Zhang, Ryland Shaw, Jacy Reese Anthis, Ashlee Milton, Emily Tseng, Jina Suh, Lama Ahmad, Ram Shankar Siva Kumar, Julian Posada, Benjamin Shestakofsky, Sarah T. Roberts, and Mary L. Gray. The human factor in ai red teaming: Perspectives from social and collaborative ...

  100. [109]

    Felm: Benchmarking factuality evaluation of large language models.Advances in Neural Information Processing Systems, 36:44502–44523, 2023

    Yiran Zhao, Jinghan Zhang, I Chern, Siyang Gao, Pengfei Liu, Junxian He, et al. Felm: Benchmarking factuality evaluation of large language models.Advances in Neural Information Processing Systems, 36:44502–44523, 2023

  101. [110]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. URL https: //arxiv.org/abs/2306.056...

  102. [2022]

    URL https://doi.org/10.6028/NIST.SP.1270

  103. [2023]

    URL https://arxiv.org/abs/2310.11986

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.