REVIEW 3 major objections 5 minor 1 cited by
Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper argues that measuring AI's second-order effects—shifts in behavior, work, and society—requires a new evaluation ecosystem built on context specification, field testing, and red teaming.
desk verdict A competent, well-cited position paper that names a real gap in AI evaluation, but the load-bearing feasibility premise—sandbox field tests yielding transferable evidence about long-term societal effects—is asserted, not demonstrated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing distinction is between first-order effects (immediate system outputs) and second-order effects (long-term outcomes and consequences of use), with third-order effects as broader societal changes. The machinery that carries the proposal is the value-sensitive design framework, whose conceptual, empirical, and technical methods organize three ecosystem activities: context specification (theory of change, systematization of real-world concepts, stakeholder engagement), field testing (structured multi-session observation of people using AI in a semi-controlled environment), and red teaming (expert, public, and automated attempts to surface failures and boundary conditions). This combination is what the paper says can capture what materializes when people use AI in context.
What would settle it
Compare the second-order outcomes predicted by a large multi-session field test in a simulated sandbox with outcomes measured in a natural longitudinal deployment of the same AI system; if the sandbox results do not track the deployment results (for example, because participants behave differently when observed or the sample is unrepresentative), the paper's central claim that such field testing is necessary for understanding real-world effects would be falsified.
Extended reading notes
Core claim
The central claim is that the unit of analysis for real-world AI evaluation must shift from the model to the contextual unit—the complex, adaptive behavior that emerges as people use AI in a setting. Benchmarks answer first-order questions about immediate system output; they abstract away the interdependencies between humans and AI, so they cannot reveal feedback loops, long-tail failures, or gradual declines in performance that only appear over repeated use. The paper proposes that evaluation should be initiated by a theory of change, systematize real-world concepts, engage stakeholders, and then gather empirical and technical evidence through field testing—multi-session observation of hundreds or thousands of human subjects in a simulated sandbox—and through red teaming of adversarial and off-label use. Together these methods supply the contextual awareness needed for downstream interpretation and decision making about second-order effects.
Load-bearing premise
The load-bearing premise is that behavior observed when hundreds or thousands of people use an AI system in a controlled sandbox carries over to real deployment; if sandbox behavior does not transfer to natural contexts, the proposed ecosystem would not actually deliver contextual awareness.
Editorial extensions
If this is right
- Adoption of the ecosystem would redirect evaluation effort from leaderboard rankings toward pre-deployment and post-deployment studies of how people actually use AI systems.
- Claims about AI safety, fairness, or societal impact would be expected to cite evidence from context-specific testing, not only static benchmark scores.
- Model developers would need to treat systematization of real-world concepts and stakeholder input as part of the evaluation lifecycle, not as optional extras.
- Field testing and red teaming would become standard components of AI risk assessment, especially in high-stakes domains such as education, healthcare, and employment.
- Evaluation results would be interpreted as hypotheses about real-world effects rather than final verdicts, with continuous feedback loops informing design and governance.
Reading between the lines
- Inference: The same argument implies that procurement and regulatory decisions should require deployment-context evidence, not just benchmark compliance, which the paper gestures at but does not fully develop.
- Inference: If field testing becomes standard, its results will themselves depend on participant sampling and observer effects; the paper's feasibility claim could be tested by comparing sandbox field-test outcomes with longitudinal data from natural deployments.
- Inference: The ecosystem could be extended to third-order effects—forecasting societal shifts such as labor-market restructuring—though the paper states these are even harder to measure and does not specify methods for them.
- Inference: A cost-effectiveness comparison between large-scale field testing and cheaper multi-turn or agent-based simulations would be a natural next step; the paper assumes field testing is necessary but does not quantify when it is worth the expense.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that conventional AI evaluation methods, particularly static single-turn benchmarks, are insufficient for understanding the second-order societal effects of AI, and that a new evaluation ecosystem combining context specification, field testing, and red teaming is necessary. The paper grounds its proposal in value-sensitive design, outlines concrete activities for establishing contextual awareness and collecting contextually informed data, and suggests governance structures such as testing hubs. It explicitly acknowledges limitations, including deferred methods for analyzing contextual information and the unresolved step from evaluation outcomes to societal insights.
Significance. If the proposed ecosystem proves feasible, it would address a significant gap in AI evaluation practice and policy. The paper is a useful interdisciplinary synthesis, drawing on metrology, social science, and red teaming literatures, and it offers a concrete, actionable roadmap rather than a vague call for more contextual evaluation. Its honesty about open problems is a strength. However, the central necessity claim is not demonstrated: the paper provides no empirical evidence that field testing in sandboxes transfers to real-world settings, no comparison with less costly alternatives, and it explicitly defers the analysis methods needed to convert collected data into societal insights. The significance is therefore conditional on feasibility premises that the paper asserts but does not support.
major comments (3)
- [Section 3.2.1, third paragraph] The load-bearing feasibility premise is stated without support: field testing in a simulated sandbox with hundreds or thousands of subjects is claimed to 'enable the collection of real world evidence about what materializes when certain AI features are deployed to the broader public.' The paper does not discuss external validity, sample representativeness, or controls for observer and novelty effects, all of which are essential if sandbox observations are to support claims about natural deployment. This premise is central to the necessity argument, so the paper should either provide empirical evidence from pilots (e.g., the NIST ARIA program) or explicitly frame the transferability of sandbox findings as an open hypothesis with a validation plan.
- [Section 3, page 5, and Section 4, page 9] The paper itself states that 'Methods for analyzing contextual information will be the focus of future directions' and that 'evaluation outcomes do not automatically ladder up to societal insights.' These admissions directly weaken the claim that the proposed ecosystem is necessary to understand second-order effects. Without at least a sketch of how annotated dialogues, surveys, and behavior logs will be aggregated and validated into societal-level conclusions, the ecosystem delivers raw data but not the promised understanding. The paper should include a provisional analytical framework or clearly narrow its claim to data collection rather than societal insight.
- [Section 1.1 and Section 4] The paper argues that a new ecosystem is 'necessary,' but it does not compare this proposal with less costly or already established alternatives such as longitudinal observational studies, natural experiments, or program evaluation methods. Section 4 concedes that contextual work is 'slow and resource-intensive' yet provides no comparative cost-benefit analysis. To support a necessity claim, the paper should demonstrate that existing methods are insufficient in kind, not merely that benchmarking alone is insufficient. A brief comparison with program evaluation and social science field methods would also clarify what genuinely new infrastructure is required.
minor comments (5)
- [Section 3.2.1] The term 'simulated sandbox and reporting environment' is introduced without a clear definition; consider adding a footnote or a more explicit description of the technical and institutional setup.
- [Table 1] The 'Testing & Evaluation' row lists effects as '1st, 2nd' and describes performance 'in silico in vitro and in situ,' but the column header is 'Type of Effects (order).' Clarify the distinction between first- and second-order effects in the table to avoid confusing measurement of immediate outputs with longer-term impacts.
- [Section 4] The statement that 'fielding qualitative research surveys and conducting ethnographies are both more expensive and time-consuming than using "found data"' would benefit from a citation or a brief example to make the cost comparison concrete.
- [References] Reference [3] is incomplete (ending at 'Which humans?'), and reference [102] has a duplicated 'https://' in the URL; these should be corrected.
- [Section 1.1 and Figure 1] Figure 1 is not referenced in the text; add a pointer to it where the interdisciplinary community is discussed so that readers understand its intended role.
Circularity Check
No significant circularity: the paper makes no fitted or derived predictions; the central argument is a position claim, and the self-citations are illustrative rather than load-bearing.
full rationale
This is a position paper arguing that static in silico benchmarking is insufficient for understanding AI's second-order effects and that a new ecosystem of context specification, field testing, and red teaming is needed. There is no fitted parameter, no equation, and no prediction that reduces to an input; the central claim is not definitionally forced, because 'second-order effects' are defined as real-world consequences and the paper argues, rather than assumes, that only in-context methods can observe them. The main self-citations occur in Section 3.2.1, where field testing is described as 'relatively nascent, with recent work in the field of AI risk assessment [83, 84]'—references that are the authors' own NIST ARIA reports—and in Sections 3.1.3 and 3.2.2, where co-authored works [59] and [47] support stakeholder engagement and red-teaming adaptation. These citations supply examples and prior-work context, but the necessity argument does not reduce to them: benchmarking's limitations are supported by independent literature, and the paper itself flags its two largest gaps: 'Methods for analyzing contextual information will be the focus of future directions' (Section 3, page 5) and 'evaluation outcomes do not automatically ladder up to societal insights' (Section 4). The feasibility premise—that sandbox field tests with hundreds or thousands of subjects produce transferable real-world evidence—is asserted rather than demonstrated, but that is a correctness or evidence gap, not circularity, because the claim is explicitly conditional and the analysis methods are deferred. No circular step is established; the score of 2 reflects only the minor, non-load-bearing self-citations.
Assumptions & free parameters
assumptions (3)
- domain assumption Static, single-turn in silico benchmarks cannot adequately capture second-order effects.
- domain assumption Large-scale field testing with human subjects yields valid evidence about real-world second-order effects.
- domain assumption Value-sensitive design's conceptual, empirical, and technical methods can be applied to AI evaluation at ecosystem scale.
invented entities (1)
-
Real-world AI evaluation ecosystem
Cite this review
Pith. "Pith review of Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects." pith.science (2026). https://pith.science/paper/B7W7W6XQ
@misc{pith2026250518893,
author = {Pith},
title = {Pith review of: Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects},
year = {2026},
howpublished = {\url{https://pith.science/paper/B7W7W6XQ}},
note = {Machine review of arXiv:2505.18893}
}
read the original abstract
Conventional AI evaluation approaches concentrated within the AI stack exhibit systemic limitations for exploring, navigating and resolving the human and societal factors that play out in real world deployment such as in education, finance, healthcare, and employment sectors. AI capability evaluations can capture detail about first-order effects, such as whether immediate system outputs are accurate, or contain toxic, biased or stereotypical content, but AI's second-order effects, i.e. any long-term outcomes and consequences that may result from AI use in the real world, have become a significant area of interest as the technology becomes embedded in our daily lives. These secondary effects can include shifts in user behavior, societal, cultural and economic ramifications, workforce transformations, and long-term downstream impacts that may result from a broad and growing set of risks. This position paper argues that measuring the indirect and secondary effects of AI will require expansion beyond static, single-turn approaches conducted in silico to include testing paradigms that can capture what actually materializes when people use AI technology in context. Specifically, we describe the need for data and methods that can facilitate contextual awareness and enable downstream interpretation and decision making about AI's secondary effects, and recommend requirements for a new ecosystem.
Figures
Forward citations
Cited by 1 Pith paper
-
The Foreign Policy AI Evaluation Gap
Public technical AI governance almost never evaluates real foreign-policy AI workflows; the paper maps that gap and proposes task-scoped, human-recombined evaluation instead of model leaderboards.
Reference graph
Works this paper leans on
-
[1]
A collaborative, human-centred taxonomy of ai, algorithmic, and automation harms
Gavin Abercrombie, Djalel Benbouzid, Paolo Giudici, Delaram Golpayegani, Julio Hernandez, Pierre Noro, Harshvardhan Pandit, Eva Paraschou, Charlie Pownall, Jyoti Prajapati, et al. A collaborative, human-centred taxonomy of ai, algorithmic, and automation harms. arXiv preprint arXiv:2407.01294, 2024
arXiv 2024
-
[2]
The simple macroeconomics of ai
Daron Acemoglu. The simple macroeconomics of ai. Working Paper 32487, National Bureau of Economic Research, May 2024. URL http://www.nber.org/papers/w32487
2024
-
[3]
Xue, Peter S
Mohammad Atari, Mona J. Xue, Peter S. Park, Damián E. Blasi, and Joseph Henrich. Which humans?
-
[4]
U. Aïvodji et al. Fairwashing: the risk of rationalization. In Proceedings of the International Conference on Machine Learning, pages 161–170, 2019. doi: 10.48550/arXiv.1901.09749. URL https://doi.org/10.48550/arXiv.1901.09749
- [5]
-
[6]
Loukas Balafoutas, Jeremy Celse, Alexandros Karakostas, and Nicholas Umashev. Incentives and the replication crisis in social sciences: A critical review of open science practices. Journal of Behavioral and Experimental Economics, 114:102327, 2025. ISSN 2214-8043. doi: https://doi.org/10.1016/j.socec.2024.102327. URL https://www.sciencedirect.com/ science...
-
[7]
On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623, 2021
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623, 2021
2021
-
[8]
Participatory ai for humani- tarian innovation: A briefing paper
Aleks Berditchevskaia, Eirini Malliaraki, and Kathy Peach. Participatory ai for humani- tarian innovation: A briefing paper. Nesta, 2021. URL https://media.nesta.org.uk/ documents/Nesta_Participatory_AI_for_humanitarian_innovation_Final.pdf
2021
Show all 111 references
-
[9]
Bertrand and S
M. Bertrand and S. Mullainathan. Are emily and greg more employable than lakisha and jamal? a field experiment on labor market discrimination. The American Economic Review, 94 (4):991–1013, 2004. URL https://www.jstor.org/stable/3592802
2004
-
[10]
A metrological framework for uncertainty evaluation in machine learning classification models, May 2025
Samuel Bilson, Maurice Cox, Anna Pustogvar, and Andrew Thompson. A metrological framework for uncertainty evaluation in machine learning classification models, May 2025. URL http://arxiv.org/abs/2504.03359. arXiv:2504.03359 [cs]
2025
-
[12]
Blaire et al
B. Blaire et al. Legal red teaming: A systematic approach to assessing legal risk of gener- ative ai models. https://www.dlapiper.com/-/media/project/dlapiper-tenant/ dlapiper/pdf/dla-piper---white-paper---ai-legal-red-teaming.pdf , 2024
2024
-
[13]
Atlas of ai risks: Enhancing public understanding of ai risks
Edyta Bogucka, Sanja Š´cepanovi´c, and Daniele Quercia. Atlas of ai risks: Enhancing public understanding of ai risks. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, volume 12, pages 33–43, 2024
2024
-
[14]
Bommasani, P
R. Bommasani, P. Liang, and T. Lee. Holistic evaluation of language models. Annals of the New York Academy of Sciences, 1525(1):140–146, 2023. doi: 10.1111/nyas.15007. URL https://nyaspubs.onlinelibrary.wiley.com/doi/abs/10.1111/nyas.15007. 10
2023 doi
-
[15]
The impact of ai on the workforce: Tasks versus jobs? Economics Letters, 244:111971, 2024
Kathryn Bonney, Cory Breaux, Catherine Buffington, Emin Dinlersoz, Lucia Foster, Nathan Goldschlag, John Haltiwanger, Zachary Kroff, and Keith Savage. The impact of ai on the workforce: Tasks versus jobs? Economics Letters, 244:111971, 2024. ISSN 0165-1765. doi: https://doi.or...
2024
-
[16]
Overcoming failures of imagination in ai infused system development and deployment, 2020
Margarita Boyarskaya, Alexandra Olteanu, and Kate Crawford. Overcoming failures of imagination in ai infused system development and deployment, 2020. URL https://arxiv. org/abs/2011.13416
2020 arXiv
-
[17]
Evaluating the replicability of social science experiments in nature and science between 2010 and 2015
Colin Camerer, Anna Dreber, Felix Holzmeister, Teck Ho, Jürgen Huber, Magnus Johannesson, Michael Kirchler, Gideon Nave, Brian Nosek, Thomas Pfeiffer, Adam Altmejd, Nick Buttrick, Taizan Chan, Yiling Chen, Eskil Forsell, Anup Gampa, Emma Heikensten, Lily Hummer, Taisuke Imai, ...
2010 doi
-
[18]
Carlini et al
N. Carlini et al. Extracting training data from large language models. In 30th USENIX Security Symposium, 2021. URL https://www.usenix.org/conference/usenixsecurity21/ presentation/carlini-extracting
2021
-
[19]
Quantifying memorization across neural language models, 2023
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. Quantifying memorization across neural language models, 2023. URL https://arxiv.org/abs/2202.07646
2023 arXiv
-
[20]
Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Itay Yona, Eric Wallace, David Rolnick, and Florian Tramèr
Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke, Jonathan Hayase, A. Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Itay Yona, Eric Wallace, David Rolnick, and Florian Tramèr. Stealing part of a production language model, ...
2024 arXiv
-
[21]
Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Bıyık, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, Tony Wang, Samuel Marks, Charbel-Raphaël Segerie, Micah Carroll, Andi Peng, Phillip Christoffersen, Mehul Damani, Stew...
2023 arXiv
-
[22]
D. Centola. The network science of collective intelligence. Trends in Cognitive Sciences, 26 (11):923–941, 2022. URL https://pubmed.ncbi.nlm.nih.gov/36180361/
2022
-
[23]
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. A survey on evaluation of large language models. ACM transactions on intelligent systems and technology, 15(3):1–45, 2024
2024
-
[24]
Pappas, and Eric Wong
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries, 2024. URL https://arxiv.org/abs/2310.08419
2024 arXiv
-
[25]
From representational harms to quality-of-service harms: A case study on llama 2 safety safeguards
Khaoula Chehbouni, Megha Roshan, Emmanuel Ma, Futian Wei, Afaf Taik, Jackie Chi Kit Cheung, and Golnoosh Farnadi. From representational harms to quality-of-service harms: A case study on llama 2 safety safeguards. In Findings of the Association for Computational Linguistics AC...
2024
-
[26]
Feder Cooper, Emily Corvi, P
Alexandra Chouldechova, Chad Atalla, Solon Barocas, A. Feder Cooper, Emily Corvi, P. Alex Dow, Jean Garcia-Gathright, Nicholas Pangakis, Stefanie Reed, Emily Sheng, Dan Vann, Matthew V ogel, Hannah Washington, and Hanna Wallach. A shared standard for valid measurement of gener...
2024 arXiv
- [27]
-
[28]
Cedric E. Dawkins. The principle of good faith: Toward substantive stakeholder engagement. J Bus Ethics, 121:283–295, 2014. doi: 10.1007/s10551-013-1697-z
2014 doi
- [29]
-
[30]
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. Show your work: Improved reporting of experimental results, 2019. URL https://arxiv.org/abs/ 1909.03004
2019 arXiv
-
[31]
Harmful speech detection by language models exhibits gender-queer dialect bias
Rebecca Dorn, Lee Kezar, Fred Morstatter, and Kristina Lerman. Harmful speech detection by language models exhibits gender-queer dialect bias. InProceedings of the 4th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, pages 1–12, 2024
2024
-
[32]
Duan et al
M. Duan et al. Do membership inference attacks work on large language models? arXiv preprint arXiv:2402.07841, 2024. Feb 12
2024 arXiv
-
[33]
On the impossible safety of large ai models, 2023
El-Mahdi El-Mhamdi, Sadegh Farhadkhani, Rachid Guerraoui, Nirupam Gupta, Lê-Nguyên Hoang, Rafael Pinot, Sébastien Rouault, and John Stephan. On the impossible safety of large ai models, 2023. URL https://arxiv.org/abs/2209.15259
2023 arXiv
-
[34]
Ethics Owners: A New Model of Organi- zational Responsibility in Data-Driven Technology Companies
Moss Emanuel and Metcalf Jacob. Ethics Owners: A New Model of Organi- zational Responsibility in Data-Driven Technology Companies . Data & Society, September 2020. URL https://datasociety.net/wp-content/uploads/2020/09/ Ethics-Owners_20200923-DataSociety.pdf
2020
-
[35]
Can we trust ai benchmarks? an interdisciplinary review of current issues in ai evaluation, 2025
Maria Eriksson, Erasmo Purificato, Arman Noroozian, Joao Vinagre, Guillaume Chaslot, Emilia Gomez, and David Fernandez-Llorca. Can we trust ai benchmarks? an interdisciplinary review of current issues in ai evaluation, 2025. URL https://arxiv.org/abs/2502. 06559
2025
-
[36]
Code-ifying the law: How disciplinary divides afflict the development of legal software
Nel Escher, Jeffrey Bilik, Nikola Banovic, and Ben Green. Code-ifying the law: How disciplinary divides afflict the development of legal software. Proc. ACM Hum.-Comput. Interact., 8(CSCW2), November 2024. doi: 10.1145/3686937. URL https://doi.org/10. 1145/3686937
2024 doi
-
[37]
The replication crisis has led to positive structural, procedural, and community changes
Max Korbmacher et al. The replication crisis has led to positive structural, procedural, and community changes. Communications Psychology, 2023. doi: 10.1038/s44271-023-00003-2
2023 doi
-
[38]
Large ai models are cultural and social technologies
Henry Farrell, Alison Gopnik, Cosma Shalizi, and James Evans. Large ai models are cultural and social technologies. Science, 387(6739):1153–1156, 2025. doi: 10.1126/science.adt9819. URL https://www.science.org/doi/abs/10.1126/science.adt9819
2025 doi
-
[39]
Value sensitive design: Theory and methods
Batya Friedman, Peter Kahn, and Alan Borning. Value sensitive design: Theory and methods. UW CSE Technical Report, 2003
2003
-
[40]
Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning, October 2016
Yarin Gal and Zoubin Ghahramani. Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning, October 2016. URL http://arxiv.org/abs/1506. 02142. arXiv:1506.02142 [stat]
2016 arXiv
-
[41]
Americans use ai in everyday prod- ucts without realizing it
Gallup and Telescope Foundation. Americans use ai in everyday prod- ucts without realizing it. URL https://www.telescopegp.com/insights/ americans-use-ai-in-everyday-products-without-realizing-it
-
[42]
Gao et al
L. Gao et al. A framework for few-shot language model evaluation. https://github.com/ EleutherAI/lm-evaluation-harness , 2021
2021
-
[43]
Real- toxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. Real- toxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462, 2020. 12
2009 arXiv
-
[44]
Trust, attitudes and use of artificial intelligence: A global study 2025, 2025
Nicole Gillespie, Steven Lockey, Tabi Ward, Alexandria Macdade, and Gerard Hassed. Trust, attitudes and use of artificial intelligence: A global study 2025, 2025. URL https://figshare.unimelb.edu.au/articles/report/Trust_attitudes_and_ use_of_artificial_intelligence_A_global_s...
2025
-
[45]
Greshake et al
K. Greshake et al. Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In 16th ACM Workshop on Artificial Intelligence and Security, pages 79–90, 2023. URL https://arxiv.org/abs/2302.12173
2023 arXiv
-
[46]
Troy, Dario Amodei, Jared Kaplan, Jack Clark, and Deep Ganguli
Kunal Handa, Alex Tamkin, Miles McCain, Saffron Huang, Esin Durmus, Sarah Heck, Jared Mueller, Jerry Hong, Stuart Ritchie, Tim Belonax, Kevin K. Troy, Dario Amodei, Jared Kaplan, Jack Clark, and Deep Ganguli. Which economic tasks are performed with ai? evidence from millions o...
2025 arXiv
-
[47]
Hoffmann and H
M. Hoffmann and H. Frase. Adding structure to ai harm: An introduc- tion to cset’s ai harm framework. https://cset.georgetown.edu/publication/ adding-structure-to-ai-harm/ , 2023
2023
-
[48]
Anna A. Ivanova. Toward best research practices in ai psychology, 2024. URL https: //arxiv.org/abs/2312.01276
2024 arXiv
-
[49]
Jansma, Anne M
Sikke R. Jansma, Anne M. Dijkstra, and Menno D.T. de Jong. Co-creation in support of responsible research and innovation: an analysis of three stakeholder workshops on nanotechnology for health. Journal of Responsible Innovation , 9(1):28–48, 2021. doi: 10.1080/23299460.2021.1994195
2021
-
[50]
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM computing surveys, 55(12):1–38, 2023
2023
-
[51]
Kapoor and A
S. Kapoor and A. Narayanan. Openai’s policies hinder reproducible research on language mod- els. https://www.aisnakeoil.com/p/openais-policies-hinder-reproducible , 2023
2023
-
[52]
Leakage and the reproducibility crisis in ml-based science, 2022
Sayash Kapoor and Arvind Narayanan. Leakage and the reproducibility crisis in ml-based science, 2022. URL https://arxiv.org/abs/2207.07048
2022 arXiv
-
[53]
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adi...
2021
-
[54]
Audiocaps: Gener- ating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. Audiocaps: Gener- ating captions for audios in the wild. In NAACL-HLT, 2019
2019
-
[55]
The history and risks of reinforcement learning and human feedback, 2023
Nathan Lambert, Thomas Krendl Gilbert, and Tom Zick. The history and risks of reinforcement learning and human feedback, 2023. URL https://arxiv.org/abs/2310.13595
2023 arXiv
-
[56]
Dempsey, Ece Kamar, Steven M
Susan Landau, James X. Dempsey, Ece Kamar, Steven M. Bellovin, and Robert Pool. Challenging the machine: Contestability in government ai systems, 2024. URL https: //arxiv.org/abs/2406.10430
2024 arXiv
-
[57]
Leask, Marlene Sandlund, Dawn A
Calum F. Leask, Marlene Sandlund, Dawn A. Skelton, Teatske M. Altenburg, Greet Cardon, et al. Framework, principles and recommendations for utilising participatory methodologies in the co-creation and evaluation of public health interventions. Res Involv Engagem, 5(2): 1153–11...
2019 doi
-
[58]
Lenaerts-Bergmans
B. Lenaerts-Bergmans. Data poisoning: The exploitation of generative ai. https://www. crowdstrike.com/cybersecurity-101/cyberattacks/data-poisoning/ , 2024
2024
-
[59]
Ai sustainability in practice part one: Foundations for sustainable ai projects
David Leslie, Cami Rincon, Morgan Briggs, et al. Ai sustainability in practice part one: Foundations for sustainable ai projects. https://aiethics.turing.ac.uk/modules/sustainability-1/, 2024
2024
- [60]
-
[61]
LLM defenses are not robust to multi-turn human jailbreaks yet
Nathaniel Li, Ziwen Han, Ian Steneker, Willow Primack, Riley Goodside, Hugh Zhang, Zifan Wang, Cristina Menghini, and Summer Yue. LLM defenses are not robust to multi-turn human jailbreaks yet. arXiv preprint arXiv:2408.15221, 2024
2024 arXiv
-
[62]
Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm- Burger, Rassin Lababidi, Lennart Justen, Andrew B
Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D. Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm- Burger, Rassin Lababidi, Lennart Justen, Andrew B. Liu, Michael Chen, Isabelle Barrass, Oliver Zhang...
2024 arXiv
-
[63]
A survey of multimodel large language models
Zijing Liang, Yanjie Xu, Yifan Hong, Penghui Shang, Qi Wang, Qiang Fu, and Ke Liu. A survey of multimodel large language models. In Proceedings of the 3rd International Conference on Computer, Artificial Intelligence and Control Engineering , pages 405–409, 2024
2024
-
[64]
Are we learn- ing yet? a meta review of evaluation failures across machine learning
Thomas Liao, Rohan Taori, Deborah Raji, and Ludwig Schmidt. Are we learn- ing yet? a meta review of evaluation failures across machine learning. In J. Vanschoren and S. Yeung, editors, Proceedings of the Neural Information Pro- cessing Systems Track on Datasets and Benchmarks ...
2021
-
[65]
Jaffe, Abhishek Nagaraj, Imke Reimers, Michael D
Brent Lutes, Joshua Gans, Shane Greenstein, Adam B. Jaffe, Abhishek Nagaraj, Imke Reimers, Michael D. Smith, Rahul Telang, Catherine E. Tucker, and Joel Waldfogel. Identifying the economic implications of artificial intelligence for copyright policy (february 12, 2025). first ...
2025 doi
-
[66]
Madaio, Jingya Chen, Hanna Wallach, and Jennifer Wortman Vaughan
Michael A. Madaio, Jingya Chen, Hanna Wallach, and Jennifer Wortman Vaughan. Tinker, tailor, configure, customize: The articulation work of contextualizing an ai fairness checklist. Proc. ACM Hum.-Comput. Interact., 8(CSCW1), April 2024. doi: 10.1145/3653705. URL https://doi.o...
2024 doi
-
[67]
Martínez
E. Martínez. Re-evaluating GPT-4’s bar exam performance. Artificial Intelligence and Law, 30:1–24, 2024. doi: 10.1177/20539517241290220. URL https://doi.org/10.1177/ 20539517241290220
2024 doi
-
[68]
Co-creation for policy: Participatory methodologies to structure multi-stakeholder policymaking processes
Cristian Matti and Gabriel Rissola. Co-creation for policy: Participatory methodologies to structure multi-stakeholder policymaking processes. https://www.eit.europa.eu/ sites/default/files/jrc128771_01.pdf, 2022
2022
-
[69]
How the u.s
Colleen McClain, Brian Kennedy, Jeffrey Gottfried, Monica Anderson, and Gi- ancarlo Pasquini. How the u.s. public and ai experts view artificial intelli- gence, April 2025. URL https://www.pewresearch.org/internet/2025/04/03/ how-the-us-public-and-ai-experts-view-artificial-in...
2025
-
[70]
A. Mislove. Red-teaming large language models to identify novel ai risks. https://www.whitehouse.gov/ostp/news-updates/2023/08/29/ red-teaming-large-language-models-to-identify-novel-ai-risks/ , 2023
2023
-
[71]
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland, May 2022
Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland, May 2022. Association for Computational Linguistics. URL https: //aclanthology.o...
2022
-
[72]
thick alignment
A. Nelson. Facct’23 keynote: "thick alignment". https://www.youtube.com/watch?v= Sq_XwqVTqvQ, 2023
2023
-
[73]
Global law and policy tracker
International Association of Privacy Professionals. Global law and policy tracker. URL https://iapp.org/resources/article/global-ai-legislation-tracker/
-
[74]
Governing with artificial intelli- gence, June 2024
OECD Artificial Intelligence Papers. Governing with artificial intelli- gence, June 2024. URL https://www.oecd.org/en/publications/ governing-with-artificial-intelligence_26324bc2-en.html
2024
-
[75]
Luona Lin Parker and Kim. U.s. workers are more worried than hopeful about future ai use in the workplace, February 2025. URL https://www.pewresearch.org/social-trends/2025/02/25/ u-s-workers-are-more-worried-than-hopeful-about-future-ai-use-in-the-workplace/
2025
-
[76]
Concrete problems in ai safety, revisited, 2023
Inioluwa Deborah Raji and Roel Dobbe. Concrete problems in ai safety, revisited, 2023. URL https://arxiv.org/abs/2401.10899
2023 arXiv
-
[77]
The fallacy of ai functionality
Inioluwa Deborah Raji, I Elizabeth Kumar, Aaron Horowitz, and Andrew Selbst. The fallacy of ai functionality. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 959–972, 2022
2022
-
[78]
Do imagenet classifiers generalize to imagenet?, 2019
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to imagenet?, 2019. URL https://arxiv.org/abs/1902.10811
2019 arXiv
-
[79]
Just what do you think you’re doing, dave?’ a checklist for responsible data use in nlp, 2021
Anna Rogers, Tim Baldwin, and Kobi Leins. Just what do you think you’re doing, dave?’ a checklist for responsible data use in nlp, 2021. URL https://arxiv.org/abs/2109. 06598
2021
-
[80]
everyone wants to do the model work, not the data work
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. “everyone wants to do the model work, not the data work”: Data cascades in high-stakes ai. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems , C...
2021
-
[81]
Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerster, Tim Rocktäschel, and Roberta Raileanu
Mikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro, Aram H. Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerster, Tim Rocktäschel, and Roberta Raileanu. Rainbow teaming: Open-ended generation of diverse adversarial prompts, 20...
2024 arXiv
-
[82]
Towards a standard for identifying and managing bias in artificial intelligence
Reva Schwartz, Apostol Vassilev, Kristen Greene, Lori Perine, Andrew Burt, and Patrick Hall. Towards a standard for identifying and managing bias in artificial intelligence. Technical Report NIST SP 1270, National Institute of Standards and Technology, Gaithersburg, MD,
-
[83]
The nist assessing risks and impacts of ai (aria) pilot evaluation plan
Reva Schwartz, Jonathan Fiscus, Kristen Greene, Gabriella Waters, Rumman Chowdhury, Theodore Jensen, Craig Greenberg, Afzal Godil, Razvan Amironesei, Patrick Hall, and Shomik Jain. The nist assessing risks and impacts of ai (aria) pilot evaluation plan. Technical report, Natio...
2024
-
[84]
The assessing risks and impacts of ai (aria) program evaluation design document
Reva Schwartz, Gabriella Waters, Razvan Amironesei, Craig Greenberg, Jon Fiscus, Patrick Hall, Anya Jones, Shomik Jain, Afzal Godil, Kristen Greene, Ted Jensen, and Noah Schulman. The assessing risks and impacts of ai (aria) program evaluation design document. Technical report...
2024
-
[85]
A. D. Selbst et al. Fairness and abstraction in sociotechnical systems. In Proceedings of the Conference on Fairness, Accountability, and Transparency , pages 59–68, 2019. doi: 10.1145/3287560.3287598. URL https://doi.org/10.1145/3287560.3287598
2019
-
[86]
Sociotech- nical harms of algorithmic systems: Scoping a taxonomy for harm reduction
Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Rostamzadeh, Paul Nicholas, N’Mah Yilla-Akbari, Jess Gallegos, Andrew Smart, Emilio Garcia, et al. Sociotech- nical harms of algorithmic systems: Scoping a taxonomy for harm reduction. In Proceedings of the 2023 ...
2023
-
[87]
Taskbench: Benchmarking large language models for task automation
Yongliang Shen, Kaitao Song, Xu Tan, Wenqi Zhang, Kan Ren, Siyu Yuan, Weiming Lu, Dongsheng Li, and Yueting Zhuang. Taskbench: Benchmarking large language models for task automation. Advances in Neural Information Processing Systems, 37:4540–4574, 2024
2024
-
[88]
Shumailov et al
I. Shumailov et al. Sponge examples: Energy-latency attacks on neural networks. In IEEE European Symposium on Security and Privacy (EuroS&P) , pages 212–231, 2021. URL https://arxiv.org/abs/2006.03463
2021 arXiv
-
[89]
The leaderboard illusion, 2025
Shivalika Singh, Yiyang Nan, Alex Wang, Daniel D’Souza, Sayash Kapoor, Ahmet Üstün, Sanmi Koyejo, Yuntian Deng, Shayne Longpre, Noah Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker. The leaderboard illusion, 2025. URL https://arxiv.org/abs/2504. 20879
2025
-
[90]
Participation is not a design fix for machine learning
Mona Sloane, Emanuel Moss, Olaitan Awomolo, and Laura Forlano. Participation is not a design fix for machine learning. EAAMO ’22, New York, NY , USA, 2022. Association for Computing Machinery. ISBN 9781450394772. doi: 10.1145/3551624.3555285. URL https://doi.org/10.1145/355162...
2022
-
[91]
Slota, Kenneth R
Stephen C. Slota, Kenneth R. Fleischmann, Sherri Greenberg, Nitin Verma, Brenna Cummings, Lan Li, and Chris Shenefiel. Many hands make many fingers to point: Challenges in creating accountable ai. AI and Society, 38(4):1287–1299, 2023. doi: 10.1007/s00146-021-01302-0
2023 doi
-
[92]
Smith, Saleema Amershi, Solon Barocas, Hanna Wallach, and Jennifer Wort- man Vaughan
Jessie J. Smith, Saleema Amershi, Solon Barocas, Hanna Wallach, and Jennifer Wort- man Vaughan. Real ml: Recognizing, exploring, and articulating limitations of machine learning research. In 2022 ACM Conference on Fairness Accountability and Transparency, FAccT ’22, page 587–5...
2022
-
[93]
Srivastava et al
A. Srivastava et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models, 2023. URL https://arxiv.org/abs/2206.04615
2023 arXiv
-
[94]
Beyond memorization: Violating privacy via inference with large language models, 2024
Robin Staab, Mark Vero, Mislav Balunovi ´c, and Martin Vechev. Beyond memorization: Violating privacy via inference with large language models, 2024. URL https://arxiv. org/abs/2310.07298
2024 arXiv
-
[95]
Stadler, T
C. Stadler, T. Rajwani, and F. Karaba. Solutions to the exploration/exploitation dilemma: networks as a new level of analysis. International Journal of Management Reviews, 16(2): 172–193, 2013. doi: 10.1111/ijmr.12015. URL https://doi.org/10.1111/ijmr.12015
2013 doi
-
[96]
The State of AI Governance Research: AI Safety and Reliability in Real World Commercial Deployment
Ilan Strauss, Isobel Moure, Tim O’Reilly, and Sruly Rosenblat. The State of AI Governance Research: AI Safety and Reliability in Real World Commercial Deployment. AI Disclosures Project, Social Science Research Council, April 2025. doi: 10.35650/aidp.4112.d.2025. URL http://dx...
2025 doi
-
[97]
Sumers, Jared Mueller, William McEachen, Wes Mitchell, Shan Carter, Jack Clark, Jared Kaplan, and Deep Ganguli
Alex Tamkin, Miles McCain, Kunal Handa, Esin Durmus, Liane Lovitt, Ankur Rathi, Saffron Huang, Alfred Mountfield, Jerry Hong, Stuart Ritchie, Michael Stern, Brian Clarke, Landon Goldberg, Theodore R. Sumers, Jared Mueller, William McEachen, Wes Mitchell, Shan Carter, Jack Clar...
2024 arXiv
-
[98]
Theofanos, Y
M. Theofanos, Y . Choong, and T. Jensen. Ai use taxonomy: A human-centered approach. Technical Report NIST AI 200-1, National Institute of Standards and Technology, Gaithersburg, MD, 2024. URL https://doi.org/10.6028/NIST.AI.200-1
2024 doi
-
[99]
Thomas and D
R. Thomas and D. Uminsky. Reliance on metrics is a fundamental challenge for ai. https: //arxiv.org/pdf/2002.08512, 2020
2002 arXiv
-
[100]
Analytical results for uncertainty propagation through trained ma- chine learning regression models, May 2024
Andrew Thompson. Analytical results for uncertainty propagation through trained ma- chine learning regression models, May 2024. URL http://arxiv.org/abs/2404.11224. arXiv:2404.11224 [cs]
2024 arXiv
-
[101]
E. L. Trist and K. W. Bamforth. Some social and psychological consequences of the long- wall method of coal-getting: An examination of the psychological situation and defences of a work group in relation to the social structure and technological content of the work system. Hum...
1951 doi
-
[102]
Prime minister sets out blueprint to turbocharge ai
UK Department for Science, Innovation and Technology. Prime minister sets out blueprint to turbocharge ai. https://https://www.gov.uk/government/news/ prime-minister-sets-out-blueprint-to-turbocharge-ai , 2025
2025
-
[103]
A Deeper Look into Aleatoric and Epistemic Uncertainty Disentanglement, April 2022
Matias Valdenegro-Toro and Daniel Saromo. A Deeper Look into Aleatoric and Epistemic Uncertainty Disentanglement, April 2022. URL http://arxiv.org/abs/2204.09308. arXiv:2204.09308 [cs]
2022 arXiv
-
[104]
Feder Cooper, Angelina Wang, Chad Atalla, Solon Baro- cas, Su Lin Blodgett, Alexandra Chouldechova, Emily Corvi, P
Hanna Wallach, Meera Desai, A. Feder Cooper, Angelina Wang, Chad Atalla, Solon Baro- cas, Su Lin Blodgett, Alexandra Chouldechova, Emily Corvi, P. Alex Dow, Jean Garcia- Gathright, Alexandra Olteanu, Nicholas Pangakis, Stefanie Reed, Emily Sheng, Dan Vann, Jennifer Wortman Vau...
2025 arXiv
-
[105]
Decodingtrust: A comprehensive assess- ment of trustworthiness in gpt models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al. Decodingtrust: A comprehensive assess- ment of trustworthiness in gpt models. In NeurIPS, 2023
2023
-
[106]
Calibration in Deep Learning: A Survey of the State-of-the-Art, May 2024
Cheng Wang. Calibration in Deep Learning: A Survey of the State-of-the-Art, May 2024. URL http://arxiv.org/abs/2308.01222. arXiv:2308.01222 [cs]
2024 arXiv
-
[107]
Sociotechnical safety evaluation of generative ai systems,
Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, and William Isaac. Sociotechnical safety evaluation of generative ai systems,
-
[108]
Roberts, and Mary L
Alice Qian Zhang, Ryland Shaw, Jacy Reese Anthis, Ashlee Milton, Emily Tseng, Jina Suh, Lama Ahmad, Ram Shankar Siva Kumar, Julian Posada, Benjamin Shestakofsky, Sarah T. Roberts, and Mary L. Gray. The human factor in ai red teaming: Perspectives from social and collaborative ...
2024 arXiv
-
[109]
Felm: Benchmarking factuality evaluation of large language models.Advances in Neural Information Processing Systems, 36:44502–44523, 2023
Yiran Zhao, Jinghan Zhang, I Chern, Siyang Gao, Pengfei Liu, Junxian He, et al. Felm: Benchmarking factuality evaluation of large language models.Advances in Neural Information Processing Systems, 36:44502–44523, 2023
2023
-
[110]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. URL https: //arxiv.org/abs/2306.056...
2023 arXiv
-
[2022]
URL https://doi.org/10.6028/NIST.SP.1270
-
[2023]
URL https://arxiv.org/abs/2310.11986
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.