Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

GPAI Evaluations Standards Taskforce: Towards Effective AI Governance

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A dedicated EU Taskforce is proposed to write adaptive standards for mandatory AI evaluations.

desk verdict A transparent, well-scoped institutional proposal for EU GPAI evaluation standards; the net-benefit claim is honestly framed as an assumption, and the paper's own caveats are both its strength and its soft spot. read the letter →

arxiv 2411.13808 v1 pith:ETZ2F6OP submitted 2024-11-21 cs.CY

classification cs.CY
keywords general-purposeAIevaluationsgovernanceEUActsystemicriskstandard-settingevaluationdesiderataBrusselseffect
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper's central claim is that the EU, as the only jurisdiction that legally mandates evaluations of general-purpose AI (GPAI) models, should create a dedicated Evaluation Standards Taskforce to write and continuously update standards for those evaluations. The authors argue that without such standards, the mandatory evaluations could lack internal validity, external validity, reproducibility, and portability, and therefore fail to support risk assessment and mitigation. They identify the scientific panel and advisory forum created by the EU AI Act as viable institutional homes, specify duties such as harmonising risk taxonomies, setting evaluation standards, and quality-controlling third-party evaluators, and spell out provider commitments on documentation and model access. A sympathetic reading: the paper is trying to show that a standing, adaptively governed expert body is the missing institutional piece that makes legally mandated AI evaluations trustworthy and globally influential.

What carries the argument

The Taskforce is the central mechanism: a vetted body of independent researchers operating inside bodies established by the EU AI Act, with three main duties—harmonizing a GPAI systemic-risk taxonomy and evaluation methodologies, regularly publishing updated evaluation standards, and vetting and auditing third-party evaluators and evaluation results. The four desiderata of internal validity, external validity, reproducibility, and portability function as design constraints that the standards must satisfy. The paper also treats the EU AI Act's Codes of Practice as the instrument that could codify provider commitments to supply documentation, white-box access, models, and training-data information.

What would settle it

A field study or natural experiment comparing EU providers' standardized evaluation results with subsequent audits or real-world incidents: if models that pass standardized evaluations are no less likely than models that fail to be implicated in systemic harms, the central claim is undercut. Alternatively, demonstrating that current provider evaluations already predict extreme-risk incidents with high reliability would argue that additional standards are unnecessary.

Watch

Extended reading notes

Core claim

The paper proposes that an EU GPAI Evaluation Standards Taskforce, housed within the EU AI Act's institutions—either the Scientific Panel of Independent Experts or the Advisory Forum—should maintain standards upholding four desiderata: internal validity (results reflect true capabilities in the test setting), external validity (results proxy real-world behavior), reproducibility (same inputs yield same results), and portability (the same evaluations run across institutions). It contends that standards must be adaptive because models, risks, and evaluation methods change quickly, and that the Taskforce's work could propagate globally via both formal adoption and the regulatory pull of the EU market. The causal claim is that institutionalised standard-setting for evaluations will increase evaluation quality and legitimacy, improving risk assessment and mitigation.

Load-bearing premise

That well-formulated evaluation standards can actually make risk assessment and mitigation more effective—if evaluations cannot reliably measure systemic risk, or if standards do not change behavior, the Taskforce has no reason to exist.

Editorial extensions

If this is right

  • If the Taskforce works as proposed, EU-mandated GPAI evaluations would have enforceable quality standards for the first time, making regulatory trigger decisions more defensible.
  • Provider commitments laid out in the paper could be codified in the EU AI Act's Code of Practice, giving the Taskforce reliable documentation and model access.
  • Adaptive standards would be updated periodically, such as annually or by qualified alert, preventing benchmarks from decaying through train-test contamination.
  • Through Brussels-effect channels, EU standards could become de facto or de jure international standards, shaping provider behavior beyond the EU.
  • The paper's failure modes, such as a false sense of security, evaluator friction, regulatory capture, and talent bottlenecks, imply that complementary approaches like outcome-based regulation, information sharing, and bounty programs should be considered.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The case ultimately rests on an empirical premise the paper does not test: that standardized evaluations actually reduce real-world systemic risk; a reader could ask for evidence linking evaluation results to incident outcomes.
  • If standards become mandatory, they may create gaming or Goodharting incentives; the four desiderata do not directly address adversarial behavior by providers, which would need ongoing monitoring.
  • The Taskforce model could be tested outside the EU, for example by piloting a similar standards body within an AI Safety Institute and comparing the reliability of its standardized evaluations against ad hoc evaluation practice.
  • The desiderata are borrowed from social-science evaluation methodology; adapting them to fast-moving model capabilities may require new measures of construct validity that do not yet exist.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This policy paper argues that general-purpose AI (GPAI) evaluations, although increasingly central to AI governance and mandated by the EU AI Act for systemic-risk models, currently lack quality and legitimacy standards. It proposes four desiderata for such evaluations—internal validity, external validity, reproducibility, and portability—and recommends the creation of a dedicated EU GPAI Evaluation Standards Taskforce housed within institutions established by the EU AI Act. The paper outlines potential institutional settings, taskforce duties, provider commitments, possible global impact through Brussels effects, and a list of failure modes and alternative approaches. It closes by stating four key assumptions on which the proposal rests and acknowledges the nascent, evidence-sparse nature of the field.

Significance. If the proposal were adopted and its assumptions held, the Taskforce would be a first formal mechanism for continuously setting and updating standards for legally mandated AI evaluations, with plausible international influence through EU regulatory leadership. The paper is a coherent, well-structured policy proposal rather than an empirical study. Its strengths include a clear desiderata framework, concrete integration with existing EU AI Act bodies, explicit provider commitments, and unusually transparent admission of the assumptions and limitations. However, the central net-benefit claim is not established: the paper's own Section 8 identifies a direct backfire pathway—false sense of security—that could make regulation worse, and assumptions 1 and 4 are asserted without supporting evidence or a proposed monitoring framework. The contribution is valuable as a position piece, but the claims need to be more carefully conditioned and the failure modes more fully integrated into the design.

major comments (3)
  1. [Section 9] Assumption 1, stated verbatim as 'well-formulated GPAI evaluations standards can support effective risk assessment, risk mitigation, and AI governance,' is load-bearing but unsupported. The paper does not provide evidence that standards improve outcomes, and Section 3 itself documents that current evaluations are highly variable (e.g., GPQA noise across ten runs in Section 3.1) and can be invalidated by fine-tuning (Section 3.2). These observations show that the desiderata are necessary, but they do not show that standards can be formulated and enforced sufficiently to deliver the claimed benefits. The paper should either provide empirical or historical evidence that similar evaluation standards have improved decision-making, or explicitly reframe the proposal as conditional on an unverified premise.
  2. [Section 8] The 'false sense of security' failure mode listed in Section 8 is a direct mechanism that can violate assumptions 1 and 4 even if the Taskforce produces well-formulated standards: the existence of standards may reduce vigilance or crowd out more promising risk-management approaches. The paper acknowledges this possibility but offers no argument that it is unlikely, no design feature to mitigate it, and no plan to monitor for it. Because the central claim is that the Taskforce 'could promote relevant and governance-enhancing standards for effective risk assessment and mitigation,' the unresolved backfire pathway means the net-benefit claim is unsupported. The proposal should specify how the Taskforce would evaluate its own effectiveness, track displacement of other risk-management efforts, and adjust or disband if backfire is detected.
  3. [Sections 5.1 and 5.2] There is an unresolved tension between the Taskforce's claimed independence and the two proposed institutional settings. Section 5.2 describes the Taskforce as 'a vetted body of independent researchers,' but IS2 (Advisory Forum) explicitly allows the involvement of GPAI industry experts, and Section 5.1 presents IS1 and IS2 as roughly equivalent options without comparing their risks of regulatory capture. Given that Section 8 itself lists regulatory capture as a standard failure mode, the paper should justify why the Taskforce would remain independent under IS2, or should state a preference and explain how conflicts of interest would be managed.
minor comments (4)
  1. [References] References [76] and [77] appear to be the same work (both list Weidinger et al., 'Holistic safety and responsibility evaluations of advanced AI models,' arXiv:2404.14068); one duplicate should be removed and the citations renumbered accordingly.
  2. [Throughout] The paper inconsistently uses 'GPAI' and 'GP AI'; one form should be chosen and used consistently throughout.
  3. [Appendix B] The sentence 'GP AI providers additionally commit to providing updated documentation to the Taskforce, including but not limited to safety cases, scaling risk management policies, capability forecast reports, and incident reporting documentation, as outlined in Article 55.1.c and Recital 115' appears twice verbatim; the duplicate should be removed.
  4. [Appendix B] There is a typo in 'number of peramet ers'; it should be 'number of parameters.'

Circularity Check

0 steps flagged · score 1.0 of 10

No meaningful circularity: the paper is a normative policy proposal, not a derivation; its few self-citations are background support and not load-bearing.

full rationale

The paper contains no fitted parameters, no equations, and no prediction that is equivalent to an input by construction. Its central claim—that an EU GPAI Evaluation Standards Taskforce could promote governance-enhancing evaluation standards—is a policy recommendation supported by stated desiderata, institutional analysis, and external evidence; the conclusion explicitly lists the four assumptions on which the proposal rests. An unproven assumption is an evidence or correctness limitation, not circular reasoning. The self-citations (e.g., [2], [61], [62], which share authors with this paper) are used only as supporting references for background claims about external scrutiny ecosystems, open problems in AI governance, and talent challenges; none of these citations supplies the load-bearing premise that the Taskforce would be effective. Section 8's 'false sense of security' backfire is an acknowledged limitation and a potential threat to the proposal's net-benefit assumption, but it is not a circular step. The paper is therefore self-contained with respect to circularity, with only minor non-load-bearing self-citations.

Assumptions & free parameters 0 free parameters · 5 assumptions · 1 invented entities

This is a policy proposal, so there are no fitted parameters. The argument rests on the four explicit assumptions listed in the conclusion plus the foundational premise that evaluations can mitigate risk. The proposed Taskforce is a new institutional entity without independent evidence.

assumptions (5)
  • domain assumption Well-formulated GPAI evaluation standards can support effective risk assessment, risk mitigation, and AI governance.
    Stated as key assumption 1 in the conclusion; the entire proposal depends on this premise.
  • domain assumption Developing standards within institutions capable of mandating them increases legitimacy and impact.
    Key assumption 2 in conclusion; motivates placing the Taskforce inside EU AI Act bodies.
  • domain assumption A dedicated multi-stakeholder Taskforce increases standards quality relative to uncoordinated interactions.
    Key assumption 3 in conclusion; the institutional design choice is not empirically validated.
  • domain assumption The benefits of the Taskforce for AI governance and public safety outweigh the costs.
    Key assumption 4 in conclusion; no cost-benefit analysis is provided.
  • domain assumption GPAI evaluations are a central tool for assessing and mitigating systemic risks.
    Foundational premise of Section 2 and the paper's motivation.
invented entities (1)
  • EU GPAI Evaluation Standards Taskforce
    purpose: To maintain and update standards for GPAI evaluations in the EU, promoting internal validity, external validity, reproducibility, and portability.
    The paper proposes this new institutional body; it has no empirical basis yet and is not an observed entity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of GPAI Evaluations Standards Taskforce: Towards Effective AI Governance." pith.science (2026). https://pith.science/paper/ETZ2F6OP

@misc{pith2026241113808,
  author       = {Pith},
  title        = {Pith review of: GPAI Evaluations Standards Taskforce: Towards Effective AI Governance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ETZ2F6OP}},
  note         = {Machine review of arXiv:2411.13808}
}
read the original abstract

General-purpose AI evaluations have been proposed as a promising way of identifying and mitigating systemic risks posed by AI development and deployment. While GPAI evaluations play an increasingly central role in institutional decision- and policy-making -- including by way of the European Union AI Act's mandate to conduct evaluations on GPAI models presenting systemic risk -- no standards exist to date to promote their quality or legitimacy. To strengthen GPAI evaluations in the EU, which currently constitutes the first and only jurisdiction that mandates GPAI evaluations, we outline four desiderata for GPAI evaluations: internal validity, external validity, reproducibility, and portability. To uphold these desiderata in a dynamic environment of continuously evolving risks, we propose a dedicated EU GPAI Evaluation Standards Taskforce, to be housed within the bodies established by the EU AI Act. We outline the responsibilities of the Taskforce, specify the GPAI provider commitments that would facilitate Taskforce success, discuss the potential impact of the Taskforce on global AI governance, and address potential sources of failure that policymakers should heed.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Preliminary suggestions for rigorous GPAI model evaluations

    cs.CY 2025-07 conditional novelty 4.0 of 10

    A RAND team turned a 64-paper literature review into a preliminary best-practice checklist for rigorously evaluating general-purpose AI models.

Reference graph

Works this paper leans on

81 extracted references · 56 canonical work pages · cited by 1 Pith paper

  1. [1]

    The Bletchley Declaration by Countrie s Attending the AI Safety Summit, 1-2 November 2023

    AI Safety Summit. The Bletchley Declaration by Countrie s Attending the AI Safety Summit, 1-2 November 2023. 2023

  2. [2]

    Towards Publicly Accountable Frontier LLMs: Building an External Scrutiny Ecosystem under the ASPIRE Framework

    Markus Anderljung, Everett Thornton Smith, Joe O’Brien , Lisa Soder, Benjamin Bucknall, Emma Bluemke, Jonas Schuett, Robert Trager, Lacey Strahm, a nd Rumman Chowdhury. To- wards publicly accountable frontier LLMs: Building an exte rnal scrutiny ecosystem under the ASPIRE framework. arXiv preprint arXiv:2311.14711 , 2023

  3. [3]

    Anthropic’s responsible scaling policy

    Anthropic. Anthropic’s responsible scaling policy. https://www-cdn.anthropic.com/1adf000c8f675958c2ee2 3805d

  4. [4]

    We need a science of evals

    Apollo Research. We need a science of evals. www.apolloresearch.ai/blog/we-need-a-science-of-eva ls,

  5. [5]

    Exploring the Relevance of Data Privacy-Enhancing Technologies for AI Governance Use Cases

    Emma Bluemke, Tantum Collins, Ben Garfinkel, and Andrew T rask. Exploring the rel- evance of data privacy-enhancing technologies for ai gover nance use cases, 2023. URL https://arxiv.org/abs/2303.08956

  6. [6]

    On the opportunities and risks of foundation models

    Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman , Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut , Emma Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258 , 2021

  7. [7]

    Holistic eva luation of language models

    Rishi Bommasani, Percy Liang, and Tony Lee. Holistic eva luation of language models. Annals of the New Y ork Academy of Sciences, 1525(1):140–146, 2023

  8. [8]

    What will it take to fix be nchmarking in natural language understanding? arXiv preprint arXiv:2104.02145 , 2021

    Samuel R Bowman and George E Dahl. What will it take to fix be nchmarking in natural language understanding? arXiv preprint arXiv:2104.02145 , 2021

Show all 81 references
  1. [9]

    The Brussels effect: How the European Union rules the world

    Anu Bradford. The Brussels effect: How the European Union rules the world . Oxford Univer- sity Press, USA, 2020

  2. [10]

    Th e malicious use of artificial intelligence: Forecasting, prevention, and mitigation

    Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, P eter Eckersley, Ben Garfinkel, Allan Dafoe, Paul Scharre, Thomas Zeitzoff, Bobby Filar, et al. Th e malicious use of artificial intelligence: Forecasting, prevention, and mitigation. arXiv preprint arXiv:1802.07228 , 2018

  3. [11]

    The dyn amics of standardization: Three per- spectives on standards in organization studies

    Nils Brunsson, Andreas Rasche, and David Seidl. The dyn amics of standardization: Three per- spectives on standards in organization studies. Organization Studies, 33(5-6):613–632, 2012. doi: 10.1177/0170840612450120. URL https://doi.org/10.1177/0170840612450120

  4. [12]

    Principled instructions are all you need for questioning LLaMa-1/2, gpt-3.5/4

    Sondos Mahmoud Bsharat, Aidar Myrzakhan, and Zhiqiang Shen. Principled instructions are all you need for questioning LLaMa-1/2, gpt-3.5/4. arXiv preprint arXiv:2312.16171 , 2023

  5. [13]

    Structured Acc ess for Third-Party Research on Frontier AI Models: Investigating Researchers’ Model Acce ss Requirements, 2023

    Benjamin S Bucknall and Robert F Trager. Structured Acc ess for Third-Party Research on Frontier AI Models: Investigating Researchers’ Model Acce ss Requirements, 2023

  6. [14]

    With little power comes great responsibility

    Dallas Card, Peter Henderson, Urvashi Khandelwal, Rob in Jia, Kyle Mahowald, and Dan Ju- rafsky. With little power comes great responsibility. arXiv preprint arXiv:2010.06595 , 2020

  7. [15]

    Black-box access is insufficient for rigorous AI audits

    Stephen Casper, Carson Ezell, Charlotte Siegmann, Noa m Kolt, Taylor Lynn Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, Jérémy Scheurer, Mariu s Hobbhahn, et al. Black-box access is insufficient for rigorous AI audits. In The 2024 ACM Conference on Fairness, Ac- countabi...

  8. [16]

    Performance-based regulation: Prospects and limitations in health, safety and environmental protection

    Cary Coglianese and Jennifer Nash. Performance-based regulation: Prospects and limitations in health, safety and environmental protection. All Faculty Scholarship, 2815, 2003. 13

  9. [17]

    Regulatory capture: A review

    Ernesto Dal Bó. Regulatory capture: A review. Oxford Review of Economic Policy , 22(2): 203–225, 2006

  10. [18]

    Frontier Safe ty Framework

    Anca Dragan, Helen King, and Allan Dafoe. Frontier Safe ty Framework. https://storage.googleapis.com/deepmind-media/DeepM ind.com/Blog/introducing-the-frontier-safety

  11. [19]

    2022 Strengthened Code of Practi ce on Disinformation, 2022

    European Commission. 2022 Strengthened Code of Practi ce on Disinformation, 2022. URL https://digital-strategy.ec.europa.eu/en/library/20 22-strengthened-code-practice-disinformati

  12. [20]

    AI Act: Participate in the drawin g-up of the first General-Purpose AI Code of Practice, 2024

    European Commission. AI Act: Participate in the drawin g-up of the first General-Purpose AI Code of Practice, 2024. URL https://digital-strategy.ec.europa.eu/en/news/ai-ac t-participate-drawing-first-general-purpos

  13. [21]

    Accessed: September 6, 2024

  14. [22]

    European Parliament and Council of the European Union. Regulation of the Euro- pean Parliament and of the Council laying down harmonised ru les on artificial intel- ligence and amending Regulations (EC) No 300/2008, (EU) No 1 67/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/...

  15. [23]

    Fro ntier AI Safety Commitments, AI Seoul Summit 2024

    Innovation Department for Science and Technology. Fro ntier AI Safety Commitments, AI Seoul Summit 2024. 2024. Accessed: September 6, 2024

  16. [24]

    European Commission, Directorate-General for Commun ications Networks, Content and Tech- nology. Commission Staff Working Document Impact Assessme nt Accompanying the Pro- posal for a Regulation of the European Parliament and of the C ouncil Laying Down Har- monised Rules on A...

  17. [25]

    Red teaming language models to reduce harms: Methods, scaling behaviors, and les sons learned

    Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda As kell, Y untao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al. Red teaming language models to reduce harms: Methods, scaling behaviors, and les sons learned. arXiv preprint arXiv:2209.0...

  18. [26]

    Challenges in Evaluating AI Systems

    Deep Ganguli, Nicholas Schiefer, Marina Favaro, and Ja ck Clark. Challenges in Evaluating AI Systems. https://www.anthropic.com/news/evaluating-ai-system s, 2023. Ac- cessed: September 6, 2024

  19. [27]

    Predictability and surprise in large generative models

    Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda As kell, Y untao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain, Nelson Elhage, et al. Predictability and surprise in large generative models. In Proceedings of the 2022 ACM Conference on Fairness, Account ability, an...

  20. [28]

    What does research repro- ducibility mean? Science Translational Medicine , 8(341):341ps312–341ps312, 2016

    Steven N Goodman, Daniele Fanelli, and John P A Ioannidi s. What does research repro- ducibility mean? Science Translational Medicine , 8(341):341ps312–341ps312, 2016. doi: 10.1126/scitranslmed.aaf5027

  21. [29]

    Promoting safety by increasing uncertai nty – implications for risk management

    Gudela Grote. Promoting safety by increasing uncertai nty – implications for risk management. Safety Science, 2015. URL https://doi.org/10.1016/j.ssci.2014.02.010

  22. [30]

    External validity: we need to do more

    Russell E Glasgow, Lawrence W Green, Lisa M Klesges, Dav id B Abrams, Edwin B Fisher, Michael G Goldstein, Laura L Hayman, Judith K Ockene, and C Tr acy Orleans. External validity: we need to do more. Annals of Behavioral Medicine , 31(2), 2006

  23. [31]

    A bias bounty for AI will help to catch unfair algorithms faster, 2022

    Melissa Heikkila. A bias bounty for AI will help to catch unfair algorithms faster, 2022. Ac- cessed: September 6, 2024. 14

  24. [32]

    Mitigating the risk of extinction fro m AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear wa r, 2023

    Geoffrey Hinton. Mitigating the risk of extinction fro m AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear wa r, 2023. Statement on AI risk

  25. [33]

    AI regulation has its own alignment problem: The technical a nd institutional feasibility of dis- closure, registration, licensing, and auditing

    Neel Guha, Christie Lawrence, Lindsey A Gailmard, Kit R odolfa, Faiz Surani, Rishi Bom- masani, Inioluwa Raji, Mariano-Florentino Cuéllar, Colle en Honigsberg, Percy Liang, et al. AI regulation has its own alignment problem: The technical a nd institutional feasibility of dis-...

  26. [34]

    AI S afety Institute approach to evalua- tions

    Innovation Department for Science and Technology. AI S afety Institute approach to evalua- tions. https://www.gov.uk/government/publications/ai-safet y-institute-approach-to-evaluations/ai

  27. [35]

    Artificial Intelligence Safety Institute

    U.S. Artificial Intelligence Safety Institute. The Uni ted States Artificial Intelligence Safety Institute: Vision, Mission, and Strategic Goals. 2024

  28. [36]

    European union: Share in global gross domestic pro duct based on purchasing-power- parity from 2017 to 2027

    IMF. European union: Share in global gross domestic pro duct based on purchasing-power- parity from 2017 to 2027. Statista, 2022. Accessed: Septemb er 6, 2024

  29. [37]

    Leakage in data mining: Formulation, detection, and avoidance

    Shachar Kaufman, Saharon Rosset, Claudia Perlich, and Ori Stitelman. Leakage in data mining: Formulation, detection, and avoidance. ACM Trans. Knowl. Discov. Data, 6(4), dec 2012. ISSN 1556-4681. doi: 10.1145/2382577.238 2579. URL https://doi.org/10.1145/2382577.2382579

  30. [38]

    Accessed: September 5, 2024

  31. [39]

    Holistic evaluation of language models

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipr as, Dilara Soylu, Michihiro Y asunaga, Yian Zhang, Deepak Narayanan, Y uhuai Wu, Ananya Kumar, et al . Holistic evaluation of language models. arXiv preprint arXiv:2211.09110 , 2022

  32. [40]

    Adm inistrative de- lay, red tape, and organizational performance

    Wesley Kauffman, Gabel Taggart, and Barry Bozeman. Adm inistrative de- lay, red tape, and organizational performance. pages 529–5 53, 2019. URL https://doi.org/10.1080/15309576.2018.1474770

  33. [41]

    Y ou need to be spending more money on evals

    Kamil ˙e Lukoši ¯ut˙e. Y ou need to be spending more money on evals. https://kamilelukosiute.com/llms/You+need+to+be+spending+more+money+on+evals, undated. Accessed: September 6, 2024

  34. [42]

    Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b

    Simon Lermen, Charlie Rogers-Smith, and Jeffrey Ladis h. Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b. arXiv preprint arXiv:2310.20624 , 2023

  35. [43]

    Encyclopedia of Evaluation

    Sandra Mathison. Encyclopedia of Evaluation . SAGE Publications, 2004

  36. [44]

    Liao, Rohan Taori, Inioluwa Deborah Raji, and Ludwig Schmidt

    Thomas I. Liao, Rohan Taori, Inioluwa Deborah Raji, and Ludwig Schmidt. Are we learning yet? a meta-review of evaluation failures across machine le arning, 2021

  37. [45]

    Portable evaluation tasks via the metr task stand ard

    METR. Portable evaluation tasks via the metr task stand ard. https://metr.org/blog/2024-02-29-metr-task-standard /, 2024. Accessed: September 12, 2024

  38. [46]

    The Necessity of AI Audit Standards Boards

    David Manheim, Sammy Martin, Mark Bailey, Mikhail Sami n, and Ross Greutzmacher. The Necessity of AI Audit Standards Boards. arXiv preprint arXiv:2404.13060 , 2024

  39. [47]

    State of what art? A call for multi-prompt LLM evaluation

    Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafn a Shahaf, and Gabriel Stanovsky. State of what art? A call for multi-prompt LLM evaluation. Transactions of the Association for Computational Linguistics, 12:933–949, 2024

  40. [48]

    Model details

    Meta. Model details. URL https://github.com/meta-llama/llama3/blob/main/MODE L_CARD.md

  41. [49]

    The O perational Risks of AI in Large-Scale Biological Attacks: Results of a R ed-Team Study

    Christopher A Mouton, Caleb Lucas, and Ella Guest. The O perational Risks of AI in Large-Scale Biological Attacks: Results of a R ed-Team Study. Technical Report RR-A2977-2, RAND Corporation, 202 4. URL https://www.rand.org/pubs/research_reports/RRA2977- 2.html

  42. [50]

    METR. Vivaria. https://vivaria.metr.org/, 2024. Accessed: September 12, 2024

  43. [51]

    Reproducibil- ity and Replicability in Science

    National Academies of Sciences, Engineering, and Medi cine. Reproducibil- ity and Replicability in Science . The National Academies Press, Washing- ton, DC, 2019. ISBN 978-0-309-48616-3. doi: 10.17226/2530 3. URL https://nap.nationalacademies.org/catalog/25303/reproducibility-...

  44. [52]

    Morris, Jr

    John B. Morris, Jr. Injecting the public interest into i nternet standards. 2011. doi: https: //doi.org/10.7551/mitpress/8066.001.0001

  45. [53]

    Preparedness framework (beta), 2023

    OpenAI. Preparedness framework (beta), 2023. URL https://cdn.openai.com/openai-preparedness-framewor k-beta.pdf

  46. [54]

    Reasons to Doubt the Impact of AI Risk Ev aluations

    Gabriel Mukobi. Reasons to Doubt the Impact of AI Risk Ev aluations. 2024. URL https://arxiv.org/pdf/2408.02565. 15

  47. [55]

    Red teaming lang uage models with language models

    Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, R oman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. Red teaming lang uage models with language models. arXiv preprint arXiv:2202.03286 , 2022

  48. [56]

    The prereg- istration revolution

    Brian A Nosek, Charles R Ebersole, Alexander C DeHaven, and David T Mellor. The prereg- istration revolution. Proceedings of the National Academy of Sciences , 115(11):2600–2606, 2018

  49. [57]

    Bender, Amandalynne Pa ullada, Emily Denton, and Alex Hanna

    Inioluwa Deborah Raji, Emily M. Bender, Amandalynne Pa ullada, Emily Denton, and Alex Hanna. AI and the Everything in the Whole Wide World Benchmar k. arXiv preprint arXiv:2111.15366, 2021. URL https://arxiv.org/pdf/2111.15366

  50. [58]

    In ternal and external validity: can you apply research study results to your patients? J Bras Pneumol , 44(3), 2018

    Cecilia Maria Patino and Juliana Carvalho Ferreira. In ternal and external validity: can you apply research study results to your patients? J Bras Pneumol , 44(3), 2018

  51. [59]

    Political standards: Corporate interest, ideology, and le adership in the shaping of accounting rules for the market economy

    Karthik Ramanna. Political standards: Corporate interest, ideology, and le adership in the shaping of accounting rules for the market economy . University of Chicago Press, 2015

  52. [60]

    Organized Uncertainty: Designing a W orld of Risk Management

    Michael Power. Organized Uncertainty: Designing a W orld of Risk Management. 2017

  53. [61]

    Open pro blems in technical ai governance

    Anka Reuel, Ben Bucknall, Stephen Casper, Tim Fist, Lis a Soder, Onni Aarne, Lewis Ham- mond, Lujain Ibrahim, Alan Chan, Peter Wills, et al. Open pro blems in technical ai governance. arXiv preprint arXiv:2407.14981 , 2024

  54. [62]

    Outsider oversight: Designing a third party audit ecosystem for ai governance

    Inioluwa Deborah Raji, Peggy Xu, Colleen Honigsberg, a nd Daniel Ho. Outsider oversight: Designing a third party audit ecosystem for ai governance. I n Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society , pages 557–571, 2022

  55. [63]

    Large Lan- guage Models (GPT) Struggle to Answer Multiple-Choice Ques tions about Code

    Jaromir Savelka, Arav Agarwal, Christopher Bogart, an d Majd Sakr. Large Lan- guage Models (GPT) Struggle to Answer Multiple-Choice Ques tions about Code. https://arxiv.org/abs/2303.08033, 2023. Accessed: October 17, 2024

  56. [64]

    Proactive risk management in a dynamic society

    Jens Rasmussen and Inge Suedung. Proactive risk management in a dynamic society. Swedish Rescue Services Agency, 2000. ISBN 9172530847

  57. [65]

    Model evaluation for extreme risks

    Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljun g, Noam Kolt, et al. Model evaluation for extreme risks. arXiv preprint arXiv:2305.15324 , 2023

  58. [66]

    Position: Technical re- search and talent is needed for effective AI governance

    Anka Reuel, Lisa Soder, Benjamin Bucknall, and Trond Ar ne Undheim. Position: Technical re- search and talent is needed for effective AI governance. In R uslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Sc arlett, and Felix Berkenkamp, ...

  59. [67]

    Interoperability is important for com- petition, consumers, and the economy, 2023

    George Slover. Interoperability is important for com- petition, consumers, and the economy, 2023. URL https://cdt.org/insights/interoperability-is-import ant-for-competition-consumers-the-economy/ 16

  60. [68]

    Structured access: an emerging paradig m for safe AI deployment

    Toby Shevlane. Structured access: an emerging paradig m for safe AI deployment. arXiv preprint arXiv:2201.05159, 2022

  61. [69]

    Beyond privacy trade-offs with structured tra nsparency

    Andrew Trask, Emma Bluemke, Ben Garfinkel, Claudia Ghez zou Cuervas-Mons, and Allan Dafoe. Beyond privacy trade-offs with structured tra nsparency. arXiv preprint arXiv:2012.08347, 2020

  62. [70]

    The Brussel s effect and artificial intelligence: How EU regulation will impact the global AI market

    Charlotte Siegmann and Markus Anderljung. The Brussel s effect and artificial intelligence: How EU regulation will impact the global AI market. arXiv preprint arXiv:2208.12645 , 2022

  63. [71]

    Emerging Processes for Frontier AI Safety

    UK Department of Science, Innovation, and Technology. Emerging Processes for Frontier AI Safety

  64. [72]

    Artificial Intelligence Risk Management Framework

    Elham Tabassi. Artificial Intelligence Risk Management Framework . NIST, 2023

  65. [73]

    Fake alignmen t: Are llms really aligned well? https://arxiv.org/abs/2311.05915s, 2024

    Yixu Wang, Y an Teng, Kexin Huang, Chengqi Lyu, SongyangZhang, Wenwei Zhang, Xingjun Ma, Y u-Gang Jiang, Y u Qiao, and Yingchun Wang. Fake alignmen t: Are llms really aligned well? https://arxiv.org/abs/2311.05915s, 2024. Accessed: October 17, 2024

  66. [74]

    UK AI Safety Institute. Inspect. https://inspect.ai-safety-institute.org.uk/,

  67. [75]

    Accessed: September 12, 2024

  68. [77]

    AI Safety I nstitutes: Can countries meet the challenge? jul 2024

    Alexandre V ariengien and Charles Martinet. AI Safety I nstitutes: Can countries meet the challenge? jul 2024. [Online; accessed CURRENT-DA TE]

  69. [78]

    In Regulating A.I., We May Be Doing Too Much

    Tim Wu. In Regulating A.I., We May Be Doing Too Much. And T oo Little., 11 2023. URL https://www.nytimes.com/2023/11/07/opinion/biden-ai -regulation.html. Ac- cessed: 2024-01-23

  70. [79]

    The icl con sistency test

    Lucas Weber, Elia Bruni, and Dieuwke Hupkes. The icl con sistency test. arXiv preprint arXiv:2312.04945, 2023

  71. [80]

    Sociotech- nical safety evaluation of generative AI systems

    Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arian na Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, et al. Sociotech- nical safety evaluation of generative AI systems. arXiv preprint arXiv:2310.11986 , 2023

  72. [82]

    Holistic safety and responsibility evaluations of advance d AI models

    Laura Weidinger, Joslyn Barnhart, Jenny Brennan, Chri stina Butterfield, Susie Y oung, Will Hawkins, Lisa Anne Hendricks, Ramona Comanescu, Oscar Chan g, Mikel Rodriguez, et al. Holistic safety and responsibility evaluations of advance d AI models. arXiv preprint arXiv:2404.14068, 2024

  73. [84]

    International Scientific Report on the Safety of Advanced AI, 2024

    Bengio Y ohsua, Privitera Daniel, Besiroglu Tamay, Bom masani Rishi, Casper Stephen, Choi Y ejin, Goldfarb Danielle, Heidari Hoda, Khalatbari Leila, Longpre Shayne, et al. International Scientific Report on the Safety of Advanced AI, 2024. 17

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.