Pith. sign in

REVIEW 4 major objections 4 minor 15 references

Expert Survey: AI Reliability & Security Research Priorities

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper claims that a 53-expert survey across 105 technical areas produces the first data-driven ranking of AI reliability and security research, placing capability forecasting and evaluations at the top.

desk verdict First quantitative expert ranking of 105 technical AI R&S subareas; useful as a directional signal, but the specific numeric order—topped by an n=3 subarea—should not be treated as a funding guide without error bars or raw data. read the letter →

arxiv 2505.21664 v1 pith:634QWUMQ submitted 2025-05-27 cs.CY cs.AI

classification cs.CYcs.AI
keywords AIreliabilitysecurityexpertsurveyresearchprioritiespromisescorecapabilityevaluationsmulti-agentsafetyemergencescaling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

To guide funding choices in AI reliability and security, the paper reports a structured survey of 53 specialists who rated 105 technical research sub-areas on importance and tractability. The central result is a promise ranking, defined as importance times tractability, topped by emergence and task-specific scaling patterns (promise 21.25) and by evaluations of dangerous capabilities such as CBRN, cyber, deception, and agent oversight. The authors read the ranking as evidence that the highest-leverage near-term work lies in practical evaluation, monitoring, and forecasting rather than theoretical frameworks, with multi-agent interactions as an emerging risk area. If this ranking reflects the field's genuine priorities, it gives funders and policymakers a directly actionable list for allocating marginal research dollars.

What carries the argument

The load-bearing object is the promise score: for each of 105 sub-areas, the mean importance rating multiplied by the mean tractability rating, computed from expert Likert responses, with a threshold of at least three raters and mean ratings of at least 4 on both dimensions for inclusion in the top list. The promise score is the mechanism that converts qualitative expert opinion into a single comparable number for resource allocation, and it depends jointly on the taxonomy, the rating definitions (severe harm; $10M over two years), and the self-selected sample.

What would settle it

Re-run the survey on a larger, pre-registered, stratified sample (e.g., 200+ experts drawn evenly from the 20 categories) and compute promise scores with the same definitions; if the Spearman rank correlation with the original top 15 is below about 0.5, the specific ordering is not stable enough to guide funding. A second check is outcome-based: track two years of funding at $10M per chosen sub-area and see whether the areas rated most tractable actually produce measurable advances.

Watch

Extended reading notes

Core claim

The paper claims to be the first to produce a comparative, data-driven ranking of technical AI reliability and security research directions across a comprehensive taxonomy. Using a 1–5 Likert scale, experts rated each sub-area's importance (would resolving it significantly reduce severe AI harms) and tractability (would roughly $10M over two years yield measurable progress); the product of the mean ratings gives a promise score. Fifteen sub-areas scored at least 4 on both dimensions, led by "Emergence and task-specific scaling patterns" (5.00 importance, 4.25 tractability). Nine of the top fifteen emphasize evaluation, detection, or monitoring, while all multi-agent interaction sub-areas ranked in the top 30; applied security engineering and deep-model understanding were rated important but not tractable at the $10M scale.

Load-bearing premise

The load-bearing premise is that the average of as few as three Likert-scale ratings from a self-selected sample of 53 experts is a stable estimate of what the broader AI reliability and security community would say about each sub-area.

Editorial extensions

If this is right

  • If the ranking is right, funders should put near-term money into dangerous-capability evaluations, scalable oversight of LLM agents, and multi-agent testbeds before expanding theoretical alignment research.
  • The top-ranked area, emergence and task-specific scaling patterns, would be a core target for forecasting infrastructure so that safety responses are ready before new capabilities appear.
  • The high ranking of multi-agent security and metrics suggests that models interacting with models should be treated as a distinct risk vector in evaluation regimes.
  • Areas such as access control, supply-chain integrity, and confidential computing should be funded as multi-year, larger-scale programs rather than $10M two-year bets.
  • Independent evaluation capacity should be expanded, since experts flagged dangerous-capability evaluations as undercapitalized despite existing work at frontier labs and third-party evaluators.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One editorial extension: the ranking says where marginal dollars have the highest expected payoff, but it does not measure current funding levels; the paper's "undercapitalized" phrasing is inferred from expert comments rather than from a funding-flow analysis.
  • A direct test of stability would be to re-administer the survey with a larger, pre-registered, stratified sample; if the top 15 promise scores shift materially, the correct reading is time-bound expert opinion rather than a durable priority list.
  • The taxonomy itself could be reused as a shared map for future elicitation, but its technical-only scope excludes governance, privacy, and fairness interventions, so the ranking should not be read as covering the full portfolio of AI risk reduction.
  • One citation in the discussion ("Hammond et al. 2025") has no matching bibliography entry, so the supporting evidence for the multi-agent risk claim should be checked before relying on it.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript reports an expert survey of 53 AI reliability and security researchers, who rated subsets of 105 technical sub-areas (in 20 categories) on 5-point Likert scales for importance and tractability. The authors compute a 'promise score' as the product of the two means and present a ranked list of the most promising sub-areas, topped by 'Emergence and task-specific scaling patterns' (promise 21.25). The paper argues that this ranking provides a data-driven basis for allocating marginal research funding, and it derives policy recommendations for funders and policymakers.

Significance. If the ranking were statistically robust, it would be a useful complement to existing qualitative taxonomies and would help prioritize evaluation, monitoring, and multi-agent research. The paper's main strengths are the transparent taxonomy (Appendix A), the complete table of results (Appendix C), the explicit listing of excluded sub-areas (Appendix D), and the detailed qualitative discussion of four example sub-areas. The finding that evaluation-focused topics dominate the top ranks, and that security implementation topics are high-importance but low-tractability, is a plausible and potentially valuable qualitative theme. However, the central ranking and its headline claims are built on extremely small samples for the top entries (n=3 for importance of the #1 area), with no uncertainty quantification or test of ranking stability; this materially limits the evidentiary value of the specific order presented.

major comments (4)
  1. [Results, Most Promising Sub-Areas (p.12); Methods, Data Processing (p.11)] The top-ranked sub-area, 'Emergence and task-specific scaling patterns', has n=3 for importance and n=4 for tractability, which is the minimum admissible sample after the paper's exclusion rule (n>2). A single respondent changing a rating by one Likert point changes the mean by 0.33, and the promise scores of the top four areas differ by less than 1.1 (21.25, 20.22, 20.19, 19.70). The paper provides no confidence intervals, standard errors, bootstrap, or sensitivity analysis for the ranking. Since the paper's central claim is a specific ranking, and the abstract and Executive Summary lead with the #1 sub-area, the absence of any uncertainty quantification is a load-bearing gap. The Limitations section (p.20) states that the sample size 'limits the survey's reliability' and that results should be 'directional', but the headline presentation does not convey that uncertainty.
  2. [Methods, Sample (p.10); Limitations (p.20)] The recruitment strategy preferentially invited the first authors of the publications cited as exemplars for each sub-area, and respondents were instructed to rate only the categories and sub-areas where they had personal expertise. This design creates a self-selection channel: experts rate their own research areas, potentially inflating the importance and tractability of those areas. The Limitations section acknowledges that 'the anonymous nature of the survey prevents confirming if experts were biased toward their own research areas,' but the paper neither tests nor adjusts for this bias. If the bias is present, the relative ranking of sub-areas is confounded by the recruitment process, and areas whose exemplar authors were more responsive or more self-interested will rise in the ranking. This is a load-bearing concern for the comparative ranking claim.
  3. [Methods, Data Processing (p.11); Appendix C] The paper computes means of 5-point ordinal Likert responses and then multiplies these means to form a 'promise score.' Treating ordinal categories as an interval scale requires justification (or a robustness check against alternative treatments, such as medians or ordinal models). The paper provides no such justification, and for n=3 the distinction between 'Agree' and 'Strongly Agree' is particularly consequential. The product of means amplifies differences in the underlying scale, so the promise scores may largely reflect the chosen arithmetic rather than a robust property of expert opinion. At minimum, the paper should report the distribution of responses and test whether the ranking is sensitive to the interval-scale assumption.
  4. [Abstract and Executive Summary vs. Limitations (p.20)] There is a mismatch between the strength of the claims made in the abstract and Executive Summary and the caveats stated in the Limitations section. The abstract asserts that the study 'quantifies expert priorities' and 'produces a data-driven ranking of their potential impact,' and the Executive Summary presents specific promise scores and a list of 'Highest-Ranking Research Areas' without the 'directional' caveat. The Limitations section, however, says that the sample size 'limits the survey's reliability' and that results 'should be read as directional.' This internal inconsistency makes it difficult for a reader to know how much weight to place on the specific ordering. The authors should either soften the headline claims to focus on the qualitative themes (e.g., evaluation and monitoring are widely seen as promising) or provide additional statistical support for the specific ranking.
minor comments (4)
  1. [Executive Summary (p.1) vs. Results (p.12) and Appendix C (p.47)] The Executive Summary lists 'Cyber-capability evaluations' while the Results table and Appendix C use 'Cyber evaluations'; this naming inconsistency should be corrected.
  2. [Discussion, Four Example Sub-Areas (p.16)] In the example for 'Emergence and Task-Specific Scaling Patterns', the text uses 'T = 4.25, I = 5' with one decimal for tractability and an integer for importance; elsewhere the paper reports two decimals. Standardize the rounding in narrative references to the survey results.
  3. [Appendix B, Introductory page (p.42)] The survey's stated title is 'AI Assurance and Reliability Research Priorities' in the introductory page, while the paper's title and running text use 'AI Reliability & Security Research Priorities'. Please harmonize the survey instrument title with the paper's terminology.
  4. [Appendix D (p.52)] The heading 'Sub-areas excluded due to insufficient response' uses a singular 'response' when multiple responses are meant; consider 'insufficient responses' for grammatical consistency.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the promise ranking is a direct summary of survey ratings, not a derived prediction.

full rationale

This paper makes no derivational claims that reduce to their inputs. The central result, a promise score equal to mean importance times mean tractability, is defined directly from the survey data (Results, p.12), so there is nothing to 'predict' and no fitted parameter is renamed as a prediction. The taxonomy is grounded in external literature (Anwar et al. 2024) and expert consultation, and the prior IAPS report (Delaney, Guest, and Williams 2024) is cited only as related work, not as the basis for any ranking. Recruitment of first authors of exemplar papers is a sampling strategy, not a derivation: the ratings are the experts' own judgment, and the paper explicitly treats the ranking as directional rather than as a theorem. The Limitations section (p.20) candidly states that the sample 'limits the survey's reliability' and that results 'should be read as directional'; acknowledging weakness is not circularity. No self-definitional step, uniqueness import, or ansatz-by-citation appears. The robustness concerns raised by small n (e.g., n=3 for the top-ranked area) are statistical validity concerns, not circularity, and therefore do not affect the circularity score.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No fitted parameters or invented entities. The central analysis rests on choices about how to define importance and tractability, how to aggregate ratings multiplicatively, and how to treat self-selected expert opinions as representative. These are modeling assumptions rather than fitted constants.

assumptions (4)
  • domain assumption Likert-scale ratings (1-5) are treated as interval data and averaged.
    Methods/Data Processing: 'we calculated mean scores for importance and tractability' from agreement levels; equal spacing between 'Strongly Disagree' and 'Strongly Agree' is assumed but not justified.
  • ad hoc to paper Research promise equals importance multiplied by tractability.
    Methods/Data Processing and Results: 'Promise = Importance × Tractability (max = 25)'; multiplicative combination is chosen without comparison to alternative aggregation functions.
  • domain assumption Self-selected expert respondents are representative of relevant expertise for each rated subarea.
    Methods/Sample and Limitations: recruitment targeted first authors of cited papers and experts were encouraged to rate only areas of personal expertise; 'the anonymous nature of the survey prevents confirming if experts were biased toward their own research areas'.
  • domain assumption The 105-subarea taxonomy is sufficiently comprehensive and unbiased for prioritization.
    Methods/Survey Design: taxonomy drawn from Anwar et al. (2024) and refined by consultations; authors note 'this selection is neither exhaustive nor perfect' and respondents suggested 16 additional unique research directions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Expert Survey: AI Reliability & Security Research Priorities." pith.science (2026). https://pith.science/paper/634QWUMQ

@misc{pith2026250521664,
  author       = {Pith},
  title        = {Pith review of: Expert Survey: AI Reliability & Security Research Priorities},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/634QWUMQ}},
  note         = {Machine review of arXiv:2505.21664}
}
read the original abstract

Our survey of 53 specialists across 105 AI reliability and security research areas identifies the most promising research prospects to guide strategic AI R&D investment. As companies are seeking to develop AI systems with broadly human-level capabilities, research on reliability and security is urgently needed to ensure AI's benefits can be safely and broadly realized and prevent severe harms. This study is the first to quantify expert priorities across a comprehensive taxonomy of AI safety and security research directions and to produce a data-driven ranking of their potential impact. These rankings may support evidence-based decisions about how to effectively deploy resources toward AI reliability and security research.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

15 extracted references · 3 canonical work pages

  1. [4]

    RedCode: Risky Code Execution and Generation Benchmark for Code Agents

    “RedCode: Risky Code Execution and Generation Benchmark for Code Agents.” arXiv. https://doi.org/10.48550/arXiv.2411.07781. Guo, Wei, and Aylin Caliskan. 2021. “Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases.” In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, 12...

  2. [5]

    Self-Destructing Models: Increasing the Costs of Harmful Dual Uses of Foundation Models

    “Self-Destructing Models: Increasing the Costs of Harmful Dual Uses of Foundation Models.” arXiv. https://doi.org/10.48550/arXiv.2211.14946. Hendrycks, Dan, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2023. “Aligning AI With Shared Human Values.” arXiv. https://doi.org/10.48550/arXiv.2008.02275. Henzinger, Thomas...

  3. [6]

    Datamodels: Predicting Predictions from Training Data

    “Datamodels: Predicting Predictions from Training Data.” arXiv. https://doi.org/10.48550/arXiv.2202.00622. Irving, Geoffrey, Paul Christiano, and Dario Amodei. 2018. “AI Safety via Debate.” arXiv. https://doi.org/10.48550/arXiv.1805.00899. Ismail, Aya Abdelsalam, Héctor Corrada Bravo, and Soheil Feizi. 2021. “Improving Deep Learning Interpretability by Sal...

  4. [7]

    A Watermark for Large Language Models

    “A Watermark for Large Language Models.” arXiv. https://doi.org/10.48550/arXiv.2301.10226. Kirk, Hannah Rose, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, et al. 2024. “The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multic...

  5. [10]

    Towards A Proactive ML Approach for Detecting Backdoor Poison Samples

    “Towards A Proactive ML Approach for Detecting Backdoor Poison Samples.” arXiv. https://doi.org/10.48550/arXiv.2205.13616. Qi, Xiangyu, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson

  6. [11]

    Fine-Tuning Aligned Language Models Compromises Safety, Even When Users Expert Survey: AI Reliability & Security Research Priorities | 66 Do Not Intend To!

    “Fine-Tuning Aligned Language Models Compromises Safety, Even When Users Expert Survey: AI Reliability & Security Research Priorities | 66 Do Not Intend To!” arXiv. https://doi.org/10.48550/arXiv.2310.03693. Qu, Qu, Zheng Ma, Anders Clausen, and Bo Nørregaard Jørgensen. 2021. “A Comprehensive Review of Machine Learning in Multi-Objective Optimization.” In...

  7. [12]

    Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 Small

    “Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 Small.” arXiv. https://doi.org/10.48550/arXiv.2211.00593. Wang, Yihan, Jatin Chauhan, Wei Wang, and Cho-Jui Hsieh. 2023. “Universality and Limitations of Prompt Tuning.” Advances in Neural Information Processing Systems 36 (December):75623–43. Ward, Francis Rhys, Francesco...

  8. [13]

    Instructional Fingerprinting of Large Language Models

    “Instructional Fingerprinting of Large Language Models.” arXiv. https://doi.org/10.48550/arXiv.2401.12255. Yang, Fan, Wenxuan Zhou, Zuxin Liu, Ding Zhao, and David Held. 2024. “Reinforcement Learning in a Safety-Embedded MDP with Trajectory Optimization.” arXiv. https://doi.org/10.48550/arXiv.2310.06903. Yang, Jiachen, Ang Li, Mehrdad Farajtabar, Peter Su...

Show all 15 references
  1. [14]

    Low-Resource Languages Jailbreak GPT-4

    https://social-dilemmas.github.io/. Yong, Zheng-Xin, Cristina Menghini, and Stephen H. Bach. 2024. “Low-Resource Languages Jailbreak GPT-4.” arXiv. https://doi.org/10.48550/arXiv.2310.02446. Expert Survey: AI Reliability & Security Research Priorities | 70 Yuan, Tongxin, Zhiwe...

  2. [15]

    Towards Fair Disentangled Online Learning for Changing Environments

    “Towards Fair Disentangled Online Learning for Changing Environments.” arXiv. https://doi.org/10.48550/arXiv.2306.01007. Zhi-Xuan, Tan, Micah Carroll, Matija Franklin, and Hal Ashton. 2024. “Beyond Preferences in AI Alignment.” Philosophical Studies, November. https://doi.org/...

  3. [2021]

    A Survey on Bias and Fairness in Machine Learning

    “A Survey on Bias and Fairness in Machine Learning.” ACM Comput. Surv. 54 (6): 115:1-115:35. https://doi.org/10.1145/3457607. Merrill, William, and Ashish Sabharwal. 2023. “The Parallelism Tradeoff: Limitations of Log-Precision Transformers.” Transactions of the Association for...

  4. [2022]

    What Does It Mean for a Language Model to Preserve Privacy?

    “What Does It Mean for a Language Model to Preserve Privacy?” arXiv. https://doi.org/10.48550/arXiv.2202.05520. Brown-Cohen, Jonah, Geoffrey Irving, and Georgios Piliouras. 2023. “Scalable AI Safety via Doubly-Efficient Debate.” arXiv. https://doi.org/10.48550/arXiv.2311.14125. B...

  5. [2023]

    Deep Reinforcement Learning from Human Preferences

    “Deep Reinforcement Learning from Human Preferences.” arXiv. https://doi.org/10.48550/arXiv.1706.03741. Christiano, Paul, Buck Shlegeris, and Dario Amodei. 2018. “Supervising Strong Learners by Amplifying Weak Experts.” arXiv. https://doi.org/10.48550/arXiv.1810.08575. Clymer,...

  6. [2024]

    LLM-Deliberation: Evaluating LLMs with Interactive Multi-Agent Negotiation Game,

    https://www.cnas.org/publications/reports/secure-governable-chips. Abdelnabi, Sahar, Amr Gomaa, Sarath Sivaprasad, Lea Schönherr, and Mario Fritz. 2023. “LLM-Deliberation: Evaluating LLMs with Interactive Multi-Agent Negotiation Game,” October. https://openreview.net/forum?id=...

  7. [2025]

    How the U.S. Public and AI Experts View Artificial Intelligence

    “How the U.S. Public and AI Experts View Artificial Intelligence.” Pew Research Center (blog). April 3, 2025. https://www.pewresearch.org/internet/2025/04/03/how-the-us-public-and-ai-experts-vi ew-artificial-intelligence/. Perez, Ethan, Sam Ringer, Kamile Lukosiute, Karina Nguye...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.