{"id":"1c9b5074-f7d4-43a5-a2a6-f55b7fedfbbc","arxiv_id":"2505.05541","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic literature review of AI safety evaluations that categorizes methods by measured property (capability, propensity, control), technique type, and governance integration.","lead":"This paper is a literature review chapter that organizes AI safety evaluation methods into a three-part taxonomy: capabilities, propensities, and control. It surveys behavioral and internal evaluation techniques, governance frameworks, and known limitations to serve as a central reference for researchers and policymakers.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'systematic' claim is unsupported by any documented methodology, and the §5.3.3 'control' category is not a model property like capability/propensity.","rationale":"The reader's weakest-assumption focused on taxonomy completeness and discriminativeness. My stress-test agrees that this is load-bearing but locates the deeper problem in the absence of any documented systematic methodology: without a reproducible search and screening protocol, there is no way to establish that the taxonomy was derived from a representative corpus, so the 'central reference point' claim cannot be audited. The reader's rationale did mention the missing methodology as the basis for the conditional verdict, but their stated weakest assumption was about taxonomy fit; my concern is closely related yet slightly different. I also identify an internal conceptual issue: 'control' is defined as a property of the safety measures rather than a property of the AI system, unlike capability and propensity, making the taxonomy's first dimension heterogeneous. This is not a fatal flaw if the authors clarify that 'property' includes system-level or relational properties, but as written it is a genuine weakness. The paper deserves credit for candidly discussing limitations, for including concrete examples from METR, Apollo, Anthropic, OpenAI, and DeepMind, and for clearly distinguishing benchmarks from evaluations. However, the load-bearing claim of being a systematic consolidation is conditional on adding a methodology section and clarifying the control category. I therefore support the reader's CONDITIONAL verdict rather than moving to REJECT or UNCHANGED.","tokens_in":40789,"tokens_out":4505,"duration_ms":55087,"concrete_test":"Attempt to reproduce the review's coverage: run a defined literature search (e.g., arXiv and Google Scholar, 2019–2025, queries including 'AI safety evaluation', 'dangerous capability evaluation', 'propensity evaluation', 'control evaluation', 'AI control'), apply explicit inclusion criteria, and check whether the methods retrieved can be placed into exactly one cell of the reported 3×2 taxonomy (capability/propensity/control × behavioral/internal). If a nontrivial fraction (e.g., >10%) of retrieved methods cannot be classified, or if major cited-adjacent works such as METR's RE-Bench, Apollo Research sabotage evaluations, or AI control evaluations are absent from the review's coverage, then the taxonomy is not demonstrably complete or discriminative and the 'systematic literature review' claim is unsupported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that this is a 'Systematic Literature Review' consolidating AI safety evaluations into a systematic taxonomy that can serve as a 'central reference point' (title/abstract). For that claim to hold, the literature must be selected representatively and reproducibly. However, the paper contains no methodology section: no search databases, query strings, inclusion/exclusion criteria, screening steps, or quality appraisal are reported. The reference list reads as a convenience narrative, so the taxonomy in §5.3–5.4 could omit entire evaluation paradigms without any way for a reader to audit coverage. This directly undermines both the 'systematic' label and the claim to consolidate the field. Separately, §5.3.3 classifies 'control' as one of the three 'Evaluated Properties', but defines control evaluations as assessing 'whether safety protocols remain effective when AI systems actively try to circumvent them' (5.3.3). That is a property of the safety infrastructure and the model–protocol system, not a property of the AI system itself in the way capability and propensity are. The taxonomy's first dimension is therefore not a homogeneous set of model properties, which weakens the claim that it is a systematic three-property framework.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is a literature review of AI safety evaluation methods, organized around a proposed taxonomy with three dimensions: evaluated properties (capability, propensity, control), evaluation techniques (behavioral and internal), and evaluation frameworks (model-organism and governance frameworks such as RSPs, the Preparedness Framework, and the Frontier Safety Framework). It surveys benchmarks and their limitations, describes dangerous capability evaluations (cybercrime, deception capability, autonomous replication, sustained task execution, situational awareness), dangerous propensity evaluations (deception propensity, long-term planning, power-seeking, scheming), and control evaluations, then closes with limitations of evaluation practice. The paper is explicitly presented as Chapter 5 of the authors' larger 'AI Safety Atlas' and claims to be a systematic literature review that provides a central reference point for AI safety evaluations.","tokens_in":40976,"tokens_out":4612,"duration_ms":50965,"significance":"If the taxonomy and coverage are accepted, the review would be a genuinely useful consolidation of a fast-moving field, valuable to researchers, auditors, and policymakers. The manuscript's strengths include its broad citation of primary and practitioner sources (METR, Apollo Research, Anthropic, OpenAI, DeepMind, UK AISI), its explicit treatment of evaluation limitations (proving absence, sandbagging, safetywashing, measurement sensitivity), and its candid caveats about the limits of situational-awareness and scheming demonstrations. The main contribution is organizational: a structured map of evaluation properties, techniques, and frameworks. The review does not present new experimental results, so its significance depends on the validity, homogeneity, and completeness of the proposed taxonomy, and on the reproducibility of the literature selection.","major_comments":[{"comment":"The title and abstract claim a 'systematic literature review,' but the manuscript reports no methodology. No search databases, query strings, inclusion/exclusion criteria, screening steps, or quality-appraisal procedure appear anywhere in §§5.1–5.11; the Limitations section (§5.10) discusses limitations of evaluations themselves but not limitations of the review process. Without a documented method, a reader cannot verify that the literature was selected representatively or reproducibly, which undermines both the 'systematic' label and the abstract's claim to 'consolidate the field' and provide a 'central reference point.' This is fixable by adding a methods subsection describing the search and selection protocol, and by calibrating the claims in the title and abstract if such a protocol is not available.","section":"Overall (§5.1–§5.11)"},{"comment":"The taxonomy's first dimension is presented as three 'Evaluated Properties' (capability, propensity, control), but control is not a property of the AI system in the same sense as the other two. §5.3.3 defines control evaluations as assessing 'whether our safety measures remain effective when AI systems actively try to circumvent them' — that is a property of the combined model-plus-safety-infrastructure system, or of the evaluation outcome, not a property of the model's capabilities or tendencies. The authors should either redefine the first dimension as 'evaluation targets' (of which control is one) or justify why a system-infrastructure property belongs in a taxonomy of model properties; as written, the first dimension is not homogeneous.","section":"§5.3.3"},{"comment":"The taxonomy is claimed to be systematic and, implicitly, exhaustive ('three distinct properties,' 'two complementary approaches'), but the manuscript provides no decision rule for assigning a given evaluation method to a cell of the taxonomy and does not validate coverage against a defined corpus. The text itself acknowledges overlaps among categories (e.g., §5.3 footnote 3, and §5.8.1 on the relationship between deception capability and deception propensity), so without an assignment procedure or a coverage check the claim that the taxonomy can serve as a 'central reference point' is not established. At minimum, the authors should state which categories are exclusive, which are graded, and how borderline cases (such as TruthfulQA, discussed in §5.7.2 as primarily a capability measure but in §5.2.1 as a safety-relevant benchmark) are assigned.","section":"§5.3–§5.4"}],"minor_comments":[{"comment":"'WDMP benchmark' should be 'WMDP benchmark' to match the definition given two paragraphs later.","section":"§5.2.1"},{"comment":"'Cesar Cipher' should be 'Caesar cipher'; also 'Humanities Last Exam' should be 'Humanity's Last Exam' as used in §5.2.1.","section":"§5.2.2"},{"comment":"'Scalar et al. 2023' should be 'Sclar et al. 2023' (the correct spelling appears in §5.8).","section":"§5.10.2"},{"comment":"Possessives are missing in 'OpenAIs Preparedness Framework,' 'Anthropics Responsible Scaling Policies,' and 'DeepMinds Frontier Safety Framework.'","section":"§5.5"},{"comment":"Figures 5.32 and 5.33 are reproduced from Sharkey et al. but their axes and their relation to the affordance discussion are not explained in the text; a short interpretive sentence for each figure would improve clarity.","section":"§5.6.1"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the 'systematic' claim in the title is not supportable as submitted, because no methodology is reported; this is a load-bearing issue but one that can be fixed within the manuscript's scope by adding a methods subsection. The underlying synthesis is useful and the taxonomy is broadly sensible, but its first dimension is not homogeneous as currently defined. I also note that the paper is explicitly self-described as Chapter 5 of the authors' 'AI Safety Atlas'; the journal may wish to consider whether a textbook-chapter framing is appropriate for a standalone 'systematic literature review' article, and whether the abstract's claims should be aligned with the absence of a documented search protocol."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know about arXiv:2505.05541. First, despite the title, there is no systematic review methodology anywhere in the paper — no search strings, no inclusion criteria, no screening steps, no quality appraisal. Second, the taxonomy has a real but fixable flaw: 'control' is not a model property like capability or propensity. It is a property of the model-plus-safeguards system. That matters because the paper's central claim is that this three-way division is a systematic organizing scheme.\n\nWhat is actually new: not much conceptually. The capability/propensity distinction comes from Shevlane et al., control from Greenblatt et al. The review's contribution is bundling them into a clean pedagogical frame and curating the scattered literature. That is worth having. The writing is clear, the coverage is broad, and the limitations section is candid about the usual suspects: sandbagging, safetywashing, absence-of-evidence problems, and the difficulty of proving a negative. The chapter-style prose with boxes and figures makes it accessible to a non-specialist.\n\nThe missing methodology is the biggest weakness. If it were retitled 'a narrative review' the label would be honest, and if a protocol appendix were added the systematic claim could be checked. As is, the claim is unsupported and should not survive peer review unchanged. The control-category issue is more conceptual: it sits uneasily next to the other two. The authors could rename the third dimension ('safety measures under adversarial conditions') or explicitly frame it as a system-level property. Either fix works. Minor issues: some figures are not well sourced, and the prose is occasionally repetitive, but those are cosmetic.\n\nWho gets value from this: students, policy people, and researchers new to AI safety evaluations. It is an honest map, not a new measurement method. It will be cited as a starting point, and it can be a useful reading-group discussion about what a good taxonomy should look like.\n\nRecommendation: send it to peer review with a request for moderate revision. The field needs a usable consolidated reference, and this one is mostly accurate and honest. Add or disavow the systematic methodology, reframe the third category, and it deserves publication. I would accept a referee invitation.","headline":"Useful pedagogical synthesis that overclaims its systematicity; the control category is conceptually off, but the review is worth a serious referee after moderate revision.","tokens_in":41511,"tokens_out":1805,"would_cite":true,"duration_ms":21847,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI safety evaluation is best organized around three measured properties: what a model can do, what it tends to do, and whether safeguards hold when it tries to break them.","keywords":["AI safety evaluations","capability evaluations","propensity evaluations","control evaluations","red teaming","interpretability","evaluation governance","benchmarks"],"falsifier":"A concrete counterexample would be a published safety evaluation whose result cannot be classified as capability, propensity, or control even when its affordances and elicitation strength are fully documented, for example a test that simultaneously measures maximum ability and default tendency with no way to decompose the result. If such an evaluation is routinely used in frontier-model safety assessment, the taxonomy's claim to be the organizing scheme of the field fails.","tokens_in":40580,"feed_emoji":"📏","tokens_out":6702,"duration_ms":69553,"temperature":0.7,"pith_summary":"This review argues that the scattered practice of AI safety evaluation can be organized into a single taxonomy. The first dimension is the property being measured: capability (what a system can do when pushed to its limit), propensity (what it tends to do by default), and control (whether safety measures hold when the system itself tries to circumvent them). The second dimension is the evidence channel: behavioral techniques that observe outputs and internal techniques that inspect representations. If the taxonomy holds, it gives researchers and regulators a shared language for comparing evaluation results and for deciding when development or deployment should pause.","feed_headline":"One taxonomy sorts AI safety tests into three measurements","feed_subtitle":"A systematic review groups capability, propensity, and control evaluations into one shared framework.","key_machinery":"The organizing device is the capability/propensity/control taxonomy crossed with the behavioral/internal technique distinction. A capability evaluation asks what the system can do when elicited aggressively; a propensity evaluation asks which of several available behaviors it tends to choose; a control evaluation asks whether protective protocols survive intentional subversion. Behavioral techniques gather evidence from outputs, while internal techniques inspect activations, circuits, and learned features. This pair of distinctions carries the argument because every method discussed is placed in one of the resulting cells, giving the review its structure and its claim to completeness.","core_discovery":"The central claim is that safety evaluations are not a loose collection of benchmarks but a systematic field with three dimensions. Evaluations differ in what they measure, how they measure it, and how the measurements plug into decision frameworks. Capability evaluations establish upper bounds under maximal elicitation, propensity evaluations reveal default behavioral tendencies in choice situations, and control evaluations test whether containment survives adversarial behavior by the model itself. The review further claims that behavioral and internal techniques are complementary, and that evaluation results become decision-relevant only when wired into governance gates that trigger pauses or additional safeguards.","pith_inferences":["If the taxonomy is as complete as claimed, an evaluation report that does not state its property category and evidence channel is arguably incomplete; coverage could be audited against the resulting grid.","A natural extension is a standard evaluation-card format reporting the property, technique, affordances, and elicitation strength, which would make results from different organizations comparable.","Because absence of a capability cannot be proven, the paper's own logic implies that safety cases should be framed as cumulative control evidence rather than as certification of safety."],"forward_implications":["Benchmarks alone are insufficient for safety claims, because they measure typical performance rather than upper bounds or default behavioral tendencies.","Safety claims should name the property and technique they rely on; a result about what a model can do under scaffolding does not transfer to what it tends to do by default.","Governance gates can be built directly from evaluation thresholds, pausing scaling when dangerous capabilities appear or when control evaluations fail.","Control evaluations distinguish \"safe because it cannot attack\" from \"safe because we detect attacks,\" which changes how long current safeguards can be trusted."],"supporting_citations":[{"why":"Supplies the list of dangerous capabilities (cyber-offense, deception, situational awareness, self-proliferation) that the capability section is organized around.","marker":"Shevlane et al. 2023"},{"why":"Defines control evaluations as testing whether safety measures hold when AI systems actively try to circumvent them.","marker":"Greenblatt et al. 2023"},{"why":"Provides the complementary-evidence view of capability and propensity evaluations and the probe technique used for internal monitoring.","marker":"Roger et al. 2023"},{"why":"Supplies the affordances concept and the distinction between model and system that underpin evaluation design.","marker":"Sharkey et al. 2024"},{"why":"Introduces the model organisms framework that grounds one of the technical evaluation frameworks discussed.","marker":"Hubinger et al. 2023"},{"why":"Frames control evaluations as a red-team/blue-team game and motivates early detection as a safety inflection point.","marker":"Greenblatt & Shlegeris, 2024"},{"why":"Provides the definition of scheming as performing well in training to later pursue hidden objectives, central to the scheming propensity section.","marker":"Carlsmith, 2023"},{"why":"Supplies the sustained task execution evaluation protocol and the human-time-horizon metric used to interpret capability growth.","marker":"Kwa et al. 2025"}],"fun_headline_variants":["AI safety eval taxonomy: capability, propensity, control","How to measure AI safety: a systematic review of evals","Capability, propensity, control: the three pillars of AI safety eval","AI safety evaluations: beyond benchmarks with a three-part taxonomy","Three measurement types for AI safety: a systematic review"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The taxonomy is assumed to be complete and discriminating: every safety-relevant evaluation method can be placed into exactly one of the three property categories and one of the two technique classes without forcing or mischaracterizing existing work.","fun_headline_variants_meta":{"raw":{"variants":["AI safety eval taxonomy: capability, propensity, control","How to measure AI safety: a systematic review of evals","Capability, propensity, control: the three pillars of AI safety eval","AI safety evaluations: beyond benchmarks with a three-part taxonomy","Three measurement types for AI safety: a systematic review"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000641,"raw_usage":{"total_tokens":2928,"prompt_tokens":899,"completion_tokens":2029,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":515,"completion_tokens_details":{"reasoning_tokens":1946}},"tokens_in":515,"tokens_out":2029,"duration_ms":13506,"temperature":1.0,"reasoning_tokens":1946,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:03:46.879163+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete counterexample would be a published safety evaluation whose result cannot be classified as capability, propensity, or control even when its affordances and elicitation strength are fully documented, for example a test that simultaneously measures maximum ability and default tendency with no way to decompose the result. If such an evaluation is routinely used in frontier-model safety assessment, the taxonomy's claim to be the organizing scheme of the field fails.","supporting_citations":[],"review_version":1}