Pith. sign in

REVIEW 4 cited by

Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.05541 v1 pith:GQHJFRVR submitted 2025-05-08 cs.AI

Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods

classification cs.AI
keywords safetycapabilitiesevaluationsreviewlikeliteraturemeasurealongside
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

As frontier AI systems advance toward transformative capabilities, we need a parallel transformation in how we measure and evaluate these systems to ensure safety and inform governance. While benchmarks have been the primary method for estimating model capabilities, they often fail to establish true upper bounds or predict deployment behavior. This literature review consolidates the rapidly evolving field of AI safety evaluations, proposing a systematic taxonomy around three dimensions: what properties we measure, how we measure them, and how these measurements integrate into frameworks. We show how evaluations go beyond benchmarks by measuring what models can do when pushed to the limit (capabilities), the behavioral tendencies exhibited by default (propensities), and whether our safety measures remain effective even when faced with subversive adversarial AI (control). These properties are measured through behavioral techniques like scaffolding, red teaming and supervised fine-tuning, alongside internal techniques such as representation analysis and mechanistic interpretability. We provide deeper explanations of some safety-critical capabilities like cybersecurity exploitation, deception, autonomous replication, and situational awareness, alongside concerning propensities like power-seeking and scheming. The review explores how these evaluation methods integrate into governance frameworks to translate results into concrete development decisions. We also highlight challenges to safety evaluations - proving absence of capabilities, potential model sandbagging, and incentives for "safetywashing" - while identifying promising research directions. By synthesizing scattered resources, this literature review aims to provide a central reference point for understanding AI safety evaluations.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI

    cs.AI 2026-04 conditional novelty 6.0

    Seven frontier LLMs showed little spontaneous power-seeking in a Linux sysadmin sandbox (bias-corrected rates roughly 0-5%), but showed more specification gaming and resistance to goal modification.

  2. Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users

    cs.AI 2025-12 unverdicted novelty 6.0

    LLM safety evaluations for personal advice must test responses against diverse user vulnerability profiles, since context-blind ratings overestimate safety and realistic prompt context does not fix the problem.

  3. Physical Foundation Models: Fixed hardware implementations of large-scale neural networks

    cs.LG 2026-04 unverdicted novelty 5.0

    Physical Foundation Models are fixed physical hardware realizations of foundation-scale neural networks that compute via inherent material dynamics, potentially delivering orders-of-magnitude gains in energy efficienc...

  4. CASE: An Agentic AI Framework for Enhancing Scam Intelligence in Digital Payments

    cs.AI 2025-08 unverdicted novelty 5.0

    CASE is a novel agentic AI system that proactively interviews scam victims using LLMs to collect detailed intelligence, which is then structured for use in scam prevention, resulting in a 21% increase in enforcements ...