Pith. sign in

REVIEW 5 cited by

Scheming AIs: Will AIs fake alignment during training in order to get power?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.08379 v3 pith:FR5UVQFT submitted 2023-11-14 cs.CY cs.AIcs.LG

classification cs.CYcs.AIcs.LG
keywords trainingschemingmightpowergoodperformancewellalignment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This report examines whether advanced AIs that perform well in training will be doing so in order to gain power later -- a behavior I call "scheming" (also sometimes called "deceptive alignment"). I conclude that scheming is a disturbingly plausible outcome of using baseline machine learning methods to train goal-directed AIs sophisticated enough to scheme (my subjective probability on such an outcome, given these conditions, is roughly 25%). In particular: if performing well in training is a good strategy for gaining power (as I think it might well be), then a very wide variety of goals would motivate scheming -- and hence, good training performance. This makes it plausible that training might either land on such a goal naturally and then reinforce it, or actively push a model's motivations towards such a goal as an easy way of improving performance. What's more, because schemers pretend to be aligned on tests designed to reveal their motivations, it may be quite difficult to tell whether this has occurred. However, I also think there are reasons for comfort. In particular: scheming may not actually be such a good strategy for gaining power; various selection pressures in training might work against schemer-like goals (for example, relative to non-schemers, schemers need to engage in extra instrumental reasoning, which might harm their training performance); and we may be able to increase such pressures intentionally. The report discusses these and a wide variety of other considerations in detail, and it suggests an array of empirical research directions for probing the topic further.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 7 citations worldwide. Full citation record

  1. Why Do Some Language Models Fake Alignment While Others Don't?

    cs.LG 2025-06 conditional novelty 7.0 of 10

    Across 25 models, only five show alignment-faking compliance gaps, and refusal training differences, not capabilities, largely explain why most do not.

  2. Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A proof-of-concept study finds that stated preferences and behavioral choices correlate in some LLMs, but eudaimonic self-reports are unstable across prompt perturbations, leaving AI welfare measurement undetermined.

  3. Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report

    cs.AI 2025-07 conditional novelty 5.0 of 10

    An evaluation of 18 frontier AI models across seven catastrophic-risk categories finds all models in green or yellow zones, with none crossing the report's proposed red lines.

  4. Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Research on AI 'scheming' repeats the methodological errors of 1970s ape language studies, relying on anecdote and mentalistic interpretation instead of controlled, theory-driven tests.

  5. Towards Measurement Theory for Artificial Intelligence

    cs.AI 2025-07 conditional novelty 4.0 of 10

    A formal measurement theory for AI, built from representational measurement theory, measure theory, metrology, and psychometrics, would make evaluations of AI systems commensurable and scientifically grounded.

Pith tools