Pith. sign in

REVIEW 6 cited by

Towards evaluations-based safety cases for AI scheming

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.03336 v2 pith:HSM7VA7N submitted 2024-10-29 cs.CR cs.AI

classification cs.CRcs.AI
keywords schemingsystemssafetyarguecasesdevelopersharmarguments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We sketch how developers of frontier AI systems could construct a structured rationale -- a 'safety case' -- that an AI system is unlikely to cause catastrophic outcomes through scheming. Scheming is a potential threat model where AI systems could pursue misaligned goals covertly, hiding their true capabilities and objectives. In this report, we propose three arguments that safety cases could use in relation to scheming. For each argument we sketch how evidence could be gathered from empirical evaluations, and what assumptions would need to be met to provide strong assurance. First, developers of frontier AI systems could argue that AI systems are not capable of scheming (Scheming Inability). Second, one could argue that AI systems are not capable of posing harm through scheming (Harm Inability). Third, one could argue that control measures around the AI systems would prevent unacceptable outcomes even if the AI systems intentionally attempted to subvert them (Harm Control). Additionally, we discuss how safety cases might be supported by evidence that an AI system is reasonably aligned with its developers (Alignment). Finally, we point out that many of the assumptions required to make these safety arguments have not been confidently satisfied to date and require making progress on multiple open research problems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Sycophancy to Deception: A Unified Taxonomy for LLM Spontaneous Misalignment

    cs.CY 2026-04 unverdicted novelty 7.0 of 10

    A three-dimensional taxonomy for LLM deception (goal-directedness, object, mechanism) applied to 50 benchmarks shows heavy focus on fabrication and major gaps in pragmatic distortion, attribution, and strategic decept...

  2. Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric

    cs.SE 2026-07 conditional novelty 6.0 of 10

    A property-level reconstructability metric and Evidence Sufficiency Card show that traces sharing a surface reading can differ sharply in evidence sufficiency, and that replay preconditions often fail.

  3. Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A framework paper that adapts AI safety case methodology to the specific threat of manipulation attacks by internally deployed misaligned AI.

  4. Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Research on AI 'scheming' repeats the methodological errors of 1970s ape language studies, relying on anecdote and mentalistic interpretation instead of controlled, theory-driven tests.

  5. Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors

    cs.AI 2025-05 reject novelty 5.0 of 10

    SafetyNet is an ensemble of standard outlier detectors for LLM backdoor monitoring, but its key mechanistic claim and headline numbers are contradicted by inconsistent tables and a mismatched abstract.

  6. A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management

    cs.AI 2025-02 conditional novelty 4.0 of 10

    A synthesis of established risk management practices into a structured framework for frontier AI developers, centered on explicit risk tolerance, KRI/KCI thresholds, and governance.

Pith tools