Pith. sign in

REVIEW 24 cited by

Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.06624 v3 pith:HNZOS2BB submitted 2024-05-10 cs.AI

classification cs.AI
keywords safetysystemsapproachescoreworldcomponentsdescriptionensuring
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Ensuring that AI systems reliably and robustly avoid harmful or dangerous behaviours is a crucial challenge, especially for AI systems with a high degree of autonomy and general intelligence, or systems used in safety-critical contexts. In this paper, we will introduce and define a family of approaches to AI safety, which we will refer to as guaranteed safe (GS) AI. The core feature of these approaches is that they aim to produce AI systems which are equipped with high-assurance quantitative safety guarantees. This is achieved by the interplay of three core components: a world model (which provides a mathematical description of how the AI system affects the outside world), a safety specification (which is a mathematical description of what effects are acceptable), and a verifier (which provides an auditable proof certificate that the AI satisfies the safety specification relative to the world model). We outline a number of approaches for creating each of these three core components, describe the main technical challenges, and suggest a number of potential solutions to them. We also argue for the necessity of this approach to AI safety, and for the inadequacy of the main alternative approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 24 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mining Verdict Boundaries for Neural Network Verification

    cs.LG 2026-07 conditional novelty 6.0 of 10

    BMiner speeds up Branch-and-Bound neural network verification by using exponential and gradient-guided search to skip subproblems on the way to each path's verdict boundary, cutting average verification time by 17–30%.

  2. Towards chemistries in dynamical systems

    cs.DM 2026-07 conditional novelty 6.0 of 10

    A finite dynamical system can be described as a Petri-net chemistry when a token map and transition map are compatible; an optional uniqueness criterion restricts when this description is least ambiguous.

  3. Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A minimal-prior pipeline with automated data curation and verifier-driven RL lets small LLMs generate verifiable Dafny specifications and beat larger proprietary models on a synthetic compositional benchmark.

  4. The Other Mind: How Language Models Exhibit Human Temporal Cognition

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Larger LLMs develop a subjective 'present' around the current date, and their year similarity judgments follow a logarithmic Weber-Fechner compression, with supporting neural and representational evidence.

  5. The Limits of Predicting Agents from Behaviour

    cs.AI 2025-06 accept novelty 6.0 of 10

    Observed behavior only weakly constrains an intentional agent's choices under distribution shift, and its perceived fairness and harm cannot be identified from behavior alone.

  6. Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training

    cs.LG 2025-05 conditional novelty 6.0 of 10

    TD-GFN uses IRL-derived edge rewards to prune the environment DAG and sample backward trajectories, training offline GFlowNets directly from ground-truth terminal rewards without a proxy reward model.

  7. Assessing confidence in frontier AI safety cases

    cs.CY 2025-02 conditional novelty 6.0 of 10

    Applying Assurance 2.0 to a cyber-misuse safety case, the authors show that high top-level confidence requires extremely high confidence in every component, and propose an LLM-based Delphi for eliciting those componen...

  8. Sequential Decision Making in Stochastic Games with Incomplete Preferences over Temporal Objectives

    cs.GT 2025-01 reject novelty 6.0 of 10

    Proposes non-dominated almost-sure winning strategies for stochastic games with incomplete LTLf preferences, via a rank-based algorithm, and claims these form Nash equilibria.

  9. BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices

    cs.AI 2024-11 conditional novelty 6.0 of 10

    A new 46-criteria assessment framework scores 24 AI benchmarks and finds that commonly used benchmarks are weak in implementation and statistical rigor.

  10. Safety case template for frontier AI: A cyber inability argument

    cs.CY 2024-11 accept novelty 6.0 of 10

    A proof-of-concept safety case template formalizes an inability argument for offensive cyber risk using risk models, proxy tasks, and evaluation results.

  11. In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?

    cs.CY 2025-04 conditional novelty 5.0 of 10

    Based on a four-risk typology, the paper concludes that verification mechanisms and codified protocols are the least risky areas for cooperation between geopolitical rivals on technical AI safety.

  12. Where AI Assurance Might Go Wrong: Initial lessons from engineering of critical systems

    cs.CY 2025-01 accept novelty 5.0 of 10

    A position paper mapping traditional critical systems engineering to AI safety frameworks and advocating Assurance 2.0-style cases, broader boundaries, and explicit risk tolerability.

  13. Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A new model quantifies how test sensitivity, capability growth, and threshold placement determine bias and detection lag in dangerous AI evaluations.

  14. Agent Safety Should Be a Runtime Contract

    cs.CR 2026-08 conditional novelty 4.0 of 10

    Agent safety should be a runtime contract, enforced by sandboxes and permission gates on the preventive side and by verifiable evidence chains on the submission side, not by model alignment alone.

  15. Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack)

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A position paper proposing that AI alignment adopt formal optimal control and a ten-layer Alignment Control Stack for organizing and interoperating control interventions.

  16. What Is AI Safety? What Do We Want It to Be?

    cs.CY 2025-05 conditional novelty 4.0 of 10

    AI safety is best understood as all research aimed at preventing or reducing harms from AI systems, covering social harms and catastrophic risks together.

  17. A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management

    cs.AI 2025-02 conditional novelty 4.0 of 10

    A synthesis of established risk management practices into a structured framework for frontier AI developers, centered on explicit risk tolerance, KRI/KCI thresholds, and governance.

  18. Revisiting Rogers' Paradox in the Context of Human-AI Interaction

    cs.AI 2025-01 conditional novelty 4.0 of 10

    A simulation of Rogers' Paradox with an AI agent that learns the population average shows that cheap AI alone does not improve collective world understanding, while critical appraisal and independent AI learning can.

  19. Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI

    cs.CY 2024-12 conditional novelty 4.0 of 10

    The paper proposes the AI-45 degree law, a Causal Ladder framework, and five trustworthiness levels as a roadmap toward trustworthy AGI.

  20. Reliability, Resilience and Human Factors Engineering for Trustworthy AI Systems

    cs.AI 2024-11 conditional novelty 4.0 of 10

    A framework applying classical reliability, resilience, and human-factors engineering to AI systems, with a small subjective case study of OpenAI status incidents.

  21. Position Paper: Bounded Alignment: What (Not) To Expect From AGI Agents

    cs.AI 2025-05 conditional novelty 3.0 of 10

    The paper argues that perfect alignment of general AI is impossible in principle and proposes 'bounded alignment' as the realistic safety goal.

  22. Responsible Artificial Intelligence (RAI) in U.S. Federal Government : Principles, Policies, and Practices

    cs.CY 2025-01 unverdicted novelty 3.0 of 10

    A U.S. federal RAI policy review that maps executive orders, memos, and frameworks onto five RAI principles and describes Census Bureau implementation tools, with no new empirical findings.

  23. Algebraic Evaluation Theorems

    cs.AI 2024-12 conditional novelty 3.0 of 10

    Under an error-independence assumption, the decisions of three binary classifiers determine their true accuracies up to exactly two alternative solutions, and one of them is the true evaluation.

  24. Report on NSF Workshop on Science of Safe AI

    cs.CY 2025-06 unverdicted novelty 2.0 of 10

    An NSF workshop report articulating a cross-disciplinary research agenda for designing and verifying safe, trustworthy AI systems.

Pith tools