REVIEW 24 cited by
Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Ensuring that AI systems reliably and robustly avoid harmful or dangerous behaviours is a crucial challenge, especially for AI systems with a high degree of autonomy and general intelligence, or systems used in safety-critical contexts. In this paper, we will introduce and define a family of approaches to AI safety, which we will refer to as guaranteed safe (GS) AI. The core feature of these approaches is that they aim to produce AI systems which are equipped with high-assurance quantitative safety guarantees. This is achieved by the interplay of three core components: a world model (which provides a mathematical description of how the AI system affects the outside world), a safety specification (which is a mathematical description of what effects are acceptable), and a verifier (which provides an auditable proof certificate that the AI satisfies the safety specification relative to the world model). We outline a number of approaches for creating each of these three core components, describe the main technical challenges, and suggest a number of potential solutions to them. We also argue for the necessity of this approach to AI safety, and for the inadequacy of the main alternative approaches.
Forward citations
Cited by 24 Pith papers
-
Mining Verdict Boundaries for Neural Network Verification
BMiner speeds up Branch-and-Bound neural network verification by using exponential and gradient-guided search to skip subproblems on the way to each path's verdict boundary, cutting average verification time by 17–30%.
-
Towards chemistries in dynamical systems
A finite dynamical system can be described as a Petri-net chemistry when a token map and transition map are compatible; an optional uniqueness criterion restricts when this description is least ambiguous.
-
Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny
A minimal-prior pipeline with automated data curation and verifier-driven RL lets small LLMs generate verifiable Dafny specifications and beat larger proprietary models on a synthetic compositional benchmark.
-
The Other Mind: How Language Models Exhibit Human Temporal Cognition
Larger LLMs develop a subjective 'present' around the current date, and their year similarity judgments follow a logarithmic Weber-Fechner compression, with supporting neural and representational evidence.
-
The Limits of Predicting Agents from Behaviour
Observed behavior only weakly constrains an intentional agent's choices under distribution shift, and its perceived fairness and harm cannot be identified from behavior alone.
-
Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training
TD-GFN uses IRL-derived edge rewards to prune the environment DAG and sample backward trajectories, training offline GFlowNets directly from ground-truth terminal rewards without a proxy reward model.
-
Assessing confidence in frontier AI safety cases
Applying Assurance 2.0 to a cyber-misuse safety case, the authors show that high top-level confidence requires extremely high confidence in every component, and propose an LLM-based Delphi for eliciting those componen...
-
Sequential Decision Making in Stochastic Games with Incomplete Preferences over Temporal Objectives
Proposes non-dominated almost-sure winning strategies for stochastic games with incomplete LTLf preferences, via a rank-based algorithm, and claims these form Nash equilibria.
-
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
A new 46-criteria assessment framework scores 24 AI benchmarks and finds that commonly used benchmarks are weak in implementation and statistical rigor.
-
Safety case template for frontier AI: A cyber inability argument
A proof-of-concept safety case template formalizes an inability argument for offensive cyber risk using risk models, proxy tasks, and evaluation results.
-
In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
Based on a four-risk typology, the paper concludes that verification mechanisms and codified protocols are the least risky areas for cooperation between geopolitical rivals on technical AI safety.
-
Where AI Assurance Might Go Wrong: Initial lessons from engineering of critical systems
A position paper mapping traditional critical systems engineering to AI safety frameworks and advocating Assurance 2.0-style cases, broader boundaries, and explicit risk tolerability.
-
Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations
A new model quantifies how test sensitivity, capability growth, and threshold placement determine bias and detection lag in dangerous AI evaluations.
-
Agent Safety Should Be a Runtime Contract
Agent safety should be a runtime contract, enforced by sandboxes and permission gates on the preventive side and by verifiable evidence chains on the submission side, not by model alignment alone.
-
Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack)
A position paper proposing that AI alignment adopt formal optimal control and a ten-layer Alignment Control Stack for organizing and interoperating control interventions.
-
What Is AI Safety? What Do We Want It to Be?
AI safety is best understood as all research aimed at preventing or reducing harms from AI systems, covering social harms and catastrophic risks together.
-
A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management
A synthesis of established risk management practices into a structured framework for frontier AI developers, centered on explicit risk tolerance, KRI/KCI thresholds, and governance.
-
Revisiting Rogers' Paradox in the Context of Human-AI Interaction
A simulation of Rogers' Paradox with an AI agent that learns the population average shows that cheap AI alone does not improve collective world understanding, while critical appraisal and independent AI learning can.
-
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
The paper proposes the AI-45 degree law, a Causal Ladder framework, and five trustworthiness levels as a roadmap toward trustworthy AGI.
-
Reliability, Resilience and Human Factors Engineering for Trustworthy AI Systems
A framework applying classical reliability, resilience, and human-factors engineering to AI systems, with a small subjective case study of OpenAI status incidents.
-
Position Paper: Bounded Alignment: What (Not) To Expect From AGI Agents
The paper argues that perfect alignment of general AI is impossible in principle and proposes 'bounded alignment' as the realistic safety goal.
-
Responsible Artificial Intelligence (RAI) in U.S. Federal Government : Principles, Policies, and Practices
A U.S. federal RAI policy review that maps executive orders, memos, and frameworks onto five RAI principles and describes Census Bureau implementation tools, with no new empirical findings.
-
Algebraic Evaluation Theorems
Under an error-independence assumption, the decisions of three binary classifiers determine their true accuracies up to exactly two alternative solutions, and one of them is the true evaluation.
-
Report on NSF Workshop on Science of Safe AI
An NSF workshop report articulating a cross-disciplinary research agenda for designing and verifying safe, trustworthy AI systems.
Discussion (0). Continue with ORCID to comment.