REVIEW 5 cited by
Safety Cases: How to Justify the Safety of Advanced AI Systems
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
As AI systems become more advanced, companies and regulators will make difficult decisions about whether it is safe to train and deploy them. To prepare for these decisions, we investigate how developers could make a 'safety case,' which is a structured rationale that AI systems are unlikely to cause a catastrophe. We propose a framework for organizing a safety case and discuss four categories of arguments to justify safety: total inability to cause a catastrophe, sufficiently strong control measures, trustworthiness despite capability to cause harm, and -- if AI systems become much more powerful -- deference to credible AI advisors. We evaluate concrete examples of arguments in each category and outline how arguments could be combined to justify that AI systems are safe to deploy.
Forward citations
Cited by 5 Pith papers
-
Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety
Changing only the consequence-allocation rule in multi-agent AI shifts collective fatality by 22–58 percentage points across seven model populations, with identity salience in rule text causally driving targeted exploitation.
-
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
SFT lessons — reason-based training, on-model replay, and wash-out robustness — transfer across toy models, model organisms, and alignment SFT, improving the capability–safety tradeoff.
-
How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements
Verification of international AI agreements will fail first at detecting hidden compute facilities, around the 10,000-H100-equivalent scale, before other enforcement mechanisms break.
-
Alignment Plausibility: A New Standard for Assuring AI in Healthcare
Alignment plausibility—evidence that an AI system's values, training, and oversight cohere with safe positive health outcomes—should be the regulatory analogue of biological plausibility for LLMs in healthcare.
-
Legal Alignment for Safe and Ethical AI
Legal alignment as a field: AI systems should (1) follow the content of law, (2) use legal-interpretation methods for reasoning, and (3) be built on legal structures like agency and fiduciary duties.
Discussion (0). Sign in to comment.