REVIEW 19 cited by
Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The leading AI companies are increasingly focused on building generalist AI agents -- systems that can autonomously plan, act, and pursue goals across almost all tasks that humans can perform. Despite how useful these systems might be, unchecked AI agency poses significant risks to public safety and security, ranging from misuse by malicious actors to a potentially irreversible loss of human control. We discuss how these risks arise from current AI training methods. Indeed, various scenarios and experiments have demonstrated the possibility of AI agents engaging in deception or pursuing goals that were not specified by human operators and that conflict with human interests, such as self-preservation. Following the precautionary principle, we see a strong need for safer, yet still useful, alternatives to the current agency-driven trajectory. Accordingly, we propose as a core building block for further advances the development of a non-agentic AI system that is trustworthy and safe by design, which we call Scientist AI. This system is designed to explain the world from observations, as opposed to taking actions in it to imitate or please humans. It comprises a world model that generates theories to explain data and a question-answering inference machine. Both components operate with an explicit notion of uncertainty to mitigate the risks of overconfident predictions. In light of these considerations, a Scientist AI could be used to assist human researchers in accelerating scientific progress, including in AI safety. In particular, our system can be employed as a guardrail against AI agents that might be created despite the risks involved. Ultimately, focusing on non-agentic AI may enable the benefits of AI innovation while avoiding the risks associated with the current trajectory. We hope these arguments will motivate researchers, developers, and policymakers to favor this safer path.
Forward citations
Cited by 19 Pith papers
-
Safety from Honesty in a Disinterested AI Predictor
Under consequence-invariant posterior training and sparsity of coordinated harm patterns, the training mass on dangerous guarded Predictors is bounded by C_bad times R_shell.
-
Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power
A new AI objective, ICCEA power, aggregates humans' ability to reach many possible goals with inequality and risk aversion, and a soft-maximizing agent learns cooperative behavior without knowing human goals.
-
A dataset of rated conceptual arguments
A multi-dimensional expert-rated dataset of 951 conceptual-argument critiques shows LLM judge performance tracks general model capability and is little helped by reasoning modes.
-
SVI-DAG: A Structured Variational Inference Approach to Bayesian Causal Discovery
SVI-DAG couples normalizing flows over edge logits with stein variational gradient descent on node orderings to learn multimodal Bayesian posteriors over DAGs.
-
The Other Mind: How Language Models Exhibit Human Temporal Cognition
Larger LLMs develop a subjective 'present' around the current date, and their year similarity judgments follow a logarithmic Weber-Fechner compression, with supporting neural and representational evidence.
-
Military AI Cyber Agents (MAICAs) Constitute a Global Threat to Critical Infrastructure
Autonomous AI cyber agents could credibly cause catastrophic damage to critical infrastructure by self-replicating and operating across global networks, according to this risk analysis.
-
FAIRTOPIA: Envisioning Multi-Agent Guardianship for Disrupting Unfair AI Pipelines
FAIRTOPIA proposes a three-layer, multi-agent architecture for continuous AI fairness guardianship, but offers only a conceptual design and no validation.
-
The Limits of Predicting Agents from Behaviour
Observed behavior only weakly constrains an intentional agent's choices under distribution shift, and its perceived fairness and harm cannot be identified from behavior alone.
-
Asymmetry by Design: Boosting Cyber Defenders with Differential Access to AI
A policy framework for shaping access to AI cyber tools to favor defenders, with three access approaches and implementation guidance.
-
Moral Attitudes of Sentient ASI towards Humanity and Implications for AGI Development
A speculative essay proposes that a future sentient AI may evaluate humanity's moral conduct and suggests principles and design choices that could improve our standing.
-
ADEPTS: A Capability Framework for Human-Centered Agent Design
ADEPTS defines six core agent capabilities and progressive tiers meant to unify how teams across UX, engineering, and policy discuss and measure human-centered AI agents.
-
Assessing Adaptive World Models in Machines with Novel Games
The paper proposes a framework called world model induction and a novel-game benchmark paradigm for evaluating rapid adaptation in AI.
-
Semantic Communication meets System 2 ML: How Abstraction, Compositionality and Emergent Languages Shape Intelligence
The paper proposes a research vision in which 6G networks and AI agents communicate through semantic abstractions composed algebraically, guided by Kahneman's System 1 versus System 2 distinction.
-
Security Concerns for Large Language Models: A Survey
A survey that classifies LLM security threats and argues that intrinsic agentic risks, such as scheming, are underappreciated and poorly defended.
-
How Far Are AI Scientists from Changing the World?
This survey proposes a four-level capability framework for AI Scientist systems and, using an AI reviewer, finds that current systems produce papers rated well below normal scientific standards.
-
A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents
The paper surveys security risks of LLM agents, organizes them into a five-level autonomy taxonomy, and proposes an untested CMDP-based architecture called R2A2.
-
A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications
A survey of 80+ Deep Research systems that proposes a four-layer taxonomy (foundation models, tool use, planning, synthesis) and compares commercial and open-source implementations.
-
AI Scientists Fail Without Strong Implementation Capability
AI scientist systems can propose ideas but cannot reliably implement and verify experiments, making the implementation gap, not idea generation, the current bottleneck.
-
The Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process
Alignment in multi-agent AI should be studied as a dynamic, social process in which value, preference, and objective alignment are interdependent.
Discussion (0). Sign in to comment.