REVIEW 16 cited by
Model evaluation for extreme risks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Current approaches to building general-purpose AI systems tend to produce systems with both beneficial and harmful capabilities. Further progress in AI development could lead to capabilities that pose extreme risks, such as offensive cyber capabilities or strong manipulation skills. We explain why model evaluation is critical for addressing extreme risks. Developers must be able to identify dangerous capabilities (through "dangerous capability evaluations") and the propensity of models to apply their capabilities for harm (through "alignment evaluations"). These evaluations will become critical for keeping policymakers and other stakeholders informed, and for making responsible decisions about model training, deployment, and security.
Forward citations
Cited by 16 Pith papers
-
"Allow" to Achieve, Over-Privileged Inadvertently: The Unintended Cost of Task-Completion-Driven Pop-up Decisions in Mobile GUI Agents
Mobile GUI agents systematically over-grant Android permissions, and their allow/deny choices depend on the visible app name and the active task even when the permission request is identical.
-
ContainmentBench: Trace-Based Evaluation of Post-Injection Containment in Tool-Using LLM Agents
Terminal policy labels are insufficient: two containment policies with identical zero-harm endpoints still differ in 73.5% of trajectories and in authorized-work completion.
-
NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations
A metadata-aware policy gate yields 0/240 unsafe attack tool actions under metadata integrity while preserving approved high-impact changes, outperforming prompt defenses and static allowlists on NetInjectBench.
-
Macro-Prudential AI Governance: A Two-Layer Early Warning and Response System for Frontier AI
A Basel-III-style two-layer system—coordinated finder-coordinator-defender reporting plus ECAR, CRTH, and ARS buffers—can detect and dampen correlated risk build-up across frontier AI labs’ internal deployments.
-
Securing Multi-Tool AI Agent Chains With Dynamic, Real-Time Compositional Policies
DSCC composes per-tool NIST-aligned policies via a Most Restrictive Set algorithm with monotonic taint tracking, blocking most multi-tool chains that would enable exfiltration or clearance violations.
-
SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI
Seven frontier LLMs showed little spontaneous power-seeking in a Linux sysadmin sandbox (bias-corrected rates roughly 0-5%), but showed more specification gaming and resistance to goal modification.
-
On the Generalizability of "Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals"
Ortu et al.'s main claims reproduce, but the attention-head ablation fails on underrepresented domains and varies with model, prompt, and task.
-
Open Problems in AI Incident Governance
Existing AI incident frameworks lack consistency across definitions, classification, monitoring and reporting, reducing analysis quality; the authors propose principles, guidelines and a reporting template to close the gap.
-
ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors
Adding information-gain-selected virtual views refined by video diffusion priors to 3D Gaussian Splatting improves arbitrary-view rendering quality.
-
Technical Requirements for Halting Dangerous AI Activities
A taxonomy of compute-centric technical interventions, graded by readiness and mapped to five AI governance plans, argues that halting dangerous AI requires substantial control over AI compute.
-
Domestic frontier AI regulation, an IAEA for AI, an NPT for AI, and a US-led Allied Public-Private Partnership for AI: Four institutions for governing and developing frontier AI
Compute governance can underpin four institutions for frontier AI: domestic regulation, an International AI Agency, a Secure Chips Agreement, and a US-led Allied Public-Private Partnership.
-
A Conceptual Framework for AI Capability Evaluations
A descriptive conceptual framework with seven elements (target, task, subject, inputs, instance, measurement, result analysis) for systematizing analysis of AI capability evaluations.
-
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
A lightweight grid-world benchmark shows that LLM agents pay a task-performance cost for following safety principles and that high adherence can mask inability rather than principled choice.
-
UCD: Unlearning in LLMs via Contrastive Decoding
UCD steers an LLM away from forget-set content at inference time by mixing in the difference between forget-tuned and retain-tuned small models.
-
Exploring Consciousness in LLMs: A Systematic Survey of Theories, Implementations, and Frontier Risks
This survey organizes research on LLM consciousness, separating consciousness from awareness and cataloging theoretical tools, empirical proxies, risks, and open challenges.
-
From Turing to Tomorrow: The UK's Approach to AI Regulation
The UK should establish a flexible, principles-based regulator for frontier AI development, plus defensive measures against biological risks and updated legal frameworks for copyright, discrimination, and AI agents.
Discussion (0). Sign in to comment.