REVIEW 15 cited by
Open Problems in Technical AI Governance
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
AI progress is creating a growing range of risks and opportunities, but it is often unclear how they should be navigated. In many cases, the barriers and uncertainties faced are at least partly technical. Technical AI governance, referring to technical analysis and tools for supporting the effective governance of AI, seeks to address such challenges. It can help to (a) identify areas where intervention is needed, (b) identify and assess the efficacy of potential governance actions, and (c) enhance governance options by designing mechanisms for enforcement, incentivization, or compliance. In this paper, we explain what technical AI governance is, why it is important, and present a taxonomy and incomplete catalog of its open problems. This paper is intended as a resource for technical researchers or research funders looking to contribute to AI governance.
Forward citations
Cited by 15 Pith papers
-
No Certificate, No Categorical Speech Act: A Brouwerian Assertibility Constraint for Public Reason
An AI may assert or deny high-stakes claims only when it can exhibit a publicly contestable certificate; otherwise it is obligated to return Undetermined.
-
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.
-
Macro-Prudential AI Governance: A Two-Layer Early Warning and Response System for Frontier AI
A Basel-III-style two-layer system—coordinated finder-coordinator-defender reporting plus ECAR, CRTH, and ARS buffers—can detect and dampen correlated risk build-up across frontier AI labs’ internal deployments.
-
The Foreign Policy AI Evaluation Gap
Public technical AI governance almost never evaluates real foreign-policy AI workflows; the paper maps that gap and proposes task-scoped, human-recombined evaluation instead of model leaderboards.
-
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
A TEE-based protocol for cryptographically verifiable AI safety benchmark results, demonstrated on Llama-3.1 with AWS Nitro Enclaves.
-
A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search
A skill-graph random-walk model gives closed-form accuracy-versus-compute formulas for four reasoning strategies and connects them to training scaling.
-
Adversarial Attacks on Robotic Vision Language Action Models
Text-based adversarial suffixes can make OpenVLA robot policies elicit chosen target actions with over 90% success on one-hot targets and persist across rollout steps.
-
Asymmetry by Design: Boosting Cyber Defenders with Differential Access to AI
A policy framework for shaping access to AI cyber tools to favor defenders, with three access approaches and implementation guidance.
-
The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It
LLM safety research at ACL venues from 2020 to 2024 is predominantly English-only, and the language gap is growing over time.
-
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
CTRAP embeds a conditional failure mode during alignment so that harmful fine-tuning degrades the model to meaningless output while benign fine-tuning is unaffected.
-
The Economics of AI Training Data: A Research Agenda
The paper synthesizes fragmented research to frame data economics around data's nonrivalry and context dependence, catalogs 2020-2025 AI training data deals, and proposes a hierarchy of data units while listing four f...
-
A Conceptual Framework for AI Capability Evaluations
A descriptive conceptual framework with seven elements (target, task, subject, inputs, instance, measurement, result analysis) for systematizing analysis of AI capability evaluations.
-
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
A lightweight grid-world benchmark shows that LLM agents pay a task-performance cost for following safety principles and that high adherence can mask inability rather than principled choice.
-
Private, Verifiable, and Auditable AI Systems
A thesis demonstrating partial prototypes for zk-verifiable model evaluation and privacy-preserving retrieval, and arguing these pieces can compose into end-to-end auditable AI systems.
-
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
A three-part taxonomy for prompt-based natural language explanations, covering context, generation and presentation, and evaluation with 15 desirable properties.
Discussion (0). Sign in to comment.