REVIEW 17 cited by
Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the recent wave of progress in artificial intelligence (AI) has come a growing awareness of the large-scale impacts of AI systems, and recognition that existing regulations and norms in industry and academia are insufficient to ensure responsible AI development. In order for AI developers to earn trust from system users, customers, civil society, governments, and other stakeholders that they are building AI responsibly, they will need to make verifiable claims to which they can be held accountable. Those outside of a given organization also need effective means of scrutinizing such claims. This report suggests various steps that different stakeholders can take to improve the verifiability of claims made about AI systems and their associated development processes, with a focus on providing evidence about the safety, security, fairness, and privacy protection of AI systems. We analyze ten mechanisms for this purpose--spanning institutions, software, and hardware--and make recommendations aimed at implementing, exploring, or improving those mechanisms.
Forward citations
Cited by 17 Pith papers
-
Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education
A co-design study shows that making LLM trustworthiness metrics visible to learning engineers modestly increases agreement when choosing between AI tutor responses.
-
Macro-Prudential AI Governance: A Two-Layer Early Warning and Response System for Frontier AI
A Basel-III-style two-layer system—coordinated finder-coordinator-defender reporting plus ECAR, CRTH, and ARS buffers—can detect and dampen correlated risk build-up across frontier AI labs’ internal deployments.
-
The Foreign Policy AI Evaluation Gap
Public technical AI governance almost never evaluates real foreign-policy AI workflows; the paper maps that gap and proposes task-scoped, human-recombined evaluation instead of model leaderboards.
-
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
A TEE-based protocol for cryptographically verifiable AI safety benchmark results, demonstrated on Llama-3.1 with AWS Nitro Enclaves.
-
HADA: Human-AI Agent Decision Alignment Architecture
HADA is a framework-agnostic architecture that uses role-specific stakeholder agents to keep LLM agents and legacy algorithms aligned with organizational KPIs and values, demonstrated in a scripted retail banking pilot.
-
Domestic frontier AI regulation, an IAEA for AI, an NPT for AI, and a US-led Allied Public-Private Partnership for AI: Four institutions for governing and developing frontier AI
Compute governance can underpin four institutions for frontier AI: domestic regulation, an International AI Agency, a Secure Chips Agreement, and a US-led Allied Public-Private Partnership.
-
Machine vs Machine: Using AI to Tackle Generative AI Threats in Assessment
Proposes a dual static-analysis and dynamic-testing framework to evaluate and reduce assessment vulnerability to generative AI, but provides no empirical validation.
-
Bottom-Up Perspectives on AI Governance: Insights from User Reviews of AI Products
Using BERTopic on 108,998 G2 reviews, the study maps governance-relevant themes in user discourse, finding overlap with official AI ethics principles plus operational topics like project management and customer interaction.
-
AGL-1: The Enterprise AI Governance Layer as a Control Plane for Trusted Enterprise Intelligence
AGL-1 frames enterprise AI trust as a shared control plane of seven governance domains spanning identity-aware retrieval through agentic execution and audit evidence.
-
Toward Secure and Compliant AI: Organizational Standards and Protocols for NLP Model Lifecycle Management
A proposed six-phase NLP lifecycle governance framework built from existing standards, illustrated with a self-authored healthcare case study but no empirical validation.
-
Strategic Alignment Patterns in National AI Policies
A policy-analysis preprint scores alignment between objectives, foresight, and instruments in 15-20 national AI strategies, claiming distinct governance-based archetypes, but ships no data, figures, or code to support...
-
Risks of AI-driven product development and strategies for their mitigation
AI-driven product development will bring technical and societal risks; the paper proposes eight mitigation principles: human control, accountability, explainable and tested design, constrained and sandboxed systems, a...
-
Trustworthy AI LLM Scalability Risk Index (LSRI): A Cybersecurity Framework Assessing Agentic-AI Security & Software Model Supply Chain Safety Boosting AI-Generated Malware Defense & Explainability Mitigating Emerging Risks of Generative AI
LSRI is a hand-calibrated weighted risk score for LLM deployments, paired with a Sigstore-based checkpoint attestation proposal; neither is empirically validated.
-
Towards Transparent Ethical AI: A Roadmap for Trustworthy Robotic Systems
The paper argues transparency is fundamental to trustworthy robotics and proposes a framework connecting technical transparency tools to ethical outcomes such as accountability and informed consent.
-
Bridging the Artificial Intelligence Governance Gap: The United States' and China's Divergent Approaches to Governing General-Purpose Artificial Intelligence
The paper compares U.S. and Chinese governance of general-purpose AI and identifies three divergences: focus of domestic regulation, key principles, and international forum choices.
-
Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies
A survey of dual-use LLM applications in cybersecurity, synthesizing defensive tools, attack vectors, governance frameworks, and the projected growth of AI-assisted malware through 2025.
-
The Hidden Costs of AI: A Review of Energy, E-Waste, and Inequality in Model Development
A review paper summarizing AI's energy use, e-waste, compute inequality, and cybersecurity energy costs, with no new data.
Discussion (0). Continue with ORCID to comment.