Pith. sign in

REVIEW 17 cited by

Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.07213 v2 pith:JESFYQN6 submitted 2020-04-15 cs.CY

classification cs.CY
keywords claimsdevelopmentmechanismssystemstheymakeneedstakeholders
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the recent wave of progress in artificial intelligence (AI) has come a growing awareness of the large-scale impacts of AI systems, and recognition that existing regulations and norms in industry and academia are insufficient to ensure responsible AI development. In order for AI developers to earn trust from system users, customers, civil society, governments, and other stakeholders that they are building AI responsibly, they will need to make verifiable claims to which they can be held accountable. Those outside of a given organization also need effective means of scrutinizing such claims. This report suggests various steps that different stakeholders can take to improve the verifiability of claims made about AI systems and their associated development processes, with a focus on providing evidence about the safety, security, fairness, and privacy protection of AI systems. We analyze ten mechanisms for this purpose--spanning institutions, software, and hardware--and make recommendations aimed at implementing, exploring, or improving those mechanisms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 125 citations worldwide. Full citation record

  1. Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education

    cs.HC 2026-08 conditional novelty 6.0 of 10

    A co-design study shows that making LLM trustworthiness metrics visible to learning engineers modestly increases agreement when choosing between AI tutor responses.

  2. Macro-Prudential AI Governance: A Two-Layer Early Warning and Response System for Frontier AI

    cs.CY 2026-07 conditional novelty 6.0 of 10

    A Basel-III-style two-layer system—coordinated finder-coordinator-defender reporting plus ECAR, CRTH, and ARS buffers—can detect and dampen correlated risk build-up across frontier AI labs’ internal deployments.

  3. The Foreign Policy AI Evaluation Gap

    cs.CY 2026-07 conditional novelty 6.0 of 10

    Public technical AI governance almost never evaluates real foreign-policy AI workflows; the paper maps that gap and proposes task-scoped, human-recombined evaluation instead of model leaderboards.

  4. Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A TEE-based protocol for cryptographically verifiable AI safety benchmark results, demonstrated on Llama-3.1 with AWS Nitro Enclaves.

  5. HADA: Human-AI Agent Decision Alignment Architecture

    cs.AI 2025-06 conditional novelty 6.0 of 10

    HADA is a framework-agnostic architecture that uses role-specific stakeholder agents to keep LLM agents and legacy algorithms aligned with organizational KPIs and values, demonstrated in a scripted retail banking pilot.

  6. Domestic frontier AI regulation, an IAEA for AI, an NPT for AI, and a US-led Allied Public-Private Partnership for AI: Four institutions for governing and developing frontier AI

    cs.CY 2025-07 accept novelty 5.0 of 10

    Compute governance can underpin four institutions for frontier AI: domestic regulation, an International AI Agency, a Secure Chips Agreement, and a US-led Allied Public-Private Partnership.

  7. Machine vs Machine: Using AI to Tackle Generative AI Threats in Assessment

    cs.CY 2025-05 conditional novelty 5.0 of 10

    Proposes a dual static-analysis and dynamic-testing framework to evaluate and reduce assessment vulnerability to generative AI, but provides no empirical validation.

  8. Bottom-Up Perspectives on AI Governance: Insights from User Reviews of AI Products

    cs.CY 2025-05 conditional novelty 5.0 of 10

    Using BERTopic on 108,998 G2 reviews, the study maps governance-relevant themes in user discourse, finding overlap with official AI ethics principles plus operational topics like project management and customer interaction.

  9. AGL-1: The Enterprise AI Governance Layer as a Control Plane for Trusted Enterprise Intelligence

    cs.SE 2026-07 conditional novelty 4.0 of 10

    AGL-1 frames enterprise AI trust as a shared control plane of seven governance domains spanning identity-aware retrieval through agentic execution and audit evidence.

  10. Toward Secure and Compliant AI: Organizational Standards and Protocols for NLP Model Lifecycle Management

    cs.CR 2025-12 conditional novelty 4.0 of 10

    A proposed six-phase NLP lifecycle governance framework built from existing standards, illustrated with a self-authored healthcare case study but no empirical validation.

  11. Strategic Alignment Patterns in National AI Policies

    cs.CY 2025-07 reject novelty 4.0 of 10

    A policy-analysis preprint scores alignment between objectives, foresight, and instruments in 15-20 national AI strategies, claiming distinct governance-based archetypes, but ships no data, figures, or code to support...

  12. Risks of AI-driven product development and strategies for their mitigation

    cs.CY 2025-05 conditional novelty 4.0 of 10

    AI-driven product development will bring technical and societal risks; the paper proposes eight mitigation principles: human control, accountability, explainable and tested design, constrained and sandboxed systems, a...

  13. Trustworthy AI LLM Scalability Risk Index (LSRI): A Cybersecurity Framework Assessing Agentic-AI Security & Software Model Supply Chain Safety Boosting AI-Generated Malware Defense & Explainability Mitigating Emerging Risks of Generative AI

    cs.CR 2026-02 reject novelty 3.0 of 10

    LSRI is a hand-calibrated weighted risk score for LLM deployments, paired with a Sigstore-based checkpoint attestation proposal; neither is empirically validated.

  14. Towards Transparent Ethical AI: A Roadmap for Trustworthy Robotic Systems

    cs.CY 2025-08 unverdicted novelty 3.0 of 10

    The paper argues transparency is fundamental to trustworthy robotics and proposes a framework connecting technical transparency tools to ethical outcomes such as accountability and informed consent.

  15. Bridging the Artificial Intelligence Governance Gap: The United States' and China's Divergent Approaches to Governing General-Purpose Artificial Intelligence

    cs.CY 2025-06 conditional novelty 3.0 of 10

    The paper compares U.S. and Chinese governance of general-purpose AI and identifies three divergences: focus of domestic regulation, key principles, and international forum choices.

  16. Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies

    cs.CR 2026-07 conditional novelty 2.0 of 10

    A survey of dual-use LLM applications in cybersecurity, synthesizing defensive tools, attack vectors, governance frameworks, and the projected growth of AI-assisted malware through 2025.

  17. The Hidden Costs of AI: A Review of Energy, E-Waste, and Inequality in Model Development

    cs.AI 2025-07 reject

    A review paper summarizing AI's energy use, e-waste, compute inequality, and cybersecurity energy costs, with no new data.

Pith tools