Pith. sign in

REVIEW 3 cited by

Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.17861 v3 pith:ROILCGCX submitted 2024-02-27 cs.CY

classification cs.CY
keywords audittoolspractitionersaccountabilityauditsbeyondeffortsevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audits are critical mechanisms for identifying the risks and limitations of deployed artificial intelligence (AI) systems. However, the effective execution of AI audits remains incredibly difficult, and practitioners often need to make use of various tools to support their efforts. Drawing on interviews with 35 AI audit practitioners and a landscape analysis of 435 tools, we compare the current ecosystem of AI audit tooling to practitioner needs. While many tools are designed to help set standards and evaluate AI systems, they often fall short in supporting accountability. We outline challenges practitioners faced in their efforts to use AI audit tools and highlight areas for future tool development beyond evaluation -- from harms discovery to advocacy. We conclude that the available resources do not currently support the full scope of AI audit practitioners' needs and recommend that the field move beyond tools for just evaluation and towards more comprehensive infrastructure for AI accountability.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. White Box Evidence Packages for Policy Audit Reports

    cs.CY 2026-07 conditional novelty 6.0 of 10

    In a 60-case controlled audit study, adding white-box model evidence to an LLM auditor increased citation volume but weakened passage grounding and raised evidence misuse, while a shuffled control showed reports can s...

  2. Impact Assessment Card: Communicating Risks and Benefits of AI Uses

    cs.HC 2025-08 conditional novelty 6.0 of 10

    In an online study of 235 people, a compact 'Impact Assessment Card' for AI systems outperformed a full written report on speed, email quality, usability, and preference.

  3. Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems

    cs.CY 2025-06 conditional novelty 6.0 of 10

    Practitioners trying to measure representational harms in LLM-based systems often cannot use public measurement instruments, either because the instruments lack validity, specificity, interpretability, or actionabilit...

Pith tools