Pith. sign in

REVIEW 3 cited by

OpenAI's Approach to External Red Teaming for AI Models and Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.16431 v1 pith:IPWNDMFU submitted 2025-01-24 cs.CY cs.AIcs.CRcs.HC

classification cs.CYcs.AIcs.CRcs.HC
keywords teamingexternalmodelsdescribedesignevaluationevaluationsexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Red teaming has emerged as a critical practice in assessing the possible risks of AI models and systems. It aids in the discovery of novel risks, stress testing possible gaps in existing mitigations, enriching existing quantitative safety metrics, facilitating the creation of new safety measurements, and enhancing public trust and the legitimacy of AI risk assessments. This white paper describes OpenAI's work to date in external red teaming and draws some more general conclusions from this work. We describe the design considerations underpinning external red teaming, which include: selecting composition of red team, deciding on access levels, and providing guidance required to conduct red teaming. Additionally, we show outcomes red teaming can enable such as input into risk assessment and automated evaluations. We also describe the limitations of external red teaming, and how it can fit into a broader range of AI model and system evaluations. Through these contributions, we hope that AI developers and deployers, evaluation creators, and policymakers will be able to better design red teaming campaigns and get a deeper look into how external red teaming can fit into model deployment and evaluation processes. These methods are evolving and the value of different methods continues to shift as the ecosystem around red teaming matures and models themselves improve as tools for red teaming.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety

    cs.CY 2026-06 accept novelty 6.5 of 10

    Legal and ethical bans on CSAM access and generation break standard AI safety techniques, creating 15 open problems that demand new methods for dataset cleaning, concept fusion prevention, fine-tuning resilience, dete...

  2. Aggregated Individual Reporting for Post-Deployment Evaluation

    cs.CY 2025-06 conditional novelty 6.0 of 10

    The authors formalize a mechanism for collecting and aggregating public reports about deployed AI systems, aiming to surface unknown harms and enable accountability.

  3. Red-Teaming Claude Opus and ChatGPT-based Security Advisors for Trusted Execution Environments

    cs.CR 2026-02 conditional novelty 5.0 of 10

    A TEE-focused red-teaming benchmark finds that up to 12% of AI security-advisor failures transfer between ChatGPT and Claude, and a proposed defense pipeline reportedly reduces failures by ~80%.

Pith tools