Pith. sign in

REVIEW 2 cited by

STAR: SocioTechnical Approach to Red Teaming Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11757 v4 pith:Z6P4AB2N submitted 2024-06-17 cs.AI cs.CLcs.CYcs.HC

classification cs.AIcs.CLcs.CYcs.HC
keywords starimprovesinstructionslanguagemodelsparameterisedqualitysignal
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This research introduces STAR, a sociotechnical framework that improves on current best practices for red teaming safety of large language models. STAR makes two key contributions: it enhances steerability by generating parameterised instructions for human red teamers, leading to improved coverage of the risk surface. Parameterised instructions also provide more detailed insights into model failures at no increased cost. Second, STAR improves signal quality by matching demographics to assess harms for specific groups, resulting in more sensitive annotations. STAR further employs a novel step of arbitration to leverage diverse viewpoints and improve label reliability, treating disagreement not as noise but as a valuable contribution to signal quality.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WeAudit: Scaffolding User Auditors and AI Practitioners in Auditing Generative AI

    cs.HC 2025-01 conditional novelty 6.0 of 10

    WeAudit scaffolds everyday users to audit generative AI through comparison, examples, discussion, verification, and structured reports, and practitioners found the resulting audit reports actionable.

  2. A theory of appropriateness with applications to generative artificial intelligence

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A theory that human and AI behavior is guided by context-dependent appropriateness implemented as predictive pattern completion, with norms as conventional sanctioning patterns.

Pith tools