REVIEW 6 cited by
Evaluating and Mitigating Discrimination in Language Model Decisions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
As language models (LMs) advance, interest is growing in applying them to high-stakes societal decisions, such as determining financing or housing eligibility. However, their potential for discrimination in such contexts raises ethical concerns, motivating the need for better methods to evaluate these risks. We present a method for proactively evaluating the potential discriminatory impact of LMs in a wide range of use cases, including hypothetical use cases where they have not yet been deployed. Specifically, we use an LM to generate a wide array of potential prompts that decision-makers may input into an LM, spanning 70 diverse decision scenarios across society, and systematically vary the demographic information in each prompt. Applying this methodology reveals patterns of both positive and negative discrimination in the Claude 2.0 model in select settings when no interventions are applied. While we do not endorse or permit the use of language models to make automated decisions for the high-risk use cases we study, we demonstrate techniques to significantly decrease both positive and negative discrimination through careful prompt engineering, providing pathways toward safer deployment in use cases where they may be appropriate. Our work enables developers and policymakers to anticipate, measure, and address discrimination as language model capabilities and applications continue to expand. We release our dataset and prompts at https://huggingface.co/datasets/Anthropic/discrim-eval
Forward citations
Cited by 6 Pith papers
-
Reference Feature Atlases for Mechanistic Auditing of Language Models
A reference feature atlas attaches a new language model with a linear decoder and surfaces panel-uncovered structure via a separate residual dictionary.
-
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
Audit format (rate vs rank vs allocate; transparent vs disguised) reverses the apparent direction of LLM demographic bias, while causal framing of need dominates allocations by roughly an order of magnitude.
-
DeFrame: Debiasing Large Language Models Against Framing Effects
LLM fairness scores shift substantially with positive vs negative framing of the same question, and DeFrame—a three-step self-revision prompt—reduces both average bias and this framing gap.
-
White-Box Sensitivity Auditing with Steering Vectors
White-box steering along gender/race concept vectors reveals larger and more consistent bias signals than black-box prompt perturbation in simulated LLM decision audits.
-
Discrimination by LLMs: Cross-lingual Bias Assessment and Mitigation in Decision-Making and Summarisation
LLMs show significant demographic bias in decision-making, favoring women, younger ages, and certain minority backgrounds; summarization shows little bias, and bias patterns largely transfer from English to Dutch.
-
Not as Sweet by Another Name: An Empirical Study of Format Robustness in LLM Document Workflows
Changing the file format of identical input content changes LLM workflow decisions in 41% of cases on average and can reduce accuracy by up to 56 percentage points, with CSV the most error-prone format.
Discussion (0). Sign in to comment.