Pith. sign in

REVIEW 8 cited by

Evaluating and Mitigating Discrimination in Language Model Decisions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03689 v1 pith:DUZUK73O submitted 2023-12-06 cs.CL

classification cs.CL
keywords discriminationcaseslanguagedecisionsmodelpotentialapplyingevaluating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As language models (LMs) advance, interest is growing in applying them to high-stakes societal decisions, such as determining financing or housing eligibility. However, their potential for discrimination in such contexts raises ethical concerns, motivating the need for better methods to evaluate these risks. We present a method for proactively evaluating the potential discriminatory impact of LMs in a wide range of use cases, including hypothetical use cases where they have not yet been deployed. Specifically, we use an LM to generate a wide array of potential prompts that decision-makers may input into an LM, spanning 70 diverse decision scenarios across society, and systematically vary the demographic information in each prompt. Applying this methodology reveals patterns of both positive and negative discrimination in the Claude 2.0 model in select settings when no interventions are applied. While we do not endorse or permit the use of language models to make automated decisions for the high-risk use cases we study, we demonstrate techniques to significantly decrease both positive and negative discrimination through careful prompt engineering, providing pathways toward safer deployment in use cases where they may be appropriate. Our work enables developers and policymakers to anticipate, measure, and address discrimination as language model capabilities and applications continue to expand. We release our dataset and prompts at https://huggingface.co/datasets/Anthropic/discrim-eval

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 10 citations worldwide. Full citation record

  1. Reference Feature Atlases for Mechanistic Auditing of Language Models

    cs.AI 2026-06 conditional novelty 7.0 of 10

    A reference feature atlas attaches a new language model with a linear decoder and surfaces panel-uncovered structure via a separate residual dictionary.

  2. Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Alignment on Llama 3 reduces explicit bias but amplifies implicit bias, because aligned models no longer represent 'black' and 'white' as racial concepts in ambiguous contexts.

  3. FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Audit format (rate vs rank vs allocate; transparent vs disguised) reverses the apparent direction of LLM demographic bias, while causal framing of need dominates allocations by roughly an order of magnitude.

  4. DeFrame: Debiasing Large Language Models Against Framing Effects

    cs.CL 2026-02 conditional novelty 6.0 of 10

    LLM fairness scores shift substantially with positive vs negative framing of the same question, and DeFrame—a three-step self-revision prompt—reduces both average bias and this framing gap.

  5. White-Box Sensitivity Auditing with Steering Vectors

    cs.CY 2026-01 unverdicted novelty 6.0 of 10

    White-box steering along gender/race concept vectors reveals larger and more consistent bias signals than black-box prompt perturbation in simulated LLM decision audits.

  6. Discrimination by LLMs: Cross-lingual Bias Assessment and Mitigation in Decision-Making and Summarisation

    cs.CL 2025-09 conditional novelty 6.0 of 10

    LLMs show significant demographic bias in decision-making, favoring women, younger ages, and certain minority backgrounds; summarization shows little bias, and bias patterns largely transfer from English to Dutch.

  7. Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)

    cs.CY 2025-05 reject novelty 6.0 of 10

    Placing demographic audience information in system prompts rather than user prompts shifts sentiment and ranking outputs across six commercial LLMs, but the design confounds position with instruction content.

  8. Not as Sweet by Another Name: An Empirical Study of Format Robustness in LLM Document Workflows

    cs.SE 2026-07 conditional novelty 5.0 of 10

    Changing the file format of identical input content changes LLM workflow decisions in 41% of cases on average and can reduce accuracy by up to 56 percentage points, with CSV the most error-prone format.

Pith tools