Pith. sign in

REVIEW 2 cited by

Few-shot Instruction Prompts for Pretrained Language Models to Detect Social Biases

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.07868 v2 pith:XNURO64G submitted 2021-12-15 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelsbiaslabeledsocialbiasesfew-shotlanguagecompared
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Detecting social bias in text is challenging due to nuance, subjectivity, and difficulty in obtaining good quality labeled datasets at scale, especially given the evolving nature of social biases and society. To address these challenges, we propose a few-shot instruction-based method for prompting pre-trained language models (LMs). We select a few class-balanced exemplars from a small support repository that are closest to the query to be labeled in the embedding space. We then provide the LM with instruction that consists of this subset of labeled exemplars, the query text to be classified, a definition of bias, and prompt it to make a decision. We demonstrate that large LMs used in a few-shot context can detect different types of fine-grained biases with similar and sometimes superior accuracy to fine-tuned models. We observe that the largest 530B parameter model is significantly more effective in detecting social bias compared to smaller models (achieving at least 13% improvement in AUC metric compared to other models). It also maintains a high AUC (dropping less than 2%) when the labeled repository is reduced to as few as $100$ samples. Large pretrained language models thus make it easier and quicker to build new bias detectors.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Prosocial Behavior Detection in Player Game Chat: From Aligning Human-AI Definitions to Efficient Annotation at Scale

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A hybrid system using a calibrated SVM plus selective GPT-4o fallback detects prosocial game chat at roughly 0.90 precision while cutting LLM inference cost by about 70%.

  2. Self-Anchored Attention Model for Sample-Efficient Classification of Prosocial Text Chat

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A self-anchored attention classifier using the whole training set as anchor features achieves 0.836 AUC for prosocial chat in Call of Duty: Modern Warfare II, 7.9% above the best benchmark.

Pith tools