REVIEW 8 cited by
Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove Them
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Word embeddings are widely used in NLP for a vast range of tasks. It was shown that word embeddings derived from text corpora reflect gender biases in society. This phenomenon is pervasive and consistent across different word embedding models, causing serious concern. Several recent works tackle this problem, and propose methods for significantly reducing this gender bias in word embeddings, demonstrating convincing results. However, we argue that this removal is superficial. While the bias is indeed substantially reduced according to the provided bias definition, the actual effect is mostly hiding the bias, not removing it. The gender bias information is still reflected in the distances between "gender-neutralized" words in the debiased embeddings, and can be recovered from them. We present a series of experiments to support this claim, for two debiasing methods. We conclude that existing bias removal techniques are insufficient, and should not be trusted for providing gender-neutral modeling.
Forward citations
Cited by 8 Pith papers
-
Understanding Undesirable Word Embedding Associations
For matrix-factorization embeddings, subspace projection can be shown to debias the reconstructed matrix, WEAT systematically overstates associations because of word-frequency and effect-size flaws, and RIPA shows SGN...
-
Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
Unsigned differential activations locate a few GLU-MLP neurons whose zeroing surgically destabilizes demographic bias while retaining ~99.5% of measured capabilities.
-
Mitigating Gender Bias in Contextual Word Embeddings
Regularized masked-language modeling and name-masking reduce gender bias in embeddings, but the contextual results rely heavily on evaluation metrics aligned with the training objective.
-
Rotate King to get Queen: Word Relationships as Orthogonal Transformations in Embedding Space
Word relationships in embedding space can be represented as orthogonal or linear transformations, not just translation vectors, with comparable or better analogy-solving accuracy.
-
Unlearn Dataset Bias in Natural Language Inference by Fitting the Residual
DRiFt, a residual-fitting debiasing algorithm, improves NLI model accuracy on challenge sets like HANS by training on examples a biased model cannot solve.
-
Debiasing Embeddings for Reduced Gender Bias in Text Classification
Standard debiasing of word embeddings increases the gender bias of a downstream occupation classifier, while a strong debiasing variant that removes the gender subspace from all words reduces bias and preserves accuracy.
-
Paired-Consistency: An Example-Based Model-Agnostic Approach to Fairness Regularization in Machine Learning
The paper introduces paired-consistency, a metric and regularizer that enforces fairness by penalizing models that give different predictions to expert-selected pairs of examples that should be treated alike.
-
A Survey on Bias and Fairness in Machine Learning
This survey catalogs types of bias, fairness definitions, and mitigation strategies across ML domains.
Discussion (0). Continue with ORCID to comment.