Pith. sign in

REVIEW 2 cited by

Recovering from Biased Data: Can Fairness Constraints Improve Accuracy?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.01094 v2 pith:QGYK5KB5 submitted 2019-12-02 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords biaseddatafairnesstrainingaccuracyclassifierconsiderconstraints
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multiple fairness constraints have been proposed in the literature, motivated by a range of concerns about how demographic groups might be treated unfairly by machine learning classifiers. In this work we consider a different motivation; learning from biased training data. We posit several ways in which training data may be biased, including having a more noisy or negatively biased labeling process on members of a disadvantaged group, or a decreased prevalence of positive or negative examples from the disadvantaged group, or both. Given such biased training data, Empirical Risk Minimization (ERM) may produce a classifier that not only is biased but also has suboptimal accuracy on the true data distribution. We examine the ability of fairness-constrained ERM to correct this problem. In particular, we find that the Equal Opportunity fairness constraint (Hardt, Price, and Srebro 2016) combined with ERM will provably recover the Bayes Optimal Classifier under a range of bias models. We also consider other recovery methods including reweighting the training data, Equalized Odds, and Demographic Parity. These theoretical results provide additional motivation for considering fairness interventions even if an actor cares primarily about accuracy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 14 citations worldwide. Full citation record

  1. CaTE Data Curation for Trustworthy AI

    cs.LG 2025-08 accept novelty 4.0 of 10

    A synthesis of data curation practices for trustworthy AI, framed around an actionable definition of trustworthiness and a decision tree.

  2. Algorithmic Approaches to Sequential Decision-Making and Social Epistemology

    cs.DS 2026-07 conditional novelty 3.0 of 10

    For improving multi-armed bandits, randomized algorithms achieve a near-tight Θ~(√k) worst-case competitive ratio, and polynomially many historical instances suffice to tune a curvature parameter; pessimism traps and ...

Pith tools