Pith. sign in

REVIEW 2 cited by

On Adversarial Bias and the Robustness of Fair Machine Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.08669 v1 pith:OFD745CG submitted 2020-06-15 stat.ML cs.CRcs.CYcs.LG

classification stat.MLcs.CRcs.CYcs.LG
keywords datafairfairnesslearningmachineadversarialattacksmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Optimizing prediction accuracy can come at the expense of fairness. Towards minimizing discrimination against a group, fair machine learning algorithms strive to equalize the behavior of a model across different groups, by imposing a fairness constraint on models. However, we show that giving the same importance to groups of different sizes and distributions, to counteract the effect of bias in training data, can be in conflict with robustness. We analyze data poisoning attacks against group-based fair machine learning, with the focus on equalized odds. An adversary who can control sampling or labeling for a fraction of training data, can reduce the test accuracy significantly beyond what he can achieve on unconstrained models. Adversarial sampling and adversarial labeling attacks can also worsen the model's fairness gap on test data, even though the model satisfies the fairness constraint on training data. We analyze the robustness of fair machine learning through an empirical evaluation of attacks on multiple algorithms and benchmark datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 37 citations worldwide. Full citation record

  1. Covert Attacks on Machine Learning Training in Passively Secure MPC

    cs.CR 2025-05 conditional novelty 7.0 of 10

    An active adversary can exploit additive error injection in passively secure MPC training to poison models, amplify membership inference, reduce fairness, and reconstruct exact training data.

  2. Hiding in Plain Sight: An Effective Physical Adversarial Patch Attack against Visual-Infrared Fused Face Detection

    cs.CR 2026-07 conditional novelty 6.0 of 10

    A jointly optimized gradient-mask plus band-aid patch reportedly bypasses visible-infrared fused face detectors with >90% attack success in both digital and physical settings.

Pith tools