Pith. sign in

REVIEW 2 cited by

It's COMPASlicated: The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.05498 v3 pith:FWOYROQG submitted 2021-06-10 cs.CY

classification cs.CY
keywords datasetsfairnessalgorithmiccontextpracticesusedwithoutassumptions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Risk assessment instrument (RAI) datasets, particularly ProPublica's COMPAS dataset, are commonly used in algorithmic fairness papers due to benchmarking practices of comparing algorithms on datasets used in prior work. In many cases, this data is used as a benchmark to demonstrate good performance without accounting for the complexities of criminal justice (CJ) processes. However, we show that pretrial RAI datasets can contain numerous measurement biases and errors, and due to disparities in discretion and deployment, algorithmic fairness applied to RAI datasets is limited in making claims about real-world outcomes. These reasons make the datasets a poor fit for benchmarking under assumptions of ground truth and real-world impact. Furthermore, conventional practices of simply replicating previous data experiments may implicitly inherit or edify normative positions without explicitly interrogating value-laden assumptions. Without context of how interdisciplinary fields have engaged in CJ research and context of how RAIs operate upstream and downstream, algorithmic fairness practices are misaligned for meaningful contribution in the context of CJ, and would benefit from transparent engagement with normative considerations and values related to fairness, justice, and equality. These factors prompt questions about whether benchmarks for intrinsically socio-technical systems like the CJ system can exist in a beneficial and ethical way.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Falsifying Discriminant Validity of Predictive Algorithms

    stat.ME 2026-01 reject novelty 6.0 of 10

    A calibrated-loss statistical test for discriminant validity that, on the paper's own examples, flags race and age as predicted as well as or better than intended outcomes but is confounded by differing outcome base rates.

  2. FairPOT: Balancing AUC Performance and Fairness with Proportional Optimal Transport

    cs.LG 2025-08 unverdicted novelty 6.0 of 10

    FairPOT selectively transports the top-lambda quantile of risk scores via optimal transport to balance AUC fairness against overall AUC performance, including partial AUC extensions.

Pith tools