Pith. sign in

REVIEW 4 major objections 6 minor 134 references

Access Denied: Meaningful Data Access for Quantitative Algorithm Audits

T0 review · 4 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read Synthetic data erases the disparities that algorithm audits are designed to detect, and should not replace real data in fairness evaluations.

desk verdict First systematic comparison of audit access levels; the synthetic-data finding is real but the strong version overreaches. read the letter →

arxiv 2502.00428 v1 pith:VGFGHKMV submitted 2025-02-01 cs.HC cs.CY

classification cs.HCcs.CY
keywords algorithmauditinggroupparitymetricssyntheticdatadifferentialprivacyminimizationfairnessevaluationaccessprivacy-enhancingtechnologies
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish which data-sharing practices let an outside auditor reliably measure whether an algorithm treats demographic groups equally, and which practices produce misleading verdicts. The authors simulate audits of two real prediction models—recidivism and public health coverage—under three levels of access, degrading the audit data the way real organizations do: smaller samples, missing features, unevenly missing values, differentially private noise, and synthetic data. The central finding is that data minimization and anonymization can push fairness metrics so far from the truth that auditors declare parity where bias exists, with synthetic data the worst offender: every generator tested erases group disparities. The paper concludes that synthetic data should not replace real data in fairness evaluations, while differentially private confusion matrices can stay reliable at privacy budgets that are realistic in practice.

What carries the argument

The machinery is a controlled audit simulation. The authors train an XGBoost model on 70% of a real dataset, hold out 30% as the audit set, compute baseline group parity metrics (Statistical Parity Difference for recidivism, Average Odds Difference for health coverage) with bootstrapped 95% confidence intervals, and then re-run the same metrics after degrading the audit data in five ways: subsampling, feature removal, disparate missingness, differentially private noise on aggregate confusion matrices, and synthetic data generated by Gaussian Copula, CT-GAN, Copula-GAN, and PrivBayes. Reliability is measured as the proportion of experimental metric values that fall inside the baseline confidence interval, plus a classification of interpretation errors (Type 1, Type 2, and reverse errors). This design lets the authors attribute shifts in audit conclusions to specific data-sharing practices rather than to metric or model choice.

What would settle it

Run the same audit protocol with a synthetic-data generator not in the paper's set, such as a modern diffusion-based tabular model, on the NIJ and ACS datasets; if the synthetic samples reproduce the baseline parity metrics' confidence intervals at high overlap rates for both low- and high-disparity cases, the claim that synthetic data systematically erases disparities would be falsified for that generator class.

Watch

Extended reading notes

Core claim

Simulating an auditor computing group parity metrics on two real-world prediction tasks (recidivism and public health coverage), the paper finds that privacy-protective data-sharing practices systematically degrade audit reliability in different ways. A differentially private confusion matrix—an aggregate—produces accurate parity estimates at privacy budgets of ε ≥ 0.5 for samples above roughly 1,000 people, meaning strong privacy protection and reliable audit metrics are compatible under Scenario A access. Individual-level data with model outputs (Scenario B) and without model outputs (Scenario C) become unreliable when the sample drops below about 1,000, when key predictive features are missing, or when missing values are concentrated in the underprivileged group; even 1% disparate missingness reduced the proportion of estimates within the baseline 95% confidence interval to 72% under Scenario B and 25% under Scenario C. The strongest and most actionable claim is that synthetic data has a systematic tendency to invisibilize disparities: metrics computed on synthetic samples consistently indicate parity regardless of the true disparity level, producing Type 2 errors. The paper concludes that synthetic data is highly misleading for auditing and should not replace real data in fairness evaluations.

Load-bearing premise

The simulations assume the auditor has the ground-truth outcome and the protected characteristic for every person in the audit dataset; if real audit data lacks these fields, the measured error rates and the synthetic-data warning do not directly apply.

Editorial extensions

If this is right

  • Regulators and platforms that suggest synthetic data as a privacy solution for auditors should treat it as unfit for quantitative fairness evaluations unless a generator is shown to preserve group disparities.
  • Differentially private confusion matrices are a viable low-disclosure transparency mechanism: organizations could release them by default for public accountability without endangering audit reliability.
  • Auditors with individual-level data should require disclosure of sample size, the list of model predictors, and group-wise missingness rates before trusting a parity estimate.
  • Audits that replicate a model (Scenario C) are as reliable as direct score access only if the audit sample is roughly 160% larger for the recidivism case, and they collapse under missing features or synthetic data, so direct access to model outputs should be prioritized.
  • Even small amounts of missing data concentrated in an underprivileged group can reverse an audit conclusion, which makes missingness diagnostics a necessary part of standard audit practice.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension is to test whether the invisibilization result holds for newer generative tabular models such as diffusion-based synthesizers; if some generators preserve disparities, the paper's blanket conclusion would need to be scoped by generator class.
  • The results imply that regulatory proposals for remote data science or sandbox access should prefer differentially private aggregates over synthetic data releases, since the former preserve audit signal while the latter systematically destroy it.
  • Because intersectional subgroups are smaller than the two-group samples studied here, the sample-size thresholds found in this paper imply that intersectional audits will be even more fragile under data minimization, strengthening the case for privileged access to such assessments.
  • The Type 2 error result suggests a concrete disclosure rule: data holders who release synthetic or minimized data should be required to document exactly what was removed or generated, so auditors can discount the data rather than be misled.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper investigates how limited data access affects the reliability of quantitative algorithm audits that estimate group parity metrics. Through simulations on two public datasets (NIJ Recidivism and ACS Public Coverage), the authors compare three access scenarios: aggregate confusion matrices only (A), individual-level data with model predictions (B), and individual-level data without model predictions (C). They examine five quality-loss factors: sample size reduction, feature removal, disparate missing values, differentially private aggregation, and synthetic data generation. The main claims are that data minimization and anonymization can substantially increase error rates, that differentially private confusion matrices are generally reliable, and that synthetic data systematically invisibilizes disparities and should not replace real data in fairness evaluations. The paper also discusses regulatory and HCI implications.

Significance. If the strong claims are supported, the paper makes a valuable contribution to the algorithmic auditing literature by providing empirical evidence on how common data-sharing practices can undermine audit integrity. The experimental design is a notable strength: 100 train/audit splits, 500 bootstrap repetitions, multiple model classes, multiple parity metrics, and two realistic public datasets. The paper also engages seriously with the legal and HCI dimensions of auditor access. However, the significance hinges on the validity of the headline conclusions about differentially private aggregates and synthetic data, which are currently not fully supported by the reported statistics.

major comments (4)
  1. [Section 4.1, Table 4] The statement that with n=1,000 a privacy budget of epsilon=0.5 or above yields estimations that are 'highly unlikely to lead to interpretation errors' is contradicted by Table 4. At n=1,000, the proportion of metric values within the baseline 95% CI is 0.08 for ACS AOD high disparity at epsilon=0.5, 0.20 at epsilon=1, and 0.41 and 0.63 for NIJ SPD high and low disparity at epsilon=0.5. These values are far below the 0.70 threshold the authors themselves use to indicate high overlap. This claim is load-bearing for the recommendation that differentially private confusion matrices are well-suited for public releases, and it should be revised or supported with direct error-rate reporting.
  2. [Section 3.5.3 and 4.2.4] The high-disparity condition is constructed by reassigning 95% of positive predictions to the underprivileged group, making the protected attribute a near-deterministic function of the model output. Synthetic generators trained on the joint distribution of features and outcomes have no feature-based signal for this disparity and are structurally unlikely to reproduce it. The observed 'invisibilization' of disparities in the high-disparity condition is therefore potentially an artifact of this construction rather than a general property of synthetic data for audits of real biased models, where disparity typically flows through features correlated with group membership. The paper's own trained models have only low disparity, so they do not independently test the strong claim that synthetic data 'has a strong tendency to invisibilize disparities' (Section 4.4). I recommend an additional experiment where high disparity is induced through feature-based mechanisms, such as strengthening group-correlated coefficients or applying group-specific base rates.
  3. [Section 3.4 and throughout Results] The paper defines Type 1, Type 2, and reverse errors based on whether the baseline and experimental confidence intervals share the same configuration, but it never reports the actual frequencies of these error types. The only quantitative measure reported is the proportion of metric values within the baseline CI (e.g., Tables 4-11), which is not the same as an error rate. Consequently, claims such as 'highly unlikely to lead to interpretation errors' (Section 4.1) and 'synthetic data generation has a strong tendency to invisibilize disparities, often leading to Type 2 errors' (Section 4.4) are not directly supported by the reported statistics. The authors should report the fraction of splits where each error type occurs for the headline conditions, at least for the main results in Section 4.
  4. [Section 3.1 and 5.5] The simulations assume the audit dataset contains ground-truth outcomes and protected characteristics for every individual. Section 5.5 acknowledges this excludes applications where sensitive features cannot be collected, but the abstract and policy recommendations (e.g., Section 5.2) draw broad conclusions about data access and anonymization without consistently carrying this caveat. Prior work cited by the authors (Kallus et al. [78]) shows that proxies for protected attributes are unreliable for disparity assessment. The measured error rates for missing features, synthetic data, and sample size therefore do not automatically transfer to settings without protected characteristics. The paper should either explicitly restrict its conclusions to the assumed setting or include an additional analysis where protected attributes must be inferred.
minor comments (6)
  1. [Table 4 and Section 4.1] The claim that Access Scenario A offers 'high metric accuracy' should be qualified by sample size and epsilon, since Table 4 shows low overlap for small n even at moderate epsilon values.
  2. [Section 3.2.5] The methods section states that PrivBayes is used with the DPART library but does not report the privacy budget used for PrivBayes; Table 10 lists two epsilon values (1 and 5). Please state the epsilon choices in the methods.
  3. [Section 3.6 and Appendix A.2] The text says 'results generalized across these metrics,' but the appendix only shows results for the ACS dataset across three metrics; the claim for the NIJ dataset is not substantiated in the appendix.
  4. [Figure 4] The cumulative feature-removal plots are described as 'at the F14 point, features 14 to 18 are missing,' but the text also refers to '5-7 low importance features' for NIJ. Please clarify the correspondence between the x-axis labels and the number of removed features.
  5. [Section 3.5.3] The high-disparity reassignment procedure should specify whether it is applied only to the audit set or also to the training set, and whether the ground-truth labels are recomputed after reassignment. This is important for interpreting the synthetic data results.
  6. [Section 5.5] The statement that intersectional assessments 'would require even more granular and higher quality data' is speculative; consider marking it explicitly as a conjecture or providing reference to supporting evidence.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity; empirical audit simulation with background-only self-citations.

full rationale

This paper is an empirical audit simulation rather than a derivational argument. The central pipeline is to split each dataset 70/30, train an XGBoost model on the training portion, compute parity metrics with bootstrapped 95% confidence intervals on the held-out 30% audit set as the baseline, then perturb that same audit set through subsampling, feature removal, missingness, differential privacy, or synthetic data generation, and compare interval overlap and interpretation error types. There is no fitted parameter that is subsequently presented as a prediction; the baseline and experimental conditions are distinct evaluations over the same underlying data, and the conclusions, such as 'Synthetic data generation has a strong tendency to invisibilize disparities, often leading to Type 2 errors,' are empirical generalizations from those comparisons. The authors cite their own prior work (Gadotti et al. [58], Houssiau et al. [71], Annamalai et al. [12], and Veale and Binns [130]) in background and discussion sections, but none of these citations supplies a load-bearing premise for the experimental design or for the central findings; they function as literature review and related prior results rather than as the justification for the paper's conclusions. The high-disparity case is constructed by reassigning protected attributes ('we reassign 95% of those who receive a positive prediction to the underprivileged group, and reassign the rest to the privileged group'), which is a potential external-validity limitation because the disparity is encoded in the protected-attribute label rather than flowing through features correlated with group membership; however, this is not circularity under the standards of this review. The synthetic-data result is an observed empirical outcome, not an identity, a definitional equivalence, or a fitted parameter renamed as a prediction. The paper also acknowledges its scope limits in Section 5.5, including the assumption that protected characteristics and ground truth are available. Overall, no derivation step reduces to its own inputs by construction, so no significant circularity is present.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new entities or theoretical constructs. Its premises are conditions on the audit setting (binary classifier, available labels and protected attributes, known model class in Scenario C) and the operational choice to treat the full audit set as the ground-truth benchmark. The high-disparity manipulation is a synthetic construction, listed as a free parameter for transparency. The design choices for sample sizes, epsilon grids, and missing-value fractions are experimental constants rather than fitted values.

free parameters (1)
  • High-disparity reassignment fraction = 0.95
    Section 3.5.3: 95% of records with positive predictions are reassigned to the underprivileged group to construct a high-disparity case. All 'high disparity' conclusions are conditional on this hand-chosen transformation.
assumptions (4)
  • domain assumption The audited model is a binary classifier.
    Section 3.1: 'we focus on binary classifiers as the target of our simulated audits'. Multi-class and risk-score models are deferred to future work, so the reliability numbers may not transfer.
  • domain assumption Audit data contains ground-truth labels and protected characteristics.
    Section 3.1: 'the audit data includes the ground-truth' and 'the protected characteristics'. Section 5.5 explicitly excludes settings where sensitive features are not collected.
  • domain assumption In Scenario C the auditor knows the model class and hyperparameters.
    Table 1 grants 'Model class & parameters (excluding weights)'. Successful replication depends on this knowledge being accurate.
  • domain assumption Baseline metrics on the full audit set approximate the model's true parity.
    Section 3.3: 95% CIs from the 30% holdout are used 'as baselines and proxies for the metrics' ground truth', treating the finite audit sample as if it defines the true value.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Access Denied: Meaningful Data Access for Quantitative Algorithm Audits." pith.science (2026). https://pith.science/paper/VGFGHKMV

@misc{pith2026250200428,
  author       = {Pith},
  title        = {Pith review of: Access Denied: Meaningful Data Access for Quantitative Algorithm Audits},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VGFGHKMV}},
  note         = {Machine review of arXiv:2502.00428}
}
read the original abstract

Independent algorithm audits hold the promise of bringing accountability to automated decision-making. However, third-party audits are often hindered by access restrictions, forcing auditors to rely on limited, low-quality data. To study how these limitations impact research integrity, we conduct audit simulations on two realistic case studies for recidivism and healthcare coverage prediction. We examine the accuracy of estimating group parity metrics across three levels of access: (a) aggregated statistics, (b) individual-level data with model outputs, and (c) individual-level data without model outputs. Despite selecting one of the simplest tasks for algorithmic auditing, we find that data minimization and anonymization practices can strongly increase error rates on individual-level data, leading to unreliable assessments. We discuss implications for independent auditors, as well as potential avenues for HCI researchers and regulators to improve data access and enable both reliable and holistic evaluations.

Figures

Figures reproduced from arXiv: 2502.00428 by the authors.

Figure 1
Figure 1. Experimental design flowchart, distinguishing between the simulated organization (in grey), baseline audit (in blue), [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Effect of differential privacy on metric reliability with full sample size and with a reduced ( [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Effect of sample size on metric reliability for Access B (left) and Access C (right). The red and yellow lines represent [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Effect of missing features on metric reliability for Access B (left) and Access C (right). On the x-axis, features are [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Effect of disparate missing values rates on metric reliability for Access B (left) and Access C (right). The red and [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: Metric reliability based on models used for synthetic data generation, for Scenario B (left) and Scenario C (right). [PITH_FULL_IMAGE:figures/full_fig_p011_6.png]
Figure 7
Figure 7. Figure 7: Results for DP-RF model audits, for the NIJ dataset (left) and ACS dataset (right). The models were trained with [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 8
Figure 8. Figure 8: Effect of sample size on metric reliability for Access B (left) and Access C (right) (ACS dataset) [PITH_FULL_IMAGE:figures/full_fig_p026_8.png]
Figure 9
Figure 9. Figure 9: Effect of missing features on metric reliability for Access B (left) and Access C (right) (ACS dataset) [PITH_FULL_IMAGE:figures/full_fig_p027_9.png]
Figure 10
Figure 10. Figure 10: Effect of disparate missing values rates on metric reliability for Access B (left) and Access C (right) (ACS dataset) [PITH_FULL_IMAGE:figures/full_fig_p028_10.png]
Figure 11
Figure 11. Figure 11: Effect of differential privacy on metric reliability for the ACS dataset, at full sample size ( [PITH_FULL_IMAGE:figures/full_fig_p029_11.png]
Figure 12
Figure 12. Figure 12: Metric reliability based on models used for synthetic data generation, for Scenario B (left) and Scenario C (right) [PITH_FULL_IMAGE:figures/full_fig_p030_12.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

134 extracted references · 77 canonical work pages

  1. [78]

    Assessing Algorithmic Fairness with Unobserved Protected Class Using Data Combination

    Kallus, N., Mao, X., and Zhou, A. Assessing Algorithmic Fairness with Unobserved Protected Class Using Data Combination. Manage. Sci. 68, 3 (Mar. 2022), 1959–1981

  2. [1]

    Technical methods for regulatory inspection of algo- rithmic systems: A survey of auditing methods for use in regulatory inspections of online harms in social media platforms, Dec

    Ada Lovelace Institute. Technical methods for regulatory inspection of algo- rithmic systems: A survey of auditing methods for use in regulatory inspections of online harms in social media platforms, Dec. 2021

  3. [2]

    Algorithmic accountability for the public sector, Aug

    Ada Lovelace Institute, AI Now Institute, and Open Government Part- nership. Algorithmic accountability for the public sector, Aug. 2021

  4. [3]

    A., Rybeck, G., Scheidegger, C., Smith, B., and Venkatasubramanian, S

    Adler, P., Falk, C., Friedler, S. A., Rybeck, G., Scheidegger, C., Smith, B., and Venkatasubramanian, S. Auditing Black-box Models for Indirect Influence, Nov. 2016

  5. [4]

    Designing and Implementing Medicaid Disease and Care Management Programs

    Agency for Healthcare Research and Quality . Designing and Implementing Medicaid Disease and Care Management Programs. Sec- tion 3: Selecting and Targeting Populations for a Care Management Pro- gram. https://ahrq.gov/patient-safety/settings/long-term-care/resource/hcbs/ medicaidmgmt/mm3.html, 2014

  6. [5]

    Aïvodji, U., Arai, H., Fortineau, O., Gambs, S., Hara, S., and Tapp, A.Fair- washing: The risk of rationalization, May 2019

  7. [6]

    Synthetic data generation

    Algorithm Audit. Synthetic data generation. https://algorithmaudit.eu/ technical-tools/sdg/

  8. [7]

    Automating Society Report, 2020

    Algorithm W atch. Automating Society Report, 2020

Show all 134 references
  1. [8]

    How Dutch activists got an invasive fraud detection algorithm banned, Apr

    Algorithm Watch. How Dutch activists got an invasive fraud detection algorithm banned, Apr. 2020

  2. [9]

    Discrimination through Optimization: How Facebook’s Ad Delivery Can Lead to Biased Outcomes

    Ali, M., Sapiezynski, P., Bogen, M., Korolova, A., Mislove, A., and Rieke, A. Discrimination through Optimization: How Facebook’s Ad Delivery Can Lead to Biased Outcomes. Proceedings of the ACM on Human-Computer Interaction 3 , CSCW (Nov. 2019), 1–30

  3. [10]

    Technical Response to Northpointe

    Angwin, J., and Larson, J. Technical Response to Northpointe. https://www. propublica.org/article/technical-response-to-northpointe, July 2016

  4. [11]

    Machine Bias: There’s software used across the country to predict future criminals

    Angwin, J., Larson, J., Mattu, S., and Kirchner, L. Machine Bias: There’s software used across the country to predict future criminals. And it’s biased against blacks. ProPublica (2016)

  5. [12]

    Annamalai, M. S. M. S., Gadotti, A., and Rocher, L. A Linear Reconstruction Approach for Attribute Inference Attacks against Synthetic Data, May 2024

  6. [13]

    Learning with Privacy at Scale

    Apple Differential Privacy Team. Learning with Privacy at Scale. https: //machinelearning.apple.com/research/learning-with-privacy-at-scale, Dec. 2017

  7. [14]

    Arawjo, I., Swoopes, C., V aithilingam, P., W attenberg, M., and Glassman, E. L. ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing. In Proceedings of the CHI Conference on Human Factors in Computing Systems (New York, NY, USA, May 2024), CHI ’24, A...

  8. [15]

    Problematic Machine Behavior: A Systematic Literature Review of Algorithm Audits, Feb

    Bandy, J. Problematic Machine Behavior: A Systematic Literature Review of Algorithm Audits, Feb. 2021

  9. [16]

    Fairness and Machine Learning: Limitations and Opportunities

    Barocas, S., Hardt, M., and Narayanan, A. Fairness and Machine Learning: Limitations and Opportunities. The MIT Press, 2023

  10. [17]

    Synthetic data protection: Towards a paradigm change in data regulation? Big Data & Society 11 , 1 (Mar

    Beduschi, A. Synthetic data protection: Towards a paradigm change in data regulation? Big Data & Society 11 , 1 (Mar. 2024), 20539517241231277

  11. [18]

    Belgodere, B., Dognin, P., Ivankay, A., Melnyk, I., Mroueh, Y., Mojsilovic, A., Navratil, J., Nitsure, A., Padhi, I., Rigotti, M., Ross, J., Schiff, Y., Vedpathak, R., and Young, R. A. Auditing and Generating Synthetic Data with Controllable Trust Trade-offs, June 2024

  12. [19]

    Besse, P., del Barrio, E., Gordaliza, P., and Loubes, J.-M.Confidence Intervals for Testing Disparate Impact in Fair Learning, July 2018

  13. [20]

    Bhat, A., Coursey, A., Hu, G., Li, S., Nahar, N., Zhou, S., Kästner, C., and Guo, J. L. Aspirations and Practice of ML Model Documentation: Moving the Needle with Nudging and Traceability. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (New Yor...

  14. [21]

    Fairness in Machine Learning: Lessons from Political Philosophy, Mar

    Binns, R. Fairness in Machine Learning: Lessons from Political Philosophy, Mar. 2021

  15. [22]

    In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (New York, NY, USA, Apr

    Binns, R., V an Kleek, M., Veale, M., Lyngs, U., Zhao, J., and Shadbolt, N.’It’s Reducing a Human Being to a Percentage’: Perceptions of Justice in Algorithmic Decisions. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (New York, NY, USA, Apr. 2...

  16. [23]

    D.AI auditing: The Broken Bus on the Road to AI Accountability, Jan

    Birhane, A., Steed, R., Ojewale, V., Vecchione, B., and Raji, I. D.AI auditing: The Broken Bus on the Road to AI Accountability, Jan. 2024

  17. [24]

    S., McFowland III, E., and Neill, D

    Boxer, K. S., McFowland III, E., and Neill, D. B. Auditing Predictive Models for Intersectional Biases, June 2023

  18. [25]

    Suspicion Machine Methodology, Mar

    Braun, J.-C., Constantaras, E., Aung, H., Geiger, G., Mehrotra, D., and Howden, D. Suspicion Machine Methodology, Mar. 2023

  19. [26]

    The algorithm audit: Scoring the algorithms that score us

    Brown, S., Davidovic, J., and Hasan, A. The algorithm audit: Scoring the algorithms that score us. Big Data & Society 8 , 1 (Jan. 2021), 205395172098386

  20. [27]

    Gender Shades: Intersectional Accuracy Dispar- ities in Commercial Gender Classification

    Buolamwini, J., and Gebru, T. Gender Shades: Intersectional Accuracy Dispar- ities in Commercial Gender Classification. In Proceedings of the 1st Conference on Fairness, Accountability and Transparency (Jan. 2018), PMLR, pp. 77–91

  21. [28]

    A., Epperson, W., Hohman, F., Kahng, M., Morgenstern, J., and Chau, D

    Cabrera, Á. A., Epperson, W., Hohman, F., Kahng, M., Morgenstern, J., and Chau, D. H. FairVis: Visual Analytics for Discovering Intersectional Bias in Machine Learning. In 2019 IEEE Conference on Visual Analytics Science and Technology (V AST)(Oct. 2019), pp. 46–56

  22. [29]

    Casper, S., Ezell, C., Siegmann, C., Kolt, N., Curtis, T. L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobbhahn, M., Sharkey, L., Krishna, S., Von Hagen, M., Alberti, S., Chan, A., Sun, Q., Gerovitch, M., Bau, D., Tegmark, M., Krueger, D., and Hadfield-Menell, D. Black-...

  23. [30]

    G., and Cosentini, A

    Castelnovo, A., Crupi, R., Greco, G., Regoli, D., Penco, I. G., and Cosentini, A. C. A clarification of the nuances in the fairness metrics landscape. Scientific Reports 12, 1 (Mar. 2022), 4209

  24. [31]

    Interim report: Review into bias in algorithmic decision-making, 2019

    Centre for Data Ethics and Innovation. Interim report: Review into bias in algorithmic decision-making, 2019

  25. [32]

    AI Transparency in practice: What was learnt from third-party audit of recommender systems at LinkedIn and Dailymotion, Oct

    Chen, J., Bandy, J., Buckley, D., and Bhatia, R. AI Transparency in practice: What was learnt from third-party audit of recommender systems at LinkedIn and Dailymotion, Oct. 2024

  26. [33]

    Beyond Fairness Metrics: Roadblocks and Challenges for Ethical AI in Practice, Aug

    Chen, J., Storchan, V., and Kurshan, E. Beyond Fairness Metrics: Roadblocks and Challenges for Ethical AI in Practice, Aug. 2021

  27. [34]

    Fair prediction with disparate impact: A study of bias in recidivism prediction instruments, Oct

    Chouldechova, A. Fair prediction with disparate impact: A study of bias in recidivism prediction instruments, Oct. 2016

  28. [35]

    Fair Prediction with Disparate Impact: A Study of Bias in Recidivism Prediction Instruments

    Chouldechova, A. Fair Prediction with Disparate Impact: A Study of Bias in Recidivism Prediction Instruments. Big Data 5, 2 (June 2017), 153–163

  29. [36]

    D., Nilforoshan, H., Shroff, R., and Goel, S

    Corbett-Davies, S., Gaebler, J. D., Nilforoshan, H., Shroff, R., and Goel, S. The Measure and Mismeasure of Fairness, Aug. 2023

  30. [37]

    In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (Halifax NS Canada, Aug

    Corbett-Davies, S., Pierson, E., Feller, A., Goel, S., and Huq, A.Algorithmic Decision Making and the Cost of Fairness. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (Halifax NS Canada, Aug. 2017), ACM, pp. 797–806

  31. [38]

    D., and Buolamwini, J.Who Audits the Auditors? Recommendations from a field scan of the algorithmic auditing ecosystem

    Costanza-Chock, S., Raji, I. D., and Buolamwini, J.Who Audits the Auditors? Recommendations from a field scan of the algorithmic auditing ecosystem. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Trans- parency (New York, NY, USA, June 2022), FAccT ...

  32. [39]

    Advancing Differential Privacy: Where We Are Now and Future Directions for Real-World Deployment

    Cummings, R., Desfontaines, D., Evans, D., Geambasu, R., Huang, Y., Jagiel- ski, M., Kairouz, P., Kamath, G., Oh, S., Ohrimenko, O., Papernot, N., Rogers, R., Shen, M., Song, S., Su, W., Terzis, A., Thakurta, A., V assilvitskii, S., W ang, Y.-X., Xiong, L., Yekhanin, S., Yu, D...

  33. [40]

    Home Office says it will abandon its racist visa algorithm - after we sued them, Aug

    Dark, M. Home Office says it will abandon its racist visa algorithm - after we sued them, Aug. 2020

  34. [41]

    H., Guo, B., Devrio, A., Shen, H., Eslami, M., and Holstein, K

    Deng, W. H., Guo, B., Devrio, A., Shen, H., Eslami, M., and Holstein, K. Understanding Practices, Challenges, and Opportunities for User-Engaged Al- gorithm Auditing in Industry Practice. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (New York...

  35. [42]

    H., Nagireddy, M., Lee, M

    Deng, W. H., Nagireddy, M., Lee, M. S. A., Singh, J., Wu, Z. S., Holstein, K., and Zhu, H. Exploring How Machine Learning Practitioners (Try To) Use Fairness Toolkits. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (New York, NY, USA, J...

  36. [43]

    Toward User-Driven Algorithm Auditing: Investigating users’ strategies for uncovering harmful algorithmic behavior

    DeVos, A., Dhabalia, A., Shen, H., Holstein, K., and Eslami, M. Toward User-Driven Algorithm Auditing: Investigating users’ strategies for uncovering harmful algorithmic behavior. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New York, NY, US...

  37. [44]

    Retiring Adult: New Datasets for Fair Machine Learning, Jan

    Ding, F., Hardt, M., Miller, J., and Schmidt, L. Retiring Adult: New Datasets for Fair Machine Learning, Jan. 2022

  38. [45]

    The accuracy, fairness, and limits of predicting recidivism

    Dressel, J., and Farid, H. The accuracy, fairness, and limits of predicting recidivism. Science Advances 4, 1 (Jan. 2018), eaao5580

  39. [46]

    Differential Privacy

    Dwork, C. Differential Privacy. In Automata, Languages and Programming (Berlin, Heidelberg, 2006), M. Bugliesi, B. Preneel, V. Sassone, and I. Wegener, Eds., Springer, pp. 1–12

  40. [47]

    InProceedings of the Forty-First Annual ACM Symposium on Theory of Computing (New York, NY, USA, May 2009), STOC ’09, Association for Computing Machinery, pp

    Dwork, C., and Lei, J.Differential privacy and robust statistics. InProceedings of the Forty-First Annual ACM Symposium on Theory of Computing (New York, NY, USA, May 2009), STOC ’09, Association for Computing Machinery, pp. 371–380

  41. [48]

    Access to data and algorithms: For an effective DMA and DSA implementation, Mar

    Edelson, L., Graef, I., and Lancieri, F. Access to data and algorithms: For an effective DMA and DSA implementation, Mar. 2023

  42. [49]

    D., Datar, A., and Coltin, K.Government jobs of the future: What will government work look like in 2025 and beyond?, 2019

    Eggers, W. D., Datar, A., and Coltin, K.Government jobs of the future: What will government work look like in 2025 and beyond?, 2019

  43. [50]

    European Commission. Proposal for a REGULATION OF THE EUROPEAN PARLIAMENT AND OF THE COUNCIL LAYING DOWN HARMONISED RULES ON ARTIFICIAL INTELLIGENCE (ARTIFICIAL INTELLIGENCE ACT) AND AMENDING CERTAIN UNION LEGISLATIVE ACTS, 2021

  44. [51]

    Report of the European Digital Media Observatory’s Working Group on Platform-to-Researcher Data Access, May 2022

    European Digital Media Observatory. Report of the European Digital Media Observatory’s Working Group on Platform-to-Researcher Data Access, May 2022

  45. [52]

    Digital Services Act, Oct

    European Union. Digital Services Act, Oct. 2022

  46. [53]

    European Union. Regulation (EU) 2024/1689 of the European Parliament Access Denied: Meaningful Data Access for Quantitative Algorithm Audits CHI ’25, April 26-May 1, 2025, Yokohama, Japan and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligenc...

  47. [54]

    Statistically Valid Inferences from Privacy-Protected Data

    Evans, G., King, G., Schwenzfeier, M., and Thakurta, A. Statistically Valid Inferences from Privacy-Protected Data. American Political Science Review 117 , 4 (Nov. 2023), 1275–1290

  48. [55]

    Machine Bias: There’s Software Used Across the Country to Predict Future Criminals. And It’s Biased Against Blacks

    Flores, A. W., and Bechtel, K. False Positives, False Negatives, and False Analyses: A Rejoinder to “Machine Bias: There’s Software Used Across the Country to Predict Future Criminals. And It’s Biased Against Blacks. ”.Federal Probation Journal (2016)

  49. [56]

    Translating Principles into Practices of Digital Ethics: Five Risks of Being Unethical

    Floridi, L. Translating Principles into Practices of Digital Ethics: Five Risks of Being Unethical. In Ethics, Governance, and Policies in Artificial Intelligence , L. Floridi, Ed. Springer International Publishing, Cham, 2021, pp. 81–90

  50. [57]

    Faking Fairness via Stealthily Biased Sampling

    Fukuchi, K., Hara, S., and Maehara, T. Faking Fairness via Stealthily Biased Sampling. Proceedings of the AAAI Conference on Artificial Intelligence 34 , 01 (Apr. 2020), 412–419

  51. [58]

    Anonymization: The imperfect science of using data while preserving privacy

    Gadotti, A., Rocher, L., Houssiau, F., Creţu, A.-M., and De Montjoye, Y.-A. Anonymization: The imperfect science of using data while preserving privacy. Science Advances 10, 29 (July 2024), eadn7053

  52. [59]

    Auditing Algorithms: On Lessons Learned and the Risks of Data Minimization

    Galdon Clavell, G., Martín Zamorano, M., Castillo, C., Smith, O., and Matic, A. Auditing Algorithms: On Lessons Learned and the Risks of Data Minimization. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (New York NY USA, Feb. 2020), ACM, pp. 265–271

  53. [60]

    W., Wallach, H., Daumé III, H., and Crawford, K

    Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., and Crawford, K. Datasheets for Datasets, Dec. 2021

  54. [61]

    2023), 104269

    Getzen, E., Ungar, L., Mowery, D., Jiang, X., and Long, Q.Mining for equitable health: Assessing the impact of missing data in electronic health records.Journal of Biomedical Informatics 139 (Mar. 2023), 104269

  55. [62]

    Characterizing Intersectional Group Fairness with Worst-Case Comparisons

    Ghosh, A., Genuit, L., and Reagan, M. Characterizing Intersectional Group Fairness with Worst-Case Comparisons. https://arxiv.org/abs/2101.01673v5, Jan. 2021

  56. [63]

    P., and Trehu, J

    Goodman, E. P., and Trehu, J. AI Audit Washing and Accountability. SSRN Electronic Journal (2022)

  57. [64]

    Heaven, W. D. Predictive policing algorithms are racist. They need to be dismantled. MIT Technology Review (July 2020)

  58. [65]

    Measuring Algorithmic Fairness

    Hellman, D. Measuring Algorithmic Fairness. Virginia Law Review 106, 4 (June 2020)

  59. [66]

    Hind, M., Houde, S., Martino, J., Mojsilovic, A., Piorkowski, D., Richards, J., and Varshney, K. R. Experiences with Improving the Transparency of AI Models and Services. In Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems (New York, NY, USA,...

  60. [67]

    Hoffmann, A. L. Where fairness fails: Data, algorithms, and the limits of antidiscrimination discourse. Information, Communication & Society 22 , 7 (June 2019), 900–915

  61. [68]

    M., and Levacher, K.Diffprivlib: The IBM Differential Privacy Library

    Holohan, N., Braghin, S., Aonghusa, P. M., and Levacher, K.Diffprivlib: The IBM Differential Privacy Library. https://dx.doi.org/10.48550/arXiv.1907.02444, July 2019

  62. [69]

    Holstein, K., Wortman Vaughan, J., Daumé, H., Dudik, M., and Wallach, H. Improving Fairness in Machine Learning Systems: What Do Industry Practi- tioners Need? In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (New York, NY, USA, May 2019), CHI ’1...

  63. [70]

    Dutch Childcare Allowance Scandal: The importance of investigation powers, Nov

    Hoogenboom, A. Dutch Childcare Allowance Scandal: The importance of investigation powers, Nov. 2022

  64. [71]

    Nature Communications 13, 1 (Jan

    Houssiau, F., Rocher, L., and de Montjoye, Y.-A.On the difficulty of achieving Differential Privacy in practice: User-level guarantees in aggregate location data. Nature Communications 13, 1 (Jan. 2022), 29

  65. [72]

    C., and Roth, A

    Hsu, J., Gaboardi, M., Haeberlen, A., Khanna, S., Narayan, A., Pierce, B. C., and Roth, A. Differential Privacy: An Economic Method for Choosing Epsilon, Feb. 2014

  66. [73]

    Measuring Misinformation in Video Search Platforms: An Audit Study on YouTube.Proc

    Hussein, E., Juneja, P., and Mitra, T. Measuring Misinformation in Video Search Platforms: An Audit Study on YouTube.Proc. ACM Hum.-Comput. Interact. 4, CSCW1 (May 2020), 48:1–48:27

  67. [74]

    Having your Privacy Cake and Eating it Too: Platform-supported Auditing of Social Media Algorithms for Public Interest

    Imana, B., Korolova, A., and Heidemann, J. Having your Privacy Cake and Eating it Too: Platform-supported Auditing of Social Media Algorithms for Public Interest. Proceedings of the ACM on Human-Computer Interaction 7 , CSCW1 (Apr. 2023), 1–33

  68. [75]

    Auditing algorithms: The existing landscape, role of regulators and future outlook, 2022

    Information Commissioner’s Office . Auditing algorithms: The existing landscape, role of regulators and future outlook, 2022

  69. [76]

    Two-Face: Ad- versarial Audit of Commercial Face Recognition Systems

    Jaiswal, S., Duggirala, K., Dash, A., and Mukherjee, A. Two-Face: Ad- versarial Audit of Commercial Face Recognition Systems. Proceedings of the International AAAI Conference on Web and Social Media 16 (May 2022), 381–392

  70. [77]

    Can I Trust My Fairness Metric? Assessing Fairness with Unlabeled Data and Bayesian Inference

    Ji, D., Smyth, P., and Steyvers, M. Can I Trust My Fairness Metric? Assessing Fairness with Unlabeled Data and Bayesian Inference. In Advances in Neural Information Processing Systems (2020), vol. 33, Curran Associates, Inc., pp. 18600– 18612

  71. [79]

    AlgorithmWatch forced to shut down Instagram monitoring project after threats from Facebook, Aug

    Kayser-Bril, N. AlgorithmWatch forced to shut down Instagram monitoring project after threats from Facebook, Aug. 2021

  72. [80]

    B., John, B

    Kery, M. B., John, B. E., O’Flaherty, P., Horvath, A., and Myers, B. A. To- wards Effective Foraging by Data Scientists to Find Past Analysis Choices. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (New York, NY, USA, May 2019), CHI ’19, Associ...

  73. [81]

    Blind Justice: Fairness with Encrypted Sensitive Attributes

    Kilbertus, N., Gascon, A., Kusner, M., Veale, M., Gummadi, K., and Weller, A. Blind Justice: Fairness with Encrypted Sensitive Attributes. In Proceedings of the 35th International Conference on Machine Learning (July 2018), PMLR, pp. 2630–2639

  74. [82]

    Inherent Trade-Offs in the Fair Determination of Risk Scores, Nov

    Kleinberg, J., Mullainathan, S., and Raghavan, M. Inherent Trade-Offs in the Fair Determination of Risk Scores, Nov. 2016

  75. [83]

    Kommiya Mothilal, R., Guha, S., and Ahmed, S. I. Towards a Non-Ideal Methodological Framework for Responsible ML. In Proceedings of the CHI Conference on Human Factors in Computing Systems (New York, NY, USA, May 2024), CHI ’24, Association for Computing Machinery, pp. 1–17

  76. [84]

    Punishing With Impunity: The Legacy of Risk Classification Assessment in Immigration Detention

    Koulish, R., and Evans, K. Punishing With Impunity: The Legacy of Risk Classification Assessment in Immigration Detention. Georgetown Immigration Law Journal 36, 1 (2021)

  77. [85]

    AI governance in the public sector: Three tales from the frontiers of automated decision-making in democratic settings

    Kuziemski, M., and Misuraca, G. AI governance in the public sector: Three tales from the frontiers of automated decision-making in democratic settings. Telecommunications Policy 44, 6 (July 2020), 101976

  78. [86]

    Scoring of welfare beneficia- ries: The indecency of CAF’s algorithm now undeniable

    La Quadrature du Net . Scoring of welfare beneficia- ries: The indecency of CAF’s algorithm now undeniable. https://www.laquadrature.net/en/2023/11/27/scoring-of-welfare-beneficiaries- the-indecency-of-cafs-algorithm-now-undeniable/, Nov. 2023

  79. [87]

    S., Pandit, A., Kalicki, C

    Lam, M. S., Pandit, A., Kalicki, C. H., Gupta, R., Sahoo, P., and Metaxa, D. Sociotechnical Audits: Broadening the Algorithm Auditing Lens to Investigate Targeted Advertising. Proc. ACM Hum.-Comput. Interact. 7 , CSCW2 (Oct. 2023), 360:1–360:37

  80. [88]

    K., Grgić-Hlača, N., Tschantz, M

    Lee, M. K., Grgić-Hlača, N., Tschantz, M. C., Binns, R., Weller, A., Carney, M., and Inkpen, K. Human-Centered Approaches to Fair and Responsible AI. In Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems (New York, NY, USA, Apr. 2020), CHI EA ’...

  81. [89]

    Lee, M. S. A., and Singh, J. The Landscape and Gaps in Open Source Fair- ness Toolkits. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (New York, NY, USA, May 2021), CHI ’21, Association for Computing Machinery, pp. 1–13

  82. [90]

    Levine, A. S. ’Chilling’: Facial recognition firm Clearview AI hits watchdog groups with subpoenas. POLITICO (Sept. 2021)

  83. [91]

    https://dx.doi.org/10.48550/arXiv.2403.04893, Mar

    Longpre, S., Kapoor, S., Klyman, K., Ramaswami, A., Bommasani, R., Blili- Hamelin, B., Huang, Y., Skowron, A., Yong, Z.-X., Kotha, S., Zeng, Y., Shi, W., Yang, X., Southen, R., Robey, A., Chao, P., Yang, D., Jia, R., Kang, D., Pentland, S., Narayanan, A., Liang, P., and Hender...

  84. [92]

    To predict and serve? Significance 13, 5 (2016), 14–19

    Lum, K., and Isaac, W. To predict and serve? Significance 13, 5 (2016), 14–19

  85. [93]

    M., and Lee, S.-I.A unified approach to interpreting model predic- tions

    Lundberg, S. M., and Lee, S.-I.A unified approach to interpreting model predic- tions. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Red Hook, NY, USA, Dec. 2017), NIPS’17, Curran Associates Inc., pp. 4768–4777

  86. [94]

    Assessing the Fairness of AI Systems: AI Practitioners’ Processes, Challenges, and Needs for Support.Proc

    Madaio, M., Egede, L., Subramonyam, H., Wortman V aughan, J., and W al- lach, H. Assessing the Fairness of AI Systems: AI Practitioners’ Processes, Challenges, and Needs for Support.Proc. ACM Hum.-Comput. Interact. 6, CSCW1 (Apr. 2022), 52:1–52:26

  87. [95]

    Experiments in Automating Immigration Sys- tems, 1 ed

    Maxwell, J., and Tomlinson, J. Experiments in Automating Immigration Sys- tems, 1 ed. Bristol University Press, 2022

  88. [96]

    Visa applications: Home Office refuses to reveal ’high risk’ countries

    McDonald, H. Visa applications: Home Office refuses to reveal ’high risk’ countries. The Guardian (Jan. 2020)

  89. [97]

    S., Robertson, R

    Metaxa, D., Park, J. S., Robertson, R. E., Karahalios, K., Wilson, C., Han- cock, J., and Sandvig, C. Auditing Algorithms: Understanding Algorithmic Systems from the Outside In. Foundations and Trends® in Human–Computer Interaction 14, 4 (2021), 272–344

  90. [98]

    D., and Gebru, T

    Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., V asserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., and Gebru, T. Model Cards for Model Reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (Jan. 2019), pp. 220–229

  91. [99]

    Algorithmic Fairness: Choices, Assumptions, and Definitions

    Mitchell, S., Potash, E., Barocas, S., D’Amour, A., and Lum, K. Algorithmic Fairness: Choices, Assumptions, and Definitions. Annual Review of Statistics and Its Application 8, 1 (Mar. 2021), 141–163

  92. [100]

    Ethics-Based Auditing to Develop Trustworthy AI, Apr

    Mokander, J., and Floridi, L. Ethics-Based Auditing to Develop Trustworthy AI, Apr. 2021. CHI ’25, April 26-May 1, 2025, Yokohama, Japan Juliette Zaccour, Reuben Binns, and Luc Rocher

  93. [101]

    R., and Floridi, L.Auditing large language models: A three-layered approach

    Mökander, J., Schuett, J., Kirk, H. R., and Floridi, L.Auditing large language models: A three-layered approach. AI and Ethics (May 2023)

  94. [102]

    Morina, G., Oliinyk, V., W aton, J., Marusic, I., and Georgatzis, K.Auditing and Achieving Intersectional Fairness in Classification Problems, June 2020

  95. [103]

    New privacy-protected Facebook data for independent research on social media’s impact on democracy

    Nayak, C. New privacy-protected Facebook data for independent research on social media’s impact on democracy. https://research.facebook.com/blog/ 2020/2/new-privacy-protected-facebook-data-for-independent-research-on- social-medias-impact-on-democracy/, Feb. 2020

  96. [104]

    Data-invisible groups and data minimization in the deployment of AI solutions: Policy brief, 2023

    Neftenov, N., Stankovic, M., and Gupta, R. Data-invisible groups and data minimization in the deployment of AI solutions: Policy brief, 2023

  97. [105]

    Dissecting racial bias in an algorithm used to manage the health of populations

    Obermeyer, Z., Powers, B., Vogeli, C., and Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. Science 366, 6464 (Oct. 2019), 447–453

  98. [106]

    Ojewale, V., Steed, R., Vecchione, B., Birhane, A., and Raji, I. D. Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling, Mar. 2024

  99. [107]

    The Synthetic Data Vault

    Patki, N., Wedge, R., and Veeramachaneni, K. The Synthetic Data Vault. In 2016 IEEE International Conference on Data Science and Advanced Analytics (DSAA) (Oct. 2016), pp. 399–410

  100. [108]

    Assessment of differentially private synthetic data for utility and fairness in end-to-end machine learning pipelines for tabular data

    Pereira, M., Kshirsagar, M., Mukherjee, S., Dodhia, R., Lavista Ferres, J., and De Sousa, R. Assessment of differentially private synthetic data for utility and fairness in end-to-end machine learning pipelines for tabular data. PLOS ONE 19, 2 (Feb. 2024), e0297271

  101. [109]

    Poland, C. M. The Right Tool for the Job: Open-Source Auditing Tools in Machine Learning, June 2022

  102. [110]

    Census Bureau’s 2020 Census Data Products and Dissemination Team

    Population Reference Bureau, and U.S. Census Bureau’s 2020 Census Data Products and Dissemination Team. Why the Census Bureau Chose Differential Privacy, Mar. 2023

  103. [111]

    D., Xu, P., Honigsberg, C., and Ho, D

    Raji, I. D., Xu, P., Honigsberg, C., and Ho, D. E.Outsider Oversight: Designing a Third Party Audit Ecosystem for AI Governance, June 2022

  104. [112]

    K., Sahay, S., and Ahammad, P

    Rogers, R., Subramaniam, S., Peng, S., Durfee, D., Lee, S., Kancha, S. K., Sahay, S., and Ahammad, P. LinkedIn’s Audience Engagements API: A Privacy Preserving Data Analytics System at Scale, Nov. 2020

  105. [113]

    Towards the Right Kind of Fairness in AI, Sept

    Ruf, B., and Detyniecki, M. Towards the Right Kind of Fairness in AI, Sept. 2021

  106. [114]

    A Tool Bundle for AI Fairness in Practice

    Ruf, B., and Detyniecki, M. A Tool Bundle for AI Fairness in Practice. In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems (New York, NY, USA, Apr. 2022), CHI EA ’22, Association for Computing Machinery, pp. 1–3

  107. [115]

    Auditing Algo- rithms : Research Methods for Detecting Discrimination on Internet Platforms

    Sandvig, C., Hamilton, K., Karahalios, K., and Langbort, C. Auditing Algo- rithms : Research Methods for Detecting Discrimination on Internet Platforms. In Data and Discrimination: Converting Critical Concerns into Productive Inquiry (Seattle, WA, 2014)

  108. [116]

    Proceedings of the ACM on Human- Computer Interaction 6, GROUP (Jan

    Seidelin, C., Moreau, T., Shklovski, I., and Holten Møller, N.Auditing Risk Prediction of Long-Term Unemployment. Proceedings of the ACM on Human- Computer Interaction 6, GROUP (Jan. 2022), 1–12

  109. [117]

    Learning to Limit Data Collection via Scaling Laws: A Computational Interpretation for the Legal Principle of Data Minimization

    Shanmugam, D., Diaz, F., Shabanian, S., Finck, M., and Biega, A. Learning to Limit Data Collection via Scaling Laws: A Computational Interpretation for the Legal Principle of Data Minimization. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transpar...

  110. [118]

    Proceedings of the ACM on Human-Computer Interaction 5 , CSCW2 (Oct

    Shen, H., DeVos, A., Eslami, M., and Holstein, K.Everyday algorithm auditing: Understanding the power of everyday users in surfacing harmful algorithmic behaviors. Proceedings of the ACM on Human-Computer Interaction 5 , CSCW2 (Oct. 2021), 1–29

  111. [119]

    J., Kämpf, N

    Sivizaca Conde, D. J., Kämpf, N. L., Sass, D. R.-v., Schurig, T., and Kliewer, N. Privacy-Preserving Data Sharing: A Systematic Review and Future Research Areas. In ECIS 2024 Proceedings (June 2024)

  112. [120]

    Why the search for a privacy-preserving data sharing mechanism is failing

    Stadler, T., and Troncoso, C. Why the search for a privacy-preserving data sharing mechanism is failing. Nature Computational Science 2 , 4 (Apr. 2022), 208–210

  113. [121]

    S., and Acqisti, A

    Steed, R., Liu, T., Wu, Z. S., and Acqisti, A. Policy impacts of statistical uncertainty and privacy. Science 377, 6609 (Aug. 2022), 928–931

  114. [122]

    Bridging healthcare gaps: A scoping review on the role of artificial intelligence, deep learning, and large language models in alleviating problems in medical deserts

    Strika, Z., Petkovic, K., Likic, R., and Batenburg, R. Bridging healthcare gaps: A scoping review on the role of artificial intelligence, deep learning, and large language models in alleviating problems in medical deserts. Postgraduate Medical Journal (Sept. 2024), qgae122

  115. [123]

    Artificial Intelligence Risk Management Framework (AI RMF 1.0), Jan

    Tabassi, E. Artificial Intelligence Risk Management Framework (AI RMF 1.0), Jan. 2023

  116. [124]

    Distill-and-Compare: Auditing Black-Box Models Using Transparent Model Distillation

    Tan, S., Caruana, R., Hooker, G., and Lou, Y. Distill-and-Compare: Auditing Black-Box Models Using Transparent Model Distillation. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society (Dec. 2018), pp. 303–310

  117. [125]

    Exploring the impact of missingness on racial disparities in predictive performance of a machine learning model for emergency department triage

    Teeple, S., Smith, A., Toerper, M., Levin, S., Halpern, S., Badaki-Makun, O., and Hinson, J. Exploring the impact of missingness on racial disparities in predictive performance of a machine learning model for emergency department triage. JAMIA Open 6, 4 (Dec. 2023), ooad107

  118. [126]

    Executive Order on the Safe, Secure, and Trustworthy Devel- opment and Use of Artificial Intelligence

    The White House. Executive Order on the Safe, Secure, and Trustworthy Devel- opment and Use of Artificial Intelligence. https://www.whitehouse.gov/briefing- room/presidential-actions/2023/10/30/executive-order-on-the-safe-secure- and-trustworthy-development-and-use-of-artifici...

  119. [127]

    Trask, A., Bluemke, E., Collins, T., Drexler, B. G. E., Cuervas-Mons, C. G., Gabriel, I., Dafoe, A., and Isaac, W.Beyond Privacy Trade-offs with Structured Transparency, Mar. 2024

  120. [128]

    van Bekkum, M., and Borgesius, F. Z. Digital welfare fraud detection and the Dutch SyRI judgment. European Journal of Social Security 23 , 4 (Dec. 2021), 323–340

  121. [129]

    van Breugel, B., Qian, Z., and van der Schaar, M.Synthetic data, real errors: How (not) to publish and use synthetic data, May 2023

  122. [130]

    Fairer machine learning in the real world: Mitigating discrimination without collecting sensitive data

    Veale, M., and Binns, R. Fairer machine learning in the real world: Mitigating discrimination without collecting sensitive data. Big Data & Society (2017)

  123. [131]

    Building and Auditing Fair Algorithms: A Case Study in Candidate Screening

    Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., and Polli, F. Building and Auditing Fair Algorithms: A Case Study in Candidate Screening. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (Virtual Event Ca...

  124. [132]

    Towards a multi- stakeholder value-based assessment framework for algorithmic systems

    Yurrita, M., Murray-Rust, D., Balayn, A., and Bozzon, A. Towards a multi- stakeholder value-based assessment framework for algorithmic systems. In2022 ACM Conference on Fairness, Accountability, and Transparency (Seoul Republic of Korea, June 2022), ACM, pp. 535–563

  125. [133]

    M., Srivastava, D., and Xiao, X

    Zhang, J., Cormode, G., Procopiuc, C. M., Srivastava, D., and Xiao, X. PrivBayes: Private Data Release via Bayesian Networks. ACM Trans. Data- base Syst. 42, 4 (Oct. 2017), 25:1–25:41

  126. [134]

    Assessing Fairness in the Presence of Missing Data

    Zhang, Y., and Long, Q. Assessing Fairness in the Presence of Missing Data. Advances in neural information processing systems 34 (Dec. 2021), 16007–16019. A APPENDICES A.1 Summary tables The tables below include results for three different group parity metrics: A verage Odds D...

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.