Pith. sign in

REVIEW 2 cited by

X Hacking: The Threat of Misguided AutoML

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.08513 v3 pith:UVW5XR22 submitted 2024-01-16 cs.LG cs.CR

classification cs.LGcs.CR
keywords x-hackingexplanationfeatureslearningmachinemethodsmetricsmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Explainable AI (XAI) and interpretable machine learning methods help to build trust in model predictions and derived insights, yet also present a perverse incentive for analysts to manipulate XAI metrics to support pre-specified conclusions. This paper introduces the concept of X-hacking, a form of p-hacking applied to XAI metrics such as SHAP values. We show how easily an automated machine learning pipeline can be adapted to exploit model multiplicity at scale: searching a Rashomon set of 'defensible' models with similar predictive performance to find a desired explanation. We formulate the trade-off between explanation and accuracy as a multi-objective optimisation problem, and illustrate empirically on familiar real-world datasets that, on average, Bayesian optimisation accelerates X-hacking 3-fold for features susceptible to it, versus random sampling. We show the vulnerability of a dataset to X-hacking can be determined by information redundancy among features. Finally, we suggest possible methods for detection and prevention, and discuss ethical implications for the credibility and reproducibility of XAI.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CleanSurvival: Automated data preprocessing for time-to-event models using reinforcement learning

    cs.LG 2025-02 reject novelty 6.0 of 10

    CleanSurvival uses Q-learning to pick imputation, outlier, and feature-selection steps for survival models, reporting C-index gains on Rotterdam and Flchain.

  2. Systemizing Multiplicity: The Curious Case of Arbitrariness in Machine Learning

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A systematic review of 80 papers on model multiplicity, with a new taxonomy of developer choices and a formal distinction between multiplicity, uncertainty, and variance.

Pith tools