REVIEW 4 cited by
Detecting and Correcting for Label Shift with Black Box Predictors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Faced with distribution shift between training and test set, we wish to detect and quantify the shift, and to correct our classifiers without test set labels. Motivated by medical diagnosis, where diseases (targets) cause symptoms (observations), we focus on label shift, where the label marginal $p(y)$ changes but the conditional $p(x| y)$ does not. We propose Black Box Shift Estimation (BBSE) to estimate the test distribution $p(y)$. BBSE exploits arbitrary black box predictors to reduce dimensionality prior to shift correction. While better predictors give tighter estimates, BBSE works even when predictors are biased, inaccurate, or uncalibrated, so long as their confusion matrices are invertible. We prove BBSE's consistency, bound its error, and introduce a statistical test that uses BBSE to detect shift. We also leverage BBSE to correct classifiers. Experiments demonstrate accurate estimates and improved prediction, even on high-dimensional datasets of natural images.
Forward citations
Cited by 4 Pith papers
-
Estimating prevalence with precision and accuracy
A new Bayesian quantifier, PQ, produces tighter and well-calibrated prediction intervals for class prevalence estimates, beating existing methods across simulated and real datasets.
-
Aligning Evaluation with Clinical Priorities: Calibration, Label Shift, and Error Costs
A new evaluation metric, the DCA log score, averages cost-weighted accuracy over a bounded, logit-uniform range of class prevalences, linking calibration, label shift, and error costs in one closed-form score.
-
Hidden-Domain Routing for All-Type Audio Deepfake Detection
A router-then-specialist audio deepfake detector, which classifies audio type first and then applies type-specific models and thresholds, achieved 96.10% Macro-F1 and first place on AT-ADD Track2.
-
Feature Engineering for Agents: An Adaptive Cognitive Architecture for Interpretable ML Monitoring
CAMA applies a three-step feature engineering procedure to LLM agents and reports 55 to 92 percent accuracy on ML monitoring report questions, outperforming six baselines.
Discussion (0). Continue with ORCID to comment.