REVIEW 3 cited by
The MCC-F1 curve: a performance evaluation technique for binary classification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Many fields use the ROC curve and the PR curve as standard evaluations of binary classification methods. Analysis of ROC and PR, however, often gives misleading and inflated performance evaluations, especially with an imbalanced ground truth. Here, we demonstrate the problems with ROC and PR analysis through simulations, and propose the MCC-F1 curve to address these drawbacks. The MCC-F1 curve combines two informative single-threshold metrics, MCC and the F1 score. The MCC-F1 curve more clearly differentiates good and bad classifiers, even with imbalanced ground truths. We also introduce the MCC-F1 metric, which provides a single value that integrates many aspects of classifier performance across the whole range of classification thresholds. Finally, we provide an R package that plots MCC-F1 curves and calculates related metrics.
Forward citations
Cited by 3 Pith papers
-
Estimating the local star formation rate density from ASKAP RACS
A machine-learning-selected sample of 11,293 ASKAP radio galaxies yields a completeness-corrected local star formation rate density of (1.4 +/- 0.5) x 10^-2 solar masses per year per cubic megaparsec, consistent with ...
-
Interactive Classification Metrics: A graphical application to build robust intuition for classification model evaluation
A free Python app lets users manipulate simulated class distributions and the threshold to see how classification metrics such as ROC AUC, MCC, and F1 change together.
-
MH-FSF: A Unified Framework for Overcoming Benchmarking and Reproducibility Limitations in Feature Selection Evaluation
A unified, publicly released benchmarking framework evaluates 17 feature selection methods across 10 Android malware datasets; LASSO, RFE, and SigAPI come out most consistent, while PCA, ReliefF, and SigPID lag.
Discussion (0). Continue with ORCID to comment.