Pith. sign in

REVIEW 1 cited by

Accounting for multiplicity in machine learning benchmark performance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.07272 v6 pith:ETNIKHN5 submitted 2023-03-10 stat.ME cs.LG

classification stat.MEcs.LG
keywords performancesampleclassifiersestimatepublicappliedbestchallenges
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

State-of-the-art (SOTA) performance refers to the highest performance achieved by some model on a test sample, preferably under controlled conditions such as public data (reproducibility) or public challenges (independent sample). Thousands of classifiers are applied, and the highest performance becomes the new reference point for a particular problem. In effect, this set-up is an estimate of the expected best performance among all classifiers applied to a random sample; a sample maximum estimate. In this paper, we argue that SOTA should instead be estimated by the expected performance of the best classifier, which can be done without knowing which classifier it is. Our contribution is the formal distinction between the two, and an investigation into the practical consequences of using the former to estimate the latter. This is done by presenting sample maximum estimator distributions for non-identical and dependent classifiers. We illustrate the impact on real world examples from public challenges.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Systemizing Multiplicity: The Curious Case of Arbitrariness in Machine Learning

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A systematic review of 80 papers on model multiplicity, with a new taxonomy of developer choices and a formal distinction between multiplicity, uncertainty, and variance.

Pith tools