Pith. sign in

REVIEW 3 cited by

Overcoming Common Flaws in the Evaluation of Selective Classification Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.01032 v2 pith:7OPN66RJ submitted 2024-07-01 cs.LG cs.CVstat.ME

classification cs.LGcs.CVstat.ME
keywords classificationsystemsmathrmselectiveaugrccurrentdataevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Selective Classification, wherein models can reject low-confidence predictions, promises reliable translation of machine-learning based classification systems to real-world scenarios such as clinical diagnostics. While current evaluation of these systems typically assumes fixed working points based on pre-defined rejection thresholds, methodological progress requires benchmarking the general performance of systems akin to the $\mathrm{AUROC}$ in standard classification. In this work, we define 5 requirements for multi-threshold metrics in selective classification regarding task alignment, interpretability, and flexibility, and show how current approaches fail to meet them. We propose the Area under the Generalized Risk Coverage curve ($\mathrm{AUGRC}$), which meets all requirements and can be directly interpreted as the average risk of undetected failures. We empirically demonstrate the relevance of $\mathrm{AUGRC}$ on a comprehensive benchmark spanning 6 data sets and 13 confidence scoring functions. We find that the proposed metric substantially changes metric rankings on 5 out of the 6 data sets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text

    cs.AI 2026-05 conditional novelty 6.0 of 10

    A 24-dataset benchmark for inducing schema graphs from raw text, plus an auditable LLM-based pipeline that reports the highest scores on the benchmark's four schema-similarity metrics.

  2. Scaling Truth: The Confidence Paradox in AI Fact-Checking

    cs.SI 2025-09 conditional novelty 6.0 of 10

    Across LLM fact-checking, model scale correlates with an inverse pattern of accuracy and decisiveness: smaller models are overconfident and less accurate, larger models are accurate but overly cautious.

  3. Calibrated Selective Prediction Using Deep Ensembles for ROI-Based Thyroid Nodule Ultrasound Classification Under Dataset Shift: A Retrospective Evaluation

    eess.IV 2026-07 conditional novelty 4.5 of 10

    Calibrated deep ensembles with MI-based selective triage achieve strong internal thyroid-nodule ROI performance but lose calibration and threshold transportability on external TN3K.

Pith tools