Pith. sign in

REVIEW 1 cited by

Multiple Testing Framework for Out-of-Distribution Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.09522 v5 pith:LJ5VQMGB submitted 2022-06-20 stat.ML cs.LG

classification stat.MLcs.LG
keywords detectionalgorithmdifferentlearningmultipleproposedtestswell
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study the problem of Out-of-Distribution (OOD) detection, that is, detecting whether a learning algorithm's output can be trusted at inference time. While a number of tests for OOD detection have been proposed in prior work, a formal framework for studying this problem is lacking. We propose a definition for the notion of OOD that includes both the input distribution and the learning algorithm, which provides insights for the construction of powerful tests for OOD detection. We propose a multiple hypothesis testing inspired procedure to systematically combine any number of different statistics from the learning algorithm using conformal p-values. We further provide strong guarantees on the probability of incorrectly classifying an in-distribution sample as OOD. In our experiments, we find that threshold-based tests proposed in prior work perform well in specific settings, but not uniformly well across different types of OOD instances. In contrast, our proposed method that combines multiple statistics performs uniformly well across different datasets and neural networks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Score Combining for Contrastive OOD Detection

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A GLRT-based score-combining rule slightly improves average AUROC and detection rate over CSI/SupCSI and classical p-value combination methods in dataset-vs-dataset and leave-one-class-out OOD experiments.

Pith tools