Pith. sign in

REVIEW 3 cited by

Sources of Uncertainty in Supervised Machine Learning -- A Statisticians' View

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.16703 v3 pith:6RMIRFLP submitted 2023-05-26 stat.ML cs.LG

classification stat.MLcs.LG
keywords uncertaintylearningmachinesourcesconceptssupervisedaleatoricepistemic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Supervised machine learning and predictive models have achieved an impressive standard today, enabling us to answer questions that were inconceivable a few years ago. Besides these successes, it becomes clear, that beyond pure prediction, which is the primary strength of most supervised machine learning algorithms, the quantification of uncertainty is relevant and necessary as well. However, before quantification is possible, types and sources of uncertainty need to be defined precisely. While first concepts and ideas in this direction have emerged in recent years, this paper adopts a conceptual, basic science perspective and examines possible sources of uncertainty. By adopting the viewpoint of a statistician, we discuss the concepts of aleatoric and epistemic uncertainty, which are more commonly associated with machine learning. The paper aims to formalize the two types of uncertainty and demonstrates that sources of uncertainty are miscellaneous and can not always be decomposed into aleatoric and epistemic. Drawing parallels between statistical concepts and uncertainty in machine learning, we emphasise the role of data and their influence on uncertainty.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 18 citations worldwide. Full citation record

  1. Probabilistic Deep Learning for Drought Forecasting: Role of Internal Climate Variability

    stat.AP 2026-08 conditional novelty 6.0 of 10

    An ensemble-informed lower bound for European drought forecasts, built from CRCM5 large-ensemble variability, is better calibrated than a reanalysis-only bound during 2020-2024, especially for extreme drought.

  2. Revisiting Active Learning under (Human) Label Variation

    cs.CL 2025-07 accept novelty 4.0 of 10

    A position paper that surveys and systematizes how active learning should change when human label variation is treated as a signal rather than noise.

  3. Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents

    cs.LG 2025-05 conditional novelty 4.0 of 10

    A position paper arguing that aleatoric/epistemic uncertainty splits fail for LLM agents and proposing underspecification, interaction, and output-based uncertainty research.

Pith tools