Pith. sign in

REVIEW 3 cited by

How to avoid machine learning pitfalls: a guide for academic researchers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.02497 v5 pith:FOW5QRNB submitted 2021-08-05 cs.LG

classification cs.LG
keywords learningmachinemodelsacademicavoidguidemistakeswhat
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Mistakes in machine learning practice are commonplace, and can result in a loss of confidence in the findings and products of machine learning. This guide outlines common mistakes that occur when using machine learning, and what can be done to avoid them. Whilst it should be accessible to anyone with a basic understanding of machine learning techniques, it focuses on issues that are of particular concern within academic research, such as the need to do rigorous comparisons and reach valid conclusions. It covers five stages of the machine learning process: what to do before model building, how to reliably build models, how to robustly evaluate models, how to compare models fairly, and how to report results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Preserving Clusters in Prompt Learning for Unsupervised Domain Adaptation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A prompt-learning method for UDA that uses source-prompt predictions to build better pseudo-labels and a Wasserstein clustering term to keep target text prompts aligned with visual embeddings.

  2. Fairness in Federated Learning: Fairness for Whom?

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A critical review of 121 federated learning fairness papers identifies five recurring pitfalls and proposes a harm-centered, lifecycle-based framework.

  3. Importance of User Control in Data-Centric Steering for Healthcare Experts

    cs.HC 2025-05 conditional novelty 5.0 of 10

    Healthcare experts who manually adjusted training data improved a diabetes prediction model more than those using automated corrections, without losing trust or understanding.

Pith tools