Pith. sign in

REVIEW 1 cited by

Deep Learning for Virtual Screening: Five Reasons to Use ROC Cost Functions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.07029 v1 pith:PG66N2LV submitted 2020-06-25 q-bio.BM cs.LGq-bio.QMstat.ML

classification q-bio.BMcs.LGq-bio.QMstat.ML
keywords druglearningcostdecisiondeepscreeningthresholdsapproaches
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Computer-aided drug discovery is an essential component of modern drug development. Therein, deep learning has become an important tool for rapid screening of billions of molecules in silico for potential hits containing desired chemical features. Despite its importance, substantial challenges persist in training these models, such as severe class imbalance, high decision thresholds, and lack of ground truth labels in some datasets. In this work we argue in favor of directly optimizing the receiver operating characteristic (ROC) in such cases, due to its robustness to class imbalance, its ability to compromise over different decision thresholds, certain freedom to influence the relative weights in this compromise, fidelity to typical benchmarking measures, and equivalence to positive/unlabeled learning. We also propose new training schemes (coherent mini-batch arrangement, and usage of out-of-batch samples) for cost functions based on the ROC, as well as a cost function based on the logAUC metric that facilitates early enrichment (i.e. improves performance at high decision thresholds, as often desired when synthesizing predicted hit compounds). We demonstrate that these approaches outperform standard deep learning approaches on a series of PubChem high-throughput screening datasets that represent realistic and diverse drug discovery campaigns on major drug target families.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 25 citations worldwide. Full citation record

  1. WelQrate: Defining the Gold Standard in Small Molecule Drug Discovery Benchmarking

    cs.LG 2024-11 conditional novelty 6.0 of 10

    WelQrate provides nine curated PubChem-derived screening datasets plus standardized splits, metrics, and formats, and benchmarks eight model families on them.

Pith tools