REVIEW 3 major objections 6 minor 58 references
WelQrate: Defining the Gold Standard in Small Molecule Drug Discovery Benchmarking
T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read WelQrate makes the case that curated screening data, not just model architecture, determines whether drug-discovery benchmarks are trustworthy.
desk verdict Useful, carefully documented benchmark with an overclaimed 'gold standard' label and an unquantified inactive-label noise problem. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the hierarchical curation pipeline. It organizes bioassays by level: a primary screen with a deliberately loose threshold, confirmatory screens that re-test putative actives, and counter screens that reject compounds with off-target or nonspecific activity; final actives are validated hits and final inactives come from primary-screen inactivity, with noted exceptions kept when follow-up readouts contradict. This pipeline is what converts raw screening data into labels the paper claims are clean and realistic. Around it, the framework adds standardized formats (isomeric SMILES and InChI, plus precomputed 2D and 3D graphs), early-enrichment metrics such as logAUC, BEDROC, EF100, and DCG100, and split protocols including nested-style random cross-validation and Bemis-Murcko scaffold splits. The central working assumption is that label quality, not model architecture alone, determines whether a benchmark ranking transfers to real screening.
What would settle it
Re-test a random sample of the final inactive compounds from one WelQrate dataset, say AID1798, in the confirmatory assay used there (AID1488) under dose-response conditions; if a substantial fraction, for example several percent, reproducibly show activity, the inactive labels are too optimistic and the clean-label premise fails. A cheaper version is to search the public bioassay records for compounds labeled inactive in WelQrate that later returned active readouts in related follow-up assays and count how often that happens.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that a benchmark built from hierarchically curated high-throughput screening data, rather than raw primary-screen readouts, can serve as a gold standard for small-molecule virtual screening. The curation works by tracing each compound through primary, confirmatory, and counter assays, keeping only actives that survive validation and only inactives from primary screens that were not contradicted by later readouts, then applying promiscuity, PAINS, druglikeness, and representation filters. Around this collection, the paper standardizes featurization, 3D conformation generation, evaluation metrics that reward early enrichment, and split schemes including scaffold splits. Its benchmarking experiments find that models improve with complexity under random splits, that a domain-expert descriptor with a simple classifier outperforms the neural models, that training on uncurated primary-screen data changes or reverses some comparisons, that predefined features beat one-hot features, and that all models struggle under scaffold splits. The paper takes these results as evidence that both data quality and evaluation design must be reported and standardized for meaningful model comparison.
Load-bearing premise
The benchmark's clean-label claim rests on the assumption that compounds labeled inactive, mostly from primary screens without confirmatory or counter-screen validation, really are inactive; if primary-screen misses are common, the inactive half of every dataset is noisier than the curation suggests.
Editorial extensions
If this is right
- Model rankings on WelQrate become a more trustworthy basis for choosing virtual-screening methods, because labels have passed confirmation and counter-screening.
- The strong performance of a simple model on domain-expert descriptors implies that architecture comparisons should include such a baseline and that featurization matters as much as model design.
- Training on uncurated primary-screen data can invert or erase performance differences, so benchmark results that ignore curation are hard to interpret.
- Scaffold splits expose distribution shift: all tested models lose accuracy, meaning claims of generalization need to be evaluated under scaffold rather than random splits.
- Adopting standardized metrics that reward early enrichment aligns benchmark scores with the real workflow of buying or synthesizing only the top-ranked compounds.
Reading between the lines
- Beyond the paper, the curation's weak point is the inactive set: most inactives are primary-screen misses never validated in confirmatory screens, so if primary-screen false negatives are frequent, the clean-label premise is weakened even though actives are well validated.
- A testable extension is to re-run the benchmark after applying a stricter inactive definition, for example requiring inactivity in a confirmatory or counter screen, and measure how rankings shift.
- The additional dose-response measurements available for three datasets invite a regression benchmark for potency prediction, which the paper mentions but does not develop.
- If the gold-standard framing is adopted widely, the field's next problem becomes split and feature variation across benchmarks; WelQrate's fixed protocols make cross-paper comparisons possible only if researchers report the version and split used.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces WelQrate, a benchmark suite for small-molecule drug discovery consisting of nine PubChem-derived datasets with a hierarchical curation pipeline, standardized data formats (SMILES, InChI, SDF, 2D/3D graphs), a proposed evaluation protocol (metrics, random and scaffold splits, adapted cross-validation), and extensive benchmarking of ten models across four research questions. The central claims are that the datasets have clean and reliable labels due to expert-designed curation, and that the framework provides a standardized, realistic basis for virtual-screening model comparison. The paper also argues that dataset quality, featurization, and split strategy materially affect model rankings, and recommends WelQrate as a new gold standard.
Significance. If the curation and evaluation claims hold, WelQrate would be a valuable community resource: the authors provide public data, curation code, and experimental scripts; the datasets are large and realistically imbalanced; the supplementary material documents the full curation hierarchy for each AID; and the benchmarking includes standard and domain-baseline models with error bars and hyperparameter tables. The inclusion of multiple realistic ranking metrics (logAUC, BEDROC, EF100, DCG100) and scaffold-split evaluation is a genuine strength. However, the significance is contingent on two load-bearing points: the reliability of the inactive labels, which are mostly unconfirmed primary-screen negatives, and the validity of the adapted cross-validation protocol. Because the paper's headline contribution is 'clean and reliable data labels,' the inactive-label issue directly affects the benchmark's core value proposition and must be addressed before the gold-standard claim is supportable.
major comments (3)
- [Sec. 3.2; Supp. A.2] The central claim of 'clean and reliable data labels' (Sec. 3) is not established for the majority class. For most datasets, the final inactive set is taken directly from primary-screen inactive readouts (e.g., AID1798, AID435034, AID1843, AID2258, AID2689, AID485290; see Supp. A.2), and the paper itself notes that primary HTS thresholds are deliberately loose to reduce false negatives (Sec. 3.2). The curation notes also document primary-inactive compounds later found active in confirmatory screens: AID435008 states that 'we found some inactive compounds in one screen that were active in the other'; AID2258 describes six primary-inactive compounds that were retested, with final active readouts causing their exclusion; AID488997 reports 17, 38, and 2 primary-inactive compounds tested in confirmatory screens, some of which were active. No dataset-level false-negative rate is reported anywhere, and Section 6 (Limitations) does not mention inactive-label noise. Since inactives constitute the vast majority of each dataset, unquantified false negatives can distort enrichment and ranking metrics, so the clean-label premise needs either explicit quantification using the available retest data or a substantially softened claim and a corresponding limitation statement.
- [Sec. 5.2; Fig. 4] The RQ2 conclusion that the results 'align with data-centric AI, highlighting the importance of dataset quality' is not supported for all model families: the 2D and 3D graph-based models trained on the less clean control data outperform models trained on WelQrate in logAUC[0.001,0.1] and BEDROC, while performing worse on EF100 and DCG100. The authors offer only an untested hypothesis about the 'range of top selected candidates.' As Fig. 4 averages across datasets and the effect is metric-dependent, the conclusion should either be restricted to the architectures and metrics where curation consistently helps (Naive, sequence-based, and Domain baselines) or be backed by a per-dataset, per-metric analysis that explains the reversal. As written, RQ2 does not provide a coherent demonstration that dataset quality improves model evaluation across the board.
- [Sec. 4.3; Supp. B.2] The adapted cross-validation protocol, in which the validation fold is fixed as the fold immediately preceding the test fold, is asserted to 'enhance computational efficiency without compromising robustness,' but no evidence is provided that this protocol yields estimates comparable to nested cross-validation or to standard k-fold cross-validation with proper hyperparameter tuning. Since all random-split results (RQ1-RQ3) rely on this protocol, the evaluation framework's reliability claim depends on this untested assumption. A small-scale comparison of the adapted protocol against nested cross-validation on one or two datasets would be sufficient to support the claim, or the assertion should be replaced with a more cautious statement.
minor comments (6)
- [Sec. 1] There is a typo in the bullet list: 'theurapeutic' should be 'therapeutic'; the author affiliation line also contains a stray 'Electical and Computer Engineering Dept„' with a formatting artifact.
- [Sec. 4.3] The text says 'a 3:1:1 training:validation ratio' but appears to mean a 3:1:1 train:validation:test ratio; please clarify the wording to avoid ambiguity about whether the test set is included.
- [Table 1] In the AID1843 row, the number of unique BM scaffolds is listed as '82,140C' with a stray 'C' suffix; this appears to be a typo.
- [Supp. B.1] The displayed BEDROC formula after 'calculated as:' appears garbled: the summation term is written as 'P n i=1 -eri/N', which omits the exponential and the alpha parameter shown in the RIE definition above it. Please correct the equation.
- [Fig. 5; Supp. C.4] The main text states that 'other metrics exhibit the same trend' for the one-hot versus predefined-feature comparison, but Supp. C.4 says the predefined features outperform 'in the majority of cases,' which is a weaker statement. Please align the main-text claim with the supplementary results.
- [Sec. 4.2; Supp. A.5] The choice of 1000 µM as a placeholder for inactive compounds in the three datasets with additional measurements is acknowledged in Supp. A.5, but the main text does not mention this artificial value or its potential effect on regression tasks. A one-sentence caveat in Sec. 4.1 or 4.2 would help readers who use the floating-value labels.
Circularity Check
No significant circularity: the curated dataset and benchmarking protocol are self-contained, and the cited same-group works are not load-bearing in any derivation.
full rationale
The paper's central contribution is a curated dataset and a standardized evaluation protocol. The curation pipeline is described algorithmically (duplicate removal, hierarchical primary/confirmatory/counter screening, PAINS and druglikeness filters, expert verification) and implemented in publicly available code, so the 'high-quality labels' claim is supported by an independent procedure rather than by definition. The benchmarking section compares existing models under a standard train/validation/test protocol; hyperparameters are tuned on validation folds and evaluated on held-out test sets, so no fitted parameter is renamed as a prediction. Self-citations are present but not load-bearing: same-group works [9,13] are cited for the general fact that primary HTS thresholds are loose and for bioassay identification, while the specific curation hierarchies, filters, and dataset statistics are produced by this paper's own pipeline; BCL [39] is used only as a comparison baseline. The 3D graph 6 Angstrom cutoff, logAUC range, and BEDROC alpha are acknowledged heuristics adopted from prior work, not results derived from the dataset. The paper explicitly acknowledges limitations such as the artificial 1000 µM inactive value (Supplement A.5), and Section 6 lists remaining limitations (imbalance, scaffold shift, conformations); these are data-quality caveats, not circular reasoning. The skeptic's concern that most inactives come from unconfirmed primary-screen negatives is a plausible correctness risk for the 'clean labels' premise, but it is an empirical claim about assay noise, not a tautology or a fit; it does not make the derivation circular. The paper's own retest examples in Supplement A.2 are treated as curation decisions rather than hidden circular constraints, and no claim in the paper reduces by construction to its inputs.
Assumptions & free parameters
free parameters (5)
- Inactive placeholder IC50/EC50 =
1000 µM (assigned)
- BEDROC alpha =
20
- logAUC FPR range =
[0.001, 0.1]
- 3D graph distance cutoff =
6 Å
- BM scaffold training assignment threshold =
>10% of dataset in a scaffold bin
assumptions (5)
- domain assumption PubChem bioassay records and their textual descriptions accurately encode the relationships among primary, confirmatory, and counter screens.
- ad hoc to paper Primary-screen inactivity is treated as true inactivity in most final inactive sets.
- domain assumption PAINS, promiscuity, and druglikeness filters remove artifacts without removing a meaningful number of real actives.
- domain assumption A low-energy 3D conformation generated by Corina is a sufficient structural representation for benchmarking 3D models.
- ad hoc to paper Adapted cross-validation with the validation fold fixed as the fold preceding the test fold gives estimates comparable to nested cross-validation.
Cite this review
Pith. "Pith review of WelQrate: Defining the Gold Standard in Small Molecule Drug Discovery Benchmarking." pith.science (2026). https://pith.science/paper/WWW42MF5
@misc{pith2026241109820,
author = {Pith},
title = {Pith review of: WelQrate: Defining the Gold Standard in Small Molecule Drug Discovery Benchmarking},
year = {2026},
howpublished = {\url{https://pith.science/paper/WWW42MF5}},
note = {Machine review of arXiv:2411.09820}
}
read the original abstract
While deep learning has revolutionized computer-aided drug discovery, the AI community has predominantly focused on model innovation and placed less emphasis on establishing best benchmarking practices. We posit that without a sound model evaluation framework, the AI community's efforts cannot reach their full potential, thereby slowing the progress and transfer of innovation into real-world drug discovery. Thus, in this paper, we seek to establish a new gold standard for small molecule drug discovery benchmarking, WelQrate. Specifically, our contributions are threefold: WelQrate Dataset Collection - we introduce a meticulously curated collection of 9 datasets spanning 5 therapeutic target classes. Our hierarchical curation pipelines, designed by drug discovery experts, go beyond the primary high-throughput screen by leveraging additional confirmatory and counter screens along with rigorous domain-driven preprocessing, such as Pan-Assay Interference Compounds (PAINS) filtering, to ensure the high-quality data in the datasets; WelQrate Evaluation Framework - we propose a standardized model evaluation framework considering high-quality datasets, featurization, 3D conformation generation, evaluation metrics, and data splits, which provides a reliable benchmarking for drug discovery experts conducting real-world virtual screening; Benchmarking - we evaluate model performance through various research questions using the WelQrate dataset collection, exploring the effects of different models, dataset quality, featurization methods, and data splitting strategies on the results. In summary, we recommend adopting our proposed WelQrate as the gold standard in small molecule drug discovery benchmarking. The WelQrate dataset collection, along with the curation codes, and experimental scripts are all publicly available at WelQrate.org.
Figures
Figures from the paper (16 more)
Reference graph
Works this paper leans on
-
[1]
MEDDRA. https://www.meddra.org/. [Online; accessed 05-June-2024]
work page 2024
-
[2]
Md G Abbas, Hirotaka Shoji, Shingo Soya, Mari Hondo, Tsuyoshi Miyakawa, and Takeshi Sakurai. Comprehensive behavioral analysis of male ox1r-/- mice showed implication of orexin receptor-1 in mood, anxiety, and social behavior. Frontiers in behavioral neuroscience, 9:324, 2015
work page 2015
-
[3]
Smitha Antony, Christophe Marchand, Andrew G Stephen, Laurent Thibaut, Keli K Agama, Robert J Fisher, and Yves Pommier. Novel high-throughput electrochemiluminescent assay for identification of human tyrosyl-dna phosphodiesterase (tdp1) inhibitors and characterization of furamidine (nsc 305831) as an inhibitor of tdp1. Nucleic acids research, 35(13):4474–...
work page 2007
-
[4]
Baskar Arumugam and Neville A McBrien. Muscarinic antagonist control of myopia: evidence for m4 and m1 receptor-based pathways in the inhibition of experimentally-induced axial myopia in the tree shrew. Investigative Ophthalmology & Visual Science, 53(9):5827–5837, 2012
work page 2012
-
[5]
Jonathan B Baell and Georgina A Holloway. New substructure filters for removal of pan assay interference compounds (pains) from screening libraries and for their exclusion in bioassays. Journal of medicinal chemistry, 53(7):2719–2740, 2010
work page 2010
-
[6]
NC Bodick, WW Offen, HE Shannon, J Satterwhite, R Lucas, R Van Lier, and SM Paul. The selective muscarinic agonist xanomeline improves both the cognitive deficits and behavioral symptoms of alzheimer disease. Alzheimer disease and associated disorders, 11:S16–22, 1997
work page 1997
-
[7]
Chemical modulation of kv7 potassium channels
Matteo Borgini, Pravat Mondal, Ruiting Liu, and Peter Wipf. Chemical modulation of kv7 potassium channels. RSC medicinal chemistry, 12(4):483–537, 2021
work page 2021
-
[8]
Bastienne Brauksiepe, Alejandro O Mujica, Harald Herrmann, and Erwin R Schmidt. The serine/threonine kinase stk33 exhibits autophosphorylation and phosphorylates the intermediate filament protein vimentin. BMC biochemistry, 9:1–12, 2008
work page 2008
Show all 58 references
-
[9]
Molecular structures: perception, autocorrelation descriptor and sar studies
P Broto, G Moreau, and C Vandycke. Molecular structures: perception, autocorrelation descriptor and sar studies. perception of molecules: topological structure and 3-dimensional structure. European journal of medicinal chemistry, 19(1):61–65, 1984
1984
-
[10]
Muscarinic m1 receptor-stimulated adenylate cyclase activity in chinese hamster ovary cells is mediated by gs α and is not a consequence of phosphoinositidase c activation
Neil T BURFORD and Stefan R NAHORSKI. Muscarinic m1 receptor-stimulated adenylate cyclase activity in chinese hamster ovary cells is mediated by gs α and is not a consequence of phosphoinositidase c activation. Biochemical Journal, 315(3):883–888, 1996
1996
-
[11]
High-throughput screening assay datasets from the pubchem database
Mariusz Butkiewicz, Yanli Wang, Stephen H Bryant, Edward W Lowe Jr, David C Weaver, and Jens Meiler. High-throughput screening assay datasets from the pubchem database. Chemical informatics (Wilmington, Del.), 3(1), 2017
2017
-
[12]
Targeting t-type/cav3
Song Cai, Kimberly Gomez, Aubin Moutal, and Rajesh Khanna. Targeting t-type/cav3. 2 channels for chronic pain. Translational Research, 234:20–30, 2021
2021
-
[13]
3d structure generator corina classic
Structure Generator CORINA Classic. 3d structure generator corina classic. Nürnberg: Molecular Networks GmbH. Available online at: www. mn-am. com, 2019
2019
-
[14]
Biopython: freely available python tools for computational molecular biology and bioinformatics
Peter JA Cock, Tiago Antao, Jeffrey T Chang, Brad A Chapman, Cymon J Cox, Andrew Dalke, Iddo Friedberg, Thomas Hamelryck, Frank Kauff, Bartek Wilczynski, et al. Biopython: freely available python tools for computational molecular biology and bioinformatics. Bioinformatics, 25(...
2009
-
[15]
Unique kir2
Amit S Dhamoon, Sandeep V Pandit, Farzad Sarmast, Keely R Parisian, Prabal Guha, You Li, Suveer Bagwe, Steven M Taffet, and Justus MB Anumonwo. Unique kir2. x properties determine regional and species differences in the cardiac inward rectifier k+ current. Circulation research...
2004
-
[16]
Blockade of orexin-1 receptors atten- uates orexin-2 receptor antagonism-induced sleep promotion in the rat
Christine Dugovic, Jonathan E Shelton, Leah E Aluisio, Ian C Fraser, Xiaohui Jiang, Steven W Sutton, Pascal Bonaventure, Sujin Yun, Xiaorong Li, Brian Lord, et al. Blockade of orexin-1 receptors atten- uates orexin-2 receptor antagonism-induced sleep promotion in the rat. Jour...
2009
-
[17]
Pytorch geometric, 2019
Matthias Fey and Jan Eric Lenssen. Pytorch geometric, 2019. URL https://github.com/rusty1s/ pytorch_geometric. Accessed: 2024-06-11
2019
-
[18]
Deep learning for virtual screening: Five reasons to use roc cost functions
Vladimir Golkov, Alexander Becker, Daniel T Plop, Daniel ˇCuturilo, Neda Davoudi, Jeffrey Mendenhall, Rocco Moretti, Jens Meiler, and Daniel Cremers. Deep learning for virtual screening: Five reasons to use roc cost functions. arXiv preprint arXiv:2007.07029, 2020. 40
2007 arXiv
-
[19]
International union of pharmacology
George A Gutman, K George Chandy, John P Adelman, Jayashree Aiyar, Douglas A Bayliss, David E Clapham, Manuel Covarriubias, Gary V Desir, Kiyoshi Furuichi, Barry Ganetzky, et al. International union of pharmacology. xli. compendium of voltage-gated ion channels: potassium chan...
2003
-
[20]
Alzheimer’s disease: targeting the cholinergic system
Talita H Ferreira-Vieira, Isabella M Guimaraes, Flavia R Silva, and Fabiola M Ribeiro. Alzheimer’s disease: targeting the cholinergic system. Current neuropharmacology, 14(1):101–115, 2016
2016
-
[21]
Orexin a activates locus coeruleus cell firing and increases arousal in the rat
Jim J Hagan, Ron A Leslie, Sara Patel, Martyn L Evans, Trevor A Wattam, Steve Holmes, Christopher D Benham, Stephen G Taylor, Carol Routledge, Panida Hemmati, et al. Orexin a activates locus coeruleus cell firing and increases arousal in the rat. Proceedings of the National Ac...
1999
-
[22]
Glide: a new approach for rapid, accurate docking and scoring
Thomas A Halgren, Robert B Murphy, Richard A Friesner, Hege S Beard, Leah L Frye, W Thomas Pollard, and Jay L Banks. Glide: a new approach for rapid, accurate docking and scoring. 2. enrichment factors in database screening. Journal of medicinal chemistry, 47(7):1750–1759, 2004
2004
-
[23]
Genetic variation of cacna1h in idiopathic generalized epilepsy.Annals of neurology, 55(4):595–596, 2004
Sarah E Heron, Hilary A Phillips, John C Mulley, Aziz Mazarib, Miriam Y Neufeld, Samuel F Berkovic, and Ingrid E Scheffer. Genetic variation of cacna1h in idiopathic generalized epilepsy.Annals of neurology, 55(4):595–596, 2004
2004
-
[24]
Therapeutics data commons: Machine learning datasets and tasks for drug discovery and development
Kexin Huang, Tianfan Fu, Wenhao Gao, Yue Zhao, Yusuf Roohani, Jure Leskovec, Connor W Coley, Cao Xiao, Jimeng Sun, and Marinka Zitnik. Therapeutics data commons: Machine learning datasets and tasks for drug discovery and development. arXiv preprint arXiv:2102.09548, 2021
2021 arXiv
-
[25]
Familial kcnq2 mutation: a psychiatric perspective
Anton Iftimovici, Angeline Charmet, Béatrice Desnous, Ana Ory, Richard Delorme, Charles Coutton, Françoise Devillard, Mathieu Milh, and Anna Maruani. Familial kcnq2 mutation: a psychiatric perspective. Psychiatric Genetics, 34(1):24–27, 2024
2024
-
[26]
Regulation of neuronal t-type calcium channels
Mircea C Iftinca and Gerald W Zamponi. Regulation of neuronal t-type calcium channels. Trends in pharmacological sciences, 30(1):32–40, 2009
2009
-
[27]
Ir evaluation methods for retrieving highly relevant documents
Kalervo Järvelin and Jaana Kekäläinen. Ir evaluation methods for retrieving highly relevant documents. In ACM SIGIR Forum, volume 51, pages 243–250. ACM New York, NY , USA, 2017
2017
-
[28]
Identification and characterization of the rat m1 muscarinic receptor promoter
Christoph PR Klett and Tom I Bonner. Identification and characterization of the rat m1 muscarinic receptor promoter. Journal of neurochemistry, 72(3):900–909, 1999
1999
-
[29]
The sider database of drugs and side effects
Michael Kuhn, Ivica Letunic, Lars Juhl Jensen, and Peer Bork. The sider database of drugs and side effects. Nucleic acids research, 44(D1):D1075–D1079, 2016
2016
-
[30]
Rdkit documentation
Greg Landrum. Rdkit documentation. Release, 1(1-79):4, 2013
2013
-
[31]
The embl-ebi bioinformatics web and programmatic tools framework
Weizhong Li, Andrew Cowley, Mahmut Uludag, Tamer Gur, Hamish McWilliam, Silvano Squizzato, Young Mi Park, Nicola Buso, and Rodrigo Lopez. The embl-ebi bioinformatics web and programmatic tools framework. Nucleic acids research, 43(W1):W580–W584, 2015
2015
-
[32]
Inhibition of human tyrosyl-dna phosphodiesterase by aminoglycoside antibiotics and ribosome inhibitors
Zhiyong Liao, Laurent Thibaut, Andrew Jobson, and Yves Pommier. Inhibition of human tyrosyl-dna phosphodiesterase by aminoglycoside antibiotics and ribosome inhibitors. Molecular pharmacology, 70 (1):366–372, 2006
2006
-
[33]
Spherical message passing for 3d graph networks
Yi Liu, Limei Wang, Meng Liu, Xuan Zhang, Bora Oztekin, and Shuiwang Ji. Spherical message passing for 3d graph networks. arXiv preprint arXiv:2102.05013, 2021
2021 arXiv
-
[34]
Interpretable chirality-aware graph neural network for quantitative structure activity relationship modeling in drug discovery
Yunchao Lance Liu, Yu Wang, Oanh Vu, Rocco Moretti, Bobby Bodenheimer, Jens Meiler, and Tyler Derr. Interpretable chirality-aware graph neural network for quantitative structure activity relationship modeling in drug discovery. In Proceedings of the AAAI Conference on Artifici...
2023
-
[35]
Effects of central muscarinic-1 receptor stimulation on blood pressure regulation
Aharon Medina, Neil Bodick, Ary L Goldberger, Margaret Mac Mahon, and Lewis A Lipsitz. Effects of central muscarinic-1 receptor stimulation on blood pressure regulation. Hypertension, 29(3):828–834, 1997
1997
-
[36]
John J Millichap and Edward C Cooper. Kcnq2 potassium channel epileptic encephalopathy syndrome: divorce of an electro-mechanical couple? kcnq2 potassium channel epileptic encephalopathy syndrome: divorce of an electro-mechanical couple? Epilepsy Currents, 12(4):150–152, 2012
2012
-
[37]
The autocorrelation of a topological structure: A new molecular descriptor
Gilles Moreau and Pierre Broto. The autocorrelation of a topological structure: A new molecular descriptor. 1980. 41
1980
-
[38]
Rapid context-dependent ligand desolvation in molecular docking
Michael M Mysinger and Brian K Shoichet. Rapid context-dependent ligand desolvation in molecular docking. Journal of chemical information and modeling, 50(9):1561–1573, 2010
2010
-
[39]
The role of t-type calcium channels in epilepsy and pain
MT Nelson, SM Todorovic, and E Perez-Reyes. The role of t-type calcium channels in epilepsy and pain. Current pharmaceutical design, 12(18):2189–2197, 2006
2006
-
[40]
High-affinity choline transporter
Takashi Okuda and Tatsuya Haga. High-affinity choline transporter. Neurochemical research, 28:483–488, 2003
2003
-
[41]
Improved scoring of ligand- protein interactions using owfeg free energy grids
David A Pearlman and Paul S Charifson. Improved scoring of ligand- protein interactions using owfeg free energy grids. Journal of medicinal chemistry, 44(4):502–511, 2001
2001
-
[42]
Molecular physiology of low-voltage-activated t-type calcium channels
Edward Perez-Reyes. Molecular physiology of low-voltage-activated t-type calcium channels. Physiologi- cal reviews, 83(1):117–161, 2003
2003
-
[43]
Acetylcholine as a neuromodulator: cholinergic signaling shapes nervous system function and behavior
Marina R Picciotto, Michael J Higley, and Yann S Mineur. Acetylcholine as a neuromodulator: cholinergic signaling shapes nervous system function and behavior. Neuron, 76(1):116–129, 2012
2012
-
[44]
Massively multitask networks for drug discovery
Bharath Ramsundar, Steven Kearnes, Patrick Riley, Dale Webster, David Konerding, and Vijay Pande. Massively multitask networks for drug discovery. arXiv preprint arXiv:1502.02072, 2015
2015 arXiv
-
[45]
Maximum unbiased validation (muv) data sets for virtual screening based on pubchem bioactivity data
Sebastian G Rohrer and Knut Baumann. Maximum unbiased validation (muv) data sets for virtual screening based on pubchem bioactivity data. Journal of chemical information and modeling, 49(2):169–184, 2009
2009
-
[46]
Synthetic lethal interaction between oncogenic kras dependency and stk33 suppression in human cancer cells
Claudia Scholl, Stefan Fröhling, Ian F Dunn, Anna C Schinzel, David A Barbie, So Young Kim, Serena J Silver, Pablo Tamayo, Raymond C Wadlow, Sridhar Ramaswamy, et al. Synthetic lethal interaction between oncogenic kras dependency and stk33 suppression in human cancer cells. Ce...
2009
-
[47]
Fast, scalable generation of high-quality protein multiple sequence alignments using clustal omega
Fabian Sievers, Andreas Wilm, David Dineen, Toby J Gibson, Kevin Karplus, Weizhong Li, Rodrigo Lopez, Hamish McWilliam, Michael Remmert, Johannes Söding, et al. Fast, scalable generation of high-quality protein multiple sequence alignments using clustal omega. Molecular system...
2011
-
[48]
Fluorescence spectroscopic profiling of compound libraries
Anton Simeonov, Ajit Jadhav, Craig J Thomas, Yuhong Wang, Ruili Huang, Noel T Southall, Paul Shinn, Jeremy Smith, Christopher P Austin, Douglas S Auld, et al. Fluorescence spectroscopic profiling of compound libraries. Journal of medicinal chemistry, 51(8):2363–2371, 2008
2008
-
[49]
Orexin/hypocretin signaling at the orexin 1 receptor regulates cue-elicited cocaine-seeking
Rachel J Smith, Ronald E See, and Gary Aston-Jones. Orexin/hypocretin signaling at the orexin 1 receptor regulates cue-elicited cocaine-seeking. European Journal of Neuroscience, 30(3):493–503, 2009
2009
-
[50]
Chronic inhibition of cardiac kir2
Haiyan Sun, Xiaodong Liu, Qiaojie Xiong, Sojin Shikano, and Min Li. Chronic inhibition of cardiac kir2. 1 and herg potassium channels by celastrol with dual effects on both ion conductivity and protein trafficking. Journal of biological chemistry, 281(9):5877–5884, 2006
2006
-
[51]
We Need Better Benchmarks for Machine Learning in Drug Discovery
Pat Walters. We Need Better Benchmarks for Machine Learning in Drug Discovery. https://practicalcheminformatics.blogspot.com/2023/08/ we-need-better-benchmarks-for-machine.html?m=1 , 2023. [Online; accessed 05-June-2024]
2023
-
[52]
Deep graph library (dgl), 2019
Minjie Wang, Jinjing Zhou, Quan Gan, and Zheng Zhang. Deep graph library (dgl), 2019. URL https://www.dgl.ai. Accessed: 2024-06-11
2019
-
[53]
Database resources of the national center for biotechnology information
David L Wheeler, Tanya Barrett, Dennis A Benson, Stephen H Bryant, Kathi Canese, Vyacheslav Chetvernin, Deanna M Church, Michael DiCuccio, Ron Edgar, Scott Federhen, et al. Database resources of the national center for biotechnology information. Nucleic acids research, 36(supp...
2007
-
[54]
Alzheimer disease: evidence for selective loss of cholinergic neurons in the nucleus basalis
Peter J Whitehouse, Donald L Price, Arthur W Clark, Joseph T Coyle, and Mahlon R DeLong. Alzheimer disease: evidence for selective loss of cholinergic neurons in the nucleus basalis. Annals of Neurology: Official Journal of the American Neurological Association and the Child N...
1981
-
[55]
Differential effects of m1 and m2 receptor antagonists in perirhinal cortex on visual recognition memory in monkeys
Wei Wu, Richard C Saunders, Mortimer Mishkin, and Janita Turchi. Differential effects of m1 and m2 receptor antagonists in perirhinal cortex on visual recognition memory in monkeys. Neurobiology of learning and memory, 98(1):41–46, 2012
2012
-
[56]
Moleculenet: a benchmark for molecular machine learning
Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande. Moleculenet: a benchmark for molecular machine learning. Chemical science, 9(2):513–530, 2018
2018
-
[57]
Frequent hitters: nuisance artifacts in high-throughput screening
Zi-Yi Yang, Jun-Hong He, Ai-Ping Lu, Ting-Jun Hou, and Dong-Sheng Cao. Frequent hitters: nuisance artifacts in high-throughput screening. Drug discovery today, 25(4):657–667, 2020
2020
-
[58]
Novel neuroprotective k+ channel inhibitor identified by high-throughput screening in yeast
Elena Zaks-Makhina, Yonjung Kim, Elias Aizenman, and Edwin S Levitan. Novel neuroprotective k+ channel inhibitor identified by high-throughput screening in yeast. Molecular pharmacology, 65(1): 214–219, 2004. 42
2004
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.