{"id":"a5447865-d0dc-48f9-b81d-c4b565e03a89","arxiv_id":"2411.09820","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"WelQrate provides nine curated PubChem-derived screening datasets plus standardized splits, metrics, and formats, and benchmarks eight model families on them.","lead":"This paper introduces WelQrate, a curated collection of nine high-throughput screening datasets and a standardized evaluation framework for benchmarking AI models in small molecule drug discovery. It also runs a multi-model benchmark to show how dataset quality, featurization, and splitting choices change measured performance.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The clean-label premise is undermined because inactives are unconfirmed primary-screen negatives; the paper's own retest notes show false negatives without quantifying them.","rationale":"The reader's weakest assumption is exactly the point that matters most. A benchmark is only as good as its labels, and WelQrate's headline contribution is label quality. The paper validates actives through hierarchical confirmatory/counter screens but leaves the vast majority of compounds—the inactives—at the primary-screen stage. The supplement contains direct evidence that this stage produces false negatives: AID435008's curation note says the two primary-screen inactive sets were intersected 'to reduce false-negative rate, since we found some inactive compounds in one screen that were active in the other'; AID2258 and AID488997 likewise report primary inactives later active in confirmatory screens. None of this is quantified, and the Limitations section omits it. Without a false-negative estimate, the 'gold standard' claim rests on an unverified assumption about the majority class. The proposed check is feasible because the paper already identifies the relevant AIDs and the retested compounds are documented in the pipeline diagrams. I also note the paper's positive features: the curation descriptions are unusually detailed, the control dataset in RQ2 is a sensible experiment, the adapted cross-validation is a reasonable practical compromise, and the data/code are stated to be public (though no commit hash or direct link appears in the text). These do not resolve the inactive-label concern. The overclaim about 'gold standard' without a head-to-head comparison against MoleculeNet/TDC is a secondary issue; it would be less damaging if the negative labels were clean. Therefore the correct outcome remains a conditional acceptance: the dataset is a valuable resource, but the clean-inactive premise must be demonstrated before the gold-standard recommendation can be fully trusted.","tokens_in":30238,"tokens_out":5507,"duration_ms":55272,"concrete_test":"For each of the nine datasets, use the PubChem AIDs already listed in Supplement A.2 to identify all primary-screen inactives that were subsequently tested in a confirmatory or dose-response assay (the compounds drawn with red dashed arrows in the curation diagrams). Compute f = (# such compounds with an active confirmatory readout) / (# such compounds retested). If f is below 0.1% for a dataset, the primary-inactive labels for that dataset are probably clean; if f is above 1% for AID1798 or AID2258, the clean-negative premise is materially weakened. As a follow-up, rerun RQ1/RQ2 with those confirmed-active compounds removed from the negative set and check whether model rankings or headline conclusions change.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central value proposition is 'clean and reliable data labels,' but the curation pipeline only validates actives through confirmatory and counter screens. For most of the nine datasets, the final inactive set is drawn primarily from primary-screen inactive readouts (e.g., AID1798: 'final inactive set was retrieved from inactive readouts in the primary screen AID626'; AID435008 uses the intersection of two primary-screen inactive sets). Primary HTS thresholds are deliberately loose to avoid false negatives, so primary inactives are unconfirmed negatives. The supplement itself documents primary-inactive compounds later found active in confirmatory screens (AID435008; AID2258's six retested compounds; AID488997's retested inactives), yet no dataset-level false-negative rate is reported anywhere, and the Limitations section does not mention inactive-label noise. If hidden actives are non-negligible, training and test negatives are contaminated, ranking metrics such as logAUC and BEDROC can be distorted, and the benchmark's core claim of clean labels fails for the majority class. This is an internal gap between the stated curation goal and the inactive-label pipeline, not a disagreement with community convention.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces WelQrate, a benchmark suite for small-molecule drug discovery consisting of nine PubChem-derived datasets with a hierarchical curation pipeline, standardized data formats (SMILES, InChI, SDF, 2D/3D graphs), a proposed evaluation protocol (metrics, random and scaffold splits, adapted cross-validation), and extensive benchmarking of ten models across four research questions. The central claims are that the datasets have clean and reliable labels due to expert-designed curation, and that the framework provides a standardized, realistic basis for virtual-screening model comparison. The paper also argues that dataset quality, featurization, and split strategy materially affect model rankings, and recommends WelQrate as a new gold standard.","tokens_in":30423,"tokens_out":4914,"duration_ms":50635,"significance":"If the curation and evaluation claims hold, WelQrate would be a valuable community resource: the authors provide public data, curation code, and experimental scripts; the datasets are large and realistically imbalanced; the supplementary material documents the full curation hierarchy for each AID; and the benchmarking includes standard and domain-baseline models with error bars and hyperparameter tables. The inclusion of multiple realistic ranking metrics (logAUC, BEDROC, EF100, DCG100) and scaffold-split evaluation is a genuine strength. However, the significance is contingent on two load-bearing points: the reliability of the inactive labels, which are mostly unconfirmed primary-screen negatives, and the validity of the adapted cross-validation protocol. Because the paper's headline contribution is 'clean and reliable data labels,' the inactive-label issue directly affects the benchmark's core value proposition and must be addressed before the gold-standard claim is supportable.","major_comments":[{"comment":"The central claim of 'clean and reliable data labels' (Sec. 3) is not established for the majority class. For most datasets, the final inactive set is taken directly from primary-screen inactive readouts (e.g., AID1798, AID435034, AID1843, AID2258, AID2689, AID485290; see Supp. A.2), and the paper itself notes that primary HTS thresholds are deliberately loose to reduce false negatives (Sec. 3.2). The curation notes also document primary-inactive compounds later found active in confirmatory screens: AID435008 states that 'we found some inactive compounds in one screen that were active in the other'; AID2258 describes six primary-inactive compounds that were retested, with final active readouts causing their exclusion; AID488997 reports 17, 38, and 2 primary-inactive compounds tested in confirmatory screens, some of which were active. No dataset-level false-negative rate is reported anywhere, and Section 6 (Limitations) does not mention inactive-label noise. Since inactives constitute the vast majority of each dataset, unquantified false negatives can distort enrichment and ranking metrics, so the clean-label premise needs either explicit quantification using the available retest data or a substantially softened claim and a corresponding limitation statement.","section":"Sec. 3.2; Supp. A.2"},{"comment":"The RQ2 conclusion that the results 'align with data-centric AI, highlighting the importance of dataset quality' is not supported for all model families: the 2D and 3D graph-based models trained on the less clean control data outperform models trained on WelQrate in logAUC[0.001,0.1] and BEDROC, while performing worse on EF100 and DCG100. The authors offer only an untested hypothesis about the 'range of top selected candidates.' As Fig. 4 averages across datasets and the effect is metric-dependent, the conclusion should either be restricted to the architectures and metrics where curation consistently helps (Naive, sequence-based, and Domain baselines) or be backed by a per-dataset, per-metric analysis that explains the reversal. As written, RQ2 does not provide a coherent demonstration that dataset quality improves model evaluation across the board.","section":"Sec. 5.2; Fig. 4"},{"comment":"The adapted cross-validation protocol, in which the validation fold is fixed as the fold immediately preceding the test fold, is asserted to 'enhance computational efficiency without compromising robustness,' but no evidence is provided that this protocol yields estimates comparable to nested cross-validation or to standard k-fold cross-validation with proper hyperparameter tuning. Since all random-split results (RQ1-RQ3) rely on this protocol, the evaluation framework's reliability claim depends on this untested assumption. A small-scale comparison of the adapted protocol against nested cross-validation on one or two datasets would be sufficient to support the claim, or the assertion should be replaced with a more cautious statement.","section":"Sec. 4.3; Supp. B.2"}],"minor_comments":[{"comment":"There is a typo in the bullet list: 'theurapeutic' should be 'therapeutic'; the author affiliation line also contains a stray 'Electical and Computer Engineering Dept„' with a formatting artifact.","section":"Sec. 1"},{"comment":"The text says 'a 3:1:1 training:validation ratio' but appears to mean a 3:1:1 train:validation:test ratio; please clarify the wording to avoid ambiguity about whether the test set is included.","section":"Sec. 4.3"},{"comment":"In the AID1843 row, the number of unique BM scaffolds is listed as '82,140C' with a stray 'C' suffix; this appears to be a typo.","section":"Table 1"},{"comment":"The displayed BEDROC formula after 'calculated as:' appears garbled: the summation term is written as 'P n i=1 -eri/N', which omits the exponential and the alpha parameter shown in the RIE definition above it. Please correct the equation.","section":"Supp. B.1"},{"comment":"The main text states that 'other metrics exhibit the same trend' for the one-hot versus predefined-feature comparison, but Supp. C.4 says the predefined features outperform 'in the majority of cases,' which is a weaker statement. Please align the main-text claim with the supplementary results.","section":"Fig. 5; Supp. C.4"},{"comment":"The choice of 1000 µM as a placeholder for inactive compounds in the three datasets with additional measurements is acknowledged in Supp. A.5, but the main text does not mention this artificial value or its potential effect on regression tasks. A one-sentence caveat in Sec. 4.1 or 4.2 would help readers who use the floating-value labels.","section":"Sec. 4.2; Supp. A.5"}],"recommendation":"major_revision","confidential_remarks":"This is a substantial and useful dataset/benchmark paper, and the authors have done considerable curation and release work. I recommend major revision rather than rejection because the main technical concern, inactive-label noise, is addressable with additional analysis and a revised discussion. The adapted cross-validation validation study is also feasible within the scope of a revision. I would not insist on new model results, but the clean-label claim and the RQ2 interpretation need to be brought in line with the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, you should know: this is a genuinely useful benchmark paper with unusually transparent curation, but the 'gold standard' label is an overclaim. The inactives are mostly unconfirmed primary-screen negatives, and the paper never quantifies the resulting label noise.\n\nWhat's new: the WelQrate collection itself—9 PubChem assays spanning five target classes, curated through hierarchical primary/confirmatory/counter screens, with PAINS and druglikeness filters, expert verification, and multi-format releases (SMILES, InChI, SDF, 2D/3D graphs). The evaluation framework is sensible: standard metrics (logAUC, BEDROC, EF100, DCG100), a practical adapted cross-validation, and scaffold splits with Bemis-Murcko. That combination doesn't exist elsewhere. The curation documentation is the paper's strongest asset—the supplement walks through every assay and shows the actual decisions, including cases where primary inactives were later found active.\n\nThe soft spots are real but not fatal. The biggest one: for most datasets, the final inactive set is just the primary-screen negatives. Primary HTS thresholds are deliberately loose, so these are unconfirmed negatives. The paper's own notes show some of them turn out active in confirmatory screens (e.g., AID2258, AID435008), yet there is no dataset-level false-negative rate, and Limitations never mentions inactive-label noise. That gap matters because the clean-label claim is load-bearing. Second, RQ2 cuts against the 'curation always helps' narrative: 2D and 3D graph models trained on the uncurated control do better on logAUC and BEDROC. The paper's explanation ('range of top selected candidates') is speculative. Third, RQ3 is run on a single dataset, so the featurization conclusion is thin. Minor issues: no commit hash or clickable link in the text, and exact 3D regeneration requires commercial Corina.\n\nWho it's for: anyone building or evaluating ligand-based virtual screening models. It deserves a serious referee: the curation work is extensive, the release is reusable, and the open questions are concrete—quantify inactive-label noise, add a head-to-head against MoleculeNet/TDC, and soften the gold-standard claim. My verdict: revise, not reject.","headline":"Useful, carefully documented benchmark with an overclaimed 'gold standard' label and an unquantified inactive-label noise problem.","tokens_in":31035,"tokens_out":2518,"would_cite":true,"duration_ms":24402,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"WelQrate makes the case that curated screening data, not just model architecture, determines whether drug-discovery benchmarks are trustworthy.","keywords":["drug discovery benchmarking","virtual screening","high-throughput screening","dataset curation","hierarchical bioassay curation","molecular representation learning","scaffold split","early enrichment metrics"],"falsifier":"Re-test a random sample of the final inactive compounds from one WelQrate dataset, say AID1798, in the confirmatory assay used there (AID1488) under dose-response conditions; if a substantial fraction, for example several percent, reproducibly show activity, the inactive labels are too optimistic and the clean-label premise fails. A cheaper version is to search the public bioassay records for compounds labeled inactive in WelQrate that later returned active readouts in related follow-up assays and count how often that happens.","tokens_in":29956,"feed_emoji":"🧪","tokens_out":7991,"duration_ms":74766,"temperature":0.7,"pith_summary":"The paper argues that machine-learning models for small-molecule drug discovery are being compared on benchmarks with noisy labels, inconsistent chemical representations, and ad hoc splits, so rankings may not reflect real ability. To fix this, it proposes WelQrate: nine datasets across five therapeutic target classes, curated by hand-inspected hierarchies of primary, confirmatory, and counter screens from the public bioassay database, followed by filters for promiscuous, interfering, and non-druglike compounds. It wraps these datasets in a standardized evaluation protocol covering molecular formats, 3D conformations, early-enrichment metrics, and random versus scaffold splits. The benchmarking results show that dataset quality, featurization, and split type all shift model rankings, with a domain-expert descriptor baseline often beating deep learning models. If the curation is sound, the community gains a common testbed on which virtual-screening claims can be compared fairly.","feed_headline":"Cleaner labels change which drug-screening models win","feed_subtitle":"WelQrate re-tests active compounds before letting models compete on the same metrics and splits.","key_machinery":"The load-bearing mechanism is the hierarchical curation pipeline. It organizes bioassays by level: a primary screen with a deliberately loose threshold, confirmatory screens that re-test putative actives, and counter screens that reject compounds with off-target or nonspecific activity; final actives are validated hits and final inactives come from primary-screen inactivity, with noted exceptions kept when follow-up readouts contradict. This pipeline is what converts raw screening data into labels the paper claims are clean and realistic. Around it, the framework adds standardized formats (isomeric SMILES and InChI, plus precomputed 2D and 3D graphs), early-enrichment metrics such as logAUC, BEDROC, EF100, and DCG100, and split protocols including nested-style random cross-validation and Bemis-Murcko scaffold splits. The central working assumption is that label quality, not model architecture alone, determines whether a benchmark ranking transfers to real screening.","core_discovery":"On the paper's own terms, the central claim is that a benchmark built from hierarchically curated high-throughput screening data, rather than raw primary-screen readouts, can serve as a gold standard for small-molecule virtual screening. The curation works by tracing each compound through primary, confirmatory, and counter assays, keeping only actives that survive validation and only inactives from primary screens that were not contradicted by later readouts, then applying promiscuity, PAINS, druglikeness, and representation filters. Around this collection, the paper standardizes featurization, 3D conformation generation, evaluation metrics that reward early enrichment, and split schemes including scaffold splits. Its benchmarking experiments find that models improve with complexity under random splits, that a domain-expert descriptor with a simple classifier outperforms the neural models, that training on uncurated primary-screen data changes or reverses some comparisons, that predefined features beat one-hot features, and that all models struggle under scaffold splits. The paper takes these results as evidence that both data quality and evaluation design must be reported and standardized for meaningful model comparison.","pith_inferences":["Beyond the paper, the curation's weak point is the inactive set: most inactives are primary-screen misses never validated in confirmatory screens, so if primary-screen false negatives are frequent, the clean-label premise is weakened even though actives are well validated.","A testable extension is to re-run the benchmark after applying a stricter inactive definition, for example requiring inactivity in a confirmatory or counter screen, and measure how rankings shift.","The additional dose-response measurements available for three datasets invite a regression benchmark for potency prediction, which the paper mentions but does not develop.","If the gold-standard framing is adopted widely, the field's next problem becomes split and feature variation across benchmarks; WelQrate's fixed protocols make cross-paper comparisons possible only if researchers report the version and split used."],"forward_implications":["Model rankings on WelQrate become a more trustworthy basis for choosing virtual-screening methods, because labels have passed confirmation and counter-screening.","The strong performance of a simple model on domain-expert descriptors implies that architecture comparisons should include such a baseline and that featurization matters as much as model design.","Training on uncurated primary-screen data can invert or erase performance differences, so benchmark results that ignore curation are hard to interpret.","Scaffold splits expose distribution shift: all tested models lose accuracy, meaning claims of generalization need to be evaluated under scaffold rather than random splits.","Adopting standardized metrics that reward early enrichment aligns benchmark scores with the real workflow of buying or synthesizing only the top-ranked compounds."],"supporting_citations":[{"why":"Supplies the high-throughput screening dataset methodology and the framing that primary screens carry a high false positive rate.","marker":"[9]"},{"why":"Supplies the bioassay identification strategy of relevance, quality, and consistency used to select WelQrate targets.","marker":"[13]"},{"why":"Supplies the PAINS substructure filters used to remove pan-assay interference compounds.","marker":"[18]"},{"why":"The earlier benchmark collection whose data-quality issues motivate the new curation effort.","marker":"[4]"},{"why":"The other earlier benchmark collection whose data-quality issues motivate the new curation effort.","marker":"[6]"},{"why":"Supplies Bemis-Murcko scaffolds, which define the standardized scaffold split for testing scaffold hopping.","marker":"[30]"},{"why":"Defines the BEDROC early-enrichment metric used in the evaluation framework.","marker":"[26]"},{"why":"Supplies the signed autocorrelation descriptors behind the domain baseline that outperforms the deep learning models.","marker":"[23]"}],"fun_headline_variants":["Curated drug data reshuffles AI model rankings","WelQrate: Cleaner labels, different winners in drug AI","Drug benchmark's curated screens redraw the podium","Gold-standard drug data flips which models win"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's clean-label claim rests on the assumption that compounds labeled inactive, mostly from primary screens without confirmatory or counter-screen validation, really are inactive; if primary-screen misses are common, the inactive half of every dataset is noisier than the curation suggests.","fun_headline_variants_meta":{"raw":{"variants":["Curated drug data reshuffles AI model rankings","WelQrate: Cleaner labels, different winners in drug AI","Drug benchmark's curated screens redraw the podium","Gold-standard drug data flips which models win"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0002,"raw_usage":{"total_tokens":1428,"prompt_tokens":1048,"completion_tokens":380,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":664,"completion_tokens_details":{"reasoning_tokens":316}},"tokens_in":664,"tokens_out":380,"duration_ms":5027,"temperature":1.0,"reasoning_tokens":316,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:17:00.478224+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-test a random sample of the final inactive compounds from one WelQrate dataset, say AID1798, in the confirmatory assay used there (AID1488) under dose-response conditions; if a substantial fraction, for example several percent, reproducibly show activity, the inactive labels are too optimistic and the clean-label premise fails. A cheaper version is to search the public bioassay records for compounds labeled inactive in WelQrate that later returned active readouts in related follow-up assays and count how often that happens.","supporting_citations":[{"cited_title":"Molecular structures: perception, autocorrelation descriptor and sar studies","cited_arxiv_id":null,"evidence_quote":"Supplies the high-throughput screening dataset methodology and the framing that primary screens carry a high false positive rate."},{"cited_title":"3d structure generator corina classic","cited_arxiv_id":null,"evidence_quote":"Supplies the bioassay identification strategy of relevance, quality, and consistency used to select WelQrate targets."},{"cited_title":"Muscarinic antagonist control of myopia: evidence for m4 and m1 receptor-based pathways in the inhibition of experimentally-induced axial myopia in the tree shrew","cited_arxiv_id":null,"evidence_quote":"The earlier benchmark collection whose data-quality issues motivate the new curation effort."},{"cited_title":"The selective muscarinic agonist xanomeline improves both the cognitive deficits and behavioral symptoms of alzheimer disease","cited_arxiv_id":null,"evidence_quote":"The other earlier benchmark collection whose data-quality issues motivate the new curation effort."},{"cited_title":"Regulation of neuronal t-type calcium channels","cited_arxiv_id":null,"evidence_quote":"Defines the BEDROC early-enrichment metric used in the evaluation framework."},{"cited_title":"Genetic variation of cacna1h in idiopathic generalized epilepsy.Annals of neurology, 55(4):595–596, 2004","cited_arxiv_id":null,"evidence_quote":"Supplies the signed autocorrelation descriptors behind the domain baseline that outperforms the deep learning models."}],"review_version":1}