{"id":"de9955e9-86d3-48c9-888b-4d68ba25720e","arxiv_id":"2504.19912","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The DO Challenge benchmark and Deep Thought multi-agent system show frontier LLM agents can roughly match non-expert humans on a synthetic virtual screening task, while remaining far behind expert-designed solutions.","lead":"This paper introduces DO Challenge, a benchmark where AI agents search one million molecular structures for the highest-scoring drug candidates under a limited labeling budget and three submission attempts. A multi-agent system called Deep Thought matched a human expert in the timed setting in its best run, but showed very high run-to-run variance and fell far short of experts when time was not constrained.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The near-expert result for Deep Thought is a single best-of-three o3 run (33.5%); the same configuration averages 14.03 ± 17.11% across three runs, which is below the best human team (16.4%). The headline comparison is therefore not robust.","rationale":"I read the paper as claiming primarily that a general-purpose LLM multi-agent system can, in a constrained virtual-screening task, approach the performance of human ML experts. For that claim to hold, the comparison in Table 1 must be a fair characterization of the system's performance. The weakest point is not the DO Score itself, although its validation is thin; it is that the 33.5% result is the maximum of three runs. The average is 14.03%, which is below the best DO Challenge 2025 team (16.4%), and the standard deviation is larger than the mean. Reporting the maximum is acceptable when documenting best-case capability, but the abstract and introduction state the result as 'Deep Thought achieved' without the best-of-n qualifier. The paper's own Table 5 provides the data needed to see this, and the paper explicitly warns about instability, so this is not a hidden flaw; it is an unaddressed framing problem. A single expert solution (33.6%) is also one human run, but the expert's pipeline (Algorithm D1) is a deterministic procedure, whereas the agent result is the best of stochastic LLM runs. Consequently, the 'near-expert' claim should be downgraded to 'one out of three runs matched expert performance,' which materially weakens the headline contribution. The external-validity question about the DO Score remains important for the 'drug discovery' framing, but the best-run selection is the more direct threat to the paper's quantitative claim. I would keep the CONDITIONAL verdict, with the condition being a median-based re-analysis of repeated runs.","tokens_in":58586,"tokens_out":5274,"duration_ms":56006,"concrete_test":"Run the same o3-based Deep Thought configuration for at least 10 independent trials under the identical 10-hour time-limited protocol, and report the median and 95% bootstrap confidence interval of the DO Challenge score alongside the current Table 1. If the median is near 7.2% (the middle of the observed 1.4/33.5/7.2 runs), the near-expert comparison should be reframed as best-case capability, not typical performance. Additionally, report how many of the 10 trials exceed the 16.4% best human team score.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—that Deep Thought's 33.5% is 'nearly identical' to the top human expert's 33.6% and 'significantly outperforms' the best DO Challenge team (16.4%)—rests on the best of three o3 runs. Table 5 reports for the o3 configuration: Best Score 33.5%, Average Score 14.03% ± 17.11%, with 0 failed runs over 3 runs. Table F13 shows the three scores are 1.4%, 33.5%, and 7.2%. Thus, the expected performance of the system under the same protocol is not near-expert; it is at or below the top competition team. The paper is transparent about run-level variance and warns about instability, but the abstract and Table 1 use the maximum as the headline figure, supporting a stronger claim than the data justify. A capability claim ('there exists a run that matches expert performance') would be defensible; a performance claim ('Deep Thought achieved 33.5%') is selection on the best outcome. The external-validity limitation of the DO Score is a separate and important concern, but this best-run selection is the immediate load-bearing weakness for the headline comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the DO Challenge, a benchmark for autonomous AI agents in a virtual-screening-style drug discovery task. Agents receive one million unlabeled molecular conformations, may query at most 100,000 DO Score labels, make up to three 3,000-molecule submissions, and are evaluated by overlap with the true top-1000. The DO Score is defined as a docking-derived binding probability to a JAK2 target minus the maximum binding probability over three ADMET anti-targets. The paper reports results from the human DO Challenge 2025 competition, two in-house expert solutions, and the Deep Thought multi-agent LLM system, together with ablation studies over LLM roles and agent configurations. The headline claim is that Deep Thought in a 10-hour setup achieved 33.5% overlap, nearly identical to the best human expert's 33.6% and far above the best competition team's 16.4%; the paper also reports an unrestricted expert result of 77.8%. The benchmark and code are released on Zenodo and GitHub.","tokens_in":58798,"tokens_out":4155,"duration_ms":46356,"significance":"If the claims were robust, the paper would be a valuable contribution to agentic AI evaluation for scientific discovery: it provides an integrated, resource-constrained task requiring planning, implementation, and execution, and it includes unusually transparent run-level reporting of agent variance, failure modes, and LLM ablations. The public release of the benchmark and the Deep Thought code is a concrete strength, as is the explicit canary string for leakage prevention. However, the central quantitative claim is weakened by best-of-three selection, and the benchmark's construct validity as a drug-discovery proxy rests on a limited validation. With appropriate re-framing and strengthened validation, the benchmark and the agent study are still publishable and useful.","major_comments":[{"comment":"The headline comparison rests on selecting the maximum of three runs for the o3 configuration. Table F13 reports the three cfg-10 o3 scores as 1.4%, 33.5%, and 7.2%, and Table 5 reports the same configuration's average as 14.03% with standard deviation 17.11%. The abstract and the Section 4 results text present 'Deep Thought achieved 33.5%' and describe this as 'nearly identical' to the expert's 33.6%, without stating that 33.5% is the best of three runs and that the expected score under the same protocol is below the top competition team's 16.4%. A capability claim ('the best run matched expert performance') is defensible, but the current wording supports a stronger performance claim than the data justify. Please either report the mean as the primary measure, or explicitly and consistently label the 33.5% as the best of three runs throughout the abstract, Tables 1 and 2, and the conclusion.","section":"4.2.4, Table 5, Supplementary F Table F13; Abstract and Section 4"},{"comment":"The external validity of the DO Score is a load-bearing assumption for the benchmark's interpretation, and the paper's own validation is weak. Table A2 shows that among 107 DUD-E JAK2 binders, DO Score ranking places only 9 in the top 1000, with EF1% = 8.41, while ranking by the single-target Score6G3C places 19 binders with EF1% = 27.10. Thus the multi-objective score enriches known binders less than the single-target score, and no validation is provided for the anti-target selectivity or for drug-likeness more broadly. The conclusion that the benchmark tests 'drug discovery pipelines' is therefore not firmly established. I recommend either substantially strengthening the validation of Eq. (2) as a drug-candidacy proxy, or explicitly reframing the benchmark as a synthetic virtual-screening task whose labels are defined by the authors' own scoring function.","section":"3.2, Eq. (2), Table A2"},{"comment":"The human expert comparator is a single run per expert, produced by authors from the same organization, with no reported variance or reproducibility check. The 33.6% expert result is therefore as fragile as the agent's best run, and the 'nearly identical to expert' conclusion compares a selected maximum from the agent against a single expert run. This is a further reason the headline comparison is not robust. Please report the expert results as single demonstrations, avoid language implying a stable expert performance level, and state the single-run limitation explicitly when the agent-vs-expert comparison is made.","section":"4, Supplementary D, Tables 1 and 2"}],"minor_comments":[{"comment":"The configuration labels are inconsistent: Table 5 lists the o3 Software Engineer configuration as cfg-11 with scores 1.4/33.5/7.2, while Supplementary F Table F13 identifies the same o3 runs as cfg-10 and cfg-11 is described as 'without Reviewer' in Section 4.2.4. Similarly, Table 5 lists Gemini 2.0 Flash as cfg-7 and Claude 3.5 Haiku as cfg-9, whereas Supplementary F assigns cfg-7 to Claude 3.5 Haiku and cfg-8 to Gemini 2.0 Flash. Please reconcile all configuration identifiers between the main text and the supplementary tables.","section":"Table 5 vs. Supplementary F"},{"comment":"The number of competition teams is stated inconsistently: the abstract and Section 4 say 20 human teams, while Section 4.1 says 24 teams were selected and 20 ultimately engaged. Please make these numbers consistent.","section":"4.1"},{"comment":"The dataset description says 500,000 molecules were randomly sampled and filtered, then 'a sample of 200,000 molecules was retained,' but it is not clear how the 200,000 were selected from the filtered set or how the 1,000,000 conformations correspond to the 200,000 molecules. Clarify the sampling procedure.","section":"3.3"},{"comment":"The enrichment validation description should state explicitly whether the 107 DUD-E binders were added to the 1M dataset as extra molecules or docked poses, and how the 'TOP 1000 Hits' overlap is defined when multiple poses of the same molecule may appear; this affects interpretation of the enrichment factors.","section":"Supplementary A, Table A2"}],"recommendation":"major_revision","confidential_remarks":"The benchmark and system release are real contributions, and the run-level transparency is commendable. The main risk for the journal is that the abstract's near-expert claim will be publicized without the best-of-three caveat; tightening that framing and the construct-validity discussion should be sufficient. I do not see a need for rejection, provided the authors revise the headline claims and reconcile the configuration tables."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, DO Challenge is a genuinely useful benchmark: a million-conformation virtual screening setup with a 100k label budget, three submissions, and a public competition—nothing else quite does this. Second, the headline result—Deep Thought at 33.5%, nearly identical to a top human expert at 33.6%—is the best of three o3 runs. The other two scored 1.4% and 7.2%; average 14.03 ± 17.11, below the best DO Challenge team (16.4%). So the capability claim \"there exists a run that matches expert performance\" is defensible; the performance claim as stated is selection on the maximum.\n\nThe paper does several things well. The benchmark design is thought through: the resource constraints, the submission feedback loop, and the open competition give a realistic testbed for agentic decision-making, not just a static prediction task. The Deep Thought system is described in enough detail to be a useful baseline, and the ablation study across models and agent roles is informative, especially the failure-mode analysis—agents ignoring position-sensitivity, missing the label budget, getting stuck in bug-fixing loops. They report run-level results and caution about instability, which is more honest than most agent papers. Code and data are public.\n\nSoft spots, in order of importance. (1) Best-run selection: the abstract and Table 1 use 33.5%, while the same configuration's average is 14.03 ± 17.11%. A reader should compare average or median, not max, when the paper itself shows high variance. (2) Expert baselines: both experts are from the same organization, and each is a single run; that's a thin reference point. (3) DO Score external validity: the enrichment test on DUD-E JAK2 binders shows EF1% of 8.41 at 1%, but only 9 of 107 binders appear in the top 1000, and the plain target score has EF1% 27.10. So the synthetic label's drug-likeness is only weakly supported. Same-lab design of benchmark and agent is a conflict, though the labels come from external docking/classifiers, so it's not circular.\n\nBottom line: the benchmark deserves serious referee time. A careful review should ask for average-run reporting, outside expert runs, and more validation of the DO Score. I'd bring it to our reading group and would cite the benchmark if I were working on agent evaluation.","headline":"Worth engaging as a benchmark paper, but the near-expert claim rests on a best-of-three run; the average o3 performance is below the top human team.","tokens_in":59484,"tokens_out":2885,"would_cite":true,"duration_ms":26503,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A multi-agent AI system matched a human expert on a virtual drug-screening task, scoring 33.5% to 33.6% in 10 hours.","keywords":["AI agents","multi-agent systems","drug discovery","virtual screening","benchmark","LLM agents","active learning","docking"],"falsifier":"Run the Deep Thought cfg-10 configuration with o3 as the primary agent 100 times under identical conditions and plot the score distribution; if the 33.5% result is a rare upper-tail event and the typical score is near the reported 14.03% mean, the claimed match with the human expert is best-of-luck, not representative agent behavior.","tokens_in":58352,"feed_emoji":"🧬","tokens_out":7329,"duration_ms":71861,"temperature":0.7,"pith_summary":"This paper attempts to show that a general-purpose multi-agent AI system can plan, write, and execute a drug-discovery pipeline from scratch, not just solve one predictive task. It introduces DO Challenge, a benchmark in which an agent must find the 1,000 highest-scoring molecules among one million docked conformations with a limited label budget, and reports that its Deep Thought system scored 33.5% in the 10-hour setup, nearly matching the top human expert's 33.6% and more than doubling the best human competition team's 16.4%. The result matters because it locates the frontier of autonomous AI in scientific discovery in end-to-end strategic decision-making, not only model accuracy. The paper also reports that the system is unstable across runs and remains far behind an unrestricted expert solution at 77.8%, so it frames the result as promising but not yet transformative.","feed_headline":"AI agent matches human expert on drug screening","feed_subtitle":"In a 10-hour virtual screen of 1M molecules, Deep Thought scored 33.5% versus the expert's 33.6%.","key_machinery":"The load-bearing object is the DO Score, defined in Eq. (2) as the mean probability of binding to the therapeutic target 6G3C minus the maximum binding probability to three ADMET-related anti-targets (1W0F, 8YXA, 8ZYQ), computed from docking poses and two logistic-regression classifiers. This single scalar label defines the ground truth behind the benchmark's top-1,000 list, so the surrounding protocol—a label-query budget of 100,000 out of one million structures, three submissions of 3,000 structures each, and an overlap metric—tests how efficiently an agent can reconstruct that ranking. The other carrying component is Deep Thought itself, a multi-agent system of role-specialized LLM agents whose code-writing-and-execution loop converts a task prompt into a submitted solution.","core_discovery":"The central claim, on the paper's terms, is that an LLM-based multi-agent system can compete with human ML experts on a constrained virtual screening task. In the DO Challenge benchmark, Deep Thought's best time-limited configuration with o3 as the primary agent achieved a 33.5% overlap with the true top 1,000 molecules, effectively tied with the top human expert's 33.6% and far ahead of the best DO Challenge 2025 team's 16.4%. The same configuration averaged only 14.03% with a standard deviation of 17.11% across three runs, and the paper documents failure modes ranging from ignoring available tools to exhausting the label budget. In the time-unrestricted setup, the best expert reached 77.8% while Deep Thought reached 33.5%, though a differently configured Deep Thought scored 50.3% in the post-challenge extension. The paper concludes that autonomous agents show real potential but still fall short of expert-designed solutions.","pith_inferences":["Editorial inference: because the same o3 configuration averaged 14.03% with a 17.11% standard deviation, benchmark comparisons should be based on score distributions over repeated runs, not best-of-three maxima, before declaring an agent competitive with a human expert.","Editorial inference: if the DO Score proxies real drug-likeness even roughly, the active-learning-plus-surrogate strategy Deep Thought discovered should transfer to other label-expensive search problems, such as materials or reaction-condition discovery, where the same planning loop can be pointed at a different oracle.","Editorial inference: the unrestricted expert's advantage suggests a concrete remedy—equipping the agent with position-aware architectures and a top-k weighted loss, since those were the decisive components in the 77.8% solution.","Editorial inference: the near-zero contribution of the Research agent group in runs with advanced models hints that web research is not the bottleneck for this task; execution reliability and budget discipline are."],"forward_implications":["If the benchmark is accepted, a general LLM-based multi-agent system can carry out an end-to-end ML drug-screening workflow—data exploration, model selection, code writing, execution, and submission—without task-specific hints.","The near-tie with the time-limited human expert implies that automated planning can reach the level of an experienced ML practitioner on this kind of constrained task.","The gap in the unrestricted setup implies that the remaining bottleneck is not basic feasibility but sustained strategy, model quality, and hyperparameter discipline.","The ablation results imply that the choice of LLM in the primary agent role is decisive: frontier models produced competitive code while smaller models mostly failed to finish.","The four correlated success factors—strategic structure selection, spatial-relational networks, position non-invariance, and strategic submission—offer a concrete design checklist for future scientific agents."],"supporting_citations":[{"why":"Supplies the JAK2 binder set from DUD-E used in the enrichment validation of the DO Score.","marker":"[35]"},{"why":"Provides the DEKOIS2 dataset on which the two logistic-regression classifiers defining target binding probabilities were trained.","marker":"[33]"},{"why":"Provides the AutoDock Vina energy scoring used by one of the two classifiers that make up the DO Score.","marker":"[34]"},{"why":"Source of the therapeutic target structure 6G3C that anchors the positive term of the DO Score.","marker":"[37]"},{"why":"Source of the CYP3A4 structure used as one of the three ADMET anti-targets in the DO Score.","marker":"[38]"},{"why":"Source of the human serum albumin structure used as an ADMET anti-target in the DO Score.","marker":"[39]"},{"why":"Source of the hERG channel structure used as the cardiac-safety anti-target in the DO Score.","marker":"[40]"},{"why":"Uni-Mol v2 is the pretrained model used by the top time-restricted human expert solution, part of the baseline the agent is compared against.","marker":"[42]"},{"why":"GENConv is the graph convolution used in the unrestricted expert solution's GNN, the strongest baseline in the paper.","marker":"[43]"},{"why":"Provides the node and edge features reused in the expert GNN that achieved 77.8%.","marker":"[45]"}],"fun_headline_variants":["AI agent ties human expert in drug screening benchmark","AI matches top human on virtual drug screen task","Deep Thought AI ties expert in drug discovery pipeline","AI agent scores 33.5% vs human 33.6% in drug screen","Autonomous AI rivals human expert in drug discovery"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison rests on the assumption that the synthetic DO Score, a single docking-derived number, is a meaningful proxy for drug candidacy; if that label does not reflect real drug-likeness, the agent-versus-human scores say little about actual drug discovery.","fun_headline_variants_meta":{"raw":{"variants":["AI agent ties human expert in drug screening benchmark","AI matches top human on virtual drug screen task","Deep Thought AI ties expert in drug discovery pipeline","AI agent scores 33.5% vs human 33.6% in drug screen","Autonomous AI rivals human expert in drug discovery"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000248,"raw_usage":{"total_tokens":1576,"prompt_tokens":1007,"completion_tokens":569,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":623,"completion_tokens_details":{"reasoning_tokens":503}},"tokens_in":623,"tokens_out":569,"duration_ms":5814,"temperature":1.0,"reasoning_tokens":503,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:39:42.165895+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the Deep Thought cfg-10 configuration with o3 as the primary agent 100 times under identical conditions and plot the score distribution; if the 33.5% result is a rare upper-tail event and the typical score is near the reported 14.03% mean, the claimed match with the human expert is best-of-luck, not representative agent behavior.","supporting_citations":[{"cited_title":"and Shoichet, B.K., 2012","cited_arxiv_id":null,"evidence_quote":"Supplies the JAK2 binder set from DUD-E used in the enrichment validation of the DO Score."},{"cited_title":"and Boeckler, F.M., 2013","cited_arxiv_id":null,"evidence_quote":"Provides the DEKOIS2 dataset on which the two logistic-regression classifiers defining target binding probabilities were trained."},{"cited_title":"and Olson, A.J., 2010","cited_arxiv_id":null,"evidence_quote":"Provides the AutoDock Vina energy scoring used by one of the two classifiers that make up the DO Score."},{"cited_title":"and Eck, M.J., 2019","cited_arxiv_id":null,"evidence_quote":"Source of the therapeutic target structure 6G3C that anchors the positive term of the DO Score."},{"cited_title":"and Jhoti, H., 2004","cited_arxiv_id":null,"evidence_quote":"Source of the CYP3A4 structure used as one of the three ADMET anti-targets in the DO Score."},{"cited_title":"and Doi, Y., 2024","cited_arxiv_id":null,"evidence_quote":"Source of the human serum albumin structure used as an ADMET anti-target in the DO Score."},{"cited_title":"and Murata, T., 2024","cited_arxiv_id":null,"evidence_quote":"Source of the hERG channel structure used as the cardiac-safety anti-target in the DO Score."},{"cited_title":"and Petrosyan, G., 2025","cited_arxiv_id":null,"evidence_quote":"Provides the node and edge features reused in the expert GNN that achieved 77.8%."}],"review_version":1}