{"id":"86935a33-e794-4c44-abb3-0b23e572ca66","arxiv_id":"2608.06961","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"CAi Copilot, a three-layer LLM agent, converts broad molecular-design requests into executed, evidence-traceable workflows and outperforms five baseline agents on 45 curated tasks plus external benchmarks.","lead":"A new AI assistant called CAi Copilot plans and executes multi-step molecular design workflows, generating, filtering, and ranking drug candidates while keeping an evidence trail. It beat five other AI agents on the authors' 45-task benchmark and several external tests, though the promised reduction in human workload is not directly measured.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CAiMD head-to-head is compromised: Appendix B.2 reveals CAi was a pre-existing reference run rescored under the rubric, not a sixth queue in the same live harness as baselines, so the 18-point lead may reflect run conditions rather than architecture.","rationale":"Read in good faith, the paper's central claim is that CAi converts intent to evidence more reliably than five baselines on CAiMD, with external benchmarks as supporting evidence. For that claim to hold, CAi and the baselines must have been evaluated under identical run conditions; otherwise the 18-point lead is not interpretable. The weakest point is execution fairness, not architecture or benchmark representativeness. Appendix B.2 explicitly distinguishes CAi as 'a previously completed reference run... rather than a sixth queue in the baseline supervisor.' That single sentence introduces an unverified asymmetry: the paper never states that the reference run was subject to the same 2,400s clock, same prompt version, same no-retry rule, and same adapter and artifact contract as the baselines. The line 'No task was rerun solely to improve a score' does not close this gap; it only rules out one kind of post-hoc selection. The reader's weakest assumption, benchmark self-authorship and judge bias, is related but distinct: it concerns the measuring instrument, whereas this concern concerns whether CAi's measurement was taken under the same conditions as the instruments it is compared against. External benchmarks provide some independent support for CAi's tool-coordination capability, but they do not test the CAiMD comparison, which is the source of the headline '84.59 versus 66.52' result. Therefore the verdict should remain conditional: release the harness, rerun CAi as a sixth queue, and make judge prompts and rubric public. If the rerun reproduces the lead, the concern is resolved; if not, the central performance claim is unsupported as stated.","tokens_in":15806,"tokens_out":4010,"duration_ms":44141,"concrete_test":"Rerun CAi as a sixth queue in the same baseline supervisor used for ChemCrow, Biomni, AgentD, Codex, and Hermes, with the identical 2,400s budget, prompt template, isolated workspace creation, adapter, and no-retry failure policy, then recompute Table 1 from the new trajectory. If the outcome score and valid-output rate change materially, for example if the 18-point lead shrinks to near zero, the current headline comparison is an artifact of run conditions. As a secondary check, have an independent group score a random sample of original and rerun trajectories with the released blinded judge prompts, verifying that the judge's scores match the paper's reported values.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 4.1 states that for CAiMD 'all agents use the same DeepSeek V4 Pro-compatible endpoint, prompts, inputs, isolated workspaces, and 2,400s budget.' Appendix B.2, however, says: 'CAi is a previously completed reference run normalized and scored under the same evaluation rubric, rather than a sixth queue in the baseline supervisor.' These two statements are not obviously compatible. If CAi's trajectory was produced before the live baseline harness existed, it is unclear whether the 2,400s budget, the no-retry policy, the prompt template, and the workspace setup were actually enforced for CAi. The appendix only rules out rerunning a task 'solely to improve a score'; it does not state that CAi was run under identical live conditions, with the same clock, the same failure handling, and the same adapter that writes the common artifact contract. Because the 84.59 versus 66.52 headline comparison comes exclusively from CAiMD, this asymmetry is load-bearing: if CAi's run benefited from additional time, manual intervention, or benchmark-aware iteration, the claimed advantage over the five baselines would not be a property of the three-layer architecture. This issue is separate from benchmark representativeness or LLM-judge bias, though it compounds them: the reference specifications and blinded judge prompts are authored by the CAi team and are not yet released, so the scoring cannot be independently audited either. The external SMDD/LIDDiA/MolBench results use native protocols and are less affected, but they do not compare CAi against the same five baselines on the same tasks.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes CAi Copilot, a three-layer agentic workflow system for early-stage molecular design, formulated as intent-to-evidence execution. The Research Interface Layer converts research intent into an executable plan, the Agent Reasoning Layer adapts the plan from interim results, and the Execution Substrate provides molecular tools and metrics. The authors evaluate CAi on their own 45-task CAiMD benchmark against five baselines, reporting an outcome score of 84.59 versus 66.52 for the next-best system, and on external benchmarks (SMDD-Bench, LIDDiA, MolBench) with native protocols. They conclude that reliable molecular-design execution requires joint strength across planning, execution, and evidence grounding, with long-horizon integrated workflows remaining a limitation.","tokens_in":16074,"tokens_out":6081,"duration_ms":62956,"significance":"If the claims hold, the paper offers a useful task formulation and a modular architecture that makes workflow planning, execution, and evidence tracing explicit. The external benchmark results (Tables 3-5) provide partial independent validation, and the paper is transparent about several limitations, including the exclusion of efficiency measurements and the absence of experimental validation. However, the headline CAiMD comparison rests on a self-authored benchmark whose semantic rubric is not yet released, and an appendix statement suggests CAi's trajectory was a pre-existing reference run rather than a live baseline-supervisor run, which could undermine the main head-to-head result. The 'operational workload reduction' in the title is also never directly measured. The core architectural ideas are promising, but the main quantitative claims require clarification and re-analysis before the paper can be accepted.","major_comments":[{"comment":"Section 4.1 states that for CAiMD 'all agents use the same DeepSeek V4 Pro-compatible endpoint, prompts, inputs, isolated workspaces, and 2,400s budget.' Appendix B.2 states that CAi is 'a previously completed reference run normalized and scored under the same evaluation rubric, rather than a sixth queue in the baseline supervisor.' These statements are not obviously compatible. If CAi's trajectory was produced before the live baseline harness existed, the 2,400s budget, the no-retry policy, the prompt template, and the artifact-contract adapter may not have been enforced for CAi. Because the headline outcome lead (84.59 vs. 66.52 in Tables 1 and 2) comes exclusively from CAiMD, this asymmetry is load-bearing. Please clarify whether CAi was run under the identical live conditions as the baselines, and if it was not, provide a rerun under those conditions or explicitly present the CAiMD head-to-head as a pilot rather than the primary comparison.","section":"§4.1 and Appendix B.2"},{"comment":"CAiMD is a benchmark whose 45 reference specifications were manually authored by the same team that designed CAi, and the semantic scoring relies on a blinded LLM judge with rubrics that are not included in the paper: Appendix B.1 says 'The final rubric and judge prompts will be released with the evaluation suite.' The rule-based metrics in Eqs. (10)-(13) are transparent, but the judge-dependent components (task decomposition, semantic trajectory precision, hallucination-free interpretation, and others) are not auditable. Because the benchmark is self-authored and the judge prompts are unavailable, the 18.07-point lead over Codex could partly reflect rubric or design alignment, or judge bias favoring CAi's reporting style. Please release the full benchmark specifications and judge prompts, report judge agreement, and provide evidence that the reference specifications were not shaped by CAi's design choices; otherwise the CAiMD comparison should be treated as a self-assessment rather than the primary evidence of superiority.","section":"§4.1, Appendix A.3, Appendix B.1"},{"comment":"The paper's title and abstract claim 'reducing operational workload,' and the introduction frames the contribution in terms of researchers needing to coordinate steps and integrate evidence. However, Table 2 states: 'Efficiency is excluded because matched manual-time measurements are not yet available.' No measurement of workload, wall-clock time, human effort, number of interventions, or cost appears anywhere in the manuscript. Thus the central stated benefit is not empirically supported. Please either add workload or efficiency measurements on at least a subset of tasks, or revise the title and abstract to claim 'workflow reliability' or 'evidence grounding' rather than workload reduction.","section":"§4.3, Table 2 note; title"},{"comment":"The results are reported as single comparisons without confidence intervals, significance tests, or multiple seeds. For example, the SMDD improvements in Table 3 (20.0% to 24.0% on 25 instances; 0.0% to 1.7% on 60 instances) correspond to one additional solved instance per task type, and the MolBench results in Table 5 are reported without variance. Given that the headline CAiMD lead is already sensitive to run-condition asymmetries, the absence of uncertainty quantification makes it difficult to assess whether the observed gaps are stable. Please report multiple seeds or bootstrapped confidence intervals for at least the main CAiMD and external benchmark comparisons, or explicitly state the number of runs and the variance.","section":"Tables 1, 3, 4, 5"}],"minor_comments":[{"comment":"The notation TA_i and TP_i in Eq. (10) is described as 'rule-based trajectory accuracy and trajectory precision proxies,' but Table 1 groups TA and TP under 'Workflow' without defining them at first use in the main text; please align the notation and define all nine metric abbreviations when they are introduced.","section":"Appendix B.1"},{"comment":"Several references contain the placeholder text 'Please verify full author list before final submission'; this must be resolved before publication.","section":"References [8], [10], [30]"},{"comment":"The figure caption and in-text citations use labels (a) through (e), but the figure image appears to repeat panel labels A and B; please renumber the panels consistently with the text.","section":"Figure 2"},{"comment":"The phrase 'all agents use the same ... prompts' should specify that this refers to task inputs rather than system prompts, since each baseline retains its native control logic and the adapters in Appendix B.2 necessarily use different system-level instructions.","section":"§4.1 first paragraph"},{"comment":"All systems score exactly 95.56% on intent understanding, which suggests a ceiling effect; please comment on whether this metric is too coarse to differentiate intent-recovery capability, or whether the near-constant score is an artifact of the scoring definition.","section":"§4.2 and Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper has a clear architecture and some independent external validation, but the CAiMD head-to-head is seriously weakened by the self-authored benchmark and by the appendix statement that CAi was a pre-existing reference run. I would not reject outright because the MolBench and LIDDiA results give some independent evidence of transfer, and the issues are addressable through rerunning CAi under live conditions and releasing the benchmark materials. However, if the run-condition discrepancy cannot be resolved, the headline claim should be withdrawn or substantially softened. The title's workload-reduction claim is also unsupported as written. I recommend major revision focused on the CAiMD protocol and the workload claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful systems paper with a real architectural idea and a reasonable external-benchmark story, but its headline result rests on a comparison that is not yet auditable, and the title's workload claim is never actually measured.\n\nWhat is new: framing molecular design as intent-to-evidence workflow execution, the three-layer division, and CAiMD as a 45-task benchmark. The paper also does the right thing by running external benchmarks (SMDD, LIDDiA, MolBench) under native protocols with shared backbones. That gives independent support for the claim that CAi coordinates molecular tools effectively, and the MolBench gains are large (MS-1 accuracy from 0.18 to 0.96). Credit where due.\n\nThe soft spots, in order of weight. First, the CAiMD head-to-head: Section 4.1 says all agents use the same endpoint, prompts, inputs, workspaces, and 2,400s budget, but Appendix B.2 says CAi was a previously completed reference run normalized and scored under the same rubric, not a sixth queue in the baseline supervisor. Those statements are not obviously compatible. If CAi's trajectory was produced before the live harness existed, then the uniform budget and retry policy may not actually apply to it. Since the 84.59 vs 66.52 lead comes exclusively from CAiMD, this asymmetry is load-bearing. The paper needs to state plainly how CAi's run was produced: same clock, same failure handling, same adapter, no manual intervention. Without that, the headline comparison is not trustworthy.\n\nSecond, the 'workload reduction' in the title is unmeasured. Table 2 explicitly excludes efficiency because matched manual-time measurements are not yet available. So the paper's central promise is asserted but not tested. That is a framing mismatch, not a fatal flaw, but it should be fixed or the title narrowed.\n\nThird, the CAiMD benchmark is authored by the same team and not yet released. That is a circularity burden, though the external benchmarks partially offset it. Release of the reference specs and judge prompts should be a condition of acceptance, not a future promise. Fourth, minor: no error bars or significance tests on the 45-task comparison; with one run per system, the 18-point gap could still be real, but we cannot tell.\n\nThe central idea survives: the architecture is sensible, the external transfer results are independently informative, and the analysis of where CAi fails (long-horizon coverage, structure-dependent docking) is honest. This deserves a serious referee. I would send it to review with a request to clarify the CAiMD run provenance, measure or drop the workload claim, and release the benchmark artifacts.","headline":"Useful agent system with a sensible architecture and credible external benchmarks, but the headline 18-point CAiMD lead rests on a run whose provenance is not yet auditable, and the workload claim is unmeasured.","tokens_in":16675,"tokens_out":2162,"would_cite":false,"duration_ms":22246,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A three-layer intent-driven agent turns open-ended molecular design requests into traceable candidate-level evidence, scoring 84.59 on 45 curated tasks and beating five baselines.","keywords":["intent-to-evidence","molecular design workflow","AI agent","three-layer architecture","CAiMD benchmark","drug discovery","evidence grounding","LLM agent"],"falsifier":"Take the released CAiMD reference specifications and judge prompts, then have an independent panel of medicinal chemists re-score every agent's final reports with the agent identities removed; separately, have the panel author 45 fresh design requests that CAi's designers never see. If CAi's outcome-score lead over the strongest general-purpose baseline shrinks to statistical insignificance or reverses on the fresh requests, the performance claim is an artifact of the benchmark rather than a consequence of the three-layer architecture.","tokens_in":15565,"feed_emoji":"🧬","tokens_out":11251,"duration_ms":102731,"temperature":0.7,"pith_summary":"Early-stage molecular design is an iterative loop of proposal, assessment, filtering, ranking, and refinement, and this paper argues that the real bottleneck is not molecule generation but translating broad research intent into an executable, adaptive, traceable workflow. The central assertion is that CAi Copilot, a three-layer agent separating intent grounding, run-time control, and a molecular execution substrate, does this reliably: on 45 curated tasks it achieves an outcome score of 84.59, exceeding the next-best baseline by 18.07 points, with its largest edge in producing valid molecular outputs from tool calls. The authors also claim the benefit transfers to external benchmarks, SMDD-Bench, LIDDiA, and MolBench, on several subtasks under a shared backbone, while acknowledging that long-horizon integrated workflows and receptor-dependent docking remain weak. A sympathetic reader would care because the paper reframes the evaluation target: instead of scoring a molecule set, it scores a workflow record that connects each rank decision to candidate-level evidence an expert can inspect.","feed_headline":"Intent-driven agent scores 84.59 on 45 molecular design tasks","feed_subtitle":"A three-layer workflow turns broad research goals into inspectable evidence, beating five baseline agents by 18 points.","key_machinery":"The load-bearing object is the combined loop of three layers and a provenance-linked evidence record. The Research Interface Layer builds the initial plan $P_0 = (V_0, D_0)$, where $V_0$ is a set of steps $v_j = (g_j, u_j, x_j, p_j, e_j)$ carrying a local goal, a molecular operation, required inputs, input needs, and expected evidence, and $D_0$ records step dependencies. The Agent Reasoning Layer maintains the trajectory $h_t$ and evidence state $E_t$, selects $a_t \\sim \\pi(\\cdot \\mid h_t, P_t, E_t)$, executes it, updates $E_{t+1} = E_t \\oplus \\mathrm{Extract}(o_{t+1}, a_t)$, then computes the evidence gap $\\Delta_{t+1} = \\mathrm{Req}(\\Gamma) \\setminus \\mathrm{Cov}(E_{t+1})$ and revises $P_t$ accordingly. The Execution Substrate gives every tool a contract $\\kappa_k = (X_k, Y_k, \\mathrm{Pre}_k, \\mathrm{Exec}_k)$ that exposes accepted inputs, outputs, prerequisites, and invocation, and returns a standard observation with status, artifacts, warnings, and errors. The final output is $Y = (E, \\tau, L)$ with $E = \\{(m_i, z_i, r_i, \\rho_i)\\}$: candidate molecule, computed evidence, rank, and a provenance link $\\rho_i$ to the supporting action, plus the trajectory $\\tau$ and recorded limits $L$. This machinery carries the argument because the evidence gap is what makes plan revision reactive, and the provenance link is what makes every reported result traceable to a specific tool run.","core_discovery":"The discovery the paper puts forward is architectural: reliable intent-to-evidence molecular design comes from separating three concerns, grounding the intent, controlling the run from interim results, and executing molecular computation through formal tool contracts, rather than from any single generator, planner, or tool. CAi's Research Interface Layer converts a design request into a plan of steps with explicit dependencies and expected evidence; the Agent Reasoning Layer executes each step, merges observations into candidate records, recomputes the remaining evidence gap $\\Delta = \\mathrm{Req}(\\Gamma) \\setminus \\mathrm{Cov}(E)$, and updates the plan; and the Execution Substrate runs molecular operations behind uniform contracts that expose inputs, outputs, prerequisites, and invocation, returning standardized observations. The paper's evidence is the 45-task CAiMD comparison, where CAi reaches 84.59 on objective satisfaction versus 66.52 for the strongest baseline, alongside external benchmarks showing transfer on executable screening, editing, and optimization while exposing limits in long-horizon execution and structure-dependent workflows. The target-specific JAK1/JAK2 case study is offered as a demonstration that the workflow prioritizes the requested predicted-affinity objective (median Vina score of -8.706 kcal/mol, with 18 of 24 selected candidates below the -8.0 threshold) while preserving explicit evidence on selectivity and property trade-offs.","pith_inferences":["The 45 CAiMD reference specifications were authored by the same team that built CAi and are scored partly by a blinded LLM judge whose prompts are slated for later release; until an independent panel re-scores the runs blind to agent identity, the 18.07-point outcome gap should be read as benchmark-relative rather than a universal capability claim.","The evidence-gap control loop $\\Delta = \\mathrm{Req}(\\Gamma) \\setminus \\mathrm{Cov}(E)$ is a generic primitive, so a natural test is to port CAi's three-layer separation to other closed-loop scientific workflows, such as materials optimization or retrosynthesis, and check whether the same execution-stage gains appear where interim observations change later steps.","The paper's own diagnosis implies a cheap ablation: freeze the plan after the Research Interface Layer so the Agent Reasoning Layer cannot revise it, and measure the drop in valid-tool-output rate; if the drop is small, most of CAi's gain comes from tool contracts and the execution substrate rather than from adaptive reasoning.","The scaffold-based case comparison of 18/24 threshold-passing candidates versus 7/80 raw baseline molecules reflects workflow-level prioritization, not controlled generator superiority; a testable extension is whether that selected portfolio improves an external oracle's true-positive rate in a prospective docking or assay setting."],"forward_implications":["On the 45 CAiMD tasks, CAi's valid-tool-output rate of 67.49% versus 34.91% for the strongest baseline implies that the execution stage, turning planned operations into usable molecular artifacts, is the main bottleneck in agentic molecular design, not intent understanding or task decomposition.","With a shared backbone, CAi's SMDD-Bench success rates improve from 20.0% to 24.0% on 2D Pharmacophore Identification and from 50.0% to 63.3% on Lead Optimization, indicating that the same workflow control transfers to independent task evaluators.","On LIDDiA's 30 targets, CAi fills 39.2 of 50 candidate slots with valid molecules versus 14.4 for the reference agent (78.5% versus 28.7% valid), suggesting its advantage is broad candidate coverage and constraint satisfaction rather than uniform per-molecule quality.","On MolBench's 190 examples, CAi raises MS-1 screening accuracy from 0.1800 to 0.9600 and overall editing correctness from 0.8718 to 0.9744, while the reference agent remains stronger on receptor-dependent virtual screening and scaffold preservation, a trade-off the paper attributes to limited cross-environment structural visibility.","Because strict success falls most sharply on integrated workflows while objective satisfaction stays stable, CAi accumulates grounded partial evidence even in long trajectories but does not yet complete every dependency; long-horizon execution is the paper's own stated remaining limit.","The target-specific case shows CAi's output is a selected portfolio rather than raw generator outputs, so the threshold-pass comparison (18 of 24 versus 7 of 80) mainly demonstrates workflow-level prioritization and leaves experimental activity outside the agent's claims."],"supporting_citations":[{"why":"Supplies the chemistry-tool agent used as the domain-specialized baseline in the CAiMD comparison.","marker":"[3]"},{"why":"Supplies the biomedical assistant baseline whose end-to-end workflow performance is compared on the 45 tasks.","marker":"[17]"},{"why":"Supplies the modular drug-discovery agent baseline that represents a predefined multi-step pipeline.","marker":"[22]"},{"why":"Supplies the general-purpose coding-agent baseline that achieves the second-best outcome score of 66.52 that CAi must exceed.","marker":"[23]"},{"why":"Supplies the general-purpose tool-use baseline whose strong task decomposition still yields the lowest overall score, anchoring the coordination argument.","marker":"[21]"},{"why":"Supplies the SMDD-Bench external benchmark with official evaluators used to test transfer of CAi's workflow control.","marker":"[15]"},{"why":"Supplies the LIDDiA language-based drug discovery agent and its native evaluation protocol for the budget-matched comparison.","marker":"[1]"},{"why":"Supplies the MolBench benchmark and the MolClaw reference agent whose released evaluators measure CAi's screening, editing, and optimization.","marker":"[31]"}],"fun_headline_variants":["CAi Copilot's three-layer agent tops molecular design with 84.59","Intent-driven CAi tops molecular design by 18.07 points on 45 tasks","CAi Copilot: Reducing molecular design workload with intent-driven agent","Three-layer CAi Copilot scores 84.59 on 45 molecular design tasks","CAi Copilot maps intent to evidence across 45 molecular design tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 45-task CAiMD benchmark, whose reference specifications were manually authored by the same team that built CAi and whose semantic scores come from a blinded LLM judge with prompts slated for later release, validly represents real expert molecular-design workflows and is not biased toward CAi's reporting style; if that premise fails, the 18.07-point outcome lead over the next-best agent would not transfer to uncurated expert requests.","fun_headline_variants_meta":{"raw":{"variants":["CAi Copilot's three-layer agent tops molecular design with 84.59","Intent-driven CAi tops molecular design by 18.07 points on 45 tasks","CAi Copilot: Reducing molecular design workload with intent-driven agent","Three-layer CAi Copilot scores 84.59 on 45 molecular design tasks","CAi Copilot maps intent to evidence across 45 molecular design tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000691,"raw_usage":{"total_tokens":3185,"prompt_tokens":1058,"completion_tokens":2127,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":674,"completion_tokens_details":{"reasoning_tokens":2024}},"tokens_in":674,"tokens_out":2127,"duration_ms":14019,"temperature":1.0,"reasoning_tokens":2024,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:32:43.107326+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the released CAiMD reference specifications and judge prompts, then have an independent panel of medicinal chemists re-score every agent's final reports with the agent identities removed; separately, have the panel author 45 fresh design requests that CAi's designers never see. If CAi's outcome-score lead over the strongest general-purpose baseline shrinks to statistical insignificance or reverses on the fresh requests, the performance claim is an artifact of the benchmark rather than a consequence of the three-layer architecture.","supporting_citations":[{"cited_title":"Aluru, Achuth Chandrasekhar, and Amir Barati Farimani","cited_arxiv_id":null,"evidence_quote":"Supplies the modular drug-discovery agent baseline that represents a predefined multi-step pipeline."},{"cited_title":"Addendum to openai o3 and o4-mini system card: Codex","cited_arxiv_id":null,"evidence_quote":"Supplies the general-purpose coding-agent baseline that achieves the second-best outcome score of 66.52 that CAi must exceed."},{"cited_title":"Hermes agent","cited_arxiv_id":null,"evidence_quote":"Supplies the general-purpose tool-use baseline whose strong task decomposition still yields the lowest overall score, anchoring the coordination argument."},{"cited_title":"SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?","cited_arxiv_id":"2605.21740","evidence_quote":"Supplies the SMDD-Bench external benchmark with official evaluators used to test transfer of CAi's workflow control."}],"review_version":1}