{"id":"d53583bf-7cda-475b-8f85-8612d19b17bc","arxiv_id":"2608.07454","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"SynthEx, an LLM agent that writes template-free atom-level graph edits, returns complete routes for 63.9% of 1,098 synthesis-free natural products (vs 13.8% for a near-exhaustive template planner), and blinded experts rated its key steps on par with human syntheses.","lead":"An AI system called SynthEx plans recipes for complex natural molecules by writing its own chemical reactions instead of looking them up, and expert chemists rated its key steps about as good as published human syntheses. The team released SynthAtlas, an open database of more than a thousand machine-planned routes.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 63.9% 'solved' criterion checks only leaf purchasability, not chemical viability; without a forward-chemistry check on sampled routes, the headline reach and expert-comparable claims may overstate what SynthEx can actually plan.","rationale":"The reader's weakest_assumption—that graph-level routes vetted only by an LLM critic and NameRXN may not be chemically sound—is the same concern I would identify. It is the single most load-bearing because it directly governs the paper's two headline numbers: the 63.9% solve rate and the expert-comparable conclusion. The paper is unusually honest: it discloses the lack of stereochemical verification, the self-assessed improvement loop, and the filtered expert comparison, and it releases code and data. ReactionJSON is a genuine deterministic representation, and the benchmark construction is careful. But the solve criterion in Section 4.4 is structural only, and the recognition analyses in Section 2.3 measure nameability and recoverability, not experimental viability. A forward-chemistry check on a sample of routes would settle whether the 63.9% is a chemically meaningful success rate or a graph-level upper bound. This does not change my recommendation relative to the reader: the paper should remain CONDITIONAL until such a check is reported, with claims phrased as graph-level planning rather than experimentally validated routes.","tokens_in":30810,"tokens_out":5944,"duration_ms":57332,"concrete_test":"Draw a stratified random sample of 100 of the 702 solved routes (spanning the control, complexity-dense, and large-complex subsets). For every reaction step, (1) run a forward reaction prediction model (e.g., Molecular Transformer or a modern equivalent trained on USPTO) to test whether the proposed reactants yield the proposed product, and (2) have two independent expert synthesis chemists, blinded to source, score each step for feasibility and stereochemical soundness using the same rubric as the paper's expert study. Run the identical procedure on 100 literature total-synthesis routes as a control. If the per-step acceptance rate for SynthEx routes is markedly below the literature control (e.g., <50% accepted), the 63.9% solve rate should be reframed as a graph-level reach estimate pending experimental validation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.4 defines a target as solved purely by leaf purchasability: a complete route in which every leaf InChIKey is in stock. Nothing in that criterion checks that the LLM-written ReactionJSON edits correspond to chemically viable transformations with correct stereochemical outcomes. The only external validation offered is (i) NameRXN naming, which labels a reaction as recognized without confirming it works as drawn, and (ii) the expert key-step study, which is restricted to 47 strategically congruent targets and rates individual key steps, not full routes. The paper itself states in the Discussion 'we have not verified stereochemical outcomes, and expert review surfaced occasional selectivity and feasibility errors', and Methods 4.3 concedes the Critic/Editor loop is 'an internal consistency procedure, not an external validation.' If a nontrivial fraction of the 33,145 steps or the 702 stitched routes contain invalid reactions—wrong regiochemistry, impossible stereochemical assignments, or incompatible conditions—the headline 63.9% solve rate overstates how many targets SynthEx can actually plan, and the 'comparable to human' conclusion inherits the same problem. The reach comparison with AiZynthFinder is a graph-level metric; it does not by itself establish chemical soundness.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces SynthEx, a multi-agent LLM retrosynthesis planner in which disconnections are written directly as atom-mapped graph edits (ReactionJSON) rather than selected from a fixed template library. A strategy generator proposes several high-level strategies, a route builder expands them into full routes, a critic/editor loop repairs flagged steps, and an analyst scores feasibility. The system is benchmarked on 1,098 NPAtlas natural products with no reported total syntheses, where it returns leaf-purchasable routes for 63.9% of targets versus 13.8% for a near-exhaustively resourced AiZynthFinder baseline; the strategic layer alone reaches 25.0%. Reaction-space analyses show that SynthEx's steps are poorly recovered by USPTO-derived classifiers and by RetroChimera, and that ring-forming, convergent chemistry is overrepresented relative to patent-derived corpora. In a blinded expert study, ten synthetic chemists rated 148 key steps from 47 targets on which SynthEx's strategy was judged congruent with a published synthesis; SynthEx was statistically indistinguishable from literature steps on feasibility, elegance, and overall quality, with a small but detectable deficit on strategic value. The routes are released as the open SynthAtlas resource.","tokens_in":31029,"tokens_out":4203,"duration_ms":42958,"significance":"The contribution is potentially significant. The ReactionJSON representation is a clean, falsifiable mechanism for stepping outside patent-derived reaction space, and the main reach comparison is benchmarked against an external open-source planner (AiZynthFinder) and external classifiers (NameRXN, RetroChimera), so the central comparison is not circular. The authors are also commendably explicit about what the pipeline does not establish: they state that stereochemical outcomes were not verified, that the critic/editor loop is an internal consistency procedure, and that the expert comparison is conditional on strategic congruence. The release of dated, public route predictions for unsynthesized natural products is a genuine strength with predictive value. The main weakness is that the headline 'solved' criterion is graph-level leaf purchasability, so the quantitative reach claims and the 'comparable to human' conclusion are both contingent on chemical validity that is asserted rather than demonstrated.","major_comments":[{"comment":"The solved criterion in §4.4 counts a target as solved when every leaf InChIKey is in the ZINC/eMolecules stock, with no check that the ReactionJSON graph edits correspond to chemically viable transformations with correct regio- and stereochemistry. The paper's own Discussion concedes that stereochemical outcomes were not verified and that expert review surfaced occasional selectivity and feasibility errors. Because the 63.9% headline and the 'beyond the reach of conventional planners' claim rest on this criterion, the metric should be relabeled as a graph-completion rate or supplemented with a forward-chemistry validation (e.g., a trained forward model or a blinded expert audit of complete routes, not just key steps) so that the reader can estimate what fraction of proposed steps are plausibly executable. As written, the abstract's phrase 'plans routes' overstates what the evaluation demonstrates.","section":"§4.4, §2.4"},{"comment":"The expert comparison is conditioned on 47 of 70 targets retained after an LLM judged SynthEx's strategy 'congruent' with the published synthesis, and the manuscript provides no reliability assessment of that congruence judgment. Since the same model family that generates the routes also selects the comparison set, the retained targets may be systematically easier or more similar to literature chemistry, and the expert ratings cannot speak to the full benchmark. In addition, the literature key steps are nominated by a different LLM and the SynthEx key steps are the tool's own nominations; the comparison therefore measures steps each source regards as pivotal, not a matched random sample of steps. I ask for a sensitivity analysis on all 70 targets, or a human audit of congruence on a random subset, and for the paper to state explicitly how the 47/70 selection affects generalization of the 'comparable to human' conclusion.","section":"§4.10, Appendix B"},{"comment":"The route-improvement evidence in Fig. 5 is generated and scored by agents sharing the same LLM backbone: the Critic labels blocking reactions, the Editor repairs them, and the Analyst re-scores feasibility. As the authors note, this is an internal consistency procedure rather than external validation, but the figure and the surrounding text present the declining blocking rate and the feasibility shift as evidence of quality improvement. Because shared blind spots could make the loop converge to a self-consistent but chemically invalid state, the improvement claim should either be validated by an independent mechanism (a different model family, a forward-reaction predictor, or human review of a random sample of before/after routes) or be explicitly reframed as measuring only convergence to the Critic's own criterion.","section":"§4.3, §2.5"}],"minor_comments":[{"comment":"The control subset is constructed from targets that initially failed a short template search, so the 80% solve rate under a generous budget cleanly demonstrates a budget-limited failure; however, the subset is small (n=123) and the manuscript does not report confidence intervals for the per-subset solve rates, which would help the reader judge the strength of the comparison.","section":"§4.1, §2.4"},{"comment":"The target-wise comparison of route lengths between SynthEx and the literature is not apples-to-apples, since literature routes are experimentally optimized and SynthEx routes are untested first proposals; the paper makes this point in the text, but the figure panel as displayed may still invite an unfavorable comparison that the SI D.2 joint-solved comparison addresses more fairly.","section":"Fig. 4f"},{"comment":"The abstract states that expert chemists 'engaged with them as genuine synthesis plans,' which is stronger than the evidence: the raters saw individual key steps, not complete routes, and only on the 47 strategically congruent targets. Please align the abstract's wording with the actual scope of the expert study.","section":"Abstract, §2.4"},{"comment":"The knowledge-cutoff argument is appropriately cautious, but the claim that SynthEx 'recovers' the Okaramine M route 'after the training cut-off' rests on API documentation rather than on a verified absence from post-training data; the paper's own caveat that post-training data is not disclosed should perhaps appear in the main-text discussion of this case, not only in the Methods.","section":"§4.2, §2.2.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the journal's scope and the core idea is strong, but the headline reach and comparability claims currently outrun the validation. The authors' own limitation statements are unusually candid; the revision should operationalize them by adding a forward-chemistry or human feasibility audit of complete routes, reporting the expert comparison on the full 70-target pool or an explicit sensitivity analysis, and re-benchmarking the claim language. This is a major revision rather than a rejection because the central weaknesses are identifiable and fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, you should know this paper is the first credible demonstration that an LLM agent writing template-free graph edits (ReactionJSON) can plan retrosynthetic routes to complex natural products at scale. The 63.9% vs 13.8% reach claim is believable as a graph-level comparison: the baseline is generously resourced, the strategic-layer-only number (25%) and the leaf-completion component are reported separately, and the control subset shows the baseline does work when the chemistry is simple. The reaction-space analysis is solid, with ring-forming steps at 16% versus 2.8% for a state-of-the-art single-step model, and the external classifiers (NameRXN, RetroChimera) give the distinctness claim independent support. The expert study is careful: cluster bootstrap over raters, Holm correction, leave-one-out AUC, all pre-specified. Finding that expert chemists rate SynthEx key steps comparable to published human steps on feasibility, elegance, and overall quality, with only a small strategic-value gap, is genuinely new.\n\nThe soft spots are mostly disclosed rather than hidden. The 'solved' criterion is leaf purchasability only; no forward-chemistry check verifies that the LLM-written edits are chemically viable. But the paper states outright: 'we have not verified stereochemical outcomes, and expert review surfaced occasional selectivity and feasibility errors,' and it explicitly frames the work as reach and per-step quality, not experimental feasibility. So the stress-test concern that the headline overstates chemical soundness is fair, but it points at a limitation the authors already own. The expert comparison is filtered to 47 targets where an LLM judged SynthEx's strategy congruent with the published one, which biases toward its better routes; that is real and acknowledged. The Critic/Editor loop is self-assessed with the same backbone, and the paper calls it 'an internal consistency procedure, not an external validation.' Correct. The training-cut-off argument is reasonable but not airtight, since post-training data are undisclosed; the authors say they treat it as evidence, not proof.\n\nWhat would strengthen the work: a forward-reaction filter on sampled routes, or a few lab-validated key steps. The Melonine case, where the Zhu group assesses SynthEx's alternative as more likely to succeed than their own failed attempt, is the most convincing external anchor and the right kind of evidence to seek more of.\n\nThis deserves a serious referee. It is a benchmark-scale, code-and-data-released contribution with real novelty and unusually candid limitation statements. The central reach claim holds up as graph-level planning, and the chemical-validity question is flagged as the next frontier by the authors themselves. I would take it to reading group and I would probably cite the SynthAtlas resource.","headline":"A well-scoped, honestly reported advance in LLM synthesis planning: the reach claim holds at graph level, chemical validity is unverified but the paper says so plainly.","tokens_in":31709,"tokens_out":2321,"would_cite":true,"duration_ms":21682,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A strategy-first LLM planner designs complete routes to complex natural products without a template library.","keywords":["retrosynthesis","natural product synthesis","large language models","agentic planning","template-free reaction representation","reaction space","total synthesis","multi-agent systems"],"falsifier":"Select a representative sample of SynthEx's proposed key steps, especially ring-forming and stereoselective steps, and attempt them in the laboratory under the abstract conditions the routes annotate; if a substantial fraction fail or give incorrect stereochemistry, the expert-comparable step quality and the 63.9% solve rate would not reflect experimentally viable planning.","tokens_in":30561,"feed_emoji":"🧪","tokens_out":5946,"duration_ms":50349,"temperature":0.7,"pith_summary":"The paper claims that a multi-agent LLM planner, SynthEx, can design full retrosynthetic routes to complex natural products that have no published synthesis, by writing each disconnection as an ordered list of atom-level graph edits instead of choosing from a fixed reaction template library. On a benchmark of 1,098 such targets, it returns complete routes to purchasable building blocks for 63.9% of them, while a near-exhaustively run template-based planner solves 13.8%. Blinded expert chemists rated SynthEx's key steps statistically indistinguishable from published human syntheses on feasibility, elegance, and overall quality, with only a small gap in strategic value. The authors release the routes as an open database, SynthAtlas, and treat experimental validation as the next frontier.","feed_headline":"LLM planner designs routes for 63.9% of unsynthesized molecules","feed_subtitle":"Template-free atom edits let it propose ring-forming chemistry that patent-trained planners rarely reach.","key_machinery":"ReactionJSON, a template-free reaction representation in which a retrosynthetic disconnection is an ordered list of atom-level graph-edit operations (break_bond, add_bond, change_bond_order, add/remove group, and stereochemistry operations) applied to an atom-mapped product to yield the precursors deterministically. Because a route is serialized as RouteJSON, a text object anchored on atom maps, it can be edited in place: a Critic agent simulates each reaction forward, flags 'blocking' steps, and an Editor performs surgical repairs such as reordering steps, inserting protections, or changing conditions without restarting the search. This representation is what converts the LLM from a selector among templates into a generator of novel disconnections, and it is what makes the iterative self-correction loop practical.","core_discovery":"The central claim is that a synthesis planner can operate inventively outside the reaction space defined by patent-derived corpora. SynthEx expresses every disconnection as ReactionJSON, an ordered list of atom-level graph edits applied to the atom-mapped product, so the language model generates novel expansions rather than selecting among templates. On 1,098 structurally complex natural products with no reported total synthesis, SynthEx produces complete routes to purchasable building blocks for 63.9% of targets versus 13.8% for a template-based planner run under a generous budget; the advantage widens with molecular complexity, and the proposed chemistry is more convergent and far richer in ring-forming steps (16.0% of steps, against 2.8% for a state-of-the-art corpus-trained single-step model's top-1 predictions). In blinded evaluation, expert chemists rated SynthEx's key steps comparable to published human syntheses on feasibility, elegance, and overall quality, while a small but detectable gap remained on strategic value. The paper is explicit that these are graph-level proposals: stereochemical outcomes were not verified, and the improvement loop is scored by the same class of model that performs the repairs.","pith_inferences":["If the 63.9% solve rate survives wet-lab testing, the bottleneck in automated synthesis shifts from proposing disconnections to predicting reaction conditions and stereochemical outcomes, which the paper itself flags as unverified.","Because SynthAtlas is atom-mapped and lies outside patent-derived corpora, it could become training and evaluation data for next-generation single-step retrosynthesis models, a use the paper mentions but does not itself pursue.","The small but detectable strategic-value gap suggests the most productive near-term arrangement is a chemist supplying the high-level strategy while the agent elaborates and critiques the route, a division the paper identifies as promising.","A direct testable extension is to apply the planner to targets where literature routes failed at a specific step, such as the Melonine Mannich cyclization, and see whether its alternative disconnections reproducibly avoid the documented failure mode."],"forward_implications":["If correct, retrosynthetic planning no longer needs to be confined to a fixed reaction library; an LLM that writes graph edits can propose chemistry that is rare in patent corpora but recognizable to expert chemists.","The solve-rate advantage grows with target complexity, reaching 56% on the large complex subset where the template baseline solves only 4%, so the method is most useful exactly where conventional planners fail hardest.","The released SynthAtlas corpus of 33,145 atom-mapped reactions provides a large, dated set of public predictions for molecules nobody has yet synthesized, enabling future concordance tracking as real syntheses appear.","On targets where both planners succeed, SynthEx routes are shorter than template-planner routes (median 5 steps versus 11), suggesting that its strategic disconnections simplify the search problem for downstream completion.","The expert-comparable ratings on feasibility, elegance, and overall quality imply that SynthEx's key steps could serve as credible starting points for experimental campaigns, with strategic guidance remaining the one axis where human input still adds value."],"supporting_citations":[{"why":"Supplies the open-source template-based planner used as the baseline and defines the solve criterion of complete routes to purchasable building blocks.","marker":"[66]"},{"why":"Provides the Synthelite architecture from which SynthEx descends, positioning the LLM as a search policy over a template library.","marker":"[47]"},{"why":"Supplies the state-of-the-art single-step model used to test whether SynthEx's disconnections are recoverable from corpus-trained predictions.","marker":"[24]"},{"why":"Reports the Okaramine M synthesis published after the model's training cutoff, used as a post-hoc test that SynthEx recovers expert strategic logic.","marker":"[51]"},{"why":"Reports the Melonine total synthesis whose failed biosynthetically inspired Mannich cyclization is compared directly with SynthEx's alternative route.","marker":"[57]"},{"why":"Documents the January 2025 knowledge cutoff for the language model backbone, supporting the claim that routes are not retrieved from post-cutoff literature.","marker":"[75]"}],"fun_headline_variants":["LLM planner designs routes for 63.9% of unsynthesized natural products","AI synthesis planner outdoes patent-trained tools on complex targets","Expert chemists find AI's synthesis steps comparable to human ones","Template-free AI planner beats patent-trained tools on complex syntheses","AI designs convergent routes for complex natural products, experts approve"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a route judged chemically plausible by an LLM critic and by name-recognition software is a genuine synthesis plan; because no reactions were run and stereochemical outcomes were not verified, a large experimental failure rate among proposed steps would mean the 63.9% solve rate and the expert-comparable ratings overstate what the planner can actually deliver.","fun_headline_variants_meta":{"raw":{"variants":["LLM planner designs routes for 63.9% of unsynthesized natural products","AI synthesis planner outdoes patent-trained tools on complex targets","Expert chemists find AI's synthesis steps comparable to human ones","Template-free AI planner beats patent-trained tools on complex syntheses","AI designs convergent routes for complex natural products, experts approve"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000975,"raw_usage":{"total_tokens":4212,"prompt_tokens":1086,"completion_tokens":3126,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":702,"completion_tokens_details":{"reasoning_tokens":3037}},"tokens_in":702,"tokens_out":3126,"duration_ms":20254,"temperature":1.0,"reasoning_tokens":3037,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:27:21.819316+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Select a representative sample of SynthEx's proposed key steps, especially ring-forming and stereoselective steps, and attempt them in the laboratory under the abstract conditions the routes annotate; if a substantial fraction fail or give incorrect stereochemistry, the expert-comparable step quality and the 63.9% solve rate would not reflect experimentally viable planning.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the open-source template-based planner used as the baseline and defines the solve criterion of complete routes to purchasable building blocks."},{"cited_title":"Organic Letters27(27), 7367–7371 (2025)","cited_arxiv_id":null,"evidence_quote":"Reports the Okaramine M synthesis published after the model's training cutoff, used as a post-hoc test that SynthEx recovers expert strategic logic."},{"cited_title":"Angewandte Chemie International Edition, 8101956 (2026)","cited_arxiv_id":null,"evidence_quote":"Reports the Melonine total synthesis whose failed biosynthetically inspired Mannich cyclization is compared directly with SynthEx's alternative route."},{"cited_title":"Gemini API documentation","cited_arxiv_id":null,"evidence_quote":"Documents the January 2025 knowledge cutoff for the language model backbone, supporting the claim that routes are not retrieved from post-cutoff literature."}],"review_version":2}