{"id":"5e1a8c0a-b5b7-46b1-9b2b-b8830294a1a6","arxiv_id":"2412.15284","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Guided by known fragments and bond connectivity, a modified single-atom R1 search can assemble small-molecule crystal structures faster than unguided searching.","lead":"This paper shows how knowing the pieces of a molecule, and how they connect, can help a computer assemble its crystal structure from X-ray data. It recommends a faster search strategy and only uses a slower fragment-fitting method when data quality is poor.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The low-resolution branch of the recommended strategy is supported by a single unreported truncation test: at 1.2 Å the pR1 orientation search is asserted to succeed, but no R1-map ranking or failure statistics are shown.","rationale":"The reader's conditional verdict is appropriate. I agree with the reader that the deepest-hole heuristic is the weakest foundational assumption, but I sharpen the concern to the low-resolution regime because that is where the central recommendation makes its strongest and least evidenced claim. The high-resolution workflow is better supported: samples 1 and 2 show connectivity-guided completion, and sample 3 gives a direct time/quality comparison, although the sample-3 comparison includes manual correction of atomic types that should be disclosed more explicitly when comparing guided versus unguided quality. The proposed test would settle whether the low-resolution pR1 branch is reliable; if the true orientation is not rank-1 at 1.2 Å, the recommendation is unsupported. Since the paper is a case-study demonstration, this gap does not warrant rejection, but it does require explicit validation before the strategy is applied.","tokens_in":9287,"tokens_out":6328,"duration_ms":61022,"concrete_test":"Re-run the sample-1 workflow on the same data truncated to 1.2, 1.5, and 1.8 Å. For each resolution, compute the free-standing benzene-star pR1 orientation map as specified in SI S4 (5° grid, then five halving steps), and report the R1 values and the non-equivalent orientation ranking. Determine whether the true orientation is ranked first; if it is not, the deepest-hole heuristic fails at that resolution. Then run the full guided workflow only at resolutions where the true orientation ranks first and compare completed models against the known structure using the SI S3 criterion. Report the same ranking for at least five unrelated CCDC structures with known fragments to obtain a success rate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central conclusion (Section 5) makes a sharp conditional recommendation: 'when the data resolution is low, it is necessary to use the pR1 method to assemble known fragments in the first step.' The only support for this branch is the sentence: 'For sample 1, when data resolution is truncated to 1.2 Å, the sR1 method alone can no longer solve the structure, but with the pR1 method’s help to orient and position four benzene-star fragments, the sR1 method can then complete the model.' No truncated-data R1 values, no orientation rankings, no timings, and no success metric are reported. The pR1 pipeline itself (SI S4) rests on the unproved 'general hypothesis' that the deepest hole of a pR1 map in 6-D orientation-location space finds the missing fragment, and on the further assumption that the 6-D search can be split into two independent 3-D searches. At low resolution the number of reflections shrinks, so the orientation R1 map is expected to have shallower false minima; whether the true orientation survives as the deepest hole is exactly what the truncated sample-1 statement asserts but does not demonstrate. Because the central workflow recommendation inherits this assertion, the low-resolution branch is currently a single anecdote, not a validated rule.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports three case studies in which pre-known chemical information (fragment identity, connectivity, approximate bond lengths) is used to guide partial-structure R1 (pR1) and single-atom R1 (sR1) calculations for solving small-molecule crystal structures. Sample 1 is a light-atom-only structure assembled by first orienting and positioning four benzene-star fragments with pR1 and then completing the model with connectivity-guided sR1. Sample 2 uses the same pR1-plus-connectivity-guided-sR1 approach for a second light-atom structure. Sample 3 is a heavy-atom-containing structure assembled by first determining the heavy-atom substructure with normal sR1 and then completing the remaining side chains with connectivity-guided sR1. The paper concludes with two workflow recommendations: for light-atom structures, use normal sR1 for the framework and connectivity-guided sR1 to complete it, except at low resolution where pR1 is said to be necessary; for heavy-atom structures, use normal sR1 for the heavy-atom substructure and then connectivity-guided sR1.","tokens_in":9492,"tokens_out":3799,"duration_ms":35862,"significance":"If the proposed strategies are reliable, the paper offers practical guidance for using pR1/sR1 in small-molecule structure assembly: it demonstrates that pre-known fragments and connectivity can be exploited in a stepwise, planned manner, and it provides timing comparisons suggesting that connectivity-guided sR1 is faster than normal sR1 in the cases shown. The manuscript is transparent about its implementation choices, such as intensity sharpening and the specific grid parameters, and it explicitly names the central algorithmic assumption (deepest-hole hypothesis in SI S4). The main limitations are that the evidence consists of three single successful runs without deposited raw data or code, the success metric is purely positional, and the low-resolution branch of the central recommendation rests on a single unreported truncation experiment. These limitations currently make the generalized conclusions stronger than the evidence supports.","major_comments":[{"comment":"The low-resolution branch of conclusion (1) is supported only by the sentence: 'For sample 1, when data resolution is truncated to 1.2 Å, the sR1 method alone can no longer solve the structure, but with the pR1 method’s help to orient and position four benzene-star fragments, the sR1 method can then complete the model.' No truncated-data R1-map values, orientation rankings, location-search rankings, timings, or success metrics are reported, and no failure statistics or additional truncation levels are shown. Because the paper states that pR1 is 'necessary' in this regime, this claim is currently a single anecdote. The authors should report the quantitative truncated experiment, including the rank of the true orientation and location for each fragment and the final model accuracy, and ideally provide more than one truncated-data trial.","section":"Section 5"},{"comment":"The entire pR1 procedure rests on the 'general hypothesis' that the deepest hole of a pR1 map in a 6-dimensional orientation-location space determines the missing fragment, together with the further assumption that the 6-D search can be split into two independent 3-D searches. The paper does not test these assumptions in the case studies; for each sample the text merely states that the pR1 method 'correctly oriented' or 'correctly positioned' the fragment. Since the recommended strategies inherit these assumptions, the paper should report, for each fragment, the rank of the true orientation among the candidate orientations and the rank of the true location among the candidate locations. Without this evidence, the workflow recommendation cannot be distinguished from a heuristic that worked in the demonstrated examples.","section":"SI S4"},{"comment":"The reported success metric counts atoms whose positions are within 0.5 Å regardless of atom type, yet sample 3's initial normal-sR1 run produced ten type misassignments (4 I as Mo, 3 Mo as I, one Mo as S, one S as I, one N as S) that were then 'corrected' manually. Because the model comparison in S3 explicitly disregards atom types, the claim that the strategy 'results in better quality of a model' is not supported with respect to chemical identity. The authors should report type-aware match counts and specify whether the type corrections are part of the algorithm or a manual crystallographer intervention, since that distinction materially affects the reproducibility of the claimed strategy.","section":"Section 4 and S3"},{"comment":"All conclusions are derived from a single successful run for each of three structures under one implementation, with no raw data or code deposited and timings reported as single measurements on one device. The paper generalizes to a 'usual strategy' and a 'correct strategy,' which requires stronger evidence than three anecdotal successes. The authors should either provide additional runs and ideally further test structures, or explicitly scope the conclusions to the demonstrated cases and present the work as a proof-of-concept rather than a validated general protocol.","section":"Sections 2, 4, 5"}],"minor_comments":[{"comment":"There is a typographical error: 'the pR1are two new model-searching techniques' should read 'the pR1 are two new model-searching techniques'.","section":"Introduction"},{"comment":"The radiation wavelengths are given as λ = 0.71073 nm and λ = 1.54178 nm, but Mo Kα is 0.71073 Å and Cu Kα is 1.54178 Å; as written these values are off by a factor of ten.","section":"SI S1"},{"comment":"The phrase 'and then uses the connectivity-guided sR1 method' in the Abstract and in Section 5 should be 'and then use' to agree with the subject 'strategy'.","section":"Abstract and Section 5"},{"comment":"The description of the model-comparison algorithm says the shift-and-overlap procedure is repeated for all atoms in both models, but it is not explicitly stated whether the inversion of model B is tried in both the original and shifted frames; please clarify.","section":"S3"},{"comment":"The figure captions summarize the steps, but the actual figures are not included in the text supplied for review; the published version should ensure that the step numbering in the figures matches the step labels used in the text.","section":"Figures 1–3"}],"recommendation":"major_revision","confidential_remarks":"The paper's core 'general hypothesis' is inherited from the author's own prior publication, and this manuscript does not provide independent validation of that hypothesis; the reported cases are all successful outcomes, with no failure cases or quantitative rankings of the search results. For a journal-level assessment, the truncated 1.2 Å experiment must be fully reported before the 'necessary' claim in Section 5 can be accepted. I would also like to see at least a minimal data/code availability statement, as the case-study claims are not reproducible without the implementation details and raw reflection data."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on Zhang's sR1/pR1 case studies. The genuinely new pieces are the connectivity-guided sR1 completion step, the Wilson-sharpening tweak that cuts ghost atoms, and an explicit decision rule for when to use pR1 versus sR1. As a demonstration, the paper works: all three structures assemble correctly, the guided runs are faster than unguided sR1 on sample 3 (1310 s vs 3840 s, 5 vs 18 misplaced atoms), and the author is honest that pR1 is slow and should be avoided when sR1 can do the job. The writing is clear, and the SI gives enough algorithmic detail to reproduce the method.\n\nThe soft spots are real but proportionate. The central strategy recommendations rest on only three structures, and the low-resolution branch rests on a single unreported truncation test: the paper says that at 1.2 Å sample 1's sR1 fails but pR1-assisted sR1 succeeds, with no R1 values, orientation rankings, or timings shown. That's a one-line anecdote carrying a general claim. The 'deepest hole' heuristic from the author's prior paper is also unproven, and the whole method is validated against known structures rather than blind tests. There's no comparison with SHELX or other standard small-molecule software, so the practical wins over existing tools are not established. The 0.5 Å match tolerance is coarse—it treats a model as correct if every atom is within half an angstrom, which is a weak standard for a structure solution.\n\nThat said, these are not fatal flaws for a methods demonstration. The paper doesn't claim a new formalism; it claims a smarter way to use one. The author cites his own prior work appropriately, and the experimental setup is transparent about what was done and what was tweaked. The main problem is overgeneralization from a small sample, not sloppiness or circularity.\n\nWho is this for? Crystallographers who work with difficult small-molecule structures where standard direct methods struggle. I'd send it to peer review, but conditionally: require release of the data and code, report the truncated-data pR1 results with proper R1-map rankings and failure statistics, and ideally add a couple more cases or a blind test. Without those, the 'necessary' and 'correct' strategy language should be softened to 'suggested'.\n\nI wouldn't cite it in my own work yet, but I'd put it in front of a reading group for discussion.","headline":"A clear, honest case study of guided sR1/pR1 assembly with a genuinely useful connectivity-guided completion step; the general strategy claims outrun the evidence, but the paper deserves peer review with conditions.","tokens_in":10071,"tokens_out":2338,"would_cite":false,"duration_ms":20938,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that small-molecule crystal structures can be assembled faster and more accurately when the single-atom R1 search is guided by pre-known connectivity and fragments, with partial-structure R1 reserved for low-resolution…","keywords":["partial structure R1","single-atom R1","molecular replacement","low data resolution","connectivity-guided search","small-molecule crystal structure","fragment assembly","heavy-atom substructure"],"falsifier":"On a previously unsolved light-atom structure with two identical known fragments and 1.2 Å data, place the first fragment by the deepest pR1 hole, define the pR1 for the second fragment, and compare the deepest hole's predicted position with the true position from a later full refinement; if the true position is not the deepest hole, the central heuristic is falsified.","tokens_in":9015,"feed_emoji":"🧩","tokens_out":7026,"duration_ms":55821,"temperature":0.7,"pith_summary":"This paper identifies practical strategies for assembling small-molecule crystal structures using the R1 search family. For light-atom-only crystals, the recommended route is normal single-atom R1 to build a framework, then connectivity-guided single-atom R1 to finish the model; only when data resolution is low should partial-structure R1 place known fragments first. For heavy-atom-containing crystals, the paper recommends normal single-atom R1 for the heavy-atom substructure followed by connectivity-guided single-atom R1. On three test structures the completed models place essentially all atoms within 0.5 Å of the correct positions, and the heavy-atom example runs in 1310 seconds with 5 misplaced atoms versus 3840 seconds and 18 misplaced atoms for the unguided route. The practical upshot is that pre-knowledge turns a blind search into an orderly, planned assembly.","feed_headline":"Pre-known fragments speed up small-molecule crystal assembly","feed_subtitle":"A guided single-atom search cuts computer time and misplaced atoms; fragment search only for low-resolution data.","key_machinery":"The load-bearing object is the connectivity-guided single-atom R1 search. It works by taking a known atom A and a target atom Q expected to bond to A at distance r, then testing only grid points inside a spherical shell of radii r−0.5 Å to r+0.5 Å around A before running the usual R1 calculation; this converts a whole-cell search into a one-bond extension. The companion pR1 method searches for a known fragment by finding the deepest hole of a partial-structure R1 map in a 6-dimensional orientation-location space, split into separate 3-dimensional orientation and location searches to save time. The paper also feeds sharpened intensities (Fo2 multiplied by exp(2Bs2), with B from a Wilson plot) into both methods, which reduces the number of ghost atoms. These pieces turn pre-known connectivity, fragments, and bond lengths into constraints that make each search step small and targeted.","core_discovery":"The paper's central claim is that the two R1 search modes should be combined according to crystal type: start with the normal single-atom R1 (sR1) to establish a framework or heavy-atom substructure, then switch to the connectivity-guided sR1, in which the search for a new atom is restricted to a spherical shell around a known bonding partner at an approximate bond length. The partial-structure R1 (pR1), which orients and positions entire known fragments by locating the deepest hole of an R1 map in a six-dimensional orientation-location space, is reserved for low-resolution cases where sR1 alone fails. The paper shows on sample 1 that when the resolution is truncated to 1.2 Å the normal sR1 can no longer solve the structure, while pR1-placed benzene-star fragments let sR1 complete it; on sample 3, the recommended strategy finishes in 1310 seconds with only 5 misplaced atoms, against 3840 seconds and 18 misplaced atoms for the unguided normal sR1 run. These results are offered as a workflow, not as a claim that pR1 is generally preferable.","pith_inferences":["At full resolution, unguided sR1 already solved sample 1 faster than the pR1-assisted workflow, so the practical value of pre-knowledge may lie mainly in low-resolution rescue and in targeting stubborn atoms rather than in speed at high resolution.","The 0.5 Å model-comparison metric ignores atom types, so reported success can coexist with type misassignments (sample 3 required manual correction of four I/Mo swaps); an automated pipeline would need a type-assignment step.","If the deepest-hole heuristic transfers, the same two-step pR1-to-connectivity-guided-sR1 plan could be tested on electron-diffraction data, where low resolution is common and fragment placement may become the standard first move.","A direct stress test would vary the assumed bond length r and shell width Δr to measure how sensitive the connectivity-guided sR1 is to imperfect pre-knowledge."],"forward_implications":["For light-atom-only structures at ordinary resolution, normal sR1 followed by connectivity-guided sR1 completes the model faster than normal sR1 alone for the finishing step (446 s vs 710 s on sample 1) and can target specific missing atoms.","At low resolution (sample 1 truncated to 1.2 Å), the normal sR1 alone can no longer solve the structure, but pR1 placement of known fragments makes completion possible.","For heavy-atom-containing structures, the normal-sR1-then-connectivity-guided-sR1 strategy cut total time from 3840 s to 1310 s and misplaced atoms from 18 to 5 on sample 3.","pR1 should be avoided whenever sR1 can do the job, because orienting and positioning fragments with pR1 took 780 s and 3700 s respectively for sample 1, making the full pR1-assisted route slower than unguided sR1 at full resolution."],"supporting_citations":[{"why":"Introduces the sR1 and pR1 methods and supplies the implementation used throughout; all three case studies build directly on it.","marker":"Zhang & Donahue, 2024"},{"why":"Establishes the rotation-function approach to using known fragments, the conceptual ancestor of pR1 orientation search.","marker":"Rossmann & Blow, 1962"},{"why":"Establishes the translation-function step that pR1 location search parallels.","marker":"Rossmann et al., 1964"},{"why":"Defines the LLGI target used in molecular replacement, the standard model-search baseline the paper compares against.","marker":"McCoy et al., 2007"},{"why":"Formulates likelihood-based molecular replacement and explains why pre-knowledge placement is effective.","marker":"Read & McCoy, 2016"},{"why":"Evaluates molecular replacement for small molecules at reduced resolution, the context that motivates the low-resolution pR1 recommendation.","marker":"Gorelik et al., 2023"}],"fun_headline_variants":["Guided single-atom search reduces misplaced atoms","Reserve fragment R1 for low-resolution data","Two-step R1 strategy speeds crystal structure assembly","Start with single-atom R1, use connectivity to finish","Pre-known fragment only when resolution is poor"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The deepest hole in the pR1 map over orientation-location space marks the true position and orientation of a missing fragment, and the split into separate three-dimensional searches preserves that signal; if either part fails on structures beyond these three cases, the recommended strategies collapse.","fun_headline_variants_meta":{"raw":{"variants":["Guided single-atom search reduces misplaced atoms","Reserve fragment R1 for low-resolution data","Two-step R1 strategy speeds crystal structure assembly","Start with single-atom R1, use connectivity to finish","Pre-known fragment only when resolution is poor"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001,"raw_usage":{"total_tokens":4268,"prompt_tokens":1016,"completion_tokens":3252,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":632,"completion_tokens_details":{"reasoning_tokens":3179}},"tokens_in":632,"tokens_out":3252,"duration_ms":22885,"temperature":1.0,"reasoning_tokens":3179,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:32:32.870469+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a previously unsolved light-atom structure with two identical known fragments and 1.2 Å data, place the first fragment by the deepest pR1 hole, define the pR1 for the second fragment, and compare the deepest hole's predicted position with the true position from a later full refinement; if the true position is not the deepest hole, the central heuristic is falsified.","supporting_citations":[],"review_version":1}