{"id":"bb05a676-7f43-4e4c-91bd-5ef363a85bf3","arxiv_id":"2602.16962","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A GPR-accelerated line integral string method makes instanton-path force evaluations nearly independent of bead count and reduces Hessian cost via selective flexible/rigid mode training.","lead":"This paper accelerates quantum tunneling rate calculations by combining a line integral string method with Gaussian process regression surrogates and selective Hessian approximations. The result is an order-of-magnitude reduction in electronic-structure evaluations for proton-transfer systems such as malonaldehyde and formic acid dimer.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Bead-independence rests on unvalidated GPR uncertainty; no calibration check against true force errors is presented.","rationale":"The reader's weakest_assumption identifies the same issue: the GPR posterior variance is treated as a calibrated measure of true force error. This is indeed the most load-bearing concern because all quantitative claims about force-evaluation counts and bead-independence derive from using that variance as a termination signal. The paper states that increasing beads from 10 to 80 costs only 8 additional evaluations, but provides no evidence that the variance threshold corresponds to actual path convergence. The supplement does not include any calibration diagnostics. The test I propose directly checks calibration by comparing predicted uncertainty with observed error and then tests whether the scaling result is robust to the stopping threshold. Since the reader already assigned a CONDITIONAL verdict due to this and other issues, and my concern is essentially the same weakest point, the appropriate verdict adjustment is UNCHANGED. The concern strengthens the conditionality but does not require moving to REJECT, because the method could still be made correct with proper calibration; it just is not demonstrated here.","tokens_in":17319,"tokens_out":1952,"duration_ms":19429,"concrete_test":"For the malonaldehyde N=80 instanton path optimized with GPR-LI-String at T=250 K, compute the standardized error zi = ||F_DFT(x_i) - F_GPR(x_i)|| / sigma_i at every bead, where sigma_i is the GPR force standard deviation (Eq. 17). If the RMS of zi deviates substantially from 1 (e.g., >2), the uncertainty is mis-calibrated. Then rerun the algorithm with a stricter threshold (e.g., reduce the allowed variance by a factor of 4) and record the number of additional force evaluations. If the bead-independence disappears (i.e., N=80 suddenly requires significantly more than 8 extra evaluations) or if the converged path changes materially, the uncertainty-based stopping rule is not reliable and the central scaling claim collapses.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 3.1 that the number of force evaluations becomes effectively independent of bead count relies entirely on using the GPR posterior force variance (Appendix B, Eq. 17) as a stopping criterion. The paper terminates path optimization when this self-reported uncertainty falls below a threshold, but never compares the predicted variance to actual force errors (e.g., the difference between GPR and DFT forces at the converged beads). If the GP hyperparameters or the active/null mode split are misspecified, the variance can be systematically overconfident (stopping prematurely on an unconverged path) or underconfident (destroying the claimed cost reduction). The stopping threshold itself is not stated, making the result irreproducible. Because the bead-independence and the 8-additional-evaluation claim are direct consequences of trusting this variance, the lack of calibration is the load-bearing weakness. No independent evidence (e.g., a scatter plot of predicted vs. actual errors) is provided, so the claim is unsupported at its foundation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a Gaussian process regression (GPR) enhanced Line Integral String (LI-String) method for ring-polymer instanton calculations of tunneling rates and tunneling splittings. The central methodological claims are: (i) using the GPR posterior force uncertainty as a stopping criterion makes the number of ab initio force evaluations effectively independent of the number of beads used to discretize the instanton path (Section 3.1 and Fig. 1); (ii) GPU-accelerated Blackbox Matrix-Matrix Multiplication (BBMM) reduces GPR hyperparameter training time by roughly an order of magnitude (Section 3.2 and Fig. 2); (iii) a selective Hessian training strategy, which separates flexible and rigid internal modes, reduces Hessian evaluation cost by 40--62.5% while keeping instanton rates within 20% of direct DFT reference rates (Section 3.3 and Tables 1--2). The method is applied to malonaldehyde and Z-3-aminopropenal for rates, and to formic acid dimer and malonaldehyde for tunneling splittings (Tables 3--4). The paper also compares cubic-spline Hessian interpolation with GPR Hessian surrogates, finding spline interpolation less robust at higher temperatures and for noisier data.","tokens_in":17479,"tokens_out":4707,"duration_ms":43985,"significance":"If the claims are substantiated, the paper would provide a practically useful acceleration of on-the-fly instanton rate and tunneling splitting calculations, which are currently expensive because each bead requires forces and Hessians. The rate benchmarks in Tables 1 and 2 report explicit percent errors against independently computed DFT reference rates, which is a strength: the surrogate rates are not evaluated against quantities used to fit the surrogate. The reported GPU speedups and selective-Hessian cost reductions are also concrete and potentially valuable. However, the central bead-independence claim rests on a single run per system and on treating the GPR posterior force variance as a calibrated stopping signal without any comparison between predicted and actual force errors. Several numerical details needed for reproducibility (GPR convergence thresholds, flexible/rigid mode assignment criteria) are not stated. The 'within 20%' accuracy claim is also not uniformly supported by the tables. The significance is therefore conditional on calibration and reproducibility checks.","major_comments":[{"comment":"The headline claim that the number of force evaluations becomes effectively independent of bead count is based on one run per system with no error bars or repeated-initialization statistics. The GPR-enhanced path search depends on the initial training set, GP hyperparameters, and the stopping threshold; a single trajectory cannot establish 'effectively independent.' The GPR force-uncertainty stopping threshold is not stated anywhere, and Eq. (12) only gives the general LI-String convergence criteria. Please report the actual threshold used for the posterior force variance, and provide statistics over multiple independent runs with different initial training configurations.","section":"Section 3.1, Fig. 1, Eq. (12)"},{"comment":"The load-bearing assumption is that the GPR posterior force variance, Eq. (17), is a faithful estimate of the true force error. The paper never compares predicted uncertainties with actual errors (e.g., |F_GPR - F_DFT| at converged beads). If the GP hyperparameters or the active/null mode decomposition are misspecified, the variance can be overconfident (premature stopping on an unconverged path) or underconfident (destroying the claimed cost reduction). The flexible/rigid mode assignment is also selected manually, with no reproducible cutoff or algorithm; the numbers of flexible modes vary (6, 8, 12, 13) across systems and temperatures. A calibration plot and a reproducible mode-selection criterion are needed before the cost claims can be considered supported.","section":"Appendix B, Eqs. (15)--(17), Section 3.1"},{"comment":"The abstract and conclusions state that tunneling rates are predicted 'within 20%' of exact values, but several entries in Tables 1 and 2 exceed 20%: Table 1, T=275 K, standard, case (a), N=40 gives 20.7%; Table 1, T=275 K, tight, case (b), N=40 gives 21.6%; Table 2, T=250 K, standard, case (b), N=320 gives 20.0%; Table 2, T=160 K, standard, case (b), N=40 gives 21.6%. Please either soften the claim to reflect the actual spread, define the statistical meaning of 'within 20%,' or explain why these specific entries are consistent with the stated accuracy claim.","section":"Tables 1 and 2, Section 4.1"},{"comment":"The reported cost reductions ('reduces the number of force evaluations by 44%, 52%, 62.5%') are not defined in terms of a quantitative cost model. For example, in Table 1 at T=275 K, case (b) uses 10 Hessian data points along 8 flexible modes and 3 Hessian data points along 13 rigid modes, while case (a) uses 10 full Hessian data points. It is unclear why adding partial Hessian evaluations (13 total vs. 10 full) reduces the number of force evaluations by 44%. If the reduction is due to the lower cost of partial Hessians for rigid modes, the cost model and the actual number of electronic-structure force evaluations should be stated explicitly.","section":"Section 4.1, Tables 1--2"}],"minor_comments":[{"comment":"The caption contains the placeholder text 'NEED to update this plot for the arxiv submission.' This must be removed before publication.","section":"Fig. 2 caption"},{"comment":"The 'within 20%' claim in the abstract should be qualified by the fact that some Table 1 and Table 2 entries exceed 20%, and the text should state whether this refers to a typical, maximum, or average error.","section":"General"},{"comment":"The text says the instanton path can be represented with 20 beads for formic acid dimer, but the tunneling-splitting tables list bead numbers 320, 640, and 1280. Please clarify the relationship between the bead count used for path optimization and the bead count used for the splitting evaluation.","section":"Tables 3 and 4"},{"comment":"For malonaldehyde, the computed splitting of 61.43 cm-1 (B3LYP/def2-TZVPP) is a factor of ~2.8 larger than the experimental value of 21.583 cm-1. Calling this 'reasonable agreement' is generous; the discussion should more explicitly quantify the deviation and its sensitivity to the DFT barrier.","section":"Section 4.2.2"},{"comment":"There are several typographical errors, e.g., 'implmentation' at the end of Section 2, 'togehter' in Section 4.1.2, and 'Z3-aminopropenal' in the Table 2 header. A careful proofread is needed.","section":"General"},{"comment":"No data or code availability statement is provided. Given the method's reliance on GP hyperparameters, convergence thresholds, and mode-assignment choices, a reproducibility statement or a documented implementation would substantially strengthen the paper.","section":"Reproducibility"}],"recommendation":"major_revision","confidential_remarks":"The paper is a follow-up to the authors' previous work and is largely incremental, but the reported accelerations are relevant to the instanton-theory community. The main concerns are calibration of the GPR uncertainty, missing thresholds and mode-selection criteria, and the 'within 20%' claim being broader than the tables support. These are fixable within the manuscript's scope, so I recommend major revision rather than rejection. I would also encourage the editor to require the authors to provide the numerical values behind Fig. 1 and the specific GPR stopping criterion, as those are central to the paper's main selling point."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper shows something potentially useful: GPR-enhanced LI-String reduces force evaluations dramatically, and selective Hessian training gives decent rates with far fewer Hessians. The rate tables, benchmarked directly against DFT, are the strongest part — explicit percent errors, two bead counts, two DFT convergence levels. That is real evidence. The tunneling splitting sections are also honest about the Z-matrix permutation issue and the spline noise problem.\n\nBut the headline claim — that force evaluations become bead-count independent — is under-supported. It rests entirely on using the GPR posterior force variance as a stopping rule. The paper never compares predicted uncertainty to actual force error (e.g., a scatter of GPR vs DFT forces at converged beads), never states the threshold, and the evidence is two single runs with no error bars. If the GP is overconfident, the path is unconverged; if underconfident, the cost advantage disappears. The stress-test note is right: the bead-independence is load-bearing and uncalibrated.\n\nThen there are consistency problems. The abstract on arXiv mentions cubic spline as efficient and lists 7,9-dinitro...; the full-text abstract and body don't. The body's spline results for aminopropenal show 60% errors, so the abstract claim is wrong. The abstract also says \"Hessian-free GPR training\" while the body uses Hessian data throughout. Section 3.2 still has the literal note \"NEED to update this plot\". The tunneling splitting tables show clear bead drift (FAD 0.021 → 0.008; malonaldehyde 6-31G* 48 → 34) with no convergence criterion. Mode classification (flexible/rigid counts change with temperature) is described qualitatively, and no hyperparameters or thresholds are given. No code or data is deposited.\n\nNone of this kills the core idea. The combination of GPR with LI-String and selective Hessian training is sensible, and the rate benchmarks are promising. But the central cost claim is not yet reproducible or calibrated.\n\nI would send it to a serious referee, with the expectation of major revision. Specifically: a calibration plot of GPR uncertainty vs true force error, the stated stopping threshold, replicates or error bars for the force-evaluation counts, convergence tests for the splittings, a corrected abstract, and code/data deposit.","headline":"Useful acceleration for instanton rates with a sensible selective-Hessian idea, but the bead-independence claim rests on an unvalidated uncertainty estimate and the manuscript shows clear signs of being unfinished.","tokens_in":18087,"tokens_out":3964,"would_cite":false,"duration_ms":36903,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Gaussian process force uncertainty makes instanton path refinement cost independent of the number of beads, while tunneling rates stay within 20% of exact.","keywords":["instanton theory","ring polymer","Gaussian process regression","tunneling rates","tunneling splitting","line integral string method","selective Hessian training","proton transfer"],"falsifier":"During a typical malonaldehyde instanton run, compute the GPR-predicted force variance at held-out beads together with the actual error |F_ML - F_DFT|. If the ratio of predicted to actual error is far from 1, or if the required number of force evaluations changes systematically with the initial bead count, the stopping criterion is not calibrated and the bead-independence claim fails.","tokens_in":17103,"feed_emoji":"⚛️","tokens_out":4856,"duration_ms":47045,"temperature":0.7,"pith_summary":"This paper is trying to establish that Gaussian process regression can remove the traditional trade-off between the number of beads in a ring polymer instanton calculation and the cost of converging the tunneling path. Using the GPR posterior force variance as a stopping criterion, the authors show that increasing the bead count from 10 to 80 requires only about 8 additional force evaluations, rather than hundreds. They further show that a selective Hessian training strategy, which treats flexible proton-coupled modes differently from rigid modes, cuts force evaluations by 40–62.5% while keeping instanton rates within 20% of the rigorous DFT result. If this works broadly, it makes high-accuracy tunneling calculations much more affordable for molecular systems.","feed_headline":"Instanton force calls drop to ~64 and stay flat as beads rise to 80","feed_subtitle":"Gaussian process uncertainty makes path refinement independent of bead number while rates stay within 20%.","key_machinery":"The central object is a Gaussian process regression surrogate of the potential energy surface that includes potential, gradient, and Hessian observations, and whose posterior force variance (the trace of the force covariance in Cartesian coordinates) serves as a convergence check for path optimization. The line integral string method provides the path optimizer, while BBMM reduces GPR hyperparameter training from cubic to quadratic complexity, with additional GPU acceleration. Selective Hessian training splits internal modes into an active subspace of flexible proton-coupled modes (modeled with GPR) and a null subspace of rigid modes (modeled with linear regression), enabling large reduction","core_discovery":"On its own terms, the paper claims that the GPR-enhanced line integral string method converges instanton paths with a number of force evaluations that is effectively independent of the number of beads: malonaldehyde needs 64 force evaluations at N=10 beads and only 8 additional at N=80, while Z-3-aminopropenal needs 70 plus 10 additional. The convergence criterion is the GPR-predicted force variance rather than a fixed bead count. For rate calculations, selective Hessian training—using only 3 beads for rigid modes and 10–20 beads for flexible modes—reduces force evaluations by 40–62.5% while keeping predicted rates within 20% of exact instanton rates, and often within 6%. GPU-accelerated Bla","pith_inferences":["If the posterior force variance is truly calibrated, the same bead-independence should extend to other chain-of-states methods, to larger molecules, and to potentials with stronger anharmonicity; a natural test is to hold the bead count fixed and verify that the stopping criterion does not depend on the initial number of beads.","The active/null mode split is selected manually with no reproducible cutoff; automating the split by ranking each mode's coupling to the path tangent would make the 40–62.5% cost savings reproducible without user judgment.","The observed dominance of electronic-structure noise suggests that combining these surrogate instanton paths with transfer-learned or higher-level corrections to the potential could push rate errors well below the current 20% envelope.","The failure of cubic splines under high noise suggests a hybrid scheme that uses GPR where local uncertainty is high and spline interpolation where data are clean could combine robustness with low cost."],"forward_implications":["Bead count can be chosen purely for discretization accuracy, since refining from N=10 to N=80 costs only about 8–10 extra electronic-structure force evaluations.","Selective Hessian training achieves tunneling rates within 20% of the rigorous instanton rate while cutting force evaluations by 40–62.5%, with errors often near 6%.","The accuracy of surrogate-predicted rates improves when the underlying DFT calculations use tighter convergence thresholds and finer grids, indicating that electronic-structure noise is a key error source.","Cubic spline interpolation is a simple and accurate alternative for Hessian fitting when noise is low, but its error grows at higher temperatures when the instanton path shortens and interpolation points cluster.","The same GPR-LI-String machinery locates tunneling-splitting instanton paths in formic acid dimer and malonaldehyde with roughly 100 potential/force evaluations, yielding splittings in reasonable agreement with experiment and high-level theory.","The paper flags that the Z-matrix descriptor is not permutation invariant, so tunneling-splitting calculations require a manual symmetry operation; a permutation-invariant descriptor would remove this step."],"fun_headline_variants":["Instanton force calls stay flat as bead count rises","Bead-independent force calls via GPR surrogates for instanton rates","Surrogate PES reduces instanton force calls regardless of bead number","GPR surrogate keeps instanton path force calls constant","Machine learning surrogates flatten instanton force evaluations"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central claim collapses if the Gaussian process uncertainty estimate is not a trustworthy measure of how wrong its force predictions really are, and the paper never compares predicted uncertainties against actual force errors.","fun_headline_variants_meta":{"raw":{"variants":["Instanton force calls stay flat as bead count rises","Bead-independent force calls via GPR surrogates for instanton rates","Surrogate PES reduces instanton force calls regardless of bead number","GPR surrogate keeps instanton path force calls constant","Machine learning surrogates flatten instanton force evaluations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000282,"raw_usage":{"total_tokens":1516,"prompt_tokens":770,"completion_tokens":746,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":662}},"tokens_in":514,"tokens_out":746,"duration_ms":7108,"temperature":1.0,"reasoning_tokens":662,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T22:22:03.561310+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"During a typical malonaldehyde instanton run, compute the GPR-predicted force variance at held-out beads together with the actual error |F_ML - F_DFT|. If the ratio of predicted to actual error is far from 1, or if the required number of force evaluations changes systematically with the initial bead count, the stopping criterion is not calibrated and the bead-independence claim fails.","supporting_citations":[],"review_version":1}