{"id":"86727533-d9b6-436b-bcc7-f0ce7fd688fd","arxiv_id":"2608.11421","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A second-derivative-corrected FANPT approximation gives more stable initial guesses for solving nonlinear coupled-cluster wavefunction equations, especially with large continuation steps.","lead":"This paper adds a second-derivative correction to the Flexible Ansatz N-body Perturbation Theory (FANPT), a method that follows a molecule's wavefunction along a path from a simple starting Hamiltonian to the real one. The correction keeps the wavefunction guess closer to the true solution at each step, reducing wild parameter jumps in difficult cases.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The qao=3 advantage rests on the same-λ FANCI benchmark; without convergence thresholds, root-uniqueness checks, and an explicit independence protocol, the large norm reductions in Tables 4–12 are not yet interpretable.","rationale":"The mathematical derivation of the qao=3 response equations is internally coherent: the second- and third-order constant vectors follow from repeated total differentiation of the residual condition, and the retained Hessian terms enter only through B(r) while preserving the response matrix. The overlap Hessian for the product-form coupled-cluster wavefunction is also consistent with the path-sum representation. The weak point is therefore not the algebra but the empirical quantity that carries the central claim. The paper's headline diagnostics compare FANPT guesses to converged FANCI solutions at the same λ, yet the workflow described in the FANPT methodology section makes the FANCI solution at each λ depend on the FANPT guess used to initialize it. Whether this dependence matters depends on root uniqueness and on the solver's convergence threshold, neither of which is reported. This is precisely the condition that would have to be true for all of Tables 2–12 to mean what the abstract says they mean. The reader's verdict already identifies this via the 'independently optimized FANCI parameter vector' assumption, the absence of convergence thresholds, and the lack of degeneracy or condition-number checks. My read agrees with that assessment and does not move the verdict: the paper should remain conditional on supplying the missing numerical protocols and, ideally, a benchmark that is demonstrably independent of the FANPT path. The proposed concrete test is the minimal experiment that would settle whether the large qao=3 norm reductions are a genuine continuation improvement or a comparison artifact.","tokens_in":18344,"tokens_out":13554,"duration_ms":120084,"concrete_test":"Reproduce the BeH2/CCSD(0) reduced-step run at the worst same-λ points (e.g., r=3.5, λ=0.7 and λ=1.0; r=20.0, λ=0.9) with a fixed, qao-independent benchmark: converge Eq. (9) from at least five random starting parameter vectors and from both the qao=2 and qao=3 FANPT guesses, using a tight residual threshold (e.g., ||G||_2 < 1e-10 or the FanPy/NLopt equivalent), and record all distinct converged solutions. Also report the condition number of the response matrix G_{n,k} at those points. If the optimized P_FANCI(λ) is unique and independent of the starting guess, recompute Tables 4 and 5; if the qao=3 reductions persist, the central claim stands, while if they vanish or change sign, the reported advantage is an artifact of solver path or stopping tolerance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is supported only by the same-λ parameter norm, which is defined as the distance from the FANPT-propagated guess to the converged FANCI parameter vector at the same λ. For this norm to measure guess quality, every P_FANCI(λ) must be a unique, converged solution of the projected equations (Eq. 9), obtained without being biased by the FANPT guess. The abstract asserts that the comparison uses 'independently optimized FANCI parameters at the same value of λ,' but the FANPT methodology section states that the FANPT prediction provides the initial guess for solving the FANCI equations at the next λ. The paper reports no residual convergence thresholds, no starting-point protocol for the 'independent' optimizations, no multiplicity or branch-continuity analysis, and no response-matrix condition numbers. Near the difficult BeH2 regions where qao=2 norms exceed 100 (Tables 5, 7, 10, 12), the projected equations may have multiple roots or a near-singular response matrix. If qao=2 and qao=3 trajectories converge to different roots, or if the FANCI solver stops at a loose threshold near the qao=3 guess, the reported reductions—including factors near 10^3—would not demonstrate a better initial guess. The paper honestly reports local degradations, but those do not resolve this ambiguity because the comparison target itself may differ between the two trajectories.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a second-derivative-corrected FANPT variant, denoted qao=3, which retains the overlap Hessian for nonlinear wavefunction ansätze while neglecting only third- and higher-order parameter derivatives. The authors derive the corresponding second-, third-, and fourth-order response equations, implement the overlap Hessian for coupled-cluster wavefunctions in the FanPy/FANCI framework, and test the method on LiH and on the BeH2 insertion coordinate with CCSD(0) and CCSDT(2)Q(0) wavefunctions in the STO-6G basis. The central claim is that, especially for large λ-steps, qao=3 yields same-λ parameter vectors much closer to the optimized FANCI solutions than the original qao=2 approximation, suppressing large parameter-space excursions, while leaving final energies essentially unchanged.","tokens_in":18605,"tokens_out":9727,"duration_ms":77145,"significance":"If the central claim is correct, the contribution is a useful practical refinement of FANPT: it includes the leading nonlinear overlap response without changing the response-matrix structure, and the authors are transparent about the additional O(MP^2) cost. The derivation is clear and the scaling analysis is explicit, and the paper honestly reports local degradations in addition to the favorable global trends. The main unresolved question is whether the same-λ norm is a reliable benchmark for guess quality, because the manuscript does not document the convergence thresholds, starting-point protocol, or root structure of the FANCI projected equations used as the comparison target.","major_comments":[{"comment":"The central diagnostic is the same-λ Frobenius norm between the FANPT-propagated guess and the 'independently optimized' FANCI solution, but the independence of the benchmark solutions is not documented. Since the FANPT prediction is described as providing the initial guess for solving the FANCI equations at the next λ, the reader cannot tell whether the benchmark optimizations used FANPT-independent starting points. Please report the residual convergence threshold used in the benchmark FANCI optimizations, the starting-point protocol (e.g., from RHF at every λ or from an explicitly FANPT-independent path), and a root-uniqueness or branch-continuity check for the projected equations at least at the geometries where qao=2 norms exceed 100 (r=3.0–4.0 and r=20.0). Without this information, the reductions of up to 10^3 in Tables 5, 7, 10, and 12 could reflect loose convergence or branch switching rather than improved initial guesses.","section":"Results and Discussion, BeH2; Computational Details and Model Systems"},{"comment":"The number of FANPT continuation steps is not reported consistently. Table 4 is labeled 'reduced number of FANPT continuation steps' without giving the number, while Table 6 explicitly states '10 continuation steps'; similarly, Figure 4 discusses 10-, 50-, and 100-step runs for CCSDT(2)Q(0), but the mean and standard-deviation values for the 50- and 100-step runs are not tabulated. Please state the step count in every table caption and provide a table of the Figure 4 data so that the claimed cost–quality improvements can be verified quantitatively.","section":"Computational Details and Model Systems; Tables 4–12"},{"comment":"The practical benefit of a better initial guess for BeH2 is not demonstrated end-to-end. Figure 4 shows that qao=3 with 10 steps costs roughly 1.5 times more FANPT wall time than qao=2 with 100 steps, and no downstream FANCI optimization effort (function evaluations, iterations, or wall time) is reported for the BeH2 calculations. The paper should either show that the reduced same-λ norms translate into fewer nonlinear FANCI solves, or be explicit that the claim is limited to the parameter-space quality of the predictor rather than total computational cost.","section":"Computational Scaling and Continuation Quality"}],"minor_comments":[{"comment":"The placeholder 'TOC ENTRY REQUIRED' remains in the manuscript and should be replaced with the actual table-of-contents graphic.","section":"TOC Graphic"},{"comment":"The notation is inconsistent: the text uses 'qao=2' and 'qao=3', while Figure 4 and the scaling section use 'QAO=2' and 'QAO=3'. Please standardize the capitalization.","section":"Throughout"},{"comment":"The rows 'Mean parameter norm' and 'Maximum parameter norm' are identical to the 'Mean Frobenius norm' and 'Maximum Frobenius norm' rows, making the table redundant. If these are meant to be different quantities, the definitions should be given.","section":"Table 4"},{"comment":"The manuscript does not provide a data availability statement or explicit software version numbers for PySCF and FanPy, and no numerical thresholds are given for the FANCI optimizations. Adding these details would aid reproducibility.","section":"Computational Details and Model Systems"}],"recommendation":"major_revision","confidential_remarks":"The derivation and implementation appear sound, but the same-λ benchmark is the linchpin of the paper's central claim. Without convergence thresholds, starting-point protocols, and root-uniqueness checks, the large norm reductions in Tables 4–12 are not yet interpretable as evidence of better initial guesses. If the authors can supply the requested diagnostics for the difficult BeH2 regions, I would support publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nThe short version: this is a solid, incremental follow-up to FANPT. The second-derivative correction works as advertised: it reduces the same-λ parameter deviation, suppresses the worst parameter-space excursions along the BeH2 continuation, and does so without overselling. The final energies are essentially unchanged, and the paper says so plainly. The derivation is coherent, and the response equations (28)–(34) are consistent with the stated approximations. The CC overlap Hessian formula is correct.\n\nWhat's genuinely new is the Hessian-corrected FANPT hierarchy: keep the overlap Hessian, drop only third- and higher derivatives, and add the resulting terms only to the constant vector while preserving the response matrix. The implementation for seniority-restricted CC in FanPy is real, and the scaling analysis is useful. The numerical tests cover two molecules and two ansatze with a small basis, which is appropriate for a proof of concept, and the paper honestly reports local degradations at r=2.5, r=1.0, and r=6.0 rather than hiding them.\n\nThe soft spots are about evidence, not math. The main diagnostic—the same-λ Frobenius norm—requires a well-defined, converged FANCI solution at every λ, and the paper does not report convergence thresholds, condition numbers, or any check for multiple roots. More importantly, the abstract says the FANCI parameters are \"independently optimized,\" but the methodology says the FANPT guess seeds the next FANCI solve. Those two statements need to be reconciled. If the reference solutions were initialized with the qao=2 trajectory, some of the large norm reductions could reflect the solver staying on the same root rather than the quality of the guess. I don't see deliberate circular fitting, but the ambiguity is real and should have been addressed. A second issue is reproducibility: all calculations used a development version of FanPy, and no code or input data are provided. For a numerical methods paper, that is a major deficiency. The wall-time comparisons also conflate step count with approximation quality, so the practical benefit is demonstrated only indirectly.\n\nI'd send this to a serious referee. The derivation is sound, the contribution is a natural and previously missing extension, and the numerical trend is consistently in the claimed direction. The referee should ask for code, input scripts, and a precise description of the FANCI solver's initialization and convergence criteria. If those checks pass, the paper will be a useful addition for people doing continuation in nonlinear wavefunction methods.\n\nFor my own work: not a citation in the next twelve months, but I'd bring it to reading group.","headline":"A solid, honest incremental extension of FANPT—the Hessian correction improves continuation stability, but the benchmark protocol and missing code make the quantitative claims conditional.","tokens_in":19164,"tokens_out":3867,"would_cite":false,"duration_ms":33802,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Retaining the overlap Hessian makes FANPT continuation stable for nonlinear wavefunctions.","keywords":["flexible ansatz for N-body configuration interaction","FANPT","overlap Hessian","nonlinear wavefunction equations","coupled cluster","adiabatic connection","continuation method","projected Schrödinger equation"],"falsifier":"Take the BeH2/CCSDT(2)Q(0) system, run both qao=2 and qao=3 with 10, 20, and 100 steps, and record the same-$\\lambda$ Frobenius norms and the number of solver failures at each geometry; if the qao=3 advantage shrinks or reverses at fine step sizes, or if the large norms simply move to different $\\lambda$ values, the claimed stabilization is limited to the coarsest steps rather than a property of the corrected predictor. Also check the condition number of the response matrix at the largest-norm points, since near-singularity would mean the norm reports parameter redundancy rather than poor guesses.","tokens_in":18114,"feed_emoji":"⚛️","tokens_out":6610,"duration_ms":52948,"temperature":0.7,"pith_summary":"This paper argues that the original quasilinear FANPT approximation, which treats the wavefunction overlap as locally linear in its parameters, can become unreliable for nonlinear ansätze such as coupled cluster, and that retaining the second derivatives of the overlap fixes the continuation path. The proposed second-derivative-corrected scheme, called qao=3, neglects only third and higher parameter derivatives, so the same response matrix as before can be reused and the new terms enter only the constant vector of the response equations. On LiH and the BeH2 insertion path with seniority-restricted coupled-cluster wavefunctions, the scheme leaves final energies essentially unchanged but sharply reduces the same-λ deviation of propagated parameters from independently optimized ones, especially for large λ-steps and near difficult geometries. The claim matters because FANPT is used mainly to generate initial guesses for solving the nonlinear projected Schrödinger equations, and a more stable guess means fewer solver iterations and fewer catastrophic parameter excursions.","feed_headline":"Curvature-corrected FANPT tames wavefunction guess blowups","feed_subtitle":"Keeping the overlap Hessian cuts parameter excursions by factors of hundreds in coarse-step BeH2 runs.","key_machinery":"The central object is the overlap Hessian, $\\partial^2 f_m/\\partial p_k\\partial p_l$, the second derivative of a determinant overlap with respect to two wavefunction parameters, which the original quasilinear approximation set to zero. For coupled-cluster wavefunctions written as a product over excitation operators, each overlap is a sum over excitation paths, and the Hessian counts paths containing both differentiated amplitudes with their fermionic signs, while diagonal elements vanish. This Hessian builds the residual tensors $G_{n,kl}$, $G_{n,klE}$, and $G_{n,kl\\lambda}$ that enter the FANPT constant vectors $B_n^{(r)}$. Because these terms appear only on the right-hand side, all orders share the same first-order response matrix as qao=2, so the correction changes the predictor rather than the linear system being solved.","core_discovery":"The central claim is that including the leading nonlinear response of the wavefunction ansatz—the second derivative of each determinant overlap with respect to the active wavefunction parameters—makes FANPT a noticeably more robust continuation method for nonlinear FANCI equations without altering the final energies. For coupled-cluster wavefunctions the overlap Hessian has a simple path-product form: a derivative of an overlap picks out excitation paths containing the differentiated amplitude, and second derivatives pick out paths containing both amplitudes, while diagonal elements vanish when an excitation appears at most once in a path. Feeding this Hessian into the FANPT response equations adds terms in $G_{n,kl}$, $G_{n,klE}$, and $G_{n,kl\\lambda}$ to the order-dependent constant vectors, while the left-hand-side response matrix is unchanged. Numerically, qao=3 reduces the mean same-$\\lambda$ Frobenius norm of the parameter correction from $1.26\\times10^{-1}$ to $1.72\\times10^{-2}$ for the 100-step LiH/CCSD(0) second-order run, and from 4.32 to 0.567 (maximum from 99.88 to 5.24) for the reduced-step BeH2/CCSD(0) run; the largest reductions occur at stretched geometries and at $\\lambda$ near 1, where qao=2 produces excursions of order $10^2$–$10^3$ in the norm. The paper therefore concludes that the second-derivative correction acts as a stabilizing term for continuation, not as an energy correction: for linear CI the Hessian vanishes, and for the tested molecules the final energies differ from qao=2 by only about $10^{-7}$ hartree (LiH) or remain within the existing energy-error trends (BeH2).","pith_inferences":["A natural next step is to include third-order overlap derivatives (qao=4) to see whether the stabilization saturates or reverses, since the constant-vector-completion pattern suggests the mechanism generalizes beyond the Hessian.","For geminal and tensor-network ansätze whose overlaps are also products or polynomial functions of parameters, the same Hessian construction should transfer directly, so the stabilization may be a general feature of product-form wavefunctions rather than a coupled-cluster-specific effect.","A sparse or symmetry-adapted overlap Hessian could lower the $\\mathcal{O}(P^2)$ overhead identified in the scaling analysis, making the correction practical for larger active spaces where the dense implementation would be the bottleneck.","The same-$\\lambda$ norm diagnostic could be complemented by measuring the basin of attraction or the number of Newton iterations required from each guess; if the Hessian correction reduces solver failures rather than just norms, the practical gain would be larger than reported."],"forward_implications":["Linear CI wavefunctions see no change: the overlap Hessian vanishes, so qao=3 reduces exactly to the original quasilinear FANPT, making the benefit specific to nonlinear ansätze.","Larger $\\lambda$-steps become usable: in the BeH2 tests, 10-step qao=3 runs attain mean same-$\\lambda$ norms comparable to or better than 50–100-step qao=2 runs, while removing nearly all norms above 10.","Continuation becomes more predictable: the LiH runs show fewer function evaluations (1299 versus 1378 in the second-order case) and lower wall time (102.4 versus 118.2 seconds), even though final energies differ by roughly $10^{-7}$ hartree.","The response-matrix structure is preserved, so existing FANPT linear solvers can be reused; the added cost is one extra power of the number of active parameters in constructing the constant vector and storing the Hessian."],"supporting_citations":[{"why":"It defines the FANCI overlap formulation and the projected Schrödinger equation on which the perturbative continuation is built.","marker":"[61]"},{"why":"It introduces the original quasilinear FANPT response equations that this work extends by retaining the overlap Hessian.","marker":"[62]"},{"why":"It provides the seniority-restricted coupled-cluster ansätze, CCSD(0) and CCSDT(2)Q(0), used in all numerical tests.","marker":"[17]"},{"why":"It generates the one- and two-electron integrals in the STO-6G basis that feed the FANCI/FANPT implementation.","marker":"[80–82]"},{"why":"It supplies the code framework in which the overlap Hessian and the modified response tensors were implemented for the test calculations.","marker":"[83–86]"},{"why":"It defines the C2v BeH2 insertion geometries used as the strong-correlation test system.","marker":"[87]"}],"fun_headline_variants":["Second-derivative FANPT fixes wavefunction guesses in BeH2","Keeping overlap Hessian curbs FANPT guess excursions","FANPT with overlap Hessian cuts guess corrections by hundreds","Including overlap second derivatives stabilizes FANPT continuation","Second-derivative FANPT yields robust wavefunction continuation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the independently optimized FANCI solution at each $\\lambda$ is the right benchmark, so the same-$\\lambda$ Frobenius norm measures the quality of the FANPT guess; this requires a unique converged FANCI solution at each $\\lambda$ and treats every parameter as equally important, which may fail if the response matrix has near-zero eigenvalues or the nonlinear equations have several solutions.","fun_headline_variants_meta":{"raw":{"variants":["Second-derivative FANPT fixes wavefunction guesses in BeH2","Keeping overlap Hessian curbs FANPT guess excursions","FANPT with overlap Hessian cuts guess corrections by hundreds","Including overlap second derivatives stabilizes FANPT continuation","Second-derivative FANPT yields robust wavefunction continuation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001118,"raw_usage":{"total_tokens":4794,"prompt_tokens":1225,"completion_tokens":3569,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":841,"completion_tokens_details":{"reasoning_tokens":3494}},"tokens_in":841,"tokens_out":3569,"duration_ms":22347,"temperature":1.0,"reasoning_tokens":3494,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:12:26.976733+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the BeH2/CCSDT(2)Q(0) system, run both qao=2 and qao=3 with 10, 20, and 100 steps, and record the same-$\\lambda$ Frobenius norms and the number of solver failures at each geometry; if the qao=3 advantage shrinks or reverses at fine step sizes, or if the large norms simply move to different $\\lambda$ values, the claimed stabilization is limited to the coarsest steps rather than a property of the corrected predictor. Also check the condition number of the response matrix at the largest-norm points, since near-singularity would mean the norm reports parameter redundancy rather than poor guesses.","supporting_citations":[{"cited_title":"D.; Miranda-Quintana, R","cited_arxiv_id":null,"evidence_quote":"It defines the FANCI overlap formulation and the projected Schrödinger equation on which the perturbative continuation is built."},{"cited_title":"A.; Kim, T","cited_arxiv_id":null,"evidence_quote":"It introduces the original quasilinear FANPT response equations that this work extends by retaining the overlap Hessian."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the seniority-restricted coupled-cluster ansätze, CCSD(0) and CCSDT(2)Q(0), used in all numerical tests."},{"cited_title":"D.; Shepard, R.; Brown, F","cited_arxiv_id":null,"evidence_quote":"It defines the C2v BeH2 insertion geometries used as the strong-correlation test system."}],"review_version":1}