{"id":"e24dfe7f-3d41-4c70-9943-8d807e3293dd","arxiv_id":"2501.12950","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A community benchmark of eleven fixed-node diffusion Monte Carlo codes shows that modern pseudopotential localization schemes give reproducible water-methane dimer interaction energies, while the old locality approximation does not.","lead":"Thirty-plus authors ran the same quantum Monte Carlo calculation for a water-methane dimer in eleven independent computer codes. The result: fixed-node diffusion Monte Carlo is reproducible across codes, provided the pseudopotential is handled with modern schemes rather than the older locality approximation.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LA spread may be confounded with differing fixed-node surfaces: the two largest LA outliers (QWalk −38.8, QMeCha −35.0 meV) are the only codes not using the shared PySCF orbitals, so the central attribution of the 25 meV LA spread to localization scheme is not yet controlled.","rationale":"The reader's weakest assumption—that identical determinantal components across codes are required for the controlled comparison—is precisely the most load-bearing concern. The SI evidence (S8.10 QWalk using GAMESS, S8.9 QMeCha using ORCA) directly contradicts the Methods claim of a common single Slater determinant from LDA DFT, and the LA outliers coincide with those two codes. This is a concrete, falsifiable confound rather than a generic worry about reproducibility. The paper has real strengths: a large community effort, tabulated raw data, explicit time-step extrapolation with reduced chi-squared and RMS residuals, and the striking TM/DLA/DTM consensus. Those strengths mean the conclusion may survive the test, but the central mechanistic attribution to the localization scheme is not yet established under a fully controlled comparison. The concern is addressable: provide the orbital files and rerun the two affected codes, or demonstrate that node differences are negligible by quantifying their effect. Because the issue is identifiable and likely fixable, the appropriate verdict remains CONDITIONAL, matching the reader's assessment. I see no need to move to ACCEPT or REJECT on the current evidence; the preprint should not be rejected, but the claim of identical determinants must be verified or the wording amended.","tokens_in":51458,"tokens_out":2805,"duration_ms":33500,"concrete_test":"Obtain the exact PySCF orbitals (e.g., the TREXIO files used by the other codes) and rerun the QWalk and QMeCha LA calculations with these identical orbitals. If their LA interaction energies move from around −38.8 and −35.0 meV into the main cluster (within about 3 meV of the LA mean), the LA spread is largely attributable to orbital/node differences rather than the localization scheme. As a complementary control, run both orbital sets (PySCF and GAMESS/ORCA) through a single code such as CASINO under both LA and DLA for the same three systems; if the LA interaction energy shifts by more than the ~3 meV reproducibility target while the DLA shift is negligible, the controlled-comparison premise fails and the conclusion must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's mechanistic conclusion is that the large LA spread across codes is attributable to the pseudopotential localization scheme, while TM/DLA/DTM are reproducible. This attribution requires that the fixed-node surfaces are identical across codes, as the Methods section (Sec. 3) asserts: all codes were required to use the same pseudopotential, basis set, and a single Slater determinant from LDA DFT, with some codes exchanging wave functions via TREXIO. However, the SI shows that QWalk (S8.10) used orbitals from GAMESS and QMeCha (S8.9) used orbitals from ORCA, not PySCF. If the orbital coefficients differ, the nodal surfaces differ, and the LA energy—which depends on the full trial wave function through the localized potential—can shift substantially. Strikingly, Table S5 shows that the two most extreme LA outliers are exactly these two codes: QWalk at −38.8 meV and QMeCha at −35.0 meV, versus the LA mean of −27.5 meV and other codes in the −17.5 to −33.3 meV range. The possibility that the LA spread is driven by nodal differences rather than the localization algorithm is therefore a concrete confound, not a hypothetical one. The TM/DLA/DTM agreement is reassuring, but those schemes may be less sensitive to small nodal differences (DLA/DTM localize on the determinant, and TM localizes only sign-violating terms), so the convergence under those schemes does not by itself rescue the LA attribution. The paper's phrase 'for the same choice of determinantal component' is not literally satisfied, and the central mechanistic conclusion rests on this premise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This community benchmark assesses the reproducibility of fixed-node diffusion Monte Carlo (FN-DMC) across eleven independently developed codes, using the water–methane dimer as a test case. All codes were asked to use the same ccECP pseudopotential, the same ccECP-ccpVTZ basis set, and a single Slater determinant of LDA DFT orbitals, with four pseudopotential localization schemes compared: locality approximation (LA), T-move (TM), determinant locality approximation (DLA), and determinant T-move (DTM). Interaction energies and total energies are extrapolated to the zero-time-step limit using polynomial fits, and the results are compared with a CCSDT(Q) reference of -26.6 meV. The main finding is that the LA shows a large cross-code spread (standard deviation about 7 meV), while TM, DLA, and DTM give much narrower distributions (standard deviations of about 2, 1.2, and 0.8 meV, respectively), leading the authors to conclude that FN-DMC is reproducible when TM, DLA, or DTM is used and when time-step convergence is carefully controlled.","tokens_in":51858,"tokens_out":8031,"duration_ms":84079,"significance":"If the result holds, this is an important and timely contribution: it is, to my knowledge, the first systematic multi-code reproducibility study of FN-DMC with a fixed pseudopotential and target system, and it provides a concrete recommendation (prefer TM/DLA/DTM over LA, and extrapolate carefully in time step) that could become standard practice. The paper's strengths include the participation of eleven independent code teams, the use of an external CCSDT(Q) reference, the publication of raw data and analysis notebooks on GitHub, and the detailed per-code supplementary notes that document algorithmic choices. The significance is moderated by two issues discussed in the major comments: the incomplete control of the determinantal component for two of the eleven codes, and the fact that no single localization scheme covers all eleven codes. These do not invalidate the core observation that TM/DLA/DTM reduce cross-code scatter, but they do require the authors either to strengthen the controlled comparison or to qualify the central mechanistic attribution.","major_comments":[{"comment":"The controlled comparison that underpins the mechanistic conclusion is incomplete for two of the eleven codes. The Methods section states that all codes used the same pseudopotential, basis set, and a single Slater determinant from LDA DFT (with TREXIO exchange for some codes), but SI S8.10 reports that QWalk used GAMESS orbitals and SI S8.9 reports that QMeCha used ORCA orbitals. Table S5 shows that these two codes are exactly the most extreme LA outliers (QWalk -38.8 meV, QMeCha -35.0 meV against an LA mean of -27.5 meV). The LA distribution therefore mixes the effect of the localization scheme with possible differences in the fixed-node surfaces. Please rerun QWalk and QMeCha with the shared PySCF orbitals for all the schemes they report, or at minimum present the LA statistics with and without these two codes and discuss whether the remaining spread still supports the attribution of the ~25 meV LA variability to the localization algorithm. In addition, please discuss explicitly why TM, DLA, and DTM are expected to be less sensitive to small nodal differences, since the current text only asserts this indirectly.","section":"Sec. 3 (Methods) and SI S8.9, S8.10, Table S5"},{"comment":"The claim that agreement is achieved 'across all eleven codes' when employing TM, DLA, and DTM is stronger than the data support. Table S5 reports the number of codes contributing to each scheme: LA 8, TM 9, DLA 9, and DTM 4; no single scheme covers all eleven codes. The statement in the Conclusions that 'agreement in the interaction energy across all eleven codes is achieved in the limit of zero time step when employing the TM, DLA, and DTM approximations' should be rephrased to specify which codes are included in each scheme and to clarify whether the eleven-code reproducibility is meant collectively (i.e., each code is represented in at least one reproducible scheme) or within each individual scheme. The Abstract's phrase 'for the same choice of determinantal component' is likewise conditional on the shared-orbital protocol that is incomplete for QWalk and QMeCha (see Major Comment 1).","section":"Abstract, Sec. 4 (Conclusions), Table S5"},{"comment":"The zero-time-step extrapolation uses a polynomial of degree 2 or 3, chosen per code and per quantity, but the selection criterion is not stated. Because the central comparison between the four localization schemes is made on the extrapolated energies, the degree choice is load-bearing. Please state the rule used to select between quadratic and cubic fits (e.g., a reduced-chi-squared threshold, an F-test, or an information criterion), and provide a stability analysis showing that the reported extrapolated interaction energies and the resulting cross-code conclusions are insensitive to the degree choice (for example, by comparing quadratic and cubic extrapolations for all fits where both are feasible). The currently reported reduced chi-squared and RMSR values are useful, but they do not by themselves justify the model selection.","section":"SI S5 and Table S1, main-text Eq. 11"}],"minor_comments":[{"comment":"The 'Mean' column reports the standard deviation of the probability distribution defined in Eq. 4 of the main text, not the standard error of the mean; please relabel this column or revise the caption to avoid ambiguity.","section":"SI Table S5 caption"},{"comment":"The construction of Pα(E) uses the single-point estimates Eα,i and σα,i without accounting for the uncertainty in these estimates. This is a reasonable practical approximation, but it likely underestimates the spread of the true distribution; a bootstrap or a t-like distribution would be a more conservative alternative. This does not affect the qualitative conclusion.","section":"Main text, Eqs. 2-4"},{"comment":"The statement that the standard deviations of the TM, DLA, and DTM total energy distributions are 'close to the theoretical minimum allowed by the precision of the performed FN-DMC simulations' would be easier to verify if the two contributions to Eq. 4 were reported separately (the statistical-error term and the deviation-from-mean term). Please add a breakdown for at least one representative case.","section":"Sec. 2, total energies discussion"},{"comment":"The broad claim 'Yes, FN-DMC is reproducible (when handled with care)' would be better qualified as 'for the water-methane dimer benchmark with single-determinant trial wave functions and the tested ccECP pseudopotential,' since the study is a single-system, single-node-surface test.","section":"Abstract and Sec. 4"},{"comment":"For QMeCha and QWalk, please confirm explicitly that the ccECP and ccECP-ccpVTZ basis set use exactly the same parameters and normalization conventions as the other codes, and whether the GAMESS/ORCA orbital coefficients have been made available in TREXIO format for reproducibility.","section":"SI S8.9 and S8.10"},{"comment":"There are several typographical and formatting issues, including ligature artifacts in the abstract ('aﬃrming'), and the phrase 'the error bar associated the the Mean' in the Table S5 caption. Please proofread the final version.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This is a valuable community study with unusually complete data sharing and a well-chosen benchmark system. My main concern is not the external validity of the TM/DLA/DTM agreement, which is robust regardless of the orbital-source issue, but the internal attribution of the LA scatter to the localization scheme. Because the two most extreme LA outliers are exactly the two codes that did not use the shared PySCF orbitals, the paper needs to either rerun those codes or carefully analyze the LA distribution without them. I also note that the conclusion 'across all eleven codes' is not met within any single scheme, which should be corrected for accuracy. The circularity concern is minor in my view: the comparison is anchored to an external CCSDT(Q) reference and to eleven independent code teams, and the fact that DLA/DTM were proposed in earlier work by some of the same authors does not by itself bias the present comparison."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is the first cross-code FN-DMC reproducibility benchmark on a common test case with controlled pseudopotential, basis, and a nominally shared Slater determinant, run across eleven independent codes. That alone makes it an important data point for the community. The empirical core holds up: after zero-time-step extrapolation, the T-move schemes (TM, DLA, DTM) give interaction energies that cluster within a few meV, while the locality approximation spreads more than 20 meV. The SI is unusually detailed, with per-code input details, raw tables, and an external CCSDT(Q) reference. Credit where due.\n\nThe weak spot is real but not fatal. The paper’s central mechanistic claim—that the LA spread is attributable to the pseudopotential localization scheme and not the trial wave function—is not fully controlled. The Methods say all codes used the same LDA/PZ orbitals from PySCF, but the SI shows QWalk used GAMESS orbitals and QMeCha used ORCA orbitals. Those two codes are precisely the most extreme LA outliers (−38.8 and −35.0 meV versus −27.5 meV mean). So nodal differences are a live confound. One can salvage the headline: excluding QWalk and QMeCha, the remaining six LA codes still span about 15 meV, far wider than the ~2 meV TM/DLA/DTM scatter. The conclusion “LA is the problem” survives; the precise attribution to localization alone does not.\n\nMinor issues: no single approximation covers all eleven codes (LA 8, TM 9, DLA 9, DTM 4), so the “all eleven codes” language in the abstract and conclusions overstates coverage. Also, the SI says all output files are on GitHub but no link or commit hash appears in the preprint.\n\nThis is a serious, careful piece of work. The right response is peer review, not desk rejection: a good referee should ask for the orbital files or a sensitivity test, a data link, and softened coverage statements. I would cite it and bring it to our next QMC reading group.","headline":"A genuinely important cross-code FN-DMC benchmark whose core finding holds up, but the LA attribution is partly confounded by two codes using different orbitals.","tokens_in":52541,"tokens_out":3988,"would_cite":true,"duration_ms":41977,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This study establishes that fixed-node diffusion Monte Carlo is reproducible across eleven community codes, provided the trial wave function's determinant is shared and the non-local pseudopotential is handled by the T-move, DLA, or DTM…","keywords":["fixed-node diffusion Monte Carlo","reproducibility","pseudopotential localization","T-move","determinant locality approximation","water-methane dimer","time-step extrapolation","quantum Monte Carlo codes"],"falsifier":"Run the LA protocol again with all codes importing one identical orbital file (or with two codes swapping orbital files) and compare the ~25 meV spread; if it persists, the localization scheme is the cause, and if it collapses, orbital or node differences were the cause.","tokens_in":66,"feed_emoji":"🧪","tokens_out":5776,"duration_ms":117382,"temperature":0.7,"pith_summary":"This community-wide study asks whether fixed-node diffusion Monte Carlo (FN-DMC), a widely used many-body electronic structure method, gives the same answer when eleven independently written codes compute the same molecule. The paper's answer, for the water-methane dimer test case, is yes, provided the trial wave function's determinantal part is the same and the non-local pseudopotential is treated with the T-move, determinant locality, or determinant T-move approximation. Under those schemes the eleven codes' interaction energies agree to within about 3 meV after extrapolation to zero time step, while the older locality approximation scatters over roughly 25 meV. The paper attributes the disagreement to the pseudopotential localization scheme, not to the codes themselves. If right, the result strengthens the case for FN-DMC as a portable reference for molecular energies.","feed_headline":"Yes, fixed-node DMC is reproducible across 11 codes","feed_subtitle":"With T-move, DLA, or DTM schemes, eleven codes agree to ~3 meV; the old locality approximation spreads ~25 meV.","key_machinery":"The central object is the treatment of non-local pseudopotentials in the DMC propagator. Non-local pseudopotentials introduce a sign-violating term into the imaginary-time Green's function, and four localization schemes replace it by an effective local operator: the locality approximation (LA) localizes on the full trial wave function; the T-move approximation (TM) localizes only the sign-violating part, restoring the upper-bound property; the determinant locality approximation (DLA) localizes on the determinantal part only; and the determinant T-move approximation (DTM) applies determinant localization to the sign-violating part while keeping the full wave function elsewhere. The quantitative machinery is a per-code probability distribution P($\\alpha$)(E) = (1/N) sum of per-code Gaussians, whose standard deviation separates stochastic error from cross-code disagreement, plus polynomial extrapolation E(tau)=A+B tau+C $tau^{2}$+D $tau^{3}$ to the zero-time-step limit.","core_discovery":"Using a single Slater determinant from LDA density functional theory as the fixed-node surface, shared ccECP pseudopotentials and basis, and only code-specific Jastrow factors, the authors find that the zero-time-step FN-DMC interaction energy of the water-methane dimer is reproducible across all eleven codes when the non-local pseudopotential is localized with the T-move (TM), determinant locality approximation (DLA), or determinant T-move (DTM) scheme. The spread across codes is about 3 meV for TM and DLA and about 1 meV for DTM (four codes), and the averaged values sit close to the CCSDT(Q) reference of about -27 meV. With the older locality approximation (LA), the same codes disagree by up to about 21 meV (standard deviation 7 meV), showing that the localization scheme, not the underlying DMC algorithm, is the dominant source of non-reproducibility. The authors also show that the time step must be small: at tau=0.04 a.u. DLA results scatter over more than 60 meV, narrowing onto the common value only near tau=0.0025 a.u.","pith_inferences":["The mechanistic attribution would be sharper if every code used bitwise identical orbitals; the supporting information records that two codes generated orbitals with different quantum-chemistry packages, so a fraction of the LA scatter could in principle originate from slightly different fixed-node surfaces rather than localization error alone.","A direct testable extension is to run the same protocol on a heavier-element dimer with a stronger non-local channel; the prediction would be that TM, DLA, and DTM still collapse the code-to-code spread while LA widens it.","For periodic solids, the same localization-scheme contrast is likely to control reproducibility, but finite-size corrections and Brillouin-zone sampling add error sources that this molecular benchmark cannot constrain.","The per-code probability-distribution formalism could serve as a general reproducibility metric for future QMC benchmark studies, cleanly separating statistical noise from systematic code disagreement."],"forward_implications":["Users of FN-DMC who employ TM, DLA, or DTM can expect published interaction energies from independent codes to be mutually consistent at the few-meV level, given identical determinants and converged time steps.","The older locality approximation should not be relied on for quantitative cross-code comparisons; its large spread is an artifact of the localization choice, not of the stochastic method.","Time-step convergence must be checked explicitly: at tau=0.04 a.u. the DLA interaction-energy spread exceeds 60 meV, while tau=0.0025 a.u. is adequate for this system.","Total energies, not just energy differences, are reproducible to sub-millihartree precision under TM, DLA, and DTM, with standard deviations below about 6 meV across codes.","The study supplies a benchmark protocol for future QMC comparisons: fix the determinant, share the pseudopotential and basis, and extrapolate to zero time step."],"supporting_citations":[{"why":"Introduces the determinant T-move scheme, the key localization variant that the paper shows gives the narrowest cross-code spread.","marker":"[116]"},{"why":"Define the T-move approximation family used in the TM and DTM schemes, including size-consistent variants and reduced time-step error.","marker":"[112, 118, 119]"},{"why":"Define the original locality approximation whose variability across codes establishes the paper's central contrast.","marker":"[111, 117]"},{"why":"Provide the ccECP pseudopotentials and basis required for identical input across all eleven codes.","marker":"[136, 137]"},{"why":"Supplies the dimer geometry and the size-consistent time-step weighting used in the calculation protocol.","marker":"[135]"},{"why":"Provides the CCSDT(Q) reference interaction energy used to compare the FN-DMC averages.","marker":"[134]"}],"fun_headline_variants":["Fixed-node DMC: 11 codes agree to 3 meV","Old LA scheme: 21 meV spread, T-move: 3 meV","T-move makes DMC reproducible across 11 codes","Water-methane dimer: DMC reproducible with T-move","11 DMC codes reproducible with T-move"],"cache_read_input_tokens":54400,"weakest_assumption_plain":"The load-bearing premise is that every code used the same determinantal component, so the fixed-node surfaces coincide; the supporting information records that two codes generated orbitals with different quantum-chemistry packages, and any orbital differences would shift the nodes and could absorb part of the LA scatter.","fun_headline_variants_meta":{"raw":{"variants":["Fixed-node DMC: 11 codes agree to 3 meV","Old LA scheme: 21 meV spread, T-move: 3 meV","T-move makes DMC reproducible across 11 codes","Water-methane dimer: DMC reproducible with T-move","11 DMC codes reproducible with T-move"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000865,"raw_usage":{"total_tokens":3827,"prompt_tokens":1099,"completion_tokens":2728,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":715,"completion_tokens_details":{"reasoning_tokens":2638}},"tokens_in":715,"tokens_out":2728,"duration_ms":20699,"temperature":1.0,"reasoning_tokens":2638,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T16:36:25.276229+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the LA protocol again with all codes importing one identical orbital file (or with two codes swapping orbital files) and compare the ~25 meV spread; if it persists, the localization scheme is the cause, and if it collapses, orbital or node differences were the cause.","supporting_citations":[],"review_version":1}