{"id":"fea06b43-8933-4e74-885b-0f803c4d62b4","arxiv_id":"1909.02487","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"The Fermionic Neural Network is an antisymmetric neural-network wavefunction which, optimized variationally, recovers most correlation energy and outperforms CCSD(T) on several strongly correlated dissociation curves.","lead":"A new type of neural network, the Fermionic Neural Network, solves the many-electron Schrödinger equation more accurately than standard coupled cluster methods on several difficult molecules, using only atomic positions and charges. It makes variational quantum Monte Carlo competitive with much more expensive quantum chemistry methods, a step toward first-principles calculations of strongly correlated systems.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central 'no data' claim conflicts with the paper's Hartree-Fock pretraining; the dissociation-curve results depend on an initialization and hyperparameter choices that are not tested for robustness.","rationale":"The paper's core technical contribution, the FermiNet ansatz and its energy results, is credible and well supported by comparisons to exact or high-accuracy references. The reader's conditional verdict is appropriate. However, the reader's weakest assumption focused on optimization reaching a near-global minimum; my read identifies a more direct discrepancy: the abstract's 'no data other than atomic positions and charges' claim is contradicted by the Hartree-Fock pretraining described in Appendix A, which is stated to be necessary for numerical stability on large systems. This is an internal inconsistency between the central claim and the method as actually executed, and it is load-bearing because the headline contribution is explicitly framed as data-free ab initio prediction. The lack of released code and the paper's own admission that convergence is highly hyperparameter-dependent reinforce the concern but are secondary. The concrete test of running with and without HF pretraining would settle whether the pretraining is a convenience or a requirement; until that is demonstrated, the claim as stated is not established. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":28857,"tokens_out":7072,"duration_ms":81295,"concrete_test":"Run the released FermiNet implementation on the N2 dissociation curve at bond lengths 2.0, 3.0, and 4.5 a0, both with the Hartree-Fock pretraining described in Appendix A and with random initialization, using Table V hyperparameters and three random seeds per condition. Record final unclipped energies, convergence curves, and numerical failures. If the pretraining-free runs converge to statistically identical energies, the 'no data' claim survives; if they diverge, fail, or land more than 5 mEh above the pretrained energies, the central claim's data-free framing and reproducibility are not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim says the method uses 'no data other than atomic positions and charges,' but Appendix A states that before local-energy optimization the network was 'pretrained to match Hartree-Fock (HF) orbitals computed using PySCF,' and that without this pretraining 'the determinants ... would often numerically underflow ... causing the optimization to fail' on large systems. This is not a cosmetic detail: the reported N2 and H10 dissociation curves are obtained after an external Hartree-Fock calculation in an STO-3G basis, which is both a computational input and a basis-set choice beyond atomic positions and charges. The same appendix concedes 'Accurate and stable convergence was highly dependent on the hyperparameters used,' and Table V's defaults were altered for bicyclobutane. Since no code is released and no hyperparameter sensitivity study is provided, the claim that one architecture with one set of training parameters transfers 'out-of-the-box' is not supported by the manuscript as written.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Fermionic Neural Network (FermiNet), a variational wavefunction ansatz for continuous-space many-electron systems. The ansatz builds antisymmetric wavefunctions from determinants whose entries are permutation-equivariant functions of all electron coordinates, computed by a neural network with one-electron and two-electron streams, plus an exponentially decaying envelope. Parameters are optimized by minimizing the variational energy via VMC with KFAC natural-gradient updates and Metropolis sampling. The authors report ground-state energies for first-row atoms, small molecules, the H4 rectangle, the N2 dissociation curve, and the H10 chain, comparing against exact/FCI, CCSD(T), DMC, AFQMC, and experimental references. The central claim is that the same architecture with a single default hyperparameter set, using no data other than atomic positions and charges, outperforms coupled cluster on strongly correlated dissociation curves and approaches projector-QMC accuracy at equilibrium.","tokens_in":29096,"tokens_out":5551,"duration_ms":58336,"significance":"If the results hold, this is a substantial advance: it demonstrates that a flexible neural-network ansatz can represent strongly correlated molecular wavefunctions in continuous space well enough to beat single-reference CCSD(T) on stretched systems and to rival DMC/AFQMC at equilibrium, using a variational method with polynomial scaling. The paper contains several strong elements worth credit: all reported energies are variational upper bounds; the results are cross-checked against exact, FCI, and experimental references; the custom reverse-mode gradient for singular determinants is a genuine technical contribution; Appendix B gives a universality argument for generalized Slater determinants; and the non-interacting-chain check in Appendix E directly addresses size consistency. The main weaknesses are the absence of released code, the acknowledged sensitivity of convergence to hyperparameters, and the overstated 'no data' wording given the Hartree-Fock pretraining step.","major_comments":[{"comment":"The abstract's claim that the method uses 'no data other than atomic positions and charges' is contradicted by the pretraining procedure in Appendix A 1, where the network is 'pretrained to match Hartree-Fock (HF) orbitals computed using PySCF' in an STO-3G basis. The HF orbitals are a fitted target for the pretraining loss, and the STO-3G basis is an external modeling choice beyond atomic positions and charges. I recommend rephrasing the claim (for example, 'no empirical data' or 'no data beyond the Hamiltonian and the atomic positions/charges') and adding a short demonstration that the final energies are insensitive to the pretraining basis.","section":"Abstract and Appendix A 1"},{"comment":"Appendix A states that 'Accurate and stable convergence was highly dependent on the hyperparameters used,' and Table V had to be altered for bicyclobutane. Because the central N2 and H10 results (Figs. 5 and 6) are obtained with one default hyperparameter set, and no sensitivity analysis or released code is provided, the manuscript does not currently support the 'out-of-the-box' claim made in Section I and the Discussion. Please add a hyperparameter-robustness study (at least on one stretched system, perturbing the Table V values) and release the code with defaults, or weaken the transferability claim proportionately.","section":"Appendix A 1 and Table V"},{"comment":"The statement in Section III B that the FermiNet is 'more accurate than CCSD(T) in the largest basis set we could practically run' should not be conflated with outperforming CCSD(T) in the complete-basis limit: the CCSD(T)/CBS column in Table II is below the FermiNet energy for most listed molecules, and CCSD(T) is non-variational. The abstract's stronger claim concerns dissociation curves, which is supported by Figs. 4-6, but the equilibrium-geometry comparisons should be worded to distinguish finite-basis CCSD(T) from basis-set-extrapolated CCSD(T).","section":"Section III B and Table II"}],"minor_comments":[{"comment":"The text refers to 'the Hamiltonian of the system as given in Eqn. I,' but the Hamiltonian is labeled Eq. (1); please fix the cross-reference.","section":"Section II B"},{"comment":"There is a typo in the second determinant of Eq. (7): 'det[φ^↓_i(r^↓_j;{r^↓_{/j}});{r^↑};])' has a misplaced semicolon and parenthesis.","section":"Eq. (7)"},{"comment":"The caption says electron affinities for Be, N, and Ne are not computed because their anions are unstable; the beryllium anion is often regarded as borderline, so a brief justification or citation would help.","section":"Table I caption"},{"comment":"The caption reports a power-law exponent with a bootstrap error bar but does not describe the bootstrap procedure; please provide details in the caption or in Appendix A.","section":"Figure 10 caption"},{"comment":"The description of the pretraining distribution uses p_pre(X) as an equal mixture of the product of Hartree-Fock orbitals and ψ²(X); it would be clearer to state explicitly how the two halves of the MCMC batch are generated in practice.","section":"Appendix A 1"}],"recommendation":"major_revision","confidential_remarks":"This is an important paper that will likely have high impact, but the reproducibility case needs strengthening before publication. The absence of released code is a serious practical barrier for a method whose central claims depend on stochastic optimization and tuned hyperparameters. The 'no data' wording in the abstract is likely to be read as a stronger claim than the Hartree-Fock pretraining actually supports; this should be resolved in revision. The overlap with concurrent works (Hermann et al. and Choo et al.) is acknowledged in footnote 27, which is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The FermiNet paper is the real thing. The core idea is clean and genuinely new: replace the one-electron orbitals in a Slater determinant with permutation-equivariant functions of all electron coordinates, evaluated in continuous space. That single change, plus a flexible neural network and KFAC optimization, gets variational Monte Carlo to energies that beat single-reference CCSD(T) on stretched N2 and H10 and match DMC/AFQMC at equilibrium. The comparisons are honest — variational upper bounds, cross-checked against FCI, exact, and experimental references where available. The H4 rectangle and the hydrogen chain results are particularly convincing because the failures of CC are qualitative, not just a few millihartrees. The appendix proof that a single generalized determinant is in principle universal is a nice touch, even if the construction is non-continuous and thus not directly learnable.\n\nThe soft spots are real but not fatal. The abstract says the method uses 'no data other than atomic positions and charges,' but Appendix A states the network is pretrained to match Hartree-Fock orbitals computed with PySCF in an STO-3G basis. That is an external computational input and a basis-set choice, not just atomic positions. It is not circular — HF is just an initialization, not a training target — but the claim as written is too strong. The stress-test note lands here. Second, the same appendix concedes that 'accurate and stable convergence was highly dependent on the hyperparameters used,' and the defaults were altered for bicyclobutane. So the 'one architecture, one set of training parameters, out of the box' claim is not fully supported by the manuscript. Third, no code is released. For a paper whose method is defined by its implementation details, that is a meaningful barrier to independent verification. These are addressable issues, not load-bearing flaws.\n\nThe reader's conditional verdict is about right. The physics and the architecture are solid; the overclaim is in the packaging. I would send this to a serious referee. The referee should ask for code, a hyperparameter sensitivity study, and a clearer statement of what the HF pretraining does and does not contribute. The paper matters for computational chemistry and for ML-guided quantum mechanics, and it deserves careful peer review.\n\nRecommendation: accept for peer review. The results are important and the core method is sound; the required revisions are about transparency, not correctness.","headline":"FermiNet is the real breakthrough: a continuous-space neural wavefunction ansatz that beats CCSD(T) on stretched systems, though the 'no data' claim is overstated by the HF pretraining and the out-of-the-box robustness is not fully demonstrated.","tokens_in":29575,"tokens_out":1392,"would_cite":true,"duration_ms":17921,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces the Fermionic Neural Network, a deep-learning wavefunction ansatz that, using only atomic positions and charges, reaches lower variational energies than CCSD(T) on stretched nitrogen and hydrogen chains and recovers…","keywords":["variational quantum Monte Carlo","FermiNet","neural network wavefunction ansatz","strongly correlated electrons","Slater determinants","coupled cluster benchmark","hydrogen chain","basis-set-free quantum chemistry"],"falsifier":"Retrain the FermiNet on stretched N2 at one bond length, say 3.5 a0, from several independent random seeds and with a different optimizer; if the spread of final variational energies exceeds the reported gap to CCSD(T), the claimed accuracy is not the ansatz's true optimum.","tokens_in":28639,"feed_emoji":"⚛️","tokens_out":7935,"duration_ms":81166,"temperature":0.7,"pith_summary":"The paper sets out to show that a neural-network wavefunction can solve the many-electron Schrödinger equation directly, without any external data or a chosen basis set. The central claim is that the Fermionic Neural Network (FermiNet), trained by variational Monte Carlo, reaches ground-state energies within chemical accuracy for first-row atoms and small molecules, recovers more than 99% of correlation energy for molecules like ethene, and on stretched nitrogen and on the hydrogen chain is more accurate than CCSD(T), the standard benchmark method. A sympathetic reader would care because this suggests variational quantum Monte Carlo, long considered less accurate than projector methods, can be competitive with or better than the best scalable quantum chemistry approaches while using the same architecture for every system.","feed_headline":"A neural network wavefunction beats coupled cluster on hard molecules","feed_subtitle":"Variational Monte Carlo with FermiNet reaches near-exact energies using only atomic positions and charges.","key_machinery":"The load-bearing object is the generalized Slater determinant: the orbital in each determinant is not a function of one electron's position but a permutation-equivariant function $\\varphi^k_i(x_j; \\{x_{/j}\\})$ of the whole electronic configuration, so exchanging two electrons swaps rows or columns of the determinant matrix and the wavefunction is antisymmetric by construction. The network feeds single-electron features (electron-nucleus vectors and distances) through one stream and pairwise electron features through another, symmetrically averages same-spin activations, and after several residual tanh layers forms spin-up and spin-down determinant blocks weighted by exponentially decaying envelopes that enforce the boundary conditions. The energy is minimized with a Kronecker-factored approximate natural-gradient optimizer, and the cusp conditions are captured because interparticle distances are included directly as inputs.","core_discovery":"The discovery is that antisymmetry, the main obstacle to using neural networks for electrons, can be built into a Slater-determinant ansatz by letting every orbital depend on all electron coordinates in a permutation-equivariant way. With this ansatz, a single network trained by energy minimization yields, from only atomic positions and charges, dissociation curves for N2 and H10 that are significantly closer to exact or experimental references than unrestricted CCSD(T), and equilibrium energies that match DMC and AFQMC on small systems. The paper further reports that the ansatz reproduces the exact FCI energy surface of the H4 rectangle where coupled cluster predicts a spurious cusp.","pith_inferences":["If the measured $O(N^{-0.395})$ decay of error with one-electron stream width continues, extrapolating FermiNet energies to infinite width would give a basis-free analogue of complete-basis-set extrapolation; the paper fits the scaling but does not perform that extrapolation.","Because the ansatz is set in continuous space with no lattice or basis, adapting it to periodic boundary conditions is a natural next step toward solids and surfaces, a direction the present paper does not address.","The paper's derivation that the optimizer is equivalent to stochastic reconfiguration suggests the practical gains come partly from the optimization algorithm, so architectural advances for the ansatz and preconditioner advances for the optimizer may compound rather than compete."],"forward_implications":["A single architecture and one hyperparameter set transfer across atoms, diatomics, and small organic molecules, so the method can be applied to a new system without system-specific ansatz design.","Since the wavefunction lives in the continuum, FermiNet results avoid basis-set extrapolation error; widening the one-electron stream provides a systematic, if polynomial, route to lower energies.","The FermiNet can be used as a trial wavefunction for projector quantum Monte Carlo, which would carry its accuracy into DMC and AFQMC calculations.","For out-of-equilibrium and multireference-like systems, variational optimization with FermiNet avoids the non-variational failures of single-reference coupled cluster, as shown by the H4 rectangle and stretched N2."],"supporting_citations":[{"why":"Supplies the conventional Slater-Jastrow and backflow variational Monte Carlo energies that the FermiNet is compared against on first-row atoms.","marker":"[44]"},{"why":"Provides the AFQMC, coupled cluster, and conventional VMC benchmark energies for the hydrogen chain that the FermiNet is measured against.","marker":"[52]"},{"why":"Gives the experimental nitrogen dissociation potential reconstructed from spectroscopy, the reference curve for the N2 comparison.","marker":"[50]"},{"why":"Provides the highly accurate r12-MR-ACPF nitrogen dissociation curve that the FermiNet matches in the strongly stretched region.","marker":"[51]"},{"why":"Supplies exact numerical ground-state energies for first-row atoms used to set the chemical-accuracy baselines.","marker":"[34]"},{"why":"Supplies exact reference energies for N2 and Li2 used to judge the FermiNet equilibrium results.","marker":"[47]"},{"why":"Introduces the Kronecker-factored approximate curvature optimization method used to train the FermiNet.","marker":"[40]"},{"why":"Provides the standard variational and diffusion Monte Carlo formalism, including Metropolis sampling from the square of the wavefunction.","marker":"[12]"}],"fun_headline_variants":["FermiNet beats CCSD(T) on N2 and H10 dissociation curves","Neural net wavefunction tops coupled cluster on hard molecules","Antisymmetric neural network yields near-exact energies for molecules","Deep learning solves many-electron equation with just atom positions","FermiNet matches exact H4 surface where coupled cluster fails"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparisons assume that the optimization routine reaches a near-global minimum of the variational energy for every system, so the reported numbers are the true FermiNet limits; the paper gives empirical convergence evidence but notes that stable convergence was highly dependent on the hyperparameters.","fun_headline_variants_meta":{"raw":{"variants":["FermiNet beats CCSD(T) on N2 and H10 dissociation curves","Neural net wavefunction tops coupled cluster on hard molecules","Antisymmetric neural network yields near-exact energies for molecules","Deep learning solves many-electron equation with just atom positions","FermiNet matches exact H4 surface where coupled cluster fails"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000391,"raw_usage":{"total_tokens":2052,"prompt_tokens":938,"completion_tokens":1114,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":1025}},"tokens_in":554,"tokens_out":1114,"duration_ms":11769,"temperature":1.0,"reasoning_tokens":1025,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:50:35.744879+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the FermiNet on stretched N2 at one bond length, say 3.5 a0, from several independent random seeds and with a different optimizer; if the spread of final variational energies exceeds the reported gap to CCSD(T), the claimed accuracy is not the ansatz's true optimum.","supporting_citations":[{"cited_title":"Motta , author D","cited_arxiv_id":null,"evidence_quote":"Provides the AFQMC, coupled cluster, and conventional VMC benchmark energies for the hydrogen chain that the FermiNet is measured against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the experimental nitrogen dissociation potential reconstructed from spectroscopy, the reference curve for the N2 comparison."},{"cited_title":"Gdanitz ,\\ @noop journal journal Chem","cited_arxiv_id":null,"evidence_quote":"Provides the highly accurate r12-MR-ACPF nitrogen dissociation curve that the FermiNet matches in the strongly stretched region."}],"review_version":1}