{"id":"0cdce49b-c6a8-4caf-b41c-e4234f494c2b","arxiv_id":"2608.11911","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A neural foundation model trained on hundreds of thousands of quadratic qubit Hamiltonians produces variational ground-state energy bounds that transfer across system sizes and topologies, though large-scale extrapolation fails on some frustrated lattices.","lead":"Hamilton-Zero is a 0.5-billion-parameter classical model that learns approximate ground states of spin-1/2 quantum Hamiltonians, generalizing across system size, topology, and interaction type. It reports variational energy upper bounds and fine-tunes to high accuracy on small systems, but its largest-scale results are much weaker than the abstract suggests.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'rigorous upper bound' claim holds only for the exact expectation under the model density |ψ|^2; with finite MCMC samples the reported energies are not certified, and the paper's own Sec. 6 concedes the bound can be violated by estimator noise or mixing bias.","rationale":"The reader's weakest assumption identifies exactly the load-bearing condition: the reported energies are variational upper bounds only if the MCMC samples are unbiased draws from |ψ|^2. The paper's own Sec. 6 weakens the main-text claim by conceding that finite-sample estimates can violate the variational inequality through estimator noise or mixing bias. This is not an internal inconsistency, but it means the central assertion is over-stated for the reported numbers, especially at 8100 qubits where no external reference exists. The proposed test would settle the question by using a large system with a rigorous exact energy and the same sampling protocol. If the test passes, the sampler concern is largely resolved at those sizes; if it fails, the 'rigorous upper bound' claim is not valid for the reported estimates. This is a reason to keep the reader's CONDITIONAL verdict rather than to reject: the theoretical framework and small-system fine-tuning results are credible, and the concern is a scalability/verification gap rather than a demonstrated contradiction of the variational principle.","tokens_in":54860,"tokens_out":16762,"duration_ms":174579,"concrete_test":"Use the released checkpoint and the identical zero-shot inference protocol (512 measurement steps, 256 walkers, 8 tempered replicas, 24 adaptive Langevin moves) on a large-N system with an exact, independently computable ground-state energy: e.g., the 1D transverse-field Ising chain H = -Σ Z_i Z_{i+1} - g Σ X_i at g=1 for N=256, 1024, and 4096, with periodic boundary conditions, whose energy can be obtained by free-fermion (Jordan-Wigner) diagonalization to machine precision. For each N, compute the signed relative gap (E_reported - E_exact)/|E_exact| and compare it with the reported Monte Carlo standard error.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central assertion of Sec. 2 — that every energy Hamilton-Zero reports is a rigorous upper bound on the physical ground-state energy — is true only if the reported number is the exact expectation of the local energy under the model density |ψθ|^2/Z. The sampling pipeline of Sec. 5.2 and SM§S4 uses a replica-exchange Langevin sampler with finite walkers, finite burn-in, and a finite number of adaptive moves per step; no proof is given that this chain is unbiased at the 8100-qubit scale. Sec. 6 explicitly states that finite-sample estimates may violate the variational inequality by estimator noise and/or mixing bias. A biased chain that undersamples regions of high local energy will produce an average below E0 even when the ansatz is exactly in the spin-1/2 sector. The 8100-qubit square-lattice row in Table 1 is precisely the case where no external reference can catch such a violation, and its reported E/N = +0.128 is far above the comparison scale, so it provides no evidence either way. The stationarity tests mentioned in Sec. 6 are heuristic and cannot certify unbiasedness; they cannot distinguish a slowly mixing chain from a converged one on a timescale shorter than the mixing time.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces Hamilton-Zero, a roughly 0.5B-parameter neural network that represents spin-1/2 ground states as centrally odd, per-site-linear functions on SU(2)^N, with the Hamiltonian acting through Lie derivatives. Using the Peter-Weyl decomposition, the authors argue that this multilinear ansatz lies in the physical spin-1/2 sector and hence satisfies the variational upper-bound inequality for exact expectations (Sec. 2, SM S1). They pretrain on hundreds of thousands of perturbed quadratic Hamiltonians, evaluate zero-shot and fine-tuned on held-out ED-referenced systems, and report large-system case studies up to 8100 qubits, along with scaling laws, a zero-variance calibration, a phase-transition witness, learned equivariance, and learned contraction routes.","tokens_in":55167,"tokens_out":6256,"duration_ms":64000,"significance":"If the technical claims hold, Hamilton-Zero would be a substantial advance in amortized quantum many-body computation: a single pretrained variational wavefunction transferring across topologies, system sizes, and interaction types at a scale far beyond prior foundation neural quantum states, with a clean Peter-Weyl-based resolution of the phase-space leakage problem. The paper is unusually explicit about limitations (Sec. 6), releases code and model weights, and the small-system exact-diagonalization comparisons support the energy claims in-distribution. The central scientific value, however, depends on whether the gap between the exact-expectation variational bound and the finite-sample Monte-Carlo estimates can be closed or the claims appropriately qualified; in its current form the 'rigorous upper bound' framing overstates what the reported numbers establish.","major_comments":[{"comment":"The variational inequality in Eq. (2.5) is derived for exact expectations under |ψθ|^2, but the abstract and Sec. 2 state that 'every energy our optimiser reports is a rigorous upper bound.' Sec. 6 itself concedes that finite-sample Monte-Carlo estimates may violate the inequality through estimator noise or mixing bias. All reported energies in Sec. 5.2 and Table 1 are finite-sample averages (e.g., 512 measurement steps, 256 walkers, 8 replicas), with no proof or diagnostic that the chain is unbiased at the largest system sizes. The 'rigorous upper bound' claim is therefore not established for any reported number; please either provide a sampler-certification argument or revise the claims to 'variational estimate subject to sampling error' and remove 'rigorous' from the abstract and Sec. 2.","section":"Sec. 2, Eq. (2.5)"},{"comment":"The zero-variance calibration constant κ=0.2825 is fitted to ED-referenced systems with N≤18 and then used in Fig. 6 to 'predict' relative gaps for N>22, including the statement that the predicted median relative gap is 3.45%. This is model-based extrapolation under the unverified assumption that q=κV is universal across system size and Hamiltonian family; it is not a measured gap. The scaling-law exponents in Fig. 4 are likewise fitted to this pretraining run. The text should clearly label these quantities as extrapolations under an assumption and should not present the 3.45% as a verified accuracy of the model without independent large-system references.","section":"Sec. 5.1, Figs. 5-6"},{"comment":"The large-system comparisons use non-certified references: MaxCut rows compare against archived feasible cuts rather than certified optima; PPP rows compare against a thermodynamic-limit literature value under a different convention; and square and triangular lattice rows compare against thermodynamic-limit scales rather than exact finite-size energies. Consequently the recovery percentages do not certify variational accuracy. The 8100-qubit square-lattice row reports E/N=+0.128 versus a reference of −0.4968, which is a clear failure; the abstract's 'evaluate on systems up to 8100 qubits' should be accompanied by the fact that the model fails at that scale, and the introduction's 'up to 8000 qubits' claim should be tempered accordingly.","section":"Table 1, Sec. 5.3"},{"comment":"No evidence is given that the replica-exchange Langevin sampler mixes on a timescale short compared with the measurement budget for the large systems. The stationarity tests mentioned in Sec. 6 are heuristic and cannot distinguish a slowly mixing chain from a converged one within a finite number of steps. Please report effective sample sizes, autocorrelation times, replica-exchange acceptance rates, or a comparison with independent references at intermediate sizes (e.g., N=1024 and 2025) before using the large-system zero-shot energies as support for the variational claim.","section":"Sec. 5.2, SM S4"}],"minor_comments":[{"comment":"The phrase 'evaluate on systems up to 8100 qubits' should be qualified with the failure and degradation reported in Sec. 5.3; as written, it invites overreading of the large-system results.","section":"Abstract and Sec. 5.3"},{"comment":"The text 'the router’s training has converge to the true optimum' should read 'has converged', and 'electric vehical charging' in the introduction should be 'vehicle charging'.","section":"Sec. 3 and Sec. 1"},{"comment":"'104 KFAC steps' should read '10^4 KFAC steps'; the exponent appears to be missing.","section":"Sec. 5.2"},{"comment":"The captions of Figs. 7 and 8 appear identical; if this is not intentional, the captions should be differentiated to describe the zero-shot and fine-tuned panels respectively.","section":"Figs. 7 and 8"},{"comment":"The spelling 'Supplimentary Material' should be 'Supplementary Material', and 'Hamilton-zero' in Sec. 6 should be consistently capitalized as 'Hamilton-Zero'.","section":"Sec. 2 and Sec. 6"},{"comment":"The operation in Eq. (S2.61) is described as a 'rank-4 quadrilinear merge' but is bilinear in the two child carriers; consider using consistent terminology such as 'rank-4 blockwise bilinear merge' to avoid confusion.","section":"SM S2.4, Eq. (S2.61)"}],"recommendation":"major_revision","confidential_remarks":"The core theoretical framework and the small-system evidence are sound and the engineering is impressive, but the central 'rigorous upper bound' claim is only valid for exact expectations, and the finite-sample MCMC estimates are not certified. The deficiencies are correctable by careful rewording, sampler diagnostics, and explicit labeling of extrapolated quantities, so I recommend major revision rather than rejection. I would also advise the editor that the abstract's 8100-qubit phrasing should not be allowed to stand without an explicit statement of the failure at that scale."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a serious paper with a real new idea, but the abstract overstates the headline capability and the \"rigorous upper bound\" claim is only as strong as the sampler. Worth a serious referee; needs revision, not rejection.\n\nThe genuinely new thing is the combination: representing spin-1/2 states as centrally odd, quaternion-multilinear functions on SU(2)^N, acting with the Hamiltonian through Lie derivatives, and using Peter-Weyl to keep the variational upper bound intact. The theory section is credible — for the exactly multilinear class, the Casimir argument pins the state to the spin-1/2 sector and the inequality follows. The architecture (featurizer, trunk, RL router, shared merge tree) is substantial, and the released code matters. Small-system fine-tuning is the strongest evidence: median signed gaps of 3e-4% to 6e-3% against exact diagonalization on 256 held-out systems, with convergence front-loaded. That is a real result.\n\nNow the soft spots, in proportion. First, the abstract says \"evaluate on systems up to 8100 qubits,\" and Table 1's 8100-qubit square-lattice row is a failure: E/N = +0.128 against a comparison scale near -0.4968. The body reports it as -25.8% recovery, which is a sign violation, not a bound. The abstract should say the model compiles and samples at that size, and that accuracy degrades and fails on frustrated lattices. Second, \"every energy is a rigorous upper bound\" holds only for the exact expectation under |psi|^2. The MCMC estimator is not certified, and the paper's own Sec 6 concedes finite-sample estimates can violate the inequality via noise or mixing bias. At 8100 qubits there is no external reference to catch a biased chain, so that row proves nothing either way. The stationarity tests are heuristic. Third, the zero-variance calibration: kappa = 0.2825 is fit to ED-referenced systems and then used to predict gaps for N > 22. That is a fit presented as a principle; the reader's circularity concern lands. Fourth, the large-system references are not certified — feasible cuts for MaxCut, thermodynamic-limit scales for PPP and lattices — so the recovery percentages are against approximate baselines.\n\nThe citation pattern looks reasonable, and the paper engages the prior FNQS lineage honestly. The Sec 6 limitations paragraph is candid; the abstract just doesn't match it.\n\nWho should read this: anyone working on neural quantum states, amortized simulation, or quantum-advantage baselines. The small-system results and the variational framework are worth taking seriously. I'd send it to review and ask for the abstract to be qualified, the sampler bias characterized at scale, and head-to-head baselines (scratch NQS, DMRG where applicable) before acceptance.","headline":"A genuinely new foundation-model approach to ground-state computation with strong small-system results, hobbled by an abstract that overstates the 8100-qubit capability and a variational-bound claim that outruns the sampler.","tokens_in":55643,"tokens_out":2891,"would_cite":true,"duration_ms":28745,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a single 0.5 billion parameter neural network, pretrained on hundreds of thousands of quadratic spin-1/2 Hamiltonians, can produce variational upper bounds for ground-state energies of unseen systems across different…","keywords":["Hamilton-Zero","foundation model","neural quantum states","variational Monte Carlo","quadratic qubit Hamiltonians","spin-1/2 ground states","SU(2) manifold","Peter-Weyl sector"],"falsifier":"Take the 8100-qubit square J1-J2 system from Table 1: the model reports +0.128 per spin while the comparison scale is about -0.497; rerunning with a much longer burn-in and more walkers and checking whether the estimate crosses below zero would settle whether the variational upper bound actually holds at that size.","tokens_in":54682,"feed_emoji":"⚛️","tokens_out":7410,"duration_ms":73107,"temperature":0.7,"pith_summary":"Hamilton-Zero is a single pretrained neural network, roughly 0.5 billion parameters, that aims to compute ground states of arbitrary quadratic spin-1/2 Hamiltonians. Rather than storing amplitudes on exponentially many spin configurations, it writes the wavefunction as a smooth odd function on a product of SU(2) spheres, so the Hamiltonian acts by differentiation. A Peter-Weyl argument restricts this function class to the physical spin-1/2 sector, making the variational upper bound exact for converged expectations. After training on hundreds of thousands of Hamiltonians that vary in size, interaction graph, and interaction type, the model transfers zero-shot to systems up to 8100 qubits, and fine-tuning only the small compiled readout brings median signed gaps to 0.00031-0.006 percent relative to exact diagonalization. The point is that ground-state computation can be amortized: one shared training run replaces per-system optimization for many future Hamiltonians.","feed_headline":"One network bounds ground states across qubit Hamiltonian families","feed_subtitle":"A single pretrained checkpoint transfers over sizes, topologies, and interaction types, while energy estimates stay variational upper…","key_machinery":"The load-bearing object is the manifold wavefunction: a centrally odd scalar function on SU(2)^N, linear in each site's quaternion. That choice lets the Hamiltonian act through Lie derivatives, computed with custom automatic-differentiation primitives, and the Peter-Weyl theorem provides the sector projection that keeps the energy variational. Around this core sits a transformer trunk that reads only the Hamiltonian's coupling tensors, a per-site leaf builder where quaternion coordinates enter, and a balanced binary merge tree with a shared quadrilinear merge tensor; a reinforcement-learned routing policy picks which sites merge early, so the contraction path adapts to each Hamiltonian's interaction structure. The same machinery yields differentiable wavefunctions and an explicit log-amplitude, so energies and observables are evaluated by variational Monte Carlo with a replica-exchange Langevin sampler on SU(2)^N.","core_discovery":"The paper's central claim is that a single checkpoint, trained once, can serve as a foundation for ground states across a universal family of spin-1/2 Hamiltonians. The construction represents each spin by a unit quaternion, so a state is a centrally odd, per-site linear function on SU(2)^N; spin operators become left-invariant Lie derivatives evaluated by automatic differentiation. By the Peter-Weyl decomposition, per-site oddness plus linearity in each quaternion confines the wavefunction to the spin-1/2 sector, so the Rayleigh quotient is a rigorous upper bound on the true ground-state energy whenever the Monte Carlo expectation is converged. Empirically, the pretrained model generalizes across system sizes, topologies, and interaction types; on held-out systems fine-tuning the compiled merge tree alone reaches median signed gaps of 3.1e-4% to 6.0e-3% relative to exact diagonalization, and zero-shot evaluation extends to systems with thousands of qubits.","pith_inferences":["A step beyond the paper: if the variational upper bound survives careful sampling at scale, the energies could be promoted to rigorous bounds in combinatorial-optimization settings, giving certified cuts or ground-state energies rather than heuristic estimates; the paper does not claim this.","One natural extension the paper leaves implicit is using the differentiable wavefunction to compute two-point correlation functions, which would allow a single checkpoint to map complete phase diagrams from correlation-based order parameters, not only fidelity susceptibility.","The positive zero-shot energy on the 8100-qubit lattice suggests the current checkpoint's size-extrapolation limit is real; pretraining at larger sizes or conditioning more explicitly on system size is the obvious next stress test."],"forward_implications":["A single pretrained checkpoint replaces per-system training from scratch for quadratic spin-1/2 ground-state problems, making the marginal cost of a new Hamiltonian the cost of sampling and optionally fine-tuning a small compiled readout.","Zero-shot transfer to system sizes far above the training range (up to 8100 qubits) is demonstrated, including nonlocal topologies where tensor-network references are hard to obtain.","Fine-tuning fewer than one percent of the model's parameters (the compiled merge tree) recovers 0.00031-0.006 percent median signed gaps on held-out systems, with 90-100 percent of systems within one percent of exact diagonalization.","A fixed checkpoint can serve as a phase-transition witness: fidelity susceptibility computed from zero-shot weights peaks near the transverse-field Ising critical point without any fine-tuning.","The learned routing policy recovers physically natural contraction hierarchies, such as nearest-neighbour pairs, plaquettes, and carbon-orbital blocks, on systems more than thirty times larger than any seen in pretraining."],"supporting_citations":[{"why":"Peter-Weyl theorem: decomposes L^2(SU(2)^N) into irreducible spin sectors, identifying the physical spin-1/2 sector that makes the variational bound exact.","marker":"[48, 49]"},{"why":"Earlier phase-space formalism for many-body spins on SU(2)^N; Hamilton-Zero extends it to scalable Lie-derivative evaluation and resolves its harmonic-leakage open problem.","marker":"[40]"},{"why":"Neural quantum states; establishes the variational-Monte-Carlo approach and the NQS function class that the manifold ansatz contains.","marker":"[16]"},{"why":"Foundation neural quantum states; the transferable-NQS baseline Hamilton-Zero generalizes from fixed interaction graphs to variable topologies, sizes, and interaction types.","marker":"[22]"},{"why":"Universality of quadratic two-qubit interactions; justifies targeting the arbitrary quadratic Pauli Hamiltonian family.","marker":"[44]"},{"why":"Variational Monte Carlo and the zero-variance principle; supplies the upper-bound guarantee and the V-score calibration used to report accuracy.","marker":"[106]"},{"why":"Kronecker-Factored Approximate Curvature; the second-order optimizer that the paper extends for sharded natural-gradient training.","marker":"[43]"}],"fun_headline_variants":["A single model bounds ground states for any qubit Hamiltonian","Pretrained qubit ground-state solver generalizes to 8100 qubits","Universal ground-state learner for spin-1/2 Hamiltonians","One checkpoint, many Hamiltonians: variational upper bounds","Neural tensor network finds ground states across qubit families"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported energies are variational upper bounds only if the Monte Carlo sampler has converged and mixed within the sampling budget; the paper states this is unverified at the largest system sizes, especially the 8100-qubit case.","fun_headline_variants_meta":{"raw":{"variants":["A single model bounds ground states for any qubit Hamiltonian","Pretrained qubit ground-state solver generalizes to 8100 qubits","Universal ground-state learner for spin-1/2 Hamiltonians","One checkpoint, many Hamiltonians: variational upper bounds","Neural tensor network finds ground states across qubit families"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000192,"raw_usage":{"total_tokens":1708,"prompt_tokens":1041,"completion_tokens":667,"prompt_tokens_details":{"cached_tokens":1024},"prompt_cache_hit_tokens":1024,"prompt_cache_miss_tokens":17,"completion_tokens_details":{"reasoning_tokens":582}},"tokens_in":17,"tokens_out":667,"duration_ms":282619,"temperature":1.0,"reasoning_tokens":582,"cache_read_input_tokens":1024,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:22:58.783559+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the 8100-qubit square J1-J2 system from Table 1: the model reports +0.128 per spin while the comparison scale is about -0.497; rerunning with a much longer burn-in and more walkers and checking whether the estimate crosses below zero would settle whether the variational upper bound actually holds at that size.","supporting_citations":[],"review_version":1}