{"id":"e5b49cc2-dd3b-488e-bf8a-c4794598fc81","arxiv_id":"2608.08758","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Analytical nuclear gradients and Hessians were implemented on an IBM quantum processor with oo-VQE and M0 error mitigation, giving an H2 vibrational frequency close to reference.","lead":"This paper computes molecular forces and vibrational frequencies on a real quantum computer, using error-corrected measurements for hydrogen and water. It shows that small molecules can approach reference accuracy, while larger ones remain limited by hardware noise.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The zero-parameter M0 calibration (Eq. 39) is not shown to model the deep transpiled circuits used for RDM and A/B measurements (Eq. 41); the corrected tensor pipeline rests on an unvalidated noise-transfer assumption.","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing point: the M0 confusion matrix is built from all-zero-parameter circuits and is assumed, without validation, to describe the noise of the deep transpiled circuits actually used for the density-matrix and linear-response measurements. This is the most fundamental concern because the paper's novelty claim is explicitly about corrected QPU measurement of tensor elements; if the correction model is untested, the corrected tensors and the Hessian built from them are not reliable, however plausible the H2 agreement looks. The H2 demonstration is real evidence, but it is not a direct test of the transfer assumption: good final numbers can arise from error cancellation, from the fact that readout errors dominate this small system, or from the selection of 35 of 46 batches. The water single-shot measurements and the discarded high-error batches are additional signs that the error model is not fully characterized. The proposed parameter-matched calibration test would settle the issue by comparing the standard M0 correction against a calibration that actually exercises the optimized circuit structure; if the two agree within errors, the concern is retired and the central claim is strengthened. The verdict remains CONDITIONAL because the existing paper does not provide that validation, but the concern is concrete and testable rather than a groundless objection.","tokens_in":32717,"tokens_out":10546,"duration_ms":136021,"concrete_test":"Run the H2 Hessian pipeline twice on a noisy simulator calibrated to IBM Pittsburgh, or on the same backend: first with M0 built from U(0) circuits (Eq. 39), then with a parameter-matched calibration obtained by preparing U(theta*)|x> for each computational basis state and least-squares fitting the confusion matrix C to the known ideal output distributions P_ideal(k|x)=|<k|U(theta*)|x>|^2. If the corrected 1-RDM, A/B matrices, or the resulting stretch frequency move by more than the reported uncertainties (about 21 cm−1), the zero-parameter transfer assumption is falsified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim depends on the M0 error-mitigation step. In Eq. (39), the confusion matrix is built from circuits U(0)|x>, with all tUPS parameters set to zero. For the tUPS generators in Eq. (9), U(0) is analytically the identity, so the calibration circuit is substantially shallower than the actual measurement circuits: for H2, the transpiled measurement circuit has depth 98 and 29 entanglers (Table 2), while the U(0) calibration circuit contains essentially no entangling gates. The inversion in Eq. (41) is therefore valid only if the dominant noise channel is independent of the ansatz parameters and identical for the shallow calibration and deep measurement circuits. No experiment in the paper validates this transfer: there is no parameter-matched calibration, no noisy-simulator cross-check, and no condition-number or bias analysis of the inverted M0 matrix. The water results, with mitigated energies 123–201 mHa below reference, are consistent with a broken or incomplete noise model. If the transfer is invalid, the corrected 1-RDM, 2-RDM, and A/B tensor elements — and hence the 5008±21 cm−1 Hessian result — carry an unquantified bias, and the close H2 agreement may be partly coincidental. The additional discarding of 11 of 46 Hessian batches (Sec. 4.2.2) compounds this by making the reported statistics conditional on a selected subset, so the quoted uncertainty does not reflect full run-to-run hardware variability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an analytical implementation of nuclear gradients and Hessians within an orbital-optimized VQE framework using the pp-tUPS ansatz, with the required 1-RDM, 2-RDM, and linear-response A/B tensor elements measured on IBM Pittsburgh hardware and corrected by the M0 confusion-matrix error-mitigation scheme with post-selection. Ideal simulations for H2/STO-3G (AS(2,2)) and H2O/STO-3G (AS(4,4)) are compared with PySCF FCI/CASSCF and Dalton references, and QPU results for H2 give mitigated energy errors below 9 mHa, gradient errors below 6 mHa/Bohr, and a stretching frequency of 5008±21 cm−1 versus a 5000 cm−1 finite-difference reference. The water QPU results are single evaluations with mitigated energies 123–201 mHa below the reference, and the Hessian experiment retains only 35 of 46 batches after discarding runs on noisy qubits.","tokens_in":33013,"tokens_out":5050,"duration_ms":53555,"significance":"If the error-mitigation pipeline is valid, the paper demonstrates a complete QPU workflow for analytical first and second nuclear derivatives in an active-space framework, including an explicit tensor-measurement protocol for the response equations. Strengths include the use of standard Helgaker–Jørgensen equations, independent classical references from PySCF and Dalton, explicit circuit resource tables, and a quantitative flagship frequency prediction that is directly falsifiable on the same hardware. The central load-bearing assumption, however, is that the M0 confusion matrix measured with all ansatz parameters equal to zero transfers to the much deeper transpiled measurement circuits; this is not validated in the manuscript. In addition, the reported hardware statistics exclude 11 of 46 Hessian runs, so the headline frequency uncertainty is conditional on a selected subset. These issues are fixable in a revision with additional control experiments and complete data reporting.","major_comments":[{"comment":"The M0 calibration circuits are structurally much shallower than the circuits actually used to measure the RDMs and the A/B matrices. Because the tUPS tile in Eq. (9) reduces to the identity when all parameters are zero, the calibration states in Eq. (39) contain essentially no entangling gates, whereas the transpiled H2 measurement circuit has depth 98 and 29 entanglers and the H2O circuit has depth 414 and 257 entanglers (Table 2). The inversion in Eq. (41) therefore assumes that gate and crosstalk noise is independent of ansatz parameters and of transpilation. No parameter-matched calibration, noisy-simulator cross-check, or condition-number/bias analysis is provided. Since every corrected density, A/B element, and hence the 5008±21 cm−1 frequency depends on this assumption, the central claim requires validation: calibrate M0 at the actual optimized parameters for at least one geometry, compare corrected expectation values against exact values on a noisy simulator with gate-dependent noise, and report the spectrum or condition number of M0.","section":"§2.5.1, Eqs. (38)–(42), and Table 2"},{"comment":"The Hessian statistics are computed after discarding 11 of 46 batches because they “were run on unreliable qubits with high error rates,” yet no pre-defined, result-independent exclusion rule is given. The reported 5008±21 cm−1 and the tensor-error statistics in Table 6 are therefore conditional on a selected subset, and the quoted uncertainty does not characterize full run-to-run variability. Please report all 46 batches, state the exclusion criterion (for example, a calibration threshold fixed before data taking), and show the sensitivity of the stretching frequency to inclusion/exclusion. The negative average translational eigenvalues (−407±171 cm−1) also indicate that the measured Hessian is not positive semidefinite; a discussion of this instability is needed to qualify the reliability of the retrieved force constants.","section":"§4.2.2, Fig. 8, and Table 5"},{"comment":"The water QPU results, although explicitly labeled as single evaluations, constitute an in-manuscript test of the M0 noise-transfer assumption and they are not consistent with it: the mitigated energies lie 123–201 mHa below the CASSCF reference, meaning the correction overshoots by an amount comparable to the raw error. Because the same M0 pipeline is used for the H2 Hessian, this overcorrection raises the possibility that the H2 agreement is partly coincidental. The manuscript should either identify a mechanism for the water overcorrection (for example, parameter-dependent gate errors, crosstalk, or state-preparation errors) or, at minimum, substantially soften the claim that the approach demonstrates good performance beyond H2.","section":"§4.1.4, Table 4"}],"minor_comments":[{"comment":"A step-size convergence test for the finite-difference CASSCF Hessian reference should be reported; 0.001 Å is reasonable, but the sensitivity of the 5000 cm−1 reference to this step is not documented.","section":"§3.1"},{"comment":"Units are missing or inconsistent across rows (for example, “∆1-RDM Max15.571±6.543” has no unit, while later rows quote mHa); add units to every row and state explicitly that both “Max” and “Distance” quantities are in the property’s own units.","section":"§4.2.2, Table 6"},{"comment":"The text states that the mitigated gradient precision is “∼ 1mHa/Bohr,” while Table 6 reports a mean gradient-modulus error of 7.261±5.894 mHa/Bohr; these numbers need to be reconciled.","section":"§4.1.3 and Table 6"},{"comment":"Since all tUPS parameters are zero, |x0⟩ is just |x⟩ for the defined ansatz; the notation should be explained or simplified, and the phrase “while considering the gate noise drifting” should be made mathematically precise.","section":"§2.5.1, Eq. (39)"},{"comment":"The first two eigenvalues are labeled x/y translations; for a finite system these should be near zero, and the mean −407 cm−1 suggests broken translational symmetry. A brief explanation of why these modes are not projected out would clarify the Hessian quality assessment.","section":"§4.2.2, Table 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for physics.chem-ph and the ideal-simulation comparisons are sound. The main risks are the unvalidated M0 noise-transfer assumption and selective reporting of the hardware Hessian data; both are addressable in revision. I would not recommend rejection on the basis of disagreement with existing error-mitigation practice, but the revision should include additional control experiments or noisy-simulator evidence for the M0 transfer, complete reporting of the discarded runs, and a mechanism or explicit caveat for the water overcorrection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nShort version: this is the first paper I know of that puts analytical nuclear Hessians on real quantum hardware, measuring the A and B matrices of linear response on the QPU. The H2 result is credible — stretching frequency 5008±21 cm^-1 against a 5000 cm^-1 reference, mitigated energies within 9 mHa — and the error tables are unusually honest. The water attempt is less flattering, and the M0 error-mitigation scheme is the soft spot.\n\nWhat's actually new: the gradient part is an incremental extension of known analytical-gradient work (Sugisaki, Mitarai, O'Brien), but the Hessian via QPU-measured tensor elements is new. They build the orbital Hessian G(0) from A and B measured on the IBM Pittsburgh device, then solve the CP-MCSCF-like response by explicit inversion. That is a real step. The ideal simulations match PySCF and Dalton well, which validates the analytical implementation before noise enters. The SI is thorough — element-wise error statistics for 1-RDM, 2-RDM, A, B, and the nuclear Hessian across 35 batches.\n\nThe main weakness is exactly what the stress-test says: the M0 confusion matrix is calibrated with U(0), which is essentially the identity, while the actual measurement circuits have depth ~98 and 29 entanglers for H2. The transfer from shallow calibration to deep circuit is assumed, not demonstrated. There is no parameter-matched calibration, no noisy-simulator cross-check, no condition-number analysis. For H2 the method happens to work — the agreement is good — but the water results, with mitigated energies 123–201 mHa below reference, look like an unvalidated noise model overcorrecting. The paper acknowledges this as a trade-off, but it means the central quantitative claim still rests on a hidden assumption.\n\nThe second issue is the data selection: 11 of 46 Hessian batches were dropped because of high qubit error rates. That is defensible in practice, but the quoted ±21 cm^-1 error bar is conditional on that selection. It is a minor-to-moderate concern, and they should state it more prominently.\n\nMinor points: the CASSCF Hessian reference is finite-difference (PySCF lacks analytic CASSCF Hessians), and water is a single evaluation per shot count. Neither undermines the H2 proof-of-principle.\n\nBottom line: the paper deserves a serious referee. The theory is sound, the H2 hardware result is a genuine benchmark for future work, and the authors are candid about limitations. The referee should ask for a validation of M0 transfer — e.g., a noisy simulation with parameter-matched calibration — and clearer reporting of the batch selection. I would cite it if I worked in NISQ quantum chemistry.\n\nBest,\n[Name]","headline":"A credible proof-of-principle for analytical Hessians on real quantum hardware, with honest error accounting; the M0 noise-transfer assumption and batch selection are the real soft spots.","tokens_in":33579,"tokens_out":2853,"would_cite":true,"duration_ms":31788,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Analytical nuclear gradients and Hessians can be computed on real quantum hardware by measuring the underlying tensor elements and correcting them with error mitigation.","keywords":["nuclear gradient","nuclear Hessian","orbital-optimized VQE","tiled unitary product state ansatz","linear response","error mitigation","quantum hardware","vibrational frequency"],"falsifier":"Build the $\\boldsymbol{M}_0$ confusion matrix from circuits whose ansatz parameters are nonzero (or from the exact transpiled gates used for the property measurements), apply the same correction to the H$_2$ densities and Hessian, and compare the resulting stretching frequency with the $5008 \\pm 21$ cm$^{-1}$ reported here; a shift larger than the quoted uncertainty, or corrected densities that move away from the ideal values, would show that the zero-parameter noise model does not transfer.","tokens_in":32515,"feed_emoji":"⚛️","tokens_out":7399,"duration_ms":71452,"temperature":0.7,"pith_summary":"This paper tries to establish that the two quantities computational chemistry most needs for geometry optimization and vibrational spectroscopy—the nuclear gradient and the nuclear Hessian—can be obtained analytically from measurements on a real quantum processor, not from finite differences or classical simulation. The strategy is to run an orbital-optimized variational quantum eigensolver with a tiled unitary product state ansatz inside an active space, then measure the reduced density matrices and the linear-response $\\boldsymbol{A}$ and $\\boldsymbol{B}$ matrices that the analytical derivative equations demand. A confusion-matrix error-mitigation scheme built from the same ansatz with all parameters set to zero, plus post-selection on electron number, brings the measured tensors close enough to the ideal values that the H$_2$ stretching frequency comes out at $5008 \\pm 21$ cm$^{-1}$ against a $5000$ cm$^{-1}$ reference. For water the same workflow still underestimates energies by about $150$ mHa, so the paper doubles as a resource and scaling analysis of what current hardware can and cannot do.","feed_headline":"Quantum hardware measures molecular forces and vibration frequencies","feed_subtitle":"No finite differences: analytical gradients and Hessians measured directly on an error-mitigated quantum processor.","key_machinery":"The machinery has three coupled pieces. First, the pp-tUPS ansatz—a tiled unitary product state circuit of spin-adapted single- and pair-double excitation gates, layered to approach CASSCF accuracy—produces the wavefunction on the quantum processing unit. Second, the measurement protocol directly evaluates the tensor elements the derivative equations need: the one- and two-particle reduced density matrices for the static gradient and static Hessian, and the $\\boldsymbol{A}$ and $\\boldsymbol{B}$ linear-response submatrices whose difference builds the electronic orbital Hessian $\\mathcal{G}^{(0)}=2(\\boldsymbol{A}-\\boldsymbol{B})$; the response vector is then obtained by inversion of $\\mathcal{G}^{(0)}$, avoiding finite differences. Third, the $\\boldsymbol{M}_0$ error mitigation constructs a confusion matrix from circuits in which all ansatz parameters are set to zero, inverts it against the raw measurement statistics, and post-selects bitstrings with the correct separate $\\alpha$ and $\\beta$ electron counts. The analytical gradient and Hessian formulas then combine these corrected tensors with classically computed integral derivatives.","core_discovery":"The central claim is that the novelty lies in the explicit and corrected measurement, on a quantum processing unit, of the tensor elements required by the analytical equations for nuclear derivatives: the one- and two-electron reduced density matrices for the static terms, and the $\\boldsymbol{A}$ and $\\boldsymbol{B}$ linear-response matrices whose combination $\\mathcal{G}^{(0)}=2(\\boldsymbol{A}-\\boldsymbol{B})$ forms the electronic orbital Hessian needed for the relaxed (response) term. The gradient follows the standard analytical-derivative expression with Pulay connection terms, and the Hessian adds the static second-derivative terms plus the response term $f^{(1)}\\lambda^{(1)}$ solved by full-space measurement and inversion of $\\mathcal{G}^{(0)}$. On the hardware side, each required expectation value is obtained from the pp-tUPS ansatz circuit and corrected with the $\\boldsymbol{M}_0$ confusion matrix built from the same ansatz with all parameters set to zero, followed by post-selection that keeps only bitstrings with the correct separate $\\alpha$ and $\\beta$ electron counts. Demonstrated on H$_2$, the mitigated energies sit within $9$ mHa of full configuration interaction, the gradient modulus within $6$ mHa/Bohr, and the retrieved H$_2$ stretching frequency is $5008 \\pm 21$ cm$^{-1}$ versus the $5000$ cm$^{-1}$ finite-difference reference; on water the same protocol still underestimates the energy by roughly $150$ mHa and demands substantially more QPU time.","pith_inferences":["If the zero-parameter confusion matrix faithfully models the noise of the optimized circuits, the same measurement pipeline should transfer to other response properties—polarizabilities, NMR shieldings, hyperfine couplings—since those share the same $\\boldsymbol{A}$ and $\\boldsymbol{B}$ building blocks; this is a testable extension rather than a claim the paper makes.","A direct experimental check of the transfer assumption would be to build the $\\boldsymbol{M}_0$ matrix from circuits with nonzero ansatz parameters and compare the corrected 1- and 2-RDMs; systematic differences would quantify the hidden bias in all mitigated Hessians.","The observed error cancellation between 2-RDM and $\\boldsymbol{A}$ errors suggests that allocating measurement shots preferentially to the tensor elements with the largest variance could improve accuracy at fixed total shot count, something the paper does not explore.","Because only 35 of 46 hardware batches survived the qubit-error selection, the reported statistics describe a filtered subset; repeating with error-mitigated selection criteria that do not discard data would test whether the selection itself inflates the apparent accuracy."],"forward_implications":["Geometry optimizations and harmonic vibrational frequencies of small molecules can be run on current noisy quantum hardware using analytical derivatives instead of numerical finite differences, with per-geometry QPU times on the order of minutes for four-qubit systems.","The same measured $\\boldsymbol{A}$ and $\\boldsymbol{B}$ matrices serve both the electronic Hessian and the static property gradients, so for small active spaces no extra circuits are needed beyond those already used for the energy.","The errors in individual tensors (especially the 2-RDM and the $\\boldsymbol{A}$ matrix) are larger than the final nuclear Hessian error, indicating beneficial error cancellation among the contributing terms.","The water experiment sets a concrete scaling boundary: an 8-qubit, 414-depth transpiled circuit with 290 Pauli groups costs hours of QPU time and still underestimates energies by about 150 mHa, so resource and noise characterization are the current bottleneck.","For larger systems, the authors propose avoiding explicit full electronic Hessian measurement via Hessian-vector products and a Davidson-style iterative solver, which would replace the exponential confusion-matrix cost with a more scalable procedure."],"supporting_citations":[{"why":"Provides the analytical gradient and Hessian equations that the QPU-measured tensors feed into.","marker":"[103]"},{"why":"Introduces the $\\boldsymbol{M}_0$ ansatz-based readout and gate error mitigation used to correct the measured expectation values.","marker":"[67]"},{"why":"Supplies the linear-response formulation and the explicit measurement strategy for the $\\boldsymbol{A}$ and $\\boldsymbol{B}$ matrices on quantum hardware.","marker":"[26]"},{"why":"Defines the pp-tUPS ansatz whose layered structure and perfect-pairing order are used for both molecules.","marker":"[96]"},{"why":"Demonstrates $\\boldsymbol{M}_0$ mitigation and post-selection on real hardware; the post-selection criterion used on the 1-RDM diagonal comes from this line of work.","marker":"[32]"},{"why":"Discusses the cost and scaling of error mitigation for tiled ansätze and motivates the resource bottleneck observed for water.","marker":"[68]"},{"why":"Software that constructs the ansatz and Hamiltonian, maps them to Pauli strings, and orchestrates the QPU jobs.","marker":"[110]"},{"why":"Classical quantum-chemistry package used to compute the FCI/CASSCF energies, gradients, and optimized reference geometries.","marker":"[107]"}],"fun_headline_variants":["Analytical molecular forces and Hessians on a quantum processor","Error-mitigated quantum gradients for molecules","Quantum hardware measures molecular forces without finite differences","Direct quantum measurement of molecular vibration frequencies","Orbital-optimized VQE yields analytical gradients on QPU"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole correction scheme assumes that the noise of the real, deep, transpiled circuits is the same as the noise seen in the shallow calibration circuits with all ansatz parameters set to zero, and no experiment in the paper checks that transfer; the reported statistics also come from only the 35 of 46 runs that were not discarded for high qubit errors.","fun_headline_variants_meta":{"raw":{"variants":["Analytical molecular forces and Hessians on a quantum processor","Error-mitigated quantum gradients for molecules","Quantum hardware measures molecular forces without finite differences","Direct quantum measurement of molecular vibration frequencies","Orbital-optimized VQE yields analytical gradients on QPU"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000428,"raw_usage":{"total_tokens":2247,"prompt_tokens":1058,"completion_tokens":1189,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":674,"completion_tokens_details":{"reasoning_tokens":1115}},"tokens_in":674,"tokens_out":1189,"duration_ms":13444,"temperature":1.0,"reasoning_tokens":1115,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:24:40.264751+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build the $\\boldsymbol{M}_0$ confusion matrix from circuits whose ansatz parameters are nonzero (or from the exact transpiled gates used for the property measurements), apply the same correction to the H$_2$ densities and Hessian, and compare the resulting stretching frequency with the $5008 \\pm 21$ cm$^{-1}$ reported here; a shift larger than the quoted uncertainty, or corrected densities that move away from the ideal values, would show that the zero-parameter noise model does not transfer.","supporting_citations":[{"cited_title":"Molecularhessiansforlarge-scalemcscfwavefunctions","cited_arxiv_id":null,"evidence_quote":"Provides the analytical gradient and Hessian equations that the QPU-measured tensors feed into."},{"cited_title":"Accurateandgate-efficientquantumansätzeforelectronicstateswithoutadaptiveoptimiza- tion","cited_arxiv_id":null,"evidence_quote":"Defines the pp-tUPS ansatz whose layered structure and perfect-pairing order are used for both molecules."},{"cited_title":"Cost-effective scalable quantum error mitigation for tiled Ans\\\"atze","cited_arxiv_id":"2511.21236","evidence_quote":"Discusses the cost and scaling of error mitigation for tiled ansätze and motivates the resource bottleneck observed for water."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Software that constructs the ansatz and Hamiltonian, maps them to Pauli strings, and orchestrates the QPU jobs."},{"cited_title":"Pyscf:Thepython-basedsimulationsofchemistryframework","cited_arxiv_id":null,"evidence_quote":"Classical quantum-chemistry package used to compute the FCI/CASSCF energies, gradients, and optimized reference geometries."}],"review_version":1}