Pith. sign in

REVIEW 3 major objections 2 minor 3 cited by

Fermi-Dirac machines as quantizations of neurons

T0 review · 3 major / 2 minor · reviewed 2026-06-30 · grok-4.3

Pith's one-line read Quantizing classical neurons via quantum Hamiltonians produces models whose decision problem is BQP-complete.

desk verdict The paper gives a clean quantization map from classical neurons to operator activations plus a BQP-completeness result, but the claimed exact reduction when operators commute only matches expectations and leaves the output stochastic. read the letter →

arxiv 2605.24386 v1 pith:3J23RP7M submitted 2026-05-23 quant-ph cond-mat.stat-mechcs.DScs.LG

classification quant-phcond-mat.stat-mechcs.DScs.LG
keywords Fermi-DiracneuronsquantumcanonicalquantizationBQP-completenesshybridquantum-classicalalgorithmsactivationfunctionsneuralnetworksmachinelearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper reinterprets Fermi-Dirac machines as quantizations of classical neurons. A classical neuron is viewed as an activation function applied to a parameterized classical Hamiltonian, which is quantized by replacing variables with operators whose eigenvalues encode the possible values. When the operators commute, the construction reduces exactly to the classical neuron. Hybrid algorithms based on sampling, Hamiltonian simulation, and the Hadamard test enable efficient evaluation and training of the resulting activation observable. Numerical experiments indicate that the quantum versions can learn functions classical neurons cannot, and a decision problem on Fermi-Dirac neurons is proven BQP-complete.

What carries the argument

The activation observable formed by applying an activation function to a parameterized quantum Hamiltonian.

What would settle it

Discovery of a classical polynomial-time algorithm that solves the defined decision problem on Fermi-Dirac neurons would falsify the BQP-completeness claim.

Watch

Extended reading notes

Core claim

Viewing a classical neuron as an activation function applied to a parameterized classical Hamiltonian allows canonical quantization by replacing the variables with operators. The resulting activation observable, when measured on an input state, yields the expected output of the quantized neuron as a random variable. When the Hamiltonian operators commute, the model recovers the classical neuron exactly. Efficient hybrid quantum-classical algorithms evaluate outputs and gradients. Numerical results indicate that neurons based on quantum Hamiltonians can learn functions inaccessible to classical neurons, while the associated decision problem is BQP-complete.

Load-bearing premise

Replacing classical variables in the neuron Hamiltonian with quantum operators whose eigenvalues encode the same values yields a valid quantization that reduces to the classical case when the operators commute and supports efficient hybrid algorithms.

Editorial extensions

If this is right

  • Various activation functions including smooth ReLU, sigmoid linear unit, and GeLU can be quantized in the same manner.
  • Hybrid algorithms using random sampling, Hamiltonian simulation, and the Hadamard test evaluate outputs and gradients of the quantized neurons.
  • The approach extends to continuous quantum variables with two sketched methods for composing neurons into networks.
  • BQP-completeness of the decision problem supplies complexity-theoretic evidence that efficient classical simulation is impossible in general.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The quantization method may connect Fermi-Dirac machines directly to neural network training on quantum hardware for semidefinite optimization tasks.
  • Practical tests on specific function approximation problems could reveal the scope of learning advantages in finite-size implementations.
  • The same operator-replacement technique might apply to other classical models such as support vector machines or decision trees.
  • BQP-completeness raises the possibility that certain trained quantum neuron behaviors cannot be replicated by any efficient classical circuit family.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The paper reinterprets Fermi-Dirac machines as canonical quantizations of classical neurons, obtained by replacing variables in a parameterized classical Hamiltonian with operators whose eigenvalues encode possible values. It asserts that the construction reduces exactly to the classical neuron when operators commute, defines an activation observable whose measurement yields the neuron output (a random variable whose expectation matches the classical activation), develops hybrid quantum-classical algorithms based on sampling, Hamiltonian simulation, and the Hadamard test for evaluation and gradients, quantizes multiple activations including smooth ReLU variants, reports numerical experiments indicating quantum neurons can learn functions inaccessible to classical ones, proves BQP-completeness of a decision problem based on Fermi-Dirac neurons, and sketches extensions to continuous variables and networks.

Significance. If the BQP-completeness proof is rigorous and the numerical results survive comparison to stochastic classical baselines, the work would supply complexity-theoretic evidence against efficient classical simulation of these models together with concrete hybrid algorithms and a quantization framework applicable to several standard activations. The explicit reduction claim and hybrid primitives are strengths when properly substantiated.

major comments (3)
  1. [Abstract and §2] Abstract and construction section: the assertion that the model 'reduces exactly to a classical neuron' when the Hamiltonian consists of commuting operators is not supported by the stated definition. The output remains a random variable obtained by measuring the activation observable; only its expectation equals the classical activation applied to eigenvalues. Classical neurons are deterministic scalar functions, so the commuting case yields a stochastic model whose sampling distribution is unaddressed. This directly affects whether the numerical experiments demonstrate advantage beyond classical stochastic neurons and whether the BQP result separates the construction from classical simulation.
  2. [Numerical experiments] Numerical experiments section: no dataset descriptions, error analysis, statistical significance, or explicit comparisons to classical stochastic or noisy neurons are provided. The claim that quantum neurons learn functions classical neurons cannot therefore cannot be evaluated; the experiments must be augmented with these controls to substantiate the advantage.
  3. [BQP-completeness] BQP-completeness section: the decision problem must be stated precisely (e.g., whether it concerns expectation values, sampling, or a promise problem) and shown to remain BQP-complete while reducing to an efficiently simulable classical problem in the commuting limit. The proof should address whether the stochastic output in the commuting case permits efficient classical simulation, as this is load-bearing for the complexity separation.
minor comments (2)
  1. [Algorithms] Clarify notation for the input state preparation and how the hybrid algorithms scale with the number of qubits and the form of the activation function.
  2. [Abstract] The abstract states the reduction is 'exact' yet immediately defines the output as a random variable; reconcile this wording for precision.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for their thorough review and insightful comments, which help clarify key aspects of the manuscript. We address each major comment below, indicating revisions where appropriate to strengthen the presentation and substantiate the claims.

read point-by-point responses
  1. Referee: [Abstract and §2] Abstract and construction section: the assertion that the model 'reduces exactly to a classical neuron' when the Hamiltonian consists of commuting operators is not supported by the stated definition. The output remains a random variable obtained by measuring the activation observable; only its expectation equals the classical activation applied to eigenvalues. Classical neurons are deterministic scalar functions, so the commuting case yields a stochastic model whose sampling distribution is unaddressed. This directly affects whether the numerical experiments demonstrate advantage beyond classical stochastic neurons and whether the BQP result separates the construction from classical simulation.

    Authors: We agree that the wording in the abstract and §2 requires clarification. The reduction holds in the sense that, when the operators commute and the input state is an eigenstate of the Hamiltonian with eigenvalues matching the classical inputs, the activation observable is diagonal and its measurement yields the classical activation value with probability 1. In general, the output is a random variable whose expectation matches the classical case. We will revise the abstract and §2 to explicitly state these conditions, describe the sampling distribution in the commuting limit (which is a delta function under the eigenstate input), and note that the model is stochastic otherwise. This clarification does not alter the hybrid algorithms or BQP result but will better contextualize the numerical experiments. revision: yes

  2. Referee: [Numerical experiments] Numerical experiments section: no dataset descriptions, error analysis, statistical significance, or explicit comparisons to classical stochastic or noisy neurons are provided. The claim that quantum neurons learn functions classical neurons cannot therefore cannot be evaluated; the experiments must be augmented with these controls to substantiate the advantage.

    Authors: The referee correctly identifies gaps in the experimental section. We will expand it to include full dataset descriptions (synthetic functions chosen to highlight non-classical behavior), error bars from multiple runs, statistical significance tests (e.g., t-tests against baselines), and direct comparisons to classical stochastic neurons obtained by adding controlled noise to the classical activation outputs. These additions will allow proper evaluation of whether the observed learning advantage exceeds what stochastic classical models can achieve. revision: yes

  3. Referee: [BQP-completeness] BQP-completeness section: the decision problem must be stated precisely (e.g., whether it concerns expectation values, sampling, or a promise problem) and shown to remain BQP-complete while reducing to an efficiently simulable classical problem in the commuting limit. The proof should address whether the stochastic output in the commuting case permits efficient classical simulation, as this is load-bearing for the complexity separation.

    Authors: We will revise the BQP section to state the decision problem precisely: given a Fermi-Dirac neuron (with promise on the Hamiltonian parameters and input state), decide whether the expectation value of the activation observable exceeds a threshold heta (a promise problem). The proof will show BQP-completeness via reduction from a known BQP-complete problem (e.g., approximating ground-state energies or similar). In the commuting limit, the problem reduces to classical evaluation of the activation on eigenvalues, which is in P; the stochastic output is addressed by noting that the decision concerns the expectation (estimable in BPP classically via sampling), while the quantum case requires BQP. The expanded proof will explicitly demonstrate the commuting reduction and why the separation holds. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; derivation is self-contained via standard quantization and external complexity result.

full rationale

The paper defines its quantization by direct substitution of operators for classical variables, following the standard canonical quantization procedure cited as external. The commuting-operator reduction is asserted as an exact property of this substitution rather than a fitted or self-referential result. The BQP-completeness claim is presented as a separate proof for a defined decision problem, with no equations or self-citations shown reducing the central claims (hybrid algorithms, numerical separation, or complexity result) to tautological inputs. No load-bearing step matches any enumerated circularity pattern.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

The abstract supplies insufficient detail to enumerate specific free parameters, axioms, or invented entities. The construction relies on the standard canonical quantization procedure and the existence of efficient hybrid algorithms for the chosen Hamiltonians, but no explicit counts or values are stated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Fermi-Dirac machines as quantizations of neurons." pith.science (2026). https://pith.science/paper/3J23RP7M

@misc{pith2026260524386,
  author       = {Pith},
  title        = {Pith review of: Fermi-Dirac machines as quantizations of neurons},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3J23RP7M}},
  note         = {Machine review of arXiv:2605.24386}
}
read the original abstract

Fermi-Dirac machines were proposed recently as an approach to solving semidefinite optimization problems on quantum computers. Here, we reinterpret them as canonical quantizations of classical neurons. By viewing a classical neuron as an activation function applied to a parameterized classical Hamiltonian, we quantize this model by replacing classical variables with operators whose eigenvalues encode their possible values. This follows the standard approach to canonical quantization in quantum mechanics. Crucially, when the Hamiltonian consists of commuting operators, our construction reduces exactly to a classical neuron. More generally, our approach yields an activation observable, defined as an activation function applied to a parameterized quantum Hamiltonian. The output of this quantized neuron is a random variable with expectation value equal to that of the activation observable with respect to an input state. We develop efficient hybrid quantum-classical algorithms for evaluating outputs and gradients of our quantized neurons, enabling evaluation and training. These algorithms rely on basic primitives that include random sampling, Hamiltonian simulation, and the Hadamard test. We also quantize a whole host of other activation functions, including the smooth rectified linear unit (ReLU), sigmoid linear unit, Gaussian-smoothed ReLU, and Gaussian error linear unit (GeLU), which are known to be useful for deep learning applications. Numerical experiments indicate that neurons based on quantum Hamiltonians can learn functions that classical neurons cannot. We further define a computational decision problem based on Fermi-Dirac neurons and prove that it is BQP-complete, providing complexity-theoretic evidence against efficient classical simulation. Finally, we generalize our approach to continuous quantum variables and sketch two different ways of composing these neurons into networks.

Figures

Figures reproduced from arXiv: 2605.24386 by the authors.

Figure 1
Figure 1. FIG. 1: This figure illustrates one of the main conceptual contributions of our paper. We quantize various common [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 3
Figure 3. FIG. 3: Quantum circuit used in the quantum convolution [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figure 2
Figure 2. FIG. 2: (a) Quantum circuit used in Algorithm [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: FIG. 4: Quantum circuit for simulating the firing of a [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: FIG. 5: Quantum circuit used in the quantum convolution [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 6
Figure 6. Figure 6: FIG. 6: This figure is similar to Figure [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]
Figure 7
Figure 7. Figure 7: FIG. 7: Results of function-approximation experiments for training using squared-loss minimization. The quantum [PITH_FULL_IMAGE:figures/full_fig_p020_7.png]
Figure 8
Figure 8. Figure 8: FIG. 8: Results of binary-classification experiments for training using logistic-loss minimization. The quantum Hamiltonian [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]
Figure 9
Figure 9. Figure 9: FIG. 9: Comparison of classical, quantum linear, and quantum nonlinear models for function approximation. Quantum linear [PITH_FULL_IMAGE:figures/full_fig_p023_9.png]
Figure 10
Figure 10. Figure 10: FIG. 10: Quantum circuit used in Algorithm [PITH_FULL_IMAGE:figures/full_fig_p041_10.png]
Figure 11
Figure 11. Figure 11: FIG. 11: Quantum circuit used in Algorithm [PITH_FULL_IMAGE:figures/full_fig_p046_11.png]
Figure 12
Figure 12. Figure 12: FIG. 12: Quantum circuit used in Algorithm [PITH_FULL_IMAGE:figures/full_fig_p059_12.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Canonical quantization of neurons

    quant-ph 2026-07 conditional novelty 7.0 of 10

    Canonical quantization turns a neuron into an activation observable of a parameterized Hamiltonian, with hybrid algorithms for training on quantum data and numerics showing advantage over classical Ising neurons.

  2. Unifying quantum measurement constructions via a relative-entropy minimum change principle

    quant-ph 2026-08 conditional novelty 6.0 of 10

    A relative-entropy minimum change principle yields a unified closed-form family of optimal measurements, including pretty good, Fermi-Dirac thermal, and new softmin thermal measurements.

  3. Quantum Spectral Anomaly Detection

    quant-ph 2026-07 conditional novelty 6.0 of 10

    QSPADE defines a smooth, temperature-controlled spectral anomaly detector on the average quantum state that recovers hard PCA scores in the zero-temperature limit and calibrates with dimension-independent sample complexity.

Reference graph

Works this paper leans on

61 extracted references · 61 canonical work pages · cited by 3 Pith papers

  1. [1]

    Proof of Theorem 2 (derivation of formula for the objective function) We can also rewrite the objective function itself in terms of its partial derivatives, by employing Theorem 1 and the fundamental theorem of calculus, leading to the following theorem. Theorem(Restatement of Theorem 2).The following equality holds: Tr[gT (H(θ))ρ] = ∥θ∥1 T Ej∼q,t∼µ, s∼υ,...

  2. [2]

    Proof of Theorem 3 (expected value of quantum convolution algorithm) Ifqis a probability density, then|q⟩is a state vector because ⟨q|q⟩= Z ∞ −∞ dp′p q(p′)⟨p′| Z ∞ −∞ dp p q(p)|p⟩ (A49) = Z ∞ −∞ dp′ Z ∞ −∞ dp p q(p′)q(p)⟨p′|p⟩(A50) = Z ∞ −∞ dp′ Z ∞ −∞ dp p q(p′)q(p)δ(p−p ′)(A51) = Z ∞ −∞ dp q(p)(A52) = 1.(A53) Let us first suppose that the state of the da...

  3. [3]

    sgn(θk)ℜ

    Proof of Theorem 4 and Equation(41) Recall thatT 1, T2 >0such thatT /2 =T 1T2, ℓT1(p) := ep/T1 T1 (ep/T1 + 1)2 ,(A77) r(p) =1 p≥0 −1 p<0,(A78) and observe that d dp fT1(p) =ℓ T1(p),(A79) wheref T1(p) = e−p/T1 + 1 −1 is the Fermi–Dirac function defined in (2). Consider that (ℓT1 ∗r)(p/T 2) = (ℓT1 ∗1 p≥0)(p/T2)−(ℓ T1 ∗1 p<0)(p/T2)(A80) = Z ∞ −∞ dp′ 1p′≥0ℓT1...

  4. [4]

    Setℓ←1, and set L←O ∥θ∥1 (Hmax +y max)H max T 2ε 2 ln 1 δ ! ,(B8) whereH max := max j∈[J] ∥Hj∥,y max := max m∈[M] |ym|,ε >0is the desired accuracy, andδ∈ (0,1)is the desired failure probability

  5. [5]

    Samplet 1, t2 ∼µ,s 1, s2 ∼υ,k∼q,λ∼υ, andm∼[M], where the probability densitiesµ andυare defined in Theorem 1,qis defined in Theorem 2, andm∼[M]indicates thatmis selected uniformly at random from[M]

  6. [6]

    41 Had|0⟩⟨0| Had ρm σZ Hj e−iH(θ)s2t2/T eiH(θ)t2/T Had|0⟩⟨0| Had ρm σZ Hk e−iH(k,λ)s1t1/T eiH(k,λ)t1/T FIG

    Prepare the statesU H(k,λ) s1t1/T (ρm)andU H(θ) s2t2/T (ρm)using two samples ofρ m and Hamiltonian simulation to realize the unitary channelsUH(k,λ) s1t1/T andU H(θ) s2t2/T. 41 Had|0⟩⟨0| Had ρm σZ Hj e−iH(θ)s2t2/T eiH(θ)t2/T Had|0⟩⟨0| Had ρm σZ Hk e−iH(k,λ)s1t1/T eiH(k,λ)t1/T FIG. 10: Quantum circuit used in Algorithm 8 for estimating thejth partial deriv...

  7. [7]

    Set Yℓ ← 2∥θ∥ 1 T 2 sgn(θk) X (1) ℓ −y m ·X (2) ℓ ·Z (1) ℓ ·Z (2) ℓ .(B9) Setℓ←ℓ+ 1

    Perform the quantum circuit depicted in Figure 10, with measurement outcomesZ(1) ℓ , Z(2) ℓ ∈ {−1,1}for theσ Z measurements,X (1) ℓ ∈spec(H k)for theH k measurement, andX (2) ℓ ∈ spec(Hj)for theH j measurement. Set Yℓ ← 2∥θ∥ 1 T 2 sgn(θk) X (1) ℓ −y m ·X (2) ℓ ·Z (1) ℓ ·Z (2) ℓ .(B9) Setℓ←ℓ+ 1

  8. [8]

    Compute the averageYL := 1 L PL ℓ=1 Yℓ and output this value as an estimate of ∂ ∂θj L(θ)

    Repeat Steps 2-4L−1more times. Compute the averageYL := 1 L PL ℓ=1 Yℓ and output this value as an estimate of ∂ ∂θj L(θ). Figure 10 depicts the quantum circuit used in Algorithm 8. By the Hoeffding inequality, we are guaranteed that Pr YL − ∂ ∂θj L(2)(θ) ≤ε ≥1−δ.(B10) Appendix C: Logistic-loss function for binary classification

Show all 61 references
  1. [9]

    Lemma 2.Letx∈R, and letx7→A(x)be a Hermitian-valued function

    Derivative of matrix logistic-loss function We begin by deriving a novel formula for the derivative of the matrix logistic-loss functionA7→ ln 1 +e −A , whereAis a Hermitian matrix. Lemma 2.Letx∈R, and letx7→A(x)be a Hermitian-valued function. Then the following equality holds...

  2. [10]

    Hybrid quantum–classical algorithm for estimating derivative of matrix logistic-loss function The first term of (63) can be easily estimated by preparing the stateρand measuringHj. Under the assumption that eachHj inH(θ)is both Hermitian and unitary (as in the common case when...

  3. [11]

    Setm←1, and set M←O ∥θ∥1 maxj∈[J] ∥Hj∥ T ε 2 ln 1 δ ! ,(C42) whereε >0is the desired accuracy andδ∈(0,1)is the desired failure probability

  4. [13]

    Prepare the stateUymH(θ) st/T (ρm)using one sample ofρ m and Hamiltonian simulation to realize the unitary channelU ymH(θ) st/T

  5. [15]

    Compute the averageWM := 1 M PM m=1 Wm and output this value as an estimate ofζj

    Repeat Steps 2-4M−1more times. Compute the averageWM := 1 M PM m=1 Wm and output this value as an estimate ofζj. By the Hoeffding inequality, we are guaranteed that Pr WM −ζ j ≤ε ≥1−δ.(C43)

  6. [16]

    46 Had|0⟩⟨0| Had ρ σZ Hj e−iymH(θ)st/T eiymH(θ)t/T Hk FIG

    Alternative formula for logistic-loss function Similar to the idea behind Theorem 2, we can use Theorem 5 and the fundamental theorem of calculus to derive an expression for the logistic-loss function that can be evaluated by a hybrid quantum–classical algorithm. 46 Had|0⟩⟨0| ...

  7. [17]

    Hybrid quantum–classical algorithm for estimating logistic-loss function The expression in (C44) then leads to a hybrid quantum–classical algorithm for estimating the logistic-loss objective function. The first term−ym 2 Tr[H(θ)ρm]in (C44) can be easily estimated by writing − ...

  8. [18]

    Setm←1, and set M←O Jθ ⋆h⋆ T ε 2 ln 1 δ ! ,(C80) θ⋆ := max λ∈[0,1],j∈[J] θ(j)(λ) 1 ,(C81) h⋆ := max j∈[J] ∥Hj∥,(C82) whereε >0is the desired accuracy andδ∈(0,1)is the desired failure probability

  9. [20]

    Prepare the stateUymH(θ (j) (λ)) st/T (ρm)using one sample ofρm and Hamiltonian simulation to realize the unitary channelU ymH(θ (j) (λ)) st/T

  10. [21]

    Set Wm ← J θ(j)(λ) 1 2T s·sgn(θ k)Zm ·X m.(C83) Setm←m+ 1

    Perform the quantum circuit depicted in Figure 11, with measurement outcomesZm ∈ {−1,1} for theσ Z measurement andX m ∈spec(H k)for theH k measurement. Set Wm ← J θ(j)(λ) 1 2T s·sgn(θ k)Zm ·X m.(C83) Setm←m+ 1

  11. [22]

    Compute the averageWM := 1 M PM m=1 Wm and output this value as an estimate ofζ

    Repeat Steps 2-4M−1more times. Compute the averageWM := 1 M PM m=1 Wm and output this value as an estimate ofζ. By the Hoeffding inequality, we are guaranteed that Pr WM −ζ ≤ε ≥1−δ.(C84) 50 Appendix D: Derivations for smooth rectified linear unit (ReLU)

  12. [23]

    Proof of Theorem 6 (derivative of smooth ReLU function) Recall from Theorem 5 that ∂ ∂θj TTr ln I+e −ymH(θ)/T ρm =− ym 2 Tr[Hjρm] + ∥θ∥1 2T Es∼υ, k∼q, t∼γ h sℜ h Tr h sgn(θk)HkHjeiymH(θ)t/T U ymH(θ) st/T (ρm) iii .(D1) Now sety m = 1andρ m =ρto get ∂ ∂θj TTr ln I+e −H(θ)/T ρ =...

  13. [24]

    Hybrid quantum–classical algorithms for smooth ReLU a. Hybrid quantum–classical algorithm for estimating gradient of smooth ReLU In this appendix, we briefly summarize a hybrid quantum–classical algorithm for estimating the formula in (71), i.e., thejth partial derivative ofTr...

  14. [25]

    Setm←1, and set M←O ∥θ∥1 maxj∈[J] ∥Hj∥ T ε 2 ln 1 δ ! ,(D9) whereε >0is the desired accuracy andδ∈(0,1)is the desired failure probability

  15. [26]

    Samples∼υ,k∼q, andt∼γ

  16. [27]

    Prepare the stateUH(θ) st/T (ρm)using one sample ofρand Hamiltonian simulation to realize the unitary channelU H(θ) st/T

  17. [28]

    SetW m ← ∥θ∥1 2T s· sgn(θk)Zm ·X m

    Perform the quantum circuit depicted in Figure 11, with measurement outcomesZm ∈ {−1,1} for theσ Z measurement andX m ∈spec(H k)for theH k measurement. SetW m ← ∥θ∥1 2T s· sgn(θk)Zm ·X m. Setm←m+ 1

  18. [29]

    Compute the averageWM := 1 M PM m=1 Wm and output this value as an estimate ofζj

    Repeat Steps 2-4M−1more times. Compute the averageWM := 1 M PM m=1 Wm and output this value as an estimate ofζj. By the Hoeffding inequality, we are guaranteed that Pr WM −ζ j ≤ε ≥1−δ.(D10) b. Hybrid quantum–classical algorithm for estimating smooth ReLU In this appendix, we d...

  19. [30]

    Setm←1, and set M←O Jθ ⋆h⋆ T ε 2 ln 1 δ ! ,(D21) θ⋆ := max λ∈[0,1],j∈[J] θ(j)(λ) 1 ,(D22) h⋆ := max j∈[J] ∥Hj∥,(D23) whereε >0is the desired accuracy andδ∈(0,1)is the desired failure probability

  20. [31]

    Samplejaccording to the uniform distribution on[J],s∼υ,λ∼υ,t∼γ, andk∼q j,λ

  21. [32]

    Prepare the stateUH(θ (j) (λ)) st/T (ρm)using one sample ofρ m and Hamiltonian simulation to realize the unitary channelU H(θ (j) (λ)) st/T

  22. [33]

    Set Wm ← J θ(j)(λ) 1 2T s·sgn(θ k)Zm ·X m.(D24) Setm←m+ 1

    Perform the quantum circuit depicted in Figure 11 (withy m set to1), with measurement outcomesZ m ∈ {−1,1}for theσ Z measurement andX m ∈spec(H k)for theH k measurement. Set Wm ← J θ(j)(λ) 1 2T s·sgn(θ k)Zm ·X m.(D24) Setm←m+ 1

  23. [34]

    Compute the averageWM := 1 M PM m=1 Wm and output this value as an estimate ofζ

    Repeat Steps 2-4M−1more times. Compute the averageWM := 1 M PM m=1 Wm and output this value as an estimate ofζ. By the Hoeffding inequality, we are guaranteed that Pr WM −ζ ≤ε ≥1−δ.(D25)

  24. [35]

    Proof of Theorem 7 (correctness of Algorithm 5 for smooth ReLU) Here we prove that rT (p) =T 2 (ReLU∗ℓ T1) p T2 ,(D26) for allp∈R, wherer T is defined in (68),ReLUis defined in (67), andℓT1 is defined in (38). Consider that, for alla∈RandT >0, (ReLU∗ℓ T ) (a) = Z ∞ −∞ dpReLU(p...

  25. [36]

    Proof of Theorem 8 (derivative of sigmoid linear unit function) Consider that SiLUT (x) := x 1 +e −x/T .(E1) The derivative of this function is given by kT (x) := ∂ ∂x x 1 +e −x/T (E2) = 1 1 +e −x/T + xe−x/T T(1 +e −x/T )2 (E3) =f T (x) +xℓ T (−x)(E4) =f T (x) +xℓ T (x).(E5) O...

  26. [37]

    Hybrid quantum–classical algorithms for SiLU a. Hybrid quantum–classical algorithm for estimating gradient of SiLU In this appendix, we briefly summarize a hybrid quantum–classical algorithm for estimating the formula in (76), i.e., thejth partial derivative ofTr[SiLUT (H(θ))ρ...

  27. [38]

    Setm←1, and set M←O ∥θ∥1 maxj∈[J] ∥Hj∥ T ε 2 ln 1 δ ! ,(E44) whereε >0is the desired accuracy andδ∈(0,1)is the desired failure probability

  28. [39]

    Samples∼υ,k∼q, andt∼ξ

  29. [40]

    Prepare the stateUH(θ) st/(2T) (ρ)using one sample ofρand Hamiltonian simulation to realize the unitary channelU H(θ) st/(2T)

  30. [41]

    SetW m ← ∥θ∥1 2T s· sgn(θk)Zm ·X m

    Perform the quantum circuit depicted in Figure 12, with measurement outcomesZm ∈ {−1,1} for theσ Z measurement andX m ∈spec(H k)for theH k measurement. SetW m ← ∥θ∥1 2T s· sgn(θk)Zm ·X m. Setm←m+ 1

  31. [42]

    Compute the averageWM := 1 M PM m=1 Wm and output this value as an estimate ofζj

    Repeat Steps 2-4M−1more times. Compute the averageWM := 1 M PM m=1 Wm and output this value as an estimate ofζj. By the Hoeffding inequality, we are guaranteed that Pr WM −ζ j ≤ε ≥1−δ.(E45) b. Hybrid quantum–classical algorithm for estimating SiLU In this appendix, we detail a...

  32. [43]

    Setm←1, and set M←O Jθ ⋆h⋆ T ε 2 ln 1 δ ! ,(E56) θ⋆ := max λ∈[0,1],j∈[J] θ(j)(λ) 1 ,(E57) h⋆ := max j∈[J] ∥Hj∥,(E58) whereε >0is the desired accuracy andδ∈(0,1)is the desired failure probability

  33. [44]

    Samplejaccording to the uniform distribution on[J],s∼υ,λ∼υ,t∼ξ, andk∼q j,λ

  34. [45]

    Prepare the stateUH(θ (j) (λ)) st/(2T) (ρ)using one sample ofρand Hamiltonian simulation to realize the unitary channelU H(θ (j) (λ)) st/(2T)

  35. [46]

    Set Wm ← J θ(j)(λ) 1 2T s·sgn(θ k)Zm ·X m.(E59) Setm←m+ 1

    Perform the quantum circuit depicted in Figure 12, with measurement outcomesZm ∈ {−1,1} for theσ Z measurement andX m ∈spec(H k)for theH k measurement. Set Wm ← J θ(j)(λ) 1 2T s·sgn(θ k)Zm ·X m.(E59) Setm←m+ 1

  36. [47]

    Compute the averageWM := 1 M PM m=1 Wm and output this value as an estimate ofζ

    Repeat Steps 2-4M−1more times. Compute the averageWM := 1 M PM m=1 Wm and output this value as an estimate ofζ. By the Hoeffding inequality, we are guaranteed that Pr WM −ζ ≤ε ≥1−δ.(E60)

  37. [48]

    As in Appendix A3, let us first suppose that the state of the data register is pure and given by|φ⟩ ⟨φ|, where|φ⟩is a state vector

    Proof of Theorem 9 (expected value of quantum convolution and multiplication algorithm) Following the same reasoning as in (A49)–(A53), we conclude that, fori∈ {1,2},|qi⟩is a state vector ifq i is a probability density. As in Appendix A3, let us first suppose that the state of...

  38. [49]

    Proof of Theorem 10 (correctness of Algorithm 7 for SiLU) This is an application of Theorem 9. Given that the first control qumode is prepared in the vacuum state, its representation in the momentum eigenbasis is |0⟩= Z dp p ϕ(p)|p⟩,(E83) 64 ϕ(p) := e−p2 √π .(E84) This means t...

  39. [50]

    Proof of Equation(90) Our goal is to prove (90), which we recall here: (ϕT1 ∗r)(p/T 2) = 2Φ 2p T −1 = : erf √ 2p T ! ,(F3) for allp∈R, withrset as in (39) andT /2 =T 1T2. Following the same approach as in (A80)–(A90) and noting that Φ p T1 = Z p −∞ dp′ ϕT1(p′),(F4) consider th...

  40. [51]

    erf √ 2H(θ) T ! ρ # = Tr

    Gradient of expectation of Gaussian error function activation observable In this appendix, we derive a formula for the gradient of∂ ∂θj Tr h erf √ 2H(θ) T ρ i that is useful for estimation on quantum computers. The development here mirrors that in Appendix A1. We begin by deri...

  41. [52]

    Proof of Equation(94) Our goal is to prove (94), which we recall here: GReLUT (x) :=xΦ x T +T ϕ x T (F36) =T 2 (ReLU∗ϕ T1) x T2 ,(F37) whereT=T 1T2. Consider that, for alla∈RandT >0, (ReLU∗ϕ T ) (a) = Z ∞ −∞ dpReLU(p)ϕ T (a−p)(F38) = Z ∞ −∞ dpReLU(p)ϕ T (p−a)(F39) = 1√ 2πT Z ∞...

  42. [53]

    Z 1 0 ds X k,ℓ (sλk + (1−s)λ ℓ)e −i(sλk+(1−s)λℓ)vtΠk ∂ ∂x A(x) Πℓ # (F75) =E t∼νT , v∼υ

    Derivative of Gaussian smoothed rectified linear unit (GReLU) We begin by proving (96). Consider that ∂ ∂x [GReLUT (x)] = ∂ ∂x h xΦ x T +T ϕ x T i (F54) = Φ x T +x ∂ ∂x Φ x T +T ∂ ∂x ϕ x T (F55) = Φ x T + x T ϕ x T −T x T ϕ x T 1 T (F56) = Φ x T .(F57) Lemma 5.The following eq...

  43. [54]

    74 Lemma 8.Letx∈R, letx7→A(x)be a Hermitian-valued function, and letT >0

    Derivative of Gaussian error linear unit (GeLU) Lemma 7.The following equality holds: ∂ ∂x h xΦ x T i = 1 2 + r 2 π x T Et∼νT ,v∼κ e−ixvt ,(F95) whereν T (t)is the following Gaussian probability density function ont∈R: νT (t) := T√ 2π e−T 2t2/2,(F96) 73 andκis the following pr...

  44. [55]

    φ2 J1X j′=1 θ(2) kj ′A(1) j′ ! ρ # = Tr h DB(2) k (θ(2)) A(1) j ρ i ,(G1) ∂ ∂θ (1) ji Tr

    Two-layer gradient formulas Inthisappendix, weshowhowtocalculatethegradientforatwo-layerquantumobservablenetwork of the form in (143)–(146), specializing to the case whenφ1 andφ 2 are bothtanh. Theorem 19.The following formulas hold for the partial derivatives of a quantum obs...

  45. [56]

    DB(3) k (θ(3)) ∂ ∂θ (2) j2j1 B(3) k θ(3) ! ρ # (G37) = Tr  DB(3) k (θ(3))   ∂ ∂θ (2) j2j1 J2X j′ 2=1 θ(3) kj ′ 2 A(2) j′ 2 θ(2)   ρ   (G38) = J2X j′ 2=1 θ(3) kj ′ 2 Tr

    Three-layer gradient formulas Now we consider the three-layer case, where the objective function is Tr h φ3 B(3) k θ(3) ρ i ,(G24) and B(3) k θ(3) := J2X j2=1 θ(3) kj2A(2) j2 θ(2) ,(G25) A(2) j2 θ(2) :=φ 2 B(2) j2 θ(2) ,(G26) B(2) j2 θ(2) := J1X j1=1 θ(2) j2j1A(1) j1 θ(1) (G27...

  46. [57]

    , θJ)) = Ξ(η∥ρ(θ1, θ2,

    Proof of Theorem 12 (derivation of formula for cross entropy) Using the notation in Theorem 12, consider that Ξ(η∥ρ(θ))−Ξ(η∥ρ(0, θ 2, . . . , θJ)) = Ξ(η∥ρ(θ1, θ2, . . . , θJ))−Ξ(η∥ρ(0, θ 2, . . . , θJ))(I1) 85 = Ξ(η∥ρ(θ (1)(1)))−Ξ(η∥ρ(θ (1)(0)))(I2) = Z 1 0 dλ1 ∂ ∂λ1 Ξ(η∥ρ(θ(1...

  47. [58]

    Hybrid quantum–classical algorithm for cross entropy and log partition function estimation This leads to the following hybrid quantum–classical algorithm for estimating the cross entropy: Algorithm 15.A hybrid quantum–classical algorithm for estimating the cross entropyΞ(η∥ρ(θ...

  48. [59]

    Setk←1, and set K←O ∥θ∥1 maxj∈[J] ∥Hj∥ ε 2 ln 1 δ ! ,(I25) whereε >0is the desired accuracy andδ∈(0,1)is the desired failure probability

  49. [60]

    Prepare the stateηand the thermal stateρ(θ(j)(λ))

  50. [61]

    Set Wk ← ∥θ∥ 1 ·sgn(θ j) (Xη k −X ρ k)

    Measure the observableHj on each state, with measurement outcomesXη k , Xρ k ∈spec(H j). Set Wk ← ∥θ∥ 1 ·sgn(θ j) (Xη k −X ρ k). Setk←k+ 1

  51. [62]

    Compute the average WK := 1 K PK k=1 Wk and output lnd+ WK as an estimate ofΞ(η∥ρ(θ))

    Repeat Steps 2-4K−1more times. Compute the average WK := 1 K PK k=1 Wk and output lnd+ WK as an estimate ofΞ(η∥ρ(θ)). By the Hoeffding inequality, we are guaranteed that Pr lnd+ WK −Ξ(η∥ρ(θ)) ≤ε ≥1−δ.(I26) Let us finally note that one can estimate the log-partition functionlnZ...

  52. [63]

    Prepare the thermal stateρ(θ(j)(λ))

  53. [64]

    Set Wk ← − ∥θ∥ 1 ·sgn(θ j)X ρ k

    Measure the observableH j on this state, with measurement outcomeX ρ k ∈spec(H j). Set Wk ← − ∥θ∥ 1 ·sgn(θ j)X ρ k. Setk←k+ 1

Pith tools

Reviewed June 30, 2026 · model on record in the stance chip above.