{"id":"47050ae9-cc7c-4003-a2f1-cf463f11908d","arxiv_id":"2504.12632","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A fixed set of four linear QAOA angle coefficients trained on one random Ising instance transfers to other instances with only a small loss in approximation ratio, eliminating per-instance optimization.","lead":"This paper shows that QAOA parameters for random Ising problems can be constrained to four linear coefficients and then reused on new problem instances with zero per-instance classical optimization. It is worth reading because it offers a practical way to cut the classical overhead of running QAOA on near-term quantum hardware, at the cost of a few percentage points of solution quality.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 80-qubit transferability result depends on an ESA-based normalization that is unavailable in practice, the suggested proxies are untested, and the large-instance experiment transfers a toy schedule rather than optimized parameters.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the large-system transfer result in Section III-C is conditioned on an energy-scale normalization that uses the target instance's ground-state energy, which is not known in practice, and the paper never validates a practical proxy. This is the point at which the strongest claim about zero-cost transferability and >0.8 approximation ratios up to 80 qubits is least secure. I additionally note that the large-instance experiment uses the unoptimized toy schedule of Eq. (6) rather than a parameter set trained on a source instance, so it isolates the effect of normalization more than it isolates transferability. The paper deserves credit for explicitly flagging the ESA limitation and for honestly reporting the 16-qubit comparisons, where LINXFER is competitive but slightly worse than INTERP/FOURIER/standard QAOA. However, that smaller-scale evidence does not by itself support the practical large-system claim, which is exactly where the untested proxy sits. The appropriate verdict remains conditional, matching the reader's assessment, and the concrete test above would settle whether the normalization gap is fatal to the practical claim.","tokens_in":13833,"tokens_out":4629,"duration_ms":51624,"concrete_test":"Recompute the Fig. 7 experiment for n_qubits in {32, 48, 64, 80} using the Goemans-Williamson SDP bound, or a strictly computable proxy such as (1/sqrt(n_edges)) * sum |J_ij|, in place of exact ESA for the normalization, keeping Eq. (6), p = 8, and the same MPS settings. If the <E>/ESA curves drop materially below 0.8, or E_best/ESA degrades by more than about 5%, the practical transferability claim is not supported. As a second arm, transfer the optimized 16-qubit parameter set Eq. (5) to the same large instances with both normalizations, to separate true parameter transfer from the effect of energy-scale rescaling.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section III-C's headline scaling result applies the toy parameters of Eq. (6) after normalizing the target Hamiltonian by |ESA|/sqrt(n_edges), where ESA is the target instance's quasi ground-state energy obtained from simulated annealing. Because the QAOA phase operator is e^{-i*gamma*H}, this normalization is equivalent to choosing the physical gamma schedule as gamma_l = (-l/p - 1)*sqrt(n_edges)/|ESA| and the beta schedule as beta_l = -l/p + 1. The approximation ratios above 0.8 in Fig. 7 are therefore contingent on knowing |ESA| for each new instance. The paper acknowledges ESA is not known in practice and suggests the Goemans-Williamson bound or n_edges as a proxy, but it reports no test of either proxy against exact ESA. If a practical proxy cannot reproduce the normalization, the 'zero-cost application to new problems' claim fails exactly in the large-system regime where it is most valuable. In addition, Eq. (6) is not a parameter set optimized on any source instance, so this experiment does not actually demonstrate transfer of learned parameters; it demonstrates that a hand-chosen linear schedule can perform well once rescaled by the optimum energy. The smaller-instance comparisons in Section III-A are more honest evidence for the transfer idea, but they do not reach 80 qubits.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes constraining QAOA parameters to linear functions of the layer index, reducing the parameter space to four dimensions regardless of depth (LINXFER), and claims these parameters transfer between problem instances without instance-specific optimization. The authors compare LINXFER against standard QAOA, INTERP, and FOURIER on 16-qubit random Ising instances using noiseless simulation; study the structure of the reduced cost landscape on both a simulator and IBM's Eagle processor; investigate energy-scale normalization as a mechanism for transferability; and report approximation ratios above 0.8 for random Ising instances up to 80 qubits using a toy linear schedule normalized by |ESA|/sqrt(n_edges). They also provide cross-problem landscape evidence for SK models and max-cut.","tokens_in":14075,"tokens_out":3501,"duration_ms":37639,"significance":"If the transferability claim holds, the paper would offer a practical reduction of QAOA's classical optimization overhead, with a clear route to deploying fixed-angle circuits on near-term hardware. The 16-qubit transfer experiment is genuinely predictive: parameters optimized on one reference instance are applied to fresh instances and compared with optimization-based baselines. The real-device landscape measurement is a useful NISQ-era datapoint. The paper is also unusually candid about its own limitations, explicitly noting that ESA is not known in practice and that the large-instance normalization is heuristic. Those caveats, however, are located precisely where the scalable-transfer claim is made, and they materially affect what the results establish.","major_comments":[{"comment":"The large-system transfer experiment is conditioned on normalizing the target Hamiltonian by |ESA|/sqrt(n_edges), where ESA is the simulated-annealing ground-state energy of each target instance. As the text acknowledges, ESA is not known in practice, and the suggested proxies (Goemans-Williamson energy or n_edges) are never tested against the exact ESA. Without such a test, the headline claim of zero-cost application to new instances in the large-system regime is not established; Fig. 6 shows that without this normalization the same parameters perform at the level of random sampling. The paper should report, for its 8-instance samples, the distribution of |ESA|/sqrt(n_edges) and how well the proposed proxies track it, or explicitly restrict the scalable-transfer claim to the oracle-normalized setting.","section":"§III-C, Eq. (6), Fig. 7"},{"comment":"The 80-qubit experiment does not actually transfer parameters that were learned on a source instance. Equation (6) is a hand-chosen toy schedule with all slopes and intercepts set to unit magnitude, and the text states it was never optimized for any problem instance. What Fig. 7 demonstrates is therefore that a particular fixed linear schedule, after rescaling with the target's exact ground-state energy, achieves high approximation ratios. It does not demonstrate transfer of the linearly optimized parameters that are the subject of the paper's central claim. To support the title and abstract, the large-system experiment should use a parameter set obtained by optimizing on a reference instance (e.g., Eq. (5) or a Bayesian-optimized set at the same p) and then apply the same normalization argument.","section":"§III-C, Eq. (6)"},{"comment":"The claim that LINXFER is within error bars of the other methods 'for most p cases' is hedged, but the p=16 data point is a clear exception: LINXFER achieves 0.91(2) while INTERP and FOURIER achieve 0.97(1) and 0.96(2), respectively, which are outside the reported error bars. This is a performance gap of roughly five to six percentage points at increased depth, and it is directly relevant to the paper's argument that LINXFER sacrifices only a few percentage points while eliminating classical overhead. The paper should discuss whether the gap grows with p and what that implies for the practical value of the zero-cost advantage on deeper circuits.","section":"Table I, p = 16 row"}],"minor_comments":[{"comment":"The sentence 'We sorely aim to demonstrate' should read 'We solely aim to demonstrate'.","section":"§III-A"},{"comment":"The objective function in Algorithm 1 is written as 'f (γslope,γintcp.,βslope,βintcp.): ⟨H⟩ using γ,β ← SET(...)', but the Hamiltonian H is not listed as an input; the pseudocode would be clearer if the Hamiltonian were passed as an argument.","section":"Algorithm 1, line 11"},{"comment":"The phrase 'an IBM's Eagle quantum processor' is grammatically awkward; 'an IBM Eagle quantum processor' or 'IBM's Eagle quantum processor' would be clearer.","section":"Figure 3 caption"},{"comment":"The sentence 'One do not necessarily know ESA' contains a grammatical error and should be rephrased, for example, 'One does not necessarily know ESA.'","section":"§III-C"},{"comment":"The sudden drop in optimization time for Standard QAOA and FOURIER at p=16 is attributed to stronger expressive power, but the explanation would be clearer if the footnote were expanded in the main text.","section":"Table II"}],"recommendation":"major_revision","confidential_remarks":"The paper's central claim is more cautiously worded in Section III-C than in the abstract and introduction. The large-system result is honest about its heuristic normalization, but the abstract's 'zero-cost application to new problems' is not supported unless the normalization can be obtained without oracle knowledge of the ground-state energy. I would encourage the authors to add the proxy tests and a transfer experiment using genuinely optimized reference parameters before accepting the scalable-transfer conclusion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Take: a useful empirical follow-up to the authors' own linear-QAOA transfer work, with honest 16-qubit evidence, but the headline 80-qubit scaling claim rests on an energy-scale normalization that users cannot actually apply and whose proposed proxies are untested.\n\nWhat's new: hardware data on IBM Eagle (landscape stability under noise), MPS simulations up to 80 qubits, and a first pass at energy-scale normalization. The 16-qubit comparison is the real contribution: LINXFER parameters optimized once on a single reference instance do transfer to fresh instances roughly within error bars of INTERP/FOURIER for most depths, and the time savings are substantial (0s vs hundreds or thousands of seconds). The paper also reports honestly that at p=16 LINXFER falls clearly below the others (0.91 vs 0.96–0.97), and it includes a bond-dimension convergence check in Appendix B.\n\nThe soft spot is Section III-C. The 'transfer to 80 qubits' figure uses toy parameters (Eq. 6) that were never optimized on any source instance, so this is not a test of transferring learned parameters. More importantly, the good approximation ratios appear only after normalizing the target Hamiltonian by |ESA|/sqrt(n_edges), where ESA is the target instance's quasi ground-state energy from simulated annealing. The paper acknowledges ESA is not known in practice and suggests the Goemans–Williamson bound or n_edges as proxies, but never tests either. Without a working proxy, the zero-cost claim fails precisely where it matters most. The authors do call this 'super heuristic' and mention iterating on the normalization factor, but the scale-up result as presented is not yet evidence for practical zero-cost deployment.\n\nMinor issues: no code or data are shipped, which hinders reproducibility; the hardware section is thin — a landscape comparison, not a transfer experiment on device.\n\nOverall, the central small-system claim is supported. The 16-qubit transfer evidence is predictive and fairly reported, and the authors are candid about the normalization caveat. But the large-system transferability and the 'zero-cost application to new problems' statements are oversold given the ESA dependence.\n\nRecommendation: send to peer review, with a request that the authors either test a practical proxy for ESA or substantially soften the scaling claims. If a proxy works, this becomes a practically useful result. If not, the contribution is the honest small-system comparison plus the hardware landscape observations, which are still worth publishing with clearer limits.","headline":"Honest small-system transfer evidence, but the 80-qubit scaling claim leans on an energy normalization that is unavailable in practice.","tokens_in":14677,"tokens_out":2733,"would_cite":false,"duration_ms":27626,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims QAOA's 2p angle parameters collapse to four linear coefficients that transfer across problem instances at zero optimization cost, with energy-scale normalization as the key.","keywords":["QAOA","parameter transfer","linear parameterization","energy scale normalization","random Ising model","approximation ratio","NISQ devices","cost landscape"],"falsifier":"Run the transfer on a fresh random Ising instance of 32 to 80 qubits, estimating the energy scale from data available before solving (edge count, mean coupling magnitude, or the semidefinite relaxation bound) instead of the exact ground-state energy, then apply the paper's toy linear angles and compare the approximation ratio on an MPS simulator. If the ratio systematically falls toward the random-sampling baseline or below 0.8, the zero-cost transfer claim fails in exactly the practical setting the paper targets.","tokens_in":13577,"feed_emoji":"⚛️","tokens_out":9927,"duration_ms":93642,"temperature":0.7,"pith_summary":"This paper argues that the $2p$ angle parameters of QAOA can be replaced by just four numbers—a slope and an intercept for the cost angle and for the mixer angle—without losing much solution quality. It presents evidence that established layer-by-layer parameter-setting schemes produce angles that lie close to such linear schedules, and that the four-parameter cost landscapes look similar across random spin-glass instances, so a parameter set tuned on one instance can be applied to others with no per-instance optimization. The paper identifies energy scale as the dominant factor controlling this transfer: after normalizing the Hamiltonian by $|E_{\\mathrm{SA}}|/\\sqrt{n_{\\mathrm{edges}}}$, even a hand-set linear schedule reaches approximation ratios typically above 0.8 for random Ising instances up to 80 qubits. It also reports that the landscape structure survives on real quantum hardware, suggesting the approach is compatible with noisy devices. If the claim holds, the practical payoff is that QAOA can be deployed on new instances at zero classical optimization cost, with all the cost pushed into a one-time pre-training.","feed_headline":"Four QAOA angles replace 2p, and they transfer for free","feed_subtitle":"With energy-scale normalization, unoptimized linear angles hit ratios above 0.8 on 80-qubit spin-glass instances.","key_machinery":"The central object is the linear-parameter constraint, called LINXFER in the paper, which fixes each layer's angles as $\\gamma_l=\\gamma_{\\mathrm{slope}}(l/p)+\\gamma_{\\mathrm{intcp.}}$ and $\\beta_l=\\beta_{\\mathrm{slope}}(l/p)+\\beta_{\\mathrm{intcp.}}$, leaving four numbers regardless of the number of layers $p$. It is supported by two observations: in the $p\\to\\infty$ limit QAOA converges to adiabatic evolution, where linear schedules are natural, and the parameters produced by INTERP and FOURIER are empirically close to linear. The mechanism that carries the transfer argument is energy-scale normalization: normalizing the Hamiltonian by a factor rescales the $\\gamma$-axis of the landscape, effectively zooming the four-parameter cost surface so that parameters tuned for one instance line up with the optimum of another. The paper also uses matrix-product-state simulation to reach system sizes beyond state-vector simulation, with the bond dimension fixed at 32 after a convergence check.","core_discovery":"The central discovery is that QAOA's optimizable angles can be constrained to strict linear functions of the layer index, $\\gamma_l=\\gamma_{\\mathrm{slope}}(l/p)+\\gamma_{\\mathrm{intcp.}}$ and $\\beta_l=\\beta_{\\mathrm{slope}}(l/p)+\\beta_{\\mathrm{intcp.}}$, collapsing the variational search from $2p$ dimensions to four. Empirically, the angles recovered by standard QAOA, INTERP, and FOURIER lie near such straight lines, and the four-parameter landscape has a consistent structure across random Ising instances of varying size and connectivity. A parameter set optimized once on one 16-qubit instance transfers to other instances with only a few percentage points of approximation-ratio loss, and with energy-scale normalization an unoptimized toy schedule ($\\gamma_l=-l/p-1$, $\\beta_l=-l/p+1$) attains $\\langle E\\rangle/E_{\\mathrm{SA}}\\gtrsim 0.8$ for random Ising instances up to 80 qubits. The same landscape pattern remains recognizable on a real superconducting processor under hardware noise. The paper concludes that energy scale is the main knob for transfer: scaling the Hamiltonian acts as a zoom on the landscape, so matching the source and target energy scales is what makes zero-cost transfer work.","pith_inferences":["The paper's normalization uses the exact ground-state energy $E_{\\mathrm{SA}}$, which it admits is unknown for a fresh instance; a direct testable extension is to replace it with a cheap classical estimate (edge count, mean coupling magnitude, or the semidefinite relaxation bound) and measure how much the transferred approximation ratio drops.","Because a strict linear schedule is a very constrained ansatz, the natural next question is whether a piecewise-linear schedule with a few breakpoints (say 8 parameters) preserves transferability while closing the residual gap to fully optimized QAOA at fixed depth.","The observation that coupling noise in the spin-glass model behaves like hardware noise on the $R_{ZZ}$ gates suggests normalization could double as a crude error-mitigation knob: rescaling the Hamiltonian to suppress landscape distortion before sampling on noisy hardware."],"forward_implications":["New random Ising instances can be attacked with zero per-instance classical optimization: pre-trained or even hand-set linear angles are plugged in directly, so the cost no longer grows with circuit depth.","The same fixed linear schedule, normalized by the energy scale, achieves typical approximation ratios above 0.8 on random Ising instances up to 80 qubits, a size where direct classical optimization of the angles is impractical.","Compared with INTERP and FOURIER, the classical overhead drops by orders of magnitude for deep circuits, at the cost of at most a few percentage points in solution quality.","Because the landscape structure persists under hardware noise, transferable linear parameters remain usable on current noisy quantum processors without error correction.","Similar landscape patterns appear across spin-glass and max-cut problems, so transfer need not be limited to instances of the same distribution; adjusting the energy scale is the main requirement."],"supporting_citations":[{"why":"Introduces QAOA and its alternating cost-mixer ansatz, the object the paper constrains.","marker":"[1]"},{"why":"Shows quantum-annealing initialization of QAOA parameters, providing the adiabatic-schedule motivation for linear ramps.","marker":"[6]"},{"why":"Defines INTERP and FOURIER; the paper shows their optimized angles are nearly linear and uses them as baselines.","marker":"[7]"},{"why":"Prior demonstration of parameter transfer for weighted MaxCut, grounding the transferability idea.","marker":"[8]"},{"why":"Shows that energy scale governs QAOA parameter setting in weighted problems, supporting the normalization argument.","marker":"[9]"},{"why":"The authors' earlier study of linearly simplified QAOA parameters and transfer on small instances, which this work extends.","marker":"[21]"},{"why":"Provides the semidefinite relaxation bound suggested as a practical proxy for the unknown ground-state energy in normalization.","marker":"[33]"}],"fun_headline_variants":["QAOA angles collapse to four, transfer across instances","Linear QAOA angles: 4 params, zero-cost transfer","Energy scale key to free QAOA transfer with 4 angles","Four linear angles: QAOA without per-instance optimization","QAOA transfer freed: just four linear parameters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The practical transfer claim depends on knowing the target problem's energy scale in advance, because without normalization derived from the ground-state energy the transferred angles perform no better than random sampling, and that energy is not generally known for a new instance.","fun_headline_variants_meta":{"raw":{"variants":["QAOA angles collapse to four, transfer across instances","Linear QAOA angles: 4 params, zero-cost transfer","Energy scale key to free QAOA transfer with 4 angles","Four linear angles: QAOA without per-instance optimization","QAOA transfer freed: just four linear parameters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000243,"raw_usage":{"total_tokens":1593,"prompt_tokens":1071,"completion_tokens":522,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":687,"completion_tokens_details":{"reasoning_tokens":442}},"tokens_in":687,"tokens_out":522,"duration_ms":5340,"temperature":1.0,"reasoning_tokens":442,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:27:08.957875+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the transfer on a fresh random Ising instance of 32 to 80 qubits, estimating the energy scale from data available before solving (edge count, mean coupling magnitude, or the semidefinite relaxation bound) instead of the exact ground-state energy, then apply the paper's toy linear angles and compare the approximation ratio on an MPS simulator. If the ratio systematically falls toward the random-sampling baseline or below 0.8, the zero-cost transfer claim fails in exactly the practical setting the paper targets.","supporting_citations":[{"cited_title":"Quantum Approximate Optimization Algorithm: Performance, Mechanism, and Implemen- tation on Near-Term Devices,","cited_arxiv_id":null,"evidence_quote":"Defines INTERP and FOURIER; the paper shows their optimized angles are nearly linear and uses them as baselines."},{"cited_title":"Pa- rameter Setting in Quantum Approximate Optimization of Weighted Problems,","cited_arxiv_id":null,"evidence_quote":"Shows that energy scale governs QAOA parameter setting in weighted problems, supporting the normalization argument."},{"cited_title":"Linearly simplified qaoa parameters and transferability,","cited_arxiv_id":null,"evidence_quote":"The authors' earlier study of linearly simplified QAOA parameters and transfer on small instances, which this work extends."},{"cited_title":"Improved approximation algorithms for maximum cut and satis- fiability problems using semidefinite programming,","cited_arxiv_id":null,"evidence_quote":"Provides the semidefinite relaxation bound suggested as a practical proxy for the unknown ground-state energy in normalization."}],"review_version":1}