{"id":"c5d413f5-85f0-4cc9-84ac-86675341eb9e","arxiv_id":"2607.11174","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Classical residual weights can be compiled into subspace quantum generators with a two-term error bound, enabling zero-shot transfer and initialization-time barren-plateau mitigation.","lead":"The paper gives an analytical map that turns classical residual-network weights into low-dimensional quantum circuits without any quantum training. If it holds, hybrid systems could start from classical models and stay trainable on large qubit registers by confining dynamics to a small subspace.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection beyond the residual-prior scope already flagged; the two-term theorem and subspace gradient floor hold as stated.","rationale":"The strongest claim is the two-term Subspace Quantization Theorem plus the polynomial gradient floor under subspace restriction. Both are proved with standard tools (polar optimality via Fan–Hoffman / SVD, Weingarten for U(k) 2-designs) and are numerically matched on the residual MNIST core. The residual prior is required only for the practical regime in which non-unitarity becomes small/second-order and the log is safely analytic; the paper states this explicitly (Sec. III.B, D.3) and already shows the first-order floor at full rank. Hardware results (Hellinger fidelity, 128-qubit gradient resolvability) are consistent with the subspace construction and do not over-claim trajectory-wide trainability. The Reader’s CONDITIONAL verdict already encodes exactly this scope caveat; no stronger load-bearing flaw (incorrect bound, gauge failure, or unsupported hardware claim) appears. Hence the verdict stays CONDITIONAL with no adjustment.","tokens_in":23409,"tokens_out":574,"duration_ms":5882,"concrete_test":"Re-run Validation II (Sec. IV.B / Fig. 3) on a deliberately non-residual dense layer (e.g., standard Glorot-initialized MNIST dense without identity residual) at the same N=64; if the measured total error still equals truncation + ||P_A-I_k||_F to machine precision and the polar/log map remains well-defined whenever σ(U_A) avoids -1, the theorem is confirmed outside the residual regime and only the practical zero-shot fidelity claim is narrowed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Reader correctly isolates residual (identity-centered) priors as the practical hinge for faithful zero-shot transfer and second-order non-unitarity. That is a scope limitation, not an internal inconsistency: Theorem 4 and Lemma 2 hold for arbitrary W (the non-unitarity term is exactly ||P_A-I_k||_F regardless of near-unitarity); the second-order claim is stated only under the additional pure-rotation condition (Appendix D.3). Theorem 2 is an initialization-time floor under a subspace 2-design and is independent of the residual prior once the active dynamics are confined to U(k). Branch-cut assumption (17) is likewise explicit. The paper already reports the first-order floor when k\to N (Fig. 3) and the soft identity anchors used in the deep experiments. No hidden algebraic gap or over-claim that would overturn the theorems was found.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript develops an analytical classical-to-quantum parameter transfer map that converts a square neural weight into a subspace hybrid operator O(Q,H)=Q e^{-iH} Q† via Stiefel SVD frame selection, polar nearest-unitary projection, and principal matrix logarithm. The Subspace Quantization Theorem (Theorem 4) bounds the Frobenius reconstruction error by geometric truncation plus exact non-unitarity ||P_A−I_k||_F (Lemma 2 / Fan–Hoffman), with a second-order rate under a near-unitary pure-rotation condition. Identity-centered residual priors keep generators near the identity, avoiding the log branch cut and seeding a warm-start whose active dynamics live in U(k); Theorem 2 then gives a Weingarten variance floor Ω(1/k) independent of ambient qubit number n at initialization. The same Lie-algebraic construction yields a covering-frame generator merge with an explicit second-order separation penalty. Minimal-model error budgets, MiniViT/Mini-DiT zero-shot transfers, and IBM ibm_kobe runs (Hellinger fidelity 0.987 at k=8; subspace gradients resolvable to 128 qubits) support the claims within the residual near-unitary regime.","tokens_in":23598,"tokens_out":1467,"duration_ms":25537,"significance":"If the stated scope is accepted, the work supplies a rare fully analytical, variational-free bridge from classical residual layers to subspace PQCs, with machine-checkable algebraic proofs (Appendices A–D), an exact two-term error identity rather than a loose inequality, and hardware evidence that subspace-restricted gradients remain above shot noise at 128 qubits while a global HEA does not. The explicit separation of truncation vs non-unitarity, the gauge-invariant hybrid operator on the associated Stiefel bundle, and the Fisher effective-dimension comparison under matched two-qubit budgets are concrete contributions. Code release further strengthens reproducibility. The main practical hinge—near-unitary residual priors—is a genuine scope limit rather than a circularity, so the theorems remain useful as stated for residual cores and as a warm-start recipe more generally.","major_comments":[{"comment":"Abstract and Theorem 4: The abstract states that non-unitarity becomes “second order in the near-unitary regime,” but Appendix D.3 shows the quadratic rate requires the additional pure-rotation (asymptotically skew-Hermitian) condition A−I_k ≈ −iH + O(||H||²_F). For generic near-unitary A the term remains first-order O(ε√k). The main-text theorem statement and abstract should state this condition explicitly so the second-order claim is not over-read.","section":"Abstract / Theorem 4 / Appendix D.3"},{"comment":"Sec. V.B (Experiment II): The authors note that damped residual coupling makes the quantum cores carry only a “modest fraction” of each block’s transformation, so the diffusion demonstration is conservative. That admission is important: the generative zero-shot claim is then only partially load-bearing. Either report a stronger residual-coupling ablation (larger core contribution) or rephrase the generative claim to match the modest-fraction regime actually tested.","section":"Sec. V.B, Experiment II"},{"comment":"Sec. III.B and practical zero-shot claim: Theorem 4 and Lemma 2 hold for arbitrary W, but faithful zero-shot fidelity without quantum-side training requires the residual near-unitary prior (W≈I+δW) so that ||P_A−I_k||_F is small and the spectrum stays away from the branch cut (17). This necessary condition should be stated as such in the abstract and conclusion, not only as a “natural” property of residual nets, so readers do not expect the same fidelity for generic dense layers.","section":"Sec. III.B / Abstract / Conclusion"}],"minor_comments":[{"comment":"Fig. 3 caption and Sec. IV.B: Clarify that the double-sided truncation ||W−Π_Q(W)||_F coincides with the Eckart–Young tail only under the near-normal residual regime; otherwise it is measured directly and need not equal the single-sided minimum.","section":"Fig. 3 / Sec. IV.B"},{"comment":"Eq. (12) vs zero-shot operator: Distinguish more sharply in the main text between the partial isometry O(Q,H) used for zero-shot compilation and the full unitary U(θ) that adds I−QQ† for the trainable warm-start; both are used but the unitarity status differs.","section":"Sec. II.F / Sec. III.A"},{"comment":"Fig. 7 / Sec. V.C: Annotate or discuss that at k=16 the 187 CX count already drives Hellinger fidelity to 0.845; the strong NISQ claim is clearest for k≤8, which the text largely acknowledges but the abstract headline (0.987 at k=8) could note the k=16 drop for balance.","section":"Fig. 7 / Abstract"},{"comment":"Notation: N is used both for ambient Hilbert dimension and (implicitly) for sample size in Appendix E; disambiguate n_data vs N.","section":"Appendix E"},{"comment":"Typos / consistency: “ibm kobe” / “ibm_kobe” / “IBMibm kobe” appear in several forms; standardize. Also “ak-dimensional” → “a k-dimensional” (abstract).","section":"Abstract / throughout"},{"comment":"Related work: A short comparison to other Lie-algebraic / DLA barren-plateau analyses (e.g. Ragone et al., Nat. Commun. 2024, already cited) on how subspace restriction differs from full DLA dimension control would help place Theorem 2.","section":"Introduction / Sec. II.F"}],"recommendation":"minor_revision","confidential_remarks":"The algebraic core (Lemma 2, Theorems 1–4, Weingarten floor) is clean and the residual-prior scope is already partially disclosed; I do not see a load-bearing inconsistency that would justify reject or major_revision. The main risk is over-reading of “zero-shot quantum learning” and “barren-plateau mitigation” beyond residual cores and initialization. Minor revision to tighten those statements should suffice. Fit for quant-ph / QML venues is good; hardware results on ibm_kobe are a genuine plus."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful core is a concrete zero-shot map: SVD frame, polar nearest unitary, principal log, then a hybrid operator O(Q,H)=Q exp(-iH) Q†. Theorem 4 is the real contribution—reconstruction error splits exactly into geometric truncation plus non-unitarity ||P_A-I_k||_F, with the second term becoming second-order only under the pure-rotation residual regime they state. Appendices A–D are standard linear algebra and Weingarten; they check out. The covering-frame merge cost is also clean (second-order in generator separation).\n\nWhat they do well: the minimal-model error budget (Fig. 3) matches the theorem to the floor when k\to N; the 4-qubit zero-shot MNIST gap is ~1%; ibm_kobe gives Hellinger 0.987 at k=8 and a subspace gradient that stays ~1.5–2.5 orders above shot noise out to 128 qubits while a global HEA dies. Code is released. That package is new relative to the usual warm-start / DLA / barren-plateau literature they cite.\n\nSoft spots are scope, not algebra. Faithful zero-shot and the second-order claim need identity-centered residual priors (soft anchors in the deep nets). Theorem 4 itself holds for arbitrary W; when the prior is far from unitary the non-unitarity term is first-order and dominant, which they already show. Theorem 2 is an initialization-time floor under a subspace 2-design, not a full-trajectory guarantee—they say so. Expressivity-per-gate (Fisher effective dimension) is a nice extra but secondary. Self-cites are their own prior BP work; not load-bearing here.\n\nThis is for people building hybrid NISQ layers or classical-to-quantum transfer who want an explicit error budget rather than another heuristic warm-start. Math and hardware data are solid enough that a serious editor should send it to referees; the residual-prior dependence and init-only trainability caveat just need to be more prominent. I would engage with it and cite the theorem and the 128-qubit gradient result.","headline":"Clean analytical compiler from residual classical weights to subspace PQCs, with a real two-term error theorem and hardware gradients that stay alive to 128 qubits.","tokens_in":24268,"tokens_out":564,"would_cite":true,"duration_ms":6246,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Classical neural weights can be turned into quantum subspace circuits without any quantum training, with reconstruction error controlled by truncation and non-unitarity, while confining gradients to a small active register.","keywords":["barren plateaus","parameterized quantum circuits","Stiefel manifold","subspace quantization","zero-shot transfer","Lie algebra","warm-start initialization","model merging"],"falsifier":"Train a non-residual dense layer far from the identity, apply the same transfer map at moderate rank k, and check whether the measured reconstruction error still collapses onto the non-unitarity floor predicted by Theorem 4 and whether the zero-shot accuracy gap remains of order 1 percent; a large first-order gap would falsify the practical claim.","tokens_in":24239,"feed_emoji":"⚛","tokens_out":1016,"duration_ms":12505,"temperature":0.7,"pith_summary":"The paper argues that the barren-plateau barrier for variational quantum circuits can be sidestepped at initialization by compiling classical residual-network weights into low-dimensional quantum evolutions through a purely algebraic map. That map selects a Stiefel frame by SVD, projects the retained block onto its nearest unitary via polar decomposition, and extracts a Hermitian generator by the principal matrix logarithm. The Subspace Quantization Theorem then bounds the reconstruction error by two explicit terms: geometric truncation of the discarded singular directions and the retained block's singular-value deviation from unitarity, the second term becoming second-order when the classical prior is near-unitary and rotational. Identity-centered residual layers keep the spectrum safely away from the logarithm branch cut and seed generators near the identity, so gradient variance scales only with the chosen subspace dimension rather than the ambient qubit count. The same generators can be transported onto a covering frame and averaged to merge models with a second-order separation penalty. On IBM hardware the compiled circuits retain high output fidelity at small rank and keep measurable subspace gradients out to 128 physical qubits.","feed_headline":"Classical weights compile to quantum circuits with no quantum training","feed_subtitle":"Error splits into truncation plus non-unitarity; gradients stay alive on 128 qubits","key_machinery":"The Subspace Quantization Theorem (Theorem 4) together with the four-step transfer map T: SVD frame extraction, subspace restriction A=Q† W Q, nearest-unitary polar factor UA, and principal logarithm H=i log(UA). It partitions reconstruction error into truncation plus singular-value deviation and supplies the generators used both for zero-shot deployment and for covering-frame model merging.","core_discovery":"An exact analytical parameter-transfer map converts a square classical weight into a hybrid subspace operator O(Q,H)=Q exp(-iH) Q† whose Frobenius reconstruction error is bounded by the sum of geometric truncation and exact non-unitarity of the retained block; when the classical prior is residual (near-identity), that non-unitarity is second-order and the same construction supplies a warm-start whose gradient variance is polynomial in the subspace dimension k and independent of ambient qubit number n.","pith_inferences":["If residual priors are the practical enabler, future hybrid compilers may deliberately regularize classical layers toward the unitary manifold rather than treating near-unitarity as an accidental property of residual nets.","The covering-frame merge construction suggests a route to classical multi-task distillation that produces a single quantum circuit without ever evaluating a joint quantum loss.","Once the discarded polar scaling component can be retained on an auxiliary register, the method could extend beyond square residual cores to general rectangular or strongly non-unitary layers."],"forward_implications":["Pre-trained residual cores can be deployed as quantum subspace circuits with no quantum-side gradient steps, converting classical layers into compact NISQ-executable unitaries.","Gradient variance at initialization is set by the classical hyperparameter k rather than ambient qubit number, so warm-start trainability is retained even when the physical register is large.","Multiple specialized models can be merged by transporting their generators onto a covering Stiefel frame and averaging in the common Lie algebra, with an explicit second-order separation cost.","The same map supplies both a high-fidelity forward circuit and a resolvable subspace gradient signal on present-day hardware up to at least 128 qubits."],"fun_headline_variants":["Classical nets map to quantum evolutions with zero quantum training","Subspace quantization enables zero-shot classical-to-quantum transfer","Near-identity generators sidestep barren plateaus via subspace map","Stiefel polar log converts weights to hybrid quantum operators","Error-bounded transfer keeps gradients alive at 128 qubits"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The classical residual weights must stay close enough to unitary that the retained block remains near the identity of the unitary group; otherwise the non-unitarity error becomes first-order and the claimed faithful zero-shot transfer collapses.","fun_headline_variants_meta":{"raw":{"variants":["Classical nets map to quantum evolutions with zero quantum training","Subspace quantization enables zero-shot classical-to-quantum transfer","Near-identity generators sidestep barren plateaus via subspace map","Stiefel polar log converts weights to hybrid quantum operators","Error-bounded transfer keeps gradients alive at 128 qubits"]},"model":"grok-4.5","effort":"low","cost_usd":0.003076,"raw_usage":{"total_tokens":1086,"prompt_tokens":764,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":30760000,"prompt_tokens_details":{"text_tokens":764,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":256,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":764,"tokens_out":66,"duration_ms":3692,"temperature":1.0,"reasoning_tokens":256,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T06:28:57.363477+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train a non-residual dense layer far from the identity, apply the same transfer map at moderate rank k, and check whether the measured reconstruction error still collapses onto the non-unitarity floor predicted by Theorem 4 and whether the zero-shot accuracy gap remains of order 1 percent; a large first-order gap would falsify the practical claim.","supporting_citations":[],"review_version":1}