{"id":"b37f8c5b-2d50-4cea-8176-d999bc60598d","arxiv_id":"2607.05774","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":4,"one_line_summary":"PLANE jointly estimates latent gene positions from a target network and proxy embeddings on a larger gene set, with provably optimal channel weighting and demonstrated gains in network recovery and imputation.","lead":"The paper proposes PLANE, a method that combines a partially observed gene-gene network with externally learned gene embeddings to improve recovery of latent biological structure. It provides theoretical guarantees on optimal weighting between network and embedding channels, validated on simulations and CRISPR perturbation data.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Block-dominance condition in Theorem 4 is unverified and may fail when n_P ≫ n_Q, undermining the target-block error bound and optimal weighting rule.","rationale":"The reader's verdict of CONDITIONAL with MODERATE confidence is appropriate. The block-dominance condition is genuinely load-bearing for the central theoretical claim (Theorem 4 / Corollary 4), and it is not independently verified from the data-generating process. This is a real gap, not a minor technicality.\n\nHowever, the verdict should remain CONDITIONAL rather than moving to REJECT for several reasons: (1) The partial-oracle analysis (Theorem B.2) independently supports the functional form of the weighting rule E(λ) and the harmonic-mean identity, even though it does not cover the full iterative procedure. (2) The simulations (Figure 2) empirically show that the weighting rule behaves as predicted—oracle λ₂ shifts toward the more informative channel, and CV tracks the oracle. (3) The zero-order result (Theorem 2) is minimax-optimal and does not require block-dominance, providing a solid foundation even if the gradient-dynamics refinement is conditional. (4) The concern is about a verifiable condition, not a fundamental flaw—if the fixed-point calculation shows the ratio vanishes, the theory is complete; if not, the bound can be modified to include B-block coupling terms.\n\nThe paper's strengths—novel problem formulation, clean identification result, minimax-optimal zero-order bound, and careful gradient-dynamics analysis—warrant a conditional acceptance pending verification of the block-dominance condition. The absence of shipped code and the small real-data sample (n_Q=99) are secondary concerns that reinforce the need for this verification but do not independently change the verdict.","tokens_in":56150,"tokens_out":4679,"duration_ms":330949,"concrete_test":"Under the Gaussian model (1.1) with n_Q fixed and n_P → ∞, analytically compute the fixed-point ratio E[s_B^(∞)]/E[s_UQ^(∞)] using the Δ_S expression from Theorem 3. Specifically: (1) derive the B-block and U_Q-block statistical error contributions from Δ_S by tracing how the aggregate error decomposes into block-level errors via the Gram normalizers G_B and G_{UQ}; (2) evaluate the ratio under the scaling ||E||²_op ~ σ²₂(n_P+d), ||D||²_op ~ σ²₁n_Q, σ²_min(U) ~ n_P, σ²_min(U_Q) ~ n_Q, σ²_min(B) ~ d. If the ratio does not vanish as n_P → ∞, the block-dominance condition fails and the target-block error bound (3.2) requires modification. A numerical complement: run Algorithm 1 on simulated data with n_Q=99, increasing n_P ∈ {200, 500, 1000, 2000}, and report the empirical ratio s_B^(T)/s_UQ^(T) at convergence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader correctly identifies the block-dominance condition sup_{0≤t≤T} s_B^(t)/s_UQ^(t) = o(1) as n_P → ∞ (Corollary 4, §B.4) as the load-bearing assumption. This condition is essential: in the proof of Theorem 4 (§B.4), it is used to drop the o(η s_B^(t)) cross-term in the U_Q-block recursion, which is what allows the clean separation of network noise ||D||_op and embedding noise ||E_Q||_op in the target-block error bound (3.2). Without it, the U_Q-block error remains coupled to the B-block error, and the harmonic-mean bound E(λ*) ≤ min{E(1), E(0)} does not follow.\n\nThe concern is not merely abstract. Under the stated data-generating model (1.1) with Gaussian noise, the B-block statistical error at the fixed point of the Theorem 3 recursion scales roughly as k·λ₂||E||²_op||U||²_op/σ²_min(U), while the U_Q-block error scales as k·[λ²₁||U_Q||²_op||D||²_op + λ²₂||E_Q||²_op||B||²_op]/[2λ₁σ²_min(U_Q)+λ₂σ²_min(B)]. When n_P grows with n_Q fixed (the paper's asymptotic regime), ||E||²_op ~ σ²₂(n_P+d) grows, while ||D||²_op ~ σ²₁n_Q stays fixed. Whether the ratio s_B/s_UQ → 0 depends on the relative growth of σ²_min(U) (which scales with n_P) versus the noise growth, but the paper does not carry out this calculation. Remark 2 only shows that block-dominance in curvature-weighted error implies Frobenius-norm dominance under an additional relative-conditioning condition on iterates—neither condition is derived from the data-generating process.\n\nThe practical relevance is high: the central use case is n_P ≫ n_Q (embeddings on a larger gene universe), which is exactly the regime where the condition is least obviously satisfied. The partial-oracle analysis (Theorem B.2) provides some support for the weighting rule's functional form, but it sidesteps the iterative coupling entirely, so it does not substitute for verifying block-dominance in the actual algorithm.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This paper proposes PLANE (Proxy-Latent Assisted Network Estimation), a method that jointly leverages a partially observed target gene–gene network and externally learned proxy gene embeddings available on a larger gene set to recover shared latent structure. The core model posits that both the target adjacency and the proxy embeddings are noisy observations of a common low-rank latent factor, and the method balances the two channels through tuning weights (λ₁, λ₂). The paper provides: (i) an identification result (Theorem 1), (ii) a zero-order minimax-optimal bound on the weighted reconstruction loss (Theorem 2), (iii) a gradient-dynamics analysis of blockwise normalized gradient descent yielding deterministic contraction with an explicit statistical tolerance (Theorem 3), and (iv) a target-block error bound that separates network and embedding noise and yields an optimal weighting rule with a harmonic-mean property (Theorem 4 / Corollary B.2). Simulations and a CRISPRa Perturb-seq analysis demonstrate practical gains over network-only baselines.","tokens_in":57207,"tokens_out":1606,"duration_ms":582207,"significance":"The paper addresses a well-motivated problem at the intersection of statistical network analysis and foundation-model embeddings. The theoretical contributions are substantial: Theorem 2 correctly identifies the minimax rate via reduction to Donoho–Gavish, and the gradient-dynamics analysis (Theorems 3–4) is technically involved, requiring careful handling of rotational alignment, blockwise preconditioning, and cross-block interactions. The harmonic-mean bound E(λ*) ≤ min{E(1), E(0)} is a clean and interpretable result. The partial-oracle analysis (Proposition B.2) provides a useful sharpness check. The simulations are reasonably comprehensive, including misspecification robustness and baseline comparisons. The real-data application is biologically grounded. The main theoretical gap—discussed below—is the block-dominance condition in Theorem 4, which is load-bearing but not verified from the data-generating process.","major_comments":[{"comment":"§3.3, Theorem 4 (Corollary 4 in the text): The block-dominance condition sup_{0≤t≤T} s_B^(t)/s_UQ^(t) = o(1) as n_P → ∞ is load-bearing for the target-block error bound (3.2) and the harmonic-mean weighting result. In the proof (§B.4), this condition is used to drop the o(η s_B^(t)) cross-term in the U_Q-block recursion, which is what enables the clean separation of ||D||_op and ||E_Q||_op in (3.2). Remark 2 (§B.5) shows that curvature-weighted dominance implies Frobenius-norm dominance under an additional relative-conditioning condition on iterates, but neither condition is derived from the data-generating model (1.1). Under the stated asymptotic regime (n_P grows, n_Q fixed), ||E||²_op ~ σ₂²(n_P + d) grows while ||D||²_op ~ σ₁² n_Q stays fixed, so whether s_B/s_UQ → 0 depends on the relative scaling of σ_min(U) (which grows with n_P) versus the noise growth. The paper does not carryout","section":null},{"comment":"§3.3 and §B.5, Remark 2: The relative-conditioning condition sup_t (4λ₁||U_Q^(t)||²_op + 2λ₂||B^(t)||²_op)/(2λ₂ σ²_min(U^(t))) = O(1) is introduced in Remark 2 to convert curvature-weighted dominance to Frobenius-norm dominance. This condition implicitly requires σ_min(U^(t)) to scale comparably to ||U_Q^(t)||²_op, which is a strong requirement when n_P grows. The authors should either verify this from the model or explicitly state it as an additional assumption on the data-generating process, separate from the algorithmic regularity conditions in §B.2. As stated, the reader cannot assess when Theorem 4 applies.","section":null},{"comment":"§3.3: The practical PLANE-CV procedure selects λ₂ by cross-validation on held-out target-network entries, while the optimal weighting rule λ* depends on population quantities (||U_Q||_op, ||B||_op, σ_min(U_Q), σ_min(B), ||D||_op, ||E_Q||_op) that are not directly observed. The paper does not discuss how to estimate these quantities or whether PLANE-CV consistently selects weights near λ*. A brief discussion of when cross-validation can be expected to approximate the oracle rule, or at least an acknowledgment of this gap and its difficulty, would strengthen the connection between theory and practice.","section":null}],"minor_comments":[{"comment":"The result in §3.3 is labeled both 'Corollary 4' and 'Theorem 4' in different places (e.g., the statement says 'Corollary 4' but the text refers to 'Theorem 4'). Please use consistent numbering.","section":null},{"comment":"§2, Eq. (2.1): The normalization λ₁ + λ₂ = 1 is introduced after the loss function, but Algorithm 1 and the theoretical results use λ₁, λ₂ without this constraint in some places (e.g., the Gram matrices in §3.1 use raw λ₁, λ₂). Clarify where the normalization applies.","section":null},{"comment":"§5: The real-data analysis uses d = 10112 concatenated embedding dimensions and n_P = 1099 genes. The ratio d/n_P is large; it would help to comment on whether Assumption 2 (full rank of B with finite condition number) is plausible in this regime, or whether the theory is intended only as an asymptotic guide.","section":null},{"comment":"§B.2, Assumption 3: The initialization condition e_0 ≤ τ is stated in terms of the curvature-weighted error metric, which depends on the iterates themselves. A brief comment on how this can be verified in practice (e.g., via spectral initialization) would help.","section":null},{"comment":"§C.1, proof of Lemma C.1: There appears to be a stray line 'U^(t)⊤ ∥²_F.' between the B-block and U_{Q^c}-block analyses that seems to be a formatting artifact.","section":null},{"comment":"Table D1: The 'Study 2, extra nodes' row lists n_Q = 30 and |Q^c| = n_R^(1), but the text in §4 refers to |Q^c| values up to 500. Clarify the sample size for this study.","section":null},{"comment":"The paper would benefit from a brief discussion of computational complexity per iteration of Algorithm 1, particularly the cost of the blockwise Gram matrix inversions.","section":null}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the block-dominance condition is well-founded and is the primary reason for the major revision recommendation. The condition is genuinely load-bearing for Theorem 4, and the paper does not verify it from the data-generating process. However, the concern is fixable: the authors could either (a) provide a calculation showing when s_B/s_UQ → 0 under model (1.1), or (b) reframe Theorem 4 as a conditional result and add a proposition verifying the condition in a simplified regime (e.g., isotropic signals, balanced growth). The rest of the paper—the zero-order analysis, the gradient-dynamics framework, the simulations, and the real-data analysis—is solid and does not depend on this condition. I would not recommend rejection on this basis."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive report. The referee correctly identifies the block-dominance condition in Theorem 4 as the main theoretical gap, and we agree that this condition and the relative-conditioning requirement in Remark 2 need to be more explicitly connected to the data-generating model. We also agree that the gap between the oracle weighting rule and the practical PLANE-CV procedure deserves discussion. We address each point below.","responses":[{"response":"The referee is correct that the block-dominance condition is not derived from the data-generating model (1.1), and we agree this gap should be addressed. We have carried out the analysis the referee requested. Under the asymptotic regime n_P → ∞ with n_Q fixed, and under Assumption 1 (Gaussian noise), we can verify the condition as follows. The B-block curvature-weighted error s_B^(t) involves the term 2λ₂||Δ_B^(t) U^(t)⊤||²_F, which is driven by the statistical error in estimating B. From the B-block gradient (Lemma C.2(ii)), the statistical contribution to s_B scales as λ₂²||E||²_op||U||²_op / σ²_min(U). Meanwhile, the U_Q-block error s_UQ^(t) scales as λ₁²||U_Q||²_op||D||²_op / (2λ₁σ²_min(U_Q) + λ₂σ²_min(B)). Under the standard factor-model scaling where σ_min(U) grows at rate √n_P (which holds when the latent positions have bounded per-row norm), the ratio s_B/s_UQ scales as λ₂²σ₂²(n_P+d)·n_P / [λ₁²σ₁²n_Q · n_P] = λ₂²σ₂²(n_P+d) / (λ₁²σ₁²n_Q), which grows rather than vanishes. This indicates that the block-dominance condition does NOT hold under the standard factor-model scaling without additional structure. However, the condition does hold under a spiked covariance regime where σ²_min(U) grows at rate n_P while ||U||²_op grows at rate n_P, so that the ratio ||E||²_op||U||²_op/σ⁴_min(U) stays bounded. We will revise §3.3 and §B.5 to (a) explicitly state the block-dominance condition as an assumption on the relative signal-to-noise scaling, (b) provide the verification under the spiked covariance regime, and (c) clearly state that the condition is not automatic under the base model and requires this relative scaling. We will also note that the practical relevance of Theorem 4 is supported by the simulation evidence, where the harmonic-mean weighting behavior is empiri","revision_made":"yes","referee_comment":"§3.3, Theorem 4: The block-dominance condition sup_{0≤t≤T} s_B^(t)/s_UQ^(t) = o(1) as n_P → ∞ is load-bearing but not verified from the data-generating process. Under the stated asymptotic regime (n_P grows, n_Q fixed), ||E||²_op ~ σ₂²(n_P + d) grows while ||D||²_op ~ σ₁² n_Q stays fixed, so whether s_B/s_UQ → 0 depends on the relative scaling of σ_min(U) versus noise growth."},{"response":"We agree with the referee that the relative-conditioning condition in Remark 2 should be stated as an explicit assumption rather than introduced only as an interpretive remark. This condition is indeed strong: it requires σ²_min(U^(t)) to be comparable to ||U_Q^(t)||²_op, which under the n_P → ∞ regime requires the smallest singular value of U to grow sufficiently fast relative to the target-block operator norm. We will revise the manuscript to: (1) elevate this condition to a formal assumption (a new Assumption 6) clearly separated from the algorithmic regularity conditions in §B.2, (2) state explicitly that it is a condition on the data-generating process, not merely on the iterates, and (3) note that it holds under the spiked covariance regime described in our response to the first comment, where σ²_min(U) ~ n_P and ||U_Q||²_op ~ n_Q (fixed). We will also add a remark that this condition is the price of converting from the curvature-weighted error metric (which is the natural quantity for the gradient-dynamics analysis) to the unweighted Frobenius-norm error that is more interpretable for practitioners. The referee is correct that without this clarification, the reader cannot assess when Theorem 4 applies.","revision_made":"yes","referee_comment":"§3.3 and §B.5, Remark 2: The relative-conditioning condition sup_t (4λ₁||U_Q^(t)||²_op + 2λ₂||B^(t)||²_op)/(2λ₂ σ²_min(U^(t))) = O(1) is introduced to convert curvature-weighted dominance to Frobenius-norm dominance. This condition implicitly requires σ_min(U^(t)) to scale comparably to ||U_Q^(t)||²_op, which is strong when n_P grows. The authors should either verify this from the model or explicitly state it as an additional assumption."},{"response":"The referee raises a valid and important point about the gap between the oracle weighting rule λ* (which depends on unobserved population quantities such as ||U_Q||_op, ||B||_op, σ_min(U_Q), σ_min(B), ||D||_op, ||E_Q||_op) and the practical PLANE-CV procedure. We acknowledge that the current manuscript does not establish a formal consistency result connecting PLANE-CV to the oracle rule λ*. This is a genuine theoretical gap that we cannot fully close in this revision. The difficulty is that the oracle rule depends on operator norms of population quantities that are themselves estimated through the iterative procedure, creating a circularity that makes direct plug-in estimation nontrivial. We will add a new paragraph in §3.3 (or §6) that: (1) explicitly acknowledges this gap, (2) explains why a formal consistency result for PLANE-CV → λ* is difficult—because the oracle quantities are functions of the latent factors being estimated, (3) notes that the simulation evidence (Figure 2A–C, D–F) provides empirical support that cross-validation tracks the oracle direction, and (4) suggests this as a direction for future work, potentially via a two-stage procedure where pilot estimates of the population quantities are used to approximate λ*. We believe this honest acknowledgment strengthens rather than weakens the paper.","revision_made":"partial","referee_comment":"§3.3: The practical PLANE-CV procedure selects λ₂ by cross-validation on held-out target-network entries, while the optimal weighting rule λ* depends on population quantities that are not directly observed. The paper does not discuss how to estimate these quantities or whether PLANE-CV consistently selects weights near λ*."}],"tokens_in":56227,"tokens_out":1495,"duration_ms":190298,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The main thing to know: this paper extends the joint latent-space model of Zhang et al. (2022) to the setting where proxy embeddings cover a larger node set than the observed network, and the real contribution is the gradient-dynamics analysis (Theorem 3) that characterizes how the (λ₁, λ₂) weighting trades off network noise against embedding noise in latent-position recovery. That piece is genuinely new and the proofs are carefully done — the blockwise Gram-normalized gradient descent scheme, the aligned error metric, and the contraction argument in Theorem 3 are all worked out in detail. The zero-order result (Theorem 2) correctly identifies the minimax rate via reduction to Donoho-Gavish, and the partial-oracle analysis (Proposition B.2) cleanly isolates the weighting trade-off. The simulations are reasonable and the CRISPRa application, while small (n_Q = 99), is a legitimate proof-of-concept with a sensible leave-one-source-out ablation design. Credit is due for the theoretical effort here — this is not a trivial extension of existing joint estimation results, and the harmonic-mean bound on the optimal weighting is a clean, interpretable result when it applies. The block-dominance condition (sup_t s_B^(t)/s_UQ^(t) = o(1) as n_P → ∞) in Theorem 4 is the load-bearing assumption that the stress-test flags, and I agree it is the right thing to flag. It is what allows the cross-term coupling between B-block and U_Q-block errors to be dropped, yielding the clean separation of network and embedding noise in the target-block bound (3.2). The concern that this condition is hardest to satisfy exactly when n_P ≫ n_Q — the paper's central use case — is legitimate. Remark 2 shows block-dominance in curvature-weighted error implies Frobenius-norm dominance under an additional relative-conditioning condition, but neither condition is derived from the data-generating process alone. That said, I do not think this is a fatal flaw. The partial-oracle result (Proposition B.2) provides independent evidence that the functional form of the weighting rule is correct, even though it sidesteps the iterative coupling. The practical PLANE-CV procedure selects λ₂ by cross-validation on held-out network entries, which is independent of the theoretical optimal λ* and does not rely on the block-dominance condition holding. So the practical method and its empirical validation stand even if the full-theory gap remains. The absence of shipped code or data is a minor weakness — the simulations are described in enough detail to reproduce, but the real-data analysis would benefit from code release. The n_Q = 99 application is small relative to the asymptotic regime, but the paper is appropriately modest about this. This paper is for methodological researchers in statistical network analysis and latent-variable models. Readers interested in the theory of joint estimation with auxiliary information will get the most value. It deserves a serious referee who can check the proof details in Appendix B and assess whether the block-dominance condition can be verified or weakened.","headline":"Gradient-dynamics analysis of network-embedding trade-off is new and mostly sound; block-dominance condition needs verification","tokens_in":57348,"tokens_out":716,"would_cite":true,"duration_ms":169956,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Optimal Weighting Proven for Mixing Gene Networks with Proxy Embeddings","keywords":["latent factor model","network estimation","proxy embeddings","normalized gradient descent","gene regulatory network","single-cell CRISPR","harmonic mean bound","low-rank factorization"],"falsifier":"If the embedding-loading estimation error does not vanish relative to the target-block error as the number of genes grows, the block-dominance condition fails, and the closed-form optimal weight and its harmonic-mean error guarantee no longer hold.","tokens_in":56213,"feed_emoji":"🧬","tokens_out":1311,"duration_ms":276709,"temperature":0.7,"pith_summary":"This paper addresses a practical asymmetry in modern genomics: gene–gene interaction networks are typically measured on a small set of target genes, while externally trained foundation models produce embedding vectors for a much larger gene universe. The authors propose PLANE (Proxy-Latent Assisted Network Estimation), a method that treats the observed network and the proxy embeddings as two noisy observations of a shared low-dimensional latent geometry. The core question is how to set the weight between these two channels. The authors show that a standard zero-order optimality analysis can bound the total reconstruction error but cannot reveal how the weighting trades off network noise against embedding noise inside the latent factors themselves. To get inside that trade-off, they analyze a blockwise normalized gradient descent scheme and prove a deterministic contraction bound on a curvature-weighted, rotation-aligned error metric. Specializing that bound to the target-gene block yields an explicit, closed-form optimal weight that separates network noise from embedding noise. When the unconstrained optimum is admissible, the resulting error satisfies a harmonic-mean identity: it is no worse than the better of the two single-channel errors. The paper validates the rule in simulations and in a CRISPR perturbation dataset, showing that informative embeddings improve latent recovery, network denoising, and link imputation, while cross-validation automatically downweights uninformative proxies.","feed_headline":"Optimal Weight Between Gene Networks and Embeddings Has a Closed Form","feed_subtitle":"A gradient-dynamics proof shows the best mix of network and proxy-embedding channels satisfies a harmonic-mean error bound, improving over","key_machinery":"The argument rests on three pieces: (1) a joint low-rank factorization in which a target adjacency matrix and a proxy embedding matrix share a common latent factor, identifiable up to orthogonal rotation under mild rank conditions; (2) a blockwise normalized gradient descent algorithm whose Gram-matrix preconditioning produces a clean recursion in an aligned, curvature-weighted error metric coupling three blocks (target latent positions, embedding loadings, and embedding-only latent positions); and (3) a block-dominance condition under which the embedding-loading error becomes asymptotically negligible relative to the target-block error, allowing the aggregate contraction to specialize to a ","core_discovery":"The central discovery is a closed-form optimal weighting formula for balancing a target network against proxy embeddings when both are noisy views of a shared latent structure. By analyzing blockwise Gram-normalized gradient descent through a curvature-weighted, Procrustes-aligned error metric, the authors prove that the target-block estimation error separates cleanly into a network-noise term and an embedding-noise term. The optimal weight is the ratio that equalizes the relative signal-to-noise contributions of the two channels, and the resulting error bound satisfies a harmonic-mean identity, guaranteeing improvement over either channel alone when the unconstrained optimum is admissible.","pith_inferences":["The harmonic-mean structure of the optimal error suggests a general principle: whenever two noisy channels observe a shared low-rank signal through different linear maps, the optimal combination error may generically take a harmonic-mean form, and the weighting formula could extend to settings beyond gene networks, such as multi-view data integration or sensor fusion with heterogeneous noise.","The block-dominance condition—that the embedding-loading estimation error must be asymptotically negligible relative to the target-block error—is the key gatekeeper for the harmonic-mean guarantee. If this condition fails, the clean separation of network and embedding noise may break down, and the optimal weighting could require a more complex formula that accounts for loading-estimation uncertain","The framework could be extended to nonlinear link functions (as the authors note) or to settings with more than two channels, where the harmonic-mean identity might generalize to a multi-channel analog, potentially connecting to multi-view canonical correlation analysis.","The leave-one-source-out ablation in the real-data analysis suggests that the practical benefit of proxy embeddings depends on complementarity across embedding sources, raising the question of whether the optimal weighting could be further refined by source-specific weights rather than a single aggregate embedding weight."],"forward_implications":["When externally trained gene embeddings carry biological signal related to the target network, they can be provably combined with the network to improve latent-position recovery beyond what either source achieves alone, with a data-adaptive weight that does not require manual tuning.","The harmonic-mean error bound provides a formal guarantee that adding an informative proxy channel never hurts target-block estimation relative to the network-only baseline, as long as the unconstrained optimum is admissible or the projected weight is used.","The framework extends covariate-assisted network estimation to a regime where auxiliary information lives on a strictly larger node set than the observed network, enabling link imputation for genes with no observed network data.","In the CRISPR perturbation application, the method recovered biologically interpretable regulatory modules (neurotransmitter transport, immune-cell development, hematopoietic control) that were less apparent from the network-only fit, suggesting the approach can surface structure that a single-channel analysis would miss."],"fun_headline_variants":["Closed-Form Weight Balances Gene Networks and Proxy Embeddings","Harmonic Mean Bound Sets Optimal Gene Network and Embedding Mix","Proven Weight Balances Target Networks Against Proxy Embeddings","Optimal Mix of Gene Networks and Embeddings Derived in Closed Form","Curvature-Weighted Gradient Dynamics Yield Network Embedding Mix"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The block-dominance condition requires that the estimation error for the embedding-loading matrix is asymptotically negligible relative to the target-block error. The paper shows this holds under an additional relative-conditioning condition on the iterates, but does not establish when that condition is met from the data-generating process alone. If the embedding-loading error is not negligible, the clean target-block error bound and the harmonic-mean optimality guarantee may","fun_headline_variants_meta":{"raw":{"variants":["Closed-Form Weight Balances Gene Networks and Proxy Embeddings","Harmonic Mean Bound Sets Optimal Gene Network and Embedding Mix","Proven Weight Balances Target Networks Against Proxy Embeddings","Optimal Mix of Gene Networks and Embeddings Derived in Closed Form","Curvature-Weighted Gradient Dynamics Yield Network Embedding Mix"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1038,"prompt_tokens":532,"completion_tokens":506,"prompt_tokens_details":null},"tokens_in":532,"tokens_out":506,"duration_ms":38163,"temperature":1.0,"reasoning_tokens":444,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T00:24:21.735389+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If the embedding-loading estimation error does not vanish relative to the target-block error as the number of genes grows, the block-dominance condition fails, and the closed-form optimal weight and its harmonic-mean error guarantee no longer hold.","supporting_citations":[],"review_version":1}