{"id":"744e0251-8214-4463-b418-31a9126e2a54","arxiv_id":"2507.23768","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A Total Risk Prior makes Bayesian transfer learning formal by placing the target near the risk-minimizing combination of source models and selecting useful sources with Gibbs sampling.","lead":"This paper introduces a new Bayesian transfer learning method that builds a single joint prior across target and source datasets, so uncertainty is quantified and unhelpful source data can be down-weighted. On gene expression data it beats a leading frequentist transfer method, especially when the source datasets are small.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that Trans-Lasso is an approximate MAP of the TRP rests on Theorem 3, whose proof is deferred to a missing appendix; as written, this load-bearing theoretical bridge cannot be checked.","rationale":"Reading in good faith, the TRP construction is coherent: Eq. (2)-(6) are internally consistent, the kernel-prior device makes the joint prior proper, and the toy Simpson's-paradox example plus the GTEx results are plausible evidence for the empirical claims. The weakest point is not the prior location assumption, which is a modeling assumption that the eta indicators honestly hedge, but rather the unverifiable theoretical bridge used to position the method as the Bayesian analogue of Trans-Lasso. Theorem 3 is the single most load-bearing result for the paper's framing, and its proof is absent. Theorem 4, on the non-concentration of eta, is similarly deferred and is needed to understand the method's negative-transfer behavior. The numerical check proposed would settle whether Theorem 3 holds as stated or requires additional assumptions. Since the computational machinery and the empirical evaluation are present and could support the paper once the proofs or checks are supplied, the appropriate verdict remains CONDITIONAL, matching the reader's verdict rather than escalating to REJECT or UNVERDICTED.","tokens_in":17958,"tokens_out":12922,"duration_ms":135446,"concrete_test":"Obtain or reconstruct the missing Appendix F and independently verify Theorem 3 numerically. Simulate P=5, K=3, N0=50 with Gaussian designs satisfying X_k^T X_k / N_k -> Sigma_k. For N_k in {100, 1e3, 1e4, 1e5}, compute the exact global MAP of the Laplace-TRP objective (Eq. 25 or its reparameterization in Eq. 30) to numerical tolerance, and compare with the modified Trans-Lasso estimator of Section 3.1 across 50 random seeds. If ||beta*_0 - beta_hat_TL_0|| does not converge to zero as min N_k -> inf with tau fixed, Theorem 3 is false as stated. If it converges only under extra conditions, those conditions must be stated explicitly and incorporated into the claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's abstract promises that minimax-frequentist transfer learning 'may be viewed as an approximate Maximum a Posteriori approach to our model.' That promise is made precise by Theorem 3 in Section 5.1, which asserts that as all source sample sizes grow with the target sample size fixed, the TRP posterior mode converges to the modified Trans-Lasso estimator. The proof is deferred to Appendix F, and Appendix G for Theorem 4, but the manuscript contains no appendices. This is not a cosmetic omission. The multivariate limit has to hold despite the nonsmooth, non-separable TRP objective for which Section 3.3 shows even block-coordinate descent fails to reach the global MAP. The univariate derivation in Eq. (28) establishes agreement only in the P=1, diagonal-precision case, and the ridge first step in Eq. (23) is not the Lasso first step of the original Trans-Lasso. If Theorem 3 requires additional conditions not stated in the theorem — for example that the ridge penalty scale as o(N_k), or that target and source designs be orthogonal — then the approximate-MAP interpretation is not established in the finite-source regime the paper targets. The GTEx experiment does not test this claim because it uses the Laplace TRP with full MCMC and finite K, not the MAP limit of Theorem 3.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a new joint Bayesian prior, the Total Risk Prior (TRP), for transfer learning in linear models. Rather than placing a prior on the target coefficients centered at an empirical minimizer computed from the source data, the prior centers β₀ on the minimizer of the expected squared loss over source datasets conditional on the source parameters βₛ, i.e., on a risk minimizer T(βₛ). This construction yields a formal joint prior over all source and target parameters, in contrast to two-stage methods. For quadratic penalties, T is linear and the resulting Gaussian or Laplace TRP leads to conjugate or auxiliary-variable Gibbs updates. The paper contributes a scalable Gibbs sampler using a rank-P update to avoid full Cholesky decompositions of size (K+1)P, along with parallel tempering for the dataset-inclusion indicators η. In the MAP analysis, the paper argues that a block-coordinate-descent view of the modified Trans-Lasso (with a ridge first step) corresponds to approximate optimization of the TRP posterior, and Theorem 3 is stated to show that the posterior mode converges to the modified Trans-Lasso as source sample sizes grow. Theorem 4 studies the asymptotic behavior of the inclusion indicators under the Gaussian TRP. The empirical section compares TRP with Trans-Lasso, pooled OLS, and target-only Lasso on GTEx gene-expression data across K = 4, 8, 16, 32 source tissues, reporting that TRP improves median out-of-sample MSE, especially for small K.","tokens_in":74,"tokens_out":5277,"duration_ms":180112,"significance":"The conceptual idea of using a risk minimizer conditional on source parameters as a prior hyperparameter is original and addresses a genuine gap between formal Bayesian hierarchical transfer and frequentist minimax transfer learning. If the stated theorems can be verified, the approximate-MAP connection to Trans-Lasso would provide a useful interpretive bridge and could help justify Bayesian uncertainty quantification in transfer settings. The computational machinery—notably the Theorem 2 sampling algorithm and the open-source JAX implementation—is a concrete asset, and the GTEx evaluation is out-of-sample with multiple baselines and repeated random splits. However, the central theoretical claims (Theorems 2–4) are unverifiable in the submitted version because their proofs are deferred to appendices that are not present in the manuscript, and the theorem statements themselves contain ambiguities. The contribution is therefore promising but not yet established.","major_comments":[{"comment":"The proofs of Theorems 3 and 4 are deferred to Appendices F and G, but the arXiv v1 text contains no appendices at all. Theorem 3 is the load-bearing result for the paper's headline claim that minimax-frequentist transfer learning can be viewed as an approximate MAP procedure; Theorem 4 is the basis for the claim about dataset-inclusion behavior. Without these proofs, neither claim can be checked. Theorem 2's proof is also deferred to a missing Appendix E. This is not a cosmetic omission: the derivation must handle a nonsmooth, non-separable objective for which the paper itself shows in Section 3.3 that block-coordinate descent fails to reach the global optimum.","section":"Section 5.1 and 5.2, Theorems 3 and 4"},{"comment":"The abstract states that 'recently proposed minimax-frequentist transfer learning techniques may be viewed as an approximate Maximum a Posteriori approach to our model,' and Section 1.2 promises 'an approximate MAP-Bayesian perspective of Li et al. (2022).' However, Theorem 3 concerns the 'modified Trans-Lasso estimator' whose first step is a ridge regression (Eq. 23), not the Lasso first step of the original Trans-Lasso. As stated, the theorem only establishes convergence to this modified estimator, and only in the limit of infinite source sample sizes with the target sample size fixed. The paper should either prove the result for the original Trans-Lasso, or clearly restrict the interpretive claims to the ridge-modified version.","section":"Section 5.1, Theorem 3 vs. abstract and Section 1.2"},{"comment":"The theorem statement lacks conditions needed for a well-defined limit. The objective in Eq. (25) is convex but not strictly convex in general, so the posterior mode may not be unique; the statement 'let β*₀ denote the argument maximizing the posterior density' implicitly assumes uniqueness. No scaling conditions are given for λ_t or τ as N_k → ∞, and no conditions are given on the design matrices other than convergence of the source Gram matrices, so it is unclear whether the limit is well-defined. The proof, even if supplied, would need to address these issues; as written, the theorem is ambiguous.","section":"Section 5.1, Theorem 3 statement"}],"minor_comments":[{"comment":"The phrase 'thistransfer learning' in the abstract is missing a space.","section":"Abstract"},{"comment":"The right-panel label 'OLS Hier TRP' is cryptic; the caption should spell out which box corresponds to which method and what the 10 repetitions are.","section":"Section 1.1, Figure 1"},{"comment":"There is a typo in the statement: 'Choleskdy decomposition' should be 'Cholesky decomposition.'","section":"Section 4.1, Theorem 2"},{"comment":"The word 'sufficienlty' should be 'sufficiently.'","section":"Section 4.2"},{"comment":"The notation '1/N_k X_k^T X_k → Σ_k^{-1}' is confusing because Σ is used elsewhere for covariance matrices; if the limit of the Gram matrix is denoted Σ_k^{-1}, then Σ_k is the asymptotic covariance of the OLS estimator, which should be stated explicitly. Also, the 'consistent OLS estimators' assumption is redundant given the convergence of the Gram matrices.","section":"Section 5.1, Theorem 3 notation"},{"comment":"The caption of Figure 4 misspells 'Auxilliary' (should be 'Auxiliary'), and the text says 'about 400 predictor variables' while the setup describes P = 399; these should be harmonized.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The complete absence of the appendices in the arXiv v1 text is unusual and is the main blocker. The core idea is interesting and the empirical contribution is reasonable, so I would not recommend rejection unless the missing proofs reveal a fatal gap. I would ask the authors to provide the full manuscript with appendices and to tighten the theorem statements and the abstract's claims about the connection to the original Trans-Lasso."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper has one solid new idea and a real application, but its central theory is presented as done when it isn't, because the appendices those results depend on are not in the arXiv v1. If you're working on Bayesian transfer learning, read it for the construction in Section 2; just don't take Theorems 3–4 on faith.\n\nWhat's new: The Total Risk Prior is a clean departure from the usual 'fit a source estimator, slap it in the prior' maneuver. Using the minimizer of expected risk conditional on source parameters keeps source estimation inside the joint model, which gives you a formal posterior and a principled way to average over source inclusion indicators. The Gibbs sampler with the auxiliary variables is explicit and checkable, and the GTEx study shows real gains over Trans-Lasso when sources are scarce. I also give them credit for the Simpson's paradox example, which clearly explains why hierarchical shrinkage can fail where the TRP works.\n\nWhere it's soft: the proofs for Theorems 3 and 4 are in Appendices F and G, which don't exist in this version. That is not cosmetic. Theorem 3 is the bridge that turns Trans-Lasso into 'an approximate MAP of our model,' and the stress-test note is right that the univariate derivation in Eq. (28) is only the P=1 case. The ridge first step of the modified Trans-Lasso is not the Lasso first step of the original Trans-Lasso, so the abstract overstates the connection. Also, the code link in the paper is broken (it points to a repo with a space in the name), and there are no convergence diagnostics for the MCMC runs. These are all fixable, but as written the strong claims are not verifiable.\n\nThe paper is honest about its own limitations—Section 7.2 names the computational cost, and Theorem 4 explicitly says the posterior on eta does not concentrate. That's good. The circularity worry I sometimes have with transfer methods doesn't apply here: T(beta_S) depends on the source designs, not the target response, so the prior is not fitting the target labels.\n\nBottom line: I'd send it to a serious referee, but the referee should be told to require the missing appendices and the code fix. The TRP idea is worth engaging with, and I expect a revised version to be a solid contribution.","headline":"The TRP construction is genuinely novel and worth a careful read, but the central asymptotic claims are uncheckable because the proofs are missing from this version.","tokens_in":18742,"tokens_out":2238,"would_cite":true,"duration_ms":21555,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62J07","62C10","65C40"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes the Total Risk Prior, a joint Bayesian prior that links the target coefficient to the expected-risk minimizer over source datasets, giving full uncertainty quantification and data-driven source selection.","keywords":["transfer learning","Total Risk Prior","Bayesian Lasso","Gibbs sampling","model averaging","negative transfer","linear regression","uncertainty quantification"],"falsifier":"Run the paper's sampler and Trans-Lasso on a synthetic linear model in which the target coefficient is deliberately placed far from the risk-minimizing source average, with limited source data; if the TRP posterior's predictive mean-squared error is worse than target-only Lasso and the credible intervals are poorly calibrated, the prior's location assumption fails in that regime.","tokens_in":17743,"feed_emoji":"🧬","tokens_out":8980,"duration_ms":83535,"temperature":0.7,"pith_summary":"The paper's central claim is that Bayesian transfer learning can be made fully formal by placing a prior on all source and target parameters together, instead of first estimating source parameters and feeding them into the target prior. The link is the Total Risk Prior: the target coefficient is encouraged to be near the minimizer of expected squared error over the source mechanisms, evaluated at the current source coefficients, rather than near the empirical source-data fit. This yields a joint posterior in which source-parameter uncertainty is propagated, source datasets can be included or excluded by Gibbs sampling over binary indicators, and the method inherits the behavior of frequentist transfer learning. The paper supports this with theory (the posterior mode approaches a modified Trans-Lasso estimator as source sample sizes grow) and with a GTEx gene-expression application in which the method's predictive mean-squared error is lower than Trans-Lasso's, especially when few source datasets are available.","feed_headline":"Risk-minimizing prior beats Trans-Lasso on limited sources","feed_subtitle":"A single joint prior propagates source uncertainty, selects informative sources, and beats Trans-Lasso on gene-expression data.","key_machinery":"The load-bearing object is the transfer operator, defined for a candidate set of source coefficients as $T_\\eta(\\beta_S) = \\arg\\min_\\beta \\frac{1}{2} \\sum_k \\eta_k (\\beta_k - \\beta)^T X_k^T X_k (\\beta_k - \\beta) + \\gamma(\\beta)$. With $\\gamma(\\beta) = \\tau\\|\\beta\\|_2^2$, it simplifies to the linear transfer matrix $T_\\eta$, a precision-weighted average of source coefficients. This operator converts the conceptual claim that the target should be located where the source mechanisms predict best in aggregate into a computable prior density $P(\\beta_0,\\ldots,\\beta_K) \\propto d(-\\lambda_t \\|\\beta_0 - T_\\eta \\beta_S\\|)$, and it is the reason the model can be both fully Bayesian and behaviorally similar to frequentist transfer learning. The same operator defines the transformed coordinates in which the Laplace prior becomes a Bayesian Lasso, and it is the object whose large-sample limit links the posterior mode to Trans-Lasso.","core_discovery":"The discovery is that the target parameter's prior location should be the regularized minimizer of expected loss conditional on source parameters, not the minimizer of empirical loss on observed source data. The paper defines the transfer operator $T_\\eta(\\beta_S)$ as this risk minimizer, and with an $\\ell^2$ ridge penalty inside the operator it becomes the linear map $T_\\eta = (\\sum_k \\eta_k X_k^T X_k + \\tau I)^{-1}[\\eta_1 X_1^T X_1 \\ldots \\eta_K X_K^T X_K]$, i.e., a precision-weighted average of the source coefficients. Placing $\\beta_0$ near $T_\\eta(\\beta_S)$ through a Laplace or Gaussian coupling gives a joint prior over all coefficients; the Laplace version is a Bayesian Lasso in the transformed coordinates $z = B\\beta_A$. Theorem 3 states that as the source sample sizes go to infinity, the posterior mode converges to a modified Trans-Lasso estimator, so minimax frequentist transfer learning can be read as an approximate maximum-a-posteriori procedure under this prior.","pith_inferences":["Beyond the paper, the same risk-minimizer principle is not tied to squared error or linear models; a Total Risk Prior for generalized linear models or neural networks would be a natural next step, but the paper does not implement it.","Beyond the paper, the non-concentration of the inclusion indicators $\\eta$ (Theorem 4) suggests that in large samples the posterior reports irreducible uncertainty about which sources are relevant; treating inclusion probabilities as selection decisions rather than estimates would miss that message.","Beyond the paper, since the sampler's cost grows with the product of source count and covariate dimension, scaling to large source collections would likely require stochastic-gradient or variational approximations, which the paper leaves open."],"forward_implications":["Uncertainty in source-parameter estimates is propagated into the target posterior, so interval estimates for the target coefficient reflect the limited size of source datasets.","Bayesian model averaging over the inclusion indicators $\\eta$ gives a principled, automatic way to down-weight or exclude sources that cause negative transfer, without refitting the model.","With a Laplace coupling and an $\\ell^2$ penalty inside the transfer operator, the model reduces to a Bayesian Lasso in a transformed coordinate system, making existing Gibbs-sampling machinery directly applicable.","As source sample sizes grow, the posterior mode approaches a modified Trans-Lasso estimator, giving a Bayesian/MAP interpretation of the minimax frequentist transfer-learning procedure.","On the GTEx benchmark, the paper reports improved out-of-sample predictive mean-squared error relative to Trans-Lasso, with the largest gains when only a few source datasets are available."],"supporting_citations":[{"why":"Defines the Trans-Lasso estimator that serves as the frequentist baseline and the estimator to which the TRP posterior mode converges in the large-source limit.","marker":"Li et al. (2022)"},{"why":"Provides the Gibbs-sampling scheme for the Bayesian Lasso that the Laplace TRP inherits after transforming to $z = B\\beta_A$.","marker":"Park & Casella (2008)"},{"why":"Establishes the Laplace distribution as an exponential scale mixture of Gaussians, which yields the auxiliary-variable sampler used for the Laplace TRP.","marker":"Andrews & Mallows (1974)"},{"why":"Supplies a fast sampling algorithm for Gaussian scale-mixture priors, adapted here to avoid forming the full $(K+1)P \\times (K+1)P$ matrix in the $\\beta_A$ update.","marker":"Bhattacharya et al. (2016)"},{"why":"Motivates Gibbs sampling over the binary source-inclusion indicators $\\eta$ as a variable-selection style update.","marker":"George & McCulloch (1993)"},{"why":"Provides the inverse-gamma hierarchy for half-Cauchy priors used to update the regularization strengths $\\lambda_t$ and $\\lambda_p$.","marker":"Makalic & Schmidt (2015)"},{"why":"Gives the convergence conditions for block coordinate descent that the paper uses to explain why simple coordinate-descent fails on the nonsmooth TRP objective.","marker":"Tseng (2001)"}],"fun_headline_variants":["Total risk prior: Bayesian transfer learning with full uncertainty","Risk-conditional prior beats Trans-Lasso on scarce data","Bayesian Lasso for transfer learning via total risk prior","Joint prior for transfer learning that outperforms Trans-Lasso","New prior turns transfer learning into a fully Bayesian problem"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the target regression coefficient is close to the transfer location $T_\\eta(\\beta_S)$, the minimizer of expected squared error over the included source datasets; if the target mechanism is not near that precision-weighted average of source mechanisms, the prior pulls $\\beta_0$ in a biased direction, and the inclusion indicators provide only partial protection.","fun_headline_variants_meta":{"raw":{"variants":["Total risk prior: Bayesian transfer learning with full uncertainty","Risk-conditional prior beats Trans-Lasso on scarce data","Bayesian Lasso for transfer learning via total risk prior","Joint prior for transfer learning that outperforms Trans-Lasso","New prior turns transfer learning into a fully Bayesian problem"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000928,"raw_usage":{"total_tokens":4007,"prompt_tokens":1008,"completion_tokens":2999,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":624,"completion_tokens_details":{"reasoning_tokens":2920}},"tokens_in":624,"tokens_out":2999,"duration_ms":24539,"temperature":1.0,"reasoning_tokens":2920,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:24:58.078003+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's sampler and Trans-Lasso on a synthetic linear model in which the target coefficient is deliberately placed far from the risk-minimizing source average, with limited source data; if the TRP posterior's predictive mean-squared error is worse than target-only Lasso and the credible intervals are poorly calibrated, the prior's location assumption fails in that regime.","supporting_citations":[],"review_version":1}