{"id":"e4c49212-0638-4a9b-a83d-cf9bd7948d44","arxiv_id":"2505.13104","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A unified identification and estimation framework transports a general class of first-moment causal effect measures from a trial to a target population under covariate shift.","lead":"This paper introduces one statistical recipe for moving treatment effects from a clinical trial to a different target population, covering both absolute measures such as risk difference and relative measures such as risk ratio and odds ratio. The framework matters because clinical guidelines recommend reporting several effect measures, and most existing transport methods handle only the risk difference.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposition 6's influence function under Assumption 5 omits the chain-rule factor ∂1Γ(τ, μ_T0)∂1Φ(μ_S1, μ_S0), so Section 4's semiparametric estimators are not the claimed efficient or doubly robust estimators when baseline risks differ.","rationale":"The reader's verdict of CONDITIONAL is appropriate, but for a different primary reason than the one stated. The reader's weakest-assumption concern is that Section 4 requires target control outcomes while Section 2 records only covariate data; this is a real scope issue, yet the paper explicitly frames the weaker assumption as requiring access to target control outcomes, so it is not an internal contradiction. The more load-bearing problem is mathematical: Proposition 6's influence function omits the chain-rule factor that arises because τΦ is a nonlinear function of source conditional means while Γ is evaluated at the target baseline mean. When μ_T0 differs from μ_S0, which is precisely the case where Assumption 5 is weaker than Assumption 3, the printed EIF is wrong. This undermines the one-step and estimating-equation estimators in Section 4.2, the claimed semiparametric efficiency, and the Appendix C.3 formulas, while leaving the identification results intact. A corrected EIF is readily derivable, so the paper is fixable rather than irredeemable; hence the verdict should remain CONDITIONAL. I partially agree with the reader because the target-outcome data availability issue is real, but the EIF defect is more central and more specific.","tokens_in":34865,"tokens_out":15286,"duration_ms":150287,"concrete_test":"Re-derive Proposition 6 by applying the chain rule to ψ_T^1 = E_T[Γ(τΦ(X), μ_T0(X))], τΦ(X)=Φ(μ_S1(X), μ_S0(X)), and verify whether the source-treated coefficient equals ∂1Γ(τ, μ_T0)/∂1Γ(τ, μ_S0) rather than 1. As a numerical check, use the RR with a binary covariate, μ_S0=0.2, μ_T0=0.4, τ=2, so the correct treated coefficient is 2 while the paper's is 1. Simulate N=100,000 under Experiment 2 and compare the paper's one-step estimator with the corrected one-step estimator: if the paper's version shows first-order bias or fails to achieve the variance of the corrected version, the EIF error is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The identification formula (12) is correct, but the semiparametric claims built on it are not. Under Assumption 5, ψ_T^1 = E_T[Γ(τΦ(X), μ_T0(X))] with τΦ(X) = Φ(μ_S1(X), μ_S0(X)). Differentiating this functional with respect to the source nuisance functions gives an influence function whose source-treated coefficient is ∂1Γ(τ, μ_T0)∂1Φ(μ_S1, μ_S0) = ∂1Γ(τ, μ_T0)/∂1Γ(τ, μ_S0), and whose source-control coefficient is ∂1Γ(τ, μ_T0)∂0Φ(μ_S1, μ_S0) = -∂1Γ(τ, μ_T0)∂0Γ(τ, μ_S0)/∂1Γ(τ, μ_S0). Proposition 6 instead uses coefficient 1 on the treated term and -∂0Γ(τ, μ_T0) on the control term. These coincide only when μ_T0 = μ_S0, i.e. in the exchangeability-in-mean regime where Assumption 5 is not weaker. In the paper's own Experiment 2 the baseline risks differ by construction, so the printed EIF is not the efficient influence function. Consequently the one-step estimator is not a valid first-order bias correction, the 'estimating equation' estimator does not solve the efficient score, and the variance/efficiency claims in Section 4.2 are unsupported. The Appendix C.3 formulas for RR and OR inherit the missing factor because they are derived from Proposition 6; this includes the OR term involving ∂0Γ(τ, μ_T0) alone. Consistency of the EE estimator may survive when all nuisance models are correct, because the misspecified moment still has mean zero, but double robustness and efficiency do not follow, and no Proposition-5-type proof is supplied for Section 4. This is a concrete mathematical defect in a central claimed contribution, not merely a scope limitation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a unified framework for transporting first-moment population causal measures, i.e., estimands of the form τ_P = Φ(E_P[Y(1)], E_P[Y(0)]) for a broad class of effect measures Φ (RD, RR, OR, NNT, and others), from an RCT source population to a target population under covariate shift. Identification is studied under two assumptions: exchangeability in mean (Assumption 3) and exchangeability in effect measure (Assumption 5). For the former, the paper derives weighted Horvitz-Thompson, weighted/transported G-formula, and one-step/estimating-equation semiparametric estimators, with asymptotic variance expressions and a double-robustness result for the estimating-equation estimator (Proposition 5). For the latter, it gives an identification formula (Eq. (12)) based on target control outcomes, proposes Γ-formula estimators, and states an influence function (Proposition 6) used to build one-step and estimating-equation estimators. The theoretical results are supplemented by simulations and a real-data application on the CRASH-3 trial and Traumabase registry.","tokens_in":35183,"tokens_out":10852,"duration_ms":106589,"significance":"The paper addresses a real and practically important gap: most generalization methods focus on the risk difference, while clinical reporting routinely uses absolute and relative measures. If the Section 4 semiparametric results were correct, the paper would provide a genuinely useful unification, including the transport of non-collapsible measures such as the odds ratio under exchangeability in effect measure. The strengths of the manuscript include the clean identification algebra in Sections 3 and 4.1, the explicit double-robustness proof for the estimating-equation estimator under exchangeability in mean (Proposition 5 with proof in Appendix B.5), the closed-form asymptotic variance computations under logistic/linear nuisance models, and the extensive simulation study. These contributions are substantive. However, the influence function stated in Proposition 6 is incorrect when baseline risks differ between source and target, and this invalidates the main semiparametric claims of Section 4. The identification formula itself appears sound, so the result is correctable, but the Section 4 estimator development and its efficiency/double-robustness claims need substantial rework.","major_comments":[{"comment":"Proposition 6 is not the influence function of ψ_T1 under Assumption 5. From Eq. (12), ψ_T1 = E_T[Γ(τΦ(X), μ_T0(X))] with τΦ(X) = Φ(μ_S1(X), μ_S0(X)). Differentiating through Γ and τΦ gives a source-treated coefficient of ∂1Γ(τ, μ_T0)∂1Φ(μ_S1, μ_S0) = ∂1Γ(τ, μ_T0)/∂1Γ(τ, μ_S0) and a source-control coefficient of ∂1Γ(τ, μ_T0)∂0Φ(μ_S1, μ_S0) = -∂1Γ(τ, μ_T0)∂0Γ(τ, μ_S0)/∂1Γ(τ, μ_S0). The printed φ1 instead uses coefficient 1 on the treated residual and -∂0Γ(τΦ, μ_T0) on the control residual. These coincide with the correct coefficients only when μ_T0 = μ_S0, i.e., exactly in the regime where Assumption 5 is not weaker than Assumption 3. The proof in Appendix C.2 actually derives the factor ∂1Γ(τ, μ_T0) multiplying IF(τΦ(x)), so the discrepancy is in the displayed formula of Proposition 6.","section":"Section 4.2, Proposition 6"},{"comment":"As a consequence of the error in Proposition 6, the one-step and estimating-equation estimators of Section 4, as well as the explicit RD/RR/OR expressions in Appendix C.3, are not first-order correct or efficient under Assumption 5 when baseline risks differ. The estimating-equation estimator may still be consistent when all nuisance models are correctly specified, because the printed moment has mean zero at the truth, but the claimed double robustness and the variance/efficiency statements in Section 4.2 are unsupported; no Proposition-5-type double-robustness proof is supplied for Section 4. The simulation setting of Experiment 2 has μ_T0 ≠ μ_S0 by construction, so the unbiasedness reported there is not evidence for the printed estimators' efficiency or double robustness. The authors should either correct Proposition 6 and all downstream formulas, or restrict the semiparametric claims to the case μ_T0 = μ_S0 and clearly label the general case as an open problem.","section":"Section 4.2 / Appendix C.3"},{"comment":"There is an internal inconsistency in the data framework. Section 2 states that the target dataset contains only covariates (X_i)_{i∈[m]}, while Eq. (12), Eq. (13), Definition 4, and all Section 4 estimators require E_T[Y(0)] and μ_T0(X). The prose says Assumption 5 'requires access to control outcomes in the target population,' but this requirement is not incorporated into the formal sampling model. The claimed weakness of Assumption 5 relative to Assumption 3 is therefore misleading: it replaces an outcome-transportability assumption with an additional data requirement. The formal setup should be amended to specify how target control outcomes arise (e.g., an additional untreated sample from the target), or Section 4 should be explicitly presented as a different data regime rather than as a weaker assumption within the same framework.","section":"Section 2 vs. Section 4.1"}],"minor_comments":[{"comment":"Several assumption references are broken, e.g., 'Assumption 1 to 1' in Propositions 12-14 and 'Assumption 1 to 3.1' in Proposition 9; these should be corrected to the intended numbered assumptions.","section":"Throughout appendices"},{"comment":"The caption 'Source values are 0.45 / 3.2 / 7.5' does not state which of the three numbers corresponds to the risk difference, risk ratio, and odds ratio; please spell this out.","section":"Figure 1 caption"},{"comment":"'SUTV A' appears to be a typo for 'SUTVA'; please fix throughout.","section":"Section 2.1, Assumption 1"},{"comment":"The sentence 'which is is related to φ1' contains a duplicated 'is'; please correct.","section":"Section 4.2, first paragraph"}],"recommendation":"major_revision","confidential_remarks":"The paper draws heavily on the authors' own prior work (Colnet et al. 2023, 2024; Boughdiri et al. 2024) and on Dahabreh et al. (2020) for the estimating-equation estimator under exchangeability in mean; the genuinely new material is the unified measure class and the Assumption 5 identification. The error in Proposition 6 is the central technical obstacle and must be corrected before any consideration of publication. Because the identification formula and the Section 3 development are sound and the fix is localized to the influence-function computation and its downstream estimators, I would not recommend rejection, but the Section 4 semiparametric claims need substantial rework and re-simulation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper does something genuinely useful. It gives a unified treatment of first-moment causal measures under covariate shift, and the identification formula (12) for transport under exchangeability in effect measure is correct and new for non-collapsible measures like the odds ratio. The Gamma-formula estimators in Section 4.1 are a real contribution. The rest is a mix of solid, careful work and one load-bearing mistake.\n\nThe mistake is in Proposition 6. The EIF for psi_T1 under Assumption 5 should carry the factor d1Gamma(tau, mu_T0)/d1Gamma(tau, mu_S0) on the source-treated term and the analogous ratio on the control term. The proof applies the identity d1Phi(a,b) d1Gamma(Phi(a,b), b)=1 with b=mu_T0, but that identity only holds at b=mu_S0, the baseline used to define tau_S_Phi. When mu_T0 != mu_S0—exactly the regime where Assumption 5 is weaker than Assumption 3—the printed EIF is not the efficient influence function. The Appendix C.3 formulas for RR and OR are internally inconsistent with Proposition 6 (they use mu_S0 where Proposition 6 says mu_T0), so the estimating equation estimators in Section 4 do not have the claimed double robustness or efficiency. Consistency with all nuisances correct may survive, but that is not what is claimed.\n\nThere is also a data-framework problem. Section 2 says the target dataset has only covariates, but Section 4's identification and estimators require target control outcomes Y(0). The prose acknowledges this when treatment has not been deployed, but the formal setup never changes. That should be stated cleanly, not left as an implicit special case. No code or data is shipped either, which is a minor issue for a methods paper.\n\nTo be fair, the generic Phi/Gamma formulation is clean. The identification under Assumption 5 for non-collapsible measures is correct and is the sort of result people will cite. Section 3 estimators are carefully derived, the variance computations are detailed, and the simulations are honest about misspecification. This deserves a serious referee and a major revision. The identification and Section 3 work should survive; Section 4.2 needs to be redone. I would not desk-reject it.","headline":"Useful unification and a correct identification formula for non-collapsible measures, but the Section 4.2 influence function is missing a chain-rule factor that breaks the efficiency and double-robustness claims.","tokens_in":35806,"tokens_out":5906,"would_cite":true,"duration_ms":52769,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that a broad class of population causal effect measures—absolute and relative, collapsible and not—can be identified and consistently estimated in a target population under covariate shift, and provides estimators with…","keywords":["transportability","generalizability","first-moment causal measures","covariate shift","effect measure exchangeability","non-collapsible measures","doubly robust estimation","semiparametric efficiency"],"falsifier":"Simulate a binary outcome with source-specific baseline outcome levels but a constant conditional odds ratio across source and target, so Assumption 3 fails while Assumption 5 holds, and compare the proposed Gamma-formula and estimating-equation estimators for the target odds ratio against the true value using correct nuisance functions; if they are biased even with correct $r$ and $\\mu_T^{(0)}$, the identification formula fails, whereas unbiasedness while standard reweighting fails would support the paper's claim.","tokens_in":34593,"feed_emoji":"📊","tokens_out":8193,"duration_ms":80924,"temperature":0.7,"pith_summary":"The paper aims to show that one unified machinery can transport almost any commonly reported causal effect measure—not just the risk difference but also risk ratios, odds ratios, number needed to treat, and similar quantities—from a randomized trial to a different target population. Under covariate shift, it identifies the target estimand $\\Phi(E_T[Y(1)], E_T[Y(0)])$ whenever either exchangeability in mean holds, or the weaker exchangeability in effect measure holds with access to target control outcomes. It then builds weighting, regression, one-step, and estimating-equation estimators, proving asymptotic normality for the classical ones and double robustness for estimating-equation estimators across the whole class. The practical payoff is that clinical and policy audiences who need several effect measures on different scales could use a single framework instead of per-measure ad hoc methods. A notable corollary is that non-collapsible measures such as the odds ratio become transportable under the weaker assumption, a result the paper says had not been derived before.","feed_headline":"RCT effects on any scale can be transported to target populations","feed_subtitle":"A single estimation recipe covers absolute and relative measures, including the odds ratio.","key_machinery":"The load-bearing object is the first-moment causal measure, a functional $\\Phi(\\psi_1,\\psi_0)$ of the two marginal potential-outcome means, paired with its effect function $\\Gamma(\\cdot,\\psi_0)$, which inverts $\\Phi$ at a fixed baseline. The effect function does the heavy lifting: under exchangeability in effect measure, the paper averages $\\Gamma(\\tau_{S,\\Phi}(X),\\mu_T^{(0)}(X))$ over the target covariate distribution, converting a conditional effect on one scale into a conditional potential-outcome mean that can be averaged despite non-collapsibility. The density ratio $r(X)=P_T(X)/P_S(X)$ carries the covariate shift in the reweighting estimators, and the efficient influence function (EIF), computed by the chain rule through $\\Phi$ and $\\Gamma$, generates both the one-step and estimating-equation estimators and their double-robustness property.","core_discovery":"The central claim is that for any first-moment population causal measure $\\tau_{T,\\Phi} = \\Phi(E_T[Y(1)], E_T[Y(0)])$, the target-population value is identifiable from a randomized trial plus target covariates under covariate shift. Under exchangeability in mean, identification runs through three equivalent formulas: $E_T[Y(a)] = E_T[\\mu_S^{(a)}(X)] = E_S[r(X)Y(a)] = E_S[r(X)\\mu_S^{(a)}(X)]$, where $r(X) = P_T(X)/P_S(X)$ is the density ratio between target and source covariate distributions. Under the weaker exchangeability in effect measure, the paper identifies $\\tau_{T,\\Phi}$ through the effect function $\\Gamma$ as $\\Phi(E_T[\\Gamma(\\tau_{S,\\Phi}(X),\\mu_T^{(0)}(X))], E_T[Y(0)])$, which lets even non-collapsible measures such as the odds ratio be transported once target control outcomes are observed. For both settings the paper provides weighting and regression estimators with closed-form asymptotic variances, plus semiparametric one-step and estimating-equation estimators, and proves the estimating-equation estimators are doubly robust for every measure in the class.","pith_inferences":["If target control outcomes are unavailable, as in the paper's Section 2 data setup where the target dataset contains only covariates, the exchangeability-in-effect-measure branch requires data that the framework does not otherwise assume; the weaker assumption is weaker statistically but not cheaper to satisfy.","Because one-step and estimating-equation estimators coincide only for linear functionals and diverge for nonlinear ones such as the odds ratio, estimator choice should depend on the reported measure, and one could construct a test of whether the one-step correction is negligible by comparing bootstrap distributions.","The paper's stated extension to multiple RCTs suggests a path toward a causal meta-analysis that transports both absolute and relative measures from several trials; formal pooled estimation with multiple source populations is left implicit.","The variance-ordering result for linear outcome models may extend to nonparametric outcome regression under oracle rates, which would give measure-specific guidance on when to prefer transported G-formula over weighting estimators."],"forward_implications":["All causal measures in the first-moment class—including risk difference, risk ratio, odds ratio, number needed to treat, excess risk ratio, survival ratio, and log-odds ratio—can be estimated in a target population under covariate shift from a single estimation recipe.","Under exchangeability in effect measure, non-collapsible measures such as the odds ratio become transportable, provided target control outcomes or their conditional mean are available.","Estimating-equation estimators are doubly robust for every first-moment measure, so consistency holds if either the outcome regression or the density-ratio model is correctly specified.","For linear outcome models, the asymptotic variances satisfy $V_{tG} \\le V_{wG} \\le V_{wHT}$, meaning the transported G-formula is the most efficient of the classical estimators in that setting.","Closed-form asymptotic variances for the weighted Horvitz-Thompson and G-formula estimators enable standard confidence intervals for all transported measures."],"supporting_citations":[{"why":"Supplies the formal transportability framework that the paper extends to general first-moment causal measures.","marker":"Pearl and Bareinboim (2011)"},{"why":"Provides the propensity-score generalization setup for moving trial results to a target population.","marker":"Stuart et al. (2011)"},{"why":"Gives the transported G-formula and estimating-equation estimator for the risk difference that the paper generalizes.","marker":"Dahabreh et al. (2020)"},{"why":"Examines collapsibility of RD, RR, and OR under distribution shift and provides prior identification results for collapsible measures that the paper extends to non-collapsible ones.","marker":"Colnet et al. (2023)"},{"why":"Establishes the reweighting framework for RCT generalization under covariate shift, including density-ratio estimation.","marker":"Colnet et al. (2024)"},{"why":"Provides risk-ratio estimation strategies and the doubly robust one-step result for the RR that the paper builds on.","marker":"Boughdiri et al. (2024)"},{"why":"Defines confounding and collapsibility, the concept used to distinguish collapsible from non-collapsible measures.","marker":"Greenland et al. (1999)"},{"why":"Supplies the efficient-influence-function machinery used to construct one-step and estimating-equation estimators.","marker":"Kennedy (2022)"},{"why":"Provides the M-estimation theory used to derive asymptotic variances for the weighting and G-formula estimators.","marker":"Stefanski and Boos (2002)"}],"fun_headline_variants":["From risk difference to odds ratio: one transport recipe","Transport any causal measure, absolute or relative, to any target","One estimation recipe for absolute and relative causal measures","A unified framework for transporting effect measures to any target","Transporting effects: unified estimation for absolute and relative measures"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that target-population control outcomes (or their conditional mean $\\mu_T^{(0)}(x)$) are available for the exchangeability-in-effect-measure branch, because Equations (12)-(13) and all Section 4 estimators require $E_T[Y(0)]$ and target control information; if the target dataset contains only covariates, that branch cannot be computed.","fun_headline_variants_meta":{"raw":{"variants":["From risk difference to odds ratio: one transport recipe","Transport any causal measure, absolute or relative, to any target","One estimation recipe for absolute and relative causal measures","A unified framework for transporting effect measures to any target","Transporting effects: unified estimation for absolute and relative measures"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001672,"raw_usage":{"total_tokens":6667,"prompt_tokens":1018,"completion_tokens":5649,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":634,"completion_tokens_details":{"reasoning_tokens":5571}},"tokens_in":634,"tokens_out":5649,"duration_ms":37815,"temperature":1.0,"reasoning_tokens":5571,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:19:42.664725+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a binary outcome with source-specific baseline outcome levels but a constant conditional odds ratio across source and target, so Assumption 3 fails while Assumption 5 holds, and compare the proposed Gamma-formula and estimating-equation estimators for the target odds ratio against the true value using correct nuisance functions; if they are biased even with correct $r$ and $\\mu_T^{(0)}$, the identification formula fails, whereas unbiasedness while standard reweighting fails would support the paper's claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the formal transportability framework that the paper extends to general first-moment causal measures."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the propensity-score generalization setup for moving trial results to a target population."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the transported G-formula and estimating-equation estimator for the risk difference that the paper generalizes."},{"cited_title":"Josse, and E","cited_arxiv_id":null,"evidence_quote":"Provides risk-ratio estimation strategies and the doubly robust one-step result for the RR that the paper builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines confounding and collapsibility, the concept used to distinguish collapsible from non-collapsible measures."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the M-estimation theory used to derive asymptotic variances for the weighting and G-formula estimators."}],"review_version":1}