{"id":"9325fec2-fba4-420e-aa38-558a68a998a0","arxiv_id":"2507.17686","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A debiased maximum-likelihood estimator for hazard ratios in a baseline-hazard-free exponential model with kernel ML adjustment recovers true causal hazard ratios in simulations, but a key convergence assumption is left as a conjecture.","lead":"This paper proposes replacing the Cox model's baseline hazard with a kernel-based machine-learning function to estimate causal hazard ratios in observational survival data. It could matter for analyzing electronic health records where treatments and covariates change over time.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 9 is a conjecture, yet Proposition 2's near-orthogonality and Appendix XI's DML rates both depend on it; without the range condition H*_fθ=H*_ffρ the Gateaux derivative is only O(n^-β), not o(n^-1/2).","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: Assumption 9 is stated as a conjecture, yet the asymptotic unbiasedness proof in Appendix XI relies on it. I agree with that assessment. My reading sharpens the issue slightly: within Assumption 9, the source condition H*_{fθ}=H*_{ff}ρ is the part that is hardest to justify and also the part that makes the near-orthogonality remainder O(n^{-α-β}) rather than O(n^{-β}). Without it, the score derivative with respect to f can be too slow relative to n^{-1/2}, and the debiasing guarantee is unsupported. This is not an internal logical inconsistency; it is an unproved but explicitly flagged assumption. The simulations are favorable but use smooth Gaussian-kernel DGPs where the density ratio is likely in the RKHS, so they provide limited evidence about the general validity of the condition. The proposed step-function test would directly probe whether the condition is essential. Since the reader already recommended CONDITIONAL and my analysis does not change that, I set verdict_should_be to UNCHANGED.","tokens_in":31031,"tokens_out":8339,"duration_ms":96196,"concrete_test":"Run the debiased estimator on a simulation satisfying Assumptions 1–8 but with the treatment-assignment odds ratio set to a discontinuous step function of X, so that dν_k/dν_0 is not in the chosen Gaussian RKHS, while keeping the true hazard f* smooth and inside the RKHS. Vary n from 1,000 to 20,000 and check whether the t-statistics of the debiased θ̂ remain centered with calibrated variance. If bias persists, the range condition in Assumption 9 is load-bearing; if not, the condition is stronger than needed and should be relaxed or proved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The core debiasing guarantee rests on Assumption 9 (Section II C), which the paper itself labels a conjecture in Remark 4-1 and again in Appendix IX. The assumption has three parts: n^{-β} convergence of (θ̂,f̂) and Ĥ to their limits, and the range/source condition H*_{fθ_k}=H*_{ff}ρ_k. The third part is not a technical ornament. In the proof of Proposition 2 (Appendix VII, Eqs. (58)-(59)), the Gateaux derivative of the score with respect to f is ζ H*_{fθ}(H*_{ff}+ζ)^{-1}(f-f*). The paper bounds this as O(n^{-α-β}) only by substituting H*_{fθ}=H*_{ff}ρ and using ‖H*_{ff}(H*_{ff}+ζ)^{-1}‖≤1. Absent that substitution, the same expression is only O(n^{-β}); because β<1/2 by assumption, this is not o(n^{-1/2}), so Neyman near-orthogonality condition (16) can fail. Appendix XI's verification of Chernozhukov et al.'s Assumption 3.4 starts from the same n^{-β} rates and derives the requirement β>63/154, so the whole asymptotic-normality argument is conditional on an unproved conjecture. Appendix IX gives only a density-ratio plausibility argument for ρ_k ∈ H_k and explicitly leaves the stochastic-process convergence of the averaged Hessians open. The simulations use smooth Gaussian-kernel DGPs for which the range condition is likely satisfied, so they do not probe this gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes abandoning the unspecified baseline hazard in favor of an exponential hazard model h(t)=exp(θ'A_t+f(X_t)) with f in a reproducing-kernel Hilbert space, and develops debiased machine-learning estimators of the log-hazard-ratio θ* based on Neyman-orthogonal scores. It gives a causal interpretation of θ* under explicit assumptions (Proposition 1), constructs two debiasing scores (Propositions 2 and 3), presents a cross-fitting algorithm with model selection by Bayesian model evidence, and reports simulation studies for observed confounders and for unobserved heterogeneity via a latent variable. The claimed asymptotic guarantees for the debiased estimator rely on Assumption 9, which the paper itself labels a conjecture in Remark 4-1 and Appendix IX.","tokens_in":31451,"tokens_out":7291,"duration_ms":79417,"significance":"If the statistical claims were fully established, this would be a substantive contribution: it directly targets the well-known uninterpretability of Cox hazard ratios in dynamic observational settings and extends double/debiased machine learning to survival models with kernel-based nuisance functions. The paper is transparent about its assumptions, provides detailed derivations in the appendices, self-identifies Assumption 9 as a conjecture, and includes reproducible simulation code as supplementary material. The core mathematical construction (Propositions 2 and 3) is plausible and non-trivial. However, because the key convergence rates and the range condition for the Hessian are not proved, the paper currently delivers a conditional framework rather than the unconditional asymptotic-unbiasedness result announced in the abstract.","major_comments":[{"comment":"Assumption 9 is introduced in Section II.C and explicitly left as a conjecture in Remark 4-1 and Appendix IX, yet the near-orthogonality of Proposition 2 depends on it. The bound in Eq. (59) is O(n^{-α-β}) only after substituting H*_{fθ_k}=H*_{ff}ρ_k; absent that range condition, the Gateaux derivative in Eq. (58) is O(n^{-β}), which is not o(n^{-1/2}) because β<1/2. Thus Neyman near-orthogonality condition (16) can fail, and the asymptotic unbiasedness of the debiased estimator is unsupported unless Assumption 9 is proved. Appendix IX provides only a density-ratio plausibility argument for ρ_k∈H_k and explicitly leaves the stochastic-process convergence of the averaged Hessians open.","section":"Section II.C, Assumption 9 and Appendix VII, Eqs. (58)-(59)"},{"comment":"The verification of the Chernozhukov et al. conditions (Assumptions 3.3 and 3.4) in Proposition 4 derives the requirement β>63/154 using the same unproved nuisance convergence rates and the range condition from Assumption 9. Consequently, the statement in the abstract and Section II.C that the paper proves 'necessary convergence results' overstates what is established. The main theorem should be stated explicitly as conditional on Assumption 9, or Assumption 9 should be replaced by a verifiable condition with a proof for a reasonable class of data-generating processes.","section":"Appendix XI, Proposition 4"},{"comment":"The numerical experiments do not probe the load-bearing parts of Assumption 9. Both simulation DGPs use smooth Gaussian-kernel components and logistic density-ratio structures, precisely the regime in which the range condition H*_{fθ}=H*_{ff}ρ is most plausible, and the nuisance functions are estimated by correctly specified kernel models. The simulations therefore cannot distinguish the proposed debiased estimator from one that would fail when the range condition is violated or when nuisance convergence is slower than n^{-β}. A stress test with a misspecified or non-smooth nuisance, or with a deliberately violated range condition, would materially strengthen the evidence.","section":"Section III, simulations"}],"minor_comments":[{"comment":"The table heading contains a typo: 'Mathmatical notations' should read 'Mathematical notations'.","section":"Table I"},{"comment":"The sentence 'On can observe whether incorporation or removal of a component function...' contains a typo; 'On' should be 'One'.","section":"Section II.E"},{"comment":"The value of P2 used for the reported t-statistic histograms in Fig. 1(G,H) is not stated; the text says different values are used without specifying which one produced the displayed results.","section":"Section III.A"},{"comment":"The cross-fitting split in Step 2 (training D\\(D(m)∪D(m+1 mod M)), validation D(m+1 mod M), held-out D(m)) is nonstandard and should be clarified: the validation set is used only for hyperparameter selection and is never used for scoring, so the text should state explicitly that this preserves the required independence.","section":"Section II.F"},{"comment":"The source codes are described as available as supplementary materials but no permanent repository or DOI is given; consider providing an archival link.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest about its main gap: the authors explicitly flag Assumption 9 as a conjecture. This is not a case of hidden circularity or misconduct. The path to publication is clear: either prove Assumption 9 (or a weaker verifiable substitute) for a nontrivial class of processes, or reframe the paper's central claims as conditional results and supplement the simulations with a stress test that violates the range condition. I see no reason to reject outright, but the current version's headline claim of proved asymptotic unbiasedness is not yet supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does something real: it drops the baseline hazard from the hazard model to dodge the Cox model's selection-induced interpretability problem, fits an exponential hazard with an RKHS nuisance f, and builds two Neyman-orthogonal score constructions for debiased ML estimation of the treatment log-hazard ratio. Proposition 1 gives a clean causal interpretation of the target under explicit assumptions, and Propositions 2 and 3 are genuinely new score constructions; the density-ratio form in Proposition 3 is elegant and practically appealing. The simulations show the debiasing removes most of the naive ML bias, which is a useful sanity check.\n\nNow the soft spots, in proportion. The load-bearing issue is Assumption 9. The paper itself labels it a conjecture in Remark 4-1 and again in Appendix IX, and the central debiasing guarantee depends on it. The range condition H*_fθ = H*_ffρ is what turns the Gateaux derivative of the score from O(n^-β) into O(n^-α-β); without it, near-orthogonality at the required o(n^-1/2) rate fails. Appendix XI then derives conditions like β > 63/154, but that entire chain inherits the conjecture. So the asymptotic normality and unbiasedness claims are conditional on an unproved regularity assumption. That is not a fatal flaw if the conjecture is true, but it does mean the paper's central theorem is not fully proven. The plausibility argument in Appendix IX (density ratio belonging to the RKHS) is suggestive, not a proof.\n\nOther concerns are smaller. The simulations do not benchmark against Cox partial likelihood or marginal structural models, which would help calibrate the efficiency tradeoff they discuss. The latent-variable simulation initializes the EM near the true values, so it does not stress the optimization. And the code is only promised as supplementary on acceptance; for this kind of method, a public repository with the code would be the norm.\n\nIs the paper worth serious peer review? Yes. The idea is novel, the causal framing is careful, and the authors are explicit about what they have not proven. A good referee could either extract a proof of Assumption 9 under reasonable data-generating conditions or force the authors to state openly that the debiasing guarantee is conditional. I would send it out rather than desk reject. I would not cite it myself until the conjecture is resolved, but I would put it on a reading group agenda to think about whether the range condition is generic.","headline":"A genuinely new debiased-ML hazard-ratio estimator whose main asymptotic guarantee currently rests on an unproved conjecture.","tokens_in":31936,"tokens_out":1574,"would_cite":false,"duration_ms":19339,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62N02","62G05","62D20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Hazard ratios become causally interpretable when the Cox baseline hazard is replaced by kernel machine learning, and debiased ML estimators make them reliably estimable.","keywords":["hazard ratio","debiased machine learning","Neyman orthogonality","kernel methods","causal inference","survival analysis","latent variable models"],"falsifier":"Simulate data where treatment assignment depends on a covariate transformation that is not representable in the chosen RKHS (so the density ratio $d\\nu_k/d\\nu_0$ lies outside the kernel space), then run the debiased estimator for increasing $n$; if the $t$-statistics drift away from $N(0,1)$, the range-inclusion half of Assumption 9 is load-bearing and false in that setting.","tokens_in":30825,"feed_emoji":"🩺","tokens_out":7000,"duration_ms":68171,"temperature":0.7,"pith_summary":"The paper proposes that the Cox model's unspecified baseline hazard is the main obstacle to giving hazard ratios a causal reading, because it lets the same parameter values describe many different risk-set selection processes. To fix this, it works with an exponential hazard model $h(t|A_t,X_t)=\\exp(\\theta'A_t+f(X_t))$ that has no baseline hazard, with $f$ learned by kernel machine learning. Under consistency, conditional randomization of treatment, time-homogeneity, and correct specification, the maximum-likelihood parameter $\\theta^*$ equals the covariate-adjusted log hazard contrast between two counterfactual treatment schedules that branch at time $t$ (Proposition 1). Since kernel ML estimators of $\\theta$ are not root-$n$ consistent, the paper builds two Neyman-orthogonal score families and a cross-fitted debiased ML estimator, proving convergence under an assumption the authors flag as a conjecture. Simulations with nonlinear confounding and with an unobserved risk factor show the debiased estimates hit the true hazard ratios with minimal bias, unlike naive ML.","feed_headline":"Drop the baseline hazard, keep the hazard ratio causal","feed_subtitle":"Kernel ML plus Neyman-orthogonal scores estimates causal hazard ratios with minimal bias.","key_machinery":"The load-bearing object is the exponential hazard model without a baseline hazard, $h(t|A_t,X_t)=\\exp(\\theta'A_t+f(X_t))$, where $f$ lies in a reproducing-kernel Hilbert space built from multiple kernels. The likelihood's Hessian operator $H$ (with blocks $H_{\\theta f}$, $H_{ff}$ and regularization $\\zeta_n$) enters Proposition 2's score $\\phi=\\partial_\\theta\\ell-H_{\\theta f}(H_{ff}+\\zeta_n)^{-1}\\partial_f\\ell$, while Proposition 3's score uses a kernel-logistic nuisance $g_k$ that plays the role of an inverse treatment-probability weight marginalized over time. These scores are Neyman-orthogonal, meaning the Gateaux derivative with respect to the nuisance parameters vanishes at the truth, so slow kernel learning rates for $f$ and $g_k$ do not bias $\\theta$.","core_discovery":"The central discovery is that the maximum-likelihood solution of a baseline-hazard-free exponential model $h(t|A_t,X_t)=\\exp(\\theta'A_t+f(X_t))$, with $f$ in a reproducing-kernel Hilbert space, is a causally interpretable quantity: under Assumptions 1–8, $\\theta^*_k$ equals the log ratio of counterfactual hazards at time $t$ for two treatment schedules that agree before $t$ and differ only on a small interval, conditional on the observed covariate history. This removes the multiplicity of contradictory interpretations that the standard semiparametric proportional-hazards model permits, because the time evolution of the risk set is modeled explicitly by $f$ (and optionally latent variables) instead of being absorbed by an unestimated baseline hazard. The paper then constructs Neyman-(near-)orthogonal scores—one based on the likelihood Hessian operator (Proposition 2) and one based on inverse-probability-weighted integral balance (Proposition 3)—and shows that, when Assumption 9 holds, the debiased ML estimator is $\\sqrt{n}$-consistent and asymptotically normal.","pith_inferences":["If the range-inclusion condition $H^*_{f\\theta}=H^*_{ff}\\rho$ fails for a data-generating process with a propensity outside the chosen kernel space, the Proposition 2 score may lose its orthogonality; the Proposition 3 density-ratio score may then be the safer default.","The time-homogeneity assumption is a strong biological claim; a testable extension would allow a smooth baseline trend in $f$ and check whether $\\theta$ estimates remain stable across model choices.","Comparing Bayesian model evidence with and without time-elapsed covariates can be read as a diagnostic for unobserved risk factors—an implicit screen that could be formalized as a model-selection rule.","For continuous or nonlinear treatment effects, Proposition 3 does not apply but Proposition 2 does, suggesting a path to dose-response analyses beyond the binary treatment setting."],"forward_implications":["Hazard ratios estimated in this framework are causal contrasts between counterfactual treatment schedules, so they can be reported with the same interpretation as randomized hazard contrasts.","The debiased estimator is asymptotically normal with a plug-in standard error, allowing significance tests even though the kernel nuisance $f$ converges slower than $\\sqrt{n}$.","Model misspecification that leaves temporal risk changes unmodeled shows up as a violation of the time-homogeneity assumption, detectable by comparing Bayesian model evidence with and without time-elapsed covariates.","The method extends to latent-variable models, so unobserved risk factors that select the risk set can be modeled explicitly while the hazard ratio stays estimable with minimal bias.","The approach provides a machine-learning-friendly alternative to marginal structural Cox models for observational data with dynamic treatment and many time-dependent covariates."],"supporting_citations":[{"why":"Supplies the concrete examples where a correctly specified proportional-hazards model yields uninterpretable hazard ratios, the problem the paper aims to solve.","marker":"[5]"},{"why":"Provides the Neyman-orthogonal/debiased machine learning framework and the cross-fitting procedure the paper adapts.","marker":"[19]"},{"why":"Marginal structural Cox models serve as the standard alternative for dynamic observational treatment, which the paper's no-baseline model is designed to improve upon.","marker":"[17]"},{"why":"Supplies the reproducing-kernel Hilbert space theory that justifies $f$ and the Hessian operator formalism.","marker":"[22]"},{"why":"Provides the kernel-based analysis tools used for the theoretical treatment of the RKHS Hessians.","marker":"[23]"},{"why":"Defines the standard proportional-hazards partial-likelihood approach whose baseline hazard motivates the no-baseline alternative.","marker":"[35]"}],"fun_headline_variants":["Abandon baseline hazard, use kernel ML for causal HRs","Drop baseline hazard, debias ML for causal hazard ratios","No baseline hazard, no contradictory HRs via kernel ML"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire debiasing argument rests on Assumption 9—that nuisance estimators converge at a polynomial $n^{-\\beta}$ rate and that the Hessian block $H^*_{f\\theta}$ lies in the range of $H^*_{ff}$—which the authors themselves leave as a conjecture in Remark 4-1 and Appendix IX; if either piece fails, the claimed asymptotic unbiasedness of the debiased estimator is unsupported.","fun_headline_variants_meta":{"raw":{"variants":["Abandon baseline hazard, use kernel ML for causal HRs","Drop baseline hazard, debias ML for causal hazard ratios","No baseline hazard, no contradictory HRs via kernel ML"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00103,"raw_usage":{"total_tokens":4336,"prompt_tokens":939,"completion_tokens":3397,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":3343}},"tokens_in":555,"tokens_out":3397,"duration_ms":27387,"temperature":1.0,"reasoning_tokens":3343,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:19:21.315070+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate data where treatment assignment depends on a covariate transformation that is not representable in the chosen RKHS (so the density ratio $d\\nu_k/d\\nu_0$ lies outside the kernel space), then run the debiased estimator for increasing $n$; if the $t$-statistics drift away from $N(0,1)$, the range-inclusion half of Assumption 9 is load-bearing and false in that setting.","supporting_citations":[{"cited_title":"Estimating heterogeneous treatment eﬀects on survi val outcomes using counterfactual censoring unbiased transformations,","cited_arxiv_id":null,"evidence_quote":"Marginal structural Cox models serve as the standard alternative for dynamic observational treatment, which the paper's no-baseline model is designed to improve upon."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the reproducing-kernel Hilbert space theory that justifies $f$ and the Hessian operator formalism."},{"cited_title":"Consistency of the group lasso and mult iple kernel learning","cited_arxiv_id":null,"evidence_quote":"Defines the standard proportional-hazards partial-likelihood approach whose baseline hazard motivates the no-baseline alternative."}],"review_version":1}