{"id":"03451884-5374-45b4-98bf-0056af6b15ed","arxiv_id":"2505.05932","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The paper introduces a diffusion-prior piecewise exponential model that combines observed survival data with expert-elicited long-term hazard assumptions, sampled via a new transdimensional PDMP algorithm.","lead":"A team of statisticians built a new Bayesian method to estimate survival beyond the end of a clinical trial, combining trial data with expert beliefs about long-term risk. The method uses a flexible piecewise-hazard model and a fast sampling algorithm, demonstrated on colon cancer and leukaemia trial data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The uncorrected splitting-scheme bias in Section 3.2 is the key unvalidated link: if the PDMP does not target the exact posterior, all reported hazard and extrapolation estimates are suspect.","rationale":"The reader's weakest assumption identifies the splitting-scheme approximation error in Section 3.2, and I agree that this is the most load-bearing concern. The paper's second contribution is explicitly an efficient and correct sampler for a transdimensional posterior; if the sampler does not target the exact posterior, then both the methodological contribution and all reported quantitative results are undermined. The paper itself concedes the approximation and gives no diagnostic, while the exact line-search scheme is available for at least some of the drifts used in the applications, so the concern is directly testable. I also considered the absence of a simulation study, which is important, but it is a broader validation gap rather than a single identifiable assumption that, if false, would break the argument. The extrapolation rescaling in Section 3.6 is a related numerical-approximation concern, but the splitting bias is the more fundamental link because it affects the observation-period posterior from which extrapolation starts. The reader's CONDITIONAL verdict already reflects the need to resolve these issues, so my stress-test does not change the verdict; it sharpens the necessary check. No ad hominem or outside-consensus objection is intended: this is an internal soundness issue flagged by the paper's own text.","tokens_in":21188,"tokens_out":11259,"duration_ms":134256,"concrete_test":"Run the Section 4.1 colon cancer analysis under identical priors, data, and candidate-knot updates using both the splitting scheme (Section 3.2) and the exact line-search scheme, which the paper states is available for Random Walk and Gaussian Langevin drifts. Compare posterior means and 95% credible intervals for h(1), h(2), h(3), and E[Y] over (0, y_infinity), using 10 independent chains per sampler to estimate Monte Carlo error. If the between-sampler difference exceeds twice the pooled Monte Carlo standard error for any reported summary, the uncorrected splitting bias is material and the central sampling claim is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central methodological claim is that the sticky PDMP sampler targets the transdimensional posterior of the diffusion piecewise exponential model. Section 3.2 states that event times are generated with the splitting scheme of Bertazzi et al. (2023) and that this introduces a small approximation error which the authors found unnecessary to correct. No bound, diagnostic, or comparison against the exact line-search scheme is provided for the settings used in Sections 4.1 and 4.2. This is load-bearing because the splitting scheme replaces the continuous-time PDMP with a discrete-time approximation; if the resulting stationary distribution is biased, every posterior hazard, credible interval, and extrapolated mean survival estimate is biased, and the comparisons in Tables 1 and 2 become unreliable. The concern is compounded by Section 3.6, where extrapolation applies a further heuristic rescaling of (gamma, sigma) to reduce first-order discretisation bias, again without a displayed derivation or sensitivity analysis in the main text. The paper does not present a simulation study with a known hazard, so there is no external check that the sampler and the extrapolation procedure together recover calibrated posterior statements. A direct comparison with the exact sampler for the drifts where it is available is therefore the most concrete way to determine whether the approximation error actually matters.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the diffusion piecewise exponential model, a Bayesian survival model in which the piecewise constant log-hazard is assigned a discretised diffusion prior and the knot locations are assigned a Poisson process prior. The drift of the diffusion is used to encode prior information about long-term hazard behaviour, with the aim of combining observation-period learning with expert opinion when extrapolating beyond administrative censoring. Posterior inference is performed with a Piecewise Deterministic Markov Process sampler: the authors extend sticky PDMP dynamics from spike-and-slab priors to general transdimensional posteriors, using a Gibbs step to resample inactive candidate knots. Extrapolations are generated by continuing the discretised diffusion from the last observation-period state, with a proposed rescaling of the discretisation parameters. The method is illustrated on colon cancer data and CLL-8 trial data, with mean survival estimates compared against parametric, independent piecewise exponential, and M-spline models.","tokens_in":21294,"tokens_out":5405,"duration_ms":56125,"significance":"If the claims hold, the paper makes a useful contribution to Bayesian survival extrapolation for health technology assessment. The prior framework is more flexible than fixed-knot spline or random-walk priors, and the explicit separation of observation-period flexibility from extrapolation-period prior information is a principled response to the administrative censoring problem. The technical contribution of extending sticky PDMP samplers to general transdimensional posteriors is also of independent interest, and Proposition 3.1 gives a clean, parameter-free comparison of expected recurrence times under two discretisations. The authors are transparent that the splitting scheme introduces an approximation error and that extrapolation is driven by the prior. The main weaknesses are the absence of any quantification of the sampler's approximation error and the lack of a simulation-based calibration check for the extrapolation procedure.","major_comments":[{"comment":"The sampler generates PDMP event times with a splitting scheme, and the paper states that this introduces a small approximation error that was found unnecessary to correct. No bound, diagnostic, or comparison against the exact line-search scheme is provided for the settings of Sections 4.1 and 4.2. This is load-bearing because the splitting scheme replaces the continuous-time PDMP with a discrete-time approximation; if the stationary distribution is biased, every posterior hazard estimate, credible interval, and extrapolated mean survival estimate in Tables 1-3 is potentially biased. Since the exact scheme is available for the random walk, Gaussian Langevin, and Gompertz drifts (Section 3.2), I ask the authors to compare posterior summaries from both schemes for at least one of the data analyses, and to report a calibration check on synthetic data with a known hazard.","section":"Section 3.2 and Sections 4.1-4.2"},{"comment":"The extrapolation period is sampled by continuing the discretised skew-symmetric diffusion from the last observation-period value, with a rescaling of (γ, σ) intended to reduce first-order discretisation bias. The rescaling is described only in the supplement, and no sensitivity analysis is reported in the main text. Because extrapolated mean survival is the primary output of the method, the reader cannot determine whether the rescaling is a principled correction or an ad hoc adjustment; please present the derivation and report extrapolations with and without the rescaling for the two data examples.","section":"Section 3.6"},{"comment":"The two case studies are informative but do not validate the extrapolation claim, because no simulation experiment with a known hazard is reported. Comparisons against M-spline and parametric models in Tables 1-3 place the method in context, but they cannot show that the model and sampler recover calibrated posterior statements for extrapolation. A simulation study with data generated from the model and from misspecified hazards, reporting coverage of the true mean survival over (0, y∞), would substantially strengthen the paper.","section":"Section 4"}],"minor_comments":[{"comment":"The relation α_j := ˇα_{jσ²} uses the random step size σ as a time index; please define explicitly how σ, the knots s_j, and the diffusion time scale are linked, since this mapping is used later in the extrapolation procedure.","section":"Section 2.2"},{"comment":"The reparameterisation γ = ωΓ is clear, but it would help to state the implied joint prior on the candidate, active, and inactive knot sets as an explicit equation, since the thinning construction is central to the Gibbs update.","section":"Section 3.4"},{"comment":"The statement that the leave-one-out criterion 'does not always sufficiently penalise overly complex models' is non-standard for a predictive criterion; please report the PSIS diagnostics (e.g., k-hat values) and the full comparison underlying this conclusion.","section":"Section 4.1"},{"comment":"Please state the time horizon y∞ used for the colon cancer and CLL-8 analyses; the M-spline comparisons depend on the final knot locations (5, 10, 15), but the corresponding horizon is not given in the main text.","section":"Section 4"},{"comment":"The estimate of E[Y_t] − E[Y_c] and its credible interval should be accompanied by a note on whether the two arms are analysed jointly or independently and how the posterior difference is computed.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"Several load-bearing implementation details, including the extrapolation rescaling, the full sampler algorithms, and prior derivations, are deferred to a supplement that was not available with this submission. I recommend that the editor require the supplement for review, or that the authors move the key derivations into the main text. I also note the absence of a data and code availability statement; for a paper whose contribution is partly computational, one should be added. The concerns above are addressable, so I do not see a need for rejection, but the approximation-error issue should be resolved before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, the paper is worth taking seriously. It combines a discretised diffusion prior for the log-hazard with a Poisson process prior for knot locations, so the drift of the diffusion can encode expert opinion about long-term hazard behaviour. That combination is new, and it is aimed at a real problem: NICE-type appraisals with heavily censored data need extrapolated mean survival, and existing spline or parametric methods either leave the final knot arbitrary or impose a parametric form. The paper also extends sticky PDMP samplers from spike-and-slab posteriors to a broader class of transdimensional posteriors, using a finite candidate set plus a Gibbs refresh of inactive knots. That extension is described clearly, and Proposition 3.1 gives a clean, correct comparison of a mixing-time proxy under the two discretisation schemes. The real-data examples show the intended behaviour: observation-period inference is data-dominated, extrapolation is prior-dominated, and the differences between drifts are sensible.\n\nSoft spots are real but not fatal. The splitting scheme in Section 3.2 is stated to introduce only a small approximation error, with no bound or diagnostic. The paper says correction was unnecessary but does not show it. For a methods paper this is a legitimate referee request: compare against the exact line-search scheme for the drifts where it exists, or provide a diagnostic. The extrapolation step in Section 3.6 also relies on a rescaling of (gamma, sigma) that is only described in the supplement; that is acceptable if the supplement is complete, but the main text should at least state the order of the correction. More importantly, there is no simulation study with a known hazard, so calibration of the whole pipeline is unverified. That is the biggest gap. The gamma selection via LOOCV plus a plausibility check is honest but ad hoc; the authors say so themselves.\n\nNone of this shakes the central construction. The paper is not hiding anything: it states that extrapolation is prior-dominated. The citation pattern is appropriate. The sampler design is thoughtful and the empirical work is useful, even if not decisive. This paper deserves a serious referee; the referee should ask for a simulation study and for evidence on the splitting bias, not for a rewrite of the core idea. I'd bring it to a reading group and would probably cite it if I worked in survival extrapolation.","headline":"Genuinely new combination of diffusion priors and random knots for survival extrapolation, with a clever PDMP sampler; needs a simulation study and an honest check of the splitting-scheme bias before I'd trust the numbers.","tokens_in":21997,"tokens_out":2316,"would_cite":true,"duration_ms":23588,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62N01","65C05"],"pacs":[],"model":"deepseek-v4-flash","headline":"The diffusion piecewise exponential model couples a discretised diffusion prior for the log-hazard with a Poisson-process prior for knots, so observed survival data and expert prior information jointly determine long-term extrapolations.","keywords":["survival analysis","piecewise exponential model","diffusion prior","extrapolation","Poisson process prior","piecewise deterministic Monte Carlo","sticky PDMP","health technology assessment"],"falsifier":"Fit the same survival model twice on the colon cancer data, once with the paper's approximate event-time generation and once with an exact event-time method known to be valid for this model, and compare the posterior hazard curves. A difference larger than Monte Carlo error would show the approximation error is material.","tokens_in":20799,"feed_emoji":"📈","tokens_out":10648,"duration_ms":101381,"temperature":0.7,"pith_summary":"The paper proposes a Bayesian model for survival data in which the log-hazard is piecewise constant, with local hazards following a discretised stochastic differential equation whose drift encodes prior belief about long-term mortality, and with knot locations drawn from a Poisson process. The goal is to make extrapolation beyond the end of a trial principled: during the observation period the data dominate, while in the extrapolation period the drift and the knot intensity control how quickly prior information takes over. This matters for health technology assessment, where mean survival beyond the trial horizon is a key input to cost-effectiveness decisions and data are often heavily administratively censored. The paper also contributes a sampling algorithm that extends sticky piecewise deterministic Monte Carlo dynamics to models with a random number of knots, so the model can be fitted without tuning between-model proposals.","feed_headline":"Mean survival beyond trials now blends data with expert prior","feed_subtitle":"Piecewise exponential hazards with Poisson knots now extrapolate long-term survival with a tuning-free sampler.","key_machinery":"The key object is the drift function $\\mu$ of the diffusion prior together with the Poisson-process intensity $\\gamma$ for the knots. The drift can encode a stationary distribution for the hazard, an underlying Gompertz hazard, or a time-varying target such as a waning treatment effect; the discretisation uses a skew-symmetric innovation scheme that remains stable for non-Lipschitz drifts. The Poisson knot process acts as a random time change between the diffusion skeleton and calendar time, so the number of knots controls both flexibility and the speed at which the prior dominates in the extrapolation period. Transdimensional sampling is achieved by sticky piecewise deterministic dynamics: each candidate knot is a switchable variable, movement onto and off the 'on' state happens at hyperplane crossings without likelihood evaluations, and inactive knots are refreshed from the prior in a Gibbs step. A proposition shows the skew-symmetric parameterisation gives a smaller expected recurrence time to the null model than the Euler-Maruyama parameterisation, supporting faster mixing.","core_discovery":"The central claim is that prior beliefs about the long-term hazard can be encoded as the drift of a diffusion and combined with the observed-data likelihood through a piecewise exponential model, without forcing a parametric hazard shape. The log-hazard values are a discretisation of $\\mathrm{d}\\alpha = \\mu(\\alpha)\\,\\mathrm{d}y + \\mathrm{d}W_y$, and the knots form a Poisson process that acts as a random time change; this makes the prior adaptive to volatility in the data and determines how fast the drift dominates beyond the end of observation. The authors show that the resulting posterior, which lives over models with different numbers of knots, can be sampled with sticky piecewise deterministic Monte Carlo by treating each candidate knot as a switchable variable, then refreshing inactive candidate knots from the prior in a Gibbs step. On colon cancer and leukaemia trial data, they demonstrate that observation-period inference is nearly unaffected by the drift, while extrapolated mean survival and its uncertainty respond strongly to the chosen drift, and that the sampler explores the posterior more fully than a reversible jump comparator.","pith_inferences":["The authors do not explore it, but the same diffusion-prior idea could be applied to the cumulative hazard or survival function directly, letting blended-survival information enter as a time-varying drift rather than as synthetic data.","A quantitative consequence that is not tested in the paper: the rate at which $\\gamma$ lets the prior overwhelm the data in the extrapolation period should be measurable in prior simulations, and could be used to calibrate $\\gamma$ to a desired 'memory' of the data.","A practical sensitivity check implied by the construction is to increase the discretisation step $\\sigma$ and re-run the colon cancer analysis; if extrapolated mean survival shifts substantially, the assumedly small discretisation bias is not small.","For health technology assessment reporting, the model suggests that extrapolated mean survival should be published together with the drift and knot intensity, since those encode the untestable assumptions that drive the decision-relevant estimate."],"forward_implications":["Analysts can encode beliefs about long-term mortality—stationary hazard levels, an external Gompertz curve, or a waning treatment effect—as a drift and obtain extrapolated mean survival with credible intervals that reflect both the data and that belief.","Because the drift and the knot intensity are separated, information criteria can pick the knot intensity without dictating the long-term hazard, breaking the usual trade-off between fit in the observation period and plausibility of the tail.","The sticky piecewise deterministic construction extends to any transdimensional posterior with a centring hyperplane where the two models' likelihoods agree, so similar transdimensional problems need no reversible-jump proposals.","Even with administrative censoring rates near 90 percent, posterior inference remains computationally feasible and the prior becomes influential exactly as the data thin out.","Reversible-jump comparisons in the paper suggest the sampler explores the posterior tails more completely for the same computational budget."],"supporting_citations":[{"why":"Supplies the sticky PDMP dynamics for spike-and-slab posteriors that the paper extends to general transdimensional posteriors.","marker":"Bierkens et al. (2023)"},{"why":"Gives the reversible jump PDMP formulation whose hyperplane-crossing construction is reused for knot activation and deactivation.","marker":"Chevallier et al. (2023)"},{"why":"Provides the forward event chain Monte Carlo scheme that removes the need to tune a refreshment rate in the bouncy particle sampler.","marker":"Michel et al. (2020)"},{"why":"Supplies the splitting scheme used to generate PDMP event times, which is the source of the claimed small approximation error.","marker":"Bertazzi et al. (2023)"},{"why":"Introduces the skew-symmetric discretisation of stochastic differential equations with non-Lipschitz drift, used for the diffusion prior and shown to speed mixing.","marker":"Livingstone et al. (2024)"},{"why":"Represents the standard piecewise exponential extrapolation approach with a constant final hazard that the new model is designed to supersede.","marker":"Cooney and White (2023)"},{"why":"Provides the flexible spline-based extrapolation comparator whose sensitivity to final-knot placement motivates the Poisson-process knot prior.","marker":"Jackson (2023)"},{"why":"Earlier latent diffusion model for survival that the drift-based prior construction generalises.","marker":"Roberts and Sangalli (2010)"}],"fun_headline_variants":["Diffusion priors and Poisson knots sharpen survival extrapolation","Expert priors meet trial data for long-term survival estimates","PDMC sampling enables flexible survival extrapolation with priors","Beyond trial data: a diffusion model for realistic hazard extrapolation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the approximate method used to generate the sampler's event times biases the posterior only negligibly; the paper gives no bound or diagnostic for that error.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion priors and Poisson knots sharpen survival extrapolation","Expert priors meet trial data for long-term survival estimates","PDMC sampling enables flexible survival extrapolation with priors","Beyond trial data: a diffusion model for realistic hazard extrapolation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000279,"raw_usage":{"total_tokens":1645,"prompt_tokens":921,"completion_tokens":724,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":656}},"tokens_in":537,"tokens_out":724,"duration_ms":7527,"temperature":1.0,"reasoning_tokens":656,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:53:03.463954+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fit the same survival model twice on the colon cancer data, once with the paper's approximate event-time generation and once with an exact event-time method known to be valid for this model, and compare the posterior hazard curves. A difference larger than Monte Carlo error would show the approximation error is material.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the sticky PDMP dynamics for spike-and-slab posteriors that the paper extends to general transdimensional posteriors."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the reversible jump PDMP formulation whose hyperplane-crossing construction is reused for knot activation and deactivation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the forward event chain Monte Carlo scheme that removes the need to tune a refreshment rate in the bouncy particle sampler."},{"cited_title":"and White, A","cited_arxiv_id":null,"evidence_quote":"Represents the standard piecewise exponential extrapolation approach with a constant final hazard that the new model is designed to supersede."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the flexible spline-based extrapolation comparator whose sensitivity to final-knot placement motivates the Poisson-process knot prior."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier latent diffusion model for survival that the drift-based prior construction generalises."}],"review_version":1}