REVIEW 2 major objections 5 minor 1 cited by
Assessing treatment efficacy for interval-censored endpoints using multistate semi-Markov models fit to multiple data streams
T0 review · 2 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper claims that a Monte Carlo expectation-maximization algorithm can fit multistate semi-Markov models to intermittently observed, multi-stream trial data, and uses it to estimate that REGEN-COV prophylaxis reduced SARS-CoV-2…
desk verdict The MCEM-with-importance-sampling machinery is a real, well-validated methodological contribution; the REGEN-2069 clinical numbers are plausible but rest on a measurement-defined infection endpoint that the paper never stress-tests. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is a five-state progressive multistate model (naive, PCR-positive asymptomatic, PCR-negative post-infection, symptomatic PCR-positive, symptomatic PCR-negative) combined with an MCEM estimation algorithm. In each E-step, complete histories are proposed from a time-homogeneous Markov surrogate conditioned on the observed data, using forward-filtering backwards-sampling to impute unknown states at observation times and uniformization to fill in endpoint-conditioned paths between visits; self-normalized importance weights correct for the discrepancy between the Markov proposal and the semi-Markov target. An ascent-based rule augments the Monte Carlo sample until changes in the Q-function are distinguishable from Monte Carlo error, and spline or Weibull baseline intensities allow semi-parametric inference. This construction sidesteps the intractable transition probabilities that make semi-Markov models difficult to fit to panel data, and a marginal-likelihood estimator built from the same importance weights enables model comparison via AIC.
What would settle it
Simulate a trial from the paper's nine-state simulation model but add a fraction of infected participants who clear virus before the first PCR assessment and never seroconvert, then fit the five-state model and check whether the estimated protective efficacy against asymptomatic infection is biased relative to the simulated truth by more than the simulation's Monte Carlo error.
Extended reading notes
Core claim
The paper's central claim is that intractable semi-Markov likelihoods for intermittently observed multistate processes can be maximized by an MCEM algorithm that samples complete sample paths from a data-conditioned Markov surrogate and reweights them by importance sampling. In the REGEN-2069 reanalysis, this yields estimates that REGEN-COV reduced all-comer infection by 60.4% (95% CI 44.9-72.5%), symptomatic infection by 83.6% (95% CI 69.4-93.1%), and asymptomatic infection by 38.7% (95% CI 10.0-60.8%), while reducing the mean duration of PCR-detectable shedding from 13.0 days (95% CI 11.5-14.6) to 6.2 days (95% CI 5.0-7.8). Simulations show that crude event tabulation and Markov model fits are biased for these functionals, whereas the semi-Markov fits achieve negligible bias and near-nominal coverage, supporting the authors' conclusion that the method recovers difficult-to-measure endpoints implicated by asymptomatic infection.
Load-bearing premise
The argument assumes that being infected is the same as being measurably affected, specifically shedding virus detectable by nasopharyngeal RT-qPCR, so a person who is infected but never tests positive, never shows symptoms, and never seroconverts is classified as uninfected.
Editorial extensions
If this is right
- If the central claim is correct, interval-censored multi-stream trial data no longer require Markov assumptions: semi-Markov models with Weibull or spline intensities give approximately unbiased estimates of transition intensities and their functionals with nominal coverage.
- For REGEN-2069 specifically, the estimates imply that REGEN-COV's protection is broad, cutting infection risk by 60.4% overall and by 83.6% for symptomatic infection, while shortening the period of detectable shedding from 13.0 to 6.2 days, a change with direct implications for transmission.
- The 38.7% estimate of protective efficacy against asymptomatic infection is positive but smaller than the symptomatic efficacy, and the model attributes part of this apparent benefit to the fact that antibody treatment shortens shedding below the detection threshold of weekly PCR sampling.
- Because seroconversion was less frequent after asymptomatic infections on the mAb arm, the data support a mechanism in which antibody treatment suppresses viral load, symptoms, and immune exposure together.
- The framework extends beyond COVID-19 to any progressive disease process whose state is partially observed through several measurement modalities, with AIC-based model choice made feasible by the marginal-likelihood estimator.
Reading between the lines
- Editorial inference: the definition of infection as detectable RT-qPCR positivity is the load-bearing assumption, and if some infected participants never shed detectable virus or seroconvert, the estimated asymptomatic-infection protective efficacy would shift; a sensitivity analysis that redefines infection or adds a fraction of undetectable infections would quantify this.
- Editorial inference: the same importance-sampling scheme could be extended to disease-driven observation processes, where sicker patients are tested more often, by modifying the proposal to condition on the observation times, an extension the paper notes but does not implement.
- Editorial inference: because the algorithm provides a Monte Carlo estimate of the marginal likelihood, it could be combined with penalized-likelihood or weighted-bootstrap Bayesian procedures for model selection and multiple-testing corrections in other panel-data trials.
- Editorial inference: the comparison with phase-type models suggests a testable roadmap: using a phase-type or discrete mixture of Markov processes as the proposal distribution could improve robustness in settings where a simple Markov surrogate yields a small effective sample size.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a Monte Carlo expectation-maximization (MCEM) framework with data-conditioned Markov proposals for fitting multistate semi-Markov models when the data are interval-censored or observed through multiple, possibly error-prone, data streams. The methods are applied to the REGEN-2069 trial of REGEN-COV prophylaxis, using symptom onset, weekly RT-qPCR, and end-of-EAP serology to estimate protective efficacy against infection (including asymptomatic infection), effects on seroconversion, and the duration of viral shedding. Simulation studies compare the proposed approach with crude tabulation, Markov models, and phase-type approximations, and report good performance for the main efficacy parameters, with some under-coverage for restricted mean infection time and PCR-positivity duration. The application estimates PE against infection of 60.4%, PE against asymptomatic infection of 38.7%, and a reduction in mean shedding duration from 13.0 to 6.2 days.
Significance. If the results are sound, the paper makes a useful methodological contribution: a general, computationally feasible inference framework for semi-Markov multistate models under complex coarsening, with importance-sampling recycling inside MCEM and semi-parametric baseline intensities. The manuscript is also notable for shipping reproducible code, for a careful simulation design that uses a richer 9-state generative model to validate a reduced 5-state inferential model, and for a direct comparison against rejection-based sampling and phase-type approximations. The application addresses a clinically important question, and the analysis illustrates that naive tabulation of panel data can be badly biased, which is a valuable message for practitioners.
major comments (2)
- [Section 2.1 / Table 1b / Section C.1] The application-level estimand is defined by detectability: infection is equated with 'being measurably affected, as evidenced by viral shedding detectable by nasopharyngeal RT-qPCR' (Section 2.1), and Table 1b assumption 1 treats participants who never become symptomatic, PCR+, or seropositive as uninfected. The simulation validation of the REGEN-2069 design does not stress this assumption because, as stated in Section C.1, 'All infected participants are assumed to have detectable virus.' Since the fitted model itself implies that mAb shortens shedding and reduces seroconversion, any infected participants who never shed detectable virus and never seroconvert would be counted as uninfected, and the missingness would plausibly be differential by arm. The reported PE against asymptomatic infection (38.7%; 95% CI 10.0-60.8) and the mAb-arm shedding duration would then mix a true biological effect with a detection effect. The paper should either consistently label the endpoint as 'detectable infection' in the abstract and Discussion, or provide a sensitivity analysis that relaxes assumption 1 (e.g., a latent infection state with assumed detection probabilities).
- [Section 3.2 / Table 2(c) / Section 4.3] The simulation study reports coverage of 0.85-0.87 for the restricted mean time to infection on the placebo arm and 0.90 for the restricted mean duration of PCR positivity on the placebo arm under the semi-parametric five-state model (Table 2(c); Tables S9-S10), with the paper attributing the under-coverage to infections detected only by serology. These are precisely the functionals used for the headline clinical estimates in Section 4.3: the placebo-arm shedding duration of 13.0 days (95% CI 11.5-14.6) and the comparison with 6.2 days on the mAb arm. The paper should either provide calibration of the bootstrap intervals for these functionals under the actual REGEN-2069 observation design, or explicitly qualify the reported confidence intervals as potentially anti-conservative for the duration endpoint.
minor comments (5)
- [Abstract] The sentence 'Our algorithm provide substantial computational improvements' has a subject-verb agreement error; it should read 'Our algorithm provides.'
- [Author affiliations] The affiliation 'Regeneron Pharmaceuticles' appears to contain a typo and should likely read 'Regeneron Pharmaceuticals.'
- [Introduction] The word 'identifiabile' should be 'identifiable'.
- [Table S6] In the table note, 'folloowup' should be 'follow-up'.
- [References] The reference to Akaike (1998) is spelled 'Aikaike' in the reference list and should be corrected.
Circularity Check
No significant circularity: the MCEM derivation and simulation benchmarks are self-contained, and the infection definition is an explicit operational choice rather than a masked equivalence.
full rationale
The paper's claimed derivations are internally non-circular. The MCEM estimator is a direct likelihood method: the Q-function (Eq. 7) is the expected complete-data log-likelihood, and the importance-sampling weights (Eq. 9; Appendix A.2) are the standard ratio of target to proposal densities, with the data-conditional indicators canceling because paths are sampled conditionally on the data. The proposal Markov model is fit separately and is used only as a sampling mechanism, not as the inferential target. The two simulation studies provide independent checks: data are generated from a 9-state semi-Markov model (Appendix C.1, Tables S4-S6) or a Weibull illness-death model, and inference uses reduced 5-state models, so the fitted models are not the same objects that generated the truth. The REGEN-2069 estimates (PE 60.4%, PE-asymptomatic 38.7%, shedding durations 13.0 vs 6.2 days) are functionals of the fitted intensity estimates, not parameters tuned to hit observed case counts. The paper is explicit that 'infection' is operationally defined as transition to PCR+ (Section 2.1; Table 1b assumption 1), and Appendix C.1 states 'All infected participants are assumed to have detectable virus' in the simulations; this is a transparent identifiability and definitional assumption, not a hidden equivalence, and it limits external interpretation but does not make the estimation circular. The Discussion also cautions that the MCEM algorithm is not a panacea for non-identifiability. Self-citations (O'Brien 2021; Follmann 2022) provide the primary trial result and a prior seroconversion hypothesis; neither is invoked as a uniqueness theorem or to forbid alternatives. No equation-level reduction of a claimed prediction to a fitted input was found.
Assumptions & free parameters
free parameters (4)
- Baseline transition intensity parameters for the four transitions (exponential rates, Weibull shape and scale…
- mAb covariate effects (log hazard ratios) on each transition
- B-spline knot locations and degrees =
Knots at 3.5, 10.5, 17.5 days for 1->2 and 2->4 transitions; one knot at 7 days for 2->3 and 4->5; degree 1
- MCEM tuning parameters (initial effective sample size, tolerance epsilon, alpha, gamma, kappa) =
Initial ESS 10 to 100; tolerance and probability bounds described in Appendix A.4.3
assumptions (7)
- domain assumption Participants who never become symptomatic, PCR+, or seropositive are uninfected.
- domain assumption PCR positivity precedes symptom onset, symptom onset precedes seroconversion, no transitions from PCR- to PCR+, and no sero-reversion on study.
- domain assumption Infection is defined as the transition to detectable PCR+.
- domain assumption Semi-Markov transition intensities depend on time since state entry and covariates, not on the full history.
- domain assumption Observation times and missingness are independent of the disease process.
- domain assumption The Markov surrogate proposal has finite importance weights and adequate effective sample size for the MCEM E-step.
- domain assumption Observed states are not misclassified: PCR, symptom, and serology data are accurate but possibly incomplete.
Cite this review
Pith. "Pith review of Assessing treatment efficacy for interval-censored endpoints using multistate semi-Markov models fit to multiple data streams." pith.science (2026). https://pith.science/paper/7QB57U2Z
@misc{pith2026250114097,
author = {Pith},
title = {Pith review of: Assessing treatment efficacy for interval-censored endpoints using multistate semi-Markov models fit to multiple data streams},
year = {2026},
howpublished = {\url{https://pith.science/paper/7QB57U2Z}},
note = {Machine review of arXiv:2501.14097}
}
read the original abstract
We introduce a computationally efficient and general approach for utilizing multiple, possibly interval-censored, data streams to study complex biomedical endpoints using multistate semi-Markov models. Our motivating application is the REGEN-2069 trial, which investigated the protective efficacy (PE) of the monoclonal antibody combination REGEN-COV against SARS-CoV-2 when administered prophylactically to individuals in households at high risk of secondary transmission. Using data on symptom onset, episodic RT-qPCR sampling, and serological testing, we estimate the PE of REGEN-COV for asymptomatic infection, its effect on seroconversion following infection, and the duration of viral shedding. We find that REGEN-COV reduced the risk of asymptomatic infection and the duration of viral shedding, and led to lower rates of seroconversion among asymptomatically infected participants. Our algorithm for fitting semi-Markov models to interval-censored data employs a Monte Carlo expectation maximization (MCEM) algorithm combined with importance sampling to efficiently address the intractability of the marginal likelihood when data are intermittently observed. Our algorithm provide substantial computational improvements over existing methods and allows us to fit semi-parametric models despite complex coarsening of the data.
Figures
Forward citations
Cited by 1 Pith paper
-
Markov Renewal Proportional Hazards is All You Need
A mostly tutorial and application paper claims semi-Markov models with the DSH estimator produce smoother transition probability curves than Aalen-Johansen in the EBMT stem cell transplant data, without quantitative v...
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION article output.bibitem format.authors "author" output.check author format.key output output.year.check new.block format.title "title" output.check new.block crossref missing format.jour.vol output format.article.crossref output.nonnull format.pages output if new.block note output fin.entry FUNCTION b...
-
[2]
Akaike , H. (1998). Information theory and an extension of the maximum likelihood principle. In: Selected papers of Hirotugu Akaike\/ . Springer, pp.\ 199--213
work page 1998
-
[3]
Andersen , P.K. and Keiding , N. (2002). Multi-state models for event history analysis. Statistical Methods in Medical Research\/ 11, 91--115
work page 2002
-
[4]
Aralis , H. and Brookmeyer , R. (2019). A stochastic estimation procedure for intermittently-observed semi- M arkov multistate models with back transitions. Statistical Methods in Medical Research\/ 28, 770--787
work page 2019
-
[5]
Barone , R. and Tancredi , A. (2022). Bayesian inference for discretely observed continuous time multi-state models. Statistics in Medicine\/ 41, 3789--3803
work page 2022
-
[6]
Bladt , M. and S rensen , M. (2005). Statistical inference for discretely observed M arkov jump processes. Journal of the Royal Statistical Society Series B: Statistical Methodology\/ 67, 395--410
work page 2005
-
[7]
Caffo , B.S., Jank , W. and Jones , G.L. (2005). Ascent-based monte carlo expectation--maximization. Journal of the Royal Statistical Society Series B\/ 67, 235--251
work page 2005
-
[8]
Cheung , L.C., Albert , P.S., Das , S. and Cook , R.J. (2022). Multistate models for the natural history of cancer progression. British Journal of Cancer\/ 127, 1279--1288
work page 2022
Show all 40 references
-
[9]
and Lawless , J.F
Cook , R.J. and Lawless , J.F. (2018). Multistate Models for the Analysis of Life History Data\/ . Chapman and Hall/CRC Press, Boca Raton
2018
-
[10]
and Lawless , J.F
Cook , R.J. and Lawless , J.F. (2021). Independence conditions and the analysis of life history studies with intermittent observation. Biostatistics\/ 22, 455--481
2021
-
[11]
and Rubin , D.B
Dempster , A.P., Laird , N.M. and Rubin , D.B. (1977). Maximum likelihood from incomplete data via the em algorithm. Journal of the Royal Statistical Society: Series B\/ 39, 1--22
1977
-
[12]
Dorai-Raj , S. (2022). binom: Binomial Confidence Intervals for Several Parameterizations\/ . R package version 1.1-1.1
2022
-
[13]
Elvira, V \' ctor, Martino, Luca, Luengo, David and Bugallo, M \'o nica F . (2019). Generalized multiple importance sampling. Statistical Science\/ 34, 129 -- 155
2019
-
[14]
and others
Follmann , D.A., Janes , H.E., Buhule , O.D., Zhou , H., Girard , B., Marks , K., Kotloff , K., Desjardins , M., Corey , L., Neuzil , K.M. and others . (2022). Antinucleocapsid antibodies after SARS-CoV-2 infection in the blinded phase of the randomized, placebo-controlled mRN...
2022
-
[15]
Food and Drug Administration . (2017). Multiple endpoints in clinical trials guidance for industry. Technical Report , Center for Biologics Evaluation and Research (CBER)
2017
-
[16]
Gong , R. (2022). Exact inference with approximate computation for differentially private data via perturbations. Journal of Privacy and Confidentiality\/ 12, 1--26
2022
-
[17]
and Stone , E.A
Hobolth , A. and Stone , E.A. (2009). Simulation from endpoint-conditioned, continuous-time M arkov chains on a finite state space, with applications to molecular evolution. The Annals of Applied Statistics\/ 3, 1204--1231
2009
-
[18]
and Lawless , J.F
Kalbfleisch , J.D. and Lawless , J.F. (1985). The analysis of panel data under a M arkov assumption. Journal of the American Statistical Association\/ 80, 863--871
1985
-
[19]
and others
Lange , J.M., Gulati , R., Leonardson , A.S., Lin , D.W., Newcomb , L.F., Trock , B.J., Carter , B.H., Cooperberg , M.R., Cowan , J.E., Klotz , L.H. and others. (2018). Estimating and comparing cancer progression risks under varying surveillance protocols. The Annals of Applie...
2018
-
[20]
and Minin , V.N
Lange , J.M., Hubbard , R.A., Inoue , L.Y.T. and Minin , V.N. (2015). A joint model for multistate disease processes and random informative observation times, with applications to electronic medical records data. Biometrics\/ 71, 90--101
2015
-
[21]
and Casella , G
Levine , R.A. and Casella , G. (2001). Implementations of the monte carlo em algorithm. Journal of Computational and Graphical Statistics\/ 10, 422--439
2001
-
[22]
Liu, Jun S and Liu, Jun S . (2001). Monte Carlo strategies in scientific computing\/ , Volume 75. Springer
2001
-
[23]
Louis , T.A. (1982). Finding the observed information matrix when using the em algorithm. Journal of the Royal Statistical Society Series B: Statistical Methodology\/ 44, 226--233
1982
-
[24]
and Marra , G
Machado , R.J.M., van den Hout , A. and Marra , G. (2021). Penalised maximum likelihood estimation in multi-state models for interval-censored data. Computational Statistics & Data Analysis\/ 153, 107057
2021
-
[25]
Mandel , M. (2013). Simulation-based confidence intervals for functions with complicated derivatives. The American Statistician\/ 67, 76--81
2013
-
[26]
and Snelling , T.L
McLeod , C., Norman , R., Litton , E., Saville , B.R., Webb , S. and Snelling , T.L. (2019). Choosing primary endpoints for clinical trials of health care interventions. Contemporary Clinical Trials Communications\/ 16, 100486
2019
-
[27]
and Xu , J
Newton , M.A., Polson , N.G. and Xu , J. (2021). Weighted B ayesian bootstrap for scalable posterior distributions. Canadian Journal of Statistics\/ 49, 421--437
2021
-
[28]
and others
O’Brien , M.P., Forleo-Neto , E., Musser , B.J., Isa , F., Chan , K., Sarkar , N., Bar , K.J., Barnabas , R.V., Barouch , D.H., Cohen , M.S. and others . (2021). Subcutaneous regen-cov antibody combination to prevent covid-19. New England Journal of Medicine\/ 385, 1184--1195
2021
-
[29]
and Sk \"o ld , M
Papaspiliopoulos , O., Roberts , G.O. and Sk \"o ld , M. (2007). A general framework for the parametrization of hierarchical models. Statistical Science\/ 22, 59--73
2007
-
[30]
Polanco , J.I. (2024). B S pline K it.jl: a collection of B -spline tools in J ulia. https://github.com/jipolanco/BSplineKit.jl
2024
-
[31]
and Eckerle , I
Puhach , O., Meyer , B. and Eckerle , I. (2023). Sars-cov-2 viral load and shedding kinetics. Nature Reviews Microbiology\/ 21, 147--161
2023
-
[32]
and Papamarkou , T
Revels , J., Lubin , M. and Papamarkou , T. (2016). Forward-mode automatic differentiation in J ulia. arXiv:1607.07892\/
2016 arXiv
-
[33]
Rubin , D.B. (1981). The bayesian bootstrap. The Annals of Statistics\/ 9, 130--134
1981
-
[34]
Scott , S.L. (2002). Bayesian methods for hidden markov models: Recursive computing in the 21st century. Journal of the American statistical Association\/ 97, 337--351
2002
-
[35]
Signorell , A. (2023). DescTools: Tools for Descriptive Statistics\/ . R package version 0.99.52
2023
-
[36]
Titman , A.C. (2011). Flexible nonhomogeneous M arkov models for panel observed data. Biometrics\/ 67, 780--787
2011
-
[37]
and Sharples , L.D
Titman , A.C. and Sharples , L.D. (2010). Semi- M arkov models with phase-type sojourn distributions. Biometrics\/ 66, 742--752
2010
-
[38]
Vehtari, Aki, Simpson, Daniel, Gelman, Andrew, Yao, Yuling and Gabry, Jonah . (2015). Pareto smoothed importance sampling. arXiv preprint arXiv:1507.02646\/
2015 arXiv
-
[39]
and Tanner , M.A
Wei , G.C.G. and Tanner , M.A. (1990). A monte carlo implementation of the em algorithm and the poor man's data augmentation algorithms. Journal of the American statistical Association\/ 85, 699--704
1990
-
[40]
Wilkinson , D.J. (2018). Stochastic Modelling for Systems Biology\/ . CRC Press, Boca Raton
2018
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.