{"id":"597026c1-b8cd-43d4-a7a8-3e034b4be2ff","arxiv_id":"2501.14097","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A Monte Carlo EM algorithm with importance sampling enables fitting semi-Markov multistate models to multiple, interval-censored data streams, yielding estimates of REGEN-COV's protective efficacy against asymptomatic infection and shorter viral shedding.","lead":"This paper introduces a computational method for fitting semi-Markov multistate models to interval-censored clinical data and uses it to reanalyze the REGEN-COV COVID-19 prophylaxis trial. The analysis estimates that REGEN-COV reduced asymptomatic infection risk, symptom development after infection, viral shedding duration, and seroconversion in asymptomatically infected participants.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Clinical estimates rely on equating infection with detectable RT-qPCR positivity; if some infections never shed or seroconvert, reported PE and shedding duration are not identified, and no sensitivity analysis is provided.","rationale":"The strongest claim has two parts: the MCEM method and the REGEN-2069 clinical findings. The method part is well supported: simulations show low bias and good coverage for the target functionals under the assumed generative model, code is available, and the comparisons with rejection and phase-type samplers are informative. The clinical part is less secure. The single assumption on which the entire application rests is Table 1b assumption 1 and the Section 2.1 operational definition of infection as detectable viral shedding. If that equivalence fails in even a modest fraction of participants, the estimated infection counts are lower than truth, and because the treatment plausibly reduces viral detectability, the between-arm comparison is differentially affected. The paper's own sensitivity to this is zero: the 9-state simulation model explicitly assumes all infected participants have detectable virus, so the favorable operating characteristics in Table 2 cannot speak to this violation. A simple simulation perturbation with a hidden never-detected infection state would settle whether the concern is material. Model-selection uncertainty is a real but secondary issue; the primary barrier to taking the clinical point estimates at face value is the untestable infection definition. Since the reader already assigned CONDITIONAL for essentially this reason, my read does not change the verdict.","tokens_in":33633,"tokens_out":7694,"duration_ms":76627,"concrete_test":"Using the authors' simulation code, generate REGEN-2069-like trials from the 9-state generative model in Appendix C.1, but add a tunable fraction q of true infections that never become PCR+ and never seroconvert (e.g., q = 0.05, 0.10, 0.20, with q possibly larger on the mAb arm to mimic reduced viral load). Fit the paper's five-state model to the resulting data and compare estimated PE for asymptomatic infection and restricted mean PCR+ duration with the true values. If either estimate changes by more than roughly the width of the reported 95% CI, the equivalence assumption is load-bearing and the REGEN-2069 point estimates should be reported as conditional on it.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central application-level claim—that REGEN-COV reduces asymptomatic infection and shortens shedding—rests on Table 1b assumption 1 and the Section 2.1 definition of infection as 'the occurrence of a participant being measurably affected, as evidenced by viral shedding detectable by nasopharyngeal RT-qPCR.' By construction, a true infection that produces no symptoms, no detectable shedding, and no seroconversion is counted as 'never infected.' This is not a harmless convention: the analysis itself estimates that mAb shortens shedding and reduces seroconversion, so the missingness mechanism is plausibly differential by arm. The reported PE for asymptomatic infection (38.7%) and the mAb-arm shedding duration (6.2 days) could therefore reflect differential detectability rather than true biological effects. The simulation validation in Section 3.2 does not address this: Section C.1 states 'All infected participants are assumed to have detectable virus,' so the simulations generate exactly the assumption whose violation is at issue. No sensitivity analysis is reported for the real data. Because the estimand is literally defined by the measurement, the headline clinical numbers are conditional on an untestable equivalence that the paper does not attempt to relax.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a Monte Carlo expectation-maximization (MCEM) framework with data-conditioned Markov proposals for fitting multistate semi-Markov models when the data are interval-censored or observed through multiple, possibly error-prone, data streams. The methods are applied to the REGEN-2069 trial of REGEN-COV prophylaxis, using symptom onset, weekly RT-qPCR, and end-of-EAP serology to estimate protective efficacy against infection (including asymptomatic infection), effects on seroconversion, and the duration of viral shedding. Simulation studies compare the proposed approach with crude tabulation, Markov models, and phase-type approximations, and report good performance for the main efficacy parameters, with some under-coverage for restricted mean infection time and PCR-positivity duration. The application estimates PE against infection of 60.4%, PE against asymptomatic infection of 38.7%, and a reduction in mean shedding duration from 13.0 to 6.2 days.","tokens_in":33843,"tokens_out":5613,"duration_ms":53321,"significance":"If the results are sound, the paper makes a useful methodological contribution: a general, computationally feasible inference framework for semi-Markov multistate models under complex coarsening, with importance-sampling recycling inside MCEM and semi-parametric baseline intensities. The manuscript is also notable for shipping reproducible code, for a careful simulation design that uses a richer 9-state generative model to validate a reduced 5-state inferential model, and for a direct comparison against rejection-based sampling and phase-type approximations. The application addresses a clinically important question, and the analysis illustrates that naive tabulation of panel data can be badly biased, which is a valuable message for practitioners.","major_comments":[{"comment":"The application-level estimand is defined by detectability: infection is equated with 'being measurably affected, as evidenced by viral shedding detectable by nasopharyngeal RT-qPCR' (Section 2.1), and Table 1b assumption 1 treats participants who never become symptomatic, PCR+, or seropositive as uninfected. The simulation validation of the REGEN-2069 design does not stress this assumption because, as stated in Section C.1, 'All infected participants are assumed to have detectable virus.' Since the fitted model itself implies that mAb shortens shedding and reduces seroconversion, any infected participants who never shed detectable virus and never seroconvert would be counted as uninfected, and the missingness would plausibly be differential by arm. The reported PE against asymptomatic infection (38.7%; 95% CI 10.0-60.8) and the mAb-arm shedding duration would then mix a true biological effect with a detection effect. The paper should either consistently label the endpoint as 'detectable infection' in the abstract and Discussion, or provide a sensitivity analysis that relaxes assumption 1 (e.g., a latent infection state with assumed detection probabilities).","section":"Section 2.1 / Table 1b / Section C.1"},{"comment":"The simulation study reports coverage of 0.85-0.87 for the restricted mean time to infection on the placebo arm and 0.90 for the restricted mean duration of PCR positivity on the placebo arm under the semi-parametric five-state model (Table 2(c); Tables S9-S10), with the paper attributing the under-coverage to infections detected only by serology. These are precisely the functionals used for the headline clinical estimates in Section 4.3: the placebo-arm shedding duration of 13.0 days (95% CI 11.5-14.6) and the comparison with 6.2 days on the mAb arm. The paper should either provide calibration of the bootstrap intervals for these functionals under the actual REGEN-2069 observation design, or explicitly qualify the reported confidence intervals as potentially anti-conservative for the duration endpoint.","section":"Section 3.2 / Table 2(c) / Section 4.3"}],"minor_comments":[{"comment":"The sentence 'Our algorithm provide substantial computational improvements' has a subject-verb agreement error; it should read 'Our algorithm provides.'","section":"Abstract"},{"comment":"The affiliation 'Regeneron Pharmaceuticles' appears to contain a typo and should likely read 'Regeneron Pharmaceuticals.'","section":"Author affiliations"},{"comment":"The word 'identifiabile' should be 'identifiable'.","section":"Introduction"},{"comment":"In the table note, 'folloowup' should be 'follow-up'.","section":"Table S6"},{"comment":"The reference to Akaike (1998) is spelled 'Aikaike' in the reference list and should be corrected.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of a statistical methodology journal, and the methodological core is sound and well evaluated. The main risk is that the abstract and Discussion overstate the biological interpretation of 'infection' relative to the detection-based definition used in the model; this is fixable with consistent language or a sensitivity analysis. I do not see concerns about citation practice or authorship disclosure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The estimating machinery is the real contribution here, and it is well built. The clinical numbers are plausible but rest on a measurement-defined notion of infection that the paper never stress-tests.\n\nWhat is actually new: combining MCEM with importance sampling from a data-conditioned Markov surrogate to fit semi-Markov multistate models under interval censoring, with spline intensities and covariate effects. That combination is not in the cited literature—Barone and Tancredi do Bayesian data augmentation with a similar surrogate, and Aralis and Brookmeyer use rejection sampling that does not scale. The importance weights simplify cleanly, the ascent MCEM recovers the ascent property, and there is a Monte Carlo estimator of the marginal likelihood for AIC comparison. That last piece is genuinely useful.\n\nThe validation is good. Two simulation studies: an illness-death model with follow-up every 1 or 3 months, and a REGEN-analog trial with weekly PCR and end-of-study serology. The semi-Markov models beat crude tabulation and Markov models by a wide margin, and the spline models are competitive with the correctly specified Weibull model. The phase-type comparison is fair and shows a real failure mode. Code is shipped in Julia and R, and the supplement is thorough.\n\nThe soft spot is real, and it sits in the clinical application rather than the algorithm. Infection is defined as transition to PCR-detectable shedding; Table 1b assumption 1 says participants who never become symptomatic, PCR+, or seropositive are uninfected. That makes the estimand literally the measured event. A true infection that never sheds and never seroconverts is invisible to the model, and because the paper estimates that mAb shortens shedding and reduces seroconversion, the missingness mechanism is plausibly differential by arm. The reported PE against asymptomatic infection (38.7%) and the mAb-arm shedding duration (6.2 days) could partly reflect differential detectability. The simulations do not address this: Section C.1 assumes all infected participants have detectable virus, so the validation builds in exactly the assumption at issue. A sensitivity analysis allowing a fraction of infections to be invisible would materially strengthen the paper.\n\nMinor issues: confidence intervals come from the AIC-selected model without propagating model-selection uncertainty, and the simulations show modest under-coverage for restricted mean time to infection and PCR+ duration (the paper is honest about both). The reproducibility artifacts lack a commit hash, and the clinical data obviously is not public.\n\nWho this is for: biostatisticians working on multistate models with panel data, and anyone analyzing trials with composite ascertainment of endpoints. The methodological core deserves a serious referee; the clinical conclusions deserve careful reading, not reflexive skepticism. I would accept it for review, with the sensitivity analysis as the main revision ask.","headline":"The MCEM-with-importance-sampling machinery is a real, well-validated methodological contribution; the REGEN-2069 clinical numbers are plausible but rest on a measurement-defined infection endpoint that the paper never stress-tests.","tokens_in":34444,"tokens_out":2775,"would_cite":true,"duration_ms":26120,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62N01","62M05","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a Monte Carlo expectation-maximization algorithm can fit multistate semi-Markov models to intermittently observed, multi-stream trial data, and uses it to estimate that REGEN-COV prophylaxis reduced SARS-CoV-2…","keywords":["multistate semi-Markov models","interval-censored data","Monte Carlo expectation-maximization","importance sampling","multiple data streams","SARS-CoV-2","protective efficacy","panel data"],"falsifier":"Simulate a trial from the paper's nine-state simulation model but add a fraction of infected participants who clear virus before the first PCR assessment and never seroconvert, then fit the five-state model and check whether the estimated protective efficacy against asymptomatic infection is biased relative to the simulated truth by more than the simulation's Monte Carlo error.","tokens_in":2107,"feed_emoji":"🦠","tokens_out":3299,"duration_ms":89496,"temperature":0.7,"pith_summary":"The paper introduces a computationally efficient Monte Carlo expectation-maximization (MCEM) framework for fitting multistate semi-Markov models when the data are interval-censored, incomplete, or subject to measurement error. The motivating application is the REGEN-2069 trial, where symptom onset, weekly RT-qPCR testing, and end-of-study serology each give a partial view of infection and recovery. Using the fitted model, the paper estimates that REGEN-COV reduced the risk of any infection by 60.4%, reduced asymptomatic infection by 38.7%, and shortened mean detectable viral shedding from 13.0 to 6.2 days. A sympathetic reader would care because the method offers a general way to synthesize multiple, imperfect data streams into estimates of protection and disease dynamics that crude tabulations or Markov models cannot reliably provide.","feed_headline":"REGEN-COV cuts infection risk 60% and shedding by half","feed_subtitle":"Fit to symptom, PCR, and serology data, this model estimates REGEN-COV's efficacy from interval-censored trial observations.","key_machinery":"The central machinery is a five-state progressive multistate model (naive, PCR-positive asymptomatic, PCR-negative post-infection, symptomatic PCR-positive, symptomatic PCR-negative) combined with an MCEM estimation algorithm. In each E-step, complete histories are proposed from a time-homogeneous Markov surrogate conditioned on the observed data, using forward-filtering backwards-sampling to impute unknown states at observation times and uniformization to fill in endpoint-conditioned paths between visits; self-normalized importance weights correct for the discrepancy between the Markov proposal and the semi-Markov target. An ascent-based rule augments the Monte Carlo sample until changes in the Q-function are distinguishable from Monte Carlo error, and spline or Weibull baseline intensities allow semi-parametric inference. This construction sidesteps the intractable transition probabilities that make semi-Markov models difficult to fit to panel data, and a marginal-likelihood estimator built from the same importance weights enables model comparison via AIC.","core_discovery":"The paper's central claim is that intractable semi-Markov likelihoods for intermittently observed multistate processes can be maximized by an MCEM algorithm that samples complete sample paths from a data-conditioned Markov surrogate and reweights them by importance sampling. In the REGEN-2069 reanalysis, this yields estimates that REGEN-COV reduced all-comer infection by 60.4% (95% CI 44.9-72.5%), symptomatic infection by 83.6% (95% CI 69.4-93.1%), and asymptomatic infection by 38.7% (95% CI 10.0-60.8%), while reducing the mean duration of PCR-detectable shedding from 13.0 days (95% CI 11.5-14.6) to 6.2 days (95% CI 5.0-7.8). Simulations show that crude event tabulation and Markov model fits are biased for these functionals, whereas the semi-Markov fits achieve negligible bias and near-nominal coverage, supporting the authors' conclusion that the method recovers difficult-to-measure endpoints implicated by asymptomatic infection.","pith_inferences":["Editorial inference: the definition of infection as detectable RT-qPCR positivity is the load-bearing assumption, and if some infected participants never shed detectable virus or seroconvert, the estimated asymptomatic-infection protective efficacy would shift; a sensitivity analysis that redefines infection or adds a fraction of undetectable infections would quantify this.","Editorial inference: the same importance-sampling scheme could be extended to disease-driven observation processes, where sicker patients are tested more often, by modifying the proposal to condition on the observation times, an extension the paper notes but does not implement.","Editorial inference: because the algorithm provides a Monte Carlo estimate of the marginal likelihood, it could be combined with penalized-likelihood or weighted-bootstrap Bayesian procedures for model selection and multiple-testing corrections in other panel-data trials.","Editorial inference: the comparison with phase-type models suggests a testable roadmap: using a phase-type or discrete mixture of Markov processes as the proposal distribution could improve robustness in settings where a simple Markov surrogate yields a small effective sample size."],"forward_implications":["If the central claim is correct, interval-censored multi-stream trial data no longer require Markov assumptions: semi-Markov models with Weibull or spline intensities give approximately unbiased estimates of transition intensities and their functionals with nominal coverage.","For REGEN-2069 specifically, the estimates imply that REGEN-COV's protection is broad, cutting infection risk by 60.4% overall and by 83.6% for symptomatic infection, while shortening the period of detectable shedding from 13.0 to 6.2 days, a change with direct implications for transmission.","The 38.7% estimate of protective efficacy against asymptomatic infection is positive but smaller than the symptomatic efficacy, and the model attributes part of this apparent benefit to the fact that antibody treatment shortens shedding below the detection threshold of weekly PCR sampling.","Because seroconversion was less frequent after asymptomatic infections on the mAb arm, the data support a mechanism in which antibody treatment suppresses viral load, symptoms, and immune exposure together.","The framework extends beyond COVID-19 to any progressive disease process whose state is partially observed through several measurement modalities, with AIC-based model choice made feasible by the marginal-likelihood estimator."],"supporting_citations":[{"why":"Supplies the primary REGEN-2069 trial results and the data streams analyzed in the application.","marker":"O'Brien and others (2021)"},{"why":"Provides the viral shedding and symptom-onset kinetics used to order the model states.","marker":"Puhach and others (2023)"},{"why":"Provides the empirical basis for mAb suppressing seroconversion, used in interpreting trial estimates.","marker":"Follmann and others (2022)"},{"why":"Defines the EM algorithm whose E-step the paper replaces with Monte Carlo sampling.","marker":"Dempster and others (1977)"},{"why":"Introduces the Monte Carlo EM algorithm that the paper builds on.","marker":"Wei and Tanner (1990)"},{"why":"Gives the ascent-based acceptance and stopping rules that make MCEM stable in practice.","marker":"Caffo and others (2005)"},{"why":"Provides the uniformization algorithm for simulating endpoint-conditioned continuous-time Markov paths.","marker":"Hobolth and Stone (2009)"},{"why":"Provides forward-filtering backwards-sampling used to impute state labels at observation times.","marker":"Scott (2002)"},{"why":"Closest prior data-augmentation approach using a Markov surrogate to propose semi-Markov paths; the paper's efficiency comparison target.","marker":"Barone and Tancredi (2022)"},{"why":"Supplies a rejection-sampling baseline that the paper's data-conditioned proposal is shown to outperform.","marker":"Aralis and Brookmeyer (2019)"}],"fun_headline_variants":["New algorithm fits interval-censored trials, cuts infection risk 60%","MCEM method estimates REGEN-COV cuts asymptomatic infection 39%","Semi-Markov model with MCEM: shedding cut from 13 to 6 days","Importance sampling unlocks semi-Markov fits for sparse trial data","Method reveals REGEN-COV's full efficacy from incomplete data"],"cache_read_input_tokens":36480,"weakest_assumption_plain":"The argument assumes that being infected is the same as being measurably affected, specifically shedding virus detectable by nasopharyngeal RT-qPCR, so a person who is infected but never tests positive, never shows symptoms, and never seroconverts is classified as uninfected.","fun_headline_variants_meta":{"raw":{"variants":["New algorithm fits interval-censored trials, cuts infection risk 60%","MCEM method estimates REGEN-COV cuts asymptomatic infection 39%","Semi-Markov model with MCEM: shedding cut from 13 to 6 days","Importance sampling unlocks semi-Markov fits for sparse trial data","Method reveals REGEN-COV's full efficacy from incomplete data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1550,"prompt_tokens":1013,"completion_tokens":537,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":438}},"tokens_in":629,"tokens_out":537,"duration_ms":4457,"temperature":1.0,"reasoning_tokens":438,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:22:16.056916+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a trial from the paper's nine-state simulation model but add a fraction of infected participants who clear virus before the first PCR assessment and never seroconvert, then fit the five-state model and check whether the estimated protective efficacy against asymptomatic infection is biased relative to the simulated truth by more than the simulation's Monte Carlo error.","supporting_citations":[{"cited_title":"and others","cited_arxiv_id":null,"evidence_quote":"Supplies the primary REGEN-2069 trial results and the data streams analyzed in the application."},{"cited_title":"and Eckerle , I","cited_arxiv_id":null,"evidence_quote":"Provides the viral shedding and symptom-onset kinetics used to order the model states."},{"cited_title":"and others","cited_arxiv_id":null,"evidence_quote":"Provides the empirical basis for mAb suppressing seroconversion, used in interpreting trial estimates."},{"cited_title":"and Rubin , D.B","cited_arxiv_id":null,"evidence_quote":"Defines the EM algorithm whose E-step the paper replaces with Monte Carlo sampling."},{"cited_title":"and Tanner , M.A","cited_arxiv_id":null,"evidence_quote":"Introduces the Monte Carlo EM algorithm that the paper builds on."},{"cited_title":"and Jones , G.L","cited_arxiv_id":null,"evidence_quote":"Gives the ascent-based acceptance and stopping rules that make MCEM stable in practice."},{"cited_title":"and Stone , E.A","cited_arxiv_id":null,"evidence_quote":"Provides the uniformization algorithm for simulating endpoint-conditioned continuous-time Markov paths."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides forward-filtering backwards-sampling used to impute state labels at observation times."},{"cited_title":"and Tancredi , A","cited_arxiv_id":null,"evidence_quote":"Closest prior data-augmentation approach using a Markov surrogate to propose semi-Markov paths; the paper's efficiency comparison target."},{"cited_title":"and Brookmeyer , R","cited_arxiv_id":null,"evidence_quote":"Supplies a rejection-sampling baseline that the paper's data-conditioned proposal is shown to outperform."}],"review_version":1}