{"id":"addecb43-54a2-4d16-ae98-8c2253eec85f","arxiv_id":"2605.03198","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"When long-term survivors are present in both groups, conventional and non-proportional hazards tests exhibit non-monotonic power as follow-up time increases, whereas correctly specified parametric cure models show monotonic power gains and highest performance at long follow-up.","lead":"This paper runs simulations to compare statistical tests for time-to-event data containing long-term survivors who never experience the event. It shows that standard tests can produce counterintuitive power patterns depending on follow-up length when survivors are present in both groups, while parametric models behave more predictably.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's conditional verdict rests on the inherent limits of simulation evidence, which the full manuscript does not overcome with external validation or real-data checks. Because the reported patterns are mechanistic consequences of the chosen DGPs and are presented as such, the assessment requires no adjustment.","tokens_in":1802,"tokens_out":331,"duration_ms":17915,"concrete_test":"Re-implement the two-group cure-model DGP with both groups having positive cure fractions (e.g., 20-40% as in the study), generate event times under the latency distributions, apply administrative censoring at a grid of follow-up times, and recompute empirical power for the log-rank test and the parametric cure model at n=200 per arm; confirm whether log-rank power peaks and then declines while parametric power continues to rise.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is an empirical observation from controlled simulations: when both groups contain long-term survivors, log-rank and one non-PH test exhibit non-monotonic power versus follow-up time, while the correctly-specified parametric model shows monotonic increase and highest power at longest follow-up. This pattern is generated directly by the reported data-generating processes (mixture cure models with varying cure fractions and latency distributions) and the implemented test statistics. No internal inconsistency, hidden assumption in the power calculation, or unsupported extrapolation beyond the simulated regimes is present. The reader's weakest assumption (DGP realism and correct parametric specification) is the standard limitation of any simulation study and does not undermine the validity of the within-simulation comparisons.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript reports a neutral Monte Carlo simulation study comparing conventional log-rank tests, selected non-proportional-hazards procedures, and a correctly-specified parametric mixture-cure model for two-sample testing of time-to-event data that contain long-term survivors. Simulations vary sample size, follow-up duration, cure fractions, and latency distributions; type-I error and power are reported for each configuration. The central empirical finding is that, when both arms contain long-term survivors, log-rank and one non-PH test exhibit non-monotonic power curves with respect to follow-up time, whereas the parametric model yields monotonically increasing power that is highest at the longest follow-up examined. A numerical diagnostic is proposed to anticipate non-monotonicity during study planning.","tokens_in":1925,"tokens_out":591,"duration_ms":22194,"significance":"If the reported patterns are robust, the work supplies immediately actionable guidance for oncology trial design where long-term survivors are increasingly common. The simulation design is explicitly varied across sample size, follow-up, and effect size and reports both error rates and power, satisfying standard reproducibility expectations for simulation studies. The provision of a planning-stage numerical check for non-monotonicity is a concrete practical contribution.","major_comments":[{"comment":"§3 (Simulation design): the data-generating process is described only at a high level in the abstract and main text. Exact parameter values for cure fractions, latency distributions, and the censoring mechanism (including how administrative censoring at the end of follow-up is implemented) are not tabulated; without these, independent verification of the non-monotonicity result is impossible.","section":"§3"},{"comment":"§4.2 (Power results when both groups contain L-TS): the claim that conventional and one non-PH test display non-monotonic power is load-bearing for the paper’s main message. The manuscript does not report the precise definition of “follow-up time” (e.g., whether it is the administrative censoring time or the maximum observed time) nor the number of Monte Carlo replicates per cell, both of which directly affect whether the observed non-monotonicity is an artifact of the chosen censoring scheme.","section":"§4.2"}],"minor_comments":[{"comment":"The specific non-PH method that exhibits the non-monotonic pattern is referred to only as “one non-PH method” in the abstract; name the procedure (and cite its reference) at first use in the results.","section":"Abstract"},{"comment":"The numerical approach for predicting non-monotonicity is introduced in the discussion but lacks an explicit algorithm, pseudocode, or worked numerical example; adding one would improve usability.","section":"Discussion"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We are grateful to the referee for their thorough review and for recognizing the potential practical value of our findings for oncology trial design. We respond to each major comment in turn and will implement revisions to enhance the manuscript's clarity and reproducibility.","responses":[{"response":"We concur that the simulation design section would benefit from greater specificity to allow independent verification. In the revised manuscript, we will add a table in §3 that tabulates all key parameters of the data-generating process, including the cure fractions for each group, the specific distributions and parameters for the latency times, and the details of the administrative censoring mechanism.","revision_made":"yes","referee_comment":"[§3] §3 (Simulation design): the data-generating process is described only at a high level in the abstract and main text. Exact parameter values for cure fractions, latency distributions, and the censoring mechanism (including how administrative censoring at the end of follow-up is implemented) are not tabulated; without these, independent verification of the non-monotonicity result is impossible."},{"response":"We acknowledge that the definition of follow-up time and the Monte Carlo sample size are important for interpreting the power results. We will clarify in the revised §4.2 that follow-up time corresponds to the administrative censoring time, and we will report the number of Monte Carlo replicates per configuration. These details will also be cross-referenced in the simulation design section to address the concern about potential artifacts.","revision_made":"yes","referee_comment":"[§4.2] §4.2 (Power results when both groups contain L-TS): the claim that conventional and one non-PH test display non-monotonic power is load-bearing for the paper’s main message. The manuscript does not report the precise definition of “follow-up time” (e.g., whether it is the administrative censoring time or the maximum observed time) nor the number of Monte Carlo replicates per cell, both of which directly affect whether the observed non-monotonicity is an artifact of the chosen censoring scheme."}],"tokens_in":1558,"tokens_out":451,"duration_ms":30548,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is straightforward: in simulations where both groups contain long-term survivors, the power of the log-rank test and one non-proportional hazards method rises then drops as follow-up lengthens, whereas the parametric cure model shows steady gains and tops out at the longest follow-up. This pattern appears consistently across sample sizes and effect sizes they tested. The paper also supplies a numerical check to flag the risk of non-monotonicity at the planning stage. That combination is the useful output for people who design oncology trials. The work is a head-to-head simulation that covers conventional tests, non-PH adaptations, and a parametric model under controlled variation in sample size, follow-up, and cure fractions. It reports both type I error and power, which keeps the comparison grounded. The non-monotonic finding is presented as an empirical observation from the chosen data-generating processes rather than a derived theorem, and the stress-test note confirms no internal contradictions in how the power curves are produced. The numerical predictor is a practical addition that builds on the simulation results without claiming a new theoretical framework. The soft spots are the standard ones for this type of study. Results depend on the mixture cure models used to generate the data; different latency distributions or dependence structures in real trials could shift the patterns. The parametric model is correctly specified by construction, which is optimistic compared with real analysis where misspecification is common. The abstract is high-level on exact implementations, so full reproducibility would require the code or more granular description. These are not load-bearing flaws for the within-simulation comparisons, but they limit how far the findings generalize without further checks. This paper is for biostatisticians who work on survival trial design or analysis in settings with improving therapies and potential cure fractions. A reader who needs to choose tests or set follow-up targets will find the reported patterns and the planning tool directly applicable. It deserves a serious referee because the question is practical, the simulations are neutral and controlled, and the central observation is reproducible from the stated data-generating processes. I would send it for review.","headline":"When both arms have long-term survivors, log-rank and one non-PH test show non-monotonic power with follow-up time while a correctly specified parametric model does not.","tokens_in":2437,"tokens_out":491,"would_cite":true,"duration_ms":24950,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"When both groups have long-term survivors, conventional log-rank and some non-PH tests show non-monotonic power as follow-up increases, while parametric models show steadily rising power.","keywords":["long-term survivors","cure models","two-sample tests","log-rank test","non-proportional hazards","power","follow-up time","simulation study"],"falsifier":"Finding that log-rank or non-PH test power increases monotonically with follow-up time, or that the parametric model does not achieve the highest power at the longest follow-up, in data generated with long-term survivors in both groups would contradict the central result.","tokens_in":2710,"feed_emoji":"📉","tokens_out":662,"duration_ms":53083,"temperature":0.7,"pith_summary":"The paper compares the performance of conventional two-sample tests, non-proportional hazards methods, and parametric cure models for time-to-event data containing long-term survivors who never experience the event. Simulations examine type I error and power across sample sizes, effect sizes, and varying lengths of follow-up. When long-term survivors appear in both groups, standard tests exhibit power that rises then falls or plateaus with longer follow-up, producing counterintuitive results. Parametric models that correctly incorporate the cure fraction display steadily increasing power that peaks at the longest follow-up examined. The authors supply a numerical method to forecast the risk of non-monotonic power during study planning.","feed_headline":"Log-rank power turns non-monotonic with longer follow-up when both groups have survivors","feed_subtitle":"Simulations show standard tests can lose power after a point while parametric cure models keep gaining at extended follow-up.","key_machinery":"Simulation study tracking power and type I error of multiple two-sample tests as functions of follow-up duration when long-term survivors are present in one or both groups.","core_discovery":"In simulations of time-to-event data with long-term survivors present in both groups, conventional log-rank tests and one non-proportional hazards method produce non-monotonic power as a function of follow-up time, whereas a correctly specified parametric cure model yields monotonic increasing power that reaches its highest value at the longest follow-up time considered.","pith_inferences":["Planners facing possible long-term survivors in both arms may need to select follow-up duration with the aid of the numerical check rather than defaulting to standard tests.","The observed non-monotonicity could influence decisions about interim analyses or maximum follow-up in trials where cure fractions are expected."],"forward_implications":["When both groups contain long-term survivors, power patterns across follow-up remain consistent regardless of sample size.","Parametric cure models achieve the highest power at the longest follow-up times examined.","A numerical approach can predict the potential for non-monotonic power before a study begins.","Conventional methods applied without adjustment can produce unexpected power behavior in the presence of long-term survivors."],"fun_headline_variants":["Log-rank power non-monotonic at longer follow-up with survivors in both groups","Tests show non-monotonic power with extended follow-up when both groups have survivors","Parametric cure models achieve top power with longest follow-up in survivor data","Follow-up duration leads to non-monotonic power in conventional tests with survivors"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The chosen simulation data-generating mechanisms match the real behavior of long-term survivors and the parametric model is correctly specified for the scenarios tested.","fun_headline_variants_meta":{"raw":{"variants":["Log-rank power non-monotonic at longer follow-up with survivors in both groups","Tests show non-monotonic power with extended follow-up when both groups have survivors","Parametric cure models achieve top power with longest follow-up in survivor data","Follow-up duration leads to non-monotonic power in conventional tests with survivors"]},"model":"grok-4.3","cost_usd":0.009505,"raw_usage":{"total_tokens":4198,"prompt_tokens":738,"num_sources_used":0,"completion_tokens":72,"cost_in_usd_ticks":95053000,"prompt_tokens_details":{"text_tokens":738,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3388,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":738,"tokens_out":72,"duration_ms":62279,"temperature":1.0,"reasoning_tokens":3388,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-08T17:29:31.761512+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Finding that log-rank or non-PH test power increases monotonically with follow-up time, or that the parametric model does not achieve the highest power at the longest follow-up, in data generated with long-term survivors in both groups would contradict the central result.","supporting_citations":[],"review_version":1}