{"id":"33b884db-5cda-4b78-9530-1292bf6ac790","arxiv_id":"2508.05472","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Jointly modeling inter-observation time and missingness alongside survival outcomes in a multi-task RNN improves clinical prediction transportability under clinical presence shift.","lead":"A new neural network learns from patient records and from when and how often patients visit the healthcare system to predict survival outcomes. The aim is to make such predictions work reliably across different hospitals, regions, or countries where patient behavior differs.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Transportability claim hinges on unstated invariance of observation–outcome dependence; abstract gives no evidence for this assumption.","rationale":"The reader's weakest assumption is the same as my primary concern: the stability of the observation-outcome dependence across settings. The abstract alone provides no support for this assumption, and the MIMIC-III evaluation cannot establish external transportability. I therefore concur that the paper remains unverified. My suggested test—external validation or a controlled decoupling of observation and outcome—would provide the evidence needed to move from UNVERDICTED to a substantive verdict. I do not see reason to change the reader's verdict because the concern is conditional and not yet shown to fail.","tokens_in":606,"tokens_out":4966,"duration_ms":51726,"concrete_test":"Obtain the full manuscript and extract the exact theoretical assumption about the observation process (likely in a formal model or Proposition). Then run a transportability experiment on a genuinely external dataset (e.g., train on MIMIC-III, test on eICU or a different year/ward) with patients' measurement frequency distributions matched. If the proposed model's AUROC or Brier score improves over baselines that ignore observation, the assumption may hold; if performance degrades relative to baselines when the observation process distribution changes, the invariance assumption is violated and the transportability claim fails. Alternatively, simulate a shift by randomly permuting observation times within each patient on the test set, breaking the link between observation and severity, and compare the joint model to a baseline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"For the joint model to improve transportability under clinical presence shift, the observation process must remain predictive of the latent health status in the new setting. The abstract claims theoretical justification but does not specify any invariance or causal assumption. If the shift alters not only the frequency of measurements but also the meaning of those measurements (e.g., different hospital protocols, different billing practices, or changes in patient willingness to seek care), the conditional dependence between observation and outcome is no longer stable. The model would then allocate capacity to learning a transferable 'observation style' that is actually an artifact of the training site, degrading external performance. The MIMIC-III evaluation, being single-site, cannot validate transportability; any test-set shift is likely simulated by resampling within the same observation mechanism, which does not mimic a true change in the conditional dependence. Thus the central claim rests on an implicit, untested stability assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a multi-task recurrent neural network for EHR prediction that jointly models the inter-observation time, the missingness process, and the survival outcome of interest. It introduces the notion of 'clinical presence shift' for deployment in new settings and claims a theoretical justification for why joint modeling of the observation process improves transportability under such shift. Empirical evidence is reported on MIMIC-III for mortality prediction, where the proposed strategy is said to outperform state-of-the-art models that ignore the observation process.","tokens_in":832,"tokens_out":2215,"duration_ms":24095,"significance":"If the claims hold, the work addresses a practically important and understudied problem: clinical prediction models are often trained on EHR data whose observation intensity reflects local healthcare practices, and transportability across hospitals, regions, or countries is a real concern. The proposed multi-task architecture is a plausible way to make the model aware of the observation process. The explicit formalization of 'clinical presence shift' and the claim of a theoretical justification are valuable contributions, and the MIMIC-III mortality prediction task provides a concrete falsifiable empirical test. However, with only the abstract available, the soundness of the theory and the validity of the empirical comparison cannot be verified, so the significance remains conditional.","major_comments":[{"comment":"The sentence 'we theoretically justify why the proposed joint modelling can improve transportability under changes in clinical presence' is load-bearing but states no assumptions. The natural sufficient condition is that the conditional dependence between the observation process and the latent health state is invariant across settings. If clinical presence shift also changes this dependence (e.g., different coding practices, protocols, or patient help-seeking behavior), the joint model may learn a site-specific 'observation style' rather than a transferable representation. Please state the formal assumptions behind the theoretical claim and, if this invariance is one of them, either prove robustness to its violation or test it explicitly with an external-site evaluation.","section":"Abstract"},{"comment":"The empirical claim 'we demonstrate, in a real-world mortality prediction task in the MIMIC-III dataset, how the proposed strategy improves performance and transportability' relies on a single-center dataset. As reported, any test-set shift is simulated within the same observation mechanism; this does not validate transportability to a setting in which the observation-outcome dependence changes. The abstract also does not specify the comparators, the evaluation metrics, or the uncertainty estimates. The full manuscript must provide these details, and ideally an external validation, before the transportability claim can be accepted.","section":"Abstract"},{"comment":"The phrase 'state-of-the-art prediction models' is not defined and no baseline names are given. Since the central empirical claim is comparative, the choice of baselines and their configurations directly affects the conclusion. The full text should name the baselines, describe how they were trained, and report statistical significance or confidence intervals for the performance differences.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract uses 'clinical presence shift' as a formal concept but does not give a compact definition. A one-sentence mathematical description (e.g., a change in the intensity or missingness process) would help readers assess the scope of the claim.","section":"Abstract"},{"comment":"Minor clarity issue: 'impacting performance and limiting the transportability of models' would read more precisely as 'degrades performance and limits transportability' if the direction of the effect is intended to be negative.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"This review is based only on the abstract; the full text was not available. The stress-test concern about the stability of the observation-outcome dependence is a live question rather than a demonstrated flaw. I recommend obtaining the full manuscript before making a final decision. If the full paper provides formal assumptions and external validation, the manuscript may be suitable for publication; otherwise, the central transportability claim is not established."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper takes on a genuine issue: clinical prediction models trained on EHR data often fail when the observation process—how often and when patients interact with the system—changes. The idea to jointly model inter-observation time, missingness, and survival in one recurrent network is sensible and not something I've seen in this exact combination. The formalization of \"clinical presence shift\" is useful vocabulary for a known but under-theorized problem. Credit where due: the authors identify a real gap and propose a coherent architectural response, plus they test on MIMIC-III, which is a reasonable benchmark. That is a solid start.\n\nThe soft spot is that the central transportability argument leans on an assumption that is nowhere stated in the abstract. To improve transportability under presence shift, the model must rely on the observation process staying predictive of the underlying health state in the new setting. That is a strong invariance condition. Different hospitals have different ordering practices, different billing incentives, and different patient help-seeking behavior; those can change not just how often measurements are taken but what those measurements mean clinically. If the full paper does not explicitly articulate this assumption and justify it, the theoretical claim will not hold. The MIMIC-III evaluation, being single-site, cannot by itself demonstrate transportability to genuinely different settings; any test-time shift is likely simulated by resampling within the same observation mechanism, which does not reproduce a change in the conditional dependence. That is a significant gap, though perhaps addressable in the full text with external validation.\n\nI want to be clear: none of this is fatal on its face. The idea is plausible, and the missing pieces could easily be in the full paper. But from the abstract alone, the strongest claim is not established. The theoretical justification is asserted, not shown, and the empirical evidence is suggestive at best.\n\nThis paper deserves a serious referee. If the full paper states its assumptions, tests them with sensitivity analyses or external data, and shows that the joint model helps when the observation mechanism truly changes, it would be a useful contribution. The area is important, and this is a legitimate attempt rather than a toy exercise. I'd send it out, but I'd ask the referees to focus sharply on the invariance assumption and on whether the evaluation actually simulates presence shift or merely reweights the same data.\n\nFor the reading group: worth a maybe, mostly to discuss what \"transportability\" should mean in this literature.","headline":"Tackles a real problem in EHR model transportability, but the abstract alone leaves the key invariance assumption unstated and the single-site evaluation cannot substantiate the transportability claim.","tokens_in":1188,"tokens_out":1330,"would_cite":false,"duration_ms":16447,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that jointly modeling the observation process—when and how patients interact with the healthcare system—alongside the survival outcome improves prediction performance and transportability under clinical presence shift.","keywords":["clinical presence shift","electronic health records","survival prediction","recurrent neural network","missingness","inter-observation time","transportability","mortality prediction"],"falsifier":"Compare the joint model with a standard survival model across two settings where the observation process shifts and the conditional association between observation pattern and outcome is known to differ—for example, a site that introduces routine telehealth check-ins, changing the meaning of a missed visit. If the joint model fails to beat the model that ignores presence, or underperforms when that association differs, the transportability claim is wrong.","tokens_in":596,"feed_emoji":"🩺","tokens_out":4670,"duration_ms":44799,"temperature":0.7,"pith_summary":"The paper claims that clinical presence—the timing and missingness of patient observations in electronic health records—is not a nuisance to be discarded but a signal that carries information about future health. It proposes a multi-task recurrent neural network that models inter-observation time and missingness jointly with the survival outcome. The authors argue this approach improves both prediction performance and transportability to new settings where clinical presence shifts, and they support this with a theoretical justification and a MIMIC-III mortality prediction experiment. If correct, this would mean that models trained on EHR data can become more reliable when deployed in hospitals or regions with different patterns of patient contact.","feed_headline":"Modeling patient visits improves survival predictions","feed_subtitle":"Models visit timing and missingness alongside outcome to make mortality predictions transferable across care settings.","key_machinery":"The key machinery is a multi-task recurrent neural network with a shared hidden state and three prediction heads: one for the survival outcome, one for the inter-observation time (time until the next clinical contact), and one for the missingness pattern of recorded variables. The shared representation forces the model to learn how the observation process relates to health status. This shared learning is what allows the model to transfer better when the observation process changes across settings.","core_discovery":"The central claim is that clinical presence shift—changes in how often and in what patterns patients are observed between development and deployment settings—can be exploited rather than ignored. By training a network with three tasks: predicting the survival outcome, predicting the time until the next observation, and predicting which variables are missing, the model learns the dependence between health status and healthcare-seeking behaviour. The paper provides a formal definition of clinical presence shift and a theoretical argument that joint modelling improves transportability under such shifts. Empirically, on the MIMIC-III in-hospital mortality task, the joint model outperforms state-","pith_inferences":["The same joint-modelling principle may extend to other EHR outcomes such as readmission, length of stay, or complication rates, wherever encounter timing and missingness are informative.","A natural sensitivity analysis would measure the strength of the presence–outcome association in the source data and test whether the transportability gain scales with it; the paper's theory implies such a relationship.","If clinical presence partly reflects patient access and choice, the model could implicitly encode disparities in healthcare utilisation; applying it without accounting for that could compound bias.","The theoretical justification appears to assume a missing-not-at-random mechanism; making that assumption explicit would help practitioners decide when the method is appropriate."],"forward_implications":["Clinical prediction models should treat the observation process as a predictive feature rather than a nuisance, because the pattern of patient contact encodes health information.","Jointly modelling inter-observation time, missingness, and outcome can improve transportability to other hospitals, regions, or countries where clinical presence differs.","On MIMIC-III mortality prediction, the proposed strategy achieves better performance than state-of-the-art baselines that ignore the observation process.","The formal definition of clinical presence shift provides a concrete target for evaluating and comparing model transportability in EHR-based prediction."],"supporting_citations":[],"fun_headline_variants":["Survival models that learn clinical presence are more portable","Modeling observation patterns makes survival predictions transferable","Joint modeling of outcomes and visit timing improves model portability","Clinical presence shift handled in survival models for better transfer","Including visit timing and missingness in survival models improves portability"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The transportability gain relies on the conditional dependence between clinical presence and underlying health status being stable across settings; if a shift in clinical presence also changes how informative presence is about health, the joint model's advantage may vanish.","fun_headline_variants_meta":{"raw":{"variants":["Survival models that learn clinical presence are more portable","Modeling observation patterns makes survival predictions transferable","Joint modeling of outcomes and visit timing improves model portability","Clinical presence shift handled in survival models for better transfer","Including visit timing and missingness in survival models improves portability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001043,"raw_usage":{"total_tokens":4192,"prompt_tokens":683,"completion_tokens":3509,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":427,"completion_tokens_details":{"reasoning_tokens":3430}},"tokens_in":427,"tokens_out":3509,"duration_ms":24764,"temperature":1.0,"reasoning_tokens":3430,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:17:18.515054+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the joint model with a standard survival model across two settings where the observation process shifts and the conditional association between observation pattern and outcome is known to differ—for example, a site that introduces routine telehealth check-ins, changing the meaning of a missed visit. If the joint model fails to beat the model that ignores presence, or underperforms when that association differs, the transportability claim is wrong.","supporting_citations":[],"review_version":1}