{"id":"ee87a779-c7cb-4263-909d-83bfc13def7a","arxiv_id":"2509.09227","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Dynamic recovery-rate features and a cross-attention multimodal network improve classification of postoperative BCVA improvement in macular hole patients, but the gain is partly from same-timepoint structural status rather than purely preoperative data.","lead":"This study adds recovery-speed features computed from repeated OCT scans to predict how much vision improves after macular hole surgery, and reports that a multimodal deep learning model beats logistic regression. It is a candidate decision-support tool, but the predictive gain may come partly from using follow-up scans at the same visit as the visual outcome being predicted.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dynamic recovery rate is derived from the same follow-up visit whose BCVA it labels; without a temporal split, the reported AUC gains are concurrent classification, not prospective prediction.","rationale":"The reader's weakest assumption is exactly the load-bearing concern: dynamic recovery rate is computed from the same follow-up visit whose visual outcome it is used to predict. The manuscript itself defines recovery rate via 'the time point at which the lesion was observed to be fully resolved,' so at the moment of prediction the rate is not yet knowable. The Methods' statement that the model predicts 'based on preoperative parameters' is therefore in direct tension with the dynamic feature construction. This makes the headline contribution—that dynamic parameters enhance prediction and form the basis of a decision-support tool—unsupported as stated. The paper still contains useful static-feature findings and a reproducible public-data pipeline, so a conditional framing (concurrent structure-function classification plus future prospective validation) is appropriate. I would not escalate to reject because the underlying association may be real and the flaw is primarily one of temporal framing and missing formal definition, not an internally impossible result. The reader's conditional verdict already captures this, so no change is needed.","tokens_in":9508,"tokens_out":3277,"duration_ms":39161,"concrete_test":"Re-run the 3-month (and all) logistic-regression and multimodal DL experiments with a strict temporal split: for target time t, admit dynamic recovery-rate features only if computed from scans no later than the previous visit (e.g., for month-3 outcome, use 2-week resolution state), and set them to missing otherwise; compare AUC/accuracy against the same models using same-visit recovery rate. If the dynamic-increment gap shrinks to noise or reverses, the prediction-enhancement claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that dynamic structural recovery parameters 'enhance prediction' depends on those parameters being available before the outcome they predict. The Methods define recovery rate as 'the ratio between the initial lesion size on preoperative OCT and the time point at which the lesion was observed to be fully resolved' (§Automated Feature Quantification). Time-to-resolution is only known by observing the same follow-up visits that supply the BCVA improvement label (e.g., 3 months). Thus a model that includes recovery rate at month 3 is not predicting month-3 BCVA from earlier data; it is classifying the same visit using outcome-derived information. The Methods statement that the model predicts 'based on preoperative parameters' contradicts the dynamic feature definition, and no feature-timing offset or equation is provided. Table 2's consistent AUC/accuracy gains with dynamic parameters and Table 3's significant EZ recovery-rate OR therefore cannot support a decision-support/prediction claim; they support a concurrent structure-function association. Reframing as concurrent classification would avoid the overclaim, but the current framing is not internally consistent.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a fully automated pipeline for longitudinal OCT segmentation and feature extraction in idiopathic full-thickness macular hole (iFTMH), introduces a category of 'dynamic recovery parameters' derived from serial OCT, and evaluates their contribution to logistic regression and a multimodal deep learning model for predicting postoperative BCVA improvement at 2 weeks, 3, 6, and 12 months. The central claims are: (i) a stage-specific nnUNet segmentation model with mean Dice 0.862; (ii) adding dynamic recovery rates improves logistic regression AUC, especially at 3 months; and (iii) a multimodal DL model combining clinical data, quantitative features, and raw OCT images outperforms logistic regression, with AUC differences up to 0.12. The paper is framed as a clinical decision-support tool for personalized postoperative management.","tokens_in":9783,"tokens_out":3146,"duration_ms":32371,"significance":"If the predictive claims are valid, the paper would make a useful contribution: it applies a fully automated segmentation/extraction pipeline to a public longitudinal dataset, introduces a clinically plausible class of dynamic structural features, and benchmarks a multimodal DL architecture against classical regression. Strengths include the use of a public dataset, explicit comparison of models with and without dynamic parameters, reporting of ORs with confidence intervals in Table 3, and ablation-style comparisons across data modalities. However, the central claim that dynamic parameters 'enhance prediction' depends critically on whether those parameters are available before the outcome they are used to predict. As written, the dynamic recovery rate is defined using the time point at which the lesion is observed to be fully resolved, which is only known at or after the same follow-up visit that supplies the BCVA label. Unless a temporal offset is introduced and documented, the reported AUC gains are concurrent structure-function associations, not prospective predictions. This is the load-bearing issue for the manuscript's stated decision-support contribution.","major_comments":[{"comment":"The key concern is temporal circularity. The recovery rate is defined as 'the ratio between the initial lesion size on preoperative OCT and the time point at which the lesion was observed to be fully resolved.' For the 3-month model, for example, time-to-resolution is determined by observing visits up to and including the 3-month visit, and the outcome label is BCVA improvement at the same 3-month visit. Thus, recovery rate at month 3 is not known before the month-3 BCVA; the model is classifying the same visit using outcome-derived information, not predicting it. The Methods statement that the model is 'based on preoperative parameters' contradicts the dynamic feature definition. No equations or feature-timing offsets are provided. Please provide the exact definitions of each dynamic parameter, specify the time window used to compute them, and either (a) enforce a strict temporal split","section":"Methods, Automated Feature Quantification; Table 3; Figure 4"},{"comment":"The DL model accepts raw OCT images and dynamic parameters from the same follow-up time point at which the BCVA-improvement label is defined. If the input at month t includes images and features from month t, the model is not predicting month-t BCVA from preoperative or earlier data. This is especially relevant because the reported AUC gain from adding dynamic parameters (0.03-0.04) and the headline DL-vs-regression gap (0.12) may largely reflect leakage of the outcome into the feature set. Please clarify, for each evaluation timepoint, exactly which visits' images and features are used as inputs and which visit's BCVA is the label. If no temporal offset is used, the DL comparison is still informative as a concurrent classification benchmark, but the decision-support claim in the Conclusion must be revised.","section":"Methods, Multimodal Deep Learning Prediction Model; Figure 4"},{"comment":"There is a direct numerical inconsistency: the Abstract states 'mean Dice > 0.89', while the Methods text reports 'the average Dice reached 0.862' and Table 1 reports mean Dice 0.8616. The supportive structures achieve Dice above 0.89 individually, but the global mean does not. Please correct the abstract or report a different summary statistic (e.g., structure-specific range).","section":"Abstract; Methods, Segmentation Model; Table 1"},{"comment":"AUC values and AUC differences (e.g., 0.94 vs 0.87; exclusion of DP reducing AUC by 0.03-0.04) are presented without confidence intervals or significance tests. Given the moderate sample size and the use of a single public dataset, it is not possible to assess whether the reported improvements are statistically reliable or within sampling variability. Please report CIs for AUCs, p-values for AUC comparisons (e.g., DeLong), and the number of eyes with complete data per timepoint.","section":"Results, Logistic Regression Model and Figure 4; Table 2"}],"minor_comments":[{"comment":"The shape-weighted factor is introduced as 'With permitted data, we further incorporated a shape-weighted factor...' but no formula, coefficient, or sensitivity analysis is provided. Please define it precisely or state that it was not used in the reported experiments.","section":"Methods, Automated Feature Quantification"},{"comment":"The number of eyes and the number of events per outcome class are not reported. This is essential for interpreting the stability of multivariable logistic regression with several predictors. Please add these numbers.","section":"Methods, Logistic Regression Model and Dataset"},{"comment":"Typo: 'inclusion of diseases recovery rates' should be 'inclusion of disease recovery rates.' Also, the univariate screening is described as p<0.10 in Methods but the Results text reports P<0.05; please clarify the threshold used.","section":"Results, Logistic Regression Model of Feature Parameters"},{"comment":"The 3-month model reports Recovery Rate (Macular Hole) with P=0.050 and Recovery Rate (EZ) with P=0.048. These are borderline and should be described as such; the abstract's statement that dynamic parameters are significant at P<0.05 is only marginally supported. Also, the text 'For each additional 1.0 μm increase in BD' is awkward for an OR of 0.968 per μm; reporting per 100 μm would be more clinically meaningful.","section":"Table 3"},{"comment":"The BCVA improvement threshold is set at 20 ETDRS letters, but the distribution of Superior vs. not-Superior outcomes is not reported. Without class counts, accuracy and AUC are difficult to interpret. Please add this information.","section":"Methods, Multimodal Deep Learning Prediction Model"},{"comment":"The claim that EZ recovery 'may begin earlier than typically appreciated' is speculative: the dynamic parameter as defined is a ratio of initial size to time of resolution, not a trajectory. Please either provide trajectory-level evidence or temper the claim.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about temporal leakage is well-posed and lands on the central claim. The paper's own definition of recovery rate implies that the dynamic features are outcome-derived at the same timepoint. This should be addressed head-on, either by a causal/temporal re-analysis or by reframing the contribution as concurrent structure-function association. The paper would also benefit from a more direct engagement with the source dataset's prior publication (Godbout et al.), which reportedly cautioned on DL with very limited data; the current comparison is thin. The Dice inconsistency in the abstract, although minor as a numerical matter, suggests a need for careful manuscript verification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nHere's my quick read of arXiv:2509.09227. The genuinely new idea is the dynamic recovery-rate category—features that encode how fast the macular hole and outer retinal defects resolve over time, rather than just their static size. The authors couple that with a multimodal deep learning model that takes clinical data, extracted features, and raw OCT images. The work uses a public dataset and compares against logistic regression both with and without dynamic features, which is the right way to isolate their contribution. The reported AUC gain from adding dynamic parameters is modest and consistent, about 0.03–0.04, and the full DL model beats the best regression by up to 0.12. I buy that as an internal comparison.\n\nThe problem is the framing. The dynamic recovery rate is defined via the time point at which the lesion is 'fully resolved.' That time is only known by looking at the same follow-up visit (or later) whose BCVA improvement the model is classifying. So at 3 months, the model is not predicting the 3-month outcome from preop data; it is using information derived from the 3-month visit itself (or beyond). The paper even says the model predicts 'based on preoperative parameters,' which directly contradicts the dynamic feature definition. This is a load-bearing issue. If you strip the prediction language and call it a concurrent structure-function association analysis, the results are plausible and interesting. But the 'clinical decision support tool' conclusion is not supported.\n\nOther soft spots: the abstract says mean Dice > 0.89, the body reports 0.862—an easy fix but not reassuring. The dynamic parameter equations and the shape-weighted factor are not given, so the novel features are not fully defined. AUCs have no confidence intervals, and there is no external validation. These are fixable in revision.\n\nCredit where due: the segmentation is standard nnUNet, not the contribution, but it is competently done and the authors cite the Godbout/Lachance work appropriately. The public dataset and clear ablations make this reproducible enough to take seriously.\n\nRecommendation: send it to peer review, but the referees should require the timing issue to be resolved. Either the authors reframe the paper as a concurrent classification/association study, or they run a proper temporal split where features at time t come only from earlier visits. If they can do that, the dynamic parameter concept has real value. As it stands, it is a solid association study with an overclaimed title.\n\nRegards.","headline":"The dynamic recovery-rate idea is genuinely new, but the main claim is undermined by using the same follow-up visit for both the feature and the outcome; this is a solid association study, not a prospective predictor.","tokens_in":10238,"tokens_out":4539,"would_cite":false,"duration_ms":48729,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding dynamic structural recovery rates from serial OCT scans improves prediction of visual recovery after macular hole surgery; a multimodal model combining them with clinical data and raw images beats logistic regression by up to 0.12 AU","keywords":["macular hole surgery","OCT segmentation","dynamic recovery parameters","visual acuity prediction","multimodal deep learning","BCVA","ellipsoid zone","longitudinal imaging"],"falsifier":"Recompute recovery rates using only the preoperative scan and earlier follow-up scans (e.g., predict 3-month BCVA from pre-op and 2-week OCT only) and rerun the logistic regression; if the AUC improvement from dynamic parameters disappears or reverses, the claimed added value is an artifact of temporal leakage rather than genuine prognosis.","tokens_in":9414,"feed_emoji":"👁️","tokens_out":3567,"duration_ms":35191,"temperature":0.7,"pith_summary":"The paper argues that the speed at which retinal structures heal after macular hole surgery carries prognostic information that static measurements miss. It introduces 'dynamic parameters'—recovery rates for the macular hole, pseudocysts, and outer retinal layers—and shows that adding them to logistic regression consistently improves prediction of best-corrected visual acuity improvement, most notably at three months. It then shows that a multimodal deep learning model using clinical variables, extracted features, and raw OCT images outperforms regression at every time point, with the largest AUC difference reaching 0.12. If true, this provides a fully automated, longitudinal framework for personalized postoperative counseling and monitoring.","feed_headline":"Healing-speed features sharpen post-surgery vision prediction","feed_subtitle":"Adding recovery rates from serial OCT scans to clinical and imaging data lifts accuracy at every follow-up visit.","key_machinery":"The central new object is the dynamic recovery rate: for each key structure, the ratio of the preoperative lesion extent to the time at which the lesion was observed to have fully resolved, optionally shape-weighted. It turns a sequence of static OCT snapshots into a single speed-of-healing descriptor. The other load-bearing component is the multimodal fusion architecture, which encodes raw OCT images with a pretrained image encoder and fuses them with clinical and feature vectors via cross-attention, letting spatial image information and structured temporal features contribute jointly.","core_discovery":"The central claim is that temporal recovery dynamics, not just static morphology, are predictive of functional vision after macular hole surgery. Using a stage-specific segmentation model on longitudinal OCT scans, the authors compute recovery rates for lesion area and outer retinal defect length, then feed these into prediction models for a binary BCVA-improvement outcome. In logistic regression, including dynamic parameters raises accuracy at all four follow-up visits and brings the ellipsoid-zone recovery rate into the significant predictors at three months (OR 1.298 per unit). The multimodal deep learning model, integrating clinical data, extracted parameters, and raw OCT volumes through","pith_inferences":["Editorial: The reported 3-month gain for dynamic parameters likely depends on recovery rate being computed from the same visit whose visual outcome is predicted; a prospective formulation using only prior scans may shrink or eliminate the gain.","Editorial: Dynamic parameters could be redefined as 'healing velocity over a fixed early window' (e.g., preoperative to 2 weeks) and then tested for predicting later outcomes; this would make the feature truly prospective and clinically actionable.","Editorial: A similar recovery-rate construction might transfer to other surgeries with serial OCT follow-up, such as retinal detachment repair or epiretinal membrane peeling, where structural healing speed may track functional recovery.","Editorial: Because the outcome threshold was set at 20 ETDRS letters rather than the usual 15, the absolute AUC values may not be comparable to studies using standard definitions; the relative ordering of models is the safer conclusion."],"forward_implications":["Including dynamic recovery rates should become standard in models predicting macular hole surgery outcomes; omitting them costs about 0.03–0.04 AUC.","The ellipsoid zone recovery rate at 3 months is a candidate biomarker for photoreceptor reconstitution and intermediate-term visual outcome.","Freely combining clinical data, OCT-derived measurements, and raw OCT images yields the strongest predictions; no single modality is sufficient.","A fully automated segmentation-to-prediction pipeline is feasible for longitudinal retinal OCT, enabling scalable decision support."],"fun_headline_variants":["Recovery-rate OCT features sharpen vision prediction","Dynamic OCT metrics boost post-surgery vision forecasts","Temporal OCT recovery ups macular hole outcome accuracy","Healing speed from scans sharpens vision recovery forecasts","Deep learning plus recovery rates lifts prediction power"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The dynamic recovery rate for a follow-up visit is assumed to be known before the visual acuity measured at that same visit; in the paper it is defined using the time point when the lesion resolved, which is the same examination that supplies the outcome label, so the extra predictive power may come from looking ahead.","fun_headline_variants_meta":{"raw":{"variants":["Recovery-rate OCT features sharpen vision prediction","Dynamic OCT metrics boost post-surgery vision forecasts","Temporal OCT recovery ups macular hole outcome accuracy","Healing speed from scans sharpens vision recovery forecasts","Deep learning plus recovery rates lifts prediction power"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1157,"prompt_tokens":783,"completion_tokens":374,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":527,"completion_tokens_details":{"reasoning_tokens":303}},"tokens_in":527,"tokens_out":374,"duration_ms":5162,"temperature":1.0,"reasoning_tokens":303,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T19:27:19.608124+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute recovery rates using only the preoperative scan and earlier follow-up scans (e.g., predict 3-month BCVA from pre-op and 2-week OCT only) and rerun the logistic regression; if the AUC improvement from dynamic parameters disappears or reverses, the claimed added value is an artifact of temporal leakage rather than genuine prognosis.","supporting_citations":[],"review_version":1}