{"id":"cdc339ad-41e3-4658-ab9c-976119c1358c","arxiv_id":"2509.06227","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Selecting the five best ENSO ensemble members by their past errors and correlations, then scoring them on the same past data, shows large in-sample gains over the all-member mean but provides no out-of-sample evidence for real forecast skill.","lead":"This paper uses past observations to pick the best members out of a 40-member ENSO forecast ensemble, then measures how well those members did on the same observations. The selected members look much better than the ensemble average, especially at long lead times, but the test is in hindsight and no out-of-sample check is provided.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"In-sample member selection inflates all reported gains; an out-of-sample holdout is needed to test whether top-member skill is stable.","rationale":"The reader's weakest assumption correctly identifies the central problem: the 1986–2017 record is used both to rank members and to score the selected subsets. For the paper's operational framing to hold, it must be true that members ranked highly in the training period remain highly skilled in unseen years. The paper provides no such test; its own caveats admit the analysis is retrospective, and the statistical procedures in §2.9 cannot rescue the inference because they resample the same selection period. A simple temporal holdout would directly separate stable member skill from selection noise. Absent that, the reported deltas are exactly what one expects from in-sample optimization, and the claim that the method 'enhances forecasting skill' is unsupported. The reader's REJECT verdict is therefore appropriate, and my read does not change it.","tokens_in":14433,"tokens_out":4056,"duration_ms":45931,"concrete_test":"Holdout validation: form Top-5-by-RMSE and Top-5-by-correlation lists using only 1986–2001 (training) and apply the same weighted Top-10 aggregation to 2002–2017 (validation) for each lead/season. Compute ΔCorrelation and ΔRMSE at L=12, 18, and 23 months on the validation years. If the long-lead gains shrink to near zero or change sign, the reported enhancements are in-sample selection bias; if they remain comparable to +0.3 in correlation and −0.15°C in RMSE, the claim would gain independent support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central inference—that Top-5/Top-10 subsets have 'substantially higher skill' and are operationally relevant—rests on member rankings computed from the same 1986–2017 verification record that is later scored (Sections 2.2, 2.4, 3.1). Because the Top-5 lists are chosen per (lead, season) by minimizing RMSE or maximizing correlation over exactly those years, the reported ΔCorrelation of +0.43 and ΔRMSE of −0.18°C at 23-month lead are in-sample selection statistics, not estimates of predictive skill. The paper repeatedly acknowledges this (e.g., §2.1: 'not an operational selection scheme'; Conclusion: 'retrospective in application'), but the abstract and §3.4 still frame the results as 'enhanced long-term prediction' and a foundation for operational gains. The bootstraps and paired t-tests in §2.9 do not repair this: they resample the same verification years after selection, so they only measure sampling noise in the already-selected deltas, not whether the selection would transfer to unseen years. The unsupported generalization 'for any large enough ensemble' also overreaches from a single, unnamed forecast system. The load-bearing missing piece is any evidence that member ranks are stable across independent time periods.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses a 40-member ensemble of Niño-3.4 forecasts at leads 1–23 months over 1986–2017, ranks members per lead–season pair against ERSST observations by RMSE and Pearson correlation, forms Top-5-by-RMSE, Top-5-by-correlation, and a weighted Top-10 union, and compares these selected forecasts with the equal-weight All-40 mean. The headline results are large gains at long leads, e.g. +0.43 correlation and −0.18°C RMSE at 23 months, with season-dependent patterns. The authors explicitly frame the exercise as retrospective/a-posteriori and disclaim operational use, but the abstract and Section 3.4 present the results as evidence for enhanced long-lead ENSO prediction and as a foundation for operational selection.","tokens_in":14745,"tokens_out":5799,"duration_ms":64048,"significance":"The question of whether selective weighting of ensemble members can improve long-lead ENSO prediction is practically relevant, and the paper clearly exposes the mechanics of ranking by RMSE and correlation, including a subset-size sensitivity analysis and a caveat that the method is retrospective. However, the central empirical claim is not supported: the top members are selected from the same 1986–2017 verification record on which they are then scored, so the reported improvements are in-sample selection statistics. A Top-5 subset chosen by minimizing RMSE or maximizing correlation on a record will almost always beat the full ensemble mean on that same record, even if the members contain no useful forecast signal. The bootstrap and paired t-tests in Section 2.9 do not repair this bias. The study therefore provides an example of retrospective optimization, not evidence of predictive skill. The forecast system is also unnamed, and no code or data are provided.","major_comments":[{"comment":"The Top-5-by-RMSE and Top-5-by-Correlation lists are formed for each (lead, season) by scoring every member against the observed Niño-3.4 series y_t, and the same y_t is then used in Section 3.1 to compute the reported ΔCorrelation and ΔRMSE. This is a direct in-sample selection loop. Selecting the five best members on a metric and then evaluating that metric on the same data is expected to produce positive improvements over the All-40 mean even under pure noise; the magnitude depends on member spread and sample size, not on predictive skill. The abstract's phrasing 'enhanced long-term prediction' and Section 3.4's operational implications therefore overstate what the analysis can show. A proper test requires an out-of-sample protocol, e.g. ranking members on 1986–2001 and scoring on 2002–2017, rolling-origin evaluation, or at least a leave-decade-out scheme, plus a null distribution fro","section":"§2.4, §2.3, §3.1"},{"comment":"The stated statistical tests do not address the selection bias. Resampling the same verification years after the Top-5 lists have been selected measures sampling noise in the fitted deltas, not whether the selected members would remain top-ranked on independent data. In addition, the paired t-test is described as comparing 'year-to-year differences' in correlation, but a per-year correlation is not well defined for a single year. The 95% bootstrap CIs and the reported percentages of significant lead–season pairs are therefore not evidence of robustness against the central circularity.","section":"§2.9"},{"comment":"The ensemble is described only as 'a new, high-resolution, state-of-the-art prediction model' with 40 members, but the model, its institution or name, and the exact forecast generation procedure are not given. The title refers to a 'CNN Ensemble', yet no CNN architecture, training data, or inference setup appears anywhere in the method. This makes it impossible to assess whether the reported behavior is system-specific, and it does not support the universal claim in the abstract, 'for any large enough ensemble of ENSO forecasts, there is a subset of members whose skill is substantially higher than that of the ensemble mean.' That claim is an overgeneralization from one unnamed system with a single 32-year verification record.","section":"§2.2, Title, Abstract"},{"comment":"The sensitivity analysis that motivates the Top-10 configuration is reported only qualitatively. No table or figure shows how skill and interannual variability vary from Top-3 to Top-20; the reader cannot verify that Top-10 is an optimum rather than a selected result. Since the same verification data are used both to choose the subset size and to evaluate it, the 'optimal' size is also at risk of in-sample overfitting. The quantitative support for this step is missing.","section":"§2.5"}],"minor_comments":[{"comment":"The sign convention for ΔRMSE is inconsistent. Section 2.7 defines ΔRMSE = RMSE_All40 − RMSE_Top5, so positive means improvement, while Section 3.1 and Figure 1 consistently describe negative ΔRMSE as improvement and report values around −0.18°C. Please harmonize the definition and the figures.","section":"§2.7 vs §3.1, Figure 1"},{"comment":"Several presentation issues: 'GGS' appears to be a typo (likely JAS or another season) in Sections 1 and 2.5; 'easonal' in the Figure 7 caption; 'degradements' in Section 2.5; inconsistent hyphenation and spacing in 'ocean atmosphere', 'a-posteriori', 'Top 10' vs 'Top-10'. A thorough language edit is needed.","section":"Throughout"},{"comment":"The reference list is inconsistently formatted and contains irrelevant entries (e.g. [31] on financial literacy) and duplicates ([18] and [20] are the same paper). The in-text citation numbering should also be checked.","section":"References"},{"comment":"No data availability or code availability statement is included. Given that the central result depends on a specific 40-member ENSO forecast ensemble, the dataset should be identified or the experiment made reproducible.","section":"Reproducibility"},{"comment":"The verification sample is described only as 1986–2017; the exact number of forecast targets per lead–season pair and how initialization months map to lead times should be stated explicitly. The statement that 'all ensemble members are verified against the same verification years' is reassuring but needs a table of sample sizes.","section":"§2.2"}],"recommendation":"reject","confidential_remarks":"For the editor: I agree with the reader that the in-sample selection issue is decisive. The paper is well-written in places and the authors are explicit about the retrospective nature of the selection, but the central claim of enhanced prediction skill is unsupported by the analysis. A resubmission that adds a genuine out-of-sample or rolling-origin evaluation, identifies the forecast system, and provides a null/random-subset baseline could be considered, but the current version does not meet the standard for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a well-written, transparent account of a known effect—if you rank ensemble members on the same verification data you later score them against, the top-ranked members look better than the ensemble mean. The authors know this. They say so in Sections 2.1 and the conclusion. But the abstract and Section 3.4 push a stronger claim, and that claim is not supported.\n\nWhat's actually new: a detailed lead-by-lead and season-by-season quantification for one 40-member ENSO system, plus a sensitivity analysis over subset size. That is legitimate descriptive material. The paper also stays honest about the limits most of the time. The paired t-tests and bootstrap are fine for sampling noise, though they don't address selection bias.\n\nThe soft spots: the biggest is the leap from 'there is a subset that did well in hindsight' to 'for any large enough ensemble there is a subset whose skill is substantially higher.' That overgeneralizes from one unnamed system. The load-bearing missing piece is any check that member ranks are stable across independent periods. Without that, all the reported deltas are in-sample selection statistics. The paper acknowledges this, then undercuts its own caveat by talking about 'enhanced long-term prediction' and 'operational benefit' in the abstract and Section 3.4. Also, the reference list includes an irrelevant financial literacy paper [31], which is sloppy.\n\nBottom line: as a descriptive case study, the paper is fine but not new. As a demonstration of enhanced predictive skill, it fails. The honest, useful version of this paper would drop the general claim and the operational framing. A referee could tell the authors that. I'd send it to review because the methodology is transparent and the flaw is instructive, but my recommendation would be reject-as-is. For a reading group, it's worth discussing as a cautionary example of in-sample selection. I wouldn't cite it in my own work.","headline":"A transparent, well-written account of the best-member hindsight effect; the abstract overclaims, but the paper itself mostly owns its retrospective nature.","tokens_in":15205,"tokens_out":2394,"would_cite":false,"duration_ms":25868,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["92.60.Ry"],"model":"deepseek-v4-flash","headline":"Any large ENSO ensemble contains a high-skill subset that beats the ensemble mean, and its advantage grows with lead time.","keywords":["ENSO forecasting","ensemble member selection","a-posteriori evaluation","Niño 3.4 index","long-lead prediction","RMSE","Pearson correlation","deep learning ensemble"],"falsifier":"Rank members on a training portion of the 1986-2017 record, then score the selected Top-5/Top-10 subset on a held-out portion (or with leave-one-out cross-validation). If the selected subset's correlation and RMSE advantage over the All-40 mean disappears or reverses out-of-sample, the claim fails. As a second check, draw many random 5-member subsets: if their average gain matches the Top-5 gain, the effect is selection noise rather than skill signal.","tokens_in":14338,"feed_emoji":"🌊","tokens_out":9284,"duration_ms":85216,"temperature":0.7,"pith_summary":"The paper tries to establish that in any large enough ensemble of ENSO forecasts, a subset of members outperforms the equal-weight ensemble mean, and that the gap grows with forecast lead. On a 40-member system verified against the 1986-2017 Niño-3.4 record, ranking members separately by RMSE and Pearson correlation and averaging the top five of each list lifts skill at every lead: at 1 month correlation gains about +0.02 (+1.7%) and RMSE drops 0.14 °C (23.3%), while at 23 months correlation rises about +0.43 (+172%) and RMSE falls 0.18 °C (22.5%). The largest correlation gains occur in SON and DJF, and the largest RMSE cuts in JJA and MJJ. The authors are explicit that this is a retrospective proof-of-concept: the same verification record is used to pick and to score the members, so the numbers are upper bounds until real-time member-skill estimators exist.","feed_headline":"Hindsight picks 5 of 40 ENSO members; long-lead correlation up 0.43","feed_subtitle":"Selecting members on past RMSE and correlation raises correlation 172% and cuts error 22.5%—but the picks use the verifying record.","key_machinery":"Rank-censored subset selection. For each lead (1-23 months) and target season, the 40 members are ranked by RMSE and by Pearson correlation against the verifying observations; the Top-5 of each list are taken, and for time-series analyses their union forms a weighted Top-10 mean in which members appearing on both lists receive double weight. This is what isolates the high-skill members in hindsight and produces the ΔCorrelation and ΔRMSE comparisons against the All-40 equal-weight mean.","core_discovery":"Using a strictly a-posteriori protocol, the paper finds that for every lead-season combination in a 40-member ENSO forecast ensemble, the five members ranked highest on RMSE and the five ranked highest on correlation each beat the All-40 equal-weight mean on their own metric, and the advantage amplifies with lead time. At one-month lead the Top-5 correlation gain is about +0.02 and RMSE reduction 0.14 °C; at 23-month lead the correlation gain reaches about +0.43 and RMSE reduction 0.18 °C. The authors interpret these deltas as evidence that equal weighting dilutes the best members exactly where skill is scarcest, and they frame the result as a demonstration of the existence of high-skill sub","pith_inferences":["The reported gains are upper bounds: the same 1986-2017 record is used both to choose the top members and to score them, so a fair operational test would select on a training period and verify on a later one.","A direct robustness check would compare the Top-5 gains with the distribution of gains from all random 5-member subsets; if random subsets show comparable gains, the 'high-skill subset' effect is largely selection noise.","The paper's Top-10 optimum was identified on the same verification record used for scoring, so its stability across leads and seasons is itself an in-sample result that needs out-of-sample confirmation.","The method's promise depends on a real-time error estimator; if member skill can be predicted from forecast properties or past performance windows, the retrospective deltas would become an operational weighting scheme."],"forward_implications":["Equal-weight ensemble means are not the best use of a large ensemble: some subset of members is consistently better, and the deficit of the all-member mean grows with lead time.","Selective member weighting can buy substantial long-lead skill—about +0.43 correlation and -0.18 °C RMSE at 23 months—without retraining models or adding computation.","A merged Top-10 (union of Top-5 by RMSE and Top-5 by correlation) balances amplitude fidelity and phase accuracy and is the most stable subset size in the paper's sensitivity analysis.","Because the selection is metric-based and model-independent, the same retrospective protocol can be applied to other large ensembles and other climate phenomena.","The paper's paired t-tests and bootstrap tests report that over 85% of lead-season pairs at leads beyond 12 months have significant correlation gains at the 95% level."],"supporting_citations":[{"why":"Supplies the ERSST Niño-3.4 observations used to verify, rank, and score all ensemble members.","marker":"[27]"},{"why":"Anchors the long-lead ENSO forecasting challenge and the multi-year forecast skill the present method seeks to improve.","marker":"[18]"},{"why":"Provides the coupled-model extended ENSO prediction context and Niño-3.4 verification conventions underlying the ensemble setup.","marker":"[11]"},{"why":"Reviews ENSO predictability and prediction progress, framing why long-lead skill gains matter.","marker":"[21]"},{"why":"Documents real-time seasonal ENSO model prediction skill that the paper's equal-weight ensemble-mean baseline reflects.","marker":"[17]"}],"fun_headline_variants":["ENSO forecasts: picking 5 best members in hindsight doubles correlation at 23-month lead","Backtested member selection lifts long-lead ENSO correlation by 172%","Hindsight ENSO ensemble: top 5 members beat mean at every lead, most at 23 months","A-posteriori ENSO member picks: +0.43 correlation, 22.5% RMSE cut at extreme lead","ENSO skill from hindsight: top 5 members cut RMSE 22.5% at 23-month lead"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The same 1986-2017 Niño-3.4 record is used both to choose the top members and to score the chosen subset, so the demonstrated gains are retrospective; the central claim transfers to real prediction only if past member skill over that record is a reliable guide to future member skill.","fun_headline_variants_meta":{"raw":{"variants":["ENSO forecasts: picking 5 best members in hindsight doubles correlation at 23-month lead","Backtested member selection lifts long-lead ENSO correlation by 172%","Hindsight ENSO ensemble: top 5 members beat mean at every lead, most at 23 months","A-posteriori ENSO member picks: +0.43 correlation, 22.5% RMSE cut at extreme lead","ENSO skill from hindsight: top 5 members cut RMSE 22.5% at 23-month lead"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000636,"raw_usage":{"total_tokens":2851,"prompt_tokens":905,"completion_tokens":1946,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":649,"completion_tokens_details":{"reasoning_tokens":1830}},"tokens_in":649,"tokens_out":1946,"duration_ms":16554,"temperature":1.0,"reasoning_tokens":1830,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T23:51:20.068679+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rank members on a training portion of the 1986-2017 record, then score the selected Top-5/Top-10 subset on a held-out portion (or with leave-one-out cross-validation). If the selected subset's correlation and RMSE advantage over the All-40 mean disappears or reverses out-of-sample, the claim fails. As a second check, draw many random 5-member subsets: if their average gain matches the Top-5 gain, the effect is selection noise rather than skill signal.","supporting_citations":[{"cited_title":"Decadal Changes in ENSO Predictability and Their Implications for Climate Forecasting,","cited_arxiv_id":null,"evidence_quote":"Supplies the ERSST Niño-3.4 observations used to verify, rank, and score all ensemble members."},{"cited_title":"Deep learning for multi-year ENSO forecasts,","cited_arxiv_id":null,"evidence_quote":"Anchors the long-lead ENSO forecasting challenge and the multi-year forecast skill the present method seeks to improve."},{"cited_title":"Extended ENSO predictions using a fully coupled ocean–atmosphere model,","cited_arxiv_id":null,"evidence_quote":"Provides the coupled-model extended ENSO prediction context and Niño-3.4 verification conventions underlying the ensemble setup."},{"cited_title":"Progress in ENSO prediction and understanding during the past decades,","cited_arxiv_id":null,"evidence_quote":"Reviews ENSO predictability and prediction progress, framing why long-lead skill gains matter."},{"cited_title":"Skill of real -time seasonal ENSO model predictions during 2002–2011,","cited_arxiv_id":null,"evidence_quote":"Documents real-time seasonal ENSO model prediction skill that the paper's equal-weight ensemble-mean baseline reflects."}],"review_version":1}