REVIEW 4 major objections 5 minor 32 references
Distillation of CNN Ensemble Results for Enhanced Long-Term Prediction of the ENSO Phenomenon
T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read Any large ENSO ensemble contains a high-skill subset that beats the ensemble mean, and its advantage grows with lead time.
desk verdict A transparent, well-written account of the best-member hindsight effect; the abstract overclaims, but the paper itself mostly owns its retrospective nature. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Rank-censored subset selection. For each lead (1-23 months) and target season, the 40 members are ranked by RMSE and by Pearson correlation against the verifying observations; the Top-5 of each list are taken, and for time-series analyses their union forms a weighted Top-10 mean in which members appearing on both lists receive double weight. This is what isolates the high-skill members in hindsight and produces the ΔCorrelation and ΔRMSE comparisons against the All-40 equal-weight mean.
What would settle it
Rank members on a training portion of the 1986-2017 record, then score the selected Top-5/Top-10 subset on a held-out portion (or with leave-one-out cross-validation). If the selected subset's correlation and RMSE advantage over the All-40 mean disappears or reverses out-of-sample, the claim fails. As a second check, draw many random 5-member subsets: if their average gain matches the Top-5 gain, the effect is selection noise rather than skill signal.
Extended reading notes
Core claim
Using a strictly a-posteriori protocol, the paper finds that for every lead-season combination in a 40-member ENSO forecast ensemble, the five members ranked highest on RMSE and the five ranked highest on correlation each beat the All-40 equal-weight mean on their own metric, and the advantage amplifies with lead time. At one-month lead the Top-5 correlation gain is about +0.02 and RMSE reduction 0.14 °C; at 23-month lead the correlation gain reaches about +0.43 and RMSE reduction 0.18 °C. The authors interpret these deltas as evidence that equal weighting dilutes the best members exactly where skill is scarcest, and they frame the result as a demonstration of the existence of high-skill sub
Load-bearing premise
The same 1986-2017 Niño-3.4 record is used both to choose the top members and to score the chosen subset, so the demonstrated gains are retrospective; the central claim transfers to real prediction only if past member skill over that record is a reliable guide to future member skill.
Editorial extensions
If this is right
- Equal-weight ensemble means are not the best use of a large ensemble: some subset of members is consistently better, and the deficit of the all-member mean grows with lead time.
- Selective member weighting can buy substantial long-lead skill—about +0.43 correlation and -0.18 °C RMSE at 23 months—without retraining models or adding computation.
- A merged Top-10 (union of Top-5 by RMSE and Top-5 by correlation) balances amplitude fidelity and phase accuracy and is the most stable subset size in the paper's sensitivity analysis.
- Because the selection is metric-based and model-independent, the same retrospective protocol can be applied to other large ensembles and other climate phenomena.
- The paper's paired t-tests and bootstrap tests report that over 85% of lead-season pairs at leads beyond 12 months have significant correlation gains at the 95% level.
Reading between the lines
- The reported gains are upper bounds: the same 1986-2017 record is used both to choose the top members and to score them, so a fair operational test would select on a training period and verify on a later one.
- A direct robustness check would compare the Top-5 gains with the distribution of gains from all random 5-member subsets; if random subsets show comparable gains, the 'high-skill subset' effect is largely selection noise.
- The paper's Top-10 optimum was identified on the same verification record used for scoring, so its stability across leads and seasons is itself an in-sample result that needs out-of-sample confirmation.
- The method's promise depends on a real-time error estimator; if member skill can be predicted from forecast properties or past performance windows, the retrospective deltas would become an operational weighting scheme.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper uses a 40-member ensemble of Niño-3.4 forecasts at leads 1–23 months over 1986–2017, ranks members per lead–season pair against ERSST observations by RMSE and Pearson correlation, forms Top-5-by-RMSE, Top-5-by-correlation, and a weighted Top-10 union, and compares these selected forecasts with the equal-weight All-40 mean. The headline results are large gains at long leads, e.g. +0.43 correlation and −0.18°C RMSE at 23 months, with season-dependent patterns. The authors explicitly frame the exercise as retrospective/a-posteriori and disclaim operational use, but the abstract and Section 3.4 present the results as evidence for enhanced long-lead ENSO prediction and as a foundation for operational selection.
Significance. The question of whether selective weighting of ensemble members can improve long-lead ENSO prediction is practically relevant, and the paper clearly exposes the mechanics of ranking by RMSE and correlation, including a subset-size sensitivity analysis and a caveat that the method is retrospective. However, the central empirical claim is not supported: the top members are selected from the same 1986–2017 verification record on which they are then scored, so the reported improvements are in-sample selection statistics. A Top-5 subset chosen by minimizing RMSE or maximizing correlation on a record will almost always beat the full ensemble mean on that same record, even if the members contain no useful forecast signal. The bootstrap and paired t-tests in Section 2.9 do not repair this bias. The study therefore provides an example of retrospective optimization, not evidence of predictive skill. The forecast system is also unnamed, and no code or data are provided.
major comments (4)
- [§2.4, §2.3, §3.1] The Top-5-by-RMSE and Top-5-by-Correlation lists are formed for each (lead, season) by scoring every member against the observed Niño-3.4 series y_t, and the same y_t is then used in Section 3.1 to compute the reported ΔCorrelation and ΔRMSE. This is a direct in-sample selection loop. Selecting the five best members on a metric and then evaluating that metric on the same data is expected to produce positive improvements over the All-40 mean even under pure noise; the magnitude depends on member spread and sample size, not on predictive skill. The abstract's phrasing 'enhanced long-term prediction' and Section 3.4's operational implications therefore overstate what the analysis can show. A proper test requires an out-of-sample protocol, e.g. ranking members on 1986–2001 and scoring on 2002–2017, rolling-origin evaluation, or at least a leave-decade-out scheme, plus a null distribution fro
- [§2.9] The stated statistical tests do not address the selection bias. Resampling the same verification years after the Top-5 lists have been selected measures sampling noise in the fitted deltas, not whether the selected members would remain top-ranked on independent data. In addition, the paired t-test is described as comparing 'year-to-year differences' in correlation, but a per-year correlation is not well defined for a single year. The 95% bootstrap CIs and the reported percentages of significant lead–season pairs are therefore not evidence of robustness against the central circularity.
- [§2.2, Title, Abstract] The ensemble is described only as 'a new, high-resolution, state-of-the-art prediction model' with 40 members, but the model, its institution or name, and the exact forecast generation procedure are not given. The title refers to a 'CNN Ensemble', yet no CNN architecture, training data, or inference setup appears anywhere in the method. This makes it impossible to assess whether the reported behavior is system-specific, and it does not support the universal claim in the abstract, 'for any large enough ensemble of ENSO forecasts, there is a subset of members whose skill is substantially higher than that of the ensemble mean.' That claim is an overgeneralization from one unnamed system with a single 32-year verification record.
- [§2.5] The sensitivity analysis that motivates the Top-10 configuration is reported only qualitatively. No table or figure shows how skill and interannual variability vary from Top-3 to Top-20; the reader cannot verify that Top-10 is an optimum rather than a selected result. Since the same verification data are used both to choose the subset size and to evaluate it, the 'optimal' size is also at risk of in-sample overfitting. The quantitative support for this step is missing.
minor comments (5)
- [§2.7 vs §3.1, Figure 1] The sign convention for ΔRMSE is inconsistent. Section 2.7 defines ΔRMSE = RMSE_All40 − RMSE_Top5, so positive means improvement, while Section 3.1 and Figure 1 consistently describe negative ΔRMSE as improvement and report values around −0.18°C. Please harmonize the definition and the figures.
- [Throughout] Several presentation issues: 'GGS' appears to be a typo (likely JAS or another season) in Sections 1 and 2.5; 'easonal' in the Figure 7 caption; 'degradements' in Section 2.5; inconsistent hyphenation and spacing in 'ocean atmosphere', 'a-posteriori', 'Top 10' vs 'Top-10'. A thorough language edit is needed.
- [References] The reference list is inconsistently formatted and contains irrelevant entries (e.g. [31] on financial literacy) and duplicates ([18] and [20] are the same paper). The in-text citation numbering should also be checked.
- [Reproducibility] No data availability or code availability statement is included. Given that the central result depends on a specific 40-member ENSO forecast ensemble, the dataset should be identified or the experiment made reproducible.
- [§2.2] The verification sample is described only as 1986–2017; the exact number of forecast targets per lead–season pair and how initialization months map to lead times should be stated explicitly. The statement that 'all ensemble members are verified against the same verification years' is reassuring but needs a table of sample sizes.
Circularity Check
The Top-5/Top-10 'skill gains' are in-sample selection statistics: members are ranked by RMSE/correlation on the same 1986–2017 observations used to score them, so the reported improvements are forced by the selection rule, not predictive evidence.
-
self definitional
[Section 2.4 (Formation of Top Lists) and Section 2.7 (Evaluation Metrics); results in Section 3.1]
"For each (𝑳, 𝑴), the Top-5-by-RMSE subset consisted of the five members with the lowest RMSE, while the Top-5-by-Correlation subset included the five members with the highest correlation. ... 𝚫𝑹𝑴𝑺𝑬𝑳,𝑴 = 𝑹𝑴𝑺𝑬𝑳,𝑴^𝑨𝒍𝒍𝟒𝟎 − 𝑹𝑴𝑺𝑬𝑳,𝑴^𝑻𝒐𝒑𝟓_𝑹𝑴𝑺𝑬; 𝚫𝒓𝑳,𝑴 = 𝒓𝑳,𝑴^𝑨𝒍𝒍𝟒𝟎 − 𝒓𝑳,𝑴^𝑻𝒐𝒑𝟓_𝑪𝒐𝒓𝒓"
The members are chosen by minimizing RMSE / maximizing correlation over the exact 1986–2017 verification years (the same y_t used in Section 2.3), and then ΔRMSE and ΔCorrelation are computed over those same years. By construction, the five lowest-RMSE members have an average RMSE no larger than the 40-member average, so a negative ΔRMSE is guaranteed; the magnitude and the correlation deltas are in-sample selection statistics. No independent forecast period is used to test whether the selected members remain superior, so the 'enhanced prediction skill' claim is an artifact of scoring the fitting target.
-
fitted input called prediction
[Section 2.9 (Statistical Testing and Uncertainty Quantification)]
"For each lead season pair, a 10,000-member non-parametric bootstrap resampling was carried out with the verification years resampled with replacement. Confidence intervals (CIs) for ΔCorrelation and ΔRMSE were derived from the bootstrap distributions. The improvements were considered to be robust if the 95% CI did not cross zero."
The bootstrap resamples the same verification years that were already used to rank and select the Top-5 members. It therefore quantifies sampling noise of the already-fitted deltas, not the stability of member ranking or the transferability of the selection to unseen years. Reporting 'statistically significant' improvements on this basis gives the appearance of validation while the selection and evaluation data are identical.
full rationale
The paper is transparent that its design is a-posteriori (Section 2.1: 'the current application is not an operational selection scheme'; Conclusion: 'the method is by its nature retrospective in application'), and it does not rely on a load-bearing self-citation chain. However, the central quantitative claim—that Top-5/Top-10 subsets have 'substantially higher skill' with ΔCorrelation up to +0.43 and ΔRMSE −0.18°C—is produced by ranking members on the same 1986–2017 Nino3.4 observations that are then used to score the subsets. Section 2.4 defines the Top-5 lists by lowest RMSE/highest correlation over those years, and Section 2.7 computes the improvements over the same years. The RMSE improvement is mathematically compelled by the selection, and the correlation improvement is an in-sample maximum of a ranking statistic. The bootstrap in Section 2.9 resamples the same verification years and cannot convert this into an out-of-sample estimate. The abstract's generalization 'for any large enough ensemble' is not supported by the single-system, single-record design. These issues make the headline skill gains reduction-by-construction, even though the paper honestly disclaims operational readiness.
Assumptions & free parameters
free parameters (3)
- Top-5 subset size per metric =
5 members
- Top-10 combined subset size =
10 members, union of two Top-5 lists
- Double weight for overlap members =
2
assumptions (4)
- domain assumption NOAA ERSST Nino3.4 index is the correct observational reference for ENSO verification
- domain assumption The 40-member ensemble is a representative and sufficiently large sample of forecast uncertainty
- domain assumption RMSE and Pearson correlation are complementary and sufficient skill metrics for member selection
- ad hoc to paper Member skill differences over 1986-2017 are systematic enough to transfer to future forecasts
Cite this review
Pith. "Pith review of Distillation of CNN Ensemble Results for Enhanced Long-Term Prediction of the ENSO Phenomenon." pith.science (2026). https://pith.science/paper/3IG7EDCO
@misc{pith2026250906227,
author = {Pith},
title = {Pith review of: Distillation of CNN Ensemble Results for Enhanced Long-Term Prediction of the ENSO Phenomenon},
year = {2026},
howpublished = {\url{https://pith.science/paper/3IG7EDCO}},
note = {Machine review of arXiv:2509.06227}
}
read the original abstract
The accurate long-term forecasting of the El Nino Southern Oscillation (ENSO) is still one of the biggest challenges in climate science. While it is true that short-to medium-range performance has been improved significantly using the advances in deep learning, statistical dynamical hybrids, most operational systems still use the simple mean of all ensemble members, implicitly assuming equal skill across members. In this study, we demonstrate, through a strictly a-posteriori evaluation , for any large enough ensemble of ENSO forecasts, there is a subset of members whose skill is substantially higher than that of the ensemble mean. Using a state-of-the-art ENSO forecast system cross-validated against the 1986-2017 observed Nino3.4 index, we identify two Top-5 subsets one ranked on lowest Root Mean Square Error (RMSE) and another on highest Pearson correlation. Generally across all leads, these outstanding members show higher correlation and lower RMSE, with the advantage rising enormously with lead time. Whereas at short leads (1 month) raises the mean correlation by about +0.02 (+1.7%) and lowers the RMSE by around 0.14 {\deg}C or by 23.3% compared to the All-40 mean, at extreme leads (23 months) the correlation is raised by +0.43 (+172%) and RMSE by 0.18 {\deg}C or by 22.5% decrease. The enhancements are largest during crucial ENSO transition periods such as SON and DJF, when accurate amplitude and phase forecasting is of greatest socio-economic benefit, and furthermore season-dependent e.g., mid-year months such as JJA and MJJ have incredibly large RMSE reductions. This study provides a solid foundation for further investigations to identify reliable clues for detecting high-quality ensemble members, thereby enhancing forecasting skill.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[31]
Moazezi Khah Tehran, A., Hassani, A., Mohajer, S., Darvishan, S., Shafiesabet, A., & Tashakkori, A. (2025). The Impact of Financial Literacy on Financial Behavior and Financial Resilience with the Mediating Role of Financial Self -Efficacy. International Jo urnal of Industrial Engineering and Operational Research, 7(2), 38 -55. https://doi.org/10.22034/ij...
-
[1]
Combined dynamical –deep learning ENSO forecasts,
Yipeng Chen, Yishuai Jin, Zhengyu Liu, Xingchen Shen, Xianyao Chen, Xiaopei Lin, Rong - Hua Zhang, Jing-Jia Luo, Wansuo Zhang, Fei Duan, Zhiming Ma, Jieming Ma, and Lu Zhou,, "Combined dynamical –deep learning ENSO forecasts," Nature Communications, p. Article number: 3845, 2025
work page 2025
-
[2]
Toward long -range ENSO prediction with an explainable deep learning model,
Qi Chen, Yinghao Cui, Guobin Hong, Karumuri Ashok, Yuchun Pu, Xiaogu Zheng, Xuanze Zhang, Wei Zhong, Peng Zhan, and Zhonglei Wang, "Toward long -range ENSO prediction with an explainable deep learning model," npj Climate and Atmospheric Science, vol. 8, p. Article number: 259, 2025
work page 2025
-
[3]
Entropic learning enables skilful forecasts of ENSO phase at up to two years lead time,
Michael Groom, Davide Bassetti, Illia Horenko, and Terence J. O’Kane,, "Entropic learning enables skilful forecasts of ENSO phase at up to two years lead time," 2025
work page 2025
-
[4]
A Hybrid Deep-Learning Model for El Niño Southern Oscillation in the Low-Data Regime,
Jakob Schlör, Michael Newman, Johann Thuemmel, Antonietta Capotondi, and Bharat Goswami,, "A Hybrid Deep-Learning Model for El Niño Southern Oscillation in the Low-Data Regime," 2024
work page 2024
-
[5]
Novel Insights in Deep Learning for Predicting Climate Phenomena,
M. Naisipour, S. Ganji, I. Saeedpanah, B. Mehrakizadeh and A. Adib, "Novel Insights in Deep Learning for Predicting Climate Phenomena," in 14th International Conference on Computer and Knowledge Engineering (ICCKE), 2024
work page 2024
-
[6]
M. Naisipour, M. H. Afshar, B. Hassani and M. Zeinali, " An error indicator for two - dimensional elasticity problems in the discrete least squares meshless method," in 8th International Congress on Civil Engineering, 2009
work page 2009
-
[7]
Multimodal Deep Learning for Two -Year ENSO Forecast,
M. Naisipour, I. Saeedpanah and A. Adib, "Multimodal Deep Learning for Two -Year ENSO Forecast," Water Resources Management, p. 3745–3775, 2025
work page 2025
Show all 32 references
-
[8]
Novel Deep Learning Method for Forecasting ENSO,
M. Naisipour, I. Saeedpanah and A. Adib, "Novel Deep Learning Method for Forecasting ENSO," Journal of Hydraulic Structures, pp. 14-25, 2025
2025
-
[9]
Forecasting El Niño Six Months in Advance Utilizing Augmented Convolutional Neural Network,
M. Naisipour, I. Saeedpanah, A. Adib and M. H. Neisi Pour, "Forecasting El Niño Six Months in Advance Utilizing Augmented Convolutional Neural Network," in 14th International Conference on Computer and Knowledge Engineering , 2024
2024
-
[10]
Advances in analysis and prediction of ENSO using deep learning,
B. Wang, J.-Y. Lee, J.-J. Luo and e. al, "Advances in analysis and prediction of ENSO using deep learning," Climate Dynamics, p. 1845–1867, 2023
2023
-
[11]
Extended ENSO predictions using a fully coupled ocean–atmosphere model,
J.-J. Luo, S. Masson, S. Behera and T. Yamagata, "Extended ENSO predictions using a fully coupled ocean–atmosphere model," Journal of Climate, p. 84–93, 2008
2008
-
[12]
ENSO as an integrating concept in Earth science,
M. J. McPhaden, S. E. Zebiak and M. H. Glantz, "ENSO as an integrating concept in Earth science," Science, p. 1740–1745, 2006
2006
-
[13]
Predicting El Niño beyond 1 -year lead: Robust statistics of the coupled model intercomparison,
J.-H. Park, S. -W. Yeh, J. -S. Kug and e. al, "Predicting El Niño beyond 1 -year lead: Robust statistics of the coupled model intercomparison," Climate Dynamics, p. 3689–3705, 2018
2018
-
[14]
Atmospheric circulation as a source of uncertainty in climate change projections,
T. G. Shepherd, "Atmospheric circulation as a source of uncertainty in climate change projections," Nature Geoscience, p. 703–708, 2014
2014
-
[15]
Causes of the 2010–2012 La Niña cooling,
Y. Gao and X. Zhang, "Causes of the 2010–2012 La Niña cooling," Climate Dynamics, p. 1861– 1873, 2017
2010
-
[16]
Atlantic Niño’s influence on ENSO,
C. Wang and e. al, "Atlantic Niño’s influence on ENSO," Nature Communications, p. 6791, 2023
2023
-
[17]
Skill of real -time seasonal ENSO model predictions during 2002–2011,
A. G. Barnston, M. K. Tippett, M. L. L’Heureux, S. Li and D. G. DeWitt, "Skill of real -time seasonal ENSO model predictions during 2002–2011," Climate Dynamics, p. 593–614, 2012
2002
-
[18]
Deep learning for multi-year ENSO forecasts,
Y. G. Ham, J. H. Kim and J. J. Luo, "Deep learning for multi-year ENSO forecasts," Nature, p. 224–228, 2019
2019
-
[19]
Efficiency test of the discrete least squares meshless method in solving heat conduction problems using error estimation,
M. Labibzadeh, R. Modaresi Rajab and M. Naisipour, "Efficiency test of the discrete least squares meshless method in solving heat conduction problems using error estimation," Sharif: Civil Engineering, no. 3.2, p. 31–40, 2015
2015
-
[20]
Deep learning for multi -year ENSO forecasts,
Y.-G. Ham, J.-H. Kim and J.-J. Luo, "Deep learning for multi -year ENSO forecasts," Nature, p. 224–228, 2019
2019
-
[21]
Progress in ENSO prediction and understanding during the past decades,
Y. Tang, W. Zhang, D. Chen and e. al, "Progress in ENSO prediction and understanding during the past decades," Nature Communications, p. 1–15, 2018
2018
-
[22]
Long-lead seasonal prediction of ENSO events using a deep learning model,
Y. Guo, H.-L. Ren, J.-J. Luo and e. al, "Long-lead seasonal prediction of ENSO events using a deep learning model," Climate Dynamics, p. 1489–1506, 2021
2021
-
[23]
Advances in Seasonal to Interannual Applications: Toward Enhanced NOAA Forecast Capabilities,
Y. Xue, "Advances in Seasonal to Interannual Applications: Toward Enhanced NOAA Forecast Capabilities," Bulletin of the American Meteorological Society, 2025
2025
-
[24]
Advances in Seasonal- to-Interannual Climate Prediction of ENSO: Current Status and Future Directions,
W. Duan, W. Zhang, X. Chen, F. Zheng, R.-H. Zhang and McPhaden, "Advances in Seasonal- to-Interannual Climate Prediction of ENSO: Current Status and Future Directions," Bulletin of the American Meteorological Society, p. E505–E529, 2025
2025
-
[25]
Hybrid Statistical –Dynamical Models for Skillful ENSO Forecasts at Long Leads,
B. Xiang, M. K. Tippett, A. G. Barnston, S. Li and B. Mu, "Hybrid Statistical –Dynamical Models for Skillful ENSO Forecasts at Long Leads," vol. 37, p. 1231–1249, 2024
2024
-
[26]
Improved Long-Lead ENSO Forecasting through Multi-Model Deep Learning Frameworks,
Y.-G. Ham, J. Kim, J.-J. Luo, J. Park and H.-K. Kim, "Improved Long-Lead ENSO Forecasting through Multi-Model Deep Learning Frameworks," vol. 52, 2025
2025
-
[27]
Decadal Changes in ENSO Predictability and Their Implications for Climate Forecasting,
L. Wang, R. -H. Zhang, D. Chen and M. J. McPhaden, "Decadal Changes in ENSO Predictability and Their Implications for Climate Forecasting," vol. 14, p. 102–110, 2024
2024
-
[28]
Advances in AI-Driven Multi-Year ENSO Prediction Using Global Climate Models,
J. Park, Y.-G. Ham, J. Kim, J.-J. Luo and S.-K. Lee, "Advances in AI-Driven Multi-Year ENSO Prediction Using Global Climate Models," vol. 62, p. 1523–1542, 2025
2025
-
[29]
A Transformer -Based Deep Learning Model for Subseasonal-to-Seasonal ENSO Forecasting,
J. Cao, G. Li, C. Fang, J. -J. Luo and W. Zhang, "A Transformer -Based Deep Learning Model for Subseasonal-to-Seasonal ENSO Forecasting," vol. 14, p. Article number: 5123, 2024
2024
-
[30]
Enhancing Long -Lead ENSO Forecast Skill via Coupled Data Assimilation and Deep Learning,
W. Zhang, W. Duan, J. -J. Luo, M. J. McPhaden and S. Li, "Enhancing Long -Lead ENSO Forecast Skill via Coupled Data Assimilation and Deep Learning," vol. 16, p. Article number: 4123, 2025
2025
-
[32]
Saghar Ganji, Ahmad Reza Labibzadeh, Alireza Hassani, Mohammad Naisipour, Leveraging GNN to Enhance MEF Method in Predicting ENSO, 2025, arXiv:2508.07410v3 [physics.ao- ph]
2025 arXiv
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.