Pith. sign in

REVIEW 3 major objections 6 minor 63 references

Paleoclimate Boundary Conditions as an Out-of-Sample Test for the Forced Response of Ocean Climate Emulators

T0 review · 3 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read An ocean emulator trained only on the control climate can reproduce the large-scale upper-ocean response to a 6,000-year-old orbital forcing, but the response is too weak and the slow deep-ocean changes are missed.

desk verdict A careful first paleoclimate testbed for ocean emulators; the main claims hold up, but the equilibration of the midHolocene reference is asserted rather than proven. read the letter →

arxiv 2608.13494 v1 pith:ZTGPOLD3 submitted 2026-08-13 physics.ao-ph

classification physics.ao-ph
keywords oceanclimateemulatormidHoloceneout-of-samplegeneralizationforcedresponseCESM2orbitalforcinglinearsuperpositionautoregressivemodel
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper's aim is to give long-term ocean climate emulators a test that is out-of-sample but still in-distribution, and it argues that the midHolocene experiment supplies exactly that. The central result is that an autoregressive, full-depth ocean emulator trained only on CESM2 piControl reproduces the spatial structure of the upper-ocean potential-temperature response to midHolocene orbital forcing, along with changes in the seasonal cycle and in variability patterns, while recovering only about two-thirds of the true amplitude. Forcing-only baselines recover much of the near-surface pattern but miss the seasonal and variability changes, which shows that some learned internal dynamics are doing real work. The same emulator fails on slow, internally driven evolution of the ocean interior, and its total response is well approximated by a linear superposition of the responses to each boundary forcing. If this holds, paleoclimate experiments offer a cheap, ground-truthed way to catch forced-response failures before emulators are trusted for warming scenarios.

What carries the argument

The load-bearing object is a ConvNEXT-UNet autoregressive emulator that steps the three-dimensional ocean state (potential temperature $\theta_O$ and salinity $S$) forward one month from a single prior state plus surface heat flux, the two wind-stress components, and explicitly computed insolation. The argument is carried by paired ensembles of rollouts: the forced response is defined as the difference between midHolocene-forced and piControl-forced rollouts sharing the same initial conditions, and the total response is then decomposed into single-forcing component responses. The linearity probe, comparing the full response with the sum of component responses, is what establishes the approximate additivity, and a dynamic channel weighting during training is what lets slowly evolving deep-ocean signals contribute to the loss at all.

What would settle it

Recompute the skill scores against a longer, fully equilibrated segment of the CESM2 midHolocene run, or against a multi-member midHolocene ensemble, so the reference is purged of spin-up drift; if the tropical Pacific correlations of 0.64-0.83 and the two-thirds amplitude ratios do not survive that cleaner ground truth, the generalization claim is tied to the chosen reference window.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that a full-depth autoregressive ocean emulator trained exclusively on the CESM2 preindustrial control run generalizes to the midHolocene: when driven by midHolocene boundary forcings it captures the pattern of the upper-1000 m temperature response (tropical Pacific correlations 0.64-0.83), the phase and amplitude changes of the seasonal cycle, and changes in the spatial structure of ocean variability, while underestimating the response amplitude (amplitude ratios about 0.6-0.7). The authors further find that noiseless checkpoints of the emulator that look equivalent on test RMSE can differ substantially in response skill, and that late-training checkpoints lose basin-wide skill in the Atlantic, so standard validation metrics do not certify dynamics. Finally, perturbing individual forcing components produces responses that superimpose nearly exactly onto the full response (correlations near 0.99 in every basin), indicating that the emulator's forcing pathways interact only weakly through the internal state; the authors are careful to note this is a property of the emulator, not demonstrated for the real ocean.

Load-bearing premise

The evaluation uses the CESM2 midHolocene-minus-piControl difference as the equilibrated forced response of the ocean, so if that difference still carries slow spin-up drift or internal variability, the reported correlations and amplitude ratios do not cleanly measure the emulator's forced-response skill.

Editorial extensions

If this is right

  • If the result is correct, the midHolocene becomes a standard intermediate test: an emulator must pass a paleoclimate out-of-sample test before its responses to warming scenarios are taken at face value.
  • Upper-ocean perturbation experiments with emulators become more defensible, since the emulator beats forcing-only baselines in both response amplitude and depth.
  • Decomposing the response by forcing component offers a cheap attribution tool for projected ocean changes, at least within the emulator's linear regime.
  • The late-training collapse in response skill implies that checkpoint selection for climate emulators should include dynamical response tests, not only held-out RMSE.
  • Slow deep-ocean and North Atlantic changes remain outside the emulator's reliability envelope, so claims about interior heat accumulation or overturning changes from such emulators should be treated as unsupported.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One natural extension, not run here, is to apply the same protocol to the Last Interglacial (~127,000 years ago): the paper's account predicts high tropical pattern correlations but a further drop in amplitude ratio as orbital anomalies grow, which would test whether the damping scales predictably.
  • Because single-forcing midHolocene experiments do not exist, the paper cannot say whether the real CESM2 response is additive; generating such runs would settle whether the near-perfect linear superposition is a genuine physical property or a learned simplification of a control climate.
  • The diagnosed failure to accumulate slow interior heat suggests a specific mechanism worth testing: if the ocean emulator is coupled to an atmospheric emulator, some of that deep response should return; retraining with explicit circulation state variables would directly probe that hypothesis.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper evaluates whether an autoregressive, full-depth ocean emulator (a ConvNEXT/UNet architecture with roughly 85 million parameters, trained on CESM2 piControl monthly data) can generalize to midHolocene orbital forcings as an out-of-sample but in-distribution test. The authors compare emulator responses to the CESM2 midHolocene minus piControl difference for upper-ocean potential temperature, seasonal-cycle changes, and variability changes, and they compare against two baselines (a local linear regression operator and a forcing-only network). They report that the emulator reproduces the large-scale spatial structure of the upper-ocean response and changes in seasonality and variability, while underestimating amplitude, and that it fails to capture slow interior evolution. They also show that the emulator's total response is approximately the linear superposition of its responses to individual forcing components, and that test RMSE on the training climate does not predict forced-response skill across epochs and seeds.

Significance. If the central claims hold, the midHolocene setup provides a valuable, controlled, ground-truthed benchmark for diagnosing forced-response failures in AI ocean emulators before they are applied to out-of-distribution climates. The experimental design has real strengths: the emulator is trained only on piControl and never fit to midHolocene data; the evaluation includes two carefully constructed baselines, multiple seeds, full epoch sweeps, and significance stippling on the CESM2 reference; and the code, preprocessing scripts, and model weights are archived on Zenodo. The component-forcing decomposition and the epoch-sensitivity analysis are also useful contributions. The main quantitative claims rest on two assumptions that the paper acknowledges but does not fully verify: that the CESM2 midHolocene reference is equilibrated, and that the sign-flip comparison between the piControl-trained and midHolocene-trained emulators is valid. Because those assumptions affect the headline correlations and amplitude ratios, the paper needs additional evidence or a more conservative presentation before the claims are fully established.

major comments (3)
  1. [Section 3.3] The reference response used as ground truth is the CESM2 midHolocene minus piControl difference, but the manuscript only states 'we expect the system to have equilibrated on the century timescales we investigate' and provides no direct test. Since each CESM2 experiment is a single realization, the difference contains internal variability and any slow drift still present in the upper 1000 m. This is load-bearing for every reported correlation and amplitude ratio in Figures 6-7 and Tables S1-S4: if the reference is not equilibrated, the Atlantic basin-wide correlation of 0.03 and the amplitude ratios do not cleanly measure the emulator's forced response. Please add a quantitative equilibration check, for example comparing early versus late 50-year windows of the overlapping piControl and midHolocene periods, or computing the trend in the upper-1000 m temperature difference over the evaluation window and showing it is small relative to the forced signal. Reporting confidence intervals on the correlations and amplitude ratios computed from the ten 25-year chunks would also help separate forced signal from internal variability.
  2. [Section 3, Section 4.3] The headline metrics are reported from 'a single representative checkpoint per emulator' that is selected after inspecting response skill ('We use an early checkpoint for the piControl emulators, taken before the late-training skill degradation'). Given that Section 4.3 demonstrates substantial epoch-to-epoch and seed-to-seed variability in response skill, and that 'no checkpoint performs best across all regions', this post hoc selection risks inflating the reported numbers and makes the quantitative claims difficult to reproduce without the same selection procedure. Please either pre-specify a checkpoint-selection rule that does not use the out-of-sample response (for example, lowest validation MSE before the degradation epoch, chosen blind to response skill), or report the distribution of response metrics across all epochs and both seeds as the primary summary, with the single-checkpoint values shown only for illustration.
  3. [Section 3.3] The comparison of the midHolocene-trained emulator R(F_mH) against the CESM2 reference uses the sign-flip convention R(F_pi) ↔ -R(F_mH), with the statement that the authors 'cannot verify the exact reversibility or linearity of the applied forcing'. The cross-emulator correlations between R(F_pi) and -R(F_mH) (0.76 in the Pacific, 0.88 in the tropics) are encouraging, but they do not establish that the CESM2 true response is antisymmetric under reversing the forcing perturbation. Since the F_mH correlation values in Figures 6-7 and throughout Section 3.3 are computed after flipping the sign of the true response, an unverified asymmetry directly affects those numbers. Please either treat the F_pi results as the primary quantitative evaluation and present the F_mH results as an internal consistency check, or add a test (e.g., comparing the spatial patterns of the two emulator responses against the flipped true response in a way that does not presuppose antisymmetry, or using single-forcing CESM2 experiments if available) to justify the sign-flipped comparison.
minor comments (6)
  1. [Section 2.1, Eqs. (1)-(3)] Equation (1) defines δ as a 'fraction' of cell overlap, but the expression with min/max and the Heaviside function yields an overlap thickness in meters; please clarify the units and terminology in the text following Eq. (3).
  2. [Section 4.1] The paragraph after Figure 8 contains a confusing contrast: it first says response magnitudes are 'all below 0.05 °C compared to the true maximum 0.16 °C' for 'other component responses', then says R(F_pi^All; τ_mH,I_mH), R(F_pi^All; hfds_mH), and R(F_pi^hfds; hfds_mH) 'reproduce more accurate values of 0.16 °C, 0.12 °C, and 0.13 °C'. Please rewrite to specify exactly which responses have weak maxima and which have accurate maxima, since the current wording appears to contradict itself.
  3. [Section 2.5.3] The reported ensemble spread of at most 6×10^-6 °C across five ensemble members seems implausibly small for a chaotic system; please state whether this is the spread under identical climatological forcing with different initial conditions, and clarify why the emulator is so insensitive to initial conditions in this diagnostic.
  4. [Figures 2-4 and associated text] Several correlations are reported without specifying the metric precisely (e.g., spatial pattern correlation versus temporal correlation) or the effective degrees of freedom. For example, the 'correlation of 0.98' in Figure 2B and the time-series correlations above 0.98 for Niño3.4 should state whether these are Pearson correlations over space, time, or both, and whether they are computed after detrending or removal of the seasonal cycle.
  5. [Section 3.2] The statement that the emulators 'capture the lagged correlation between indices, though the spread between subsets of the data remains large' would benefit from a quantitative measure of that spread, such as a confidence interval or a null expectation, since the large spread weakens the strength of the claim as presented.
  6. [Section 2.4] The dynamic weighting in Eq. (5) is described as using the reciprocal root-mean-square error per channel and rollout step, with clipping at 500:1; a brief justification for the clipping value and the smoothing period N_smooth=100 would help readers assess how sensitive the results are to these hyperparameters.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the midHolocene target is external to training and the additivity result is an internal consistency check; self-citations are not load-bearing.

full rationale

The central evaluation is genuinely out-of-sample: F_pi is trained only on CESM2 piControl monthly samples, and no midHolocene data enter the loss, dynamic weighting, checkpoint selection, or any fitted parameter. The target response (CESM2 midHolocene minus piControl) is external to the emulator, and the baselines are likewise fit on piControl and applied to midHolocene forcings. The linear-superposition claim in Section 4.2 is explicitly an internal property of the emulator, not of the underlying CESM2 experiments: the paper states 'we note that this probes additivity of the emulator's response, not the additivity of the underlying CESM2 experiments, which would require single-forcing CESM2 experiments that are not available.' The self-citations (Samudra for architecture, Subel & Zanna for prior MSE issues) provide architecture and context, not the target response or the evaluation metrics, and no uniqueness theorem or ansatz is imported to force the conclusions. The paper's caveat in Section 3.3 that 'we expect the system to have equilibrated on the century timescales we investigate' is a validity concern about using the CESM2 difference as ground truth, not a circular step, because it does not reduce the emulator's prediction to its inputs. No fitted parameter is renamed as a prediction, and no equation defines the target in terms of the emulator output. Therefore no circularity is exhibited; score 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No physical constants or per-result parameters were fitted to the midHolocene target. The neural network weights and training hyperparameters, including channel widths, rollout schedule, and the dynamic loss weighting W, are fitted to piControl data only and are not used to force the out-of-sample response. The central evaluation relies on three domain assumptions: equilibrium of the CESM2 difference, fidelity of the reconstructed surface stress, and the sign-flip convention for F_mH. No new physical entities are introduced.

assumptions (3)
  • domain assumption CESM2 midHolocene and piControl experiments are equilibrated, so their difference is a valid estimate of the forced response on the evaluated timescales.
    Section 2.5.2 and Section 3.3: the paper says 'we expect the system to have equilibrated on the century timescales we investigate.' All skill correlations and amplitude ratios score emulators against this difference.
  • domain assumption The surface stress field reconstructed by weighting atmospheric and sub-ice stress with sea ice concentration matches CESM2's own convention.
    Section 2.1: the reconstructed total stress is used to drive the emulator. If the weighting is not faithful, forcing errors propagate into the response metrics.
  • ad hoc to paper The piControl and midHolocene emulator responses are approximately opposite, so R(F_mH) can be compared to -R(F_pi) after flipping the sign of the true response.
    Section 3.3: 'we expect much of the large-scale upper ocean response to be mirrored... comparisons between R(F_pi) and -R(F_mH) are made under this convention and are not designed to establish the validity of this assumption.' This assumption is load-bearing for all F_mH correlation and RMSE values.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Paleoclimate Boundary Conditions as an Out-of-Sample Test for the Forced Response of Ocean Climate Emulators." pith.science (2026). https://pith.science/paper/ZTGPOLD3

@misc{pith2026260813494,
  author       = {Pith},
  title        = {Pith review of: Paleoclimate Boundary Conditions as an Out-of-Sample Test for the Forced Response of Ocean Climate Emulators},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZTGPOLD3}},
  note         = {Machine review of arXiv:2608.13494}
}
read the original abstract

AI weather emulators benefit from clear objectives and metrics, which have led to the rapid development of models that outperform traditional benchmarks. In contrast, long-term climate emulators must reliably reproduce forced responses over months to centuries, while relying on training objectives that span a small number of model time steps. We assess autoregressive, full-depth ocean emulators using data from the midHolocene experiment of a numerical climate model to examine their skill in responding to surface forcings from an in-distribution, out-of-sample climate. We demonstrate that these emulators generalize to new orbital forcings, reproducing the spatial structure of the large-scale response as well as changes in seasonal patterns and in the spatial structure of ocean variability, while underestimating their amplitude. Baselines that infer the ocean state directly from the boundary forcings also recover much of the large-scale pattern, but only near the surface, and capture neither the seasonal nor the variability changes, indicating that these require some representation of dynamics. Despite these successes, the emulators fail to reproduce the slow, internally driven evolution of the ocean interior. We then show that the emulators' total forced response is well reconstructed by linearly composing their independent responses to each forcing component. Tracking response across training epochs, we find that convergence on mean state metrics in the training climate does not guarantee that the emulators capture the dynamics necessary for a skillful response. Together, these experiments establish the midHolocene as a controlled, ground-truthed setting for diagnosing forced-response failures before emulators are pushed to out-of-distribution climates.

Figures

Figures reproduced from arXiv: 2608.13494 by the authors.

Figure 1
Figure 1. Differences between the piControl and midHolocene CESM2 data. Panels A-C show the difference between midHolocene and piControl in the time-mean boundary forcing compo￾nents for net heat flux, hfds (A), surface zonal stress, τu (B), and surface meridional stress, τv (C), respectively. The spatiotemporal standard deviation of each component in the piControl experiment is 91.17 [W/m2 ], 0.085 [N/m2 ], and 0.065 [N/m2 ]… view at source ↗
Figure 2
Figure 2. Comparison of the tropical Pacific seasonal SST anomaly changes between the mid￾Holocene and piControl. We define SST as the uppermost layer (0-5m) and compute anomalies as the difference between the monthly and annual climatologies. We then compute the meridional average of these anomalies between 1.5 ◦S − 1.5 ◦N. A: difference between the midHolocene and piControl CESM2 numerical experiments. For B-D, we use the t… view at source ↗
Figure 3
Figure 3. Depth profiles detailing the change in the seasonal range of potential temperature (SON minus MAM) between the midHolocene and piControl climates (Equation 7) in the Pacific Ocean. A: change in range between the two CESM2 experiments. For B-D: we use the title to indicate the networks and rollout parameters using abbreviated forms of the notation defined in Section 2.5.3. For each rollout, we use initial conditions … view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Comparison of dominant modes of ocean variability across the piControl and midHolocene experiments. A-C: the timeseries of Ni˜no3.4 in the piControl climate (A), the mid￾Holocene climate (B), and the lagged autocorrelation (C). D-F: similar results for the DMI, a measu…
Figure 5
Figure 5. Figure 5: Comparison of the change in the potential temperature variability over the upper 200m. A: true difference between the two CESM2 experiments. For B-D, we use the title to in￾dicate the networks and rollout parameters, using abbreviated notation defined in Section 2.5.3.…
Figure 6
Figure 6. Figure 6: Comparison of the zonally averaged potential temperature response in each ocean basin when perturbing an emulator using climatological forcings taken from mid￾Holocene/piControl CESM2 data. A, D, G, and J: difference between the CESM2 midHolocene and piControl experime…
Figure 7
Figure 7. Figure 7: Comparison of the depth-averaged potential temperature response over the upper 200m (A-C) and 200-1000m (D-F) when perturbing an emulator with climatological forcings from the midHolocene/piControl CESM2 experiments. A and D: difference between the CESM2 midHolocene an…
Figure 8
Figure 8. Figure 8: Comparison of the zonally averaged potential temperature response in the Southern Ocean when perturbing piControl trained emulators with boundary forcing components of clima￾tological values from the midHolocene CESM2 data. The title of each panel follows the notation …
Figure 9
Figure 9. Figure 9: Comparison of the total Southern Ocean potential temperature response to mid￾Holocene boundary forcings against the linear superposition of individual component responses. A: Difference between the CESM2 midHolocene and piControl numerical experiments for ref￾erence. B…
Figure 10
Figure 10. Figure 10: Epoch sensitivity of out-of-sample response skill and spread as a function of epoch. In all panels, response metrics are plotted against the RMSE of the upper 1000m depth￾averaged mean state in the referenced basin. The mean states are computed over a 100 year roll￾ou…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

63 extracted references · 51 canonical work pages

  1. [1]

    Journal of Physical Oceanography , volume =

    Wunsch, Carl and Roemmich, Dean , title =. Journal of Physical Oceanography , volume =. 1985 , doi =

  2. [2]

    and Roemmich, Dean H

    Hautala, Susan L. and Roemmich, Dean H. and Schmitz, William J., Jr. , title =. Journal of Geophysical Research: Oceans , volume =. 1994 , doi =

  3. [3]

    Journal of Oceanography , volume =

    Aoki, Kunihiro and Kutsuwada, Kunio , title =. Journal of Oceanography , volume =. 2008 , doi =

  4. [4]

    and Riser, Stephen C

    Gray, Alison R. and Riser, Stephen C. , title =. Journal of Physical Oceanography , volume =. 2014 , doi =

  5. [5]

    and Storer, Benjamin A

    Khatri, Hemant and Griffies, Stephen M. and Storer, Benjamin A. and Buzzicotti, Michele and Aluie, Hussein and Sonnewald, Maike and Dussin, Raphael and Shao, Andrew , title =. Journal of Advances in Modeling Earth Systems , volume =. 2024 , doi =

  6. [6]

    2025 , publisher=

    Dheeshjith, Surya and Subel, Adam and Adcroft, Alistair and Busecke, Julius and Fernandez-Granda, Carlos and Gupta, Shubham and Zanna, Laure , journal=. 2025 , publisher=

  7. [7]

    Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

    Transformers without normalization , author=. Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

  8. [8]

    Kageyama, Masa and Braconnot, Pascale and Harrison, Sandy P and Haywood, Alan M and Jungclaus, Johann H and Otto-Bliesner, Bette L and Peterschmitt, Jean-Yves and Abe-Ouchi, Ayako and Albani, Samuel and Bartlein, Patrick J and others , journal=. The. 2018 , publisher=

Show all 63 references
  1. [9]

    Rose, Brian EJ , journal=

  2. [10]

    Getting started with

    McDougall, Trevor J and Barker, Paul M , journal=. Getting started with

  3. [11]

    Structure and performance of

    Held, IM and Guo, H and Adcroft, A and Dunne, JP and Horowitz, LW and Krasting, J and Shevliakova, E and Winton, M and Zhao, M and Bushuk, M and others , journal=. Structure and performance of. 2019 , publisher=

  4. [12]

    Danabasoglu, Gokhan and Lamarque, J-F and Bacmeister, J and Bailey, DA and DuVivier, AK and Edwards, Jim and Emmons, LK and Fasullo, John and Garcia, R and Gettelman, Andrew and others , journal=. The. 2020 , publisher=

  5. [13]

    Loose, Nora and Abernathey, Ryan and Grooms, Ian and Busecke, Julius and Guillaumin, Arthur and Yankovsky, Elizabeth and Marques, Gustavo and Steinberg, Jacob and Ross, Andrew and Khatri, Hemant and others , journal=

  6. [14]

    doi:10.5281/zenodo.8356796 , url =

    Zhuang, Jiawei and dussin, raphael and Huard, David and Bourgault, Pascal and Banihirwe, Anderson and Raynaud, Stephane and Malevich, Brewster and Schupfner, Martin and Filipe and Levang, Sam and Gauthier, Charles and Jüling, André and Almansi, Mattia and RichardScottOZ and Ro...

  7. [15]

    arXiv preprint arXiv:2402.04342 , year=

    Building ocean climate emulators , author=. arXiv preprint arXiv:2402.04342 , year=

  8. [16]

    Liu, Zhuang and Mao, Hanzi and Wu, Chao-Yuan and Feichtenhofer, Christoph and Darrell, Trevor and Xie, Saining , booktitle=. A

  9. [17]

    Transfer learning for emulating ocean climate variability across

    Dheeshjith, Surya and Subel, Adam and Gupta, Shubham and Adcroft, Alistair and Fernandez-Granda, Carlos and Busecke, Julius and Zanna, Laure , journal=. Transfer learning for emulating ocean climate variability across

  10. [18]

    A comparison of the

    Otto-Bliesner, Bette L and Brady, Esther C and Tomas, Robert A and Albani, Samuel and Bartlein, Patrick J and Mahowald, Natalie M and Shafer, Sarah L and Kluzek, Erik and Lawrence, Peter J and Leguy, Gunter and others , journal=. A comparison of the. 2020 , publisher=

  11. [19]

    Otto-Bliesner, Bette L and Braconnot, Pascale and Harrison, Sandy P and Lunt, Daniel J and Abe-Ouchi, Ayako and Albani, Samuel and Bartlein, Patrick J and Capron, Emilie and Carlson, Anders E and Dutton, Andrea and others , journal=. The. 2017 , publisher=

  12. [20]

    Simpson, Isla R and Rosenbloom, Nan and Danabasoglu, Gokhan and Deser, Clara and Yeager, Stephen G and McCluskey, Christina S and Yamaguchi, Ryohei and Lamarque, Jean-Francois and Tilmes, Simone and Mills, Michael J and others , journal=. The. 2023 , publisher=

  13. [21]

    Sun, Y Qiang and Hassanzadeh, Pedram and Zand, Mohsen and Chattopadhyay, Ashesh and Weare, Jonathan and Abbot, Dorian S , journal=. Can. 2025 , publisher=

  14. [22]

    Predicting Beyond Training Data via Extrapolation versus Translocation:

    Sun, Y Qiang and Hassanzadeh, Pedram and Shaw, Tiffany and Pahlavan, Hamid A , journal=. Predicting Beyond Training Data via Extrapolation versus Translocation:

  15. [23]

    npj Climate and Atmospheric Science , volume=

    Skilful global seasonal predictions from a machine learning weather model trained on reanalysis data , author=. npj Climate and Atmospheric Science , volume=. 2025 , publisher=

  16. [24]

    Numerical models outperform

    Zhang, Zhongwei and Fischer, Erich and Zscheischler, Jakob and Engelke, Sebastian , journal=. Numerical models outperform

  17. [25]

    Science , volume=

    Learning skillful medium-range global weather forecasting , author=. Science , volume=. 2023 , publisher=

  18. [26]

    2025 , publisher=

    Watt-Meyer, Oliver and Henn, Brian and McGibbon, Jeremy and Clark, Spencer K and Kwa, Anna and Perkins, W Andre and Wu, Elynn and Harris, Lucas and Bretherton, Christopher S , journal=. 2025 , publisher=

  19. [27]

    Chapman, William E and Schreck, John S and Sha, Yingkai and Gagne II, David John and Kimpara, Dhamma and Zanna, Laure and Mayer, Kirsten J and Berner, Judith , journal=

  20. [28]

    Accurate medium-range global weather forecasting with

    Bi, Kaifeng and Xie, Lingxi and Zhang, Hengheng and Chen, Xin and Gu, Xiaotao and Tian, Qi , journal=. Accurate medium-range global weather forecasting with. 2023 , publisher=

  21. [29]

    Artificial Intelligence for the Earth Systems , volume=

    Using neural networks to learn the jet stream forced response from natural variability , author=. Artificial Intelligence for the Earth Systems , volume=

  22. [30]

    arXiv preprint arXiv:2506.22552 , year=

    Neural models of multiscale systems: conceptual limitations, stochastic parametrizations, and a climate application , author=. arXiv preprint arXiv:2506.22552 , year=

  23. [31]

    Guan, Haiwen and Arcomano, Troy and Chattopadhyay, Ashesh and Maulik, Romit , journal=

  24. [32]

    Artificial Intelligence for the Earth Systems , volume=

    Dynamical tests of a deep learning weather prediction model , author=. Artificial Intelligence for the Earth Systems , volume=. 2024 , publisher=

  25. [33]

    2024 , publisher=

    Chattopadhyay, Ashesh and Gray, Michael and Wu, Tianning and Lowe, Anna B and He, Ruoying , journal=. 2024 , publisher=

  26. [34]

    Science Advances , volume=

    Data-driven global ocean modeling for seasonal to decadal prediction , author=. Science Advances , volume=. 2025 , publisher=

  27. [35]

    Ocean emulation with

    Bire, Suyash and L. Ocean emulation with. Journal of Advances in Modeling Earth Systems , volume=. 2025 , publisher=

  28. [36]

    Nature , volume=

    Neural general circulation models for weather and climate , author=. Nature , volume=. 2024 , publisher=

  29. [37]

    Bonev, Boris and Kurth, Thorsten and Mahesh, Ankur and Bisson, Mauro and Kossaifi, Jean and Kashinath, Karthik and Anandkumar, Anima and Collins, William D and Pritchard, Michael S and Keller, Alexander , journal=

  30. [38]

    arXiv preprint arXiv:2412.15832 , year=

    Lang, Simon and Alexe, Mihai and Clare, Mariana CA and Roberts, Christopher and Adewoyin, Rilwan and Bouall. arXiv preprint arXiv:2412.15832 , year=

  31. [39]

    Nature , volume=

    Probabilistic weather forecasting with machine learning , author=. Nature , volume=. 2025 , publisher=

  32. [40]

    Clark, Spencer K and Watt-Meyer, Oliver and Kwa, Anna and McGibbon, Jeremy and Henn, Brian and Perkins, W Andre and Wu, Elynn and Harris, Lucas M and Bretherton, Christopher S , journal=

  33. [41]

    and Crotwell, A.M

    Thoning, K.W. and Crotwell, A.M. and Mund, J.W. Atmospheric Carbon Dioxide Dry Air Mole Fractions from continuous measurements at Mauna Loa , Hawaii , Barrow , Alaska , American Samoa and South Pole. 2025. doi:10.15138/yaf1-bk21

  34. [42]

    arXiv preprint arXiv:2501.19374 , year=

    Fixing the double penalty in data-driven weather forecasting through a modified spherical harmonic loss function , author=. arXiv preprint arXiv:2501.19374 , year=

  35. [43]

    Journal of the American Statistical Association , volume=

    Strictly proper scoring rules, prediction, and estimation , author=. Journal of the American Statistical Association , volume=. 2007 , publisher=

  36. [44]

    Improving

    Sha, Yingkai and Schreck, John S and Chapman, William and Gagne II, David John , journal=. Improving

  37. [45]

    Duncan, James PC and Wu, Elynn and Dheeshjith, Surya and Subel, Adam and Arcomano, Troy and Clark, Spencer K and Henn, Brian and Kwa, Anna and McGibbon, Jeremy and Perkins, W Andre and others , journal=

  38. [46]

    Couairon, Guillaume and Singh, Renu and Charantonis, Anastase and Lessig, Christian and Monteleoni, Claire , journal=

  39. [47]

    Applying the

    Wu, Elynn and Rebassoo, Finn and Paul, Pappu and Proistosescu, Cristian and Nugent, Jacqueline and McCoy, Daniel and Caldwell, Peter and Bretherton, Christopher S , journal=. Applying the

  40. [48]

    Nature Climate Change , volume=

    Evaluation of climate models using palaeoclimatic data , author=. Nature Climate Change , volume=. 2012 , publisher=

  41. [49]

    Huang, Qiusheng and Niu, Yuan and Zhong, Xiaohui and Guo, Anboyu and Chen, Lei and Zhang, Dianjun and Zhang, Xuefeng and Li, Hao , journal=

  42. [50]

    Nature Communications , volume=

    Forecasting the eddying ocean with a deep neural network , author=. Nature Communications , volume=. 2025 , publisher=

  43. [51]

    A dipole mode in the tropical

    Saji, NH and Goswami, Bhupendra Nath and Vinayachandran, PN and Yamagata, Toshio , journal=. A dipole mode in the tropical. 1999 , publisher=

  44. [52]

    Science , volume=

    Past climates inform our future , author=. Science , volume=. 2020 , publisher=

  45. [53]

    Journal of Advances in Modeling Earth Systems , volume=

    Increasingly sophisticated climate models need the out-of-sample tests paleoclimates provide , author=. Journal of Advances in Modeling Earth Systems , volume=. 2022 , publisher=

  46. [54]

    AGU Advances , volume=

    A deep learning earth system model for efficient simulation of the observed climate , author=. AGU Advances , volume=. 2025 , publisher=

  47. [55]

    Gregory, William and Bushuk, Mitchell and Duncan, James and Wu, Elynn and Subel, Adam and Clark, Spencer K and Hurlin, Bill and Watt-Meyer, Oliver and Adcroft, Alistair and Bretherton, Chris and others , journal=

  48. [56]

    Communication of the role of natural variability in future

    Deser, Clara and Knutti, Reto and Solomon, Susan and Phillips, Adam S , journal=. Communication of the role of natural variability in future. 2012 , publisher=

  49. [57]

    Earth System Dynamics , volume=

    How large does a large ensemble need to be? , author=. Earth System Dynamics , volume=. 2020 , publisher=

  50. [58]

    Reanalysis-based global radiative response to sea surface temperature patterns: Evaluating the

    Van Loon, Senne and Rugenstein, Maria and Barnes, Elizabeth A , journal=. Reanalysis-based global radiative response to sea surface temperature patterns: Evaluating the. 2025 , publisher=

  51. [59]

    2006 , publisher=

    McPhaden, Michael J and Zebiak, Stephen E and Glantz, Michael H , journal=. 2006 , publisher=

  52. [60]

    2019 , publisher =

    Danabasoglu, Gokhan , title =. 2019 , publisher =. doi:10.22033/ESGF/CMIP6.7674 , url =

  53. [61]

    2019 , publisher =

    Danabasoglu, Gokhan and Lawrence, David and Lindsay, Keith and Lipscomb, William and Strand, Gary , title =. 2019 , publisher =. doi:10.22033/ESGF/CMIP6.7733 , url =

  54. [62]

    2019 , publisher =

    Danabasoglu, Gokhan and Lawrence, David and Lindsay, Keith and Lipscomb, William and Strand, Gary , title =. 2019 , publisher =. doi:10.22033/ESGF/CMIP6.7497 , url =

  55. [63]

    2026 , publisher =

    Subel, Adam and Zanna, Laure , title =. 2026 , publisher =. doi:10.5281/zenodo.21891296 , url =

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.