Pith. sign in

REVIEW 3 major objections 7 minor 51 references

Machine learning trained only on sparse Earth observations can produce multi-decade global reanalyses without physics models.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-10 16:08 UTC pith:UZMG4F3Z

load-bearing objection Solid prototype: observation-only multi-decade reanalysis that is fast, independent of NWP, and competitive on held-out winds/surface checks, with residual physics and MSE-smoothing limits the authors already flag. the 3 major comments →

arxiv 2607.07879 v1 pith:UZMG4F3Z submitted 2026-07-08 physics.ao-ph

Global reanalysis from observations alone with machine learning

classification physics.ao-ph
keywords reanalysismachine learningobservation-drivendata assimilationAIFS-DOPEarth system observationsglobal atmosphere
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Traditional reanalyses fill sparse observations with physics-based numerical models and take years to produce. This paper shows a prototype that instead trains a machine-learning model end-to-end solely on quality-controlled satellite and conventional observations, then generates dense six-hourly global fields by cycling short predictions conditioned on recent data. The resulting 42-year gridded product recovers large-scale atmospheric structure, seasonal and interannual variability, and several dynamical balances, while matching or approaching leading reanalyses on independent wind and surface checks. Because inference is cheap, the entire multi-decade archive can be written in a working day. If the approach holds, reanalysis becomes an iterative, observation-only reconstruction rather than a multi-year physics-assimilation project.

Core claim

A machine-learning model trained exclusively on sparse Earth-system observations, with no reanalysis targets and no physics-based forecast model, can generate multi-decade global gridded reanalyses that capture mean atmospheric structure, multi-timescale variability and key dynamical relationships, and that achieve upper-level wind errors close to ERA5 at matched resolution and surface errors between ERA-Interim and ERA5.

What carries the argument

AIFS-DOP: an encoder–processor–decoder graph/transformer model that maps sparse observations on a regular O96 grid through a short cycling of six-hour predictions conditioned on the previous 30 hours of data, trained only with a masked mean-squared-error loss on the next observation window.

Load-bearing premise

That agreement with independent held-out observations and with large-scale ERA5 patterns is enough to prove the dense multi-variable fields are physically coherent reconstructions rather than sophisticated interpolations of the dense modern observing system.

What would settle it

A systematic comparison of the generated fields against a dense, never-used observing system (for example independent radiosonde or campaign profiles) in data-sparse regions and periods, checking whether dynamical balances and small-scale variance degrade when the modern satellite network is thinned or removed.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The manuscript presents a prototype multi-decadal global atmospheric reanalysis (1981–2022, O96, ~112 km) generated by AIFS-DOP, a graph/transformer model trained end-to-end solely on sparse conventional and satellite observations, with no reanalysis fields as inputs or targets and no physics-based NWP model. Analyses are produced by independent short cycling of one-step (6 h) predictions conditioned on the previous 30 h of observations. The authors show that the gridded fields recover large-scale mean structure (zonal jets, thermal stratification, ITCZ migration), multi-timescale variability (storm tracks, ENSO and teleconnections, volcanic and surface temperature anomalies, selected extremes), and signs of dynamical coherence (effective Coriolis parameter; cross-variable linear regression patterns). Held-out MISR cloud-motion winds give upper-level RMS vector differences close to ERA5 at matched resolution; independent surface-land stations yield error standard deviations between ERA-Interim and ERA5. Production of the full 42-year product is reported to take a single working day after a short GPU training run. The paper is framed as a method demonstration rather than a finished climate product.

Significance. If the result holds under the authors’ carefully hedged reading, this is a genuine new production pathway for reanalysis: dense multi-variable gridded states from observations alone, independent of NWP priors, at a cost that enables iterative refinement, ensembles, and rapid nesting. Strengths that should be credited include (i) a training setup that excludes reanalysis targets, (ii) genuine held-out verification (MISR never assimilated in ERA5 or AIFS-DOP; surface-land stations excluded by O96 spatio-temporal matching against the ECMWF archive), (iii) multi-diagnostic physical-consistency checks beyond point skill (effective Coriolis; cross-variable regression), and (iv) an explicit computational demonstration. These place the work well above a pure interpolation exercise and make it of clear interest to the reanalysis and ML-weather communities, provided claims remain matched to the evidence.

major comments (3)
  1. [Discussion; Fig. 9] Discussion (paragraph on ENSO teleconnections) and Fig. 9: the authors correctly note that teleconnection patterns “may simply indicate that the observations are sufficiently dense to constrain these features at initialisation time” rather than that dynamics were learned. That caveat is load-bearing for the central claim of a “physically coherent” multi-variable reanalysis from observations alone. Please either (a) add a diagnostic that tests dynamical consistency preferentially in data-sparse regions/eras (e.g., SH midlatitudes or pre-1990s windows; residual balance errors stratified by observation density), or (b) systematically scope the abstract, introduction, and conclusions to “observation-constrained gridded state estimates with emergent large-scale balance,” so the stronger dynamical-reconstruction reading is not the default.
  2. [Evaluation against independent observations; Fig. 13; Fig. 12] Evaluation against independent observations / Fig. 13: the headline that upper-level wind RMSVD is “close to that of ERA5” is undercut by the paper’s own spectral and double-penalty discussion (Fig. 12; text noting unconstrained small-scale energy in ERA5). AIFS-DOP’s smoother fields can improve RMSVD without implying equal analysis quality. Please report at least one activity- or scale-aware comparison (e.g., RMSVD after common spectral filtering to the effective AIFS-DOP resolution, or scores stratified by spatial scale / against the EDA mean as the primary ERA5 reference) so the abstract claim is not inflated by smoothness.
  3. [Atmospheric structure and mean state; Figs. 4–5] Figs. 4–5 and Physical consistency: mid-level tropical meridional circulation and polar/stratospheric relative humidity show clear, physically implausible departures from ERA5 (deeper mid-level V cells; unrealistically high RH in dry polar/stratospheric air). These are not peripheral cosmetics; they speak directly to multi-variable 3D coherence. Either demonstrate that these defects do not contaminate the variables and applications for which skill is claimed, or state more prominently (including near the abstract skill statements) which components of the 3D state are not yet reliable and why MSE-on-specific-humidity is the suspected cause.
minor comments (7)
  1. [Abstract; Introduction] Abstract and Introduction: “without using physics-based numerical models” is accurate for the analysis step but could be misread as “no physical information of any kind.” A short clause that balance emerges from observation-trained representations (not from an NWP prior) would reduce ambiguity.
  2. [Model and datasets] Model and datasets: the independent cycling of each analysis (no serial long-window assimilation) is important and well motivated; please state explicitly whether temporal discontinuities at cycle boundaries were checked (e.g., 6-hourly jump statistics vs ERA5).
  3. [Fig. 3] Fig. 3: island-scale convergence spots are noted as possible station artifacts; a brief sensitivity test (masking nearby SYNOP) or a clearer caveat in the caption would help readers not over-interpret those features.
  4. [Fig. 14; Evaluation against independent observations] Fig. 14: evaluation on the 15th of each month only is pragmatic but underspecified for reproducibility; state the exact matching rules and sample sizes per period in the Methods or caption.
  5. [Model and datasets; Table 1] Table 1 / Methods: training ends 2020, reanalysis runs through 2022; a short skill split for 2021–2022 vs the training decades (even for MISR or surface) would reassure readers on memorisation for the product period.
  6. [Throughout] Typos/clarity: “betweensparse” (Introduction); “Asanexampleofvariability” and similar missing spaces in Multi-scale variability; “1European” affiliation formatting; ensure consistent ERA5 vs ERA5.1 labelling in Fig. 7.
  7. [References] References to AIFS-DOP and GraphDOP arXiv preprints are appropriate; if any have been peer-reviewed by acceptance, update citations.

Circularity Check

1 steps flagged

No load-bearing circularity: skill claims rest on held-out independent observations (MISR, independent surface stations); ERA5 is only a non-training structural reference; self-citations describe the prior DOP architecture but do not force the reanalysis results.

specific steps
  1. self citation load bearing [Model and datasets section; citations [28], [24], [14]]
    "This paper uses the AIFS-DOP model introduced in Pinnington et al. (2026) [28]. AIFS-DOP builds on the AIFS model... It draws on experience gained from GraphDOP [24]. ... This dataset is the same as that described in Pinnington et al. (2026) [28]."

    The architecture, training objective, and curated observation dataset are imported from the authors’ own contemporaneous/prior DOP papers. This is ordinary self-citation for a methods extension and is not load-bearing for the novel claim (that the resulting multi-decade gridded fields achieve ERA5-comparable skill on held-out independent observations). The skill numbers themselves are computed against external data never used in those prior works.

full rationale

The paper trains AIFS-DOP end-to-end exclusively on sparse observations (no reanalysis fields as inputs or targets) and generates the multi-decade gridded product by independent six-hour cycling. Primary quantitative claims (upper-air RMSVD close to ERA5 at matched O96 resolution; surface error SDs between ERA-Interim and ERA5) are evaluated against held-out MISR cloud-motion winds never assimilated in ERA5 or the training set, and against surface-land stations excluded by O96 spatio-temporal matching from the ECMWF archive. ERA5 appears only as a qualitative structural reference (explicitly not a training target or ground truth). Self-citations to the authors’ prior DOP/GraphDOP/AIFS papers supply the model architecture and training dataset description; they are not uniqueness theorems, do not define the evaluation metrics, and do not make the reanalysis skill claims true by construction. Physical-consistency diagnostics (effective Coriolis, cross-variable regressions) and multi-scale variability plots are post-hoc checks, not fitted inputs renamed as predictions. No equation or procedure reduces a claimed prediction to its own inputs. Residual scientific caveats (MSE smoothing, possible observation-constrained rather than dynamically learned teleconnections) are correctness/weak-assumption issues, not circularity. Score 1 reflects only the presence of non-load-bearing self-citations that are normal for a methods paper building on prior work by the same group.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 0 invented entities

The claim rests on standard ML and DA-inspired engineering choices plus domain assumptions that sparse multi-modal observations plus a learned latent dynamics model can replace an NWP prior for dense state estimation. No new physical entities are postulated. Free parameters are architectural and procedural choices that define the prototype rather than fitted physical constants.

free parameters (5)
  • Horizontal resolution (O96 octahedral grid) = O96 (~112 km)
    Fixed prototype mesh (~112 km) that sets resolvable scales and comparison fairness to ERA5/ERA-Interim at O96; not derived from first principles.
  • Cycling window length and steps = 4 × 6 h (~30 h context)
    Four independent six-hour steps conditioned on ~30 hours of recent observations; design choice inspired by DA cycling, not optimised or derived in the paper.
  • MSE loss with missing-value mask = MSE (masked)
    Training objective that drives conditional-mean predictions and small-scale smoothing; acknowledged as limiting mesoscale fidelity.
  • Processor depth and attention design = 16 layers
    16 pre-norm transformer layers with sliding-window longitudinal attention; architectural hyperparameters of AIFS-DOP.
  • Training period split = 1981–2020 train
    Train 1981–2020, validate Jan–May 2021, generate 1981–2022; choice affects memorisation risk and claimed multi-decade coverage.
axioms (4)
  • domain assumption Sparse conventional and satellite observations quality-controlled and mapped to a regular 6-hourly O96 grid with missing values imputed as zeros after normalisation are a sufficient training signal for global state estimation.
    Model and datasets section; entire pipeline depends on this curated Anemoi observation cube.
  • ad hoc to paper A learned encoder–processor–decoder with independent short cycling can substitute for a physics-based forecast model and background-error covariances in producing dense, multi-variable analyses.
    Core methodological premise of the reanalysis procedure; not proven, tested empirically against ERA5 and held-out obs.
  • domain assumption Geostrophic balance and linear cross-variable regressions against Z500 anomalies are informative diagnostics of physical coherence for ML-generated fields.
    Physical consistency section and Methods (effective Coriolis derivation); standard dynamical meteorology tools applied to ML output.
  • domain assumption Held-out MISR stereoscopic cloud-motion winds and C3S land surface stations excluded by O96 archive matching are independent enough to rank reanalysis skill without circular use of training data.
    Evaluation against independent observations section; load-bearing for the ERA-Interim/ERA5 skill ranking claim.

pith-pipeline@v1.1.0-grok45 · 27059 in / 3514 out tokens · 47508 ms · 2026-07-10T16:08:48.015859+00:00 · methodology

0 comments
read the original abstract

Earth system reanalysis datasets are foundational for weather and climate research and provide the gridded training data used by most machine learning weather prediction systems. Here we show results from a prototype system that suggest that machine learning models trained only on Earth system observations can potentially be used to generate multi-decade global reanalyses without using physics-based numerical models. The resulting gridded fields capture large-scale atmospheric structure and variability across multiple timescales, while exhibiting signs of physical coherence in several key dynamical diagnostics. Evaluations of the prototype against held-out independent atmospheric observations indicate that the root mean square vector error of upper-level winds is close to that of ERA5 when compared at a consistent resolution, and that the standard deviation of the error at the surface is between that of 4th- and 5th-generation ECMWF reanalyses (ERA-Interim and ERA5). Furthermore, while traditional reanalysis production is computationally expensive, typically taking several years to produce, the reanalysis presented here was generated during the course of a single working day. These results suggest that observation-trained machine learning models offer a promising new approach for reanalysis production from observations alone.

Figures

Figures reproduced from arXiv: 2607.07879 by Anthony McNally, Eulalie Boucher, Ewan Pinnington, Hans Hersbach, Matthew Chantry, Mihai Alexe, Niels Bormann, Patrick Laloyaux, Paul Poli, Peter Lean, Simon Lang, Tomas Kral.

Figure 1
Figure 1. Figure 1: Zonal mean cross-sections of temperature (shaded, K) and zonal wind (black [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Cross-sections of the zonal mean differences between AIFS-DOP and ERA5 (cal [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Seasonal climatology of tropical surface wind convergence (2000–2019). 20-year [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Zonal mean meridional wind climatology (2000–2019). Cross-sections of the merid [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Zonal mean relative humidity climatology (2000–2019). Cross-sections of the rela [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: (a, b) Extratropical storm-track structure for the Northern and Southern Hemi [PITH_FULL_IMAGE:figures/full_fig_p010_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Global temperature anomalies comparing the AIFS-DOP reanalysis with ERA5.1 [PITH_FULL_IMAGE:figures/full_fig_p011_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: El Niño–Southern Oscillation (ENSO) sea surface temperature (SST) anomalies. [PITH_FULL_IMAGE:figures/full_fig_p012_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: El Niño minus La Niña atmospheric composites (2013–2022). Global composite [PITH_FULL_IMAGE:figures/full_fig_p013_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Extreme event case studies comparing AIFS-DOP (left panels) and ERA5 (right [PITH_FULL_IMAGE:figures/full_fig_p014_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Physical-consistency diagnostics; (a) time-mean (2000–2019) effective Coriolis pa [PITH_FULL_IMAGE:figures/full_fig_p016_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: One-dimensional zonal spectra of horizontal kinetic energy at 250 hPa (top row) [PITH_FULL_IMAGE:figures/full_fig_p018_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Validation of reanalysis wind fields against MISR AMV satellite observations [PITH_FULL_IMAGE:figures/full_fig_p020_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Assessment of AIFS-DOP (orange) against surface-land observations, showing [PITH_FULL_IMAGE:figures/full_fig_p022_14.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

51 extracted references · 51 canonical work pages · 15 internal anchors

  1. [2]

    Integration of space and in situ observations to study global climate change.Bulletin of the American Meteorological Society, 69(10): 1130–1143, 1988

    Lennart Bengtsson and Jagadish Shukla. Integration of space and in situ observations to study global climate change.Bulletin of the American Meteorological Society, 69(10): 1130–1143, 1988. doi: 10.1175/1520-0477(1988)069<1130:IOSAIS>2.0.CO;2

  2. [3]

    Hans Hersbach, Bill Bell, Paul Berrisford, Shoji Hirahara, András Horányi, Joaquín Muñoz-Sabater, Julien Nicolas, Carole Peubey, Raluca Radu, Dinand Schepers, Adrian 27 Simmons, Cornel Soci, Saleh Abdalla, Xavier Abellan, Gianpaolo Balsamo, Peter Bech- told, Gionata Biavati, Jean Bidlot, Massimo Bonavita, Giovanna De Chiara, Per Dahlgren, Dick Dee, Michai...

  3. [4]

    Suárez, Ricardo Todling, Andrea Molod, Lawrence Takacs, Cynthia A

    Ronald Gelaro, Will McCarty, Max J. Suárez, Ricardo Todling, Andrea Molod, Lawrence Takacs, Cynthia A. Randles, Anton Darmenov, Michael G. Bosilovich, Rolf Reichle, Krzysztof Wargan, Lawrence Coy, Richard Cullather, Clara Draper, Santha Akella, Virginie Buchard, Austin Conaty, Arlindo M. da Silva, Wei Gu, Gi-Kong Kim, Randal Koster, Robert Lucchesi, Dagma...

  4. [5]

    The JRA-3Q reanalysis.Journal of the Meteorological Society of Japan

    Yuki Kosaka, Shinya Kobayashi, Yayoi Harada, Chiaki Kobayashi, Hiroaki Naoe, Koichi Yoshimoto, Masashi Harada, Naochika Goto, Jotaro Chiba, Kengo Miyaoka, Ryohei Sekiguchi, MakotoDeushi, HirotakaKamahori, TosiyukiNakaegawa, TaichuY.Tanaka, Takayuki Tokuhiro, Yoshiaki Sato, Yasuhiro Matsushita, and Kazutoshi Onogi. The JRA-3Q reanalysis.Journal of the Mete...

  5. [6]

    Suranjana Saha, Shrinivas Moorthi, Hua-Lu Pan, Xingren Wu, Jiande Wang, Sudhir Nadiga, Patrick Tripp, Robert Kistler, John Woollen, David Behringer, Haixia Liu, Di- ane Stokes, Robert Grumbine, George Gayno, Jun Wang, Yu-Tai Hou, Hui-ya Chuang, Hann-Ming H. Juang, Joe Sela, Mark Iredell, Russ Treadon, Daryl Kleist, Paul van Delst, Dennis Keyser, John Derb...

  6. [7]

    Value generated by ERA5: Full report.https: //climate.copernicus.eu/sites/default/files/2024-12/Value-generated-b y-ERA5-full-report.pdf, December 2024

    Copernicus Climate Change Service. Value generated by ERA5: Full report.https: //climate.copernicus.eu/sites/default/files/2024-12/Value-generated-b y-ERA5-full-report.pdf, December 2024. Accessed 2026-04-17

  7. [8]

    Cam- bridge University Press, Cambridge, 2002

    Eugenia Kalnay.Atmospheric Modeling, Data Assimilation and Predictability. Cam- bridge University Press, Cambridge, 2002. doi: 10.1017/CBO9780511802270. 28

  8. [9]

    Klinker, Jean-François Mahfouf, and Adrian Sim- mons

    Florence Rabier, Heikki Järvinen, E. Klinker, Jean-François Mahfouf, and Adrian Sim- mons. The ECMWF operational implementation of four-dimensional variational assim- ilation. I: Experimental results with simplified physics.Quarterly Journal of the Royal Meteorological Society, 126(564):1143–1170, 2000. doi: 10.1002/qj.49712656415

  9. [10]

    Forecasting Global Weather with Graph Neural Networks

    Ryan Keisler. Forecasting global weather with graph neural networks.arXiv preprint arXiv:2202.07575, 2022. doi: 10.48550/arXiv.2202.07575. URLhttps://arxiv.org/ abs/2202.07575

  10. [11]

    FourCastNet: A Global Data-driven High-resolution Weather Model using Adaptive Fourier Neural Operators

    Jaideep Pathak, Shashank Subramanian, Peter Harrington, Sanjeev Raja, Ashesh Chat- topadhyay, Morteza Mardani, Thorsten Kurth, David Hall, Zongyi Li, Kamyar Aziz- zadenesheli, Pedram Hassanzadeh, Karthik Kashinath, and Animashree Anandkumar. Fourcastnet: A global data-driven high-resolution weather model using adaptive fourier neural operators.arXiv prepr...

  11. [12]

    URLhttps://arxiv.org/abs/2202.11214

  12. [13]

    Accurate medium-range global weather forecasting with 3d neural networks.Nature, 619:533–538,

    Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi Tian. Accurate medium-range global weather forecasting with 3d neural networks.Nature, 619:533–538,

  13. [14]

    doi: 10.1038/s41586-023-06185-3

  14. [15]

    Learning skillful medium-range global weather forecasting.Science, 382(6677):1416–1421, 2023

    Remi Lam, Alvaro Sanchez-Gonzalez, Matthew Willson, Peter Wirnsberger, Meire Fortunato, Ferran Alet, Suman Ravuri, Timo Ewalds, Zach Eaton-Rosen, Weihua Hu, Alexander Merose, Stephan Hoyer, George Holland, Oriol Vinyals, Jacklynn Stott, Alexander Pritzel, Shakir Mohamed, and Peter Battaglia. Learning skillful medium-range global weather forecasting.Scienc...

  15. [16]

    Simon Lang, Mihai Alexe, Matthew Chantry, Jesper Dramsch, Florian Pinault, Bau- douin Raoult, Mariana C. A. Clare, Christian Lessig, Michael Maier-Gerber, Linus Mag- nusson, Zied Ben Bouallègue, Ana Prieto Nemesio, Peter D. Dueben, Andrew Brown, Florian Pappenberger, and Florence Rabier. AIFS — ECMWF’s data-driven forecast- ing system.arXiv preprint arXiv...

  16. [17]

    AIFS-CRPS: ensemble forecasting using a model trained with a loss function based on the continuous ranked probability score.npj Artificial Intelligence, 2(1):18, 2026

    Simon Lang, Mihai Alexe, Mariana CA Clare, Christopher Roberts, Rilwan Adewoyin, ZiedBenBouallègue, MatthewChantry, JesperDramsch, PeterDDueben, SaraHahner, et al. AIFS-CRPS: ensemble forecasting using a model trained with a loss function based on the continuous ranked probability score.npj Artificial Intelligence, 2(1):18, 2026. doi: https://doi.org/10.1...

  17. [18]

    Deep Learning for Day Forecasts from Sparse Observations

    Marcin Andrychowicz, Lasse Espeholt, Di Li, Samier Merchant, Alexander Merose, Fred Zyda, Shreya Agrawal, and Nal Kalchbrenner. Deep learning for day forecasts from sparse observations.arXiv preprint arXiv:2306.06079, 2023. doi: 10.48550/arXiv .2306.06079. URLhttps://arxiv.org/abs/2306.06079

  18. [19]

    Bruinsma, Tom R

    Anna Allen, Stratis Markou, Will Tebbutt, James Requeima, Wessel P. Bruinsma, Tom R. Andersson, Michael Herzog, Nicholas D. Lane, Matthew Chantry, J. Scott 29 Hosking, and Richard E. Turner. End-to-end data-driven weather prediction.Nature, 641(8065):1172–1179, 2025. doi: 10.1038/s41586-025-08897-0

  19. [20]

    Brenowitz

    Peter Manshausen, Yair Cohen, Peter Harrington, Jaideep Pathak, Mike Pritchard, Piyush Garg, Morteza Mardani, Karthik Kashinath, Simon Byrne, and Noah D. Brenowitz. Generative data assimilation of sparse weather station observations at kilo- meter scales.Journal of Advances in Modeling Earth Systems, 17(10):e2024MS004505,

  20. [21]

    URLhttps://agupubs.onlinelibrary.wiley

    doi: 10.1029/2024MS004505. URLhttps://agupubs.onlinelibrary.wiley. com/doi/10.1029/2024MS004505

  21. [22]

    Huracan: A skillful end-to-end data-driven system for ensemble data assimilation and weather prediction

    Zekun Ni, Jonathan Weyn, Hang Zhang, Yanfei Xiang, Jiang Bian, Weixin Jin, Kit Thambiratnam, Qi Zhang, Haiyu Dong, and Hongyu Sun. Huracan: A skillful end-to- end data-driven system for ensemble data assimilation and weather prediction.arXiv preprint arXiv:2508.18486, 2025. doi: 10.48550/arXiv.2508.18486. URLhttps: //arxiv.org/abs/2508.18486

  22. [23]

    HealDA: Highlighting the importance of initial errors in end-to-end AI weather forecasts

    Aayush Gupta, Akshay Subramaniam, Michael S. Pritchard, Karthik Kashinath, Sergey Frolov, Kelsey Lieberman, Christopher Miller, Nicholas Silverman, and Noah D. Brenowitz. HealDA: Highlighting the importance of initial errors in end-to-end AI weather forecasts.arXiv preprint arXiv:2601.17636, 2026. doi: 10.48550/arXiv.2601. 17636. URLhttps://arxiv.org/abs/...

  23. [24]

    Learning accurate storm-scale evolution from observations.arXiv preprint arXiv:2601.17268, 2026

    Jaideep Pathak, Mohammad Shoaib Abbas, Peter Harrington, Zeyuan Hu, Noah Brenowitz, Suman Ravuri, Alberto Carpentieri, Jussi Leinonen, Corey Adams, Oliver Hennigh, Nicholas Geneva, Dale Durran, and Mike Pritchard. Learning accurate storm-scale evolution from observations.arXiv preprint arXiv:2601.17268, 2026. doi: 10.48550/arXiv.2601.17268. URLhttps://arx...

  24. [25]

    Red sky at night

    Anthony McNally, Christian Lessig, Peter Lean, Matthew Chantry, Mihai Alexe, and Simon Lang. Red sky at night... producing weather forecasts directly from observations. ECMWF Newsletter, pages 30–34, 2024. doi: 10.21957/tmc81jo4c7. URLhttps: //www.ecmwf.int/en/elibrary/81544-red-sky-night-producing-weather-forec asts-directly-observations. No. 178

  25. [26]

    Data driven weather forecasts trained and initialised directly from observations

    Anthony McNally, Christian Lessig, Peter Lean, Eulalie Boucher, Mihai Alexe, Ewan Pinnington, Matthew Chantry, Simon Lang, Chris Burrows, Marcin Chrust, Florian Pinault, Ethel Villeneuve, Niels Bormann, and Sean Healy. Data driven weather forecasts trained and initialised directly from observations.arXiv preprint arXiv:2407.15586, 2024. doi: 10.48550/arXi...

  26. [27]

    GraphDOP: Towards skilful data-driven medium-range weather forecasts learnt and initialised directly from observations

    Mihai Alexe, Eulalie Boucher, Peter Lean, Ewan Pinnington, Patrick Laloyaux, An- thony McNally, Simon Lang, Matthew Chantry, Chris Burrows, Marcin Chrust, Florian Pinault, Ethel Villeneuve, Niels Bormann, and Sean Healy. GraphDOP: Towards skilful data-driven medium-range weather forecasts learnt and initialised directly from obser- vations.arXiv preprint ...

  27. [28]

    Learning from nature: insights into GraphDOP's representations of the Earth System

    Peter Lean, Mihai Alexe, Eulalie Boucher, Ewan Pinnington, Simon Lang, Patrick Laloyaux, Niels Bormann, and Anthony McNally. Learning from nature: insights into graphdop’s representations of the earth system.arXiv preprint arXiv:2508.18018, 2025. doi: 10.48550/arXiv.2508.18018. URLhttps://arxiv.org/abs/2508.18018

  28. [29]

    Learning coupled earth system dynamics with graphdop.arXiv preprint arXiv:2510.20416, 2025

    Eulalie Boucher, Mihai Alexe, Peter Lean, Ewan Pinnington, Simon Lang, Patrick Laloyaux, Lorenzo Zampieri, Patricia de Rosnay, Niels Bormann, and Anthony Mc- Nally. Learning coupled earth system dynamics with graphdop.arXiv preprint arXiv:2510.20416, 2025. doi: 10.48550/arXiv.2510.20416. URLhttps://arxiv. org/abs/2510.20416

  29. [30]

    Using data assimilation tools to dissect graphdop.arXiv preprint arXiv:2510.27388, 2025

    Patrick Laloyaux, Mihai Alexe, Eulalie Boucher, Peter Lean, Ewan Pinnington, Simon Lang, Tobias Necker, and Anthony McNally. Using data assimilation tools to dissect graphdop.arXiv preprint arXiv:2510.27388, 2025. doi: 10.48550/arXiv.2510.27388. URLhttps://arxiv.org/abs/2510.27388

  30. [31]

    AIFS-DOP: End-to-End Medium-Range Weather Prediction from Observations Alone with Machine Learning

    Ewan Pinnington, Peter Lean, Mihai Alexe, Eulalie Boucher, Simon Lang, Patrick Laloyaux, Gert Mertes, Tomas Kral, Patricia de Rosnay, Matthew Chantry, and An- thony McNally. AIFS-DOP: End-to-end medium-range weather prediction from ob- servations alone with machine learning.arXiv preprint arXiv:2606.19093, 2026. doi: 10.48550/arXiv.2606.19093. URLhttps://...

  31. [32]

    Global atmospheric data assimilation with multi-modal masked autoencoders

    Thomas J. Vandal, Kate Duffy, Daniel McDuff, Yoni Nachmany, and Chris Hartshorn. Global atmospheric data assimilation with multi-modal masked autoencoders.arXiv preprint arXiv:2407.11696, 2024. doi: 10.48550/arXiv.2407.11696. URLhttps: //arxiv.org/abs/2407.11696

  32. [33]

    Earth-o1: A Grid-free Observation-native Atmospheric World Model

    Junchao Gong, Kaiyi Xu, Wangxu Wei, Siwei Tu, Jingyi Xu, Zili Liu, Hang Fan, Zhi- wang Zhou, Tao Han, Yi Xiao, Xinyu Gu, Zhangrui Li, Wenlong Zhang, Hao Chen, Xiaokang Yang, Yaqiang Wang, Lijing Cheng, Pierre Gentine, Wanli Ouyang, Feng Zhang, Zhe-Min Tan, Bowen Zhou, Fenghua Ling, Ben Fei, and Lei Bai. Earth-o1: A grid-free observation-native atmospheric...

  33. [34]

    Earth-o1: A Grid-free Observation-native Atmospheric World Model

    doi: 10.48550/arXiv.2605.06337. URLhttps://arxiv.org/abs/2605.06337

  34. [35]

    Skillful high-resolution weather forecasting independent of physical models

    Pengcheng Zhao, Siqi Xiang, Weixin Jin, Zekun Ni, Jiang Bian, Zuliang Fang, Hongyu Sun, Bin Zhang, Richard E. Turner, Jonathan Weyn, Haiyu Dong, Kit Thambiratnam, and Qi Zhang. Skillful high-resolution weather forecasting independent of physical models.arXiv preprint arXiv:2605.28153, 2026. doi: 10.48550/arXiv.2605.28153. URL https://arxiv.org/abs/2605.28153

  35. [36]

    Aifs single 1.1.0: an update to ecmwf’s machine-learned weather forecast model aifs.Geoscientific Model Development, 19:4703–4724, 2026

    GabrielMoldovan, EwanPinnington, AnaPrietoNemesio, SimonLang, ZiedBenBoual- lègue, Jesper Dramsch, Mihai Alexe, Mario Santa Cruz, Sara Hahner, Harrison Cook, Helen Theissen, Mariana Clare, Cathal O’Brien, Jan Polster, Linus Magnusson, Gert Mertes, Florian Pinault, Baudouin Raoult, Patricia de Rosnay, Richard Forbes, and Matthew Chantry. Aifs single 1.1.0:...

  36. [37]

    N. P. Wedi. Increasing the horizontal resolution in numerical weather prediction and cli- mate simulations: illusion or panacea?Philosophical Transactions of the Royal Society A, 372, 2014. doi: 10.1098/rsta.2013.0289

  37. [38]

    Hakim and Sanjit Masanam

    Gregory J. Hakim and Sanjit Masanam. Dynamical tests of a deep learning weather prediction model.Artificial Intelligence for the Earth Systems, 3, 2024. doi: 10.1175/ aies-d-23-0090.1

  38. [39]

    Mueller, Dong L

    Kevin J. Mueller, Dong L. Wu, Ákos Horváth, Veljko M. Jovanovic, Jan-Peter Muller, Larry Di Girolamo, Michael J. Garay, David J. Diner, Catherine M. Moroney, and Steve Wanzong. Assessment of MISR cloud motion vectors (CMVs) relative to GOES and MODIS atmospheric motion vectors (AMVs).Journal of Applied Meteorology and Climatology, 56(3):555–572, 2017. doi...

  39. [40]

    Global land surface atmospheric variables from 1718 to present from comprehensive in-situ observations, version 3.0.0, 2026

    Copernicus Climate Change Service. Global land surface atmospheric variables from 1718 to present from comprehensive in-situ observations, version 3.0.0, 2026. URL https://cds.climate.copernicus.eu/doi/10.24381/cds.cf5f3bac

  40. [41]

    J. K. Gibson, P. Kållberg, S. Uppala, A. Hernandez, A. Nomura, and E. Serrano. ERA description. ECMWF Re-Analysis Project Report Series 1, European Centre for Medium-Range Weather Forecasts, Shinfield Park, Reading, UK, July 1997. URL https://www.ecmwf.int/sites/default/files/elibrary/1997/9584-era-descrip tion.pdf

  41. [42]

    S. M. Uppala, P. W. KÅllberg, A. J. Simmons, U. Andrae, V. Da Costa Bechtold, M. Fiorino, J. K. Gibson, J. Haseler, A. Hernandez, G. A. Kelly, X. Li, K. Onogi, S. Saarinen, N. Sokka, R. P. Allan, E. Andersson, K. Arpe, M. A. Balmaseda, A. C. M. Beljaars, L. Van De Berg, J. Bidlot, N. Bormann, S. Caires, F. Chevallier, A. Dethof, M.Dragosavac, M.Fisher, M....

  42. [43]

    D. P. Dee, S. M. Uppala, A. J. Simmons, P. Berrisford, P. Poli, S. Kobayashi, U. Andrae, M.A.Balmaseda, G.Balsamo, P.Bauer, P.Bechtold, A.C.M.Beljaars, L.VanDeBerg, J. Bidlot, N. Bormann, C. Delsol, R. Dragani, M. Fuentes, A. J. Geer, L. Haimberger, S. B. Healy, H. Hersbach, E. V. Hólm, L. Isaksen, P. Kållberg, M. Köhler, M. Ma- tricardi, A. P. McNally, B...

  43. [44]

    Morice, John J

    Colin P. Morice, John J. Kennedy, Nick A. Rayner, Jonathan Winn, Emma Hogan, Rachel E. Killick, Timothy J. Osborn, Philip D. Jones, Ian R. Simpson, and Jonathan Harle. An updated assessment of near-surface temperature change from 1850: The had- crut5 data set.Journal of Geophysical Research: Atmospheres, 126(4):e2019JD032361,

  44. [45]

    doi: 10.1029/2019JD032361

  45. [46]

    Masayoshi Ishii, Asako Shouji, Satoshi Sugimoto, and Takao Matsumoto. Objective analyses of sea-surface temperature and marine meteorological variables for the 20th century using icoads and the kobe collection.International Journal of Climatology, 25 (7):865–879, 2005. doi: 10.1002/joc.1169

  46. [47]

    Nathan J. L. Lenssen, Gavin A. Schmidt, James E. Hansen, Matthew J. Menne, Allan Persin, Reto Ruedy, and David Zyss. Improvements in the gistemp uncertainty model. Journal of Geophysical Research: Atmospheres, 124(12):6307–6326, 2019. doi: 10.102 9/2018JD029522

  47. [48]

    Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson

    Ilan Price, Alvaro Sanchez-Gonzalez, Ferran Alet, Tom R. Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson. Probabilistic weather forecasting with machine learn- ing.Nature, 637(8044):84–90, January 2025. doi: 10.1038/s41586-024-08252-9. URL https://doi.org/10.1038/s4158...

  48. [49]

    HIRS Level 1C Fundamental Data Record Release 2 - Multimission - Global, 2024

    EUMETSAT. HIRS Level 1C Fundamental Data Record Release 2 - Multimission - Global, 2024. URLhttps://doi.org/10.15770/EUM_SEC_CLM_0036

  49. [50]

    NOAA Fundamental Cli- mate Data Record (FCDR) of MSU Level 1c Brightness Temperature, Version 1.0, 2013

    Cheng-Zhi Zou, Wenhui Wang, and NOAA CDR Program. NOAA Fundamental Cli- mate Data Record (FCDR) of MSU Level 1c Brightness Temperature, Version 1.0, 2013. URLhttps://doi.org/10.7289/V51Z429F. Accessed: 2026-01-30

  50. [51]

    SSM/T-2 Microwave Humidity Sounder Climate Data Record Release 1 - DMSP, 2020

    EUMETSAT. SSM/T-2 Microwave Humidity Sounder Climate Data Record Release 1 - DMSP, 2020. URLhttps://doi.org/10.15770/EUM_SEC_CLM_0046

  51. [52]

    Knapp, S

    Kenneth R. Knapp, S. Ansari, C. L. Bain, M. A. Bourassa, M. J. Dickinson, C. Funk, C. N. Helms, C. C. Hennon, C. D. Holmes, G. J. Huffman, J. P. Kossin, H.-T. Lee, A. Loew, and G. Magnusdottir. Globally gridded satellite (GridSat) observations for climate studies.Bulletin of the American Meteorological Society, 92:893–907, 2011. doi: 10.1175/2011BAMS3039....