{"id":"7cd51df4-d589-4c9a-ac3f-6305cb48b164","arxiv_id":"2508.15724","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"For record-breaking weather extremes, the numerical HRES model shows smaller forecast errors than GraphCast, Pangu-Weather, and Fuxi, which also underestimate the frequency and intensity of such events.","lead":"A new evaluation finds that Europe's high-resolution numerical weather model still beats leading AI forecasters on record-breaking heat, cold, and wind. The result matters because AI weather models are being proposed for early warning systems, but they appear to extrapolate poorly beyond their training data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim hinges on record-definition baseline and error metric; if records are defined from the same reanalysis used to train AI models, the 'extrapolation' conclusion is partly circular and metric-dependent.","rationale":"The reader's UNVERDICTED verdict is appropriate because only the abstract is available. My stress-test sharpens the reader's weakest assumption: the record definition and metric are not just missing details; they could be confounded with the AI models' training objective. The proposed test would resolve whether the observed AI underperformance reflects a fundamental limitation in forecasting extremes or an artifact of using the training distribution as the verification baseline and an L2-like error metric. I therefore keep the verdict as UNVERDICTED rather than UNCHANGED because the concern, while not disproving the claim, underscores that the current abstract does not provide enough information to reach any verdict. I set agreement_with_reader to 'partial' because the reader identified a similar assumption but I go further in specifying the circularity and metric-dependence.","tokens_in":695,"tokens_out":3402,"duration_ms":43909,"concrete_test":"Recompute the error comparison after redefining 'record-breaking' events relative to a climatological baseline that ends before the AI training period (e.g., 1979–2000 for models trained on ERA5 through 2018), and evaluate with a threshold-weighted CRPS instead of RMSE. If AI models no longer consistently underperform HRES across all lead times, the headline conclusion depends on the record baseline and metric.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The paper's central claim is that HRES consistently outperforms AI models for record-breaking extremes. This claim rests on two unverified choices: (1) how 'record-breaking' is defined (reference dataset and climatological baseline) and (2) which error metric is used. The abstract notes AI models 'underpredict hot records and overestimate cold records' with growing errors for larger exceedance—a signature of regression-to-the-mean under an L2 training loss when evaluated selectively on extremes. If the record baseline is ERA5 or a similar reanalysis that also served as training data for the AI models, then the test is partly circular: it asks whether AI models can output values outside the range of their training labels, rather than whether they are less skillful for extremes in general. An L2-based metric (e.g., RMSE) will heavily penalize a smooth model for missing an extreme, while a physical model like HRES may have different error correlations. The abstract provides no event counts, verification dataset, lead-time alignment details, or uncertainty intervals, so the possibility that the result is an artifact of metric choice, sample selection, or small-sample noise cannot be excluded. Without the full methods, the strongest claim is not load-bearing in its current form.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript addresses whether AI-based weather forecast models can extrapolate to record-breaking extremes, using the numerical HRES model as a benchmark. The abstract claims that HRES consistently outperforms five state-of-the-art AI models (GraphCast, GraphCast operational, Pangu-Weather, Pangu-Weather operational, and Fuxi) for record-breaking heat, cold, and wind across nearly all lead times, and that AI models underestimate both the frequency and intensity of such events, with errors growing for larger record exceedance. The abstract concludes that AI models are currently limited in high-stakes early-warning applications. No verification methodology, event definitions, sample sizes, uncertainty intervals, or lead-time details are presented in the abstract.","tokens_in":948,"tokens_out":3562,"duration_ms":40759,"significance":"If the central claim holds, the result has substantial practical significance for operational weather forecasting and for the safe deployment of AI emulators, particularly in disaster warning where record-breaking extremes are the most consequential. The paper is valuable for naming specific AI models and for focusing on extrapolation beyond the training distribution, an issue often overlooked in benchmark comparisons. However, the evaluation of significance is impossible from the abstract alone: the claim must be substantiated with a transparent verification protocol. The abstract gives no way to assess event counts, statistical confidence, or whether the result is robust to metric and reference-dataset choices.","major_comments":[{"comment":"The central claim, that HRES 'consistently outperforms' the AI models across 'nearly all lead times', lacks any supporting verification methodology. The abstract does not define how record-breaking events are identified, what observational or reanalysis reference is used, what error metric is computed, how many events are considered, or what the uncertainty intervals are. Without this information, the result could be an artifact of a particular metric or a small, non-representative sample. The full paper must report this information for the claim to be load-bearing.","section":"Abstract"},{"comment":"The comparison may be confounded by the training data used for AI models. If the reference baseline for defining record-breaking events is a reanalysis (e.g., ERA5) that also served as the training target for these models, then 'record-breaking' means that the verification target lies beyond the maximum label seen during training. The reported pattern—AI models underpredict hot records and overestimate cold records, with growing errors for larger exceedance—is a signature of regression to the mean under an L2 training loss when evaluated selectively on extremes. The authors should state the reference dataset explicitly and demonstrate that the conclusion holds when using an independent observational reference and alternative climatological baselines.","section":"Abstract"},{"comment":"The claim of 'growing errors for larger record exceedance' is reported without quantitative support or confidence intervals. If the number of record-breaking events is small, a few large outliers could dominate the averages. The abstract does not report event counts, the distribution of errors, or any significance testing. The full paper must show that the consistency across models and lead times is statistically robust rather than driven by a handful of events.","section":"Abstract"}],"minor_comments":[{"comment":"The term 'consistent' is vague. Does it mean that HRES outperforms on the majority of lead times, on all regions, or with statistical significance? A precise consistency criterion should be defined.","section":"Abstract"},{"comment":"The abstract lists both operational and non-operational versions of GraphCast and Pangu-Weather. The difference between these versions should be clarified, including whether the same initial conditions and lead times are used for fair comparison.","section":"Abstract"},{"comment":"The verification reference is not mentioned. Including a short phrase such as 'verified against ERA5 and station observations' would help the reader interpret the claims.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"This report is based solely on the abstract because the full text was not provided for review. The omission makes it impossible to verify the central claim. The potential circularity with the training reanalysis, if present in the full paper, would be a serious issue that needs to be addressed with an independent reference dataset and additional metrics. I recommend sending the manuscript back for full-text review before making a decision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper makes a clear, testable claim: for record-breaking heat, cold, and wind, ECMWF's HRES beats five state-of-the-art AI forecast models at nearly all lead times. That is worth knowing. Standard skill-score comparisons don't target extrapolation beyond the training domain, so this benchmark fills a real gap, and the finding is consequential for anyone thinking about AI models in early warning systems.\n\nWhat the paper does well, even from the abstract, is frame the question sharply and report the failure mode honestly: AI models underpredict hot records, overestimate cold records, and the errors grow with record exceedance. That is a concrete, falsifiable signature, and the authors don't oversell—they explicitly say further rigorous verification is needed before these models are relied on for high-stakes use.\n\nThe soft spot is that we're working from the abstract alone. The central claim depends on two load-bearing choices we can't see: how a \"record-breaking event\" is defined and which error metric is used. The stress-test note is right to flag the possible circularity: if the record baseline is the same reanalysis used to train the AI models, the test is less about skill and more about whether the models can output values outside their training labels. That's still a legitimate extrapolation test, but it's not the same as showing they're worse in general. Also, record-breaking events are rare by definition; without event counts and uncertainty intervals, small-sample noise could be doing work. And an L2-based metric will naturally hammer a smooth model on extremes—that's a real concern, not a manufactured one.\n\nI can't tell from the abstract whether those issues are handled properly. The paper might well do everything right; the caution is about what's missing, not a known flaw. If the full manuscript shows that records are defined from an independent reference or a holdout baseline, and that the metric comparison is robust to choice, then this is a solid result. If not, the strongest claim shrinks to a narrower statement about extrapolation limits under specific training conditions.\n\nWho should read this: anyone using or deploying AI weather models for extremes, and ML folks working on distribution shift in forecasting. It deserves a serious referee—the question is important and the comparative evidence is directly relevant. I'd send it to peer review rather than desk reject, with the explicit request that referees check the event definition, sample sizes, and metric sensitivity.\n\nFor the record: I'd bring it to a reading group because the topic is timely and the claim is checkable. I wouldn't cite it yet, but I would after seeing the full methods.","headline":"Plausible, well-scoped benchmark result that AI models trail HRES on record-breaking extremes, but the abstract alone can't rule out metric or baseline artifacts.","tokens_in":1337,"tokens_out":1277,"would_cite":false,"duration_ms":14823,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Against the five leading AI weather models, the operational numerical model HRES still produces smaller errors for record-breaking heat, cold, and wind extremes at nearly all lead times.","keywords":["AI weather forecasting","record-breaking extremes","numerical weather prediction","HRES","GraphCast","Pangu-Weather","extrapolation","forecast verification"],"falsifier":"A decisive test: initialise HRES, GraphCast, and Pangu-Weather from the same analysis time for a well-documented record event not in the AI training data (e.g., the June 2021 Pacific Northwest heatwave) and compare 2-metre temperature forecast errors at the record location across lead times. The paper's claim predicts the AI models' errors are larger than HRES's and grow with record exceedance; observing the opposite would refute it.","tokens_in":639,"feed_emoji":"🌡️","tokens_out":5312,"duration_ms":49204,"temperature":0.7,"pith_summary":"This paper claims that the supposed AI revolution in weather forecasting does not yet extend to the most dangerous events: record-breaking heat, cold, and wind. Comparing five state-of-the-art AI models (GraphCast, GraphCast operational, Pangu-Weather, Pangu-Weather operational, and Fuxi) against the European Centre's high-resolution numerical model HRES, the authors find that HRES consistently produces smaller forecast errors for record-breaking extremes across nearly all lead times. The AI models also underestimate the frequency and intensity of these records, underpredict hot records and overpredict cold records, with errors growing as the record exceedance gets larger. If true, the result identifies a systematic extrapolation failure: AI models trained on past weather cannot generalize beyond the largest events in their training data. This matters because record-breaking events are becoming more frequent in a warming climate and are exactly the events that trigger early warnings and disaster response.","feed_headline":"AI weather models lose to HRES on record-breaking extremes","feed_subtitle":"State-of-the-art AI models under-predict the frequency and intensity of record heat, cold, and wind.","key_machinery":"The central object is the record-breaking extreme: a weather event that exceeds the previous observed record at a given location and time of year for temperature or wind. The key mechanism is to condition forecast-verification statistics on whether the observation sets a new record, and on how much it exceeds the old record, instead of averaging over all weather. This conditioning exposes the AI models' tendency to underpredict the frequency and intensity of the most extreme events, with errors growing as the record exceedance grows.","core_discovery":"Using a unified verification framework, the authors show that when forecasts are evaluated on record-breaking extremes rather than on all weather, the numerical model HRES outperforms all five AI models. The AI models' forecast errors are larger for record-breaking heat, cold, and wind at nearly all lead times; the AI models tend to underforecast both how often records occur and how far records are exceeded; and they show an asymmetric bias: hot records are underpredicted and cold records overpredicted, with errors increasing with the magnitude of the record exceedance. The authors interpret this as evidence that AI models extrapolate poorly beyond their training domain, and conclude that th","pith_inferences":["Beyond the paper: the asymmetric bias (underpredicting hot records, overpredicting cold records) is consistent with AI models anchoring to the mean of their training distribution; a direct test would compare forecast bias to the climatological anomaly distribution in each model's training data.","Beyond the paper: because AI models are trained on historical data, their extrapolation failure may shrink as training sets are updated to include recent record events; the paper's findings may describe the current generation of training data rather than an eternal limitation.","Beyond the paper: the result implies a public-safety standard: AI forecasts should be used only where their tail error profile is independently verified, not on the basis of average skill metrics."],"forward_implications":["AI weather models should not be deployed as stand-alone forecast systems for extreme-event early warning until their extrapolation behaviour is redesigned and verified.","Aggregate benchmark skill (e.g., RMSE over all dates) can conceal a systematic failure on the tail of the distribution, so operational readiness must be evaluated on record-breaking subsets.","Hybrid systems that use AI to post-process or emulate numerical models may be safer than pure AI forecasts for extremes, because the numerical component anchors the tail.","Because climate change makes hot records more frequent, the AI models' underprediction of hot records will cause more missed warnings as warming continues.","Growing error with record exceedance magnitude means the severity of the most extreme events will be underestimated most, which can understate impact."],"supporting_citations":[],"fun_headline_variants":["Numerical models outclass AI on record-breaking weather","AI weather models fall short on record-breaking extremes","Record-breaking events expose AI forecast weaknesses","AI misses mark on record heat, cold, wind extremes","For unprecedented weather, numerical models beat AI"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The comparison assumes that record-breaking events are defined and identified identically for every model, using the same verification reference and event-detection rules, and that the AI models' error metrics are not inflated by their training climatology; the abstract provides no detail on the event definition, verification data, or uncertainty intervals.","fun_headline_variants_meta":{"raw":{"variants":["Numerical models outclass AI on record-breaking weather","AI weather models fall short on record-breaking extremes","Record-breaking events expose AI forecast weaknesses","AI misses mark on record heat, cold, wind extremes","For unprecedented weather, numerical models beat AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000498,"raw_usage":{"total_tokens":2272,"prompt_tokens":733,"completion_tokens":1539,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":477,"completion_tokens_details":{"reasoning_tokens":1468}},"tokens_in":477,"tokens_out":1539,"duration_ms":17009,"temperature":1.0,"reasoning_tokens":1468,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:41:34.357111+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test: initialise HRES, GraphCast, and Pangu-Weather from the same analysis time for a well-documented record event not in the AI training data (e.g., the June 2021 Pacific Northwest heatwave) and compare 2-metre temperature forecast errors at the record location across lead times. The paper's claim predicts the AI models' errors are larger than HRES's and grow with record exceedance; observing the opposite would refute it.","supporting_citations":[],"review_version":1}