{"id":"6907d5b3-caf7-40a8-b206-11c92b27fdb7","arxiv_id":"2607.11457","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"HourGlass probabilistically reconstructs hourly weather evolution between 6-hourly forecast states using CRPS training on NWP trajectories, preserving skill and small-scale variability better than deterministic downscalers.","lead":"HourGlass is a probabilistic ML model that fills in hourly weather states between 6-hourly data-driven forecasts, trained with CRPS losses on NWP trajectories. It keeps the skill of systems like AIFS and Bris while producing coherent, less-smoothed hourly fields for operational use.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The paper’s strongest claim is modest and well-supported: a CRPS-trained temporal downscaler applied to existing 6-hourly AI forecasts yields usable hourly products that retain skill and improve temporal/spatial realism relative to MSE interpolators or cubic splines. The two methodological novelties (probabilistic training with min/max/mean/difference aggregates; training on continuous NWP segments) are ablated and motivated by concrete failure modes of prior work (smoothing, ERA5 jumpiness). Verification is against independent observations, not reanalysis, and case studies of Storm Amy and US Plains convection show physically plausible evolution. The residual precipitation issues and mild boundary artifacts are disclosed rather than hidden. The reader correctly identified the NWP-target assumption as the softest point, yet the manuscript already supplies enough comparative evidence (ERA5 vs IFS training, MEPS-analysis vs Bris-forecast boundaries) that the assumption does not collapse the claim. No stronger load-bearing concern emerges from a second pass; therefore the ACCEPT verdict with high confidence stands unchanged.","tokens_in":18878,"tokens_out":553,"duration_ms":6779,"concrete_test":"Re-run the 2025 SYNOP verification of AIFS-HourGlass Single (Figs. 6, 8) after replacing the IFS-trained model with an otherwise identical model trained only on continuous ERA5 segments that deliberately straddle 09/21 UTC assimilation boundaries; if hourly RMSE or precipitation intensity distributions degrade by more than ~10 % relative to the published IFS-trained curves, the forecast-trajectory choice is load-bearing; otherwise the assumption is robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that HourGlass preserves upstream 6-hourly skill while producing temporally coherent hourly fields with realistic small-scale variability—is supported by the evidence presented. The reader’s weakest assumption (NWP forecast segments as training targets) is real but already partially stress-tested: Fig. 5 shows ERA5-only vs IFS-only training yields comparable RMSE; Fig. 14 and §3.2 openly document residual precipitation bias and boundary effects without overturning skill retention for T2m, wind and MSLP (Figs. 6–7, 11). Spectral diagnostics (Figs. 9–10), temporal-aggregate CRPS ablations (Figs. 6–8), and independent SYNOP verification further corroborate coherence and skill preservation. Precipitation extremes remain under-estimated, but this is acknowledged as a shared data-driven limitation rather than a flaw unique to the downscaler. No hidden inconsistency or untested premise undermines the operational bridging claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript introduces HourGlass, a probabilistic GNN-based temporal downscaler that reconstructs intermediate hourly states between two 6-hourly forecast anchors. Models are trained with almost-fair CRPS plus min/max/mean/difference aggregate terms (and a regional spectral term for Bris) on continuous NWP forecast segments (IFS and MEPS) rather than reanalysis, to avoid assimilation-window jumps. Two applications are presented: global AIFS-HourGlass (applied to AIFS-Single and AIFS-ENS) and regional stretched-grid Bris-HourGlass. Verification against independent 2025 SYNOP stations, spectral diagnostics of fields and temporal increments, loss ablations, cubic-spline and NWP baselines, and case studies of Storm Amy and US Southern Plains convection support the claim that skill of the upstream 6-hourly systems is retained while producing temporally coherent hourly trajectories with improved small-scale spatial variability. Precipitation extremes remain underestimated, which the authors acknowledge as a shared data-driven limitation.","tokens_in":19051,"tokens_out":1034,"duration_ms":9668,"significance":"If the results hold, HourGlass supplies a practical, operationally relevant bridge from the dominant 6-hourly data-driven forecast paradigm to the hourly products required by many applications, without the error accumulation of direct hourly forecasting or the smoothing of deterministic temporal interpolators. Strengths include independent SYNOP verification for 2025, spectral analysis of increments (Fig. 10), explicit loss ablations (Figs. 6–8), residual connections that leave 6-hour anchors unchanged, training on forecast trajectories to avoid ERA5 jumpiness (Figs. 4–5), dual global/regional demonstration, and open code in the Anemoi framework. These elements make the contribution concrete and reusable rather than purely methodological.","major_comments":[{"comment":"§3.2 and Fig. 14: When Bris-HourGlass is driven by Bris (or by MEPS analysis), mean hourly precipitation exhibits clear mid-window relaxation toward the training distribution and a dry bias inherited from the upstream model near the 6-hour anchors. The paper correctly notes that 6-hourly accumulated precipitation CRPS improves, but the residual boundary dependence and spin-up imprint remain load-bearing for the claim of seamless hourly products. A short quantitative summary of how much of the mid-window skill gain is pure spread versus genuine error reduction (already hinted at by Fig. 13) would strengthen the operational interpretation.","section":null},{"comment":"§2.2–2.3 and Fig. 5: The decision to train exclusively on IFS/MEPS forecast segments (lead times 12–29 h / 6–23 h) is well motivated by ERA5 jumpiness, yet the manuscript does not fully quantify how much parent-NWP bias or spin-up structure is transferred into the downscaler when it is later applied to analysis-trained AI forecasts. Fig. 5 shows comparable RMSE for ERA5-only vs IFS-only training on a few variables, but a parallel diagnostic for precipitation intensity distributions (analogous to Fig. 8) and for the regional model would make the weakest assumption more transparent.","section":null}],"minor_comments":[{"comment":"§2.1: The switch from a pure transformer processor (prior AIFS) to a graph transformer is stated but not motivated; a sentence on why this choice was made for multi-time-step decoding would help.","section":null},{"comment":"Table 1: Asterisks and daggers for diagnostic vs prognostic variables and for global-only fields are dense; a short legend or footnote clarifying which variables are residual-connected would improve readability.","section":null},{"comment":"Fig. 2 caption: “Hours 4-6 share the same random noise” is useful; stating explicitly that the same noise seed is used across the ablation would make the visual comparison clearer.","section":null},{"comment":"§3.1: The small RMSE jumps attributed to diurnal observation density are plausible; a brief note on the number of stations per hour (or a supplementary plot) would remove residual ambiguity.","section":null},{"comment":"Code availability: The GitHub link points to a feature branch; a commit hash or release tag would improve long-term reproducibility.","section":null},{"comment":"Typographical: “availabilty” in the Code availability heading; “Code availabilty” should be corrected.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a solid, well-executed contribution that fits operational ML-weather journals. The precipitation-boundary issue is real but already partially diagnosed by the authors; it does not overturn the central skill-retention claim for the core surface variables. I see no novelty or citation concerns that would require editorial intervention."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean methods paper that solves a real operational problem. Most AI weather models still spit out 6-hourly fields; services and impact models want hourly. Direct hourly training accumulates error and inherits reanalysis jumps. HourGlass instead takes the two bounding 6-hour states and fills the intermediate hours with a single forward pass that shares latent noise, so the trajectory stays coherent.\n\nWhat is actually new is the combination: almost-fair CRPS plus explicit min/max/mean/difference aggregates over the window, noise injection for probabilistic samples, simultaneous multi-hour decoding, and training exclusively on continuous IFS/MEPS forecast segments rather than ERA5. Prior temporal downscalers (Leinonen, Zhong) were deterministic MSE and used reanalysis; they smooth and jump. The residual connection that leaves the final anchor unchanged is a sensible engineering choice so skill at the 6-hour marks is not re-litigated.\n\nThey do the verification properly. Independent SYNOP for 2025, spectral diagnostics of increments (no artificial small-scale power at the hand-off), ablation of the temporal loss terms, cubic-spline and MEPS/IFS baselines, and two hard case studies (Storm Amy, Southern Plains convection). Skill for T2m, wind and MSLP is retained; precipitation spatial structure improves but extremes remain under-estimated, which they acknowledge as a shared data-driven limitation. Fig. 14 honestly shows residual precipitation bias and boundary effects when the upstream model is dry. Code and configs are public in Anemoi.\n\nSoft spots are real but secondary. Precipitation is still the weak variable; the NWP-trajectory training assumption is only partially stress-tested (ERA5 vs IFS RMSE looks similar, but bias inheritance is visible); free parameters (epsilon, spectral weight, tendency weights) are not exhaustively ablated. None of that breaks the central claim.\n\nThis is for people building or deploying AI NWP systems who need hourly products without re-training the whole forecaster. It deserves a serious referee. I would engage with it and expect it to be useful.","headline":"Solid operational methods paper: probabilistic CRPS temporal downscaling trained on continuous NWP trajectories, verified on independent SYNOP, with code released.","tokens_in":19840,"tokens_out":517,"would_cite":true,"duration_ms":5755,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"HourGlass reconstructs the hours between 6-hourly data-driven weather forecasts as skillful, temporally coherent probabilistic trajectories with realistic small-scale structure.","keywords":["temporal downscaling","data-driven weather forecasts","probabilistic forecasting","CRPS","hourly weather prediction","ensemble forecasting","graph neural networks"],"falsifier":"Hourly verification against surface stations would show larger errors or broken temporal-increment spectra inside each 6-hour window relative to cubic-spline interpolation of the same parent forecast, or storm case studies would exhibit unphysical jumps or loss of frontal structure at mid-window.","tokens_in":19700,"feed_emoji":"⏱️","tokens_out":926,"duration_ms":26639,"temperature":0.7,"pith_summary":"Most leading data-driven weather systems still step only every six hours, yet operations and extreme-event work need hourly detail. Direct hourly forecasting piles up error and can learn unphysical time correlations, so HourGlass instead reconstructs the intermediate hours from the two bounding forecast states. It is trained probabilistically with continuous-ranked-probability losses plus extra penalties on window min, max, mean and successive differences, so uncertainty appears as spread rather than as spatial smoothing, and successive hours evolve coherently. Training targets are continuous segments of numerical weather prediction forecast trajectories rather than reanalysis, avoiding assimilation-window jumps that earlier methods absorbed. Applied both globally and regionally, the models keep the skill of their parent 6-hourly systems, produce spectra of temporal increments that match numerical benchmarks, and track storms and organised convection in a physically consistent way, while intense hourly precipitation extremes remain underestimated.","feed_headline":"6-hour AI weather becomes hourly without losing skill","feed_subtitle":"Probabilistic reconstruction keeps small-scale realism and coherent storm timing for operations","key_machinery":"Almost-fair CRPS on each output hour, augmented by the same score applied to min, max, mean and successive differences over the six-hour window (plus a regional spectral term), with shared latent noise so one forward pass yields a single coherent trajectory; residual connections force the final hour of prognostic fields to match the parent forecast.","core_discovery":"A probabilistic temporal downscaler trained with almost-fair CRPS plus window-wise min/max/mean and difference terms on continuous NWP forecast segments can fill the hours between 6-hourly data-driven forecasts while preserving upstream skill, realistic small-scale spatial variability, and coherent physical evolution, rather than producing the temporally inconsistent smoothing of deterministic MSE-trained interpolators.","pith_inferences":["The same boundary-anchored residual design should generalise to other fixed temporal gaps (for example 3-hourly to hourly) provided the parent anchors remain skilful.","Diagnostic precipitation’s over-confidence near the window edges suggests multi-scale or boundary-aware probabilistic losses will be needed before hourly rain products match operational needs.","Training only on post-spin-up forecast segments may systematically under-represent the first hours after analysis, limiting direct use for nowcasting-style applications.","Once upstream AI models reduce dry bias and raise extreme rain rates, HourGlass should inherit those gains at hourly resolution with little additional training cost."],"forward_implications":["Existing 6-hourly data-driven forecast systems can supply operational hourly products without being retrained for hourly steps.","Skill gains that come from analysis-trained parent models (for example better 2 m temperature) can be carried into the intermediate hours even though the downscaler itself never saw analyses as targets.","Deterministic or MSE-trained temporal interpolators become inferior defaults whenever small-scale spatial realism and coherent storm timing matter.","Each ensemble member can be downscaled independently, preserving a probabilistic hourly product at modest extra cost.","Improvements to the 6-hourly parent forecasts, especially of intensity extremes, are expected to translate directly into better hourly fields."],"fun_headline_variants":["HourGlass turns 6-hour AI forecasts into coherent hourly weather","Probabilistic downscaler fills hours between AI weather states","HourGlass rebuilds hourly evolution while keeping upstream skill","Temporal downscaling yields hourly forecasts from 6-hour AI systems","Data-driven HourGlass bridges 6-hour outputs to hourly products"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"Continuous stretches of numerical weather prediction forecast trajectories teach the true physical hour-to-hour evolution without imprinting the parent model’s biases and spin-up so strongly that skill collapses when the downscaler is later applied to analysis-trained AI forecasts.","fun_headline_variants_meta":{"raw":{"variants":["HourGlass turns 6-hour AI forecasts into coherent hourly weather","Probabilistic downscaler fills hours between AI weather states","HourGlass rebuilds hourly evolution while keeping upstream skill","Temporal downscaling yields hourly forecasts from 6-hour AI systems","Data-driven HourGlass bridges 6-hour outputs to hourly products"]},"model":"grok-4.5","effort":"low","cost_usd":0.007442,"raw_usage":{"total_tokens":1831,"prompt_tokens":838,"num_sources_used":0,"completion_tokens":88,"cost_in_usd_ticks":74420000,"prompt_tokens_details":{"text_tokens":838,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":905,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":838,"tokens_out":88,"duration_ms":7700,"temperature":1.0,"reasoning_tokens":905,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T05:27:03.050004+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Hourly verification against surface stations would show larger errors or broken temporal-increment spectra inside each 6-hour window relative to cubic-spline interpolation of the same parent forecast, or storm case studies would exhibit unphysical jumps or loss of frontal structure at mid-window.","supporting_citations":[],"review_version":1}