Pith. sign in

REVIEW 4 major objections 5 minor 16 references

Improving Tropical Cyclone Forecasting With Video Diffusion Models

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Video diffusion models that generate ten storm frames at once outperform frame-by-frame diffusion for tropical cyclone forecasting, extending the reliable horizon from 36 to 50 hours.

desk verdict Useful workshop paper on video diffusion for TC forecasting, but the headline 36-to-50-hour horizon gain rests on eyeballing SSIM curves without a defined threshold. read the letter →

arxiv 2501.16003 v5 pith:4PBHMC37 submitted 2025-01-27 cs.CV physics.ao-ph

classification cs.CVphysics.ao-ph
keywords videodiffusionmodelstropicalcycloneforecastingtemporalcoherencetwo-stagetrainingnowcastingFréchetDistancesatelliteimagerygenerativeweatherprediction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to show that tropical cyclone nowcasting improves when a diffusion model generates several forecast frames at once instead of predicting frames independently. It applies a video diffusion model with temporal layers to satellite and reanalysis inputs, and reports that it beats the prior frame-by-frame diffusion baseline by 19.3% in MAE, 16.2% in PSNR, 36.1% in SSIM, and 45.6% in Fréchet Video Distance while keeping single-frame FID roughly unchanged. If the results hold, emergency planners gain a longer reliable warning window, because the claimed forecast horizon grows from 36 to 50 hours, and low-data basins get a more efficient training recipe. The paper uses a two-stage curriculum—single-frame supervision before multi-frame training—because it finds that this preserves frame quality while learning motion.

What carries the argument

The load-bearing object is a video diffusion model built on a 3D U-Net with temporal convolutions and temporal attention, following the Video Diffusion Models design of [11], with classifier-free guidance and dynamic thresholding. It denoises a noisy 10-frame sequence conditioned on the first observed IR frame and ERA5 meteorological fields, learning the joint distribution of all frames rather than a per-frame marginal. Two-stage training first runs single-frame denoising (100–200 epochs) to anchor spatial structure, then shifts to 10-frame sequences (200–300 epochs) to learn temporal dynamics. Fréchet Video Distance serves as the new evaluation metric that is sensitive to temporal coherence.

What would settle it

Fix a threshold such as SSIM $= 0.6$, recompute the hourly SSIM curves for every cyclone in the test split, and measure the last hour at which the video model and the baseline each stay above it; the 36-to-50-hour claim fails if the two crossing hours do not differ by about 14 hours. A second check would replace the 10-frame inference with a single-frame cascade and confirm that the quality drop returns.

Watch

Extended reading notes

Core claim

The central claim is that treating cyclone evolution as a video generation problem yields forecasts that are both more accurate and more temporally coherent than treating it as a collection of independent stills. A 3D U-Net video diffusion model, conditioned on the first infrared frame and ERA5 fields, generates ten 10.8 µm IR frames simultaneously. In the test split of 335 ten-frame clips, the video model reports MAE $0.1781$ versus $0.2209$, PSNR $26.13$ versus $22.49$, SSIM $0.7123$ versus $0.5235$, and FVD $242.41$ versus $445.83$ for the prior baseline; FID is essentially tied. The paper further claims that when inserted into the prior cascaded long-horizon pipeline, the reliable forecast horizon grows from 36 to 50 hours, with minimum SSIM values staying above the baseline for each cyclone.

Load-bearing premise

The claim of a longer reliable horizon rests on reading 'acceptable' forecast quality from SSIM charts, but the paper never fixes a numerical SSIM threshold before inspecting those charts.

Editorial extensions

If this is right

  • Forecasters could issue warnings roughly 14 hours earlier than the previous model allowed, assuming the reported reliability holds.
  • The two-stage training recipe gives a concrete path for applying diffusion models to other small weather datasets, not just cyclones.
  • The large FVD reduction suggests the video model is less likely to generate temporally inconsistent or flickering cloud structures over time.
  • Because FID stays about the same while other metrics improve, the gains come from temporal structure rather than from sharper individual frames.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 36-to-50-hour horizon is only as strong as the SSIM cutoff used to define 'reliable,' and the paper chooses that cutoff by inspecting charts rather than fixing a threshold in advance; a test with a preset numerical threshold could shrink the gain.
  • Part of the improvement in MAE and SSIM may come from the model producing smoother, temporally averaged motion rather than from genuinely better cyclone physics; checking against a physics-based consistency metric would separate the two.
  • The two-stage curriculum is inexpensive and basin-agnostic, so the same recipe should transfer to other sparse weather-video tasks such as precipitation nowcasting in regions with short radar records.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a video diffusion model for tropical cyclone (TC) nowcasting, extending the frame-by-frame diffusion baseline of Nath et al. by adding temporal layers and generating 10 frames simultaneously. A two-stage training curriculum (single-frame pretraining followed by multi-frame training) is introduced to preserve single-frame quality. Experiments on a satellite/ERA5 dataset report large improvements in MAE, PSNR, SSIM, and FVD over the baseline on a held-out test split, as well as an extension of the 'reliable forecasting horizon' from 36 to 50 hours. The authors also propose FVD as a metric for evaluating temporal coherence in TC forecasts and release code.

Significance. If the results hold, this is a useful demonstration that video diffusion models provide better temporal consistency than frame-by-frame diffusion for TC nowcasting, with an interesting two-stage training trick for low-data settings. The paper ships open code, evaluates on a held-out test set, and reports a wide range of metrics, including the temporal-coherence metric FVD, which is a strength. However, the headline long-horizon claim (36 to 50 hours) is not quantitatively defined, and the main comparison lacks uncertainty quantification. The significance is therefore conditional: the core methodology is plausible and likely of interest to the climate-ML community, but the central quantitative claim about horizon extension needs an operational definition before the contribution can be fully assessed.

major comments (4)
  1. [Section 3.3 and Appendix C] The central claim of extending the reliable forecasting horizon from 36 to 50 hours is not operationalized. The text says Nath et al. identified 36 hours from 'sharp declines in forecast accuracy, as indicated by SSIM charts,' and that the current model extends this to 50 hours 'as shown by SSIM charts in Appendix C.' However, no threshold or rule is ever specified: the paper does not define a minimum acceptable SSIM value, a drop magnitude, or a sustained-decline criterion. The dashed lines in Figure C.2 only mark the hourly point of minimum SSIM for each cyclone, not a pre-specified acceptability level. As written, the numbers 36 and 50 appear to be chosen after seeing the curves, which would make the claimed 14-hour gain a post hoc interpretation of a vertical shift in SSIM rather than a measured extension of a meaningful horizon. Please provide an explicit, pre-registered or at least pre-specified rule (e.g., 'the horizon is the first hour where SSIM falls below γ' or 'the first hour of a sustained δ drop from the peak SSIM'), apply the same rule to both models, and report the resulting horizons and their sensitivity to the choice of γ or δ.
  2. [Table 2, Section 3.2] The main comparison in Table 2 is presented without any uncertainty quantification. With a single test split and no seeds, the reported differences (MAE 0.2209 vs 0.1781; PSNR 22.49 vs 26.13; SSIM 0.5235 vs 0.7123) cannot be distinguished from training stochasticity, and the 2.4% FID degradation is only acknowledged in passing. Please report means and standard deviations across multiple seeds (or, at minimum, bootstrapped confidence intervals on the test set) for both Table 1 and Table 2, and discuss whether the FID difference is within noise. Without this, the strength of the claimed improvements is not established.
  3. [Section 3.1, Table 1] The paper states that two-stage training 'significantly improves individual frame quality,' but the evidence in Table 1 is mixed. While FID improves dramatically (1.2633 to 0.4955), PSNR and SSIM are slightly worse for the two-stage variant (20.72 vs 20.62 and 0.6522 vs 0.6387, respectively). Since PSNR and SSIM are standard single-frame quality metrics, the claim as stated is not uniformly supported. Please restrict the claim to FID (or distribution-level fidelity) or show that the PSNR/SSIM differences are within noise.
  4. [Section 3.3, long-horizon protocol] The long-horizon experiment is not reproducible from the description. The text says the model is integrated into the authors' cascaded pipeline and used to 'forecast the entire duration of all cyclones,' but it does not describe how the cascade is initialized beyond the first 10-frame block, how the conditioning frames are obtained at subsequent steps, or whether the baseline (Nath et al.) uses the identical cascade protocol in this comparison. These details are necessary to confirm that the 36- vs 50-hour difference is due to the model architecture and not to a difference in how the cascade is run.
minor comments (5)
  1. [Figure C.2 caption] The caption says the dashed lines mark 'the hourly marks at which the minimum SSIM values are obtained,' while Section 3.3 uses these charts to claim the reliable horizon. Please clarify whether the dashed lines represent the proposed reliable-horizon cutoff or simply the minimum-SSIM hour, as the two are not the same.
  2. [Figure 1] The 'difference' rows in Figure 1 are shown without a colormap or scale, making it impossible to judge the magnitude of the errors visually. Please add a colorbar and state the intensity scale.
  3. [Appendix B.1] The hyperparameter table lists batch size, sequence length, learning rate, guidance scale, and epoch splits, but omits diffusion-specific details such as the number of diffusion timesteps, noise schedule, and UNet channel dimensions. Please add these for reproducibility.
  4. [Section 2.1] The data-processing section states that NaN values are replaced with zeros and a mask is applied, but it does not describe how the mask is used during training and inference. Please specify the masking mechanism.
  5. [References] The baseline Nath et al. [10] is cited by arXiv number; please provide the version and, if applicable, the official publication venue, since the comparison depends on the exact code and checkpoint used.

Circularity Check

1 steps flagged · score 2.0 of 10

Held-out metric gains are independent and not circular; the 'reliable horizon' claim is under-specified because 'reliable' is read from the same SSIM curves that define the claimed extension.

  1. self definitional [Section 3.3 (Long-Horizon Forecasting); also Introduction Contribution 3]
    "Nath et al. [10] identified a reliable forecasting horizon of 36 hours, beyond which sharp declines in forecast accuracy, as indicated by SSIM charts, are observed. Integrating our model into their cascaded pipeline extends this reliable horizon from 36 to 50 hours, as shown by SSIM charts (see charts in Appendix C)."

    The claimed output is the 'reliable forecasting horizon,' but 'reliable' is defined in the same passage by the shape of the model's own SSIM curves: a horizon is where sharp SSIM declines begin, and the extension is demonstrated by those same charts. No fixed threshold is specified (no numerical SSIM floor, drop magnitude, or sustained-decline rule), so the horizon is not a measurable quantity independent of the curves. The 36-hour anchor is imported by self-citation from the authors' prior work, and the 50-hour endpoint is read post hoc from the new model's generally higher SSIM curves.

full rationale

The core comparisons are external evaluations on a held-out test split: the MAE, PSNR, SSIM, FID, and FVD numbers in Tables 1 and 2 are computed from model outputs, and none is obtained by plugging a fitted parameter back into its own definition. The two-stage training claim is a curriculum choice whose benefit is tested by the same metrics, not guaranteed by construction. FVD is an established video metric rather than a renamed version of the authors' own outputs. The baseline is prior work by two of the current authors, so self-citation is present, but the paper re-runs comparable experiments and does not rely on the baseline's numbers except for the 36-hour anchor. The one genuinely fragile step is the reliable-horizon claim in Section 3.3: because 'reliable' is never operationally defined and is read from the same SSIM charts that produce the headline number, the 36-to-50 extension is not quantitatively established. That is a measurement-definition weakness and it is tied to self-citation of the 36-hour result, but it does not make the metric-table improvements circular. Hence a score of 2 rather than 0.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

No new physical entities or forces are introduced; the model and training recipe are composed from existing components. The central claim rests on standard diffusion machinery, several domain assumptions about the data and metrics, and tuned hyperparameters.

free parameters (4)
  • Classifier-free guidance scale = 3.0
    Chosen by hand for both data regimes; no sensitivity analysis is reported, and it strongly affects sample quality.
  • Stage-1/stage-2 epoch split = 100/300 (full data), 200/200 (low data)
    The split was tuned after observing results (Section 2.3), so the reported gains depend on this hand-chosen schedule.
  • Video sequence length = 10 frames
    The model generates exactly 10 frames per pass; no ablation on sequence length is provided.
  • Learning rate = 3e-4
    Fixed hyperparameter for all runs; a standard choice but not derived or tuned.
assumptions (5)
  • standard math Diffusion denoising and classifier-free guidance work as described in Ho et al. when implemented as a 3D UNet.
    The paper inherits the correctness of the diffusion framework without deriving it.
  • domain assumption Ten consecutive IR frames with NaN values replaced by zeros and a mask are a faithful representation of cyclone evolution for training.
    Section 2.1 states this preprocessing; no validation is given that it preserves physical fields.
  • ad hoc to paper SSIM curves with an unspecified sharp-decline threshold define the reliable forecast horizon.
    Section 3.3 and Appendix C use visual markers for minimum SSIM; the threshold is not defined and appears selected for the 36-to-50-hour claim.
  • domain assumption FVD based on I3D features is a valid metric for temporal coherence of TC forecasts.
    The paper asserts FVD is more suitable but gives no comparison to other temporal metrics.
  • domain assumption The train/test split of the 1,092 and 335 video sequences prevents temporal leakage between clips.
    Section 2.1 reports sequence counts but does not state whether splits are by cyclone event or whether overlapping clip windows cross the train/test boundary.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Improving Tropical Cyclone Forecasting With Video Diffusion Models." pith.science (2026). https://pith.science/paper/4PBHMC37

@misc{pith2026250116003,
  author       = {Pith},
  title        = {Pith review of: Improving Tropical Cyclone Forecasting With Video Diffusion Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4PBHMC37}},
  note         = {Machine review of arXiv:2501.16003}
}
read the original abstract

Tropical cyclone (TC) forecasting is crucial for disaster preparedness and mitigation. While recent deep learning approaches have shown promise, existing methods often treat TC evolution as a series of independent frame-to-frame predictions, limiting their ability to capture long-term dynamics. We present a novel application of video diffusion models for TC forecasting that explicitly models temporal dependencies through additional temporal layers. Our approach enables the model to generate multiple frames simultaneously, better capturing cyclone evolution patterns. We introduce a two-stage training strategy that significantly improves individual-frame quality and performance in low-data regimes. Experimental results show our method outperforms the previous approach of Nath et al. by 19.3% in MAE, 16.2% in PSNR, and 36.1% in SSIM. Most notably, we extend the reliable forecasting horizon from 36 to 50 hours. Through comprehensive evaluation using both traditional metrics and Fr\'echet Video Distance (FVD), we demonstrate that our approach produces more temporally coherent forecasts while maintaining competitive single-frame quality. Code accessible at https://github.com/Ren-creater/forecast-video-diffmodels.

Figures

Figures reproduced from arXiv: 2501.16003 by the authors.

Figure 1
Figure 1. Qualitative comparison of TC forecasting results on the first four frames generated. From [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

16 extracted references · 9 canonical work pages

  1. [1]

    Tropical Cy- clones and Climate Change Assessment: Part II: Projected Response to Anthropogenic Warming

    Knutson T, Camargo SJ, Chan JCL, Emanuel K, Ho CH, Kossin J, et al. Tropical Cy- clones and Climate Change Assessment: Part II: Projected Response to Anthropogenic Warming. Bulletin of the American Meteorological Society. 2020 Mar;101(3):E303-22. Available from: https://journals.ametsoc.org/view/journals/bams/101/ 3/bams-d-18-0194.1.xml

  2. [2]

    The social costs of tropical cyclones

    Krichene H, V ogt T, Piontek F, Geiger T, Schötz C, Otto C. The social costs of tropical cyclones. Nature Communications. 2023 Nov;14(1):7294. Available from: https://www.nature. com/articles/s41467-023-43114-4

  3. [3]

    Response of Global Tropical Cyclone Activity to Increasing CO2: Re- sults from Downscaling CMIP6 Models

    Emanuel K. Response of Global Tropical Cyclone Activity to Increasing CO2: Re- sults from Downscaling CMIP6 Models. Journal of Climate. 2021 Jan;34(1):57-70. Available from: https://journals.ametsoc.org/view/journals/clim/34/ 1/jcliD200367.xml

  4. [4]

    Weather Forecasting Using GPU-Based Large-Eddy Simulations

    Schalkwijk J, Jonker HJJ, Siebesma AP, Van Meijgaard E. Weather Forecasting Using GPU-Based Large-Eddy Simulations. Bulletin of the American Meteorological Society. 2015 May;96(5):715-23. Available from: https://journals.ametsoc.org/doi/ 10.1175/BAMS-D-14-00114.1

  5. [5]

    TCP-Diffusion: A Multi-modal Diffusion Model for Global Tropical Cyclone Precipitation Forecasting with Change Awareness

    Huang C, Mu P, Bai C, Watson PA. TCP-Diffusion: A Multi-modal Diffusion Model for Global Tropical Cyclone Precipitation Forecasting with Change Awareness. arXiv; 2024. ArXiv:2410.13175. Available from: http://arxiv.org/abs/2410.13175

  6. [6]

    A Survey on Deep Learning: Algorithms, Techniques, and Applications

    Pouyanfar S, Sadiq S, Yan Y , Tian H, Tao Y , Reyes MP, et al. A Survey on Deep Learning: Algorithms, Techniques, and Applications. ACM Computing Surveys. 2019 Sep;51(5):1-36. Available from: https://dl.acm.org/doi/10.1145/3234150

  7. [7]

    An Ensemble Machine Learn- ing Approach for Tropical Cyclone Detection Using ERA5 Reanalysis Data

    Accarino G, Donno D, Immorlano F, Elia D, Aloisio G. An Ensemble Machine Learn- ing Approach for Tropical Cyclone Detection Using ERA5 Reanalysis Data. arXiv; 2023. ArXiv:2306.07291. Available from: http://arxiv.org/abs/2306.07291

  8. [8]

    Cyclone track forecasting based on satellite images using artificial neural networks

    Kovordányi R, Roy C. Cyclone track forecasting based on satellite images using artificial neural networks. ISPRS Journal of Photogrammetry and Remote Sensing. 2009 Nov;64(6):513-

Show all 16 references
  1. [9]

    Latent diffusion models for generative precipitation nowcasting with accurate uncertainty quantification

    Leinonen J, Hamann U, Nerini D, Germann U, Franch G. Latent diffusion models for generative precipitation nowcasting with accurate uncertainty quantification. arXiv; 2023. ArXiv:2304.12891. Available from: http://arxiv.org/abs/2304.12891

  2. [10]

    Forecasting Tropical Cyclones with Cascaded Diffusion Models

    Nath P, Shukla P, Wang S, Quilodrán-Casas C. Forecasting Tropical Cyclones with Cascaded Diffusion Models. arXiv; 2024. ArXiv:2310.01690. Available from: http://arxiv.org/ abs/2310.01690

  3. [11]

    Video Diffusion Models

    Ho J, Salimans T, Gritsenko A, Chan W, Norouzi M, Fleet DJ. Video Diffusion Models. arXiv

  4. [12]

    Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

    Blattmann A, Dockhorn T, Kulal S, Mendelevitch D, Kilian M, Lorenz D, et al.. Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets. arXiv; 2023. ArXiv:2311.15127. Available from: http://arxiv.org/abs/2311.15127

  5. [13]

    Towards Accurate Generative Models of Video: A New Metric & Challenges

    Unterthiner T, Steenkiste Sv, Kurach K, Marinier R, Michalski M, Gelly S. Towards Accurate Generative Models of Video: A New Metric & Challenges. arXiv; 2019. ArXiv:1812.01717. Available from: http://arxiv.org/abs/1812.01717

  6. [14]

    Tackling Climate Change with Machine Learning

    Ho J, Salimans T. Classifier-Free Diffusion Guidance. arXiv; 2022. ArXiv:2207.12598. Avail- able from: http://arxiv.org/abs/2207.12598. 6 Published as a workshop paper at "Tackling Climate Change with Machine Learning", ICLR 2025 APPENDIX A M ODEL ARCHITECTURE Figure A.1: Illu...

  7. [21]

    Available from: https://linkinghub.elsevier.com/retrieve/pii/ S0924271609000434

  8. [2022]

    Available from: http://arxiv.org/abs/2204.03458

    ArXiv:2204.03458. Available from: http://arxiv.org/abs/2204.03458

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.