Pith. sign in

REVIEW 3 major objections 3 minor 88 references

RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts

T0 review · 3 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper claims that RadarQA, a domain-tuned multi-modal language model trained on the RQA-70K dataset, outperforms general multimodal language models across all tested settings of radar forecast quality analysis.

desk verdict Abstract-only look at a plausible MLLM forecast QA benchmark; the contribution is real but the central validity claims are uncheckable from the abstract alone. read the letter →

arxiv 2508.12291 v1 pith:UDL22MJG submitted 2025-08-17 cs.AI

classification cs.AI
keywords weatherradarforecastsmulti-modallargelanguagemodelsforecastqualityanalysisratingandassessmentRQA-70Kdatasethybridannotationpipelinemulti-stagetraining
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that a multi-modal large language model can perform credible quality analysis of weather radar forecasts, offering both numeric ratings and written assessments for single frames and for whole forecast sequences. It reports that a domain-tuned model, RadarQA, trained on a new 70,000-sample dataset called RQA-70K, outperforms general-purpose MLLMs in every evaluation setting that was tested. A sympathetic reader should care because current forecast verification produces numbers but little explanation; if the claim holds, automated evaluation could describe what went wrong in a forecast and why, in terms a forecaster can use.

What carries the argument

The central machinery is a combination of three objects: a task paradigm that crosses single-frame versus sequence input with rating versus assessment output to define four quality-analysis settings; RQA-70K, a dataset of roughly 70,000 radar-forecast quality annotations produced by human experts plus automated heuristics and spanning easy to hard cases; and RadarQA itself, an MLLM trained with a multi-stage strategy in which each stage iteratively improves performance before the next begins. The load-bearing mechanism is the interaction between the hybrid labels and that training schedule: the model learns to ground its numeric ratings and written assessment reports in physical attributes of the radar forecasts rather than in generic language priors.

What would settle it

Have an independent panel of operational meteorologists, blinded to the original annotations, re-score a random sample of the radar forecasts in RQA-70K; if their labels agree with the dataset's hybrid labels no better than chance, the ground truth is not measuring forecast quality and the outperformance claim loses its foundation.

Watch

Extended reading notes

Core claim

The paper's central claim is that RadarQA, a multi-modal large language model fine-tuned on the RQA-70K dataset, outperforms existing general MLLMs across all evaluation settings for multi-modal forecast quality analysis, covering single-frame and sequence inputs and both rating and open-ended assessment outputs. The dataset is constructed through a hybrid annotation pipeline that combines human expert labeling with automated heuristics and is built to include varying difficulty levels. On the paper's own terms, this demonstrates that domain-specific MLLMs can give weather forecast verification descriptive, interpretable quality analysis that goes beyond score-based metrics, integrating key physical attributes of the radar forecasts into the model's ratings and reports.

Load-bearing premise

The result rests on a single assumption: the quality labels created by combining human expert judgment with automated rules are trustworthy, and if those labels are biased or inconsistent, both the dataset and the model built on it will not actually measure forecast quality.

Editorial extensions

If this is right

  • If the claim is right, forecast verification can offer readable explanations of why a radar forecast was good or poor, not just a single numeric error score.
  • RQA-70K provides a shared benchmark for future multi-modal quality analysis of weather radar forecasts.
  • The multi-stage training result implies that domain-specific data and staged fine-tuning are worthwhile for applying MLLMs to meteorological tasks.
  • RadarQA's output format combines ratings with written assessments, so it could be used to flag specific failure modes in radar forecasts in a way forecasters can scan quickly.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the hybrid annotation pipeline could be transferred to other forecast products, such as precipitation nowcasts, satellite-based retrievals, or climate model output, although the paper does not test those settings.
  • A test the paper leaves implicit is whether RadarQA's ratings and written assessments correlate with established skill scores on the same events, such as the critical success index or fractions skill score.
  • Because the annotation style used for training and evaluation is the same, an independent human study comparing RadarQA's reports with general MLLM reports on fresh forecasts would clarify whether the advantage generalizes beyond the RQA-70K label distribution.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper introduces RadarQA, a multi-modal large language model (MLLM) for weather radar forecast quality analysis. It defines a task paradigm covering single-frame and sequence inputs under rating and assessment scenarios, constructs a large dataset RQA-70K through a hybrid human-expert and automated-heuristic annotation pipeline, and proposes a multi-stage training strategy. The central claim, stated in the abstract, is that RadarQA outperforms existing general MLLMs across all evaluation settings, implying that domain-specialized MLLMs can provide descriptive forecast-quality assessments beyond score-based metrics.

Significance. If the central claim holds, the work could advance weather forecast verification by adding interpretable, descriptive, and dynamic-evolution-aware assessments to traditional score-based metrics. The proposed dataset and task paradigm would be useful resources for the community, and the multi-stage training approach could inform future domain-specific MLLM development. The authors are to be credited for explicitly proposing a new task paradigm and for assembling a large-scale dataset, though the abstract alone does not yet provide the quantitative evidence needed to assess the strength of these contributions.

major comments (3)
  1. [Abstract] The central claim that 'RadarQA outperforms existing general MLLMs across all evaluation settings' is not accompanied by any quantitative result, such as accuracy, correlation, or agreement scores, nor by error bars, significance tests, or a description of the evaluation protocol. Because this claim is the paper's main load-bearing assertion, the abstract must at least report the headline numbers and state, for example, how many evaluation settings were tested and whether the reported gains are statistically significant.
  2. [Abstract] The hybrid annotation pipeline combining human expert labeling and automated heuristics is described without any evidence of label validity or reliability. The abstract does not mention inter-annotator agreement, validation of the heuristics against expert judgment, or comparison with established meteorological skill scores such as FSS or CSI. Without such evidence, it is unclear whether RQA-70K labels faithfully represent forecast quality or whether the model learns artifacts of the annotation heuristics.
  3. [Abstract] The statement that RQA-70K has 'varying difficulty levels' raises the question of how difficulty is defined and whether evaluation is stratified by difficulty. The abstract does not clarify whether the evaluation sets are held out from training, whether human-expert labels used for evaluation are independent from the labels used for training, and whether the reported 'outperforms' result is consistent across difficulty levels. These details are necessary to rule out circularity and to interpret the claim as a general capability rather than a single favorable split.
minor comments (3)
  1. [Abstract] The phrase 'multi-modal quality analysis' would benefit from specifying which modalities are involved (e.g., radar images plus text or numerical metadata), as this is central to understanding the task paradigm.
  2. [Abstract] The term 'rating and assessment scenarios' is introduced without examples; a brief clarification of the difference (e.g., numeric scoring versus free-text evaluation) would improve readability.
  3. [Abstract] The abstract would be strengthened by naming one or two baselines explicitly (which general MLLMs were compared) and by giving the dataset size split (e.g., train/validation/test) for RQA-70K.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity identified from the abstract-only evidence; benchmark evaluation on held-out data is a normal, non-circular protocol.

full rationale

The manuscript under review is abstract-only, so the derivation chain cannot be walked in full. The central claim, 'RadarQA outperforms existing general MLLMs across all evaluation settings,' is an empirical benchmark statement, not a mathematical or definitional derivation. The abstract describes training RadarQA on the RQA-70K dataset and evaluating against general MLLMs; training and evaluating on the same benchmark is standard practice when the evaluation uses held-out splits, and nothing in the abstract indicates that the evaluation labels are constructed from the model's own outputs or from the target quantities being predicted. The hybrid annotation pipeline combining human expert labeling with automated heuristics could in principle introduce construct-validity concerns, but the abstract does not specify that the heuristics are defined in terms of the model's outputs, nor that the human labels are fitted to the model, nor that a cited prior work supplies the load-bearing conclusion. No self-definitional step, fitted-input-called-prediction step, load-bearing self-citation, imported uniqueness theorem, ansatz smuggled by citation, or renamed known result can be quoted from the available text. Concerns about inter-annotator agreement, heuristic validation, and statistical significance are legitimate correctness and validity risks, but under the hard rules they are not circularity arguments without quoted evidence of a specific reduction. The honest finding is therefore no significant circularity, score 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

Only the abstract was available. The ledger records assumptions stated or implied in the abstract. No free parameters or invented entities could be identified without the full text.

assumptions (4)
  • domain assumption The hybrid annotation pipeline produces valid ground-truth labels for radar forecast quality.
    Introduced in the abstract as the method for constructing RQA-70K; if labels are noisy or biased, the model's performance cannot be trusted.
  • domain assumption The RQA-70K dataset covers a representative range of radar forecast quality issues.
    The abstract claims varying difficulty levels but does not demonstrate coverage of real-world operational conditions.
  • standard math Standard supervised learning assumptions hold, including independent train/test splits and no data leakage.
    Required for the claim that RadarQA outperforms general MLLMs; the abstract gives no details on evaluation protocol.
  • domain assumption MLLMs can effectively process radar imagery and associated physical attributes to produce meaningful assessments.
    The entire method depends on the ability of a multi-modal large language model to interpret radar data, which is plausible but unverified in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts." pith.science (2026). https://pith.science/paper/UDL22MJG

@misc{pith2026250812291,
  author       = {Pith},
  title        = {Pith review of: RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UDL22MJG}},
  note         = {Machine review of arXiv:2508.12291}
}
read the original abstract

Quality analysis of weather forecasts is an essential topic in meteorology. Although traditional score-based evaluation metrics can quantify certain forecast errors, they are still far from meteorological experts in terms of descriptive capability, interpretability, and understanding of dynamic evolution. With the rapid development of Multi-modal Large Language Models (MLLMs), these models become potential tools to overcome the above challenges. In this work, we introduce an MLLM-based weather forecast analysis method, RadarQA, integrating key physical attributes with detailed assessment reports. We introduce a novel and comprehensive task paradigm for multi-modal quality analysis, encompassing both single frame and sequence, under both rating and assessment scenarios. To support training and benchmarking, we design a hybrid annotation pipeline that combines human expert labeling with automated heuristics. With such an annotation method, we construct RQA-70K, a large-scale dataset with varying difficulty levels for radar forecast quality evaluation. We further design a multi-stage training strategy that iteratively improves model performance at each stage. Extensive experiments show that RadarQA outperforms existing general MLLMs across all evaluation settings, highlighting its potential for advancing quality analysis in weather prediction.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

88 extracted references · 38 canonical work pages

  1. [1]

    Agrawal, K

    H. Agrawal, K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson. Nocaps: Novel object captioning at scale. In ICCV, 2019

  2. [2]

    Antol, A

    S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh. VQA : Visual question answering. In ICCV, 2015

  3. [3]

    S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923, 2025

  4. [4]

    Banerjee and A

    S. Banerjee and A. Lavie. METEOR : An automatic metric for mt evaluation with improved correlation with human judgments. In ACL Workshops, 2005

  5. [5]

    K. Bi, L. Xie, H. Zhang, X. Chen, X. Gu, and Q. Tian. Pangu-weather: A 3d high-resolution model for fast and accurate global weather forecast. arXiv preprint arXiv:2211.02556, 2022

  6. [6]

    X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Dong, Q. Du, Z. Fu, et al. Deepseek LLM : Scaling open-source language models with longtermism. arXiv preprint arXiv:2401.02954, 2024

  7. [7]

    J. Chen, P. Zhou, Y. Hua, D. Chong, M. Cao, Y. Li, Z. Yuan, B. Zhu, and J. Liang. Vision-language models meet meteorology: Developing models for extreme weather events detection with heatmaps. arXiv preprint arXiv:2406.09838, 2024 a

  8. [8]

    K. Chen, T. Han, J. Gong, L. Bai, F. Ling, J.-J. Luo, X. Chen, L. Ma, T. Zhang, R. Su, et al. Fengwu: Pushing the skillful global medium-range weather forecast beyond 10 days lead. arXiv preprint arXiv:2304.02948, 2023

Show all 88 references
  1. [9]

    X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Doll \'a r, and C. L. Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015

  2. [10]

    Z. Chen, W. Wang, Y. Cao, Y. Liu, Z. Gao, E. Cui, J. Zhu, S. Ye, H. Tian, Z. Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271, 2024 b

  3. [11]

    Davis, B

    C. Davis, B. Brown, and R. Bullock. Object-based verification of precipitation forecasts. part i: Methodology and application to mesoscale rain areas. Monthly Weather Review, 134 0 (7): 0 1772--1784, 2006 a

  4. [12]

    Davis, B

    C. Davis, B. Brown, and R. Bullock. Object-based verification of precipitation forecasts. part ii: Application to convective rain systems. Monthly Weather Review, 134 0 (7): 0 1785--1795, 2006 b

  5. [13]

    Donaldson, R

    R. Donaldson, R. M. Dyer, and M. J. Kraus. An objective evaluator of techniques for predicting severe weather events. Preprints, Ninth Conf. on Severe Local Storms, Norman, OK, Amer. Meteor. Soc, 1975

  6. [14]

    J. P. Finley. Tornado predictions. American Meteorological Journal. A Monthly Review of Meteorology and Allied Branches of Study (1884-1896), 1884

  7. [15]

    Y. Gao, H. Wu, R. Shu, H. Dong, F. Xu, R. Chen, Y. Yan, Q. Wen, X. Hu, K. Wang, et al. Oneforecast: A universal framework for global and regional weather forecasting. arXiv preprint arXiv:2502.00338, 2025

  8. [16]

    Z. Gao, X. Shi, H. Wang, Y. Zhu, Y. B. Wang, M. Li, and D.-Y. Yeung. EarthFormer : Exploring space-time transformers for earth system forecasting. In NeurIPS, 2022 a

  9. [17]

    Z. Gao, C. Tan, L. Wu, and S. Z. Li. Simvp: Simpler yet better video prediction. In CVPR, 2022 b

  10. [18]

    Q. Ge, W. Sun, Y. Zhang, Y. Li, Z. Ji, F. Sun, S. Jui, X. Min, and G. Zhai. LMM-VQA : Advancing video quality assessment with large multimodal models. arXiv preprint arXiv:2408.14008, 2024

  11. [19]

    T. GLM, A. Zeng, B. Xu, B. Wang, C. Zhang, D. Yin, D. Zhang, D. Rojas, G. Feng, H. Zhao, et al. ChatGLM : A family of large language models from glm-130b to glm-4 all tools. arXiv preprint arXiv:2406.12793, 2024

  12. [20]

    F. Gofa, D. Boucouvala, P. Louka, and H. Flocas. Spatial verification approaches as a tool to evaluate the performance of high resolution precipitation forecasts. Atmospheric Research, 2018

  13. [21]

    J. Gong, L. Bai, P. Ye, W. Xu, N. Liu, J. Dai, X. Yang, and W. Ouyang. Cascast: Skillful high-resolution precipitation nowcasting via cascaded modelling. arXiv preprint arXiv:2402.04290, 2024

  14. [22]

    Goyal, T

    Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In CVPR, 2017

  15. [23]

    D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025

  16. [24]

    X. He, Z. Zhou, W. Zhang, X. Zhao, H. Chen, S. Chen, and L. Bai. DiffSR : Learning radar reflectivity synthesis via diffusion model from satellite observations. In ICASSP, 2025

  17. [25]

    K. A. Hilburn, I. Ebert-Uphoff, and S. D. Miller. Development and interpretation of a neural-network-based synthetic radar reflectivity estimator using goes-r satellite observations. Journal of Applied Meteorology and Climatology, 2020

  18. [26]

    E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al. Lora: Low-rank adaptation of large language models. In ICLR, 2022

  19. [27]

    Hurst, A

    A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al. GPT -4o system card. arXiv preprint arXiv:2410.21276, 2024

  20. [28]

    Jaech, A

    A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024

  21. [29]

    Z. Jia, Z. Zhang, J. Qian, H. Wu, W. Sun, C. Li, X. Liu, W. Lin, G. Zhai, and X. Min. VQA ^2 : Visual question answering for video quality assessment. arXiv preprint arXiv:2411.03795, 2024

  22. [30]

    I. T. Jolliffe and D. B. Stephenson. Forecast verification: a practitioner's guide in atmospheric science. John Wiley & Sons, 2012

  23. [31]

    B. Li, Y. Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y. Li, Z. Liu, et al. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024 a

  24. [32]

    F. Li, R. Zhang, H. Zhang, Y. Zhang, B. Li, W. Li, Z. Ma, and C. Li. Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models. arXiv preprint arXiv:2407.07895, 2024 b

  25. [33]

    W. Li, X. Zhang, S. Zhao, Y. Zhang, J. Li, L. Zhang, and J. Zhang. Q-Insight : Understanding image quality via visual reinforcement learning. arXiv preprint arXiv:2503.22679, 2025

  26. [34]

    B. Lin, Y. Ye, B. Zhu, J. Cui, M. Ning, P. Jin, and L. Yuan. Video-llava: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122, 2023

  27. [35]

    C.-Y. Lin. Rouge: A package for automatic evaluation of summaries. In ACL, 2004

  28. [36]

    F. Liu, Y. Wang, T. Wang, and V. Ordonez. Visual news: Benchmark and challenges in news image captioning. arXiv preprint arXiv:2010.03743, 2020

  29. [37]

    H. Liu, C. Li, Q. Wu, and Y. J. Lee. Visual instruction tuning. In NeurIPS, 2023

  30. [38]

    Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, et al. MMBench : Is your multi-modal model an all-around player? In ECCV, 2024

  31. [39]

    H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yang, et al. Deepseek-vl: towards real-world vision-language understanding. arXiv preprint arXiv:2403.05525, 2024

  32. [40]

    C. Ma, Z. Hua, A. Anderson-Frey, V. Iyer, X. Liu, and L. Qin. WeatherQA : Can multimodal language models reason about severe weather? arXiv preprint arXiv:2406.11217, 2024

  33. [41]

    Marino, M

    K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi. OK-VQA : A visual question answering benchmark requiring external knowledge. In CVPR, 2019

  34. [42]

    A. H. Murphy. The finley affair: A signal event in the history of forecast verification. Weather and forecasting, 1996

  35. [43]

    Palmer and R

    W. Palmer and R. Allen. Note on the accuracy of forecasts concerning the rain problem. US Weather Bureau, 1949

  36. [44]

    H. A. Panofsky and G. W. Brier. Some applications of statistics to meteorology. Mineral Industries Extension Services, College of Mineral Industries, Pennsylvania State University, 1958

  37. [45]

    Papineni, S

    K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu. Bleu: a method for automatic evaluation of machine translation. In ACL, 2002

  38. [46]

    Ravuri, K

    S. Ravuri, K. Lenc, M. Willson, D. Kangin, R. Lam, P. Mirowski, M. Fitzsimons, M. Athanassiadou, S. Kashem, S. Madge, et al. Skilful precipitation nowcasting using deep generative models of radar. Nature, 2021

  39. [47]

    Rempel, F

    M. Rempel, F. Senf, and H. Deneke. Object-based metrics for forecast verification of convective development with geostationary satellite data. Monthly Weather Review, 145 0 (8): 0 3161--3178, 2017

  40. [48]

    Robinson, J

    M. Robinson, J. Evans, and B. Crowe. En route weather depiction benefits of the nexrad vertically integrated liquid water product utilized by the corridor integrated weather system. In Conference on aviation, range and aerospace meteorology, american meteorological society, 2002

  41. [49]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017

  42. [50]

    Schwenk, A

    D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi. A-OKVQA : A benchmark for visual question answering using world knowledge. In ECCV, 2022

  43. [51]

    H. Shao, S. Qian, H. Xiao, G. Song, Z. Zong, L. Wang, Y. Liu, and H. Li. Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning. In NeurIPS, 2024 a

  44. [52]

    Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024 b

  45. [53]

    D. B. Stephenson, B. Casati, C. Ferro, and C. Wilson. The extreme dependency score: A non-vanishing measure for forecasts of rare events. Meteorological Applications: A journal of forecasting, practical applications, training techniques and modelling, 2008

  46. [54]

    Stock, K

    J. Stock, K. Hilburn, I. Ebert-Uphoff, and C. Anderson. Srvit: Vision transformers for estimating radar reflectivity from satellite observations at scale. arXiv preprint arXiv:2406.16955, 2024

  47. [55]

    K. Sun, J. Pan, Y. Ge, H. Li, H. Duan, X. Wu, R. Zhang, A. Zhou, Z. Qin, Y. Wang, et al. Journeydb: A benchmark for generative image understanding. In NeurIPS, 2023

  48. [56]

    Touvron, T

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi \`e re, N. Goyal, E. Hambro, F. Azhar, et al. LLaMA : Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023

  49. [57]

    Veillette, S

    M. Veillette, S. Samsi, and C. Mattioli. SEVIR : A storm event imagery dataset for deep learning applications in radar and satellite meteorology. In NeurIPS, 2020

  50. [58]

    F. Wang, M. Chen, X. He, Y. Zhang, F. Liu, Z. Guo, Z. Hu, J. Wang, J. Xu, Z. Li, et al. Omniearth-bench: Towards holistic evaluation of earth's six spheres and cross-spheres interactions with multimodal observational earth data. arXiv preprint arXiv:2505.23522, 2025

  51. [59]

    W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y. Wang, J. Ji, Z. Yang, L. Zhao, S. XiXuan, et al. Cogvlm: Visual expert for pretrained language models. In NeurIPS, 2024 a

  52. [60]

    Y. Wang, M. Long, J. Wang, Z. Gao, and P. S. Yu. PredRNN : Recurrent neural networks for predictive learning using spatiotemporal lstms. In NeurIPS, 2017

  53. [61]

    Y. Wang, Y. Zeng, J. Zheng, X. Xing, J. Xu, and X. Xu. VideoCoT : A video chain-of-thought dataset with active annotation tool. arXiv preprint arXiv:2407.05355, 2024 b

  54. [62]

    Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE TIP, 2004

  55. [63]

    C. J. Willmott and K. Matsuura. Advantages of the mean absolute error (mae) over the root mean square error (rmse) in assessing average model performance. Climate research, 2005

  56. [64]

    H. Wu, Z. Zhang, E. Zhang, C. Chen, L. Liao, A. Wang, C. Li, W. Sun, Q. Yan, G. Zhai, et al. Q-Bench : A benchmark for general-purpose foundation models on low-level vision. arXiv preprint arXiv:2309.14181, 2023

  57. [65]

    H. Wu, Z. Zhang, E. Zhang, C. Chen, L. Liao, A. Wang, K. Xu, C. Li, J. Hou, G. Zhai, et al. Q-Instruct : Improving low-level visual abilities for multi-modality foundation models. In CVPR, 2024 a

  58. [66]

    H. Wu, H. Zhu, Z. Zhang, E. Zhang, C. Chen, L. Liao, C. Li, A. Wang, W. Sun, Q. Yan, et al. Towards open-ended visual quality comparison. In ECCV, 2024 b

  59. [67]

    Z. Wu, X. Chen, Z. Pan, X. Liu, W. Liu, D. Dai, H. Gao, Y. Ma, C. Wu, B. Wang, et al. Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding. arXiv preprint arXiv:2412.10302, 2024 c

  60. [68]

    J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y. Fan, K. Dang, et al. Qwen2.5-omni technical report. arXiv preprint arXiv:2503.20215, 2025

  61. [69]

    A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, et al. Qwen2.5 technical report. arXiv preprint arXiv:2412.15115, 2024

  62. [70]

    J. Ye, H. Xu, H. Liu, A. Hu, M. Yan, Q. Qian, J. Zhang, F. Huang, and J. Zhou. mplug-owl3: Towards long image-sequence understanding in multi-modal large language models. arXiv preprint arXiv:2408.04840, 2024 a

  63. [71]

    Q. Ye, H. Xu, J. Ye, M. Yan, A. Hu, H. Liu, Q. Qian, J. Zhang, and F. Huang. mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration. In CVPR, 2024 b

  64. [72]

    Z. You, J. Gu, Z. Li, X. Cai, K. Zhu, C. Dong, and T. Xue. Descriptive image quality assessment in the wild. arXiv preprint arXiv:2405.18842, 2024 a

  65. [73]

    Z. You, Z. Li, J. Gu, Z. Yin, T. Xue, and C. Dong. Depicting beyond scores: Advancing image quality assessment through multi-modal language models. In ECCV, 2024 b

  66. [74]

    Z. You, X. Cai, J. Gu, T. Xue, and C. Dong. Teaching large language models to regress accurate image quality scores using score distribution. In CVPR, 2025

  67. [75]

    Young, A

    P. Young, A. Lai, M. Hodosh, and J. Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2014

  68. [76]

    D. Yu, X. Li, Y. Ye, B. Zhang, C. Luo, K. Dai, R. Wang, and X. Chen. Diffcast: A unified framework via residual diffusion for precipitation nowcasting. In CVPR, 2024

  69. [77]

    Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, T. Fan, G. Liu, L. Liu, X. Liu, et al. Dapo: An open-source llm reinforcement learning system at scale. arXiv preprint arXiv:2503.14476, 2025

  70. [78]

    Zhang, X

    P. Zhang, X. Dong, B. Wang, Y. Cao, C. Xu, L. Ouyang, Z. Zhao, H. Duan, S. Zhang, S. Ding, et al. Internlm-xcomposer: A vision-language large model for advanced text-image comprehension and composition. arXiv preprint arXiv:2309.15112, 2023 a

  71. [79]

    Zhang, V

    T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675, 2019

  72. [80]

    Zhang, H

    Y. Zhang, H. Yu, M. Zhang, Y. Yang, and Z. Meng. Uncertainties and error growth in forecasting the record-breaking rainfall in zhengzhou, henan on 19--20 july 2021. Science China Earth Sciences, 2022

  73. [81]

    Zhang, M

    Y. Zhang, M. Long, K. Chen, L. Xing, R. Jin, M. I. Jordan, and J. Wang. Skilful nowcasting of extreme precipitation with nowcastnet. Nature, 2023 b

  74. [82]

    Zhang, Z

    Z. Zhang, Z. Jia, H. Wu, C. Li, Z. Chen, Y. Zhou, W. Sun, X. Liu, X. Min, W. Lin, et al. Q-Bench-Video : Benchmarking the video quality understanding of lmms. arXiv preprint arXiv:2409.20063, 2024 a

  75. [83]

    Zhang, H

    Z. Zhang, H. Wu, Y. Zhou, C. Li, W. Sun, C. Chen, X. Min, X. Liu, W. Lin, and G. Zhai. LMM-PCQA : Assisting point cloud quality assessment with LMM . In ACM MM, 2024 b

  76. [84]

    Zhang, T

    Z. Zhang, T. Kou, S. Wang, C. Li, W. Sun, W. Wang, X. Li, Z. Wang, X. Cao, X. Min, et al. Q-Eval-100K : Evaluating visual quality and alignment level for text-to-vision content. arXiv preprint arXiv:2503.02357, 2025

  77. [85]

    X. Zhao, W. Xu, B. Liu, Y. Zhou, F. Ling, B. Fei, X. Yue, L. Bai, W. Zhang, and X.-M. Wu. Msearth: A benchmark for multimodal scientific comprehension of earth science. arXiv preprint arXiv:2505.20740, 2025

  78. [86]

    Zhong, Z

    Q. Zhong, Z. Sun, H. Chen, J. Li, and L. Shen. Multi model forecast biases of the diurnal variations of intense rainfall in the beijing-tianjin-hebei region. Science China Earth Sciences, 2022

  79. [87]

    M. Zhou, J. Wu, M. Chen, and L. Han. Comparative study on the performance of convlstm and convgru in classification problems—taking early warning of short-duration heavy rainfall as an example. Atmospheric and Oceanic Science Letters, 2024

  80. [88]

    Y. Zhou, Y. Wang, X. He, R. Xiao, Z. Li, Q. Feng, Z. Guo, Y. Yang, H. Wu, W. Huang, et al. Scientists' first exam: Probing cognitive abilities of mllm via perception, understanding, and reasoning. arXiv preprint arXiv:2506.10521, 2025

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.