REVIEW 3 major objections 3 minor 88 references
RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts
T0 review · 3 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper claims that RadarQA, a domain-tuned multi-modal language model trained on the RQA-70K dataset, outperforms general multimodal language models across all tested settings of radar forecast quality analysis.
desk verdict Abstract-only look at a plausible MLLM forecast QA benchmark; the contribution is real but the central validity claims are uncheckable from the abstract alone. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is a combination of three objects: a task paradigm that crosses single-frame versus sequence input with rating versus assessment output to define four quality-analysis settings; RQA-70K, a dataset of roughly 70,000 radar-forecast quality annotations produced by human experts plus automated heuristics and spanning easy to hard cases; and RadarQA itself, an MLLM trained with a multi-stage strategy in which each stage iteratively improves performance before the next begins. The load-bearing mechanism is the interaction between the hybrid labels and that training schedule: the model learns to ground its numeric ratings and written assessment reports in physical attributes of the radar forecasts rather than in generic language priors.
What would settle it
Have an independent panel of operational meteorologists, blinded to the original annotations, re-score a random sample of the radar forecasts in RQA-70K; if their labels agree with the dataset's hybrid labels no better than chance, the ground truth is not measuring forecast quality and the outperformance claim loses its foundation.
Extended reading notes
Core claim
The paper's central claim is that RadarQA, a multi-modal large language model fine-tuned on the RQA-70K dataset, outperforms existing general MLLMs across all evaluation settings for multi-modal forecast quality analysis, covering single-frame and sequence inputs and both rating and open-ended assessment outputs. The dataset is constructed through a hybrid annotation pipeline that combines human expert labeling with automated heuristics and is built to include varying difficulty levels. On the paper's own terms, this demonstrates that domain-specific MLLMs can give weather forecast verification descriptive, interpretable quality analysis that goes beyond score-based metrics, integrating key physical attributes of the radar forecasts into the model's ratings and reports.
Load-bearing premise
The result rests on a single assumption: the quality labels created by combining human expert judgment with automated rules are trustworthy, and if those labels are biased or inconsistent, both the dataset and the model built on it will not actually measure forecast quality.
Editorial extensions
If this is right
- If the claim is right, forecast verification can offer readable explanations of why a radar forecast was good or poor, not just a single numeric error score.
- RQA-70K provides a shared benchmark for future multi-modal quality analysis of weather radar forecasts.
- The multi-stage training result implies that domain-specific data and staged fine-tuning are worthwhile for applying MLLMs to meteorological tasks.
- RadarQA's output format combines ratings with written assessments, so it could be used to flag specific failure modes in radar forecasts in a way forecasters can scan quickly.
Reading between the lines
- Beyond the paper, the hybrid annotation pipeline could be transferred to other forecast products, such as precipitation nowcasts, satellite-based retrievals, or climate model output, although the paper does not test those settings.
- A test the paper leaves implicit is whether RadarQA's ratings and written assessments correlate with established skill scores on the same events, such as the critical success index or fractions skill score.
- Because the annotation style used for training and evaluation is the same, an independent human study comparing RadarQA's reports with general MLLM reports on fresh forecasts would clarify whether the advantage generalizes beyond the RQA-70K label distribution.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces RadarQA, a multi-modal large language model (MLLM) for weather radar forecast quality analysis. It defines a task paradigm covering single-frame and sequence inputs under rating and assessment scenarios, constructs a large dataset RQA-70K through a hybrid human-expert and automated-heuristic annotation pipeline, and proposes a multi-stage training strategy. The central claim, stated in the abstract, is that RadarQA outperforms existing general MLLMs across all evaluation settings, implying that domain-specialized MLLMs can provide descriptive forecast-quality assessments beyond score-based metrics.
Significance. If the central claim holds, the work could advance weather forecast verification by adding interpretable, descriptive, and dynamic-evolution-aware assessments to traditional score-based metrics. The proposed dataset and task paradigm would be useful resources for the community, and the multi-stage training approach could inform future domain-specific MLLM development. The authors are to be credited for explicitly proposing a new task paradigm and for assembling a large-scale dataset, though the abstract alone does not yet provide the quantitative evidence needed to assess the strength of these contributions.
major comments (3)
- [Abstract] The central claim that 'RadarQA outperforms existing general MLLMs across all evaluation settings' is not accompanied by any quantitative result, such as accuracy, correlation, or agreement scores, nor by error bars, significance tests, or a description of the evaluation protocol. Because this claim is the paper's main load-bearing assertion, the abstract must at least report the headline numbers and state, for example, how many evaluation settings were tested and whether the reported gains are statistically significant.
- [Abstract] The hybrid annotation pipeline combining human expert labeling and automated heuristics is described without any evidence of label validity or reliability. The abstract does not mention inter-annotator agreement, validation of the heuristics against expert judgment, or comparison with established meteorological skill scores such as FSS or CSI. Without such evidence, it is unclear whether RQA-70K labels faithfully represent forecast quality or whether the model learns artifacts of the annotation heuristics.
- [Abstract] The statement that RQA-70K has 'varying difficulty levels' raises the question of how difficulty is defined and whether evaluation is stratified by difficulty. The abstract does not clarify whether the evaluation sets are held out from training, whether human-expert labels used for evaluation are independent from the labels used for training, and whether the reported 'outperforms' result is consistent across difficulty levels. These details are necessary to rule out circularity and to interpret the claim as a general capability rather than a single favorable split.
minor comments (3)
- [Abstract] The phrase 'multi-modal quality analysis' would benefit from specifying which modalities are involved (e.g., radar images plus text or numerical metadata), as this is central to understanding the task paradigm.
- [Abstract] The term 'rating and assessment scenarios' is introduced without examples; a brief clarification of the difference (e.g., numeric scoring versus free-text evaluation) would improve readability.
- [Abstract] The abstract would be strengthened by naming one or two baselines explicitly (which general MLLMs were compared) and by giving the dataset size split (e.g., train/validation/test) for RQA-70K.
Circularity Check
No circularity identified from the abstract-only evidence; benchmark evaluation on held-out data is a normal, non-circular protocol.
full rationale
The manuscript under review is abstract-only, so the derivation chain cannot be walked in full. The central claim, 'RadarQA outperforms existing general MLLMs across all evaluation settings,' is an empirical benchmark statement, not a mathematical or definitional derivation. The abstract describes training RadarQA on the RQA-70K dataset and evaluating against general MLLMs; training and evaluating on the same benchmark is standard practice when the evaluation uses held-out splits, and nothing in the abstract indicates that the evaluation labels are constructed from the model's own outputs or from the target quantities being predicted. The hybrid annotation pipeline combining human expert labeling with automated heuristics could in principle introduce construct-validity concerns, but the abstract does not specify that the heuristics are defined in terms of the model's outputs, nor that the human labels are fitted to the model, nor that a cited prior work supplies the load-bearing conclusion. No self-definitional step, fitted-input-called-prediction step, load-bearing self-citation, imported uniqueness theorem, ansatz smuggled by citation, or renamed known result can be quoted from the available text. Concerns about inter-annotator agreement, heuristic validation, and statistical significance are legitimate correctness and validity risks, but under the hard rules they are not circularity arguments without quoted evidence of a specific reduction. The honest finding is therefore no significant circularity, score 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The hybrid annotation pipeline produces valid ground-truth labels for radar forecast quality.
- domain assumption The RQA-70K dataset covers a representative range of radar forecast quality issues.
- standard math Standard supervised learning assumptions hold, including independent train/test splits and no data leakage.
- domain assumption MLLMs can effectively process radar imagery and associated physical attributes to produce meaningful assessments.
Cite this review
Pith. "Pith review of RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts." pith.science (2026). https://pith.science/paper/UDL22MJG
@misc{pith2026250812291,
author = {Pith},
title = {Pith review of: RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts},
year = {2026},
howpublished = {\url{https://pith.science/paper/UDL22MJG}},
note = {Machine review of arXiv:2508.12291}
}
read the original abstract
Quality analysis of weather forecasts is an essential topic in meteorology. Although traditional score-based evaluation metrics can quantify certain forecast errors, they are still far from meteorological experts in terms of descriptive capability, interpretability, and understanding of dynamic evolution. With the rapid development of Multi-modal Large Language Models (MLLMs), these models become potential tools to overcome the above challenges. In this work, we introduce an MLLM-based weather forecast analysis method, RadarQA, integrating key physical attributes with detailed assessment reports. We introduce a novel and comprehensive task paradigm for multi-modal quality analysis, encompassing both single frame and sequence, under both rating and assessment scenarios. To support training and benchmarking, we design a hybrid annotation pipeline that combines human expert labeling with automated heuristics. With such an annotation method, we construct RQA-70K, a large-scale dataset with varying difficulty levels for radar forecast quality evaluation. We further design a multi-stage training strategy that iteratively improves model performance at each stage. Extensive experiments show that RadarQA outperforms existing general MLLMs across all evaluation settings, highlighting its potential for advancing quality analysis in weather prediction.
Reference graph
Works this paper leans on
-
[1]
Agrawal, K
H. Agrawal, K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson. Nocaps: Novel object captioning at scale. In ICCV, 2019
2019
-
[2]
Antol, A
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh. VQA : Visual question answering. In ICCV, 2015
2015
-
[3]
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923, 2025
arXiv 2025
-
[4]
Banerjee and A
S. Banerjee and A. Lavie. METEOR : An automatic metric for mt evaluation with improved correlation with human judgments. In ACL Workshops, 2005
2005
-
[5]
K. Bi, L. Xie, H. Zhang, X. Chen, X. Gu, and Q. Tian. Pangu-weather: A 3d high-resolution model for fast and accurate global weather forecast. arXiv preprint arXiv:2211.02556, 2022
arXiv 2022
-
[6]
X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Dong, Q. Du, Z. Fu, et al. Deepseek LLM : Scaling open-source language models with longtermism. arXiv preprint arXiv:2401.02954, 2024
arXiv 2024
-
[7]
J. Chen, P. Zhou, Y. Hua, D. Chong, M. Cao, Y. Li, Z. Yuan, B. Zhu, and J. Liang. Vision-language models meet meteorology: Developing models for extreme weather events detection with heatmaps. arXiv preprint arXiv:2406.09838, 2024 a
arXiv 2024
-
[8]
K. Chen, T. Han, J. Gong, L. Bai, F. Ling, J.-J. Luo, X. Chen, L. Ma, T. Zhang, R. Su, et al. Fengwu: Pushing the skillful global medium-range weather forecast beyond 10 days lead. arXiv preprint arXiv:2304.02948, 2023
arXiv 2023
Show all 88 references
-
[9]
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Doll \'a r, and C. L. Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015
2015 arXiv
-
[10]
Z. Chen, W. Wang, Y. Cao, Y. Liu, Z. Gao, E. Cui, J. Zhu, S. Ye, H. Tian, Z. Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271, 2024 b
2024 arXiv
-
[11]
Davis, B
C. Davis, B. Brown, and R. Bullock. Object-based verification of precipitation forecasts. part i: Methodology and application to mesoscale rain areas. Monthly Weather Review, 134 0 (7): 0 1772--1784, 2006 a
2006
-
[12]
Davis, B
C. Davis, B. Brown, and R. Bullock. Object-based verification of precipitation forecasts. part ii: Application to convective rain systems. Monthly Weather Review, 134 0 (7): 0 1785--1795, 2006 b
2006
-
[13]
Donaldson, R
R. Donaldson, R. M. Dyer, and M. J. Kraus. An objective evaluator of techniques for predicting severe weather events. Preprints, Ninth Conf. on Severe Local Storms, Norman, OK, Amer. Meteor. Soc, 1975
1975
-
[14]
J. P. Finley. Tornado predictions. American Meteorological Journal. A Monthly Review of Meteorology and Allied Branches of Study (1884-1896), 1884
-
[15]
Y. Gao, H. Wu, R. Shu, H. Dong, F. Xu, R. Chen, Y. Yan, Q. Wen, X. Hu, K. Wang, et al. Oneforecast: A universal framework for global and regional weather forecasting. arXiv preprint arXiv:2502.00338, 2025
2025
-
[16]
Z. Gao, X. Shi, H. Wang, Y. Zhu, Y. B. Wang, M. Li, and D.-Y. Yeung. EarthFormer : Exploring space-time transformers for earth system forecasting. In NeurIPS, 2022 a
2022
-
[17]
Z. Gao, C. Tan, L. Wu, and S. Z. Li. Simvp: Simpler yet better video prediction. In CVPR, 2022 b
2022
-
[18]
Q. Ge, W. Sun, Y. Zhang, Y. Li, Z. Ji, F. Sun, S. Jui, X. Min, and G. Zhai. LMM-VQA : Advancing video quality assessment with large multimodal models. arXiv preprint arXiv:2408.14008, 2024
2024 arXiv
-
[19]
T. GLM, A. Zeng, B. Xu, B. Wang, C. Zhang, D. Yin, D. Zhang, D. Rojas, G. Feng, H. Zhao, et al. ChatGLM : A family of large language models from glm-130b to glm-4 all tools. arXiv preprint arXiv:2406.12793, 2024
2024 arXiv
-
[20]
F. Gofa, D. Boucouvala, P. Louka, and H. Flocas. Spatial verification approaches as a tool to evaluate the performance of high resolution precipitation forecasts. Atmospheric Research, 2018
2018
-
[21]
J. Gong, L. Bai, P. Ye, W. Xu, N. Liu, J. Dai, X. Yang, and W. Ouyang. Cascast: Skillful high-resolution precipitation nowcasting via cascaded modelling. arXiv preprint arXiv:2402.04290, 2024
2024 arXiv
-
[22]
Goyal, T
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In CVPR, 2017
2017
-
[23]
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025
2025 arXiv
-
[24]
X. He, Z. Zhou, W. Zhang, X. Zhao, H. Chen, S. Chen, and L. Bai. DiffSR : Learning radar reflectivity synthesis via diffusion model from satellite observations. In ICASSP, 2025
2025
-
[25]
K. A. Hilburn, I. Ebert-Uphoff, and S. D. Miller. Development and interpretation of a neural-network-based synthetic radar reflectivity estimator using goes-r satellite observations. Journal of Applied Meteorology and Climatology, 2020
2020
-
[26]
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al. Lora: Low-rank adaptation of large language models. In ICLR, 2022
2022
-
[27]
Hurst, A
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al. GPT -4o system card. arXiv preprint arXiv:2410.21276, 2024
2024 arXiv
-
[28]
Jaech, A
A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024
2024 arXiv
-
[29]
Z. Jia, Z. Zhang, J. Qian, H. Wu, W. Sun, C. Li, X. Liu, W. Lin, G. Zhai, and X. Min. VQA ^2 : Visual question answering for video quality assessment. arXiv preprint arXiv:2411.03795, 2024
2024 arXiv
-
[30]
I. T. Jolliffe and D. B. Stephenson. Forecast verification: a practitioner's guide in atmospheric science. John Wiley & Sons, 2012
2012
-
[31]
B. Li, Y. Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y. Li, Z. Liu, et al. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024 a
2024 arXiv
-
[32]
F. Li, R. Zhang, H. Zhang, Y. Zhang, B. Li, W. Li, Z. Ma, and C. Li. Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models. arXiv preprint arXiv:2407.07895, 2024 b
2024 arXiv
-
[33]
W. Li, X. Zhang, S. Zhao, Y. Zhang, J. Li, L. Zhang, and J. Zhang. Q-Insight : Understanding image quality via visual reinforcement learning. arXiv preprint arXiv:2503.22679, 2025
2025 arXiv
-
[34]
B. Lin, Y. Ye, B. Zhu, J. Cui, M. Ning, P. Jin, and L. Yuan. Video-llava: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122, 2023
2023 arXiv
-
[35]
C.-Y. Lin. Rouge: A package for automatic evaluation of summaries. In ACL, 2004
2004
-
[36]
F. Liu, Y. Wang, T. Wang, and V. Ordonez. Visual news: Benchmark and challenges in news image captioning. arXiv preprint arXiv:2010.03743, 2020
2010 arXiv
-
[37]
H. Liu, C. Li, Q. Wu, and Y. J. Lee. Visual instruction tuning. In NeurIPS, 2023
2023
-
[38]
Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, et al. MMBench : Is your multi-modal model an all-around player? In ECCV, 2024
2024
-
[39]
H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yang, et al. Deepseek-vl: towards real-world vision-language understanding. arXiv preprint arXiv:2403.05525, 2024
2024 arXiv
-
[40]
C. Ma, Z. Hua, A. Anderson-Frey, V. Iyer, X. Liu, and L. Qin. WeatherQA : Can multimodal language models reason about severe weather? arXiv preprint arXiv:2406.11217, 2024
2024 arXiv
-
[41]
Marino, M
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi. OK-VQA : A visual question answering benchmark requiring external knowledge. In CVPR, 2019
2019
-
[42]
A. H. Murphy. The finley affair: A signal event in the history of forecast verification. Weather and forecasting, 1996
1996
-
[43]
Palmer and R
W. Palmer and R. Allen. Note on the accuracy of forecasts concerning the rain problem. US Weather Bureau, 1949
1949
-
[44]
H. A. Panofsky and G. W. Brier. Some applications of statistics to meteorology. Mineral Industries Extension Services, College of Mineral Industries, Pennsylvania State University, 1958
1958
-
[45]
Papineni, S
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu. Bleu: a method for automatic evaluation of machine translation. In ACL, 2002
2002
-
[46]
Ravuri, K
S. Ravuri, K. Lenc, M. Willson, D. Kangin, R. Lam, P. Mirowski, M. Fitzsimons, M. Athanassiadou, S. Kashem, S. Madge, et al. Skilful precipitation nowcasting using deep generative models of radar. Nature, 2021
2021
-
[47]
Rempel, F
M. Rempel, F. Senf, and H. Deneke. Object-based metrics for forecast verification of convective development with geostationary satellite data. Monthly Weather Review, 145 0 (8): 0 3161--3178, 2017
2017
-
[48]
Robinson, J
M. Robinson, J. Evans, and B. Crowe. En route weather depiction benefits of the nexrad vertically integrated liquid water product utilized by the corridor integrated weather system. In Conference on aviation, range and aerospace meteorology, american meteorological society, 2002
2002
-
[49]
Schulman, F
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017
2017 arXiv
-
[50]
Schwenk, A
D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi. A-OKVQA : A benchmark for visual question answering using world knowledge. In ECCV, 2022
2022
-
[51]
H. Shao, S. Qian, H. Xiao, G. Song, Z. Zong, L. Wang, Y. Liu, and H. Li. Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning. In NeurIPS, 2024 a
2024
-
[52]
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024 b
2024 arXiv
-
[53]
D. B. Stephenson, B. Casati, C. Ferro, and C. Wilson. The extreme dependency score: A non-vanishing measure for forecasts of rare events. Meteorological Applications: A journal of forecasting, practical applications, training techniques and modelling, 2008
2008
-
[54]
Stock, K
J. Stock, K. Hilburn, I. Ebert-Uphoff, and C. Anderson. Srvit: Vision transformers for estimating radar reflectivity from satellite observations at scale. arXiv preprint arXiv:2406.16955, 2024
2024 arXiv
-
[55]
K. Sun, J. Pan, Y. Ge, H. Li, H. Duan, X. Wu, R. Zhang, A. Zhou, Z. Qin, Y. Wang, et al. Journeydb: A benchmark for generative image understanding. In NeurIPS, 2023
2023
-
[56]
Touvron, T
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi \`e re, N. Goyal, E. Hambro, F. Azhar, et al. LLaMA : Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[57]
Veillette, S
M. Veillette, S. Samsi, and C. Mattioli. SEVIR : A storm event imagery dataset for deep learning applications in radar and satellite meteorology. In NeurIPS, 2020
2020
-
[58]
F. Wang, M. Chen, X. He, Y. Zhang, F. Liu, Z. Guo, Z. Hu, J. Wang, J. Xu, Z. Li, et al. Omniearth-bench: Towards holistic evaluation of earth's six spheres and cross-spheres interactions with multimodal observational earth data. arXiv preprint arXiv:2505.23522, 2025
2025
-
[59]
W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y. Wang, J. Ji, Z. Yang, L. Zhao, S. XiXuan, et al. Cogvlm: Visual expert for pretrained language models. In NeurIPS, 2024 a
2024
-
[60]
Y. Wang, M. Long, J. Wang, Z. Gao, and P. S. Yu. PredRNN : Recurrent neural networks for predictive learning using spatiotemporal lstms. In NeurIPS, 2017
2017
-
[61]
Y. Wang, Y. Zeng, J. Zheng, X. Xing, J. Xu, and X. Xu. VideoCoT : A video chain-of-thought dataset with active annotation tool. arXiv preprint arXiv:2407.05355, 2024 b
2024 arXiv
-
[62]
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE TIP, 2004
2004
-
[63]
C. J. Willmott and K. Matsuura. Advantages of the mean absolute error (mae) over the root mean square error (rmse) in assessing average model performance. Climate research, 2005
2005
-
[64]
H. Wu, Z. Zhang, E. Zhang, C. Chen, L. Liao, A. Wang, C. Li, W. Sun, Q. Yan, G. Zhai, et al. Q-Bench : A benchmark for general-purpose foundation models on low-level vision. arXiv preprint arXiv:2309.14181, 2023
2023 arXiv
-
[65]
H. Wu, Z. Zhang, E. Zhang, C. Chen, L. Liao, A. Wang, K. Xu, C. Li, J. Hou, G. Zhai, et al. Q-Instruct : Improving low-level visual abilities for multi-modality foundation models. In CVPR, 2024 a
2024
-
[66]
H. Wu, H. Zhu, Z. Zhang, E. Zhang, C. Chen, L. Liao, C. Li, A. Wang, W. Sun, Q. Yan, et al. Towards open-ended visual quality comparison. In ECCV, 2024 b
2024
-
[67]
Z. Wu, X. Chen, Z. Pan, X. Liu, W. Liu, D. Dai, H. Gao, Y. Ma, C. Wu, B. Wang, et al. Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding. arXiv preprint arXiv:2412.10302, 2024 c
2024 arXiv
-
[68]
J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y. Fan, K. Dang, et al. Qwen2.5-omni technical report. arXiv preprint arXiv:2503.20215, 2025
2025 arXiv
-
[69]
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, et al. Qwen2.5 technical report. arXiv preprint arXiv:2412.15115, 2024
2024 arXiv
-
[70]
J. Ye, H. Xu, H. Liu, A. Hu, M. Yan, Q. Qian, J. Zhang, F. Huang, and J. Zhou. mplug-owl3: Towards long image-sequence understanding in multi-modal large language models. arXiv preprint arXiv:2408.04840, 2024 a
2024 arXiv
-
[71]
Q. Ye, H. Xu, J. Ye, M. Yan, A. Hu, H. Liu, Q. Qian, J. Zhang, and F. Huang. mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration. In CVPR, 2024 b
2024
-
[72]
Z. You, J. Gu, Z. Li, X. Cai, K. Zhu, C. Dong, and T. Xue. Descriptive image quality assessment in the wild. arXiv preprint arXiv:2405.18842, 2024 a
2024
-
[73]
Z. You, Z. Li, J. Gu, Z. Yin, T. Xue, and C. Dong. Depicting beyond scores: Advancing image quality assessment through multi-modal language models. In ECCV, 2024 b
2024
-
[74]
Z. You, X. Cai, J. Gu, T. Xue, and C. Dong. Teaching large language models to regress accurate image quality scores using score distribution. In CVPR, 2025
2025
-
[75]
Young, A
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2014
2014
-
[76]
D. Yu, X. Li, Y. Ye, B. Zhang, C. Luo, K. Dai, R. Wang, and X. Chen. Diffcast: A unified framework via residual diffusion for precipitation nowcasting. In CVPR, 2024
2024
-
[77]
Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, T. Fan, G. Liu, L. Liu, X. Liu, et al. Dapo: An open-source llm reinforcement learning system at scale. arXiv preprint arXiv:2503.14476, 2025
2025 arXiv
-
[78]
Zhang, X
P. Zhang, X. Dong, B. Wang, Y. Cao, C. Xu, L. Ouyang, Z. Zhao, H. Duan, S. Zhang, S. Ding, et al. Internlm-xcomposer: A vision-language large model for advanced text-image comprehension and composition. arXiv preprint arXiv:2309.15112, 2023 a
2023 arXiv
-
[79]
Zhang, V
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675, 2019
1904 arXiv
-
[80]
Zhang, H
Y. Zhang, H. Yu, M. Zhang, Y. Yang, and Z. Meng. Uncertainties and error growth in forecasting the record-breaking rainfall in zhengzhou, henan on 19--20 july 2021. Science China Earth Sciences, 2022
2021
-
[81]
Zhang, M
Y. Zhang, M. Long, K. Chen, L. Xing, R. Jin, M. I. Jordan, and J. Wang. Skilful nowcasting of extreme precipitation with nowcastnet. Nature, 2023 b
2023
-
[82]
Zhang, Z
Z. Zhang, Z. Jia, H. Wu, C. Li, Z. Chen, Y. Zhou, W. Sun, X. Liu, X. Min, W. Lin, et al. Q-Bench-Video : Benchmarking the video quality understanding of lmms. arXiv preprint arXiv:2409.20063, 2024 a
2024 arXiv
-
[83]
Zhang, H
Z. Zhang, H. Wu, Y. Zhou, C. Li, W. Sun, C. Chen, X. Min, X. Liu, W. Lin, and G. Zhai. LMM-PCQA : Assisting point cloud quality assessment with LMM . In ACM MM, 2024 b
2024
-
[84]
Zhang, T
Z. Zhang, T. Kou, S. Wang, C. Li, W. Sun, W. Wang, X. Li, Z. Wang, X. Cao, X. Min, et al. Q-Eval-100K : Evaluating visual quality and alignment level for text-to-vision content. arXiv preprint arXiv:2503.02357, 2025
2025 arXiv
-
[85]
X. Zhao, W. Xu, B. Liu, Y. Zhou, F. Ling, B. Fei, X. Yue, L. Bai, W. Zhang, and X.-M. Wu. Msearth: A benchmark for multimodal scientific comprehension of earth science. arXiv preprint arXiv:2505.20740, 2025
2025 arXiv
-
[86]
Zhong, Z
Q. Zhong, Z. Sun, H. Chen, J. Li, and L. Shen. Multi model forecast biases of the diurnal variations of intense rainfall in the beijing-tianjin-hebei region. Science China Earth Sciences, 2022
2022
-
[87]
M. Zhou, J. Wu, M. Chen, and L. Han. Comparative study on the performance of convlstm and convgru in classification problems—taking early warning of short-duration heavy rainfall as an example. Atmospheric and Oceanic Science Letters, 2024
2024
-
[88]
Y. Zhou, Y. Wang, X. He, R. Xiao, Z. Li, Q. Feng, Z. Guo, Y. Yang, H. Wu, W. Huang, et al. Scientists' first exam: Probing cognitive abilities of mllm via perception, understanding, and reasoning. arXiv preprint arXiv:2506.10521, 2025
2025
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.