REVIEW 3 major objections 3 minor 88 references
Flow reorganization and transport enhancement in two-dimensional horizontal convection near a density extremum
T0 review · 3 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read In two-dimensional horizontal convection, combining a density extremum with a nonlinear equation of state reorganizes the flow into a single-roll circulation and, when mixing plumes reach full depth, enhances heat transport from the…
desk verdict The attached full text is a different paper (RadarQA), so this verdict rests entirely on the abstract: the density-extremum horizontal convection result is plausible and likely novel, but the central z-hat scaling argument is too under-supported to accept as established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is $\Phi_{i2}$, an additional potential-energy transfer term appearing in the total energy budget when the equation of state is nonlinear, alongside the usual Oberbeck-Boussinesq terms. Its magnitude is argued to scale with the characteristic mixing-plume height $\hat{z}$; full-depth plumes ($\hat{z} \sim H$) make the standard OB horizontal-convection kinetic-energy dissipation closure incomplete, which is how the paper moves the heat-transport exponent from $1/5$ to $1/4$–$1/3$.
What would settle it
Run the EXT-NELT configuration in a taller cavity, or at a higher Rayleigh number, where $\hat{z}$ no longer reaches $H$; if $Nu$ still follows a $Ra^{1/4}$–$Ra^{1/3}$ scaling, the plume-depth link fails. Conversely, a laboratory experiment with water near $4^\circ\mathrm{C}$ that directly measures plume penetration depth and Nusselt number would test whether $\hat{z} \sim H$ is necessary for the enhanced scaling.
Extended reading notes
Core claim
The central claim is that the nonlinearity of the equation of state, not the density-extremum boundary condition alone, is what changes horizontal convection. In the EXT-NELT configuration, the large-scale circulation shifts from two cells to a single roll, with central mixing plumes carrying the transport; when these plumes span the whole height $H$ ($\hat{z} \sim H$), the Nusselt number grows as $Ra^{1/4}$ to $Ra^{1/3}$, distinctly steeper than the Rossby $Ra^{1/5}$ seen in the other three configurations. The paper's mechanism is a new term $\Phi_{i2}$ in the global kinetic-energy balance, a potential-energy transfer produced by the nonlinear equation of state, whose magnitude is controlled by $\hat{z}$. With $\hat{z} \sim H$, this term changes the dissipation closure, and a scaling model built from it reproduces the main trends of the DNS data.
Load-bearing premise
The explanation assumes that the extra energy term $\Phi_{i2}$ is controlled by the plume height $\hat{z}$ measured from the same simulations whose heat-transport scaling it is then used to reproduce, so the mechanism is calibrated to the data rather than independently predicting them.
Editorial extensions
If this is right
- For density-extremum horizontal convection with full-depth plumes, the standard Oberbeck-Boussinesq closure under-predicts heat transport; global energy budgets for such flows should include $\Phi_{i2}$.
- The flow structure itself changes: the bicellular circulation becomes a single-roll, central-plume regime, with transitional anomalies in the Reynolds-number scaling.
- The heat-transport exponent is not fixed: it can lie between $1/4$ and $1/3$ depending on plume depth, so a single power law may not describe the full parameter range.
- Models of lakes, ice-ocean settings, or industrial systems where water near $4^\circ\mathrm{C}$ is the working fluid should not assume the Rossby $1/5$ scaling once plumes become depth-penetrating.
- The energy-budget model offers a diagnostic route: computing $\Phi_{i2}$ and $\hat{z}$ from DNS or experiments could predict when enhanced transport begins.
Reading between the lines
- If the plume-height control holds, the transition from $Ra^{1/5}$ to $Ra^{1/3}$ should be continuous in $\hat{z}/H$; a testable prediction is that the effective Nusselt exponent is a monotone function of plume penetration depth.
- The single-roll reorganization suggests that three-dimensional simulations may show a qualitatively different large-scale flow, since two-dimensional central plumes are often sensitive to confinement; the scaling claim should be checked in three dimensions.
- One can use the $\Phi_{i2}$ budget term to design controlled experiments: changing the temperature of the cold boundary relative to $4^\circ\mathrm{C}$ should tune $\hat{z}$ and produce a predictable shift in the Nusselt scaling.
- Analogous potential-energy transfer terms may appear for other non-Oberbeck-Boussinesq nonlinearities, such as salinity or compressibility effects, so the same energy-budget analysis could identify enhanced transport in those settings.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript as identified by its title and abstract reports a two-dimensional direct numerical simulation study of horizontal convection with a nonlinear equation of state exhibiting a density extremum near 4°C. The abstract claims that in the EXT-NELT configuration the flow reorganizes from a bicellular structure into a single-roll circulation driven by full-depth mixing plumes, that heat transport is enhanced from the Rossby scaling Nu ~ Ra^{1/5} to Nu ~ Ra^{1/4}–Ra^{1/3}, and that this enhancement arises from an additional potential-energy transfer term Phi_{i2} introduced by the nonlinear equation of state, with the term's magnitude controlled by the characteristic plume height z-hat. The submitted full text, however, is not this paper: it is an unrelated manuscript on multi-modal quality analysis of weather radar forecasts (RadarQA, arXiv:2508.12291), containing no governing equations, simulation setup, energy-budget analysis, or numerical results for horizontal convection. The abstract-level claims are therefore the only reviewable content of the physics submission.
Significance. If the abstract's claims were fully substantiated, the result would be of genuine interest to the horizontal-convection and geophysical fluid dynamics communities: it would indicate that the standard Oberbeck-Boussinesq energy closure underpredicts heat transport in density-extremum configurations once plumes penetrate the full cavity depth, and it would identify a specific energy-budget term, Phi_{i2}, responsible for the change. The reported reorganization from a bicellular to a single-roll circulation is a plausible and potentially impactful observation for flows near the density maximum of water. However, the paper as submitted offers no machine-checked proofs, reproducible code, derivations, figures, or data tables for any of these claims; the actual manuscript body is a different work. The only quantitative statement available, that the model 'captures the main trends of the numerical data,' is too weak to constitute validation, and the claimed exponent range 1/4–1/3 is presented without fitting statistics or a stated Ra window.
major comments (3)
- [Full text] The body of the submission is a different manuscript: the entire 'Full Text' section is the RadarQA paper on weather radar forecast quality analysis, with different authors, different subject matter, and no equations or results relating to horizontal convection, the density extremum, Phi_{i2}, z-hat, or Nu-Ra scaling. None of the central claims of the title and abstract can be checked because the supporting derivation, DNS setup, boundary conditions, equation-of-state form, and data-analysis procedure are absent. This is a load-bearing defect: the manuscript is internally inconsistent and cannot be reviewed as a physics paper.
- [Abstract, second paragraph] The explanatory mechanism appears circular: the new potential-energy transfer term Phi_{i2} is asserted to be controlled by the characteristic plume height z-hat, and z-hat is diagnosed from the same EXT-NELT DNS runs whose heat-transport scaling the model then reproduces. The abstract states this only as a 'scaling argument' with 'suggests,' and no independent derivation of the Phi_{i2}(z-hat) relation is provided. As presented, the model is calibrated to the data it explains rather than offering an out-of-sample prediction; a test on a different Rayleigh-number range or an alternative equation-of-state parametrization would be needed to establish the claimed 1/4-to-1/3 exponent.
- [Abstract, first paragraph] The claim of 'enhanced heat transport scaling ranging from Nu ~ Ra^{1/4} to Nu ~ Ra^{1/3}' spans two distinct power-law exponents, and the abstract does not say whether this is a single fitted exponent that drifts with Rayleigh number, a crossover between two asymptotic regimes, or a fit with substantial uncertainty. Without the figures and fitting details that would appear in the full text, the reader cannot assess whether the data support a new dissipation closure or merely a transitional, non-asymptotic effect.
minor comments (3)
- [Abstract, second paragraph] The term Phi_{i2} is introduced without a definition or an equation number, so its physical content is not accessible from the abstract alone.
- [Abstract, first paragraph] The abbreviations EXT, MON, LENT, and NELT are used without expansion, and the nonlinear equation-of-state form is never stated.
- [Full text] The author list and reference list of the full text belong to the RadarQA manuscript, which is unrelated to the title and abstract of the submission; this mismatch should be resolved by the authors before any further review.
Circularity Check
The Phi_i2-vs-z-hat scaling argument is calibrated to the same DNS data whose Nu exponent it explains.
-
fitted input called prediction
[Abstract, second paragraph (energy-budget discussion of Phi_i2 and z-hat)]
"The scaling argument suggests that the magnitude of this contribution is controlled by the characteristic plume height ($\hat{z}$). Specifically, when plumes penetrate the entire cavity depth ($\hat{z} \sim H$), as observed in the EXT-NELT case, the global kinetic energy dissipation is no longer described by the standard OB HC energy closure alone. The resulting model captures the main trends of the numerical data and provides a possible energy budget interpretation of the enhanced transport observed in this two-dimensional configuration."
The explanatory variable z-hat is diagnosed from the same DNS runs whose heat-transport scaling the model then reproduces. The magnitude of the new energy term Phi_i2 is said to be controlled by z-hat, and the full-depth condition z-hat ~ H is itself 'observed in the EXT-NELT case' from the same simulations. Thus the claimed agreement between the resulting model and the numerical data is a consistency check rather than an independent prediction; the scaling argument is calibrated to the data it explains. The phrase 'captures the main trends of the numerical data' reinforces that the model is fitted to the observed Nu(Ra) behavior rather than predicting it from first principles.
full rationale
The central empirical finding—that EXT-NELT reorganizes into a single roll and exhibits enhanced Nu scaling—is a DNS observation and is not circular by itself. No self-citation chains or imported uniqueness theorems appear in the abstract. However, the proposed energy-budget mechanism is not independent: Phi_i2 is stated to be controlled by z-hat, and z-hat is taken from the same simulations whose Nu trend is then reproduced. The abstract's own wording ('scaling argument suggests', 'as observed in the EXT-NELT case', 'model captures the main trends of the numerical data') shows that the explanatory step uses diagnosed flow quantities rather than an a priori prediction. The Nu~Ra^{1/4} to Nu~Ra^{1/3} range also spans two power laws, consistent with a non-asymptotic transition rather than a new dissipation closure, though that concern is about correctness rather than circularity. Because the load-bearing control variable is internal to the data being explained, the mechanistic account is partially circular, warranting a score of 5. The supplied full text is an unrelated RadarQA manuscript, so no additional physics equations could be checked beyond the abstract.
Assumptions & free parameters
free parameters (3)
- Heat transport scaling exponent gamma =
1/4 to 1/3 for EXT-NELT
- Scaling prefactor C in Nu = C Ra^gamma =
not stated
- Characteristic plume height z-hat =
up to ~H (full cavity depth) in EXT-NELT
assumptions (4)
- domain assumption Standard Oberbeck-Boussinesq horizontal convection energy closure and Rossby scaling Nu ~ Ra^{1/5} are the correct reference behavior for the MON/LENT cases.
- domain assumption Two-dimensional DNS captures the reorganization and scaling that matter for this configuration.
- domain assumption A specific nonlinear equation of state approximating the 4°C density extremum (functional form and parameters not given in the abstract) drives the EXT-NELT results.
- ad hoc to paper The new potential-energy transfer term Phi_{i2} is controlled by the characteristic plume height z-hat, and z-hat ~ H in the EXT-NELT case.
invented entities (1)
-
Phi_{i2}, an additional potential-energy transfer term from the nonlinear equation of state
Cite this review
Pith. "Pith review of Flow reorganization and transport enhancement in two-dimensional horizontal convection near a density extremum." pith.science (2026). https://pith.science/paper/7LOBAJL3
@misc{pith2026250812289,
author = {Pith},
title = {Pith review of: Flow reorganization and transport enhancement in two-dimensional horizontal convection near a density extremum},
year = {2026},
howpublished = {\url{https://pith.science/paper/7LOBAJL3}},
note = {Machine review of arXiv:2508.12289}
}
abstract
Horizontal convection serves as a canonical model for geophysical and industrial flows. While the Oberbeck-Boussinesq approximation is well established, the impact of a nonlinear equation of state, specifically the density extremum of water near $4^\circ\mathrm{C}$, remains underexplored. Here we investigate this effect using two-dimensional direct numerical simulations over the Rayleigh number range $10^6 \le Ra \le 5\times 10^{10}$. We examine four configurations, contrasting extremum (EXT) and monotonic (MON) buoyancy boundary conditions against linear (LENT) and nonlinear (NELT) equations of state. Our results reveal that the EXT-NELT case undergoes a pronounced reorganization of the large-scale flow, evolving from a bicellular structure to a single-roll circulation driven by central `mixing plumes'. This reorganization manifests as transitional anomalies in the $Re$ scaling, while the emergence of full-depth plumes alters the heat transport mechanism. Consequently, distinct from the Rossby scaling ($Nu \sim Ra^{1/5}$) observed in the reference cases, the EXT-NELT case exhibits an enhanced heat transport scaling ranging from $Nu \sim Ra^{1/4}$ to $Nu \sim Ra^{1/3}$. To interpret this behaviour, we examine the total energy budget and identify an additional potential-energy transfer term, \(\Phi_{i2}\), arising from the nonlinear equation of state. The scaling argument suggests that the magnitude of this contribution is controlled by the characteristic plume height ($\hat{z}$). Specifically, when plumes penetrate the entire cavity depth ($\hat{z} \sim H$), as observed in the EXT-NELT case, the global kinetic energy dissipation is no longer described by the standard OB HC energy closure alone. The resulting model captures the main trends of the numerical data and provides a possible energy budget interpretation of the enhanced transport observed in this two-dimensional configuration.
Reference graph
Works this paper leans on
-
[1]
Agrawal, K
H. Agrawal, K. Desai, Y . Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson. Nocaps: Novel object captioning at scale. In ICCV, 2019
2019
-
[2]
Antol, A
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh. VQA: Visual question answering. In ICCV, 2015
2015
-
[3]
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923, 2025
arXiv 2025
-
[4]
Banerjee and A
S. Banerjee and A. Lavie. METEOR: An automatic metric for mt evaluation with improved correlation with human judgments. In ACL Workshops, 2005
2005
-
[5]
K. Bi, L. Xie, H. Zhang, X. Chen, X. Gu, and Q. Tian. Pangu-weather: A 3d high-resolution model for fast and accurate global weather forecast. arXiv preprint arXiv:2211.02556, 2022
arXiv 2022
-
[6]
X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Dong, Q. Du, Z. Fu, et al. Deepseek LLM: Scaling open-source language models with longtermism. arXiv preprint arXiv:2401.02954, 2024
arXiv 2024
-
[7]
J. Chen, P. Zhou, Y . Hua, D. Chong, M. Cao, Y . Li, Z. Yuan, B. Zhu, and J. Liang. Vision-language models meet meteorology: Developing models for extreme weather events detection with heatmaps. arXiv preprint arXiv:2406.09838, 2024
arXiv 2024
-
[8]
K. Chen, T. Han, J. Gong, L. Bai, F. Ling, J.-J. Luo, X. Chen, L. Ma, T. Zhang, R. Su, et al. Fengwu: Pushing the skillful global medium-range weather forecast beyond 10 days lead. arXiv preprint arXiv:2304.02948, 2023
arXiv 2023
Show all 88 references
-
[9]
X. Chen, H. Fang, T.-Y . Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015
2015 arXiv
-
[10]
Z. Chen, W. Wang, Y . Cao, Y . Liu, Z. Gao, E. Cui, J. Zhu, S. Ye, H. Tian, Z. Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271, 2024
2024 arXiv
-
[11]
Davis, B
C. Davis, B. Brown, and R. Bullock. Object-based verification of precipitation forecasts. part i: Methodology and application to mesoscale rain areas. Monthly Weather Review, 134(7):1772–1784, 2006
2006
-
[12]
Davis, B
C. Davis, B. Brown, and R. Bullock. Object-based verification of precipitation forecasts. part ii: Application to convective rain systems. Monthly Weather Review, 134(7):1785–1795, 2006
2006
-
[13]
Donaldson, R
R. Donaldson, R. M. Dyer, and M. J. Kraus. An objective evaluator of techniques for predicting severe weather events. Preprints, Ninth Conf. on Severe Local Storms, Norman, OK, Amer. Meteor. Soc, 1975
1975
-
[14]
J. P. Finley. Tornado predictions. American Meteorological Journal. A Monthly Review of Meteorology and Allied Branches of Study (1884-1896), 1884
-
[15]
Y . Gao, H. Wu, R. Shu, H. Dong, F. Xu, R. Chen, Y . Yan, Q. Wen, X. Hu, K. Wang, et al. Oneforecast: A universal framework for global and regional weather forecasting. arXiv preprint arXiv:2502.00338, 2025
2025
-
[16]
Z. Gao, X. Shi, H. Wang, Y . Zhu, Y . B. Wang, M. Li, and D.-Y . Yeung. EarthFormer: Exploring space-time transformers for earth system forecasting. In NeurIPS, 2022
2022
-
[17]
Z. Gao, C. Tan, L. Wu, and S. Z. Li. Simvp: Simpler yet better video prediction. In CVPR, 2022. 10
2022
-
[18]
Q. Ge, W. Sun, Y . Zhang, Y . Li, Z. Ji, F. Sun, S. Jui, X. Min, and G. Zhai. LMM-VQA: Advancing video quality assessment with large multimodal models. arXiv preprint arXiv:2408.14008, 2024
2024 arXiv
-
[19]
T. GLM, A. Zeng, B. Xu, B. Wang, C. Zhang, D. Yin, D. Zhang, D. Rojas, G. Feng, H. Zhao, et al. ChatGLM: A family of large language models from glm-130b to glm-4 all tools. arXiv preprint arXiv:2406.12793, 2024
2024 arXiv
-
[20]
F. Gofa, D. Boucouvala, P. Louka, and H. Flocas. Spatial verification approaches as a tool to evaluate the performance of high resolution precipitation forecasts. Atmospheric Research, 2018
2018
-
[21]
J. Gong, L. Bai, P. Ye, W. Xu, N. Liu, J. Dai, X. Yang, and W. Ouyang. Cascast: Skillful high-resolution precipitation nowcasting via cascaded modelling. arXiv preprint arXiv:2402.04290, 2024
2024 arXiv
-
[22]
Goyal, T
Y . Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In CVPR, 2017
2017
-
[23]
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025
2025 arXiv
-
[24]
X. He, Z. Zhou, W. Zhang, X. Zhao, H. Chen, S. Chen, and L. Bai. DiffSR: Learning radar reflectivity synthesis via diffusion model from satellite observations. In ICASSP, 2025
2025
-
[25]
K. A. Hilburn, I. Ebert-Uphoff, and S. D. Miller. Development and interpretation of a neural-network-based synthetic radar reflectivity estimator using goes-r satellite observations. Journal of Applied Meteorology and Climatology, 2020
2020
-
[26]
E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, W. Chen, et al. Lora: Low-rank adaptation of large language models. In ICLR, 2022
2022
-
[27]
Hurst, A
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al. GPT-4o system card. arXiv preprint arXiv:2410.21276, 2024
2024 arXiv
-
[28]
Jaech, A
A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024
2024 arXiv
-
[29]
Z. Jia, Z. Zhang, J. Qian, H. Wu, W. Sun, C. Li, X. Liu, W. Lin, G. Zhai, and X. Min. VQA � : Visual question answering for video quality assessment. arXiv preprint arXiv:2411.03795, 2024
2024 arXiv
-
[30]
I. T. Jolliffe and D. B. Stephenson. Forecast verification: a practitioner’s guide in atmospheric science. John Wiley & Sons, 2012
2012
-
[31]
B. Li, Y . Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y . Li, Z. Liu, et al. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024
2024 arXiv
-
[32]
F. Li, R. Zhang, H. Zhang, Y . Zhang, B. Li, W. Li, Z. Ma, and C. Li. Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models. arXiv preprint arXiv:2407.07895, 2024
2024 arXiv
-
[33]
W. Li, X. Zhang, S. Zhao, Y . Zhang, J. Li, L. Zhang, and J. Zhang. Q-Insight: Understanding image quality via visual reinforcement learning. arXiv preprint arXiv:2503.22679, 2025
2025 arXiv
-
[34]
B. Lin, Y . Ye, B. Zhu, J. Cui, M. Ning, P. Jin, and L. Yuan. Video-llava: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122, 2023
2023 arXiv
-
[35]
C.-Y . Lin. Rouge: A package for automatic evaluation of summaries. In ACL, 2004
2004
-
[36]
F. Liu, Y . Wang, T. Wang, and V . Ordonez. Visual news: Benchmark and challenges in news image captioning. arXiv preprint arXiv:2010.03743, 2020
2010 arXiv
-
[37]
H. Liu, C. Li, Q. Wu, and Y . J. Lee. Visual instruction tuning. In NeurIPS, 2023
2023
-
[38]
Y . Liu, H. Duan, Y . Zhang, B. Li, S. Zhang, W. Zhao, Y . Yuan, J. Wang, C. He, Z. Liu, et al. MMBench: Is your multi-modal model an all-around player? In ECCV, 2024
2024
-
[39]
H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yang, et al. Deepseek-vl: towards real-world vision-language understanding. arXiv preprint arXiv:2403.05525, 2024
2024 arXiv
-
[40]
C. Ma, Z. Hua, A. Anderson-Frey, V . Iyer, X. Liu, and L. Qin. WeatherQA: Can multimodal language models reason about severe weather? arXiv preprint arXiv:2406.11217, 2024. 11
2024 arXiv
-
[41]
Marino, M
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi. OK-VQA: A visual question answering benchmark requiring external knowledge. In CVPR, 2019
2019
-
[42]
A. H. Murphy. The finley affair: A signal event in the history of forecast verification. Weather and forecasting, 1996
1996
-
[43]
Palmer and R
W. Palmer and R. Allen. Note on the accuracy of forecasts concerning the rain problem. US Weather Bureau, 1949
1949
-
[44]
H. A. Panofsky and G. W. Brier. Some applications of statistics to meteorology . Mineral Industries Extension Services, College of Mineral Industries, Pennsylvania State University, 1958
1958
-
[45]
Papineni, S
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu. Bleu: a method for automatic evaluation of machine translation. In ACL, 2002
2002
-
[46]
Ravuri, K
S. Ravuri, K. Lenc, M. Willson, D. Kangin, R. Lam, P. Mirowski, M. Fitzsimons, M. Athanassiadou, S. Kashem, S. Madge, et al. Skilful precipitation nowcasting using deep generative models of radar.Nature, 2021
2021
-
[47]
Rempel, F
M. Rempel, F. Senf, and H. Deneke. Object-based metrics for forecast verification of convective development with geostationary satellite data. Monthly Weather Review, 145(8):3161–3178, 2017
2017
-
[48]
Robinson, J
M. Robinson, J. Evans, and B. Crowe. En route weather depiction benefits of the nexrad vertically integrated liquid water product utilized by the corridor integrated weather system. In Conference on aviation, range and aerospace meteorology, american meteorological society, 2002
2002
-
[49]
Schulman, F
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017
2017 arXiv
-
[50]
Schwenk, A
D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi. A-OKVQA: A benchmark for visual question answering using world knowledge. In ECCV, 2022
2022
-
[51]
H. Shao, S. Qian, H. Xiao, G. Song, Z. Zong, L. Wang, Y . Liu, and H. Li. Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning. In NeurIPS, 2024
2024
-
[52]
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y . Li, Y . Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024
2024 arXiv
-
[53]
D. B. Stephenson, B. Casati, C. Ferro, and C. Wilson. The extreme dependency score: A non-vanishing measure for forecasts of rare events. Meteorological Applications: A journal of forecasting, practical applications, training techniques and modelling, 2008
2008
-
[54]
Stock, K
J. Stock, K. Hilburn, I. Ebert-Uphoff, and C. Anderson. Srvit: Vision transformers for estimating radar reflectivity from satellite observations at scale. arXiv preprint arXiv:2406.16955, 2024
2024 arXiv
-
[55]
K. Sun, J. Pan, Y . Ge, H. Li, H. Duan, X. Wu, R. Zhang, A. Zhou, Z. Qin, Y . Wang, et al. Journeydb: A benchmark for generative image understanding. In NeurIPS, 2023
2023
-
[56]
Touvron, T
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[57]
Veillette, S
M. Veillette, S. Samsi, and C. Mattioli. SEVIR: A storm event imagery dataset for deep learning applications in radar and satellite meteorology. In NeurIPS, 2020
2020
-
[58]
F. Wang, M. Chen, X. He, Y . Zhang, F. Liu, Z. Guo, Z. Hu, J. Wang, J. Xu, Z. Li, et al. Omniearth- bench: Towards holistic evaluation of earth’s six spheres and cross-spheres interactions with multimodal observational earth data. arXiv preprint arXiv:2505.23522, 2025
2025
-
[59]
W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y . Wang, J. Ji, Z. Yang, L. Zhao, S. XiXuan, et al. Cogvlm: Visual expert for pretrained language models. In NeurIPS, 2024
2024
-
[60]
Y . Wang, M. Long, J. Wang, Z. Gao, and P. S. Yu. PredRNN: Recurrent neural networks for predictive learning using spatiotemporal lstms. In NeurIPS, 2017
2017
-
[61]
Y . Wang, Y . Zeng, J. Zheng, X. Xing, J. Xu, and X. Xu. VideoCoT: A video chain-of-thought dataset with active annotation tool. arXiv preprint arXiv:2407.05355, 2024. 12
2024 arXiv
-
[62]
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE TIP, 2004
2004
-
[63]
C. J. Willmott and K. Matsuura. Advantages of the mean absolute error (mae) over the root mean square error (rmse) in assessing average model performance. Climate research, 2005
2005
-
[64]
H. Wu, Z. Zhang, E. Zhang, C. Chen, L. Liao, A. Wang, C. Li, W. Sun, Q. Yan, G. Zhai, et al. Q-Bench: A benchmark for general-purpose foundation models on low-level vision. arXiv preprint arXiv:2309.14181, 2023
2023 arXiv
-
[65]
H. Wu, Z. Zhang, E. Zhang, C. Chen, L. Liao, A. Wang, K. Xu, C. Li, J. Hou, G. Zhai, et al. Q-Instruct: Improving low-level visual abilities for multi-modality foundation models. In CVPR, 2024
2024
-
[66]
H. Wu, H. Zhu, Z. Zhang, E. Zhang, C. Chen, L. Liao, C. Li, A. Wang, W. Sun, Q. Yan, et al. Towards open-ended visual quality comparison. In ECCV, 2024
2024
-
[67]
Z. Wu, X. Chen, Z. Pan, X. Liu, W. Liu, D. Dai, H. Gao, Y . Ma, C. Wu, B. Wang, et al. Deepseek- vl2: Mixture-of-experts vision-language models for advanced multimodal understanding. arXiv preprint arXiv:2412.10302, 2024
2024 arXiv
-
[68]
J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y . Fan, K. Dang, et al. Qwen2.5-omni technical report. arXiv preprint arXiv:2503.20215, 2025
2025 arXiv
-
[69]
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, et al. Qwen2.5 technical report. arXiv preprint arXiv:2412.15115, 2024
2024 arXiv
-
[70]
J. Ye, H. Xu, H. Liu, A. Hu, M. Yan, Q. Qian, J. Zhang, F. Huang, and J. Zhou. mplug-owl3: Towards long image-sequence understanding in multi-modal large language models. arXiv preprint arXiv:2408.04840, 2024
2024 arXiv
-
[71]
Q. Ye, H. Xu, J. Ye, M. Yan, A. Hu, H. Liu, Q. Qian, J. Zhang, and F. Huang. mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration. In CVPR, 2024
2024
-
[72]
Z. You, J. Gu, Z. Li, X. Cai, K. Zhu, C. Dong, and T. Xue. Descriptive image quality assessment in the wild. arXiv preprint arXiv:2405.18842, 2024
2024
-
[73]
Z. You, Z. Li, J. Gu, Z. Yin, T. Xue, and C. Dong. Depicting beyond scores: Advancing image quality assessment through multi-modal language models. In ECCV, 2024
2024
-
[74]
Z. You, X. Cai, J. Gu, T. Xue, and C. Dong. Teaching large language models to regress accurate image quality scores using score distribution. In CVPR, 2025
2025
-
[75]
Young, A
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2014
2014
-
[76]
D. Yu, X. Li, Y . Ye, B. Zhang, C. Luo, K. Dai, R. Wang, and X. Chen. Diffcast: A unified framework via residual diffusion for precipitation nowcasting. In CVPR, 2024
2024
-
[77]
Q. Yu, Z. Zhang, R. Zhu, Y . Yuan, X. Zuo, Y . Yue, T. Fan, G. Liu, L. Liu, X. Liu, et al. Dapo: An open-source llm reinforcement learning system at scale. arXiv preprint arXiv:2503.14476, 2025
2025 arXiv
-
[78]
Zhang, X
P. Zhang, X. Dong, B. Wang, Y . Cao, C. Xu, L. Ouyang, Z. Zhao, H. Duan, S. Zhang, S. Ding, et al. Internlm- xcomposer: A vision-language large model for advanced text-image comprehension and composition. arXiv preprint arXiv:2309.15112, 2023
2023 arXiv
-
[79]
Zhang, V
T. Zhang, V . Kishore, F. Wu, K. Q. Weinberger, and Y . Artzi. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675, 2019
1904 arXiv
-
[80]
Zhang, H
Y . Zhang, H. Yu, M. Zhang, Y . Yang, and Z. Meng. Uncertainties and error growth in forecasting the record-breaking rainfall in zhengzhou, henan on 19–20 july 2021. Science China Earth Sciences, 2022
2021
-
[81]
Zhang, M
Y . Zhang, M. Long, K. Chen, L. Xing, R. Jin, M. I. Jordan, and J. Wang. Skilful nowcasting of extreme precipitation with nowcastnet. Nature, 2023
2023
-
[82]
Zhang, Z
Z. Zhang, Z. Jia, H. Wu, C. Li, Z. Chen, Y . Zhou, W. Sun, X. Liu, X. Min, W. Lin, et al. Q-Bench-Video: Benchmarking the video quality understanding of lmms. arXiv preprint arXiv:2409.20063, 2024
2024 arXiv
-
[83]
Zhang, H
Z. Zhang, H. Wu, Y . Zhou, C. Li, W. Sun, C. Chen, X. Min, X. Liu, W. Lin, and G. Zhai. LMM-PCQA: Assisting point cloud quality assessment with LMM. In ACM MM, 2024. 13
2024
-
[84]
Zhang, T
Z. Zhang, T. Kou, S. Wang, C. Li, W. Sun, W. Wang, X. Li, Z. Wang, X. Cao, X. Min, et al. Q-Eval-100K: Evaluating visual quality and alignment level for text-to-vision content. arXiv preprint arXiv:2503.02357, 2025
2025 arXiv
-
[85]
X. Zhao, W. Xu, B. Liu, Y . Zhou, F. Ling, B. Fei, X. Yue, L. Bai, W. Zhang, and X.-M. Wu. Msearth: A benchmark for multimodal scientific comprehension of earth science. arXiv preprint arXiv:2505.20740, 2025
2025 arXiv
-
[86]
Zhong, Z
Q. Zhong, Z. Sun, H. Chen, J. Li, and L. Shen. Multi model forecast biases of the diurnal variations of intense rainfall in the beijing-tianjin-hebei region. Science China Earth Sciences, 2022
2022
-
[87]
M. Zhou, J. Wu, M. Chen, and L. Han. Comparative study on the performance of convlstm and convgru in classification problems—taking early warning of short-duration heavy rainfall as an example. Atmospheric and Oceanic Science Letters, 2024
2024
-
[88]
high value retain
Y . Zhou, Y . Wang, X. He, R. Xiao, Z. Li, Q. Feng, Z. Guo, Y . Yang, H. Wu, W. Huang, et al. Scientists’ first exam: Probing cognitive abilities of mllm via perception, understanding, and reasoning. arXiv preprint arXiv:2506.10521, 2025. 14 Appendix A Overview This Appendix i...
2025
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.