Pith. sign in

REVIEW 2 cited by

CLLMate: A Multimodal Benchmark for Weather and Climate Events Forecasting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.19058 v2 pith:M2GBWFYQ submitted 2024-09-27 cs.LG cs.AIcs.CLphysics.ao-ph

classification cs.LGcs.AIcs.CLphysics.ao-ph
keywords climatecllmatedataeventsforecastingweatherenvironmentalexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Forecasting weather and climate events is crucial for making appropriate measures to mitigate environmental hazards and minimize losses. However, existing environmental forecasting research focuses narrowly on predicting numerical meteorological variables (e.g., temperature), neglecting the translation of these variables into actionable textual narratives of events and their consequences. To bridge this gap, we proposed Weather and Climate Event Forecasting (WCEF), a new task that leverages numerical meteorological raster data and textual event data to predict weather and climate events. This task is challenging to accomplish due to difficulties in aligning multimodal data and the lack of supervised datasets. To address these challenges, we present CLLMate, the first multimodal dataset for WCEF, using 26,156 environmental news articles aligned with ERA5 reanalysis data. We systematically benchmark 23 existing MLLMs on CLLMate, including closed-source, open-source, and our fine-tuned models. Our experiments reveal the advantages and limitations of existing MLLMs and the value of CLLMate for the training and benchmarking of the WCEF task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Domain Specific Benchmarks for Evaluating Multimodal Large Language Models

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.

  2. Can We Predict the Unpredictable? Leveraging DisasterNet-LLM for Multimodal Disaster Classification

    cs.LG 2025-06 reject novelty 2.0 of 10

    A multimodal transformer combining GPT, CLIP, and a geospatial network is claimed to outperform prior disaster classifiers, but the supporting experiments are under-specified.

Pith tools