{"id":"452e9918-e257-4fef-a7ba-df0909ff1806","arxiv_id":"2411.13724","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"GPT-4o rainfall forecasts are conservative and close to historical averages; adding expert model inputs does not improve accuracy and often dampens predicted peaks.","lead":"This paper tests whether GPT-4o can forecast rainfall for 15 US cities over a 15-day and a 12-month horizon, with and without inputs from an LSTM expert model. It finds that GPT tends to generate conservative forecasts close to the 30-year historical average, even when given expert predictions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim of consistent conservatism rests on a single forecast window and single stochastic LLM runs; without multiple periods and repeated sampling, the conclusion generalizes beyond the evidence.","rationale":"The paper has a clear empirical design and the LSTM expert model provides a reasonable baseline. The authors also deserve credit for testing four prompt/input conditions and for reporting RMSE, correlation, and Nash-Sutcliffe efficiency. However, the central claim is about a stable behavioral tendency of GPT-4o, and the evidence covers one temporal realization without repeated sampling. The reader's weakest assumption (representativeness of the single test period) is exactly the load-bearing limitation here, and it is compounded by the absence of any uncertainty quantification around the LLM outputs. A copy-paste artifact in the Figure 4 caption further reduces confidence in the manuscript's polish, though it does not by itself invalidate the quantitative comparison. The proposed test, running additional rolling windows with multiple repeats, would directly test whether the conservative behavior is consistent or an artifact of the chosen period. Since the reader's verdict is already CONDITIONAL and this concern reinforces that conditionality rather than overturning the paper entirely, no verdict change is needed.","tokens_in":7099,"tokens_out":3163,"duration_ms":33790,"concrete_test":"Re-run the full protocol on 12 independent 15-day forecast windows (e.g., the first 15 days of each month from September 2023 through August 2024) and also on 12 overlapping 12-month windows, with 5 independent GPT-4o calls per condition and city. For each window, compute the mean RMSE and the correlation with the 30-year historical average. If the Exp1/Exp2 outputs are systematically closer to the 30-year average than EM across all windows, the 'consistent conservative tendency' claim is supported; if the effect varies by season or disappears outside October 2023, the current evidence is insufficient to support the general conclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The conclusion that GPT-4o 'consistently prioritizes stable predictions closely aligned with historical averages' is supported by exactly one short-term window (October 1-15, 2023) and one long-term window (October 2023-September 2024), with one set of outputs per condition per city. Section III fixes the baseline at September 30, 2023, and all RMSE, correlation, and Nash-Sutcliffe comparisons are computed over these same two windows. This is a single realization of a stochastic system: ChatGPT-4o is not deterministic, repeated sampling is not reported, and no confidence intervals or significance tests accompany the reported differences (e.g., short-term RMSE 0.23 vs. 0.20; long-term correlation with the 30-year average 0.86 vs. 0.59). The observed closeness to the 30-year average could be a property of this particular El-Nino-adjacent period, of the specific prompts, or of one lucky/unlucky draw from the model, rather than a stable behavioral tendency. Because the paper's central claim is explicitly about GPT's general forecasting behavior, the single-window, single-draw design is the load-bearing point that must be tested before the conclusion can be accepted.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper investigates whether GPT-4o can generate useful rainfall forecasts at short-term (15-day) and long-term (12-month) horizons for 15 U.S. cities. It compares four prompting conditions: GPT-only (Exp1), GPT with direct LSTM expert model rainfall predictions (Exp2), GPT with indirect temperature predictions (Exp3), and GPT with teleconnection indices (Exp4). The authors report that GPT-4o's outputs are closer to a 30-year historical average than to the expert model's outputs, and they conclude that GPT-4o consistently prioritizes stable, conservative predictions aligned with historical averages, even when provided with expert information. An additional experiment in the Discussion adds standard deviation as uncertainty information and shows improved agreement with the EM, but this experiment is not part of the core evaluation.","tokens_in":7304,"tokens_out":5384,"duration_ms":49309,"significance":"If the central finding were robust, it would provide a useful empirical characterization of a widely used LLM's behavior in a specialized forecasting domain: GPT-4o tends to smooth toward climatology rather than following expert-model signals. The paper's strengths include its multi-city design, the use of publicly available data, a transparent LSTM baseline, and an honest acknowledgment of the look-ahead nature of the teleconnection experiment. The four-way experimental design is a sensible framework for probing how LLMs respond to different types of auxiliary information. However, the evidence base is a single forecast window with no repeated sampling or uncertainty quantification, so the generalizability of the main claim is not yet established; the paper reads as an exploratory case study rather than a definitive evaluation.","major_comments":[{"comment":"The central claim of consistent conservative behavior is based on exactly one test period: October 1–15, 2023 for short-term and October 2023–September 2024 for long-term forecasts, with a single set of GPT-4o outputs per condition per city. GPT-4o is stochastic; no repeated sampling, confidence intervals, or significance tests accompany the reported differences (e.g., short-term RMSE 0.23 vs. 0.20; long-term correlations with the 30-year average of 0.86, 0.82, 0.76, and 0.62 vs. the EM's 0.59). The observed closeness to the 30-year average could be an artifact of this particular El Niño-adjacent period, the specific prompts, or one draw from the model. Because the Conclusion states that GPT \"consistently\" prioritizes stable predictions, the paper must evaluate multiple independent forecast periods and multiple stochastic draws before that generalization can be supported.","section":"Sections III and IV"},{"comment":"The teleconnection experiment provides GPT-4o with actual values of Nino3.4, PDO, and NAO for the forecast period itself (October 2023–September 2024). These values would not be known to a forecaster operating at the stated timestamp of September 30, 2023, so Exp4 is an oracle-input condition, not a realistic forecast condition. The paper acknowledges this in a note, but the subsequent interpretation treats Exp4 as evidence about how GPT uses teleconnection information in forecasting (e.g., \"adding global teleconnection factors, GPT's results declined\"). This framing conflates a data-leakage condition with a legitimate input scenario and undermines the comparative claims. The experiment should be either reframed as an oracle study or run with predicted/forecasted teleconnection indices.","section":"Section III, Experiment 4"},{"comment":"The claim that GPT-generated predictions \"closely align with the 30-year average\" is supported primarily by Pearson correlations between the predictions and the 30-year average (values of 0.86, 0.82, 0.76, and 0.62 for Exp1–Exp4, and 0.59 for the EM). A high correlation with a smooth climatological seasonal cycle may simply reflect the strong annual periodicity in rainfall, rather than a deliberate conservative strategy. To substantiate the interpretation that GPT is reverting to historical averages, the paper should also report direct distances (e.g., RMSE, MAE) between each method's outputs and the 30-year baseline, and ideally a skill score such as Nash-Sutcliffe efficiency computed against that baseline. Without such metrics, the \"alignment\" claim is not quantitatively established.","section":"Section IV, correlations with 30-year average"}],"minor_comments":[{"comment":"The paper inconsistently refers to the model as \"ChatGPT-4\" in some places and \"GPT-4o\" in others, including within the abstract text of the manuscript; choose one name and use it throughout.","section":"Abstract and Introduction"},{"comment":"The prompt sample for Experiment 4 says the prediction period is \"October 1, 2023, to October 15, 2023,\" but the experiment is described as a 12-month forecast; correct the period in the prompt to match the intended monthly-scale evaluation.","section":"Section III, Experiment 4 prompt"},{"comment":"The caption for Figure 4 appears to be a leftover from another manuscript (\"Comparison of predicted results and observation for different token setup and different lead time...\") and does not describe the short-term time series comparisons shown; replace it with a proper caption.","section":"Figure 4 caption"},{"comment":"The paper lists Nash-Sutcliffe efficiency as an evaluation metric but no NSE values are reported in the text or figures; either report NSE results in tables or remove the mention to avoid an unmet expectation.","section":"Section II, evaluation metrics"},{"comment":"The \"30-year historical average\" baseline is not precisely defined; specify the exact period (e.g., 1993–2022) used to compute the daily and monthly climatological means.","section":"Section II, baseline definition"},{"comment":"Reference [10] has an incomplete DOI string with a trailing period embedded in the text; provide the full DOI for the dataset citation.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The single-window, single-draw design is the most serious issue; the authors need to expand the evaluation to multiple periods and repeated samples, or explicitly reframe the paper as a case study. The look-ahead in Experiment 4 is a correctness concern that should be addressed directly. The paper has some interesting observations, but the current evidence does not support the general claim of \"consistent\" conservative behavior. A major revision with additional evaluation or substantially softened claims would make it acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a legitimate, clearly written evaluation of GPT-4o as a rainfall forecaster. The headline result is that for the one period tested, GPT's outputs sit very close to 30-year climatology, and feeding it LSTM expert predictions doesn't pull it away from that. That is a useful data point for anyone building LLM-based climate tools.\n\nWhat's new: the experiment design—comparing GPT alone, GPT with direct expert rainfall, GPT with indirect temperature inputs, and GPT with teleconnection indices—is not in the papers it cites. The comparison against the 30-year baseline is a sensible way to expose the averaging behavior. The time-series figures make the smoothing concrete, e.g., Atlanta's Oct 11-12 peak being damped to climatology. The discussion is appropriately measured and the proposed uncertainty-aware combination framework is a reasonable direction.\n\nWeaknesses, in proportion. The big one: the generality claim. 'Consistently prioritizes stable predictions' is supported by exactly one 15-day window and one 12-month window, one set of GPT outputs per condition per city. GPT-4o is stochastic; there are no repeated draws and no confidence intervals or significance tests, so we can't rule out a lucky/unlucky draw. The test period is also a single climate state. So the evidence supports 'GPT-4o did this in October 2023-September 2024,' not 'GPT-4o consistently does this.'\n\nExperiment 4 uses actual future teleconnection values as inputs—that is not a forecast setup, it's a leak. The poor performance there shouldn't be taken as evidence about GPT's forecasting skill; it's arguably testing something else. The copy-paste text in the Figure 4 caption looks like it's from another paper; I assume that's an artifact, but it needs fixing. Code is not provided.\n\nIf I were editing, I'd send this to review with major revision. The core observation is worth publishing as a cautionary result, but the conclusions need to be scaled back to the specific conditions, and the experiment needs more windows and proper sampling. It's not a desk reject.","headline":"Useful cautionary negative result about GPT-4o rainfall forecasts, but the 'consistent conservatism' claim needs more than one test window.","tokens_in":7850,"tokens_out":2097,"would_cite":false,"duration_ms":19994,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GPT-4o rainfall forecasts cling to 30-year averages in climate test.","keywords":["large language models","climate forecasting","rainfall prediction","GPT-4o","historical average baseline","conservative forecasts","LSTM expert model","extreme rainfall events"],"falsifier":"Run the same four prompts over a period containing a well-documented rainfall extreme, such as the 2015-2016 El Nino winter in the southern United States, with the model's timestamp set before that winter, and check whether GPT-4o's forecast deviations from the 30-year average grow large enough to track observed rainfall. The paper's claim predicts it will still damp the peaks; if instead the model responds to the extreme signal, the claim of consistent conservatism is false.","tokens_in":6852,"feed_emoji":"🌧️","tokens_out":6458,"duration_ms":56454,"temperature":0.7,"pith_summary":"The paper tests whether ChatGPT-4o can act as a climate forecaster by asking it to predict rainfall for 15 U.S. cities at 15-day and 12-month horizons. Across four prompt conditions, the model's forecasts stayed close to the 30-year historical average and never outperformed a two-layer LSTM expert model. Adding expert rainfall predictions, regional temperature hints, or global teleconnection indices did not help; GPT-4o smoothed the expert's peaks back toward average history. The authors conclude that GPT-4o is inherently biased toward conservative, stable outputs and is poorly suited, by itself, to capturing extreme rainfall events. This matters because the public increasingly turns to LLMs for accessible climate information, and a forecast that always reverts to average will hide exactly the anomalies that matter.","feed_headline":"GPT-4o rainfall forecasts cling to 30-year averages in climate test","feed_subtitle":"Expert model inputs don't help ChatGPT predict rain; it smooths extremes back toward history.","key_machinery":"The carrying mechanism is the 30-year historical average as an implicit anchoring baseline, made visible by comparing GPT-4o's outputs under four prompt conditions against that baseline. The expert model is a two-layer LSTM, a recurrent neural network that learns temporal dependencies from 60-day or 60-month input windows, used to generate the expert rainfall and temperature predictions that GPT-4o was asked to incorporate in some experiments. Correlations with the 30-year average diagnose the smoothing tendency: the higher GPT's correlation with history, the more it damped the expert model's peaks. The standard-deviation experiment adds historical monthly rainfall variability as an uncertainty signal, and when GPT-4o receives it, its forecasts move closer to the expert model rather than to the average.","core_discovery":"The central claim is that GPT-4o, when asked to produce numerical rainfall forecasts, consistently chooses stable predictions close to historical averages regardless of what additional information it is given. In short-term tests, average RMSE stayed around 0.20-0.23, far above the LSTM expert model's 0.06; in long-term tests, feeding expert data raised error and lowered correlation instead of improving results. Correlations between GPT-4o's outputs and the 30-year average were 0.86 for GPT alone, 0.82 with expert rainfall, 0.76 with regional temperatures, and 0.62 with teleconnection indices, while the expert model's correlation with the average was only 0.59. The paper interprets this as evidence that GPT substitutes text-pattern common sense for physical climate reasoning, reverting to safe historical norms whenever no strong trend signal is visible.","pith_inferences":["A testable extension the paper leaves implicit: if the behavior is a general pretraining artifact rather than a prompt effect, other LLM families should show the same averaging bias under the same protocol.","The paper's results imply a calibration strategy: measure how far an LLM's forecast sits from climatology and treat that distance as a trust score.","Since the evidence covers one test window, a strong El Nino or hurricane-season repeat would clarify whether the conservative bias is absolute or limited to quiet periods."],"forward_implications":["If GPT-4o always reverts to historical averages, LLM-based public climate tools will understate extreme rainfall risk unless they are explicitly constrained or fine-tuned.","Adding expert model outputs through prompts is not enough to change GPT-4o's behavior, so direct integration strategies need more than a prompt.","Using historical variability as an uncertainty cue can move GPT-4o's forecasts closer to expert-model results, suggesting a cheap improvement path.","Any practical LLM climate service should treat LLM outputs as baseline summaries, not event forecasts, until the averaging bias is overcome.","The smoothing tendency is stronger at longer horizons and more visible in high-rainfall cities."],"supporting_citations":[{"why":"Specifies the identity and architecture of the model whose forecasting behavior is under test.","marker":"[3]"},{"why":"Contrasts with LLM climate services that focus on descriptive output, setting up the gap this study addresses.","marker":"[7]"},{"why":"A representative specialized climate LLM for text generation, used to frame predictive ability as the missing component.","marker":"[8]"},{"why":"Cited as the source of the compiled daily temperature and precipitation data for U.S. cities that trains the expert model and forms the historical baseline.","marker":"[10]"},{"why":"Supplies the standard-deviation-of-historical-rainfall measure used as an uncertainty signal in the improved experiment.","marker":"[11]"}],"fun_headline_variants":["GPT-4o rain forecasts stick to history, ignore expert data","LLM climate prediction: ChatGPT default to 30-year averages","GPT-4o can't outguess rainfall history, even with expert help","ChatGPT climate forecasts play safe, nudging toward the mean","LLM rain forecasting fails to beat simple historical averages"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the single test window (October 1-15, 2023, and October 2023-September 2024) represents GPT-4o's general forecasting behavior; if that period is unusual or too short, the conclusion that GPT always reverts to historical averages does not follow.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4o rain forecasts stick to history, ignore expert data","LLM climate prediction: ChatGPT default to 30-year averages","GPT-4o can't outguess rainfall history, even with expert help","ChatGPT climate forecasts play safe, nudging toward the mean","LLM rain forecasting fails to beat simple historical averages"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00016,"raw_usage":{"total_tokens":1224,"prompt_tokens":932,"completion_tokens":292,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":204}},"tokens_in":548,"tokens_out":292,"duration_ms":3444,"temperature":1.0,"reasoning_tokens":204,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:56:44.177183+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same four prompts over a period containing a well-documented rainfall extreme, such as the 2015-2016 El Nino winter in the southern United States, with the model's timestamp set before that winter, and check whether GPT-4o's forecast deviations from the 30-year average grow large enough to track observed rainfall. The paper's claim predicts it will still damp the peaks; if instead the model responds to the extreme signal, the claim of consistent conservatism is false.","supporting_citations":[{"cited_title":"Thus spoke GPT-3: Interviewing a large-language model on climate finance,","cited_arxiv_id":null,"evidence_quote":"A representative specialized climate LLM for text generation, used to frame predictive ability as the missing component."},{"cited_title":"Figure 1 The locations of selected cities in the United States and their corresponding annual rainfall amounts","cited_arxiv_id":null,"evidence_quote":"Cited as the source of the compiled daily temperature and precipitation data for U.S. cities that trains the expert model and forms the historical baseline."},{"cited_title":"Use of Historical Data to Assess Regional Climate Change","cited_arxiv_id":null,"evidence_quote":"Supplies the standard-deviation-of-historical-rainfall measure used as an uncertainty signal in the improved experiment."}],"review_version":1}