{"id":"ed3004a4-6bff-4463-a0cb-66fa773e5d9e","arxiv_id":"2508.03590","paper_version":2,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A satellite-image-only deep learning model forecasts 24-hour solar irradiance across the US with lower RMSE than NOAA's HRRR numerical weather prediction model.","lead":"SolarSeer is an AI model that forecasts 24-hour cloud cover and solar irradiance across the US directly from satellite images, without running a weather simulation. It reports 15 to 27 percent lower forecast error than the operational HRRR model while running about 1,500 times faster.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline 27.28% RMSE gain mixes 24-hour SolarSeer forecasts with 18-hour HRRR forecasts at most initialization hours; the claimed advantage must be re-evaluated on matched horizons.","rationale":"The reader's verdict is already CONDITIONAL and its rationale mentions the mixed 24h-vs-18h comparison, but the reader's weakest_assumption focused on whether 6 hours of satellite imagery carries enough information for 24-hour cloud prediction. My stress-test identifies a more direct, protocol-level threat to the strongest claim: the reported 27.28% RMSE reduction is an average over hours where HRRR is only being asked to forecast 18 hours ahead. Since error typically increases with lead time, even a perfect 24h-vs-24h comparison could show a much smaller advantage. The paper openly discloses this design, which is why the issue is internally checkable rather than a matter of outside consensus. A horizon-matched re-analysis is the single most informative check: it can either salvage the central claim or show that the headline number is an artifact of unequal forecast lengths. I do not propose changing the overall CONDITIONAL verdict, because the paper may still be right after re-analysis, and the station-level comparison (15.35%) is not scrutinized in the available text; but the condition for acceptance should explicitly include a matched-horizon benchmark. Hence UNCHANGED, with this condition sharpened.","tokens_in":4773,"tokens_out":4914,"duration_ms":53549,"concrete_test":"Recompute the CONUS-average RMSE/MAE reductions for solar irradiance two ways: (1) using only initializations at UTC 00, 06, 12, and 18 where HRRR has true 24-hour forecasts; and (2) using all 24 initialization hours but truncating SolarSeer forecasts to the common 18-hour lead. If the horizon-matched reduction is substantially smaller than 27.28% (for example, below 10% or no longer statistically separable), the headline overstates the 24-hour skill advantage and the paper should be revised to report matched-horizon metrics.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—27.28% lower RMSE than HRRR for 24-hour solar irradiance forecasts—rests on a comparison that is not consistently 24h-vs-24h. The paper states in 'Cloud cover forecast results' that HRRR provides 48-hour forecasts only at UTC 00/06/12/18 and 18-hour forecasts at all other hours, while SolarSeer always provides 24-hour forecasts. The authors then 'evaluated the 24-hour forecasts of SolarSeer and the 18-hour forecasts of HRRR-NWP at other initial times throughout the day.' Thus the reported CONUS-average reductions (29.13% cloud cover RMSE, 27.28% irradiance RMSE, 36.08% irradiance MAE) aggregate a majority of hours where HRRR is being scored at a shorter, easier lead time. Because forecast error grows with lead time, this protocol mechanically inflates SolarSeer's apparent advantage. The claim that SolarSeer outperforms HRRR for true 24-hour forecasts is therefore not established by the headline numbers; only the four synoptic initialization hours provide a clean 24h-vs-24h test. This is an internal fairness issue, independent of any consensus about AI weather models.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces SolarSeer, an end-to-end AI model that maps six hours of satellite imagery to 24-hour cloud-cover and solar-irradiance forecasts over the contiguous United States at roughly 5-km resolution, bypassing data assimilation and PDE solving. Against NOAA's operational HRRR NWP, the authors report large average RMSE reductions (29.13% for cloud cover, 27.28% for irradiance on ERA5 ground truth, and 15.35% across 1,800 stations) and claim inference under 3 seconds, more than 1,500 times faster than HRRR. The central claim is that SolarSeer is the first end-to-end AI model to outperform state-of-the-art NWP for 24-hour solar irradiance forecasting.","tokens_in":4932,"tokens_out":4939,"duration_ms":58617,"significance":"If substantiated, the result is significant for solar energy forecasting: a fast, end-to-end, satellite-based model that beats an operational NWP at kilometer scale would be practically valuable for day-ahead electricity markets and would support the broader trend of learned weather emulators. The paper deserves credit for benchmarking against a real operational model (HRRR), for reporting station-based validation in addition to reanalysis, and for emphasizing inference speed. However, the headline comparison currently mixes forecast horizons across initialization hours, so the quantitative claims need to be re-established on matched lead times before the significance can be assessed.","major_comments":[{"comment":"The evaluation protocol is explicitly mismatched for most initialization hours. The text states that at non-synoptic hours the authors evaluate \"the 24-hour forecasts of SolarSeer and the 18-hour forecasts of HRRR-NWP,\" while at UTC 00/06/12/18 they compare 24-hour forecasts of both. Because forecast error generally grows with lead time, comparing a 24-hour SolarSeer forecast against an 18-hour HRRR forecast at 20 of 24 initialization hours systematically inflates SolarSeer's apparent advantage. The reported CONUS averages (29.13% cloud-cover RMSE, 27.28% irradiance RMSE, 36.08% irradiance MAE) therefore do not establish the claim that SolarSeer outperforms HRRR at the same 24-hour lead time. The authors should recompute all aggregate metrics using only matched lead times (e.g., lead times 1-18 h for all hours and 24 h only for the four synoptic hours) and report the 24-hour-vs-24-hour comparison separately. Without this, the central claim of \"outperforming NWP for 24-hour forecasts\" is not supported by the current numbers.","section":"Results: Cloud cover forecast results / Solar irradiance forecast results"},{"comment":"The repeated use of \"significantly\" (e.g., \"SolarSeer significantly outperforms HRRR-NWP,\" \"significantly reduces the root mean squared error\") is not accompanied by any uncertainty intervals or statistical significance tests. Given the strong spatial and temporal correlations in the data, the very large domain fractions (99.99% of CONUS) do not by themselves establish that the differences are beyond sampling noise. The authors should provide confidence intervals (for example, via block bootstrap over forecast initialization dates) for the headline RMSE/MAE reductions and for the 15.35% station-level result, or otherwise state the uncertainty associated with each point estimate.","section":"Results: Solar irradiance forecast results"}],"minor_comments":[{"comment":"Typos: \"start-of-the-art\" should be \"state-of-the-art\" in the sentence describing HRRR-NWP.","section":"Solar irradiance forecast results"},{"comment":"The abstract claims that SolarSeer \"significantly enhances the first-order irradiance difference forecasting accuracy,\" but no results or metrics for this claim appear in the main text; please add the corresponding analysis or remove the claim from the abstract.","section":"Abstract"},{"comment":"The station-based evaluation across 1,800 stations is mentioned in the abstract and results overview, but the main text does not describe the station data source, the evaluation period, or whether lead times are matched there; please provide these details.","section":"Abstract / Results"},{"comment":"The statement that SolarSeer produces forecasts \"in under 3 seconds\" would benefit from specifying the hardware and whether the time includes input/output and preprocessing; if this is given in the Methods, please cite the relevant subsection here.","section":"SolarSeer overview"}],"recommendation":"major_revision","confidential_remarks":"The paper is promising and the engineering contribution is real, but the headline comparison must be fixed before publication. If the matched-horizon analysis still shows substantial improvements, the paper could become acceptable. I would also recommend that the editor check whether the journal's statistical standards are satisfied by the current use of \"significantly\" without uncertainty quantification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"SolarSeer is a serious engineering result, but the headline 27.28% improvement over HRRR is not based on a clean 24h-vs-24h comparison at most initialization hours.\n\nWhat's genuinely new: an end-to-end model mapping six hours of satellite imagery to 24-hour cloud and irradiance forecasts over CONUS at 5 km, running in under three seconds. The two-block architecture (AFNO for clouds, Swin for irradiance) with clear-sky irradiance as a physical prior is sensible, and the reported station-level 15.35% RMSE reduction across 1,800 stations is a meaningful external benchmark. I believe this is the first large AI model aimed specifically at solar irradiance with a demonstrated NWP comparison, so the paper fills a real gap.\n\nThe main soft spot is the evaluation protocol. The authors state that at non-synoptic initial times they evaluate SolarSeer's 24-hour forecasts against HRRR's 18-hour forecasts. That covers 20 of 24 initialization hours. For solar irradiance, the extra hours in SolarSeer's forecast are frequently nighttime, where error is trivially zero, so aggregating these hours can make SolarSeer look better than a matched 24h-vs-24h comparison would. The stress-test note is right to flag this. The clean synoptic-hour comparison should be reported separately, and the authors should provide a matched valid-time or lead-time analysis. This is fixable, but it means the 27.28% and 29.13% headline numbers are not yet supported.\n\nThe paper also lacks uncertainty intervals or significance tests. The station comparison is the most convincing piece but is underspecified in the text I've seen: station selection, data period, and whether stations were held out from training aren't described. The ERA5-based 27.28% is harder to trust if ERA5-like outputs were part of the training target, though the station result partially addresses that.\n\nOverall, the direction of the result is credible, and the engineering contribution is real. The evaluation needs another pass before the claims are rigorous. I'd send it to peer review with the request for a matched-horizon analysis, confidence intervals, and full dataset/code release.\n\nThis is worth bringing to a reading group; the architecture and training approach are worth studying, and the evaluation critique is instructive.","headline":"SolarSeer is a serious engineering result, but the headline 27.28% improvement over HRRR is not a clean 24h-vs-24h comparison at most initialization hours.","tokens_in":5607,"tokens_out":5695,"would_cite":true,"duration_ms":61240,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SolarSeer maps six hours of satellite imagery to a 24-hour, 5-km solar irradiance forecast, and the paper reports it beats HRRR by 27.28% in RMSE on reanalysis data and 15.35% at 1,800 stations, while running over 1,500 times faster.","keywords":["solar irradiance forecasting","cloud cover forecasting","satellite imagery","numerical weather prediction","deep learning","transformer","renewable energy","HRRR"],"falsifier":"Hold out a set of days with rapid convective initiation where the six hours of preceding imagery shows little sign of cloud growth, and compare SolarSeer's 24-hour cloud RMSE against HRRR's; if SolarSeer's advantage disappears or reverses on those days, the satellite-only premise fails.","tokens_in":1557,"feed_emoji":"☀️","tokens_out":3722,"duration_ms":102044,"temperature":0.7,"pith_summary":"SolarSeer is an end-to-end AI model that claims to beat the best operational numerical weather prediction in the contiguous United States at forecasting solar irradiance. It does this by mapping just six hours of satellite imagery directly onto a 24-hour cloud cover and irradiance forecast at 5 km resolution, skipping data assimilation and partial differential equation (PDE) based atmosphere simulation entirely. The paper reports a 27.28% lower RMSE than HRRR on ERA5 reanalysis and 15.35% lower RMSE across 1,800 ground stations, while running more than 1,500 times faster, in under 3 seconds. A sympathetic reader should care because day-ahead irradiance accuracy is what lets grid operators trust solar power in electricity markets, and prior AI methods only worked for short lead times.","feed_headline":"AI beats top US weather model at 24-hour solar forecasts","feed_subtitle":"From six hours of satellite images, SolarSeer cuts next-day irradiance error by up to 27 percent, in seconds.","key_machinery":"The carrying mechanism is a two-block network: a cloud block of Adaptive Fourier Neural Operator (AFNO) Transformer layers, a frequency-domain Transformer that learns cloud motion and evolution, maps the six-hour satellite image sequence to the future 24-hour cloud cover. An irradiance block of Swin Transformer layers, a shifted-window vision Transformer for spatial detail, then combines that cloud forecast with the deterministic clear-sky irradiance from the Ineichen-Perez model, which supplies the Sun-Earth geometry as a physical prior. Bypassing data assimilation is what removes the supercomputer-level cost: the model goes from raw satellite images to forecasts end to end.","core_discovery":"The paper's central claim is that a learned mapping from historical satellite observations to future cloud cover can replace the data-assimilation-plus-PDE pipeline for solar irradiance forecasting and still outperform the operational NWP baseline. SolarSeer outputs both cloud cover and irradiance; its cloud block reduces 24-hour cloud-cover RMSE by 29.13% and MAE by 25.24% against HRRR, with 99.99% of CONUS grid cells improved. The irradiance block then reduces RMSE by 27.28% against ERA5 and 15.35% against 1,800 station observations, and improves forecast accuracy of first-order irradiance differences, meaning it captures ramps in solar output.","pith_inferences":["Editorial inference: If the satellite-only premise holds under weather regimes absent from training, such as wildfire smoke or rapid convection, the same architecture could be retrained for any region with geostationary satellite coverage, effectively exporting skill without a local high-resolution NWP.","Editorial inference: A natural stress test is to hold out entire seasons or extreme-event years; if skill drops sharply on rare cloud regimes, the model may be partly memorizing climatological cloud patterns rather than learning cloud processes from the images.","Editorial inference: The six-hour input window is a tunable design choice; varying it or adding coarse atmospheric state fields would reveal how much of the predictive signal actually comes from the imagery alone.","Editorial inference: Because the clear-sky prior fixes the irradiance geometry, the model's irradiance errors should concentrate where cloud attribution is wrong; separating errors on cloudy versus clear grid cells could identify whether the cloud block or the irradiance block is the limiting component."],"forward_implications":["With 24-hour irradiance forecasts produced in seconds, day-ahead market participants could update solar positions far more frequently than NWP currently allows.","The satellite-only design means operational solar forecasting no longer must wait for an NWP initialization cycle; a fresh satellite image can trigger a forecast at any hour.","Because SolarSeer improves cloud-cover forecasts at all 24 lead times and across nearly all CONUS regions, the irradiance error reductions follow in the same areas, with the largest gains in the western and southern United States.","Better first-order irradiance difference accuracy means rapid ramp events, a main threat to grid stability, are tracked more faithfully than by HRRR.","SolarSeer produces 24-hour forecasts at every initialization hour of the day, whereas HRRR only offers 24-hour forecasts at four initialization times, extending forecast availability for grid operators."],"supporting_citations":[{"why":"Define HRRR-NWP, the operational NWP baseline that SolarSeer is compared against for cloud cover and irradiance forecasts.","marker":"24, 25"},{"why":"Provides the RTMA reanalysis dataset used as ground truth for the total cloud cover evaluation.","marker":"26"},{"why":"Support the claim that HRRR-NWP is the state-of-the-art solar irradiance forecast in CONUS, making it the baseline to beat.","marker":"27–29"},{"why":"Provides the ERA5 reanalysis dataset used as ground truth for the solar irradiance evaluation.","marker":"30"}],"fun_headline_variants":["AI solar forecast beats HRRR by 27% in 3 seconds flat","SolarSeer: AI cuts solar irradiance error 27% vs top US model","1500x faster than NWP: AI predicts 24h solar output accurately","From satellite images, AI nails 24-hour solar irradiance forecasts","AI beats physics-based weather model for 24-hour solar irradiance"],"cache_read_input_tokens":7552,"weakest_assumption_plain":"The load-bearing premise is that six hours of satellite imagery contains enough information to predict the next 24 hours of cloud cover at 5 km resolution; any cloud development driven by variables invisible in the images, such as temperature and humidity, is something the model can only learn statistically from training data.","fun_headline_variants_meta":{"raw":{"variants":["AI solar forecast beats HRRR by 27% in 3 seconds flat","SolarSeer: AI cuts solar irradiance error 27% vs top US model","1500x faster than NWP: AI predicts 24h solar output accurately","From satellite images, AI nails 24-hour solar irradiance forecasts","AI beats physics-based weather model for 24-hour solar irradiance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000236,"raw_usage":{"total_tokens":1512,"prompt_tokens":962,"completion_tokens":550,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":450}},"tokens_in":578,"tokens_out":550,"duration_ms":6127,"temperature":1.0,"reasoning_tokens":450,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:20:24.782899+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Hold out a set of days with rapid convective initiation where the six hours of preceding imagery shows little sign of cloud growth, and compare SolarSeer's 24-hour cloud RMSE against HRRR's; if SolarSeer's advantage disappears or reverses on those days, the satellite-only premise fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the RTMA reanalysis dataset used as ground truth for the total cloud cover evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ERA5 reanalysis dataset used as ground truth for the solar irradiance evaluation."}],"review_version":1}