{"id":"92ea38da-1846-4f98-9a85-c7e8f9347a04","arxiv_id":"2508.14107","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"SuryaBench provides a full-resolution, machine-learning-ready SDO solar image dataset with six benchmark tasks for space weather prediction.","lead":"The authors release SuryaBench, a machine-learning-ready dataset of full-resolution solar images from NASA's SDO mission, spanning 2010 to 2024. It includes six benchmark tasks for solar and space weather prediction with baseline deep-learning results.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Saturation clamping after AIA degradation correction is unvalidated and likely truncates the flare/AR pixels central to the DS1/DS4 benchmarks, threatening the high-fidelity claim.","rationale":"The reader's weakest assumption already names the saturation clamp as a concern, and additional flagged issues (data-volume inconsistency, label validation deferred to supplementary) support a conditional verdict. I focus on the saturation clamp because it is the most load-bearing for the paper's central high-fidelity claim: the dataset's core value is full-resolution, homogenized AIA data for flare/AR science, and the clamp acts directly on those regions. The paper honestly describes the procedure but supplies no validation, so the concern is concrete and testable. The numerical inconsistencies (467,400 files at ~600 MB implies ~280 TB, not 360 TB; missing fraction ~27% versus stated 6%) further weaken the Data Records section but are secondary to the physical-information question. Because the concern could be resolved by a straightforward audit and does not by itself prove the dataset unusable, the conditional verdict remains appropriate: the paper should be accepted after the authors provide saturation statistics and, if needed, a saturation-aware version of the affected benchmarks. No independent verification of the claimed HuggingFace release or label-generation code is possible from the manuscript alone, which also justifies keeping the conditional status.","tokens_in":10462,"tokens_out":5252,"duration_ms":69539,"concrete_test":"Scan all AIA timestamps in the released dataset and compute the fraction of pixels whose degradation-corrected value exceeds 16,383 DN before clamping, both globally and within active-region masks from HEK/SPoCA and within GOES flare-time windows. Then retrain or rerun the provided DS1 segmentation and DS4 flare baselines on an unclamped (or saturation-masked) version and compare TSS/IoU with the released clamped version. If the affected pixel fraction is below ~0.01% of AR pixels and baseline metrics shift by less than ~1%, the concern is resolved; otherwise the benchmark validity is compromised and the saturation treatment must be documented and corrected.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.1.1 states that degradation-corrected AIA images are clamped to the 16,383 DN saturation limit, but no statistics are given for how many pixels are affected or how this truncation interacts with the benchmark labels. The central claim emphasizes full native spatial resolution and high fidelity; yet in active regions and during flares, EUV intensities are highest and multiplicative degradation factors are largest, so these are precisely the pixels most likely to exceed the clamp threshold. Clamping creates flat-topped intensity plateaus and removes dynamic-range information that DS1 (AR segmentation) and DS4 (flare prediction) are designed to exploit. Figure 2 reports only mean pixel values, which are insensitive to saturation because clamping reduces the mean only weakly while strongly truncating the upper tail. The paper itself acknowledges the issue ('may lead to issues') but does not validate the assumption that no scientifically relevant information is lost. If even a small fraction of active-region pixels are clamped, the benchmark baselines and the dataset's stated fitness for flare/AR tasks are called into question. This is a load-bearing gap in the dataset's own quality story, not a disagreement with community consensus.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"SuryaBench introduces a curated, ML-ready dataset from NASA's Solar Dynamics Observatory (SDO), covering May 2010 through December 2024. The core dataset provides AIA extreme-ultraviolet/UV images (eight channels) and HMI line-of-sight/vector magnetograms at the native 4096×4096 pixel scale and a homogenized 12-minute cadence. Preprocessing includes roll-angle correction, plate-scale alignment, exposure normalization, degradation compensation, and fixed solar-disk-radius rescaling. Six auxiliary benchmark datasets target active-region segmentation, active-region emergence forecasting, coronal magnetic-field extrapolation, solar flare prediction, solar EUV spectra prediction, and solar wind speed estimation; baseline model results are reported for several tasks. The manuscript claims that SuryaBench is the largest curated and homogenized SDO dataset to date and is designed to support machine-learning benchmarking, self-supervised learning, and operational space-weather workflows.","tokens_in":10699,"tokens_out":6692,"duration_ms":82823,"significance":"If the data-quality claims hold, SuryaBench would be a valuable community resource: it avoids the spatial-resolution loss of the existing 512×512 SDO ML dataset, covers more than a solar cycle at a consistent 12-minute cadence, and packages six concrete benchmark tasks with public code and baselines. The use of standard tools (SunPy/aiapy), public JSOC calibration tables, explicit QUALITY-flag filtering, and a clear train/validation/test chronology are strengths. However, the paper's central 'high-fidelity' claim is not yet fully established because the degradation-corrected AIA images are clamped at the 16,383 DN saturation limit without quantitative validation of the affected pixels. A numerical inconsistency in the reported storage volume and the thin main-text description of label-generation protocols also need attention before the dataset can serve as a trustworthy benchmark resource.","major_comments":[{"comment":"The degradation-corrected AIA images are clamped to the 16,383 DN saturation limit. The paper itself states that this 'may lead to issues' but provides no statistics on the fraction of pixels affected, their spatial distribution, or the interaction with the DS1 and DS4 labels. Figure 2 reports only full-disk mean intensities, which are insensitive to upper-tail clamping. Because saturated pixels are expected in active regions and flare kernels—precisely the structures targeted by DS1 and DS4—the central high-fidelity claim is not yet supported. Please quantify saturated-pixel rates per channel and year, compare active-region versus quiet-Sun rates, and assess whether the benchmark labels or baseline results change under an alternative treatment (e.g., retaining pre-clamp values, using log-scaled intensities, or publishing a saturation mask).","section":"Sec. 2.1.1, Fig. 2"},{"comment":"The reported data volume is internally inconsistent. With ~600 MB per hourly netCDF file and 379,920 training files, the training set is ~228 TB, not the stated ~360 TB; adding the validation and test splits brings the total to roughly 280 TB. Please report exact per-file sizes (including data type and compression), the total size of each split, and the overall collection size. These figures are part of the public dataset record and directly affect usability and storage planning.","section":"Sec. 3"},{"comment":"The construction of the six benchmark labels is only summarized in the main text; the actual protocols (AR/PIL mask generation, flare and emergence event definitions, target windows, solar-wind target alignment, EVE spectral preprocessing, etc.) are deferred to the Supplementary Information. Since the correctness of DS1–DS6 is central to the benchmark package, the main manuscript should either include the label-generation protocols in sufficient detail or clearly point to the exact online documentation. It should also report key validation statistics, such as label distributions, event counts, and class balance, so that users can judge benchmark difficulty and potential label errors.","section":"Sec. 2.2 / Supplementary"},{"comment":"The claim that SuryaBench is 'the largest curated and homogenized dataset to date' is not quantitatively substantiated. The only explicit comparison is to the 512×512 resolution of Galvez et al. (2019). For a dataset paper, please add a comparison table covering spatial resolution, temporal cadence, temporal coverage, number of channels, and total size against existing SDO ML datasets, so that the 'largest' claim can be verified.","section":"Sec. 1 / Abstract"}],"minor_comments":[{"comment":"Baseline results are reported without error bars, confidence intervals, or the number of independent runs. Please state the number of seeds and the standard deviation across runs; otherwise it is difficult to judge whether the differences between architectures are meaningful.","section":"Sec. 4"},{"comment":"Please clarify units in the dynamic-range row: AIA values are DN after clamping, while the text also refers to DN/sec after exposure normalization; HMI vector components should specify their coordinate frame (e.g., radial, heliographic).","section":"Table 1"},{"comment":"The global '~6% missing data' figure should be disaggregated by year, channel, and instrument. Missing data is rarely uniform over a solar cycle, and users of the 12-minute sequence need to know where gaps concentrate.","section":"Sec. 3"},{"comment":"The name of the specific SunPy/aiapy function used for the level-1.5 promotion is garbled in the typeset text. Please cite the exact function and version so that the preprocessing is reproducible.","section":"Sec. 2.1.1"},{"comment":"DS3 has H=4,186 rather than 4,096, and DS6 has 1,343 output channels. These are unusual dimensions; a one-sentence explanation in the main text would help users understand the data layout.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The saturation-clamping validation is the main gate for this paper. If the authors can quantify the affected pixels and show that the DS1/DS4 labels and baselines are robust, or alternatively release an unclamped/saturation-masked version, the dataset would meet the stated quality bar. The storage-volume inconsistency should also be corrected before final acceptance. I recommend that the editor also ask to see the Supplementary Information document in the review loop, because the main text currently does not stand alone for the benchmark label protocols. Several benchmarks build directly on the authors' own prior work (Pandey et al., Arge et al.); that is not a problem per se, but the distinct contribution of SuryaBench should be made explicit in the revised manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a real new artifact: full 4096×4096 AIA/HMI data with a 12-minute cadence spanning most of a solar cycle, plus six benchmark tasks. That fills a genuine gap left by SDOML's 512×512 product, and the preprocessing pipeline is standard and sensible: roll correction, exposure normalization, orbit normalization, and quality filtering using public tools and calibration tables. I give them credit for putting it on HuggingFace with code and baselines, and for acknowledging the saturation clamp rather than hiding it.\n\nThe soft spots are real but fixable. The biggest is the saturation clamping in Section 2.1.1. The paper says degradation-corrected values can exceed the 16,383 DN limit and clamps them, but gives zero statistics on how many pixels are affected. Figure 2 only shows means, which are insensitive to upper-tail truncation. This matters because DS1 (active region segmentation) and DS4 (flare prediction) depend on exactly the bright EUV pixels most likely to be clamped. It's not a fatal flaw, but it is a load-bearing gap in the dataset's quality story, and the paper should quantify it before claiming high fidelity for flare/AR work.\n\nThe arithmetic also doesn't hold up: ~600 MB per file times 467,400 files is roughly 280 TB, not the stated 360 TB for training. Minor but sloppy.\n\nThird, the benchmark labels and baselines are under-reported in the main text. Label generation for DS1–DS6 lives in the supplementary, and the main-text baseline results lack error bars. For a benchmark paper, that's a reproducibility issue. The fact that some tasks build on prior works (e.g., Pandey et al. for flares, Arge et al. for solar wind) is fine; that's normal lineage, not fatal circularity.\n\nAll told, the central claim—a full-resolution, homogenized, cycle-spanning SDO dataset—holds up. The paper deserves a serious referee, not a desk reject. The referee should ask for saturation statistics, error bars, and corrected data volumes before acceptance. I'd bring this to a reading group; it's a useful reference point for heliophysics ML and foundation-model work.","headline":"Genuinely useful full-resolution SDO dataset, but the saturation clamp needs validation before the high-fidelity claims are trusted.","tokens_in":11312,"tokens_out":2404,"would_cite":true,"duration_ms":28208,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SuryaBench offers the full SDO solar archive at native resolution for machine learning","keywords":["SuryaBench","solar physics","space weather","SDO","AIA","HMI","machine learning dataset","benchmark"],"falsifier":"Count the fraction of pixels in the released AIA files that hit the 16,383 DN clamp for each timestamp and channel, then re-run one benchmark (e.g., flare prediction) on a sample after masking or un-clamping those pixels; if the skill scores change materially, the clamping assumption fails.","tokens_in":10359,"feed_emoji":"☀️","tokens_out":4605,"duration_ms":49095,"temperature":0.7,"pith_summary":"SuryaBench is presented as the largest curated, homogenized machine-learning-ready dataset of NASA's Solar Dynamics Observatory, preserving the full 4096×4096 resolution of AIA and HMI images at a steady 12-minute cadence from May 2010 to July 2024. The paper's aim is to remove the preprocessing burden that has kept SDO data out of mainstream ML workflows and to standardize six core heliophysics tasks—flare prediction, active-region segmentation and emergence, coronal field extrapolation, EUV spectra, and solar wind—with baseline results. If this dataset works as claimed, it gives model developers a single, reproducible floor for comparing AI models on space-weather forecasting.","feed_headline":"Full-resolution solar dataset released for space-weather AI","feed_subtitle":"SuryaBench packs 14 years of SDO images at 4096×4096 with six benchmark tasks, from flares to solar wind.","key_machinery":"The homogenization pipeline is the central object: it promotes AIA data from level 1 to level 1.5, removes spacecraft roll, rescales to a common 0.6 arcsec/pixel grid, fixes the solar disk to a radius of 976 arcsec, normalizes exposure times, applies degradation correction clamped at 16,383 DN, re-projects HMI to the same grid, and matches all channels to HMI's 12-minute magnetogram cadence. This pipeline is what makes the simultaneous full-disk, multi-wavelength snapshot at each timestamp possible.","core_discovery":"The central claim is that SDO's roughly 1.5 TB/day of raw data can be converted into a fixed-grid, AI-ready collection through a pipeline of level-1.5 promotion, exposure normalization, degradation compensation with saturation clamping, disk-radius normalization to 976 arcsec, and HMI re-projection to 0.6 arcsec/pixel, all synchronized to 12-minute timestamps. On top of this core collection, the authors construct six labelled benchmark datasets (DS1–DS6) with specified evaluation protocols and baseline results from ResNet, U-Net, and Transformer architectures. The resulting collection totals about 360 TB, divided into training (2010–2018), validation (2019), and test (2020) subsets with roug","pith_inferences":["The 6% data loss through the QUALITY-flag filter likely concentrates around eclipse seasons and instrument anomalies; training on the remaining timestamps may introduce a subtle temporal bias that could affect flare statistics. This is an inference, not a claim in the paper.","Because the dataset fixes the solar disk radius and centralizes the Sun, it should support direct registration across wavelengths without further alignment, which in turn enables multimodal fusion models that combine EUV and magnetogram channels—an extension the paper notes but does not demonstrate.","A testable extension would be to compare model skill for flare prediction using the 512×512 SDOML dataset against SuryaBench's 4096×4096 version on the same held-out flares; the paper does not run this comparison, so the resolution benefit remains a hypothesis."],"forward_implications":["Full-resolution training inputs (4096×4096) become feasible for self-supervised and foundation models, potentially improving detection of small-scale features such as emerging flux and polarity inversion lines.","A uniform 12-minute cadence across a full solar cycle enables consistent temporal models for flare and solar-wind forecasting without per-instrument alignment work.","The six benchmark datasets give researchers reference baselines (ResNet, U-Net, SpatioTemporal Transformer) so new models can be compared on identical splits and metrics.","The public Hugging Face collection and open code make the preprocessing reproducible, lowering the entry barrier for ML researchers without heliophysics domain expertise."],"supporting_citations":[{"why":"Defines the SDO mission and its instrument suite, establishing the source of all data in SuryaBench.","marker":"[2]"},{"why":"Documents the AIA instrument whose EUV/UV channels form the core imaging collection.","marker":"[21]"},{"why":"Documents the HMI instrument whose magnetograms set the 12-minute cadence and provide the magnetic field data.","marker":"[22]"},{"why":"The prior 512×512 SDO ML dataset that SuryaBench explicitly extends and improves by preserving full native resolution.","marker":"[5]"},{"why":"SunPy provides the reprojection and coordinate-transformation functions used in the HMI alignment step.","marker":"[24]"},{"why":"The SpatioTemporal Transformer model used as the baseline for forecasting experiments on the SuryaBench core data.","marker":"[4]"}],"fun_headline_variants":["14 years of SDO imagery, AI-ready for space weather","SuryaBench: 360TB solar dataset for forecasting flares and wind","Six benchmark tasks on 4096×4096 solar images for ML","Sun's 14-year archive now a machine-learning playground","New solar benchmark dataset targets space weather AI"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The central claim relies on the assumption that clamping degradation-corrected AIA values at the 16,383 DN saturation limit does not remove information needed for flare and active-region benchmarks, and that the benchmark labels themselves are correct—label generation is only summarized in the main text and deferred to the Supplementary Information.","fun_headline_variants_meta":{"raw":{"variants":["14 years of SDO imagery, AI-ready for space weather","SuryaBench: 360TB solar dataset for forecasting flares and wind","Six benchmark tasks on 4096×4096 solar images for ML","Sun's 14-year archive now a machine-learning playground","New solar benchmark dataset targets space weather AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000831,"raw_usage":{"total_tokens":3469,"prompt_tokens":752,"completion_tokens":2717,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":496,"completion_tokens_details":{"reasoning_tokens":2645}},"tokens_in":496,"tokens_out":2717,"duration_ms":20889,"temperature":1.0,"reasoning_tokens":2645,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:25:13.941998+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Count the fraction of pixels in the released AIA files that hit the 16,383 DN clamp for each timestamp and channel, then re-run one benchmark (e.g., flare prediction) on a sample after masking or un-clamping those pixels; if the skill scores change materially, the clamping assumption fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The prior 512×512 SDO ML dataset that SuryaBench explicitly extends and improves by preserving full native resolution."},{"cited_title":"AI Foundation Model for Heliophysics: Applications, Design, and Implementation","cited_arxiv_id":"2410.10841","evidence_quote":"The SpatioTemporal Transformer model used as the baseline for forecasting experiments on the SuryaBench core data."}],"review_version":1}