{"id":"20373a10-e6f2-4eae-9ce8-b4d4552e73d4","arxiv_id":"2607.25940","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Four turbulence regimes emerge from K-means clustering of CNN embeddings of Canary Islands SCIDAR heatmaps, with distinct vertical structure, seasonality, and wind roses.","lead":"The authors ran Canary Islands turbulence profiles through a pretrained image-recognition network, then clustered the resulting embeddings into four atmospheric regimes. This shows that standard deep-learning tools can digest a ~300-night SCIDAR archive, but the seasonal signal is partly baked in and no code was released.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Wind-rose 'independent validation' is confounded by season: day-of-year is two of three CNN inputs, clusters are seasonal, and seasonal wind climatology can produce the same signatures without turbulence information.","rationale":"The reader already identified the seasonal-channel circularity as a weak assumption and assigned CONDITIONAL. My concern sharpens this: the seasonal confound also undermines the paper’s one supposedly independent validation, the wind-rose analysis. Because wind direction at these sites is seasonally climatological, clustering by the injected sine/cosine channels can produce distinct wind-roses even if the turbulence channel contributes nothing. This directly attacks the strongest claim, so it is load-bearing. However, I do not think it demands rejection: the analysis is explicitly preliminary, the null/ablation test is straightforward, and the paper’s honest statement that additional validation is required is appropriate. The existing CONDITIONAL verdict already covers this; no change to the verdict is needed. I also note secondary concerns (heatmap overlap/independence not reported, no code/data release, no quantitative wind-rose significance test), but the seasonal confound of the wind-rose evidence is the single most decisive issue.","tokens_in":6881,"tokens_out":4471,"duration_ms":49612,"concrete_test":"Construct a season-only null model: take the same 238 high-quality tensors, replace the first channel (normalized turbulence intensity) with its global mean (or i.i.d. noise with the same marginal distribution), and keep the two seasonal channels unchanged. Run the identical frozen-EfficientNet feature extraction, k-means with k=4, and compute the same cross-cluster wind-rose separation metric used in Figures 6/7 (e.g., mean pairwise chi-square distance or a KS statistic). If the null model reproduces the real pipeline’s wind-rose separation, the Section 5 independent-validation argument fails; if separation is clearly smaller, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 5) is that the turbulence classes are physically meaningful, and the paper’s only independent evidence is that wind-rose signatures differ between clusters even though wind was not used in clustering. This independence is weaker than it appears. Two of the three input channels are sine/cosine of day of year (Section 3.5), and Figure 5 shows the resulting clusters are strongly seasonal. Surface wind direction at the Canary Islands observatories also has a strong seasonal climatology. Therefore, wind-rose differences between clusters can arise from season as a common cause: the CNN may simply be reading the seasonal channels, clustering by calendar time, and the wind-rose separation then follows from seasonal wind patterns without the turbulence channel contributing any physical information. No ablation removing the seasonal channels is provided, and no null model that clusters using season alone is tested. The paper itself notes in Section 4 that 'additional validation is required,' but the specific missing validation is this control. Without it, the wind-rose analysis does not establish that the embeddings preserve turbulence physics beyond seasonal labels.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a machine-learning reanalysis of the Generalized SCIDAR C_n^2(h) database at the two Canary Islands observatories (ORM and OT). The authors build an ETL pipeline, aggregate individual turbulence profiles into time-altitude heatmaps, and convert each accepted heatmap into a three-channel tensor containing normalized turbulence intensity plus sine and cosine encodings of the day of the year. These tensors are passed through a frozen pre-trained EfficientNet CNN to produce embeddings, which are clustered with K-means. A four-cluster solution is adopted and visualized in terms of mean turbulence maps, seasonal occurrence, and surface wind roses. The central claim is that the resulting clusters capture physically meaningful atmospheric circulation regimes, based mainly on the observation that wind-rose patterns differ across clusters even though wind information was not used in the clustering.","tokens_in":7154,"tokens_out":4037,"duration_ms":44491,"significance":"If validated, this approach would provide a novel data-driven way to extract recurrent turbulence regimes from large astronomical site-characterization archives, with potential utility for adaptive-optics scheduling and site monitoring. The paper builds an original end-to-end processing framework and applies modern representation-learning tools to a valuable public-domain dataset. The central claim, however, is not yet established. The analysis currently lacks controls for a strong seasonal confound, quantitative cluster validation, and an explicit treatment of sample non-independence. The methodology is promising, but the physical interpretation goes beyond what the presented evidence supports.","major_comments":[{"comment":"The analysis uses 238 high-quality heatmaps, but the paper never reports the temporal stride or overlap between heatmaps, nor the quantitative thresholds for the selection criteria listed in §3.5 ('temporal coverage, minimum profile density, maximum interpolated interval'). If the heatmaps are generated from overlapping windows, the 238 samples are not independent. This would inflate the apparent significance of the cluster structure and of the wind-rose differences. Please report the full sampling schedule of the heatmaps, the effective number of independent samples, and repeat the clustering and wind-rose analysis with non-overlapping windows as a robustness check.","section":"§3.5 and §5, Figs. 5-7"},{"comment":"The wind-rose differences in Figures 6 and 7 are presented visually, without any statistical test or statement of the number of wind measurements per cluster. Even if the seasonal confound were removed, the visual differences could be within sampling noise, especially if some clusters contain only a small number of heatmaps. A formal comparison of circular distributions, such as a permutation test on cluster labels or a two-sample Kuiper test, is needed to quantify the strength of the wind-rose separation.","section":"§3.6 and §4"}],"minor_comments":[{"comment":"The text introducing Figure 5 says that seasonal patterns 'emerge naturally from the underlying turbulence structure'. This phrasing is misleading, because two of the three input channels explicitly encode the day of year. Consider rewording to acknowledge the role of the seasonal channels in the input representation.","section":"§4"},{"comment":"The description of the wind stations notes that the JKT anemometer is on a rooftop and less than 5 m from the telescope dome on the prevailing leeward side, and the GONG station is also on a roof. These local obstructions may affect the measured wind directions. This limitation should be stated in the interpretation of the wind-rose results.","section":"Figures 6-7"},{"comment":"The 'normalized arbitrary units' used for the turbulence intensity colour scale in Figure 4 are not defined. Please specify how the normalization was performed so that quantitative comparisons are meaningful.","section":"§3.4"},{"comment":"No data or code availability statement is included. Since the paper introduces a complete processing framework, sharing the ETL and heatmap-generation code would significantly improve reproducibility and would also allow readers to test alternative clustering choices.","section":"General"},{"comment":"The reference list correctly identifies previous statistical analyses of the OCAN database, but the reader would benefit from a more explicit statement of what those analyses found that the present clustering approach is meant to complement or improve.","section":"§1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable preliminary methodological study, but the central physical-interpretation claim is not yet supported. The seasonal-channel issue is the crux: the wind-rose 'independent validation' is not independent as long as season is a common cause. The revision should add the missing ablations and null models, or substantially soften the conclusions. I do not think this requires rejection, because the requested controls are feasible with the existing dataset and pipeline."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is an honest, preliminary ML reanalysis of a substantial SCIDAR archive. The genuinely new thing is the application of frozen CNN embeddings plus K-means to the OCAN database, yielding four turbulence regimes with distinct vertical structure and seasonal patterns. That is a legitimate new application, not a new measurement or physical mechanism.\n\nWhat the paper does well: the ETL pipeline is described in enough detail to be reproduced, the heatmap generation is sensible, and the use of a frozen EfficientNet as a feature extractor is standard but appropriate. The cluster mean maps look coherent, and the authors are upfront that this is preliminary and that additional validation is needed. The wind-rose analysis is a good instinct for external validation, and they correctly note wind was not used in clustering.\n\nThe soft spots are real. Two of the three input channels are sine and cosine of day of year, so the seasonal separation in Figure 5 is partly a direct product of the input encoding. That means the wind-rose differences, while independent of the turbulence channel in the narrow sense, can be explained by season as a common cause: the CNN could be clustering by calendar time, and the seasonal wind climatology then produces distinct wind roses without the turbulence channel contributing anything. No ablation removing the seasonal channels, and no null model clustering on season alone, is provided. The paper says \"additional validation is required,\" but it does not name this specific control.\n\nAlso missing: no baseline against simpler features like raw profile statistics, no error bars or stability analysis, no cross-validation of the clustering, and no code or data release. The 238 high-quality heatmaps is a small sample, and the acceptance criteria are not fully specified. These are addressable.\n\nThe central regime claim is not destroyed. The mean maps do show different vertical structures, and it is plausible that recurrent turbulence states exist. But the evidence does not yet rule out the seasonal-label explanation, so CONDITIONAL is the right verdict.\n\nWho benefits: observational astronomers interested in site characterization and ML practitioners looking for a template for unsupervised analysis of atmospheric data. The paper deserves a serious referee because the underlying archive is valuable and the approach is transferable, but the revision must include ablations and null models.","headline":"A promising but under-validated CNN reanalysis of Cn2 profiles; the four-cluster result is plausible, but the only independent check is confounded by season.","tokens_in":7625,"tokens_out":1642,"would_cite":false,"duration_ms":16581,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper tries to establish that four turbulence regimes found by unsupervised clustering of CNN embeddings of Canary Islands SCIDAR heatmaps are physically meaningful atmospheric states, because each regime has a distinct surface wind si","keywords":["optical turbulence","Generalized SCIDAR","site characterization","deep learning","convolutional neural networks","feature extraction","clustering","adaptive optics"],"falsifier":"Run the same pipeline on the same 238 heatmaps after randomly permuting the turbulence values within each heatmap (or blanking the first channel, leaving only the sine/cosine channels); if K-means still returns four clusters with the same seasonal distributions and the same wind-rose separations, the clustering is driven by the seasonal channels, not by turbulence structure. Conversely, if clusters collapse or lose wind signatures when the two seasonal channels are removed, that would confirm the turbulence structure is necessary.","tokens_in":6781,"feed_emoji":"🌬️","tokens_out":5020,"duration_ms":43362,"temperature":0.7,"pith_summary":"By reprocessing the ~300-night Canary Islands Generalized SCIDAR database into 60-minute time–altitude turbulence heatmaps, feeding them through a frozen EfficientNetB1 backbone, and K-means clustering the resulting embeddings, this paper aims to show that the data contain recurrent, physically distinct turbulence regimes. The four clusters differ in the altitude and strength of turbulent layers, show sharp seasonal transitions, and — critically — have clearly different wind-rose patterns at both ORM and OT, even though wind information was excluded from the embeddings and clustering. The paper argues that these wind signatures independently validate the physical relevance of the classes, since only genuine atmospheric states would co-vary with circulation. A sympathetic reader would care because this offers an unsupervised, data-driven way to decompose a large turbulence archive into interpretable states, with potential uses in adaptive optics scheduling and site climatology.","feed_headline":"CNN finds four turbulence regimes at Canary observatories","feed_subtitle":"Unsupervised clustering of telescope turbulence maps yields four regimes with distinct wind signatures.","key_machinery":"The central object is the turbulence heatmap: a time–altitude matrix of log C_n^2 values (60-minute window, 180-second bins) made into a three-channel tensor by appending sine and cosine of day-of-year as two extra channels. The tensor is passed through the frozen convolutional backbone of EfficientNetB1 (pre-trained on a large image corpus), which yields a compact embedding per heatmap. K-means clustering on these embeddings partitions the 238 high-quality heatmaps into four groups. The key simplifying mechanism is transfer learning: without any fine-tuning on atmospheric data, the pre-trained filters, as padding-preserving 'feature extractors,' are claimed to retain enough of the heatmaps'","core_discovery":"The central discovery claimed is that unsupervised clustering of deep embeddings extracted from turbulence heatmaps recovers four atmospheric turbulence regimes with distinct vertical-layer structures and distinct associated wind-direction and speed distributions. At ORM, clusters show different prevailing wind directions; at OT, the regimes also separate in wind space despite both sites sharing similar large-scale turbulence behaviour. Because wind data were not included in any stage of the CNN feature extraction or K-means clustering, the cluster-specific wind roses are presented as independent evidence that the embeddings preserve physically meaningful information about the atmospheric st","pith_inferences":["The sine/cosine day-of-year channels are two of the three input channels, so the strong seasonal separation of clusters in Fig. 5 is at least partly a product of the input representation; an ablation that removes these channels would show how much genuinely turbulence-structural information the embeddings retain.","The 238 heatmaps are treated as independent samples, but if consecutive heatmaps overlap temporally, cluster assignments are autocorrelated and the wind-rose statistics may overstate separation; reporting the temporal stride would clarify effective sample size.","A direct null-model test — e.g., shuffling the turbulence pixels within each heatmap while keeping the seasonal channels intact — would determine whether the clusters and their wind signatures depend on vertical turbulence structure at all, or only on season.","The paper's own statement that embedding quality depends on matching heatmap resolution to the CNN's receptive scales suggests a tunable hyperparameter; this could be used to search for an optimal heatmap aspect ratio for other sites."],"forward_implications":["Each identified regime provides a compact label for a full 60-minute turbulence state, enabling night-by-night or hour-by-hour classification from C_n^2 profiles alone.","Regime membership can be correlated with integrated AO parameters (seeing, isoplanatic angle, coherence time), potentially allowing regime-aware scheduling of adaptive-optics instruments.","The same pipeline can be applied to other long-term SCIDAR or DIMM archives to check whether analogous turbulence regimes exist at other sites.","The link between clusters and wind roses suggests that regime transitions track synoptic circulation changes, offering a data-driven connection from turbulence profiles to meteorology.","Because clusters emerge without fixed calendar bins, the analysis defines 'natural' seasons of turbulence rather than imposing predefined ones."],"fun_headline_variants":["CNN plus clustering reveals four turbulence regimes at Canary telescopes","Four distinct turbulence regimes found in Canary sky data via CNN","Unsupervised CNN uncovers four wind-linked turbulence patterns","Deep learning splits Canary turbulence into four regimes"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole interpretation rests on the assumption that the CNN embeddings encode the physical vertical-turbulence structure rather than the injected day-of-year channels, because two of the three input channels are pure seasonality and no ablation or null model is provided to rule out that the clusters merely recapitulate the seasonal input.","fun_headline_variants_meta":{"raw":{"variants":["CNN plus clustering reveals four turbulence regimes at Canary telescopes","Four distinct turbulence regimes found in Canary sky data via CNN","Unsupervised CNN uncovers four wind-linked turbulence patterns","Deep learning splits Canary turbulence into four regimes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000305,"raw_usage":{"total_tokens":1604,"prompt_tokens":775,"completion_tokens":829,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":519,"completion_tokens_details":{"reasoning_tokens":772}},"tokens_in":519,"tokens_out":829,"duration_ms":6024,"temperature":1.0,"reasoning_tokens":772,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T01:02:17.896200+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same pipeline on the same 238 heatmaps after randomly permuting the turbulence values within each heatmap (or blanking the first channel, leaving only the sine/cosine channels); if K-means still returns four clusters with the same seasonal distributions and the same wind-rose separations, the clustering is driven by the seasonal channels, not by turbulence structure. Conversely, if clusters collapse or lose wind signatures when the two seasonal channels are removed, that would confirm the turbulence structure is necessary.","supporting_citations":[],"review_version":1}