{"id":"cd2991df-7560-4644-b615-cbbb01ce0431","arxiv_id":"2608.13494","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An ocean emulator trained on CESM2's control climate generalizes to midHolocene orbital forcing in the upper ocean but under-shoots amplitude and misses slow interior dynamics.","lead":"An AI ocean emulator trained on one simulated climate was tested on a different, bygone climate from the same model, and it reproduced much of the upper-ocean response to the changed sunlight and winds, though too weakly and not in the deep ocean. The paper proposes using such paleoclimate experiments as a controlled test bed before trusting emulators on future warming scenarios.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Equilibration of the CESM2 midHolocene reference is asserted, not shown; if the mH−piC difference contains transient drift or large internal variability, the reported skill correlations and amplitude ratios do not cleanly measure the emulator's forced response.","rationale":"The reader identified the equilibration/internal-variability assumption as the weakest link, and I agree that it is the most load-bearing concern. The emulator's forced-response skill is scored entirely against the CESM2 midHolocene−piControl difference. If that difference includes transient adjustment or substantial internal variability, then every reported correlation and amplitude ratio in Sections 3.1 and 3.3 is not a clean measure of forced-response skill; it is a measure of similarity to a single noisy realization. The paper explicitly acknowledges the assumption ('we expect the system to have equilibrated') but provides no quantitative check, such as comparing successive chunks of the midHolocene run to verify stationarity of the upper-ocean response. I considered the other concerns raised by the reader. Checkpoint selection and the sign-flip assumption for F_mH are real but secondary: the checkpoint issue is partially mitigated by the reported epoch ranges and by the paper treating epoch sensitivity as a finding in itself, and the sign-flip assumption affects only the F_mH cross-check, not the primary piControl-to-midHolocene generalization. The equilibration concern, by contrast, underpins all ground-truth comparisons in the paper. The proposed test would settle it: if the upper-1000 m difference is stable across successive 50-year windows, the concern is resolved and the verdict should remain as the reader set it; if it drifts, the central claim needs to be reformulated or re-quantified. A CONDITIONAL verdict is therefore appropriate until that check is reported.","tokens_in":38980,"tokens_out":7818,"duration_ms":335705,"concrete_test":"Using the available CESM2 midHolocene and piControl monthly data, compute the zonally averaged upper-1000 m potential temperature difference (mH−piC) over successive non-overlapping 50-year windows (e.g., years 350–400, 450–500, 550–600, and 600–650) and compare the inter-window mean differences to the internal-variability spread estimated from ten 25-year chunks. If the inter-window shift is comparable to or larger than the chunk spread, the reference is not equilibrated; then recompute the Figures 6–7 metrics against a longer or de-drifted reference to see whether the reported correlations and amplitude ratios survive.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central evaluation compares emulator responses R(F_pi) and R(F_mH) to the CESM2 midHolocene−piControl difference, treating that difference as the equilibrated forced response. Section 3.3 states 'we expect the system to have equilibrated on the century timescales we investigate' but provides no direct test. The two CESM2 experiments are single realizations, so their difference contains both the forced signal and internal variability; any slow adjustment still in progress in the upper 1000 m will be attributed to the forced response. The significance stippling in Figures 6–7 is computed from ten 25-year chunks, but the reported correlations and amplitude ratios are scored against the full noisy difference, which may be dominated by non-forced variability in regions where the forced signal is weak (e.g., the Atlantic basin-wide correlation of 0.03). If the reference drifts across the available century-scale windows, then the central claim that the emulator captures the spatial structure of the large-scale forced response is not quantitatively established; it is only established that the emulator resembles one noisy realization of the CESM2 difference.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper evaluates whether an autoregressive, full-depth ocean emulator (a ConvNEXT/UNet architecture with roughly 85 million parameters, trained on CESM2 piControl monthly data) can generalize to midHolocene orbital forcings as an out-of-sample but in-distribution test. The authors compare emulator responses to the CESM2 midHolocene minus piControl difference for upper-ocean potential temperature, seasonal-cycle changes, and variability changes, and they compare against two baselines (a local linear regression operator and a forcing-only network). They report that the emulator reproduces the large-scale spatial structure of the upper-ocean response and changes in seasonality and variability, while underestimating amplitude, and that it fails to capture slow interior evolution. They also show that the emulator's total response is approximately the linear superposition of its responses to individual forcing components, and that test RMSE on the training climate does not predict forced-response skill across epochs and seeds.","tokens_in":39177,"tokens_out":4535,"duration_ms":55914,"significance":"If the central claims hold, the midHolocene setup provides a valuable, controlled, ground-truthed benchmark for diagnosing forced-response failures in AI ocean emulators before they are applied to out-of-distribution climates. The experimental design has real strengths: the emulator is trained only on piControl and never fit to midHolocene data; the evaluation includes two carefully constructed baselines, multiple seeds, full epoch sweeps, and significance stippling on the CESM2 reference; and the code, preprocessing scripts, and model weights are archived on Zenodo. The component-forcing decomposition and the epoch-sensitivity analysis are also useful contributions. The main quantitative claims rest on two assumptions that the paper acknowledges but does not fully verify: that the CESM2 midHolocene reference is equilibrated, and that the sign-flip comparison between the piControl-trained and midHolocene-trained emulators is valid. Because those assumptions affect the headline correlations and amplitude ratios, the paper needs additional evidence or a more conservative presentation before the claims are fully established.","major_comments":[{"comment":"The reference response used as ground truth is the CESM2 midHolocene minus piControl difference, but the manuscript only states 'we expect the system to have equilibrated on the century timescales we investigate' and provides no direct test. Since each CESM2 experiment is a single realization, the difference contains internal variability and any slow drift still present in the upper 1000 m. This is load-bearing for every reported correlation and amplitude ratio in Figures 6-7 and Tables S1-S4: if the reference is not equilibrated, the Atlantic basin-wide correlation of 0.03 and the amplitude ratios do not cleanly measure the emulator's forced response. Please add a quantitative equilibration check, for example comparing early versus late 50-year windows of the overlapping piControl and midHolocene periods, or computing the trend in the upper-1000 m temperature difference over the evaluation window and showing it is small relative to the forced signal. Reporting confidence intervals on the correlations and amplitude ratios computed from the ten 25-year chunks would also help separate forced signal from internal variability.","section":"Section 3.3"},{"comment":"The headline metrics are reported from 'a single representative checkpoint per emulator' that is selected after inspecting response skill ('We use an early checkpoint for the piControl emulators, taken before the late-training skill degradation'). Given that Section 4.3 demonstrates substantial epoch-to-epoch and seed-to-seed variability in response skill, and that 'no checkpoint performs best across all regions', this post hoc selection risks inflating the reported numbers and makes the quantitative claims difficult to reproduce without the same selection procedure. Please either pre-specify a checkpoint-selection rule that does not use the out-of-sample response (for example, lowest validation MSE before the degradation epoch, chosen blind to response skill), or report the distribution of response metrics across all epochs and both seeds as the primary summary, with the single-checkpoint values shown only for illustration.","section":"Section 3, Section 4.3"},{"comment":"The comparison of the midHolocene-trained emulator R(F_mH) against the CESM2 reference uses the sign-flip convention R(F_pi) ↔ -R(F_mH), with the statement that the authors 'cannot verify the exact reversibility or linearity of the applied forcing'. The cross-emulator correlations between R(F_pi) and -R(F_mH) (0.76 in the Pacific, 0.88 in the tropics) are encouraging, but they do not establish that the CESM2 true response is antisymmetric under reversing the forcing perturbation. Since the F_mH correlation values in Figures 6-7 and throughout Section 3.3 are computed after flipping the sign of the true response, an unverified asymmetry directly affects those numbers. Please either treat the F_pi results as the primary quantitative evaluation and present the F_mH results as an internal consistency check, or add a test (e.g., comparing the spatial patterns of the two emulator responses against the flipped true response in a way that does not presuppose antisymmetry, or using single-forcing CESM2 experiments if available) to justify the sign-flipped comparison.","section":"Section 3.3"}],"minor_comments":[{"comment":"Equation (1) defines δ as a 'fraction' of cell overlap, but the expression with min/max and the Heaviside function yields an overlap thickness in meters; please clarify the units and terminology in the text following Eq. (3).","section":"Section 2.1, Eqs. (1)-(3)"},{"comment":"The paragraph after Figure 8 contains a confusing contrast: it first says response magnitudes are 'all below 0.05 °C compared to the true maximum 0.16 °C' for 'other component responses', then says R(F_pi^All; τ_mH,I_mH), R(F_pi^All; hfds_mH), and R(F_pi^hfds; hfds_mH) 'reproduce more accurate values of 0.16 °C, 0.12 °C, and 0.13 °C'. Please rewrite to specify exactly which responses have weak maxima and which have accurate maxima, since the current wording appears to contradict itself.","section":"Section 4.1"},{"comment":"The reported ensemble spread of at most 6×10^-6 °C across five ensemble members seems implausibly small for a chaotic system; please state whether this is the spread under identical climatological forcing with different initial conditions, and clarify why the emulator is so insensitive to initial conditions in this diagnostic.","section":"Section 2.5.3"},{"comment":"Several correlations are reported without specifying the metric precisely (e.g., spatial pattern correlation versus temporal correlation) or the effective degrees of freedom. For example, the 'correlation of 0.98' in Figure 2B and the time-series correlations above 0.98 for Niño3.4 should state whether these are Pearson correlations over space, time, or both, and whether they are computed after detrending or removal of the seasonal cycle.","section":"Figures 2-4 and associated text"},{"comment":"The statement that the emulators 'capture the lagged correlation between indices, though the spread between subsets of the data remains large' would benefit from a quantitative measure of that spread, such as a confidence interval or a null expectation, since the large spread weakens the strength of the claim as presented.","section":"Section 3.2"},{"comment":"The dynamic weighting in Eq. (5) is described as using the reciprocal root-mean-square error per channel and rollout step, with clipping at 500:1; a brief justification for the clipping value and the smoothing period N_smooth=100 would help readers assess how sensitive the results are to these hyperparameters.","section":"Section 2.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong candidate for this journal if the authors can provide direct evidence for equilibration of the CESM2 reference and a principled checkpoint-selection rule. The two issues are fixable within the manuscript's scope, so I would not reject. The supplementary material is unusually thorough, and the archival of code and weights is commendable. My main editorial concern is that the current presentation of the F_mH sign-flip comparison could be misread as a validated symmetry of the underlying CESM2 response; I recommend the authors either demote it to a consistency check or add explicit caveats in the main text."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First: this is the first paper to train an autoregressive ocean emulator on control climate and test it on midHolocene forcings, and the test is genuinely out-of-sample since midHolocene never touches training. That makes it a real contribution. The experimental design is careful: held-out midHolocene, two baselines that isolate direct forcing imprint, multiple seeds, full epoch sweeps, and code and weights on Zenodo. The main results—upper-ocean pattern of the forced response is captured with amplitude underestimated, seasonal cycle and variability changes reproduced, slow interior evolution missed, and MSE gains do not track response skill—are supported by the evidence. The component additivity result (r≈0.99 reconstruction) is a nice diagnostic even if it only probes the emulator's linearity, which the authors state.\n\nThe soft spots are real but mostly acknowledged in the text. The biggest is the reference itself: the paper treats the mH−piC CESM2 difference as the equilibrated forced response, with the sentence 'we expect the system to have equilibrated on the century timescales we investigate' doing the work. Nothing tests that. Single realizations mean internal variability and any slow drift are baked into the target, so skill numbers in weak-signal regions (Atlantic r=0.03) aren't a clean measure of forced-response skill. I'd want a drift check or a statement of what the chunk-to-chunk spread does to the headline correlations. Second, headline metrics come from a checkpoint chosen after looking at response skill; the authors disclose this and report ranges, which mitigates it, but the abstract's 'reproducing' is stronger than the checkpoint-robust evidence warrants. Third, the F_mH comparison flips sign with an explicit, honest caveat that reversibility is unverified; I'd like that called a limitation in the abstract or conclusions rather than buried in methods.\n\nThe equilibration issue is the one that could move the central claim: if the mH−piC difference is contaminated by internal variability, the paper has shown the emulator resembles one noisy realization. But in the tropical Pacific and Southern Ocean, where the forced signal is large and the paper's strongest claims live, the variance argument is less concerning, and the convergence table (S1) shows metrics stabilize over rollout length. So I think the central argument holds for the upper ocean; the high latitudes and deep interior are honestly reported as failures.\n\nWho is this for: anyone building or using ocean emulators for climate perturbation experiments. It deserves a serious referee: the testbed is reusable, the baselines are sensible, and the epoch-sensitivity result alone is worth publishing. I'd send it to review with a request to address the equilibration and checkpoint-selection issues.","headline":"A careful first paleoclimate testbed for ocean emulators; the main claims hold up, but the equilibration of the midHolocene reference is asserted rather than proven.","tokens_in":39701,"tokens_out":2399,"would_cite":true,"duration_ms":24986,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An ocean emulator trained only on the control climate can reproduce the large-scale upper-ocean response to a 6,000-year-old orbital forcing, but the response is too weak and the slow deep-ocean changes are missed.","keywords":["ocean climate emulator","midHolocene","out-of-sample generalization","forced response","CESM2","orbital forcing","linear superposition","autoregressive model"],"falsifier":"Recompute the skill scores against a longer, fully equilibrated segment of the CESM2 midHolocene run, or against a multi-member midHolocene ensemble, so the reference is purged of spin-up drift; if the tropical Pacific correlations of 0.64-0.83 and the two-thirds amplitude ratios do not survive that cleaner ground truth, the generalization claim is tied to the chosen reference window.","tokens_in":38720,"feed_emoji":"🌊","tokens_out":8989,"duration_ms":79638,"temperature":0.7,"pith_summary":"The paper's aim is to give long-term ocean climate emulators a test that is out-of-sample but still in-distribution, and it argues that the midHolocene experiment supplies exactly that. The central result is that an autoregressive, full-depth ocean emulator trained only on CESM2 piControl reproduces the spatial structure of the upper-ocean potential-temperature response to midHolocene orbital forcing, along with changes in the seasonal cycle and in variability patterns, while recovering only about two-thirds of the true amplitude. Forcing-only baselines recover much of the near-surface pattern but miss the seasonal and variability changes, which shows that some learned internal dynamics are doing real work. The same emulator fails on slow, internally driven evolution of the ocean interior, and its total response is well approximated by a linear superposition of the responses to each boundary forcing. If this holds, paleoclimate experiments offer a cheap, ground-truthed way to catch forced-response failures before emulators are trusted for warming scenarios.","feed_headline":"Ocean emulator generalizes to 6,000-year-old orbital forcing","feed_subtitle":"Upper-ocean patterns and seasonality transfer; slow deep-ocean changes do not.","key_machinery":"The load-bearing object is a ConvNEXT-UNet autoregressive emulator that steps the three-dimensional ocean state (potential temperature $\\theta_O$ and salinity $S$) forward one month from a single prior state plus surface heat flux, the two wind-stress components, and explicitly computed insolation. The argument is carried by paired ensembles of rollouts: the forced response is defined as the difference between midHolocene-forced and piControl-forced rollouts sharing the same initial conditions, and the total response is then decomposed into single-forcing component responses. The linearity probe, comparing the full response with the sum of component responses, is what establishes the approximate additivity, and a dynamic channel weighting during training is what lets slowly evolving deep-ocean signals contribute to the loss at all.","core_discovery":"On the paper's own terms, the discovery is that a full-depth autoregressive ocean emulator trained exclusively on the CESM2 preindustrial control run generalizes to the midHolocene: when driven by midHolocene boundary forcings it captures the pattern of the upper-1000 m temperature response (tropical Pacific correlations 0.64-0.83), the phase and amplitude changes of the seasonal cycle, and changes in the spatial structure of ocean variability, while underestimating the response amplitude (amplitude ratios about 0.6-0.7). The authors further find that noiseless checkpoints of the emulator that look equivalent on test RMSE can differ substantially in response skill, and that late-training checkpoints lose basin-wide skill in the Atlantic, so standard validation metrics do not certify dynamics. Finally, perturbing individual forcing components produces responses that superimpose nearly exactly onto the full response (correlations near 0.99 in every basin), indicating that the emulator's forcing pathways interact only weakly through the internal state; the authors are careful to note this is a property of the emulator, not demonstrated for the real ocean.","pith_inferences":["One natural extension, not run here, is to apply the same protocol to the Last Interglacial (~127,000 years ago): the paper's account predicts high tropical pattern correlations but a further drop in amplitude ratio as orbital anomalies grow, which would test whether the damping scales predictably.","Because single-forcing midHolocene experiments do not exist, the paper cannot say whether the real CESM2 response is additive; generating such runs would settle whether the near-perfect linear superposition is a genuine physical property or a learned simplification of a control climate.","The diagnosed failure to accumulate slow interior heat suggests a specific mechanism worth testing: if the ocean emulator is coupled to an atmospheric emulator, some of that deep response should return; retraining with explicit circulation state variables would directly probe that hypothesis."],"forward_implications":["If the result is correct, the midHolocene becomes a standard intermediate test: an emulator must pass a paleoclimate out-of-sample test before its responses to warming scenarios are taken at face value.","Upper-ocean perturbation experiments with emulators become more defensible, since the emulator beats forcing-only baselines in both response amplitude and depth.","Decomposing the response by forcing component offers a cheap attribution tool for projected ocean changes, at least within the emulator's linear regime.","The late-training collapse in response skill implies that checkpoint selection for climate emulators should include dynamical response tests, not only held-out RMSE.","Slow deep-ocean and North Atlantic changes remain outside the emulator's reliability envelope, so claims about interior heat accumulation or overturning changes from such emulators should be treated as unsupported."],"supporting_citations":[{"why":"Provides the CESM2 model whose piControl and midHolocene experiments supply all training and evaluation data.","marker":"Danabasoglu et al. (2020)"},{"why":"Supplies the CESM2 midHolocene simulation used as the out-of-sample target climate.","marker":"Danabasoglu (2019)"},{"why":"Supplies the CESM2 piControl simulation that defines the training distribution and the base climate.","marker":"Danabasoglu et al. (2019b)"},{"why":"Supplies the ConvNEXT-UNet architecture and regridding choices that the emulator adapts.","marker":"Dheeshjith et al. (2025)"},{"why":"Supplies the orbital parameters used to compute the insolation input channel for both experiments.","marker":"Otto-Bliesner et al. (2017)"},{"why":"Provides the CESM2 midHolocene analysis, especially tropical Pacific seasonality, that the emulator's response is measured against.","marker":"Otto-Bliesner et al. (2020)"},{"why":"Supplies the conceptual case that paleoclimate experiments are out-of-sample tests for climate models, which the paper applies to emulators.","marker":"Burls & Sagoo (2022)"}],"fun_headline_variants":["Paleoclimate test: ocean emulator transfers, but amplitude off","Ocean emulator passes paleo pattern test, fails deep ocean","Emulator generalizes to 6,000-year-old climate, with caveats","MidHolocene forcing: emulator captures pattern, underestimates response","Deep-ocean changes elude climate emulator in paleo test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation uses the CESM2 midHolocene-minus-piControl difference as the equilibrated forced response of the ocean, so if that difference still carries slow spin-up drift or internal variability, the reported correlations and amplitude ratios do not cleanly measure the emulator's forced-response skill.","fun_headline_variants_meta":{"raw":{"variants":["Paleoclimate test: ocean emulator transfers, but amplitude off","Ocean emulator passes paleo pattern test, fails deep ocean","Emulator generalizes to 6,000-year-old climate, with caveats","MidHolocene forcing: emulator captures pattern, underestimates response","Deep-ocean changes elude climate emulator in paleo test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000267,"raw_usage":{"total_tokens":1674,"prompt_tokens":1061,"completion_tokens":613,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":677,"completion_tokens_details":{"reasoning_tokens":518}},"tokens_in":677,"tokens_out":613,"duration_ms":6459,"temperature":1.0,"reasoning_tokens":518,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:45:08.543532+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the skill scores against a longer, fully equilibrated segment of the CESM2 midHolocene run, or against a multi-member midHolocene ensemble, so the reference is purged of spin-up drift; if the tropical Pacific correlations of 0.64-0.83 and the two-thirds amplitude ratios do not survive that cleaner ground truth, the generalization claim is tied to the chosen reference window.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the CESM2 model whose piControl and midHolocene experiments supply all training and evaluation data."},{"cited_title":"2025 , publisher=","cited_arxiv_id":null,"evidence_quote":"Supplies the ConvNEXT-UNet architecture and regridding choices that the emulator adapts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the orbital parameters used to compute the insolation input channel for both experiments."},{"cited_title":"A comparison of the","cited_arxiv_id":null,"evidence_quote":"Provides the CESM2 midHolocene analysis, especially tropical Pacific seasonality, that the emulator's response is measured against."}],"review_version":1}