{"id":"bc27ed5a-b05c-4c03-867e-5c110cbbf0e5","arxiv_id":"2412.04418","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"ACE2-SOM, an ML atmospheric emulator with a slab ocean, accurately emulates equilibrium climate sensitivity to CO2 doubling, tripling, and quadrupling, but mishandles non-equilibrium transitions.","lead":"A machine-learning climate emulator was coupled to a slab ocean and trained on simulations at several CO2 levels; it reproduces equilibrium temperature and precipitation changes, including for a CO2 level it never saw. It fails to capture the pace of change in non-equilibrium scenarios, warming too quickly and showing unphysical jumps in the stratosphere.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 3xCO2 'out-of-sample' skill may be interpolation between the 2x and 4x training climates, not a generalizable CO2 sensitivity; Section 3.2.3 supports this concern.","rationale":"I agree with the reader's weakest assumption and make it the single load-bearing concern. The paper is well executed and transparent; the three equilibrium training climates plus the 3x test are a sensible pilot. But the word 'out-of-sample' is doing important work in the abstract, and 3x is the weakest possible out-of-sample test: it lies inside the convex hull of the training CO2 values, and the training set contains the two adjacent equilibrium climates. A model that has learned a smooth mapping from CO2 to equilibrium state would pass this test by interpolation. The radiation multi-call experiments in Section 3.2.3 are the key internal evidence: the model has not learned the physical radiative transfer, since it gives spurious CO2 sensitivity to shortwave and surface-upward longwave fluxes and misses the logarithmic dependence. Those unphysical sensitivities would be expected to corrupt an extrapolation outside the training range, even if they cancel or are masked in the bracketed 3x equilibrium test. Section 3.2.1's regime shifts in stratospheric fields reinforce the interpretation that the model partly maps CO2 to climatological values rather than evolving physically. A leave-one-out training experiment would settle this: hold out 2x, train on 1x, 3x, 4x, and see whether the held-out 2x skill is comparable to the original 3x skill. If it is not, the equilibrium skill is tied to having bracketing training climates, and the central claim should be reframed as interpolation within the training range rather than generalization. This does not change the reader's CONDITIONAL verdict; it sharpens the condition that should be met before the out-of-sample claim is accepted.","tokens_in":24384,"tokens_out":7212,"duration_ms":104128,"concrete_test":"Train ACE2-SOM with the same data, architecture, and checkpoint-selection protocol on 1x, 3x, and 4x equilibrium data only (holding out 2x), and evaluate the 2x-minus-1x equilibrium pattern RMSE for surface temperature and precipitation against C96 SHiELD-SOM, reporting all four seeds. If the held-out 2x skill is substantially worse than the original held-out 3x skill (e.g., pattern RMSE no longer falls near the noise floor), the original 3x result depended on bracketing training CO2 concentrations and the out-of-sample generalization claim is not established; if 2x skill matches, interpolation alone is unlikely to explain the 3x skill.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim 'ACE2-SOM performs well in equilibrium-climate inference with both in-sample and out-of-sample CO2 concentrations' depends on the 3xCO2 equilibrium test demonstrating transferable CO2 sensitivity. The 3x concentration is bracketed by the 2x and 4x training concentrations, and training pairs each CO2 level with a consistent, near-equilibrium atmospheric and surface state. Thus ACE2-SOM can attain low 3x pattern RMSE by effectively interpolating between the neighboring equilibrium climates in input space, without learning a physically generalizable forcing mechanism. The paper's own diagnostics point this way: Section 3.2.3 shows ACE2-SOM predicts spurious CO2 dependence of shortwave fluxes and upward longwave flux at the surface, and misses the logarithmic dependence of longwave fluxes; Section 3.2.1 shows stratospheric variables jumping between values associated with the quantized training CO2 levels. In an equilibrium 3x run, these artifacts can be masked because CO2 and the rest of the state co-vary as in training. The reported skill therefore does not establish the 'out-of-sample' part of the claim, nor that the learned sensitivity would transfer to 1.5x, 5x, other greenhouse gases, or combined forcings. This is not an internal inconsistency or a suggestion of dishonesty; it is a question of what the 3x result actually demonstrates.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript couples the Ai2 Climate Emulator version 2 (ACE2) to a slab ocean model, trains the resulting ACE2-SOM on equilibrium SHiELD-SOM output at 1x, 2x, and 4x CO2, and evaluates it in equilibrium (3x CO2) and non-equilibrium (2%/yr gradual increase, abrupt quadrupling) scenarios. The authors report that ACE2-SOM reproduces equilibrium surface temperature and precipitation change patterns, the vertical structure of warming and stratospheric cooling, and extreme precipitation distributions up to the 99.9999th percentile at 3x CO2, with smaller pattern RMSE than a coarse physics-based baseline in most cases. For non-equilibrium scenarios, ACE2-SOM captures global-mean surface temperature and precipitation trends in the gradual-increase case but shows stratospheric regime shifts, too-rapid adjustment under abrupt 4x CO2, violation of the global moist static energy budget, and spurious CO2 sensitivity of radiative fluxes in radiation multi-call experiments. The paper is transparent about these limitations and frames them as motivation for future work.","tokens_in":24619,"tokens_out":6434,"duration_ms":66472,"significance":"If the equilibrium results transfer beyond the specific 3x test, this is a noteworthy proof-of-concept: a learned 6-hourly atmospheric emulator coupled to a slab ocean can emulate the equilibrium climate response to CO2 changes at about 1500 simulated years per day, with public code, processed data, and model checkpoints provided. The paper also includes multiple ensemble members, a noise-floor estimate, and an explicit out-of-sample test, which are strengths. However, the central 'out-of-sample' claim is not yet fully supported, because 3x CO2 is bracketed by the 2x and 4x training concentrations and the paper's own diagnostics in Section 3.2.3 show that the learned CO2 sensitivity is partly unphysical. The honest reporting of non-equilibrium failures is a strength, but it also sharpens the need to determine whether the equilibrium skill reflects interpolation among training climates rather than a generalizable forcing mechanism.","major_comments":[{"comment":"The radiation multi-call experiments show that ACE2-SOM predicts CO2-dependent shortwave fluxes and surface-upward longwave flux that are physically spurious, and that it misses the logarithmic dependence of longwave fluxes on CO2. Section 3.2.1 additionally shows stratospheric fields jumping among values correlated with the quantized training CO2 levels. Because 3x CO2 lies between the 2x and 4x training concentrations, and because the equilibrium test states co-vary with CO2 as in training, the low 3x pattern RMSE is consistent with interpolation among neighboring training climates and does not by itself establish that the learned sensitivity is physically generalizable. I ask the authors to add an exterior-extrapolation equilibrium test (for example 1.5x or 5x CO2, or retraining on 1x/3x/4x and reporting 2x), and to either temper the abstract's 'out-of-sample' wording or qualify it as 'unseen intermediate concentration' until such a test is provided.","section":"Section 3.2.3 and Figure 10; Section 3.2.1"},{"comment":"The primary equilibrium results in Figures 1-5 are shown for the single best of four random seeds, selected by validation inference. Section 3.3 and Figure 11 reveal non-negligible seed-to-seed spread in the equilibrium climate-change-pattern RMSE for the equilibrium-trained models, but the main text does not quantify this spread for the 3x case or report whether all seeds beat the C24 baseline and remain close to the noise floor. For a central claim about equilibrium skill, the manuscript should report the median and range across seeds for the 3x temperature and precipitation change-pattern RMSE, or present all-seed results in the main text.","section":"Section 2.2.2 and Section 3.3/Figure 11"},{"comment":"The paper labels the gradual-increase and abrupt-4x runs as 'out-of-sample,' but the 3x equilibrium run is out-of-sample only in its CO2 value, whereas the non-equilibrium runs are out-of-sample in the combination of CO2 and atmospheric/surface state. This distinction matters because the abrupt-4x run violates global energy conservation (Figure 9) and shows radiative-flux sensitivities that are not physically consistent (Figure 10d). The authors should state explicitly in the abstract and conclusions that the equilibrium skill is not yet evidence of generalizable transient sensitivity, and should reserve 'out-of-sample' for tests that do not lie inside the convex hull of the training forcings or should clearly define the term if they keep the current usage.","section":"Section 2.2.3 and Section 3.2.2/Figure 9"}],"minor_comments":[{"comment":"The sentence 'they violate global energy conservation and exhibit unphysical sensitivities of and surface and top of atmosphere radiative fluxes' contains a typo; it should read 'unphysical sensitivities of surface and top-of-atmosphere radiative fluxes.'","section":"Abstract"},{"comment":"The caption states 'relative that for C96 SHiELD-SOM'; the phrase should read 'relative to that for C96 SHiELD-SOM.'","section":"Figure 3 caption"},{"comment":"The noise-floor estimate in Equation (4) is described briefly; it would aid reproducibility to state explicitly that the windows are drawn from the same 50-year reference used as the target, and to note any caveat about non-independence between the emulator's internal variability and the reference variability.","section":"Section 3.1.1"},{"comment":"The panels labeled 'air_temperature_0' and 'specific_total_water_0' refer to the top atmospheric layer of ACE2's vertical coordinate; the caption should define the numbering convention for the layer index.","section":"Figure 7 caption"},{"comment":"The footnote describing the increasing-CO2 run says 'CO2 increases at a rate of 2% year-1 thereafter up to about 4x CO2 in 2100'; since the run starts in 2030 and rises for 70 years, the compound increase is approximately 3.9x, so 'about 4x' is fine but the arithmetic could be stated explicitly to avoid confusion.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the scope of JGR: Machine Learning and Computation and the code/data/checkpoint release is exemplary. My main concern is that the headline 'out-of-sample' claim is currently supported by a single bracketed 3x test whose interpretation is weakened by the paper's own diagnostics showing spurious CO2-radiation sensitivities. An exterior extrapolation test or a clearly qualified claim would resolve this. I would be happy to see a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, honest pilot study, and it deserves a serious referee. It is the first autoregressive ML atmospheric emulator coupled to a slab ocean and trained across multiple CO2 levels, with an out-of-sample 3xCO2 equilibrium test and unusually candid diagnostics of non-equilibrium failure. The data, code, and best checkpoint are public, and the citation pattern looks appropriate—the ACE2 self-citations are legitimate follow-on work.\n\nThat said, I partly side with the stress-test. The 3xCO2 equilibrium result does not establish a generalizable CO2 sensitivity. 3x is bracketed by the 2x and 4x training climatologies, and the model is trained on consistent equilibrium state–CO2 pairs, so it can pass that test by interpolation. Section 3.2.3 reinforces the worry: radiation multi-call experiments show spurious CO2 dependence in shortwave fluxes and surface-upward longwave, and the logarithmic longwave scaling is missed. Section 3.2.1 shows stratospheric fields jumping between values correlated with the quantized training CO2 levels. So the equilibrium skill is real, but the “out-of-sample” label is doing more work than it should.\n\nWhat is genuinely good: the equilibrium tests are clean. Five-member ensembles, a noise-floor estimate, fair comparison against a coarser physics-based baseline, stable 10-year rollouts at 3x, and precipitation tails through the 99.9999th percentile. The non-equilibrium experiments are reported with real candor: the energy budget shows the abrupt 4x transition violates conservation, and the authors trace it to learned covariances rather than hiding it. That is the right way to run this kind of study.\n\nSoft spots, in proportion. First, the abstract and parts of Section 3.1.2 run the in-sample 2x and 4x results together with the out-of-sample 3x result under “performs well,” so a casual reader can miss that only 3x is a prediction on unseen CO2. Second, the headline figures come from the best of four random seeds; seed spread appears later in the paper but not on the primary metrics, so robustness is under-quantified. Third, the phrase “learned sensitivity” invites a stronger reading than the evidence supports—transfer to 1.5x, 5x, other greenhouse gases, or combined forcings is untested. These are framing and quantification fixes, not fatal flaws.\n\nWho this is for: people working on ML climate emulators and anyone tracking whether learned atmospheric models can be used for climate sensitivity studies. It is a proof-of-concept with clearly stated limitations, not a finished benchmark.\n\nRecommendation: send it to peer review. Ask the authors to fix the in-sample/out-of-sample framing, quantify seed dependence on the main claims, and state plainly that the 3x result is interpolation between training concentrations with no demonstrated transfer beyond that range.","headline":"Good, honest pilot showing an ML atmospheric emulator can hit equilibrium CO2 response patterns, but the flagship 3xCO2 result is bracketed interpolation between training climates, not proof of transferable sensitivity.","tokens_in":25258,"tokens_out":3435,"would_cite":true,"duration_ms":42120,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ACE2-SOM—a machine-learned atmospheric emulator coupled to a slab ocean—reproduces equilibrium temperature and precipitation change patterns under CO2 doubling, tripling, and quadrupling, including the unseen tripling case.","keywords":["machine-learning climate emulator","slab ocean model","CO2 climate sensitivity","equilibrium climate response","autoregressive emulator","precipitation extremes","energy conservation","out-of-sample generalization"],"falsifier":"Run the trained emulator at a CO2 level between 2x and 4x that was not used in training, for example 2.5x, and compare its equilibrium temperature and precipitation change patterns against a new physics-based reference run at that level; if the pattern errors are comparable to the 3x case, the sensitivity is generalizing, whereas a sharp error spike would indicate interpolation between the quantized training concentrations. A second decisive check is the paper's own radiation multi-call: hold the atmospheric state fixed and vary only CO2; if upward longwave flux at the surface or any shortwave flux changes with CO2, the learned radiative sensitivity is unphysical.","tokens_in":24121,"feed_emoji":"🌡️","tokens_out":6674,"duration_ms":60766,"temperature":0.7,"pith_summary":"ACE2-SOM—a machine-learned atmospheric emulator coupled to a simplified slab ocean—is trained on equilibrium climates at 1x, 2x, and 4x CO2, then tested at 3x CO2, a concentration it never saw. The paper's central claim is that the emulator accurately reproduces the time-mean spatial patterns of surface temperature and precipitation change with CO2 doubling, tripling, or quadrupling, as well as the vertical profile of warming and changes in extreme precipitation up to the 99.9999th percentile. It also claims that non-equilibrium inference is more fragile: gradual CO2 increase yields correct surface and lower-atmosphere trends but unphysical jumps in the stratosphere, and an abrupt CO2 quadrupling warms the atmosphere too quickly, violating global energy conservation. The result matters because it is the first demonstration that an autoregressive ML emulator can be trained to respond to a substantial greenhouse-gas forcing rather than only to the historical climate.","feed_headline":"ML emulator reproduces unseen CO2 climate response","feed_subtitle":"Trained at 1x, 2x, and 4x CO2, it matches temperature and precipitation change at 3x.","key_machinery":"The central object is the differentiable coupling of the ACE2 neural atmospheric model to a slab ocean: each 6-hour step the ML model predicts the surface fluxes that make up $F_{\\mathrm{net}}$, and the mixed-layer temperature is updated by $\\rho_o C_o h\\,\\partial T_s/\\partial t = F_{\\mathrm{net}} + Q$, where $Q$ is a prescribed climatological ocean-heat-convergence (Q-flux). This lets the sea surface temperature respond to CO2-driven changes in surface energy fluxes while keeping the whole system trainable by backpropagation. The argument for CO2 sensitivity rides on the model learning to use the prescribed CO2 input to produce radiation and flux responses that match the equilibrium physics-model states it was trained on; the paper's multi-call radiation experiments probe whether those sensitivities are physical, and show that some are not.","core_discovery":"On its own terms, the paper establishes that ACE2-SOM, trained on equilibrium output from a 100-km-resolution physics-based model coupled to a slab ocean, can generalize its learned dynamics to a CO2 concentration between the training values. In the out-of-sample 3xCO2 equilibrium, time-mean surface temperature and precipitation biases are small, and the model reproduces the well-known pattern of greenhouse warming—tropical upper-tropospheric maximum, stratospheric cooling, land warming more than ocean—closely. Precipitation extremes up to the 99.9999th percentile of daily rates also match the target model, and the emulator's climate-change-pattern errors are smaller than or comparable to those of a 400-km physics-based baseline at a fraction of the cost. The same skill does not carry over to transitions: with gradually increasing CO2, stratospheric temperature and water content shift abruptly between values correlated with the quantized training CO2 levels, and with abrupt quadrupling, ML-predicted fields relax to the 4x regime faster than the slab ocean, producing a moist static energy budget imbalance and spurious sensitivity of surface and top-of-atmosphere radiative fluxes to CO2.","pith_inferences":["A natural extension the paper does not pursue is to train on continuously varying CO2 trajectories instead of quantized equilibrium levels; the stratospheric regime shifts suggest this would break the spurious association between CO2 bins and slowly varying fields.","The spurious radiative sensitivities point toward a hybrid fix: letting a lightweight, differentiable radiation scheme handle CO2-dependent longwave transfer while the ML model learns the rest of the dynamics.","The same equilibrium-training protocol could be extended to other forcing agents—methane, aerosols, or land-use—as long as enough equilibrium reference climates are generated to span the forcing range.","One testable consequence of the paper's analysis is that a model trained on only 1x and 4x should fail at 3x if the sensitivity is interpolative; a per-CO2 multi-call diagnostic (net flux versus CO2) could serve as a cheap screening test for physicality."],"forward_implications":["If the central claim is correct, equilibrium climate sensitivity experiments—which currently require years of physics-based simulation—could be approximated by a trained ML emulator at roughly 100 times faster simulation speed.","The skill extends to precipitation extremes up to the 99.9999th percentile, suggesting the emulator captures not just mean shifts but the distributional response of the hydrological cycle to warming.","The out-of-sample 3x result implies the model has learned some generalizable relationship between CO2 forcing and the equilibrium state, rather than simply memorizing the three training climates.","The energy-conservation violation in abrupt-CO2 runs shows that autoregressive ML emulators need explicit physical constraints before they can be trusted for transient climate change scenarios."],"supporting_citations":[{"why":"Describes the ACE2 architecture, grid, variables, normalization, and loss function that ACE2-SOM inherits.","marker":"Watt-Meyer, Henn, et al. (2024)"},{"why":"Introduces the Spherical Fourier Neural Operator architecture used as the ML backbone.","marker":"Bonev et al. (2023)"},{"why":"Documents the SHiELD physics-based model that supplies the reference training and testing data.","marker":"Harris et al. (2020)"},{"why":"Provides the slab ocean equation and Q-flux formulation used in the coupled model.","marker":"Kiehl et al. (2006)"},{"why":"Supplies the mixed-layer-depth climatology used to prescribe the slab ocean layer.","marker":"de Boyer Montégut et al. (2004)"},{"why":"Defines the CMIP-style CO2 experiment protocols (abrupt quadrupling, gradual increase) that structure the test cases.","marker":"Eyring et al. (2016)"},{"why":"Provides the physical scaling expectation for increases in precipitation extremes used to contextualize the model's tail behavior.","marker":"O'Gorman & Schneider (2009)"}],"fun_headline_variants":["ML emulator reproduces unseen CO2 climate response","AI climate model learns 3x CO2 equilibrium from 1x, 2x, 4x","Machine learning emulator matches CO2 warming patterns, not transitions","ACE2-SOM predicts climate-forcing patterns, misses rapid shifts","Emulator generalizes to 3x CO2, but fails non-equilibrium scenarios"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The equilibrium skill depends on ACE2-SOM's CO2 response being a physically generalizable forcing mechanism rather than an interpolation among the three training CO2 levels; the paper's own evidence of spurious shortwave and upward-longwave flux sensitivities, and of stratospheric values that jump between quantized training levels, shows this assumption can be violated.","fun_headline_variants_meta":{"raw":{"variants":["ML emulator reproduces unseen CO2 climate response","AI climate model learns 3x CO2 equilibrium from 1x, 2x, 4x","Machine learning emulator matches CO2 warming patterns, not transitions","ACE2-SOM predicts climate-forcing patterns, misses rapid shifts","Emulator generalizes to 3x CO2, but fails non-equilibrium scenarios"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00102,"raw_usage":{"total_tokens":4395,"prompt_tokens":1128,"completion_tokens":3267,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":744,"completion_tokens_details":{"reasoning_tokens":3167}},"tokens_in":744,"tokens_out":3267,"duration_ms":23733,"temperature":1.0,"reasoning_tokens":3167,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:25:18.878999+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained emulator at a CO2 level between 2x and 4x that was not used in training, for example 2.5x, and compare its equilibrium temperature and precipitation change patterns against a new physics-based reference run at that level; if the pattern errors are comparable to the 3x case, the sensitivity is generalizing, whereas a sharp error spike would indicate interpolation between the quantized training concentrations. A second decisive check is the paper's own radiation multi-call: hold the atmospheric state fixed and vary only CO2; if upward longwave flux at the surface or any shortwave flux changes with CO2, the learned radiative sensitivity is unphysical.","supporting_citations":[{"cited_title":", Shields, C A","cited_arxiv_id":null,"evidence_quote":"Provides the slab ocean equation and Q-flux formulation used in the coupled model."}],"review_version":1}