{"id":"14785e83-f3af-4a18-870c-0e7a1e5b8f62","arxiv_id":"2604.09661","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"A featurize-and-cluster workflow with optimal feature selection and a new intermingledness metric recovers and ranks alternate states in three climate datasets.","lead":"The paper gives a practical workflow that finds alternate states in high-dimensional climate-style simulations and ranks which observables separate those states. It also defines intermingledness, a per-variable score of how mixed the states and their basins are, aimed at monitoring and early-warning design.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged ad-hoc Q; the central claim holds as a methods contribution.","rationale":"The paper is a methods contribution whose strongest claim is algorithmic recovery of distinguishable alternate states plus identification of informative observables via an optimized featurize-and-group loop and a new sparse-data intermingledness metric. That claim is supported by (i) recovery of expert-known structure on three independent climate ensembles, (ii) open source code, and (iii) explicit sensitivity checks on the empirical quality Q. The only genuine fragility is the lack of theoretical foundation for Q, which the authors themselves state in Appendix B; the reader already identified it correctly. Because the method is computational and the recovery experiments succeed under the reported settings, that fragility does not rise to a load-bearing failure of the central claim. No stronger internal inconsistency or untested assumption was found. Verdict therefore remains ACCEPT; no adjustment is warranted.","tokens_in":21754,"tokens_out":472,"duration_ms":4056,"concrete_test":"Re-run the Veros analysis (§3.1) with the Iterative option of IA-DBSCAN disabled and with the three weight combinations at the corners of the stable plateau in Fig. 7; confirm that the same five attractors and the same top-ranked feature triple (AMOC strength + NA surface/subsurface temperature) are recovered. If any of those runs yields a different number of groups or a completely different optimal feature set, the claimed optimality of the observables would weaken.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption (ad-hoc Q in Eq. 1 / Appendix B) is the real soft spot, but it is not load-bearing enough to overturn the central claim. The paper recovers expert-known attractors on three climate ensembles, ships open code, and treats Q as an empirical, weight-tunable objective whose sensitivity is shown in Fig. 7. Intermingledness is presented with explicit caveats on absolute values and sampling dependence (Appendix C). No hidden contradiction, circular derivation, or failure of the recovery experiments is present. The workflow's claim is therefore that of a usable, reproducible computational method, not a theorem of uniqueness of the recovered states.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes a computational workflow for identifying and analysing alternate states in finite, high-dimensional simulation ensembles. Building on a featurize-and-group approach with an iterative advanced DBSCAN (IA-DBSCAN), it optimizes over feature subspaces using an empirically defined grouping quality Q (Eq. 1) to decide whether clearly distinguishable alternate states exist and which observables best separate them. It further introduces intermingledness, a per-group, per-dimension ratio of pairwise-averaged intra- to inter-group distances, to quantify similarity of attractor features and of basins of attraction. The method is demonstrated on three climate ensembles (Veros AMOC multistability, MAOOAM midlatitude flow with continuation in emissivity, and ExoPlaSim exoplanet habitability classes), recovers expert-known groupings where available, and is released with open-source code.","tokens_in":22026,"tokens_out":1104,"duration_ms":26675,"significance":"If the workflow performs as claimed, it fills a genuine methodological gap: objective, transferable analysis of multistability and alternate operating regimes in sparse, high-dimensional climate and complex-system data where traditional dense basin methods and simple clustering fail. Recovering expert attractor counts on Veros and MAOOAM, reproducing a known destabilization under continuation, and providing open code and an intermingledness implementation in DynamicalSystems.jl are concrete strengths. The indicator and feature-ranking outputs could usefully guide monitoring, early-warning design, and model intercomparison (e.g. TipMIP, habitability MIPs). The contribution is primarily methodological rather than a uniqueness theorem; its value rests on practical reliability and reproducibility, which the manuscript largely supports.","major_comments":[{"comment":"Eq. (1) and Appendix B: Grouping quality Q is acknowledged as empirically derived with no theoretical foundation, and silhouette mean S is shown to be nearly independent of cluster count A and feature count F. Fig. 7 further shows that weight choices can change the recovered number of attractors (e.g. 4 vs 5 when iteration is off). Because the central claim is that the workflow algorithmically decides which alternate states exist and which observables are optimal, the manuscript needs clearer default weight recommendations, a short protocol for sensitivity checks practitioners must run, and explicit discussion of failure modes when Q mis-ranks feature sets. Without that, the claimed objectivity of the recovered states remains partly practitioner-dependent.","section":null},{"comment":"Sections 3.1–3.3 and 4.2: Validation is performed only on ensembles already known (or pre-labeled) to contain alternate states. The workflow’s claim to decide whether alternate states are “clearly distinguishable, if any” therefore lacks a demonstrated null or false-positive case (e.g. a monostable ensemble or pure noise features). A brief controlled test or synthetic monostable example would substantially strengthen confidence that the pipeline does not invent structure when none exists, which is load-bearing for applications where the answer is unknown a priori.","section":null}],"minor_comments":[{"comment":"Notation for Q in Eq. (1) uses A, F, O without fully consistent symbols in the surrounding text (sometimes n_outliers, etc.); unify symbols and define n_min explicitly in the main text, not only in Appendix B.","section":null},{"comment":"Figure 1 step labels and spine highlighting are helpful but dense; a short legend or caption sentence listing what each panel’s axes represent would improve readability for non-specialists.","section":null},{"comment":"Section 2.2: The list of possible features is good, but the paper only uses late-time mean/std. A one-sentence justification why more complex features (entropy, spectral) were unnecessary for these three datasets would help transferability claims.","section":null},{"comment":"Appendix C: Boundary intermingledness is interesting but under-used in the main applications. Either apply it once to Veros/MAOOAM basins or move the full comparison to a shorter note so the main text stays focused.","section":null},{"comment":"Typos and style: “posses” → “possess” (abstract); “intermingledess” in §6; occasional missing articles; “MOAOOAM” in Fig. 4 caption should be MAOOAM.","section":null},{"comment":"Data availability: MAOOAM trajectories “available upon request” weakens reproducibility relative to Veros and ExoPlaSim; if possible, deposit a reduced feature table or a subset of trajectories.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a solid methods contribution appropriate for a physics/climate-methods venue. The ad-hoc Q is the soft spot but is handled honestly and is not fatal for a computational workflow paper. I would not block on it if the authors add practitioner guidance and a null-case check. Fit with physics.ao-ph / nonlinear dynamics methods is good; novelty relative to DLW23 is incremental but the high-dimensional climate packaging, intermingledness, and three real applications justify publication after minor revision."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a methods paper that does what it claims on the data it has. The authors take their earlier featurize-and-group / ADBSCAN work, add iterative re-clustering, a combinatorial search over feature subspaces, and a new sparse-data intermingledness measure, then show that the pipeline recovers the expert-known attractor counts on Veros AMOC and MAOOAM mid-latitude flow, finds physically sensible best features (North Atlantic surface/subsurface temperature with AMOC strength; particular streamfunction modes), and reproduces a known destabilization under emissivity continuation. On the exoplanet habitability set they skip clustering and just map intermingledness onto pre-labeled classes, which is still informative. Open code and NetCDF-friendly tooling are shipped; that matters for TipMIP-style work and for anyone sitting on expensive ensembles who currently decides attractors by eye.\n\nWhat is new is the full sparse-data workflow plus intermingledness (and a boundary variant) that can be computed per diagnostic and per group without dense basin covering. The ranking of which observables best and least separate states and basins is the practical payoff, and the three applications make that concrete rather than abstract.\n\nThe soft spot is exactly the one the reader flagged: grouping quality Q is empirical, weights are free parameters, and silhouette mean alone does not prefer more clusters. Appendix B shows robustness over a range of weights and the recovery experiments still match expert labels, so Q is not load-bearing for the central claim, but anyone applying this will need to re-run the sensitivity. Absolute intermingledness values are also hard to interpret without sampling context; the paper is honest about that. Simulation count and transient handling remain practical limits, again stated.\n\nMath and citations look solid; self-cites to DLW23/DW22 are the base algorithms, not circular. This is for people who run or re-analyze high-dimensional multistable ensembles (climate, power grids, networks) and want an objective, reproducible first pass. I would send it to referees; it is a usable advance with clear limitations, not a theorem of uniqueness.","headline":"Usable open workflow that recovers expert multistability on three climate ensembles and ranks which observables actually separate the states; soft spot is the ad-hoc Q objective, not the recovery results.","tokens_in":22598,"tokens_out":539,"would_cite":true,"duration_ms":7120,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A featurize-and-group workflow decides which alternate states exist in high-dimensional simulations and which observables best separate them.","keywords":["multistability","alternate states","high-dimensional data","DBSCAN clustering","intermingledness","climate tipping","basins of attraction","feature selection"],"falsifier":"On a held-out ensemble whose attractors are already known by exhaustive long control runs, re-run the full optimization with the same candidate features and check whether the recovered partition and the top-ranked diagnostics match the known classification.","tokens_in":22644,"feed_emoji":"🌊","tokens_out":618,"duration_ms":5958,"temperature":0.7,"pith_summary":"Complex systems such as climate models, power grids, and biological networks often admit multiple alternative long-term states, yet high-dimensional finite ensembles are hard to classify without subjective judgment. This paper supplies an algorithmic workflow that extracts statistical features from diagnostic timeseries, clusters them with an iterative advanced density-based method, and optimizes over feature subsets so that the number of clearly distinguishable states and the observables that best separate them are chosen together. Once the states are found, a new indicator called intermingledness measures how mixed the states (or their basins of attraction) are along each diagnostic, revealing which variables are useful for monitoring or early warning and which are not. The method recovers expert-known attractors on three climate ensembles—Atlantic overturning, mid-latitude ocean–atmosphere flow, and exoplanet habitability—and is released as open-source code that reads ordinary NetCDF files.","feed_headline":"Algorithm finds alternate climate states and the best observables","feed_subtitle":"Open workflow recovers known attractors and ranks which diagnostics separate them","key_machinery":"Intermingledness: for each group and each one-dimensional diagnostic, the ratio of mean intra-group pairwise distance to mean inter-group pairwise distance; values near 1 mean the groups cannot be told apart along that diagnostic.","core_discovery":"Finite high-dimensional simulation ensembles can be partitioned into objectively distinguishable alternate states by projecting trajectories onto diagnostic features, clustering those features with iterative advanced DBSCAN, and ranking candidate feature subspaces by a grouping-quality score that rewards more states, fewer dimensions, and fewer outliers; the same partition yields an intermingledness matrix that quantifies, per diagnostic, how similar or mixed the states and their basins are.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Workflow partitions high-dim ensembles into alternate states","Ranks diagnostics that best separate multistable attractors","Intermingledness scores mixed states and basins per diagnostic","Algorithm recovers alternate climate states from finite data","Finds distinguishable states and top separating observables"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The ad-hoc grouping-quality score that multiplies silhouette mean by powers of the number of states, features, and outliers is assumed to select the scientifically correct feature subspaces; if that score mis-ranks, the claimed optimal observables change.","fun_headline_variants_meta":{"raw":{"variants":["Workflow partitions high-dim ensembles into alternate states","Ranks diagnostics that best separate multistable attractors","Intermingledness scores mixed states and basins per diagnostic","Algorithm recovers alternate climate states from finite data","Finds distinguishable states and top separating observables"]},"model":"grok-4.5","effort":"low","cost_usd":0.003566,"raw_usage":{"total_tokens":1209,"prompt_tokens":826,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":35660000,"prompt_tokens_details":{"text_tokens":826,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":309,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":826,"tokens_out":74,"duration_ms":3670,"temperature":1.0,"reasoning_tokens":309,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T16:05:44.588755+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a held-out ensemble whose attractors are already known by exhaustive long control runs, re-run the full optimization with the same candidate features and check whether the recovered partition and the top-ranked diagnostics match the known classification.","supporting_citations":[],"review_version":2}