{"id":"0423274a-c470-4149-a7ea-c30027366689","arxiv_id":"2504.17578","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"LCC-CMAES learns when to use random, min-variance, or max-variance decomposition in cooperative coevolution, improving CMA-ES on large-scale benchmarks and transferring to unseen problems.","lead":"This paper trains a reinforcement learning agent to choose how a large optimization problem is split into smaller pieces during optimization, replacing hand-designed grouping rules. On benchmarks and robot-control tasks, the learned selector often uses fewer evaluations and finds better solutions than existing methods.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LCC-CMAES's 'excels in optimization results' claim is budget-dependent: at 3E6 FEs the paper's own Fig. 4 shows non-CC baselines overtaking it, so the headline advantage may be an artifact of the 2E5 budget.","rationale":"The reader's weakest assumption is that the 58 state features are sufficient to determine the optimal decomposition strategy. That is a real concern, but the paper's ablation (removing GO/SD/AH degrades performance) and transfer results (BNS, Mujoco) provide at least partial evidence for feature informativeness. The budget dependence is more direct: the authors themselves report in Sec. 5.2.2 that at 3E6 FEs the non-CC baselines overtake LCC-CMAES on most problems. Since all headline tables use 2E5 FEs, the 'excels in optimization results' contribution is asserted for exactly the budget where it holds, without showing that it persists. The reader already listed 'unusually low budget' as a reason for conditional acceptance, so my concern does not change the verdict; it identifies the specific experiment that would settle the scope of the claim. I keep the verdict at CONDITIONAL/UNCHANGED pending that test.","tokens_in":22511,"tokens_out":10342,"duration_ms":102719,"concrete_test":"Re-run the full CEC2013 LSGO experiment with MaxFEs=3E6 (the benchmark standard), using the same trained policy and hyperparameters, and report final mean/median values plus Wilcoxon win/tie/loss against each baseline. If LCC-CMAES is not significantly better than the best non-CC baseline (e.g., LM-CMA, Sep-CMAES) on a majority of the 15 problems, then the headline claim must be explicitly scoped to short FE budgets and the abstract/contribution wording revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 2, Table 4, and Table 5 all use MaxFEs=2E5 (CEC/BNS) or 1,000 (neuroevolution), with D=1,000 or higher. The CEC 2013 LSGO standard is 3E6 FEs. Under this larger budget, Section 5.2.2 and Figure 4 state that 'algorithms without CC exhibit more prominent optimization performance' on most problems and that LCC-CMAES's advantage 'gradually diminishes and may even be surpassed.' No final statistics or significance tests are reported at 3E6, so the reader cannot see whether the 'excels' claim survives. The no-extra-decomposition-FEs and low wall-clock properties are budget-independent, but the optimization-quality component of the central claim is not. Because the paper advertises LCC-CMAES as a general drop-in replacement that 'excels... in optimization results' (Sec. 1, Conclusion), the absence of a quantified large-budget comparison is a load-bearing gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes LCC-CMAES, a cooperative coevolution (CC) framework in which a PPO-trained neural network selects, at each epoch, one of three CMA-ES decomposition strategies (MiVD, RD, MaVD) based on 58 hand-crafted features comprising global optimization statistics, per-subgroup statistics, and action history. The policy is trained on a subset of the CEC 2013 LSGO benchmark with a budget of 2E5 function evaluations and evaluated on the remaining CEC 2013 functions, the BNS suite, and four Mujoco neuroevolution tasks, with additional ablations on state features and reward designs. The central claims are that LCC-CMAES improves optimization effectiveness and resource consumption relative to CC and non-CC baselines and that the learned strategy selector transfers to unseen problems without additional decomposition FEs.","tokens_in":22811,"tokens_out":7579,"duration_ms":68086,"significance":"The core idea—replacing expert-designed decomposition selection with a learned, dynamic strategy scheduler—is timely and potentially useful for low-budget, high-dimensional regimes. Strengths include 25-run comparisons with Wilcoxon tests, a dedicated ablation of state features and reward formulations, and zero-shot transfer experiments to BNS and Mujoco tasks; these make the main positive results concrete. The paper is also honest about the extended-budget behavior in Figure 4, which is important because that evidence substantially qualifies the headline 'excels in optimization results' claim. If the authors quantify the 3E6-FE comparison, address the privileged f* in the reward, and soften the transferability claims to match the reported tables, the contribution would be a solid addition to the MetaBBO and CC literature.","major_comments":[{"comment":"The headline claim that LCC-CMAES \"excels ... in optimization results\" is supported only at 2E5 FEs, which is 15x smaller than the standard 3E6 budget for CEC 2013 LSGO. In §5.2.2 the authors state that at 3E6 FEs the advantage \"gradually diminishes and may even be surpassed\" and that non-CC algorithms show \"more prominent optimization performance\" on most problems, but no final means, standard deviations, or Wilcoxon statistics are given for the 3E6 budget. Because the conclusion is unqualified and the low budget is motivated by resource consumption rather than by the optimization-quality claim, the authors should report full significance-tested comparisons at 3E6 and either show the advantage persists or explicitly restrict the claim to limited-budget settings.","section":"§5.2.2, Fig. 4, Tables 2, 4, and 5"},{"comment":"The reward in Eq. (4) divides by f*0 - f*, where f* is the true optimal fitness. This quantity is unknown in real-world applications and in the BNS/Mujoco transfer experiments, and the paper does not state that f* is used only in training or analyze what happens when f* is inaccurate. Since the learned policy is trained exclusively against this reward, the dependence on a privileged benchmark quantity should be acknowledged and the robustness of the policy to mis-specified f* should be tested, for instance by training with a biased estimate of f* and re-running the transfer evaluations.","section":"§4.2.3, Eq. (4)"},{"comment":"The transferability conclusion is stronger than the evidence in Table 4: on the BNS suite, LCC-CMAES is significantly worse than CMA-ES, Sep-CMAES, LM-MA-ES, and LM-CMA on several problems (e.g., problems 1, 9, and 12), and the summary row shows LCC-CMAES winning only five of ten pairwise comparisons against each of several non-CC baselines. The conclusion that LCC-CMAES holds \"a distinct advantage\" on BNS is therefore not supported; the authors should report the loss counts as prominently as the wins and phrase the transferability claim in terms of \"competitive transfer\" or \"advantages on CC-type baselines\" rather than blanket superiority.","section":"§5.3, Table 4, and Conclusion"}],"minor_comments":[{"comment":"There is a typo in the bullet list: \"LCC-CAMES\" should be \"LCC-CMAES\".","section":"§5.2.1"},{"comment":"The entries for problems 3 and 6 contain stray \"+\" and \"\\\" symbols that make the table hard to read; these formatting artifacts should be removed.","section":"Table 2"},{"comment":"The sentence \"we zero-shot the pre-trained LCC-CMAES agent in Section 5.2\" should refer to the training setup in Section 5.1.2, since Section 5.2 is the main comparison rather than the training description.","section":"§5.4"},{"comment":"The BNS experiment omits MetaES and L-BFGS from Table 4 without explanation, although both are listed as comparison algorithms in Section 5.1.1; a sentence explaining the omission would be helpful.","section":"§5.3"},{"comment":"The formula for reward1 should include parentheses to be mathematically unambiguous: r_t = (f*0 - f*t) / f*0.","section":"§5.5.2"},{"comment":"The state features require the full global covariance matrix C_t at each step, which costs O(D^2) memory and computation; since the paper targets scalable optimization, a brief comment on the practical dimension limit and the cost of computing Corrcoef(C_t) for D > 10^4 would be useful.","section":"Table 1"},{"comment":"The sentence \"F13 and F14 in CEC2013LSGO is 905\" should read \"are 905\" for grammatical agreement.","section":"§5.1.2"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope and the core mechanism is worth publishing after a revision that aligns the claims with the evidence. The most important points to convey to the authors are the need for a quantified 3E6-FE comparison with significance tests, an explicit discussion of the f* assumption in the reward, and a more calibrated transferability claim based on Table 4. I would also encourage the authors to release code and seeds, as the experimental setup is otherwise not fully reproducible from the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: this is the first paper I've seen that treats CC decomposition-strategy selection as an RL problem and trains a PPO policy to switch among MiVD, RD, and MaVD during optimization. That is a real step beyond RL-DAS (operator selection for DE) and CC-CMAES (hand-designed selection rules). The state design is thoughtful, and the transfer results to BNS and Mujoco are genuine predictions, not fitted constants.\n\nWhat the paper does well: the MDP formulation is clean; the 58 hand-crafted features are motivated and the ablation shows that removing any of the three feature groups hurts performance. The resource-consumption advantage is also real: LCC-CMAES needs no extra FEs for decomposition and has low wall-clock time. That part of the contribution is budget-independent.\n\nThe main soft spot is the budget. The headline tables use MaxFEs=2E5, while the standard budget for CEC 2013 LSGO is 3E6. The paper itself says that at 3E6, algorithms without CC show more prominent optimization performance and that LCC-CMAES's advantage gradually diminishes or is surpassed. Figure 4 appears to confirm this, but no end-of-run numbers or significance tests are given. So the claim that LCC-CMAES \"excels in optimization results\" is only supported at the small budget. That is not fatal if the paper is positioned as a low-FE method, but the abstract and conclusion do not make that qualification.\n\nSecond, the reward in Eq. 4 uses the true optimum f*, so training requires privileged information that is unavailable in real applications. This is acceptable for benchmark-driven training, but it weakens the practical-deployment story. Third, no code or data are released, and the tables have enough stray formatting errors to slow a careful reader. Minor issues, but they matter for reproducibility.\n\nOverall, the central idea is solid and the paper is honest enough to report the large-budget behavior even though it complicates the headline. It deserves a serious referee. My recommendation: send it to peer review, ask for code and data, and require the authors to either restrict the optimization-quality claims to low-budget settings or provide full statistics at 3E6.","headline":"A genuinely new RL-based decomposition scheduler for CC-CMAES with real zero-shot transfer results, but the headline optimization-quality claim is budget-dependent and needs stronger large-budget evidence and code before I'd fully trust it.","tokens_in":23315,"tokens_out":2457,"would_cite":true,"duration_ms":25989,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A PPO-trained policy can pick the right variable decomposition for CMA-ES.","keywords":["cooperative coevolution","CMA-ES","large-scale global optimization","reinforcement learning","decomposition strategy selection","meta-black-box optimization","transferability","PPO"],"falsifier":"Evaluate the trained LCC-CMAES policy on a suite built from entirely new base functions with known, time-varying interaction graphs; if a fixed strategy such as always-MiVD or a cheap random switch matches its final fitness, then the state features are not capturing the scheduling-relevant information.","tokens_in":22346,"feed_emoji":"🧩","tokens_out":7969,"duration_ms":74240,"temperature":0.7,"pith_summary":"The paper sets out to show that the most expertise-heavy step in cooperative coevolution for large-scale optimization—deciding how to split variables into subgroups—can be automated by learning. It trains a Proximal Policy Optimization agent to choose, at every epoch, among three existing decomposition strategies (random, min-variance, and max-variance splitting) for CMA-ES, using a 58-dimensional state built from statistics of the covariance, subgroup populations, and action history. On the CEC 2013 LSGO benchmark the resulting LCC-CMAES significantly beats the expert-designed CC-CMAES in 11 of 15 problems, uses no extra function evaluations for decomposition, and keeps an edge on overlapping and non-separable problems. The authors also report zero-shot transfer to the BNS suite and to Mujoco neuroevolution tasks. If these results hold, cooperative coevolution becomes usable without expert decomposition knowledge and with a smaller evaluation budget.","feed_headline":"Trained agent picks the right variable split for CMA-ES","feed_subtitle":"LCC-CMAES removes expert decomposition design, spends no extra evaluations, and transfers to unseen problems.","key_machinery":"The load-bearing object is the learned policy $\\pi_\\theta$: a three-layer MLP with structure $58 \\times 64 \\times 64 \\times 3$ that maps the decision vector $\\mathrm{DV} = s^{GO} \\oplus s^{SD} \\oplus s^{AH} \\in \\mathbb{R}^{58}$ to a softmax distribution over the three decomposition strategies. MiVD, RD, and MaVD are covariance-based partitions of the variables into $m=10$ subgroups, ranging from low-diversity (exploitative) to high-diversity (exploratory) grouping; after each subgroup run, the subgroup covariance and mean are merged back into the global CMA-ES distribution. The reward $r_t = (f^*_{t-1} - f^*_t)/(f^*_0 - f^*)$ normalizes per-step fitness improvement and is the training signal for PPO. This mechanism is what lets the framework drop both the expert selection table of CC-CMAES and the extra function-evaluation cost of interaction-detection decomposition methods.","core_discovery":"On the paper's own terms, the discovery is that a learned strategy scheduler can replace the expert-designed selection rule inside CC-CMAES and improve both final fitness and resource use. The policy observes global optimization statistics (12 features), per-subgroup statistics (40 features), and action-history statistics (6 features), and outputs a probability distribution over MiVD, RD, and MaVD; PPO trains it with a reward equal to the normalized best-fitness decrease per step. The evidence is a 25-run comparison on CEC 2013 LSGO in which LCC-CMAES outperforms CC-CMAES on 11 of 15 problems, outperforms most non-CC large-scale optimizers, and wins clearly on the overlapping functions F12-F14 and the fully non-separable F15. Additional experiments show the zero-shot policy retains advantages on the BNS benchmark and on four Mujoco neuroevolution tasks at dimensions up to 2312.","pith_inferences":["A direct extension is to enlarge the action set beyond the three covariance-based splits, for example the number of subgroups or the offspring size; the MDP formalism in the paper would carry over unchanged.","Part of the reported transfer may come from shared base functions: the paper's training subset contains Elliptic and Rastrigin variants, so hold-out tests built from entirely new base-function families would separate landscape familiarity from learned grouping ability.","The current reward charges no cost for switching strategies or for time; if the strategy pool ever includes decomposition methods that consume function evaluations, the reward would need a budget term, which is a modification the paper does not address."],"forward_implications":["Cooperative coevolution loses its main knowledge barrier: the same trained policy can decide decomposition on a new large-scale problem without an expert tuning the grouping rule.","Because LCC-CMAES spends no function evaluations on detecting variable interactions, its whole budget goes to optimization, giving it a concrete resource advantage over CSG, ERDG, MDG, and FII.","The scheduling policy transfers zero-shot: trained on a subset of CEC 2013 LSGO, it still improves over CC-CMAES on the BNS suite and on Mujoco tasks with up to 2312 dimensions.","Dynamic scheduling is most valuable exactly where static CC methods fail—overlapping and fully non-separable problems—so the approach targets the hardest cases rather than only separable ones.","The agent's choices are informed by recent strategy effectiveness through the action-history features, meaning the same policy can behave exploitatively at one stage and exploratively at another within a single run."],"supporting_citations":[{"why":"Supplies CC-CMAES and its three decomposition strategies (MiVD, RD, MaVD), the expert-designed baseline to be replaced.","marker":"[33]"},{"why":"Provides the PPO objective used to train the actor and critic networks.","marker":"[59]"},{"why":"Defines CMA-ES, the underlying optimizer whose mean, covariance, and step size produce the state features and subgroup updates.","marker":"[19]"},{"why":"Defines the CEC 2013 LSGO benchmark used for training and in-distribution evaluation.","marker":"[26]"},{"why":"Defines the BNS benchmark of additively and non-additively separable problems used for the transferability study.","marker":"[5]"},{"why":"Precedent for using deep reinforcement learning in operator selection for evolutionary algorithms, motivating the learned scheduling formulation.","marker":"[16]"},{"why":"Exploratory landscape analysis is cited as the basis for designing the optimization-status state features.","marker":"[43]"},{"why":"The Delta method inspires the subgroup mean-difference features (features 23-32) that capture variable interactions.","marker":"[48]"}],"fun_headline_variants":["Neural network schedules CMA-ES splits, beating expert rules","RL-trained decomposition policy improves CMA-ES on large-scale","Learning-based CC-CMAES: no expert design, no extra evaluations","CMA-ES learns to split variables, transfers to unseen tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 58 hand-picked statistics of the current search state carry enough information to decide which of the three decompositions is best at every step, even on problems whose variable interactions were never seen in training.","fun_headline_variants_meta":{"raw":{"variants":["Neural network schedules CMA-ES splits, beating expert rules","RL-trained decomposition policy improves CMA-ES on large-scale","Learning-based CC-CMAES: no expert design, no extra evaluations","CMA-ES learns to split variables, transfers to unseen tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0003,"raw_usage":{"total_tokens":1703,"prompt_tokens":887,"completion_tokens":816,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":503,"completion_tokens_details":{"reasoning_tokens":744}},"tokens_in":503,"tokens_out":816,"duration_ms":7873,"temperature":1.0,"reasoning_tokens":744,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:36:33.567876+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Evaluate the trained LCC-CMAES policy on a suite built from entirely new base functions with known, time-varying interaction graphs; if a fixed strategy such as always-MiVD or a cheap random switch matches its final fitness, then the state features are not capturing the scheduling-relevant information.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies CC-CMAES and its three decomposition strategies (MiVD, RD, MaVD), the expert-designed baseline to be replaced."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the CEC 2013 LSGO benchmark used for training and in-distribution evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Precedent for using deep reinforcement learning in operator selection for evolutionary algorithms, motivating the learned scheduling formulation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Exploratory landscape analysis is cited as the basis for designing the optimization-status state features."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Delta method inspires the subgroup mean-difference features (features 23-32) that capture variable interactions."}],"review_version":1}