{"id":"b7f04ac8-ecf1-4f5e-b459-2ebdc84fe81f","arxiv_id":"2506.11332","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A multi-objective genetic algorithm with adaptive space group diversity control improves polymorphic and regular crystal structure prediction over two prior GA/ML baselines.","lead":"ParetoCSP2 is a genetic algorithm that predicts multiple crystal structures (polymorphs) of the same chemical formula by forcing the search to explore many different symmetry groups. It reports large accuracy gains over two earlier algorithms, but the gains rest on a benchmark and an energy model that share the same training data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Polymorphism benchmark is drawn from the same Materials Project database used to train the surrogate ML potentials, so the reported coverage rates may reflect memorization rather than predictive search; no holdout or exclusion is described.","rationale":"The reader's weakest_assumption was the accuracy of ML-potential energy ranking, which is a legitimate limitation and is explicitly supported by the paper's own failure case. However, the more load-bearing concern is the circular evaluation: the benchmark and the surrogate are both derived from Materials Project, and no exclusion of benchmark structures from training is described. This threatens the validity of the absolute performance numbers, not just the search direction in difficult cases. The reader did mention the training/test overlap in their rationale, so there is partial agreement, but they did not elevate it to the primary issue. If the overlap is confirmed, the central claim of 'excellent performance in polymorphism prediction' is not established, because the algorithm may be retrieving structures it has effectively seen during pretraining. The concept of the algorithm (adaptive space-group diversity control) remains plausible and useful, but the current evidence is insufficient to verify the claimed predictive capability. A holdout-based evaluation would settle the matter; depending on the outcome, the verdict could return to conditional acceptance or move to rejection. Therefore, the current evidence warrants 'UNVERDICTED' rather than the reader's 'CONDITIONAL.'","tokens_in":22316,"tokens_out":9760,"duration_ms":114353,"concrete_test":"Check whether the 50 benchmark polymorphs appear in the CHGNet MPtrj and M3GNet training sets (e.g., by Pymatgen structure matching). If any are present, retrain or fine-tune an ML potential on a version of Materials Project with all benchmark polymorphs removed (or use a leave-one-out split) and rerun ParetoCSP2 on the same 50 formulas; compare the coverage rates with Table 2. A substantial drop would confirm that the reported performance is due to training overlap. Also run at least five independent seeds to assess variance, since the current tables report no stochastic error bars.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The evaluation of polymorphism prediction (Sec. 'Benchmark dataset for polymorphism CSP') uses ground-truth polymorphs from the Materials Project database, while the energy and relaxation surrogate is either M3GNet or CHGNet, both trained on Materials Project data (CHGNet on MPtrj, ~1.5M structures). The paper nowhere states that the benchmark compositions or their polymorphs were excluded from the training data. If the target structures are in the training set, the surrogate's energy landscape is biased toward these known polymorphs, so the high coverage rates in Table 2 (e.g., 96.67% space-group coverage for two-polymorph formulas) may be an artifact of memorization rather than evidence of a general polymorphism search. This concern is exacerbated by the paper's own failure case (Supplementary Fig. S11), where the potential assigns an intermediate structure an energy below the ground truth, demonstrating that the surrogate is not a reliable oracle. Because the central claim is 'excellent performance in polymorphism prediction,' an evaluation on training data cannot establish that claim. The missing baseline comparison for the polymorphism task (the reported 2.46-8.62x improvements come from regular CSP, Fig. 5, not from Table 2) is a related but distinct gap; the training/evaluation overlap is the more fundamental validity threat.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes ParetoCSP2, an evolutionary crystal structure prediction algorithm aimed at polymorphism prediction. The key novelty is a third optimization objective, added to energy and genotypic age, that penalizes over-representation of any single space group in the NSGA-III/AFPO population, combined with PyXtal-based initialization and per-generation structural relaxation using pretrained ML interatomic potentials (M3GNet or CHGNet). The authors evaluate polymorphism recovery on 50 formulas from the Materials Project with equal unit-cell atom counts and up to 10 polymorphs, and regular CSP on 120 CSPBench structures. They report near-perfect space-group and StructureMatcher coverage for two-polymorph formulas, declining coverage for higher polymorph counts, and large improvements over ParetoCSP and GN-OA on the regular-CSP benchmark metrics.","tokens_in":22558,"tokens_out":5287,"duration_ms":56137,"significance":"If the results are reproducible and the evaluation is not confounded by training-data overlap, the space-group diversity objective is a simple and plausible mechanism for maintaining multiple polymorphic candidates, and the per-generation relaxation plus PyXtal initialization are practical accelerations. The paper ships source code and adopts quantitative structural metrics (ED, SD, CD, HD, FP) that make comparisons transparent. However, the polymorphism claim currently rests on a benchmark drawn from the same database that trained the surrogate potentials, no polymorphism baseline is provided, and the stochastic GA is evaluated without repeated runs or error bars. These issues must be resolved before the claimed 'excellent performance in polymorphism prediction' can be accepted.","major_comments":[{"comment":"The polymorphism benchmark ground-truth structures are taken from the Materials Project, while the search and relaxation are guided by M3GNet and CHGNet, both of which were trained on Materials Project data (CHGNet on MPtrj, ~1.5 million structures). The paper nowhere states that the 50 benchmark formulas or their polymorphs were excluded from the training data. Recovering a structure that is effectively a training example is not the same as discovering it de novo, so the 96.67% space-group coverage for two-polymorph formulas in Table 2 cannot currently be attributed to the search algorithm. I ask the authors to (i) state explicitly which of the 50 formulas overlap with the M3GNet/CHGNet training sets, (ii) add a holdout evaluation or re-validate final candidates with DFT, and (iii) include a control experiment with random structure generation to estimate how much of the coverage is attributable to the surrogate's energy landscape alone.","section":"Benchmark dataset for polymorphism CSP; Table 2; Methods: ML IAPs"},{"comment":"ParetoCSP2 is a stochastic genetic algorithm (random initialization via PyXtal, random parent selection, mutation), but no repeated runs, random seeds, or error bars are reported for any result in Table 2 or Fig. 5. A single run of a GA can produce coverage rates that differ substantially from the expected value, particularly for the small samples in the 6-10 polymorph classes, which contain only one or two formulas each. The authors should report mean and standard deviation over at least 5-10 independent seeds for the central metrics and specify the population size, number of generations, and j/k values used for each benchmark formula.","section":"Performance evaluation for polymorphism prediction; Fig. 5"},{"comment":"The abstract states that ParetoCSP2 'outperforms baseline algorithms by factors of 2.46-8.62 for these accuracies,' but the 2.46-8.62 factors are the improvement in space-group and StructureMatcher success rates on the regular-CSP benchmark (Fig. 5), not on the polymorphism benchmark. For the polymorphism benchmark itself, no comparison against a baseline algorithm is provided, even though CALYPSO's symmetry-orientated divide-and-conquer method (Ref. 34) is described as the closest prior work. The claim of 'excellent performance in polymorphism prediction' therefore lacks comparative support. I suggest either adding a polymorphism-task baseline (e.g., CALYPSO or a random-sampling control) or rephrasing the abstract so the factor improvements are attributed to regular CSP only.","section":"Abstract; Results: Performance evaluation for regular CSP (Fig. 5)"},{"comment":"The selection of the 50 polymorphism benchmark formulas is not transparent: the text says 'We chose a total of 50 formulas' but gives no deterministic selection rule or sampling procedure beyond the constraints of ≤20 atoms and ≤10 polymorphs, and the higher-polymorph classes contain only one or two formulas. Because the average coverage rates in Fig. 3 are the primary evidence for the polymorphism claim, the authors should (i) document the exact selection criteria, (ii) state clearly whether Supplementary Table S1 is exhaustive or a subset, and (iii) either add more samples for the 6-10 polymorph classes or refrain from drawing quantitative trends from those one- or two-sample averages.","section":"Benchmark dataset for polymorphism CSP; Fig. 3"}],"minor_comments":[{"comment":"In the paragraph reporting success rates, the sentence 'Evidently, ParetoCSP outperformed ParetoCSP and GN-OA by 162.50% and 687.11%' should read 'ParetoCSP2 outperformed ParetoCSP and GN-OA...'.","section":"Results: Performance evaluation for regular CSP"},{"comment":"The flowchart contains typos such as 'Enerrgy', 'Calcualate', and 'corrensponding'; these should be corrected.","section":"Fig. 2 and flowchart text"},{"comment":"The terms 'shallow relax' and 'deep relax' are used without quantitative definition; please specify the number of relaxation steps or convergence criteria for each.","section":"Methods: ML IAPs and Discussion"},{"comment":"The claim that ParetoCSP2 converges within 1-10 generations is based on a selected sample of crystals; please clarify whether this holds across the full 120-crystal benchmark or add per-crystal convergence data.","section":"Comparative analysis of convergence speed; Fig. 10"},{"comment":"The StructureMatcher thresholds (ftol=0.2, stol=0.3, angle tolerance 5 degrees) are introduced without sensitivity analysis; a short ablation or a discussion of how the reported rates depend on these choices would strengthen the evaluation.","section":"Evaluation metrics"}],"recommendation":"major_revision","confidential_remarks":"The central algorithmic idea is plausible and the regular-CSP comparison, if confirmed with repeated runs, would be a useful contribution. The main validity threat is the training/evaluation overlap: the polymorphism benchmark is drawn from the same Materials Project data used to train M3GNet/CHGNet, and the paper does not describe any exclusion. I would ask the editor to require a holdout or DFT-validation component, repeated-run statistics, and a polymorphism-specific baseline before acceptance. The manuscript is within scope for the journal, but the current abstract overstates the polymorphism results relative to the evidence provided."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: ParetoCSP2 is a real, incremental improvement over ParetoCSP and GN-OA. Adding a shared space group count as a third objective in an NSGA-III framework, plus PyXtal initialization and per-generation ML relaxation, is a sensible combination that appears to work. The space group diversity plots are convincing: the algorithm consistently maintains many more distinct space groups than the baselines, and the regular CSP benchmark results (120 crystals from CSPBench) show large gains across all five metrics. The code and data are public, and the failure case study (Supplementary Fig. S11) shows the authors know the surrogate is not an oracle. Credit where due: the central mechanism is plausible and the engineering is transparent.\n\nThe soft spots are real and mostly fixable. First, the polymorphism evaluation is drawn from the same Materials Project database that trained M3GNet/CHGNet. The paper never states that the benchmark compositions or their polymorphs were excluded from training. With surrogates trained on ~1.5M MP structures, high coverage rates on MP polymorphs could reflect memorization as much as search. This is the load-bearing validity threat, and it needs either a clean holdout, an explicitly excluded composition set, or a demonstration that the same coverage holds on polymorphs only discovered after the training cutoff. Second, the abstract says \"evaluated on a benchmark dataset, it outperforms baseline algorithms by factors of 2.46-8.62 for these accuracies\" — but those factors come from the regular CSP comparison, not the polymorphism benchmark, where no baseline is run. That's a misleading presentation, not just a wording glitch. Third, for a stochastic GA there are no repeated runs, seeds, or error bars, so we can't tell how stable the high numbers are. Fourth, the selection of the 50 polymorph formulas is not fully described (e.g., how the random choice was made, and whether it biases toward easy cases).\n\nNone of this kills the idea. The third-objective diversity control is likely useful beyond this specific framework, and the paper is worth a careful review. But the strong polymorphism claims need to be backed by an evaluation that rules out training-set leakage, and the abstract needs to be rewritten to separate what was measured in the polymorph task from what was measured in the regular CSP task.\n\nFor a peer review, I'd send it out with a request for major revision. The core method is sound enough that a referee can meaningfully evaluate it; the gaps are identifiable and fixable.","headline":"Solid incremental advance in polymorph CSP, but the evaluation has a train/test overlap problem and missing polymorphism baselines that need addressing before the strong claims are credible.","tokens_in":23132,"tokens_out":2173,"would_cite":true,"duration_ms":21922,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adaptive space group diversity control lets one genetic algorithm recover multiple crystal polymorphs, including nearly all two-polymorph cases.","keywords":["polymorphism prediction","crystal structure prediction","genetic algorithm","space group diversity","neural network interatomic potential","multi-objective optimization","materials discovery"],"falsifier":"Take a two-polymorph formula where the ML potential's energy ranking of the two relaxed structures is known to contradict a DFT ranking, then run ParetoCSP2 with the ML potential as the fitness oracle; if space-group coverage stays high on such formulas, the diversity control alone compensates for misrankings, but if coverage drops, the method's polymorphism guarantees depend on the potential's accuracy. A direct test would rerun the benchmark with DFT energies substituted for ML energies at the final selection step and compare the recovered polymorph counts.","tokens_in":22078,"feed_emoji":"💎","tokens_out":5434,"duration_ms":53816,"temperature":0.7,"pith_summary":"This paper claims that crystal structure prediction can recover multiple polymorphs of an inorganic formula if the search explicitly prevents any one space group from dominating the genetic-algorithm population. The proposed method, ParetoCSP2, adds the count of structures sharing a space group as a third optimization objective alongside energy and genotypic age, initializes populations with a symmetry-aware structure generator, and relaxes all candidates every generation with a neural-network interatomic potential. On a benchmark of 50 formulas with equal-atom-count polymorphs, it reports near-perfect space-group and structural-similarity coverage for two-polymorph cases and sizable gains over two baseline algorithms. The broader stakes are that materials design could screen polymorphs computationally without expensive density-functional-theory sweeps.","feed_headline":"Adaptive diversity control recovers nearly all crystal polymorphs","feed_subtitle":"A genetic algorithm that penalizes space-group crowding beats baselines by 2.5–8.6× and identifies near-all two-polymorph structures.","key_machinery":"The load-bearing mechanism is the adaptive space group diversity control: each candidate structure is assigned a shared space group count equal to how many other individuals in the population share its space group, and this count is minimized as a third objective in an NSGA-III/AFPO multi-objective genetic algorithm. This counteracts the natural tendency of the search to converge onto a handful of low-energy, high-symmetry space groups. Supporting components are a PyXtal-based initial population that produces physically valid, symmetry-diverse crystals; iterative shallow relaxation of all structures after each generation using a neural-network interatomic potential; and tracking the best $j$ structures per $k$ distinct space groups (here $j=3$, $k=10$) to output a polymorph set.","core_discovery":"ParetoCSP2's central claim is that polymorphism prediction becomes tractable when space-group diversity is treated as an explicit objective rather than an emergent property of the search. Each individual is assigned a shared space group count — the number of population members with the same space group — and the multi-objective optimizer minimizes that count together with energy and age. The paper reports that this adaptive space group diversity control, combined with PyXtal-based initialization and per-generation relaxation using a neural-network interatomic potential, yields an average space-group coverage of 96.67% and complete StructureMatcher coverage for formulas with two polymorphs, and improvements of 44.8%–87.04% over baselines on regular crystal structure prediction.","pith_inferences":["If the space-group crowding penalty is the true driver of the gains, then varying the penalty strength or replacing the fixed budget $j,k$ with a temperature-like schedule could tune exploration versus exploitation; the paper does not study this trade-off.","The reliance on a single pretrained potential means the method's polymorph coverage for near-degenerate low-symmetry structures is likely bounded by the potential's energy resolution rather than by the search; substituting DFT energies for the final candidates would test this directly.","Because the benchmark only includes same-atom-count polymorphs from the Materials Project, the method's performance on variable-cell polymorphs or on polymorphs separated by larger energy gaps remains open; the paper's silica case suggests such cases are harder."],"forward_implications":["For formulas with two known polymorphs of equal unit-cell atom count, ParetoCSP2 is claimed to recover essentially all of them (96.67% average space-group coverage, 100% structural similarity), suggesting that routine polymorph enumeration is feasible for simple systems.","Because space-group diversity is preserved throughout the run rather than only at the end, the algorithm is claimed to find the optimal structure within 1–10 generations, which would make high-throughput polymorph screening practical with ML potentials.","The same diversity control improves standard crystal structure prediction, with reported gains of 44.8%–87.04% over ParetoCSP and GN-OA across energy and structural-distance metrics.","The method handles some variable-unit-cell polymorphism, such as ZnS wurtzite and zincblende, by running separate searches per cell size, though it fails on $\\alpha$-quartz in the case study."],"supporting_citations":[{"why":"The base algorithm ParetoCSP that this work extends with space-group diversity control.","marker":"[30]"},{"why":"The age-fitness Pareto optimization technique that provides the multi-objective selection framework.","marker":"[32]"},{"why":"NSGA-III, the reference-point-based many-objective genetic optimizer used for population selection.","marker":"[31]"},{"why":"M3GNet, the neural-network interatomic potential used for energy evaluation and relaxation.","marker":"[27]"},{"why":"CHGNet, an alternative neural-network interatomic potential used for energy and relaxation.","marker":"[28]"},{"why":"PyXtal, the symmetry-aware structure generation library that supplies valid, diverse initial crystals.","marker":"[48]"},{"why":"CSPBench, the benchmark dataset used to evaluate regular crystal structure prediction performance.","marker":"[57]"},{"why":"The Materials Project database, source of both the polymorph test set and the regular CSP test crystals.","marker":"[55]"}],"fun_headline_variants":["Space-group diversity keeps polymorph predictions on target","Genetic algorithm with diversity control nails polymorph prediction","Adaptive diversity control boosts crystal structure prediction accuracy","ParetoCSP2: diverse space groups lead to near-perfect polymorph hits","Diversity-aware genetic search finds all two-polymorph structures"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The neural-network interatomic potential's energy rankings are accurate enough to guide the search toward all relevant polymorphs, especially low-symmetry structures whose energies differ by only a few meV per atom; the paper's own failure case shows the potential assigning an intermediate structure an energy below the ground truth, which would mislead any search.","fun_headline_variants_meta":{"raw":{"variants":["Space-group diversity keeps polymorph predictions on target","Genetic algorithm with diversity control nails polymorph prediction","Adaptive diversity control boosts crystal structure prediction accuracy","ParetoCSP2: diverse space groups lead to near-perfect polymorph hits","Diversity-aware genetic search finds all two-polymorph structures"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000768,"raw_usage":{"total_tokens":3416,"prompt_tokens":970,"completion_tokens":2446,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":586,"completion_tokens_details":{"reasoning_tokens":2366}},"tokens_in":586,"tokens_out":2446,"duration_ms":19517,"temperature":1.0,"reasoning_tokens":2366,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:10:57.230803+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a two-polymorph formula where the ML potential's energy ranking of the two relaxed structures is known to contradict a DFT ranking, then run ParetoCSP2 with the ML potential as the fitness oracle; if space-group coverage stays high on such formulas, the diversity control alone compensates for misrankings, but if coverage drops, the method's polymorphism guarantees depend on the potential's accuracy. A direct test would rerun the benchmark with DFT energies substituted for ML energies at the final selection step and compare the recovered polymorph counts.","supporting_citations":[{"cited_title":"S.; Wei, L.; Hu, M.; Hu, J","cited_arxiv_id":null,"evidence_quote":"The base algorithm ParetoCSP that this work extends with space-group diversity control."},{"cited_title":"D.; Lipson, H","cited_arxiv_id":null,"evidence_quote":"The age-fitness Pareto optimization technique that provides the multi-objective selection framework."},{"cited_title":"An evolutionary many-objective optimization algorithm using reference-point-based nondominated sorting approach, part I: solving problems with box constraints","cited_arxiv_id":null,"evidence_quote":"NSGA-III, the reference-point-based many-objective genetic optimizer used for population selection."},{"cited_title":"PyXtal: A Python library for crystal structure generation and symmetry analysis","cited_arxiv_id":null,"evidence_quote":"PyXtal, the symmetry-aware structure generation library that supplies valid, diverse initial crystals."},{"cited_title":"CSPBench: a benchmark and critical evaluation of Crystal Structure Prediction","cited_arxiv_id":"2407.00733","evidence_quote":"CSPBench, the benchmark dataset used to evaluate regular crystal structure prediction performance."}],"review_version":1}