{"id":"aea8b2c1-2b80-4938-95cb-450a26f12cc5","arxiv_id":"2412.16736","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"autoplex automates iterative random-structure-search data generation and fitting of machine-learned potentials, yielding usable models for Si, Ti-O, SiO2, water, and phase-change materials.","lead":"autoplex is an open-source framework that automates the iterative loop of random structure searching, density-functional theory labelling, and machine-learned potential fitting. It aims to remove the manual dataset-curation bottleneck that slows the development of interatomic potentials for atomistic simulation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Robustness claim rests on unquantified force accuracy and held-out transfer to liquid/amorphous systems; the MD demonstrations are not validated against held-out DFT data.","rationale":"The reader's weakest assumption—that small-cell RSS generates a sufficiently diverse training distribution for transfer to large bulk liquid and amorphous systems—is closely related to my concern, but I sharpen it: even within the demonstrated systems, there is no held-out energy or force validation. The paper's strongest claim is that the potentials are 'robust and useful', and the most impressive evidence is qualitative MD behaviour. MD is governed by forces, yet the paper only reports energy errors, and those are on crystal-like structures near the RSS distribution. The Methods explicitly note the small unit cells used in RSS, and Fig. 4d documents a clear extrapolation failure of GAP to ice polymorphs, which underscores the risk. This is not an internal inconsistency but a missing measurement, and it is directly testable. Because the software contribution, breadth of demonstrations, and candid discussion of limitations are real strengths, conditional acceptance remains appropriate; however, the condition should explicitly require held-out force/energy evaluation on the target liquid/amorphous systems. My verdict is therefore unchanged from the reader's CONDITIONAL, with the concern made more precise and actionable.","tokens_in":15509,"tokens_out":5652,"duration_ms":53254,"concrete_test":"Take the final GAP-RSS and NequIP autoplex models used for Figs. 4–5 and evaluate them on 100–200 held-out configurations drawn from SCAN DFT-MD trajectories of liquid water and amorphous Ge1Sb2Te4, ensuring these configurations are not in the training set. Compute energy and force RMSE per atom. If the force RMSE exceeds roughly 0.3 eV/Å or the energy RMSE exceeds roughly 30 meV/atom, or if these errors are substantially larger than on RSS-derived validation structures, the transferability/robustness claim is not supported. Additionally, recompute the liquid-water O–O RDF from a SCAN DFT-MD trajectory under identical simulation conditions, rather than experimental data, to isolate model error from functional error.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that autoplex yields 'robust and useful' potentials is supported mainly by energy RMSEs on crystal polymorphs (Fig. 2, Tables 1–2), which are close to or within the RSS training distribution. The most compelling demonstrations—liquid water RDFs (Fig. 4a–c) and Ge1Sb2Te4 crystallisation (Fig. 5g)—are MD simulations, which depend on force accuracy, yet no force-error metric is reported anywhere in the paper. Furthermore, no held-out energy or force evaluation is performed on the actual application systems: the water RDF is compared to experiment, not to SCAN DFT-MD at the same conditions, and the crystallisation simulation is compared qualitatively to literature AIMD. The Methods state that RSS searches 'typically use small unit cells', and Fig. 4d itself shows that the GAP-RSS model extrapolates poorly to ice polymorphs. Therefore, the load-bearing assumption that small-cell RSS training data transfer to large bulk liquid and amorphous systems in both energy and force space is unmeasured. If this assumption fails, the headline demonstrations are anecdotal rather than evidence of robustness.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces autoplex, an automated workflow framework that combines random structure searching (RSS) with iterative fitting of machine-learned interatomic potentials, primarily GAP but with interfaces to other architectures such as NequIP, MACE, and ACE. The workflow is demonstrated on elemental silicon, TiO2 and the full Ti-O binary system, SiO2 at PBE and SCAN levels, liquid water and ice using GAP and NequIP, and the phase-change materials Ge1Sb2Te4 and In3Sb1Te2. The central claim is that MLIP fitting can be carried out in a largely automated, high-throughput manner and that the resulting potentials are robust and useful, especially given the ease of creating them from scratch.","tokens_in":15680,"tokens_out":7435,"duration_ms":63177,"significance":"If the claims are substantiated, this is a valuable contribution: it packages iterative RSS-plus-MLIP fitting in an open-source, workflow-based tool that integrates with atomate2 and jobflow, potentially lowering the barrier to building potentials from scratch. The demonstrations span several chemistries and multiple MLIP architectures, and the SiO2/SCAN example shows that the automated pipeline can access higher-rung functionals at moderate cost. Named strengths include the open code with a Zenodo-deposited version, the interoperability with existing workflow ecosystems, and the demonstration that an RSS-generated dataset can be reused with a different architecture (NequIP). The main gap is validation: the robustness claim currently rests on energy errors that are partly in-sample and on qualitative MD comparisons without force-error metrics or explicit held-out tests on the application systems.","major_comments":[{"comment":"The error metrics are reported without an explicit statement of train/test separation. Fig. 2 shows the evolution of energy errors for selected crystalline phases, but it is not stated whether those phases, or rattled versions of them, are included in the training set at each iteration. Table 1 evaluates errors on 10 rattled structures per polymorph; because these are generated from the ground-state structures that the RSS pipeline is designed to find, they are likely inside or very close to the training distribution. As a result, the reported RMSEs do not by themselves establish that the potentials are robust for unseen configurations. Please report held-out errors and state explicitly which structures enter the training set at each stage.","section":"Results; Fig. 2 and Table 1"},{"comment":"The MD demonstrations are the strongest evidence for transferability, but no force-error metric is reported anywhere in the manuscript. The liquid-water RDFs and hydrogen-bond analysis (Fig. 4a-c) and the Ge1Sb2Te4 crystallisation simulation (Fig. 5g) are driven by forces, yet the paper only reports energy RMSEs. Please add force RMSEs on held-out configurations representative of the application systems, for example SCAN DFT-MD snapshots of liquid water and AIMD snapshots of amorphous GST/IST. Without such metrics, the conclusion that small-cell RSS data transfer to bulk liquid and amorphous systems in force space is unsubstantiated.","section":"Describing water; Application to chalcogenide memory materials; Figs. 4 and 5"},{"comment":"Fig. 4d shows that the GAP-RSS model's energies for ice polymorphs are 'highly scattered', and the authors attribute this to low-density phases outside the training distribution. This is a direct counterexample to the general claim that RSS-derived potentials are robust across phases. Please either soften the robustness claim in the abstract and introduction or provide an analysis of when the workflow produces a transferable GAP, for example by detecting extrapolation through uncertainty estimates or active learning. The current text acknowledges the limitation in one sentence but does not reconcile it with the central claim.","section":"Describing water; Fig. 4d"},{"comment":"The Data Availability statement says that raw data and notebooks 'will be made available via GitHub upon journal publication', so the numerical results in Figs. 2-5 and Tables 1-2 are not currently available to reviewers or readers. Since the contribution is an automated framework that others should be able to reproduce, please deposit the datasets, trained potentials, and plotting notebooks, or at least the train/test splits and error tables, in a public repository with the submitted version.","section":"Data availability"}],"minor_comments":[{"comment":"The statement that RSS searches 'typically use small unit cells' should be quantified: please give the cell sizes and number of molecules used in the water RSS runs so the reader can judge how far the liquid-water simulation extrapolates from the training data.","section":"Describing water"},{"comment":"The simulation protocols for Figs. 4a-c and 5c-g (ensemble, thermostat, system size, simulation length, and initial configurations) are not given in the Methods; please add these details for reproducibility.","section":"Describing water; Application to chalcogenide memory materials"},{"comment":"The license is described only as 'permissive'; please specify the exact open-source license in the text or in the repository.","section":"The autoplex framework"},{"comment":"For the entries marked '— a', the footnote states that errors exceed 1 eV/atom and are therefore not meaningful to report; giving approximate numerical values would be more informative for assessing the failure mode.","section":"Table 1"},{"comment":"Please clarify whether the 0.001 eV/atom floor is a plotting cutoff or a stopping criterion used during the iterative fitting.","section":"Fig. 2 caption"},{"comment":"Equation (1) defines the total-energy decomposition but the force expression is not given; please state that forces are obtained as derivatives of the total energy and give the relevant implementation detail or reference.","section":"Methods; ML potentials"},{"comment":"The phrase 'standard DFT and GAP fitting settings' should be replaced by the actual settings or a pointer to repository defaults, since these settings are not standard across the community.","section":"Results; Ti-O system"}],"recommendation":"major_revision","confidential_remarks":"This is a solid software/automation contribution, and I do not see grounds for rejection. The main gap is validation: the load-bearing claims about robustness and transferability to liquid and amorphous systems would be materially strengthened by force-error metrics and explicit held-out evaluations, and the in-sample nature of the energy-error tables needs to be addressed. The data availability statement should also be fulfilled before acceptance so that the numerical results can be checked."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe short version: this is a genuine software contribution, not a scientific breakthrough. The iterative RSS-plus-MLIP idea is prior work (Refs 44 and 48), and wfl already offered custom workflows (Ref 64). What autoplex adds is integration with the atomate2/jobflow ecosystem, support for multiple buildcell parameter sets, a Hookean repulsion term, and interfaces to NequIP, MACE, ACE, and M3GNet, with a versioned open release on Zenodo. That is real, and it directly addresses the data-generation bottleneck that dominates MLIP development.\n\nThe paper does well. Demonstrations span silicon, Ti–O, SiO2 at PBE and SCAN levels, liquid water with two MLIP architectures, and phase-change materials including a crystallisation trajectory. The authors are candid: they flag the GAP model's poor extrapolation to ice polymorphs and the absence of nuclear quantum effects in the water simulations. Code is open, methodology is reproducible in principle, and the claimed cost savings (around $100 for a quartz-level potential) are plausible.\n\nThe soft spots are real but not disqualifying. The robustness claim rests mostly on energy RMSEs for crystal polymorphs that are close to or inside the RSS training distribution. No force-error metric appears anywhere, although the most convincing demonstrations — liquid water RDFs and Ge1Sb2Te4 crystallisation — are MD simulations whose correctness depends on forces. The water RDF is compared to experiment, not to SCAN DFT-MD at the same conditions; the crystallisation is compared qualitatively to literature AIMD. The Methods admit that RSS searches typically use small unit cells, and Fig. 4d shows the GAP model struggling on ice. So the load-bearing assumption — that small-cell random structures transfer to bulk liquid and amorphous systems — is not directly measured. Raw data and notebooks are promised only for after publication, and the main error figures lack uncertainty estimates and explicit train/test separation. All of that is fixable in revision.\n\nThe central automation claim holds. The paper deserves a serious referee. I would send it to peer review with requests for held-out energy and force errors on the actual application systems, a force-error table, and better uncertainty reporting.","headline":"Useful, honest automation paper; the transferability claim needs held-out force metrics.","tokens_in":16275,"tokens_out":2807,"would_cite":true,"duration_ms":23331,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An open-source workflow turns random structure searches into ready-to-use machine-learned interatomic potentials.","keywords":["automated machine-learned interatomic potentials","random structure searching","GAP-RSS","potential-energy surface exploration","high-throughput workflows","atomate2","Gaussian approximation potentials","active learning"],"falsifier":"Train an autoplex GAP-RSS model on a new small-cell random structure set for water and run molecular dynamics at 300 K; if the predicted O–O radial distribution function first peak deviates by more than roughly 0.05 Å from the experimental value, or the hydrogen-bond count falls outside the 3.48–3.84 range, the claim that RSS alone provides transferable robustness would fail.","tokens_in":15302,"feed_emoji":"⚛️","tokens_out":5862,"duration_ms":47666,"temperature":0.7,"pith_summary":"The paper presents autoplex, an open-source software framework that automates the construction of machine-learned interatomic potentials (MLIPs) from scratch, replacing hand-curated training sets with an iterative loop of random structure searches, DFT single-point evaluations, and model refitting. The central claim is that this largely automated, high-throughput pipeline can produce potentials that are accurate and robust enough for real materials applications, demonstrated on silicon, the titanium–oxygen system, silica, liquid and crystalline water, and phase-change memory materials. If true, it means the most labor-intensive part of building an MLIP—deciding which configurations to compute and curating them—can be handed to the computer. The authors argue that automation of data generation is the next frontier for making ML-driven atomistic simulation a mainstream tool.","feed_headline":"Random searches automate machine-learning potential fitting from scratch","feed_subtitle":"Autoplex turns random structures plus DFT labels into usable potentials for water, oxides, and phase-change materials.","key_machinery":"The load-bearing machinery is the iterative GAP-RSS loop: the buildcell code generates random structures under user-chosen constraints, the current MLIP relaxes them, a subset is labelled with single-point DFT, the new data are added to the training set, a new potential is fitted, and the cycle repeats. A Hookean repulsion term keeps atoms from approaching unphysically close during relaxation. The automation layer wraps this loop in workflow-management infrastructure, so thousands of tasks can be submitted and monitored on high-performance computing systems without manual intervention. This loop is what transfers the burden from hand-curated datasets to the random-search parameters.","core_discovery":"The central claim is that a GAP-RSS-style workflow—randomly generating small-cell structures, relaxing them with a progressively improved Gaussian approximation potential, labelling only with single-point DFT, and refitting—can be implemented as a modular, automated workflow system and still yield potentials that describe not just the training region but also bulk liquids, amorphous phases, and crystallisation dynamics. The demonstrations show energy errors near 0.01 eV per atom for held-out polymorphs, correct polymorph stability ordering at the SCAN level for SiO2, qualitatively correct liquid-water structure and hydrogen-bond counts, and a 350 ps crystallisation simulation of Ge1Sb2Te4. The paper presents this as evidence that automated random searching can serve as a general starting point for MLIP construction, including for materials with little existing domain knowledge.","pith_inferences":["The same automation could be extended to multi-element and disordered systems beyond the demonstrated binaries and ternaries, potentially enabling high-throughput screening of MLIPs across many chemistries.","Small-cell random searches may under-sample long-ranged or slowly relaxing degrees of freedom, so adding uncertainty-based active learning to the loop could further improve transfer to large amorphous systems.","Autoplex-generated RSS datasets could serve as a cheap pre-training or synthetic-data source for foundational MLIP fine-tuning, extending the paper's observation that a NequIP model benefits from the same dataset.","The paper's cost estimates for SiO2 suggest that automated MLIP construction could become a routine pre-screening tool in computational materials discovery, not just a specialist technique."],"forward_implications":["MLIP models for a new material system can be created from scratch in an automated run, with the user's main choice being the random-search constraints and the DFT functional.","Automated RSS datasets serve as training data not only for GAP but also for other architectures: the paper shows a NequIP model fitted to the same dataset improves extrapolation to ice polymorphs.","Because only single-point DFT is needed, higher-rung functionals such as SCAN become affordable for building potentials, enabling qualitatively correct stability ordering where PBE fails.","Materials with little domain knowledge, such as In3Sb1Te2, can get a first usable potential without months of hand curation.","The workflow integrates with existing high-throughput infrastructure, making iterative MLIP fitting accessible on large HPC systems."],"supporting_citations":[{"why":"Establishes the GAP-RSS idea: iterative random structure searching with MLIP fitting using only single-point DFT labels.","marker":"Ref. 44"},{"why":"Supplies the Gaussian approximation potential framework used for most of the paper's fits.","marker":"Ref. 15"},{"why":"Provides the AIRSS random structure generation approach and the buildcell code used in the workflow.","marker":"Ref. 41,42"},{"why":"Earlier GAP-RSS implementation with structural and energetic selection steps that autoplex builds on and benchmarks against.","marker":"Ref. 48"},{"why":"The atomate2 workflow ecosystem that autoplex interfaces with for automated DFT and job execution.","marker":"Ref. 68"},{"why":"A related workflow toolkit (wfl) mentioned as a contrast in design choices for automating MLIP dataset generation.","marker":"Ref. 64"},{"why":"The hand-built GST-GAP-22 model for phase-change materials, serving as the domain-specific baseline that autoplex aims to automate.","marker":"Ref. 26"}],"fun_headline_variants":["Autoplex automates ML potentials via random search","Random structure searches yield MLIPs from scratch","Automated potential fitting from random structures","Autoplex: exploring potential surfaces on autopilot"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that small-cell random structure searches, with user-chosen buildcell constraints and a Hookean repulsion term, generate a training distribution diverse enough for the resulting potential to transfer to large bulk liquid and amorphous systems.","fun_headline_variants_meta":{"raw":{"variants":["Autoplex automates ML potentials via random search","Random structure searches yield MLIPs from scratch","Automated potential fitting from random structures","Autoplex: exploring potential surfaces on autopilot"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1275,"prompt_tokens":855,"completion_tokens":420,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":471,"completion_tokens_details":{"reasoning_tokens":360}},"tokens_in":471,"tokens_out":420,"duration_ms":3876,"temperature":1.0,"reasoning_tokens":360,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:15:19.010609+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train an autoplex GAP-RSS model on a new small-cell random structure set for water and run molecular dynamics at 300 K; if the predicted O–O radial distribution function first peak deviates by more than roughly 0.05 Å from the experimental value, or the hydrogen-bond count falls outside the 3.48–3.84 range, the claim that RSS alone provides transferable robustness would fail.","supporting_citations":[],"review_version":1}