REVIEW 2 major objections 5 minor 5 references
Benchmarking Universal Interatomic Potentials on Zeolite Structures
T0 review · 2 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read Universal ML potentials match DFT on zeolites; eSEN leads the pack.
desk verdict Solid, reproducible benchmark of universal IPs on zeolites; the eSEN result is useful, but the abstract overstates what is tested for guest-containing zeolites—only energies at DFT geometries, not MLIP-relaxed geometries. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the pretrained universal machine-learning interatomic potential: a graph neural network trained on large DFT datasets and used as-is without retraining. The comparative engine is the relative-energy benchmark: all energies are referenced to the most stable structure of the same composition, so elemental-reference offsets cancel and the RMSE isolates each model's physical accuracy for polymorphs and cation arrangements.
What would settle it
Recalculate a subset of the benchmark (for example, 50 of the Cu/CHA structures and 50 of the K-OSDA/ERI structures) with a higher-level reference such as r2SCAN+D4 or a wavefunction-based method, and compare the model rankings; if eSEN's RMSE advantage shrinks or reverses, its lead is an artifact of the shared PBE training reference.
Extended reading notes
Core claim
The paper claims that modern pretrained universal machine-learning interatomic potentials have reached the point where, for zeolites, they act as practical surrogates for PBE+D3 DFT. Across all test sets—eight experimentally characterized silica structures, dozens of pure-silica topologies, 347 Cu/CHA structures with varied aluminum distributions and copper siting, and 1,190 K/OSDA-containing ERI structures—every MLIP reproduced DFT relative energies with small errors, and eSEN-30M-OAM had the lowest RMSE in every comparison (0.44 kJ/mol per Si for pure silicas; 0.14 and 0.02 kJ/mol per atom for Cu/CHA and K-OSDA/ERI, respectively). GFN-FF is the best universal analytic IP, matching ClayFF o
Load-bearing premise
The benchmark treats PBE+D3 DFT as the reference truth, and all the MLIPs are trained on similar PBE-level DFT, so the test measures how faithfully the potentials imitate that particular DFT rather than how physically accurate they are.
Editorial extensions
If this is right
- Pretrained universal MLIPs, especially eSEN-30M-OAM, can replace DFT for screening relative stabilities of zeolite frameworks and guest-containing catalysts at a fraction of the cost.
- SLC remains the fastest reliable choice for pure-silica zeolite geometries, so the practical hierarchy is SLC for simple frameworks and MLIPs for compositionally complex systems.
- Universal analytic potentials such as UFF, Dreiding, and even GFN-FF should not be used blindly for zeolites: strained three-membered rings and aluminosilicate guests trigger large structural and energetic errors.
- The newly generated Cu/CHA and K-OSDA/ERI structures are unlikely to be in current MLIP training sets, making them reusable benchmarks for future universal potentials without data leakage.
- Transition-metal-containing Cu/CHA systems are consistently harder for all MLIPs than alkali/organic-cation systems, indicating where future model improvements will matter most.
Reading between the lines
- Beyond the paper, the ranking should be read as a measure of how well these models imitate Materials-Project-style PBE DFT, not of absolute physical accuracy; a reference change (e.g., r2SCAN, hybrid functionals, or experiment) could reshuffle the leaders.
- A natural extension the paper does not test is whether eSEN also reproduces adsorption energies, reaction barriers, and dynamics in zeolites, since accurate relaxed geometries and energies do not guarantee accurate forces over long trajectories.
- The Cu/CHA set acts as a severe transferability probe: future universal MLIPs can be stress-tested on these 347 structures as a quick gauge of transition-metal chemistry before deployment in catalysis screening.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a benchmark of universal interatomic potentials (UFF, Dreiding, GFN-FF) and pretrained universal machine-learning interatomic potentials (CHGNet, ORB-v3, MatterSim, eSEN-30M-OAM, PFP-v7, EQUIVARIANT-v2/OC22) against tailor-made potentials (SLC, ClayFF, BSFF) for zeolite structures. Pure silica frameworks are evaluated against experimental structures and thermochemical data and against PBE+D3 DFT for a wider set of topologies. Guest-containing zeolites (347 Cu/CHA and 1,190 K-OSDA/ERI structures) are evaluated only through single-point energies computed at DFT-relaxed geometries. The main findings are that GFN-FF is the best universal analytic IP but fails for strained silica rings and aluminosilicates, whereas all tested MLIPs reproduce PBE+D3 energetics well, with eSEN-30M-OAM showing the lowest RMSE overall. The paper concludes that modern universal MLIPs are practical tools for zeolite screening.
Significance. If the claims hold, the study provides a useful, current benchmark for practitioners choosing interatomic potentials for zeolite modeling. The methodology is generally clean: consistent DFT settings (PBE+D3, 520 eV cutoff), a large and diverse structure set, multiple reference types, and public data release on Zenodo. The eSEN result is a concrete, falsifiable ranking and the analysis of training-data consistency across models is thoughtful. However, the central practical claim—that universal MLIPs are ready for screening workflows involving diverse compositions—rests on energy predictions at DFT geometries for the guest-containing systems, not on the MLIPs' ability to relax those structures. This gap in evidence is the main concern.
major comments (2)
- [Reproducibility of DFT results for guests containing zeolites (Fig. 4, Table 4)] For Cu/CHA and K-OSDA/ERI, the text states: 'We relax all structures using DFT with PBE+D3 and perform single-point calculations using GFN-FF and several universal MLIPs.' Thus the reported RMSE values and ranking describe energy prediction at fixed DFT geometries; the MLIPs are never used to relax or generate these guest-containing structures. The abstract's claim that 'All MLIPs can well reproduce experimental or DFT-level geometries and energetics' and the conclusion's recommendation of universal MLIPs as 'practical tools' for high-throughput screening across 'various compositions' overstate the evidence, because a screening workflow would normally require structure relaxation for new compositions. Please either add a subset of MLIP geometry relaxations for these systems and compare structures (e.g., Al distributions, cation siting, cell parameters), or explicitly limit the claim to s
- [Tables 2-4, RMSE metric] The headline ranking of eSEN relies on RMSE values that are extremely small for the guest-containing systems, especially K-OSDA/ERI (0.02-0.11 kJ mol^-1 atom^-1 ≈ 0.0002-0.001 eV/atom). The paper does not discuss whether these differences are statistically significant or practically meaningful, and the reader cannot tell whether the reported ordering is robust to the choice of reference configurations or to small perturbations in the DFT-relaxed geometries. I recommend adding an uncertainty estimate (e.g., bootstrap over the 1,190 configurations, or comparison on a held-out subsample) before asserting that eSEN 'excels all the other MLIPs' in both guest systems. This is a load-bearing point for the central ranking.
minor comments (5)
- [Abstract/Introduction] The phrase 'well reproduce experimental or DFT-level geometries and energetics' conflates the two reference levels. For pure silica, geometries are compared with experiment; for guest-containing zeolites, only single-point energies are compared with DFT. Please clarify in the abstract which claim applies to which system class.
- [Section 2 (first paragraph)] The sentence 'DFT with PBE+D3 best predicts experimental relative energies, and with slightly smaller error by eSEN' is awkward and potentially misleading. Table 2 shows eSEN has a slightly larger RMSE than DFT, not 'smaller error.' Rewrite for clarity.
- [Table 1] Table 1's formatting is confusing: 'DFT PBE, SCAN*' appears in the Tailor-made column, and the table header is not self-explanatory. Consider reorganizing so that DFT is clearly listed as a reference method rather than a potential, or move it to a separate row/column.
- [Supplementary Information / Methods] EqV2(OC22) is described in the SI as not supporting direct stress calculations, which is why it was excluded from pure-silica relaxations. This is an important methodological detail and should be mentioned in the main text where EqV2 is introduced, not only in the SI.
- [Conclusion] The statement 'since our guest-containing structures are unlikely to be used in training data for present and future universal MLIPs, they can also be utilized for evaluating future potentials, avoiding data leakage' is plausible but unverifiable. Please soften to 'we expect to be unlikely' and note that the exact composition of proprietary training sets is not publicly known.
Circularity Check
No circular derivation: the paper is an empirical benchmark with no fitted parameters, no load-bearing self-citation, and no equation that reduces to its inputs; the only mild issue is the acknowledged DFT-training/DFT-reference overlap and a guest-system geometry-testing gap.
full rationale
This is a benchmarking study rather than a derivation. The authors introduce no fitted parameters and no equations whose outputs are defined by their inputs; the central results are RMSE/MAE comparisons between pre-trained universal IPs and experimental or DFT references. There is no pattern of fitted-input-called-prediction: the models are used as-is, and the eSEN ranking is an observed outcome, not a fit. There is no load-bearing self-citation: the cited prior work by the same group (refs. 36 and 78) is contextual, not used to justify the benchmark's premise or to forbid alternatives. There is no imported uniqueness theorem and no ansatz smuggled in via citation. Two limitations are present and should be flagged, but neither is circular. First, the guest-containing zeolite claim in the abstract ('well reproduce experimental or DFT-level geometries and energetics') is only partly tested for those systems: the paper states, in 'Reproducibility of DFT results for guests containing zeolites,' 'We relax all structures using DFT with PBE+D3 and perform single-point calculations using GFN-FF and several universal MLIPs.' Thus for Cu/CHA and K-OSDA/ERI, geometry reproduction is not tested; only energies at DFT-relaxed geometries are compared. This narrows the evidence but is not circular. Second, the benchmark uses PBE+D3 DFT as reference while the universal MLIPs were trained on PBE-family DFT data; the paper acknowledges this in the Conclusion: 'the training data for universal MLIPs, as well as part of the reference data in our benchmark, are based on DFT calculations, which themselves involve intrinsic errors.' This makes the 'MLIPs reproduce DFT' finding partly a consistency check on the training-data family, but the specific test structures and relative-energy rankings are not determined by construction, and the pure-silica comparisons against experimental geometries and enthalpies provide an external anchor. The paper is therefore self-contained against external benchmarks and shows no significant circularity; the appropriate score is 1 due only to the mild, explicitly acknowledged DFT training/reference overlap.
Assumptions & free parameters
assumptions (4)
- domain assumption DFT using PBE with D3 correction is an accurate enough reference for zeolite geometries and relative energies
- domain assumption Experimental thermochemical data for pure silica zeolites (Navrotsky et al.) are reliable references
- domain assumption The training data of the universal MLIPs is consistent with the benchmark DFT settings (PBE, similar cutoff and pseudopotentials)
- domain assumption The generated Cu/CHA and K-OSDA/ERI structures are representative of realistic zeolite catalyst systems
Cite this review
Pith. "Pith review of Benchmarking Universal Interatomic Potentials on Zeolite Structures." pith.science (2026). https://pith.science/paper/CEE66VSR
@misc{pith2026250907417,
author = {Pith},
title = {Pith review of: Benchmarking Universal Interatomic Potentials on Zeolite Structures},
year = {2026},
howpublished = {\url{https://pith.science/paper/CEE66VSR}},
note = {Machine review of arXiv:2509.07417}
}
read the original abstract
Interatomic potentials (IPs) with wide elemental coverage and high accuracy are powerful tools for high-throughput materials discovery. While the past few years witnessed the development of multiple new universal IPs that cover wide ranges of the periodic table, their applicability to target chemical systems should be carefully investigated. We benchmark several universal IPs using equilibrium zeolite structures as testbeds. We select a diverse set of universal IPs encompassing two major categories: (i) universal analytic IPs, including GFN-FF, UFF, and Dreiding; (ii) pretrained universal machine learning IPs (MLIPs), comprising CHGNet, ORB-v3, MatterSim, eSEN-30M-OAM, PFP-v7, and EquiformerV2-lE4-lF100-S2EFS-OC22. We compare them with established tailor-made IPs, SLC, ClayFF, and BSFF using experimental data and density functional theory (DFT) calculations with dispersion correction as the reference. The tested zeolite structures comprise pure silica frameworks and aluminosilicates containing copper species, potassium, and organic cations. We found that GFN-FF is the best among the tested universal analytic IPs, but it does not achieve satisfactory accuracy for highly strained silica rings and aluminosilicate systems. All MLIPs can well reproduce experimental or DFT-level geometries and energetics. Among the universal MLIPs, the eSEN-30M-OAM model shows the most consistent performance across all zeolite structures studied. These findings show that the modern pretrained universal MLIPs are practical tools in zeolite screening workflows involving various compositions.
Reference graph
Works this paper leans on
-
[18]
Tran, R. et al. The Open Catalyst 2022 (OC22) Dataset and Challenges for Oxide Electrocatalysts. ACS Catal. 13, 3066–3084 (2023). 19. Fischer, M., Evers, F. O., Formalik, F. & Olejniczak, A. Benchmarking DFT-GGA calculations for the structure optimisation of neutral-framework zeotypes. Theor. Chem. Acc. 135, 257 (2016). 20. Navrotsky, A., Trofymluk, O. & ...
work page 2022
-
[22]
Lewis, G. V. & Catlow, C. R. A. Potential models for ionic oxides. J. Phys. C Solid State Phys. 18, 1149 (1985). 23. Rappe, A. K., Casewit, C. J., Colwell, K. S., Goddard, W. A. I. & Skiff, W. M. UFF, a full periodic table force field for molecular mechanics and molecular dynamics simulations. J. Am. Chem. Soc. 114, 10024–10035 (1992). 24. Barlow, S., Roh...
-
[38]
Ojih, J., Al-Fahdi, M., Yao, Y., Hu, J. & Hu, M. Graph theory and graph neural network assisted high-throughput crystal structure prediction and screening for energy conversion and storage. J. Mater. Chem. A 12, 8502–8515 (2024). 39. Wines, D. & Choudhary, K. CHIPS-FF: Evaluating Universal Machine Learning Force Fields for Material Properties. Preprint at...
work page Pith review arXiv doi:10.48550/arxiv.2412.10516 2024
-
[57]
Gale, J. D. GULP: A computer program for the symmetry-adapted simulation of solids. J. Chem. Soc. Faraday Trans. 93, 629–637 (1997). 58. Gale, J. D. & Rohl, A. L. The General Utility Lattice Program (GULP). Mol. Simul. 29, 291–341 (2003). 59. Tran, R. et al. The Open Catalyst 2022 (OC22) Dataset and Challenges for Oxide Electrocatalysts. ACS Catal. 13, 30...
work page 1997
-
[75]
Lewis, D. W., Willock, D. J., Catlow, C. R. A., Thomas, J. M. & Hutchings, G. J. De novo design of structure-directing agents for the synthesis of microporous solids. Nature 382, 604–606 (1996). 76. Schwalbe-Koda, D. & Gómez-Bombarelli, R. Benchmarking binding energy calculations for organic structure-directing agents in pure-silica zeolites. J. Chem. Phy...
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.