REVIEW 3 major objections 5 minor 36 references
Improving robustness and training efficiency of machine-learned potentials by incorporating short-range empirical potentials
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Bolting an empirical short-range repulsive potential (ZBL) onto a machine-learned force field prevents unphysical atomic clustering and lets a 25-configuration training set match a ~2000-configuration model for the solid electrolyte LLZO.
desk verdict Solid qualitative demonstration that ZBL short-range repulsion fixes a real MLFF clustering failure in LLZO; the headline '25 configurations' claim is over-claimed because it rests on surrogate labels and a distillation-style selection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the ZBL potential, an empirical screened-Coulomb repulsion originally designed for high-energy atom collisions, spliced onto the learned potential below an inner cutoff and smoothly switched off by a cosine function between 0.9 and 1.8 Å. The ZBL form supplies an essentially insurmountable repulsive barrier that prevents atoms from sampling the short distances where the learned potential is arbitrary. The paper stresses that neither the ZBL parameters nor its functional form need be exact; the mechanism is simply that a physical repulsive wall occupies the data-sparse short-range region.
What would settle it
Run a set of short-range Li–Li dimer calculations and compare NEP802, NEP802-ZBL, and NEPstd directly against density functional theory energies and forces at separations between 0.5 and 2.0 Å; if DFT shows that NEP802's attractive well is physically correct or that NEP802-ZBL's repulsive wall is too stiff, the paper's central claim fails. A second check is to simulate NEPstd itself at higher temperatures or longer times to see whether the residual breakdown at 1.15 Å eventually produces clustering, which would indicate the fix only postpones the same failure.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the failure mode of MLFFs in LLZO is not a global accuracy problem but a localized short-range one: NEP802, trained through conventional active learning, develops an artificial attractive well for Li$^+$–Li$^+$ pairs below about 1.32 Å, with spurious energy minima near 0.15 and 0.7 Å, exactly where unphysical RDF peaks appear. Adding the universal ZBL repulsive potential between inner and outer cutoffs of 0.9 and 1.8 Å removes the attractive well entirely, restores correct phase-transition, RDF, and MSD behavior, and allows a far smaller training set. The authors also show that active-learning iterations only slowly push the breakdown distance to shorter values, so sampling alone is an inefficient cure; the empirical repulsive wall is the load-bearing fix.
Load-bearing premise
The paper's results rest on treating the previously published NEPstd potential as a stand-in for density functional theory: it generates all training labels and is the benchmark that NEP802, NEP802-ZBL, and NEP25-ZBL are compared against, so if NEPstd is wrong in the short-range or high-energy regime, the clustering failure and its cure could be artifacts of that surrogate.
Editorial extensions
If this is right
- NEP802-ZBL reproduces the NEPstd phase-transition temperatures (around 900 K heating, 880 K cooling), RDF profiles, and Li$^+$ diffusion slopes, eliminating the sub-angstrom clustering peaks seen in NEP802.
- NEP25-ZBL, trained on 25 configurations selected by farthest-point sampling, yields diffusion slopes of 0.026 and 0.583 Å$^2$/ps at 800 and 1000 K, within a few percent of NEPstd's 0.025 and 0.575 Å$^2$/ps.
- The active-learning loop for the hybrid framework converges in 3 iterations (207 structures) instead of 13 (802 structures), with a 0% failure ratio versus 19.9% and 57.9% for the non-hybrid model in later iterations.
- The ZBL parameters and functional form are not sensitive: the paper states that any reasonable empirical short-range repulsion would serve, because the mechanism is a physical barrier in a data-sparse region.
- Because the ZBL modification acts below a cutoff and leaves the learned potential untouched at larger distances, the same strategy can be layered onto other MLFF architectures that suffer from rare-event short-range extrapolation.
Reading between the lines
- Editorial inference: if the surrogate caveat is set aside, the result suggests that for ionic conductors the dominant MLFF failure is not descriptor coverage but missing short-range physics; other physics-informed constraints, such as long-range dispersion corrections, could yield similar data-efficiency gains.
- Editorial inference: the descriptor-coverage analysis implies a concrete selection rule for future active learning: prioritize configurations that fill short-range descriptor space rather than temporal rarities, which could shorten the iterative loop further.
- Editorial inference: a direct test of the mechanism would be to replace ZBL with a simpler hard-core or Lennard-Jones repulsive wall; if the same robustness and data efficiency appear, the specific nuclear-screening form is irrelevant and any sufficiently stiff repulsion works.
- Editorial inference: the approach may transfer to radiation-damage and high-pressure simulations, where short-range excursions are frequent, making the hybrid potential's barrier the difference between a stable long trajectory and a catastrophic one.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Yan, Fan, and Zhu propose a hybrid machine-learned potential that adds an empirical ZBL short-range repulsive term to a NEP model for LLZO. The authors retrain NEP on an 802-configuration active-learning set whose labels are predicted by their previously published NEPstd potential, train a ZBL-augmented version on the same set, and also train NEP25-ZBL on 25 farthest-point-selected structures. They compare all models against NEPstd for the tetragonal-to-cubic phase transition, Li-Li RDFs, and Li MSDs, and use Li-Li dimer energy/force curves to show that NEP802 develops an artificial attractive well below about 1.32 Å, which is removed by NEP802-ZBL. The paper claims that ZBL integration eliminates unphysical clustering, reduces active-learning iterations from 13 to 3, and allows comparable performance with only 25 training configurations.
Significance. The paper's mechanistic diagnosis is valuable: the dimer scan (Fig. 2) cleanly identifies an artificial short-range attraction in a purely data-driven NEP, and the RDF peaks near 0.13 Å and 0.65 Å in Fig. 1(e) provide a concrete observable fingerprint of that failure. The demonstrated cure by the ZBL term, with MD comparisons showing restored phase-transition, RDF, and MSD behavior relative to NEPstd, is a useful practical recipe for MLFF robustness and is compatible with existing MLFF architectures. The authors also make reproducibility a strength by sharing GPUMDkit and source files. However, the quantitative data-efficiency and '13 to 3 iteration' claims are not yet established as stated, because the retraining pipeline and benchmarks are closed on the authors' own NEPstd surrogate rather than on DFT.
major comments (3)
- [II (Methods), 'Machine learning force fields training'; Fig. 2] The training labels for NEP802, NEP802-ZBL, and NEP25-ZBL are all energies, forces, and stresses predicted by NEPstd, and the benchmarks are comparisons to NEPstd. Because Fig. 2 shows NEPstd itself has an unphysical attractive branch below about 1.15 Å, matching NEPstd does not certify physical correctness at the short distances that the ZBL term is introduced to control. The manuscript should either provide direct DFT reference values for a small set of short-range or dimer configurations (and for the clustered configurations removed by ZBL) or explicitly restrict the claims to agreement with NEPstd; in its present form the statements in the abstract and conclusion that the hybrid model 'eliminates unphysical configurations' and 'ensures robust simulations' are stronger than the evidence supports.
- [III, NEP25-ZBL training and Fig. 5] The claim that a model trained on only 25 configurations matches a model trained on roughly 2000 configurations conflates distillation with data efficiency. The 25 structures were selected by farthest-point sampling from the full 802-structure NEP802 set, so the selection used global information about the explored descriptor space, and both the labels and the benchmark are NEPstd predictions. The UMAP coverage shown in Fig. 5 is then a largely expected consequence of the selection procedure and cannot rule out errors common to NEPstd. To support the training-efficiency claim, the authors should test with 25 independently generated configurations, benchmark against DFT on a handful of configurations, or explicitly reframe the result as a distillation experiment.
- [III, Table I and active-learning comparison] The reduction from 13 active-learning iterations in the NEP802 campaign to 3 iterations in Table I is not a controlled comparison: the two campaigns differ in seed structures, temperature and duration schedules, and sampling details (Table I versus Tab. S2), and the NEP column in Table I is a separate no-ZBL protocol rather than the NEP802 protocol rerun without ZBL. The table does show that the NEP-ZBL protocol avoids the failure ratios seen in the no-ZBL protocol, but the '13 to 3' statement in the abstract and conclusion should be supported by a matched-protocol comparison or removed.
minor comments (5)
- [Eq. (3)] The formula for the screening length is garbled; it should read a = 0.46848 / (Z_i^{0.23} + Z_j^{0.23}), with the division by the sum made explicit.
- [Eq. (4)] The notation for the inner and outer cutoff radii, written as r^a_c and r^b_c, is broken in the typeset equations; please define them in a clear form such as r_c^a and r_c^b.
- [Section III heading] The section heading reads 'RESUL TS' in the full text; this should be corrected to 'RESULTS'.
- [Fig. 2] The caption and text refer to blue and yellow dashed lines for breakdown distances, but the colors are not identified in the text or legend; specify which curve and color correspond to which breakdown distance.
- [Data availability] The data availability statement says 'partial source files' are on GitHub while the code availability statement presents GPUMDkit as fully available; please clarify which files are omitted and why.
Circularity Check
The central data-efficiency and accuracy claims are benchmarked against the same surrogate (NEPstd) that generated all training labels, making the '25 configurations' result a distillation check; the ZBL robustness mechanism itself is independently supported.
-
fitted input called prediction
[Methods (Machine learning force fields training) and Results first paragraph / Conclusion]
"To optimize computational efficiency, the energies, forces, and stresses for the NEP802 dataset were predicted by the validated NEPstd model using calorine package [30], thereby avoiding costly DFT calculations while maintaining accuracy. ... This NEPstd model ... served as the surrogate model for DFT and the reference point for our performance benchmark. ... the NEP25-ZBL model, trained on only 25 configurations, achieves performance comparable to the NEPstd model trained on nearly 2000 configurations."
NEP802, NEP802-ZBL, NEP25-ZBL, and NEP207-ZBL are all trained on energies, forces, and stresses generated by the authors' own NEPstd potential, and the reported benchmarks (phase transition, RDF, MSD) are NEPstd MD results. Agreement with one's own label generator is a distillation check, not an independent accuracy test. The 25 structures are farthest-point samples from the 802 NEPstd-labeled structures, so '25 configurations suffice' reduces to 'a small model reproduces its teacher after informed selection.' Moreover NEPstd itself has an acknowledged unphysical short-range attractive branch (Fig. 2), so matching NEPstd cannot certify short-range robustness.
full rationale
The qualitative ZBL mechanism is not circular: the dimer scan directly shows NEP802 and NEPstd develop attractive branches below ~1.3 Å while NEP-ZBL is monotonically repulsive, and this follows from the ZBL switch by construction rather than from fitting. The self-citation to prior work [22] is transparent and is not itself the flaw. The circular part is the closed validation loop for the quantitative claims: the same NEPstd model is used to generate the labels for all new potentials and to define the 'performance' benchmark, so the headline result that NEP25-ZBL is comparable to NEPstd is a statement about distillation fidelity and interpolation capacity, not about physical accuracy relative to DFT. The paper's own Fig. 2 passage 'NEP std also exhibits unphysical short-range Li-Li attractive interactions' concedes the reference is unreliable in exactly the short-range region the paper aims to fix. The 13- vs 3-iteration active-learning comparison is a protocol confound rather than a circularity. On balance, the central data-efficiency claim partly reduces by construction (score 6), while the robustness mechanism has independent support.
Assumptions & free parameters
free parameters (2)
- ZBL inner cutoff radius =
0.9 Å
- ZBL outer cutoff radius =
1.8 Å
assumptions (4)
- domain assumption NEPstd is an accurate surrogate for DFT for generating training labels and as reference truth.
- domain assumption The ZBL potential, developed for high-energy ion collisions, correctly describes short-range Li-Li repulsion in condensed-phase LLZO.
- domain assumption Descriptor-space coverage (UMAP/PCA) implies predictive accuracy in the covered regions.
- domain assumption The NEP-ZBL switching function smoothly connects the two potentials without introducing artifacts.
Cite this review
Pith. "Pith review of Improving robustness and training efficiency of machine-learned potentials by incorporating short-range empirical potentials." pith.science (2026). https://pith.science/paper/DKD7TR6C
@misc{pith2026250415925,
author = {Pith},
title = {Pith review of: Improving robustness and training efficiency of machine-learned potentials by incorporating short-range empirical potentials},
year = {2026},
howpublished = {\url{https://pith.science/paper/DKD7TR6C}},
note = {Machine review of arXiv:2504.15925}
}
abstract
Machine learning force fields (MLFFs) are powerful tools for materials modeling, but their performance is often limited by training dataset quality, particularly the lack of rare event configurations. This limitation undermines their accuracy and robustness in long-time and large-scale molecular dynamics simulations. In this work, we present a hybrid MLFF framework that integrates an empirical short-range repulsive potential and demonstrates improved robustness and training efficiency. Using solid electrolyte Li$_7$La$_3$Zr$_2$O$_{12}$ (LLZO) as a model system, we show that purely data-driven MLFFs fail to prevent unphysical atomistic clustering in extended simulations due to inadequate short-range repulsion. In contrast, the hybrid force field eliminates these artifacts, enabling stable long-time simulations, which are critical for studying various properties of LLZO. The hybrid framework also reduces the need for extensive active learning and performs well with just 25 training configurations. By combining physics-driven constraints with data-driven flexibility, this approach is compatible with most existing MLFF architectures and establishes a universal paradigm for developing robust, training-efficient force fields for complex material systems.
Figures
Reference graph
Works this paper leans on
-
[1]
J. Behler, Perspective: Machine learning potentials for atomistic simulations, The Journal of chemical physics 145 (2016)
work page 2016
-
[2]
V. Botu, R. Batra, J. Chapman, and R. Ramprasad, Ma- chine learning force fields: construction, validation, and outlook, J. Phys. Chem. C 121, 511 (2017)
work page 2017
-
[3]
O. T. Unke, S. Chmiela, H. E. Sauceda, M. Gastegger, I. Poltavsky, K. T. Sch¨ utt, A. Tkatchenko, and K.-R. M¨ uller, Machine learning force fields, Chem. Rev. 121, 10142 (2021)
2021
-
[4]
P. Ying, C. Qian, R. Zhao, Y. Wang, K. Xu, F. Ding, S. Chen, and Z. Fan, Advances in modeling complex materials: The rise of neuroevolution potentials, Chem. 9 Phys. Rev. 6 (2025)
work page 2025
-
[5]
A. P. Bart´ ok, M. C. Payne, R. Kondor, and G. Cs´ anyi, Gaussian approximation potentials: The accuracy of quantum mechanics, without the electrons, Phys. Rev. Lett. 104, 136403 (2010)
2010
-
[6]
A. V. Shapeev, Moment tensor potentials: A class of sys- tematically improvable interatomic potentials, Multiscale Model. Simul. 14, 1153 (2016)
work page 2016
-
[7]
I. S. Novikov, K. Gubaev, E. V. Podryabinkin, and A. V. Shapeev, The mlip package: moment tensor potentials with mpi and active learning, Mach. Learn.: Sci. Technol. 2, 025002 (2020)
work page 2020
- [8]
Show all 36 references
-
[9]
Z. Fan, Z. Zeng, C. Zhang, Y. Wang, K. Song, H. Dong, Y. Chen, and T. Ala-Nissila, Neuroevolution machine learning potentials: Combining high accuracy and low cost in atomistic simulations and application to heat transport, Phys. Rev. B 104, 104309 (2021)
2021
-
[10]
Drautz, Atomic cluster expansion for accurate and transferable interatomic potentials, Phys
R. Drautz, Atomic cluster expansion for accurate and transferable interatomic potentials, Phys. Rev. B 99, 014104 (2019)
2019
-
[11]
Batatia, D
I. Batatia, D. P. Kovacs, G. N. C. Simm, C. Ortner, and G. Csanyi, MACE: Higher order equivariant message passing neural networks for fast and accurate force fields, in Advances in Neural Information Processing Systems , edited by A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho (2022)
2022
-
[12]
Batzner, A
S. Batzner, A. Musaelian, L. Sun, M. Geiger, J. P. Mailoa, M. Kornbluth, N. Molinari, T. E. Smidt, and B. Kozinsky, E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials, Nat. Commun. 13, 2453 (2022)
2022
-
[13]
Musaelian, S
A. Musaelian, S. Batzner, A. Johansson, L. Sun, C. J. Owen, M. Kornbluth, and B. Kozinsky, Learning local equivariant representations for large-scale atomistic dy- namics, Nat. Commun. 14, 579 (2023)
2023
-
[14]
Cheng, Cartesian atomic cluster expansion for ma- chine learning interatomic potentials, npj Comput
B. Cheng, Cartesian atomic cluster expansion for ma- chine learning interatomic potentials, npj Comput. Mater. 10, 157 (2024)
2024
-
[15]
Finkbeiner, S
J. Finkbeiner, S. Tovey, and C. Holm, Generating min- imal training sets for machine learned potentials, Phys. Rev. Lett. 132, 167301 (2024)
2024
-
[16]
J. C. Bachman, S. Muy, A. Grimaud, H.-H. Chang, N. Pour, S. F. Lux, O. Paschos, F. Maglia, S. Lupart, P. Lamp, et al. , Inorganic solid-state electrolytes for lithium batteries: mechanisms and properties governing ion conduction, Chem. Rev. 116, 140 (2016)
2016
-
[17]
Manthiram, X
A. Manthiram, X. Yu, and S. Wang, Lithium battery chemistries enabled by solid-state electrolytes, Nat. Rev. Mater. 2, 1 (2017)
2017
-
[18]
Q. Zhao, S. Stalin, C.-Z. Zhao, and L. A. Archer, Design- ing solid-state electrolytes for safe, energy-dense batter- ies, Nat. Rev. Mater. 5, 229 (2020)
2020
-
[19]
Janek and W
J. Janek and W. G. Zeier, Challenges in speeding up solid-state battery development, Nat. Energy 8, 230 (2023)
2023
-
[20]
S. Wang, Y. Liu, and Y. Mo, Frustration in super-ionic conductors unraveled by the density of atomistic states, Angew. Chem. Int. Ed. 62, e202215544 (2023)
2023
-
[21]
J. Geng, Z. Yan, and Y. Zhu, Elucidating anisotropic ionic diffusion mechanism in Li3YCl6 with molecular dy- namics simulations, ACS Appl. Energy Mater. 7, 7019 (2024)
2024
-
[22]
Yan and Y
Z. Yan and Y. Zhu, Impact of lithium nonstoi- chiometry on ionic diffusion in tetragonal garnet-type Li7La3Zr2O12, Chem. Mater. 36, 11551 (2024)
2024
-
[23]
J. F. Ziegler and J. P. Biersack, The stopping and range of ions in matter, in Treatise on Heavy-Ion Science: Vol- ume 6: Astrophysics, Chemistry, and Condensed Mat- ter, edited by D. A. Bromley (Springer US, Boston, MA,
-
[24]
J. Liu, J. Byggm¨ astar, Z. Fan, P. Qian, and Y. Su, Large- scale machine-learning molecular dynamics simulation of primary radiation damage in tungsten, Phys. Rev. B108, 054312 (2023)
2023
-
[25]
K. Song, R. Zhao, J. Liu, Y. Wang, E. Lindgren, Y. Wang, S. Chen, K. Xu, T. Liang, P. Ying, et al. , General-purpose machine-learned potential for 16 ele- mental metals and their alloys, Nat. Commun. 15, 10208 (2024)
2024
-
[26]
X. He, Y. Zhu, and Y. Mo, Origin of fast ion diffusion in super-ionic conductors, Nat. Commun. 8, 15893 (2017)
2017
-
[27]
Bernstein, M
N. Bernstein, M. Johannes, and K. Hoang, Origin of the structural phase transition in Li 7La3Zr2O12, Phys. Rev. Lett. 109, 205702 (2012)
2012
-
[28]
H. Wang, X. Guo, L. Zhang, H. Wang, and J. Xue, Deep learning inter-atomic potential model for accurate irradi- ation damage simulations, Appl. Phys. Lett. 114 (2019)
2019
-
[29]
Byggm¨ astar, A
J. Byggm¨ astar, A. Hamedani, K. Nordlund, and F. Djurabekova, Machine-learning interatomic potential for radiation damage and defects in tungsten, Phys. Rev. B 100, 144105 (2019)
2019
-
[30]
Lindgren, M
E. Lindgren, M. Rahm, E. Fransson, F. Eriksson, N. ¨Osterbacka, Z. Fan, and P. Erhart, calorine: A python package for constructing and sampling neuroevolution potential models, J. Open Source Softw. 9, 6264 (2024)
2024
-
[31]
Z. Fan, W. Chen, V. Vierimaa, and A. Harju, Efficient molecular dynamics simulations with many-body poten- tials on graphics processing units, Comput. Phys. Com- mun. 218, 10 (2017)
2017
-
[32]
J. Zeng, D. Zhang, D. Lu, P. Mo, Z. Li, Y. Chen, M. Rynik, L. Huang, Z. Li, S. Shi, et al., Deepmd-kit v2: A software package for deep potential models, J. Chem. Phys. 159 (2023)
2023
-
[33]
Wang and W
Y. Wang and W. Lai, Phase transition in lithium garnet oxide ionic conductors Li7La3Zr2O12: The role of ta sub- stitution and H2O/CO2 exposure, J. Power Sources 275, 612 (2015)
2015
-
[34]
Y. Chen, E. Rangasamy, C. R. dela Cruz, C. Liang, and K. An, A study of suppressed formation of low- conductivity phases in doped Li 7La3Zr2O12 garnets by in situ neutron diffraction, J. Mater. Chem. A 3, 22868 (2015)
2015
-
[35]
Jalem, M
R. Jalem, M. Rushton, W. Manalastas Jr, M. Nakayama, T. Kasuga, J. A. Kilner, and R. W. Grimes, Effects of gallium doping in garnet-type Li 7La3Zr2O12 solid elec- trolytes, Chem. Mater. 27, 2821 (2015)
2015
-
[36]
Z. Fan, Y. Wang, P. Ying, K. Song, J. Wang, Y. Wang, Z. Zeng, K. Xu, E. Lindgren, J. M. Rahm, et al., Gpumd: A package for constructing accurate machine-learned po- tentials and performing highly efficient atomistic simula- tions, J. Chem. Phys. 157, 114801 (2022)
2022
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.