Pith. sign in

REVIEW 4 major objections 5 minor 3 cited by

LAMBench: A Benchmark for Large Atomistic Models

T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read No current atomistic AI model behaves as a universal simulator, a new ten-model benchmark finds.

desk verdict A genuinely useful, open benchmark for large atomistic models whose headline gap-to-universal claim holds up, but the top-rank claim needs a leakage check given the authors' own models lead the board. read the letter →

arxiv 2504.19578 v2 pith:CMBJ4YV4 submitted 2025-04-28 physics.comp-ph cond-mat.mtrl-sci

classification physics.comp-phcond-mat.mtrl-sci
keywords largeatomisticmodelsmachinelearninginteratomicpotentialsbenchmarkpotentialenergysurfacegeneralizabilitymoleculardynamicsdensityfunctionaltheoryfoundation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to measure whether large atomistic models (LAMs) are approaching a universal potential energy surface: the single energy function that density functional theory, in principle, defines for any arrangement of nuclei. To do this it introduces LAMBench, a benchmarking system that scores ten LAMs released before August 2025 on three capabilities: generalizability to out-of-distribution systems, adaptability to property-prediction tasks, and applicability in real simulations. The headline finding is a substantial gap: no model comes close to the ideal universal surface, and domain-specific models remain more accurate inside their own domains. The paper argues that closing the gap requires multi-domain pretraining, inference-time multi-fidelity support, and conservative, differentiable models.

What carries the argument

LAMBench is a modular workflow paired with dimensionless error metrics. The central object is the ratio of a model's raw error to a dummy baseline that predicts energy solely from the chemical formula; after truncation, log-averaging over datasets, and weighting by prediction type, it yields $\bar{M}^m_{\mathrm{FF}}$ for force-field tasks and $\bar{M}^m_{\mathrm{PC}}$ for property-calculation tasks. Efficiency is measured as normalized inverse inference time, and stability as the log-scale magnitude of total-energy drift in 10 ps NVE simulations. The workflow automates job submission, result aggregation, and leaderboard updates so new models and tasks can be added.

What would settle it

Search each of the twelve test datasets against the ten models' training corpora for configurations with the same chemical composition and near-identical local environments, then re-run the force-field task after removing overlapping frames. If a top-scoring model loses its lead, the out-of-distribution claim is refuted; if the rankings are unchanged, the generalizability ranking stands.

Watch

Extended reading notes

Core claim

The paper claims that no currently released LAM approximates the universal potential energy surface well enough to serve as an out-of-the-box simulator. Running ten models without fine-tuning on twelve downstream force-field datasets across inorganic materials, catalysis, and molecules, it finds the best generalist, DPA-3.1-3M, still records a dimensionless force-field error of 0.175 and a property-calculation error of 0.322 on a scale where 0 is perfect and 1 is a chemistry-formula-only dummy. DPA-3.1-3M's lead is attributed to multi-task training on datasets spanning several domains. The same measurements show domain-specific models beating generalists on their home territory, non-conservative models drifting badly in long molecular-dynamics runs, and catalysis transition states as the clearest shared weakness.

Load-bearing premise

The ranking assumes that the twelve force-field test sets are genuinely outside every model's training data, but the paper reports no overlap check against the models' training corpora, so any hidden overlap would inflate the affected model's generalizability score.

Editorial extensions

If this is right

  • If the out-of-distribution scores are taken at face value, no LAM released before August 2025 is a drop-in universal simulator; deployments should expect accuracy losses outside each model's training domain.
  • Domain-specific models remain the accuracy ceiling in their own domains: a molecular specialist beats the best LAM on torsion and conformer energy profiles, and OC20-trained models beat all LAMs on reaction-barrier prediction.
  • Multi-task training with cross-domain data is the ingredient most strongly associated with generalizability, supporting the paper's call for more balanced training-data distributions.
  • Conservativeness and differentiability are requirements, not options: non-conservative force prediction made Orb-v2 the fastest model but unstable in NVE simulations and poor at phonon and elasticity calculations.
  • Multi-fidelity modeling is needed because LAMs trained at the PBE exchange-correlation level cannot be compared directly with CCSD(T) or RPBE references; matching task heads cut molecular property error from 0.31 to 0.10.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The out-of-distribution ranking would be confounded if any of the twelve test sets overlaps a model's training corpus; the paper reports no near-duplicate check, so the leaderboard is provisional until overlap is examined.
  • The dimensionless relative-error metric rewards models on high-variance datasets, so the headline ordering is partly a statement about metric choice rather than absolute physical accuracy.
  • If LAMBench becomes standard, it will create pressure to add transition-state and hybrid-functional molecular data to pretraining corpora, likely shifting data acquisition toward domains the benchmark exposes as weak.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces LAMBench, an open-source benchmarking system for Large Atomistic Models (LAMs), and uses it to evaluate ten LAMs released before August 1, 2025. The benchmark measures three capabilities: generalizability (force-field predictions on twelve datasets across inorganic, molecular, and catalytic domains, plus property-calculation tasks), adaptability (fine-tuning on eight Matbench regression tasks), and applicability (inference efficiency and NVE stability). Results are aggregated into dimensionless metrics normalized by a dummy model, and a leaderboard is presented in which DPA-3.1-3M ranks first in generalizability. The authors conclude that there remains a substantial gap between current LAMs and an ideal universal potential energy surface, and they argue for cross-domain training, multi-fidelity support, and conservative/differentiable models.

Significance. If the reported results hold, LAMBench is a valuable community resource: it is open-sourced, modular, and accompanied by an interactive leaderboard; the metric definitions are explicit and anchored to a transparent dummy-model baseline; and the study covers a broader range of capabilities than most existing benchmarks. The main substantive findings—that no current LAM performs as a universal simulator, that domain-specific models still beat LAMs on their own domains, and that non-conservative models trade stability for speed—are plausible and practically relevant. The strengths include machine-checkable workflow automation, clear computational details for relabeled datasets, and the inclusion of efficiency and stability alongside accuracy. However, the central comparative claims about out-of-distribution generalizability and the leaderboard ordering rest on unsupported assumptions about dataset disjointness and on metrics that are sensitive to a small number of runs.

major comments (4)
  1. [II A and Table S-2] The label 'OOD' is asserted without a leakage check. Section II A defines OOD generalizability as performance on datasets whose distribution is distinct from training data, but the manuscript reports no analysis of overlap between the twelve force-field test sets (Table I) and the training corpora of the benchmarked models (Table IV and Table S-2). This is not a formal concern only: OpenLAM contains SPICE2, Yang2023ab, OC20M, OC22, and OMat24, which overlap in chemical and catalytic space with the molecular and catalysis test sets; ANI-1x itself is a training dataset for parts of the ANI family. If DPA-3.1-3M has higher overlap with these test sets than its competitors, its leading M_FF is inflated by training coverage, and the claim of 'substantially greater generalizability' (Section II B) is confounded. The Discussion concedes that some OOD cases may become in-distribution, but no quantitative support is given. Please add per-model and per-dataset overlap or distributional-similarity checks, or qualify the OOD and ranking claims accordingly.
  2. [II B, Table II, and IV E (Eqs. 6-7)] The instability metric M_IS is dominated by a single failed NVE simulation for DPA-3.1-3M. Equation (7) averages over nine structures, and Eq. (6) assigns a penalty of 5 to any failed run; the text states that DPA-3.1-3M's relatively high value of 0.572 is due to one failure, and that using the OMat24 task head reduces the instability metric to zero. Because M_IS is a component of the leaderboard, the applicability ordering of DPA-3.1-3M relative to models with M_IS=0 is effectively an artifact of a one-in-nine event. Please report per-structure instability values in the main text, show the sensitivity of the leaderboard to removing or reweighting failed runs, and consider a metric that is less sensitive to a single failure.
  3. [Table III and Sec. II B (adaptability)] The adaptability conclusion relies on unconverged runs. The footnote to Table III states that for the MP Eform and MP Gap tasks 'the reported accuracy may not reflect the fully converged results due to insufficient training epochs.' The claim that better force-field generalizability translates into better adaptability is based on comparing DPA-3.1-3M and DPA-2.4-7M across all eight tasks; two unconverged tasks weaken that inference. Please either complete these runs, provide learning curves or convergence diagnostics, or restrict the adaptability claim to the tasks that are demonstrably converged.
  4. [Table II (general)] No uncertainty estimates or significance tests are provided for any leaderboard metric. Several adjacent entries are close (e.g., Orb-v3 M_FF=0.215 vs DPA-2.4-7M=0.241, or MACE-MPA-0=0.308 vs SevenNet-l3i5=0.326), yet the leaderboard implies an exact ordering without any indication of run-to-run or test-set variability. At a minimum, bootstrap confidence intervals over test frames or repeated evaluations would establish which differences are statistically meaningful.
minor comments (5)
  1. [Table IV] The training-set entry for MatterSim-v1-5M is listed as 'MattterSim'; this appears to be a typo for 'MatterSim'.
  2. [Section II B] The sentence comparing Matbench Discovery rankings contains 'DPA2-2.4-7M' instead of 'DPA-2.4-7M'; please fix this typo.
  3. [Section IV A] The sentence beginning 'The energies, interatomic forces and virials labels of the Lopanitsyna2023Modeling and Mazitov2024Surface data on were obtained' contains a grammatical error ('data on were obtained'); it should read 'data were obtained'.
  4. [Table II] Reporting M_IS values as 0.000 may be misleading because the metric is a log-ratio against a tolerance; readers cannot tell whether this indicates zero drift or drift below the tolerance. Please consider reporting the underlying energy-drift values, as in Table S-8, alongside the dimensionless metric.
  5. [II A] The definition of OOD as 'downstream datasets designed to address specific scientific challenges' is operational but weak; in addition to the overlap analysis requested above, a more formal statement of the distributional distance used (or a citation to one) would improve reproducibility.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: LAMBench is an empirical benchmark; author-model overlap is a conflict-of-interest caveat, not a circularity.

full rationale

LAMBench's central claims are empirical measurements, not derived quantities. The generalizability metrics (Eqs. 1-4) are defined as normalized prediction errors against external reference labels, with a dummy-model baseline; no parameter is fitted to the test labels and then reported as a prediction. The force-field and property-calculation leaderboards are obtained by zero-shot inference of frozen models on independently published datasets, and the 'substantial gap' conclusion is a direct reading of those errors. The only notable author-related issue is that DPA-3.1-3M and DPA-2.4-7M, which rank highly, are developed by overlapping authors, and the paper cites its own DPA-2 and DPA-3.1 papers when explaining their training strategy. However, those citations document training data and architecture rather than supplying the benchmark's evidence; the ranking itself is reproducible from the open-sourced toolkit and the reported raw tables. One flagged limitation is in Section III, where the authors write that 'This enhancement necessitates dynamically adjusting the force field generalizability test datasets to accommodate emerging trends in training dataset development, which may render some OOD generalizability test cases in-distribution,' acknowledging possible OOD contamination. This is a dataset-validity concern, not a circularity, because the benchmark scores are still independent measurements. No step in the paper's derivation chain equates a fitted parameter with a prediction, imports a uniqueness theorem from the authors, or smuggles an ansatz via self-citation. The central finding therefore stands on independent empirical content, with only a minor self-citation and design-overlap caveat.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The benchmark introduces no new physical entities. The main ledger entries are the hand-chosen weights in the aggregated error metrics, the arbitrary efficiency reference, and the stability tolerance calibrated on one model. The strongest unstated assumption is that the OOD test sets do not overlap with LAM training data.

free parameters (4)
  • Force-field prediction weights w_E, w_F, w_V = 0.5/0.5 without virial; 0.45/0.45/0.1 with virial
    Hand-chosen weights in Eq. (3) control how energy, force, and virial errors are combined into the leaderboard metric.
  • Property task weights = 1/6 per property in Inorganic Materials, 1/4 in Molecules, 1/5 in Catalysis
    Equal-weight choices in the property calculation metric; changing them changes rankings.
  • Efficiency reference eta_0 = 100 microseconds per atom
    Arbitrary normalization in Eq. (5) that rescales the efficiency metric.
  • Stability tolerance Phi_tol = 5e-4 eV/atom/ps
    Set as three times the statistical uncertainty of the slope from the MACE-MPA-0 model; calibrating the tolerance to one model is a fitted threshold.
assumptions (4)
  • domain assumption Born-Oppenheimer approximation defines a universal potential energy surface that LAMs can approximate.
    Section I; the benchmark's goal assumes a single universal surface exists and LAMs can be evaluated against it.
  • domain assumption DFT labels at GGA, RPBE, and CCSD(T) levels are adequate ground truth for comparing LAM performance.
    Sections II and IV; the benchmark treats these quantum chemistry labels as the reference for energy, forces, and properties.
  • domain assumption The twelve OOD test datasets are not represented in the training corpora of the ten benchmarked LAMs.
    This is load-bearing for the OOD generalizability interpretation and is not verified against training set membership.
  • ad hoc to paper Relative errors, log-averaged and truncated at the dummy model error, yield a meaningful comparability metric.
    Eqs. (1) to (4); a design choice with no external validation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LAMBench: A Benchmark for Large Atomistic Models." pith.science (2026). https://pith.science/paper/CMBJ4YV4

@misc{pith2026250419578,
  author       = {Pith},
  title        = {Pith review of: LAMBench: A Benchmark for Large Atomistic Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CMBJ4YV4}},
  note         = {Machine review of arXiv:2504.19578}
}
read the original abstract

Large Atomistic Models (LAMs) have undergone remarkable progress recently, emerging as universal or fundamental representations of the potential energy surface defined by the first-principles calculations of atomistic systems. However, our understanding of the extent to which these models achieve true universality, as well as their comparative performance across different models, remains limited. This gap is largely due to the lack of comprehensive benchmarks capable of evaluating the effectiveness of LAMs as approximations to the universal potential energy surface. In this study, we introduce LAMBench, a benchmarking system designed to evaluate LAMs in terms of their generalizability, adaptability, and applicability. These attributes are crucial for deploying LAMs as ready-to-use tools across a diverse array of scientific discovery contexts. We benchmark ten state-of-the-art LAMs released prior to August 1, 2025, using LAMBench. Our findings reveal a significant gap between the current LAMs and the ideal universal potential energy surface. They also highlight the need for incorporating cross-domain training data, supporting multi-fidelity modeling, and ensuring the models' conservativeness and differentiability. As a dynamic and extensible platform, LAMBench is intended to continuously evolve, thereby facilitating the development of robust and generalizable LAMs capable of significantly advancing scientific research. The LAMBench code is open-sourced at https://github.com/deepmodeling/lambench, and an interactive leaderboard is available at https://www.aissquare.com/openlam?tab=Benchmark.

Figures

Figures reproduced from arXiv: 2504.19578 by the authors.

Figure 1
Figure 1. FIG. 1. The schematic plot of the LAMBench benchmark. [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2. Dimensionless error metrics for generalizability tasks across different domains. (a) Dimen [PITH_FULL_IMAGE:figures/full_fig_p015_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3. Distribution of inference time, normalized by the number of atoms, measured across 900 [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pushing the limits of unconstrained machine-learned interatomic potentials

    physics.chem-ph 2026-01 conditional novelty 7.0 of 10

    Unconstrained non-equivariant and direct-force neural interatomic potentials scale to 730M parameters and match or beat equivariant state-of-the-art models on several atomistic benchmarks.

  2. VASP Plugins: Linking the Vienna ab-initio Simulation Package with Python

    cond-mat.mtrl-sci 2026-07 accept novelty 5.5 of 10

    A C++/pybind11 shared-memory plugin layer exposes VASP SCF and ionic data as NumPy arrays so Python can modify structure, forces, local potential, and occupancies in place.

  3. How Far Can You Grow? Characterizing the Extrapolation Frontier of Graph Generative Models for Materials Science

    cond-mat.mtrl-sci 2026-02 conditional novelty 5.0 of 10

    A size-resolved benchmark shows well-behaved crystal generative models degrade by ~13% beyond training radii and follow RMSD ~ N^{1/3}, but the task is deterministic and the headline 'all models' claim is not supporte...

Reference graph

Works this paper leans on

103 extracted references · 50 canonical work pages · cited by 3 Pith papers

  1. [1]

    Naveed, A

    H. Naveed, A. U. Khan, S. Qiu, M. Saqib, S. Anwar, M. Usman, N. Akhtar, N. Barnes, and A. Mian, A comprehensive overview of large language models (2024), arXiv:2307.06435 [cs.CL]

  2. [2]

    Schr¨ odinger, Quantisierung als eigenwertproblem, Annalen der physik 386, 109 (1926)

    E. Schr¨ odinger, Quantisierung als eigenwertproblem, Annalen der physik 386, 109 (1926)

  3. [3]

    Born and W

    M. Born and W. Heisenberg, Zur quantentheorie der molekeln, Original Scientific Papers Wissenschaftliche Originalarbeiten , 216 (1985)

  4. [4]

    Zhang, X

    D. Zhang, X. Liu, X. Zhang, C. Zhang, C. Cai, H. Bi, Y. Du, X. Qin, A. Peng, J. Huang, et al., Dpa-2: a large atomic model as a multi-task learner, npj Computational Materials 10, 293 (2024)

  5. [5]

    B. M. Austin, D. Y. Zubarev, and W. A. Lester Jr, Quantum monte carlo and related approaches, Chemical reviews 112, 263 (2012)

  6. [6]

    Hohenberg and W

    P. Hohenberg and W. Kohn, Inhomogeneous electron gas, Physical review 136, B864 (1964)

  7. [7]

    Kohn and L

    W. Kohn and L. J. Sham, Self-consistent equations including exchange and correlation effects, Physical review 140, A1133 (1965)

  8. [8]

    J. P. Perdew, K. Burke, and M. Ernzerhof, Generalized gradient approximation made simple, Physical review letters 77, 3865 (1996)

Show all 103 references
  1. [9]

    A. D. Becke, A new mixing of hartree-fock and local density-functional theories, Journal of chemical Physics 98, 1372 (1993)

  2. [10]

    Mardirossian and M

    N. Mardirossian and M. Head-Gordon, Thirty years of density functional theory in computa- tional chemistry: an overview and extensive assessment of 200 density functionals, Molecular physics 115, 2315 (2017)

  3. [11]

    Batatia, P

    I. Batatia, P. Benner, Y. Chiang, A. M. Elena, D. P. Kov´ acs, J. Riebesell, X. R. Advincula, M. Asta, M. Avaylon, W. J. Baldwin, et al. , A foundation model for atomistic materials chemistry, arXiv preprint arXiv:2401.00096 (2023)

  4. [12]

    Y. Park, J. Kim, S. Hwang, and S. Han, Scalable parallel algorithm for graph neural network interatomic potentials in molecular dynamics simulations, Journal of chemical theory and computation 20, 4857 (2024)

  5. [13]

    B. Deng, P. Zhong, K. Jun, J. Riebesell, K. Han, C. J. Bartel, and G. Ceder, Chgnet as a pretrained universal neural network potential for charge-informed atomistic modelling, 45 Nature Machine Intelligence 5, 1031 (2023)

  6. [14]

    Zubatyuk, J

    R. Zubatyuk, J. S. Smith, J. Leszczynski, and O. Isayev, Accurate and transferable multi- task prediction of chemical properties with an atoms-in-molecules neural network, Science Advances 5, eaav6490 (2019), https://www.science.org/doi/pdf/10.1126/sciadv.aav6490

  7. [15]

    Eastman, B

    P. Eastman, B. P. Pritchard, J. D. Chodera, and T. E. Markland, Nutmeg and spice: Models and data for biomolecular machine learning (2024), arXiv:2406.13112 [physics.chem-ph]

  8. [16]

    Shoghi, A

    N. Shoghi, A. Kolluru, J. R. Kitchin, Z. W. Ulissi, C. L. Zitnick, and B. M. Wood, From molecules to materials: Pre-training large generalizable models for atomic property prediction (2024), arXiv:2310.16802 [cs.LG]

  9. [17]

    B. M. Wood, M. Dzamba, X. Fu, M. Gao, M. Shuaibi, L. Barroso-Luque, K. Abdelmaqsoud, V. Gharakhanyan, J. R. Kitchin, D. S. Levine, K. Michel, A. Sriram, T. Cohen, A. Das, A. Rizvi, S. J. Sahoo, Z. W. Ulissi, and C. L. Zitnick, Uma: A family of universal models for atoms (2025)...

  10. [18]

    Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, T. Li, M. Ku, K. Wang, A. Zhuang, R. Fan, X. Yue, and W. Chen, Mmlu-pro: A more ro- bust and challenging multi-task language understanding benchmark (2024), arXiv:2406.01574 [cs.CL]

  11. [19]

    Lightman, V

    H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe, Let’s verify step by step, arXiv preprint arXiv:2305.20050 (2023)

  12. [20]

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, ImageNet: A Large-Scale Hierarchical Image Database, in CVPR09 (2009)

  13. [21]

    Kryshtafovych, T

    A. Kryshtafovych, T. Schwede, M. Topf, K. Fidelis, and J. Moult, Critical assessment of methods of protein structure prediction (casp)—round xiii, Proteins: Structure, Function, and Bioinformatics 87, 1011 (2019)

  14. [22]

    Jumper, R

    J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvu- nakool, R. Bates, A. ˇZ´ ıdek, A. Potapenko,et al., Highly accurate protein structure prediction with alphafold, nature 596, 583 (2021)

  15. [23]

    Ramakrishnan, P

    R. Ramakrishnan, P. O. Dral, M. Rupp, and O. A. Von Lilienfeld, Quantum chemistry structures and properties of 134 kilo molecules, Scientific data 1, 1 (2014)

  16. [24]

    Chmiela, A

    S. Chmiela, A. Tkatchenko, H. E. Sauceda, I. Poltavsky, K. T. Sch¨ utt, and K.-R. M¨ uller, Machine learning of accurate energy-conserving molecular force fields, Science advances 3, 46 e1603015 (2017)

  17. [25]

    Riebesell, R

    J. Riebesell, R. E. A. Goodall, P. Benner, Y. Chiang, B. Deng, G. Ceder, M. Asta, A. A. Lee, A. Jain, and K. A. Persson, Matbench discovery – a framework to evaluate machine learning crystal stability predictions (2024), arXiv:2308.14920 [cond-mat.mtrl-sci]

  18. [26]

    Chanussot*, A

    L. Chanussot*, A. Das*, S. Goyal*, T. Lavril*, M. Shuaibi*, M. Riviere, K. Tran, J. Heras- Domingo, C. Ho, W. Hu, A. Palizhati, A. Sriram, B. Wood, J. Yoon, D. Parikh, C. L. Zitnick, and Z. Ulissi, Open catalyst 2020 (oc20) dataset and community challenges, ACS Catalysis 10.10...

  19. [27]

    X. Fu, B. M. Wood, L. Barroso-Luque, D. S. Levine, M. Gao, M. Dzamba, and C. L. Zit- nick, Learning smooth and expressive interatomic potentials for physical property prediction (2025), arXiv:2502.12147 [physics.comp-ph]

  20. [28]

    Neumann, J

    M. Neumann, J. Gin, B. Rhodes, S. Bennett, Z. Li, H. Choubisa, A. Hussey, and J. Godwin, Orb: A fast, scalable neural network potential, arXiv preprint arXiv:2410.22570 (2024)

  21. [29]

    X. Fu, Z. Wu, W. Wang, T. Xie, S. Keten, R. Gomez-Bombarelli, and T. Jaakkola, Forces are not enough: Benchmark and critical evaluation for machine learning force fields with molec- ular simulations, Transactions on Machine Learning Research (2023), survey Certification

  22. [30]

    Chiang, T

    Y. Chiang, T. Kreiman, E. Weaver, I. Amin, M. Kuner, C. Zhang, A. Kaplan, D. Chrzan, S. M. Blau, A. S. Krishnapriyan, and M. Asta, MLIP arena: Advancing fairness and trans- parency in machine learning interatomic potentials through an open and accessible bench- mark platform, ...

  23. [31]

    The UMA model license (2025), https://huggingface.co/facebook/UMA#license

  24. [32]

    J. D. Morrow, J. L. A. Gardner, and V. L. Deringer, How to validate machine-learned inter- atomic potentials, The Journal of Chemical Physics 158, 10.1063/5.0139611 (2023)

  25. [33]

    B. Deng, Y. Choi, P. Zhong, J. Riebesell, S. Anand, Z. Li, K. Jun, K. A. Persson, and G. Ceder, Systematic softening in universal machine learning interatomic potentials, npj Computational Materials 11, 9 (2025)

  26. [34]

    K. Li, A. N. Rubungo, X. Lei, D. Persaud, K. Choudhary, B. DeCost, A. B. Dieng, and J. Hattrick-Simpers, Probing out-of-distribution generalization in machine learning for ma- terials (2024), arXiv:2406.06489 [cond-mat.mtrl-sci]

  27. [35]

    Mazitov, M

    A. Mazitov, M. A. Springer, N. Lopanitsyna, G. Fraux, S. De, and M. Ceriotti, Surface segregation in high-entropy alloys from alchemical machine learning, Journal of Physics: 47 Materials 7, 025007 (2024)

  28. [36]

    Lopanitsyna, G

    N. Lopanitsyna, G. Fraux, M. A. Springer, S. De, and M. Ceriotti, Modeling high-entropy transition metal alloys with alchemical compression, Phys. Rev. Mater. 7, 045802 (2023)

  29. [37]

    Torres, F

    A. Torres, F. Luque, J. Tortajada, and M. Arroyo-de Dompablo, Analysis of minerals as electrode materials for ca-based rechargeable batteries, Scientific Reports 9, 9644 (2019)

  30. [38]

    T. G. Sours and A. R. Kulkarni, Predicting structural properties of pure silica zeolites using deep neural network potentials, The Journal of Physical Chemistry C 127, 1455 (2023), https://doi.org/10.1021/acs.jpcc.2c08429

  31. [39]

    Batzner, A

    S. Batzner, A. Musaelian, L. Sun, et al. , E(3)-equivariant graph neural networks for data- efficient and accurate interatomic potentials, Nature Communications 13, 2453 (2022), re- ceived: 15 February 2021; Accepted: 07 April 2022; Published: 04 May 2022

  32. [40]

    Y. Gao, F. Deng, R. He, and Z. Zhong, Spontaneous curvature in two-dimensional van der waals heterostructures, Nature Communications 16, 717 (2025)

  33. [41]

    J. S. Smith, B. Nebgen, N. Lubbers, O. Isayev, and A. E. Roitberg, Less is more: Sampling chemical space with active learning, The Journal of Chemical Physics 148, 241733 (2018), https://pubs.aip.org/aip/jcp/article- pdf/doi/10.1063/1.5023802/16656391/241733 1 online.pdf

  34. [42]

    Chmiela, V

    S. Chmiela, V. Vassilev-Galindo, O. T. Unke, A. Kabylda, H. E. Sauceda, A. Tkatchenko, and K.-R. M¨ uller, Accurate global machine learning force fields for molecules with hundreds of atoms, Science Advances 9, eadf0873 (2023), https://www.science.org/doi/pdf/10.1126/sciadv.adf0873

  35. [43]

    T. Wang, X. He, M. Li, et al., Aimd-chig: Exploring the conformational space of a 166-atom protein chignolin with ab initio molecular dynamics, Scientific Data 10, 549 (2023), received: 13 February 2023; Accepted: 11 August 2023; Published: 22 August 2023

  36. [44]

    Zhang, X

    Y. Zhang, X. Zhou, and B. Jiang, Bridging the gap between direct dynamics and globally accurate reactive potential energy surfaces using neural networks, The Journal of Physical Chemistry Letters 10, 1185 (2019)

  37. [45]

    Vandermause, Y

    J. Vandermause, Y. Xie, J. Lim, et al. , Active learning of reactive bayesian force fields ap- plied to heterogeneous catalysis dynamics of h/pt, Nature Communications 13, 5183 (2022), received: 16 December 2021; Accepted: 21 July 2022; Published: 02 September 2022. 48

  38. [46]

    Fern´ andez-Villanueva, P

    E. Fern´ andez-Villanueva, P. G. Lustemberg, M. Zhao, J. Soriano Rodriguez, P. Concepci´ on, and M. V. Ganduglia-Pirovano, Water and cu+ synergy in selective co2 hydrogenation to methanol over cu-mgo-al2o3 catalysts, Journal of the American Chemical Society 146, 2024 (2024), p...

  39. [47]

    Grimme, J

    S. Grimme, J. Antony, S. Ehrlich, and H. Krieg, A consistent and accurate ab initio parametrization of density functional dispersion correction (dft-d) for the 94 elements h-pu, The Journal of Chemical Physics 132, 154104 (2010), https://pubs.aip.org/aip/jcp/article- pdf/doi/1...

  40. [48]

    Grimme, S

    S. Grimme, S. Ehrlich, and L. Goerigk, Effect of the damping function in dispersion cor- rected density functional theory, Journal of Computational Chemistry 32, 1456 (2011), https://onlinelibrary.wiley.com/doi/pdf/10.1002/jcc.21759

  41. [49]

    A. Loew, D. Sun, H.-C. Wang, S. Botti, and M. A. Marques, Universal machine learning interatomic potentials are ready for phonons, arXiv preprint arXiv:2412.16551 (2024)

  42. [50]

    de Jong, W

    M. de Jong, W. Chen, T. Angsten, A. Jain, R. Notestine, A. Gamst, M. Sluiter, C. Kr- ishna Ande, S. van der Zwaag, J. J. Plata, C. Toher, S. Curtarolo, G. Ceder, K. A. Persson, and M. Asta, Charting the complete elastic properties of inorganic crystalline compounds, Scientific...

  43. [51]

    R. Liu, E. Liu, J. Riebesell, J. Qi, S. P. Ong, and T. W. Ko, MatCalc (2024)

  44. [52]

    B. Rai, V. Sresht, Q. Yang, R. J. Unwalla, M. Tu, A. M. Mathiowetz, et al. , Torsionnet: A deep neural network to rapidly predict small molecule torsion energy profiles with the accuracy of quantum mechanics, ChemRxiv 10.26434/chemrxiv.13483185.v1 (2020)

  45. [53]

    R. R. Brew, I. A. Nelson, M. Binayeva, A. S. Nayak, W. J. Simmons, J. J. Gair, and C. C. Wagen, Wiggle150: Benchmarking density functionals and neural network potentials on highly strained conformers, Journal of Chemical Theory and Computation21, 3922 (2025)

  46. [54]

    Wander, M

    B. Wander, M. Shuaibi, J. R. Kitchin, Z. W. Ulissi, and C. L. Zitnick, Cattsunami: Ac- celerating transition state energy calculations with pretrained graph neural networks, ACS Catalysis 15, 5283 (2025)

  47. [55]

    J. Zeng, D. Zhang, A. Peng, X. Zhang, S. He, Y. Wang, X. Liu, H. Bi, Y. Li, C. Cai, C. Zhang, Y. Du, J.-X. Zhu, P. Mo, Z. Huang, Q. Zeng, S. Shi, X. Qin, Z. Yu, C. Luo, Y. Ding, Y.-P. Liu, R. Shi, Z. Wang, S. L. Bore, J. Chang, Z. Deng, Z. Ding, S. Han, W. Jiang, G. Ke, Z. Liu...

  48. [56]

    A. Dunn, Q. Wang, A. Ganose, D. Dopp, and A. Jain, Benchmarking materials property pre- diction methods: The matbench test set and automatminer reference algorithm, npj Com- putational Materials 6, 138 (2020)

  49. [57]

    D. L. Lynch, A. Pavlova, Z. Fan, and J. C. Gumbart, Understanding virus structure and dynamics through molecular simulations, Journal of Chemical Theory and Computation 19, 3025 (2023)

  50. [58]

    A. Jain, S. P. Ong, G. Hautier, W. Chen, W. D. Richards, S. Dacek, S. Cholia, D. Gunter, D. Skinner, G. Ceder, and K. a. Persson, The Materials Project: A materials genome ap- proach to accelerating materials innovation, APL Materials 1, 011002 (2013)

  51. [59]

    Eastman, P

    P. Eastman, P. K. Behara, D. Dotson, R. Galvelis, J. Herr, J. Horton, Y. Mao, J. Chodera, B. Pritchard, Y. Wang, G. De Fabritiis, and T. Markland, Spice 2.0.1 (2024)

  52. [60]

    C. L. Zitnick, L. Chanussot, A. Das, S. Goyal, J. Heras-Domingo, C. Ho, W. Hu, T. Lavril, A. Palizhati, M. Riviere, M. Shuaibi, A. Sriram, K. Tran, B. Wood, J. Yoon, D. Parikh, and Z. Ulissi, An introduction to electrocatalyst design using machine learning for renewable energy...

  53. [61]

    Zhang, A

    D. Zhang, A. Peng, C. Cai, W. Li, Y. Zhou, J. Zeng, M. Guo, C. Zhang, B. Li, H. Jiang, T. Zhu, W. Jia, L. Zhang, and H. Wang, A graph neural network for the era of large atomistic models (2025), arXiv:2506.01686 [physics.comp-ph]

  54. [62]

    D. S. Levine, M. Shuaibi, E. W. C. Spotte-Smith, M. G. Taylor, M. R. Hasyim, K. Michel, I. Batatia, G. Cs´ anyi, M. Dzamba, P. Eastman, N. C. Frey, X. Fu, V. Gharakhanyan, A. S. Krishnapriyan, J. A. Rackers, S. Raja, A. Rizvi, A. S. Rosen, Z. Ulissi, S. Vargas, C. L. Zitnick, ...

  55. [63]

    D. P. Kov´ acs, J. H. Moore, N. J. Browning, I. Batatia, J. T. Horton, V. Kapil, W. C. Witt, I.-B. Magd˘ au, D. J. Cole, and G. Cs´ anyi, Mace-off23: Transferable machine learning force fields for organic molecules (2023), arXiv:2312.15211

  56. [64]

    Y.-L. Liao, B. Wood, A. Das, and T. Smidt, Equiformerv2: Improved equivariant transformer for scaling to higher-degree representations (2024), arXiv:2306.12059 [cs.LG]. 50

  57. [65]

    Chanussot, A

    L. Chanussot, A. Das, S. Goyal, T. Lavril, M. Shuaibi, M. Riviere, K. Tran, J. Heras- Domingo, C. Ho, W. Hu, et al., Open catalyst 2020 (oc20) dataset and community challenges, Acs Catalysis 11, 6059 (2021)

  58. [66]

    Passaro and C

    S. Passaro and C. L. Zitnick, Reducing so(3) convolutions to so(2) for efficient equivariant gnns (2023), arXiv:2302.03655 [cs.LG]

  59. [67]

    H. Yang, C. Hu, Y. Zhou, X. Liu, Y. Shi, J. Li, G. Li, Z. Chen, S. Chen, C. Zeni, M. Horton, R. Pinsler, A. Fowler, D. Z¨ ugner, T. Xie, J. Smith, L. Sun, Q. Wang, L. Kong, C. Liu, H. Hao, and Z. Lu, Mattersim: A deep learning atomistic model across elements, temperatures and ...

  60. [68]

    J. Kim, J. Kim, J. Kim, J. Lee, Y. Park, Y. Kang, and S. Han, Data-efficient multifidelity training for high-fidelity machine learning interatomic potentials, J. Am. Chem. Soc. 147, 1042 (2024)

  61. [69]

    J. P. Perdew, K. Burke, and M. Ernzerhof, Generalized gradient approximation made simple, Phys. Rev. Lett. 77, 3865 (1996)

  62. [70]

    J. P. Perdew, K. Burke, and M. Ernzerhof, Generalized gradient approximation made simple [phys. rev. lett. 77, 3865 (1996)], Phys. Rev. Lett. 78, 1396 (1997)

  63. [71]

    M. J. Frisch, G. W. Trucks, H. B. Schlegel, G. E. Scuseria, M. A. Robb, J. R. Cheeseman, G. Scalmani, V. Barone, G. A. Petersson, H. Nakatsuji, X. Li, M. Caricato, A. V. Marenich, J. Bloino, B. G. Janesko, R. Gomperts, B. Mennucci, H. P. Hratchian, J. V. Ortiz, A. F. Izmaylov,...

  64. [72]

    Barroso-Luque, M

    L. Barroso-Luque, M. Shuaibi, X. Fu, B. M. Wood, M. Dzamba, M. Gao, A. Rizvi, C. L. Zitnick, and Z. W. Ulissi, Open materials 2024 (omat24) inorganic materials dataset and models (2024), arXiv:2410.12771 [cond-mat.mtrl-sci]. 51

  65. [73]

    Rhodes, S

    B. Rhodes, S. Vandenhaute, V. ˇSimkus, J. Gin, J. Godwin, T. Duignan, and M. Neumann, Orb-v3: atomistic simulation at scale (2025), arXiv:2504.06231 [cond-mat.mtrl-sci]

  66. [74]

    Bochkarev, Y

    A. Bochkarev, Y. Lysogorskiy, and R. Drautz, Graph atomic cluster expansion for semilocal interactions beyond equivariant message passing, Phys. Rev. X 14, 021036 (2024)

  67. [75]

    A. H. Larsen, J. J. Mortensen, J. Blomqvist, I. E. Castelli, R. Christensen, M. Du lak, J. Friis, M. N. Groves, B. Hammer, C. Hargus, et al., The atomic simulation environment—a python library for working with atoms, Journal of Physics: Condensed Matter 29, 273002 (2017)

  68. [76]

    X. Liu, Y. Han, Z. Li, J. Fan, C. Zhang, J. Zeng, Y. Shan, Y. Yuan, W.-H. Xu, Y.-P. Liu, et al., Dflow, a python framework for constructing cloud-native ai-for-science workflows, arXiv preprint arXiv:2404.18392 (2024)

  69. [77]

    H.-C. Wang, J. Schmidt, M. A. Marques, L. Wirtz, and A. H. Romero, Symmetry-based computational search for novel binary and ternary 2d materials, 2D Materials 10, 035007 (2023)

  70. [78]

    R. Tran, J. Lan, M. Shuaibi, B. M. Wood, S. Goyal, A. Das, J. Heras-Domingo, A. Kolluru, A. Rizvi, N. Shoghi, A. Sriram, F. Therrien, J. Abed, O. Voznyy, E. H. Sargent, Z. Ulissi, and C. L. Zitnick, The open catalyst 2022 (oc22) dataset and challenges for oxide electrocatalyst...

  71. [79]

    Sriram, S

    A. Sriram, S. Choi, X. Yu, L. M. Brabson, A. Das, Z. Ulissi, M. Uyttendaele, A. J. Medford, and D. S. Sholl, The open dac 2023 dataset and challenges for sorbent discovery in direct air capture (2024)

  72. [80]

    Eastman, B

    P. Eastman, B. P. Pritchard, J. D. Chodera, and T. E. Markland, Nutmeg and spice: models and data for biomolecular machine learning, Journal of chemical theory and computation 20, 8583 (2024)

  73. [81]

    Schreiner, A

    M. Schreiner, A. Bhowmik, T. Vegge, J. Busk, and O. Winther, Transition1x-a dataset for building generalizable reactive machine learning potentials, Scientific Data 9, 779 (2022)

  74. [82]

    J. Wu, J. Yang, Y.-J. Liu, D. Zhang, Y. Yang, Y. Zhang, L. Zhang, and S. Liu, Universal interatomic potential for perovskite oxides, Physical Review B 108, L180104 (2023)

  75. [83]

    Dai and W

    F. Dai and W. Jiang, Alloy dpa v1 0, https://aissquare.com/datasets/detail? pageType=datasets&name=Alloy_DPA_v1_0&id=147 (2023), accessed: 2025-04-07

  76. [84]

    Zhang and J

    L. Zhang and J. Liu, Cathode(anode) dpa v1 0, https://aissquare.com/datasets/ detail?pageType=datasets&name=Cathode%28Anode%29_DPA_v1_0&id=130 (2023), ac- 52 cessed: 2025-04-07

  77. [85]

    Gong, Cluster dpa v1 0, https://aissquare.com/datasets/detail?pageType= datasets&name=Cluster_DPA_v1_0&id=131 (2023), accessed: 2025-04-07

    F. Gong, Cluster dpa v1 0, https://aissquare.com/datasets/detail?pageType= datasets&name=Cluster_DPA_v1_0&id=131 (2023), accessed: 2025-04-07

  78. [86]

    Z. Li, T. Wen, Y. Zhang, X. Liu, C. Zhang, A. S. Pattamatta, X. Gong, B. Ye, H. Wang, L. Zhang, et al. , Apex: an automated cloud-native material property explorer, npj Compu- tational Materials 11, 88 (2025)

  79. [87]

    Shi and Y

    M. Shi and Y. Zhang, Electrolyte, https://www.aissquare.com/datasets/detail?name= Electrolyte&id=216&pageType=datasets (2023), accessed: 2025-04-07

  80. [88]

    M. Shi, R. Wang, and Y. Gao, Sse-abacus, https://aissquare.com/datasets/detail? pageType=datasets&name=SSE-abacus&id=260 (2024), accessed: 2025-04-07

  81. [89]

    M. Yang, D. Zhang, X. Wang, L. Zhang, T. Zhu, and H. Wang, Ab initio accuracy neural network potential for drug-like molecules, ChemRxiv (2024), this content is a preprint and has not been peer-reviewed

  82. [90]

    B. L. et al., General reactive machine learning potentials for chon elements (2025), in prepa- ration

  83. [91]

    Huang, L

    J. Huang, L. Zhang, H. Wang, J. Zhao, J. Cheng, et al., Deep potential generation scheme and simulation protocol for the li10gep2s12-type superionic conductors, The Journal of Chemical Physics 154 (2021)

  84. [92]

    J. Liu, X. Zhang, T. Chen, Y. Zhang, D. Zhang, L. Zhang, and M. Chen, Machine-learning- based interatomic potentials for group iib to via semiconductors: Toward a universal model, Journal of Chemical Theory and Computation 20, 5717 (2024)

  85. [93]

    Zhang, H

    L. Zhang, H. Wang, R. Car, and W. E, Phase diagram of a deep potential water model, Physical review letters 126, 236001 (2021)

  86. [94]

    Jiang, Y

    W. Jiang, Y. Zhang, L. Zhang, and H. Wang, Accurate deep potential model for the al–cu–mg alloy in the full concentration space, Chinese Physics B 30, 050706 (2021)

  87. [95]

    T. Chen, F. Yuan, J. Liu, H. Geng, L. Zhang, H. Wang, and M. Chen, Modeling the high- pressure solid and liquid phases of tin from deep potentials with ab initio accuracy, Physical Review Materials 7, 053603 (2023)

  88. [96]

    O. T. Unke and M. Meuwly, Physnet: A neural network for predicting energies, forces, dipole moments, and partial charges, Journal of chemical theory and computation 15, 3678 (2019). 53

  89. [97]

    T. Wen, R. Wang, L. Zhu, L. Zhang, H. Wang, D. J. Srolovitz, and Z. Wu, Specialising neural network potentials for accurate properties and application to the mechanical response of titanium, npj Computational Materials 7, 206 (2021)

  90. [98]

    R. Wang, X. Ma, L. Zhang, H. Wang, D. J. Srolovitz, T. Wen, and Z. Wu, Classical and machine learning interatomic potentials for bcc vanadium, Physical Review Materials 6, 113603 (2022)

  91. [99]

    X. Wang, Y. Wang, L. Zhang, F. Dai, and H. Wang, A tungsten deep neural-network potential for simulating mechanical property degradation under fusion service environment, Nuclear Fusion 62, 126013 (2022)

  92. [100]

    J. Wu, Y. Zhang, L. Zhang, and S. Liu, Deep learning of accurate force field of ferroelectric hfo 2, Physical Review B 103, 024108 (2021)

  93. [101]

    Y. Wang, L. Zhang, B. Xu, X. Wang, and H. Wang, A generalizable machine learning po- tential of ag–au nanoalloys and its application to surface reconstruction, segregation and diffusion, Modelling and Simulation in Materials Science and Engineering 30, 025003 (2021)

  94. [102]

    J. Wu, L. Bai, J. Huang, L. Ma, J. Liu, and S. Liu, Accurate force field of two-dimensional ferroelectrics from deep learning, Physical Review B 104, 174107 (2021)

  95. [103]

    P. Tuo, L. Li, X. Wang, J. Chen, Z. Zhong, B. Xu, and F.-Z. Dai, Spontaneous hybrid nano- domain behavior of the organic–inorganic hybrid perovskites, Advanced Functional Materials 33, 2301663 (2023). 54

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.