REVIEW 8 cited by
Forces are not Enough: Benchmark and Critical Evaluation for Machine Learning Force Fields with Molecular Simulations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Molecular dynamics (MD) simulation techniques are widely used for various natural science applications. Increasingly, machine learning (ML) force field (FF) models begin to replace ab-initio simulations by predicting forces directly from atomic structures. Despite significant progress in this area, such techniques are primarily benchmarked by their force/energy prediction errors, even though the practical use case would be to produce realistic MD trajectories. We aim to fill this gap by introducing a novel benchmark suite for learned MD simulation. We curate representative MD systems, including water, organic molecules, a peptide, and materials, and design evaluation metrics corresponding to the scientific objectives of respective systems. We benchmark a collection of state-of-the-art (SOTA) ML FF models and illustrate, in particular, how the commonly benchmarked force accuracy is not well aligned with relevant simulation metrics. We demonstrate when and how selected SOTA methods fail, along with offering directions for further improvement. Specifically, we identify stability as a key metric for ML models to improve. Our benchmark suite comes with a comprehensive open-source codebase for training and simulation with ML FFs to facilitate future work.
Forward citations
Cited by 8 Pith papers
-
Girsanov Reweighting for Uncertainty Propagation in Rare-Event Kinetics
Girsanov reweighting of AMS-sampled reactive trajectories propagates machine-learned interatomic potential parameter uncertainty to committor probabilities and, under extra assumptions, to reaction rates.
-
From Evaluation to Design: Using Potential Energy Surface Smoothness Metrics to Guide Machine Learning Interatomic Potential Architectures
A bond-deformation benchmark plus a force-smoothness metric is proposed to detect PES artifacts and guide MLIP architecture design, with improvements shown on a new Transformer-style model.
-
Knowledge Distillation of a Protein Language Model Yields a Foundational Implicit Solvent Model
A 45,000-parameter GNN distilled from a protein language model's secondary-structure predictions can drive MD simulations and, combined with GBn2 electrostatics, approximately reproduces explicit-solvent folding profi...
-
chemtrain-deploy: A parallel and scalable framework for machine learning potentials in million-atom MD simulations
A model-agnostic JAX-to-LAMMPS framework runs machine learning potentials in million-atom multi-GPU molecular dynamics with near-ideal strong and weak scaling.
-
Global Universal Scaling and Ultra-Small Parameterization in Machine Learning Interatomic Potentials with Super-Linearity
By rescaling atomic pair distances with element-pair-specific parameters, the authors make one shared radial function serve all elements, yielding an ultra-small machine learning interatomic potential with accuracy cl...
-
Uncertainty Quantification for Misspecified Machine Learned Interatomic Potentials
POPS uncertainty bounds from misspecification-aware parameter sampling envelop DFT reference values for a broad range of tungsten properties and for MACE-MPA-0 energies.
-
Toward Exascale AI for Science: A Scalable AI Skill for Autonomous Microkinetics Discovery
Introduces a scalable AI skill framework for autonomous microkinetics discovery that automates workflows and evaluates surrogate reliability.
-
An Iterative Framework for Generative Backmapping of Coarse Grained Proteins
A two-stage generative backmapping framework substantially improves atomistic reconstruction from ultra-coarse-grained protein representations compared to a one-step baseline.
Discussion (0). Continue with ORCID to comment.