Pith. sign in

REVIEW 9 cited by

Forces are not Enough: Benchmark and Critical Evaluation for Machine Learning Force Fields with Molecular Simulations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.07237 v2 pith:H6XEB5T2 submitted 2022-10-13 physics.comp-ph cs.LGphysics.chem-ph

classification physics.comp-phcs.LGphysics.chem-ph
keywords benchmarkforcesimulationmodelsbenchmarkedevaluationforceslearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Molecular dynamics (MD) simulation techniques are widely used for various natural science applications. Increasingly, machine learning (ML) force field (FF) models begin to replace ab-initio simulations by predicting forces directly from atomic structures. Despite significant progress in this area, such techniques are primarily benchmarked by their force/energy prediction errors, even though the practical use case would be to produce realistic MD trajectories. We aim to fill this gap by introducing a novel benchmark suite for learned MD simulation. We curate representative MD systems, including water, organic molecules, a peptide, and materials, and design evaluation metrics corresponding to the scientific objectives of respective systems. We benchmark a collection of state-of-the-art (SOTA) ML FF models and illustrate, in particular, how the commonly benchmarked force accuracy is not well aligned with relevant simulation metrics. We demonstrate when and how selected SOTA methods fail, along with offering directions for further improvement. Specifically, we identify stability as a key metric for ML models to improve. Our benchmark suite comes with a comprehensive open-source codebase for training and simulation with ML FFs to facilitate future work.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 163 citations worldwide. Full citation record

  1. Girsanov Reweighting for Uncertainty Propagation in Rare-Event Kinetics

    physics.chem-ph 2026-07 conditional novelty 6.0 of 10

    Girsanov reweighting of AMS-sampled reactive trajectories propagates machine-learned interatomic potential parameter uncertainty to committor probabilities and, under extra assumptions, to reaction rates.

  2. From Evaluation to Design: Using Potential Energy Surface Smoothness Metrics to Guide Machine Learning Interatomic Potential Architectures

    cs.LG 2026-02 conditional novelty 6.0 of 10

    A bond-deformation benchmark plus a force-smoothness metric is proposed to detect PES artifacts and guide MLIP architecture design, with improvements shown on a new Transformer-style model.

  3. Knowledge Distillation of a Protein Language Model Yields a Foundational Implicit Solvent Model

    physics.bio-ph 2026-01 conditional novelty 6.0 of 10

    A 45,000-parameter GNN distilled from a protein language model's secondary-structure predictions can drive MD simulations and, combined with GBn2 electrostatics, approximately reproduces explicit-solvent folding profi...

  4. chemtrain-deploy: A parallel and scalable framework for machine learning potentials in million-atom MD simulations

    physics.comp-ph 2025-06 conditional novelty 6.0 of 10

    A model-agnostic JAX-to-LAMMPS framework runs machine learning potentials in million-atom multi-GPU molecular dynamics with near-ideal strong and weak scaling.

  5. Global Universal Scaling and Ultra-Small Parameterization in Machine Learning Interatomic Potentials with Super-Linearity

    cond-mat.mtrl-sci 2025-02 conditional novelty 6.0 of 10

    By rescaling atomic pair distances with element-pair-specific parameters, the authors make one shared radial function serve all elements, yielding an ultra-small machine learning interatomic potential with accuracy cl...

  6. Uncertainty Quantification for Misspecified Machine Learned Interatomic Potentials

    cond-mat.mtrl-sci 2025-02 conditional novelty 6.0 of 10

    POPS uncertainty bounds from misspecification-aware parameter sampling envelop DFT reference values for a broad range of tungsten properties and for MACE-MPA-0 energies.

  7. Universal machine learning interatomic potentials poised to supplant DFT in modeling general defects in metals and random alloys

    cond-mat.mtrl-sci 2025-02 conditional novelty 6.0 of 10

    EquiformerV2 universal machine learning potentials predict energies and forces of metal and alloy defects with errors below 5 meV/atom and 100 meV/A on most benchmark datasets, approaching DFT accuracy.

  8. Toward Exascale AI for Science: A Scalable AI Skill for Autonomous Microkinetics Discovery

    cs.CE 2026-06 unverdicted novelty 5.0 of 10

    Introduces a scalable AI skill framework for autonomous microkinetics discovery that automates workflows and evaluates surrogate reliability.

  9. An Iterative Framework for Generative Backmapping of Coarse Grained Proteins

    cs.LG 2025-05 reject novelty 5.0 of 10

    A two-stage generative backmapping framework substantially improves atomistic reconstruction from ultra-coarse-grained protein representations compared to a one-step baseline.

Pith tools