REVIEW 5 cited by
Towards Fast, Specialized Machine Learning Force Fields: Distilling Foundation Models via Energy Hessians
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
The foundation model (FM) paradigm is transforming Machine Learning Force Fields (MLFFs), leveraging general-purpose representations and scalable training to perform a variety of computational chemistry tasks. Although MLFF FMs have begun to close the accuracy gap relative to first-principles methods, there is still a strong need for faster inference speed. Additionally, while research is increasingly focused on general-purpose models which transfer across chemical space, practitioners typically only study a small subset of systems at a given time. This underscores the need for fast, specialized MLFFs relevant to specific downstream applications, which preserve test-time physical soundness while maintaining train-time scalability. In this work, we introduce a method for transferring general-purpose representations from MLFF foundation models to smaller, faster MLFFs specialized to specific regions of chemical space. We formulate our approach as a knowledge distillation procedure, where the smaller "student" MLFF is trained to match the Hessians of the energy predictions of the "teacher" foundation model. Our specialized MLFFs can be up to 20 $\times$ faster than the original foundation model, while retaining, and in some cases exceeding, its performance and that of undistilled models. We also show that distilling from a teacher model with a direct force parameterization into a student model trained with conservative forces (i.e., computed as derivatives of the potential energy) successfully leverages the representations from the large-scale teacher for improved accuracy, while maintaining energy conservation during test-time molecular dynamics simulations. More broadly, our work suggests a new paradigm for MLFF development, in which foundation models are released along with smaller, specialized simulation "engines" for common chemical subsets.
Forward citations
Cited by 5 Pith papers
-
Distillation of atomistic foundation models across architectures and chemical domains
Distillation of atomistic foundation models via synthetic data yields 10x-100x faster student potentials with near-teacher accuracy.
-
Teacher-student training improves accuracy and efficiency of machine learning interatomic potentials
A teacher-student training method that transfers per-atom energy knowledge makes smaller machine learning interatomic potentials faster, smaller, and comparable or better in training-set accuracy.
-
Universal machine learning interatomic potentials poised to supplant DFT in modeling general defects in metals and random alloys
EquiformerV2 universal machine learning potentials predict energies and forces of metal and alloy defects with errors below 5 meV/atom and 100 meV/A on most benchmark datasets, approaching DFT accuracy.
-
Fine-Tuning Universal Machine-Learned Interatomic Potentials: A Tutorial on Methods and Applications
Fine-tuning universal MLIPs improves accuracy and data efficiency across electrolytes, defects, and interfaces, with some evidence of implicit long-range behavior that is not conclusive.
-
Energy & Force Regression on DFT Trajectories is Not Enough for Universal Machine Learning Interatomic Potentials
A perspective arguing that current MLIP training on DFT data is insufficient, and proposing CCSD(T)-quality data, metrology, and efficient inference as research priorities.
Discussion (0). Continue with ORCID to comment.