REVIEW 12 cited by
Matbench Discovery -- A framework to evaluate machine learning crystal stability predictions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The rapid adoption of machine learning (ML) in domain sciences necessitates best practices and standardized benchmarking for performance evaluation. We present Matbench Discovery, an evaluation framework for ML energy models, applied as pre-filters for high-throughput searches of stable inorganic crystals. This framework addresses the disconnect between thermodynamic stability and formation energy, as well as retrospective vs. prospective benchmarking in materials discovery. We release a Python package to support model submissions and maintain an online leaderboard, offering insights into performance trade-offs. To identify the best-performing ML methodologies for materials discovery, we benchmarked various approaches, including random forests, graph neural networks (GNNs), one-shot predictors, iterative Bayesian optimizers, and universal interatomic potentials (UIP). Our initial results rank models by test set F1 scores for thermodynamic stability prediction: EquiformerV2 + DeNS > Orb > SevenNet > MACE > CHGNet > M3GNet > ALIGNN > MEGNet > CGCNN > CGCNN+P > Wrenformer > BOWSR > Voronoi fingerprint random forest. UIPs emerge as the top performers, achieving F1 scores of 0.57-0.82 and discovery acceleration factors (DAF) of up to 6x on the first 10k stable predictions compared to random selection. We also identify a misalignment between regression metrics and task-relevant classification metrics. Accurate regressors can yield high false-positive rates near the decision boundary at 0 eV/atom above the convex hull. Our results demonstrate UIPs' ability to optimize computational budget allocation for expanding materials databases. However, their limitations remain underexplored in traditional benchmarks. We advocate for task-based evaluation frameworks, as implemented here, to address these limitations and advance ML-guided materials discovery.
Forward citations
Cited by 12 Pith papers
-
Pushing the limits of unconstrained machine-learned interatomic potentials
Unconstrained non-equivariant and direct-force neural interatomic potentials scale to 730M parameters and match or beat equivariant state-of-the-art models on several atomistic benchmarks.
-
MP-ALOE: An r2SCAN dataset for universal machine learning interatomic potentials
MP-ALOE provides 909,792 r2SCAN DFT frames of mostly off-equilibrium structures; a MACE potential trained on it improves molecular dynamics stability and pressure robustness over MatPES-trained models.
-
From Evaluation to Design: Using Potential Energy Surface Smoothness Metrics to Guide Machine Learning Interatomic Potential Architectures
A bond-deformation benchmark plus a force-smoothness metric is proposed to detect PES artifacts and guide MLIP architecture design, with improvements shown on a new Transformer-style model.
-
MiAD: Mirage Atom Diffusion for De Novo Crystal Generation
Mirage infusion lets crystal diffusion models vary atom counts during generation and raises the S.U.N. rate on MP-20 to 8.2%.
-
OpenCSP: A Deep Learning Framework for Crystal Structure Prediction from Ambient to High Pressure
OpenCSP is an open pressure-diverse dataset and model suite that matches or beats larger universal atomistic models on high-pressure crystal structure prediction with far fewer training data.
-
Universal Machine Learning Potentials under Pressure
Universal machine learning interatomic potentials systematically lose accuracy under pressure up to 150 GPa, and fine-tuning on high-pressure DFT data recovers most of the lost performance.
-
Universal Machine Learning Potential for Systems with Reduced Dimensionality
Benchmarking 11 universal machine learning interatomic potentials on a new 40,000-structure, 0D-3D dataset shows energy and geometry errors grow as dimensionality falls, with eSEN the most transferable.
-
Enhancing Materials Discovery with Valence Constrained Design in Generative Modeling
CrysVCD generates valence-balanced compositions with an elemental language model and then constructs their crystal structures with a diffusion model, reporting improved stability and functional property targeting.
-
Comparing classical and machine learning force fields for modeling deformation of solid sorbents relevant for direct air capture
No tested force field, classical or machine-learned, achieves the 0.1 eV accuracy target for adsorption energies in deformable MOFs relevant to direct air capture.
-
Benchmarking Universal Machine Learning Interatomic Potentials for Real-Time Analysis of Inelastic Neutron Scattering Data
Several universal machine learning interatomic potentials, especially ORB v3, MatterSim, and MACE-OFF, reach near-DFT phonon accuracy and match many experimental neutron spectra, with important caveats about test-set ...
-
Toward Exascale AI for Science: A Scalable AI Skill for Autonomous Microkinetics Discovery
Introduces a scalable AI skill framework for autonomous microkinetics discovery that automates workflows and evaluates surrogate reliability.
-
Toward Greater Autonomy in Materials Discovery Agents: Unifying Planning, Physics, and Scientists
MAPPS combines LLM workflow planning, code generation, and human intuition with machine-learned force fields to discover crystal structures, reporting high stability and novelty rates on MP-20 and Matbench.
Discussion (0). Sign in to comment.