Pith. sign in

REVIEW 12 cited by

Matbench Discovery -- A framework to evaluate machine learning crystal stability predictions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.14920 v3 pith:3HOYXYA7 submitted 2023-08-28 cond-mat.mtrl-sci cs.LG

classification cond-mat.mtrl-scics.LG
keywords discoverymaterialsevaluationframeworkrandomstabilitybenchmarkingcgcnn
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid adoption of machine learning (ML) in domain sciences necessitates best practices and standardized benchmarking for performance evaluation. We present Matbench Discovery, an evaluation framework for ML energy models, applied as pre-filters for high-throughput searches of stable inorganic crystals. This framework addresses the disconnect between thermodynamic stability and formation energy, as well as retrospective vs. prospective benchmarking in materials discovery. We release a Python package to support model submissions and maintain an online leaderboard, offering insights into performance trade-offs. To identify the best-performing ML methodologies for materials discovery, we benchmarked various approaches, including random forests, graph neural networks (GNNs), one-shot predictors, iterative Bayesian optimizers, and universal interatomic potentials (UIP). Our initial results rank models by test set F1 scores for thermodynamic stability prediction: EquiformerV2 + DeNS > Orb > SevenNet > MACE > CHGNet > M3GNet > ALIGNN > MEGNet > CGCNN > CGCNN+P > Wrenformer > BOWSR > Voronoi fingerprint random forest. UIPs emerge as the top performers, achieving F1 scores of 0.57-0.82 and discovery acceleration factors (DAF) of up to 6x on the first 10k stable predictions compared to random selection. We also identify a misalignment between regression metrics and task-relevant classification metrics. Accurate regressors can yield high false-positive rates near the decision boundary at 0 eV/atom above the convex hull. Our results demonstrate UIPs' ability to optimize computational budget allocation for expanding materials databases. However, their limitations remain underexplored in traditional benchmarks. We advocate for task-based evaluation frameworks, as implemented here, to address these limitations and advance ML-guided materials discovery.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 30 citations worldwide. Full citation record

  1. Pushing the limits of unconstrained machine-learned interatomic potentials

    physics.chem-ph 2026-01 conditional novelty 7.0 of 10

    Unconstrained non-equivariant and direct-force neural interatomic potentials scale to 730M parameters and match or beat equivariant state-of-the-art models on several atomistic benchmarks.

  2. MP-ALOE: An r2SCAN dataset for universal machine learning interatomic potentials

    cond-mat.mtrl-sci 2025-07 conditional novelty 7.0 of 10

    MP-ALOE provides 909,792 r2SCAN DFT frames of mostly off-equilibrium structures; a MACE potential trained on it improves molecular dynamics stability and pressure robustness over MatPES-trained models.

  3. From Evaluation to Design: Using Potential Energy Surface Smoothness Metrics to Guide Machine Learning Interatomic Potential Architectures

    cs.LG 2026-02 conditional novelty 6.0 of 10

    A bond-deformation benchmark plus a force-smoothness metric is proposed to detect PES artifacts and guide MLIP architecture design, with improvements shown on a new Transformer-style model.

  4. MiAD: Mirage Atom Diffusion for De Novo Crystal Generation

    cs.LG 2025-11 conditional novelty 6.0 of 10

    Mirage infusion lets crystal diffusion models vary atom counts during generation and raises the S.U.N. rate on MP-20 to 8.2%.

  5. OpenCSP: A Deep Learning Framework for Crystal Structure Prediction from Ambient to High Pressure

    cond-mat.mtrl-sci 2025-09 conditional novelty 6.0 of 10

    OpenCSP is an open pressure-diverse dataset and model suite that matches or beats larger universal atomistic models on high-pressure crystal structure prediction with far fewer training data.

  6. Universal Machine Learning Potentials under Pressure

    cond-mat.mtrl-sci 2025-08 conditional novelty 6.0 of 10

    Universal machine learning interatomic potentials systematically lose accuracy under pressure up to 150 GPa, and fine-tuning on high-pressure DFT data recovers most of the lost performance.

  7. Universal Machine Learning Potential for Systems with Reduced Dimensionality

    cond-mat.mtrl-sci 2025-08 conditional novelty 6.0 of 10

    Benchmarking 11 universal machine learning interatomic potentials on a new 40,000-structure, 0D-3D dataset shows energy and geometry errors grow as dimensionality falls, with eSEN the most transferable.

  8. Enhancing Materials Discovery with Valence Constrained Design in Generative Modeling

    cond-mat.mtrl-sci 2025-07 conditional novelty 6.0 of 10

    CrysVCD generates valence-balanced compositions with an elemental language model and then constructs their crystal structures with a diffusion model, reporting improved stability and functional property targeting.

  9. Comparing classical and machine learning force fields for modeling deformation of solid sorbents relevant for direct air capture

    cond-mat.mtrl-sci 2025-06 conditional novelty 6.0 of 10

    No tested force field, classical or machine-learned, achieves the 0.1 eV accuracy target for adsorption energies in deformable MOFs relevant to direct air capture.

  10. Benchmarking Universal Machine Learning Interatomic Potentials for Real-Time Analysis of Inelastic Neutron Scattering Data

    physics.comp-ph 2025-06 conditional novelty 6.0 of 10

    Several universal machine learning interatomic potentials, especially ORB v3, MatterSim, and MACE-OFF, reach near-DFT phonon accuracy and match many experimental neutron spectra, with important caveats about test-set ...

  11. Toward Exascale AI for Science: A Scalable AI Skill for Autonomous Microkinetics Discovery

    cs.CE 2026-06 unverdicted novelty 5.0 of 10

    Introduces a scalable AI skill framework for autonomous microkinetics discovery that automates workflows and evaluates surrogate reliability.

  12. Toward Greater Autonomy in Materials Discovery Agents: Unifying Planning, Physics, and Scientists

    cs.AI 2025-06 reject novelty 5.0 of 10

    MAPPS combines LLM workflow planning, code generation, and human intuition with machine-learned force fields to discover crystal structures, reporting high stability and novelty rates on MP-20 and Matbench.

Pith tools