Pith. sign in

REVIEW 5 cited by

GemNet-OC: Developing Graph Neural Networks for Large and Diverse Molecular Simulation Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.02782 v3 pith:DX2TFTK2 submitted 2022-04-06 cs.LG cond-mat.mtrl-sciphysics.chem-phphysics.comp-ph

classification cs.LGcond-mat.mtrl-sciphysics.chem-phphysics.comp-ph
keywords datasetsdatasetmodelgemnet-ococ20developinglargemolecular
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent years have seen the advent of molecular simulation datasets that are orders of magnitude larger and more diverse. These new datasets differ substantially in four aspects of complexity: 1. Chemical diversity (number of different elements), 2. system size (number of atoms per sample), 3. dataset size (number of data samples), and 4. domain shift (similarity of the training and test set). Despite these large differences, benchmarks on small and narrow datasets remain the predominant method of demonstrating progress in graph neural networks (GNNs) for molecular simulation, likely due to cheaper training compute requirements. This raises the question -- does GNN progress on small and narrow datasets translate to these more complex datasets? This work investigates this question by first developing the GemNet-OC model based on the large Open Catalyst 2020 (OC20) dataset. GemNet-OC outperforms the previous state-of-the-art on OC20 by 16% while reducing training time by a factor of 10. We then compare the impact of 18 model components and hyperparameter choices on performance in multiple datasets. We find that the resulting model would be drastically different depending on the dataset used for making model choices. To isolate the source of this discrepancy we study six subsets of the OC20 dataset that individually test each of the above-mentioned four dataset aspects. We find that results on the OC-2M subset correlate well with the full OC20 dataset while being substantially cheaper to train on. Our findings challenge the common practice of developing GNNs solely on small datasets, but highlight ways of achieving fast development cycles and generalizable results via moderately-sized, representative datasets such as OC-2M and efficient models such as GemNet-OC. Our code and pretrained model weights are open-sourced.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Replacing explicit neural network stacks with self-consistent fixed-point iterations, and warm-starting the solver across timesteps, gives 2-5x cheaper molecular dynamics force evaluation at matched accuracy.

  2. Global Plane Waves From Local Gaussians: Periodic Charge Densities in a Blink

    cond-mat.mtrl-sci 2026-01 conditional novelty 6.0 of 10

    ELECTRAFI predicts periodic electron densities by analytically Fourier-transforming a Gaussian mixture, reaching near-SOTA accuracy with up to 633× faster inference and ~20% end-to-end DFT speedups.

  3. Insights into CO dimerization at electrified Cu interfaces from large-scale machine learning simulations

    cond-mat.mtrl-sci 2025-09 reject novelty 6.0 of 10

    OC25 is a large open dataset and baseline models for solid-liquid interfaces, but the claimed CO dimerization insights are absent from the manuscript body.

  4. Machine Learning Interatomic Potentials: library for efficient training, model development and simulation of molecular systems

    physics.chem-ph 2025-05 conditional novelty 6.0 of 10

    InstaDeep's mlip library ports MACE, NequIP, and ViSNet to JAX with a JAX-MD backend, ships SPICE2-trained organics models, reports faster MD steps than its own Torch routes, and proposes a faster gated MACE variant i...

  5. Global Universal Scaling and Ultra-Small Parameterization in Machine Learning Interatomic Potentials with Super-Linearity

    cond-mat.mtrl-sci 2025-02 conditional novelty 6.0 of 10

    By rescaling atomic pair distances with element-pair-specific parameters, the authors make one shared radial function serve all elements, yielding an ultra-small machine learning interatomic potential with accuracy cl...

Pith tools