REVIEW 4 major objections 5 minor 2 cited by
Massive Atomic Diversity: a compact universal dataset for atomistic machine learning
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A dataset of under 100,000 distorted structures can train universal interatomic potentials as well as datasets 100–1000 times larger.
desk verdict MAD is a serious, well-documented data contribution, but the headline claim of competitive training performance is borrowed from the companion paper and the diversity map is built on features trained on MAD itself; neither flaw is fatal. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the MAD dataset itself: 95,595 structures spanning 85 elements, organized into eight subsets (MC3D, MC3D-rattled, MC3D-random, MC3D-surface, MC3D-cluster, MC2D, SHIFTML-molcrys, SHIFTML-molfrags), all computed with a single consistent plane-wave DFT protocol using the PBEsol functional, no spin polarization, and uniform pseudopotential and smearing choices. The transformations that generate diversity—Gaussian rattling scaled to covalent radii, random elemental substitution with volume rescaling, surface cleavage, and cluster cutting—are the mechanism by which the dataset escapes the stable-structure bias of conventional databases. A secondary mechanism is the latent representation obtained from the last-layer features of the PET-MAD model, projected to two or three dimensions with sketch-map and approximated by a neural network, which serves both to measure the dataset's coverage and to compare it with other datasets.
What would settle it
Compute the formation energy and magnetic ordering energy of a strongly magnetic transition-metal oxide such as NiO with the MAD protocol (PBEsol, no spin polarization) and with a spin-polarized DFT+U protocol, then compare a MAD-trained universal potential against a model trained on a dataset that includes such corrections; if the MAD-trained model's errors on these quantities exceed its typical accuracy on non-magnetic systems, the claim of universality across the full chemical space is falsified for that domain. Alternatively, test for selection bias by recomputing the discarded non-converged MC3D-random structures (55% convergence rate) with more robust settings; if their energies are systematically higher than the surviving structures, the dataset's advertised high-energy diversity is partly an artifact of filtering.
Extended reading notes
Core claim
The discovery the paper is trying to establish is that deliberate diversity and computational consistency can substitute for sheer dataset size in training universal machine-learning interatomic potentials. Starting from stable structures in the MC3D, MC2D, and SHIFTML databases, the authors generate rattled, randomized-composition, surface, cluster, and molecular-fragment structures, and compute energies and forces for all of them with one set of DFT settings chosen for consistency rather than per-system accuracy. The resulting 95,595-structure dataset is claimed to enable training of the PET-MAD universal potential to a level competitive with models trained on datasets that are two to three orders of magnitude larger. The paper backs this with energy and force distributions, a structural cartography comparing MAD with MPtrj, Alexandria, SPICE, MD22, and OC2020, and a benchmark of structures recomputed under both MAD and MPtrj-style settings.
Load-bearing premise
The dataset's utility rests on the premise that one approximate DFT protocol, PBEsol without spin polarization, dispersion, or Hubbard corrections, gives a reliable enough reference for transferable potentials across all 85 elements and their distorted configurations, even though the paper acknowledges that magnetism, correlations, and dispersion are neglected.
Editorial extensions
If this is right
- Universal interatomic potentials can be trained with a fraction of the compute and data currently considered necessary.
- Dataset construction should prioritize diversity of configurations over sampling stable or plausible structures.
- A single consistent DFT protocol can serve as a common reference across organic and inorganic materials, enabling models that interpolate across both domains.
- The MAD benchmark subsets, recomputed under both MAD and MPtrj settings, provide a controlled way to compare models that would otherwise be confounded by differing reference calculations.
- The structural cartography offers a quantitative way to audit the coverage of any atomistic dataset before committing to expensive calculations.
Reading between the lines
- The MAD philosophy naturally extends to active learning: a model trained on MAD could select new distorted structures where its uncertainty is highest, growing coverage on demand rather than by blind random distortion.
- Because MAD deliberately omits spin polarization and dispersion, a hybrid strategy becomes plausible: train on MAD for general chemical space, then fine-tune on small corrected datasets for specific magnetic or van der Waals systems, combining broad coverage with targeted accuracy.
- The paper's own convergence statistics (55% for MC3D-random) point to a testable selection bias: if non-converged random structures tend to have higher energies, the surviving subset may under-represent the most extreme high-energy configurations, making the diversity tail thinner than intended.
- The latent-space cartography could be repurposed as a pre-training diagnostic: projecting a candidate dataset onto the MAD map would reveal whether it covers the high-energy regions needed for a universal potential, or whether it is concentrated near stable minima.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces the Massive Atomic Diversity (MAD) dataset, a collection of 95,595 structures built by aggressively distorting stable crystals and molecules from several existing databases, divided into eight subsets (bulk, rattled, random-composition, surfaces, clusters, 2D, molecular crystals, molecular fragments). All structures are recomputed with a deliberately consistent DFT protocol (PBEsol, no spin polarization, no dispersion, cold smearing, fixed cutoffs) using Quantum ESPRESSO, and the dataset is released in FAIR form with AiiDA provenance and extended XYZ files, including a benchmark set of 322 structures calculated with both MAD and MPtrj-like settings. The paper also proposes a low-dimensional structural latent space based on last-layer features of the PET-MAD model, projected with sketch-map and an MLP approximation, and uses this representation to compare the coverage of MAD with Alexandria, MPtrj, SPICE, MD22, and OC2020. The abstract and introduction claim that models trained on MAD are competitive with models trained on much larger traditional datasets, with this performance result attributed to the companion PET-MAD paper.
Significance. The dataset itself is a valuable and carefully documented resource: it is compact, openly released with full provenance, and its construction philosophy is transparent. The paired MAD/MPtrj benchmark subset is a strong contribution that enables direct comparisons of DFT settings. If the training-performance claim from the companion paper holds and the diversity analysis is robust, MAD could become a standard lightweight training set for universal machine-learning interatomic potentials. However, the manuscript's own contribution is primarily the dataset and its analysis; the headline performance claim is not demonstrated in this paper, and the diversity metric is potentially self-referential, so the significance as presented is conditional on resolving these points.
major comments (4)
- [Abstract; Section I] The central claim that MAD enables training universal interatomic potentials competitive with models trained on two to three orders of magnitude more structures is stated in the abstract and Section I but is not demonstrated anywhere in this manuscript; it is attributed to the companion paper [18]. Because this is the principal scientific motivation for the dataset, the paper should either include a direct benchmarking result (for example using the MAD-benchmark set described in Section IV A) or clearly mark the claim as a result of the companion paper and summarize its relevant benchmarks. As written, the abstract asserts the finding as established in the present work.
- [Section III A; Figure 7] The diversity analysis uses the last-layer features of the PET-MAD model, which is itself trained on the MAD dataset. This introduces a circularity: the feature metric used to demonstrate MAD's broad coverage is fitted to MAD, which may inflate its apparent diversity relative to datasets the model has never seen. To support the claim that MAD covers a broader chemical space than MPtrj or Alexandria, the authors should validate the comparison with an independent structural descriptor (e.g., SOAP or ACSF) or show that the PET-MAD features are not significantly adapted to MAD (for instance by comparing with features from a model trained on a different dataset). Without such validation, the coverage comparison in Figure 7 is not conclusive.
- [Section IV B; Section IV A] The DFT protocol deliberately neglects spin polarization, Hubbard-type corrections, and dispersion, and the paper acknowledges that this introduces errors for magnetic, correlated, and dispersion-bound systems. However, no quantitative estimate of the magnitude of these systematic shifts is provided. Since the MAD-benchmark set in Section IV A contains 322 structures computed under both MAD and MPtrj-like settings, the authors could directly report energy and force differences to quantify the reference-level offset. This would allow readers to judge whether the consistency choice introduces errors that are relevant for the target accuracy of modern machine-learning potentials, and would substantially strengthen the universality claim.
- [Section IV A; Section IV B] The MC3D-random subset, which the paper identifies as the most diverse (Figures 2, 3, and 6), has a DFT convergence rate of only about 55%. The non-converged structures are likely to be the most highly strained or chemically extreme configurations that the subset was specifically designed to provide. The paper should analyze the discarded structures—for example in terms of element combinations, strain, and energy distribution—and discuss how the 45% loss affects the coverage and diversity claims for this subset. As written, the filtering could selectively remove the very configurations that justify the subset's presence in the dataset.
minor comments (5)
- [Figure 7 caption] The caption spells 'Two-dimentional'; this should be 'Two-dimensional'.
- [Section IV C] The sentence 'The idea of is to project' is ungrammatical and should be 'The idea is to project'.
- [Section IV C] The phrase 'an simple Multi-Layer Perceptron' should be 'a simple Multi-Layer Perceptron'.
- [Section III C] The text refers to 'MD17 dataset' in the comparison, but the datasets plotted in Figure 7 and described in the caption are SPICE and MD22; this appears to be a typo and should be corrected to 'MD22'.
- [Section IV A] The acronym 'MPtraj' is used in the text while 'MPtrj' is used elsewhere; the spelling should be made consistent throughout (the Materials Project trajectory dataset is usually abbreviated MPtrj).
Circularity Check
Dataset construction is self-contained, but the headline utility claim is outsourced to a same-group companion paper and the diversity map is built from features of a model trained on the same dataset.
-
self citation load bearing
[Abstract; Section II, paragraph 2]
"The MAD dataset we present here, despite containing fewer than 100k structures, has already been shown to enable training universal interatomic potentials that are competitive with models trained on traditional datasets with two to three orders of magnitude more structures."
This sentence is the central value proposition of the paper, yet the 'shown' refers to Ref. [18], the companion PET-MAD paper by the same group, and no training benchmark appears in the present manuscript. The claim that the dataset enables competitive universal potentials therefore reduces to a self-citation: the reader is asked to accept the dataset's utility on the authors' own prior work, rather than on evidence contained in or independently verifiable from this paper.
-
self definitional
[Section III A (PET-MAD latent features); Section V (dataset split)]
"For the high-dimensional description, we use the last-layer features of the trained PET-MAD model [18], that provide a 512-sized token describing each i-atom-centered environment in a given structure, xi(A_i)."
PET-MAD is the universal potential trained on this MAD dataset, as stated in Section V: 'The MAD dataset, with an 80:10:10 train:validation:test split as used in training the PET-MAD model'. Thus the high-dimensional descriptors used to construct the distance histograms (Fig. 3) and maps (Fig. 7) are themselves fitted to the dataset whose diversity they are used to demonstrate. Distances and projections in this space are not independent external coordinates; they are a compressed view produced by a model whose parameters encode the MAD training set. The conclusion that MAD 'covers a broader portion of chemical space' is therefore self-referential, not an external measurement.
full rationale
The paper's core dataset construction is not circular: structures are generated from MC3D, MC2D, and SHIFTML inputs by well-defined distortions and recomputed at a consistent DFT level, and the energy/force spread in Fig. 2 follows directly from that protocol. However, two load-bearing parts of the argument are self-referential. First, the abstract claims that the dataset 'has already been shown to enable training universal interatomic potentials' that are competitive with much larger datasets, but the demonstration is not in this paper: it is attributed to Ref. [18], the PET-MAD paper by the same group, making the central utility claim a self-citation. Second, Section III uses PET-MAD last-layer features to define the chemical-space map and diversity histograms; since PET-MAD was trained on this very dataset, the metric used to establish MAD's broad coverage is an in-sample representation rather than an independent benchmark. These issues make the diversity and competitiveness claims partially circular, although the dataset itself, its DFT consistency rationale, and its acknowledged approximations are independent content. The PBEsol/no-spin/no-dispersion reference is an accuracy concern, not a circularity, and the paper itself flags the limitations.
Assumptions & free parameters
free parameters (3)
- Rattling Gaussian noise amplitude =
20% of covalent radii
- Outlier force thresholds =
100 eV/A for MC3D-rattled and MC3D-random; 10 eV/A for other subsets
- DFT protocol parameters =
PBEsol; SSSP v1.2 efficiency; 110/1320 Ry cutoffs; 0.01 Ry smearing; 0.125 per Angstrom k-grid; no spin polarization
assumptions (3)
- domain assumption Aggressively distorted structures with randomized compositions are representative of the out-of-equilibrium configurations encountered in real atomistic simulations.
- domain assumption PBEsol without spin polarization and without explicit dispersion is an adequate reference for training transferable universal potentials.
- ad hoc to paper The last-layer features of PET-MAD define a fair metric for comparing chemical and structural coverage across datasets.
Cite this review
Pith. "Pith review of Massive Atomic Diversity: a compact universal dataset for atomistic machine learning." pith.science (2026). https://pith.science/paper/JKJ3MVMT
@misc{pith2026250619674,
author = {Pith},
title = {Pith review of: Massive Atomic Diversity: a compact universal dataset for atomistic machine learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/JKJ3MVMT}},
note = {Machine review of arXiv:2506.19674}
}
read the original abstract
The development of machine-learning models for atomic-scale simulations has benefited tremendously from the large databases of materials and molecular properties computed in the past two decades using electronic-structure calculations. More recently, these databases have made it possible to train universal models that aim at making accurate predictions for arbitrary atomic geometries and compositions. The construction of many of these databases was however in itself aimed at materials discovery, and therefore targeted primarily to sample stable, or at least plausible, structures and to make the most accurate predictions for each compound - e.g. adjusting the calculation details to the material at hand. Here we introduce a dataset designed specifically to train machine learning models that can provide reasonable predictions for arbitrary structures, and that therefore follows a different philosophy. Starting from relatively small sets of stable structures, the dataset is built to contain massive atomic diversity (MAD) by aggressively distorting these configurations, with near-complete disregard for the stability of the resulting configurations. The electronic structure details, on the other hand, are chosen to maximize consistency rather than to obtain the most accurate prediction for a given structure, or to minimize computational effort. The MAD dataset we present here, despite containing fewer than 100k structures, has already been shown to enable training universal interatomic potentials that are competitive with models trained on traditional datasets with two to three orders of magnitude more structures. We describe in detail the philosophy and details of the construction of the MAD dataset. We also introduce a low-dimensional structural latent space that allows us to compare it with other popular datasets and that can be used as a general-purpose materials cartography tool.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 2 Pith papers
-
Score-based diffusion models for accurate crystal-structure inpainting and reconstruction of hydrogen positions
Adapting TD-Paint to crystal diffusion models reconstructs hydrogen positions with a LES success rate above 97%, beating unconditioned diffusion and DFT-based inpainting.
-
VASP Plugins: Linking the Vienna ab-initio Simulation Package with Python
A C++/pybind11 shared-memory plugin layer exposes VASP SCF and ionic data as NumPy arrays so Python can modify structure, forces, local potential, and occupancies in place.
Reference graph
Works this paper leans on
-
[18]
A. Mazitov, F. Bigi, M. Kellner, P. Pegolo, D. Tisi, G. Fraux, S. Pozdnyakov, P. Loche, and M. Ceri- otti, Pet-mad, a universal interatomic potential for advanced materials modeling (2025), arXiv:2503.14118 [cond-mat.mtrl-sci]
arXiv 2025
-
[1]
A. Jain, S. P. Ong, G. Hautier, W. Chen, W. D. Richards, S. Dacek, S. Cholia, D. Gunter, D. Skinner, G. Ceder, and K. A. Persson, Commentary: The Materials Project: A materials genome approach to accelerating materials innovation, APL Materials1, 011002 (2013)
2013
-
[2]
B. Deng, P. Zhong, K. Jun, J. Riebesell, K. Han, C. J. Bartel, and G. Ceder, Chgnet: Pretrained universal neu- ral network potential for charge-informed atomistic mod- eling (2023), arXiv:2302.14231 [cond-mat.mtrl-sci]
arXiv 2023
-
[3]
L. Talirz, S. Kumbhar, E. Passaro, A. V. Yakutovich, V. Granata, F. Gargiulo, M. Borelli, M. Uhrin, S. P. Huber, S. Zoupanos, C. S. Adorf, C. W. Andersen, O. Sch¨ utt, C. A. Pignedoli, D. Passerone, J. VandeVon- dele, T. C. Schulthess, B. Smit, G. Pizzi, and N. Marzari, Materials cloud, a platform for open computational sci- ence, Scientific Data7, 299 (2020)
work page 2020
-
[4]
M. Scheidgen, L. M. Ghiringhelli, F. Dietrich, D. Lehm- berg, T. Denell, A. Albino, H. N¨ asstr¨ om, S. Shabih, F. Dobener, M. K¨ uhbach, R. Mozumder, J. F. Rudzin- ski, N. Daelman, J. M. Pizarro, M. Kuban, C. Salazar, P. Ondraˇ cka, H.-J. Bungartz, and C. Draxl, Nomad: A distributed web-based platform for managing materials science research data, Journal...
work page 2023
-
[5]
K. Choudhary, K. F. Garrity, A. C. E. Reid, B. DeCost, A. J. Biacchi, A. R. Hight Walker, Z. Trautt, J. Hattrick- Simpers, A. G. Kusne, A. Centrone, A. Davydov, J. Jiang, R. Pachter, G. Cheon, E. Reed, A. Agrawal, X. Qian, V. Sharma, H. Zhuang, S. V. Kalinin, B. G. Sumpter, G. Pilania, P. Acar, S. Mandal, K. Haule, D. Vanderbilt, K. Rabe, and F. Tavazza, ...
work page 2020
-
[6]
S. Chmiela, V. Vassilev-Galindo, O. T. Unke, A. Kabylda, H. E. Sauceda, A. Tkatchenko, and K.-R. M¨ uller, Accurate global machine learning force fields for molecules with hundreds of atoms (2022), arXiv:2209.14865 [physics.chem-ph]
arXiv 2022
-
[7]
R. Tran, J. Lan, M. Shuaibi, B. M. Wood, S. Goyal, A. Das, J. Heras-Domingo, A. Kolluru, A. Rizvi, N. Shoghi, A. Sriram, F. Therrien, J. Abed, O. Voznyy, E. H. Sargent, Z. Ulissi, and C. L. Zitnick, The open cat- alyst 2022 (oc22) dataset and challenges for oxide elec- trocatalysts, ACS Catalysis13, 3066–3084 (2023)
work page 2023
Show all 45 references
-
[8]
Schmidt, N
J. Schmidt, N. Hoffmann, H.-C. Wang, P. Borlido, P. J. M. A. Carri¸ co, T. F. T. Cerqueira, S. Botti, and M. A. L. Marques, Machine-learning-assisted determination of the global zero-temperature phase diagram of materials, Ad- vanced Materials35, 2210788 (2023)
2023
-
[9]
H.-C. Wang, J. Schmidt, M. A. L. Marques, L. Wirtz, and A. H. Romero, Symmetry-based computational search for novel binary and ternary 2d materials, 2D Materials10, 035007 (2023)
2023
-
[10]
Barroso-Luque, M
L. Barroso-Luque, M. Shuaibi, X. Fu, B. M. Wood, M. Dzamba, M. Gao, A. Rizvi, C. L. Zitnick, and Z. W. Ulissi, Open materials 2024 (omat24) inorganic materi- als dataset and models (2024), arXiv:2410.12771 [cond- mat.mtrl-sci]
2024 arXiv
-
[11]
Huber, M
S. Huber, M. Bercx, N. H¨ ormann, M. Uhrin, G. Pizzi, and N. Marzari, Materials cloud three-dimensional crys- tals database (mc3d), Materials Cloud Archive (2022), version v1, publication date: March 12, 2022
2022
-
[12]
Mounet, M
N. Mounet, M. Gibertini, P. Schwaller, D. Campi, A. Merkys, A. Marrazzo, T. Sohier, I. E. Castelli, A. Ce- pellotti, G. Pizzi, and N. Marzari, Two-dimensional ma- terials from high-throughput computational exfoliation of experimentally known compounds, Nature Nanotech- nology1...
2018
-
[13]
Campi, N
D. Campi, N. Mounet, M. Gibertini, G. Pizzi, and N. Marzari, Expansion of the materials cloud 2d database, ACS Nano17, 11268 (2023)
2023
-
[14]
Cordova, E
M. Cordova, E. A. Engel, A. Stefaniuk, F. Paruzzo, A. Hofstetter, M. Ceriotti, and L. Emsley, A machine learning model of chemical shifts for chemically and structurally diverse molecular solids, The Journal of Physical Chemistry C126, 16710 (2022)
2022
-
[15]
C. R. Groom, I. J. Bruno, M. P. Lightfoot, and S. C. Ward, The Cambridge Structural Database, Acta Crys- tallographica Section B72, 171 (2016)
2016
-
[16]
R. K. Cersonsky, M. Pakhnova, E. A. Engel, and M. Ce- riotti, A data-driven interpretation of the stability of or- ganic molecular crystals, Chem. Sci.14, 1272 (2023)
2023
-
[17]
Mindless
M. Korth and S. Grimme, “Mindless” DFT Benchmark- ing, J. Chem. Theory Comput.5, 993 (2009)
2009
-
[19]
Imbalzano, A
G. Imbalzano, A. Anelli, D. Giofr´ e, S. Klees, J. Behler, and M. Ceriotti, Automatic selection of atomic finger- prints and reference configurations for machine-learning potentials, Journal of Chemical Physics148, 241730 (2018), arXiv:1804.02150
2018 arXiv
-
[20]
McInnes, J
L. McInnes, J. Healy, and J. Melville, Umap: Uniform manifold approximation and projection for dimension re- duction (2020), arXiv:1802.03426 [stat.ML]
2020 arXiv
-
[21]
van der Maaten and G
L. van der Maaten and G. Hinton, Visualizing data us- ing t-sne, Journal of Machine Learning Research9, 2579 (2008)
2008
-
[22]
Wattenberg, F
M. Wattenberg, F. Vi´ egas, and I. Johnson, How to use t-sne effectively, Distill 10.23915/distill.00002 (2016)
2016 doi
-
[23]
Ceriotti, G
M. Ceriotti, G. A. Tribello, and M. Parrinello, Simplify- ing the representation of complex free-energy landscapes using sketch-map, Proceedings of the National Academy of Sciences of the United States of America108, 13023 (2011). 10
2011
-
[24]
Eastman, P
P. Eastman, P. K. Behara, D. L. Dotson, R. Galvelis, J. E. Herr, J. T. Horton, Y. Mao, J. D. Chodera, B. P. Pritchard, Y. Wang, G. D. Fabritiis, and T. E. Mark- land, Spice, a dataset of drug-like molecules and pep- tides for training machine learning potentials (2022), arXiv:...
2022 arXiv
-
[25]
Chanussot, A
L. Chanussot, A. Das, S. Goyal, T. Lavril, M. Shuaibi, M. Riviere, K. Tran, J. Heras-Domingo, C. Ho, W. Hu, A. Palizhati, A. Sriram, B. Wood, J. Yoon, D. Parikh, C. L. Zitnick, and Z. Ulissi, Open catalyst 2020 (oc20) dataset and community challenges, ACS Catalysis11, 6059–6072 (2021)
2021
-
[26]
H.-C. Wang, S. Botti, and M. A. Marques, Predicting stable crystalline compounds using chemical similarity, npj Computational Materials7, 12 (2021)
2021
-
[27]
Riebesell, R
J. Riebesell, R. E. Goodall, P. Benner, Y. Chiang, B. Deng, A. A. Lee, A. Jain, and K. A. Persson, Matbench discovery–a framework to evaluate machine learning crystal stability predictions, arXiv preprint arXiv:2308.14920 (2023)
2023 arXiv
-
[28]
Giannozzi, S
P. Giannozzi, S. Baroni, N. Bonini, M. Calandra, R. Car, C. Cavazzoni, D. Ceresoli, G. L. Chiarotti, M. Cococ- cioni, I. Dabo, A. D. Corso, S. de Gironcoli, S. Fabris, G. Fratesi, R. Gebauer, U. Gerstmann, C. Gougoussis, A. Kokalj, M. Lazzeri, L. Martin-Samos, N. Marzari, F. M...
2009
-
[29]
Zhang, A
L. Zhang, A. Kozhevnikov, T. Schulthess, S. B. Trickey, and H.-P. Cheng, All-electron apw+localculation of mag- netic molecules with the sirius domain-specific package (2022), arXiv:2105.07363 [cond-mat.mtrl-sci]
2022 arXiv
-
[30]
Pizzi, A
G. Pizzi, A. Cepellotti, R. Sabatini, N. Marzari, and B. Kozinsky, AiiDA: automated interactive infrastructure and database for computational science, Computational Materials Science111, 218 (2016)
2016
-
[31]
S. P. Huber, S. Zoupanos, M. Uhrin, L. Talirz, L. Kahle, R. H¨ auselmann, D. Gresch, T. M¨ uller, A. V. Yakutovich, C. W. Andersen, F. F. Ramirez, C. S. Adorf, F. Gargiulo, S. Kumbhar, E. Passaro, C. Johnston, A. Merkys, A. Ce- pellotti, N. Mounet, N. Marzari, B. Kozinsky, and...
2020
-
[32]
Uhrin, S
M. Uhrin, S. P. Huber, J. Yu, N. Marzari, and G. Pizzi, Workflows in AiiDA: Engineering a high-throughput, event-based engine for robust and modular computa- tional workflows, Computational Materials Science187, 110086 (2021)
2021
-
[33]
J. P. Perdew, A. Ruzsinszky, G. I. Csonka, O. A. Vydrov, G. E. Scuseria, L. A. Constantin, X. Zhou, and K. Burke, Restoring the density-gradient expansion for exchange in solids and surfaces, Physical Review Letters100, 136406 (2008)
2008
-
[34]
Prandini, A
G. Prandini, A. Marrazzo, I. E. Castelli, N. Mounet, and N. Marzari, Precision and efficiency in solid-state pseu- dopotential calculations, npj Computational Materials4, 72 (2018)
2018
-
[35]
Marzari, D
N. Marzari, D. Vanderbilt, A. De Vita, and M. C. Payne, Thermal Contraction and Disordering of the Al(110) Sur- face, Phys. Rev. Lett.82, 3296 (1999)
1999
-
[36]
G. J. Martyna and M. E. Tuckerman, A reciprocal space based method for treating long range interactions inab initioand force-field-based calculations in clusters, The Journal of Chemical Physics110, 2810 (1999)
1999
-
[37]
Ceriotti, G
M. Ceriotti, G. A. Tribello, and M. Parrinello, Sim- plifying the representation of complex free-energy landscapes using sketch-map, Proceedings of the National Academy of Sciences108, 13023 (2011), https://www.pnas.org/doi/pdf/10.1073/pnas.1108486108
2011 doi
-
[38]
W. S. Torgerson, Multidimensional scaling: I. theory and method, Psychometrika17, 401–419 (1952)
1952
-
[39]
Ceriotti, G
M. Ceriotti, G. A. Tribello, and M. Parrinello, Demon- strating the transferability and the descriptive power of sketch-map, Journal of Chemical Theory and Computa- tion9, 1521 (2013)
2013
-
[40]
Girshick, Fast r-cnn (2015), arXiv:1504.08083 [cs.CV]
R. Girshick, Fast r-cnn (2015), arXiv:1504.08083 [cs.CV]
2015 arXiv
-
[41]
Fraux, R
G. Fraux, R. Cersonsky, and M. Ceriotti, Chemiscope: Interactive structure-property explorer for materials and molecules, JOSS5, 2117 (2020)
2020
-
[42]
Fraux, R
G. Fraux, R. K. Cersonski, and M. Ceriotti, Chemiscope software (2020)
2020
-
[43]
A. H. Larsen, J. J. Mortensen, J. Blomqvist, I. E. Castelli, R. Christensen, M. Du lak, J. Friis, M. N. Groves, B. Hammer, C. Hargus, E. D. Hermes, P. C. Jennings, P. B. Jensen, J. Kermode, J. R. Kitchin, E. L. Kolsbjerg, J. Kubal, K. Kaasbjerg, S. Lysgaard, J. B. Maronsson, T...
2017
-
[44]
Mazitov, S
A. Mazitov, S. Chorna, G. Fraux, M. Bercx, G. Pizzi, S. De, and M. Ceriotti, Massive Atomic Diversity: a compact universal dataset for atomistic machine learn- ing (2025)
2025
-
[45]
Pizzi, A
G. Pizzi, A. Cepellotti, R. Sabatini, N. Marzari, and B. Kozinsky, AiiDA: Automated interactive infrastruc- ture and database for computational science, Computa- tional Materials Science111, 218 (2016)
2016
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.