{"id":"006376c9-ae33-470a-9de2-8fe21b28ec13","arxiv_id":"2506.19674","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"MAD is a compact dataset of distorted, randomized, and diverse atomic structures computed with consistent DFT settings, designed to train universal interatomic potentials with far fewer samples.","lead":"A new database called MAD packs fewer than 100,000 atomic structures, many deliberately distorted or chemically randomized, all computed with one consistent electronic-structure protocol. It is meant to train general-purpose machine-learning interatomic potentials that match models trained on millions of structures.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unquantified systematic errors in the consistent PBEsol/no-spin/no-dispersion DFT reference are the key risk to the universality claim; the released MAD-vs-MPtrj paired benchmark data can quantify them directly.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the consistent but approximate DFT protocol may not provide a reliable reference for transferable potentials across the full chemical space, and the paper acknowledges but does not quantify the neglect of magnetism, correlations, and dispersion. This is the physical foundation of the central claim; if the reference PES is systematically biased for important chemistries, no amount of structural diversity or benchmark tuning can make the trained potential genuinely universal. The diversity-map circularity (features trained on MAD) and the reliance on a companion paper for the headline training result are real but secondary: the former affects a supporting visualization, and the latter can be resolved by checking the cited work. The proposed test uses data the authors have already released, so it is a low-cost, decisive check on the weakest premise. Since the reader already recommended a conditional verdict, my independent assessment does not change that verdict.","tokens_in":13196,"tokens_out":7688,"duration_ms":95034,"concrete_test":"Analyze the released paired files mad-bench-mad-settings.xyz and mad-bench-mptrj-settings.xyz: compute per-structure energy and per-atom force differences between the two DFT protocols, stratified by subset and by composition (especially magnetic elements Fe/Co/Ni, lanthanides, and dispersion-dominated molecular crystals). If the RMS energy/force differences for these strata exceed the companion paper's reported test errors for PET-MAD (e.g., approximately 50 meV/atom or 0.1 eV/A), then the consistent-DFT premise does not support the universality claim for those chemistries; if the differences are within the model's test errors, the premise holds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The dataset's central value proposition is that a consistent-but-approximate DFT protocol (PBEsol, no spin polarization, no dispersion, cold smearing; Section IV B) provides a reliable reference PES for training universal potentials. The paper explicitly acknowledges that magnetism, correlations, and dispersion are neglected (Section II), but it provides no quantitative estimate of how much these omissions shift energies and forces relative to the settings used by the large traditional datasets it claims to compete with. For magnetic transition-metal oxides, lanthanide-containing MC3D-random structures, and dispersion-bound molecular crystals, spin polarization, Hubbard U, and van der Waals corrections can change energies by hundreds of meV/atom—far larger than the target accuracy of modern machine-learning interatomic potentials. If a MAD-trained potential inherits this PBEsol/no-spin PES, its 'competitive' performance may be confined to benchmarks that avoid exactly these chemistries, undermining the universality claim. The paper's own MAD-benchmark release (Section IV A, 322 paired structures computed with both MAD and MPtrj-like settings) makes this testable: the same structures were computed under both protocols, so the systematic shift can be quantified directly. The 55% convergence rate of the MC3D-random subset is a secondary risk, as it may preferentially discard the most exotic high-energy configurations that subset was designed to provide.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces the Massive Atomic Diversity (MAD) dataset, a collection of 95,595 structures built by aggressively distorting stable crystals and molecules from several existing databases, divided into eight subsets (bulk, rattled, random-composition, surfaces, clusters, 2D, molecular crystals, molecular fragments). All structures are recomputed with a deliberately consistent DFT protocol (PBEsol, no spin polarization, no dispersion, cold smearing, fixed cutoffs) using Quantum ESPRESSO, and the dataset is released in FAIR form with AiiDA provenance and extended XYZ files, including a benchmark set of 322 structures calculated with both MAD and MPtrj-like settings. The paper also proposes a low-dimensional structural latent space based on last-layer features of the PET-MAD model, projected with sketch-map and an MLP approximation, and uses this representation to compare the coverage of MAD with Alexandria, MPtrj, SPICE, MD22, and OC2020. The abstract and introduction claim that models trained on MAD are competitive with models trained on much larger traditional datasets, with this performance result attributed to the companion PET-MAD paper.","tokens_in":13482,"tokens_out":5386,"duration_ms":51870,"significance":"The dataset itself is a valuable and carefully documented resource: it is compact, openly released with full provenance, and its construction philosophy is transparent. The paired MAD/MPtrj benchmark subset is a strong contribution that enables direct comparisons of DFT settings. If the training-performance claim from the companion paper holds and the diversity analysis is robust, MAD could become a standard lightweight training set for universal machine-learning interatomic potentials. However, the manuscript's own contribution is primarily the dataset and its analysis; the headline performance claim is not demonstrated in this paper, and the diversity metric is potentially self-referential, so the significance as presented is conditional on resolving these points.","major_comments":[{"comment":"The central claim that MAD enables training universal interatomic potentials competitive with models trained on two to three orders of magnitude more structures is stated in the abstract and Section I but is not demonstrated anywhere in this manuscript; it is attributed to the companion paper [18]. Because this is the principal scientific motivation for the dataset, the paper should either include a direct benchmarking result (for example using the MAD-benchmark set described in Section IV A) or clearly mark the claim as a result of the companion paper and summarize its relevant benchmarks. As written, the abstract asserts the finding as established in the present work.","section":"Abstract; Section I"},{"comment":"The diversity analysis uses the last-layer features of the PET-MAD model, which is itself trained on the MAD dataset. This introduces a circularity: the feature metric used to demonstrate MAD's broad coverage is fitted to MAD, which may inflate its apparent diversity relative to datasets the model has never seen. To support the claim that MAD covers a broader chemical space than MPtrj or Alexandria, the authors should validate the comparison with an independent structural descriptor (e.g., SOAP or ACSF) or show that the PET-MAD features are not significantly adapted to MAD (for instance by comparing with features from a model trained on a different dataset). Without such validation, the coverage comparison in Figure 7 is not conclusive.","section":"Section III A; Figure 7"},{"comment":"The DFT protocol deliberately neglects spin polarization, Hubbard-type corrections, and dispersion, and the paper acknowledges that this introduces errors for magnetic, correlated, and dispersion-bound systems. However, no quantitative estimate of the magnitude of these systematic shifts is provided. Since the MAD-benchmark set in Section IV A contains 322 structures computed under both MAD and MPtrj-like settings, the authors could directly report energy and force differences to quantify the reference-level offset. This would allow readers to judge whether the consistency choice introduces errors that are relevant for the target accuracy of modern machine-learning potentials, and would substantially strengthen the universality claim.","section":"Section IV B; Section IV A"},{"comment":"The MC3D-random subset, which the paper identifies as the most diverse (Figures 2, 3, and 6), has a DFT convergence rate of only about 55%. The non-converged structures are likely to be the most highly strained or chemically extreme configurations that the subset was specifically designed to provide. The paper should analyze the discarded structures—for example in terms of element combinations, strain, and energy distribution—and discuss how the 45% loss affects the coverage and diversity claims for this subset. As written, the filtering could selectively remove the very configurations that justify the subset's presence in the dataset.","section":"Section IV A; Section IV B"}],"minor_comments":[{"comment":"The caption spells 'Two-dimentional'; this should be 'Two-dimensional'.","section":"Figure 7 caption"},{"comment":"The sentence 'The idea of is to project' is ungrammatical and should be 'The idea is to project'.","section":"Section IV C"},{"comment":"The phrase 'an simple Multi-Layer Perceptron' should be 'a simple Multi-Layer Perceptron'.","section":"Section IV C"},{"comment":"The text refers to 'MD17 dataset' in the comparison, but the datasets plotted in Figure 7 and described in the caption are SPICE and MD22; this appears to be a typo and should be corrected to 'MD22'.","section":"Section III C"},{"comment":"The acronym 'MPtraj' is used in the text while 'MPtrj' is used elsewhere; the spelling should be made consistent throughout (the Materials Project trajectory dataset is usually abbreviated MPtrj).","section":"Section IV A"}],"recommendation":"major_revision","confidential_remarks":"The dataset release itself is a solid contribution with good provenance and a useful paired benchmark set. The main concern is that the paper's central selling point, the competitive training performance, is outsourced to companion paper [18], and the diversity analysis in Section III is based on features of a model trained on the same dataset. Both issues are addressable in revision, but they affect the paper's scientific claims and should be resolved before acceptance. I would also encourage the editor to ensure that the relationship with Ref. [18] is disclosed clearly enough that the reader understands which results are established where."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a useful dataset paper and the data release is a real asset. The central claim that a <100k-structure set trains universal potentials competitive with models trained on millions of structures is not demonstrated here. It is cited to the companion PET-MAD paper by the same group. That does not make the claim false, but this manuscript asks the reader to take it on faith, and the companion paper needs to be part of the assessment.\n\nWhat is genuinely new: the Korth–Grimme \"mindless\" distortion idea is extended from isolated molecules to a mixed organic/inorganic materials dataset with surfaces, clusters, random compositions, and consistent PBEsol recalculation. The construction details are concrete. The AiiDA archives and xyz files are released with a recommended split. The MAD-benchmark set of 322 structures computed under both MAD and MPtrj-like settings is a nice, potentially decisive addition—it makes the systematic DFT-protocol question directly answerable.\n\nSoft spots, in proportion:\n- The diversity analysis uses last-layer features of PET-MAD, a model trained on MAD itself. So the map is not an independent measure of coverage. The authors are open about this; it still weakens the comparison with other datasets.\n- The consistent-PBEsol/no-spin/no-dispersion reference is a deliberate trade-off, and the paper says so. But the magnitude of the systematic shift is not quantified. For magnetic oxides, lanthanide-containing randomized structures, and dispersion-bound molecular crystals, hundreds of meV/atom errors are plausible. The stress-test concern lands: the paired benchmark data should be used to show the shift, or the universality claim should be scoped.\n- The 55% convergence of MC3D-random is disclosed; it may preferentially kill the most exotic high-energy structures, which is exactly what that subset was for. Minor-to-moderate concern, and partly unavoidable.\n\nWho gets value: MLIP developers choosing training data, and anyone building or benchmarking universal potentials. The paper deserves serious peer review. I would send it out with a request that the authors either include a direct training comparison or make the dependence on the companion paper explicit and verify the DFT-protocol shift with the released paired structures.","headline":"MAD is a serious, well-documented data contribution, but the headline claim of competitive training performance is borrowed from the companion paper and the diversity map is built on features trained on MAD itself; neither flaw is fatal.","tokens_in":13982,"tokens_out":1843,"would_cite":true,"duration_ms":19274,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A dataset of under 100,000 distorted structures can train universal interatomic potentials as well as datasets 100–1000 times larger.","keywords":["massive atomic diversity","atomistic machine learning","interatomic potentials","dataset design","density functional theory consistency","structural diversity","latent space cartography","universal potential"],"falsifier":"Compute the formation energy and magnetic ordering energy of a strongly magnetic transition-metal oxide such as NiO with the MAD protocol (PBEsol, no spin polarization) and with a spin-polarized DFT+U protocol, then compare a MAD-trained universal potential against a model trained on a dataset that includes such corrections; if the MAD-trained model's errors on these quantities exceed its typical accuracy on non-magnetic systems, the claim of universality across the full chemical space is falsified for that domain. Alternatively, test for selection bias by recomputing the discarded non-converged MC3D-random structures (55% convergence rate) with more robust settings; if their energies are systematically higher than the surviving structures, the dataset's advertised high-energy diversity is partly an artifact of filtering.","tokens_in":13021,"feed_emoji":"⚛️","tokens_out":7209,"duration_ms":60589,"temperature":0.7,"pith_summary":"The paper introduces the Massive Atomic Diversity (MAD) dataset, which contains fewer than 100,000 structures built by aggressively distorting stable crystals and molecules so that it covers high-energy, out-of-equilibrium configurations that conventional materials databases miss. Every structure is recomputed with a single, deliberately consistent density-functional theory protocol (PBEsol, no spin polarization) rather than per-material tuned settings. The core claim is that a universal interatomic potential trained on this compact set is competitive with models trained on traditional datasets with two to three orders of magnitude more structures. The paper also presents a low-dimensional latent space, built from PET-MAD features and projected with sketch-map, for comparing datasets and visualizing chemical-space coverage. If the claim holds, the bottleneck for universal potentials is not raw data volume but the diversity of configurations and the consistency of the reference calculations.","feed_headline":"A 96k-structure dataset rivals million-structure training sets","feed_subtitle":"Deliberately distorted structures plus one consistent DFT setting beat raw dataset size.","key_machinery":"The central object is the MAD dataset itself: 95,595 structures spanning 85 elements, organized into eight subsets (MC3D, MC3D-rattled, MC3D-random, MC3D-surface, MC3D-cluster, MC2D, SHIFTML-molcrys, SHIFTML-molfrags), all computed with a single consistent plane-wave DFT protocol using the PBEsol functional, no spin polarization, and uniform pseudopotential and smearing choices. The transformations that generate diversity—Gaussian rattling scaled to covalent radii, random elemental substitution with volume rescaling, surface cleavage, and cluster cutting—are the mechanism by which the dataset escapes the stable-structure bias of conventional databases. A secondary mechanism is the latent representation obtained from the last-layer features of the PET-MAD model, projected to two or three dimensions with sketch-map and approximated by a neural network, which serves both to measure the dataset's coverage and to compare it with other datasets.","core_discovery":"The discovery the paper is trying to establish is that deliberate diversity and computational consistency can substitute for sheer dataset size in training universal machine-learning interatomic potentials. Starting from stable structures in the MC3D, MC2D, and SHIFTML databases, the authors generate rattled, randomized-composition, surface, cluster, and molecular-fragment structures, and compute energies and forces for all of them with one set of DFT settings chosen for consistency rather than per-system accuracy. The resulting 95,595-structure dataset is claimed to enable training of the PET-MAD universal potential to a level competitive with models trained on datasets that are two to three orders of magnitude larger. The paper backs this with energy and force distributions, a structural cartography comparing MAD with MPtrj, Alexandria, SPICE, MD22, and OC2020, and a benchmark of structures recomputed under both MAD and MPtrj-style settings.","pith_inferences":["The MAD philosophy naturally extends to active learning: a model trained on MAD could select new distorted structures where its uncertainty is highest, growing coverage on demand rather than by blind random distortion.","Because MAD deliberately omits spin polarization and dispersion, a hybrid strategy becomes plausible: train on MAD for general chemical space, then fine-tune on small corrected datasets for specific magnetic or van der Waals systems, combining broad coverage with targeted accuracy.","The paper's own convergence statistics (55% for MC3D-random) point to a testable selection bias: if non-converged random structures tend to have higher energies, the surviving subset may under-represent the most extreme high-energy configurations, making the diversity tail thinner than intended.","The latent-space cartography could be repurposed as a pre-training diagnostic: projecting a candidate dataset onto the MAD map would reveal whether it covers the high-energy regions needed for a universal potential, or whether it is concentrated near stable minima."],"forward_implications":["Universal interatomic potentials can be trained with a fraction of the compute and data currently considered necessary.","Dataset construction should prioritize diversity of configurations over sampling stable or plausible structures.","A single consistent DFT protocol can serve as a common reference across organic and inorganic materials, enabling models that interpolate across both domains.","The MAD benchmark subsets, recomputed under both MAD and MPtrj settings, provide a controlled way to compare models that would otherwise be confounded by differing reference calculations.","The structural cartography offers a quantitative way to audit the coverage of any atomistic dataset before committing to expensive calculations."],"supporting_citations":[{"why":"Supplies the PET-MAD universal potential whose training on MAD demonstrates the dataset's competitiveness.","marker":"[18]"},{"why":"Source of the bulk crystal seed structures for the MC3D and derived subsets.","marker":"[11]"},{"why":"Source of the two-dimensional crystal structures used in the MC2D subset.","marker":"[12, 13]"},{"why":"Source of the curated molecular crystals used in the SHIFTML-molcrys subset.","marker":"[14]"},{"why":"Source of the neutral molecular fragments used in the SHIFTML-molfrags subset.","marker":"[16]"},{"why":"Provides the 'mindless' aggressive-distortion philosophy that MAD adapts to periodic and universal settings.","marker":"[17]"},{"why":"Largest-scale comparison dataset (MPtrj/CHGNet) whose trained models MAD-trained potentials are claimed to rival.","marker":"[2]"},{"why":"Alexandria datasets used as comparison benchmarks in the coverage analysis.","marker":"[8, 9]"},{"why":"Sketch-map method used to build the low-dimensional structural cartography for comparing datasets.","marker":"[23]"}],"fun_headline_variants":["96k diverse structures beat million-structure datasets","Massive diversity, not size, powers universal atomistic ML","Small dataset, big diversity: MAD challenges big data","Consistent DFT + distortion beats raw dataset scale","Rattled structures make a compact universal ML dataset"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's utility rests on the premise that one approximate DFT protocol, PBEsol without spin polarization, dispersion, or Hubbard corrections, gives a reliable enough reference for transferable potentials across all 85 elements and their distorted configurations, even though the paper acknowledges that magnetism, correlations, and dispersion are neglected.","fun_headline_variants_meta":{"raw":{"variants":["96k diverse structures beat million-structure datasets","Massive diversity, not size, powers universal atomistic ML","Small dataset, big diversity: MAD challenges big data","Consistent DFT + distortion beats raw dataset scale","Rattled structures make a compact universal ML dataset"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000605,"raw_usage":{"total_tokens":2853,"prompt_tokens":1005,"completion_tokens":1848,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":1772}},"tokens_in":621,"tokens_out":1848,"duration_ms":13516,"temperature":1.0,"reasoning_tokens":1772,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:27:38.343111+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the formation energy and magnetic ordering energy of a strongly magnetic transition-metal oxide such as NiO with the MAD protocol (PBEsol, no spin polarization) and with a spin-polarized DFT+U protocol, then compare a MAD-trained universal potential against a model trained on a dataset that includes such corrections; if the MAD-trained model's errors on these quantities exceed its typical accuracy on non-magnetic systems, the claim of universality across the full chemical space is falsified for that domain. Alternatively, test for selection bias by recomputing the discarded non-converged MC3D-random structures (55% convergence rate) with more robust settings; if their energies are systematically higher than the surviving structures, the dataset's advertised high-energy diversity is partly an artifact of filtering.","supporting_citations":[{"cited_title":"Huber, M","cited_arxiv_id":null,"evidence_quote":"Source of the bulk crystal seed structures for the MC3D and derived subsets."},{"cited_title":"Cordova, E","cited_arxiv_id":null,"evidence_quote":"Source of the curated molecular crystals used in the SHIFTML-molcrys subset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the neutral molecular fragments used in the SHIFTML-molfrags subset."},{"cited_title":"Mindless","cited_arxiv_id":null,"evidence_quote":"Provides the 'mindless' aggressive-distortion philosophy that MAD adapts to periodic and universal settings."},{"cited_title":"Ceriotti, G","cited_arxiv_id":null,"evidence_quote":"Sketch-map method used to build the low-dimensional structural cartography for comparing datasets."}],"review_version":2}