{"id":"988e33a3-f5b8-4249-b773-5f39bdbfb43e","arxiv_id":"2607.18092","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Zero-shot r2SCAN energies from the PET-OMATPES MLIP cut GGA formation-energy errors by >40%, and delta-learning on the MLIP's latent features pushes the MAE below 50 meV/atom.","lead":"This paper adds formation-energy data to the MC3D crystal database and shows that a machine-learned foundational potential, plus small correction models trained on its internal features, brings calculated formation energies to about 50 meV/atom from experiment. The practical payoff: existing GGA-level databases can get near-meta-GGA stability estimates without running any new DFT calculations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The <50 meV/atom claim rests on a filtered, experimentally characterized subset with random splits; transfer to the full MC3D database is not demonstrated.","rationale":"The reader's weakest assumption—that the filtered experimental reference set and random 80/20 splits estimate performance on the full MC3D database—is also the most load-bearing condition for the central claim. The paper's headline MAE of 49 meV/atom is obtained on experimentally characterized compounds that passed several quality filters and are split randomly, so the test compounds are chemically similar to training compounds. The intended application is to correct the entire MC3D database, including many compounds without experimental references; the reported evaluation cannot establish that the learned corrections transfer. The paper's own SI acknowledges split-dependent artifacts and calls for larger experimental datasets, supporting the concern. I agree with the reader's identification of this assumption, and the recommended check—a chemical-system hold-out—would directly quantify whether random splits are misleading. This does not change the conditional verdict: the evidence is suggestive but incomplete, and the requested validation is needed before the claim is taken as definitive.","tokens_in":29361,"tokens_out":7434,"duration_ms":437966,"concrete_test":"Run a compositional leave-out analysis: group compounds by chemical system (same set of elements), assign whole systems to training or test folds, train the KRR-LAP-LF model with α=0.1 on 80% of systems, and evaluate on the held-out systems, repeated over 30 folds. Compare the mean test MAE to the reported 49 meV/atom. If the system-held-out MAE exceeds ~70 meV/atom (or is more than ~20 meV/atom higher than the random-split value), the random 80/20 split overestimates transferability to the full MC3D database.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that GGA/MC3D formation energies can be corrected to ~50 meV/atom, comparable to experimental uncertainty—rests on a test set of 1384 compounds that survive four exclusion filters (elemental phases; uncertainty >10%; DFT–experiment disagreement >0.5 eV/atom; cross-source disagreement >150 meV/atom) and are split randomly, stratified only by number of elements. This evaluation does not emulate the intended use case: correcting the full MC3D database, where most entries lack experimental references and include metastable or less-characterized chemistries. Random splits keep the test distribution close to the training distribution; a composition-based or chemical-system hold-out would be a stricter test. The paper itself acknowledges in SI §S4.D that 'few residual structure-specific artifacts persist depending on which structures appear in the training set' and that larger experimental reference datasets are needed, effectively conceding that the 1384-compound set is the bottleneck. If the filters preferentially retain well-measured, DFT-friendly compounds, the reported 49 meV/atom MAE is a lower bound for true database-wide performance. Data and code are not yet released, so the magnitude of this bias cannot currently be audited.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces formation energies for the Materials Cloud three-dimensional crystals database (MC3D), compares them with OQMD and Materials Project, and proposes a two-step correction towards experimental formation enthalpies: first, zero-shot r2SCAN energies from the PET-OMATPES foundational MLIP, and second, learned delta corrections using classical ML models (RF, GPR, KRR) with compositional (Magpie) or MLIP latent features. On a filtered experimental reference set of 1384 PBEsol compounds, the best model (KRR-LAP with latent features) reaches a test MAE of 49 meV/atom for the r2SCAN target, down from ~160 meV/atom for uncorrected GGA, and the paper analyses stability flip rates and regularization trade-offs. The authors also validate the MLIP zero-shot energies against the Alexandria database.","tokens_in":29709,"tokens_out":8997,"duration_ms":95583,"significance":"If the <50 meV/atom accuracy transfers beyond the filtered experimental subset, the method provides a practical route to correct existing GGA databases to near-experimental accuracy without additional DFT. The paper is transparent about its evaluation: 30 train-test splits, 5-fold CV, two feature sets, four model classes, and a comparison with FERE. The introduction of MC3D formation energies is a useful resource. However, the central claim rests on a small, filtered, experimentally characterized set with random splits, and the FERE comparison is confounded by in-sample fitting; these issues need to be resolved before the general claim is supported.","major_comments":[{"comment":"The evaluation protocol uses 30 random 80/20 splits stratified only by element count on the 1384-compound filtered experimental subset. This subset passes four exclusion filters (elemental phases; uncertainty >10%; DFT–experiment disagreement >0.5 eV/atom; cross-source disagreement >150 meV/atom). Random splits keep test structures close to training structures, so the reported test MAE is an in-distribution estimate. The central claim of the paper—that GGA/MC3D formation energies can be corrected to <50 meV/atom—is used to motivate correcting the full MC3D database, for which most entries lack experimental references. The paper does not provide a composition-based or chemical-family hold-out, nor a validation on compounds excluded by the filters. The SI (§S4.D) itself concedes that 'few residual structure-specific artifacts persist depending on which structures appear in the training set","section":"Methods, 'Reference data a'; 'Machine learning models and training'"},{"comment":"The FERE corrections are fitted on the full dataset (1511 compounds, MC3D PBE version; Table S2) and then evaluated on the same test splits used for ML, which are disjoint from the ML training data. Thus the FERE results in Fig. 6 and Figs. S15–S19 are in-sample (or at least use the test labels for fitting), while the ML results are out-of-sample. Moreover, the FERE corrections are fitted on PBE data but applied to the PBEsol baseline in the comparison. This confounds the claimed superiority of the ML approach. Refit FERE within each training split (or a nested CV) and on the same functional/target as the ML models before comparing.","section":"Comparing with existing approaches; SI §S5"},{"comment":"The computation of energies above the convex hull and stability flip rates is not described. It is unclear whether the hull is constructed from all MC3D structures or only from the 1384-compound experimental subset. If only the subset is used, the hull is incomplete and the flip rates (Fig. 3b, 4, 5) may be artifacts of a truncated hull. Specify the hull construction, including how corrected formation energies of non-experimental MC3D entries are handled.","section":"Balance accurate formation energies...; Methods"},{"comment":"The final KRR-LAP-LF model uses α=0.1 selected based on the test-set trade-off in Fig. 4a. This is a post-selection choice; the reported MAE of 49 meV/atom is therefore not a purely out-of-sample estimate. Use a validation set for α selection or report the MAE across all α values and discuss the selection bias.","section":"Machine learning models and training; Fig. 4"}],"minor_comments":[{"comment":"The phrase 'reducing the mean absolute error by more than 40% relative to GGA' should specify which comparison (zero-shot vs. pure DFT) and which GGA (PBE/PBEsol), since the 40% reduction refers to the zero-shot step and the later ML correction is additional.","section":"Abstract"},{"comment":"The claim that the MAE is 'comparable to the experimental uncertainty itself' is somewhat overstated: the achieved 49 meV/atom is above the commonly cited ~25 meV/atom typical uncertainty, although within the 70 meV/atom range across databases. Qualify the statement accordingly.","section":"Discussion / Fig. 3a"},{"comment":"The main-text figure uses the test split with the best ML MAE; the worst split is only in the SI. Label the figure as 'best split' or show both splits in the main text to avoid cherry-picking concerns.","section":"Fig. 6"},{"comment":"The statement says 'will be made publicly available upon publication'. For a fully reproducible claim, provide a repository link or temporary access during review.","section":"Data/Code Availability"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest about limitations and the data/code promise is good. The main issues—generalization to the full database and the FERE comparison protocol—are fixable with additional validation and refitting. I recommend major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one. The genuinely new result is that PET-OMATPES backbone latent features, fed into a regularized KRR delta model, beat Magpie compositional features on both MAE and stability-flip rate, reaching test MAEs around 49-52 meV/atom on the r2SCAN and PBEsol targets. The zero-shot r2SCAN MLIP part is a clean confirmation of an established recipe rather than a surprise, but it is done carefully. The paper is transparent: 30 stratified splits, hyperparameters chosen by CV, two kernels, FERE and Alexandria comparisons, and a thoughtful analysis of the regularization/flip-rate trade-off. The MC3D formation-energy data itself is a useful new resource.\n\nThe soft spot is the evaluation. The headline numbers come from 1384 compounds that survive four filters, and the splits are random, stratified only by element count. That gives a lower bound for how the correction will behave on the full MC3D database, which includes less-characterized chemistries and metastable structures. The paper itself concedes in SI S4.D that residual structure-specific artifacts persist and that larger experimental reference datasets are needed. Data and code are not yet released, so the magnitude of any filter bias cannot currently be audited. These are real limitations, but they are not fatal: the paper mostly claims a benchmark on the available experimental reference, and the random-split MAE is what it is. I would like to see a compositional or chemical-system hold-out before believing the 50 meV/atom number transfers far out of distribution.\n\nThe citation pattern looks fine—Gong et al. and Adhikari et al. are the right prior work, and the Alexandria comparison is a sensible external anchor. The honest reading is that this is a solid, useful engineering result with a clear caveat about evaluation scope. It deserves a serious referee; if I were the editor I would send it out. I would ask reviewers to verify the code and data once released and to probe out-of-distribution behavior.","headline":"Solid, transparent benchmark of latent-feature delta-learning with PET-OMATPES, hitting <50 meV/atom on a filtered experimental set; the main caveat is evaluation scope.","tokens_in":30139,"tokens_out":1880,"would_cite":true,"duration_ms":21568,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A delta-learning model built on the latent features of a foundational interatomic potential corrects DFT formation energies to within 49 meV/atom of experiment, matching experimental uncertainty.","keywords":["formation energies","delta learning","foundational MLIP","latent features","r2SCAN","GGA corrections","thermodynamic stability","kernel ridge regression"],"falsifier":"Collect new experimental formation enthalpies for chemistries and structure families underrepresented in the current reference set (for example, chalcogenides beyond oxides or compounds with large DFT-experiment disagreements) and evaluate the trained delta model on them. An MAE clearly above 50 to 75 meV/atom, or a stability flip rate outside the 3 to 4 percent band, would show that the corrections do not transfer beyond the reference distribution.","tokens_in":29295,"feed_emoji":"⚛️","tokens_out":10716,"duration_ms":133434,"temperature":0.7,"pith_summary":"The paper sets out to show that formation energies in GGA-level crystal-structure databases can be pushed to near-experimental accuracy without any new density-functional calculations. Its recipe is two-stage: evaluate a foundational machine-learning interatomic potential trained on r2SCAN reference data at existing GGA-relaxed geometries (this alone cuts the mean absolute error versus experiment by more than 40%), then train a small kernel-ridge model on the potential's latent structural features to learn and subtract the remaining residual. On the merged experimental reference set, the best model reaches a mean absolute error below 50 meV/atom, which the authors equate with experimental uncertainty itself. The paper also introduces formation-energy values for its underlying database and shows that differences between public DFT databases are dominated by their empirical correction schemes rather than by DFT codes or settings. A sympathetic reader would care because formation energies and convex-hull stability are the standard filter in computational materials discovery, and this recipe promises to upgrade large existing databases at almost no cost.","feed_headline":"Delta learning hits 49 meV/atom in DFT energy correction","feed_subtitle":"A two-step ML correction brings computed formation energies in line with measured values, no costly DFT reruns.","key_machinery":"The load-bearing mechanism is delta-learning on the backbone latent features of the PET-OMATPES foundational machine-learning interatomic potential, a representation extracted after the message-passing layers and before the readout layers, which the paper finds to be richer than the final-layer features. A kernel ridge regression model with a tuned Laplacian kernel learns the difference between the potential's zero-shot r2SCAN formation energy and the experimental formation enthalpy, and that learned correction is added on top. Regularization strength plays a double role: it controls both overfitting of the small reference set and the distortion of relative phase stability, letting the model","core_discovery":"On the authors' own terms, the central discovery is that the latent-space representation of a foundational machine-learning interatomic potential is a better feature set for delta-learning DFT-to-experiment residuals than composition-only descriptors. Kernel ridge regression with a Laplacian kernel trained on those latent features predicts the residual between the potential's zero-shot r2SCAN formation energies and experimental formation enthalpies with a held-out MAE of 49 meV/atom, and the same approach applied to the directly corrected GGA target reaches 52 meV/atom. The structure-aware latent features lower the MAE and, once regularization is tuned, lower the stability flip rate compared","pith_inferences":["If the error transfers beyond the curated reference distribution, the same two-step recipe could re-score entire GGA databases into a meta-GGA-quality stability layer for the cost of one MLIP evaluation and a small kernel model.","The structure-aware latent features are not limited to formation energies; the same delta-learning construction could be pointed at other experimental targets such as band gaps, magnetic ordering temperatures, or adsorption energies wherever a labelled reference set exists.","The corrections cluster into a few discrete modes across train-test splits, which suggests that a committee or majority-vote ensemble of delta models would be a more stable production choice than any single split; this is a testable extension implied by the paper's variance analysis."],"forward_implications":["Existing GGA databases can have their formation energies brought to about 50 meV/atom agreement with experiment without any additional DFT, making their stability filters materially more reliable.","Pairing PBEsol-relaxed geometries with r2SCAN-level MLIP energies is almost as accurate as fully relaxing with the MLIP, extending a known DFT practice to foundational potentials.","Structure-aware latent features outperform composition-only features for delta corrections, lowering both the formation-energy MAE and the stability flip rate.","With tuned regularization, the learned correction keeps stability flip rates at 3 to 4 percent at typical energy windows above the convex hull, with most flips inherited from the physically grounded MLIP rather than introduced by the correction model.","The ML correction improves more compounds than fitted elemental-reference corrections and degrades the compounds it misses less severely."],"fun_headline_variants":["MLIP latent features cut DFT energy error to 49 meV/atom","Delta-learning with MLIP features reaches 49 meV/atom MAE","Structure-aware features from MLIP correct DFT to 49 meV/atom","MLIP latent features beat composition in DFT energy correction","Correcting DFT to experimental accuracy with MLIP latent features"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the filtered set of experimental formation enthalpies used for training and testing (2,356 compounds after filtering, 1,384 after merging with the studied database) is a fair and unbiased sample of a GGA database's contents, so the sub-50 meV/atom error measured on that set transfers to new materials; if the reference set skews toward well-measured, easier compounds, the claim does not generalize.","fun_headline_variants_meta":{"raw":{"variants":["MLIP latent features cut DFT energy error to 49 meV/atom","Delta-learning with MLIP features reaches 49 meV/atom MAE","Structure-aware features from MLIP correct DFT to 49 meV/atom","MLIP latent features beat composition in DFT energy correction","Correcting DFT to experimental accuracy with MLIP latent features"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000667,"raw_usage":{"total_tokens":2919,"prompt_tokens":822,"completion_tokens":2097,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":566,"completion_tokens_details":{"reasoning_tokens":2015}},"tokens_in":566,"tokens_out":2097,"duration_ms":15537,"temperature":1.0,"reasoning_tokens":2015,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T16:00:48.786374+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect new experimental formation enthalpies for chemistries and structure families underrepresented in the current reference set (for example, chalcogenides beyond oxides or compounds with large DFT-experiment disagreements) and evaluate the trained delta model on them. An MAE clearly above 50 to 75 meV/atom, or a stability flip rate outside the 3 to 4 percent band, would show that the corrections do not transfer beyond the reference distribution.","supporting_citations":[],"review_version":1}