{"id":"1cf77d9f-8704-4b8b-8c16-3c07db1d31c8","arxiv_id":"2504.19578","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"LAMBench evaluates ten large atomistic models on out-of-distribution accuracy, property prediction, fine-tuning, speed, and stability, finding a large gap to a universal potential and DPA-3.1-3M at the top.","lead":"This paper introduces LAMBench, a benchmark that tests ten large AI models for simulating atoms and molecules on how well they generalize, adapt, and run in practice. A generalist might read it to see which materials-science foundation models actually work outside their training data and where the field stands.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"OOD status of the twelve force-field test sets is unverified: no overlap check against the training corpora in Table S-2 is reported, and the Discussion concedes test cases may become in-distribution, so DPA-3.1-3M's top ranking may partly reflect training coverage.","rationale":"The reader's weakest assumption is exactly the same load-bearing concern: the OOD status of the test sets is asserted but not verified, and this can confound the comparative leaderboard. I agree with that identification. The concern is real and should be disclosed, but it does not by itself overturn the central 'gap to the universal PES' conclusion, because even the best model, including DPA-3.1-3M-DomainMatch, exhibits large errors in catalysis and molecules. The main risk is the comparative ranking and the strength of the 'substantially greater generalizability' phrasing, not the existence of a universality gap. The paper is otherwise careful: the benchmark is open-sourced, the dimensionless metrics are explicitly defined in Section IV D, the domain-wise inorganic ranking correlates with Matbench Discovery, and the multi-fidelity and conservativeness findings are internally consistent. Therefore the reader's CONDITIONAL verdict is appropriate: the paper should be accepted only with a leakage analysis (or a clear statement that OOD is defined functionally rather than distributionally) and with any necessary softening of the ranking claim. This does not move the verdict, so I recommend UNCHANGED.","tokens_in":32317,"tokens_out":5320,"duration_ms":52717,"concrete_test":"For each of the twelve force-field test sets in Table I and each model training corpus in Table S-2 (plus MPtrj, OMat24, SPICE2, OC20/OC22 for the non-DPA models), quantify overlap using structure matching (e.g., pymatgen StructureMatcher on the primitive/relaxed cells) and a local-environment fingerprint nearest-neighbor search (e.g., SOAP or ACE descriptors) to flag near-duplicate or closely related configurations. Then recompute M_FF for every model after excluding test configurations whose nearest training neighbor falls below a pre-specified similarity threshold. If DPA-3.1-3M's M_FF increases by more than about 0.02 relative to competitors, or if its rank changes, the claim of 'substantially greater generalizability' must be weakened or qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparative claim that DPA-3.1-3M shows 'substantially greater generalizability compared to other LAMs' (Section II B) rests on the untested premise that the twelve force-field test datasets in Table I are out-of-distribution for every benchmarked model. Section II A defines OOD as 'downstream datasets designed to address specific scientific challenges,' but no overlap check against the training sets in Table S-2 (OpenLAM, MPtrj, OMat24, SPICE2, OC20/OC22, etc.) is reported. This matters concretely: ANI-1x is explicitly the training data of the ANI-1x potential, while OpenLAM contains SPICE2, Yang2023ab, OC20M, OC22, and OMat24, which cover the same chemical and catalytic domains as several of the test sets. Distributional overlap does not require exact duplicate frames; similar molecules, surfaces, or local environments can already make a test set partially in-distribution. The Discussion itself concedes that 'some OOD generalizability test cases' may become in-distribution and will need replacement. If DPA-3.1-3M has higher overlap with these test sets than its competitors, its leading M_FF is inflated by training coverage rather than by a fundamentally more universal representation. The headline 'gap to the universal PES' is less affected, because even the best model shows large errors, but the leaderboard ranking and the 'substantially greater' wording are directly confounded.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces LAMBench, an open-source benchmarking system for Large Atomistic Models (LAMs), and uses it to evaluate ten LAMs released before August 1, 2025. The benchmark measures three capabilities: generalizability (force-field predictions on twelve datasets across inorganic, molecular, and catalytic domains, plus property-calculation tasks), adaptability (fine-tuning on eight Matbench regression tasks), and applicability (inference efficiency and NVE stability). Results are aggregated into dimensionless metrics normalized by a dummy model, and a leaderboard is presented in which DPA-3.1-3M ranks first in generalizability. The authors conclude that there remains a substantial gap between current LAMs and an ideal universal potential energy surface, and they argue for cross-domain training, multi-fidelity support, and conservative/differentiable models.","tokens_in":32674,"tokens_out":6133,"duration_ms":64491,"significance":"If the reported results hold, LAMBench is a valuable community resource: it is open-sourced, modular, and accompanied by an interactive leaderboard; the metric definitions are explicit and anchored to a transparent dummy-model baseline; and the study covers a broader range of capabilities than most existing benchmarks. The main substantive findings—that no current LAM performs as a universal simulator, that domain-specific models still beat LAMs on their own domains, and that non-conservative models trade stability for speed—are plausible and practically relevant. The strengths include machine-checkable workflow automation, clear computational details for relabeled datasets, and the inclusion of efficiency and stability alongside accuracy. However, the central comparative claims about out-of-distribution generalizability and the leaderboard ordering rest on unsupported assumptions about dataset disjointness and on metrics that are sensitive to a small number of runs.","major_comments":[{"comment":"The label 'OOD' is asserted without a leakage check. Section II A defines OOD generalizability as performance on datasets whose distribution is distinct from training data, but the manuscript reports no analysis of overlap between the twelve force-field test sets (Table I) and the training corpora of the benchmarked models (Table IV and Table S-2). This is not a formal concern only: OpenLAM contains SPICE2, Yang2023ab, OC20M, OC22, and OMat24, which overlap in chemical and catalytic space with the molecular and catalysis test sets; ANI-1x itself is a training dataset for parts of the ANI family. If DPA-3.1-3M has higher overlap with these test sets than its competitors, its leading M_FF is inflated by training coverage, and the claim of 'substantially greater generalizability' (Section II B) is confounded. The Discussion concedes that some OOD cases may become in-distribution, but no quantitative support is given. Please add per-model and per-dataset overlap or distributional-similarity checks, or qualify the OOD and ranking claims accordingly.","section":"II A and Table S-2"},{"comment":"The instability metric M_IS is dominated by a single failed NVE simulation for DPA-3.1-3M. Equation (7) averages over nine structures, and Eq. (6) assigns a penalty of 5 to any failed run; the text states that DPA-3.1-3M's relatively high value of 0.572 is due to one failure, and that using the OMat24 task head reduces the instability metric to zero. Because M_IS is a component of the leaderboard, the applicability ordering of DPA-3.1-3M relative to models with M_IS=0 is effectively an artifact of a one-in-nine event. Please report per-structure instability values in the main text, show the sensitivity of the leaderboard to removing or reweighting failed runs, and consider a metric that is less sensitive to a single failure.","section":"II B, Table II, and IV E (Eqs. 6-7)"},{"comment":"The adaptability conclusion relies on unconverged runs. The footnote to Table III states that for the MP Eform and MP Gap tasks 'the reported accuracy may not reflect the fully converged results due to insufficient training epochs.' The claim that better force-field generalizability translates into better adaptability is based on comparing DPA-3.1-3M and DPA-2.4-7M across all eight tasks; two unconverged tasks weaken that inference. Please either complete these runs, provide learning curves or convergence diagnostics, or restrict the adaptability claim to the tasks that are demonstrably converged.","section":"Table III and Sec. II B (adaptability)"},{"comment":"No uncertainty estimates or significance tests are provided for any leaderboard metric. Several adjacent entries are close (e.g., Orb-v3 M_FF=0.215 vs DPA-2.4-7M=0.241, or MACE-MPA-0=0.308 vs SevenNet-l3i5=0.326), yet the leaderboard implies an exact ordering without any indication of run-to-run or test-set variability. At a minimum, bootstrap confidence intervals over test frames or repeated evaluations would establish which differences are statistically meaningful.","section":"Table II (general)"}],"minor_comments":[{"comment":"The training-set entry for MatterSim-v1-5M is listed as 'MattterSim'; this appears to be a typo for 'MatterSim'.","section":"Table IV"},{"comment":"The sentence comparing Matbench Discovery rankings contains 'DPA2-2.4-7M' instead of 'DPA-2.4-7M'; please fix this typo.","section":"Section II B"},{"comment":"The sentence beginning 'The energies, interatomic forces and virials labels of the Lopanitsyna2023Modeling and Mazitov2024Surface data on were obtained' contains a grammatical error ('data on were obtained'); it should read 'data were obtained'.","section":"Section IV A"},{"comment":"Reporting M_IS values as 0.000 may be misleading because the metric is a log-ratio against a tolerance; readers cannot tell whether this indicates zero drift or drift below the tolerance. Please consider reporting the underlying energy-drift values, as in Table S-8, alongside the dimensionless metric.","section":"Table II"},{"comment":"The definition of OOD as 'downstream datasets designed to address specific scientific challenges' is operational but weak; in addition to the overlap analysis requested above, a more formal statement of the distributional distance used (or a citation to one) would improve reproducibility.","section":"II A"}],"recommendation":"major_revision","confidential_remarks":"Several authors are developers of the DPA models and of the DeepMD-kit/OpenLAM ecosystem, and the paper's interactive leaderboard is hosted on the OpenLAM page. This is a clear potential conflict of interest. I recommend requesting a formal competing-interests statement and, if possible, an independent check of the DPA entries before publication. The scientific claims are otherwise plausible and the framework is useful, but the OOD and leaderboard claims need the revisions described above."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nLAMBench is the first integrated benchmark I've seen that puts ten released LAMs on the same cross-domain platform: force-field accuracy on twelve OOD datasets, property calculations, fine-tuning adaptability, and MD stability/efficiency, with the code and leaderboard open. That is a real contribution. The dimensionless relative-error metric with a dummy-model baseline is transparent and sensible, and the relabeling of five datasets to PBE is a thoughtful attempt to control for XC mismatch. The measurements themselves are reproducible; the paper ships the toolkit. Credit where earned: this will be a useful platform for the MLIP community.\n\nThe central finding, that even the best current LAMs are still far from a universal PES, is supported by the reported numbers and is robust. DPA-3.1-3M's M_FF of 0.175 versus dummy 1.0, and the large errors on Catalysis barriers, make that point convincingly.\n\nWhere the paper is softer: the OOD claim. \"OOD\" is defined as \"downstream datasets designed to address specific scientific challenges,\" but there is no reported overlap check against the training corpora listed in Table S-2. The Discussion itself concedes some OOD test cases may become in-distribution. This matters most for the leaderboard ranking, specifically the \"substantially greater generalizability\" wording for DPA-3.1-3M. It is reasonable to think multitask training helps, but if OpenLAM's SPICE2, OC20/22, and OMat24 contents overlap with the test distributions more than other models' training sets do, the gap is partly coverage. The ANI-1x test-set entry is less relevant than the stress-test suggests, since ANI-1x is not in OpenLAM, but the broader point stands.\n\nTwo smaller caveats: leaderboard values have no uncertainty estimates, and the instability metric for DPA-3.1-3M is dominated by one failed MD run (they disclose this and show an alternate head fixes it). Adaptability is only tested for DPA models, so the cross-model comparison there is thin. And several authors are DPA developers; that doesn't invalidate the measurements, but a COI statement and a training-overlap analysis should be added before this becomes a community standard.\n\nVerdict: worth a serious referee. I would send it to review, and I would ask for the leakage analysis and uncertainties before acceptance. The headline message survives; the ranking claim needs more evidence.\n\nBest.","headline":"A genuinely useful, open benchmark for large atomistic models whose headline gap-to-universal claim holds up, but the top-rank claim needs a leakage check given the authors' own models lead the board.","tokens_in":33204,"tokens_out":2871,"would_cite":true,"duration_ms":27587,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"No current atomistic AI model behaves as a universal simulator, a new ten-model benchmark finds.","keywords":["large atomistic models","machine learning interatomic potentials","benchmark","potential energy surface","generalizability","molecular dynamics","density functional theory","foundation models"],"falsifier":"Search each of the twelve test datasets against the ten models' training corpora for configurations with the same chemical composition and near-identical local environments, then re-run the force-field task after removing overlapping frames. If a top-scoring model loses its lead, the out-of-distribution claim is refuted; if the rankings are unchanged, the generalizability ranking stands.","tokens_in":32141,"feed_emoji":"⚛️","tokens_out":7662,"duration_ms":71609,"temperature":0.7,"pith_summary":"The paper sets out to measure whether large atomistic models (LAMs) are approaching a universal potential energy surface: the single energy function that density functional theory, in principle, defines for any arrangement of nuclei. To do this it introduces LAMBench, a benchmarking system that scores ten LAMs released before August 2025 on three capabilities: generalizability to out-of-distribution systems, adaptability to property-prediction tasks, and applicability in real simulations. The headline finding is a substantial gap: no model comes close to the ideal universal surface, and domain-specific models remain more accurate inside their own domains. The paper argues that closing the gap requires multi-domain pretraining, inference-time multi-fidelity support, and conservative, differentiable models.","feed_headline":"Benchmark: no atomistic AI model is yet a universal simulator","feed_subtitle":"Ten large atomistic models tested across three domains still trail domain-specific models, a new benchmark reports.","key_machinery":"LAMBench is a modular workflow paired with dimensionless error metrics. The central object is the ratio of a model's raw error to a dummy baseline that predicts energy solely from the chemical formula; after truncation, log-averaging over datasets, and weighting by prediction type, it yields $\\bar{M}^m_{\\mathrm{FF}}$ for force-field tasks and $\\bar{M}^m_{\\mathrm{PC}}$ for property-calculation tasks. Efficiency is measured as normalized inverse inference time, and stability as the log-scale magnitude of total-energy drift in 10 ps NVE simulations. The workflow automates job submission, result aggregation, and leaderboard updates so new models and tasks can be added.","core_discovery":"The paper claims that no currently released LAM approximates the universal potential energy surface well enough to serve as an out-of-the-box simulator. Running ten models without fine-tuning on twelve downstream force-field datasets across inorganic materials, catalysis, and molecules, it finds the best generalist, DPA-3.1-3M, still records a dimensionless force-field error of 0.175 and a property-calculation error of 0.322 on a scale where 0 is perfect and 1 is a chemistry-formula-only dummy. DPA-3.1-3M's lead is attributed to multi-task training on datasets spanning several domains. The same measurements show domain-specific models beating generalists on their home territory, non-conservative models drifting badly in long molecular-dynamics runs, and catalysis transition states as the clearest shared weakness.","pith_inferences":["The out-of-distribution ranking would be confounded if any of the twelve test sets overlaps a model's training corpus; the paper reports no near-duplicate check, so the leaderboard is provisional until overlap is examined.","The dimensionless relative-error metric rewards models on high-variance datasets, so the headline ordering is partly a statement about metric choice rather than absolute physical accuracy.","If LAMBench becomes standard, it will create pressure to add transition-state and hybrid-functional molecular data to pretraining corpora, likely shifting data acquisition toward domains the benchmark exposes as weak."],"forward_implications":["If the out-of-distribution scores are taken at face value, no LAM released before August 2025 is a drop-in universal simulator; deployments should expect accuracy losses outside each model's training domain.","Domain-specific models remain the accuracy ceiling in their own domains: a molecular specialist beats the best LAM on torsion and conformer energy profiles, and OC20-trained models beat all LAMs on reaction-barrier prediction.","Multi-task training with cross-domain data is the ingredient most strongly associated with generalizability, supporting the paper's call for more balanced training-data distributions.","Conservativeness and differentiability are requirements, not options: non-conservative force prediction made Orb-v2 the fastest model but unstable in NVE simulations and poor at phonon and elasticity calculations.","Multi-fidelity modeling is needed because LAMs trained at the PBE exchange-correlation level cannot be compared directly with CCSD(T) or RPBE references; matching task heads cut molecular property error from 0.31 to 0.10."],"supporting_citations":[{"why":"Supplies the DPA multi-task training strategy whose pretraining on many domain datasets is credited for DPA-3.1-3M's generalizability.","marker":"[4]"},{"why":"Defines DPA-3.1-3M, the model that achieves the lowest errors in both force-field and property-calculation tasks.","marker":"[61]"},{"why":"Defines the non-conservative Orb models used as the efficiency-versus-stability contrast in the applicability tests.","marker":"[28]"},{"why":"Establishes the earlier finding that conservative, smooth models are needed for physical property prediction, which LAMBench confirms.","marker":"[27]"},{"why":"Supplies the Matbench Discovery benchmark whose inorganic-materials ranking correlates with LAMBench's, used as a coherence check.","marker":"[25]"},{"why":"Provides the OC20 dataset used for catalysis pretraining and the OC20NEB transition-state test set.","marker":"[65]"},{"why":"Provides the multi-fidelity training approach that the paper cites as the way to handle exchange-correlation functional mismatch across domains.","marker":"[68]"},{"why":"Provides the implementation that currently limits adaptability tests to the DPA models.","marker":"[55]"},{"why":"Supplies the MPtrj training dataset that most of the benchmarked single-domain LAMs are trained on and that anchors the inorganic-materials comparisons.","marker":"[13]"}],"fun_headline_variants":["Atomistic AI models fail universal test","No universal atomistic simulator yet, benchmark finds","Benchmark: LAMs still far from universal","Atomistic models trail domain-specific rivals"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ranking assumes that the twelve force-field test sets are genuinely outside every model's training data, but the paper reports no overlap check against the models' training corpora, so any hidden overlap would inflate the affected model's generalizability score.","fun_headline_variants_meta":{"raw":{"variants":["Atomistic AI models fail universal test","No universal atomistic simulator yet, benchmark finds","Benchmark: LAMs still far from universal","Atomistic models trail domain-specific rivals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00052,"raw_usage":{"total_tokens":2540,"prompt_tokens":986,"completion_tokens":1554,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":602,"completion_tokens_details":{"reasoning_tokens":1498}},"tokens_in":602,"tokens_out":1554,"duration_ms":12744,"temperature":1.0,"reasoning_tokens":1498,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:48:46.669488+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Search each of the twelve test datasets against the ten models' training corpora for configurations with the same chemical composition and near-identical local environments, then re-run the force-field task after removing overlapping frames. If a top-scoring model loses its lead, the out-of-distribution claim is refuted; if the rankings are unchanged, the generalizability ranking stands.","supporting_citations":[{"cited_title":"Chanussot, A","cited_arxiv_id":null,"evidence_quote":"Provides the OC20 dataset used for catalysis pretraining and the OC20NEB transition-state test set."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the multi-fidelity training approach that the paper cites as the way to handle exchange-correlation functional mismatch across domains."}],"review_version":1}