{"id":"6b5eb559-c462-4a9e-9744-ba1145c36fd6","arxiv_id":"2607.22903","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Among seven DFT-based machine-learning force fields for water, RPBE-D3 best matches experimental density, structure, entropy, diffusion, and viscosity, and translational and orientational entropy scale linearly across models.","lead":"Researchers trained seven machine-learning force fields for liquid water, each from a different quantum chemistry functional, and compared their structure, entropy, and flow properties. They found predicted water behavior depends strongly on the functional, with dispersion-corrected RPBE-D3 matching experiments best.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Functional ranking may rest on unvalidated ML-FF training convergence; R2SCAN+rVV10's structural failure suggests training artifacts.","rationale":"The reader's weakest assumption identifies exactly the point on which the paper's main comparison depends: all seven ML-FFs are assumed converged for their respective DFT functionals, but the training protocol is very short and no validation against held-out DFT data is provided in the main text. My reading of the manuscript confirms this is the least secure condition for the central claim. The peculiar failure of R2SCAN+rVV10 in Fig. 3 supports the suspicion that training artifacts, not XC physics, are influencing the results. Because the reader's verdict is already CONDITIONAL, my stress test does not move it; it strengthens the condition. I therefore recommend UNCHANGED, meaning the verdict stands as conditional acceptance pending independent validation of ML-FF convergence. No other concern—such as entropy binning, finite-size corrections, or the small sample for the scaling fits—appears as load-bearing as training convergence, since all models share those methodology features and the central ranking could still survive them. The proposed retraining test directly targets the weakest link.","tokens_in":16592,"tokens_out":5903,"duration_ms":69004,"concrete_test":"Retrain the R2SCAN+rVV10 and RPBE-D3 ML-FFs with the same on-the-fly protocol but extend training to at least 200 ps at 300 K after 50 ps equilibration, using physical hydrogen masses; then recompute gOO(r), excess entropy, self-diffusion, and viscosity. If R2SCAN+rVV10 then develops a well-defined second hydration shell, or if RPBE-D3's agreement ranking changes relative to SPC/E or other functionals, the reported functional hierarchy is not robust to training convergence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that RPBE-D3 is the most consistent water model among seven XC functionals, and that the observed spread in structure, entropy, and transport reflects XC physics. This requires each ML-FF to faithfully represent its parent DFT functional under production conditions. That requirement is weakly supported: training used only 50 ps of on-the-fly AIMD on 64 molecules, hydrogen masses set to 8 amu, and a 100–400 K heating ramp (Computational Details). The main text does not report force/energy RMSE or a held-out DFT test; it refers only to SI for training errors. The complete loss of a second hydration shell for R2SCAN+rVV10 (Fig. 3) is physically implausible for a meta-GGA with nonlocal vdW, which suggests an ML fitting or sampling failure rather than a genuine XC deficiency. If the best or worst model is not converged, the ranking and the fitted entropy–transport correlations (Figs. 7b, 8c) could be dominated by ML training artifacts. This is the load-bearing weak point.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains seven machine-learned force fields (ML-FFs) for water, each derived from a different DFT exchange-correlation functional (PBE, PBE-D3, PBE-TS, RPBE, RPBE-D3, R2SCAN+rVV10, vdW-DF-cx), then compares their predicted structure, excess entropy, and transport against experiment and against the classical SPC/E model. Using the six-dimensional pair correlation function, three-body angular distributions, tetrahedral order, and H-bond statistics, the authors report that RPBE-D3 gives the most consistent agreement with experiment, that translational and orientational excess entropies are linearly related, and that reduced diffusivity follows an exponential excess-entropy scaling relation. The paper also argues that SPC/E and RPBE-D3 converge to similar effective electrostatics and dispersion physics. The computational pipeline includes on-the-fly training in VASP, 512-molecule production runs, finite-size correction of diffusion, Green-Kubo viscosity with biexponential extrapolation, and entropy extrapolation to infinite sampling.","tokens_in":16906,"tokens_out":4302,"duration_ms":53567,"significance":"If the conclusions hold, the paper provides a practically useful benchmark of DFT-based ML-FFs for water and advances the use of structural/entropic descriptors to diagnose ML-FF quality. The study is well posed and mostly clearly executed: transport properties carry error bars, entropy is extrapolated to infinite sampling, and the authors explicitly discuss error cancellation. The data and analysis scripts are made available, which is a real strength. The central claim, however, rests on the assumption that each ML-FF faithfully represents its parent DFT functional. That assumption is not independently validated in the main text, and one of the functionals (R2SCAN+rVV10) produces a physically implausible loss of the second hydration shell, which could indicate a training artifact. Because the functional ranking is the paper's main conclusion, this issue is load-bearing and should be addressed before publication.","major_comments":[{"comment":"The central claim that the differences among the seven models reflect the underlying XC functional requires each ML-FF to be a converged surrogate for its parent DFT method. The training protocol is very short (64 molecules, 50 ps, H mass 8 amu, 100–400 K ramp) relative to the production state (512 molecules, 300 K, physical masses), and the main text does not report held-out force/energy RMSEs or a comparison against DFT test configurations in the production region. The complete absence of a second hydration shell for R2SCAN+rVV10 in Fig. 3 is difficult to attribute to the functional itself and is more plausibly the signature of a fitting or sampling failure. Without per-functional train/test errors and a convergence check (e.g., retraining with longer on-the-fly sampling and comparing the resulting RDFs), the RPBE-D3 ranking and the fitted entropy–transport correlations may be artifact","section":"Computational Details, Machine learning force fields; Fig. 3"},{"comment":"The linear sor–str relation and the exponential D*–stot relation are fitted to eight models with two empirical parameters each, and no uncertainty is reported for the fit parameters or R² values. The text states that the D*–stot agreement 'confirms' the excess-entropy scaling link. Since these are in-sample fits, they are empirical correlations, not independent validations. Adding bootstrap or leave-one-out errors and stating explicitly that the relations are fitted would make the claim proportionate. This is especially important because the abstract and conclusion present the linear sor–str relation as a central finding.","section":"Results & Discussion, Coupled Translational and Orientational Excess Entropy; Fig. 7(b); Fig. 8(c)"}],"minor_comments":[{"comment":"The dashed line is labeled as the experimental density, but the x-axis is a temperature ramp from 100 to 400 K. The experimental density is temperature-dependent; please either show the full experimental curve or restrict the label to the 300 K point.","section":"Fig. 2(a)"},{"comment":"The RDF RMSE values (ε) and H-bond numbers are reported without uncertainties across the independent trajectories, unlike the transport and entropy values. Since three independent NVT runs were used for the PCFs, error bars or a statement that the values are pooled single estimates would improve clarity.","section":"Fig. 3; Table 1"},{"comment":"The text says 'V ASPs on-the-fly-training scheme' and 'V ASP MD engine'; these are likely typographical artifacts of the VASP name. Also, the sentence about the Nosé-Hoover thermostat gives 'a Nosé-mass of 5 in VASP and 0.1 ps damping parameter in LAMMPS'; please specify the units and which thermostat the SPC/E runs used.","section":"Computational Details, ML-FFs"},{"comment":"The Born effective charges are described as 'isotropic' but the averaging procedure is not defined. Please state whether the values are the average of the diagonal components or an isotropic projection of the Born charge tensor.","section":"Results & Discussion, Classical and ML-FFs Converge through Effective Interactions; Table 2"},{"comment":"The entropy extrapolation to infinite sampling is essential to the reported values but is only mentioned in one sentence. If the SI is the only source, this is acceptable; otherwise, please add one sentence summarizing the extrapolation procedure in the main text.","section":"Supporting Information / Excess entropy"}],"recommendation":"major_revision","confidential_remarks":"The paper is well structured and the data/code availability is a strength. My main concern is internal validity: the functional ranking is only as reliable as the training convergence of each ML-FF, and the R2SCAN+rVV10 result raises a red flag. If the Supporting Information does not already contain per-functional train/test RMSEs against held-out DFT data, that is the decisive gap. The entropy–transport fits are presented as confirmatory when they are in-sample correlations; this should be softened or supported by cross-validation. I would encourage the editor to ask for the SI validation material to be made available to reviewers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a well-executed benchmark paper that will be useful to anyone running ML-FF production simulations of water. The genuinely new bits are the full 6D pair-correlation entropy decomposition across seven XC functionals and the clean linear sor–str relationship, plus the excess-entropy scaling of diffusivity across those models. The methods are standard and the transport numbers come with error bars and finite-size corrections. The RPBE-D3 result is credible and consistent with prior work on dispersion corrections.\n\nThe soft spot is exactly where the stress test points. Each ML-FF was trained on only 50 ps of on-the-fly AIMD with 64 molecules, a hydrogen mass of 8 amu, and a 100–400 K ramp. The main text gives no force/energy RMSE on a held-out DFT test set. That matters because the ranking is the whole point. The R2SCAN+rVV10 model completely loses its second hydration shell (Fig. 3), which is not a plausible physical prediction for that functional. That looks like an ML fitting or sampling failure, not a genuine XC deficiency. If one model is broken, the comparison is no longer a clean test of XC physics; it's partly a test of training convergence. The paper should either provide held-out validation for all seven or explicitly reframe the conclusions as being about these particular ML-FFs, not about the functionals per se.\n\nThe fitted relations in Figs. 7(b) and 8(c) lack uncertainty intervals, and with only eight points the R² values are not very informative. That is a minor issue, because the fits are clearly labeled as empirical, and the excess-entropy trend is consistent with established theory. Same for the viscosity fit parameters.\n\nCredit where due: the paper handles error cancellation honestly (RPBE's 'good' transport is explained as a density artifact), and the SPC/E comparison via Born effective charges is a nice touch. Nothing is oversold; this is a practical map rather than a claim about fundamental water physics.\n\nWho should read it: anyone choosing a water ML-FF for production, and people working on entropy–transport relations. It deserves a serious referee. The main requested revision would be a head-to-head DFT validation of the ML-FFs (even a short test set) and a clear discussion of the R2SCAN failure. With that, it would be a solid contribution.\n\nMy recommendation: send to peer review, conditional on adding that validation.","headline":"A practical functional benchmark for water ML-FFs; the ranking is probably right but needs a held-out DFT check before it can be trusted.","tokens_in":17341,"tokens_out":3573,"would_cite":true,"duration_ms":41413,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The density functional used to train a machine-learning water model controls whether the simulated liquid is ice-like or realistic; among seven functionals, dispersion-corrected RPBE-D3 comes closest to experiment across structure, entropy,","keywords":["water","machine learning force fields","exchange-correlation functionals","six-dimensional pair correlation function","excess entropy","self-diffusion","viscosity","SPC/E"],"falsifier":"Compare each trained machine-learning force field to direct density-functional single-point energies and forces on a held-out set of water configurations drawn from the production runs; if the model with the best structural agreement (RPBE-D3) is not also the most accurate surrogate, the functional ranking is at least partly a training artifact.","tokens_in":16532,"feed_emoji":"💧","tokens_out":6223,"duration_ms":59370,"temperature":0.7,"pith_summary":"The paper asks whether machine-learning force fields for liquid water inherit their quality from the density-functional approximation used to produce their training data. By training seven separate models with different exchange-correlation functionals and comparing the full six-dimensional pair correlation function, three-body structure, excess entropy, viscosity, and self-diffusion against experiment, it establishes that the functional choice changes water from a nearly ice-like tetrahedral network into a disordered, mobile liquid. Dispersion corrections are essential; translational and orientational excess entropy are linearly coupled; and the reduced self-diffusion of all models collapses onto one exponential excess-entropy scaling curve. RPBE-D3 gives the most consistent agreement with experiment across every property tested, while the classical SPC/E model behaves similarly because of comparable effective electrostatic charges and long-range dispersion physics. A sympathetic reader would care because machine-learning potentials are only as trustworthy as the electronic-structure level they imitate, and this work shows how to diagnose that trustworthiness from structure and entropy alone.","feed_headline":"RPBE-D3 water model best matches experiment across structure and flow","feed_subtitle":"Across seven DFT-trained water models, RPBE-D3 best matches measured structure, entropy, and diffusion.","key_machinery":"The central object is the full six-dimensional molecular pair correlation function g(r,ω), which records how the relative position and five orientational degrees of freedom of two water molecules are correlated. From it the paper obtains the oxygen-oxygen radial distribution function and the conditional orientational distribution; these are integrated into translational and orientational excess entropy using pair-correlation entropy expressions. The key identities are the linear excess-entropy relation and the excess-entropy scaling law D* = A exp(β stot/R) for the reduced diffusivity, which together connect measurable structure to viscosity and diffusion. A supporting diagnostic is the comp","core_discovery":"The paper's central claim is that the exchange-correlation functional used to generate training data propagates through a machine-learned water potential and controls whether the simulated liquid reproduces experiment. Evaluated through the full six-dimensional molecular pair correlation function, the predicted water structure varies strongly with functional: PBE-class functionals without dispersion over-structure water into an almost tetrahedral, ice-like network, whereas dispersion-corrected RPBE-D3 reproduces the measured radial and orientational correlations, including the population of interstitial, non-tetrahedral water molecules. The paper further claims a linear coupling between orie","pith_inferences":["Editorial inference: the linear translational-orientational entropy relation, if it holds for other hydrogen-bonded or tetrahedral liquids, would turn the cheaply computed radial distribution function into a screening tool for force-field quality, without requiring expensive six-dimensional sampling.","Editorial inference: because the ranking rests on 50-picosecond training runs with only 64 molecules, an untested possibility is that some functionals simply train harder; a held-out DFT validation of each surrogate would separate functional physics from training error.","Editorial inference: the RPBE-D3/SPC/E coincidence suggests a practical design rule for classical models — matching the Born effective charges of a dispersion-corrected functional may reproduce its liquid behavior without machine learning; this could be tested by re-parameterizing a classical model to RPBE-D3 charges."],"forward_implications":["Neglecting dispersion in the training functional over-structures water and produces ice-like hydrogen-bond networks; adding D3 or TS corrections moves the predicted radial and orientational structure toward experiment.","RPBE-D3 yields the closest density, oxygen-oxygen radial distribution function, excess entropy, viscosity, and self-diffusion to experiment among the seven functionals, while plain RPBE only appears accurate through cancellation of a low density against an over-structured liquid.","Orientational excess entropy dominates the total for every model, and it scales linearly with translational excess entropy, so the radial distribution function alone is a practical proxy for how well a water model captures structure.","Across all models, reduced self-diffusion follows the excess-entropy scaling law, so a model's diffusion and viscosity errors are directly traceable to its structural ordering.","The near-agreement of SPC/E and RPBE-D3 arises from similar effective electrostatics and similar long-range dispersion, not from identical construction."],"fun_headline_variants":["DFT choice dictates if ML water freezes or flows","RPBE-D3 tops seven machine-learned water models","Dispersion fixes ML water's over-structured ice","Which functional makes ML water real? RPBE-D3 wins"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The ranking of the seven functionals assumes that each 50-picosecond, 64-molecule training run (with heavy hydrogen atoms and a 100–400 K ramp) converged to a reliable potential for its functional; if some models are undertrained, the comparison reflects training protocol, not the physics of the functional.","fun_headline_variants_meta":{"raw":{"variants":["DFT choice dictates if ML water freezes or flows","RPBE-D3 tops seven machine-learned water models","Dispersion fixes ML water's over-structured ice","Which functional makes ML water real? RPBE-D3 wins"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00071,"raw_usage":{"total_tokens":2978,"prompt_tokens":634,"completion_tokens":2344,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":378,"completion_tokens_details":{"reasoning_tokens":2286}},"tokens_in":378,"tokens_out":2344,"duration_ms":15745,"temperature":1.0,"reasoning_tokens":2286,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T01:46:29.339056+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare each trained machine-learning force field to direct density-functional single-point energies and forces on a held-out set of water configurations drawn from the production runs; if the model with the best structural agreement (RPBE-D3) is not also the most accurate surrogate, the functional ranking is at least partly a training artifact.","supporting_citations":[],"review_version":2}