{"id":"697a71d5-6a11-4f76-98d6-e4a0a2c60e29","arxiv_id":"2607.13261","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A CCSD(T)-accurate machine-learned potential for aqueous toluene shows that popular force fields and hybrid DFT distort the balance between hydrophobic and π-interaction solvation.","lead":"Scientists trained a machine-learning model of water and toluene to match very accurate quantum chemistry calculations, then used it to watch how water molecules arrange around toluene. The simulations show that common force fields and density functional methods get the balance between hydrophobic and π-hydrogen-bonding interactions wrong.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Finite-cluster upfit may not span the bulk CCSD(T) ensemble; the FF/DFT comparison inherits this risk.","rationale":"The paper's most important assertions are benchmark-quality claims, so the weakest point is not any individual force-field/DFT comparison but the validity of the benchmark itself. The reader's weakest-assumption identification matches my own: the finite-cluster-to-bulk transferability is the load-bearing assumption. The main text points to SI S4-S5 for validation, but from the main text we cannot evaluate whether that validation is independent of the lower-level sampling distribution. The proposed active-learning check directly tests whether adding configurations from the model's own bulk ensemble changes predictions. If it does not, the concern is resolved; if it does, the headline claims about force-field/DFT failure would need re-scoping. I do not see an internal inconsistency, and the upfitting strategy is credible enough to warrant the conditional status. The verdict should therefore remain CONDITIONAL pending this test, so no change from the reader's assessment is needed.","tokens_in":18739,"tokens_out":9134,"duration_ms":111288,"concrete_test":"Run a self-consistency active-learning test: use the final CCSD(T) GRACE potential to generate a 3 ns bulk MD; select ~500 uncorrelated snapshots by farthest-point sampling in the GRACE descriptor space; compute DLPNO-CCSD(T)/def2-TZVPP energies and forces for finite clusters of these snapshots (toluene + nearest 32 water, with boundary treatment as in Fig. 1); retrain a new CCSD(T) GRACE potential on the union of original and new data. If the new potential reproduces the original bulk RDFs and π-H-bond population (Fig. 5(f)) within statistical error, the finite-cluster training set was adequate. If the retrained potential shifts the Ph-Hw first-peak height by >0.05 or changes the single-π-H-bond population by >3%, the original coverage was incomplete and the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the CCSD(T)-upfitted GRACE potential reproduces CCSD(T) energies and forces in bulk, and therefore that the force-field/DFT failures are real. For this, the finite-cluster training set of Fig. 7(a) must cover the same local configurations the MLIP visits in bulk. This is not guaranteed. The clusters are sampled from periodic MD driven by lower-level MP2 (or PBE) potentials, then distance-scaled; CCSD(T) upfitting uses only energies. The paper's own results show MP2 and revPBE0-D3 have π-H-bond populations and free-energy barriers (Figs. 5(f), 6) that differ substantially from CCSD(T). If the MP2 ensemble underweights the CCSD(T) barrierless Ph-Ow pathway or overweights two-H-bond configurations, CCSD(T) energy labels on MP2-sampled clusters cannot repair the missing regions; the MLIP must extrapolate there. The claimed end-to-end validation in SI S4-S5 compares against clusters drawn from a similar sampling distribution, not against the CCSD(T) bulk ensemble, which cannot be computed directly. A second, related gap: CCSD(T) energies of isolated clusters include boundary and long-range interactions that a 5 Å cutoff GRACE model cannot represent; without demonstrating that the local energy is converged with respect to cluster size/buffer, the reference labels may not equal the bulk local energies the model can express. The downstream conclusions about force fields and DFT all rest on this transferability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a two-step procedure to train a GRACE machine-learned interatomic potential (MLIP) for aqueous toluene at CCSD(T) accuracy: first a lower-level MP2 model is trained on energies and forces of finite clusters sampled from periodic MD (with distance scaling for subsets), then a small CCSD(T) energy-only training set is used to upfit the potential. The authors claim that this CCSD(T)-upfitted GRACE potential reproduces coupled cluster energies and forces in bulk liquid water, and use it to benchmark three force fields (OPLS/AA, GAFF2, CGenFF) and two electronic-structure-based GRACE potentials (revPBE0-D3, MP2). The comparisons focus on radial and spatial distribution functions, interfacial water orientations, π-hydrogen-bond populations, and free-energy profiles along Ph–O/H and Cm–O/H distances. The central conclusion is that all tested force fields and the hybrid DFT functional fail to capture the balance between hydrophobic solvation and hydrophilic π-interactions relative to the CCSD(T) reference.","tokens_in":19131,"tokens_out":3122,"duration_ms":37289,"significance":"If the transferability claim holds, the manuscript offers a practical, data-efficient route to CCSD(T)-accuracy condensed-phase MLIPs from finite cluster data and provides a much-needed benchmark for force-field and DFT descriptions of aromatic solvation. The explicit two-step training architecture, the use of only finite clusters, and the inclusion of nuclear quantum effects via path integral simulation are sensible and extend prior water-focused CCSD(T)-accuracy MLIP work. The paper also usefully identifies specific qualitative failures (e.g., overestimated π-H-bond populations in OPLS/AA and GAFF2, understructured hydrophobic shells) that are informative for developers. However, the strength of the central claim depends on validation that is largely relegated to the Supporting Information, and the quantitative comparisons in the main text lack statistical uncertainties, so the significance as it stands is conditional on that supporting evidence being complete and decisive.","major_comments":[{"comment":"The central transferability premise is not directly demonstrated in the main text. The CCSD(T)-upfitted potential is trained on finite clusters sampled from periodic MD driven by MP2 or PBE, with distance scaling for subsets. If the lower-level ensembles underweight the π-H-bond configurations that dominate the CCSD(T) bulk ensemble (and the paper's own Fig. 5(f) shows MP2 and revPBE0-D3 differ substantially from CCSD(T) in p(N)), then CCSD(T) labels on those clusters cannot repair the missing regions, and the MLIP must extrapolate in bulk. The claim that the potential 'reproduces CCSD(T) energies and forces in bulk' is supported only by cluster-based end-to-end validation in SI S4–S5, not by any CCSD(T)-labeled configurations drawn from the final bulk trajectory. Please add a self-consistency test in which CCSD(T) energies (and, if feasible, forces) are computed for configurations extra","section":"Fig. 7(a), Methods 'Machine learning interatomic potential training', SI S4–S5"},{"comment":"No statistical uncertainties are reported for any simulation observable. RDF peak heights, p(N) values, straddling-angle distributions, and free-energy profiles are plotted without error bars or block-averaging estimates. This is load-bearing because several headline claims rest on small differences: the CCSD(T) barrier in Ph–Hw is ≈0.57 kJ/mol versus ≈1.12 kJ/mol for MP2 (Fig. 6(a,c)), and p(N) RMSDs of 2–6% are quoted in the text. With 3 ns of production data for one solute in 256 waters, these differences may be within sampling uncertainty. Please add block-averaged error bars or otherwise quantify the statistical precision of the key quantities, and temper claims that are not statistically resolvable.","section":"Figs. 2, 4, 5, 6; Results 'Solvation free energy barriers due to π-water interactions'"},{"comment":"The DLPNO-CCSD(T) reference is used with def2-TZVPP (with a def2-QZVPP HF energy), and the validation of this choice is described only in SI Section S5. Since the entire benchmark chain rests on this reference, the main text should at least summarize the convergence tests with respect to DLPNO cutoffs, basis set, and cluster size/buffer. Without this information, the reader cannot judge whether the reference energies are converged to the claimed 'CCSD(T) accuracy' for the bulk local configurations. This is a missing-support issue that must be addressed before the central claim can be accepted.","section":"Methods 'Reference calculations', SI Section S5"},{"comment":"The manuscript repeatedly asserts that the upfitted GRACE potential 'reproduces CCSD(T) energies and forces in bulk' and that 'only interaction potentials that recover the CCSD(T) balance ... can be considered predictive.' These claims are stronger than what the main text demonstrates: the upfitting uses only energies, and the force accuracy is claimed to follow from the mock-training and end-to-end tests in SI S4–S5, which are not reproduced or summarized numerically in the paper. Please state explicitly what the SI validation shows (e.g., force RMSE on held-out clusters, RDF agreement with periodic reference) and, if possible, include a figure or table of energy/force errors in the main text. This would make the reasoning auditable without requiring the reader to access the SI.","section":"Results 'Toluene solvation structure' and 'Discussion and Conclusions'"}],"minor_comments":[{"comment":"Typo: 'nucethe' should be 'nucleate' or similar; also 'in nucethe role' is garbled.","section":"Introduction, 'nucethe'"},{"comment":"The figure caption states 'The RDFs use a binwidth of 0.04 Å' but the inset panels are not labeled; please add explicit axis labels and legend entries for the zoomed panels.","section":"Fig. 2 caption"},{"comment":"The term 'barrier to breaking of the H-bond' is used for both Ph–Ow and Ph–Hw coordinates, but the physical meaning differs (radial escape vs. rotational opening). Please define the coordination criteria more explicitly in the text.","section":"Results, 'Solvation free energy barriers'"},{"comment":"The statement 'All data needed to evaluate the conclusions in the paper are present in the paper and/or the Supporting Information' conflicts with the fact that the SI is not included with the manuscript. Please either include the SI as supplementary material or specify where it can be obtained.","section":"Data Availability"}],"recommendation":"major_revision","confidential_remarks":"The central idea is valuable and the two-step upfitting strategy is worth publishing if the transferability and reference-quality concerns are properly addressed. My recommendation of major_revision hinges on the expectation that SI S4–S5 contain the missing validation; if they do not, the manuscript should not be accepted. The lack of error bars is a recurring theme that affects almost every quantitative comparison, and I would treat that as a mandatory revision rather than a stylistic request. The paper also tends to overstate its conclusions with claims such as 'only interaction potentials that recover the CCSD(T) balance ... can be considered predictive' — even if the GRACE model is correct for toluene, extrapolating to all biomolecular simulations is unsupported. On scope, the manuscript fits well in a physical-chemistry journal; the benchmarking aspect is of interest to the force-field community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nFirst thing you should know: this is the first CCSD(T)-quality MLIP for an aromatic solute in bulk water, and the benchmark results against force fields and DFT are the real payload. The upfitting machinery is inherited from prior water work, so the novelty is in the transfer to a molecule with a π-system and in the specific findings: OPLS/AA and GAFF2 overestimate π-H-bond populations and water-exchange barriers, CGenFF gets π contacts roughly right but understructures the hydrophobic shell, and revPBE0-D3/MP2 overestimate the barrier by about a factor of two. If the CCSD(T) potential is genuinely accurate in bulk, those are results worth having.\n\nWhat the paper does well: the training scheme is clearly described, the choice of toluene as a two-face probe (aromatic + aliphatic) is smart, and the authors are upfront that the CCSD(T) forces in bulk are an upfitting consequence rather than a direct fit. The mock validation against periodic revPBE0-D3 (Section S1, cited in the main text) is the right kind of transferability test, and it gives me some confidence that the finite-cluster-to-bulk step is not handwaving.\n\nSoft spots, in order of seriousness. First, the central bulk CCSD(T) validation lives in the SI, and the main text gives no statistical uncertainties on any RDF, SDF, H-bond population, or free-energy barrier. Claims like ΔF = 0.57 vs 1.12 kJ/mol need error bars to support a benchmarking conclusion. Second, the stress-test point is real: the finite training clusters are sampled from MP2-driven MD and distance-scaled, and MP2 mis-describes the very π-H-bond configurations that CCSD(T) corrects. The mock DFT test suggests the upfitting procedure can survive a lower-level reference, but it does not settle whether the CCSD(T) labels on MP2-sampled clusters cover the bulk CCSD(T) ensemble. Third, the 5 Å cutoff GRACE model is being asked to reproduce total CCSD(T) energies of clusters including long-range contributions; the authors need to show local-energy convergence with cluster size. Fourth, no code or data are shipped, which matters for a paper whose value is largely as a benchmark.\n\nNone of these are load-bearing flaws; they are addressable. This deserves a serious referee and a careful reading of the SI. If the validation holds, it is a citation-worthy benchmark. If I were the editor, I would send it to review but ask for error bars and a clear statement of what the SI demonstrates about bulk transferability.","headline":"First CCSD(T)-quality MLIP for an aromatic solute in water, with useful benchmark findings, but the bulk validation and missing error bars need to survive review.","tokens_in":19637,"tokens_out":2675,"would_cite":true,"duration_ms":29889,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A machine-learned potential trained only on finite clusters reproduces CCSD(T) accuracy for toluene in bulk water and shows standard force fields and hybrid DFT get the hydrophilic–hydrophobic balance wrong.","keywords":["toluene hydration","CCSD(T) benchmark","machine learning interatomic potential","graph atomic cluster expansion","upfitting","pi-hydrogen bonding","hydrophobic solvation","force field validation"],"falsifier":"Take independent snapshots from the CCSD(T)-MLIP bulk simulation, compute explicit DLPNO-CCSD(T) single-point energies and forces on solute-plus-nearby-water clusters, and check that the MLIP errors stay at coupled-cluster tolerance across both π-hydrogen-bonded and hydrophobic-contact configurations; a large error spike near the ring or methyl group would refute the transferability claim.","tokens_in":18639,"feed_emoji":"💧","tokens_out":4777,"duration_ms":47792,"temperature":0.7,"pith_summary":"The paper aims to establish that condensed-phase simulations of aromatic molecules in water can reach coupled-cluster (CCSD(T)) accuracy by training a graph-based machine-learned potential on finite molecular clusters alone, using a two-step upfitting scheme: first fit to MP2 energies and forces, then refine to CCSD(T) energies. Applied to toluene in liquid water, the resulting potential reproduces coupled-cluster energies and forces in the bulk and serves as a benchmark. Against this benchmark, three widely used biomolecular force-field families and a dispersion-corrected hybrid DFT functional all fail in distinct ways: the hydrophobic shell near the methyl group is understructured, interfacial water is misoriented, the number of water–π hydrogen bonds is over- or underestimated, and the free-energy barrier to breaking a water–π bond is roughly doubled. A sympathetic reader would care because these interactions are thought to govern protein folding, ligand binding, and DNA solvation, and the paper offers a practical route to accurate reference data for them.","feed_headline":"Coupled-cluster accuracy comes to aromatic molecules in liquid water","feed_subtitle":"Standard force fields and hybrid DFT misjudge the balance of π-hydrogen bonds and hydrophobic hydration.","key_machinery":"The central object is the MLIP built on the graph atomic cluster expansion (GRACE), a many-body interatomic potential parameterized on local atomic environments. The training strategy is upfitting: a GRACE potential is first fit to MP2 energies and forces on finite clusters, then a smaller set of clusters is labeled with CCSD(T) energies and the potential is refined using energies only, without any DFT training data. To make finite clusters cover the bulk configurational space, clusters are sampled from periodic molecular dynamics and interatomic distances are scaled for subsets of homo- and hetero-clusters. This combination is what transfers CCSD(T) accuracy from small cluster calculations","core_discovery":"The central claim is that only interaction potentials that recover the CCSD(T) balance between π-interactions and hydrophobic interactions can be considered predictive for biomolecular simulations. Concretely, the paper reports that its CCSD(T)-quality MLIP yields a free-energy barrier of about 0.57 kJ/mol for breaking a water–π hydrogen bond, with essentially no barrier along the oxygen coordinate, and that configurations with one accepted π-hydrogen bond are the most probable, about half of all configurations. None of the compared force fields or the hybrid DFT functional reproduces this pattern: force fields either overbind water to the ring or understructure the hydrophobic shell, while","pith_inferences":["The finite-cluster transferability premise deserves independent scrutiny: an MLIP that matches cluster energies may still miss collective bulk effects, such as long-range electrostatics or cooperative hydrogen-bond networks that small clusters cannot represent.","A direct test of the 'toluene as archetype' claim would be to apply the same recipe to benzene alone, isolating ring-specific π-effects from the methyl group's hydrophobic perturbation.","The reported insensitivity of solvation structure to nuclear quantum effects predicts that classical simulations suffice for structure but does not rule out isotope effects on dynamic quantities such as π-H-bond lifetimes; that is a testable extension.","If the hydrophilic–hydrophobic trade-off is intrinsic to isotropic pairwise force fields, then anisotropic or polarizable models—or MLIPs fitted to correlated wavefunction data—are required for reliable biomolecular simulations, rather than further reparameterization of existing fixed-charge forms."],"forward_implications":["CCSD(T)-quality condensed-phase simulations of aqueous aromatic solutes become feasible without periodic coupled-cluster calculations, using only finite cluster data.","The specific benchmark numbers—the ~0.57 kJ/mol barrier, the dominance of one-π-H-bond configurations, and the straddling orientation of hydrophobic water—become reference targets for force fields and DFT.","OPLS/AA, GAFF2, and CGenFF cannot simultaneously describe hydrophobic and π-interactions; improving one appears to degrade the other.","revPBE0-D3 and MP2 overestimate water–π hydrogen-bond breaking barriers by about a factor of two, implying too slow hydrophilic water exchange near aromatic rings.","The training workflow is general and transferable to larger solutes such as drug-like molecules, amino-acid residues, and nucleic acid fragments, provided explicit electrostatics are added where long-range charge transfer matters."],"fun_headline_variants":["Force fields fail: MLIP nails aromatic solvation at CCSD(T) level","Aromatic solvation: MLIP achieves gold-standard CCSD(T) where others don't","Coupled-cluster MLIP reveals hydrophobic and π-bond balance in water","MLIP matches CCSD(T) for aromatic solvation, unlike force fields and DFT","CCSD(T)-quality MLIP reveals true balance of π-bonds and hydrophobicity in water"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that CCSD(T) energies computed on finite clusters sampled from lower-level periodic molecular dynamics, with scaled distances for subsets of clusters, cover the same configurational space that the MLIP visits in bulk; if the cluster set omits the π-hydrogen-bond and hydrophobic-contact geometries that dominate the bulk ensemble, the fitted potential can look accurate on training-like clusters yet fail in bulk.","fun_headline_variants_meta":{"raw":{"variants":["Force fields fail: MLIP nails aromatic solvation at CCSD(T) level","Aromatic solvation: MLIP achieves gold-standard CCSD(T) where others don't","Coupled-cluster MLIP reveals hydrophobic and π-bond balance in water","MLIP matches CCSD(T) for aromatic solvation, unlike force fields and DFT","CCSD(T)-quality MLIP reveals true balance of π-bonds and hydrophobicity in water"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001618,"raw_usage":{"total_tokens":6296,"prompt_tokens":787,"completion_tokens":5509,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":5398}},"tokens_in":531,"tokens_out":5509,"duration_ms":39874,"temperature":1.0,"reasoning_tokens":5398,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T05:41:11.293164+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take independent snapshots from the CCSD(T)-MLIP bulk simulation, compute explicit DLPNO-CCSD(T) single-point energies and forces on solute-plus-nearby-water clusters, and check that the MLIP errors stay at coupled-cluster tolerance across both π-hydrogen-bonded and hydrophobic-contact configurations; a large error spike near the ring or methyl group would refute the transferability claim.","supporting_citations":[],"review_version":1}