{"id":"b3c23fc4-b81e-4996-8d83-dcd7e764d8af","arxiv_id":"2509.10872","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A transferable machine learning potential trained on a new 3,119-configuration unrestricted CCSD(T) dataset outperforms DFT-trained potentials for gas-phase organic reaction barriers and forces.","lead":"This paper builds an automated workflow to compute energies and forces for 3,119 organic molecular configurations at the unrestricted CCSD(T) level, then trains a machine learning potential on those calculations. The resulting potential reproduces activation energies and forces more accurately than a potential trained on DFT data, pointing toward cheap reactive simulations at coupled cluster accuracy.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Force basis-set correction in Eq. (2) is validated only on near-equilibrium six-atom molecules, yet its 0.073 eV/Å error is the same order as the claimed 0.1 eV/Å improvement; failure at stretched or radical geometries would bias every force label and the resulting MLIP comparison.","rationale":"The reader's weakest_assumption and my own review converge on Eq. (2): the force basis-set correction is the single most load-bearing component of the workflow. The paper is transparent about the 0.073 eV/Å validation set, and the filtering protocol is a reasonable safeguard, but neither establishes that the correction holds at the stretched, radical, and transition-state geometries that dominate the reactive dataset. Because the claimed improvement (0.1 eV/Å) is comparable to the validation error, even a moderate degradation of the correction at those geometries could invert or manufacture the headline result. The concrete test is feasible because the problematic structures are small and PySCF supports analytic UCCSD(T) forces; it would settle the concern without requiring new method development. A controlled DFT-versus-UCCSD(T) MLIP comparison with matched training set sizes would also be valuable, but it is secondary: if the force labels themselves are biased, every downstream comparison inherits the bias. Therefore the existing CONDITIONAL verdict is appropriate, and no verdict change is needed on the basis of this stress-test pass.","tokens_in":19775,"tokens_out":3381,"duration_ms":31836,"concrete_test":"Select a stratified sample of about 50 configurations from the 1953 bond-stretched structures and 50 from the NEB/dimer/SEGS subset, restricting to molecules with at most 8 atoms so that true UCCSD(T)/cc-pVQZ analytic forces are feasible, and apply the paper's own filtering criteria first. Compute full UCCSD(T)/cc-pVQZ forces with PySCF and compare component-wise against the Eq. (2) QZ* predictions. If the RMSE per component exceeds roughly 0.1 eV/Å, or if the signed mean error becomes non-negligible, the force labels are systematically biased at reactive geometries and the MLIP comparison is compromised. As a cross-check, recompute activation energies for five held-out reactions using true UCCSD(T)/cc-pVQZ energies at the same TS/R/P geometries; if the QZ* versus QZ shift exceeds 0.1 eV, the headline activation-energy improvement is within the label error.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that training on UCCSD(T)/QZ* forces gives more than 0.1 eV/Å improvement over DFT-trained MLIPs. Eq. (2) constructs every force label as F_UCCSD(T)/DZ + (F_UCCSD/TZ - F_UCCSD/DZ) + (F_UMP2/QZ - F_UMP2/TZ). The paper's own validation (SI S1.2, Fig. S2) is limited to 300K MD configurations of CO2, H2O, C2H2, and H2CO, all near-equilibrium six-atom molecules, and reports an RMSE of 0.073 eV/Å per force component against true UCCSD(T)/QZ forces. That error is the same order as the claimed 0.1 eV/Å benefit. The dataset deliberately contains 1953 bond-stretched structures plus NEB/dimer/SEGS configurations with force components up to tens of eV/Å (Fig. S11), where the additivity of MP2 and UCCSD basis-set corrections and the consistency of the UHF reference across DZ/TZ/QZ are not established. The paper itself documents that the Coulson-Fischer point shifts with basis set and that ⟨S2⟩ varies strongly with basis (0.51/0.22/0.13 in the worked example), so the correction's central assumption—that basis-set incompleteness errors in forces are transferable across methods and geometries—is weakest exactly where reactive chemistry matters most. Since all 3119 force labels share this correction, a systematic failure would transfer directly into the MLIP training signal and could create or mask the reported improvement. The filtering protocol removes some problematic structures, but it does not validate the correction on the structures that remain.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents an automated workflow for generating unrestricted CCSD(T) energies and forces for gas-phase organic reactions, using UHF stability analysis, a composite basis-set correction (Eqs. (1)–(2)), and a filtering protocol. The resulting dataset contains 3119 configurations drawn from transition states, reactants, products, NEB/dimer/SEGS sampling, and bond-stretched structures. Using transfer learning from a DFT-trained HIP-HOP-NN model, the authors fine-tune an MLIP on this UCCSD(T) dataset and report force RMSE improvements over a DFT-trained MLIP of more than 0.1 eV/Å and an activation-energy RMSE improvement from 0.377 eV to 0.252 eV. The paper also analyzes DFT vs UCCSD(T) differences across the dataset and fine-tunes MACE-MP as a second baseline architecture.","tokens_in":20162,"tokens_out":6201,"duration_ms":53184,"significance":"The significance is high if the central claim holds: a transferable reactive MLIP at CCSD(T) accuracy would be a substantial advance over DFT-trained potentials, and the dataset plus ALF code release would be a community resource. The paper has clear strengths: the workflow is described in detail; internal consistency checks include a 0.073 eV/Å validation of the force correction on small molecules; active learning and transfer learning are well motivated; activation-energy evaluations are performed on held-out reaction paths; and a second model family (MACE-MP) is fine-tuned as a comparison. The main risk is that the load-bearing basis-set correction is validated only on near-equilibrium six-atom molecules, with an error comparable to the claimed improvement, while the reactive geometries in the dataset are precisely where the correction's transferability is least tested. This concern does not invalidate the approach, but it must be addressed before the quantitative claims can be accepted.","major_comments":[{"comment":"The central force-accuracy claim depends on Eq. (2), whose key assumption—that basis-set incompleteness errors in forces are transferable between MP2, UCCSD, and UCCSD(T)—is validated in SI S1.2 only on 300 K MD configurations of CO2, H2O, C2H2, and H2CO. These are near-equilibrium six-atom species, and the reported correction error of 0.073 eV/Å per force component is the same order of magnitude as the >0.1 eV/Å improvement claimed in the abstract. Since the dataset deliberately includes 1953 bond-stretched structures and NEB/dimer/SEGS geometries with force components up to tens of eV/Å (Fig. S11), the transferability of the correction at stretched, radical, and transition-state geometries is a load-bearing assumption. Please add validation of Eq. (2) on representative reactive geometries, or provide a bound on the correction error for those geometries.","section":"II.C.2 / SI S1.2, Eq. (2)"},{"comment":"The filtering protocol removes structures near the Coulson–Fischer point and structures with large spin-state differences across bases, but it does not establish the accuracy of Eq. (2) on the structures that remain. The manuscript itself documents large basis-set dependence of the UHF reference for a removed transition state (⟨S2⟩ = 0.51/0.22/0.13 for DZ/TZ/QZ and an orbital Hessian eigenvalue below 10^-3 Ha), and the filter thresholds of 1 eV/Å and 5 eV/Å are much larger than the claimed 0.1 eV/Å force improvement. A systematic failure of the correction on retained reactive structures would propagate into all 3119 force labels and could create or mask the reported MLIP improvement. Please quantify how the final training labels and the MLIP comparison depend on these thresholds and on the correction assumptions.","section":"SI S1.3 / S1.4"},{"comment":"The same transferability assumption underlies the energy correction in Eq. (1), and the activation-energy results in Figs. 5 and 7 are based on these composite energies. No validation of Eq. (1) at transition-state geometries is provided; the SI energy validation shown in Fig. S3 is limited to the same near-equilibrium small molecules. Please either validate Eq. (1) on representative transition-state and stretched geometries or state the expected uncertainty in the activation-energy RMSE values.","section":"II.C.2, Eq. (1)"}],"minor_comments":[{"comment":"The title contains a line-break typo, \"high-throug hput\"; please correct it.","section":"Title / throughout"},{"comment":"The caption reports the RMSE as \"0.073 eV/\" with an incomplete unit; please specify eV/Å per force component.","section":"SI Fig. S2"},{"comment":"The text refers to \"UCCD(T)/QZ\" in one passage; this should be \"UCCSD(T)/QZ*\" to avoid ambiguity.","section":"SI S1.2"},{"comment":"The SI figure numbering is inconsistent: both S1.3.1 and S1.4 reference \"Fig. S10\" for different content; please renumber the figures and update all references.","section":"SI S1.3.1 and S1.4"},{"comment":"The main text does not give the numeric test-set force RMSE values for the DFT-trained and UCCSD(T)-trained HIP-HOP models; please include these values in the caption or text so the claimed >0.1 eV/Å improvement can be verified directly.","section":"Fig. 6 / Section III.B"},{"comment":"The activation-energy numbers are reported inconsistently: the text states 0.377 eV RMSE for the DFT-trained MLIP, while the Fig. 7 caption reports 0.297 eV for the ωB97X functional; please clarify which comparison supports the abstract statement of \"over 0.1 eV\" improvement.","section":"Fig. 7 / Section III.B"},{"comment":"In the bond-dissociation discussion the DFT reference is described only as \"singlet DFT\"; please specify the exact functional and basis set used in Fig. S13.","section":"SI S6.1 / Fig. S13"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does something that has needed doing: it builds a transferable machine learning potential for reactive organic chemistry trained on unrestricted CCSD(T) energies and forces, not DFT. Prior CCSD(T) potentials were molecule-specific; this one is a real step toward coupled-cluster-quality forces over a broad reaction space. The workflow—automated UHF stability analysis, spin-contamination filtering, composite force basis correction, active learning sampling—is practical and the SI is refreshingly detailed. They even fine-tune MACE-MP to the dataset as an independent check, which strengthens the conclusion that the data, not the architecture, is the point.\n\nThe central claim is that training on UCCSD(T) beats training on DFT by more than 0.1 eV/Å in forces and 0.1 eV in activation energies. I think the qualitative conclusion is right: the held-out activation energy test is genuinely out-of-sample, and the comparison between DFT-trained and UCCSD(T)-trained models is fair in that both are evaluated on the same reference. But the reference itself is UCCSD(T)/QZ*, built through Eq. (2), and that composite correction is the soft spot. It is validated on 300 K MD of four six-atom molecules, giving an RMSE of 0.073 eV/Å per component. That is the same order as the claimed improvement, and the dataset deliberately includes bond-stretched and NEB/dimer structures with force components up to tens of eV/Å. At those geometries the additivity of MP2 and UCCSD basis-set corrections, and even the consistency of the UHF reference, is not established. The paper itself shows a problematic structure where ⟨S²⟩ jumps from 0.51 to 0.13 across DZ/TZ/QZ. The filtering protocol removes the worst offenders, but it does not validate the correction on the structures that remain.\n\nSo the magnitude of the headline improvement could be inflated by a systematic bias in the force labels that the UCCSD(T)-trained MLIP learns and the DFT-trained MLIP does not. That is a real concern, but it is a caveat, not a fatal flaw. The MLIP still performs well on held-out reaction paths, and the dataset plus workflow are valuable even if some force labels carry hidden uncertainty. Minor additional issues: the dataset is not yet released, and some extrapolation tests (bond dissociation RMSE 0.485 eV, HCN barrier off by 0.467 eV) show the model is not yet uniformly at gold-standard accuracy.\n\nThis deserves a serious referee. The reviewers should push for explicit validation of Eq. (2) on stretched and radical geometries, or a frank analysis of how the correction error propagates into the MLIP comparison, and for the dataset to be released with the paper. I would bring it to a reading group and would cite it, with the basis-set caveat noted.","headline":"A genuinely useful UCCSD(T) reactive dataset and transferable MLIP, but the force basis-set correction is the load-bearing assumption and it is only validated near equilibrium.","tokens_in":20718,"tokens_out":3305,"would_cite":true,"duration_ms":29828,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Unrestricted CCSD(T) can be automated to supply a 3,119-structure reactive dataset, and potentials trained on it beat DFT-trained potentials in force and activation-energy accuracy.","keywords":["unrestricted coupled cluster","CCSD(T)","machine learning interatomic potential","active learning","basis set correction","activation energy","transfer learning","reactive chemistry"],"falsifier":"Compute explicit UCCSD(T)/QZ forces for a sample of the dataset's bond-stretched and radical transition-state structures and compare them component-by-component with the UCCSD(T)/QZ* values; if the per-component RMSE grows well beyond the 0.073 eV/Å reported for small equilibrium molecules, or if the errors correlate with bond length or spin contamination, the basis-set correction—and therefore every MLIP label—is systematically biased.","tokens_in":19576,"feed_emoji":"⚛️","tokens_out":12571,"duration_ms":98782,"temperature":0.7,"pith_summary":"This paper tries to establish that unrestricted CCSD(T)—the coupled-cluster method with singles, doubles, and perturbative triples—can be automated well enough to generate thousands of energies and forces for reactive organic molecules, and that machine-learned interatomic potentials trained on those labels are more accurate than potentials trained on DFT data. The authors build a 3,119-configuration gas-phase dataset of C/H/N/O molecules at a composite UCCSD(T)/QZ* level, using active learning, transition-state searches, and bond-stretching sampling. The concrete payoff is quantitative: switching the training data from DFT to UCCSD(T) improves force accuracy by more than 0.1 eV/Å and activation-energy reproduction by more than 0.1 eV, with the UCCSD(T)-trained potential reproducing held-out reaction barriers at 0.252 eV RMSE. A sympathetic reader would care because reactive chemistry—bond breaking, radicals, transition states—is exactly where DFT is known to be weakest, so a transferable coupled-cluster-level potential would make high-fidelity reaction simulation affordable.","feed_headline":"Coupled-cluster data beats DFT for training reactive ML models","feed_subtitle":"Training on 3,119 gold-standard quantum structures cuts force error by 0.1 eV/Å and barrier error to 0.25 eV.","key_machinery":"The load-bearing identity is the composite basis-set-corrected UCCSD(T) force, $$F_{\\mathrm{UCCSD(T)/QZ^*}} = F_{\\mathrm{UCCSD(T)/DZ}} + (F_{\\mathrm{UCCSD/TZ}} - F_{\\mathrm{UCCSD/DZ}}) + (F_{\\mathrm{UMP2/QZ}} - F_{\\mathrm{UMP2/TZ}})$$. It is what makes thousands of coupled-cluster force labels affordable, because it replaces a direct UCCSD(T)/QZ force evaluation with a DZ calculation plus two cheaper corrections, and the paper shows the corrected forces track explicit UCCSD(T)/QZ forces with a per-component RMSE of 0.073 eV/Å. The companion mechanism is an automated pipeline that selects a stable unrestricted Hartree-Fock reference with stability analysis and discards structures near Coulson-Fischer points, where the energy is continuous but the force is not; this filtering keeps inaccurate labels out of the training set.","core_discovery":"The central discovery is that the expensive UCCSD(T) level is a practical training source for reactive machine-learned interatomic potentials once three obstacles are automated: choosing a stable unrestricted Hartree-Fock reference by stability analysis, correcting basis-set incompleteness in both energies and forces with a composite scheme, and filtering structures close to Hartree-Fock instability where forces are undefined. With that workflow, the paper computes energies and forces for 3,119 gas-phase organic configurations and fine-tunes a neural-network interatomic potential to them. That potential reproduces held-out UCCSD(T) forces and activation energies more accurately than the same model trained on 270,720 DFT structures, and for activation energies it also beats the ωB97X DFT functional itself (0.252 eV versus 0.297 eV RMSE). The authors conclude that the level of theory of the training data, rather than the model family, is the current bottleneck limiting the accuracy of reactive machine-learned potentials.","pith_inferences":["Extension: a direct audit of the composite basis-set correction on the hardest geometries in the dataset would be the cheapest stress test; computing true UCCSD(T)/QZ forces for a sample of stretched and radical structures and comparing them with UCCSD(T)/QZ* would show whether the 0.073 eV/Å benchmark holds where it matters most.","Extension: the paper's filtering of near-Coulson-Fischer structures delimits the potential's reliable domain to single-reference territory; automated multireference workflows, which the paper names as future work, would be the natural route to cover the excluded bond-breaking regions.","Extension: an ablation that removes the 1,953 bond-stretching structures from the training set could reveal whether the activation-energy advantage comes from transition-state labels or from stretched-geometry labels, telling future dataset builders which sampling earns its cost.","Extension: the two-stage recipe—cheap DFT exploration with active learning, then UCCSD(T) labels only on uncertain structures—should transfer to condensed-phase problems if cluster-extraction methods mature, which is the paper's stated open challenge."],"forward_implications":["The UCCSD(T)-trained potential reproduces forces on transition states, reactants, and products with an RMSE more than 0.1 eV/Å lower than the DFT-trained potential.","Activation energies of reactions outside the training set are reproduced at 0.252 eV RMSE, better than the 0.297 eV RMSE of the ωB97X DFT functional and the 0.377 eV RMSE of the DFT-trained potential.","The UCCSD(T)-trained potential can drive NEB and dimer transition-state searches, converging to transition states and minimum-energy paths in minutes, where direct UCCSD(T)/QZ* searches would be prohibitively expensive.","Fine-tuning a pre-trained foundation model to the same 3,119-structure dataset yields a force error comparable to the transfer-learned model, suggesting the dataset itself—not the particular architecture—is what delivers the accuracy gain."],"supporting_citations":[{"why":"Supplies the Hartree-Fock stability analysis used to automate selection of the unrestricted reference.","marker":"[24]"},{"why":"Supplies the basis-set-correction framework that the composite energy and force corrections extend.","marker":"[25]"},{"why":"Provides the initial reaction paths, reactants, products, and transition states that seed the UCCSD(T) dataset.","marker":"[27]"},{"why":"Provides the Transition-1x DFT reactive dataset used both to initialize active learning and as the DFT-trained baseline.","marker":"[34]"},{"why":"Supplies the single-ended growing string method used to generate new reaction pathways beyond the seed set.","marker":"[30]"},{"why":"Supplies the HIP-NN-TS architecture whose ensemble uncertainty drives active-learning selection.","marker":"[41]"},{"why":"Supplies the HIP-HOP-NN architecture used for the final transfer-learned potentials.","marker":"[43]"},{"why":"Supplies the electronic-structure code used for the unrestricted MP2, UCCSD, and UCCSD(T) calculations.","marker":"[51]"},{"why":"Supplies the QM9 equilibrium-molecule pool used to propose reactants for new reaction-path searches.","marker":"[45]"}],"fun_headline_variants":["Gold-standard coupled cluster now feasible for reactive ML training","3,119 UCCSD(T) structures beat 270k DFT for reactive ML","Theory level, not model family, limits reactive ML accuracy","Automating UCCSD(T) yields reactive force fields that beat DFT"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the basis-set incompleteness error in forces is transferable between MP2, UCCSD, and UCCSD(T), so the composite correction built from cheap DZ/TZ/QZ calculations faithfully reproduces UCCSD(T)/QZ forces; the paper validates this only on a few small near-equilibrium molecules, not on the stretched, radical, and transition-state geometries that dominate the reactive dataset.","fun_headline_variants_meta":{"raw":{"variants":["Gold-standard coupled cluster now feasible for reactive ML training","3,119 UCCSD(T) structures beat 270k DFT for reactive ML","Theory level, not model family, limits reactive ML accuracy","Automating UCCSD(T) yields reactive force fields that beat DFT"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000559,"raw_usage":{"total_tokens":2709,"prompt_tokens":1046,"completion_tokens":1663,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":662,"completion_tokens_details":{"reasoning_tokens":1588}},"tokens_in":662,"tokens_out":1663,"duration_ms":10181,"temperature":1.0,"reasoning_tokens":1588,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:52:21.659251+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute explicit UCCSD(T)/QZ forces for a sample of the dataset's bond-stretched and radical transition-state structures and compare them component-by-component with the UCCSD(T)/QZ* values; if the per-component RMSE grows well beyond the 0.073 eV/Å reported for small equilibrium molecules, or if the errors correlate with bond length or spin contamination, the basis-set correction—and therefore every MLIP label—is systematically biased.","supporting_citations":[{"cited_title":"Seeger \\ and\\ author J","cited_arxiv_id":null,"evidence_quote":"Supplies the Hartree-Fock stability analysis used to automate selection of the unrestricted reference."},{"cited_title":"Fan , author J","cited_arxiv_id":null,"evidence_quote":"Supplies the basis-set-correction framework that the composite energy and force corrections extend."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the initial reaction paths, reactants, products, and transition states that seed the UCCSD(T) dataset."},{"cited_title":"Schreiner , author A","cited_arxiv_id":null,"evidence_quote":"Provides the Transition-1x DFT reactive dataset used both to initialize active learning and as the DFT-trained baseline."},{"cited_title":"Chigaev , author J","cited_arxiv_id":null,"evidence_quote":"Supplies the HIP-NN-TS architecture whose ensemble uncertainty drives active-learning selection."},{"cited_title":"Optimal Invariant Bases for Atomistic Machine Learning","cited_arxiv_id":"2503.23515","evidence_quote":"Supplies the HIP-HOP-NN architecture used for the final transfer-learned potentials."},{"cited_title":"Ramakrishnan , author P","cited_arxiv_id":null,"evidence_quote":"Supplies the QM9 equilibrium-molecule pool used to propose reactants for new reaction-path searches."}],"review_version":2}