{"id":"485119a3-406c-45c3-a58d-ad87d5bb604d","arxiv_id":"2501.18876","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new 186k-molecule dataset with B3LYP/TZVP spectra and dynamics that improves learned IR, Raman and NMR simulation over the prior QM9S dataset.","lead":"This paper introduces QMe14S, a dataset of 186,102 small organic molecules with DFT-computed energies, multipole moments, Hessians, and IR, Raman and NMR spectra across 14 elements and 47 functional groups. It reports that a neural network trained on this dataset reproduces DFT spectra better than one trained on the earlier QM9S dataset.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 14-element comprehensiveness claim is unvalidated: the ML comparison is computed only on HCNOF molecules, and no higher-level reference is provided for the newly added elements.","rationale":"The paper's central value proposition is a spectral dataset that extends QM9S from 5 to 14 elements. To support that, the reference calculations for Al, Si, P, Cl, As, Se, and Br must be trustworthy. The text never checks this: Technical Validation and ML-Predicted Spectra restrict the aggregate QM9S-vs-QMe14S comparison to 7,038 HCNOF-only test molecules, and the few spectral examples in Fig. 4 are hydrazine/hydrazone molecules without heavy elements. The reader identified this as the weakest assumption, and I agree. The concern is not that B3LYP/TZVP is necessarily wrong for these elements, but that the headline claim is conditional on a level-of-theory validation that is absent. Secondary issues, such as inconsistent Table 1 totals and missing error bars on the modest IR/Raman improvements, strengthen the need for conditional acceptance, but they do not move the verdict beyond what the reader already assigned.","tokens_in":10147,"tokens_out":8628,"duration_ms":88606,"concrete_test":"Access the Figshare HDF5 release and draw a stratified random sample of 15 molecules per newly added element (Al, Si, P, Cl, As, Se, Br), plus 15 HCNOF controls. Re-optimize the geometries and recompute harmonic frequencies, NMR shieldings, and polarizability-derived Raman intensities at CCSD(T)-F12/def2-TZVP or G4. Compare per-element mean absolute deviations against the HCNOF controls; if the heavy-element deviations exceed the HCNOF baseline by more than the 10–15 cm−1 Lorentzian broadening used for spectra, the 'comprehensive' claim is not supported as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To make the claim 'QMe14S is comprehensive across 14 elements' true, the B3LYP/TZVP labels on the 56,285 PubChem supplement must be reliable for elements QM9S lacked. The paper provides no such evidence. The quantitative benchmark in ML-Predicted Spectra is explicitly computed for '7038 molecules containing only HCNOF elements in the test set', so the new elemental coverage never enters the headline accuracy numbers. The only check offered is that calculations were run; no CCSD(T), G4, or experimental reference is given, and no per-element ML error table is shown. This matters because the dataset's distinctive properties, first hyperpolarizabilities, Raman intensities, and NMR shieldings, are exactly the quantities where B3LYP/TZVP behavior for Br, Se, As, Al, Si, P, and Cl requires validation; an unvalidated systematic bias would make the added molecules a source of confidently wrong labels. The accompanying inventory inconsistencies, such as Table 1 listing 'Atomization Energy' total as 5,907,000 for 186,102 molecules and 6 million MD frames, reinforce the need to audit the released records rather than take the 14-element coverage on faith.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces QMe14S, a dataset of 186,102 small organic molecules covering 14 elements (H, B, C, N, O, F, Al, Si, P, S, Cl, As, Se, Br) and 47 functional groups, with geometries and properties computed at the B3LYP/TZVP level. The dataset augments the authors' earlier QM9S dataset with 56,285 molecules from PubChem, and includes static properties (energy, forces, charges, multipole moments, polarizabilities, hyperpolarizability), harmonic IR, Raman and NMR spectra, and 6 million nonequilibrium configurations from ab initio molecular dynamics. Using their E(3)-equivariant neural network DetaNet, the authors report that models trained on QMe14S outperform models trained on QM9S in predicting IR and Raman spectra, and they demonstrate accurate prediction of NMR spectra on a few examples.","tokens_in":10241,"tokens_out":6424,"duration_ms":60910,"significance":"If the underlying DFT labels are reliable, QMe14S is a valuable resource for machine-learned molecular simulation. Its distinctive contributions are the inclusion of nine elements absent from QM9S, a wider functional-group coverage, and the unique availability of high-order tensors (first hyperpolarizability, octupole moment, Hessian, dipole and polarizability derivatives) together with dynamic configurations. The dataset is released with a reading script, and the ML benchmark is a useful sanity check. However, the central claim that the 14-element dataset improves spectral prediction is only tested on HCNOF molecules, and no higher-level validation is provided for the newly added elements, so the significance is conditional on closing those gaps.","major_comments":[{"comment":"The quantitative comparison between DetaNet-QMe14S and DetaNet-QM9S is restricted to the 7,038 test molecules containing only H, C, N, O, and F; no spectral prediction accuracy is reported for molecules containing any of the nine elements added in QMe14S (Al, Si, P, Cl, As, Se, Br, and others). Because the central claim is that QMe14S improves spectral simulation for a 14-element chemical space, this test set does not substantiate that claim. Please report per-element or per-new-element errors on a test set that includes the added elements, or explicitly restrict the 'outperform' claim to HCNOF molecules.","section":"ML-Predicted Spectra"},{"comment":"No benchmark against higher-level theory or experiment is provided for the B3LYP/TZVP calculations on the newly added elements (Al, Si, P, Cl, As, Se, Br). The properties that make QMe14S distinctive — first hyperpolarizabilities, Raman intensities, and NMR shieldings — are precisely those where B3LYP/TZVP may have significant systematic errors for heavier elements. Please add a validation subsection that compares B3LYP/TZVP results against a higher-level method (e.g., CCSD(T), G4) or available experimental data for a few representative molecules per added element; otherwise the 'comprehensive' claim is not supported.","section":"Quantum mechanical calculation details / Technical Validation"},{"comment":"Table 1's 'Total Numbers' column contains internal inconsistencies: Atomization Energy is listed as 5,907,000 for 186,102 molecules, Atomic Force as 5,907,000 although it should have the same 3N dimension as Atomic positions (6,093,102), and Dipole Moment as 6,093,102 although only three components per molecule are expected (186,102 × 3 = 558,306). These errors suggest the array sizes were not checked against the actual HDF5 files. Please correct the table and verify all entry counts against the released files.","section":"Table 1"},{"comment":"The reported improvement of DetaNet-QMe14S over DetaNet-QM9S on HCNOF test molecules conflates two effects: the addition of new elements and the increase in total training data (including additional HCNOF molecules from PubChem). Because the test set is HCNOF-only, the improvement could stem entirely from having more HCNOF training examples rather than from the elemental diversity. Please include an ablation experiment training on an HCNOF-only subset of QMe14S of comparable size to QM9S, or otherwise disentangle these factors.","section":"Training details / ML-Predicted Properties"}],"minor_comments":[{"comment":"The text contains a typo: 'DateNet-QMe14S' should be 'DetaNet-QMe14S'.","section":"ML-Predicted Spectra"},{"comment":"The statement that ADMP simulations with a 1 fs step and 100 fs total time yield 6 million configurations is inconsistent with 186,102 molecules at 100 steps each (18.6 million); please clarify whether MD was run on a subset and how the 6 million figure is obtained.","section":"Quantum mechanical calculation details"},{"comment":"The abstract claims NMR spectra for the dataset, but Table 1 lists shielding tensors for only 59,260 molecules and Table 2 describes NMR for '60k molecules selected from PubChem'; please reconcile this with the unqualified abstract statement.","section":"Abstract / Data Records"},{"comment":"The sentence 'we also filtered several token that are commonly exist in SMILES' has grammatical errors and should be rephrased, e.g., 'we also filtered several token types that commonly occur in SMILES'.","section":"Methods, Data collection and preprocessing"}],"recommendation":"major_revision","confidential_remarks":"The dataset is potentially useful, but the paper's headline claim is not yet supported by the evidence: the ML benchmark excludes the newly added elements, and the DFT level is unvalidated for those elements. The Table 1 inconsistencies, though fixable, currently undermine trust in the data records. I would encourage the editor to request an ablation and a higher-level validation before acceptance, as these are within the scope of a data-descriptor revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First: the dataset is real work and useful. QMe14S extends QM9S with 56k PubChem molecules, pushing to 14 elements and 47 functional groups, and adds Raman, NMR, and MD-derived Hessians. If the data are sound, this is a practical resource for anyone training spectral ML models who doesn't want to deal with QMugs or PubChemQC size. The diversity analysis is fine. The ML comparison is honest in one respect: it is supervised regression against DFT labels, so no circularity.\n\nBut the headline claim — comprehensiveness across 14 elements — is not supported by the evidence presented. The only quantitative benchmark compares DetaNet trained on QMe14S vs QM9S on 7,038 test molecules containing only H, C, N, O, F. So the new elemental coverage never enters the accuracy numbers. The examples in Fig. 4 are drawn from functional groups that were added, which flatters the comparison. There is no higher-level reference (CCSD(T), G4, experiment) for Al, Si, P, Cl, As, Se, Br, and the properties that are distinctive here — hyperpolarizabilities, Raman intensities, NMR shieldings — are exactly where B3LYP/TZVP can drift for heavier elements. The stress-test note is right about this.\n\nSecond: Table 1 is internally inconsistent. Atomization energy is listed as 5,907,000 entries for 186,102 molecules; APT charges 539,737, dipole 6,093,102. Either the column means something else (e.g., total number of scalar components per property summed over all molecules) or there are real counting errors. This sloppiness matters for a dataset paper because the whole product is the records. It is fixable but needs an audit.\n\nThe ML metrics show a small but consistent gain (IR cosine 93.84 to 94.63, Raman 91.46 to 92.34). No error bars or repeated splits are reported, and the training code is not released. That is minor relative to the data-quality issue, but worth asking.\n\nOverall: this deserves peer review, not desk rejection. The resource is genuinely useful and the flaws are addressable. I would ask for a corrected inventory, a per-element breakdown of ML errors, a validation subset against higher-level theory, and the Figshare URL plus training code. The 14-element comprehensiveness claim should be toned down until the new elements are actually tested.","headline":"A genuinely useful dataset extension, but the 14-element claim is untested: the ML benchmark only covers HCNOF molecules.","tokens_in":10930,"tokens_out":2414,"would_cite":false,"duration_ms":24382,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"QMe14S, a dataset of 186,102 small organic molecules spanning 14 elements and 47 functional groups with DFT-level IR, Raman and NMR spectra, improves machine-learned molecular spectrum simulation compared with training on QM9S.","keywords":["QMe14S dataset","molecular spectra","density functional theory","E(3)-equivariant neural network","IR Raman NMR spectroscopy","small organic molecules","chemical space diversity","ab initio molecular dynamics"],"falsifier":"Compare B3LYP/TZVP vibrational frequencies and IR/Raman intensities for a set of Se- and Br-containing molecules against CCSD(T)-level reference calculations; if the heavy-element errors are systematically larger than for HCNOF molecules, the uniform-quality assumption behind QMe14S fails. Alternatively, retrain DetaNet on QMe14S but test only on molecules containing elements absent from QM9S; a large accuracy drop would indicate the dataset's diversity claim is overstated.","tokens_in":9800,"feed_emoji":"🧪","tokens_out":5183,"duration_ms":47320,"temperature":0.7,"pith_summary":"The paper introduces QMe14S, a dataset of 186,102 small organic molecules built to train machine-learning models that predict molecular spectra. It argues that previous datasets such as QM9S cover too few elements and functional groups, so the authors supplement them with 56,285 PubChem molecules to ensure every one of 14 elements and 47 functional groups appears at least 500 times. All molecules are calculated at the B3LYP/TZVP level, yielding IR, Raman and NMR spectra plus static and dynamic molecular properties. Using their E(3)-equivariant network DetaNet, the authors show that models trained on QMe14S reproduce DFT spectra more closely than models trained on QM9S, particularly for molecules with hydrazine- and nitrogen-rich groups. A sympathetic reader would take the paper's central claim to be that a balanced, diverse but still computationally affordable spectral dataset improves machine-learned spectrum simulation.","feed_headline":"186,102-molecule spectral dataset beats QM9S baseline","feed_subtitle":"A balanced 14-element, 47-group benchmark makes machine-learned IR, Raman and NMR spectra match DFT more closely.","key_machinery":"The load-bearing piece is the dataset's construction itself: the RDKit substructure search on PubChem used to top up underrepresented elements and functional groups, followed by geometry optimization and harmonic frequency analysis at B3LYP/TZVP in Gaussian 16, and ADMP molecular dynamics for nonequilibrium configurations. On the modelling side, DetaNet — an E(3)-equivariant message-passing neural network that predicts scalar and high-order tensorial properties — converts predicted Hessians, dipole derivatives and polarizability derivatives into IR and Raman intensities via normal-mode analysis, and predicted shielding tensors into NMR chemical shifts. The uniform element and functional-group coverage, rather than raw dataset size, is what the paper claims makes the dataset efficient for machine learning.","core_discovery":"QMe14S is a spectral and property dataset spanning 186,102 molecules, 14 elements (H, B, C, N, O, F, Al, Si, P, S, Cl, As, Se, Br) and 47 functional groups, with static and dynamic properties computed at the B3LYP/TZVP level. The dataset adds 56,285 molecules from PubChem to the earlier QM9S dataset so that every element and functional group appears in over 500 instances, producing a much flatter distribution of element and functional-group frequencies than QMugs or PubChemQC. In addition to equilibrium geometries, energies, charges, multipole moments, polarizabilities, Hessians and derivative tensors, QMe14S contains roughly 6 million ab initio molecular dynamics configurations with energies, forces and dipole moments, of which 10 per molecule have additional Hessian and polarizability data. The authors demonstrate with DetaNet that models trained on QMe14S predict IR and Raman spectra closer to DFT than models trained on QM9S, with cosine similarities of about 95% for IR and 92% for Raman on the HCNOF test subset, and that predicted 13C and 1H NMR spectra match DFT calculations closely.","pith_inferences":["Editorially, the 500-instance minimum per element or functional group is a plausible rule of thumb for dataset coverage, but the paper does not test whether this threshold is optimal; measuring spectra accuracy as a function of that threshold would tell whether 500 is enough or overkill.","Editorially, if B3LYP/TZVP holds up for the heavy elements, QMe14S could serve as the spectral analogue of QM9 for ML benchmark comparisons, and the same balancing strategy could be applied to other property databases.","Editorially, the reported comparison to QM9S is limited to molecules containing only H, C, N, O, F in the test set; a direct test on molecules containing the newly added elements is needed to confirm the transferability gain the title implies."],"forward_implications":["Models trained on QMe14S should generalize to molecular spectra for compounds containing heavier main-group elements (Al, Si, P, Cl, As, Se, Br) better than models trained on QM9S.","The dataset's Hessian and derivative tensors allow ML models to simulate IR and Raman spectra without additional quantum-chemistry calculations for each new molecule.","The 6 million dynamic configurations, with forces and Hessians, provide training data for machine-learned force fields that work away from equilibrium geometries.","Unique higher-order tensors (first hyperpolarizability, octupole moment) and their derivatives in QMe14S enable prediction of nonlinear optical properties from machine-learned models."],"supporting_citations":[{"why":"Defines the DetaNet architecture and the QM9S baseline dataset that QMe14S extends.","marker":"[21]"},{"why":"Supplies the QM9 molecules from which QM9S is built and on which QMe14S adds coverage.","marker":"[35]"},{"why":"GDB-17 enumeration that defines the chemical space of the base molecules.","marker":"[32]"},{"why":"PubChem is the source of the 56,285 supplemental molecules that add elements and functional groups.","marker":"[45]"},{"why":"Gaussian 16 is the software used for all B3LYP/TZVP geometry optimization, frequency and property calculations.","marker":"[46]"},{"why":"ADMP is the method used for the ab initio molecular dynamics that generate the dataset's dynamic configurations.","marker":"[47]"},{"why":"Defines the TZVP basis set used in the B3LYP/TZVP level of theory.","marker":"[39]"},{"why":"Cited with [39] to specify the B3LYP/TZVP level of theory used throughout.","marker":"[38]"}],"fun_headline_variants":["186k-molecule spectral dataset beats QM9S in ML","QMe14S: 14-element, 47-group benchmark for spectra","Balanced quantum dataset yields better IR, Raman, NMR","New dataset pushes ML spectra closer to DFT","QMe14S improves machine-learned spectra over QM9S"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claimed breadth and transferability depend on B3LYP/TZVP being accurate enough for the heavier elements (Al, Si, P, Cl, As, Se, Br), which the paper does not check against higher-level theory.","fun_headline_variants_meta":{"raw":{"variants":["186k-molecule spectral dataset beats QM9S in ML","QMe14S: 14-element, 47-group benchmark for spectra","Balanced quantum dataset yields better IR, Raman, NMR","New dataset pushes ML spectra closer to DFT","QMe14S improves machine-learned spectra over QM9S"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1551,"prompt_tokens":1022,"completion_tokens":529,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":638,"completion_tokens_details":{"reasoning_tokens":442}},"tokens_in":638,"tokens_out":529,"duration_ms":5878,"temperature":1.0,"reasoning_tokens":442,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T22:06:26.032004+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare B3LYP/TZVP vibrational frequencies and IR/Raman intensities for a set of Se- and Br-containing molecules against CCSD(T)-level reference calculations; if the heavy-element errors are systematically larger than for HCNOF molecules, the uniform-quality assumption behind QMe14S fails. Alternatively, retrain DetaNet on QMe14S but test only on molecules containing elements absent from QM9S; a large accuracy drop would indicate the dataset's diversity claim is overstated.","supporting_citations":[],"review_version":1}