REVIEW 4 major objections 4 minor 1 cited by
QMe14S, A Comprehensive and Efficient Spectral Dataset for Small Organic Molecules
T0 review · 4 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read QMe14S, a dataset of 186,102 small organic molecules spanning 14 elements and 47 functional groups with DFT-level IR, Raman and NMR spectra, improves machine-learned molecular spectrum simulation compared with training on QM9S.
desk verdict A genuinely useful dataset extension, but the 14-element claim is untested: the ML benchmark only covers HCNOF molecules. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing piece is the dataset's construction itself: the RDKit substructure search on PubChem used to top up underrepresented elements and functional groups, followed by geometry optimization and harmonic frequency analysis at B3LYP/TZVP in Gaussian 16, and ADMP molecular dynamics for nonequilibrium configurations. On the modelling side, DetaNet — an E(3)-equivariant message-passing neural network that predicts scalar and high-order tensorial properties — converts predicted Hessians, dipole derivatives and polarizability derivatives into IR and Raman intensities via normal-mode analysis, and predicted shielding tensors into NMR chemical shifts. The uniform element and functional-group coverage, rather than raw dataset size, is what the paper claims makes the dataset efficient for machine learning.
What would settle it
Compare B3LYP/TZVP vibrational frequencies and IR/Raman intensities for a set of Se- and Br-containing molecules against CCSD(T)-level reference calculations; if the heavy-element errors are systematically larger than for HCNOF molecules, the uniform-quality assumption behind QMe14S fails. Alternatively, retrain DetaNet on QMe14S but test only on molecules containing elements absent from QM9S; a large accuracy drop would indicate the dataset's diversity claim is overstated.
Extended reading notes
Core claim
QMe14S is a spectral and property dataset spanning 186,102 molecules, 14 elements (H, B, C, N, O, F, Al, Si, P, S, Cl, As, Se, Br) and 47 functional groups, with static and dynamic properties computed at the B3LYP/TZVP level. The dataset adds 56,285 molecules from PubChem to the earlier QM9S dataset so that every element and functional group appears in over 500 instances, producing a much flatter distribution of element and functional-group frequencies than QMugs or PubChemQC. In addition to equilibrium geometries, energies, charges, multipole moments, polarizabilities, Hessians and derivative tensors, QMe14S contains roughly 6 million ab initio molecular dynamics configurations with energies, forces and dipole moments, of which 10 per molecule have additional Hessian and polarizability data. The authors demonstrate with DetaNet that models trained on QMe14S predict IR and Raman spectra closer to DFT than models trained on QM9S, with cosine similarities of about 95% for IR and 92% for Raman on the HCNOF test subset, and that predicted 13C and 1H NMR spectra match DFT calculations closely.
Load-bearing premise
The claimed breadth and transferability depend on B3LYP/TZVP being accurate enough for the heavier elements (Al, Si, P, Cl, As, Se, Br), which the paper does not check against higher-level theory.
Editorial extensions
If this is right
- Models trained on QMe14S should generalize to molecular spectra for compounds containing heavier main-group elements (Al, Si, P, Cl, As, Se, Br) better than models trained on QM9S.
- The dataset's Hessian and derivative tensors allow ML models to simulate IR and Raman spectra without additional quantum-chemistry calculations for each new molecule.
- The 6 million dynamic configurations, with forces and Hessians, provide training data for machine-learned force fields that work away from equilibrium geometries.
- Unique higher-order tensors (first hyperpolarizability, octupole moment) and their derivatives in QMe14S enable prediction of nonlinear optical properties from machine-learned models.
Reading between the lines
- Editorially, the 500-instance minimum per element or functional group is a plausible rule of thumb for dataset coverage, but the paper does not test whether this threshold is optimal; measuring spectra accuracy as a function of that threshold would tell whether 500 is enough or overkill.
- Editorially, if B3LYP/TZVP holds up for the heavy elements, QMe14S could serve as the spectral analogue of QM9 for ML benchmark comparisons, and the same balancing strategy could be applied to other property databases.
- Editorially, the reported comparison to QM9S is limited to molecules containing only H, C, N, O, F in the test set; a direct test on molecules containing the newly added elements is needed to confirm the transferability gain the title implies.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces QMe14S, a dataset of 186,102 small organic molecules covering 14 elements (H, B, C, N, O, F, Al, Si, P, S, Cl, As, Se, Br) and 47 functional groups, with geometries and properties computed at the B3LYP/TZVP level. The dataset augments the authors' earlier QM9S dataset with 56,285 molecules from PubChem, and includes static properties (energy, forces, charges, multipole moments, polarizabilities, hyperpolarizability), harmonic IR, Raman and NMR spectra, and 6 million nonequilibrium configurations from ab initio molecular dynamics. Using their E(3)-equivariant neural network DetaNet, the authors report that models trained on QMe14S outperform models trained on QM9S in predicting IR and Raman spectra, and they demonstrate accurate prediction of NMR spectra on a few examples.
Significance. If the underlying DFT labels are reliable, QMe14S is a valuable resource for machine-learned molecular simulation. Its distinctive contributions are the inclusion of nine elements absent from QM9S, a wider functional-group coverage, and the unique availability of high-order tensors (first hyperpolarizability, octupole moment, Hessian, dipole and polarizability derivatives) together with dynamic configurations. The dataset is released with a reading script, and the ML benchmark is a useful sanity check. However, the central claim that the 14-element dataset improves spectral prediction is only tested on HCNOF molecules, and no higher-level validation is provided for the newly added elements, so the significance is conditional on closing those gaps.
major comments (4)
- [ML-Predicted Spectra] The quantitative comparison between DetaNet-QMe14S and DetaNet-QM9S is restricted to the 7,038 test molecules containing only H, C, N, O, and F; no spectral prediction accuracy is reported for molecules containing any of the nine elements added in QMe14S (Al, Si, P, Cl, As, Se, Br, and others). Because the central claim is that QMe14S improves spectral simulation for a 14-element chemical space, this test set does not substantiate that claim. Please report per-element or per-new-element errors on a test set that includes the added elements, or explicitly restrict the 'outperform' claim to HCNOF molecules.
- [Quantum mechanical calculation details / Technical Validation] No benchmark against higher-level theory or experiment is provided for the B3LYP/TZVP calculations on the newly added elements (Al, Si, P, Cl, As, Se, Br). The properties that make QMe14S distinctive — first hyperpolarizabilities, Raman intensities, and NMR shieldings — are precisely those where B3LYP/TZVP may have significant systematic errors for heavier elements. Please add a validation subsection that compares B3LYP/TZVP results against a higher-level method (e.g., CCSD(T), G4) or available experimental data for a few representative molecules per added element; otherwise the 'comprehensive' claim is not supported.
- [Table 1] Table 1's 'Total Numbers' column contains internal inconsistencies: Atomization Energy is listed as 5,907,000 for 186,102 molecules, Atomic Force as 5,907,000 although it should have the same 3N dimension as Atomic positions (6,093,102), and Dipole Moment as 6,093,102 although only three components per molecule are expected (186,102 × 3 = 558,306). These errors suggest the array sizes were not checked against the actual HDF5 files. Please correct the table and verify all entry counts against the released files.
- [Training details / ML-Predicted Properties] The reported improvement of DetaNet-QMe14S over DetaNet-QM9S on HCNOF test molecules conflates two effects: the addition of new elements and the increase in total training data (including additional HCNOF molecules from PubChem). Because the test set is HCNOF-only, the improvement could stem entirely from having more HCNOF training examples rather than from the elemental diversity. Please include an ablation experiment training on an HCNOF-only subset of QMe14S of comparable size to QM9S, or otherwise disentangle these factors.
minor comments (4)
- [ML-Predicted Spectra] The text contains a typo: 'DateNet-QMe14S' should be 'DetaNet-QMe14S'.
- [Quantum mechanical calculation details] The statement that ADMP simulations with a 1 fs step and 100 fs total time yield 6 million configurations is inconsistent with 186,102 molecules at 100 steps each (18.6 million); please clarify whether MD was run on a subset and how the 6 million figure is obtained.
- [Abstract / Data Records] The abstract claims NMR spectra for the dataset, but Table 1 lists shielding tensors for only 59,260 molecules and Table 2 describes NMR for '60k molecules selected from PubChem'; please reconcile this with the unqualified abstract statement.
- [Methods, Data collection and preprocessing] The sentence 'we also filtered several token that are commonly exist in SMILES' has grammatical errors and should be rephrased, e.g., 'we also filtered several token types that commonly occur in SMILES'.
Circularity Check
No significant circularity: held-out DFT benchmark and external label evaluation; only self-citations are prior DetaNet/QM9S work, not load-bearing.
full rationale
The paper's derivation chain is: select molecules, compute B3LYP/TZVP DFT properties and spectra, train DetaNet on a 90/5/5 split, and evaluate predictions against held-out DFT reference data. No predicted quantity is defined in terms of the quantity it is claimed to predict, and no fitted parameter is fed back into the definition of the benchmark target. The DetaNet model and the QM9S dataset come from the authors' prior work (ref 21), but these self-citations are not used as unverified premises that force the conclusion: the central claim that QMe14S-trained models outperform QM9S-trained models is tested by comparing both against independent, held-out B3LYP/TZVP reference spectra. Although the spectral comparison is restricted to 7,038 HCNOF molecules, this is a limitation in validating the new elemental coverage rather than a circularity. Likewise, the lack of higher-level benchmarks for the added heavier elements is a correctness/completeness risk, not a circularity. Overall, the evaluation is a standard supervised regression benchmark against externally computed DFT labels, so no circular step is present; the only mild self-reference is invoking the authors' own prior DetaNet architecture and QM9S dataset, which does not make the comparison circular.
Assumptions & free parameters
free parameters (2)
- Minimum element/functional-group coverage threshold =
500 instances
- Lorentzian broadening half-widths =
15 cm-1 (IR), 10 cm-1 (Raman)
assumptions (4)
- domain assumption B3LYP/TZVP is sufficiently accurate for all 14 elements included
- domain assumption RDKit substructure search correctly enumerates the 47 functional groups
- domain assumption ADMP trajectories at 1 fs steps over 100 fs sample chemically relevant conformations
- domain assumption The 90/5/5 stratified split prevents leakage of molecular conformations
Cite this review
Pith. "Pith review of QMe14S, A Comprehensive and Efficient Spectral Dataset for Small Organic Molecules." pith.science (2026). https://pith.science/paper/ZP7SL4WN
@misc{pith2026250118876,
author = {Pith},
title = {Pith review of: QMe14S, A Comprehensive and Efficient Spectral Dataset for Small Organic Molecules},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZP7SL4WN}},
note = {Machine review of arXiv:2501.18876}
}
read the original abstract
Developing machine learning protocols for molecular simulations requires comprehensive and efficient datasets. Here we introduce the QMe14S dataset, comprising 186,102 small organic molecules featuring 14 elements (H, B, C, N, O, F, Al, Si, P, S, Cl, As, Se, Br) and 47 functional groups. Using density functional theory at the B3LYP/TZVP level, we optimized the geometries and calculated properties including energy, atomic charge, atomic force, dipole moment, quadrupole moment, polarizability, octupole moment, first hyperpolarizability, and Hessian. At the same level, we obtained the harmonic IR, Raman and NMR spectra. Furthermore, we conducted ab initio molecular dynamics simulations to generate dynamic configurations and extract nonequilibrium properties, including energy, forces, and Hessians. By leveraging our E(3)-equivariant message-passing neural network (DetaNet), we demonstrated that models trained on QMe14S outperform those trained on the previously developed QM9S dataset in simulating molecular spectra. The QMe14S dataset thus serves as a comprehensive benchmark for molecular simulations, offering valuable insights into structure-property relationships.
Forward citations
Cited by 1 Pith paper
-
Leveraging active learning-enhanced machine-learned interatomic potential for efficient infrared spectra prediction
PALIRS combines active learning with MACE neural network potentials and a dipole moment model to predict infrared spectra of small organic molecules at a fraction of the DFT cost.
Reference graph
Works this paper leans on
-
[1]
M., Bruna, J., LeCun, Y ., Szlam, A
1 Bronstein, M. M., Bruna, J., LeCun, Y ., Szlam, A. & Vandergheynst, P. J. I. S. P. M. Geometric deep learning: going beyond euclidean data. IEEE Signal Process. Mag. 34, 18-42 (2017). 2 Atz, K., Grisoni, F. & Schneider, G. J. N. M. I. Geometric deep learning on molecular representations. Nat. Mach. Intell. 3, 1023-1032 (2021). 3 Gilmer, J., Schoenholz, ...
arXiv 2017
-
[2]
Introduction to methodology and encoding rules. J. Chem. Inf. Comp. Sci. 28, 31-36 (1988). 38 Perdew, J. P., Burke, K. & Ernzerhof, M. J. P. r. l. Generalized gradient approximation made simple. Phys. Rev. Lett. 77, 3865 (1996). 39 Schäfer, A., Huber, C. & Ahlrichs, R. J. T. J. o. c. p. Fully optimized contracted Gaussian basis sets of triple zeta valence...
work page 1988
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.