Pith. sign in

REVIEW 4 major objections 4 minor 1 cited by

QMe14S, A Comprehensive and Efficient Spectral Dataset for Small Organic Molecules

T0 review · 4 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read QMe14S, a dataset of 186,102 small organic molecules spanning 14 elements and 47 functional groups with DFT-level IR, Raman and NMR spectra, improves machine-learned molecular spectrum simulation compared with training on QM9S.

desk verdict A genuinely useful dataset extension, but the 14-element claim is untested: the ML benchmark only covers HCNOF molecules. read the letter →

arxiv 2501.18876 v1 pith:ZP7SL4WN submitted 2025-01-31 physics.chem-ph cs.LG

classification physics.chem-phcs.LG
keywords QMe14SdatasetmolecularspectradensityfunctionaltheoryE(3)-equivariantneuralnetworkIRRamanNMRspectroscopysmallorganicmoleculeschemicalspacediversityabinitiodynamics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces QMe14S, a dataset of 186,102 small organic molecules built to train machine-learning models that predict molecular spectra. It argues that previous datasets such as QM9S cover too few elements and functional groups, so the authors supplement them with 56,285 PubChem molecules to ensure every one of 14 elements and 47 functional groups appears at least 500 times. All molecules are calculated at the B3LYP/TZVP level, yielding IR, Raman and NMR spectra plus static and dynamic molecular properties. Using their E(3)-equivariant network DetaNet, the authors show that models trained on QMe14S reproduce DFT spectra more closely than models trained on QM9S, particularly for molecules with hydrazine- and nitrogen-rich groups. A sympathetic reader would take the paper's central claim to be that a balanced, diverse but still computationally affordable spectral dataset improves machine-learned spectrum simulation.

What carries the argument

The load-bearing piece is the dataset's construction itself: the RDKit substructure search on PubChem used to top up underrepresented elements and functional groups, followed by geometry optimization and harmonic frequency analysis at B3LYP/TZVP in Gaussian 16, and ADMP molecular dynamics for nonequilibrium configurations. On the modelling side, DetaNet — an E(3)-equivariant message-passing neural network that predicts scalar and high-order tensorial properties — converts predicted Hessians, dipole derivatives and polarizability derivatives into IR and Raman intensities via normal-mode analysis, and predicted shielding tensors into NMR chemical shifts. The uniform element and functional-group coverage, rather than raw dataset size, is what the paper claims makes the dataset efficient for machine learning.

What would settle it

Compare B3LYP/TZVP vibrational frequencies and IR/Raman intensities for a set of Se- and Br-containing molecules against CCSD(T)-level reference calculations; if the heavy-element errors are systematically larger than for HCNOF molecules, the uniform-quality assumption behind QMe14S fails. Alternatively, retrain DetaNet on QMe14S but test only on molecules containing elements absent from QM9S; a large accuracy drop would indicate the dataset's diversity claim is overstated.

Watch

Extended reading notes

Core claim

QMe14S is a spectral and property dataset spanning 186,102 molecules, 14 elements (H, B, C, N, O, F, Al, Si, P, S, Cl, As, Se, Br) and 47 functional groups, with static and dynamic properties computed at the B3LYP/TZVP level. The dataset adds 56,285 molecules from PubChem to the earlier QM9S dataset so that every element and functional group appears in over 500 instances, producing a much flatter distribution of element and functional-group frequencies than QMugs or PubChemQC. In addition to equilibrium geometries, energies, charges, multipole moments, polarizabilities, Hessians and derivative tensors, QMe14S contains roughly 6 million ab initio molecular dynamics configurations with energies, forces and dipole moments, of which 10 per molecule have additional Hessian and polarizability data. The authors demonstrate with DetaNet that models trained on QMe14S predict IR and Raman spectra closer to DFT than models trained on QM9S, with cosine similarities of about 95% for IR and 92% for Raman on the HCNOF test subset, and that predicted 13C and 1H NMR spectra match DFT calculations closely.

Load-bearing premise

The claimed breadth and transferability depend on B3LYP/TZVP being accurate enough for the heavier elements (Al, Si, P, Cl, As, Se, Br), which the paper does not check against higher-level theory.

Editorial extensions

If this is right

  • Models trained on QMe14S should generalize to molecular spectra for compounds containing heavier main-group elements (Al, Si, P, Cl, As, Se, Br) better than models trained on QM9S.
  • The dataset's Hessian and derivative tensors allow ML models to simulate IR and Raman spectra without additional quantum-chemistry calculations for each new molecule.
  • The 6 million dynamic configurations, with forces and Hessians, provide training data for machine-learned force fields that work away from equilibrium geometries.
  • Unique higher-order tensors (first hyperpolarizability, octupole moment) and their derivatives in QMe14S enable prediction of nonlinear optical properties from machine-learned models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorially, the 500-instance minimum per element or functional group is a plausible rule of thumb for dataset coverage, but the paper does not test whether this threshold is optimal; measuring spectra accuracy as a function of that threshold would tell whether 500 is enough or overkill.
  • Editorially, if B3LYP/TZVP holds up for the heavy elements, QMe14S could serve as the spectral analogue of QM9 for ML benchmark comparisons, and the same balancing strategy could be applied to other property databases.
  • Editorially, the reported comparison to QM9S is limited to molecules containing only H, C, N, O, F in the test set; a direct test on molecules containing the newly added elements is needed to confirm the transferability gain the title implies.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces QMe14S, a dataset of 186,102 small organic molecules covering 14 elements (H, B, C, N, O, F, Al, Si, P, S, Cl, As, Se, Br) and 47 functional groups, with geometries and properties computed at the B3LYP/TZVP level. The dataset augments the authors' earlier QM9S dataset with 56,285 molecules from PubChem, and includes static properties (energy, forces, charges, multipole moments, polarizabilities, hyperpolarizability), harmonic IR, Raman and NMR spectra, and 6 million nonequilibrium configurations from ab initio molecular dynamics. Using their E(3)-equivariant neural network DetaNet, the authors report that models trained on QMe14S outperform models trained on QM9S in predicting IR and Raman spectra, and they demonstrate accurate prediction of NMR spectra on a few examples.

Significance. If the underlying DFT labels are reliable, QMe14S is a valuable resource for machine-learned molecular simulation. Its distinctive contributions are the inclusion of nine elements absent from QM9S, a wider functional-group coverage, and the unique availability of high-order tensors (first hyperpolarizability, octupole moment, Hessian, dipole and polarizability derivatives) together with dynamic configurations. The dataset is released with a reading script, and the ML benchmark is a useful sanity check. However, the central claim that the 14-element dataset improves spectral prediction is only tested on HCNOF molecules, and no higher-level validation is provided for the newly added elements, so the significance is conditional on closing those gaps.

major comments (4)
  1. [ML-Predicted Spectra] The quantitative comparison between DetaNet-QMe14S and DetaNet-QM9S is restricted to the 7,038 test molecules containing only H, C, N, O, and F; no spectral prediction accuracy is reported for molecules containing any of the nine elements added in QMe14S (Al, Si, P, Cl, As, Se, Br, and others). Because the central claim is that QMe14S improves spectral simulation for a 14-element chemical space, this test set does not substantiate that claim. Please report per-element or per-new-element errors on a test set that includes the added elements, or explicitly restrict the 'outperform' claim to HCNOF molecules.
  2. [Quantum mechanical calculation details / Technical Validation] No benchmark against higher-level theory or experiment is provided for the B3LYP/TZVP calculations on the newly added elements (Al, Si, P, Cl, As, Se, Br). The properties that make QMe14S distinctive — first hyperpolarizabilities, Raman intensities, and NMR shieldings — are precisely those where B3LYP/TZVP may have significant systematic errors for heavier elements. Please add a validation subsection that compares B3LYP/TZVP results against a higher-level method (e.g., CCSD(T), G4) or available experimental data for a few representative molecules per added element; otherwise the 'comprehensive' claim is not supported.
  3. [Table 1] Table 1's 'Total Numbers' column contains internal inconsistencies: Atomization Energy is listed as 5,907,000 for 186,102 molecules, Atomic Force as 5,907,000 although it should have the same 3N dimension as Atomic positions (6,093,102), and Dipole Moment as 6,093,102 although only three components per molecule are expected (186,102 × 3 = 558,306). These errors suggest the array sizes were not checked against the actual HDF5 files. Please correct the table and verify all entry counts against the released files.
  4. [Training details / ML-Predicted Properties] The reported improvement of DetaNet-QMe14S over DetaNet-QM9S on HCNOF test molecules conflates two effects: the addition of new elements and the increase in total training data (including additional HCNOF molecules from PubChem). Because the test set is HCNOF-only, the improvement could stem entirely from having more HCNOF training examples rather than from the elemental diversity. Please include an ablation experiment training on an HCNOF-only subset of QMe14S of comparable size to QM9S, or otherwise disentangle these factors.
minor comments (4)
  1. [ML-Predicted Spectra] The text contains a typo: 'DateNet-QMe14S' should be 'DetaNet-QMe14S'.
  2. [Quantum mechanical calculation details] The statement that ADMP simulations with a 1 fs step and 100 fs total time yield 6 million configurations is inconsistent with 186,102 molecules at 100 steps each (18.6 million); please clarify whether MD was run on a subset and how the 6 million figure is obtained.
  3. [Abstract / Data Records] The abstract claims NMR spectra for the dataset, but Table 1 lists shielding tensors for only 59,260 molecules and Table 2 describes NMR for '60k molecules selected from PubChem'; please reconcile this with the unqualified abstract statement.
  4. [Methods, Data collection and preprocessing] The sentence 'we also filtered several token that are commonly exist in SMILES' has grammatical errors and should be rephrased, e.g., 'we also filtered several token types that commonly occur in SMILES'.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: held-out DFT benchmark and external label evaluation; only self-citations are prior DetaNet/QM9S work, not load-bearing.

full rationale

The paper's derivation chain is: select molecules, compute B3LYP/TZVP DFT properties and spectra, train DetaNet on a 90/5/5 split, and evaluate predictions against held-out DFT reference data. No predicted quantity is defined in terms of the quantity it is claimed to predict, and no fitted parameter is fed back into the definition of the benchmark target. The DetaNet model and the QM9S dataset come from the authors' prior work (ref 21), but these self-citations are not used as unverified premises that force the conclusion: the central claim that QMe14S-trained models outperform QM9S-trained models is tested by comparing both against independent, held-out B3LYP/TZVP reference spectra. Although the spectral comparison is restricted to 7,038 HCNOF molecules, this is a limitation in validating the new elemental coverage rather than a circularity. Likewise, the lack of higher-level benchmarks for the added heavier elements is a correctness/completeness risk, not a circularity. Overall, the evaluation is a standard supervised regression benchmark against externally computed DFT labels, so no circular step is present; the only mild self-reference is invoking the authors' own prior DetaNet architecture and QM9S dataset, which does not make the comparison circular.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claim rests on standard DFT assumptions and on hand-chosen design thresholds, but not on newly invented physical entities. The dataset itself is a resource, not a postulated entity.

free parameters (2)
  • Minimum element/functional-group coverage threshold = 500 instances
    Chosen by hand to decide how many PubChem molecules to add; directly determines dataset composition and the claimed diversity.
  • Lorentzian broadening half-widths = 15 cm-1 (IR), 10 cm-1 (Raman)
    Chosen by hand for spectral visualization and affects the spectral similarity metrics used in the ML benchmark.
assumptions (4)
  • domain assumption B3LYP/TZVP is sufficiently accurate for all 14 elements included
    Used for all QM calculations; no comparison to higher-level theory is provided (Methods, Quantum mechanical calculation details).
  • domain assumption RDKit substructure search correctly enumerates the 47 functional groups
    The supplement selection and the claim of coverage rest on this; no manual curation or validation of the group assignments is described (Methods, Data collection and preprocessing).
  • domain assumption ADMP trajectories at 1 fs steps over 100 fs sample chemically relevant conformations
    Used to generate dynamic configurations; no validation of sampling convergence is shown (Methods, Quantum mechanical calculation details).
  • domain assumption The 90/5/5 stratified split prevents leakage of molecular conformations
    All 100 MD configurations of a molecule are kept in the same split, but the random seed and split code are not released (Training details).

how reviews work

0 comments
Cite this review

Pith. "Pith review of QMe14S, A Comprehensive and Efficient Spectral Dataset for Small Organic Molecules." pith.science (2026). https://pith.science/paper/ZP7SL4WN

@misc{pith2026250118876,
  author       = {Pith},
  title        = {Pith review of: QMe14S, A Comprehensive and Efficient Spectral Dataset for Small Organic Molecules},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZP7SL4WN}},
  note         = {Machine review of arXiv:2501.18876}
}
read the original abstract

Developing machine learning protocols for molecular simulations requires comprehensive and efficient datasets. Here we introduce the QMe14S dataset, comprising 186,102 small organic molecules featuring 14 elements (H, B, C, N, O, F, Al, Si, P, S, Cl, As, Se, Br) and 47 functional groups. Using density functional theory at the B3LYP/TZVP level, we optimized the geometries and calculated properties including energy, atomic charge, atomic force, dipole moment, quadrupole moment, polarizability, octupole moment, first hyperpolarizability, and Hessian. At the same level, we obtained the harmonic IR, Raman and NMR spectra. Furthermore, we conducted ab initio molecular dynamics simulations to generate dynamic configurations and extract nonequilibrium properties, including energy, forces, and Hessians. By leveraging our E(3)-equivariant message-passing neural network (DetaNet), we demonstrated that models trained on QMe14S outperform those trained on the previously developed QM9S dataset in simulating molecular spectra. The QMe14S dataset thus serves as a comprehensive benchmark for molecular simulations, offering valuable insights into structure-property relationships.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Leveraging active learning-enhanced machine-learned interatomic potential for efficient infrared spectra prediction

    physics.chem-ph 2025-06 conditional novelty 5.0 of 10

    PALIRS combines active learning with MACE neural network potentials and a dipole moment model to predict infrared spectra of small organic molecules at a fraction of the DFT cost.

Reference graph

Works this paper leans on

2 extracted references · 1 canonical work pages · cited by 1 Pith paper

  1. [1]

    M., Bruna, J., LeCun, Y ., Szlam, A

    1 Bronstein, M. M., Bruna, J., LeCun, Y ., Szlam, A. & Vandergheynst, P. J. I. S. P. M. Geometric deep learning: going beyond euclidean data. IEEE Signal Process. Mag. 34, 18-42 (2017). 2 Atz, K., Grisoni, F. & Schneider, G. J. N. M. I. Geometric deep learning on molecular representations. Nat. Mach. Intell. 3, 1023-1032 (2021). 3 Gilmer, J., Schoenholz, ...

  2. [2]

    Introduction to methodology and encoding rules. J. Chem. Inf. Comp. Sci. 28, 31-36 (1988). 38 Perdew, J. P., Burke, K. & Ernzerhof, M. J. P. r. l. Generalized gradient approximation made simple. Phys. Rev. Lett. 77, 3865 (1996). 39 Schäfer, A., Huber, C. & Ahlrichs, R. J. T. J. o. c. p. Fully optimized contracted Gaussian basis sets of triple zeta valence...

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.