REVIEW 4 major objections 5 minor 8 references
Constructing and explaining machine learning models for chemistry: example of the exploration and design of boron-based Lewis acids
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read For boron Lewis acids built on a fixed molecular scaffold, fluoride ion affinity can be predicted to within about 5 kJ/mol by a linear model using Hammett-style substituent descriptors, and the model's interpretations yield concrete…
desk verdict A solid, reproducible low-data ML benchmark, but the design rules and the 'oracle' confirmation never leave the M062X label space, so the physical claims need an independent high-level check. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Hammett-extended substituent descriptor vector, a set of 36 features per molecule encoding the ortho, meta, and para substituents through computed benzoic-acid proxies: NBO partial charges, infrared carbonyl stretching and COH bending frequencies and intensities, Sterimol steric parameters B1, B5, and L, and the carbonyl–ring torsion angle. These descriptors make the electronic demand of each substituent position explicit, so a linear model trained on them is interpretable by construction. For the cross-scaffold and physical-interpretation parts, the quantum descriptors, 43 DFT-derived features including frontier orbital energies and the boron NPA charge, carry the argument.
What would settle it
Compute FIA at a higher level of theory, such as CCSD(T)/CBS or DSD-PBEP86/aug-cc-pVTZ, for a stratified sample of the ONO test set and compare those labels with the M062X/6-31G(d) labels; if the oracle's errors against the high-level labels exceed the reported 5.39 kJ/mol by more than the known method bias around 13 kJ/mol, or if the decision tree's para NBO charge threshold near -0.59e stops separating strong from medium acids, then the reported accuracy and design rules are artifacts of the label level.
Extended reading notes
Core claim
Using M062X/6-31G(d) isodesmic FIA values as labels and concatenated Hammett-extended plus cheminformatics descriptors filtered to 126 features, a linear ridge regression oracle predicts FIA for the ONO scaffold with a mean absolute error of 5.39 kJ/mol on the test set (R² = 0.98), while the same test set gives a mean absolute error of 23 kJ/mol for the graph neural network baseline. Interpretability of this oracle and of decision trees indicates the para substituent's electron demand is the dominant lever: mesomeric electron-withdrawing groups (CN, NO₂) at the para position set the accessible FIA range, while ortho and meta substituents fine-tune it. For Lewis acidity itself, linear models on quantum descriptors point to absolute electronegativity, a combination of frontier orbital energies, and the boron partial charge as the two controlling features, with electronegativity leading, which the authors read as orbital interactions dominating Coulombic ones.
Load-bearing premise
The quantitative claims rest on treating M062X/6-31G(d) isodesmic FIA values as ground truth; that level differs from CCSD(T)/CBS references by a mean absolute error of about 13.1 kJ/mol (SI Table S3), larger than the 5.39 kJ/mol test-set error, so a systematic scaffold-dependent label bias would change the accuracy, the comparison with the GNN, and the derived design rules.
Editorial extensions
If this is right
- For ONO-scaffold boron Lewis acids, FIA can be predicted from fast substituent descriptors with a test-set mean absolute error near 5 kJ/mol, so screening the full 2197-molecule space becomes practical without new DFT calculations.
- A para mesomeric electron-withdrawing group sets the accessible FIA range, and ortho and meta substituents then tune within that range, giving a recipe for targeting a specific Lewis acidity band.
- The linear dependence of FIA on absolute electronegativity and boron partial charge implies that Lewis acidity in these constrained boron acids is dominated by orbital interactions rather than purely Coulombic ones.
- After task-specific feature selection, the same simple linear model extrapolates from ONO to the related NNN scaffold with a mean absolute error around 14 kJ/mol, showing that scaffold transfer is possible with simple models.
- The success of linear models over graph neural networks in this low-data regime suggests that, for well-defined scaffolds, physically meaningful substituent descriptors can replace large black-box models.
Reading between the lines
- Because the Hammett-extended descriptors are scaffold-independent in their feature definitions, the same workflow could be applied to other reactivity endpoints, such as hydride affinity or electrophilicity, on other p-block Lewis acids whenever substituent parameters exist.
- The decision tree's threshold on the para NBO charge (-0.59e) is a concrete, testable design rule: synthesizing ONO compounds with para substituents straddling that threshold and measuring FIA or Gutmann–Beckett shifts would show whether the suggested boundary is physical or an artifact of the DFT label set.
- The paper reports a strong FIA–HIA correlation (Pearson's r = 0.95) for the Lewis acids it benchmarks; a high-level comparison of FIA and hydride affinity on the constrained ONO scaffold would sharpen the claim that these compounds behave as soft Lewis acids.
- The reported comparison against a 49k-molecule graph neural network is a low-data comparison on one scaffold; a fairer test would train the GNN on a subset of the same scaffold or use a scaffold-aware split, which the paper does not do.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents an explainable machine-learning workflow for predicting fluoride ion affinity (FIA) of boron-based Lewis acids, using a restricted chemical space of four scaffolds and low-data regimes. The authors benchmark multiple descriptor types (Morgan fingerprints, RDKit descriptors, quantum descriptors, and Hammett-extended descriptors) with several regression algorithms on an ONO-scaffold dataset, and report an optimized linear model combining RDKit and Hammett-extended descriptors with a test MAE of 5.39 kJ/mol (R²=0.98). They also study cross-scaffold extrapolation via feature selection, interpret the models to relate Lewis acidity to electronegativity and boron partial charge, and derive decision-tree rules for molecular design, notably that a para mesomeric electron-withdrawing group sets the FIA range, with ortho and meta substituents refining it. The oracle is then used to screen the full 2197-molecule ONO space to support the design rules.
Significance. If the reported accuracy transfers to physically meaningful FIA values, the paper offers a practical interpretable workflow for designing boron Lewis acids with targeted acidity, with the notable strengths of a publicly available code repository, reproducible 10-fold cross-validation, and a clear comparison across descriptor families. The interpretability analysis is thoughtful and connects to existing Hammett/Sigman substituent parameters. However, the central quantitative claim currently refers to prediction of M062X/6-31G(d) isodesmic FIA labels, and the reported label error against CCSD(T)/CBS references (MAE 13.1 kJ/mol, SI Table S3) is larger than the model test error of 5.39 kJ/mol; moreover, the design-rule confirmation uses the same oracle labels as the tree that generated the rule. These issues cap the significance until independent validation is provided.
major comments (4)
- [SI Table S3; Results, 'Lewis acidity scale'] The label method M062X/6-31G(d) isodesmic FIA has a mean absolute error of 13.1 kJ/mol against CCSD(T)/CBS references (SI Table S3), while the oracle's test MAE is 5.39 kJ/mol. The headline claim of 'highly accurate predictions (<6 kJ/mol)' is therefore accurate only with respect to the DFT labels, not to the physical FIA. The authors should quantify how the 13.1 kJ/mol label uncertainty propagates to the reported model error, and check whether the M062X bias is systematic or position/substituent-dependent within the ONO space by computing high-level FIA values for a representative subset of ONO molecules, especially those with para CN/NO2 groups. Without this, the design rules could reflect a DFT artifact.
- [Results, 'Chemometrics on ONO chemical space'] The confirmation of the decision-tree rule is circular: the oracle was trained on the same M062X FIA labels that were used to grow the tree in Scheme 1, and this oracle is then used to screen the entire ONO space and produce the violin plots in Figure 7 that are interpreted as confirming the para-substituent rule. Screening with the same label source cannot provide independent evidence about true Lewis acidity. The authors should validate the rule with high-level reference calculations (e.g., CCSD(T)/CBS) on a designed set of ONO molecules, or with experimental measurements, and should rephrase the current text so that the oracle screen is presented as an interpolation/enumeration of the model, not as independent confirmation.
- [Results, 'Constructing models' (GNN comparison)] The comparison with the graph neural network of Greb and co-workers is not balanced as reported. The GNN gives MAE = 23 kJ/mol on the ONO testing set, but this appears to be a zero-shot evaluation of a pretrained model rather than a model trained or fine-tuned on the ONO training split. To support the claim that the proposed approach 'surpasses conventional black-box deep learning models in low-data regimes', the authors should either train or fine-tune the GNN on the ONO data with the same split, or clearly state and justify the zero-shot comparison.
- [SI S5, 'Extrapolation' (quantum descriptors)] The feature-selection procedure for extrapolating from ONO to NNN appears to use the target scaffold's data to select features: features are ranked by differences between ONO and NNN and by correlation with FIA, and then models are assessed on the NNN scaffold after systematic feature removal. This leaks information from the test scaffold into model construction and can lead to optimistic MAE values. The authors should use a nested cross-validation procedure or a separate held-out scaffold to demonstrate that the selected features generalize, or explicitly state that the NNN scaffold is used only as a feature-selection criterion rather than as an independent test.
minor comments (5)
- [Abstract] The abstract contains a typo ('electron-ccepting') and should be corrected.
- [General] Several figure cross-references appear as broken placeholders, e.g., 'Figure 1Error: Reference source not found', 'Figure 2Error: Reference source not found', and 'Figure 3.Error: Reference source not foundA/B/C'; these should be repaired before publication.
- [Results, 'Constructing models'] The heading 'Chemometrics on ONO chemical space' is run together with the preceding text and should be separated into a clear subsection title.
- [Results, 'Interpretability'] The linear relationship FIA = 243 σm + 91 σp + 351 is reported without units for the coefficients; specifying the units (kJ/mol per σ unit) would improve clarity.
- [SI S5, 'Hammett-extended descriptors'] The statement that 'no feature selection could improve the MAE of 86.4 kJ.mol-1' for ONO-to-NNN prediction is useful, but the authors could briefly explain why descriptor sets that ignore the scaffold cannot transfer, since this is directly related to the extrapolation discussion.
Circularity Check
Design-rule "confirmation" via oracle screen is a same-label consistency check, not independent evidence; the core MAE<6 kJ/mol predictive claim is not circular.
-
fitted input called prediction
[Section 'Molecular design – Interpretable ML' and 'Chemometrics on ONO chemical space'; Scheme 1, Figure 7.]
"To confirm these results, we need more data for statistical analyses on the diverse substituents. For that, we have used the oracle previously developed to screen the entire ONO chemical space and provide precise FIA values to enrich our database. ... Using the root node criterion from the tree in Scheme 1, we compared the FIA distributions for ONO molecules with a mesomeric electron-withdrawing group in the para position to those without."
The oracle is a linear-regression model trained on the same M062X/6-31G(d) isodesmic FIA labels used to grow the decision tree, and its 126 selected features include the Hammett-extended NBO=O parameters that also define the tree's root split (NBO=O_p threshold). Screening the full 2197-molecule space with this oracle and then showing that the tree's own root criterion separates the oracle-predicted FIA distributions is therefore not independent confirmation: both models share the same training labels and a key input feature. The "precise FIA values" are model predictions, not new ab initio data, so the violin plots in Figure 7 demonstrate self-consistency between two models fitted to the same labels, not independent evidence about true Lewis acidity.
full rationale
Most of the derivation chain is self-contained. FIA labels are computed with a stated isodesmic M062X/6-31G(d) protocol, benchmarked against CCSD(T)/CBS references in SI Table S3; the main predictive claim (MAE 5.39 kJ/mol on a held-out test set) is evaluated against held-out labels of the same type and compared to an external GNN baseline, so it is not circular. The Hammett-extended descriptors are external substituent parameters from Sigman et al., not derived from the target property. The interpretability equation FIA = 60.0 chi + 8.15 NPA_charge + 161 is presented as a fitted regression summary, not as a deduction. No load-bearing self-citations occur; reference [5] is an example citation only. The M062X versus CCSD(T)/CBS label error (13.1 kJ/mol) is an accuracy and validity risk, not a circularity. The one genuine circular step is the oracle-based 'confirmation' of the decision-tree design rule: the oracle is a model fitted to the same DFT labels and sharing the key NBO=O_p feature with the tree, so screening the chemical space with it and observing separation along the tree's own root threshold is a consistency check between two models trained on identical ground truth, not independent evidence. This affects the design-rule confirmation but not the central regression accuracy claim, giving a partial circularity score of 5.
Assumptions & free parameters
free parameters (4)
- Oracle linear regression coefficients and selected feature subset =
126 selected features; coefficient values in GitHub repository
- FIA-energy linear relationship coefficients =
60.0 kJ/(mol·eV) for electronegativity, 8.15 kJ/(mol·e) for NPA charge, intercept 161 kJ/mol
- Decision tree split thresholds =
NBO=O para threshold at -0.59 e, plus subsequent ortho and meta splits
- Feature selection choices for cross-scaffold extrapolation =
Exclusion of NPA_Rydberg, dipole moment, LUMO energy, and other features for ONO-to-NNN transfer
assumptions (4)
- domain assumption M062X/6-31G(d) isodesmic FIA values are accurate enough to serve as ground-truth labels.
- domain assumption Gas-phase FIA is an adequate proxy for Lewis acidity.
- domain assumption Hammett-extended substituent parameters computed on benzoic acid are transferable to boron scaffolds.
- domain assumption Auto-QChem quantum descriptors derived from DFT-optimized geometries capture the electronic features relevant to Lewis acidity.
Cite this review
Pith. "Pith review of Constructing and explaining machine learning models for chemistry: example of the exploration and design of boron-based Lewis acids." pith.science (2026). https://pith.science/paper/SWCQ6M7F
@misc{pith2026250101576,
author = {Pith},
title = {Pith review of: Constructing and explaining machine learning models for chemistry: example of the exploration and design of boron-based Lewis acids},
year = {2026},
howpublished = {\url{https://pith.science/paper/SWCQ6M7F}},
note = {Machine review of arXiv:2501.01576}
}
read the original abstract
The integration of machine learning (ML) into chemistry offers transformative potential in the design of molecules with targeted properties. However, the focus has often been on creating highly efficient predictive models, sometimes at the expense of interpretability. In this study, we leverage explainable AI techniques to explore the rational design of boron-based Lewis acids, which play a pivotal role in organic reactions due to their electron-ccepting properties. Using Fluoride Ion Affinity as a proxy for Lewis acidity, we developed interpretable ML models based on chemically meaningful descriptors, including ab initio computed features and substituent-based parameters derived from the Hammett linear free-energy relationship. By constraining the chemical space to well-defined molecular scaffolds, we achieved highly accurate predictions (mean absolute error < 6 kJ/mol), surpassing conventional black-box deep learning models in low-data regimes. Interpretability analyses of the models shed light on the origin of Lewis acidity in these compounds and identified actionable levers to modulate it through the nature and positioning of substituents on the molecular scaffold. This work bridges ML and chemist's way of thinking, demonstrating how explainable models can inspire molecular design and enhance scientific understanding of chemical reactivity.
Reference graph
Works this paper leans on
-
[1]
(1) Erdmann, P.; Leitner, J.; Schwarz, J.; Greb, L. An Extensive Set of Accurate Fluoride Ion Affinities for p Block‐ Element Lewis Acids and Basic Design Principles for Strong Fluoride Ion Acceptors. ChemPhysChem 2020, 21 (10), 987–
work page 2020
-
[3]
https://doi.org/10.1021/om3011195
Organometallics 2013, 32 (1), 317–322. https://doi.org/10.1021/om3011195. (16) Herrington, T. J.; Thom, A. J. W.; White, A. J. P.; Ashley, A. E. Novel H2 Activation by a Tris[3,5- Bis(Trifluoromethyl)Phenyl]Borane Frustrated Lewis Pair. Dalton Trans. 2012, 41 (30),
-
[4]
http://www.daylight.com/dayhtml/doc/theory/theory.smarts.html
SMARTS—a Language for Describing Molecular Patterns. http://www.daylight.com/dayhtml/doc/theory/theory.smarts.html. (35) XAI_boron_LA. https://github.com/jfenogli/XAI_boron_LA. (36) Santiago, C. B.; Milo, A.; Sigman, M. S. Developing a Modern Approach To Account for Steric Effects in Hammett-Type Correlations. J. Am. Chem. Soc. 2016, 138 (40), 13424–13430...
-
[7]
Introduction to Methodology and Encoding Rules. J. Chem. Inf. Model. 1988, 28 (1), 31–36. https://doi.org/10.1021/ci00057a005. (30) Rogers, D.; Hahn, M. Extended-Connectivity Fingerprints. J. Chem. Inf. Model. 2010, 50 (5), 742–754. https://doi.org/10.1021/ci100050t. (31) RDKit. https://www.rdkit.org/ (accessed 2022-11-09). (32) DeepChem. https://deepchem...
-
[117]
(8) Ashley, A. E.; Herrington, T. J.; Wildgoose, G. G.; Zaher, H.; Thompson, A. L.; Rees, N. H.; Krämer, T.; O’Hare, D. Separating Electrophilicity and Lewis Acidity: The Synthesis, Characterization, and Electrochemistry of the Electron Deficient Tris (Aryl)Boranes B(C 6 F 5 ) 3– n (C 6 Cl 5 ) n ( n = 1–3). J. Am. Chem. Soc. 2011, 133 (37), 14727–14740. h...
-
[994]
https://doi.org/10.1002/cphc.202000244. (2) Greb, L. Lewis Superacids: Classifications, Candidates, and Applications. Chem. – Eur. J. 2018, 24 (68), 17881– 17896. https://doi.org/10.1002/chem.201802698. (3) Parr, R. G.; Pearson, R. G. Absolute Hardness: Companion Parameter to Absolute Electronegativity. J. Am. Chem. Soc. 1983, 105 (26), 7512–7516. https:/...
-
[2010]
(11) Beckett, M. A.; Rugen-Hankey, M. P.; Strickland, G. C.; Varma, K. S. Lewis Acidity in Haloalkyl Orthoborate and Metaborate Esters. Phosphorus Sulfur Silicon Relat. Elem. 2001, 169 (1), 113–116. https://doi.org/10.1080/10426500108546603. (12) Beckett, M. A.; Owen, P.; Varma, K. S. Synthesis and Lewis Acidity of Triorganosilyl and Triorganostannyl Este...
-
[9019]
https://doi.org/10.1039/c2dt30384a. (17) Lu, Z.; Cheng, Z.; Chen, Z.; Weng, L.; Li, Z. H.; Wang, H. Heterolytic Cleavage of Dihydrogen by “Frustrated Lewis Pairs” Comprising Bis(2,4,6 tris(Trifluoromethyl)Phenyl)Borane and Amines: Stepwise versus Concerted Mechanism.‐ Angew. Chem. Int. Ed. 2011, 50 (51), 12227–12231. https://doi.org/10.1002/anie.201104999...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.