Pith. sign in

REVIEW 5 major objections 6 minor 11 references

Toward Routine CSP of Pharmaceuticals: A Fully Automated Protocol Using Neural Network Potentials

T0 review · 5 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A fully automated cloud-parallel crystal structure prediction protocol driven by the Lavo-NN neural network potential generates and correctly ranks all known $Z'=1$ polymorphs of a 49-molecule pharmaceutical benchmark at an average cost…

desk verdict A serious and useful CSP automation paper whose headline generalization claim needs a holdout check before it can be taken at face value. read the letter →

arxiv 2507.16218 v1 pith:Q2ZO2OWW submitted 2025-07-22 physics.chem-ph cs.LG

classification physics.chem-phcs.LG
keywords crystalstructurepredictionneural-networkpotentialspolymorphismpharmaceuticalsolidformscreeningmany-bodyexpansioncloudcomputingretrospectivebenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that crystal structure prediction (CSP) for pharmaceutical molecules can be made fully automatic, fast enough to run alongside experimental screening, without losing the ability to find and correctly rank every experimentally known polymorph. The authors build a neural network potential specialized for molecular crystals, Lavo-NN, and wrap it in a cloud-parallel protocol that takes only a SMILES string and a chirality flag as input. On a retrospective benchmark of 49 drug-like molecules with 110 known $Z'=1$ polymorphs, the protocol generated structures matching all of them and ranked most near the bottom of the energy landscape, at an average cost of about 8.4k CPU hours per molecule. If this holds, CSP stops being a specialist, high-cost exercise and becomes a routine de-risking tool in drug development.

What carries the argument

Lavo-NN is an equivariant message-passing neural network whose crystal energy is decomposed through the many-body expansion truncated at second order: an intramolecular term plus a sum of dimer interaction energies. Message passing runs only along covalent intramolecular edges, and the intermolecular readout augments a short-range neural term with a long-range polarizable force field, with some force-field parameters predicted by the network; this decomposition is what lets the model reach near-DFT accuracy at low cost. Around the potential sits a protocol that generates dense crystal candidates geometrically, refines them with a semi-local Monte Carlo search using interpolated moves, tracks convergence per space group with a Poisson-lognormal species-abundance model, and optionally re-ranks the lowest-energy structures with periodic PBE-D3(BJ) plus a monomer PBE0 correction.

What would settle it

Search the 10k-molecule training set and the extracted conformers and dimers for any of the 49 benchmark molecules or their close tautomers, then retrain Lavo-NN with those molecules strictly excluded and rerun the benchmark to see whether all 110 polymorphs are still generated and ranked near the bottom of the landscape.

Watch

Extended reading notes

Core claim

The central claim is that a single purpose-built neural network potential, trained once on a large corpus of pharmaceutical-like crystal structures and dimer interactions, can both enumerate and rank the experimentally observable polymorphs of new drug molecules without system-specific re-training or manual specification. Specifically, the paper reports generating structures that match all 110 known $Z'=1$ experimental polymorphs of its 49-molecule benchmark, ranking 87% of the matched structures within the top 50 of their predicted landscapes, and doing so at an average cost of about 8.4k CPU hours per molecule. The authors argue, through case studies and a semi-blinded exercise using only powder diffraction patterns, that the predicted landscapes can resolve ambiguities in experimental data and identify the structures of forms that lack single-crystal data.

Load-bearing premise

The benchmark molecules must be absent from Lavo-NN's training data; the paper trains on structures built from 10k SMILES drawn from a public molecular dataset and never states that the 49 test molecules were held out, so the headline result could be retrieval from training data rather than generalization.

Editorial extensions

If this is right

  • CSP can be run routinely on real drug candidates: a typical molecule costs about 8.4k CPU hours and finishes in days of wall time, making solid-form screens practical during lead optimization.
  • The protocol can flag thermodynamically metastable marketed forms before launch; the rotigotine case shows a form that later caused a recall being correctly ranked above the stable form.
  • When only powder diffraction data exist, the protocol can propose full three-dimensional crystal structures and rank their stability, as demonstrated for three recent drugs.
  • Lavo-NN's speed makes it suitable for the generation phase of CSP, where hundreds of millions of energy evaluations are needed, while its accuracy (54% Top-10 on the benchmark) means optional DFT re-ranking can be limited to single-point calculations.
  • As ranking costs drop, the bottleneck shifts to structure generation, so future gains for multi-component systems ($Z'>1$, hydrates, salts, co-crystals) will need new generative methods.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A prospective blind test on molecules guaranteed absent from the training corpus is the decisive check on generalization; the current retrospective benchmark, lacking a stated holdout split, cannot rule out memorization of some structures.
  • If the second-order many-body truncation fails to cancel between polymorphs of some molecule, rankings could reverse; testing on molecules with strong three-body dispersion would probe this.
  • The per-space-group convergence scheme transfers to other sampling problems where one wants to estimate missed low-energy configurations from repeat counts.
  • If the cost claims hold in outside use, routine early-stage CSP could change how polymorph risk is priced in drug development, reducing late-appearing-form surprises like the one that forced reformulation of ritonavir.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The manuscript introduces Lavo-NN, an equivariant neural network potential designed for pharmaceutical crystal structure prediction, and packages it into a cloud-based, near-automatic CSP workflow. The protocol is validated on a benchmark of 49 drug-like molecules, with claimed success in generating and ranking all 110 reported Z' = 1 experimental polymorphs at roughly 8.4k CPU hours per molecule, roughly 438k CPU hours in total. Additional case studies address rotigotine, mebendazole, fenofibrate, and progesterone, and a semi-blinded PXRD study assigns structures to polymorphs of three marketed drugs. The paper also compares Lavo-NN with other NNPs on a Top-10 ranking metric, reporting 54% Top-10 accuracy.

Significance. If the central claims are ultimately supported, this would be a substantial advance for pharmaceutical CSP: the benchmark is the largest of its kind, the reported cost reduction is striking, and the PXRD-to-structure case studies illustrate a practical route to solving powder-only forms. The paper's explicit architectural choices (intramolecular-only message passing, a polarizable force-field readout, and a truncated many-body expansion training target) are well motivated and clearly described. The inclusion of a multi-model comparison on a common ranking task is a useful contribution to the community. However, the validity of the headline claims currently depends on the resolution of several load-bearing issues: the possibility of training/benchmark overlap, the post-hoc exclusion of a known polymorph, an unexplained discrepancy in the total polymorph count, and an unvalidated convergence heuristic.

major comments (5)
  1. [§2.4 vs §5.1] The paper does not establish that the 49 benchmark molecules were held out from Lavo-NN's training data. Section 2.4 describes iterative training on roughly 10k SMILES strings sampled from the SPICE dataset, plus crystal structures, conformers, and dimers generated from those SMILES, with no statement that the benchmark molecules (or their close analogs, tautomers, or protonation states) were excluded. Because the benchmark consists largely of marketed drugs and SPICE is a drug-like dataset, overlap is plausible. The absence of a documented holdout means the claimed generalization to new pharmaceutical molecules is not supported by the presented evidence. The authors should either report an overlap analysis between the training SMILES and the benchmark molecules, or rerun the benchmark on a subset that is provably unseen during training; without this, the central 'successful generation of all experimental polymorphs' result may reflect memorization rather than generalization.
  2. [§5.1 vs Abstract] The 'all 110' claim is internally inconsistent. The abstract and Section 5 state that the benchmark contains 110 experimental Z' = 1 polymorphs across 49 molecules. Section 5.1 then reports 104 CSD-sourced polymorphs for 46 molecules, and Section 5.2 adds seven PXRD-only forms (omaveloxolone: 2, deucravacitinib: 2, zuranolone: 3), which sums to 111. Moreover, Section 5.1 says both that 'All 104 polymorphs are both generated and ranked' and that 'One polymorph, galunisertib Form I, is excluded from analysis.' The reader cannot determine the exact number of forms attempted and successfully generated. The authors must reconcile these numbers and present a single, unambiguous success count that accounts for all inclusions and exclusions.
  3. [§3.4] The completeness claim — that the protocol generates all low-energy polymorphs — rests on a convergence estimator that is not validated. Section 3.4 models the counts of unique crystal structures with a Poisson-lognormal distribution and uses the estimated number of unseen structures as the stopping criterion. No evidence is given that this species-abundance analogy is reliable for CSP landscapes, nor is the estimator tested against a known answer (e.g., by subsampling a fully enumerated landscape or by comparing predictions for molecules with exhaustive prior CSP studies). Because the headline 'all 110 polymorphs generated' requires that generation is truly complete, the authors should provide a validation of the convergence metric, for example by showing that the estimator correctly predicts the number of held-out known structures in a subsampling experiment, or by benchmarking against an independent exhaustive search for a small test molecule.
  4. [§5.1, §5.2, §3] The 'fully automated' claim is qualified by several manual interventions in the benchmark. Specifically, mebendazole required the user to identify three tautomers and run separate CSPs (§5.1.2), progesterone required separate runs in Sohncke and non-Sohncke space groups to cover enantiopure and racemic scenarios (§5.1.4), and the ripretinib landscape required manual addition of a Z' = 2 structure that is outside the stated Z' = 1 scope (Figure 6 caption). These cases contradict the protocol description in Section 3, which states that the only inputs are a SMILES string and a chirality choice, and they undermine the claim of minimal manual input. The authors should either quantify how many of the 49 molecules needed such expert intervention beyond the stated inputs, or explicitly limit the 'fully automated' claim to the subset that required no additional specification.
  5. [§5.1 vs §5.3] The strong ranking results reported in Section 5.1 are produced by the full protocol, which includes periodic PBE-D3(BJ) and a monomer PBE0 correction (Section 3.5), not by Lavo-NN alone. Section 5.3 shows that Lavo-NN's own Top-10 accuracy is 54%, whereas the benchmark statement 'All 104 polymorphs are both generated and ranked near the bottom' describes the DFT-requalified landscape. The abstract and conclusions should state clearly that the NNP is responsible for structure generation and preliminary ranking, and that the final polymorph rankings in the headline benchmark reflect the additional DFT re-ranking step. As written, the reader could reasonably attribute the ranking success to the NNP, which is not what the data show.
minor comments (6)
  1. [§2.4] The basis set 'def2-TZPPD' appears to be a typo; the standard basis set is def2-TZVPPD. Please confirm the correct basis set name.
  2. [Table 1] Compound names are inconsistently capitalized in the table: 'GSk-269984B', 'Mk-2022', 'Mk-8876', and 'TIk-301' should be 'GSK-269984B', 'MK-2022', 'MK-8876', and 'TIK-301'.
  3. [Figure 1 caption] The caption contains a typo: 'T op' should be 'Top'.
  4. [§5.1] The sentence 'a dramatic reduction in the typical the computational cost' contains a duplicated article and should be corrected.
  5. [§5.3] The Top-10 accuracy is evaluated on the landscapes generated by the protocol, which may themselves be biased by Lavo-NN used during generation. The metric is described as a ranking test, but the authors should acknowledge that generation and ranking are not fully decoupled in this evaluation.
  6. [Appendix A.1] The data availability statement says the landscapes will be made available 'upon publication.' For a benchmark paper, it would strengthen reproducibility if the CIFs and energy files were accessible to reviewers or available as supporting information at submission.

Circularity Check

0 steps flagged · score 0.0 of 10

No demonstrated circularity: the benchmark is external and the central claims do not reduce to training data, fitted parameters, or self-citations by construction.

full rationale

The paper's central derivation chain is self-contained with respect to the benchmark: Lavo-NN is trained on DFT labels for monomers and dimers, the CSP protocol generates and optimizes crystal structures with Lavo-NN and optionally re-ranks them with periodic PBE-D3(BJ) plus a monomer PBE0 correction, and the benchmark outcomes are comparisons against experimental CSD structures and experimental PXRD patterns. None of the stated equations (Eqs. 1-5) define the benchmark answer in terms of the training labels, and no fitted parameter is renamed as a prediction. The self-citations to AP-Net and Splinter are prior-architecture and training-data references, not load-bearing uniqueness theorems or ansatz-smuggling citations. The absence of an explicit holdout split between the 10k SPICE-derived training molecules and the 49 benchmark molecules is a legitimate generalization-risk concern, but the paper text does not establish that any benchmark molecule was actually in the training set, so this cannot be exhibited as a specific circular reduction. Similarly, the exclusions and manual additions (galunisertib Form I, ripretinib, mebendazole, progesterone) qualify the strength of the headline claims but do not make the derivation circular. The core validation is against external experimental structures, so the appropriate finding is no significant circularity.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

Lavo-NN is a computational model, not a new physical entity. The free parameters listed are the ones the central claims depend on. The axioms capture the unproved modeling choices that support the protocol's ranking and completeness claims.

free parameters (4)
  • Ai, Bi, Ci force-field elemental parameters = not specified
    Section 2.2: 'either borrowed from other force fields or fit to a small number of DFT interaction energy calculations.' These parameters enter the intermolecular force field and affect the NNP's energies.
  • Poisson-lognormal convergence parameters (mu, sigma) = fit per space group per molecule
    Section 3.4: 'The parameters of the distribution mu and sigma are continually fit to best reproduce the empirical frequencies of unique crystal counts.' Used to decide when to stop generating structures.
  • RMSD20 similarity cutoff = 0.8 angstrom
    Section 3.4: 'a heavy-atom-only similarity cutoff of 0.8 Å' determines which generated structures count as duplicates, thus affecting the convergence estimate.
  • NNP weights = learned from DFT datasets
    The neural network's parameters are fitted to DFT energies and gradients during training (Section 2.4). These are the model's core fitted parameters.
assumptions (5)
  • domain assumption Experimental crystal structures correspond to minima of the lattice energy landscape.
    Section 1.1: 'experimentally observed crystal structures correspond to minima on the free energy (or to an approximation, lattice energy) landscape of crystal structures.' This is the foundational assumption of CSP.
  • domain assumption Second-order many-body expansion is sufficient for polymorph ranking.
    Section 2.3: 'Lavo-NN is trained to the MBE, truncated at second order' and 'higher order MBE terms can be non-zero, but they usually cancel between polymorphs.' This is an unproved approximation that the protocol's ranking depends on.
  • ad hoc to paper The Poisson-lognormal distribution models the distribution of undiscovered crystal structures.
    Section 3.4: 'We found that the Poisson-lognormal distribution works well for modeling counts of unique crystals.' The convergence criterion relies on this empirical fit with no derivation.
  • domain assumption PBE-D3(BJ) with monomer PBE0 correction is accurate enough for final polymorph ranking.
    Section 3.5: 'a pragmatic level of theory avoids the cost of hybrid periodic DFT while still eliminating much of the intramolecular delocalization error.' The protocol's final rankings inherit this assumption.
  • domain assumption The benchmark molecules are representative of pharmaceutical CSP targets and are not in the training set.
    Implicit in the validation (Section 5.1). The paper does not demonstrate a holdout split from the training set described in Section 2.4.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Toward Routine CSP of Pharmaceuticals: A Fully Automated Protocol Using Neural Network Potentials." pith.science (2026). https://pith.science/paper/Q2ZO2OWW

@misc{pith2026250716218,
  author       = {Pith},
  title        = {Pith review of: Toward Routine CSP of Pharmaceuticals: A Fully Automated Protocol Using Neural Network Potentials},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q2ZO2OWW}},
  note         = {Machine review of arXiv:2507.16218}
}
abstract

Crystal structure prediction (CSP) is a useful tool in pharmaceutical development for identifying and assessing risks associated with polymorphism, yet widespread adoption has been hindered by high computational costs and the need for both manual specification and expert knowledge to achieve useful results. Here, we introduce a fully automated, high-throughput CSP protocol designed to overcome these barriers. The protocol's efficiency is driven by Lavo-NN, a novel neural network potential (NNP) architected and trained specifically for pharmaceutical crystal structure generation and ranking. This NNP-driven crystal generation phase is integrated into a scalable cloud-based workflow. We validate this CSP protocol on an extensive retrospective benchmark of 49 unique molecules, almost all of which are drug-like, successfully generating structures that match all 110 $Z' = 1$ experimental polymorphs. The average CSP in this benchmark is performed with approximately 8.4k CPU hours, which is a significant reduction compared to other protocols. The practical utility of the protocol is further demonstrated through case studies that resolve ambiguities in experimental data and a semi-blinded challenge that successfully identifies and ranks polymorphs of three modern drugs from powder X-ray diffraction patterns alone. By significantly reducing the required time and cost, the protocol enables CSP to be routinely deployed earlier in the drug discovery pipeline, such as during lead optimization. Rapid turnaround times and high throughput also enable CSP that can be run in parallel with experimental screening, providing chemists with real-time insights to guide their work in the lab.

Figures

Figures reproduced from arXiv: 2507.16218 by the authors.

Figure 1
Figure 1. Relationship between the success rate (number of successfully predicted experimental [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Overview of the Lavo-NN architecture. Center: high-level representation of the archi￾tecture, a NNP designed to efficiently and accurately predict energies and gradients of molecular crystals. Left: The intramolecular module predicts the gas-phase energy of the molecule in the crystal as well as atom-in-molecule properties such as atomic charges and dipoles. Right: The in￾termolecular module predicts interaction ene… view at source ↗
Figure 3
Figure 3. Lavo-NN calculates a crystal lattice energy as a combination of an intramolecular energy [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Visual overview of the software architecture. The coordinator process initiates jobs and [PITH_FULL_IMAGE:figures/full_fig_p016_4.png]
Figure 5
Figure 5. Figure 5: Benchmark molecules used to assess the CSP protocol. [PITH_FULL_IMAGE:figures/full_fig_p019_5.png]
Figure 6
Figure 6. Figure 6: Rankings (A) and relative energies (B) of all experimental crystals found in the study. [PITH_FULL_IMAGE:figures/full_fig_p022_6.png]
Figure 7
Figure 7. Figure 7: Density energy plot for Rotigotine. Form I is the originally marketed form, but was later [PITH_FULL_IMAGE:figures/full_fig_p024_7.png]
Figure 8
Figure 8. Figure 8: Density energy plot of the three polymorphs of mebendazole. The variant of Form B [PITH_FULL_IMAGE:figures/full_fig_p025_8.png]
Figure 9
Figure 9. Figure 9: A qualitative summary of experimental forms of fenofibrate and their relative rankings, [PITH_FULL_IMAGE:figures/full_fig_p026_9.png]
Figure 10
Figure 10. Figure 10: Summary of semi-blinded powder X-ray diffraction study. Rows correspond to each [PITH_FULL_IMAGE:figures/full_fig_p028_10.png]
Figure 11
Figure 11. Figure 11: Cost-accuracy comparison of various NNPs (blue circles), semi-empirical methods [PITH_FULL_IMAGE:figures/full_fig_p031_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

11 extracted references · 6 canonical work pages

  1. [1]

    International tables for crystallography: Space-group symmetry

    (2016). International tables for crystallography: Space-group symmetry. http://dx.doi.org/10.1107/97809553602060000114 Adamo, C. & Barone, V. (1999). The Journal of Chemical Physics , 110(13), 6158–6170. http://dx.doi.org/10.1063/1.478522 Agarwal, P., Huckle, J., Newman, J. & Reid, D. L. (2022). Drug Discovery Today, 27(12), 103366. https://www.sciencedir...

  2. [2]

    P., Burke, K

    Perdew, J. P., Burke, K. & Ernzerhof, M. (1996). Phys. Rev. Lett. 77, 3865–3868. https://link.aps.org/doi/10.1103/PhysRevLett.77.3865 Perdew, J. P. & Schmidt, K. (2001). AIP Conference Proceedings, 577(1), 1–20. https://doi.org/10.1063/1.1390175 Perry, C., Ramos, S., Phelps, M., Mueller, L. & Beran, G. (2025). ChemRxiv. https://doi.org/10.26434/chemrxiv-2...

  3. [579]

    & Grimme, S

    M¨ uller, M., Hansen, A. & Grimme, S. (2023).Journal of Chemical Physics , 158(1), 014103. Nelson, P. M. & Sherrill, C. D. (2024). The Journal of Chemical Physics , 161(21). Neumann, M. A. (2008). The Journal of Physical Chemistry B , 112(32), 9810–9829. http://dx.doi.org/10.1021/jp710575h Newman, J. A., Iuzzolino, L., Tan, M., Orth, P., Bruhn, J. & Lee, ...

  4. [1986]

    Hellweg, A

    Siam. Hellweg, A. & Rappoport, D. (2015). Physical Chemistry Chemical Physics , 17, 1010–1017. Heo, Y.-A. (2023). Drugs, 83(16), 1559–1567. Herman, K. M. & Xantheas, S. S. (2023). Journal of Physical Chemistry Letters , 14(4), 989–999. Hilfiker, R. e. (2006). Polymorphism in the Pharmaceutical Industry . Weinheim: Wiley-VCH. Chaps. 2–4 survey solvent/temp...

  5. [2018]

    & Tripp, J

    https://patents.google.com/patent/WO2018018618A1/en Sidhu, G. & Tripp, J. (2023). Fenofibrate. Treasure Island (FL): StatPearls Publishing. Last update March 13,

  6. [2023]

    C., Pickard, F

    https://www.ncbi.nlm.nih.gov/books/NBK559219/ Simmonett, A. C., Pickard, F. C., Shao, Y., Cheatham, T. E. & Brooks, B. R. (2015). The Journal of chemical physics , 143(7). Smith, D. G. A., Burns, L. A., Simmonett, A. C., Parrish, R. M., Schieber, M. C., Galvelis, R., Kraus, P., Kruse, H., Di Remigio, R. et al. (2020). Journal of Chemical Physics , 152(18)...

  7. [2025]

    & Dideberg, O

    https://www.ccdc.cam.ac.uk/structures Campsteyn, H., Dupont, L. & Dideberg, O. (1972). Acta Crystallographica Section B , 28, 3032–

  8. [2991]

    & Toennies, J

    38 Tang, K. & Toennies, J. P. (1986). Zeitschrift f¨ ur Physik D Atoms, Molecules and Clusters, 1(1), 91–101. Taylor, C. R., Butler, P. W. & Day, G. M. (2025 a). Faraday Discussions, 256, 434–458. Taylor, C. R., Butler, P. W. V. & Day, G. M. (2025 b). Faraday Discussions, 256, 434–458. Thirunahari, S., Aitipamula, S., Chow, P. S. & Tan, R. B. (2010). Jour...

Show all 11 references
  1. [3042]

    C., Kuang, J., Wang, L., Zhang, C., Carbone, M

    https://doi.org/10.1107/S0567740872007393 Cao, C., Kingan, A., Hill, R. C., Kuang, J., Wang, L., Zhang, C., Carbone, M. R., van Dam, H., Yoo, S., Marschilok, A. C. & Lu, D. (2025). PRX Energy, 4, 023004. Catlow, C. R. A. (2023). IUCrJ, 10(2), 143–144. http://dx.doi.org/10.1107...

  2. [5148]

    D., Markland, T

    Saban´ es Zariquiey, F., Galvelis, R., Gallicchio, E., Chodera, J. D., Markland, T. E. & De Fabritiis, G. (2024). Journal of Chemical Information and Modeling , 64(5), 1481–1485. Sargent, C. T., Metcalf, D. P., Glick, Z. L., Borca, C. H. & Sherrill, C. D. (2023). The Journal o...

  3. [6615]

    & L´ opez, N

    Lian, Z., Dattila, F. & L´ opez, N. (2024).Nature Catalysis, 7, 401–411. Liang, Y. H., Ye, H.-Z. & Berkelbach, T. C. (2023). The Journal of Physical Chemistry Letters , 14(46), 10435–10441. Lin, L., Pan, L. & Liu, S. (2022). Computers in Industry , 141, 103718. https://www.sci...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.