Pith. sign in

REVIEW 3 major objections 7 minor 24 references

Survey on Recent Progress of AI for Chemistry: Methods, Applications, and Opportunities

T0 review · 3 major / 7 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read A computational map of AI for chemistry, organized by data, representation, and model.

desk verdict A decent orientation map for newcomers, but 'comprehensive' overstates it: the survey skips machine-learned force fields, a major active area, and has several accuracy slips in its tables. read the letter →

arxiv 2502.17456 v1 pith:G6IHY3RI submitted 2025-02-09 physics.chem-ph cs.AIcs.LG

classification physics.chem-phcs.AIcs.LG
keywords AIforchemistrymachinelearningmolecularrepresentationgraphneuralnetworksdesignretrosynthesislargelanguagemodelsself-drivinglaboratories
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey aims to give computer scientists a reliable map of the rapidly growing field of AI for chemistry. Its central claim is that the field can be understood through three components—data, representation, and model—and that nearly all current work, from property prediction to retrosynthesis to laboratory automation, fits into that frame. The paper catalogues the available datasets and the main ways molecules are encoded, reviews the model families used for prediction, molecular design, and reaction planning, and closes with four challenges it says must be tackled before these methods become routine. A sympathetic reader would take this as a starting point: a structured overview that tells you what the field looks like and where the open problems are.

What carries the argument

The machinery that carries the survey is the data–representation–model triad, used as a classification scheme for the whole field. Within that scheme, the load-bearing mechanisms are message passing in graph neural networks, the attention mechanism in SMILES-based Transformers, and the task-specific engines of molecular design—generative models, genetic algorithms, and reinforcement learning—plus Monte Carlo tree search and template-based reasoning for retrosynthesis. The triad lets the paper place every method under one of three design questions: what data feeds it, how molecules are encoded, and which model family learns from that encoding.

What would settle it

Run a systematic literature search on AI-for-chemistry with explicit inclusion criteria and compare the resulting set of methods and datasets with those covered here. If significant methodological families (for example equivariant neural network potentials or physics-informed models for molecular dynamics) or widely used benchmarks turn out to be absent without explanation, the survey's claim of comprehensive coverage would be shown to be too strong.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that the state of AI for chemistry is best organised by a data–representation–model pipeline. It shows that datasets split naturally into molecular-level (QM9, MD17, Tox21, BBBP, HIV, MoleculeNet) and reaction-level (USPTO, Reaxys, HTE) records; that representations range from fingerprints and identifiers such as SMILES and InChI to learned embeddings from graph neural networks and Transformers; and that the major application areas are property prediction, de novo molecular design, retrosynthesis, and self-driving laboratories. The review also finds that large language models are entering chemistry as agents, as fine-tuned downstream encoders, and as objects of evaluation, but that they currently fail at precise SMILES-to-IUPAC conversion and can hallucinate chemically unreasonable molecules. It concludes that data scarcity, data bias, interpretability, and the mismatch between general-purpose generative models and chemical precision are the four obstacles that warrant further attention.

Load-bearing premise

The claim of comprehensiveness rests on the unstated assumption that the papers chosen for discussion fairly represent the field and that the data–representation–model taxonomy captures the full research landscape, since no systematic search strategy or inclusion criteria are given.

Editorial extensions

If this is right

  • Newcomers can treat the data–representation–model triad as a working map for locating any AI-for-chemistry method and understanding the design choices behind it.
  • The four listed challenges—data scarcity, data bias, interpretability, and domain-specific generative models—define a concrete research agenda for the field.
  • For large language models, the reviewed evidence implies that they are most useful when paired with external chemical tools, not when asked to reason about molecular strings directly.
  • Template-free retrosynthesis is converging on synthon-based two-stage and string-editing architectures, which should keep improving the synthesizability of AI-designed molecules.
  • Self-driving laboratories that close the loop between machine learning and robotics have already cut discovery times for specific materials, making data generation itself a machine-learning problem.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the taxonomy is as complete as the survey claims, a natural extension is a unified molecular foundation model that fuses the GNN and Transformer strands, since the review repeatedly notes that graphs and sequences each capture complementary information.
  • The emphasis on omitted failed reactions suggests a concrete test: including negative data in standard yield-prediction benchmarks would probably shift model rankings, a hypothesis the survey does not itself test.
  • Because the review does not state how papers were selected, the reliability of the map depends on the authors' coverage; a versioned, community-maintained update of the taxonomy would be a direct way to test and refresh that map.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. This manuscript is a survey of recent AI techniques for chemistry, organized by the three components of machine learning: data, representation, and model. It reviews common chemical datasets (molecular and reaction level), representation learning with GNNs and Transformers, prediction tasks (properties, reaction products, yields), molecular design (generative, genetic/RL, combinatorial optimization), retrosynthesis (template-based, template-free, multi-step), self-driving laboratories, and LLM applications in chemistry. It closes with a discussion of challenges such as data scarcity, bias, interpretability, and the limitations of general-purpose generative models. The paper is positioned as a comprehensive review from a computational perspective.

Significance. If the comprehensiveness claim were substantiated, the survey would serve as a useful entry point for computer scientists entering AI for chemistry, and its organization around data, representation, and models is pedagogically attractive. The paper's explicit coverage of recent LLM-based agents and self-driving laboratories is a strength, as is the inclusion of several recent high-profile works (e.g., ChemCrow, Coscientist, A-Lab). However, the claimed comprehensiveness is undermined by the omission of an entire major methodological family—machine-learned interatomic potentials / equivariant force fields—that the paper itself flags as important in §2.2.3. Because the selection of topics is not justified by any stated inclusion criteria, the paper as it stands is better described as a selective overview than a comprehensive review. The technical descriptions that are included are mostly accurate, and the paper does not make internally inconsistent claims; the main weakness is a mismatch between the abstract's promise and the delivered coverage.

major comments (3)
  1. [§2.2.3 and §4] The paper explicitly identifies ML-accelerated molecular dynamics as an important direction and cites the machine-learned force-field work of Chmiela et al. [2017, 2018, 2020], yet no section of §4 surveys machine-learned interatomic potentials or equivariant force-field models (e.g., SchNet, ANI, NequIP, PaiNN, MACE). The application coverage in §4 is limited to property prediction, molecular design, retrosynthesis, and self-driving labs. Since the abstract claims a 'comprehensive review of current AI techniques in chemistry from a computational perspective,' this omission is load-bearing. A reader using this paper as a map would miss an active and central research program. I recommend either adding a substantive subsection on ML force fields and their benchmarks, or revising the abstract and conclusion to describe the review as selective rather than comprehensive.
  2. [Introduction / Abstract] The 'comprehensive' claim is not accompanied by any methodology statement: the paper gives no search strategy, inclusion criteria, time window, or coverage statistics. Without such criteria, a reader cannot verify that the selected works are representative or that the taxonomy in Section 1 fully captures the research landscape. Even if the force-field omission were repaired, the absence of an explicit scope statement makes the comprehensiveness claim unfalsifiable. Please either add a short methodology paragraph describing the selection process and any intentional exclusions, or soften the claim to 'a broad overview' or 'a selective survey.'
  3. [Table 3] The row for Segler et al. [2018a] describes the work as 'The first sequence-to-sequence model for molecular generation,' but the cited paper (Segler et al., ACS Central Science 2018) uses a recurrent neural network language model to generate SMILES strings; it is not a sequence-to-sequence (encoder-decoder) model. The 'Methodology' column correctly says 'RNN,' so the highlight text is internally inconsistent and mischaracterizes a widely cited contribution. This should be corrected to something like 'The first RNN-based generative model for de novo SMILES generation.'
minor comments (7)
  1. [§2.2.2] 'Quantum-machine qua is an open project' appears to be a typo for 'Quantum Machine' (the quantum-machine.org platform), and 'a weath of' should be 'a wealth of.' Please fix these errors.
  2. [§2.1.2] The surname 'Wpolber et al.' should be 'Wolber et al.' (corresponding to reference Wolber et al. [2008]).
  3. [§2.1.1] 'SMILES Arbitrary Target Specification (SMART)Systems' should be 'SMARTS' (SMILES Arbitrary Target Specification), not 'SMART.'
  4. [§3.3] The word 'imapcts' should be 'impacts' in the first sentence.
  5. [§4.3.1] 'formly known as' should be 'formerly known as' in the description of Synthia.
  6. [References] Two entries in the reference list appear orphaned: an unattributed 2023 Computers in Biology and Medicine article on BiGRU and GraphSAGE toxicity prediction, and a URL-only entry for quantum-machine.org. Both should either be cited in the text or removed from the reference list.
  7. [§2.2.4] The BBBP dataset is attributed to Sakiyama et al. [2021], but the original BBBP benchmark is commonly traced to Martins et al. (2012); please verify that the intended citation is the one describing the dataset rather than a subsequent prediction study.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey summarizes external literature and contains no self-referential derivation or fitted prediction.

full rationale

This is a review paper with no equations, no fitted parameters, and no derived predictions. Its content is an organized summary of external, published work on datasets, representations, models, applications, LLMs, and challenges in AI for chemistry. The paper introduces no self-referential argument whose conclusion is presupposed by its inputs, and no self-citations by the present authors were identified in the reference list. The closest issue is the gap between the abstract's 'comprehensive review' claim and the omission of a full treatment of machine-learned interatomic potentials even though §2.2.3 flags ML-accelerated molecular dynamics as important; that is a coverage/completeness concern for the survey's accuracy, not a circularity concern, because the statements made about the works that are cited remain externally checkable and do not reduce to the paper's own assumptions. The circularity burden is therefore zero.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

As a review, the paper introduces no free parameters or invented entities. It relies on two domain assumptions: that the chosen organizing taxonomy is sufficient and that the cited references are representative and correctly described.

assumptions (2)
  • domain assumption The three-component framing of machine learning (data, representation, model) is a sufficient organizing principle for AI-for-chemistry research.
    Section 1 defines the paper's scope through these three components; the entire survey structure depends on this framing.
  • domain assumption The selected references are representative and accurately described.
    The paper claims comprehensiveness, but provides no systematic search or inclusion criteria; correctness of descriptions of third-party results is assumed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Survey on Recent Progress of AI for Chemistry: Methods, Applications, and Opportunities." pith.science (2026). https://pith.science/paper/G6IHY3RI

@misc{pith2026250217456,
  author       = {Pith},
  title        = {Pith review of: Survey on Recent Progress of AI for Chemistry: Methods, Applications, and Opportunities},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/G6IHY3RI}},
  note         = {Machine review of arXiv:2502.17456}
}
read the original abstract

The development of artificial intelligence (AI) techniques has brought revolutionary changes across various realms. In particular, the use of AI-assisted methods to accelerate chemical research has become a popular and rapidly growing trend, leading to numerous groundbreaking works. In this paper, we provide a comprehensive review of current AI techniques in chemistry from a computational perspective, considering various aspects in the design of methods. We begin by discussing the characteristics of data from diverse sources, followed by an overview of various representation methods. Next, we review existing models for several topical tasks in the field, and conclude by highlighting some key challenges that warrant further attention.

Figures

Figures reproduced from arXiv: 2502.17456 by the authors.

Figure 1
Figure 1. Three components in machine learning: data, representation and model. To provide a thorough introduction to the advancement of AI for Chemistry from the lens of computer science, it is essential to first clarify the paradigm of AI research, particularly in machine learning. The goal of ML is to extract knowledge from data to assist machines in performing certain tasks. Generally, ML involves three important componen… view at source ↗
Figure 2
Figure 2. Different types of chemical datasets. 2 Datasets and Descriptions Machine learning is a data-driven discipline, and its rapid progress has been largely propelled by the increasing availability of data Zhou et al. [2017]. In the field of chemistry, addressing various data challenges is particularly demanding Strieth-Kalthoff et al. [2022]. High-quality, sufficiently large datasets are essential for machine learning t… view at source ↗
Figure 3
Figure 3. An example of the mapping between SMILES and structure. The molecule in the figure is “melatonin”. Hydrogen atoms are omitted, and atoms in different colors correspond to the text in the SMILES notation, representing the main chain and side chains, respectively. The orange bonds highlight the edges of the main chain split by SMILES. unique label for a molecule but also encodes basic chemical information. There are s… view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: A common framework of molecular graph neural networks. Through message-passing mechanisms, GNNs progressively refine molecular representations, starting from the initial molecular embed￾dings to produce graph-level representations of the entire molecule. In this figure…
Figure 5
Figure 5. Figure 5: Simplified transformer model in chemistry. Based on SMILES, reactions and molecules can be represented in text form. The Transformer architecture can perform various chemical tasks such as representation learning, property prediction, and retrosynthesis analysis by pro…
Figure 6
Figure 6. Figure 6: Overview of the prediction problems in chemistry aided by AI. The overall framework begins with data, encompassing various molecular and reaction datasets (first column). Subsequently, different representation methods can be employed, such as molecular fingerprints, GN…
Figure 7
Figure 7. Figure 7: Overview of the molecular design methods. 13 [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: Overview of the retrosynthesis task. In single-step retrosynthesis, the task is to identify the reactants corresponding to a target product. In multi-step retrosynthesis, the goal is to construct a complete synthesis tree where all leaf nodes correspond to available ch…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

24 extracted references · 13 canonical work pages

  1. [1]

    Accessed: 2024-12-29

    Quantum machine learning platform.http://quantum-machine.org/. Accessed: 2024-12-29. The prediction of molecular toxicity based on bigru and graphsage.Computers in Biology and Medicine, 153: 106524,

  2. [4]

    Language models are few-shot learners.arXiv preprint arXiv:2005.14165,

    Tom B Brown. Language models are few-shot learners.arXiv preprint arXiv:2005.14165,

  3. [9]

    Connor W Coley, Luke Rogers, William H Green, and Klavs F Jensen

    doi: 10.1007/978-3-030-40245-7\_7. Connor W Coley, Luke Rogers, William H Green, and Klavs F Jensen. Computer-assisted retrosynthesis based on molecular similarity.ACS central science, 3(12):1237–1245,

  4. [11]

    Sourabh Katoch, Sumit Singh Chauhan, and Vijay Kumar

    ISSN 2041-1723. Sourabh Katoch, Sumit Singh Chauhan, and Vijay Kumar. A review on genetic algorithm: past, present, and future. Multimedia tools and applications, 80:8091–8126,

  5. [14]

    URLhttps://arxiv.org/abs/2404.01475. David L. Mobley and J. Peter Guthrie. FreeSolv: a database of experimental and calculated hydration free energies, with input files.Journal of Computer-Aided Molecular Design, 28(7):711–720, July

  6. [15]

    AkshatKumar Nigam, Pascal Friederich, Mario Krenn, and Alan Aspuru-Guzik

    arXiv:1504.04909 [cs]. AkshatKumar Nigam, Pascal Friederich, Mario Krenn, and Alan Aspuru-Guzik. Augmenting genetic algorithms with deep neural networks for exploring the chemical space. InInternational Conference on Learning Representations,

  7. [18]

    overview.The Journal of Organic Chemistry, 45(11):2043–2051,

  8. [19]

    Predicting retrosynthetic pathways using transformer-based models and a hyper-graph exploration strategy.Chemical science, 11(12):3316–3325, 2020a

    Philippe Schwaller, Riccardo Petraglia, Valerio Zullo, Vishnu H Nair, Rico Andreas Haeuselmann, Riccardo Pisoni, Costas Bekas, Anna Iuliano, and Teodoro Laino. Predicting retrosynthetic pathways using transformer-based models and a hyper-graph exploration strategy.Chemical science, 11(12):3316–3325, 2020a. Philippe Schwaller, Alain C Vaucher, Teodoro Lain...

Show all 24 references
  1. [21]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al

    ISSN 0009-2665, 1520-6890. Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971,

  2. [22]

    Graph attention networks.arXiv preprint arXiv:1710.10903,

    Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. Graph attention networks.arXiv preprint arXiv:1710.10903,

  3. [23]

    The chembl database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods.Nucleic Acids Research, 52(D1):D1180–D1192, 11

    Barbara Zdrazil, Eloy Felix, Fiona Hunter, Emma J Manners, James Blackshaw, Sybilla Corbett, Marleen de Veij, Harris Ioannidis, David Mendez Lopez, Juan F Mosquera, Maria Paula Magarinos, Nicolas Bosc, Ricardo Arcila, Tevfik Kizilören, Anna Gaulton, A Patrícia Bento, Melissa F...

  4. [25]

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al

    URLhttps://arxiv.org/abs/2402.06852. Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. A survey of large language models.arXiv preprint arXiv:2303.18223,

  5. [26]

    Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong, and Qing He

    ISSN 2095-5138. Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong, and Qing He. A comprehensive survey on transfer learning.Proceedings of the IEEE, 109(1):43–76,

  6. [2002]

    Reaxys: A leading chemistry database.www.reaxys.com

    Elsevier. Reaxys: A leading chemistry database.www.reaxys.com. Accessed: 2024-12-29. 24 David Feller. The role of databases in support of computational chemistry calculations.Journal of computa- tional chemistry, 17(13):1571–1586,

  7. [2015]

    Jaakkola, and Regina Barzilay

    Benson Chen, Tianxiao Shen, Tommi S. Jaakkola, and Regina Barzilay. Learning to make generalizable and diverse predictions for retrosynthesis.CoRR, abs/1910.09688,

  8. [2016]

    Pubchem 2025 update

    Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, Leonid Zaslavsky, Jian Zhang, and Evan E Bolton. Pubchem 2025 update. Nucleic Acids Research, 53(D1):D1516–D1525, 11

  9. [2017]

    Smiles arbitrary target specification.https://www.daylight.com/ dayhtml_tutorials/languages/smarts/index.html

    Daylight Chemical Information Systems. Smiles arbitrary target specification.https://www.daylight.com/ dayhtml_tutorials/languages/smarts/index.html. Accessed: 2025-01-13. Attila Szabo and Neil S Ostlund.Modern quantum chemistry: introduction to advanced electronic structure t...

  10. [2018]

    Stefan Chmiela, Huziel E

    doi: 10.1038/s41467-018-06169-2. Stefan Chmiela, Huziel E. Sauceda, Alexandre Tkatchenko, and Klaus-Robert Müller.Accurate molecular dynamics enabled by efficient physically-constrained machine learning approaches, pages 129–154. Springer International Publishing,

  11. [2019]

    URLhttp://arxiv.org/abs/1910. 09688. Binghong Chen, Chengtao Li, Hanjun Dai, and Le Song. Retro*: learning retrosynthetic planning with neural guided a* search. InInternational conference on machine learning, pages 1608–1616. PMLR, 2020a. Shuan Chen and Yousung Jung. Deep retr...

  12. [2020]

    ISSN 1476-4687. Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, Xiao wen Dong, Haodong Duan, Qi Fan, Zhaoye Fei, Yang Gao, Jiaye Ge, Chenya Gu, Yuzhe Gu, Tao Gui, Aijia Guo, Qipeng Guo, Conghui He, Yingfan Hu, Ting...

  13. [2021]

    IAM graph database repository for graph based pattern recognition and machine learning

    Kaspar Riesen and Horst Bunke. IAM graph database repository for graph based pattern recognition and machine learning. InStructural, Syntactic, and Statistical Pattern Recognition: Joint IAPR International Workshop, SSPR & SPR 2008, Orlando, USA, December 4-6,

  14. [2022]

    Distributionally robust optimization: A review

    Hamed Rahimian and Sanjay Mehrotra. Distributionally robust optimization: A review. ArXiv, abs/1908.05659,

  15. [2023]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al

    ISSN 0010-4825. Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report.arXiv preprint arXiv:2303.08774,

  16. [2024]

    Molgpt: molecular generation using a transformer-decoder model

    Viraj Bagal, Rishal Aggarwal, PK Vinod, and U Deva Priyakumar. Molgpt: molecular generation using a transformer-decoder model. Journal of Chemical Information and Modeling, 62(9):2064–2076,

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.