REVIEW 4 major objections 4 minor 3 cited by
A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools
T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This survey argues that foundation models and LLM agents are moving materials science from task-specific machine learning to general-purpose, transferable AI, and maps the field into six task areas.
desk verdict A broad, well-organized survey whose central claim of comprehensive multiscale coverage is undercut by an internal contradiction about Marcato FM; useful as a map, but needs a revision before it can be trusted as a reference. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the survey's six-category task taxonomy (T1–T6), together with its three-way model classification: unimodal foundation models, multimodal foundation models, and LLM agents. The taxonomy is what converts a scattered set of papers, datasets, and tools into a map: each model is assigned to the tasks it can perform, and each dataset or benchmark is tied to the tasks it serves. The Transformer architecture supplies the common engine underneath most of the surveyed models, and the taxonomy is what lets the authors make comparative statements about coverage, gaps, and future directions.
What would settle it
A reader could re-run a systematic search with documented queries and inclusion criteria over the same period and check whether significant foundation models, datasets, or tools are missing from the six categories, or whether any quoted success figure, such as GNoME's 2.2 million materials or A-Lab's 71 percent rate, is misattributed; a major omission or misattribution would refute the survey's comprehensiveness claim.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that foundation models are catalyzing a transformative shift in materials science by replacing narrow, task-specific machine-learning pipelines with general-purpose systems that transfer across modalities and tasks. The survey's contribution is a task-driven taxonomy of six application areas (data extraction, interpretation and Q&A; atomistic simulation; property prediction; materials structure, design and discovery; process planning, discovery and optimization; and multiscale modeling), a division of the models into unimodal, multimodal, and LLM-agent categories, and a curated inventory of datasets, benchmarks, tools, and autonomous experimental platforms. Read sympathetically, the paper is saying that the pieces for generalist AI-driven materials research now exist and are converging, while acknowledging that long-range interactions, data imbalance, interpretability, and experimental integration remain unsolved.
Load-bearing premise
The survey's map is only as reliable as its unstated choice of which papers and tools to include, since the authors say they selectively incorporated high-quality research from Google Scholar and relied on prior experience for tools, without reporting search queries, inclusion criteria, or a verification protocol.
Editorial extensions
If this is right
- If the taxonomy is accepted, researchers can locate a model, dataset, or tool for a given task by consulting the corresponding category, which lowers the barrier to adopting AI in materials labs.
- The documented early successes imply that universal machine-learned potentials and generative models can operate at scales that conventional DFT screening cannot, at least for crystalline inorganic materials.
- The survey's limitations discussion implies that progress on polymers, disordered solids, and multiscale modeling depends on new kinds of data rather than on new architectures alone.
- The appearance of benchmarks such as LLM4Mat-Bench and MACBENCH suggests that evaluating LLMs and multimodal models on materials tasks is becoming standard practice.
- If agentic systems like MatPilot mature, experimental workflows could shift toward closed-loop, human-in-the-loop discovery, with LLMs planning and robots executing.
Reading between the lines
- My inference: the same six-task taxonomy could serve as a checklist for evaluating any new 'materials foundation model' claim, since a model that cannot be placed in at least one category is probably not yet a foundation model for this field.
- I infer that the lack of reported search queries and inclusion criteria means the survey should be read as a curated snapshot rather than an exhaustive census, and that quantitative statements like GNoME's 2.2 million materials are inherited from the cited papers rather than independently verified here.
- I infer that the next test of the field will be whether multimodal, co-registered datasets appear that pair structure with synthesis history and spectra; the survey identifies their absence as a bottleneck, and their arrival would likely accelerate the agentic systems it describes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey reviews foundation models (FMs), large language model (LLM) agents, datasets, and computational tools in materials science. It proposes a six-task taxonomy (T1–T6: data extraction/Q&A, atomistic simulation, property prediction, structure design/discovery, process planning/optimization, and multiscale modeling) and organizes the model landscape into unimodal FMs, multimodal FMs, and LLM agents. Tables summarize representative models, datasets, and tools, and the text discusses early successes (e.g., GNoME, MatterSim, A-Lab), limitations, and future directions. The central claimed contribution is a comprehensive, task-driven map of the field that goes beyond prior surveys by including atomistic simulation and multiscale modeling.
Significance. If accurate, this survey would be a useful reference for researchers entering the area, and the task-driven taxonomy is a sensible organizing principle. The tables are dense and potentially valuable. The paper does not make empirical claims of its own, so its correctness rests entirely on faithful reporting of cited work and on the representativeness of its selection. That dependence makes the internal inconsistencies and factual errors identified below consequential: they directly affect the survey's reliability as a map of the field. The coverage of tools and datasets is broad, and the discussion of limitations is balanced, but the undocumented selection protocol and the T6 contradiction must be resolved before the 'comprehensive' claim can be accepted.
major comments (4)
- [§3.1.6, Table 1, §3.2] Section 3.1.6 states, 'there are no existing foundation model architectures specifically designed for multiscale modeling problems in materials science.' This directly contradicts Table 1, which lists Marcato FM as a multimodal foundation model with task T6, and Section 3.2, which describes Marcato FM as a cross-modal encoder-decoder model for material failure prediction combining simulation outputs, structure, and grid-based fields. Since the survey uses the inclusion of multiscale modeling (T6) as one of its differentiators from prior work in Section 1, this is a load-bearing contradiction: either Marcato FM qualifies as a T6 foundation model and the Section 3.1.6 statement is false, or it does not and the table and Section 3.2 misclassify it. The authors must reconcile these statements, ideally by defining more precisely what counts as a 'foundation model' for T6 and then applying that definition consistently.
- [Abstract and §5.1] The abstract and Section 5.1 state that GNoME 'discovered over 2.2 million new stable materials' (or '2.2M new stable inorganic materials'). The cited source (Merchant et al., Nature 624, 2023) reports 2.2 million new crystals, of which roughly 380,000 are predicted to be stable. The survey's phrasing inflates the number of stable materials by about a factor of six, overstating one of the field's headline successes. This should be corrected in both the abstract and Section 5.1 to match the cited numbers.
- [§1] The paper claims to provide 'a comprehensive overview' and to have conducted 'a systematic literature search using Google Scholar,' but it does not report the search queries, date range, inclusion/exclusion criteria, or the number of papers screened versus included. The statement that the authors 'selectively incorporated high-quality research published in top-tier venues, as well as influential preprints' and chose tools 'based on our prior experience' is not a reproducible protocol. Given that the survey's value depends on the representativeness of its selection, the authors should add a methodology subsection (or appendix) specifying search strings, screening steps, and a clear selection policy so readers can assess coverage and potential omissions.
- [Table 1 and §3.1] Table 1, labeled 'Representative unimodal, multimodal, and agent-based foundation models,' includes models that do not fit the paper's own definition of foundation models as 'large, pretrained models trained on broad, diverse datasets' (Section 1 and Section 3). For example, CGCNN, MEGNet, CDVAE, and DiffCSP are task-specific architectures trained for particular prediction or generation tasks, not large pretrained general-purpose models. Including them under the 'Unimodal Foundation Models' heading conflates general-purpose pretrained models with conventional deep learning models, which misrepresents the maturity of the foundation-model landscape. The authors should either restrict the table to genuine foundation models or add a clear distinction (e.g., a separate category for 'task-specific deep learning models') so the taxonomy is not misleading.
minor comments (4)
- [§3 (first paragraph)] There is a duplicated word in the sentence 'The input is tokenized from the model's vocabulary and and each token is represented as a vector.'
- [Figure 3] In the timeline figure, 'ALIGN-FF' should be 'ALIGNN-FF' and 'FORCE' should be 'FORGE'; also 'Nach0' appears with inconsistent capitalization relative to the text's 'nach0'.
- [Table 2] The OQMD row lists 'T6' (multiscale modeling) among its relevant tasks, but OQMD is an atomistic DFT database; T6 (multiscale) appears to be an error. The same may apply to the Materials Project row, where T6 is listed without evident support.
- [§4.2.2] The text mentions 'Geom3D' and 'JARVIS-Leaderboard' as platforms but provides no citations for them. Please add appropriate references or remove these mentions if they are not accessible.
Circularity Check
No circularity found: the survey makes no derived predictions and its taxonomy is an organizational scheme, not a fitted or self-referential result.
full rationale
This manuscript is a literature survey rather than a derivation or empirical study. It introduces a six-category task taxonomy (T1–T6) as an organizing framework and summarizes models, datasets, and tools from external work. There is no fitted parameter, no predicted quantity that is statistically forced by an input subset, and no derivation chain in which an output is equivalent to an input by construction. The paper's claims about model performance, such as GNoME's 2.2M materials and A-Lab's 71% success rate, are reported from cited external sources, not produced by the survey itself. The authors do not invoke a load-bearing uniqueness theorem from their own prior work, and the taxonomy is not defined circularly in terms of the models it categorizes. The only notable inconsistency is that Section 3.1.6 states there are no existing foundation model architectures specifically designed for multiscale modeling, while Table 1 and Section 3.2 list Marcato FM with task T6; however, this is an internal consistency and correctness issue, not circularity, because neither statement is derived from the other or from the survey's own fitted inputs. Accordingly, the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Six-task taxonomy (T1-T6) is a meaningful, comprehensive partition of AI-for-materials work.
- domain assumption Selected references and quantitative success claims are accurately reported from primary sources.
- ad hoc to paper The 'systematic' Google Scholar search, restricted to top-tier venues and influential preprints, reflects the field's important work.
Cite this review
Pith. "Pith review of A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools." pith.science (2026). https://pith.science/paper/2IQCX44V
@misc{pith2026250620743,
author = {Pith},
title = {Pith review of: A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools},
year = {2026},
howpublished = {\url{https://pith.science/paper/2IQCX44V}},
note = {Machine review of arXiv:2506.20743}
}
read the original abstract
Foundation models (FMs) are catalyzing a transformative shift in materials science (MatSci) by enabling scalable, general-purpose, and multimodal AI systems for scientific discovery. Unlike traditional machine learning models, which are typically narrow in scope and require task-specific engineering, FMs offer cross-domain generalization and exhibit emergent capabilities. Their versatility is especially well-suited to materials science, where research challenges span diverse data types and scales. This survey provides a comprehensive overview of foundation models, agentic systems, datasets, and computational tools supporting this growing field. We introduce a task-driven taxonomy encompassing six broad application areas: data extraction, interpretation and Q\&A; atomistic simulation; property prediction; materials structure, design and discovery; process planning, discovery, and optimization; and multiscale modeling. We discuss recent advances in both unimodal and multimodal FMs, as well as emerging large language model (LLM) agents. Furthermore, we review standardized datasets, open-source tools, and autonomous experimental platforms that collectively fuel the development and integration of FMs into research workflows. We assess the early successes of foundation models and identify persistent limitations, including challenges in generalizability, interpretability, data imbalance, safety concerns, and limited multimodal fusion. Finally, we articulate future research directions centered on scalable pretraining, continual learning, data governance, and trustworthiness.
Figures
Forward citations
Cited by 3 Pith papers
-
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
SciAgent-8B, fine-tuned on trajectories synthesized from a tool dependency graph, outperforms Qwen3-VL-235B-Instruct on SciAgentBench, a new 259-task benchmark for multi-step scientific tool-use.
-
An Encoder-Decoder Foundation Chemical Language Model for Generative Polymer Design
A T5-based polymer language model pre-trained on 100 million hypothetical polymers predicts thermal, electronic, and solubility properties and generates polymers conditioned on a target glass-transition temperature, w...
-
The evolution of AI from image interpretation toward scientific inference in nanoparticle electron microscopy
AI for nanoparticle TEM/STEM has progressed from detection and segmentation to physics-informed restoration, 2D-to-3D inference, and spatiotemporal analysis of in situ dynamics, with remaining gaps in benchmarking and...
Reference graph
Works this paper leans on
-
[1]
Foundation models for materials discovery–current state and future directions.npj Computational Materials, 11(1):61, 2025
Edward O Pyzer-Knapp, Matteo Manica, Peter Staar, Lucas Morin, Patrick Ruch, Teodoro Laino, John R Smith, and Alessandro Curioni. Foundation models for materials discovery–current state and future directions.npj Computational Materials, 11(1):61, 2025
2025
-
[2]
Zhenzhong Wang, Haowei Hua, Wanyu Lin, Ming Yang, and Kay Chen Tan. Crystalline material discovery in the era of artificial intelligence.arXiv preprint arXiv:2408.08044, 2024
arXiv 2024
-
[3]
Ai-driven inverse design of materials: Past, present and future.Chinese Physics Letters, 2024
Xiao-Qi Han, Xin-De Wang, Meng-Yuan Xu, Zhen Feng, Bo-Wen Yao, Peng-Jie Guo, Ze-Feng Gao, and Zhong-Yi Lu. Ai-driven inverse design of materials: Past, present and future.Chinese Physics Letters, 2024
2024
-
[4]
A review of large language models and autonomous agents in chemistry.Chemical Science, 2025
Mayk Caldas Ramos, Christopher J Collison, and Andrew D White. A review of large language models and autonomous agents in chemistry.Chemical Science, 2025
2025
-
[5]
A strategic approach to machine learning for material science: how to tackle real-world challenges and avoid pitfalls.Chemistry of Materials, 34(17): 7650–7665, 2022
Piyush Karande, Brian Gallagher, and Thomas Yong-Jin Han. A strategic approach to machine learning for material science: how to tackle real-world challenges and avoid pitfalls.Chemistry of Materials, 34(17): 7650–7665, 2022
2022
-
[6]
Scope of machine learning in materials research—a review.Applied Surface Science Advances, 18:100523, 2023
Md Hosne Mobarak, Mariam Akter Mimona, Md Aminul Islam, Nayem Hossain, Fatema Tuz Zohura, Ibnul Imtiaz, and Md Israfil Hossain Rimon. Scope of machine learning in materials research—a review.Applied Surface Science Advances, 18:100523, 2023
2023
-
[7]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. InProceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), pages 4171–4186, 2019
2019
-
[8]
Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019
2019
Show all 188 references
-
[9]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020
1901
-
[10]
Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023
2023 arXiv
-
[11]
Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1–113, 2023
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1–113, 2023
2023
-
[12]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InInternational conference on machine learning, pag...
2021
-
[13]
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. InProceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021
2021
-
[14]
Scaling deep learning for materials discovery.Nature, 624(7990):80–85, 2023
Amil Merchant, Simon Batzner, Samuel S Schoenholz, Muratahan Aykol, Gowoon Cheon, and Ekin Dogus Cubuk. Scaling deep learning for materials discovery.Nature, 624(7990):80–85, 2023
2023
-
[15]
Mattersim: A deep learning atomistic model across elements, temperatures and pressures
Han Yang, Chenxi Hu, Yichi Zhou, Xixian Liu, Yu Shi, Jielan Li, Guanzhi Li, Zekun Chen, Shuizhou Chen, Claudio Zeni, et al. Mattersim: A deep learning atomistic model across elements, temperatures and pressures. arXiv preprint arXiv:2405.04967, 2024. 22 Survey of AI for MSA PREPRINT
2024 arXiv
-
[16]
A foundation model for atomistic materials chemistry.arXiv preprint arXiv:2401.00096, 2023
Ilyes Batatia, Philipp Benner, Yuan Chiang, Alin M Elena, Dávid P Kovács, Janosh Riebesell, Xavier R Advincula, Mark Asta, Matthew Avaylon, William J Baldwin, et al. A foundation model for atomistic materials chemistry.arXiv preprint arXiv:2401.00096, 2023
2023 arXiv
-
[17]
Mattergen: a generative model for inorganic materials design
Claudio Zeni, Robert Pinsler, Daniel Zügner, Andrew Fowler, Matthew Horton, Xiang Fu, Sasha Shysheya, Jonathan Crabbé, Lixin Sun, Jake Smith, et al. Mattergen: a generative model for inorganic materials design. arXiv preprint arXiv:2312.03687, 2023
2023 arXiv
-
[18]
Space group constrained crystal generation.arXiv preprint arXiv:2402.03992, 2024
Rui Jiao, Wenbing Huang, Yu Liu, Deli Zhao, and Yang Liu. Space group constrained crystal generation.arXiv preprint arXiv:2402.03992, 2024
2024 arXiv
-
[19]
Space group informed transformer for crystalline materials generation.arXiv preprint arXiv:2403.15734, 2024
Zhendong Cao, Xiaoshan Luo, Jian Lv, and Lei Wang. Space group informed transformer for crystalline materials generation.arXiv preprint arXiv:2403.15734, 2024
2024
-
[20]
nach0: multimodal natural and chemical languages foundation model.Chemical Science, 15(22):8380–8389, 2024
Micha Livne, Zulfat Miftahutdinov, Elena Tutubalina, Maksim Kuznetsov, Daniil Polykovskiy, Annika Brundyn, Aastha Jhunjhunwala, Anthony Costa, Alex Aliper, Alán Aspuru-Guzik, et al. nach0: multimodal natural and chemical languages foundation model.Chemical Science, 15(22):8380...
2024
-
[21]
Multimodal foundation models for material property prediction and discovery.Newton, 2025
Viggo Moro, Charlotte Loh, Rumen Dangovski, Ali Ghorashi, Andrew Ma, Zhuo Chen, Samuel Kim, Peter Y Lu, Thomas Christensen, and Marin Soljaˇci´c. Multimodal foundation models for material property prediction and discovery.Newton, 2025
2025
-
[22]
Matterchat: A multi-modal llm for material science.arXiv preprint arXiv:2502.13107, 2025
Yingheng Tang, Wenbin Xu, Jie Cao, Jianzhu Ma, Weilu Gao, Steve Farrell, Benjamin Erichson, Michael W Mahoney, Andy Nonaka, and Zhi Yao. Matterchat: A multi-modal llm for material science.arXiv preprint arXiv:2502.13107, 2025
2025 arXiv
-
[23]
Atlantic: Structure-aware retrieval-augmented language model for interdisciplinary science.arXiv preprint arXiv:2311.12289, 2023
Sai Munikoti, Anurag Acharya, Sridevi Wagle, and Sameera Horawalavithana. Atlantic: Structure-aware retrieval-augmented language model for interdisciplinary science.arXiv preprint arXiv:2311.12289, 2023
2023 arXiv
-
[24]
Crystal structure generation with autoregressive large language modeling.Nature Communications, 15(1):1–16, 2024
Luis M Antunes, Keith T Butler, and Ricardo Grau-Crespo. Crystal structure generation with autoregressive large language modeling.Nature Communications, 15(1):1–16, 2024
2024
-
[25]
Accelerating material design with the generative toolkit for scientific discovery.npj Computational Materials, 9(1):69, 2023
Matteo Manica, Jannis Born, Joris Cadow, Dimitrios Christofidellis, Ashish Dave, Dean Clarke, Yves Gae- tan Nana Teukam, Giorgio Giannone, Samuel C Hoffman, Matthew Buchan, et al. Accelerating material design with the generative toolkit for scientific discovery.npj Computation...
2023
-
[26]
An autonomous laboratory for the accelerated synthesis of novel materials.Nature, 624(7990):86–91, 2023
Nathan J Szymanski, Bernardus Rendy, Yuxing Fei, Rishi E Kumar, Tanjin He, David Milsted, Matthew J McDermott, Max Gallant, Ekin Dogus Cubuk, Amil Merchant, et al. An autonomous laboratory for the accelerated synthesis of novel materials.Nature, 624(7990):86–91, 2023
2023
-
[27]
Developing a foundation model for predicting material failure
Agnese Marcato, Javier E Santos, Aleksandra Pachalieva, Kai Gao, Ryley Hill, Esteban Rougier, Qinjun Kang, Jeffrey Hyman, Abigail Hunter, Janel Chua, et al. Developing a foundation model for predicting material failure. arXiv preprint arXiv:2411.08354, 2024
2024 arXiv
-
[28]
Foundation models for the process industry: Challenges and opportunities.Engineering, 2025
Lei Ren, Haiteng Wang, Yuqing Wang, Keke Huang, Lihui Wang, and Bohu Li. Foundation models for the process industry: Challenges and opportunities.Engineering, 2025
2025
-
[29]
A perspective on foundation models in chemistry.JACS Au, 5(4):1499–1518, 2025
Junyoung Choi, Gunwook Nam, Jaesik Choi, and Yousung Jung. A perspective on foundation models in chemistry.JACS Au, 5(4):1499–1518, 2025
2025
-
[30]
Towards foundation models for materials science: The open matsci ml toolkit
Kin Long Kelvin Lee, Carmelo Gonzales, Matthew Spellings, Mikhail Galkin, Santiago Miret, and Nalini Kumar. Towards foundation models for materials science: The open matsci ml toolkit. InProceedings of the SC’23 Workshops of the International Conference on High Performance Com...
2023
-
[31]
Atomgpt: Atomistic generative pretrained transformer for forward and inverse materials design.The Journal of Physical Chemistry Letters, 15(27):6909–6917, 2024
Kamal Choudhary. Atomgpt: Atomistic generative pretrained transformer for forward and inverse materials design.The Journal of Physical Chemistry Letters, 15(27):6909–6917, 2024
2024
-
[32]
Multi-view mixture-of-experts for predicting molecular properties using smiles, selfies, and graph-based representations
Eduardo Soares, Indra Priyadarsini, Emilio Vital Brazil, Victor Yukio Shirasuna, and Seiji Takeda. Multi-view mixture-of-experts for predicting molecular properties using smiles, selfies, and graph-based representations. In Neurips 2024 Workshop Foundation Models for Science: ...
2024
-
[33]
Chemdfm: A large language foundation model for chemistry.arXiv preprint arXiv:2401.14818, 2024
Zihan Zhao, Da Ma, Lu Chen, Liangtai Sun, Zihao Li, Yi Xia, Bo Chen, Hongshen Xu, Zichen Zhu, Su Zhu, et al. Chemdfm: A large language foundation model for chemistry.arXiv preprint arXiv:2401.14818, 2024
2024 arXiv
-
[34]
Foundational large language models for materials research.arXiv preprint arXiv:2412.09560, 2024
Vaibhav Mishra, Somaditya Singh, Dhruv Ahlawat, Mohd Zaki, Vaibhav Bihani, Hargun Singh Grover, Biswajit Mishra, Santiago Miret, NM Krishnan, et al. Foundational large language models for materials research.arXiv preprint arXiv:2412.09560, 2024
2024 arXiv
-
[35]
Scitune: Aligning large language models with scientific multimodal instructions.arXiv preprint arXiv:2307.01139, 2023
Sameera Horawalavithana, Sai Munikoti, Ian Stewart, and Henry Kvinge. Scitune: Aligning large language models with scientific multimodal instructions.arXiv preprint arXiv:2307.01139, 2023. 23 Survey of AI for MSA PREPRINT
2023 arXiv
-
[36]
Honeycomb: A flexible llm-based agent system for materials science.arXiv preprint arXiv:2409.00135, 2024
Huan Zhang, Yu Song, Ziyu Hou, Santiago Miret, and Bang Liu. Honeycomb: A flexible llm-based agent system for materials science.arXiv preprint arXiv:2409.00135, 2024
2024 arXiv
-
[37]
Llmatdesign: Autonomous materials discovery with large language models.arXiv preprint arXiv:2406.13163, 2024
Shuyi Jia, Chao Zhang, and Victor Fung. Llmatdesign: Autonomous materials discovery with large language models.arXiv preprint arXiv:2406.13163, 2024
2024 arXiv
-
[38]
Chatmof: an artificial intelligence system for predicting and generating metal-organic frameworks using large language models.Nature communications, 15(1):4705, 2024
Yeonghun Kang and Jihan Kim. Chatmof: an artificial intelligence system for predicting and generating metal-organic frameworks using large language models.Nature communications, 15(1):4705, 2024
2024
-
[39]
Matagent: A human-in-the-loop multi-agent llm framework for accelerating the material science discovery cycle
Adib Bazgir, Yuwen Zhang, et al. Matagent: A human-in-the-loop multi-agent llm framework for accelerating the material science discovery cycle. InAI for Accelerated Materials Design-ICLR 2025
2025
-
[40]
Matpilot: an llm-enabled ai materials scientist under the framework of human-machine collaboration.arXiv preprint arXiv:2411.08063, 2024
Ziqi Ni, Yahao Li, Kaijia Hu, Kunyuan Han, Ming Xu, Xingyu Chen, Fengqi Liu, Yicong Ye, and Shuxin Bai. Matpilot: an llm-enabled ai materials scientist under the framework of human-machine collaboration.arXiv preprint arXiv:2411.08063, 2024
2024 arXiv
-
[41]
Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller
Andres M. Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. Augmenting large language models with chemistry tools.Nature Machine Intelligence, 6(5):525–535, 2024
2024
-
[42]
The open matsci ml toolkit: A flexible framework for machine learning in materials science.arXiv preprint arXiv:2210.17484, 2022
Santiago Miret, Kin Long Kelvin Lee, Carmelo Gonzales, Marcel Nassar, and Matthew Spellings. The open matsci ml toolkit: A flexible framework for machine learning in materials science.arXiv preprint arXiv:2210.17484, 2022
2022 arXiv
-
[43]
Forge: Pre-training open foundation models for science
Junqi Yin, Sajal Dash, Feiyi Wang, and Mallikarjun Shankar. Forge: Pre-training open foundation models for science. InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–13, 2023
2023
-
[44]
Attention is all you need.Advances in neural information processing systems, 30, 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need.Advances in neural information processing systems, 30, 2017
2017
-
[45]
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. 2018
2018
-
[46]
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalk- wyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023
2023 arXiv
-
[47]
Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[48]
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension.arXiv preprint arXiv:1910.13461, 2019
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension.arXiv preprint arXiv:1910.13461, 2019
1910 arXiv
-
[49]
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. InInternational Conference on Learning Representations (ICLR), 2023
2023
-
[50]
Llm-planner: Few-shot grounded planning for embodied agents with large language models
Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M Sadler, Wei-Lun Chao, and Yu Su. Llm-planner: Few-shot grounded planning for embodied agents with large language models. InProceedings of the IEEE/CVF international conference on computer vision, pages 2998–3009, 2023
2023
-
[51]
Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023
2023
-
[52]
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. InInternational conference on machine learning, pages 9118–9147. PMLR, 2022
2022
-
[53]
Theory of mind may have spontaneously emerged in large language models.arXiv preprint arXiv:2302.02083, 4:169, 2023
Michal Kosinski. Theory of mind may have spontaneously emerged in large language models.arXiv preprint arXiv:2302.02083, 4:169, 2023
2023 arXiv
-
[54]
Role play with large language models.Nature, 623 (7987):493–498, 2023
Murray Shanahan, Kyle McDonell, and Laria Reynolds. Role play with large language models.Nature, 623 (7987):493–498, 2023
2023
-
[55]
A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6):186345, 2024
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. A survey on large language model based autonomous agents.Frontiers of Computer Science, 18(6):186345, 2024
2024
-
[56]
Honeybee: Progressive instruction finetuning of large language models for materials science.arXiv preprint arXiv:2310.08511, 2023
Yu Song, Santiago Miret, Huan Zhang, and Bang Liu. Honeybee: Progressive instruction finetuning of large language models for materials science.arXiv preprint arXiv:2310.08511, 2023. 24 Survey of AI for MSA PREPRINT
2023 arXiv
-
[57]
Crystal graph convolutional neural networks for an accurate and interpretable prediction of material properties.Physical review letters, 120(14):145301, 2018
Tian Xie and Jeffrey C Grossman. Crystal graph convolutional neural networks for an accurate and interpretable prediction of material properties.Physical review letters, 120(14):145301, 2018
2018
-
[58]
Graph networks as a universal machine learning framework for molecules and crystals.Chemistry of Materials, 31(9):3564–3572, 2019
Chi Chen, Weike Ye, Yunxing Zuo, Chen Zheng, and Shyue Ping Ong. Graph networks as a universal machine learning framework for molecules and crystals.Chemistry of Materials, 31(9):3564–3572, 2019
2019
-
[59]
Catalyst energy prediction with catberta: Unveiling feature exploration strategies through large language models.ACS Catalysis, 13(24):16032–16044, 2023
Janghoon Ock, Chakradhar Guntuboina, and Amir Barati Farimani. Catalyst energy prediction with catberta: Unveiling feature exploration strategies through large language models.ACS Catalysis, 13(24):16032–16044, 2023
2023
-
[60]
Molxpt: Wrapping molecules with text for generative pre-training.arXiv preprint arXiv:2305.10688, 2023
Zequn Liu, Wei Zhang, Yingce Xia, Lijun Wu, Shufang Xie, Tao Qin, Ming Zhang, and Tie-Yan Liu. Molxpt: Wrapping molecules with text for generative pre-training.arXiv preprint arXiv:2305.10688, 2023
2023 arXiv
-
[61]
Crystal diffusion variational autoencoder for periodic material generation.arXiv preprint arXiv:2110.06197, 2021
Tian Xie, Xiang Fu, Octavian-Eugen Ganea, Regina Barzilay, and Tommi Jaakkola. Crystal diffusion variational autoencoder for periodic material generation.arXiv preprint arXiv:2110.06197, 2021
2021 arXiv
-
[62]
Crystal structure prediction by joint equivariant diffusion.Advances in Neural Information Processing Systems, 36:17464–17497, 2023
Rui Jiao, Wenbing Huang, Peijia Lin, Jiaqi Han, Pin Chen, Yutong Lu, and Yang Liu. Crystal structure prediction by joint equivariant diffusion.Advances in Neural Information Processing Systems, 36:17464–17497, 2023
2023
-
[63]
Gp-molformer: A foundation model for molecular generation.arXiv preprint arXiv:2405.04912, 2024
Jerret Ross, Brian Belgodere, Samuel C Hoffman, Vijil Chenthamarakshan, Jiri Navratil, Youssef Mroueh, and Payel Das. Gp-molformer: A foundation model for molecular generation.arXiv preprint arXiv:2405.04912, 2024
2024 arXiv
-
[64]
Mattergpt: A generative transformer for multi-property inverse design of solid-state materials.arXiv preprint arXiv:2408.07608, 2024
Yan Chen, Xueru Wang, Xiaobin Deng, Yilun Liu, Xi Chen, Yunwei Zhang, Lei Wang, and Hang Xiao. Mattergpt: A generative transformer for multi-property inverse design of solid-state materials.arXiv preprint arXiv:2408.07608, 2024
2024 arXiv
-
[65]
Chemformer: a pre-trained transformer for computational chemistry.Machine Learning: Science and Technology, 3(1):015022, 2022
Ross Irwin, Spyridon Dimitriadis, Jiazhen He, and Esben Jannik Bjerrum. Chemformer: a pre-trained transformer for computational chemistry.Machine Learning: Science and Technology, 3(1):015022, 2022
2022
-
[66]
Flowllm: Flow matching for material generation with large language models as base distributions.Advances in Neural Information Processing Systems, 37:46025–46046, 2024
Anuroop Sriram, Benjamin Miller, Ricky TQ Chen, and Brandon Wood. Flowllm: Flow matching for material generation with large language models as base distributions.Advances in Neural Information Processing Systems, 37:46025–46046, 2024
2024
-
[67]
Is large language model all you need to predict the synthesizability and precursors of crystal structures?arXiv preprint arXiv:2407.07016, 2024
Zhilong Song, Shuaihua Lu, Minggang Ju, Qionghua Zhou, and Jinlan Wang. Is large language model all you need to predict the synthesizability and precursors of crystal structures?arXiv preprint arXiv:2407.07016, 2024
2024 arXiv
-
[68]
Llm-fusion: A novel multimodal fusion model for accelerated material discovery.arXiv preprint arXiv:2503.01022, 2025
Onur Boyar, Indra Priyadarsini, Seiji Takeda, and Lisa Hamada. Llm-fusion: A novel multimodal fusion model for accelerated material discovery.arXiv preprint arXiv:2503.01022, 2025
2025 arXiv
-
[69]
Unifying molecular and textual representations via multi-task language modelling
Dimitrios Christofidellis, Giorgio Giannone, Jannis Born, Ole Winther, Teodoro Laino, and Matteo Manica. Unifying molecular and textual representations via multi-task language modelling. InInternational Conference on Machine Learning, pages 6140–6157. PMLR, 2023
2023
-
[70]
Quantifying the advantage of domain-specific pre-training on named entity recognition tasks in materials science.Patterns, 3(4), 2022
Amalie Trewartha, Nicholas Walker, Haoyan Huo, Sanghoon Lee, Kevin Cruse, John Dagdelen, Alexander Dunn, Kristin A Persson, Gerbrand Ceder, and Anubhav Jain. Quantifying the advantage of domain-specific pre-training on named entity recognition tasks in materials science.Patter...
2022
-
[71]
Matscibert: A materials domain language model for text mining and information extraction.npj Computational Materials, 8(1):102, 2022
Tanishq Gupta, Mohd Zaki, NM Anoop Krishnan, and Mausam. Matscibert: A materials domain language model for text mining and information extraction.npj Computational Materials, 8(1):102, 2022
2022
-
[72]
A general-purpose material property data extraction pipeline from large polymer corpora using natural language processing.npj Computational Materials, 9(1):52, 2023
Pranav Shetty, Arunkumar Chitteth Rajan, Chris Kuenneth, Sonakshi Gupta, Lakshmi Prerana Panchumarti, Lauren Holm, Chao Zhang, and Rampi Ramprasad. A general-purpose material property data extraction pipeline from large polymer corpora using natural language processing.npj Com...
2023
-
[73]
Molscribe: robust molecular structure recognition with image-to-graph generation.Journal of Chemical Information and Modeling, 63(7):1925–1934, 2023
Yujie Qian, Jiang Guo, Zhengkai Tu, Zhening Li, Connor W Coley, and Regina Barzilay. Molscribe: robust molecular structure recognition with image-to-graph generation.Journal of Chemical Information and Modeling, 63(7):1925–1934, 2023
1925
-
[74]
Deplot: One-shot visual language reasoning by plot-to- table translation.arXiv preprint arXiv:2212.10505, 2022
Fangyu Liu, Julian Martin Eisenschlos, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Wenhu Chen, Nigel Collier, and Yasemin Altun. Deplot: One-shot visual language reasoning by plot-to- table translation.arXiv preprint arXiv:2212.10505, 2022
2022 arXiv
-
[75]
Mascqa: investigating materials science knowledge of large language models.Digital Discovery, 3(2):313–327, 2024
Mohd Zaki, NM Anoop Krishnan, et al. Mascqa: investigating materials science knowledge of large language models.Digital Discovery, 3(2):313–327, 2024
2024
-
[76]
Assessment of fine-tuned large language models for real-world chemistry and material science applications.Chemical science, 16(2):670–684, 2025
Joren Van Herck, María Victoria Gil, Kevin Maik Jablonka, Alex Abrudan, Andy S Anker, Mehrdad Asgari, Ben Blaiszik, Antonio Buffo, Leander Choudhury, Clemence Corminboeuf, et al. Assessment of fine-tuned large language models for real-world chemistry and material science appli...
2025
-
[77]
Gpt-j-6b: A 6 billion parameter autoregressive language model, 2021
Ben Wang and Aran Komatsuzaki. Gpt-j-6b: A 6 billion parameter autoregressive language model, 2021. 25 Survey of AI for MSA PREPRINT
2021
-
[78]
The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024
2024 arXiv
-
[79]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...
2023 arXiv
-
[80]
Smiles, a chemical language and information system
David Weininger. Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules.Journal of chemical information and computer sciences, 28(1):31–36, 1988
1988
-
[81]
Uni-smart: Universal science multimodal analysis and research transformer
Hengxing Cai, Xiaochen Cai, Shuwen Yang, Jiankun Wang, Lin Yao, Zhifeng Gao, Junhan Chang, Sihang Li, Mingjun Xu, Changxin Wang, et al. Uni-smart: Universal science multimodal analysis and research transformer. arXiv preprint arXiv:2403.10301, 2024
2024 arXiv
-
[82]
The claude 3 model family: Opus, sonnet, haiku
Anthropic. The claude 3 model family: Opus, sonnet, haiku. 2025. URL https://api.semanticscholar. org/CorpusID:268232499
2025
-
[83]
Do transformers really perform badly for graph representation?Advances in neural information processing systems, 34:28877–28888, 2021
Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu. Do transformers really perform badly for graph representation?Advances in neural information processing systems, 34:28877–28888, 2021
2021
-
[84]
Mace: Higher order equivariant message passing neural networks for fast and accurate force fields.Advances in neural information processing systems, 35:11423–11436, 2022
Ilyes Batatia, David P Kovacs, Gregor Simm, Christoph Ortner, and Gábor Csányi. Mace: Higher order equivariant message passing neural networks for fast and accurate force fields.Advances in neural information processing systems, 35:11423–11436, 2022
2022
-
[85]
Ani-1: an extensible neural network potential with dft accuracy at force field computational cost.Chemical science, 8(4):3192–3203, 2017
Justin S Smith, Olexandr Isayev, and Adrian E Roitberg. Ani-1: an extensible neural network potential with dft accuracy at force field computational cost.Chemical science, 8(4):3192–3203, 2017
2017
-
[86]
Generalized neural-network representation of high-dimensional potential- energy surfaces.Physical review letters, 98(14):146401, 2007
Jörg Behler and Michele Parrinello. Generalized neural-network representation of high-dimensional potential- energy surfaces.Physical review letters, 98(14):146401, 2007
2007
-
[87]
Accurate and transferable multitask prediction of chemical properties with an atoms-in-molecules neural network.Science advances, 5(8):eaav6490, 2019
Roman Zubatyuk, Justin S Smith, Jerzy Leszczynski, and Olexandr Isayev. Accurate and transferable multitask prediction of chemical properties with an atoms-in-molecules neural network.Science advances, 5(8):eaav6490, 2019
2019
-
[88]
Aimnet2: a neural network potential to meet your neutral, charged, organic, and elemental-organic needs.Chemical Science, 2025
Dylan M Anstine, Roman Zubatyuk, and Olexandr Isayev. Aimnet2: a neural network potential to meet your neutral, charged, organic, and elemental-organic needs.Chemical Science, 2025
2025
-
[89]
The graph neural network model.IEEE transactions on neural networks, 20(1):61–80, 2008
Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. The graph neural network model.IEEE transactions on neural networks, 20(1):61–80, 2008
2008
-
[90]
Self-referencing embedded strings (selfies): A 100% robust molecular string representation.Machine Learning: Science and Technology, 1(4):045024, 2020
Mario Krenn, Florian Häse, AkshatKumar Nigam, Pascal Friederich, and Alan Aspuru-Guzik. Self-referencing embedded strings (selfies): A 100% robust molecular string representation.Machine Learning: Science and Technology, 1(4):045024, 2020
2020
-
[91]
Zinc 15–ligand discovery for everyone.Journal of chemical information and modeling, 55(11):2324–2337, 2015
Teague Sterling and John J Irwin. Zinc 15–ligand discovery for everyone.Journal of chemical information and modeling, 55(11):2324–2337, 2015
2015
-
[92]
Zinc20—a free ultralarge-scale chemical database for ligand discovery.Journal of chemical information and modeling, 60(12):6065–6073, 2020
John J Irwin, Khanh G Tang, Jennifer Young, Chinzorig Dandarchuluun, Benjamin R Wong, Munkhzul Khurel- baatar, Yurii S Moroz, John Mayfield, and Roger A Sayle. Zinc20—a free ultralarge-scale chemical database for ligand discovery.Journal of chemical information and modeling, 6...
2020
-
[93]
The chembl database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods.Nucleic acids research, 52(D1): D1180–D1192, 2024
Barbara Zdrazil, Eloy Felix, Fiona Hunter, Emma J Manners, James Blackshaw, Sybilla Corbett, Marleen de Veij, Harris Ioannidis, David Mendez Lopez, Juan F Mosquera, et al. The chembl database in 2023: a drug discovery platform spanning multiple bioactivity data types and time ...
2023
-
[94]
Chembl web services: streamlining access to drug discovery data and utilities
Mark Davies, Michał Nowotka, George Papadatos, Nathan Dedman, Anna Gaulton, Francis Atkinson, Louisa Bellis, and John P Overington. Chembl web services: streamlining access to drug discovery data and utilities. Nucleic acids research, 43(W1):W612–W620, 2015
2015
-
[95]
Schnet–a deep learning architecture for molecules and materials.The Journal of Chemical Physics, 148(24), 2018
Kristof T Schütt, Huziel E Sauceda, P-J Kindermans, Alexandre Tkatchenko, and K-R Müller. Schnet–a deep learning architecture for molecules and materials.The Journal of Chemical Physics, 148(24), 2018
2018
-
[96]
Commentary: The materials project: A materials genome approach to accelerating materials innovation.APL materials, 1(1), 2013
Anubhav Jain, Shyue Ping Ong, Geoffroy Hautier, Wei Chen, William Davidson Richards, Stephen Dacek, Shreyas Cholia, Dan Gunter, David Skinner, Gerbrand Ceder, et al. Commentary: The materials project: A materials genome approach to accelerating materials innovation.APL materia...
2013
-
[97]
Enumeration of 166 billion organic small molecules in the chemical universe database gdb-17.Journal of chemical information and modeling, 52(11):2864–2875, 2012
Lars Ruddigkeit, Ruud Van Deursen, Lorenz C Blum, and Jean-Louis Reymond. Enumeration of 166 billion organic small molecules in the chemical universe database gdb-17.Journal of chemical information and modeling, 52(11):2864–2875, 2012
2012
-
[98]
Quantum chemistry structures and properties of 134 kilo molecules.Scientific data, 1(1):1–7, 2014
Raghunathan Ramakrishnan, Pavlo O Dral, Matthias Rupp, and O Anatole V on Lilienfeld. Quantum chemistry structures and properties of 134 kilo molecules.Scientific data, 1(1):1–7, 2014
2014
-
[99]
A compact review of molecular property prediction with graph neural networks.Drug Discovery Today: Technologies, 37:1–12, 2020
Oliver Wieder, Stefan Kohlbacher, Mélaine Kuenemann, Arthur Garon, Pierre Ducrot, Thomas Seidel, and Thierry Langer. A compact review of molecular property prediction with graph neural networks.Drug Discovery Today: Technologies, 37:1–12, 2020
2020
-
[100]
Chemberta: large-scale self-supervised pretraining for molecular property prediction.arXiv preprint arXiv:2010.09885, 2020
Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. Chemberta: large-scale self-supervised pretraining for molecular property prediction.arXiv preprint arXiv:2010.09885, 2020
2010 arXiv
-
[101]
Mol-bert: An effective molecular representation with bert for molecular property prediction.Wireless Communications and Mobile Computing, 2021(1):7181815, 2021
Juncai Li and Xiaofei Jiang. Mol-bert: An effective molecular representation with bert for molecular property prediction.Wireless Communications and Mobile Computing, 2021(1):7181815, 2021
2021
-
[102]
Roberta: A robustly optimized bert pretraining approach.arXiv preprint arXiv:1907.11692, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach.arXiv preprint arXiv:1907.11692, 2019
1907 arXiv
-
[103]
Pubchem 2025 update.Nucleic Acids Research, 53(D1):D1516– D1525, 2025
Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, et al. Pubchem 2025 update.Nucleic Acids Research, 53(D1):D1516– D1525, 2025
2025
-
[104]
Open catalyst 2020 (oc20) dataset and community challenges.Acs Catalysis, 11(10):6059–6072, 2021
Lowik Chanussot, Abhishek Das, Siddharth Goyal, Thibaut Lavril, Muhammed Shuaibi, Morgane Riviere, Kevin Tran, Javier Heras-Domingo, Caleb Ho, Weihua Hu, et al. Open catalyst 2020 (oc20) dataset and community challenges.Acs Catalysis, 11(10):6059–6072, 2021
2020
-
[105]
Selformer: molecular representation learning via selfies language models.Machine Learning: Science and Technology, 4(2):025035, 2023
Atakan Yüksel, Erva Ulusoy, Atabey Ünlü, and Tunca Do˘gan. Selformer: molecular representation learning via selfies language models.Machine Learning: Science and Technology, 4(2):025035, 2023
2023
-
[106]
Periodic graph transformers for crystal material property prediction.Advances in Neural Information Processing Systems, 35:15066–15080, 2022
Keqiang Yan, Yi Liu, Yuchao Lin, and Shuiwang Ji. Periodic graph transformers for crystal material property prediction.Advances in Neural Information Processing Systems, 35:15066–15080, 2022
2022
-
[107]
Complete and efficient graph transformers for crystal material property prediction.arXiv preprint arXiv:2403.11857, 2024
Keqiang Yan, Cong Fu, Xiaofeng Qian, Xiaoning Qian, and Shuiwang Ji. Complete and efficient graph transformers for crystal material property prediction.arXiv preprint arXiv:2403.11857, 2024
2024 arXiv
-
[108]
Generative pre-training from molecules
Sanjar Adilov. Generative pre-training from molecules. 2021
2021
-
[109]
PubMed Database
National Library of Medicine (US). PubMed Database. https://pubmed.ncbi.nlm.nih.gov/, 2024. Ac- cessed: 2025-06-09
2024
-
[110]
A smile is all you need: predicting limiting activity coefficients from smiles with natural language processing.Digital Discovery, 1(6):859–869, 2022
Benedikt Winter, Clemens Winter, Johannes Schilling, and André Bardow. A smile is all you need: predicting limiting activity coefficients from smiles with natural language processing.Digital Discovery, 1(6):859–869, 2022
2022
-
[111]
Leveraging large language models for predictive chemistry.Nature Machine Intelligence, 6(2):161–169, 2024
Kevin Maik Jablonka, Philippe Schwaller, Andres Ortega-Guerrero, and Berend Smit. Leveraging large language models for predictive chemistry.Nature Machine Intelligence, 6(2):161–169, 2024
2024
-
[112]
Symmetry-adapted generation of 3d point sets for the targeted discovery of molecules.Advances in neural information processing systems, 32, 2019
Niklas Gebauer, Michael Gastegger, and Kristof Schütt. Symmetry-adapted generation of 3d point sets for the targeted discovery of molecules.Advances in neural information processing systems, 32, 2019
2019
-
[113]
Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
2020
-
[114]
Auto-encoding variational bayes, 2013
Diederik P Kingma, Max Welling, et al. Auto-encoding variational bayes, 2013
2013
-
[115]
An introduction to variational autoencoders.Foundations and Trends® in Machine Learning, 12(4):307–392, 2019
Diederik P Kingma, Max Welling, et al. An introduction to variational autoencoders.Foundations and Trends® in Machine Learning, 12(4):307–392, 2019
2019
-
[116]
Con-cdvae: A method for the conditional generation of crystal structures.Computational Materials Today, 1:100003, 2024
Cai-Yuan Ye, Hong-Ming Weng, and Quan-Sheng Wu. Con-cdvae: A method for the conditional generation of crystal structures.Computational Materials Today, 1:100003, 2024
2024
-
[117]
An invertible, invariant crystal representation for inverse design of solid-state materials using generative deep learning.Nature Communications, 14(1):7027, 2023
Hang Xiao, Rong Li, Xiaoyang Shi, Yan Chen, Liangliang Zhu, Xi Chen, and Lei Wang. An invertible, invariant crystal representation for inverse design of solid-state materials using generative deep learning.Nature Communications, 14(1):7027, 2023
2023
-
[118]
Fine-tuned language models generate stable inorganic materials as text.arXiv preprint arXiv:2402.04379, 2024
Nate Gruver, Anuroop Sriram, Andrea Madotto, Andrew Gordon Wilson, C Lawrence Zitnick, and Zachary Ulissi. Fine-tuned language models generate stable inorganic materials as text.arXiv preprint arXiv:2402.04379, 2024
2024 arXiv
-
[119]
Language models can generate molecules, materials, and protein binding sites directly in three dimensions as xyz, cif, and pdb files.arXiv preprint arXiv:2305.05708, 2023
Daniel Flam-Shepherd and Alán Aspuru-Guzik. Language models can generate molecules, materials, and protein binding sites directly in three dimensions as xyz, cif, and pdb files.arXiv preprint arXiv:2305.05708, 2023. 27 Survey of AI for MSA PREPRINT
2023 arXiv
-
[120]
Jarvis: An integrated infrastructure for data-driven materials design.Preprint at, https://arxiv
Kamal Choudhary, Kevin F Garrity, Andrew CE Reid, Brian DeCost, Adam J Biacchi, Angela R Hight Walker, Zachary Trautt, Jason Hattrick-Simpers, A Gilad Kusne, Andrea Centrone, et al. Jarvis: An integrated infrastructure for data-driven materials design.Preprint at, https://arxi...
2007 arXiv
-
[121]
Large language models are innate crystal structure generators
Jingru Gan, Peichen Zhong, Yuanqi Du, Yanqiao Zhu, Chenru Duan, Haorui Wang, Carla P Gomes, Kristin A Persson, Daniel Schwalbe-Koda, and Wei Wang. Large language models are innate crystal structure generators. arXiv preprint arXiv:2502.20933, 2025
2025
-
[122]
Molecular transformer: a model for uncertainty-calibrated chemical reaction prediction.ACS central science, 5(9):1572–1583, 2019
Philippe Schwaller, Teodoro Laino, Théophile Gaudin, Peter Bolgar, Christopher A Hunter, Costas Bekas, and Alpha A Lee. Molecular transformer: a model for uncertainty-calibrated chemical reaction prediction.ACS central science, 5(9):1572–1583, 2019
2019
-
[123]
Uspto – open data portal.https://data.uspto.gov/home, 2025
USPTO. Uspto – open data portal.https://data.uspto.gov/home, 2025
2025
-
[124]
Esol: estimating aqueous solubility directly from molecular structure.Journal of chemical information and computer sciences, 44(3):1000–1005, 2004
John S Delaney. Esol: estimating aqueous solubility directly from molecular structure.Journal of chemical information and computer sciences, 44(3):1000–1005, 2004
2004
-
[125]
Matchat: A large language model and application service platform for materials science.Chinese Physics B, 32(11):118104, 2023
Zi-Yi Chen, Fan-Kai Xie, Meng Wan, Yang Yuan, Miao Liu, Zong-Guo Wang, Sheng Meng, and Yan-Gang Wang. Matchat: A large language model and application service platform for materials science.Chinese Physics B, 32(11):118104, 2023
2023
-
[126]
Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022
2022
-
[127]
Synask: unleashing the power of large language models in organic synthesis.Chemical Science, 16(1):43–56, 2025
Chonghuan Zhang, Qianghua Lin, Biwei Zhu, Haopeng Yang, Xiao Lian, Hao Deng, Jiajun Zheng, and Kuangbiao Liao. Synask: unleashing the power of large language models in organic synthesis.Chemical Science, 16(1):43–56, 2025
2025
-
[128]
Are llms ready for real-world materials discovery?arXiv preprint arXiv:2402.05200, 2024
Santiago Miret and Nandan M Krishnan. Are llms ready for real-world materials discovery?arXiv preprint arXiv:2402.05200, 2024
2024 arXiv
-
[129]
Mesoscopic and multiscale modelling in materials.Nature materials, 20(6):774–786, 2021
Jacob Fish, Gregory J Wagner, and Sinan Keten. Mesoscopic and multiscale modelling in materials.Nature materials, 20(6):774–786, 2021
2021
-
[130]
Machine learning-driven multiscale modeling: bridging the scales with a next-generation simulation infrastructure.Journal of Chemical Theory and Computation, 19(9):2658–2675, 2023
Helgi I Ingólfsson, Harsh Bhatia, Fikret Aydin, Tomas Oppelstrup, Cesar A López, Liam G Stanton, Timothy S Carpenter, Sergio Wong, Francesco Di Natale, Xiaohua Zhang, et al. Machine learning-driven multiscale modeling: bridging the scales with a next-generation simulation infr...
2023
-
[131]
Efficient multiscale modeling of heterogeneous materials using deep neural networks.Computational Mechanics, 72(1):155–171, 2023
Fadi Aldakheel, Elsayed S Elsayed, Tarek I Zohdi, and Peter Wriggers. Efficient multiscale modeling of heterogeneous materials using deep neural networks.Computational Mechanics, 72(1):155–171, 2023
2023
-
[132]
Pistachio: Search and faceting of large reaction databases
John Mayfield, Daniel Lowe, and Roger Sayle. Pistachio: Search and faceting of large reaction databases. In ABSTRACTS OF PAPERS OF THE AMERICAN CHEMICAL SOCIETY, volume 254. AMER CHEMICAL SOC 1155 16TH ST, NW, W ASHINGTON, DC 20036 USA, 2017
2017
-
[133]
Text2mol: Cross-modal molecule retrieval with natural language queries
Carl Edwards, ChengXiang Zhai, and Heng Ji. Text2mol: Cross-modal molecule retrieval with natural language queries. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 595–607, 2021
2021
-
[134]
Regression transformer enables concurrent sequence regression and generation for molecular language modelling.Nature Machine Intelligence, 5(4):432–444, 2023
Jannis Born and Matteo Manica. Regression transformer enables concurrent sequence regression and generation for molecular language modelling.Nature Machine Intelligence, 5(4):432–444, 2023
2023
-
[135]
Moleculenet: a benchmark for molecular machine learning.Chemical science, 9(2): 513–530, 2018
Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande. Moleculenet: a benchmark for molecular machine learning.Chemical science, 9(2): 513–530, 2018
2018
-
[136]
Evaluating protein transfer learning with tape.Advances in neural information processing systems, 32, 2019
Roshan Rao, Nicholas Bhattacharya, Neil Thomas, Yan Duan, Peter Chen, John Canny, Pieter Abbeel, and Yun Song. Evaluating protein transfer learning with tape.Advances in neural information processing systems, 32, 2019
2019
-
[137]
A large encoder-decoder family of foundation models for chemical language.arXiv preprint arXiv:2407.20267, 2024
Eduardo Soares, Victor Shirasuna, Emilio Vital Brazil, Renato Cerqueira, Dmitry Zubarev, and Kristin Schmidt. A large encoder-decoder family of foundation models for chemical language.arXiv preprint arXiv:2407.20267, 2024
2024 arXiv
-
[138]
Self-bart: A transformer-based molecular representation model using selfies.arXiv preprint arXiv:2410.12348, 2024
Indra Priyadarsini, Seiji Takeda, Lisa Hamada, Emilio Vital Brazil, Eduardo Soares, and Hajime Shinohara. Self-bart: A transformer-based molecular representation model using selfies.arXiv preprint arXiv:2410.12348, 2024
2024 arXiv
-
[139]
Mhg-gnn: Combination of molecular hypergraph grammar with graph neural network.arXiv preprint arXiv:2309.16374, 2023
Akihiro Kishimoto, Hiroshi Kajino, Masataka Hirose, Junta Fuchiwaki, Indra Priyadarsini, Lisa Hamada, Hajime Shinohara, Daiju Nakano, and Seiji Takeda. Mhg-gnn: Combination of molecular hypergraph grammar with graph neural network.arXiv preprint arXiv:2309.16374, 2023. 28 Surv...
2023 arXiv
-
[140]
Improvements to bm25 and language models examined
Andrew Trotman, Antti Puurula, and Blake Burgess. Improvements to bm25 and language models examined. In Proceedings of the 19th Australasian Document Computing Symposium, pages 58–65, 2014
2014
-
[141]
Unsupervised dense information retrieval with contrastive learning.arXiv preprint arXiv:2112.09118, 2021
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. Unsupervised dense information retrieval with contrastive learning.arXiv preprint arXiv:2112.09118, 2021
2021 arXiv
-
[142]
Benchmarking graph neural networks for materials chemistry.npj Computational Materials, 7(1):84, 2021
Victor Fung, Jiaxin Zhang, Eric Juarez, and Bobby G Sumpter. Benchmarking graph neural networks for materials chemistry.npj Computational Materials, 7(1):84, 2021
2021
-
[143]
Torchmd-net: equivariant transformers for neural network based molecular potentials.arXiv preprint arXiv:2202.02541, 2022
Philipp Thölke and Gianni De Fabritiis. Torchmd-net: equivariant transformers for neural network based molecular potentials.arXiv preprint arXiv:2202.02541, 2022
2022 arXiv
-
[144]
Recent developments in the inorganic crystal structure database: theoretical crystal structure data and related features.Applied Crystallography, 52(5): 918–925, 2019
Dejan Zagorac, H Müller, S Ruehl, J Zagorac, and Silke Rehme. Recent developments in the inorganic crystal structure database: theoretical crystal structure data and related features.Applied Crystallography, 52(5): 918–925, 2019
2019
-
[145]
The open quantum materials database (oqmd): assessing the accuracy of dft formation energies.npj Computational Materials, 1(1):1–15, 2015
Scott Kirklin, James E Saal, Bryce Meredig, Alex Thompson, Jeff W Doak, Muratahan Aykol, Stephan Rühl, and Chris Wolverton. The open quantum materials database (oqmd): assessing the accuracy of dft formation energies.npj Computational Materials, 1(1):1–15, 2015
2015
-
[146]
Materials design and discovery with high-throughput density functional theory: the open quantum materials database (oqmd).Jom, 65:1501–1509, 2013
James E Saal, Scott Kirklin, Muratahan Aykol, Bryce Meredig, and Christopher Wolverton. Materials design and discovery with high-throughput density functional theory: the open quantum materials database (oqmd).Jom, 65:1501–1509, 2013
2013
-
[147]
The nomad laboratory: from data sharing to artificial intelligence.Journal of Physics: Materials, 2(3):036001, 2019
Claudia Draxl and Matthias Scheffler. The nomad laboratory: from data sharing to artificial intelligence.Journal of Physics: Materials, 2(3):036001, 2019
2019
-
[148]
Open materials 2024 (omat24) inorganic materials dataset and models.arXiv preprint arXiv:2410.12771, 2024
Luis Barroso-Luque, Muhammed Shuaibi, Xiang Fu, Brandon M Wood, Misko Dzamba, Meng Gao, Ammar Rizvi, C Lawrence Zitnick, and Zachary W Ulissi. Open materials 2024 (omat24) inorganic materials dataset and models.arXiv preprint arXiv:2410.12771, 2024
2024 arXiv
-
[149]
Snumat.https://www.snumat.com/apis, 2025
SNU MDIL. Snumat.https://www.snumat.com/apis, 2025
2025
-
[150]
Chgnet as a pretrained universal neural network potential for charge-informed atomistic modelling
Bowen Deng, Peichen Zhong, KyuJung Jun, Janosh Riebesell, Kevin Han, Christopher J Bartel, and Gerbrand Ceder. Chgnet as a pretrained universal neural network potential for charge-informed atomistic modelling. Nature Machine Intelligence, 5(9):1031–1041, 2023
2023
-
[151]
Improving machine-learning models in materials science through large datasets
Jonathan Schmidt, Tiago FT Cerqueira, Aldo H Romero, Antoine Loew, Fabian Jäger, Hai-Chen Wang, Silvana Botti, and Miguel AL Marques. Improving machine-learning models in materials science through large datasets. Materials Today Physics, 48:101560, 2024
2024
-
[152]
Guacamol: benchmarking models for de novo molecular design.Journal of chemical information and modeling, 59(3):1096–1108, 2019
Nathan Brown, Marco Fiscato, Marwin HS Segler, and Alain C Vaucher. Guacamol: benchmarking models for de novo molecular design.Journal of chemical information and modeling, 59(3):1096–1108, 2019
2019
-
[153]
Molecular sets (moses): a benchmarking platform for molecular generation models.Frontiers in pharmacology, 11:565644, 2020
Daniil Polykovskiy, Alexander Zhebrak, Benjamin Sanchez-Lengeling, Sergey Golovanov, Oktai Tatanov, Stanislav Belyaev, Rauf Kurbanov, Aleksey Artamonov, Vladimir Aladinskiy, Mark Veselov, et al. Molecular sets (moses): a benchmarking platform for molecular generation models.Fr...
2020
-
[154]
Unsupervised word embeddings capture latent knowledge from materials science literature.Nature, 571(7763):95–98, 2019
Vahe Tshitoyan, John Dagdelen, Leigh Weston, Alexander Dunn, Ziqin Rong, Olga Kononova, Kristin A Persson, Gerbrand Ceder, and Anubhav Jain. Unsupervised word embeddings capture latent knowledge from materials science literature.Nature, 571(7763):95–98, 2019
2019
-
[155]
Leigh Weston, Vahe Tshitoyan, John Dagdelen, Olga Kononova, Amalie Trewartha, Kristin A Persson, Gerbrand Ceder, and Anubhav Jain. Named entity recognition and normalization applied to large-scale information extraction from the materials science literature.Journal of chemical...
2019
-
[156]
Matsci-nlp: Evaluating scientific language models on materials science language tasks using text-to-schema modeling.arXiv preprint arXiv:2305.08264, 2023
Yu Song, Santiago Miret, and Bang Liu. Matsci-nlp: Evaluating scientific language models on materials science language tasks using text-to-schema modeling.arXiv preprint arXiv:2305.08264, 2023
2023 arXiv
-
[157]
Llm4mat-bench: Benchmarking large language models for materials property prediction.Machine Learning: Science and Technology, 2024
Andre Niyongabo Rubungo, Kangming Li, Jason Hattrick-Simpers, and Adji Bousso Dieng. Llm4mat-bench: Benchmarking large language models for materials property prediction.Machine Learning: Science and Technology, 2024
2024
-
[158]
Benchmarking large language models for molecule prediction tasks.arXiv preprint arXiv:2403.05075, 2024
Zhiqiang Zhong, Kuangyu Zhou, and Davide Mottin. Benchmarking large language models for molecule prediction tasks.arXiv preprint arXiv:2403.05075, 2024
2024 arXiv
-
[159]
MaCBench: A multimodal chemistry and materi- als science benchmark
Nawaf Alampara, Indrajeet Mandal, Pranav Khetarpal, Hargun Singh Grover, Mara Schilling-Wilhelmi, N M Anoop Krishnan, and Kevin Maik Jablonka. MaCBench: A multimodal chemistry and materi- als science benchmark. InAI for Accelerated Materials Design - NeurIPS 2024, 2024. URL ht...
2024
-
[160]
Redpajama: an open dataset for training large language models.Advances in neural information processing systems, 37:116462–116492, 2024
Maurice Weber, Dan Fu, Quentin Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xiaozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, et al. Redpajama: an open dataset for training large language models.Advances in neural information processing systems, 37:116462–116492, 2024
2024
-
[161]
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE international conference on computer ...
2015
-
[162]
Pointer sentinel mixture models.arXiv preprint arXiv:1609.07843, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models.arXiv preprint arXiv:1609.07843, 2016
2016 arXiv
-
[163]
Common crawl dataset, 2025
Common Crawl Team. Common crawl dataset, 2025. URLhttps://commoncrawl.org/
2025
-
[164]
Specter: Document-level representation learning using citation-informed transformers.arXiv preprint arXiv:2004.07180, 2020
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S Weld. Specter: Document-level representation learning using citation-informed transformers.arXiv preprint arXiv:2004.07180, 2020
2004 arXiv
-
[165]
Scicap: Generating captions for scientific figures
Ting-Yao Hsu, C Lee Giles, and Ting-Hao’Kenneth’ Huang. Scicap: Generating captions for scientific figures. arXiv preprint arXiv:2110.11624, 2021
2021 arXiv
-
[166]
Orca: Progressive learning from complex explanation traces of gpt-4.arXiv preprint arXiv:2306.02707, 2023
Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Awadallah. Orca: Progressive learning from complex explanation traces of gpt-4.arXiv preprint arXiv:2306.02707, 2023
2023 arXiv
-
[167]
Mammoth2: Scaling instructions from the web
Xiang Yue, Tianyu Zheng, Ge Zhang, and Wenhu Chen. Mammoth2: Scaling instructions from the web. Advances in Neural Information Processing Systems, 37:90629–90660, 2024
2024
-
[168]
Mathqa: Towards interpretable math word problem solving with operation-based formalisms.arXiv preprint arXiv:1905.13319, 2019
Aida Amini, Saadia Gabriel, Peter Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. Mathqa: Towards interpretable math word problem solving with operation-based formalisms.arXiv preprint arXiv:1905.13319, 2019
1905 arXiv
-
[169]
Think you have solved question answering? try arc, the ai2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018
2018 arXiv
-
[170]
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical commonsense in natural language. InProceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432–7439, 2020
2020
-
[171]
Crowdsourcing multiple choice science questions.arXiv preprint arXiv:1707.06209, 2017
Johannes Welbl, Nelson F Liu, and Matt Gardner. Crowdsourcing multiple choice science questions.arXiv preprint arXiv:1707.06209, 2017
2017 arXiv
-
[172]
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35:2507–2521, 2022
2022
-
[173]
Optimade, an api for exchanging materials data.Scientific data, 8(1):217, 2021
Casper W Andersen, Rickard Armiento, Evgeny Blokhin, Gareth J Conduit, Shyam Dwaraknath, Matthew L Evans, Ádám Fekete, Abhijith Gopakumar, Saulius Gražulis, Andrius Merkys, et al. Optimade, an api for exchanging materials data.Scientific data, 8(1):217, 2021
2021
-
[174]
Python materials genomics (pymatgen): A robust, open-source python library for materials analysis.Computational Materials Science, 68:314–319, 2013
Shyue Ping Ong, William Davidson Richards, Anubhav Jain, Geoffroy Hautier, Michael Kocher, Shreyas Cholia, Dan Gunter, Vincent L Chevrier, Kristin A Persson, and Gerbrand Ceder. Python materials genomics (pymatgen): A robust, open-source python library for materials analysis.C...
2013
-
[175]
M 2hub: Unlocking the potential of machine learning for materials discovery
Yuanqi Du, Yingheng Wang, Yining Huang, Jianan Canal Li, Yanqiao Zhu, Tian Xie, Chenru Duan, John Gregoire, and Carla P Gomes. M 2hub: Unlocking the potential of machine learning for materials discovery. Advances in Neural Information Processing Systems, 36:77359–77378, 2023
2023
-
[176]
Fireworks: a dynamic workflow system designed for high-throughput applications.Concurrency and Computation: Practice and Experience, 27(17):5037–5059, 2015
Anubhav Jain, Shyue Ping Ong, Wei Chen, Bharat Medasani, Xiaohui Qu, Michael Kocher, Miriam Brafman, Guido Petretto, Gian-Marco Rignanese, Geoffroy Hautier, et al. Fireworks: a dynamic workflow system designed for high-throughput applications.Concurrency and Computation: Pract...
2015
-
[177]
Maggma Toolkit.https://materialsproject.github.io/maggma/, 2025
The Materials Project. Maggma Toolkit.https://materialsproject.github.io/maggma/, 2025
2025
-
[178]
Benchmarking materials property prediction methods: the matbench test set and automatminer reference algorithm.npj Computational Materials, 6(1):138, 2020
Alexander Dunn, Qi Wang, Alex Ganose, Daniel Dopp, and Anubhav Jain. Benchmarking materials property prediction methods: the matbench test set and automatminer reference algorithm.npj Computational Materials, 6(1):138, 2020
2020
-
[179]
Unified graph neural network force-field for the periodic table: solid state applications.Digital Discovery, 2(2): 346–355, 2023
Kamal Choudhary, Brian DeCost, Lily Major, Keith Butler, Jeyan Thiyagalingam, and Francesca Tavazza. Unified graph neural network force-field for the periodic table: solid state applications.Digital Discovery, 2(2): 346–355, 2023. 30 Survey of AI for MSA PREPRINT
2023
-
[180]
Ollama tool.https://ollama.com/, 2025
Ollama. Ollama tool.https://ollama.com/, 2025
2025
-
[181]
Langchain tool.https://github.com/langchain-ai/langchain, 2025
LangChain. Langchain tool.https://github.com/langchain-ai/langchain, 2025
2025
-
[182]
Crystal toolkit: A web app framework to improve usability and accessibility of materials science research algorithms.arXiv preprint arXiv:2302.06147, 2023
Matthew Horton, Jimmy-Xuan Shen, Jordan Burns, Orion Cohen, François Chabbey, Alex M Ganose, Rishabh Guha, Patrick Huck, Hamming Howard Li, Matthew McDermott, et al. Crystal toolkit: A web app framework to improve usability and accessibility of materials science research algor...
2023 arXiv
-
[183]
Atomate: A high-level interface to generate, execute, and analyze computational materials science workflows.Computational Materials Science, 139:140–152, 2017
Kiran Mathew, Joseph H Montoya, Alireza Faghaninia, Shyam Dwarakanath, Muratahan Aykol, Hanmei Tang, Iek-heng Chu, Tess Smidt, Brandon Bocklund, Matthew Horton, et al. Atomate: A high-level interface to generate, execute, and analyze computational materials science workflows.C...
2017
-
[184]
Jobflow: Computational workflows made simple.Journal of Open Source Software, 9(93):5995, 2024
Andrew S Rosen, Max Gallant, Janine George, Janosh Riebesell, Hrushikesh Sahasrabuddhe, Jimmy-Xuan Shen, Mingjian Wen, Matthew L Evans, Guido Petretto, David Waroquiers, et al. Jobflow: Computational workflows made simple.Journal of Open Source Software, 9(93):5995, 2024
2024
-
[185]
Emmet Toolkit.https://materialsproject.github.io/emmet/, 2025
The Materials Project. Emmet Toolkit.https://materialsproject.github.io/emmet/, 2025
2025
-
[186]
Autogen tool.https://github.com/microsoft/autogen, 2025
Microsoft. Autogen tool.https://github.com/microsoft/autogen, 2025
2025
-
[187]
Crewai tool.https://github.com/crewAIInc/crewAI, 2025
CrewAI. Crewai tool.https://github.com/crewAIInc/crewAI, 2025
2025
-
[188]
Llamaindex tool.https://github.com/run-llama/llama_index, 2025
LlamaIndex. Llamaindex tool.https://github.com/run-llama/llama_index, 2025. 31
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.