REVIEW 2 major objections 5 minor 255 references
Language Models for Materials Discovery and Sustainability: Progress, Challenges, and Opportunities
T0 review · 2 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Large language models can guide materials discovery and sustainable alloy design by turning corpus statistics into candidate rankings.
desk verdict Useful but imperfect review: broad, well-cited coverage; overclaimed uniqueness and an unvalidated sustainability-screening proposal. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the word-embedding vector space: tokens representing elements, alloys, and concepts become high-dimensional vectors whose cosine similarity encodes how often and how they are discussed together in the scientific corpus. That similarity is used three ways: to rank known materials for new applications, to pick element combinations for alloys not yet reported, and, in the paper's sustainability proposal, to rank materials against sustainability keywords. Around this core sit retrieval-augmented generation, which grounds model answers in external documents; knowledge graphs, which standardize entity names so different strings for the same alloy resolve identically; and fine-tuned language models for property regression.
What would settle it
A direct test would be to compute cosine similarity to sustainability keywords from a materials corpus and compare the top-ranked candidates against life-cycle assessment data; if a substantial fraction of the top-ranked materials shows high CO2 footprint or poor recyclability, the screening premise fails.
Extended reading notes
Core claim
On its own terms, the paper's contribution is a synthesis: it identifies a pipeline in which language models act at four levels, from retrieving and summarizing literature, to providing word-vector features for regression models, to generating hypotheses and validation plans, and finally to autonomous execution through AI agents. The load-bearing evidence it assembles is that unsupervised word embeddings trained on materials abstracts can rank candidate elements and alloys by context similarity, that fine-tuned language models can extract structured data and predict properties with accuracy comparable to dedicated machine-learning models, and that domain knowledge graphs can standardize alloy naming for reliable retrieval. The paper then extends this machinery to sustainability, proposing that materials be screened by cosine similarity to keywords such as 'sustainability', 'recycling', and 'CO2 footprint', either directly or as features in a second model.
Load-bearing premise
The proposal that sustainability can be screened by cosine similarity to keywords like 'sustainability' and 'CO2 footprint' assumes that the word-embedding spaces of scientific corpora encode reliable environmental and recyclability semantics, and the paper offers no validation of that specific mapping.
Editorial extensions
If this is right
- If context-similarity design holds, materials discovery can start from corpus statistics before any simulation, narrowing candidates like the reported search of 2.6 million alloys to a few hundred for further screening.
- If fine-tuned language models match dedicated machine-learning models on property prediction, small labeled datasets become less of a bottleneck because foundation knowledge transfers.
- If sustainability keywords encode environmental semantics, a multi-step cosine-similarity filter can be added to any word-embedding design loop to bias toward recyclable, low-footprint materials.
- If language-model agents with retrieval augmentation and tool use mature, the five-step discovery loop from requirements through literature check, candidate generation, synthesis, and property measurement can be automated with humans writing only the prompts.
- Knowledge-graph standardization implies that non-standard alloy names cease to be a retrieval barrier, so any permutation of a composition returns the same publications.
Reading between the lines
- The sustainability-by-cosine-similarity proposal is testable now: one could construct a small benchmark corpus with human-annotated environmental impact of alloys and check whether the top-ranked cosine neighbors are genuinely greener.
- Word-embedding proximity can be brittle to corpus bias: a material frequently discussed alongside 'CO2' in a synthesis context may be a producer rather than a low-footprint option, so the direction of association needs disambiguation.
- The four-level language-model application ladder suggests a natural evaluation metric: measure how far each level can proceed without human intervention, which would quantify progress toward autonomous discovery.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a review/perspective on the use of natural language processing (NLP) and large language models (LLMs) for materials science, with an emphasis on materials discovery and sustainability. It covers the fundamentals of word embeddings, transformers, and LLMs; surveys applications such as information extraction, prompt engineering, knowledge graphs, and word-embedding-based design; and discusses specific domains including structural materials, inorganic and organic materials, and additive manufacturing. A substantial portion is devoted to sustainability, proposing that cosine similarity between material names and sustainability keywords can screen for sustainable materials, and to forward-looking topics such as autonomous AI agents, small language models, and quantum computing for LLMs. The paper is framed as a timely and unique review, and its conclusions call for combining language models with ICME and other computational methods to accelerate materials design.
Significance. If the paper were positioned more carefully, it would be a useful, up-to-date map of a fast-moving field. Its strengths include broad coverage of recent literature, a helpful taxonomy of NLP tasks in materials science (Figure 10), and explicit flagging of unconfirmed claims, such as the 72-qubit LLM fine-tuning result in Section 7.6. However, the significance is undercut by two load-bearing problems: the claimed uniqueness of the review is contradicted by the paper's own cited prior reviews, and the sustainability-screening proposal rests on an unvalidated assumption about the semantics of word embeddings. Because the title and abstract advertise 'sustainability' as a central contribution, the unsupported nature of that proposal is a substantive concern rather than a mere presentation issue. With revision, the manuscript could serve as a balanced perspective; in its current form, its main claims are overstated.
major comments (2)
- [Section 4.4 and Section 7.5] The sustainability screening pipeline is presented as a feasible method ('we can train an NLP model for sustainability... The model can determine a list of materials based on their cosine similarity with keywords like "sustainability", "recycling", and "CO2 footprint"') without any validation, benchmark, or error analysis. The only cited support, Ref. [46], demonstrates that word-embedding context similarity can identify chemically interchangeable elements from co-occurrence in 6.4 million abstracts; it does not establish that cosine similarity between material names and abstract sustainability keywords encodes recyclability, carbon footprint, or other environmental-impact semantics. General scientific corpora frequently use 'sustainability' in policy, economic, and social contexts, so high cosine similarity to that keyword may reflect co-occurrence patterns rather than material sustainability. Please either provide preliminary evidence (for example, a case study ranking known sustainable versus known unsustainable materials) or explicitly reframe this as an open research hypothesis with a discussion of known risks. Given that the title and abstract advertise sustainability, this unsupported assertion is load-bearing for the paper's central contribution.
- [Section 1, 'Uniqueness of this review'] The claim that 'we have not noticed review articles on the same topic' is contradicted by the paper's own reference list. Ref. [74] (Olivetti et al., Applied Physics Reviews, 2020) is a review of data-driven materials research enabled by NLP and information extraction; Ref. [75] (Kononova et al., iScience, 2021) reviews opportunities and challenges of text mining in materials research; Ref. [73] (Smith et al., Chemistry of Materials, 2022) addresses challenges in information-mining the materials literature; and Refs. [175] (Yu et al., 2024) and [247] (Lei et al., 2024) are recent reviews/perspectives on large language models in materials science. Please substantiate the specific novelty (for example, the sustainability angle or the emphasis on metallic materials) or remove the blanket uniqueness claim, which is factually inaccurate as stated.
minor comments (5)
- [Section 3.1] 'Exacted information' should be 'Extracted information'.
- [Section 2.1] 'Transformed has been used in foundational models' is a typographical error; it should read 'Transformers have been used in foundational models.'
- [Figure 2 caption] The caption labels ViT as an example of 'Encoder-Decoder'; ViT (Vision Transformer) is an encoder-only architecture. Please correct this classification.
- [Section 4.4 and throughout] The rendering of carbon dioxide is inconsistent ('CO2' in some places and 'CO 2' in others). Please use a consistent subscript notation.
- [Table 3] The entry for DeepSeek describes the model as 'dedicated for multimodal and coding'; this phrasing is awkward, and the vendor name appears as both 'Deepseek' and 'DeepSeek'. Please use a consistent name and clearer phrasing.
Circularity Check
Sustainability screening via keyword cosine similarity is self-definitional: the model scores materials by the same sustainability vocabulary used to upweight its training corpus.
-
self definitional
[Section 4.4, 'NLP for sustainability in materials design' (extended in Section 7.5)]
"Since the sustainability information of materials can be extracted from scientific corpora, we can train an NLP model for sustainability by giving papers on sustainability and environmental materials a higher weight. The model can determine a list of materials based on their cosine similarity with keywords like “sustainability”, “recycling”, and “CO2 footprint”."
The training signal and the scoring variable are the same semantic construct. Upweighting papers flagged as sustainability-related biases the embedding so terms co-occurring with 'sustainability' in those papers move close to the keyword; ranking by cosine similarity to 'sustainability', 'recycling', and 'CO2 footprint' then recovers that same co-occurrence pattern. No independent property (emissions, recyclability, toxicity, lifetime) is measured. The paper nevertheless calls this screening for sustainable materials and says it 'can realize the target of designing sustainable materials' (Sec. 4.4), so the proposed model's output is its own keyword axis. The analogy to Ref. [46] does not help, as that work validated element-context similarity, not environmental semantics.
full rationale
The review is essentially a survey, and most of its content is independent of the authors' own prior work. The materials-discovery examples are drawn from published external studies (Tshitoyan et al., Jablonka et al., AlphaFold, CrystaLLM, etc.) and are not derived from the present paper's assumptions, so no circularity is present there. References [46] and [181], which share authors with this review, are used frequently as illustrative examples, but they are not invoked as uniqueness theorems or as the sole support for a derived result; they summarize previously published, externally validated findings, so the self-citation by itself is not load-bearing. The single construction-level problem is the sustainability screening proposal in Section 4.4 (extended in Section 7.5). The paper proposes to train an NLP model by upweighting papers on sustainability and then to rank materials by cosine similarity to 'sustainability', 'recycling', and 'CO2 footprint'. Because those keywords are the same semantic axis that defines the upweighted training selection, the resulting material ranking is determined by the input construct rather than by any independent measure of environmental impact. The text then treats this ranking as a route to 'designing sustainable materials' and as a 'sustainability model'. That step is partially circular; it does not, however, infect the rest of the review, which contains no other derivation whose conclusion equals its premises. The score is therefore 6 rather than higher.
Assumptions & free parameters
assumptions (2)
- domain assumption Cosine similarity in word-embedding space is a valid proxy for materials-science relevance and sustainability semantics.
- domain assumption Current LLMs can serve as reliable control centers for automated materials discovery agents.
Cite this review
Pith. "Pith review of Language Models for Materials Discovery and Sustainability: Progress, Challenges, and Opportunities." pith.science (2026). https://pith.science/paper/ONMSUSPY
@misc{pith2026250414849,
author = {Pith},
title = {Pith review of: Language Models for Materials Discovery and Sustainability: Progress, Challenges, and Opportunities},
year = {2026},
howpublished = {\url{https://pith.science/paper/ONMSUSPY}},
note = {Machine review of arXiv:2504.14849}
}
read the original abstract
Significant advancements have been made in one of the most critical branches of artificial intelligence: natural language processing (NLP). These advancements are exemplified by the remarkable success of OpenAI's GPT-3.5/4 and the recent release of GPT-4.5, which have sparked a global surge of interest akin to an NLP gold rush. In this article, we offer our perspective on the development and application of NLP and large language models (LLMs) in materials science. We begin by presenting an overview of recent advancements in NLP within the broader scientific landscape, with a particular focus on their relevance to materials science. Next, we examine how NLP can facilitate the understanding and design of novel materials and its potential integration with other methodologies. To highlight key challenges and opportunities, we delve into three specific topics: (i) the limitations of LLMs and their implications for materials science applications, (ii) the creation of a fully automated materials discovery pipeline, and (iii) the potential of GPT-like tools to synthesize existing knowledge and aid in the design of sustainable materials.
Figures
Figures from the paper (19 more)
Reference graph
Works this paper leans on
-
[46]
Pei, Z., Yin, J., Liaw, P. K. & Raabe, D. Toward the design of ultrahigh-entropy alloys via mining six million texts.Nature Communications14, 54 (2023)
2023
-
[74]
A.et al.Data-driven materials research enabled by natural language process- ing and information extraction.Applied Physics Reviews7, 041317 (2020)
Olivetti, E. A.et al.Data-driven materials research enabled by natural language process- ing and information extraction.Applied Physics Reviews7, 041317 (2020)
2020
-
[75]
Iscience24, 102155 (2021)
Kononova, O.et al.Opportunities and challenges of text mining in materials research. Iscience24, 102155 (2021)
2021
-
[73]
& Risko, C
Smith, A., Bhat, V ., Ai, Q. & Risko, C. Challenges in information-mining the materials literature: A case study and perspective.Chemistry of Materials34, 4821–4827 (2022)
2022
-
[175]
& Liu, J
Yu, S., Ran, N. & Liu, J. Large-language models: The game-changers for materials science research.Artificial Intelligence Chemistry2, 100076 (2024)
2024
-
[247]
Lei, G., Docherty, R. & Cooper, S. J. Materials science in the era of large language models: a perspective.Digital Discovery3, 1257–1272 (2024). 87
work page 2024
-
[1]
& Schutze, H.Foundations of statistical natural language processing(MIT press, 1999)
Manning, C. & Schutze, H.Foundations of statistical natural language processing(MIT press, 1999)
1999
-
[2]
& Damerau, F
Indurkhya, N. & Damerau, F. J.Handbook of natural language processing(Chapman and Hall/CRC, 2010)
2010
Show all 255 references
-
[3]
& Manning, C
Hirschberg, J. & Manning, C. D. Advances in natural language processing.Science349, 261–266 (2015)
2015
-
[4]
& Chowdhary, K
Chowdhary, K. & Chowdhary, K. Natural language processing.Fundamentals of artifi- cial intelligence603–649 (2020)
2020
-
[5]
& Hinton, G
LeCun, Y ., Bengio, Y . & Hinton, G. Deep learning.nature521, 436–444 (2015)
2015
-
[6]
Deep learning in neural networks: An overview.Neural networks61, 85–117 (2015)
Schmidhuber, J. Deep learning in neural networks: An overview.Neural networks61, 85–117 (2015)
2015
-
[7]
& Hinton, G
Krizhevsky, A., Sutskever, I. & Hinton, G. E. Imagenet classification with deep convo- lutional neural networks.Communications of the ACM60, 84–90 (2017)
2017
-
[8]
URLhttps://doi.org/10
Tshitoyan, V .et al.Unsupervised word embeddings capture latent knowledge from ma- terials science literature.Nature571, 95–98 (2019). URLhttps://doi.org/10. 1038/s41586-019-1335-8. 63
2019
-
[9]
& Pan, F
Nie, Z., Liu, Y ., Yang, L., Li, S. & Pan, F. Construction and application of materials knowledge graph based on author disambiguation: Revisiting the evolution of lifepo4. Advanced Energy Materials2003580 (2021)
2021
-
[10]
& Ginebra, M.-P
Hakimi, O., Krallinger, M. & Ginebra, M.-P. Time to kick-start text mining for biomate- rials.Nature Reviews Materials5, 553–556 (2020)
2020
-
[11]
Court, C. J. & Cole, J. M. Magnetic and superconducting phase diagrams and transi- tion temperatures predicted using text mining and machine learning.npj Computational Materials6, 1–9 (2020)
2020
-
[12]
& Zeilinger, A
Krenn, M. & Zeilinger, A. Predicting research trends with semantic and neural net- works with an application in quantum physics.Proceedings of the National Academy of Sciences(2020). URLhttps://www.pnas.org/content/early/2020/01/ 13/1914370116.https://www.pnas.org/content/earl...
2020
-
[13]
& Stewart, B
Grimmer, J. & Stewart, B. M. Text as data: The promise and pitfalls of automatic content analysis methods for political texts.Political analysis21, 267–297 (2013)
2013
-
[14]
& Ausloos, M
Ficcadenti, V ., Cerqueti, R. & Ausloos, M. A joint text mining-rank size investigation of the rhetoric structures of the us presidents’ speeches.Expert Systems with Applications 123, 127–142 (2019)
2019
-
[15]
Krallinger, M., Erhardt, R. A.-A. & Valencia, A. Text-mining approaches in molecular biology and biomedicine.Drug discovery today10, 439–445 (2005)
2005
-
[16]
Zheng, S., Dharssi, S., Wu, M., Li, J. & Lu, Z. Text mining for drug discovery.Bioinfor- matics and Drug Discovery231–252 (2019)
2019
-
[17]
Hirschman, L.et al.Text mining for the biocuration workflow.Database2012(2012). 64
2012
-
[18]
Wang, X.et al.Cross-type biomedical named entity recognition with deep multi-task learning.Bioinformatics35, 1745–1752 (2019)
2019
-
[19]
J.et al.Large language models in medicine.Nature medicine29, 1930–1940 (2023)
Thirunavukarasu, A. J.et al.Large language models in medicine.Nature medicine29, 1930–1940 (2023)
2023
-
[20]
Birgmeier, J.et al.Amelie speeds mendelian diagnosis by matching patient phenotype and genotype to primary literature.Science translational medicine12(2020)
2020
-
[21]
URLhttps://www.science
Hoffmann, R.et al.Text mining for metabolic pathways, signaling cascades, and protein networks.Science’s STKE2005, pe21–pe21 (2005). URLhttps://www.science. org/doi/abs/10.1126/stke.2832005pe21.https://www.science. org/doi/pdf/10.1126/stke.2832005pe21
2005 doi
-
[22]
& Liao, S
Cheng, X., Cao, Q. & Liao, S. S. An overview of literature on covid-19, mers and sars: Using text mining and latent dirichlet allocation.Journal of Information Science(2020)
2020
-
[23]
& Hope, T
Mani, G. & Hope, T. Viral science: Masks, speed bumps, and guard rails.Patterns1, 100101 (2020)
2020
-
[24]
L.et al.Cord-19: The covid-19 open research dataset.ArXiv(2020)
Wang, L. L.et al.Cord-19: The covid-19 open research dataset.ArXiv(2020)
2020
-
[25]
Wang, L. L. & Lo, K. Text mining approaches for dealing with the rapidly expanding literature on covid-19.Briefings in Bioinformatics22, 781–799 (2021)
2021
-
[26]
E.et al.Language models for the prediction of sars-cov-2 inhibitors.The International Journal of High Performance Computing Applications36, 587–602 (2022)
Blanchard, A. E.et al.Language models for the prediction of sars-cov-2 inhibitors.The International Journal of High Performance Computing Applications36, 587–602 (2022)
2022
-
[27]
White, A. D. The future of chemistry is language.Nature Reviews Chemistry1–2 (2023)
2023
-
[28]
D.et al.Assessment of chemistry knowledge in large language models that generate code.Digital Discovery2, 368–376 (2023)
White, A. D.et al.Assessment of chemistry knowledge in large language models that generate code.Digital Discovery2, 368–376 (2023). 65
2023
-
[29]
van de Schoot, R.et al.An open source machine learning framework for efficient and transparent systematic reviews.Nature Machine Intelligence3, 125–133 (2021)
2021
-
[30]
& Dean, J
Mikolov, T., Chen, K., Corrado, G. & Dean, J. Efficient estimation of word representa- tions in vector space (2013).1301.3781
2013 arXiv
-
[31]
& Dean, J
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. & Dean, J. Distributed representations of words and phrases and their compositionality (2013).1310.4546
2013 arXiv
-
[32]
& Manning, C
Pennington, J., Socher, R. & Manning, C. D. Glove: Global vectors for word represen- tation. InEmpirical Methods in Natural Language Processing (EMNLP), 1532–1543 (2014). URLhttp://www.aclweb.org/anthology/D14-1162
2014
-
[33]
& Toutanova, K
Devlin, J., Chang, M., Lee, K. & Toutanova, K. BERT: pre-training of deep bidirectional transformers for language understanding.CoRRabs/1810.04805(2018). URLhttp: //arxiv.org/abs/1810.04805.1810.04805
2018 arXiv
-
[34]
A., Pereira, F
Grand, G., Blank, I. A., Pereira, F. & Fedorenko, E. Semantic projection recovers rich human knowledge of multiple object features from word embeddings.Nature human behaviour6, 975–987 (2022)
2022
-
[35]
Zhang, Y ., Chen, Q., Yang, Z., Lin, H. & Lu, Z. Biowordvec, improving biomedical word embeddings with subword information and mesh.Scientific data6, 52 (2019)
2019
-
[36]
& Wachter, S
Birhane, A., Kasirzadeh, A., Leslie, D. & Wachter, S. Science in the age of large language models.Nature Reviews Physics5, 277–280 (2023)
2023
-
[37]
Accessed: 2023- 04-14
Introducing ChatGPT.https://openai.com/blog/chatgpt. Accessed: 2023- 04-14
2023
-
[38]
Touvron, H.et al.Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971(2023). 66
2023 arXiv
-
[39]
Swain, M. C. & Cole, J. M. Chemdataextractor: a toolkit for automated extraction of chemical information from the scientific literature.Journal of chemical information and modeling56, 1894–1904 (2016)
2016
-
[40]
Court, C. J. & Cole, J. M. Magnetic and superconducting phase diagrams and transition temperatures predicted using text mining and machine learning.npj Computational Ma- terials6, 18 (2020). URLhttps://doi.org/10.1038/s41524-020-0287-8
2020 doi
-
[41]
URLhttps://doi.org/10.1021/acs
Weston, L.et al.Named entity recognition and normalization applied to large-scale infor- mation extraction from the materials science literature.Journal of Chemical Information and Modeling59, 3692–3702 (2019). URLhttps://doi.org/10.1021/acs. jcim.9b00470. PMID: 31361962,https...
2019 doi
-
[42]
Kim, E.et al.Inorganic materials synthesis planning with literature-trained neural net- works.Journal of chemical information and modeling60, 1194–1201 (2020)
2020
-
[43]
K.et al.Semantic text mining in early drug discovery for type 2 diabetes
Hansson, L. K.et al.Semantic text mining in early drug discovery for type 2 diabetes. Plos one15, e0233956 (2020)
2020
-
[44]
Trewartha, A.et al.Quantifying the advantage of domain-specific pre-training on named entity recognition tasks in materials science.Patterns3, 100488 (2022)
2022
-
[45]
Mehr, S. H. M., Craven, M., Leonov, A. I., Keenan, G. & Cronin, L. A universal system for digitization and automatic execution of the chemical synthesis literature.Science370, 101–108 (2020)
2020
-
[47]
& Pei, Z
Liu, X., Zhang, J. & Pei, Z. Machine learning for high-entropy alloys: Progress, chal- lenges and opportunities.Progress in Materials Science101018 (2022). 67
2022
-
[48]
Yeh, J.-W.et al.Nanostructured high-entropy alloys with multiple principal elements: novel alloy design concepts and outcomes.Advanced Engineering Materials6, 299–303 (2004)
2004
-
[49]
& Vincent, A
Cantor, B., Chang, I., Knight, P. & Vincent, A. Microstructural development in equiatomic multicomponent alloys.Materials Science and Engineering: A375, 213– 218 (2004)
2004
-
[50]
Zhang, Y .et al.Microstructures and properties of high-entropy alloys.Progress in materials science61, 1–93 (2014)
2014
-
[51]
Miracle, D. B. & Senkov, O. N. A critical review of high entropy alloys and related concepts.Acta Materialia122, 448–511 (2017)
2017
-
[52]
P., Raabe, D
George, E. P., Raabe, D. & Ritchie, R. O. High-entropy alloys.Nature Reviews Materials 4, 515–534 (2019)
2019
-
[53]
Shi, P.et al.Hierarchical crack buffering triples ductility in eutectic herringbone high- entropy alloys.Science373, 912–918 (2021)
2021
-
[54]
Huang, E.-W.et al.Machine-learning and high-throughput studies for high-entropy ma- terials.Materials Science and Engineering: R: Reports147, 100645 (2022)
2022
-
[55]
Multicomponent high-entropy cantor alloys.Progress in Materials Science 100754 (2020)
Cantor, B. Multicomponent high-entropy cantor alloys.Progress in Materials Science 100754 (2020)
2020
-
[56]
Eswarappa Prameela, S.et al.Materials for extreme environments.Nature Reviews Materials8, 81–88 (2023)
2023
-
[57]
Raabe, D., Tasan, C. C. & Olivetti, E. A. Strategies for improving the sustainability of structural metals.Nature575, 64–74 (2019)
2019
-
[58]
The materials science behind sustainable metals and alloys.Chemical Reviews (2023)
Raabe, D. The materials science behind sustainable metals and alloys.Chemical Reviews (2023). 68
2023
-
[59]
Raabe, D., Mianroodi, J. R. & Neugebauer, J. Accelerating the design of compositionally complex materials via physics-informed artificial intelligence.Nature Computational Science3, 198–209 (2023)
2023
-
[60]
Rao, Z.et al.Machine learning–enabled high-entropy alloy discovery.Science378, 78–85 (2022)
2022
-
[61]
Scientific reports7, 10458 (2017)
Sandlöbes, S.et al.A rare-earth free magnesium alloy with improved intrinsic ductility. Scientific reports7, 10458 (2017)
2017
-
[62]
A., Alman, D
Pei, Z., Yin, J., Hawk, J. A., Alman, D. E. & Gao, M. C. Machine-learning informed prediction of high-entropy solid solution formation: Beyond the hume-rothery rules.npj Computational Materials6, 1–8 (2020)
2020
-
[63]
Feng, R.et al.High-throughput design of high-performance lightweight high-entropy alloys.Nature Communications12, 1–10 (2021)
2021
-
[64]
Pei, Z.et al.Rapid theory-guided prototyping of ductile mg alloys: from binary to multi- component materials.New Journal of Physics17, 093009 (2015)
2015
-
[65]
& Yin, J
Pei, Z. & Yin, J. Machine learning as a contributor to physics: Understanding mg alloys. Materials & Design172, 107759 (2019)
2019
-
[66]
& Yin, J
Pei, Z. & Yin, J. The relation between two ductility mechanisms for mg alloys revealed by high-throughput simulations.Materials & Design186, 108286 (2020)
2020
-
[67]
Pei, Z.et al.Machine-learning microstructure for inverse material design.Advanced Science8, 2101207 (2021)
2021
-
[68]
Mechanisms and machine learning for magnesium alloys design
Pei, Z. Mechanisms and machine learning for magnesium alloys design. InMagnesium Technology 2021, 61–66 (Springer, 2021)
2021
-
[69]
Li, Y .et al.Machine learning-enabled tomographic imaging of chemical short-range atomic ordering.Advanced materials36, 2407564 (2024). 69
2024
-
[70]
Joshi, A. K. Natural language processing.Science253, 1242–1249 (1991)
1991
-
[71]
E., Cui, H
Thessen, A. E., Cui, H. & Mozzherin, D. Applications of natural language processing in biodiversity science.Advances in bioinformatics2012(2012)
2012
-
[72]
Unsal, S.et al.Learning functional properties of proteins with language models.Nature Machine Intelligence4, 227–245 (2022)
2022
-
[76]
Jones, K. S. Natural language processing: a historical review.Current issues in compu- tational linguistics: in honour of Don Walker3–16 (1994)
1994
-
[77]
Deep learning for natural language processing: advantages and challenges.Na- tional Science Review5, 24–26 (2018)
Li, H. Deep learning for natural language processing: advantages and challenges.Na- tional Science Review5, 24–26 (2018)
2018
-
[78]
Jensen, Z.et al.A machine learning approach to zeolite synthesis enabled by automatic literature data extraction.ACS central science5, 892–899 (2019)
2019
-
[79]
Kim, E.et al.Materials synthesis insights from scientific literature via text extraction and machine learning.Chemistry of Materials29, 9436–9444 (2017)
2017
-
[80]
& Ling, C
Huang, L. & Ling, C. Representing multiword chemical terms through phrase-level preprocessing and word embedding.ACS omega4, 18510–18519 (2019)
2019
-
[81]
& Krishnan, N
Gupta, T., Zaki, M. & Krishnan, N. A. Matscibert: A materials domain language model for text mining and information extraction.npj Computational Materials8, 102 (2022). 70
2022
-
[82]
Hakimi, O.et al.The devices, experimental scaffolds, and biomaterials ontology (deb): a tool for mapping, annotation, and analysis of biomaterials data.Advanced Functional Materials30, 1909910 (2020)
2020
-
[83]
Cruse, K.et al.Text-mined dataset of gold nanoparticle synthesis procedures, morpholo- gies, and size entities.Scientific Data9, 234 (2022)
2022
-
[84]
& Zhang, L
He, M. & Zhang, L. Prediction of solar-chargeable battery materials: A text-mining and first-principles investigation.International Journal of Energy Research45, 15521–15533 (2021)
2021
-
[85]
L., Guo, Y
Porter, A. L., Guo, Y . & Chiavatta, D. Tech mining: Text mining and visualization tools, as applied to nanoenhanced solar cells.Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery1, 172–181 (2011)
2011
-
[86]
& Huang, L
Li, X., Xie, Q., Daim, T. & Huang, L. Forecasting technology trends using text mining of the gaps between science and technology: The case of perovskite solar cell technology. Technological Forecasting and Social Change146, 432–449 (2019)
2019
-
[87]
Park, S.et al.Text mining metal–organic framework papers.Journal of chemical infor- mation and modeling58, 244–251 (2018)
2018
-
[88]
El-Bousiydy, H.et al.What can text mining tell us about lithium-ion battery researchers’ habits?Batteries & Supercaps4, 758–766 (2021)
2021
-
[89]
J., Jain, A
Court, C. J., Jain, A. & Cole, J. M. Inverse design of materials that exhibit the magne- tocaloric effect by text-mining of the scientific literature and generative deep learning. Chemistry of Materials33, 7217–7231 (2021)
2021
-
[90]
& Altman, R
Percha, B., Garten, Y . & Altman, R. B. Discovery and explanation of drug-drug interac- tions via text mining. InBiocomputing 2012, 410–421 (World Scientific, 2012). 71
2012
-
[91]
Collier, N.et al.Biocaster: detecting public health rumors with a web-based text mining system.Bioinformatics24, 2940–2941 (2008)
2008
-
[92]
& Collier, N
Conway, M., Doan, S., Kawazoe, A. & Collier, N. Classifying disease outbreak reports using n-grams and semantic features.International journal of medical informatics78, e47–e58 (2009)
2009
-
[93]
Kloptchenko, A.et al.Combining data and text mining techniques for analysing finan- cial reports.Intelligent Systems in Accounting, Finance & Management: International Journal12, 29–41 (2004)
2004
-
[94]
& Chung, K
Kim, J.-C. & Chung, K. Associative feature information extraction using text mining from health big data.Wireless Personal Communications105, 691–707 (2019)
2019
-
[95]
& Zhang, Y
Lu, J. & Zhang, Y . Unified deep learning model for multitask reaction predictions with explanation.Journal of Chemical Information and Modeling62, 1376–1387 (2022)
2022
-
[96]
Raffel, C.et al.Exploring the limits of transfer learning with a unified text-to-text trans- former.The Journal of Machine Learning Research21, 5485–5551 (2020)
2020
-
[97]
Jumper, J.et al.Highly accurate protein structure prediction with alphafold.Nature596, 583–589 (2021)
2021
-
[98]
Baek, M.et al.Accurate prediction of protein structures and interactions using a three- track neural network.Science373, 871–876 (2021)
2021
-
[99]
The use of ai-supported chatbot in psychology.Available at SSRN 4331367 (2023)
Uludag, K. The use of ai-supported chatbot in psychology.Available at SSRN 4331367 (2023)
2023
-
[100]
& Karaarslan, E
Aydın, Ö. & Karaarslan, E. Openai chatgpt generated literature review: Digital twin in healthcare.Available at SSRN 4308687(2022)
2022
-
[101]
H.et al.Performance of chatgpt on usmle: Potential for ai-assisted medical education using large language models.PLoS digital health2, e0000198 (2023)
Kung, T. H.et al.Performance of chatgpt on usmle: Potential for ai-assisted medical education using large language models.PLoS digital health2, e0000198 (2023). 72
2023
-
[102]
Alhawiti, K. M. Natural language processing and its use in education.International Journal of Advanced Computer Science and Applications5(2014)
2014
-
[103]
Chatgpt: The end of online exam integrity?arXiv preprint arXiv:2212.09292 (2022)
Susnjak, T. Chatgpt: The end of online exam integrity?arXiv preprint arXiv:2212.09292 (2022)
2022 arXiv
-
[104]
Clark, T. M. Investigating the use of an artificial intelligence chatbot with general chem- istry exam questions.Journal of Chemical Education(2023)
2023
-
[105]
& Ren, Z
Zhu, J.-J., Jiang, J., Yang, M. & Ren, Z. J. Chatgpt and environmental research.Envi- ronmental Science & Technology(2023)
2023
-
[106]
URLhttps://arxiv.org/abs/2001.08361.2001.08361
Kaplan, J.et al.Scaling laws for neural language models.CoRRabs/2001.08361(2020). URLhttps://arxiv.org/abs/2001.08361.2001.08361
2020 arXiv
-
[107]
Lee, J.et al.Biobert: a pre-trained biomedical language representation model for biomedical text mining.Bioinformatics36, 1234–1240 (2020)
2020
-
[108]
& Cohan, A
Beltagy, I., Lo, K. & Cohan, A. Scibert: A pretrained language model for scientific text. arXiv preprint arXiv:1903.10676(2019)
2019 arXiv
-
[109]
Alsentzer, E.et al.Publicly available clinical bert embeddings.arXiv preprint arXiv:1904.03323(2019)
2019 arXiv
-
[110]
& Fraser, A
Libovick `y, J., Rosa, R. & Fraser, A. On the language neutrality of pre-trained multilin- gual representations.arXiv preprint arXiv:2004.05160(2020)
2020 arXiv
-
[111]
& Hsiang, J
Lee, J.-S. & Hsiang, J. Patent classification by fine-tuning bert language model.World Patent Information61, 101965 (2020)
2020
-
[112]
Finbert: Financial sentiment analysis with pre-trained language models.arXiv preprint arXiv:1908.10063(2019)
Araci, D. Finbert: Financial sentiment analysis with pre-trained language models.arXiv preprint arXiv:1908.10063(2019). 73
2019 arXiv
-
[113]
Wang, W.et al.Automated pipeline for superalloy data by text mining.npj Computa- tional Materials8(2022)
2022
-
[114]
Sutskever, I., Vinyals, O. & Le, Q. V . Sequence to sequence learning with neural net- works.Advances in neural information processing systems27(2014)
2014
-
[115]
K., Stephens, M
Pritchard, J. K., Stephens, M. & Donnelly, P. Inference of population structure using multilocus genotype data.Genetics155, 945–959 (2000)
2000
-
[116]
M., Ng, A
Blei, D. M., Ng, A. Y . & Jordan, M. I. Latent dirichlet allocation.Journal of machine Learning research3, 993–1022 (2003)
2003
-
[117]
& Jeffrey David, U.Mining of massive datasets(Cambridge University Press, 2011)
Anand, R. & Jeffrey David, U.Mining of massive datasets(Cambridge University Press, 2011)
2011
-
[118]
& Breitinger, C
Beel, J., Gipp, B., Langer, S. & Breitinger, C. Paper recommender systems: a literature survey.International Journal on Digital Libraries17, 305–338 (2016)
2016
-
[119]
& Dhillon, I
Sra, S. & Dhillon, I. Generalized nonnegative matrix approximations with bregman divergences.Advances in neural information processing systems18(2005)
2005
-
[121]
Poli, M.et al.Mechanistic design and scaling of hybrid architectures (2024).2403. 17844
2024
-
[122]
Brown, T.et al.Language models are few-shot learners.Advances in neural information processing systems33, 1877–1901 (2020)
2020
-
[123]
2205.11342
Hong, Z.et al.The diminishing returns of masked language models to science (2023). 2205.11342. 74
2023 arXiv
-
[124]
& Dao, T
Gu, A. & Dao, T. Mamba: Linear-time sequence modeling with selective state spaces (2023).2312.00752
2023 arXiv
-
[125]
Peng, B.et al.Rwkv: Reinventing rnns for the transformer era (2023).2305.13048
2023 arXiv
-
[126]
URLhttps://arxiv.org/abs/2108.07258.2108
Bommasani, R.et al.On the opportunities and risks of foundation models.CoRR abs/2108.07258(2021). URLhttps://arxiv.org/abs/2108.07258.2108. 07258
2021 arXiv
-
[127]
& Shankar, M
Yin, J., Dash, S., Wang, F. & Shankar, M. Forge: Pre-training open foundation models for science. InProceedings of the International Conference for High Performance Com- puting, Networking, Storage and Analysis, SC ’23 (Association for Computing Machin- ery, New York, NY , USA...
2023 doi
-
[128]
URLhttps://doi.org/10
Luo, R.et al.BioGPT: generative pre-trained transformer for biomedical text generation and mining.Briefings in Bioinformatics23(2022). URLhttps://doi.org/10. 1093%2Fbib%2Fbbac409
2022
-
[129]
Pubmedgpt (2022)
HAI, S. Pubmedgpt (2022). URLhttps://hai.stanford.edu/news/ stanford-crfm-introduces-pubmedgpt-27b. [Online; posted 15- December-2022]
2022
-
[130]
& Grover, A
Nguyen, T., Brandstetter, J., Kapoor, A., Gupta, J. & Grover, A. Climax: A foundation model for weather and climate (2023)
2023
-
[131]
Accessed: 2021-10-24
Elsevier research products APIs.https://dev.elsevier.com(2021). Accessed: 2021-10-24. [132]https://journals.aps.org/(Accessed: 2023-04-18). [133]https://www.nature.com/siteindex(Accessed: 2023-04-18). 75
2021
-
[134]
InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 1255–1268 (Association for Compu- tational Linguistics, Online, 2020)
Friedrich, A.et al.The SOFC-exp corpus and neural approaches to information ex- traction in the materials science domain. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 1255–1268 (Association for Compu- tational Linguistics, Online, ...
2020
-
[135]
[136]https://scikit-learn.org/(Accessed: 2023-04-18)
Mysore, S.et al.The materials science procedural text corpus: Annotating materials synthesis procedures with shallow semantic structures.arXiv preprint arXiv:1905.06939 (2019). [136]https://scikit-learn.org/(Accessed: 2023-04-18). [137]https://www.tensorflow.org(Accessed: 2023...
2019 arXiv
-
[145]
O’Reilly Media, Inc
Bird, S., Klein, E. & Loper, E.Natural language processing with Python: analyzing text with the natural language toolkit(" O’Reilly Media, Inc.", 2009)
2009
-
[146]
Rasley, J., Rajbhandari, S., Ruwase, O. & He, Y . Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. InProceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery and Data 76 Mining, KDD ’20, 3505–3506 ...
2020
-
[147]
Narayanan, D.et al.Efficient large-scale language model training on gpu clusters using megatron-lm. InProceedings of the International Conference for High Perfor- mance Computing, Networking, Storage and Analysis, SC ’21 (Association for Comput- ing Machinery, New York, NY , U...
2021
-
[148]
& Tourassi, G
Yin, J., Dash, S., Gounley, J., Wang, F. & Tourassi, G. Evaluation of pre-training large language models on leadership-class supercomputers.The Journal of Supercomputing 1–22 (2023)
2023
-
[149]
Supercom- put.81(2024)
Yin, J.et al.chathpc: Empowering hpc users with large language models.J. Supercom- put.81(2024). URLhttps://doi.org/10.1007/s11227-024-06637-1
2024 doi
-
[150]
Gupta, T., Zaki, M., Krishnan, N. M. A. & Mausam. MatSciBERT: A materials domain language model for text mining and information extraction.npj Computa- tional Materials8, 102 (2022). URLhttps://www.nature.com/articles/ s41524-022-00784-w
2022
-
[151]
& Anthony, Q
Yin, J., Bose, A., Cong, G., Lyngaas, I. & Anthony, Q. Comparative study of large language model architectures on frontier. In2024 IEEE International Parallel and Dis- tributed Processing Symposium (IPDPS), 556–569 (IEEE Computer Society, Los Alami- tos, CA, USA, 2024). URLhtt...
2024
-
[152]
Huang, L.et al.A survey on hallucination in large language models: Principles, taxon- omy, challenges, and open questions.arXiv preprint arXiv:2311.05232(2023)
2023 arXiv
-
[153]
arXiv preprint arXiv:2011.02593(2020)
Zhou, C.et al.Detecting hallucinated content in conditional neural sequence generation. arXiv preprint arXiv:2011.02593(2020). 77
2020 arXiv
-
[154]
& Cheung, J
Cao, M., Dong, Y . & Cheung, J. C. K. Hallucinated but factual! inspecting the fac- tuality of hallucinations in abstractive summarization. InProceedings of the 60th An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3340–3354 (2022)
2022
-
[155]
& Srihari, R
Das, S., Saha, S. & Srihari, R. K. Diving deep into modes of fact hallucinations in dialogue systems.arXiv preprint arXiv:2301.04449(2023)
2023 arXiv
-
[156]
& Smith, N
Zhang, M., Press, O., Merrill, W., Liu, A. & Smith, N. A. How language model halluci- nations can snowball.arXiv preprint arXiv:2305.13534(2023)
2023 arXiv
-
[157]
& Chen-Chuan Chang, K
Zheng, S., Huang, J. & Chen-Chuan Chang, K. Why does chatgpt fall short in providing truthful answers?arXiv preprint arXiv:2304.10513(2023)
2023 arXiv
-
[158]
Dhuliawala, S.et al.Chain-of-verification reduces hallucination in large language mod- els.arXiv preprint arXiv:2309.11495(2023)
2023 arXiv
-
[159]
Ji, Z.et al.Survey of hallucination in natural language generation.ACM Computing Surveys55, 1–38 (2023)
2023
-
[160]
Zhang, Y .et al.Siren’s song in the ai ocean: A survey on hallucination in large language models.arXiv preprint arXiv:2309.01219(2023)
2023 arXiv
-
[161]
& Jia, W
Ye, H., Liu, T., Zhang, A., Hua, W. & Jia, W. Cognitive mirage: A review of hallucina- tions in large language models.arXiv preprint arXiv:2309.06794(2023)
2023 arXiv
-
[162]
Liu, T.et al.A token-level reference-free hallucination detection benchmark for free- form text generation.arXiv preprint arXiv:2104.08704(2021)
2021 arXiv
-
[163]
X., Nie, J.-Y
Li, J., Cheng, X., Zhao, W. X., Nie, J.-Y . & Wen, J.-R. Halueval: A large-scale hal- lucination evaluation benchmark for large language models.arXiv e-printsarXiv–2305 (2023). 78
2023
-
[164]
& Wan, X
Yang, S., Sun, R. & Wan, X. A new benchmark and reverse validation method for passage-level hallucination detection.arXiv preprint arXiv:2310.06498(2023)
2023 arXiv
-
[165]
& Wang, W
Xiao, Y . & Wang, W. Y . On hallucination and predictive uncertainty in conditional lan- guage generation.arXiv preprint arXiv:2103.15025(2021)
2021 arXiv
-
[166]
Varshney, N., Yao, W., Zhang, H., Chen, J. & Yu, D. A stitch in time saves nine: Detect- ing and mitigating hallucinations of llms by validating low-confidence generation.arXiv preprint arXiv:2307.03987(2023)
2023 arXiv
-
[167]
& Mueller, J
Chen, J. & Mueller, J. Quantifying uncertainty in answers from any language model via intrinsic and extrinsic confidence assessment.arXiv preprint arXiv:2308.16175(2023)
2023 arXiv
-
[168]
& Farquhar, S
Kuhn, L., Gal, Y . & Farquhar, S. Semantic uncertainty: Linguistic invariances for un- certainty estimation in natural language generation.arXiv preprint arXiv:2302.09664 (2023)
2023 arXiv
-
[169]
R.et al.Selectively answering ambiguous questions.arXiv preprint arXiv:2305.14613(2023)
Cole, J. R.et al.Selectively answering ambiguous questions.arXiv preprint arXiv:2305.14613(2023)
2023 arXiv
-
[170]
& Kalai, A
Agrawal, A., Mackey, L. & Kalai, A. T. Do language models know when they’re hallu- cinating references?arXiv preprint arXiv:2305.18248(2023)
2023 arXiv
-
[171]
& Wattenberg, M
Li, K., Patel, O., Viégas, F., Pfister, H. & Wattenberg, M. Inference-time intervention: Eliciting truthful answers from a language model.arXiv preprint arXiv:2306.03341 (2023)
2023 arXiv
-
[172]
arXiv preprint arXiv:2312.10997(2023)
Gao, Y .et al.Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997(2023)
2023 arXiv
-
[173]
Llm-powered autonomous agents.lilianweng.github.io(2023)
Weng, L. Llm-powered autonomous agents.lilianweng.github.io(2023). URLhttps: //lilianweng.github.io/posts/2023-06-23-agent/. 79
2023
-
[174]
Wang, L.et al.A survey on large language model based autonomous agents.Frontiers of Computer Science18, 186345 (2024)
2024
-
[176]
Guo, T.et al.What can large language models do in chemistry? a comprehensive bench- mark on eight tasks.Advances in Neural Information Processing Systems36, 59662– 59688 (2023)
2023
-
[177]
Dagdelen, J.et al.Structured information extraction from scientific text with large lan- guage models.Nature Communications15, 1418 (2024)
2024
-
[178]
P.et al.Flexible, model-agnostic method for materials data extraction from text using general purpose language models.Digital Discovery3, 1221–1235 (2024)
Polak, M. P.et al.Flexible, model-agnostic method for materials data extraction from text using general purpose language models.Digital Discovery3, 1221–1235 (2024)
2024
-
[179]
Polak, M. P. & Morgan, D. Extracting accurate materials data from research papers with conversational language models and prompt engineering.Nature Communications15, 1569 (2024)
2024
-
[180]
Schilling-Wilhelmi, M.et al.From text to insight: large language models for materials science data extraction.arXiv preprint arXiv:2407.16867(2024)
2024 arXiv
-
[181]
& Jain, A
Pei, Z., Yin, J., Neugebauer, J. & Jain, A. Towards the holistic design of alloys with large language models.Nature Reviews Materials1–2 (2024)
2024
-
[182]
A., MacKnight, R., Kline, B
Boiko, D. A., MacKnight, R., Kline, B. & Gomes, G. Autonomous chemical research with large language models.Nature624, 570–578 (2023)
2023
-
[183]
& Riebesell, J
Chiang, Y ., Hsieh, E., Chou, C.-H. & Riebesell, J. Llamp: Large language model made powerful for high-fidelity materials knowledge retrieval and distillation.arXiv preprint arXiv:2401.17244(2024). 80
2024 arXiv
-
[184]
N.et al.Enhancing corrosion-resistant alloy design through natural lan- guage processing and deep learning.Science Advances9, eadg7992 (2023)
Sasidhar, K. N.et al.Enhancing corrosion-resistant alloy design through natural lan- guage processing and deep learning.Science Advances9, eadg7992 (2023)
2023
-
[185]
M., Schwaller, P., Ortega-Guerrero, A
Jablonka, K. M., Schwaller, P., Ortega-Guerrero, A. & Smit, B. Leveraging large lan- guage models for predictive chemistry.Nature Machine Intelligence6, 161–169 (2024)
2024
-
[186]
Jacobs, R.et al.Regression with large language models for materials and molecular property prediction.arXiv preprint arXiv:2409.06080(2024)
2024 arXiv
-
[187]
Liu, S., Wen, T., Pattamatta, A. S. & Srolovitz, D. J. A prompt-engineered large language model, deep learning workflow for materials classification.Materials Today80, 240–249 (2024)
2024
-
[188]
& Aspuru-Guzik, A
Sanchez-Lengeling, B. & Aspuru-Guzik, A. Inverse molecular design using machine learning: Generative models for matter engineering.Science361, 360–365 (2018)
2018
-
[189]
Inverse design in search of materials with target functionalities.Nature Reviews Chemistry2, 1–16 (2018)
Zunger, A. Inverse design in search of materials with target functionalities.Nature Reviews Chemistry2, 1–16 (2018)
2018
-
[190]
Popper, bayes and the inverse problem.Nature physics2, 492–494 (2006)
Tarantola, A. Popper, bayes and the inverse problem.Nature physics2, 492–494 (2006)
2006
-
[191]
S., Keszler, D
Yu, L., Kokenyesi, R. S., Keszler, D. A. & Zunger, A. Inverse design of high absorption thin-film photovoltaic materials.Advanced Energy Materials3, 43–48 (2013)
2013
-
[192]
& Fung, V
Zhang, J. & Fung, V . Efficient inverse learning for materials design and discovery. In ICLR 2021 Workshop on Science and Engineering of Deep Learning(2021)
2021
-
[193]
& Sumpter, B
Fung, V ., Zhang, J., Hu, G., Ganesh, P. & Sumpter, B. G. Inverse design of two- dimensional materials with invertible neural networks.npj Computational Materials7, 1–9 (2021)
2021
-
[194]
URLhttps://www
Zhang, J.et al.Robust data-driven approach for predicting the configurational energy of high entropy alloys.Materials & Design185, 108247 (2020). URLhttps://www. sciencedirect.com/science/article/pii/S0264127519306859. 81
2020
-
[195]
URLhttps://www.sciencedirect.com/science/ article/pii/S0927025620306261
Liu, X.et al.Monte carlo simulation of order-disorder transition in refractory high entropy alloys: A data-driven approach.Computational Materials Science 187, 110135 (2021). URLhttps://www.sciencedirect.com/science/ article/pii/S0927025620306261
2021
-
[196]
& Gao, M
Yin, J., Pei, Z. & Gao, M. C. Neural network-based order parameter for phase transitions and its applications in high-entropy alloys.Nature Computational Science1, 686–693 (2021)
2021
-
[197]
M.et al.Materials informatics for the screening of multi-principal elements and high-entropy alloys.Nature Communications10, 2618 (2019)
Rickman, J. M.et al.Materials informatics for the screening of multi-principal elements and high-entropy alloys.Nature Communications10, 2618 (2019). URLhttps:// doi.org/10.1038/s41467-019-10533-1
2019 doi
-
[198]
Ha, M.-Q.et al.Evidence-based recommender system and experimental validation for high-entropy alloys.Nature Computational Science1, 470–478 (2021)
2021
-
[199]
& Johnson, D
Singh, R., Sharma, A., Singh, P., Balasubramanian, G. & Johnson, D. D. Ac- celerating computational modeling and design of high-entropy alloys.Nature Computational Science1, 54–61 (2021). URLhttps://doi.org/10.1038/ s43588-020-00006-7
2021
-
[200]
URLhttp://www.sciencedirect.com/science/article/pii/ S1359645419308158
Zhang, Y .et al.Phase prediction in high entropy alloys with a rational selection of materials descriptors and machine learning models.Acta Materialia185, 528 – 539 (2020). URLhttp://www.sciencedirect.com/science/article/pii/ S1359645419308158
2020
-
[201]
& Zhuang, H
Huang, W., Martin, P. & Zhuang, H. L. Machine-learning phase prediction of high-entropy alloys.Acta Materialia169, 225–236 (2019). URLhttps://www. sciencedirect.com/science/article/pii/S1359645419301454
2019
-
[202]
& Zhuang, H
Islam, N., Huang, W. & Zhuang, H. L. Machine learning for phase selection in multi-principal element alloys.Computational Materials Science150, 230– 82 235 (2018). URLhttps://www.sciencedirect.com/science/article/ pii/S0927025618302386
2018
-
[203]
& Rivera Díaz-Del-Castillo, P
Tancret, F., Toda-Caraballo, I., Menou, E. & Rivera Díaz-Del-Castillo, P. E. J. Designing high entropy alloys employing thermodynamics and gaussian process statistical analysis. Materials & Design115, 486–497 (2017). URLhttps://www.sciencedirect. com/science/article/pii/S02641...
2017
-
[204]
& Vecchio, K
Kaufmann, K. & Vecchio, K. S. Searching for high entropy alloys: A machine learning approach.Acta Materialia198, 178–222 (2020). URLhttps://www. sciencedirect.com/science/article/pii/S1359645420305814
2020
-
[205]
& Guo, W
Li, Y . & Guo, W. Machine-learning model for predicting phase formations of high- entropy alloys.Phys. Rev. Materials3, 095005 (2019). URLhttps://link.aps. org/doi/10.1103/PhysRevMaterials.3.095005
2019 doi
-
[206]
Y ., Byeon, S., Kim, H
Lee, S. Y ., Byeon, S., Kim, H. S., Jin, H. & Lee, S. Deep learning-based phase predic- tion of high-entropy alloys: Optimization, generation, and explanation.Materials & De- sign197, 109260 (2021). URLhttps://www.sciencedirect.com/science/ article/pii/S0264127520307954
2021
-
[207]
URLhttps://doi
Zhou, Z.et al.Machine learning guided appraisal and exploration of phase design for high entropy alloys.npj Computational Materials5, 128 (2019). URLhttps://doi. org/10.1038/s41524-019-0265-1
2019 doi
-
[208]
URLhttp://www.sciencedirect
Wen, C.et al.Machine learning assisted design of high entropy alloys with desired prop- erty.Acta Materialia170, 109 – 117 (2019). URLhttp://www.sciencedirect. com/science/article/pii/S1359645419301430
2019
-
[209]
A., Lara-Curzio, E
Peng, J., Yamamoto, Y ., Hawk, J. A., Lara-Curzio, E. & Shin, D. Coupling physics in machine learning to predict properties of high-temperatures alloys.npj Computational Materials6, 1–7 (2020). 83
2020
-
[210]
& Yoon, D
Kumar, S., Jaafreh, R., Singh, N., Hamad, K. & Yoon, D. H. Introducing magbert: A language model for magnesium textual data mining and analysis.Journal of Magnesium and Alloys12, 3216–3228 (2024)
2024
-
[211]
Wang, W.et al.Alloy synthesis and processing by semi-supervised text mining.npj Computational Materials9, 183 (2023)
2023
-
[212]
Introducing pre-trained transformers for high entropy alloy informatics.Ma- terials Letters358, 135871 (2024)
Kamnis, S. Introducing pre-trained transformers for high entropy alloy informatics.Ma- terials Letters358, 135871 (2024)
2024
-
[213]
Yan, Z.et al.Pdgpt: A large language model for acquiring phase diagram information in magnesium alloys.Materials Genome Engineering Advancese77 (2024)
2024
-
[214]
Scientific data4, 1–9 (2017)
Kim, E.et al.Machine-learned and codified synthesis parameters of oxide materials. Scientific data4, 1–9 (2017)
2017
-
[215]
Kononova, O.et al.Text-mined dataset of inorganic materials synthesis recipes.Scientific data6, 203 (2019)
2019
-
[216]
Wang, Z.et al.Dataset of solution-based inorganic materials synthesis procedures ex- tracted from the scientific literature.Scientific Data9, 231 (2022)
2022
-
[217]
Venugopal, V .et al.Looking through glass: Knowledge discovery from materials science literature using natural language processing.Patterns2, 100290 (2021)
2021
-
[218]
Su, Y .et al.Automation and machine learning augmented by large language models in a catalysis study.Chemical Science15, 12200–12233 (2024)
2024
-
[219]
& Schrier, J
Kim, S., Jung, Y . & Schrier, J. Large language models for inorganic synthesis predictions. Journal of the American Chemical Society146, 19654–19659 (2024)
2024
-
[220]
Shetty, P.et al.A general-purpose material property data extraction pipeline from large polymer corpora using natural language processing.npj Computational Materials9, 52 (2023). 84
2023
-
[221]
Zaki, M., Krishnan, N. A.et al.Extracting processing and testing parameters from materials science literature for improved property prediction of glasses.Chemical Engi- neering and Processing-Process Intensification180, 108607 (2022)
2022
-
[222]
& Machiraju, R
Kulkarni, C., Xu, W., Ritter, A. & Machiraju, R. An annotated corpus for machine read- ing of instructions in wet lab protocols. InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lan- guage Technologies, ...
2018
-
[223]
S.et al.Eras: Improving the quality control in the annotation process for natural language processing tasks.Information Systems93, 101553 (2020)
Grosman, J. S.et al.Eras: Improving the quality control in the annotation process for natural language processing tasks.Information Systems93, 101553 (2020)
2020
-
[224]
Guha, S.et al.Matscie: An automated tool for the generation of databases of methods and parameters used in the computational materials science literature.Computational Materials Science192, 110325 (2021)
2021
-
[225]
M., Butler, K
Antunes, L. M., Butler, K. T. & Grau-Crespo, R. Crystal structure generation with au- toregressive large language modeling.Nature Communications15, 1–16 (2024)
2024
-
[226]
Atomgpt: Atomistic generative pretrained transformer for forward and in- verse materials design.The Journal of Physical Chemistry Letters15, 6909–6917 (2024)
Choudhary, K. Atomgpt: Atomistic generative pretrained transformer for forward and in- verse materials design.The Journal of Physical Chemistry Letters15, 6909–6917 (2024)
2024
-
[227]
InMethods in enzymology, vol
Leaver-Fay, A.et al.Rosetta3: an object-oriented software suite for the simulation and design of macromolecules. InMethods in enzymology, vol. 487, 545–574 (Elsevier, 2011)
2011
-
[228]
Abramson, J.et al.Accurate structure prediction of biomolecular interactions with al- phafold 3.Nature1–3 (2024). 85
2024
-
[229]
Computer-aided drug discovery: From traditional simulation methods to language models and quantum computing.Cell Reports Physical Science(2024)
Pei, Z. Computer-aided drug discovery: From traditional simulation methods to language models and quantum computing.Cell Reports Physical Science(2024)
2024
-
[230]
& Aspuru-Guzik, A
Flam-Shepherd, D., Zhu, K. & Aspuru-Guzik, A. Language models can learn complex molecular distributions.Nature Communications13, 3293 (2022)
2022
-
[231]
& Jaakkola, T
Jin, W., Barzilay, R. & Jaakkola, T. Junction tree variational autoencoder for molecular graph generation. InInternational conference on machine learning, 2323–2332 (PMLR, 2018)
2018
-
[232]
& Gaunt, A
Liu, Q., Allamanis, M., Brockschmidt, M. & Gaunt, A. Constrained graph variational autoencoders for molecule design.Advances in neural information processing systems 31(2018)
2018
-
[233]
& Brinson, L
Circi, D., Khalighinejad, G., Chen, A., Dhingra, B. & Brinson, L. C. How well do large language models understand tables in materials science?Integrating Materials and Manufacturing Innovation13, 669–687 (2024)
2024
-
[234]
& Dahotre, N
Parsazadeh, M., Sharma, S. & Dahotre, N. Towards the next generation of machine learn- ing models in additive manufacturing: A review of process dependent material evolution. Progress in Materials Science135, 101102 (2023)
2023
-
[235]
& Pugliese, R
Badini, S., Regondi, S., Frontoni, E. & Pugliese, R. Assessing the capabilities of chatgpt to improve additive manufacturing troubleshooting.Advanced Industrial and Engineer- ing Polymer Research6, 278–287 (2023)
2023
-
[236]
& Farimani, A
Chandrasekhar, A., Chan, J., Ogoke, F., Ajenifujah, O. & Farimani, A. B. Amgpt: a large language model for contextual querying in additive manufacturing.arXiv preprint arXiv:2406.00031(2024)
2024 arXiv
-
[237]
Eslaminia, A.et al.Fdm-bench: A comprehensive benchmark for evaluating large lan- guage models in additive manufacturing tasks.arXiv preprint arXiv:2412.09819(2024). 86
2024 arXiv
-
[238]
Jignasu, A.et al.Towards foundational ai models for additive manufacturing: Lan- guage models for g-code debugging, manipulation, and comprehension.arXiv preprint arXiv:2309.02465(2023)
2023 arXiv
-
[239]
A., Fuh, J
Liu, X., Erkoyuncu, J. A., Fuh, J. Y . H., Lu, W. F. & Li, B. Knowledge extraction for additive manufacturing process via named entity recognition with llms.Robotics and Computer-Integrated Manufacturing93, 102900 (2025)
2025
-
[240]
C.et al.Computational design and manufacturing of sustainable materials through first-principles and materiomics.Chemical Reviews(2023)
Shen, S. C.et al.Computational design and manufacturing of sustainable materials through first-principles and materiomics.Chemical Reviews(2023)
2023
-
[241]
& Burgert, I
Schubert, M., Panzarasa, G. & Burgert, I. Sustainability in wood products: A new per- spective for handling natural diversity.Chemical Reviews(2022)
2022
-
[242]
& Mohammadi, P
Miserez, A., Yu, J. & Mohammadi, P. Protein-based biological materials: Molecular design and artificial production.Chemical Reviews123, 2049–2111 (2023)
2023
-
[243]
& Harrington, M
Rising, A. & Harrington, M. J. Biological materials processing: Time-tested tricks for sustainable fiber fabrication.Chemical Reviews(2022)
2022
-
[244]
Council, N. R.et al. Integrated computational materials engineering: a transformational discipline for improved competitiveness and national security(National Academies Press, 2008)
2008
-
[245]
& Fries, S.Computational thermodynamics: the Calphad method(Cambridge university press New York, 2007)
Sundman, B., Lukas, H. & Fries, S.Computational thermodynamics: the Calphad method(Cambridge university press New York, 2007)
2007
-
[246]
J.et al.New frontiers for the materials genome initiative.npj Computational Materials5, 1–23 (2019)
de Pablo, J. J.et al.New frontiers for the materials genome initiative.npj Computational Materials5, 1–23 (2019)
2019
-
[248]
& Ghosh, S
Satpute, P., Tiwari, S., Gupta, M. & Ghosh, S. Exploring large language models for mi- crostructure evolution in materials.Materials Today Communications40, 109583 (2024)
2024
-
[249]
Liu, Y .et al.Generative artificial intelligence and its applications in materials science: Current situation and future perspectives.Journal of Materiomics9, 798–816 (2023)
2023
-
[250]
Deb, J., Saikia, L., Dihingia, K. D. & Sastry, G. N. Chatgpt in the material design: Selected case studies to assess the potential of chatgpt.Journal of Chemical Information and Modeling64, 799–811 (2024)
2024
-
[251]
Chatgpt for computational materials science: a perspective.Energy Material Advances4, 0026 (2023)
Hong, Z. Chatgpt for computational materials science: a perspective.Energy Material Advances4, 0026 (2023)
2023
-
[252]
Schütze, H., Manning, C. D. & Raghavan, P.Introduction to information retrieval, vol. 39 (Cambridge University Press Cambridge, 2008)
2008
-
[253]
Introducing the knowledge graph: things, not strings
Singhal, A. Introducing the knowledge graph: things, not strings. Google Official Blog (2021). Accessed: 2021-10-24
2021
-
[254]
& Liu, X
Fang, Y ., Chen, M., Liang, W., Zhou, Z. & Liu, X. Knowledge graph learning for vehicle additive manufacturing of recycled metal powder.World Electric Vehicle Journal14, 289 (2023)
2023
-
[255]
Taylor, R.et al.Galactica: A large language model for science.arXiv preprint arXiv:2211.09085(2022)
2022 arXiv
-
[256]
npj Computational Materials10, 58 (2024)
Qu, J.et al.Leveraging language representation for materials exploration and discovery. npj Computational Materials10, 58 (2024)
2024
-
[257]
Azamfirei, R., Kudchadkar, S. R. & Fackler, J. Large language models and the perils of their hallucinations.Critical Care27, 120 (2023)
2023
-
[258]
& Singla, S
Banerjee, S., Agarwal, A. & Singla, S. Llms will always hallucinate, and we need to live with this.arXiv preprint arXiv:2409.05746(2024). 88
2024 arXiv
-
[259]
Training a 1 trillion parameter model with pytorch fully sharded data parallel on aws (2022)
team, P. Training a 1 trillion parameter model with pytorch fully sharded data parallel on aws (2022). URLhttps://medium.com/pytorch/ training-a-1-trillion-parameter-model-with-pytorch-fully-sharded-data-parallel-on-aws-3ac13aa96cff. [Online; posted 15-March-2022]
2022
-
[260]
& Wolf, T
Sanh, V ., Debut, L., Chaumond, J. & Wolf, T. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter.arXiv preprint arXiv:1910.01108(2019)
2019 arXiv
-
[261]
M.et al.Chemcrow: Augmenting large-language models with chemistry tools (2023).2304.05376
Bran, A. M.et al.Chemcrow: Augmenting large-language models with chemistry tools (2023).2304.05376
2023 arXiv
-
[262]
Li, J.et al.Ai applications through the whole life cycle of material discovery.Matter3, 393–432 (2020)
2020
-
[263]
Yao, Z.et al.Machine learning for a sustainable energy future.Nature Reviews Materials 1–14 (2022)
2022
-
[264]
Han, L., Mu, W., Wei, S., Liaw, P. K. & Raabe, D. Sustainable high-entropy materials? Science Advances10, eads3926 (2024)
2024
-
[265]
Meichanetzidis, K.et al.Quantum natural language processing on near-term quantum computers.arXiv preprint arXiv:2005.04147(2020)
2020
-
[266]
& Wootton, J
Karamlou, A., Pfaffhauser, M. & Wootton, J. Quantum natural language generation on near-term devices.arXiv preprint arXiv:2211.00727(2022)
2022 arXiv
-
[267]
& Pei, Z
Del Castillo, J., Zhao, D. & Pei, Z. Comparative study of the ans\" atze in quantum language models.arXiv preprint arXiv:2502.20744(2025). 89
2025 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.