Pith. sign in

REVIEW 3 major objections 4 minor 4 cited by

Reflections from the 2024 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read 34 teams produced working LLM prototypes across seven areas of materials science and chemistry, and the organizers read this as evidence that LLMs now serve as both general-purpose predictors and rapid-prototyping platforms.

desk verdict A useful, honest catalogue of 34 LLM prototypes in materials and chemistry; the abstract's 'significant improvements' trend claim is unsupported and should be qualified. read the letter →

arxiv 2411.15221 v2 pith:PGFBIFJG submitted 2024-11-20 cs.LG cond-mat.mtrl-sciphysics.chem-ph

Yoel Zimmermann , Adib Bazgir , Zartashia Afzal , Fariha Agbere , Qianxiang Ai , Nawaf Alampara , Alexander Al-Feghali , Mehrad Ansari
show 135 more authors
Dmytro Antypov Amro Aswad Jiaru Bai Viktoriia Baibakova Devi Dutta Biswajeet Erik Bitzek Joshua D. Bocarsly Anna Borisova Andres M Bran L. Catherine Brinson Marcel Moran Calderon Alessandro Canalicchio Victor Chen Yuan Chiang Defne Circi Benjamin Charmes Vikrant Chaudhary Zizhang Chen Min-Hsueh Chiu Judith Clymo Kedar Dabhadkar Nathan Daelman Archit Datar Wibe A. de Jong Matthew L. Evans Maryam Ghazizade Fard Giuseppe Fisicaro Abhijeet Sadashiv Gangan Janine George Jose D. Cojal Gonzalez Michael Götte Ankur K. Gupta Hassan Harb Pengyu Hong Abdelrahman Ibrahim Ahmed Ilyas Alishba Imran Kevin Ishimwe Ramsey Issa Kevin Maik Jablonka Colin Jones Tyler R. Josephson Greg Juhasz Sarthak Kapoor Rongda Kang Ghazal Khalighinejad Sartaaj Khan Sascha Klawohn Suneel Kuman Alvin Noe Ladines Sarom Leang Magdalena Lederbauer Sheng-Lun (Mark) Liao Hao Liu Xuefeng Liu Stanley Lo Sandeep Madireddy Piyush Ranjan Maharana Shagun Maheshwari Soroush Mahjoubi José A. Márquez Rob Mills Trupti Mohanty Bernadette Mohr Seyed Mohamad Moosavi Alexander Mo{ss}hammer Amirhossein D. Naghdi Aakash Naik Oleksandr Narykov Hampus Näsström Xuan Vu Nguyen Xinyi Ni Dana O'Connor Teslim Olayiwola Federico Ottomano Aleyna Beste Ozhan Sebastian Pagel Chiku Parida Jaehee Park Vraj Patel Elena Patyukova Martin Hoffmann Petersen Luis Pinto José M. Pizarro Dieter Plessers Tapashree Pradhan Utkarsh Pratiush Charishma Puli Andrew Qin Mahyar Rajabi Francesco Ricci Elliot Risch Martiño Ríos-García Aritra Roy Tehseen Rug Hasan M Sayeed Markus Scheidgen Mara Schilling-Wilhelmi Marcel Schloz Fabian Schöppach Julia Schumann Philippe Schwaller Marcus Schwarting Samiha Sharlin Kevin Shen Jiale Shi Pradip Si Jennifer D'Souza Taylor Sparks Suraj Sudhakar Leopold Talirz Dandan Tang Olga Taran Carla Terboven Mark Tropin Anastasiia Tsymbal Katharina Ueltzen Pablo Andres Unzueta Archit Vasan Tirtha Vinchurkar Trung Vo Gabriel Vogel Christoph Völker Jan Weinreich Faradawn Yang Mohd Zaki Chi Zhang Sylvester Zhang Weijie Zhang Ruijie Zhu Shang Zhu Jan Janssen Calvin Li Ian Foster Ben Blaiszik
This is my paper · ORCID
classification cs.LGcond-mat.mtrl-sciphysics.chem-ph
keywords largelanguagemodelsmaterialssciencechemistryhackathonmolecularpropertypredictionretrieval-augmentedgenerationscientificworkflowsrapidprototyping
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This report tries to establish that modern LLMs have become practical tools for materials science and chemistry, not just research curiosities. It summarizes 34 projects built within a 24-hour hackathon, covering molecular property prediction, molecular design, automation, education, data management, hypothesis evaluation, and literature mining. If the report's reading is right, the same class of models can serve as general-purpose prediction engines and as a fast prototyping layer for custom scientific tools. The paper also argues that the hybrid hackathon format, with physical hubs and an online community, supports durable new scientific collaborations.

What carries the argument

The carrying mechanism is the time-boxed, open-ended hybrid challenge itself: 556 registered participants, 34 completed team submissions, each producing code and a short project report. The organizers use those reports as their dataset, group them into seven application areas, and compare the demonstrated capabilities against those of the previous year's event.

What would settle it

Inspect the linked repositories and rerun the key claims in a controlled setting: five-fold cross-validate the phonon peak predictor, re-run the XRD schema-filling pipeline on the same raw files and count hallucinated values, and test the novelty-scoring model on a larger abstract set. If the flagship prototypes do not reproduce or their error rates equal random baselines, the paper's evidence for expanded LLM capability collapses.

Watch

Extended reading notes

Core claim

The paper claims that second-generation LLMs now have a dual utility in materials and chemistry: they are multipurpose models that can be pointed at regression, generation, and extraction tasks, and they are platforms on which researchers can prototype custom applications in under a day. The evidence is the set of 34 hackathon submissions, each with code and a written report, which the organizers sorted into seven application areas. They observe that performance across this application space improved substantially since the previous year's hackathon, in many cases simply because newer model versions were released, and they take this as a sign that LLM use in these fields will keep expanding.

Load-bearing premise

The conclusion rests on treating 34 self-selected, largely unvalidated team reports as representative evidence; many entries lack cross-validation, statistical significance, or quantitative checks, so if the reports flatter reality, the case for broad LLM utility is not made.

Editorial extensions

If this is right

  • Low-data property prediction appears within reach: teams improved phonon peak and lithium-ion conductivity predictions by enriching composition inputs with bonding or literature context.
  • Design tasks can be bootstrapped with zero-shot LLM suggestions, and small locally hosted models beat random baselines in concrete formulation design, suggesting that private and data-sensitive laboratories could use them.
  • Natural-language interfaces for instruments and simulations are feasible: teams automated bulk-modulus calculations, microscope parameter estimation, and DFT setup, lowering the need for specialist operators.
  • Text-native workflows, such as knowledge-graph construction, schema filling, hypothesis evaluation, and question answering, are where LLMs most readily plug in, because these tasks align with the models' core strengths.
  • If the trend continues, newer model generations alone will broaden the range of research tasks that LLMs can handle without custom adaptation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The report's emphasis on breadth over benchmarks suggests the fastest near-term gains will come in text-native data work, such as literature extraction, electronic lab notebooks, and report drafting, where success is judged by workflow completion rather than scientific accuracy.
  • Because many prototypes lack error bars, the honest reading is that LLMs widen what can be attempted, not yet what can be trusted; a reproducibility pass across the 34 repositories would separate stable tools from one-off demonstrations.
  • The low-data property-prediction results hint that context augmentation, such as bonding analysis and literature summaries, may matter more than model scale, a claim the paper supports only partially and that deserves direct comparison against classical machine-learning baselines.
  • If small locally hosted models continue to approach larger ones on design tasks, the practical ceiling may shift from model capability to prompt and workflow engineering, which would make data privacy less of an obstacle.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper reports on the second Large Language Model Hackathon for Applications in Materials Science and Chemistry, held in May 2024 with hybrid physical hubs and a global online hub. It describes 34 team submissions, of which 32 have written reports in the appendix, organizes them into seven application areas (property prediction, design, automation and interfaces, communication and education, data management, hypothesis generation, and knowledge extraction), and highlights exemplar projects in each area. The paper also discusses the event format and concludes that the event demonstrated the dual utility of LLMs as multipurpose predictive models and as rapid-prototyping platforms, and that LLM capabilities improved significantly since the previous year's hackathon. Individual appendix reports provide code links and, in several cases, candid statements of limitations.

Significance. The paper's main value is as an archival record of a large, open, community-driven prototyping event: it collects 32 project reports with public code repositories, spans a broad application taxonomy, and includes unusually honest reporting of failures and non-significant results. If its broader claims were supported, the paper would provide useful evidence that LLMs can be productively applied across a wide range of materials chemistry tasks within a 24-hour hackathon format. However, the headline claims of the abstract and conclusion go beyond what the evidence supports: no matched comparison with the previous hackathon is provided, and several reported projects show hallucinations, model failure, or non-significant results. These issues are local to the framing and can be corrected by softening the claims; the descriptive core of the paper remains valuable.

major comments (3)
  1. [Abstract; Conclusion; Appendix §§5, 8, 18, 24] The abstract states that 'the event highlighted significant improvements in LLM capabilities since the previous year's hackathon', and the conclusion attributes this improvement to the release of new model versions. No matched comparison between the 2024 and 2023 hackathons is provided: the teams, tasks, datasets, and evaluation protocols differ, and no project benchmarks current models against the models used in the previous event. Moreover, several projects in the appendix report negative or mixed outcomes, including LLMSpectrometry solving only 3 of 19 NMR spectra (§5), the Phi-3 model failing after the fifth development cycle in the small-LLM concrete study (§8), hallucinated values in LLMads (§18), and a non-significant p=0.07 in G-Peer-T (§24). The improvement claim should either be removed or replaced by a carefully hedged statement that the breadth of prototypes is consistent with, but does not measure, year-over-year capability growth.
  2. [Abstract; Event Overview; Table 1; Appendix TOC] The manuscript reports '34 team submissions' but Table 1 and the appendix table of contents list only 32 projects. The abstract also says that each team submission is presented in a summary table and as brief papers in the appendix, which is inconsistent with the '32 submissions with a written description included here' formulation in the Event Overview. Please reconcile the total: either list the two additional submissions (even if they lack write-ups or code), or change the stated total to 32.
  3. [Conclusion; Abstract] The claim that the event 'demonstrated the dual utility of LLMs' is stronger than the evidence warrants. The paper is a collection of self-selected, time-constrained prototypes, and several appendix reports explicitly note missing validation: Learning LOBSTERs states that no five-fold cross-validation was implemented (§1), NOMAD Query Reporter reports hallucinations on heterogeneous data (§19), and the small-LLM study reports a complete failure mode for Phi-3 (§8). The conclusion should be phrased in terms of demonstrated breadth of prototyping and potential utility, not demonstrated utility, unless the authors add a systematic validation criterion applied across projects.
minor comments (4)
  1. [Table 1 caption] The caption reads 'Overview of the tools developed by the various tools'; 'tools' should presumably be 'teams'.
  2. [§8.4, Table 2] The sentence 'all investigated models generated designs that outperformed the statistical baseline' is immediately followed by an exception for Phi-3 with design context, which failed after the fifth development cycle; please rephrase to state that all models outperformed the baseline in the first round, with the noted exception in later rounds.
  3. [§19, references] In the NOMAD Query Reporter section, the in-text citations appear swapped: 'Query Reporter [1]' and 'NOMAD [2]' should refer to the GitHub repository and the NOMAD paper respectively, but the reference list assigns [1] to the NOMAD paper and [2] to the GitHub repository.
  4. [Overview; Appendix §17] The project name is spelled inconsistently as 'yeLLowhaMmer' in the overview and 'yeLLowhaMMer' in the appendix title; please standardize the spelling.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: the report's conclusions are inductive summaries of self-selected prototypes; self-citations are contextual, and the unsupported year-over-year improvement claim is an external attribution rather than a circular inference.

full rationale

This paper is a hackathon proceedings rather than a derivation: it reports 34 prototype submissions and summarizes their application areas. The abstract's comparative claim, 'significant improvements in LLM capabilities since the previous year's hackathon,' is not derived from any fitted parameter in this paper; the Conclusion attributes it to an external event ('the performance across the diverse application space was improved simply via the release of new versions of Gemini, ChatGPT, Claude, Llama, and other models'). That claim lacks a matched comparison with the 2023 hackathon and is therefore unsupported, but unsupported inference from model-version releases is an evidence gap, not circularity. Self-citations appear (the 2023 hackathon report [53]; several appendix projects citing the authors' own datasets or methods, e.g., Learning LOBSTERs uses its earlier LobsterPy dataset and the MOF agent adapts the dZiner approach), but they are contextual or methodological starting points: the reported outcomes are evaluated against external resources (MatBench phonon DOS, QM9, published geopolymer strengths, MaScQA, SDBS NMR spectra) or are honestly reported as failures or non-significant (LLMads hallucinations, Phi-3 breakdown, G-Peer-T p=0.07, LLMSpectrometry 3/19). The closest thing to a closed loop is the MOF agent's use of its own surrogate to accept candidates with lower predicted band gaps and then plot those predictions, but the text labels these as 'inferred band gap' and claims no external validation; that is a transparency and validity limitation of a 24-hour prototype, not a hidden reduction of the paper's central claim to its inputs.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The central report introduces no new physical or mathematical entities. The only notable assumptions are the reliability of self-reported project outcomes and the validity of the author-chosen application taxonomy.

assumptions (2)
  • domain assumption Team self-reports in the appendix accurately describe the capabilities and limitations of the submitted projects.
    The overview and conclusion generalize from the 34 project reports without independent verification. Several reports contain explicit caveats (project 1 lacks cross-validation, project 18 reports hallucination, project 8 reports Phi-3 failure), so this assumption is load-bearing for the paper's conclusions.
  • domain assumption The seven application categories are meaningful and the assignment of projects to categories is correct.
    The analysis, including trends and exemplar projects, depends on this author-defined taxonomy. The categories are conventional and reasonable, but they are not independently derived.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Reflections from the 2024 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry." pith.science (2026). https://pith.science/paper/PGFBIFJG

@misc{pith2026241115221,
  author       = {Pith},
  title        = {Pith review of: Reflections from the 2024 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PGFBIFJG}},
  note         = {Machine review of arXiv:2411.15221}
}
read the original abstract

Here, we present the outcomes from the second Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry, which engaged participants across global hybrid locations, resulting in 34 team submissions. The submissions spanned seven key application areas and demonstrated the diverse utility of LLMs for applications in (1) molecular and material property prediction; (2) molecular and material design; (3) automation and novel interfaces; (4) scientific communication and education; (5) research data management and automation; (6) hypothesis generation and evaluation; and (7) knowledge extraction and reasoning from scientific literature. Each team submission is presented in a summary table with links to the code and as brief papers in the appendix. Beyond team results, we discuss the hackathon event and its hybrid format, which included physical hubs in Toronto, Montreal, San Francisco, Berlin, Lausanne, and Tokyo, alongside a global online hub to enable local and virtual collaboration. Overall, the event highlighted significant improvements in LLM capabilities since the previous year's hackathon, suggesting continued expansion of LLMs for applications in materials science and chemistry research. These outcomes demonstrate the dual utility of LLMs as both multipurpose models for diverse machine learning tasks and platforms for rapid prototyping custom applications in scientific research.

Figures

Figures reproduced from arXiv: 2411.15221 by the authors.

Figure 1
Figure 1. LLM Hackathon for Applications in Materials and Chemistry hybrid hackathon. Researchers were [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Schematic depicting the prompt for fine-tuning the LLM with Alpaca prompt format. [PITH_FULL_IMAGE:figures/full_fig_p015_2.png] view at source ↗
Figure 3
Figure 3. Model architecture and the schema of the second experiment. Material composition is encoded [PITH_FULL_IMAGE:figures/full_fig_p017_3.png] view at source ↗
Figures from the paper (32 more)
Figure 6
Figure 6. Figure 6: Model comparison of the pre-trained and fine-tuned T5-Chem on zero-point vibrational energy [PITH_FULL_IMAGE:figures/full_fig_p020_6.png]
Figure 7
Figure 7. Figure 7: Model comparison of the pre-trained and fine-tuned ChemBERTa on zero-point vibrational energy [PITH_FULL_IMAGE:figures/full_fig_p020_7.png]
Figure 8
Figure 8. Figure 8: Schematic representation of the training process for a regression model illustrating the novel [PITH_FULL_IMAGE:figures/full_fig_p022_8.png]
Figure 9
Figure 9. Figure 9: Performance of an LLM in predicting total energies of benzene and ethanol structures, where the [PITH_FULL_IMAGE:figures/full_fig_p023_9.png]
Figure 10
Figure 10. Figure 10: Scheme of the system. Data is first converted into text format, with peak positions and intensities [PITH_FULL_IMAGE:figures/full_fig_p026_10.png]
Figure 11
Figure 11. Figure 11: MC-Peptide: Pipeline implemented in this work. The example illustrates a user request, followed [PITH_FULL_IMAGE:figures/full_fig_p028_11.png]
Figure 12
Figure 12. Figure 12: a) Workflow overview. The ReAct agent looks up guidelines for designing low band gap MOFs [PITH_FULL_IMAGE:figures/full_fig_p031_12.png]
Figure 13
Figure 13. Figure 13: LLM-based material design workflow (left) and diagram showing evaluation metric (right). [PITH_FULL_IMAGE:figures/full_fig_p036_13.png]
Figure 15
Figure 15. Figure 15: Schematic overview of the LLMicroscopilot assistant. The microscope user interface allows the [PITH_FULL_IMAGE:figures/full_fig_p041_15.png]
Figure 16
Figure 16. Figure 16: Retrieval Augmented Generation [RAG] architecture with LLM interface [PITH_FULL_IMAGE:figures/full_fig_p043_16.png]
Figure 17
Figure 17. Figure 17: Workflow of a response generated to a user prompt by Materials Agent using a tool based on [PITH_FULL_IMAGE:figures/full_fig_p046_17.png]
Figure 18
Figure 18. Figure 18: Niche purpose tools with LLM. (a) Illustration of the RDF computing tool. (b) Code snippet to [PITH_FULL_IMAGE:figures/full_fig_p046_18.png]
Figure 19
Figure 19. Figure 19: Illustration of RAG for interacting with a MSDS. [PITH_FULL_IMAGE:figures/full_fig_p047_19.png]
Figure 20
Figure 20. Figure 20: Workflow for integrating chemical encoders with large language models. Molecular data from [PITH_FULL_IMAGE:figures/full_fig_p049_20.png]
Figure 21
Figure 21. Figure 21: Modified code of the function get embeddings. References [1] S. Chithrananda, G. Grand, and B. Ramsundar, “ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction,” arXiv preprint arXiv:2010.09885, 2020. Available: https://arxiv. org/abs/2…
Figure 22
Figure 22. Figure 22: Modified forward function which allows for the molecular token to be added. [PITH_FULL_IMAGE:figures/full_fig_p051_22.png]
Figure 24
Figure 24. Figure 24: WaterLLM approach: custom chatGPT with RAG from scientific papers, context and chain-of [PITH_FULL_IMAGE:figures/full_fig_p058_24.png]
Figure 25
Figure 25. Figure 25: Sample of the WaterLLM communication with the User. [PITH_FULL_IMAGE:figures/full_fig_p059_25.png]
Figure 26
Figure 26. Figure 26: The yeLLowhaMmer multimodal agent can be used for a variety of data management tasks. Here, [PITH_FULL_IMAGE:figures/full_fig_p061_26.png]
Figure 27
Figure 27. Figure 27: Flowchart of the Query Reporter usage, including the back-end interaction with external resources, [PITH_FULL_IMAGE:figures/full_fig_p064_27.png]
Figure 28
Figure 28. Figure 28: a) Part of a JSON Schema defining a data structure for a solution preparation. b) The schema [PITH_FULL_IMAGE:figures/full_fig_p067_28.png]
Figure 29
Figure 29. Figure 29: Likelihood of accepting the hypothesis “LK-99 is a room-temperature superconductor” via three [PITH_FULL_IMAGE:figures/full_fig_p069_29.png]
Figure 30
Figure 30. Figure 30: Multi-Agent Hypothesis Generation and Verification Pipeline [PITH_FULL_IMAGE:figures/full_fig_p071_30.png]
Figure 31
Figure 31. Figure 31: A Schematic illustration of ActiveScience architecture and its potential applications. Code snippet [PITH_FULL_IMAGE:figures/full_fig_p076_31.png]
Figure 32
Figure 32. Figure 32: Performance of Gemini Pro, GPT-4 Turbo, and Claude3 Opus on text, visual, and text+visual [PITH_FULL_IMAGE:figures/full_fig_p079_32.png]
Figure 33
Figure 33. Figure 33: Summary of our LLM hackathon project. Future Work One major challenge is the low recall of the extracted information, as only 46 out of 152 labeled pieces of information were retrieved. Upon investigating the papers, we found that much of the Coulombic Efficiency was …
Figure 34
Figure 34. Figure 34: KnowMat Workflow. The graphical abstract illustrates the KnowMat workflow, which begins [PITH_FULL_IMAGE:figures/full_fig_p085_34.png]
Figure 36
Figure 36. Figure 36: Creating Knowledge Graph Retrieval-Augmented Generation (KGRAG) for Polymer Simulation. [PITH_FULL_IMAGE:figures/full_fig_p089_36.png]
Figure 37
Figure 37. Figure 37: Insightful machine learning for HEA hydrides [PITH_FULL_IMAGE:figures/full_fig_p092_37.png]
Figure 38
Figure 38. Figure 38: Comparison of the alignment of the different LLMs with the SMILES (left) and IUPAC (right) [PITH_FULL_IMAGE:figures/full_fig_p093_38.png]
Figure 39
Figure 39. Figure 39: Overview of (left) the graphical user interface (GUI) protoype and (right) the generated Neo4J [PITH_FULL_IMAGE:figures/full_fig_p096_39.png]
Figure 40
Figure 40. Figure 40: Overview of the GlossaryGenerator class, responsible for processing text chunks and extracting [PITH_FULL_IMAGE:figures/full_fig_p097_40.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models

    cs.AI 2025-08 unverdicted novelty 6.0 of 10

    SEAM measures VLM reasoning consistency across modalities using paired semantically equivalent textual and visual notations, and finds systematic vision-language imbalance.

  2. Foundational Large Language Models for Materials Research

    cond-mat.mtrl-sci 2024-12 conditional novelty 5.0 of 10

    Domain-adapted LLaMA models (LLaMat) outperform commercial LLMs on materials NLP and structured extraction tasks and generate M3GNet-predicted stable crystals, with LLaMA-2-based variants beating LLaMA-3-based ones.

  3. Multicrossmodal Automated Agent for Integrating Diverse Materials Science Data

    cond-mat.mtrl-sci 2025-05 reject novelty 4.0 of 10

    A prompt-only multi-agent LLM system claims to fuse video, image, table, and text data for materials-science questions, reporting 85% recall and 35% coverage gains, but with unverifiable evaluation.

  4. From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines

    cs.DL 2026-06 unverdicted novelty 3.0 of 10

    LLMs accelerate research workflows from idea generation to writing but introduce challenges like hallucination, bias, opacity, and ten systemic risks requiring new governance frameworks.

Reference graph

Works this paper leans on

235 extracted references · 25 canonical work pages · cited by 4 Pith papers

  1. [1]

    Nolte, L

    A. Nolte, L. B. Hayden and J. D. Herbsleb, Proc. ACM Hum.-Comput. Interact., 2020, 4, 1--23. https://doi.org/10.1145/3392830

  2. [2]

    E. P. P. Pe-Than and J. D. Herbsleb, in Lecture Notes in Computer Science, Springer, 2019, vol. 11546, pp. 27--37. https://doi.org/10.1007/978-3-030-15742-5_3

  3. [3]

    Heller, A

    B. Heller, A. Amir, R. Waxman and Y. Maaravi, J. Innov. Entrep., 2023, 12, 1. https://doi.org/10.1186/s13731-023-00269-0

  4. [5]

    C. Qian, H. Tang, Z. Yang, H. Liang and Y. Liu, arXiv, 2023. https://arxiv.org/abs/2307.07443

  5. [6]

    Jacobs, M

    R. Jacobs, M. P. Polak, L. E. Schultz, H. Mahdavi, V. Honavar and D. Morgan, arXiv, 2024. https://arxiv.org/abs/2409.06080

  6. [7]

    Vacareanu, V

    R. Vacareanu, V. A. Negru, V. Suciu and M. Surdeanu, in First Conference on Language Modeling, 2024. https://openreview.net/forum?id=LzpaUxcNFK

  7. [8]

    A. K. Gupta and K. Raghavachari, J. Chem. Theory Comput., 2022, 18, 2132--2143. https://doi.org/10.1021/acs.jctc.1c00504

  8. [10]

    Lu and Y

    J. Lu and Y. Zhang, J. Chem. Inf. Model., 2022, 62, 1376--1387. https://doi.org/10.1021/acs.jcim.1c01467

Show all 235 references
  1. [11]

    Bhattacharya, H

    D. Bhattacharya, H. J. Cassady, M. A. Hickner and W. F. Reinhart, J. Chem. Inf. Model., 2024, 64, 7086--7096. https://doi.org/10.1021/acs.jcim.4c01396

  2. [12]

    G. Liu, M. Sun, W. Matusik, M. Jiang and J. Chen, arXiv, 2024. https://arxiv.org/abs/2410.04223

  3. [13]

    S. Jia, C. Zhang and V. Fung, arXiv, 2024. https://arxiv.org/abs/2406.13163

  4. [14]

    H. Jang, Y. Jang, J. Kim and S. Ahn, arXiv, 2024. https://arxiv.org/abs/2410.03138

  5. [15]

    J. Lu, Z. Song, Q. Zhao, Y. Du, Y. Cao, H. Jia and C. Duan, arXiv, 2024. https://arxiv.org/abs/2410.18136

  6. [16]

    Kristiadi et al., in Proceedings of the 41st International Conference on Machine Learning, PMLR, 2024, vol

    A. Kristiadi et al., in Proceedings of the 41st International Conference on Machine Learning, PMLR, 2024, vol. 235, pp. 25603--25622. https://proceedings.mlr.press/v235/kristiadi24a.html

  7. [17]

    Miret and N

    S. Miret and N. M. A. Krishnan, arXiv, 2024. https://arxiv.org/abs/2402.05200

  8. [18]

    X. Ji, A. L. Nielsen and C. Heinis, Angew. Chem. Int. Ed., 2023, 63, 3. https://doi.org/10.1002/anie.202308251

  9. [21]

    Dubey et al., arXiv, 2024

    A. Dubey et al., arXiv, 2024. https://arxiv.org/abs/2407.21783

  10. [22]

    Abdin et al., arXiv, 2024

    M. Abdin et al., arXiv, 2024. https://arxiv.org/abs/2404.14219

  11. [23]

    A. M. Bran, S. Cox, O. Schilter, C. Baldassari, A. D. White and P. Schwaller, arXiv, 2023. https://arxiv.org/abs/2304.05376

  12. [24]

    Song et al., arXiv, 2023

    Y. Song et al., arXiv, 2023. https://arxiv.org/abs/2306.06624

  13. [25]

    Zhang, Y

    H. Zhang, Y. Song, Z. Hou, S. Miret and B. Liu, arXiv, 2024. https://arxiv.org/abs/2409.00135

  14. [27]

    Darvish et al., arXiv, 2024

    K. Darvish et al., arXiv, 2024. https://arxiv.org/abs/2401.06949

  15. [28]

    Tom et al., Chem

    G. Tom et al., Chem. Rev., 2024, 124, 9633--9732. https://doi.org/10.1021/acs.chemrev.4c00055

  16. [29]

    Chase, Langchain, 2024

    H. Chase, Langchain, 2024. https://github.com/langchain-ai/langchain

  17. [30]

    RDKit: Open-source cheminformatics, http://www.rdkit.org

  18. [31]

    Yan et al., Br

    L. Yan et al., Br. J. Educ. Technol., 2023, 55, 90--112. https://doi.org/10.1111/bjet.13370

  19. [32]

    Wang et al., arXiv, 2024

    S. Wang et al., arXiv, 2024. https://arxiv.org/abs/2403.18105

  20. [33]

    Kasneci et al., Learn

    E. Kasneci et al., Learn. Individ. Differ., 2023, 103, 102274. https://doi.org/10.1016/j.lindif.2023.102274

  21. [34]

    M. S. Schäfer, J. Sci. Commun., 2023, 22, 2. https://doi.org/10.22323/2.22020402

  22. [35]

    Zaki, Jayadeva, Mausam and N

    M. Zaki, Jayadeva, Mausam and N. M. A. Krishnan, arXiv, 2023. https://arxiv.org/abs/2308.09115

  23. [36]

    https://assets.anthropic.com/m/61e7d27f8c8f5919/original/Claude-3-Model-Card.pdf

    Anthropic, The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024. https://assets.anthropic.com/m/61e7d27f8c8f5919/original/Claude-3-Model-Card.pdf

  24. [37]

    A. Q. Jiang et al., arXiv, 2024. https://arxiv.org/abs/2401.04088

  25. [38]

    Draxl and M

    C. Draxl and M. Scheffler, J. Phys. Mater., 2019, 2, 036001. https://doi.org/10.1088/2515-7639/ab13bb

  26. [39]

    Radford et al., arXiv, 2022

    A. Radford et al., arXiv, 2022. https://arxiv.org/abs/2212.04356

  27. [40]

    Y. Zhou, H. Liu, T. Srivastava, H. Mei and C. Tan, arXiv, 2024. https://arxiv.org/abs/2404.04326

  28. [41]

    Abdel-Rehim et al., arXiv, 2024

    A. Abdel-Rehim et al., arXiv, 2024. https://arxiv.org/abs/2405.12258

  29. [42]

    S. Tong, K. Mao, Z. Huang, Y. Zhao and K. Peng, Humanit. Soc. Sci. Commun., 2024, 11, 1. https://doi.org/10.1057/s41599-024-03407-5

  30. [43]

    Ciucă, Y.-S

    I. Ciucă, Y.-S. Ting, S. Kruk and K. Iyer, arXiv, 2023. https://arxiv.org/abs/2306.11648

  31. [44]

    Liu et al., arXiv, 2024

    Q. Liu et al., arXiv, 2024. https://arxiv.org/abs/2409.06756

  32. [45]

    Shir, ChemRxiv, 2024

    O. Shir, ChemRxiv, 2024. https://doi.org/10.26434/chemrxiv-2024-lf2xx

  33. [46]

    S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao and K. Narasimhan, in Advances in Neural Information Processing Systems, Curran Associates, Inc., 2023, vol. 36, pp. 11809--11822. https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac7...

  34. [47]

    Shamsabadi, J

    M. Shamsabadi, J. D'Souza and S. Auer, arXiv, 2024. https://arxiv.org/abs/2401.10040

  35. [48]

    Dagdelen et al., Nat

    J. Dagdelen et al., Nat. Commun., 2024, 15, 1. https://doi.org/10.1038/s41467-024-45563-x

  36. [49]

    Xu et al., arXiv, 2024

    D. Xu et al., arXiv, 2024. https://arxiv.org/abs/2312.17617

  37. [50]

    J. Li, M. Zhang, N. Li, D. Weyns, Z. Jin and K. Tei, ACM Trans. Auton. Adapt. Syst., 2024, 19, 1--60. https://doi.org/10.1145/3686803

  38. [51]

    Ma et al., arXiv, 2024

    Y. Ma et al., arXiv, 2024. https://arxiv.org/abs/2402.11451

  39. [52]

    https://arxiv.org/abs/2303.08774

    OpenAI et al., arXiv, 2023. https://arxiv.org/abs/2303.08774

  40. [54]

    K. M. Jablonka, P. Schwaller, A. Ortega-Guerrero, B. Smit, Nat Mach Intell, 2024, 6, 161–169

  41. [55]

    Choudhary, J

    K. Choudhary, J. Phys. Chem. Lett., 2024, 6909–6917

  42. [56]

    Petretto, S

    G. Petretto, S. Dwaraknath, H. P.C. Miranda, D. Winston, M. Giantomassi, M. J. van Setten, X. Gonze, K. A. Persson, G. Hautier, G.-M. Rignanese, Sci Data, 2018, 5, 180065

  43. [57]

    A. Dunn, Q. Wang, A. Ganose, D. Dopp, A. Jain, npj Comput Mater, 2020, 6, 1–10

  44. [58]

    A. M. Ganose, A. Jain, MRS Communications, 2019, 9, 874–881

  45. [59]

    H. M. Sayeed, S. G. Baird, T. D. Sparks, 2023, DOI 10.26434/chemrxiv-2023-3q8wj

  46. [60]

    A. N. Rubungo, C. Arnold, B. P. Rand, A. B. Dieng, 2023, DOI 10.48550/arXiv.2310.14029

  47. [61]

    V. Moro, C. Loh, R. Dangovski, A. Ghorashi, A. Ma, Z. Chen, S. Kim, P. Y. Lu, T. Christensen, M. Soljačić, 2024, DOI 10.48550/arXiv.2312.00111

  48. [62]

    A. A. Naik, K. Ueltzen, C. Ertural, A. J. Jackson, J. George, Journal of Open Source Software, 2024, 9, 6286

  49. [63]

    A. A. Naik, C. Ertural, N. Dhamrait, P. Benner, J. George, 2023, DOI 10.5281/zenodo.8091844

  50. [64]

    The Matbench Test Suite, Phonon dataset as per 12.07.2024, https://matbench.materialsproject.org/Leaderboards

  51. [65]

    Raffel, N

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, P. J. Liu, Journal of Machine Learning Research, 2020, 21, 1–67

  52. [66]

    The Unsloth package, https://github.com/unslothai/unsloth, 2024

  53. [67]

    Goodall, A

    R. Goodall, A. J. et al., ``Predicting materials properties without crystal structure: deep representation learning from stoichiometry'', Nature Communications, vol. 11, no. 6280, 2020

  54. [68]

    Trewartha, et al., 'Quantifying the advantage of domain-specific pre-training on named entity recognition tasks in materials science', Patterns, vol

    A. Trewartha, et al., 'Quantifying the advantage of domain-specific pre-training on named entity recognition tasks in materials science', Patterns, vol. 3, no. 8, 2022

  55. [69]

    Gupta, et al., MatSciBERT: A materials domain language model for text mining and information extraction, npj Computational Materials, vol

    V. Gupta, et al., MatSciBERT: A materials domain language model for text mining and information extraction, npj Computational Materials, vol. 8, no. 1, 2022

  56. [70]

    Matbench Perovskites dataset provided by the Materials Project, https://ml.materialsproject.org/projects/matbench_perovskites.json.gz

  57. [71]

    Li-ion Conductors Database, https://pcwww.liv.ac.uk/ msd30/lmds/LiIonDatabase.html

  58. [72]

    PyMuPDF, https://pypi.org/project/PyMuPDF

  59. [73]

    AI@Meta, Llama 3 Model Card, 2024, https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md

  60. [74]

    Hargreaves et al., The Earth Mover’s Distance as a Metric for the Space of Inorganic Compositions, Chem. Mater. 2020

  61. [75]

    Hargreaves et al., A Database of Experimentally Measured Lithium Solid Electrolyte Conductivities Evaluated with Machine Learning, npj Computational Materials 2022

  62. [77]

    M., Ai, Q., Al-Feghali, A., Badhwar, S., Bocarsly, J

    Jablonka, K. M., Ai, Q., Al-Feghali, A., Badhwar, S., Bocarsly, J. D., Bran, A. M., ... & Blaiszik, B. (2023). 14 examples of how LLMs can transform materials science and chemistry: a reflection on a large language model hackathon. Digital Discovery, 2(5), 1233-1250

  63. [78]

    Alampara, N., Miret, S., & Jablonka, K. M. (2024). MatText: Do Language Models Need More than Text & Scale for Materials Modeling? arXiv preprint arXiv:2406.17295

  64. [79]

    P., Kornbluth, M.,

    Batzner, S., Musaelian, A., Sun, L., Geiger, M., Mailoa, J. P., Kornbluth, M., ... & Kozinsky, B. (2022). E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials. Nature Communications, 13(1), 2453

  65. [80]

    K., & Raghavachari, K

    Gupta, A. K., & Raghavachari, K. (2022). Three-dimensional convolutional neural networks utilizing molecular topological features for accurate atomization energy predictions. Journal of Chemical Theory and Computation, 18(4), 2132-2143

  66. [81]

    Grzegorz Kaszuba, Amirhossein D

    In-Context Learning of Physical Properties: Few-Shot Adaptation to Out-of-Distribution Molecular Graphs. Grzegorz Kaszuba, Amirhossein D. Naghdi, Dario Massa, Stefanos Papanikolaou, Andrzej Jaszkiewicz, Piotr Sankowski, https://arxiv.org/abs/2406.01808

  67. [82]

    Khan, D., Heinen, S., & von Lilienfeld, O. A. (2023). Kernel based quantum machine learning at record rate: Many-body distribution functionals as compact representations. Journal of Chemical Physics, 159(034106)

  68. [83]

    T., Kabylda, A., Sauceda, H

    Chmiela, S., Vassilev-Galindo, V., Unke, O. T., Kabylda, A., Sauceda, H. E., Tkatchenko, A., & Müller, K. R. (2023). Accurate global machine learning force fields for molecules with hundreds of atoms. Science Advances, 9(2), https://doi.org/10.1126/sciadv.adf0873

  69. [84]

    M., Qu, C., Conte, R., Nandi, A., Houston, P

    Bowman, J. M., Qu, C., Conte, R., Nandi, A., Houston, P. L., & Yu, Q. (2022). The MD17 datasets from the perspective of datasets for gas-phase “small” molecule potentials. The Journal of chemical physics, 156(24)

  70. [85]

    Weinreich, J., & Probst, D. (2023). Parameter-Free Molecular Classification and Regression with Gzip. ChemRxiv

  71. [86]

    P., Simm, G., Ortner, C., & Csányi, G

    Batatia, I., Kovacs, D. P., Simm, G., Ortner, C., & Csányi, G. (2022). MACE: Higher order equivariant message passing neural networks for fast and accurate force fields. Advances in Neural Information Processing Systems, 35, 11423-11436

  72. [87]

    A., ACS Central Sci., 2019, Vol

    Schwaller, P., Laino, T., Gaudin, T., Bolgar, P., Hunter, C., Bekas, C., Lee, A. A., ACS Central Sci., 2019, Vol. 5, No. 9, 1572-1583. https://pubs.acs.org/doi/10.1021/acscentsci.9b00576

  73. [88]

    C., ChemRxiv Preprint

    Alberts, M., Zipoli, F., Vaucher, A. C., ChemRxiv Preprint. https://doi.org/10.26434/chemrxiv-2023-8wxcz

  74. [90]

    https://sdbs.db.aist.go.jp

    Yamaji, T., Saito, T., Hayamizu, K., Yanagisawa, M., Yamamoto, O., Wasada, N., Someno, K., Kinugasa, S., Tanabe, K., Tamura, T., Hiraishi, J., 2024. https://sdbs.db.aist.go.jp

  75. [91]

    Socha, O., Osifova, Z., Dracinsky, M., J. Chem. Educ., 2023, Vol. 100, No. 2, 962-968. https://pubs.acs.org/doi/10.1021/acs.jchemed.2c01067

  76. [92]

    https://github.com/ATOMSLab/LLMSpectroscopy

  77. [93]

    Cyclic peptides for drug development,

    Ji, X., Nielsen, A. L., Heinis, C., "Cyclic peptides for drug development," Angewandte Chemie International Edition, 2024, 63(3), e202308251

  78. [94]

    De novo development of small cyclic peptides that are orally bioavailable,

    Merz, M.L., Habeshian, S., Li, B. et al., "De novo development of small cyclic peptides that are orally bioavailable," Nat Chem Biol, 2024, 20, 624–633. https://doi.org/10.1038/s41589-023-01496-y

  79. [95]

    Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation,

    Beurer-Kellner, L., et al., "Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation," ArXiv, 2024, abs/2403.06988

  80. [96]

    A survey on in-context learning,

    Dong, Q., et al., "A survey on in-context learning," ArXiv, 2022, arXiv:2301.00234

  81. [97]

    Many-shot in-context learning,

    Agarwal, R., et al., "Many-shot in-context learning," ArXiv, 2024, arXiv:2404.11018

  82. [98]

    A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?,

    Kristiadi, A., et al., "A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?," ArXiv, 2024, arXiv:2402.05015

  83. [99]

    A Detailed Investigation on Conformation, Permeability and PK Properties of Two Related Cyclohexapeptides,

    Lewis, I., Schaefer, M., Wagner, T. et al., "A Detailed Investigation on Conformation, Permeability and PK Properties of Two Related Cyclohexapeptides," Int J Pept Res Ther, 2015, 21, 205–221. https://doi.org/10.1007/s10989-014-9447-3

  84. [100]

    Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,

    Lewis, P., et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks," ArXiv, 2020, abs/2005.11401

  85. [101]

    Reducing hallucination in structured outputs via Retrieval-Augmented Generation,

    Béchard, P., Ayala, O. M., "Reducing hallucination in structured outputs via Retrieval-Augmented Generation," ArXiv, 2024, abs/2404.08189

  86. [102]

    Review on applications of metal–organic frameworks for co2 capture and the performance enhancement mechanisms

    Lirong Li, Han Sol Jung, Jae Won Lee, and Yong Tae Kang. Review on applications of metal–organic frameworks for co2 capture and the performance enhancement mechanisms. Renewable and Sustainable Energy Reviews, 162: 112441, 2022

  87. [103]

    React: Synergizing reasoning and acting in language models

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. arXiv preprint, arXiv:2210.03629, 2022

  88. [104]

    Brown, and Joseph S

    Mehrad Ansari, Jeffrey Watchorn, Carla E. Brown, and Joseph S. Brown. dZiner: Rational Inverse Design of Materials with AI Agents. arXiv, 2410.03963, 2024. URL: https://arxiv.org/abs/2410.03963

  89. [105]

    Semiconductor metal–organic frameworks: future low- g"bandgap materials

    Muhammad Usman, Shruti Mendiratta, and Kuang-Lieh Lu. Semiconductor metal–organic frameworks: future low- g"bandgap materials. Advanced Materials, 29(6):1605071, 2017

  90. [106]

    Band gap modulations in uio metal–organic frameworks

    Espen Flage-Larsen, Arne Røyset, Jasmina Hafizovic Cavka, and Knut Thorshaug. Band gap modulations in uio metal–organic frameworks. The Journal of Physical Chemistry C, 117(40):20610–20616, 2013

  91. [107]

    Band gap engineering of paradigm mof-5

    Li-Ming Yang, Guo-Yong Fang, Jing Ma, Eric Ganz, and Sang Soo Han. Band gap engineering of paradigm mof-5. Crystal growth & design, 14(5):2532–2541, 2014

  92. [108]

    Theoretical investigations on the chemical bonding, electronic structure, and optical properties of the metal- organic framework mof-5

    Li-Ming Yang, Ponniah Vajeeston, Ponniah Ravindran, Helmer Fjellvag, and Mats Tilset. Theoretical investigations on the chemical bonding, electronic structure, and optical properties of the metal- organic framework mof-5. Inorganic chemistry, 49(22):10283–10290, 2010

  93. [109]

    Recent advancements in mof-based catalysts for applications in electrochemical and photoelectrochemical water splitting: A review

    Maryum Ali, Erum Pervaiz, Tayyaba Noor, Osama Rabi, Rubab Zahra, and Minghui Yang. Recent advancements in mof-based catalysts for applications in electrochemical and photoelectrochemical water splitting: A review. International Journal of Energy Research, 45(2):1190–1226, 2021

  94. [110]

    Tuning electrical and mechanical properties of metal–organic frameworks by metal substitution

    Yabin Yan, Chunyu Wang, Zhengqing Cai, Xiaoyuan Wang, and Fuzhen Xuan. Tuning electrical and mechanical properties of metal–organic frameworks by metal substitution. ACS Applied Materials & Interfaces, 15(36):42845–42853, 2023

  95. [111]

    Tunability of band gaps in metal–organic frameworks

    Chi-Kai Lin, Dan Zhao, Wen-Yang Gao, Zhenzhen Yang, Jingyun Ye, Tao Xu, Qingfeng Ge, Shengqian Ma, and Di-Jia Liu. Tunability of band gaps in metal–organic frameworks. Inorganic chemistry, 51(16):9039–9044, 2012

  96. [112]

    New and improved embedding model, 2022

    Ryan Greene, Ted Sanders, Lilian Weng, and Arvind Neelakantan. New and improved embedding model, 2022

  97. [113]

    Agent-based learning of materials datasets from scientific literature

    Mehrad Ansari and Seyed Mohamad Moosavi. Agent-based learning of materials datasets from scientific literature. arXiv preprint, arXiv:2312.11690, 2023

  98. [114]

    Moformer: self-supervised transformer model for metal–organic framework property prediction

    Zhonglin Cao, Rishikesh Magar, Yuyang Wang, and Amir Barati Farimani. Moformer: self-supervised transformer model for metal–organic framework property prediction. Journal of the American Chemical Society, 145(5): 2958–2967, 2023

  99. [115]

    Barlow twins: Self-supervised learning via redundancy reduction

    Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. Barlow twins: Self-supervised learning via redundancy reduction. In International Conference on Machine Learning, pages 12310–12320. PMLR, 2021

  100. [116]

    Grossman

    Tian Xie and Jeffrey C. Grossman. Crystal graph convolutional neural networks for an accurate and interpretable prediction of material properties. Physical Review Letters, 120(14):145301, 2018

  101. [117]

    Understanding the diversity of the metal-organic framework ecosystem

    Seyed Mohamad Moosavi, Aditya Nandy, Kevin Maik Jablonka, Daniele Ongari, Jon Paul Janet, Peter G Boyd, Yongjin Lee, Berend Smit, and Heather J Kulik. Understanding the diversity of the metal-organic framework ecosystem. Nature communications, 11(1):1–10, 2020

  102. [118]

    Rosen, Shaelyn M

    Andrew S. Rosen, Shaelyn M. Iyer, Debmalya Ray, Zhenpeng Yao, Alan Aspuru-Guzik, Laura Gagliardi, Justin M. Notestein, and Randall Q. Snurr. Machine learning the quantum-chemical properties of metal–organic frameworks for accelerated materials discovery. Matter, 4(5):1578–1597, 2021

  103. [119]

    RDKit documentation

    Greg Landrum. RDKit documentation. Release, 1(1-79):4, 2013

  104. [120]

    Gpt-4 technical report, 2023

    OpenAI. Gpt-4 technical report, 2023

  105. [121]

    LangChain, 10 2022

    Harrison Chase. LangChain, 10 2022. URL https://github.com/langchain-ai/langchain

  106. [122]

    Eco-efficient cements: Potential economically viable solutions for a low-CO2 cement-based materials industry,

    U. Environment, K. L. Scrivener, V. M. John and E. M. Gartner, "Eco-efficient cements: Potential economically viable solutions for a low-CO2 cement-based materials industry," Cement and Concrete Research, vol. 114, pp. 2-26; DOI: https://doi.org/10.1016/j.cemconres.2018.03.015, 2018

  107. [123]

    Advances in understanding alkali-activated materials,

    J. L. Provis, A. Palomo and C. Shi, "Advances in understanding alkali-activated materials," Cement and Concrete Research, pp. 110-125, 2015

  108. [124]

    Development of Eco-Efficient Fly Ash–Based Alkali-Activated and Geopolymer Composites with Reduced Alkaline Activator Dosage,

    H. S. Gökçe, M. Tuyan, K. Ramyar and M. L. Nehdi, "Development of Eco-Efficient Fly Ash–Based Alkali-Activated and Geopolymer Composites with Reduced Alkaline Activator Dosage," Journal of Materials in Civil Engineering, vol. 32, no. 2, pp. 04019350; DOI: 10.1061/(ASCE)MT.1943...

  109. [125]

    Synthesis and characterization of red mud and rice husk ash-based geopolymer composites,

    J. He, Y. Jie, J. Zhang, Y. Yu and G. Zhang, "Synthesis and characterization of red mud and rice husk ash-based geopolymer composites," Cement and Concrete Composites, vol. 37, pp. 108-118; DOI: http://dx.doi.org/10.1016/j.cemconcomp.2012.11.010, 2013

  110. [126]

    LLMs can Design Sustainable Concrete – a Systematic Benchmark,

    C. Völker, T. Rug, K. M. Jablonbka and S. Kurschwitz, "LLMs can Design Sustainable Concrete – a Systematic Benchmark," (Preprint), pp. 1-12; DOI: 10.21203/rs.3.rs-3913272/v1, 2023

  111. [127]

    A quantitative method of approach in designing the mix proportions of fly ash and GGBS-based geopolymer concrete,

    G. M. Rao and T. D. G. Rao, "A quantitative method of approach in designing the mix proportions of fly ash and GGBS-based geopolymer concrete," Australian Journal of Civil Engineering, vol. 16, no. 1, pp. 53-63; DOI: 10.1080/14488353.2018.1450716, 2018

  112. [128]

    Boiko, D.A., MacKnight, R., Kline, B. et al. Autonomous chemical research with large language models. Nature 624, 570–578 (2023). https://doi.org/10.1038/s41586-023-06792-0

  113. [129]

    Bran, A., Cox, S., Schilter, O

    M. Bran, A., Cox, S., Schilter, O. et al. Augmenting large language models with chemistry tools. Nat Mach Intell 6, 525–535 (2024). https://doi.org/10.1038/s42256-024-00832-8

  114. [130]

    E., Christensen, R., Dułak, M., Friis, J., Groves, M

    Hjorth Larsen, A., Jørgen Mortensen, J., Blomqvist, J., Castelli, I. E., Christensen, R., Dułak, M., Friis, J., Groves, M. N., Hammer, B., Hargus, C., Hermes, E. D., Jennings, P. C., Bjerre Jensen, P., Kermode, J., Kitchin, J. R., Leonhard Kolsbjerg, E., Kubal, J., Kaasbjerg, ...

  115. [131]

    W., Stoltze, P., Nørskov, J

    Jacobsen, K. W., Stoltze, P., Nørskov, J. K. (1996). A semi-empirical effective medium theory for metals and alloys. Surface Science, 366(2), 394–402. https://doi.org/10.1016/0039-6028(96)00816-3

  116. [132]

    M., Kovács, D

    Batatia, I., Benner, P., Chiang, Y., Elena, A. M., Kovács, D. P., Riebesell, J., Advincula, X. R., Asta, M., Avaylon, M., Baldwin, W. J., Berger, F., Bernstein, N., Bhowmik, A., Blau, S. M., Cărare, V., Darby, J. P., De, S., Della Pia, F., Deringer, V. L. et al. (2024). A foun...

  117. [133]

    https://github.com/langchain-ai/langchain

  118. [134]

    https://pydantic.dev/

  119. [135]

    Ye, H., Liu, T., Zhang, A., Hua, W., Jia, W. (2023). Cognitive mirage: A review of hallucinations in large language models. arXiv. https://arxiv.org/abs/2309.06794

  120. [136]

    Stefan Bauer et al, Roadmap on data-centric materials science, Modelling Simul. Mater. Sci. Eng., 2024, 32, 063301

  121. [137]

    Leveraging Large Language Models and Social Media for Automation in Scanning Probe Microscopy

    Diao, Zhuo, Hayato Yamashita, and Masayuki Abe. "Leveraging Large Language Models and Social Media for Automation in Scanning Probe Microscopy." arXiv preprint arXiv:2405.15490 (2024)

  122. [138]

    Synergizing Human Expertise and AI Efficiency with Language Model for Microscopy Operation and Automated Experiment Design

    Liu, Yongtao, Marti Checa, and Rama K. Vasudevan. "Synergizing Human Expertise and AI Efficiency with Language Model for Microscopy Operation and Automated Experiment Design." Machine Learning: Science and Technology (2024)

  123. [139]

    The abTEM code: transmission electron microscopy from first principles

    Madsen, Jacob, and Toma Susi. "The abTEM code: transmission electron microscopy from first principles." Open Research Europe 1 (2021)

  124. [140]

    Nion Swift: Open Source Image Processing Software for Instrument Control, Data Acquisition, Organization, Visualization, and Analysis Using Python

    Meyer, Chris, et al. "Nion Swift: Open Source Image Processing Software for Instrument Control, Data Acquisition, Organization, Visualization, and Analysis Using Python." Microscopy and Microanalysis 25.S2 (2019): 122-123

  125. [142]

    Ab‐initio simulations of materials using VASP: Density‐functional theory and beyond

    Hafner, Jürgen, "Ab‐initio simulations of materials using VASP: Density‐functional theory and beyond." Journal of computational chemistry 29.13 (2008): 2044-2078

  126. [143]

    Ab initio molecular dynamics for liquid metals

    Kresse, Georg, and Jürgen Hafner, "Ab initio molecular dynamics for liquid metals." Physical review B 47.1 (1993): 558

  127. [144]

    Real-space grid implementation of the projector augmented wave method

    Mortensen, Jens Jørgen, Lars Bruno Hansen, and Karsten Wedel Jacobsen, "Real-space grid implementation of the projector augmented wave method." Physical Review B—Condensed Matter and Materials Physics 71.3 (2005): 035109

  128. [146]

    Materials modelling using density functional theory: properties and predictions

    Giustino, Feliciano. Materials modelling using density functional theory: properties and predictions. Oxford University Press, 2014

  129. [147]

    LlamaIndex, https://docs.llamaindex.ai/en/stable/examples/llm/llama_2_llama_cpp/

  130. [148]

    Mistral 7B, the model used, https://huggingface.co/TheBloke/Mistral-7B-Instruct-v0.1-GGUF

  131. [149]

    OpenAI plugins, https://openai.com/index/chatgpt-plugins/

  132. [151]

    LangChain, https://www.langchain.com/

  133. [152]

    OpenAI models, https://platform.openai.com/docs/models

  134. [153]

    Fast Dash, https://docs.fastdash.app/

  135. [154]

    RDKit: Open-source cheminformatics; http://www.rdkit.org

  136. [155]

    Embedchain, https://github.com/embedchain/embedchain

    Singh, Taranjeet. Embedchain, https://github.com/embedchain/embedchain

  137. [156]

    PubChem programmatic access, https://pubchem.ncbi.nlm.nih.gov/docs/programmatic-access

  138. [157]

    PubChemPy, https://pubchempy.readthedocs.io/en/latest/guide/introduction.html

  139. [158]

    E., Jalalypour, F., Jordan, J., Kutzner, C., Lemkul, J

    Abraham, M., Alekseenko, A., Basov, V., Bergh, C., Briand, E., Brown, A., Doijade, M., Fiorin, G., Fleischmann, S., Gorelov, S., Gouaillardet, G., Grey, A., Irrgang, M. E., Jalalypour, F., Jordan, J., Kutzner, C., Lemkul, J. A., Lundborg, M., Merz, P., … Lindahl, E. (2024). GR...

  140. [159]

    E., & Snurr, R

    Dubbeldam, D., Calero, S., Ellis, D. E., & Snurr, R. Q. (2015). RASPA: molecular simulation software for adsorption and diffusion in flexible nanoporous materials. Molecular Simulation, 42(2), 81–101. https://doi.org/10.1080/08927022.2015.1010082

  141. [160]

    L., Cococcioni, M., Dabo, I., Dal Corso, A., de Gironcoli, S., Fabris, S., Fratesi, G., Gebauer, R., Gerstmann, U., Gougoussis, C., Kokalj, A., Lazzeri, M., … Wentzcovitch, R

    Giannozzi, P., Baroni, S., Bonini, N., Calandra, M., Car, R., Cavazzoni, C., Ceresoli, D., Chiarotti, G. L., Cococcioni, M., Dabo, I., Dal Corso, A., de Gironcoli, S., Fabris, S., Fratesi, G., Gebauer, R., Gerstmann, U., Gougoussis, C., Kokalj, A., Lazzeri, M., … Wentzcovitch,...

  142. [161]

    Chithrananda, G

    S. Chithrananda, G. Grand, and B. Ramsundar, ``ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction,'' arXiv preprint arXiv:2010.09885, 2020. Available: https://arxiv.org/abs/2010.09885

  143. [162]

    G. Zhou, Z. Gao, Q. Ding, H. Zheng, H. Xu, Z. Wei, et al., ``Uni-Mol: A Universal 3D Molecular Representation Learning Framework,'' ChemRxiv, 2022, doi:10.26434/chemrxiv-2022-jjm0j

  144. [163]

    B. Yu, F. N. Baker, Z. Chen, X. Ning, and H. Sun, ``LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset,'' arXiv preprint arXiv:2402.09391, 2024. Available: https://arxiv.org/abs/2402.09391

  145. [164]

    S. Liu, J. Wang, Y. Yang, C. Wang, L. Liu, H. Guo, and C. Xiao, ``ChatGPT-powered Conversational Drug Editing Using Retrieval and Domain Feedback,'' arXiv preprint arXiv:2305.18090, 2023. Available: https://arxiv.org/abs/2305.18090

  146. [165]

    A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. Le Scao, T. Lavril, T. Wang, T. Lacroix, and W. El Sayed, ``Mistral 7B,'' arXiv preprint arXiv:2310.0...

  147. [166]

    Dettmers, A

    T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, ``QLoRA: Efficient Finetuning of Quantized LLMs,'' arXiv preprint arXiv:2305.14314, 2023. Available: https://arxiv.org/abs/2305.14314

  148. [167]

    Zaki, M., & Krishnan, N. A. (2024). MaScQA: investigating materials science knowledge of large language models. Digital Discovery, 3(2), 313-327

  149. [168]

    OpenAI. (2024). GPT-3.5-turbo. https://openai.com/api/

  150. [169]

    Marp. (2024). Markdown Presentation Ecosystem. https://marp.app/

  151. [170]

    K.; Mentha, S

    Mishra, R. K.; Mentha, S. S.; Misra, Y.; Dwivedi, N. Emerging Pollutants of Severe Environmental Concern in Water and Wastewater: A Comprehensive Review on Current Developments and Future Research. Water-Energy Nexus 2023, 6, 74–95. https://doi.org/10.1016/j.wen.2023.08.002

  152. [171]

    A Microscopic Survey on Microplastics in Beverages: The Case of Beer, Mineral Water and Tea

    Li, Y.; Peng, L.; Fu, J.; Dai, X.; Wang, G. A Microscopic Survey on Microplastics in Beverages: The Case of Beer, Mineral Water and Tea. Analyst 2022, 147 (6), 1099–1105. https://doi.org/10.1039/D2AN00083K

  153. [172]

    M. L. Evans and J. D. Bocarsly. datalab, July 2024. URL https://github.com/datalab-org doi:10.5281/zenodo.12545475

  154. [173]

    14 examples of how LLMs can transform materials science and chemistry: a reflection on a large language model hackathon

    Jablonka et al. 14 examples of how LLMs can transform materials science and chemistry: a reflection on a large language model hackathon. Digital Discovery, 2023. doi:10.1039/D3DD00113J

  155. [175]

    https://github.com/ndaelman-hu/nomad_query_reporter

  156. [176]

    Gao, Y., et al., arXiv (2024), doi.org/10.48550/arXiv.2312.10997

    Retrieval-Augmented Generation for Large Language Models: A Survey. Gao, Y., et al., arXiv (2024), doi.org/10.48550/arXiv.2312.10997

  157. [177]

    Llama: open and efficient foundation language models. H. Touvron, et al., arXiv (2023), doi.org/10.48550/arXiv.2302.13971

  158. [178]

    Evans, M

    Development and applications of the OPTIMADE API for materials discovery, design, and data exchange. Evans, M. L., et al., Digital Discovery (2024), DOI: doi.org/10.1039/D4DD00039K

  159. [179]

    Wilkinson, M., Dumontier, M., Aalbersberg, I. et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data 3, 160018 (2016) https://doi.org/10.1038/sdata.2016.18

  160. [180]

    https://json-schema.org/

  161. [181]

    NOMAD: A distributed web-based platform for managing materials science research data

    Scheidgen et al., (2023). NOMAD: A distributed web-based platform for managing materials science research data. Journal of Open Source Software, 8(90), 5388, https://doi.org/10.21105/joss.05388

  162. [182]

    https://pypi.org/project/SpeechRecognition/

  163. [183]

    https://pypi.org/project/openai-whisper/

  164. [184]

    https://pypi.org/project/langchain-experimental/

  165. [185]

    Active learning literature survey

    Settles, Burr. "Active learning literature survey." (2009)

  166. [186]

    Romano, and George Casella

    Lehmann, Erich Leo, Joseph P. Romano, and George Casella. Testing statistical hypotheses. Vol. 3. New York: springer, 1986

  167. [187]

    The logic of scientific discovery

    Popper, Karl. The logic of scientific discovery. Routledge, 2005

  168. [188]

    The first room-temperature ambient-pressure superconductor

    Lee, Sukbae, Ji-Hoon Kim, and Young-Wan Kwon. "The first room-temperature ambient-pressure superconductor." arXiv preprint arXiv:2307.12008 (2023)

  169. [189]

    LK-99 Is the Superconductor of the Summer

    Chang, Kenneth. “LK-99 Is the Superconductor of the Summer.” New York Times (2023)

  170. [190]

    Claimed superconductor LK-99 is an online sensation—But replication efforts fall short

    Garisto, Dan. "Claimed superconductor LK-99 is an online sensation—But replication efforts fall short." Nature 620, no. 7973 (2023): 253-253

  171. [191]

    Natural language inference in context-investigating contextual reasoning over long texts

    Liu, Hanmeng, Leyang Cui, Jian Liu, and Yue Zhang. "Natural language inference in context-investigating contextual reasoning over long texts." In Proceedings of the AAAI conference on artificial intelligence, vol. 35, no. 15, pp. 13388-13396. 2021

  172. [192]

    Conjugate Bayesian analysis of the Gaussian distribution

    Murphy, Kevin P. "Conjugate Bayesian analysis of the Gaussian distribution." def 1, no. 2 2 (2007): 16

  173. [193]

    Will the LK-99 room temp superconductivity pre-print replicate in 2023

    “Will the LK-99 room temp superconductivity pre-print replicate in 2023”. Manifold Markets (2024). https://manifold.markets/Ernie/will-the-lk99-room-temp-ambient-pre-17fc7cb7a2a0

  174. [194]

    Large Language Models for Automated Open-domain Scientific Hypotheses Discovery

    Yang, Zonglin, et al. "Large Language Models for Automated Open-domain Scientific Hypotheses Discovery." arXiv preprint arXiv:2309.02726 (2023). https://arxiv.org/pdf/2309.02726

  175. [195]

    Tree of thoughts: Deliberate problem solving with large language models

    Yao, Shunyu, et al. "Tree of thoughts: Deliberate problem solving with large language models." Advances in Neural Information Processing Systems 36 (2024)

  176. [196]

    https://huggingface.co/datasets/AtlasUnified/Atlas-Reasoning/commits/main

  177. [197]

    https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2

  178. [198]

    K. M. Jablonka, P. Schwaller, A. Ortega-Guerrero, B. Smit, Nat. Mach. Intell., 2024, 6, 161–169

  179. [199]

    D. A. Boiko, R. MacKnight, B. Kline, G. Gomes, Nature, 2023, 624, 570–578

  180. [200]

    Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, C. Zhu, Proc. 2023 Conf. Empir. Methods Nat. Lang. Process., 2023, 2511–2522

  181. [201]

    Mangrulkar, S

    S. Mangrulkar, S. Gugger, L. Debut, Y. Belkada, S. Paul, B. Bossan, PEFT: State-of-the-art Parameter-Efficient Fine-Tuning Methods; GitHub: https://github.com/huggingface/peft, 2022

  182. [202]

    Zhang, G

    P. Zhang, G. Zeng, T. Wang, W. Lu, TinyLlama: An Open-Source Small Language Model; arXiv:2401.02385

  183. [203]

    Zhang, S

    S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, L. Zettlemoyer, OPT: Open Pre-trained Transformer Language Models; arXiv:2205.01068

  184. [204]

    Bethesda (MD): National Library of Medicine (US), National Center for Biotechnology Information; https://www.ncbi.nlm.nih.gov/, 1998

    National Center for Biotechnology Information (NCBI) [Internet]. Bethesda (MD): National Library of Medicine (US), National Center for Biotechnology Information; https://www.ncbi.nlm.nih.gov/, 1998

  185. [205]

    Al-Feghali, S

    A. Al-Feghali, S. Zhang, G-Peer-T; GitHub: https://github.com/alxfgh/G-Peer-T, 2024

  186. [206]

    Translation between molecules and natural language

    Carl Edwards, Tuan Lai, Kevin Ros, Garrett Honke, Kyunghyun Cho, and Heng Ji. Translation between molecules and natural language. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing...

  187. [207]

    Isobench: Benchmarking multimodal foundation models on isomorphic representations, 2024

    Deqing Fu, Ghazal Khalighinejad, Ollie Liu, Bhuwan Dhingra, Dani Yogatama, Robin Jia, and Willie Neiswanger. Isobench: Benchmarking multimodal foundation models on isomorphic representations, 2024. URL https://arxiv.org/abs/2404.01266

  188. [208]

    Chawla, Olaf Wiest, and Xiangliang Zhang

    Taicheng Guo, Kehan Guo, Bozhao Nan, Zhenwen Liang, Zhichun Guo, Nitesh V. Chawla, Olaf Wiest, and Xiangliang Zhang. What can large language models do in chemistry? a comprehensive benchmark on eight tasks, 2023. URL https://arxiv.org/abs/2305.18365

  189. [209]

    Chemformer: a pre-trained transformer for computational chemistry

    Ross Irwin, Spyridon Dimitriadis, Jiazhen He, and Esben Jannik Bjerrum. Chemformer: a pre-trained transformer for computational chemistry. Machine Learning: Science and Technology, 3(1):015022, January 2022. doi: 10.1088/2632-2153/ac3ffb. URL https://dx.doi.org/10.1088/2632-21...

  190. [210]

    ChemQA: a multimodal question-and-answering dataset on chemistry reasoning

    Shang Zhu, Xuefeng Liu, and Ghazal Khalighinejad. ChemQA: a multimodal question-and-answering dataset on chemistry reasoning. https://huggingface.co/datasets/shangzhu/ChemQA, 2024

  191. [211]

    Data-driven electrolyte design for lithium metal anodes

    Kim, S.C., et al. "Data-driven electrolyte design for lithium metal anodes." Proceedings of the National Academy of Sciences 120.10 (2023): e2214357120. https://www.pnas.org/doi/full/10.1073/pnas.2214357120

  192. [212]

    https://unstructured.io/

  193. [213]

    https://platform.openai.com/docs/models

  194. [214]

    https://llama.meta.com/llama3/

  195. [215]

    Ai, Q.; Meng, F.; Shi, J.; Pelkie, B.; Coley, C. W. Extracting Structured Data from Organic Synthesis Procedures Using a Fine-Tuned Large Language Model. ChemRxiv April 8, 2024. https://doi.org/10.26434/chemrxiv-2024-979fz

  196. [216]

    https://github.com/qai222/ontosynthesis (accessed 2024-07-07)

    Ontosynthesis, 2024. https://github.com/qai222/ontosynthesis (accessed 2024-07-07)

  197. [217]

    J.; Karan, D.; Lee, K

    Bai, J.; Mosbach, S.; Taylor, C. J.; Karan, D.; Lee, K. F.; Rihm, S. D.; Akroyd, J.; Lapkin, A. A.; Kraft, M. A Dynamic Knowledge Graph Approach to Distributed Self-Driving Laboratories. Nat. Commun. 2024, 15 (1), 462. https://doi.org/10.1038/s41467-023-44599-9

  198. [218]

    Synthesis Operation Ontology

    Ai, Q.; Klein, C. Synthesis Operation Ontology. GitHub. https://github.com/qai222/ontosynthesis/blob/main/ontologies/soo/soo.md (accessed 2024-07-07)

  199. [219]

    https://dash.plotly.com/cytoscape (accessed 2024-07-06)

    Dash-Cytoscape: A Component Library for Dash Aimed at Facilitating Network Visualization in Python, Wrapped around Cytoscape.Js. https://dash.plotly.com/cytoscape (accessed 2024-07-06)

  200. [220]

    E., III; Jayaraman, A

    Gartner, T. E., III; Jayaraman, A. Modeling and Simulations of Polymers: A Roadmap. Macromolecules 2019, 52 (3), 755-786. DOI: 10.1021/acs.macromol.8b01836

  201. [221]

    GPT-3.5 Turbo. 2024. https://platform.openai.com/docs/models/gpt-3-5-turbo (accessed 2024/07/10)

  202. [222]

    Welcome to GraphRAG. 2024. https://microsoft.github.io/graphrag/ (accessed 2024/07/10)

  203. [223]

    From Local to Global: A Graph RAG Approach to Query-Focused Summarization

    Edge, D.; Trinh, H.; Cheng, N.; Bradley, J.; Chao, A.; Mody, A.; Truitt, S.; Larson, J. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. 2024. (acccessed 2024/07/10)

  204. [224]

    LlamaIndex. 2024. https://docs.llamaindex.ai/en/stable/ (accessed 2024/07/10)

  205. [225]

    Focassio, B., Freitas, M., Schleder, G.R. (2024). Performance assessment of universal machine learning interatomic potentials: Challenges and directions for materials’ surfaces. ACS Applied Materials & Interfaces

  206. [226]

    Marques, F., Balcerzak, M., Winkelmann, F., Zepon, G., Felderhoff, M. (2021). Review and outlook on high-entropy alloys for hydrogen storage. Energy & Environmental Science, 14(10), 5191-5227

  207. [227]

    Jain, S.M. (2022). Hugging face: Introduction to transformers for NLP with the hugging face library and models to solve problems. Berkeley, CA: Apress

  208. [228]

    Lewis, P., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459-9474

  209. [229]

    OpenAI. (2020). OpenAI GPT-3: Language models are few-shot learners. Retrieved from https://openai.com/blog/openai-api

  210. [230]

    OpenAI. (2023). New embedding models and API updates. Retrieved from https://openai.com/index/new-embedding-models-and-api-updates

  211. [231]

    Generative models for molecular discovery: Recent advances and challenges

    Bilodeau, Camille et al. (2022). “Generative models for molecular discovery: Recent advances and challenges”. In: Wiley Interdisciplinary Reviews: Computational Molecular Science 12.5, e1608

  212. [232]

    Chennakesavalu, Shriram et al. (2024). Energy Rank Alignment: Using Preference Optimization to Search Chemical Space at Scale. DOI: 10.48550/ARXIV.2405.12961. URL: https://arxiv. org/abs/2405.12961

  213. [233]

    Extracting medicinal chemistry intuition via preference machine learning

    Choung, Oh-Hyeon et al. (Oct. 2023). “Extracting medicinal chemistry intuition via preference machine learning”. In: Nature Communications 14.1. ISSN: 2041-1723. DOI: 10.1038/s41467- 023-42242-1. URL: http://dx.doi.org/10.1038/s41467-023-42242-1

  214. [234]

    Leveraging large language models for predictive chemistry

    Jablonka, Kevin Maik et al. (Feb. 2024). “Leveraging large language models for predictive chemistry”. In: Nature Machine Intelligence 6.2, pp. 161–169. ISSN: 2522-5839. DOI: 10.1038/s42256-023-00788-1. URL: http://dx.doi.org/10.1038/s42256-023-00788-1

  215. [235]

    Are large language models superhuman chemists?

    Mirza, Adrian et al. (2024). “Are large language models superhuman chemists?” In: arXiv preprint. DOI: 10.48550/arXiv.2404.01475. arXiv: 2404.01475 [cs.LG]

  216. [236]

    Yang, Kaiqi et al. (2024). Are Large Language Models (LLMs) Good Social Predictors? DOI: 10.48550/ARXIV.2402.12620. URL: https://arxiv.org/abs/2402.12620

  217. [237]

    Copier template: Available at https://github.com/copier-org/copier

  218. [238]

    T.; Moazam, H.; Miller, H.; Zaharia, M.; Potts, C

    DSPy: Khattab, O.; Singhvi, A.; Maheshwari, P.; Zhang, Z.; Santhanam, K.; Vardhamanan, S.; Haq, S.; Sharma, A.; Joshi, T. T.; Moazam, H.; Miller, H.; Zaharia, M.; Potts, C. DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines. preprint arXiv:2310.03714. 2023

  219. [239]

    PyMuPDF: Available at ttps://github.com/pymupdf/PyMuPDF

  220. [240]

    Available at \\ https://platform.openai.com/docs/models/gpt-3-5-turbo

    GPT-3.5-Turbo: OpenAI. Available at \\ https://platform.openai.com/docs/models/gpt-3-5-turbo

  221. [241]

    Available at \\ https://platform.openai.com/docs/models/gpt-4-turbo-and-gpt-4

    GPT-4-Turbo: OpenAI. Available at \\ https://platform.openai.com/docs/models/gpt-4-turbo-and-gpt-4

  222. [242]

    DSPy Typed Predictors: Documentation at \\ https://dspy-docs.vercel.app/docs/building-blocks/typed_predictors

  223. [243]

    Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

    Chain-of-Thought prompting: Wei, J.; Wang X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.; Zhou, D. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. preprint arXiv: 2201.11903. 2023

  224. [244]

    M.; Heyer, A

    Test review article on zeolites: Rhoda, H. M.; Heyer, A. J.; Snyder, B. E. R.; Plessers, D.; Bols, M. L.; Schoonheydt, R. A.; Sels, B. F.; Solomon, E. I. Second-Sphere Lattice Effects in Copper and Iron Zeolite Catalysis. Chem. Rev. 2022, 122, 12207–12243

  225. [245]

    Neo4J: Documentation at https://neo4j.com/

  226. [246]

    Graph Maker: Available at https://github.com/rahulnyk/graph_maker

  227. [247]

    Gradio: Available at https://github.com/gradio-app/gradio

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.