REVIEW 3 major objections 4 minor 4 cited by
Reflections from the 2024 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read 34 teams produced working LLM prototypes across seven areas of materials science and chemistry, and the organizers read this as evidence that LLMs now serve as both general-purpose predictors and rapid-prototyping platforms.
desk verdict A useful, honest catalogue of 34 LLM prototypes in materials and chemistry; the abstract's 'significant improvements' trend claim is unsupported and should be qualified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the time-boxed, open-ended hybrid challenge itself: 556 registered participants, 34 completed team submissions, each producing code and a short project report. The organizers use those reports as their dataset, group them into seven application areas, and compare the demonstrated capabilities against those of the previous year's event.
What would settle it
Inspect the linked repositories and rerun the key claims in a controlled setting: five-fold cross-validate the phonon peak predictor, re-run the XRD schema-filling pipeline on the same raw files and count hallucinated values, and test the novelty-scoring model on a larger abstract set. If the flagship prototypes do not reproduce or their error rates equal random baselines, the paper's evidence for expanded LLM capability collapses.
Extended reading notes
Core claim
The paper claims that second-generation LLMs now have a dual utility in materials and chemistry: they are multipurpose models that can be pointed at regression, generation, and extraction tasks, and they are platforms on which researchers can prototype custom applications in under a day. The evidence is the set of 34 hackathon submissions, each with code and a written report, which the organizers sorted into seven application areas. They observe that performance across this application space improved substantially since the previous year's hackathon, in many cases simply because newer model versions were released, and they take this as a sign that LLM use in these fields will keep expanding.
Load-bearing premise
The conclusion rests on treating 34 self-selected, largely unvalidated team reports as representative evidence; many entries lack cross-validation, statistical significance, or quantitative checks, so if the reports flatter reality, the case for broad LLM utility is not made.
Editorial extensions
If this is right
- Low-data property prediction appears within reach: teams improved phonon peak and lithium-ion conductivity predictions by enriching composition inputs with bonding or literature context.
- Design tasks can be bootstrapped with zero-shot LLM suggestions, and small locally hosted models beat random baselines in concrete formulation design, suggesting that private and data-sensitive laboratories could use them.
- Natural-language interfaces for instruments and simulations are feasible: teams automated bulk-modulus calculations, microscope parameter estimation, and DFT setup, lowering the need for specialist operators.
- Text-native workflows, such as knowledge-graph construction, schema filling, hypothesis evaluation, and question answering, are where LLMs most readily plug in, because these tasks align with the models' core strengths.
- If the trend continues, newer model generations alone will broaden the range of research tasks that LLMs can handle without custom adaptation.
Reading between the lines
- The report's emphasis on breadth over benchmarks suggests the fastest near-term gains will come in text-native data work, such as literature extraction, electronic lab notebooks, and report drafting, where success is judged by workflow completion rather than scientific accuracy.
- Because many prototypes lack error bars, the honest reading is that LLMs widen what can be attempted, not yet what can be trusted; a reproducibility pass across the 34 repositories would separate stable tools from one-off demonstrations.
- The low-data property-prediction results hint that context augmentation, such as bonding analysis and literature summaries, may matter more than model scale, a claim the paper supports only partially and that deserves direct comparison against classical machine-learning baselines.
- If small locally hosted models continue to approach larger ones on design tasks, the practical ceiling may shift from model capability to prompt and workflow engineering, which would make data privacy less of an obstacle.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports on the second Large Language Model Hackathon for Applications in Materials Science and Chemistry, held in May 2024 with hybrid physical hubs and a global online hub. It describes 34 team submissions, of which 32 have written reports in the appendix, organizes them into seven application areas (property prediction, design, automation and interfaces, communication and education, data management, hypothesis generation, and knowledge extraction), and highlights exemplar projects in each area. The paper also discusses the event format and concludes that the event demonstrated the dual utility of LLMs as multipurpose predictive models and as rapid-prototyping platforms, and that LLM capabilities improved significantly since the previous year's hackathon. Individual appendix reports provide code links and, in several cases, candid statements of limitations.
Significance. The paper's main value is as an archival record of a large, open, community-driven prototyping event: it collects 32 project reports with public code repositories, spans a broad application taxonomy, and includes unusually honest reporting of failures and non-significant results. If its broader claims were supported, the paper would provide useful evidence that LLMs can be productively applied across a wide range of materials chemistry tasks within a 24-hour hackathon format. However, the headline claims of the abstract and conclusion go beyond what the evidence supports: no matched comparison with the previous hackathon is provided, and several reported projects show hallucinations, model failure, or non-significant results. These issues are local to the framing and can be corrected by softening the claims; the descriptive core of the paper remains valuable.
major comments (3)
- [Abstract; Conclusion; Appendix §§5, 8, 18, 24] The abstract states that 'the event highlighted significant improvements in LLM capabilities since the previous year's hackathon', and the conclusion attributes this improvement to the release of new model versions. No matched comparison between the 2024 and 2023 hackathons is provided: the teams, tasks, datasets, and evaluation protocols differ, and no project benchmarks current models against the models used in the previous event. Moreover, several projects in the appendix report negative or mixed outcomes, including LLMSpectrometry solving only 3 of 19 NMR spectra (§5), the Phi-3 model failing after the fifth development cycle in the small-LLM concrete study (§8), hallucinated values in LLMads (§18), and a non-significant p=0.07 in G-Peer-T (§24). The improvement claim should either be removed or replaced by a carefully hedged statement that the breadth of prototypes is consistent with, but does not measure, year-over-year capability growth.
- [Abstract; Event Overview; Table 1; Appendix TOC] The manuscript reports '34 team submissions' but Table 1 and the appendix table of contents list only 32 projects. The abstract also says that each team submission is presented in a summary table and as brief papers in the appendix, which is inconsistent with the '32 submissions with a written description included here' formulation in the Event Overview. Please reconcile the total: either list the two additional submissions (even if they lack write-ups or code), or change the stated total to 32.
- [Conclusion; Abstract] The claim that the event 'demonstrated the dual utility of LLMs' is stronger than the evidence warrants. The paper is a collection of self-selected, time-constrained prototypes, and several appendix reports explicitly note missing validation: Learning LOBSTERs states that no five-fold cross-validation was implemented (§1), NOMAD Query Reporter reports hallucinations on heterogeneous data (§19), and the small-LLM study reports a complete failure mode for Phi-3 (§8). The conclusion should be phrased in terms of demonstrated breadth of prototyping and potential utility, not demonstrated utility, unless the authors add a systematic validation criterion applied across projects.
minor comments (4)
- [Table 1 caption] The caption reads 'Overview of the tools developed by the various tools'; 'tools' should presumably be 'teams'.
- [§8.4, Table 2] The sentence 'all investigated models generated designs that outperformed the statistical baseline' is immediately followed by an exception for Phi-3 with design context, which failed after the fifth development cycle; please rephrase to state that all models outperformed the baseline in the first round, with the noted exception in later rounds.
- [§19, references] In the NOMAD Query Reporter section, the in-text citations appear swapped: 'Query Reporter [1]' and 'NOMAD [2]' should refer to the GitHub repository and the NOMAD paper respectively, but the reference list assigns [1] to the NOMAD paper and [2] to the GitHub repository.
- [Overview; Appendix §17] The project name is spelled inconsistently as 'yeLLowhaMmer' in the overview and 'yeLLowhaMMer' in the appendix title; please standardize the spelling.
Circularity Check
No circular derivation: the report's conclusions are inductive summaries of self-selected prototypes; self-citations are contextual, and the unsupported year-over-year improvement claim is an external attribution rather than a circular inference.
full rationale
This paper is a hackathon proceedings rather than a derivation: it reports 34 prototype submissions and summarizes their application areas. The abstract's comparative claim, 'significant improvements in LLM capabilities since the previous year's hackathon,' is not derived from any fitted parameter in this paper; the Conclusion attributes it to an external event ('the performance across the diverse application space was improved simply via the release of new versions of Gemini, ChatGPT, Claude, Llama, and other models'). That claim lacks a matched comparison with the 2023 hackathon and is therefore unsupported, but unsupported inference from model-version releases is an evidence gap, not circularity. Self-citations appear (the 2023 hackathon report [53]; several appendix projects citing the authors' own datasets or methods, e.g., Learning LOBSTERs uses its earlier LobsterPy dataset and the MOF agent adapts the dZiner approach), but they are contextual or methodological starting points: the reported outcomes are evaluated against external resources (MatBench phonon DOS, QM9, published geopolymer strengths, MaScQA, SDBS NMR spectra) or are honestly reported as failures or non-significant (LLMads hallucinations, Phi-3 breakdown, G-Peer-T p=0.07, LLMSpectrometry 3/19). The closest thing to a closed loop is the MOF agent's use of its own surrogate to accept candidates with lower predicted band gaps and then plot those predictions, but the text labels these as 'inferred band gap' and claims no external validation; that is a transparency and validity limitation of a 24-hour prototype, not a hidden reduction of the paper's central claim to its inputs.
Assumptions & free parameters
assumptions (2)
- domain assumption Team self-reports in the appendix accurately describe the capabilities and limitations of the submitted projects.
- domain assumption The seven application categories are meaningful and the assignment of projects to categories is correct.
Cite this review
Pith. "Pith review of Reflections from the 2024 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry." pith.science (2026). https://pith.science/paper/PGFBIFJG
@misc{pith2026241115221,
author = {Pith},
title = {Pith review of: Reflections from the 2024 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry},
year = {2026},
howpublished = {\url{https://pith.science/paper/PGFBIFJG}},
note = {Machine review of arXiv:2411.15221}
}
read the original abstract
Here, we present the outcomes from the second Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry, which engaged participants across global hybrid locations, resulting in 34 team submissions. The submissions spanned seven key application areas and demonstrated the diverse utility of LLMs for applications in (1) molecular and material property prediction; (2) molecular and material design; (3) automation and novel interfaces; (4) scientific communication and education; (5) research data management and automation; (6) hypothesis generation and evaluation; and (7) knowledge extraction and reasoning from scientific literature. Each team submission is presented in a summary table with links to the code and as brief papers in the appendix. Beyond team results, we discuss the hackathon event and its hybrid format, which included physical hubs in Toronto, Montreal, San Francisco, Berlin, Lausanne, and Tokyo, alongside a global online hub to enable local and virtual collaboration. Overall, the event highlighted significant improvements in LLM capabilities since the previous year's hackathon, suggesting continued expansion of LLMs for applications in materials science and chemistry research. These outcomes demonstrate the dual utility of LLMs as both multipurpose models for diverse machine learning tasks and platforms for rapid prototyping custom applications in scientific research.
Figures
Figures from the paper (32 more)
Forward citations
Cited by 4 Pith papers
-
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
SEAM measures VLM reasoning consistency across modalities using paired semantically equivalent textual and visual notations, and finds systematic vision-language imbalance.
-
Foundational Large Language Models for Materials Research
Domain-adapted LLaMA models (LLaMat) outperform commercial LLMs on materials NLP and structured extraction tasks and generate M3GNet-predicted stable crystals, with LLaMA-2-based variants beating LLaMA-3-based ones.
-
Multicrossmodal Automated Agent for Integrating Diverse Materials Science Data
A prompt-only multi-agent LLM system claims to fuse video, image, table, and text data for materials-science questions, reporting 85% recall and 35% coverage gains, but with unverifiable evaluation.
-
From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines
LLMs accelerate research workflows from idea generation to writing but introduce challenges like hallucination, bias, opacity, and ten systemic risks requiring new governance frameworks.
Reference graph
Works this paper leans on
-
[1]
A. Nolte, L. B. Hayden and J. D. Herbsleb, Proc. ACM Hum.-Comput. Interact., 2020, 4, 1--23. https://doi.org/10.1145/3392830
-
[2]
E. P. P. Pe-Than and J. D. Herbsleb, in Lecture Notes in Computer Science, Springer, 2019, vol. 11546, pp. 27--37. https://doi.org/10.1007/978-3-030-15742-5_3
-
[3]
B. Heller, A. Amir, R. Waxman and Y. Maaravi, J. Innov. Entrep., 2023, 12, 1. https://doi.org/10.1186/s13731-023-00269-0
-
[5]
C. Qian, H. Tang, Z. Yang, H. Liang and Y. Liu, arXiv, 2023. https://arxiv.org/abs/2307.07443
arXiv 2023
- [6]
-
[7]
Vacareanu, V
R. Vacareanu, V. A. Negru, V. Suciu and M. Surdeanu, in First Conference on Language Modeling, 2024. https://openreview.net/forum?id=LzpaUxcNFK
2024
-
[8]
A. K. Gupta and K. Raghavachari, J. Chem. Theory Comput., 2022, 18, 2132--2143. https://doi.org/10.1021/acs.jctc.1c00504
-
[10]
J. Lu and Y. Zhang, J. Chem. Inf. Model., 2022, 62, 1376--1387. https://doi.org/10.1021/acs.jcim.1c01467
Show all 235 references
-
[11]
Bhattacharya, H
D. Bhattacharya, H. J. Cassady, M. A. Hickner and W. F. Reinhart, J. Chem. Inf. Model., 2024, 64, 7086--7096. https://doi.org/10.1021/acs.jcim.4c01396
2024 doi
-
[12]
G. Liu, M. Sun, W. Matusik, M. Jiang and J. Chen, arXiv, 2024. https://arxiv.org/abs/2410.04223
2024 arXiv
-
[13]
S. Jia, C. Zhang and V. Fung, arXiv, 2024. https://arxiv.org/abs/2406.13163
2024 arXiv
-
[14]
H. Jang, Y. Jang, J. Kim and S. Ahn, arXiv, 2024. https://arxiv.org/abs/2410.03138
2024 arXiv
-
[15]
J. Lu, Z. Song, Q. Zhao, Y. Du, Y. Cao, H. Jia and C. Duan, arXiv, 2024. https://arxiv.org/abs/2410.18136
2024 arXiv
-
[16]
Kristiadi et al., in Proceedings of the 41st International Conference on Machine Learning, PMLR, 2024, vol
A. Kristiadi et al., in Proceedings of the 41st International Conference on Machine Learning, PMLR, 2024, vol. 235, pp. 25603--25622. https://proceedings.mlr.press/v235/kristiadi24a.html
2024
-
[17]
Miret and N
S. Miret and N. M. A. Krishnan, arXiv, 2024. https://arxiv.org/abs/2402.05200
2024 arXiv
-
[18]
X. Ji, A. L. Nielsen and C. Heinis, Angew. Chem. Int. Ed., 2023, 63, 3. https://doi.org/10.1002/anie.202308251
2023 doi
- [21]
- [22]
-
[23]
A. M. Bran, S. Cox, O. Schilter, C. Baldassari, A. D. White and P. Schwaller, arXiv, 2023. https://arxiv.org/abs/2304.05376
2023 arXiv
- [24]
-
[25]
Zhang, Y
H. Zhang, Y. Song, Z. Hou, S. Miret and B. Liu, arXiv, 2024. https://arxiv.org/abs/2409.00135
2024 arXiv
-
[27]
Darvish et al., arXiv, 2024
K. Darvish et al., arXiv, 2024. https://arxiv.org/abs/2401.06949
2024 arXiv
-
[28]
Tom et al., Chem
G. Tom et al., Chem. Rev., 2024, 124, 9633--9732. https://doi.org/10.1021/acs.chemrev.4c00055
2024 doi
-
[29]
Chase, Langchain, 2024
H. Chase, Langchain, 2024. https://github.com/langchain-ai/langchain
2024
-
[30]
RDKit: Open-source cheminformatics, http://www.rdkit.org
-
[31]
Yan et al., Br
L. Yan et al., Br. J. Educ. Technol., 2023, 55, 90--112. https://doi.org/10.1111/bjet.13370
2023 doi
- [32]
-
[33]
Kasneci et al., Learn
E. Kasneci et al., Learn. Individ. Differ., 2023, 103, 102274. https://doi.org/10.1016/j.lindif.2023.102274
2023
-
[34]
M. S. Schäfer, J. Sci. Commun., 2023, 22, 2. https://doi.org/10.22323/2.22020402
2023 doi
-
[35]
Zaki, Jayadeva, Mausam and N
M. Zaki, Jayadeva, Mausam and N. M. A. Krishnan, arXiv, 2023. https://arxiv.org/abs/2308.09115
2023 arXiv
-
[36]
https://assets.anthropic.com/m/61e7d27f8c8f5919/original/Claude-3-Model-Card.pdf
Anthropic, The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024. https://assets.anthropic.com/m/61e7d27f8c8f5919/original/Claude-3-Model-Card.pdf
2024
-
[37]
A. Q. Jiang et al., arXiv, 2024. https://arxiv.org/abs/2401.04088
2024 arXiv
-
[38]
Draxl and M
C. Draxl and M. Scheffler, J. Phys. Mater., 2019, 2, 036001. https://doi.org/10.1088/2515-7639/ab13bb
2019 doi
-
[39]
Radford et al., arXiv, 2022
A. Radford et al., arXiv, 2022. https://arxiv.org/abs/2212.04356
2022 arXiv
-
[40]
Y. Zhou, H. Liu, T. Srivastava, H. Mei and C. Tan, arXiv, 2024. https://arxiv.org/abs/2404.04326
2024 arXiv
-
[41]
Abdel-Rehim et al., arXiv, 2024
A. Abdel-Rehim et al., arXiv, 2024. https://arxiv.org/abs/2405.12258
2024 arXiv
-
[42]
S. Tong, K. Mao, Z. Huang, Y. Zhao and K. Peng, Humanit. Soc. Sci. Commun., 2024, 11, 1. https://doi.org/10.1057/s41599-024-03407-5
2024 doi
-
[43]
Ciucă, Y.-S
I. Ciucă, Y.-S. Ting, S. Kruk and K. Iyer, arXiv, 2023. https://arxiv.org/abs/2306.11648
2023 arXiv
- [44]
- [45]
-
[46]
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao and K. Narasimhan, in Advances in Neural Information Processing Systems, Curran Associates, Inc., 2023, vol. 36, pp. 11809--11822. https://proceedings.neurips.cc/paper_files/paper/2023/file/271db9922b8d1f4dd7aaef84ed5ac7...
2023
-
[47]
Shamsabadi, J
M. Shamsabadi, J. D'Souza and S. Auer, arXiv, 2024. https://arxiv.org/abs/2401.10040
2024 arXiv
-
[48]
Dagdelen et al., Nat
J. Dagdelen et al., Nat. Commun., 2024, 15, 1. https://doi.org/10.1038/s41467-024-45563-x
2024 doi
- [49]
-
[50]
J. Li, M. Zhang, N. Li, D. Weyns, Z. Jin and K. Tei, ACM Trans. Auton. Adapt. Syst., 2024, 19, 1--60. https://doi.org/10.1145/3686803
2024 doi
- [51]
-
[52]
https://arxiv.org/abs/2303.08774
OpenAI et al., arXiv, 2023. https://arxiv.org/abs/2303.08774
2023 arXiv
-
[54]
K. M. Jablonka, P. Schwaller, A. Ortega-Guerrero, B. Smit, Nat Mach Intell, 2024, 6, 161–169
2024
-
[55]
Choudhary, J
K. Choudhary, J. Phys. Chem. Lett., 2024, 6909–6917
2024
-
[56]
Petretto, S
G. Petretto, S. Dwaraknath, H. P.C. Miranda, D. Winston, M. Giantomassi, M. J. van Setten, X. Gonze, K. A. Persson, G. Hautier, G.-M. Rignanese, Sci Data, 2018, 5, 180065
2018
-
[57]
A. Dunn, Q. Wang, A. Ganose, D. Dopp, A. Jain, npj Comput Mater, 2020, 6, 1–10
2020
-
[58]
A. M. Ganose, A. Jain, MRS Communications, 2019, 9, 874–881
2019
-
[59]
H. M. Sayeed, S. G. Baird, T. D. Sparks, 2023, DOI 10.26434/chemrxiv-2023-3q8wj
2023 doi
- [60]
- [61]
-
[62]
A. A. Naik, K. Ueltzen, C. Ertural, A. J. Jackson, J. George, Journal of Open Source Software, 2024, 9, 6286
2024
-
[63]
A. A. Naik, C. Ertural, N. Dhamrait, P. Benner, J. George, 2023, DOI 10.5281/zenodo.8091844
2023 doi
-
[64]
The Matbench Test Suite, Phonon dataset as per 12.07.2024, https://matbench.materialsproject.org/Leaderboards
2024
-
[65]
Raffel, N
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, P. J. Liu, Journal of Machine Learning Research, 2020, 21, 1–67
2020
-
[66]
The Unsloth package, https://github.com/unslothai/unsloth, 2024
2024
-
[67]
Goodall, A
R. Goodall, A. J. et al., ``Predicting materials properties without crystal structure: deep representation learning from stoichiometry'', Nature Communications, vol. 11, no. 6280, 2020
2020
-
[68]
Trewartha, et al., 'Quantifying the advantage of domain-specific pre-training on named entity recognition tasks in materials science', Patterns, vol
A. Trewartha, et al., 'Quantifying the advantage of domain-specific pre-training on named entity recognition tasks in materials science', Patterns, vol. 3, no. 8, 2022
2022
-
[69]
Gupta, et al., MatSciBERT: A materials domain language model for text mining and information extraction, npj Computational Materials, vol
V. Gupta, et al., MatSciBERT: A materials domain language model for text mining and information extraction, npj Computational Materials, vol. 8, no. 1, 2022
2022
-
[70]
Matbench Perovskites dataset provided by the Materials Project, https://ml.materialsproject.org/projects/matbench_perovskites.json.gz
-
[71]
Li-ion Conductors Database, https://pcwww.liv.ac.uk/ msd30/lmds/LiIonDatabase.html
-
[72]
PyMuPDF, https://pypi.org/project/PyMuPDF
-
[73]
AI@Meta, Llama 3 Model Card, 2024, https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md
2024
-
[74]
Hargreaves et al., The Earth Mover’s Distance as a Metric for the Space of Inorganic Compositions, Chem. Mater. 2020
2020
-
[75]
Hargreaves et al., A Database of Experimentally Measured Lithium Solid Electrolyte Conductivities Evaluated with Machine Learning, npj Computational Materials 2022
2022
-
[77]
M., Ai, Q., Al-Feghali, A., Badhwar, S., Bocarsly, J
Jablonka, K. M., Ai, Q., Al-Feghali, A., Badhwar, S., Bocarsly, J. D., Bran, A. M., ... & Blaiszik, B. (2023). 14 examples of how LLMs can transform materials science and chemistry: a reflection on a large language model hackathon. Digital Discovery, 2(5), 1233-1250
2023
-
[78]
Alampara, N., Miret, S., & Jablonka, K. M. (2024). MatText: Do Language Models Need More than Text & Scale for Materials Modeling? arXiv preprint arXiv:2406.17295
2024 arXiv
-
[79]
P., Kornbluth, M.,
Batzner, S., Musaelian, A., Sun, L., Geiger, M., Mailoa, J. P., Kornbluth, M., ... & Kozinsky, B. (2022). E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials. Nature Communications, 13(1), 2453
2022
-
[80]
K., & Raghavachari, K
Gupta, A. K., & Raghavachari, K. (2022). Three-dimensional convolutional neural networks utilizing molecular topological features for accurate atomization energy predictions. Journal of Chemical Theory and Computation, 18(4), 2132-2143
2022
-
[81]
Grzegorz Kaszuba, Amirhossein D
In-Context Learning of Physical Properties: Few-Shot Adaptation to Out-of-Distribution Molecular Graphs. Grzegorz Kaszuba, Amirhossein D. Naghdi, Dario Massa, Stefanos Papanikolaou, Andrzej Jaszkiewicz, Piotr Sankowski, https://arxiv.org/abs/2406.01808
-
[82]
Khan, D., Heinen, S., & von Lilienfeld, O. A. (2023). Kernel based quantum machine learning at record rate: Many-body distribution functionals as compact representations. Journal of Chemical Physics, 159(034106)
2023
-
[83]
T., Kabylda, A., Sauceda, H
Chmiela, S., Vassilev-Galindo, V., Unke, O. T., Kabylda, A., Sauceda, H. E., Tkatchenko, A., & Müller, K. R. (2023). Accurate global machine learning force fields for molecules with hundreds of atoms. Science Advances, 9(2), https://doi.org/10.1126/sciadv.adf0873
2023 doi
-
[84]
M., Qu, C., Conte, R., Nandi, A., Houston, P
Bowman, J. M., Qu, C., Conte, R., Nandi, A., Houston, P. L., & Yu, Q. (2022). The MD17 datasets from the perspective of datasets for gas-phase “small” molecule potentials. The Journal of chemical physics, 156(24)
2022
-
[85]
Weinreich, J., & Probst, D. (2023). Parameter-Free Molecular Classification and Regression with Gzip. ChemRxiv
2023
-
[86]
P., Simm, G., Ortner, C., & Csányi, G
Batatia, I., Kovacs, D. P., Simm, G., Ortner, C., & Csányi, G. (2022). MACE: Higher order equivariant message passing neural networks for fast and accurate force fields. Advances in Neural Information Processing Systems, 35, 11423-11436
2022
-
[87]
A., ACS Central Sci., 2019, Vol
Schwaller, P., Laino, T., Gaudin, T., Bolgar, P., Hunter, C., Bekas, C., Lee, A. A., ACS Central Sci., 2019, Vol. 5, No. 9, 1572-1583. https://pubs.acs.org/doi/10.1021/acscentsci.9b00576
2019 doi
-
[88]
C., ChemRxiv Preprint
Alberts, M., Zipoli, F., Vaucher, A. C., ChemRxiv Preprint. https://doi.org/10.26434/chemrxiv-2023-8wxcz
2023 doi
-
[90]
https://sdbs.db.aist.go.jp
Yamaji, T., Saito, T., Hayamizu, K., Yanagisawa, M., Yamamoto, O., Wasada, N., Someno, K., Kinugasa, S., Tanabe, K., Tamura, T., Hiraishi, J., 2024. https://sdbs.db.aist.go.jp
2024
-
[91]
Socha, O., Osifova, Z., Dracinsky, M., J. Chem. Educ., 2023, Vol. 100, No. 2, 962-968. https://pubs.acs.org/doi/10.1021/acs.jchemed.2c01067
2023 doi
-
[92]
https://github.com/ATOMSLab/LLMSpectroscopy
-
[93]
Cyclic peptides for drug development,
Ji, X., Nielsen, A. L., Heinis, C., "Cyclic peptides for drug development," Angewandte Chemie International Edition, 2024, 63(3), e202308251
2024
-
[94]
De novo development of small cyclic peptides that are orally bioavailable,
Merz, M.L., Habeshian, S., Li, B. et al., "De novo development of small cyclic peptides that are orally bioavailable," Nat Chem Biol, 2024, 20, 624–633. https://doi.org/10.1038/s41589-023-01496-y
2024 doi
-
[95]
Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation,
Beurer-Kellner, L., et al., "Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation," ArXiv, 2024, abs/2403.06988
2024 arXiv
-
[96]
A survey on in-context learning,
Dong, Q., et al., "A survey on in-context learning," ArXiv, 2022, arXiv:2301.00234
2022 arXiv
-
[97]
Many-shot in-context learning,
Agarwal, R., et al., "Many-shot in-context learning," ArXiv, 2024, arXiv:2404.11018
2024 arXiv
-
[98]
A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?,
Kristiadi, A., et al., "A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?," ArXiv, 2024, arXiv:2402.05015
2024 arXiv
-
[99]
A Detailed Investigation on Conformation, Permeability and PK Properties of Two Related Cyclohexapeptides,
Lewis, I., Schaefer, M., Wagner, T. et al., "A Detailed Investigation on Conformation, Permeability and PK Properties of Two Related Cyclohexapeptides," Int J Pept Res Ther, 2015, 21, 205–221. https://doi.org/10.1007/s10989-014-9447-3
2015 doi
-
[100]
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,
Lewis, P., et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks," ArXiv, 2020, abs/2005.11401
2020 arXiv
-
[101]
Reducing hallucination in structured outputs via Retrieval-Augmented Generation,
Béchard, P., Ayala, O. M., "Reducing hallucination in structured outputs via Retrieval-Augmented Generation," ArXiv, 2024, abs/2404.08189
2024 arXiv
-
[102]
Review on applications of metal–organic frameworks for co2 capture and the performance enhancement mechanisms
Lirong Li, Han Sol Jung, Jae Won Lee, and Yong Tae Kang. Review on applications of metal–organic frameworks for co2 capture and the performance enhancement mechanisms. Renewable and Sustainable Energy Reviews, 162: 112441, 2022
2022
-
[103]
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. arXiv preprint, arXiv:2210.03629, 2022
2022 arXiv
-
[104]
Brown, and Joseph S
Mehrad Ansari, Jeffrey Watchorn, Carla E. Brown, and Joseph S. Brown. dZiner: Rational Inverse Design of Materials with AI Agents. arXiv, 2410.03963, 2024. URL: https://arxiv.org/abs/2410.03963
2024 arXiv
-
[105]
Semiconductor metal–organic frameworks: future low- g"bandgap materials
Muhammad Usman, Shruti Mendiratta, and Kuang-Lieh Lu. Semiconductor metal–organic frameworks: future low- g"bandgap materials. Advanced Materials, 29(6):1605071, 2017
2017
-
[106]
Band gap modulations in uio metal–organic frameworks
Espen Flage-Larsen, Arne Røyset, Jasmina Hafizovic Cavka, and Knut Thorshaug. Band gap modulations in uio metal–organic frameworks. The Journal of Physical Chemistry C, 117(40):20610–20616, 2013
2013
-
[107]
Band gap engineering of paradigm mof-5
Li-Ming Yang, Guo-Yong Fang, Jing Ma, Eric Ganz, and Sang Soo Han. Band gap engineering of paradigm mof-5. Crystal growth & design, 14(5):2532–2541, 2014
2014
-
[108]
Theoretical investigations on the chemical bonding, electronic structure, and optical properties of the metal- organic framework mof-5
Li-Ming Yang, Ponniah Vajeeston, Ponniah Ravindran, Helmer Fjellvag, and Mats Tilset. Theoretical investigations on the chemical bonding, electronic structure, and optical properties of the metal- organic framework mof-5. Inorganic chemistry, 49(22):10283–10290, 2010
2010
-
[109]
Recent advancements in mof-based catalysts for applications in electrochemical and photoelectrochemical water splitting: A review
Maryum Ali, Erum Pervaiz, Tayyaba Noor, Osama Rabi, Rubab Zahra, and Minghui Yang. Recent advancements in mof-based catalysts for applications in electrochemical and photoelectrochemical water splitting: A review. International Journal of Energy Research, 45(2):1190–1226, 2021
2021
-
[110]
Tuning electrical and mechanical properties of metal–organic frameworks by metal substitution
Yabin Yan, Chunyu Wang, Zhengqing Cai, Xiaoyuan Wang, and Fuzhen Xuan. Tuning electrical and mechanical properties of metal–organic frameworks by metal substitution. ACS Applied Materials & Interfaces, 15(36):42845–42853, 2023
2023
-
[111]
Tunability of band gaps in metal–organic frameworks
Chi-Kai Lin, Dan Zhao, Wen-Yang Gao, Zhenzhen Yang, Jingyun Ye, Tao Xu, Qingfeng Ge, Shengqian Ma, and Di-Jia Liu. Tunability of band gaps in metal–organic frameworks. Inorganic chemistry, 51(16):9039–9044, 2012
2012
-
[112]
New and improved embedding model, 2022
Ryan Greene, Ted Sanders, Lilian Weng, and Arvind Neelakantan. New and improved embedding model, 2022
2022
-
[113]
Agent-based learning of materials datasets from scientific literature
Mehrad Ansari and Seyed Mohamad Moosavi. Agent-based learning of materials datasets from scientific literature. arXiv preprint, arXiv:2312.11690, 2023
2023 arXiv
-
[114]
Moformer: self-supervised transformer model for metal–organic framework property prediction
Zhonglin Cao, Rishikesh Magar, Yuyang Wang, and Amir Barati Farimani. Moformer: self-supervised transformer model for metal–organic framework property prediction. Journal of the American Chemical Society, 145(5): 2958–2967, 2023
2023
-
[115]
Barlow twins: Self-supervised learning via redundancy reduction
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. Barlow twins: Self-supervised learning via redundancy reduction. In International Conference on Machine Learning, pages 12310–12320. PMLR, 2021
2021
-
[116]
Grossman
Tian Xie and Jeffrey C. Grossman. Crystal graph convolutional neural networks for an accurate and interpretable prediction of material properties. Physical Review Letters, 120(14):145301, 2018
2018
-
[117]
Understanding the diversity of the metal-organic framework ecosystem
Seyed Mohamad Moosavi, Aditya Nandy, Kevin Maik Jablonka, Daniele Ongari, Jon Paul Janet, Peter G Boyd, Yongjin Lee, Berend Smit, and Heather J Kulik. Understanding the diversity of the metal-organic framework ecosystem. Nature communications, 11(1):1–10, 2020
2020
-
[118]
Rosen, Shaelyn M
Andrew S. Rosen, Shaelyn M. Iyer, Debmalya Ray, Zhenpeng Yao, Alan Aspuru-Guzik, Laura Gagliardi, Justin M. Notestein, and Randall Q. Snurr. Machine learning the quantum-chemical properties of metal–organic frameworks for accelerated materials discovery. Matter, 4(5):1578–1597, 2021
2021
-
[119]
RDKit documentation
Greg Landrum. RDKit documentation. Release, 1(1-79):4, 2013
2013
-
[120]
Gpt-4 technical report, 2023
OpenAI. Gpt-4 technical report, 2023
2023
-
[121]
LangChain, 10 2022
Harrison Chase. LangChain, 10 2022. URL https://github.com/langchain-ai/langchain
2022
-
[122]
Eco-efficient cements: Potential economically viable solutions for a low-CO2 cement-based materials industry,
U. Environment, K. L. Scrivener, V. M. John and E. M. Gartner, "Eco-efficient cements: Potential economically viable solutions for a low-CO2 cement-based materials industry," Cement and Concrete Research, vol. 114, pp. 2-26; DOI: https://doi.org/10.1016/j.cemconres.2018.03.015, 2018
-
[123]
Advances in understanding alkali-activated materials,
J. L. Provis, A. Palomo and C. Shi, "Advances in understanding alkali-activated materials," Cement and Concrete Research, pp. 110-125, 2015
2015
-
[124]
Development of Eco-Efficient Fly Ash–Based Alkali-Activated and Geopolymer Composites with Reduced Alkaline Activator Dosage,
H. S. Gökçe, M. Tuyan, K. Ramyar and M. L. Nehdi, "Development of Eco-Efficient Fly Ash–Based Alkali-Activated and Geopolymer Composites with Reduced Alkaline Activator Dosage," Journal of Materials in Civil Engineering, vol. 32, no. 2, pp. 04019350; DOI: 10.1061/(ASCE)MT.1943...
-
[125]
Synthesis and characterization of red mud and rice husk ash-based geopolymer composites,
J. He, Y. Jie, J. Zhang, Y. Yu and G. Zhang, "Synthesis and characterization of red mud and rice husk ash-based geopolymer composites," Cement and Concrete Composites, vol. 37, pp. 108-118; DOI: http://dx.doi.org/10.1016/j.cemconcomp.2012.11.010, 2013
-
[126]
LLMs can Design Sustainable Concrete – a Systematic Benchmark,
C. Völker, T. Rug, K. M. Jablonbka and S. Kurschwitz, "LLMs can Design Sustainable Concrete – a Systematic Benchmark," (Preprint), pp. 1-12; DOI: 10.21203/rs.3.rs-3913272/v1, 2023
2023 doi
-
[127]
A quantitative method of approach in designing the mix proportions of fly ash and GGBS-based geopolymer concrete,
G. M. Rao and T. D. G. Rao, "A quantitative method of approach in designing the mix proportions of fly ash and GGBS-based geopolymer concrete," Australian Journal of Civil Engineering, vol. 16, no. 1, pp. 53-63; DOI: 10.1080/14488353.2018.1450716, 2018
-
[128]
Boiko, D.A., MacKnight, R., Kline, B. et al. Autonomous chemical research with large language models. Nature 624, 570–578 (2023). https://doi.org/10.1038/s41586-023-06792-0
2023 doi
-
[129]
Bran, A., Cox, S., Schilter, O
M. Bran, A., Cox, S., Schilter, O. et al. Augmenting large language models with chemistry tools. Nat Mach Intell 6, 525–535 (2024). https://doi.org/10.1038/s42256-024-00832-8
2024 doi
-
[130]
E., Christensen, R., Dułak, M., Friis, J., Groves, M
Hjorth Larsen, A., Jørgen Mortensen, J., Blomqvist, J., Castelli, I. E., Christensen, R., Dułak, M., Friis, J., Groves, M. N., Hammer, B., Hargus, C., Hermes, E. D., Jennings, P. C., Bjerre Jensen, P., Kermode, J., Kitchin, J. R., Leonhard Kolsbjerg, E., Kubal, J., Kaasbjerg, ...
2017 doi
-
[131]
W., Stoltze, P., Nørskov, J
Jacobsen, K. W., Stoltze, P., Nørskov, J. K. (1996). A semi-empirical effective medium theory for metals and alloys. Surface Science, 366(2), 394–402. https://doi.org/10.1016/0039-6028(96)00816-3
1996 doi
-
[132]
M., Kovács, D
Batatia, I., Benner, P., Chiang, Y., Elena, A. M., Kovács, D. P., Riebesell, J., Advincula, X. R., Asta, M., Avaylon, M., Baldwin, W. J., Berger, F., Bernstein, N., Bhowmik, A., Blau, S. M., Cărare, V., Darby, J. P., De, S., Della Pia, F., Deringer, V. L. et al. (2024). A foun...
2024 arXiv
-
[133]
https://github.com/langchain-ai/langchain
-
[134]
https://pydantic.dev/
-
[135]
Ye, H., Liu, T., Zhang, A., Hua, W., Jia, W. (2023). Cognitive mirage: A review of hallucinations in large language models. arXiv. https://arxiv.org/abs/2309.06794
2023 arXiv
-
[136]
Stefan Bauer et al, Roadmap on data-centric materials science, Modelling Simul. Mater. Sci. Eng., 2024, 32, 063301
2024
-
[137]
Leveraging Large Language Models and Social Media for Automation in Scanning Probe Microscopy
Diao, Zhuo, Hayato Yamashita, and Masayuki Abe. "Leveraging Large Language Models and Social Media for Automation in Scanning Probe Microscopy." arXiv preprint arXiv:2405.15490 (2024)
2024 arXiv
-
[138]
Synergizing Human Expertise and AI Efficiency with Language Model for Microscopy Operation and Automated Experiment Design
Liu, Yongtao, Marti Checa, and Rama K. Vasudevan. "Synergizing Human Expertise and AI Efficiency with Language Model for Microscopy Operation and Automated Experiment Design." Machine Learning: Science and Technology (2024)
2024
-
[139]
The abTEM code: transmission electron microscopy from first principles
Madsen, Jacob, and Toma Susi. "The abTEM code: transmission electron microscopy from first principles." Open Research Europe 1 (2021)
2021
-
[140]
Nion Swift: Open Source Image Processing Software for Instrument Control, Data Acquisition, Organization, Visualization, and Analysis Using Python
Meyer, Chris, et al. "Nion Swift: Open Source Image Processing Software for Instrument Control, Data Acquisition, Organization, Visualization, and Analysis Using Python." Microscopy and Microanalysis 25.S2 (2019): 122-123
2019
-
[142]
Ab‐initio simulations of materials using VASP: Density‐functional theory and beyond
Hafner, Jürgen, "Ab‐initio simulations of materials using VASP: Density‐functional theory and beyond." Journal of computational chemistry 29.13 (2008): 2044-2078
2008
-
[143]
Ab initio molecular dynamics for liquid metals
Kresse, Georg, and Jürgen Hafner, "Ab initio molecular dynamics for liquid metals." Physical review B 47.1 (1993): 558
1993
-
[144]
Real-space grid implementation of the projector augmented wave method
Mortensen, Jens Jørgen, Lars Bruno Hansen, and Karsten Wedel Jacobsen, "Real-space grid implementation of the projector augmented wave method." Physical Review B—Condensed Matter and Materials Physics 71.3 (2005): 035109
2005
-
[146]
Materials modelling using density functional theory: properties and predictions
Giustino, Feliciano. Materials modelling using density functional theory: properties and predictions. Oxford University Press, 2014
2014
-
[147]
LlamaIndex, https://docs.llamaindex.ai/en/stable/examples/llm/llama_2_llama_cpp/
-
[148]
Mistral 7B, the model used, https://huggingface.co/TheBloke/Mistral-7B-Instruct-v0.1-GGUF
-
[149]
OpenAI plugins, https://openai.com/index/chatgpt-plugins/
-
[151]
LangChain, https://www.langchain.com/
-
[152]
OpenAI models, https://platform.openai.com/docs/models
-
[153]
Fast Dash, https://docs.fastdash.app/
-
[154]
RDKit: Open-source cheminformatics; http://www.rdkit.org
-
[155]
Embedchain, https://github.com/embedchain/embedchain
Singh, Taranjeet. Embedchain, https://github.com/embedchain/embedchain
-
[156]
PubChem programmatic access, https://pubchem.ncbi.nlm.nih.gov/docs/programmatic-access
-
[157]
PubChemPy, https://pubchempy.readthedocs.io/en/latest/guide/introduction.html
-
[158]
E., Jalalypour, F., Jordan, J., Kutzner, C., Lemkul, J
Abraham, M., Alekseenko, A., Basov, V., Bergh, C., Briand, E., Brown, A., Doijade, M., Fiorin, G., Fleischmann, S., Gorelov, S., Gouaillardet, G., Grey, A., Irrgang, M. E., Jalalypour, F., Jordan, J., Kutzner, C., Lemkul, J. A., Lundborg, M., Merz, P., … Lindahl, E. (2024). GR...
2024 doi
-
[159]
E., & Snurr, R
Dubbeldam, D., Calero, S., Ellis, D. E., & Snurr, R. Q. (2015). RASPA: molecular simulation software for adsorption and diffusion in flexible nanoporous materials. Molecular Simulation, 42(2), 81–101. https://doi.org/10.1080/08927022.2015.1010082
2015
-
[160]
L., Cococcioni, M., Dabo, I., Dal Corso, A., de Gironcoli, S., Fabris, S., Fratesi, G., Gebauer, R., Gerstmann, U., Gougoussis, C., Kokalj, A., Lazzeri, M., … Wentzcovitch, R
Giannozzi, P., Baroni, S., Bonini, N., Calandra, M., Car, R., Cavazzoni, C., Ceresoli, D., Chiarotti, G. L., Cococcioni, M., Dabo, I., Dal Corso, A., de Gironcoli, S., Fabris, S., Fratesi, G., Gebauer, R., Gerstmann, U., Gougoussis, C., Kokalj, A., Lazzeri, M., … Wentzcovitch,...
2009 doi
-
[161]
Chithrananda, G
S. Chithrananda, G. Grand, and B. Ramsundar, ``ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction,'' arXiv preprint arXiv:2010.09885, 2020. Available: https://arxiv.org/abs/2010.09885
2010 arXiv
-
[162]
G. Zhou, Z. Gao, Q. Ding, H. Zheng, H. Xu, Z. Wei, et al., ``Uni-Mol: A Universal 3D Molecular Representation Learning Framework,'' ChemRxiv, 2022, doi:10.26434/chemrxiv-2022-jjm0j
2022 doi
-
[163]
B. Yu, F. N. Baker, Z. Chen, X. Ning, and H. Sun, ``LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset,'' arXiv preprint arXiv:2402.09391, 2024. Available: https://arxiv.org/abs/2402.09391
2024 arXiv
-
[164]
S. Liu, J. Wang, Y. Yang, C. Wang, L. Liu, H. Guo, and C. Xiao, ``ChatGPT-powered Conversational Drug Editing Using Retrieval and Domain Feedback,'' arXiv preprint arXiv:2305.18090, 2023. Available: https://arxiv.org/abs/2305.18090
2023 arXiv
-
[165]
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. Le Scao, T. Lavril, T. Wang, T. Lacroix, and W. El Sayed, ``Mistral 7B,'' arXiv preprint arXiv:2310.0...
-
[166]
Dettmers, A
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, ``QLoRA: Efficient Finetuning of Quantized LLMs,'' arXiv preprint arXiv:2305.14314, 2023. Available: https://arxiv.org/abs/2305.14314
2023 arXiv
-
[167]
Zaki, M., & Krishnan, N. A. (2024). MaScQA: investigating materials science knowledge of large language models. Digital Discovery, 3(2), 313-327
2024
-
[168]
OpenAI. (2024). GPT-3.5-turbo. https://openai.com/api/
2024
-
[169]
Marp. (2024). Markdown Presentation Ecosystem. https://marp.app/
2024
-
[170]
K.; Mentha, S
Mishra, R. K.; Mentha, S. S.; Misra, Y.; Dwivedi, N. Emerging Pollutants of Severe Environmental Concern in Water and Wastewater: A Comprehensive Review on Current Developments and Future Research. Water-Energy Nexus 2023, 6, 74–95. https://doi.org/10.1016/j.wen.2023.08.002
2023 doi
-
[171]
A Microscopic Survey on Microplastics in Beverages: The Case of Beer, Mineral Water and Tea
Li, Y.; Peng, L.; Fu, J.; Dai, X.; Wang, G. A Microscopic Survey on Microplastics in Beverages: The Case of Beer, Mineral Water and Tea. Analyst 2022, 147 (6), 1099–1105. https://doi.org/10.1039/D2AN00083K
2022 doi
-
[172]
M. L. Evans and J. D. Bocarsly. datalab, July 2024. URL https://github.com/datalab-org doi:10.5281/zenodo.12545475
2024 doi
-
[173]
14 examples of how LLMs can transform materials science and chemistry: a reflection on a large language model hackathon
Jablonka et al. 14 examples of how LLMs can transform materials science and chemistry: a reflection on a large language model hackathon. Digital Discovery, 2023. doi:10.1039/D3DD00113J
2023 doi
-
[175]
https://github.com/ndaelman-hu/nomad_query_reporter
- [176]
- [177]
-
[178]
Evans, M
Development and applications of the OPTIMADE API for materials discovery, design, and data exchange. Evans, M. L., et al., Digital Discovery (2024), DOI: doi.org/10.1039/D4DD00039K
2024 doi
-
[179]
Wilkinson, M., Dumontier, M., Aalbersberg, I. et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data 3, 160018 (2016) https://doi.org/10.1038/sdata.2016.18
2016 doi
-
[180]
https://json-schema.org/
-
[181]
NOMAD: A distributed web-based platform for managing materials science research data
Scheidgen et al., (2023). NOMAD: A distributed web-based platform for managing materials science research data. Journal of Open Source Software, 8(90), 5388, https://doi.org/10.21105/joss.05388
2023 doi
-
[182]
https://pypi.org/project/SpeechRecognition/
-
[183]
https://pypi.org/project/openai-whisper/
-
[184]
https://pypi.org/project/langchain-experimental/
-
[185]
Active learning literature survey
Settles, Burr. "Active learning literature survey." (2009)
2009
-
[186]
Romano, and George Casella
Lehmann, Erich Leo, Joseph P. Romano, and George Casella. Testing statistical hypotheses. Vol. 3. New York: springer, 1986
1986
-
[187]
The logic of scientific discovery
Popper, Karl. The logic of scientific discovery. Routledge, 2005
2005
-
[188]
The first room-temperature ambient-pressure superconductor
Lee, Sukbae, Ji-Hoon Kim, and Young-Wan Kwon. "The first room-temperature ambient-pressure superconductor." arXiv preprint arXiv:2307.12008 (2023)
2023 arXiv
-
[189]
LK-99 Is the Superconductor of the Summer
Chang, Kenneth. “LK-99 Is the Superconductor of the Summer.” New York Times (2023)
2023
-
[190]
Claimed superconductor LK-99 is an online sensation—But replication efforts fall short
Garisto, Dan. "Claimed superconductor LK-99 is an online sensation—But replication efforts fall short." Nature 620, no. 7973 (2023): 253-253
2023
-
[191]
Natural language inference in context-investigating contextual reasoning over long texts
Liu, Hanmeng, Leyang Cui, Jian Liu, and Yue Zhang. "Natural language inference in context-investigating contextual reasoning over long texts." In Proceedings of the AAAI conference on artificial intelligence, vol. 35, no. 15, pp. 13388-13396. 2021
2021
-
[192]
Conjugate Bayesian analysis of the Gaussian distribution
Murphy, Kevin P. "Conjugate Bayesian analysis of the Gaussian distribution." def 1, no. 2 2 (2007): 16
2007
-
[193]
Will the LK-99 room temp superconductivity pre-print replicate in 2023
“Will the LK-99 room temp superconductivity pre-print replicate in 2023”. Manifold Markets (2024). https://manifold.markets/Ernie/will-the-lk99-room-temp-ambient-pre-17fc7cb7a2a0
2024
-
[194]
Large Language Models for Automated Open-domain Scientific Hypotheses Discovery
Yang, Zonglin, et al. "Large Language Models for Automated Open-domain Scientific Hypotheses Discovery." arXiv preprint arXiv:2309.02726 (2023). https://arxiv.org/pdf/2309.02726
2023 arXiv
-
[195]
Tree of thoughts: Deliberate problem solving with large language models
Yao, Shunyu, et al. "Tree of thoughts: Deliberate problem solving with large language models." Advances in Neural Information Processing Systems 36 (2024)
2024
-
[196]
https://huggingface.co/datasets/AtlasUnified/Atlas-Reasoning/commits/main
-
[197]
https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2
-
[198]
K. M. Jablonka, P. Schwaller, A. Ortega-Guerrero, B. Smit, Nat. Mach. Intell., 2024, 6, 161–169
2024
-
[199]
D. A. Boiko, R. MacKnight, B. Kline, G. Gomes, Nature, 2023, 624, 570–578
2023
-
[200]
Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, C. Zhu, Proc. 2023 Conf. Empir. Methods Nat. Lang. Process., 2023, 2511–2522
2023
-
[201]
Mangrulkar, S
S. Mangrulkar, S. Gugger, L. Debut, Y. Belkada, S. Paul, B. Bossan, PEFT: State-of-the-art Parameter-Efficient Fine-Tuning Methods; GitHub: https://github.com/huggingface/peft, 2022
2022
-
[202]
Zhang, G
P. Zhang, G. Zeng, T. Wang, W. Lu, TinyLlama: An Open-Source Small Language Model; arXiv:2401.02385
-
[203]
Zhang, S
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, L. Zettlemoyer, OPT: Open Pre-trained Transformer Language Models; arXiv:2205.01068
-
[204]
Bethesda (MD): National Library of Medicine (US), National Center for Biotechnology Information; https://www.ncbi.nlm.nih.gov/, 1998
National Center for Biotechnology Information (NCBI) [Internet]. Bethesda (MD): National Library of Medicine (US), National Center for Biotechnology Information; https://www.ncbi.nlm.nih.gov/, 1998
1998
-
[205]
Al-Feghali, S
A. Al-Feghali, S. Zhang, G-Peer-T; GitHub: https://github.com/alxfgh/G-Peer-T, 2024
2024
-
[206]
Translation between molecules and natural language
Carl Edwards, Tuan Lai, Kevin Ros, Garrett Honke, Kyunghyun Cho, and Heng Ji. Translation between molecules and natural language. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing...
2022 doi
-
[207]
Isobench: Benchmarking multimodal foundation models on isomorphic representations, 2024
Deqing Fu, Ghazal Khalighinejad, Ollie Liu, Bhuwan Dhingra, Dani Yogatama, Robin Jia, and Willie Neiswanger. Isobench: Benchmarking multimodal foundation models on isomorphic representations, 2024. URL https://arxiv.org/abs/2404.01266
2024 arXiv
-
[208]
Chawla, Olaf Wiest, and Xiangliang Zhang
Taicheng Guo, Kehan Guo, Bozhao Nan, Zhenwen Liang, Zhichun Guo, Nitesh V. Chawla, Olaf Wiest, and Xiangliang Zhang. What can large language models do in chemistry? a comprehensive benchmark on eight tasks, 2023. URL https://arxiv.org/abs/2305.18365
2023 arXiv
-
[209]
Chemformer: a pre-trained transformer for computational chemistry
Ross Irwin, Spyridon Dimitriadis, Jiazhen He, and Esben Jannik Bjerrum. Chemformer: a pre-trained transformer for computational chemistry. Machine Learning: Science and Technology, 3(1):015022, January 2022. doi: 10.1088/2632-2153/ac3ffb. URL https://dx.doi.org/10.1088/2632-21...
2022 doi
-
[210]
ChemQA: a multimodal question-and-answering dataset on chemistry reasoning
Shang Zhu, Xuefeng Liu, and Ghazal Khalighinejad. ChemQA: a multimodal question-and-answering dataset on chemistry reasoning. https://huggingface.co/datasets/shangzhu/ChemQA, 2024
2024
-
[211]
Data-driven electrolyte design for lithium metal anodes
Kim, S.C., et al. "Data-driven electrolyte design for lithium metal anodes." Proceedings of the National Academy of Sciences 120.10 (2023): e2214357120. https://www.pnas.org/doi/full/10.1073/pnas.2214357120
2023 doi
-
[212]
https://unstructured.io/
-
[213]
https://platform.openai.com/docs/models
-
[214]
https://llama.meta.com/llama3/
-
[215]
Ai, Q.; Meng, F.; Shi, J.; Pelkie, B.; Coley, C. W. Extracting Structured Data from Organic Synthesis Procedures Using a Fine-Tuned Large Language Model. ChemRxiv April 8, 2024. https://doi.org/10.26434/chemrxiv-2024-979fz
2024 doi
-
[216]
https://github.com/qai222/ontosynthesis (accessed 2024-07-07)
Ontosynthesis, 2024. https://github.com/qai222/ontosynthesis (accessed 2024-07-07)
2024
-
[217]
J.; Karan, D.; Lee, K
Bai, J.; Mosbach, S.; Taylor, C. J.; Karan, D.; Lee, K. F.; Rihm, S. D.; Akroyd, J.; Lapkin, A. A.; Kraft, M. A Dynamic Knowledge Graph Approach to Distributed Self-Driving Laboratories. Nat. Commun. 2024, 15 (1), 462. https://doi.org/10.1038/s41467-023-44599-9
2024 doi
-
[218]
Synthesis Operation Ontology
Ai, Q.; Klein, C. Synthesis Operation Ontology. GitHub. https://github.com/qai222/ontosynthesis/blob/main/ontologies/soo/soo.md (accessed 2024-07-07)
2024
-
[219]
https://dash.plotly.com/cytoscape (accessed 2024-07-06)
Dash-Cytoscape: A Component Library for Dash Aimed at Facilitating Network Visualization in Python, Wrapped around Cytoscape.Js. https://dash.plotly.com/cytoscape (accessed 2024-07-06)
2024
-
[220]
E., III; Jayaraman, A
Gartner, T. E., III; Jayaraman, A. Modeling and Simulations of Polymers: A Roadmap. Macromolecules 2019, 52 (3), 755-786. DOI: 10.1021/acs.macromol.8b01836
2019 doi
-
[221]
GPT-3.5 Turbo. 2024. https://platform.openai.com/docs/models/gpt-3-5-turbo (accessed 2024/07/10)
2024
-
[222]
Welcome to GraphRAG. 2024. https://microsoft.github.io/graphrag/ (accessed 2024/07/10)
2024
-
[223]
From Local to Global: A Graph RAG Approach to Query-Focused Summarization
Edge, D.; Trinh, H.; Cheng, N.; Bradley, J.; Chao, A.; Mody, A.; Truitt, S.; Larson, J. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. 2024. (acccessed 2024/07/10)
2024
-
[224]
LlamaIndex. 2024. https://docs.llamaindex.ai/en/stable/ (accessed 2024/07/10)
2024
-
[225]
Focassio, B., Freitas, M., Schleder, G.R. (2024). Performance assessment of universal machine learning interatomic potentials: Challenges and directions for materials’ surfaces. ACS Applied Materials & Interfaces
2024
-
[226]
Marques, F., Balcerzak, M., Winkelmann, F., Zepon, G., Felderhoff, M. (2021). Review and outlook on high-entropy alloys for hydrogen storage. Energy & Environmental Science, 14(10), 5191-5227
2021
-
[227]
Jain, S.M. (2022). Hugging face: Introduction to transformers for NLP with the hugging face library and models to solve problems. Berkeley, CA: Apress
2022
-
[228]
Lewis, P., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459-9474
2020
-
[229]
OpenAI. (2020). OpenAI GPT-3: Language models are few-shot learners. Retrieved from https://openai.com/blog/openai-api
2020
-
[230]
OpenAI. (2023). New embedding models and API updates. Retrieved from https://openai.com/index/new-embedding-models-and-api-updates
2023
-
[231]
Generative models for molecular discovery: Recent advances and challenges
Bilodeau, Camille et al. (2022). “Generative models for molecular discovery: Recent advances and challenges”. In: Wiley Interdisciplinary Reviews: Computational Molecular Science 12.5, e1608
2022
-
[232]
Chennakesavalu, Shriram et al. (2024). Energy Rank Alignment: Using Preference Optimization to Search Chemical Space at Scale. DOI: 10.48550/ARXIV.2405.12961. URL: https://arxiv. org/abs/2405.12961
2024 doi
-
[233]
Extracting medicinal chemistry intuition via preference machine learning
Choung, Oh-Hyeon et al. (Oct. 2023). “Extracting medicinal chemistry intuition via preference machine learning”. In: Nature Communications 14.1. ISSN: 2041-1723. DOI: 10.1038/s41467- 023-42242-1. URL: http://dx.doi.org/10.1038/s41467-023-42242-1
2023 doi
-
[234]
Leveraging large language models for predictive chemistry
Jablonka, Kevin Maik et al. (Feb. 2024). “Leveraging large language models for predictive chemistry”. In: Nature Machine Intelligence 6.2, pp. 161–169. ISSN: 2522-5839. DOI: 10.1038/s42256-023-00788-1. URL: http://dx.doi.org/10.1038/s42256-023-00788-1
2024 doi
- [235]
- [236]
-
[237]
Copier template: Available at https://github.com/copier-org/copier
-
[238]
T.; Moazam, H.; Miller, H.; Zaharia, M.; Potts, C
DSPy: Khattab, O.; Singhvi, A.; Maheshwari, P.; Zhang, Z.; Santhanam, K.; Vardhamanan, S.; Haq, S.; Sharma, A.; Joshi, T. T.; Moazam, H.; Miller, H.; Zaharia, M.; Potts, C. DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines. preprint arXiv:2310.03714. 2023
-
[239]
PyMuPDF: Available at ttps://github.com/pymupdf/PyMuPDF
-
[240]
Available at \\ https://platform.openai.com/docs/models/gpt-3-5-turbo
GPT-3.5-Turbo: OpenAI. Available at \\ https://platform.openai.com/docs/models/gpt-3-5-turbo
-
[241]
Available at \\ https://platform.openai.com/docs/models/gpt-4-turbo-and-gpt-4
GPT-4-Turbo: OpenAI. Available at \\ https://platform.openai.com/docs/models/gpt-4-turbo-and-gpt-4
-
[242]
DSPy Typed Predictors: Documentation at \\ https://dspy-docs.vercel.app/docs/building-blocks/typed_predictors
-
[243]
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Chain-of-Thought prompting: Wei, J.; Wang X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.; Zhou, D. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. preprint arXiv: 2201.11903. 2023
2023 arXiv
-
[244]
M.; Heyer, A
Test review article on zeolites: Rhoda, H. M.; Heyer, A. J.; Snyder, B. E. R.; Plessers, D.; Bols, M. L.; Schoonheydt, R. A.; Sels, B. F.; Solomon, E. I. Second-Sphere Lattice Effects in Copper and Iron Zeolite Catalysis. Chem. Rev. 2022, 122, 12207–12243
2022
-
[245]
Neo4J: Documentation at https://neo4j.com/
-
[246]
Graph Maker: Available at https://github.com/rahulnyk/graph_maker
-
[247]
Gradio: Available at https://github.com/gradio-app/gradio
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.