REVIEW 5 major objections 6 minor 1 cited by
Computational Protein Science in the Era of Large Language Models (LLMs)
T0 review · 5 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Protein language models—LLMs trained on amino acid sequences, structures, and scientific text—can grasp the “grammar” and “semantics” of proteins and be adapted to a broad range of structure prediction, function prediction, and protein…
desk verdict A useful, well-organized survey of pLMs whose abstract overclaims generalization; worth refereeing after revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the pLM itself, treated as a foundation model: a Transformer (or state-space) network pre-trained on large protein corpora so that amino acid tokens behave like words, with “grammar” in residue patterns and “semantics” in encoded structure and function. The paper identifies three technical routes that carry downstream transfer: (1) representation extraction, where frozen or fine-tuned pLM encodings feed prediction heads; (2) likelihood inference, where the model's probability of a mutant versus wild-type sequence is used as a zero-shot fitness score; and (3) prompting and instruction tuning, where a unified decoder answers protein questions or generates sequences under text or function control. Structure tokenization (VQ-VAE-derived tokens such as 3Di) is a notable sub-mechanism that lets 3D structure enter language-model training as discrete tokens.
What would settle it
A controlled benchmark that runs sequence-only, structure-enhanced, and multimodal pLMs on a fixed set of structure, function, and design tasks would settle the central claim: for example, if SaProt and ProLLaMA do not systematically beat ESM-2 on held-out tasks, the paper's hierarchy of protein knowledge fails, and if no pLM category transfers beyond task-specific models trained from scratch, the generalization story collapses.
Extended reading notes
Core claim
The paper's central claim is that protein language models “skillfully grasp the foundational knowledge of proteins and can be effectively generalized to solve a diversity of sequence-structure-function reasoning problems.” The evidence it assembles is a taxonomy: sequence-only pLMs (ESM-2, ProtGPT2, xTrimoPGLM) capture evolutionarily favored amino acid patterns; structure- and function-enhanced pLMs (SaProt, ESM-3) add explicit 3D and annotation knowledge; multimodal pLMs (ProLLaMA, BioT5) bridge protein sequences with natural language and molecule languages. On top of these foundations, the paper reports pLM-based single-sequence structure prediction comparable to MSA-based methods, zero-shot fitness and mutation-effect prediction, text-guided and condition-tagged protein generation, and ChatGPT-like protein question answering. The intended conclusion is that pLMs, not bespoke per-task models, now carry the main line of computational protein science.
Load-bearing premise
The survey's usefulness rests on the assumption that its chosen models and papers give a representative map of the field; no search protocol or inclusion criteria are stated, so the taxonomy could miss important pLMs or misweight the landscape.
Editorial extensions
If this is right
- Single-sequence structure prediction methods such as ESMFold can replace slow MSA searches for many proteins, making structure inference practical for orphan and fast-evolving proteins.
- Zero-shot mutation-effect scoring by pLMs gives experimental labs a cheap first pass at fitness landscapes before deep mutational scanning.
- Function prediction moves from many task-specific models to unified question-answering systems that answer property, annotation, site, and interaction questions with one model.
- Conditional and text-guided pLMs extend protein design beyond redesign, generating de novo sequences, antibodies targeting new variants, and enzymes with improved stability.
- Structure- and function-enhanced pLMs, including multi-track models like ESM-3, suggest that sequence, structure, and function can be treated as interchangeable token tracks in one generative model.
Reading between the lines
- Beyond the survey's claims, the taxonomy suggests that the next wave of pLMs will blur its three categories: sequence-only models will acquire structure tokens, and multimodal models will absorb molecule and text modalities, so the field's frontier becomes integration cost rather than architecture.
- A testable extension of the paper's generalizability claim: on fixed benchmarks, pLM-based methods should dominate task-specific models trained from scratch whenever labeled data are scarce; a controlled comparison across ProteinGym-like tasks would settle this.
- Because pLM likelihood is used as a proxy for evolutionary plausibility, species bias in training databases may leak into fitness predictions; adjusting training data composition is a direct extension of the redesign workflow the paper describes.
- The survey's coverage assumption can be checked by re-running its categorization against a systematic search of recent pLMs; if major models fall outside the three categories, the proposed map needs revision.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of protein language models (pLMs) and their applications in computational protein science. It proposes a taxonomy that divides pLMs into sequence-based, structure-and-function-enhanced, and multimodal models; reviews how pLMs are used for structure prediction, function prediction, and protein design; and discusses applications in antibody design, enzyme design, and drug discovery. The Abstract states that pLMs 'skillfully grasp the foundational knowledge of proteins and can be effectively generalized to solve a diversity of sequence-structure-function reasoning problems,' and the survey is organized to support this claim by mapping pLMs to downstream tasks.
Significance. If the survey's claims hold, it would provide a valuable entry point for researchers from both AI and biology backgrounds. The taxonomy is thoughtful and the coverage is unusually broad, including recent models such as ESM-3, DPLM-2, and xTrimoPGLM, as well as practical applications with wet-lab validation. The detailed tables (Tables 1–4) that list corpora, architectures, parameter counts, and pre-training objectives are a useful resource. The paper does not present original machine-checked proofs or reproducible code, but it cites primary sources for most claims, which is appropriate for a survey. The main contribution is organizational rather than experimental, and its usefulness depends on whether the selected literature is representative and whether the claims in the Abstract are calibrated to the evidence.
major comments (5)
- [Abstract and Section 4] The central claim that pLMs are 'effectively generalized to solve a diversity of sequence-structure-function reasoning problems' is broader than the evidence presented in Section 4. In Section 4.1, all pLM-based structure prediction methods (ESMFold, HelixFold-Single, OmegaFold, trRosettaX-Single, RGN2, IgFold) require an additional trained folding trunk or prediction head. In Section 4.2, most function prediction methods use the LM-as-encoder scheme with a learned classifier, fine-tuning, or parameter-efficient fine-tuning; only likelihood-based fitness and mutation-effect prediction are described as zero-shot (Section 4.2.1). In Section 4.3, ProGen and ZymCTRL require control tags or fine-tuning for controllable generation. The survey thus conflates 'adaptable via task-specific training' with 'generalization.' I recommend softening the Abstract to 'adaptable across tasks' or explicitly scoping the generalization claim to the zero-shot settings where evidence exists.
- [Section 3, first paragraph] The survey does not state its literature selection method. No databases searched, search dates, keywords, or inclusion/exclusion criteria are provided. This makes the taxonomy non-reproducible and leaves open the possibility that important pLMs were omitted or that the relative emphasis of topics is not representative. As the survey's central contribution is its categorization, a short methodology paragraph describing the search protocol and time window is needed.
- [Section 4.1, single-sequence structure prediction paragraph] The sentence 'In investigations, OmegaFold, trRosettaX-Single, and RGN2 are all observed to outperform AlphaFold2 and RoseTTAFold on those orphan proteins and de novo designed proteins' is a specific quantitative claim with no citation. Please provide references for these comparisons or remove the claim. As written, it is a load-bearing assertion for the section's argument that pLM-based methods address the limitations of MSA-based approaches.
- [Section 4.2.1] The statement 'the likelihoods inferred from pLMs correlate well with protein fitness [72, 253, 254]' is central to the zero-shot generalization narrative, yet no quantitative evidence is given. The survey should report representative correlation values (e.g., Spearman rho) from ProteinGym or the cited works, or state the range across benchmarks. Without those numbers, the claim is too vague to evaluate.
- [Section 4 (general)] The survey seldom reports how pLM-based methods compare with strong non-pLM baselines. For structure prediction, comparisons to AlphaFold2 and RoseTTAFold appear in Section 4.1, but for function prediction (Section 4.2) and protein design (Section 4.3) the text rarely mentions results from BLAST, HMMER, profile-based predictors, or classical machine-learning methods. Without this comparative context, the reader cannot judge whether pLMs are necessary or superior for these tasks. I recommend adding a comparative synthesis from the cited benchmarks (e.g., ProteinGym, FLIP, TAPE) or explicitly softening claims of pLM advantage.
minor comments (6)
- [Figure 8 caption] The caption contains the typo 'Workfolw'; it should read 'Workflow'.
- [Table 4] The table header spells 'Functional' as 'Funcitional', and the scheme abbreviation 'Mull-Model Fine-Tuning' should be 'Full-Model Fine-Tuning'.
- [Figure 9 caption] The second framework in the caption is labeled 'LM-as-Encoder' but the text and Table 4 use 'LM-as-Predictor'; the caption should be corrected for consistency.
- [Section 3.1.1 and Table 1] The model 'paired-IgGen' is abbreviated 'p-IgGen' in the text but 'g-IgGen' in Table 1; the abbreviation should be unified.
- [Section 3.1.1, BiMamba-S and Table 1] The Mamba architecture is cited via references [96] and [97], which are a survey and a recommendation paper; the original Mamba paper (Gu & Dao, reference [139]) should be cited at first mention.
- [Section 6.5] The phrase 'reached the unanimous conclusion of non-optimal' is too strong for two cited studies; a softer formulation such as 'several recent studies have concluded' would be more accurate.
Circularity Check
No significant circularity: this is a literature survey whose taxonomy and claims rest on external, independently published methods.
full rationale
The paper is a survey, not a derivation or prediction pipeline. Its central claim—that protein language models learn foundational protein knowledge and can be adapted to structure prediction, function prediction, and design—is supported by citations to externally published models (ESM-2, AlphaFold2, ProGen, ProteinGym, etc.) and to benchmarks with their own evaluations. There is no fitted parameter that is later renamed as a prediction, no equation that is assumed and then re-derived, and no uniqueness theorem invoked to force the survey's organizational choice. The paper's three-way taxonomy (sequence-based, structure/function-enhanced, multimodal pLMs) is a literature-organizing scheme, and its validity does not depend on any of the cited results being equivalent to the taxonomy itself. A few references authored by the survey's own team (e.g., prior surveys on LLMs in recommender systems and retrieval-augmented generation) appear only as background citations and are not load-bearing for the paper's claims about protein language models. The skeptic's concern that 'effectively generalized' overstates the evidence is a legitimate correctness/calibration critique, not a circularity critique: the survey itself documents that most applications use task-specific heads, fine-tuning, or architectural additions, but overstatement of a literature-based conclusion does not make the argument circular. Therefore, no circular step can be exhibited, and the honest finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption The sequence-structure-function paradigm holds: amino acid sequence determines 3D structure, and structure determines function.
- domain assumption Sequence-only pLMs capture implicit structural and functional knowledge from large-scale pre-training.
- domain assumption pLM likelihood scores correlate with protein fitness.
- domain assumption Scaling laws observed for natural-language LLMs transfer to protein language models.
Cite this review
Pith. "Pith review of Computational Protein Science in the Era of Large Language Models (LLMs)." pith.science (2026). https://pith.science/paper/DQZLB4BP
@misc{pith2026250110282,
author = {Pith},
title = {Pith review of: Computational Protein Science in the Era of Large Language Models (LLMs)},
year = {2026},
howpublished = {\url{https://pith.science/paper/DQZLB4BP}},
note = {Machine review of arXiv:2501.10282}
}
read the original abstract
Considering the significance of proteins, computational protein science has always been a critical scientific field, dedicated to revealing knowledge and developing applications within the protein sequence-structure-function paradigm. In the last few decades, Artificial Intelligence (AI) has made significant impacts in computational protein science, leading to notable successes in specific protein modeling tasks. However, those previous AI models still meet limitations, such as the difficulty in comprehending the semantics of protein sequences, and the inability to generalize across a wide range of protein modeling tasks. Recently, LLMs have emerged as a milestone in AI due to their unprecedented language processing & generalization capability. They can promote comprehensive progress in fields rather than solving individual tasks. As a result, researchers have actively introduced LLM techniques in computational protein science, developing protein Language Models (pLMs) that skillfully grasp the foundational knowledge of proteins and can be effectively generalized to solve a diversity of sequence-structure-function reasoning problems. While witnessing prosperous developments, it's necessary to present a systematic overview of computational protein science empowered by LLM techniques. First, we summarize existing pLMs into categories based on their mastered protein knowledge, i.e., underlying sequence patterns, explicit structural and functional information, and external scientific languages. Second, we introduce the utilization and adaptation of pLMs, highlighting their remarkable achievements in promoting protein structure prediction, protein function prediction, and protein design studies. Then, we describe the practical application of pLMs in antibody design, enzyme design, and drug discovery. Finally, we specifically discuss the promising future directions in this fast-growing field.
Figures
Figures from the paper (10 more)
Forward citations
Cited by 1 Pith paper
-
HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous Structure Tokens
HD-Prot shows that a protein language model can jointly generate sequences and structures using continuous structure tokens instead of quantized tokens, reaching competitive performance on four protein design tasks.
Reference graph
Works this paper leans on
-
[1]
Controllable protein design with language models
Noelia Ferruz and Birte Höcker. Controllable protein design with language models. Nature Machine Intelligence, 4(6):521–532, 2022
2022
-
[2]
Learning the protein language: Evolution, structure, and function
Tristan Bepler and Bonnie Berger. Learning the protein language: Evolution, structure, and function. Cell systems , 12(6):654–669, 2021
2021
-
[3]
Principles that govern the folding of protein chains
Christian B Anfinsen. Principles that govern the folding of protein chains. Science, 181(4096):223–230, 1973
1973
-
[4]
Exploring the structure and function paradigm
Oliver C Redfern, Benoit Dessailly, and Christine A Orengo. Exploring the structure and function paradigm. Current opinion in structural biology, 18(3):394–402, 2008
2008
-
[5]
Natural selection and the concept of a protein space
John Maynard Smith. Natural selection and the concept of a protein space. Nature, 225(5232):563–564, 1970
1970
-
[6]
The language of pro- teins: Nlp, machine learning & protein sequences
Dan Ofer, Nadav Brandes, and Michal Linial. The language of pro- teins: Nlp, machine learning & protein sequences. Computational and Structural Biotechnology Journal, 19:1750–1758, 2021
2021
-
[7]
Unified rational protein engineering with sequence-based deep representation learning
Ethan C Alley, Grigory Khimulya, Surojit Biswas, Mohammed AlQuraishi, and George M Church. Unified rational protein engineering with sequence-based deep representation learning. Nature methods, 16(12):1315–1322, 2019
2019
-
[8]
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021
2021
Show all 300 references
-
[9]
Accurate prediction of protein structures and interactions using a three-track neural network
Minkyung Baek, Frank DiMaio, Ivan Anishchenko, Justas Dau- paras, Sergey Ovchinnikov, Gyu Rie Lee, Jue Wang, Qian Cong, Lisa N Kinch, R Dustin Schaeffer, et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science, 373(6557):87...
2021
-
[10]
Deepgoplus: im- proved protein function prediction from sequence
Maxat Kulmanov and Robert Hoehndorf. Deepgoplus: im- proved protein function prediction from sequence. Bioinformatics, 36(2):422–429, 2020
2020
-
[11]
Prediction of designer-recombinases for dna editing with generative deep learning
Lukas Theo Schmitt, Maciej Paszkowski-Rogacz, Florian Jug, and Frank Buchholz. Prediction of designer-recombinases for dna editing with generative deep learning. Nature Communications, 13(1):7966, 2022
2022
-
[12]
Ig-vae: Generative modeling of protein structure by direct 3d coordinate generation
Raphael R Eguchi, Christian A Choe, and Po-Ssu Huang. Ig-vae: Generative modeling of protein structure by direct 3d coordinate generation. PLoS computational biology, 18(6):e1010271, 2022
2022
-
[13]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018
2018 arXiv
-
[14]
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67, 2020
2020
-
[15]
Improving language understanding by generative pre- training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre- training. 2018
2018
-
[16]
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019
2019
-
[17]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020
1901
-
[18]
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 , 2023
2023 arXiv
-
[19]
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 2023
2023 arXiv
-
[20]
Recommender systems in the era of large language models (llms)
Zihuai Zhao, Wenqi Fan, Jiatong Li, Yunqing Liu, Xiaowei Mei, Yiqi Wang, Zhen Wen, Fei Wang, Xiangyu Zhao, Jiliang Tang, et al. Recommender systems in the era of large language models (llms). IEEE Transactions on Knowledge and Data Engineering, 2024
2024
-
[21]
Tokenrec: Learning to tokenize id for llm-based generative recommendation
Haohao Qu, Wenqi Fan, Zihuai Zhao, and Qing Li. Tokenrec: Learning to tokenize id for llm-based generative recommendation. arXiv preprint arXiv:2406.10450, 2024
2024 arXiv
-
[22]
A survey of large language models for healthcare: from data, technology, and applications to accountabil- ity and ethics
Kai He, Rui Mao, Qika Lin, Yucheng Ruan, Xiang Lan, Mengling Feng, and Erik Cambria. A survey of large language models for healthcare: from data, technology, and applications to accountabil- ity and ethics. arXiv preprint arXiv:2310.05694, 2023
-
[23]
To transformers and beyond: Large language models for the genome
Micaela E Consens, Cameron Dufault, Michael Wainberg, Duncan Forster, Mehran Karimzadeh, Hani Goodarzi, Fabian J Theis, Alan Moses, and Bo Wang. To transformers and beyond: Large language models for the genome. arXiv preprint arXiv:2311.07621, 2023
2023 arXiv
-
[24]
Empowering molecule discovery for molecule-caption translation with large language models: A chatgpt perspective
Jiatong Li, Yunqing Liu, Wenqi Fan, Xiao-Yong Wei, Hui Liu, Jiliang Tang, and Qing Li. Empowering molecule discovery for molecule-caption translation with large language models: A chatgpt perspective. IEEE Transactions on Knowledge and Data Engineering, 2024
2024
-
[25]
Evolutionary-scale prediction of atomic-level protein structure with a language model
Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637):1123– 1130, 2023
2023
-
[26]
Protgpt2 is a deep unsupervised language model for protein design
Noelia Ferruz, Steffen Schmidt, and Birte Höcker. Protgpt2 is a deep unsupervised language model for protein design. Nature communications, 13(1):4348, 2022
2022
-
[27]
xtrimopglm: unified 100b-scale pre-trained transformer for deciphering the language of protein
Bo Chen, Xingyi Cheng, Pan Li, Yangli-ao Geng, Jing Gong, Shen Li, Zhilei Bei, Xu Tan, Boyan Wang, Xin Zeng, et al. xtrimopglm: unified 100b-scale pre-trained transformer for deciphering the language of protein. arXiv preprint arXiv:2401.06199, 2024
2024 arXiv
-
[28]
Msa transformer
Roshan M Rao, Jason Liu, Robert Verkuil, Joshua Meier, John Canny, Pieter Abbeel, Tom Sercu, and Alexander Rives. Msa transformer. In International Conference on Machine Learning, pages 8844–8856. PMLR, 2021
2021
-
[29]
Saprot: protein language modeling with structure- aware vocabulary
Jin Su, Chenchen Han, Yuyang Zhou, Junjie Shan, Xibin Zhou, and Fajie Yuan. Saprot: protein language modeling with structure- aware vocabulary. bioRxiv, pages 2023–10, 2023
2023
-
[30]
Simulating 500 million years of evolution with a language model
Thomas Hayes, Roshan Rao, Halil Akin, Nicholas J Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q Tran, Jonathan Deaton, Marius Wiggert, et al. Simulating 500 million years of evolution with a language model. Science, page eads0018, 2025
2025
-
[31]
Prollama: A protein large language model for multi-task protein language processing
Liuzhenghao Lv, Zongying Lin, Hao Li, Yuyang Liu, Jiaxi Cui, Calvin Yu-Chian Chen, Li Yuan, and Yonghong Tian. Prollama: A protein large language model for multi-task protein language processing. arXiv preprint arXiv:2402.16445, 2024
2024 arXiv
-
[32]
Biot5: Enriching cross-modal integration in biology with chemical knowledge and natural language associations
Qizhi Pei, Wei Zhang, Jinhua Zhu, Kehan Wu, Kaiyuan Gao, Lijun Wu, Yingce Xia, and Rui Yan. Biot5: Enriching cross-modal integration in biology with chemical knowledge and natural language associations. arXiv preprint arXiv:2310.07276, 2023. JOURNAL OF LATEX CLASS FILES, VOL. ...
-
[33]
Protein- chat: Towards achieving chatgpt-like functionalities on protein 3d structures
Han Guo, Mingjia Huo, Ruiyi Zhang, and Pengtao Xie. Protein- chat: Towards achieving chatgpt-like functionalities on protein 3d structures. Authorea Preprints, 2023
2023
-
[34]
Proteinnpt: improving protein property prediction and design with non-parametric transformers
Pascal Notin, Ruben Weitzman, Debora Marks, and Yarin Gal. Proteinnpt: improving protein property prediction and design with non-parametric transformers. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[35]
Large language models generate functional protein sequences across diverse families
Ali Madani, Ben Krause, Eric R Greene, Subu Subramanian, Benjamin P Mohr, James M Holton, Jose Luis Olmos, Caiming Xiong, Zachary Z Sun, Richard Socher, et al. Large language models generate functional protein sequences across diverse families. Nature Biotechnology, 41(8):1099...
2023
-
[36]
Scientific large language models: A survey on biological & chemical domains
Qiang Zhang, Keyang Ding, Tianwen Lyv, Xinda Wang, Qingyu Yin, Yiwen Zhang, Jing Yu, Yuhao Wang, Xiaotong Li, Zhuoyi Xiang, et al. Scientific large language models: A survey on biological & chemical domains. arXiv preprint arXiv:2401.14656, 2024
2024 arXiv
-
[37]
Protein language models and structure prediction: Connection and progression
Bozhen Hu, Jun Xia, Jiangbin Zheng, Cheng Tan, Yufei Huang, Yongjie Xu, and Stan Z Li. Protein language models and structure prediction: Connection and progression. arXiv preprint arXiv:2211.16742, 2022
2022 arXiv
-
[38]
Learning functional properties of proteins with language models
Serbulent Unsal, Heval Atas, Muammer Albayrak, Kemal Turhan, Aybar C Acar, and Tunca Do˘ gan. Learning functional properties of proteins with language models. Nature Machine Intelligence , 4(3):227–245, 2022
2022
-
[39]
Designing proteins with language models
Jeffrey A Ruffolo and Ali Madani. Designing proteins with language models. Nature Biotechnology, 42(2):200–202, 2024
2024
-
[40]
Mass spectrometry: principles and applications
Edmond De Hoffmann and Vincent Stroobant. Mass spectrometry: principles and applications. John Wiley & Sons, 2007
2007
-
[41]
A new generation of crystallographic validation tools for the protein data bank
Randy J Read, Paul D Adams, W Bryan Arendall, Axel T Brunger, Paul Emsley, Robbie P Joosten, Gerard J Kleywegt, Eugene B Krissinel, Thomas Lütteke, Zbyszek Otwinowski, et al. A new generation of crystallographic validation tools for the protein data bank. Structure, 19(10):139...
2011
-
[42]
Outcome of the first electron microscopy validation task force meeting
Richard Henderson, Andrej Sali, Matthew L Baker, Bridget Carragher, Batsal Devkota, Kenneth H Downing, Edward H Egelman, Zukang Feng, Joachim Frank, Nikolaus Grigorieff, et al. Outcome of the first electron microscopy validation task force meeting. Structure, 20(2):205–214, 2012
2012
-
[43]
Deep mutational scanning: a new style of protein science
Douglas M Fowler and Stanley Fields. Deep mutational scanning: a new style of protein science. Nature methods, 11(8):801–807, 2014
2014
-
[44]
Uniref clusters: a comprehensive and scalable alternative for improving sequence similarity searches
Baris E Suzek, Yuqi Wang, Hongzhan Huang, Peter B McGarvey, Cathy H Wu, and UniProt Consortium. Uniref clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics, 31(6):926–932, 2015
2015
-
[45]
Nucleic acids research, 47(D1):D520–D528, 2019
Protein data bank: the single global archive for 3d macromolecular structure data. Nucleic acids research, 47(D1):D520–D528, 2019
2019
-
[46]
Gene ontology: tool for the unification of biology
Michael Ashburner, Catherine A Ball, Judith A Blake, David Botstein, Heather Butler, J Michael Cherry, Allan P Davis, Kara Dolinski, Selina S Dwight, Janan T Eppig, et al. Gene ontology: tool for the unification of biology. Nature genetics, 25(1):25–29, 2000
2000
-
[47]
Central dogma of molecular biology
Francis Crick. Central dogma of molecular biology. Nature, 227(5258):561–563, 1970
1970
-
[48]
Origin and evolution of the genetic code: the universal enigma
Eugene V Koonin and Artem S Novozhilov. Origin and evolution of the genetic code: the universal enigma. IUBMB life, 61(2):99–111, 2009
2009
-
[49]
One thousand families for the molecular biologist
Cyrus Chothia. One thousand families for the molecular biologist. Nature, 357(6379), 1992
1992
-
[50]
The protein-folding problem, 50 years on
Ken A Dill and Justin L MacCallum. The protein-folding problem, 50 years on. science, 338(6110):1042–1046, 2012
2012
-
[51]
Learning from protein structure with geometric vector perceptrons
Bowen Jing, Stephan Eismann, Patricia Suriana, Raphael John Lamarre Townshend, and Ron Dror. Learning from protein structure with geometric vector perceptrons. In International Conference on Learning Representations, 2020
2020
-
[52]
Pifold: Toward effective and efficient protein inverse folding
Zhangyang Gao, Cheng Tan, Pablo Chacón, and Stan Z Li. Pifold: Toward effective and efficient protein inverse folding. arXiv preprint arXiv:2209.12643, 2022
2022 arXiv
-
[53]
Exploring protein fitness landscapes by directed evolution
Philip A Romero and Frances H Arnold. Exploring protein fitness landscapes by directed evolution. Nature reviews Molecular cell biology, 10(12):866–876, 2009
2009
-
[54]
A survey on protein representation learning: Retrospect and prospect
Lirong Wu, Yufei Huang, Haitao Lin, and Stan Z Li. A survey on protein representation learning: Retrospect and prospect. arXiv preprint arXiv:2301.00813, 2022
2022 arXiv
-
[55]
Convolutions are competitive with transformers for protein sequence pretraining
Kevin K Yang, Nicolo Fusi, and Alex X Lu. Convolutions are competitive with transformers for protein sequence pretraining. Cell Systems, 15(3):286–294, 2024
2024
-
[56]
Prottrans: Toward understanding the language of life through self-supervised learning
Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, et al. Prottrans: Toward understanding the language of life through self-supervised learning. IEEE transactions on pattern ...
2021
-
[57]
Protein representation learning by geometric structure pretraining
Zuobai Zhang, Minghao Xu, Arian Jamasb, Vijil Chenthama- rakshan, Aurelie Lozano, Payel Das, and Jian Tang. Protein representation learning by geometric structure pretraining. arXiv preprint arXiv:2203.06125, 2022
2022 arXiv
-
[58]
Lm-gvp: an extensible sequence and structure informed deep learning framework for protein property prediction
Zichen Wang, Steven A Combs, Ryan Brand, Miguel Romero Calvo, Panpan Xu, George Price, Nataliya Golovach, Emmanuel O Salawu, Colby J Wise, Sri Priya Ponnapalli, et al. Lm-gvp: an extensible sequence and structure informed deep learning framework for protein property prediction...
2022
-
[59]
A systematic study of joint representation learning on protein sequences and structures
Z Zhang, C Wang, M Xu, V Chenthamarakshan, AC Lozano, P Das, and J Tang. A systematic study of joint representation learning on protein sequences and structures. Preprint at http://arxiv. org/abs/2303.06275, 2023
2023 arXiv
-
[60]
Colabfold: making protein folding accessible to all
Milot Mirdita, Konstantin Schütze, Yoshitaka Moriwaki, Lim Heo, Sergey Ovchinnikov, and Martin Steinegger. Colabfold: making protein folding accessible to all. Nature methods, 19(6):679–682, 2022
2022
-
[61]
Single-sequence protein structure prediction using supervised transformer protein language models
Wenkai Wang, Zhenling Peng, and Jianyi Yang. Single-sequence protein structure prediction using supervised transformer protein language models. Nature Computational Science , 2(12):804–814, 2022
2022
-
[62]
A method for multiple-sequence-alignment-free protein structure prediction using a protein language model
Xiaomin Fang, Fan Wang, Lihang Liu, Jingzhou He, Dayong Lin, Yingfei Xiang, Kunrui Zhu, Xiaonan Zhang, Hua Wu, Hui Li, et al. A method for multiple-sequence-alignment-free protein structure prediction using a protein language model. Nature Machine Intelligence, 5(10):1087–1096, 2023
2023
-
[63]
Flip: Benchmark tasks in fitness landscape inference for proteins
Christian Dallago, Jody Mou, Kadina E Johnston, Bruce J Wittmann, Nicholas Bhattacharya, Samuel Goldman, Ali Madani, and Kevin K Yang. Flip: Benchmark tasks in fitness landscape inference for proteins. bioRxiv, pages 2021–11, 2021
2021
-
[64]
Exploring machine learning algorithms and protein language models strategies to develop enzyme classi- fication systems
Diego Fernández, Álvaro Olivera-Nappa, Roberto Uribe-Paredes, and David Medina-Ortiz. Exploring machine learning algorithms and protein language models strategies to develop enzyme classi- fication systems. In International Work-Conference on Bioinformatics and Biomedical Engi...
2023
-
[65]
Protein–dna binding sites prediction based on pre-trained protein language model and contrastive learning
Yufan Liu and Boxue Tian. Protein–dna binding sites prediction based on pre-trained protein language model and contrastive learning. Briefings in Bioinformatics, 25(1):bbad488, 2024
2024
-
[66]
Genome-scale annotation of protein binding sites via language model and geometric deep learning
Qianmu Yuan, Chong Tian, and Yuedong Yang. Genome-scale annotation of protein binding sites via language model and geometric deep learning. Elife, 13:RP93695, 2024
2024
-
[67]
Contrastive learning in protein language space predicts interactions between drugs and protein targets
Rohit Singh, Samuel Sledzieski, Bryan Bryson, Lenore Cowen, and Bonnie Berger. Contrastive learning in protein language space predicts interactions between drugs and protein targets. Proceedings of the National Academy of Sciences , 120(24):e2220778120, 2023
2023
-
[68]
Unikp: a unified framework for the prediction of enzyme kinetic parameters
Han Yu, Huaxiang Deng, Jiahui He, Jay D Keasling, and Xiaozhou Luo. Unikp: a unified framework for the prediction of enzyme kinetic parameters. Nature Communications, 14(1):8211, 2023
2023
-
[69]
Protchatgpt: Towards understanding proteins with large language models
Chao Wang, Hehe Fan, Ruijie Quan, and Yi Yang. Protchatgpt: Towards understanding proteins with large language models. arXiv preprint arXiv:2402.09649, 2024
2024 arXiv
-
[70]
Prot2text: Multimodal protein’s function generation with gnns and transformers
Hadi Abdine, Michail Chatzianastasis, Costas Bouyioukos, and Michalis Vazirgiannis. Prot2text: Multimodal protein’s function generation with gnns and transformers. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 10757–10765, 2024
2024
-
[71]
Machine learning for functional protein design
Pascal Notin, Nathan Rollins, Yarin Gal, Chris Sander, and Debora Marks. Machine learning for functional protein design. Nature Biotechnology, 42(2):216–228, 2024
2024
-
[72]
Language models enable zero-shot prediction of the effects of mutations on protein function
Joshua Meier, Roshan Rao, Robert Verkuil, Jason Liu, Tom Sercu, and Alex Rives. Language models enable zero-shot prediction of the effects of mutations on protein function. Advances in neural information processing systems, 34:29287–29303, 2021
2021
-
[73]
De novo protein design—from new structures to programmable functions
Tanja Kortemme. De novo protein design—from new structures to programmable functions. Cell, 187(3):526–544, 2024
2024
-
[74]
De novo design of protein structure and function with rfdiffusion
Joseph L Watson, David Juergens, Nathaniel R Bennett, Brian L Trippe, Jason Yim, Helen E Eisenach, Woody Ahern, Andrew J Borst, Robert J Ragotte, Lukas F Milles, et al. De novo design of protein structure and function with rfdiffusion. Nature, 620(7976):1089–1100, 2023. JOURNA...
2023
-
[75]
Illuminating protein space with a programmable generative model
John B Ingraham, Max Baranov, Zak Costello, Karl W Barber, Wujie Wang, Ahmed Ismail, Vincent Frappier, Dana M Lord, Christopher Ng-Thow-Hing, Erik R Van Vlack, et al. Illuminating protein space with a programmable generative model. Nature, 623(7989):1070–1078, 2023
2023
-
[76]
Learning inverse folding from millions of predicted structures
Chloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin, Brian Hie, Tom Sercu, Adam Lerer, and Alexander Rives. Learning inverse folding from millions of predicted structures. In International conference on machine learning, pages 8946–8970. PMLR, 2022
2022
-
[77]
Robust deep learning–based protein sequence design using proteinmpnn
Justas Dauparas, Ivan Anishchenko, Nathaniel Bennett, Hua Bai, Robert J Ragotte, Lukas F Milles, Basile IM Wicky, Alexis Courbet, Rob J de Haas, Neville Bethel, et al. Robust deep learning–based protein sequence design using proteinmpnn. Science, 378(6615):49– 56, 2022
2022
-
[78]
Protein generation with evolutionary diffusion: sequence is all you need
Sarah Alamdari, Nitya Thakkar, Rianne van den Berg, Alex Xijie Lu, Nicolo Fusi, Ava Pardis Amini, and Kevin K Yang. Protein generation with evolutionary diffusion: sequence is all you need. bioRxiv, pages 2023–09, 2023
2023
-
[79]
Graph machine learning in the era of large language models (llms)
Wenqi Fan, Shijie Wang, Jiani Huang, Zhikai Chen, Yu Song, Wenzhuo Tang, Haitao Mao, Hui Liu, Xiaorui Liu, Dawei Yin, et al. Graph machine learning in the era of large language models (llms). arXiv preprint arXiv:2404.14928, 2024
2024 arXiv
-
[80]
Moleculargpt: Open large language model (llm) for few-shot molecular property prediction
Yuyan Liu, Sirui Ding, Sheng Zhou, Wenqi Fan, and Qiaoyu Tan. Moleculargpt: Open large language model (llm) for few-shot molecular property prediction. arXiv preprint arXiv:2406.12950 , 2024
2024 arXiv
-
[81]
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020
2001 arXiv
-
[82]
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021
2021 arXiv
-
[83]
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettle- moyer. Qlora: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314, 2023
2023 arXiv
-
[84]
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022
2022
-
[85]
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691, 2021
2021 arXiv
-
[86]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pa...
2021
-
[87]
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023
2023 arXiv
-
[88]
Leveraging biomolecule and natural language through multi-modal learning: A survey
Qizhi Pei, Lijun Wu, Kaiyuan Gao, Jinhua Zhu, Yue Wang, Zun Wang, Tao Qin, and Rui Yan. Leveraging biomolecule and natural language through multi-modal learning: A survey. arXiv preprint arXiv:2403.01528, 2024
2024
-
[89]
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730–19742. PMLR, 2023
2023
-
[90]
A survey on rag meeting llms: Towards retrieval-augmented large language models
Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. A survey on rag meeting llms: Towards retrieval-augmented large language models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages...
2024
-
[91]
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C Lawrence Zitnick, Jerry Ma, et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National ...
2021
-
[92]
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019
1907 arXiv
-
[93]
ESM Cambrian: Revealing the mysteries of proteins with unsupervised learning
ESM Team. ESM Cambrian: Revealing the mysteries of proteins with unsupervised learning. https://evolutionaryscale.ai/blog/ esm-cambrian, December 2024
2024
-
[94]
Pre-training co-evolutionary protein representation via a pairwise masked language model
Liang He, Shizhuo Zhang, Lijun Wu, Huanhuan Xia, Fusong Ju, He Zhang, Siyuan Liu, Yingce Xia, Jianwei Zhu, Pan Deng, et al. Pre-training co-evolutionary protein representation via a pairwise masked language model. arXiv preprint arXiv:2110.15527, 2021
-
[95]
Long-context protein language model
Yingheng Wang, Zichen Wang, Gil Sadeh, Luca Zancato, Alessan- dro Achille, George Karypis, and Huzefa Rangwala. Long-context protein language model. bioRxiv, pages 2024–10, 2024
2024
-
[96]
A survey of mamba
Haohao Qu, Liangbo Ning, Rui An, Wenqi Fan, Tyler Derr, Xin Xu, and Qing Li. A survey of mamba. arXiv preprint arXiv:2408.01129, 2024
2024 arXiv
-
[97]
Ssd4rec: a structured state space duality model for efficient sequential recommendation
Haohao Qu, Yifeng Zhang, Liangbo Ning, Wenqi Fan, and Qing Li. Ssd4rec: a structured state space duality model for efficient sequential recommendation. arXiv preprint arXiv:2409.01192, 2024
2024 arXiv
-
[98]
Deciphering the protein landscape with protflash, a lightweight language model
Lei Wang, Hui Zhang, Wei Xu, Zhidong Xue, and Yan Wang. Deciphering the protein landscape with protflash, a lightweight language model. Cell Reports Physical Science, 4(10), 2023
2023
-
[99]
Distilprotbert: a distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts
Yaron Geffen, Yanay Ofran, and Ron Unger. Distilprotbert: a distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts. Bioinformatics, 38(Supplement_2):ii95–ii98, 2022
2022
-
[100]
Mixture of experts enable efficient and effective protein understanding and design
Ning Sun, Shuxian Zou, Tianhua Tao, Sazan Mahbub, Dian Li, Yonghao Zhuang, Hongyi Wang, Xingyi Cheng, Le Song, and Eric P Xing. Mixture of experts enable efficient and effective protein understanding and design. bioRxiv, pages 2024–11, 2024
2024
-
[101]
Toward ai-driven digital organ- ism: Multiscale foundation models for predicting, simulating and programming biology at all levels
Le Song, Eran Segal, and Eric Xing. Toward ai-driven digital organ- ism: Multiscale foundation models for predicting, simulating and programming biology at all levels. arXiv preprint arXiv:2412.06993, 2024
2024 arXiv
-
[102]
Diffusion language models are versatile protein learners
Xinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue, Shujian Huang, and Quanquan Gu. Diffusion language models are versatile protein learners. arXiv preprint arXiv:2402.18567, 2024
2024 arXiv
-
[103]
Generative diffusion models on graphs: methods and applications
Chengyi Liu, Wenqi Fan, Yunqing Liu, Jiatong Li, Hang Li, Hui Liu, Jiliang Tang, and Qing Li. Generative diffusion models on graphs: methods and applications. In Proceedings of the Thirty- Second International Joint Conference on Artificial Intelligence , pages 6702–6711, 2023
2023
-
[104]
Rita: a study on scaling up generative protein sequence models
Daniel Hesslow, Niccoló Zanichelli, Pascal Notin, Iacopo Poli, and Debora Marks. Rita: a study on scaling up generative protein sequence models. arXiv preprint arXiv:2205.05789, 2022
2022 arXiv
-
[105]
Progen2: exploring the boundaries of protein language models
Erik Nijkamp, Jeffrey A Ruffolo, Eli N Weinstein, Nikhil Naik, and Ali Madani. Progen2: exploring the boundaries of protein language models. Cell systems, 14(11):968–978, 2023
2023
-
[106]
Ankh: Optimized protein language model unlocks general-purpose modelling
Ahmed Elnaggar, Hazem Essam, Wafaa Salah-Eldin, Walid Moustafa, Mohamed Elkerdawy, Charlotte Rochereau, and Burkhard Rost. Ankh: Optimized protein language model unlocks general-purpose modelling. arXiv preprint arXiv:2301.06568, 2023
2023 arXiv
-
[107]
Glm: General language model pretraining with autoregressive blank infilling
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. Glm: General language model pretraining with autoregressive blank infilling. arXiv preprint arXiv:2103.10360, 2021
2021 arXiv
-
[108]
Deciphering antibody affinity maturation with language models and weakly supervised learning
Jeffrey A Ruffolo, Jeffrey J Gray, and Jeremias Sulam. Deciphering antibody affinity maturation with language models and weakly supervised learning. arXiv preprint arXiv:2112.07782, 2021
2021 arXiv
-
[109]
Deciphering the language of antibodies using self-supervised learning
Jinwoo Leem, Laura S Mitchell, James HR Farmery, Justin Barton, and Jacob D Galson. Deciphering the language of antibodies using self-supervised learning. Patterns, 3(7), 2022
2022
-
[110]
Ablang: an antibody language model for completing antibody sequences
Tobias H Olsen, Iain H Moal, and Charlotte M Deane. Ablang: an antibody language model for completing antibody sequences. Bioinformatics Advances, 2(1):vbac046, 2022
2022
-
[111]
Pre-training antibody language models for antigen-specific compu- tational antibody design
Kaiyuan Gao, Lijun Wu, Jinhua Zhu, Tianbo Peng, Yingce Xia, Liang He, Shufang Xie, Tao Qin, Haiguang Liu, Kun He, et al. Pre-training antibody language models for antigen-specific compu- tational antibody design. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Di...
2023
-
[112]
Reprogram- ming pretrained language models for antibody sequence infilling
Igor Melnyk, Vijil Chenthamarakshan, Pin-Yu Chen, Payel Das, Amit Dhurandhar, Inkit Padhi, and Devleena Das. Reprogram- ming pretrained language models for antibody sequence infilling. In International Conference on Machine Learning , pages 24398–24419. PMLR, 2023
2023
-
[113]
Iglm: Infilling language modeling for antibody sequence design
Richard W Shuai, Jeffrey A Ruffolo, and Jeffrey J Gray. Iglm: Infilling language modeling for antibody sequence design. Cell Systems, 14(11):979–989, 2023
2023
-
[114]
p-iggen: A paired antibody generative language model
Oliver Marcus Turnbull, Dino Oglic, Rebecca Croasdale-Wood, and Charlotte M Deane. p-iggen: A paired antibody generative language model. bioRxiv, pages 2024–08, 2024. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 29
2024
-
[115]
Generative antibody design for complementary chain pairing sequences through encoder-decoder language model
Simon KS Chu and Kathy Y Wei. Generative antibody design for complementary chain pairing sequences through encoder-decoder language model. arXiv preprint arXiv:2301.02748, 2023
2023 arXiv
-
[116]
Hhblits: lightning-fast iterative protein sequence searching by hmm-hmm alignment
Michael Remmert, Andreas Biegert, Andreas Hauser, and Jo- hannes Söding. Hhblits: lightning-fast iterative protein sequence searching by hmm-hmm alignment. Nature methods, 9(2):173–175, 2012
2012
-
[117]
Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets
Martin Steinegger and Johannes Söding. Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nature biotechnology, 35(11):1026–1028, 2017
2017
-
[118]
Few shot protein generation
Soumya Ram and Tristan Bepler. Few shot protein generation. arXiv preprint arXiv:2204.01168, 2022
2022 arXiv
-
[119]
Enhancing the protein tertiary structure prediction by multiple sequence alignment generation
Le Zhang, Jiayang Chen, Tao Shen, Yu Li, and Siqi Sun. Enhancing the protein tertiary structure prediction by multiple sequence alignment generation. arXiv preprint arXiv:2306.01824, 2023
2023 arXiv
-
[120]
Msagpt: Neural prompting protein structure prediction via msa generative pre-training
Bo Chen, Zhilei Bei, Xingyi Cheng, Pan Li, Jie Tang, and Le Song. Msagpt: Neural prompting protein structure prediction via msa generative pre-training. arXiv preprint arXiv:2406.05347, 2024
2024 arXiv
-
[121]
Poet: A generative model of protein families as sequences-of-sequences
Timothy Truong Jr and Tristan Bepler. Poet: A generative model of protein families as sequences-of-sequences. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[122]
Protmamba: a homology-aware but alignment-free protein state space model
Damiano Sgarbossa, Cyril Malbranke, and Anne-Florence Bitbol. Protmamba: a homology-aware but alignment-free protein state space model. bioRxiv, pages 2024–05, 2024
2024
-
[123]
Codon language embeddings provide strong signals for use in protein engineering
Carlos Outeiral and Charlotte M Deane. Codon language embeddings provide strong signals for use in protein engineering. Nature Machine Intelligence, 6(2):170–179, 2024
2024
-
[124]
cdsbert- extending protein language models with codon awareness
Logan Hallee, Nikolaos Rafailidis, and Jason P Gleghorn. cdsbert- extending protein language models with codon awareness. bioRxiv, 2023
2023
-
[125]
Ptm-mamba: A ptm-aware protein language model with bidirec- tional gated mamba blocks
Zhangzhi Peng, Benjamin Schussheim, and Pranam Chatterjee. Ptm-mamba: A ptm-aware protein language model with bidirec- tional gated mamba blocks. bioRxiv, pages 2024–02, 2024
2024
-
[126]
Mgnify: the microbiome sequence data analysis resource in 2023
Lorna Richardson, Ben Allen, Germana Baldi, Martin Bera- cochea, Maxwell L Bileschi, Tony Burdett, Josephine Burgin, Juan Caballero-Pérez, Guy Cochrane, Lucy J Colwell, et al. Mgnify: the microbiome sequence data analysis resource in 2023. Nucleic Acids Research, 51(D1):D753–D...
2023
-
[127]
The img/m data management and analysis system v
I-Min A Chen, Ken Chu, Krishnaveni Palaniappan, Anna Ratner, Jinghua Huang, Marcel Huntemann, Patrick Hajek, Stephan J Ritter, Cody Webb, Dongying Wu, et al. The img/m data management and analysis system v. 7: content updates and new features. Nucleic acids research, 51(D1):D7...
2023
-
[128]
Protein- level assembly increases protein sequence recovery from metage- nomic samples manyfold
Martin Steinegger, Milot Mirdita, and Johannes Söding. Protein- level assembly increases protein sequence recovery from metage- nomic samples manyfold. Nature methods, 16(7):603–606, 2019
2019
-
[129]
Albert: A lite bert for self-supervised learning of language representations
Z Lan. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942, 2019
1909 arXiv
-
[130]
Electra: Pre-training text encoders as discriminators rather than generators
K Clark. Electra: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555, 2020
2003 arXiv
-
[131]
Mixtral of experts
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024
2024 arXiv
-
[132]
Transformer-xl: Attentive lan- guage models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive lan- guage models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019
1901 arXiv
-
[133]
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang. Xlnet: Generalized autoregressive pretraining for language understanding. arXiv preprint arXiv:1906.08237, 2019
1906 arXiv
-
[134]
Observed antibody space: a resource for data mining next-generation sequencing of antibody repertoires
Aleksandr Kovaltsuk, Jinwoo Leem, Sebastian Kelm, James Snowden, Charlotte M Deane, and Konrad Krawczyk. Observed antibody space: a resource for data mining next-generation sequencing of antibody repertoires. The Journal of Immunology , 201(8):2502–2509, 2018
2018
-
[135]
Sabdab: the structural antibody database
James Dunbar, Konrad Krawczyk, Jinwoo Leem, Terry Baker, Angelika Fuchs, Guy Georges, Jiye Shi, and Charlotte M Deane. Sabdab: the structural antibody database. Nucleic acids research, 42(D1):D1140–D1146, 2014
2014
-
[136]
Pfam: The protein families database in 2021
Jaina Mistry, Sara Chuguransky, Lowri Williams, Matloob Qureshi, Gustavo A Salazar, Erik LL Sonnhammer, Silvio CE Tosatto, Lisanna Paladin, Shriya Raj, Lorna J Richardson, et al. Pfam: The protein families database in 2021. Nucleic acids research , 49(D1):D412–D419, 2021
2021
-
[137]
Uniclust databases of clustered and deeply annotated protein sequences and alignments
Milot Mirdita, Lars Von Den Driesch, Clovis Galiez, Maria J Martin, Johannes Söding, and Martin Steinegger. Uniclust databases of clustered and deeply annotated protein sequences and alignments. Nucleic acids research, 45(D1):D170–D176, 2017
2017
-
[138]
Openproteinset: Training data for structural biology at scale
Gustaf Ahdritz, Nazim Bouatta, Sachin Kadyan, Lukas Jarosch, Dan Berenberg, Ian Fisk, Andrew Watkins, Stephen Ra, Richard Bonneau, and Mohammed AlQuraishi. Openproteinset: Training data for structural biology at scale. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[139]
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023
2023 arXiv
-
[140]
Consensus coding sequence (ccds) database: a standardized set of human and mouse protein-coding regions supported by expert curation
Shashikant Pujar, Nuala A O’Leary, Catherine M Farrell, Jane E Loveland, Jonathan M Mudge, Craig Wallin, Carlos G Girón, Mark Diekhans, If Barnes, Ruth Bennett, et al. Consensus coding sequence (ccds) database: a standardized set of human and mouse protein-coding regions suppo...
2018
-
[141]
Uniprotkb/swiss-prot: the man- ually annotated section of the uniprot knowledgebase
Emmanuel Boutet, Damien Lieberherr, Michael Tognolli, Michel Schneider, and Amos Bairoch. Uniprotkb/swiss-prot: the man- ually annotated section of the uniprot knowledgebase. In Plant bioinformatics: methods and protocols, pages 89–112. Springer, 2007
2007
-
[142]
Petribert: Augmenting bert with tridimensional encoding for inverse protein folding and design
Baldwin Dumortier, Antoine Liutkus, Clément Carré, and Gabriel Krouk. Petribert: Augmenting bert with tridimensional encoding for inverse protein folding and design. BioRxiv, pages 2022–08, 2022
2022
-
[143]
Structure-informed protein language model
Zuobai Zhang, Jiarui Lu, Vijil Chenthamarakshan, Aurélie Lozano, Payel Das, and Jian Tang. Structure-informed protein language model. arXiv preprint arXiv:2402.05856, 2024
2024 arXiv
-
[144]
Biophysics-based protein language models for protein engineering
Sam Gelman, Bryce Johnson, Chase Freschlin, Sameer D’Costa, Anthony Gitter, and Philip A Romero. Biophysics-based protein language models for protein engineering. bioRxiv, pages 2024–03, 2024
2024
-
[145]
The rosetta all-atom energy function for macromolecular modeling and design
Rebecca F Alford, Andrew Leaver-Fay, Jeliazko R Jeliazkov, Matthew J O’Meara, Frank P DiMaio, Hahnbeom Park, Maxim V Shapovalov, P Douglas Renfrew, Vikram K Mulligan, Kalli Kappel, et al. The rosetta all-atom energy function for macromolecular modeling and design. Journal of c...
2017
-
[146]
Transformer protein language models are unsupervised structure learners
Roshan Rao, Joshua Meier, Tom Sercu, Sergey Ovchinnikov, and Alexander Rives. Transformer protein language models are unsupervised structure learners. Biorxiv, pages 2020–12, 2020
2020
-
[147]
Endowing protein language models with structural knowledge
Dexiong Chen, Philip Hartout, Paolo Pellizzoni, Carlos Oliver, and Karsten Borgwardt. Endowing protein language models with structural knowledge. arXiv preprint arXiv:2401.14819, 2024
2024 arXiv
-
[148]
Structure-informed language models are protein designers
Zaixiang Zheng, Yifan Deng, Dongyu Xue, Yi Zhou, Fei Ye, and Quanquan Gu. Structure-informed language models are protein designers. In International Conference on Machine Learning, pages 42317–42338. PMLR, 2023
2023
-
[149]
Adapting protein language models for structure- conditioned design
Jeffrey A Ruffolo, Aadyot Bhatnagar, Joel Beazer, Stephen Nayfach, Jordan Russ, Emily Hill, Riffat Hussain, Joseph Gallagher, and Ali Madani. Adapting protein language models for structure- conditioned design. bioRxiv, pages 2024–08, 2024
2024
-
[150]
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017
2017
-
[151]
Fast and accurate protein structure search with foldseek
Michel Van Kempen, Stephanie S Kim, Charlotte Tumescheit, Milot Mirdita, Jeongjae Lee, Cameron LM Gilchrist, Johannes Söding, and Martin Steinegger. Fast and accurate protein structure search with foldseek. Nature Biotechnology, 42(2):243–246, 2024
2024
-
[152]
Pro- tokens: A machine-learned language for compact and informative encoding of protein 3d structures
Xiaohan Lin, Zhenyu Chen, Yanheng Li, Xingyu Lu, Chuanliu Fan, Ziqiang Cao, Shihao Feng, Yi Qin Gao, and Jun Zhang. Pro- tokens: A machine-learned language for compact and informative encoding of protein 3d structures. bioRxiv, pages 2023–11, 2023
2023
-
[153]
Foldtoken4: Consistent & hierarchical fold language
Zhangyang Gao, Cheng Tan, and Stan Z Li. Foldtoken4: Consistent & hierarchical fold language. bioRxiv, pages 2024–08, 2024
2024
-
[154]
Balancing locality and reconstruction in protein structure tokenizer
Barthelemy Meynard-Piganeau, Jiayou Zhang, James Gong, Xingyi Cheng, Yingtao Luo, Hugo Ly, Le Song, and Eric P Xing. Balancing locality and reconstruction in protein structure tokenizer. bioRxiv, pages 2024–12, 2024
2024
-
[155]
Prostt5: Bilingual language model for protein sequence and structure
Michael Heinzinger, Konstantin Weissenow, Joaquin Gomez Sanchez, Adrian Henkel, Martin Steinegger, and Burkhard Rost. Prostt5: Bilingual language model for protein sequence and structure. bioRxiv, pages 2023–07, 2023
2023
-
[156]
Saprothub: Making protein modeling accessible to all biologists
Jin Su, Zhikai Li, Chenchen Han, Yuyang Zhou, Yan He, Junjie Shan, Xibin Zhou, Xing Chang, Dacheng Ma, OPMC, et al. Saprothub: Making protein modeling accessible to all biologists. bioRxiv, pages 2024–05, 2024
2024
-
[157]
Deprot: A protein language model with quantizied structure and disentangled attention
Mingchen Li, Yang Tan, Bozitao Zhong, Ziyi Zhou, Huiqun Yu, Xinzhu Ma, Wanli Ouyang, Liang Hong, Bingxin Zhou, and Pan Tan. Deprot: A protein language model with quantizied structure and disentangled attention. bioRxiv, pages 2024–04, 2024. JOURNAL OF LATEX CLASS FILES, VOL. 1...
2024
-
[158]
Evaluating protein transfer learning with tape
Roshan Rao, Nicholas Bhattacharya, Neil Thomas, Yan Duan, Peter Chen, John Canny, Pieter Abbeel, and Yun Song. Evaluating protein transfer learning with tape. Advances in neural information processing systems, 32, 2019
2019
-
[159]
Peer: a comprehensive and multi-task benchmark for protein sequence understanding
Minghao Xu, Zuobai Zhang, Jiarui Lu, Zhaocheng Zhu, Yangtian Zhang, Ma Chang, Runcheng Liu, and Jian Tang. Peer: a comprehensive and multi-task benchmark for protein sequence understanding. Advances in Neural Information Processing Systems , 35:35156–35173, 2022
2022
-
[160]
Proteingym: Large-scale benchmarks for protein fitness prediction and design
Pascal Notin, Aaron Kollasch, Daniel Ritter, Lood Van Niekerk, Steffanie Paul, Han Spinner, Nathan Rollins, Ada Shaw, Rose Orenbuch, Ruben Weitzman, et al. Proteingym: Large-scale benchmarks for protein fitness prediction and design. Advances in Neural Information Processing S...
2024
-
[161]
Dplm-2: A multimodal diffusion protein language model
Xinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue, Shujian Huang, and Quanquan Gu. Dplm-2: A multimodal diffusion protein language model. arXiv preprint arXiv:2410.13782, 2024
2024 arXiv
-
[162]
Proteinbert: a universal deep-learning model of protein sequence and function
Nadav Brandes, Dan Ofer, Yam Peleg, Nadav Rappoport, and Michal Linial. Proteinbert: a universal deep-learning model of protein sequence and function. Bioinformatics, 38(8):2102–2110, 2022
2022
-
[163]
Multi-level protein structure pre-training via prompt learning
Zeyuan Wang, Qiang Zhang, HU Shuang-Wei, Haoran Yu, Xurui Jin, Zhichen Gong, and Huajun Chen. Multi-level protein structure pre-training via prompt learning. In The Eleventh International Conference on Learning Representations, 2022
2022
-
[164]
Zymctrl: a conditional language model for the controllable generation of artificial enzymes
Geraldene Munsamy, Sebastian Lindner, Philipp Lorenz, and Noelia Ferruz. Zymctrl: a conditional language model for the controllable generation of artificial enzymes. In NeurIPS Machine Learning in Structural Biology Workshop, 2022
2022
-
[165]
Regression transformer enables concurrent sequence regression and generation for molecular language modelling
Jannis Born and Matteo Manica. Regression transformer enables concurrent sequence regression and generation for molecular language modelling. Nature Machine Intelligence , 5(4):432–444, 2023
2023
-
[166]
A text-guided protein design framework
Shengchao Liu, Yutao Zhu, Jiarui Lu, Zhao Xu, Weili Nie, Anthony Gitter, Chaowei Xiao, Jian Tang, Hongyu Guo, and Anima Anandkumar. A text-guided protein design framework. arXiv preprint arXiv:2302.04611, 2023
2023 arXiv
-
[167]
Protst: Multi-modality learning of protein sequences and biomedical texts
Minghao Xu, Xinyu Yuan, Santiago Miret, and Jian Tang. Protst: Multi-modality learning of protein sequences and biomedical texts. In International Conference on Machine Learning , pages 38749–38767. PMLR, 2023
2023
-
[168]
Multi-modal clip-informed protein editing
Mingze Yin, Hanjing Zhou, Yiheng Zhu, Miao Lin, Yixuan Wu, Jialu Wu, Hongxia Xu, Chang-Yu Hsieh, Tingjun Hou, Jintai Chen, et al. Multi-modal clip-informed protein editing. bioRxiv, pages 2024–07, 2024
2024
-
[169]
Proteinclip: en- hancing protein language models with natural language
Kevin E Wu, Howard Chang, and James Zou. Proteinclip: en- hancing protein language models with natural language. bioRxiv, pages 2024–05, 2024
2024
-
[170]
Protrek: Navigating the protein universe through tri-modal contrastive learning
Jin Su, Xibin Zhou, Xuting Zhang, and Fajie Yuan. Protrek: Navigating the protein universe through tri-modal contrastive learning. bioRxiv, pages 2024–05, 2024
2024
-
[171]
Ontoprotein: Protein pretraining with gene ontology embedding
Ningyu Zhang, Zhen Bi, Xiaozhuan Liang, Siyuan Cheng, Haosen Hong, Shumin Deng, Jiazhang Lian, Qiang Zhang, and Huajun Chen. Ontoprotein: Protein pretraining with gene ontology embedding. arXiv preprint arXiv:2201.11147, 2022
2022 arXiv
-
[172]
Protein representation learning via knowledge enhanced primary structure reasoning
Hong-Yu Zhou, Yunxiang Fu, Zhicheng Zhang, Bian Cheng, and Yizhou Yu. Protein representation learning via knowledge enhanced primary structure reasoning. InThe Eleventh International Conference on Learning Representations, 2022
2022
-
[173]
Benchmarking text-integrated protein language model embeddings and embed- ding fusion on diverse downstream tasks
Young Su Ko, Jonathan Parkinson, and Wei Wang. Benchmarking text-integrated protein language model embeddings and embed- ding fusion on diverse downstream tasks. bioRxiv, pages 2024–08, 2024
2024
-
[174]
Ssemb: A joint embedding of protein sequence and structure enables robust variant effect predictions
Lasse M Blaabjerg, Nicolas Jonsson, Wouter Boomsma, Amelie Stein, and Kresten Lindorff-Larsen. Ssemb: A joint embedding of protein sequence and structure enables robust variant effect predictions. Nature Communications, 15(1):9646, 2024
2024
-
[175]
Pre- training sequence, structure, and surface features for comprehen- sive protein representation learning
Youhan Lee, Hasun Yu, Jaemyung Lee, and Jaehoon Kim. Pre- training sequence, structure, and surface features for comprehen- sive protein representation learning. In The Twelfth International Conference on Learning Representations, 2023
2023
-
[176]
Contrasting sequence with structure: Pre-training graph representations with plms
Louis Robinson, Timothy Atkinson, Liviu Copoiu, Patrick Bordes, Thomas Pierrot, and Thomas D Barrett. Contrasting sequence with structure: Pre-training graph representations with plms. bioRxiv, pages 2023–12, 2023
2023
-
[177]
A multimodal protein representation framework for quantifying transferability across biochemical downstream tasks
Fan Hu, Yishen Hu, Weihong Zhang, Huazhen Huang, Yi Pan, and Peng Yin. A multimodal protein representation framework for quantifying transferability across biochemical downstream tasks. Advanced Science, 10(22):2301223, 2023
2023
-
[178]
Learning sequence, structure, and function representations of proteins with language models
Tymor Hamamsy, Meet Barot, James T Morton, Martin Steinegger, Richard Bonneau, and Kyunghyun Cho. Learning sequence, structure, and function representations of proteins with language models. bioRxiv, 2023
2023
-
[179]
Alphafold protein structure database: massively expanding the structural coverage of protein-sequence space with high-accuracy models
Mihaly Varadi, Stephen Anyango, Mandar Deshpande, Sreenath Nair, Cindy Natassia, Galabina Yordanova, David Yuan, Oana Stroe, Gemma Wood, Agata Laydon, et al. Alphafold protein structure database: massively expanding the structural coverage of protein-sequence space with high-a...
2022
-
[180]
Deepsf: deep convolutional neural network for mapping protein sequences to folds
Jie Hou, Badri Adhikari, and Jianlin Cheng. Deepsf: deep convolutional neural network for mapping protein sequences to folds. Bioinformatics, 34(8):1295–1303, 2018
2018
-
[181]
Cath–a hierarchic classification of protein domain structures
Christine A Orengo, Alex D Michie, Susan Jones, David T Jones, Mark B Swindells, and Janet M Thornton. Cath–a hierarchic classification of protein domain structures. Structure, 5(8):1093– 1109, 1997
1997
-
[182]
Interpro in
Typhaine Paysan-Lafosse, Matthias Blum, Sara Chuguransky, Tiago Grego, Beatriz Lázaro Pinto, Gustavo A Salazar, Maxwell L Bileschi, Peer Bork, Alan Bridge, Lucy Colwell, et al. Interpro in
-
[183]
Interproscan 5: genome-scale protein function classification
Philip Jones, David Binns, Hsin-Yu Chang, Matthew Fraser, Weizhong Li, Craig McAnulla, Hamish McWilliam, John Maslen, Alex Mitchell, Gift Nuka, et al. Interproscan 5: genome-scale protein function classification. Bioinformatics, 30(9):1236–1240, 2014
2014
-
[184]
String v11: protein–protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets
Damian Szklarczyk, Annika L Gable, David Lyon, Alexander Junge, Stefan Wyder, Jaime Huerta-Cepas, Milan Simonovic, Nadezhda T Doncheva, John H Morris, Peer Bork, et al. String v11: protein–protein association networks with increased coverage, supporting functional discovery in...
2019
-
[185]
The ncbi taxonomy database
Scott Federhen. The ncbi taxonomy database. Nucleic acids research, 40(D1):D136–D143, 2012
2012
-
[186]
Scibert: A pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. Scibert: A pretrained language model for scientific text. arXiv preprint arXiv:1903.10676, 2019
1903 arXiv
-
[187]
Domain-specific language model pretraining for biomedical natural language processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH), 3(1):1–23, 2021
2021
-
[188]
Instructprotein: Aligning human and protein language via knowledge instruction
Zeyuan Wang, Qiang Zhang, Keyan Ding, Ming Qin, Xiang Zhuang, Xiaotong Li, and Huajun Chen. Instructprotein: Aligning human and protein language via knowledge instruction. arXiv preprint arXiv:2310.03269, 2023
2023 arXiv
-
[189]
Protllm: An interleaved protein-language llm with protein-as-word pre- training
Le Zhuo, Zewen Chi, Minghao Xu, Heyan Huang, Heqi Zheng, Conghui He, Xian-Ling Mao, and Wentao Zhang. Protllm: An interleaved protein-language llm with protein-as-word pre- training. arXiv preprint arXiv:2403.07920, 2024
2024 arXiv
-
[190]
Decoding the molecular language of proteins with evolla
Xibin Zhou, Chenchen Han, Yingqi Zhang, Jin Su, Kai Zhuang, Shiyu Jiang, Zichen Yuan, Wei Zheng, Fengyuan Dai, Yuyang Zhou, et al. Decoding the molecular language of proteins with evolla. bioRxiv, pages 2025–01, 2025
2025
-
[191]
Druggpt: A gpt-based strategy for designing potential ligands targeting specific proteins
Yuesen Li, Chengyi Gao, Xin Song, Xiangyu Wang, Yungang Xu, and Suxia Han. Druggpt: A gpt-based strategy for designing potential ligands targeting specific proteins. bioRxiv, pages 2023– 06, 2023
2023
-
[192]
Multi- scale protein language model for unified molecular modeling
Kangjie Zheng, Siyu Long, Tianyu Lu, Junwei Yang, Xinyu Dai, Ming Zhang, Zaiqing Nie, Wei-Ying Ma, and Hao Zhou. Multi- scale protein language model for unified molecular modeling. bioRxiv, pages 2024–03, 2024
2024
-
[193]
Pubmed: the bibliographic database
Kathi Canese and Sarah Weis. Pubmed: the bibliographic database. The NCBI handbook, 2(1), 2013
2013
-
[194]
Mol- instructions: A large-scale biomolecular instruction dataset for large language models
Yin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu, Rui Huang, Zhuo Chen, Xiaohui Fan, and Huajun Chen. Mol- instructions: A large-scale biomolecular instruction dataset for large language models. arXiv preprint arXiv:2306.08018, 2023
2023 arXiv
-
[195]
The llama 3 herd of models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024
2024 arXiv
-
[196]
Zinc20—a free ultralarge- scale chemical database for ligand discovery
John J Irwin, Khanh G Tang, Jennifer Young, Chinzorig Dandarchu- luun, Benjamin R Wong, Munkhzul Khurelbaatar, Yurii S Moroz, John Mayfield, and Roger A Sayle. Zinc20—a free ultralarge- scale chemical database for ligand discovery. Journal of chemical information and modeling,...
2020
-
[197]
Uni- mol: A universal 3d molecular representation learning framework
Gengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng, Hongteng Xu, Zhewei Wei, Linfeng Zhang, and Guolin Ke. Uni- mol: A universal 3d molecular representation learning framework. 2023
2023
-
[198]
Biomedgpt: Open multimodal generative pre-trained transformer for biomedicine
Yizhen Luo, Jiahuan Zhang, Siqi Fan, Kai Yang, Yushuai Wu, Mu Qiao, and Zaiqing Nie. Biomedgpt: Open multimodal generative pre-trained transformer for biomedicine. arXiv preprint arXiv:2308.09442, 2023
2023 arXiv
-
[199]
Pubchem 2023 update
Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, et al. Pubchem 2023 update. Nucleic acids research, 51(D1):D1373–D1380, 2023
2023
-
[200]
Pre-training molecular graph representation with 3d geometry
Shengchao Liu, Hanchen Wang, Weiyang Liu, Joan Lasenby, Hongyu Guo, and Jian Tang. Pre-training molecular graph representation with 3d geometry. arXiv preprint arXiv:2110.07728, 2021
2021 arXiv
-
[201]
Biot5+: Towards generalized biological understanding with iupac integration and multi-task tuning
Qizhi Pei, Lijun Wu, Kaiyuan Gao, Xiaozhuan Liang, Yin Fang, Jinhua Zhu, Shufang Xie, Tao Qin, and Rui Yan. Biot5+: Towards generalized biological understanding with iupac integration and multi-task tuning. arXiv preprint arXiv:2402.17810, 2024
2024 arXiv
-
[202]
Instructbiomol: Advancing biomolecule understand- ing and design following human instructions
Xiang Zhuang, Keyan Ding, Tianwen Lyu, Yinuo Jiang, Xiaotong Li, Zhuoyi Xiang, Zeyuan Wang, Ming Qin, Kehua Feng, Jike Wang, et al. Instructbiomol: Advancing biomolecule understand- ing and design following human instructions. arXiv preprint arXiv:2410.07919, 2024
-
[203]
Translation between molecules and natural language
Carl Edwards, Tuan Lai, Kevin Ros, Garrett Honke, Kyunghyun Cho, and Heng Ji. Translation between molecules and natural language. arXiv preprint arXiv:2204.11817, 2022
2022 arXiv
-
[204]
Bindingdb in 2015: a public database for medicinal chemistry, computational chemistry and systems pharmacology
Michael K Gilson, Tiqing Liu, Michael Baitaluk, George Nicola, Linda Hwang, and Jenny Chong. Bindingdb in 2015: a public database for medicinal chemistry, computational chemistry and systems pharmacology. Nucleic acids research, 44(D1):D1045–D1053, 2016
2015
-
[205]
Rhea, the reaction knowledgebase in 2022
Parit Bansal, Anne Morgat, Kristian B Axelsen, Venkatesh Muthukrishnan, Elisabeth Coudert, Lucila Aimo, Nevila Hyka- Nouspikel, Elisabeth Gasteiger, Arnaud Kerhornou, Teresa Batista Neto, et al. Rhea, the reaction knowledgebase in 2022. Nucleic acids research, 50(D1):D693–D700, 2022
2022
-
[206]
Multilingual translation for zero-shot biomedical classification using biotranslator
Hanwen Xu, Addie Woicik, Hoifung Poon, Russ B Altman, and Sheng Wang. Multilingual translation for zero-shot biomedical classification using biotranslator. Nature Communications, 14(1):738, 2023
2023
-
[207]
The human phenotype ontology in 2021
Sebastian Köhler, Michael Gargano, Nicolas Matentzoglu, Leigh C Carmody, David Lewis-Smith, Nicole A Vasilevsky, Daniel Danis, Ganna Balagura, Gareth Baynam, Amy M Brower, et al. The human phenotype ontology in 2021. Nucleic acids research , 49(D1):D1207–D1217, 2021
2021
-
[208]
Galactica: A large language model for science
Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. Galactica: A large language model for science. arXiv preprint arXiv:2211.09085, 2022
2022 arXiv
-
[209]
Bio- bridge: Bridging biomedical foundation models via knowledge graph
Zifeng Wang, Zichen Wang, Balasubramaniam Srinivasan, Vas- silis N Ioannidis, Huzefa Rangwala, and Rishita Anubhai. Bio- bridge: Bridging biomedical foundation models via knowledge graph. arXiv preprint arXiv:2310.03320, 2023
-
[210]
Recommendations of the wwpdb nmr validation task force
Gaetano T Montelione, Michael Nilges, Ad Bax, Peter Güntert, Torsten Herrmann, Jane S Richardson, Charles D Schwieters, Wim F Vranken, Geerten W Vuister, David S Wishart, et al. Recommendations of the wwpdb nmr validation task force. Structure, 21(9):1563–1570, 2013
2013
-
[211]
Protein structure prediction beyond alphafold
Guo-Wei Wei. Protein structure prediction beyond alphafold. Nature Machine Intelligence, 1(8):336–337, 2019
2019
-
[212]
High-resolution de novo structure prediction from primary sequence
Ruidong Wu, Fan Ding, Rui Wang, Rui Shen, Xiwen Zhang, Shitong Luo, Chenpeng Su, Zuofan Wu, Qi Xie, Bonnie Berger, et al. High-resolution de novo structure prediction from primary sequence. BioRxiv, pages 2022–07, 2022
2022
-
[213]
Single- sequence protein structure prediction using a language model and deep learning
Ratul Chowdhury, Nazim Bouatta, Surojit Biswas, Christina Floristean, Anant Kharkar, Koushik Roy, Charlotte Rochereau, Gustaf Ahdritz, Joanna Zhang, George M Church, et al. Single- sequence protein structure prediction using a language model and deep learning. Nature Biotechno...
2022
-
[214]
Fast, accurate antibody structure prediction from deep learn- ing on massive set of natural antibodies
Jeffrey A Ruffolo, Lee-Shin Chu, Sai Pooja Mahajan, and Jeffrey J Gray. Fast, accurate antibody structure prediction from deep learn- ing on massive set of natural antibodies. Nature communications, 14(1):2389, 2023
2023
-
[215]
Accurate structure prediction of biomolecular interactions with alphafold 3
Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans, Tim Green, Alexander Pritzel, Olaf Ronneberger, Lindsay Willmore, Andrew J Ballard, Joshua Bambrick, et al. Accurate structure prediction of biomolecular interactions with alphafold 3. Nature, pages 1–3, 2024
2024
-
[216]
Nucleic acids research, 51(D1):D523–D531, 2023
Uniprot: the universal protein knowledgebase in 2023. Nucleic acids research, 51(D1):D523–D531, 2023
2023
-
[217]
Fast and accurate protein intrinsic disorder prediction by using a pretrained language model
Yidong Song, Qianmu Yuan, Sheng Chen, Ken Chen, Yaoqi Zhou, and Yuedong Yang. Fast and accurate protein intrinsic disorder prediction by using a pretrained language model. Briefings in bioinformatics, 24(4):bbad173, 2023
2023
-
[218]
Protein stability prediction by fine-tuning a protein language model on a mega-scale dataset
Kit Sang Chu and Justin B Siegel. Protein stability prediction by fine-tuning a protein language model on a mega-scale dataset. bioRxiv, pages 2023–11, 2023
2023
-
[219]
Vish-pred: an ensemble of fine- tuned esm models for protein toxicity prediction
Raghvendra Mall, Ankita Singh, Chirag N Patel, Gregory Guir- imand, and Filippo Castiglione. Vish-pred: an ensemble of fine- tuned esm models for protein toxicity prediction. Briefings in Bioinformatics, 25(4), 2024
2024
-
[220]
Com- prehensive prediction and analysis of human protein essentiality based on a pretrained large language model
Boming Kang, Rui Fan, Chunmei Cui, and Qinghua Cui. Com- prehensive prediction and analysis of human protein essentiality based on a pretrained large language model. Nature Computational Science, pages 1–11, 2024
2024
-
[221]
Peptidebert: A language model based on transformers for peptide property prediction
Chakradhar Guntuboina, Adrita Das, Parisa Mollaei, Seongwon Kim, and Amir Barati Farimani. Peptidebert: A language model based on transformers for peptide property prediction. The Journal of Physical Chemistry Letters, 14(46):10427–10434, 2023
2023
-
[222]
Lassoesm: A tailored language model for enhanced lasso peptide property prediction
Xuenan Mi, Susanna E Barrett, Douglas A Mitchell, and Diwakar Shukla. Lassoesm: A tailored language model for enhanced lasso peptide property prediction. bioRxiv, pages 2024–10, 2024
2024
-
[223]
Approaching optimal ph enzyme prediction with large language models
Mark Zaretckii, Pavel Buslaev, Igor Kozlovskii, Alexander Moro- zov, and Petr Popov. Approaching optimal ph enzyme prediction with large language models. ACS Synthetic Biology , 13(9):3013– 3021, 2024
2024
-
[224]
Deep learning prediction of enzyme optimum ph
Japheth E Gado, Matthew Knotts, Ada Y Shaw, Debora Marks, Nicholas P Gauthier, Chris Sander, and Gregg T Beckham. Deep learning prediction of enzyme optimum ph. bioRxiv, pages 2023– 06, 2023
2023
-
[225]
Genome-wide prediction of disease variant effects with a deep protein language model
Nadav Brandes, Grant Goldman, Charlotte H Wang, Chun Jimmie Ye, and Vasilis Ntranos. Genome-wide prediction of disease variant effects with a deep protein language model. Nature Genetics, 55(9):1512–1522, 2023
2023
-
[226]
Enhancing efficiency of protein language models with minimal wet-lab data through few-shot learning
Ziyi Zhou, Liang Zhang, Yuanxi Yu, Banghao Wu, Mingchen Li, Liang Hong, and Pan Tan. Enhancing efficiency of protein language models with minimal wet-lab data through few-shot learning. Nature Communications, 15(1):5566, 2024
2024
-
[227]
Deplm: Denoising protein language models for property optimization
Zeyuan Wang, Keyan Ding, Ming Qin, Xiaotong Li, Xiang Zhuang, Yu Zhao, Jianhua Yao, Qiang Zhang, and Huajun Chen. Deplm: Denoising protein language models for property optimization. In The Thirty-eighth Annual Conference on Neural Information Processing Systems
-
[228]
Large language models improve annotation of prokaryotic viral proteins
Zachary N Flamholz, Steven J Biller, and Libusha Kelly. Large language models improve annotation of prokaryotic viral proteins. Nature Microbiology, pages 1–13, 2024
2024
-
[229]
Accurately predicting enzyme functions through geometric graph learning on esmfold-predicted structures
Yidong Song, Qianmu Yuan, Sheng Chen, Yuansong Zeng, Huiy- ing Zhao, and Yuedong Yang. Accurately predicting enzyme functions through geometric graph learning on esmfold-predicted structures. Nature Communications, 15(1):8180, 2024
2024
-
[230]
Netgo 3.0: protein language model improves large-scale functional annotations
Shaojun Wang, Ronghui You, Yunjia Liu, Yi Xiong, and Shanfeng Zhu. Netgo 3.0: protein language model improves large-scale functional annotations. Genomics, Proteomics & Bioinformatics , 21(2):349–358, 2023
2023
-
[231]
Protein function prediction as approximate semantic entailment
Maxat Kulmanov, Francisco J Guzmán-Vega, Paula Duek Roggli, Lydie Lane, Stefan T Arold, and Robert Hoehndorf. Protein function prediction as approximate semantic entailment. Nature Machine Intelligence, pages 1–9, 2024
2024
-
[232]
Fast and accurate protein function prediction from sequence through pretrained language model and homology- based label diffusion
Qianmu Yuan, Junjie Xie, Jiancong Xie, Huiying Zhao, and Yuedong Yang. Fast and accurate protein function prediction from sequence through pretrained language model and homology- based label diffusion. Briefings in bioinformatics , 24(3):bbad117, 2023
2023
-
[233]
Identifying b- cell epitopes using alphafold2 predicted structures and pretrained language model
Yuansong Zeng, Zhuoyi Wei, Qianmu Yuan, Sheng Chen, Weijiang Yu, Yutong Lu, Jianzhao Gao, and Yuedong Yang. Identifying b- cell epitopes using alphafold2 predicted structures and pretrained language model. Bioinformatics, 39(4):btad187, 2023
2023
-
[234]
Signalp 6.0 predicts all five types of signal peptides using protein language models
Felix Teufel, José Juan Almagro Armenteros, Alexander Rosenberg Johansen, Magnús Halldór Gíslason, Silas Irby Pihl, Konstanti- nos D Tsirigos, Ole Winther, Søren Brunak, Gunnar von Heijne, and Henrik Nielsen. Signalp 6.0 predicts all five types of signal peptides using protein...
2022
-
[235]
Unbiased organism-agnostic and highly sensitive signal peptide predictor with deep protein language model.Nature Computational Science, 4(1):29–42, 2024
Junbo Shen, Qinze Yu, Shenyang Chen, Qingxiong Tan, Jingchen Li, and Yu Li. Unbiased organism-agnostic and highly sensitive signal peptide predictor with deep protein language model.Nature Computational Science, 4(1):29–42, 2024
2024
-
[236]
Protein language models are performant in structure-free virtual screening
Hilbert Yuen In Lam, Jia Sheng Guan, Xing Er Ong, Robbe Pincket, and Yuguang Mu. Protein language models are performant in structure-free virtual screening. Briefings in Bioinformatics , 25(6):bbae480, 2024
2024
-
[237]
Smiles transformer: Pre-trained molecular fingerprint for low data drug discovery
Shion Honda, Shoi Shi, and Hiroki R Ueda. Smiles transformer: Pre-trained molecular fingerprint for low data drug discovery. arXiv preprint arXiv:1911.04738, 2019
1911 arXiv
-
[238]
Learning binding affinities via fine-tuning of protein and ligand language models
Rohan Gorantla, Aryo Pradipta Gema, Ian Xi Yang, Álvaro Serrano-Morrás, Benjamin Suutari, Jordi Juárez Jiménez, and Antonia SJS Mey. Learning binding affinities via fine-tuning of protein and ligand language models. bioRxiv, pages 2024–11, 2024
2024
-
[239]
Chemberta-2: Towards chemical founda- tion models
Walid Ahmad, Elana Simon, Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. Chemberta-2: Towards chemical founda- tion models. arXiv preprint arXiv:2209.01712, 2022
2022 arXiv
-
[240]
Reactzyme: A benchmark for enzyme-reaction prediction
Chenqing Hua, Bozitao Zhong, Sitao Luan, Liang Hong, Guy Wolf, Doina Precup, and Shuangjia Zheng. Reactzyme: A benchmark for enzyme-reaction prediction. arXiv preprint arXiv:2408.13659, 2024
2024 arXiv
-
[241]
Multi-modal deep learning enables efficient and accurate annotation of enzymatic active sites
Xiaorui Wang, Xiaodan Yin, Dejun Jiang, Huifeng Zhao, Zhenxing Wu, Odin Zhang, Jike Wang, Yuquan Li, Yafeng Deng, Huanxiang Liu, et al. Multi-modal deep learning enables efficient and accurate annotation of enzymatic active sites. Nature Communications , 15(1):7348, 2024
2024
-
[242]
Enhancing protein language model with structure-based encoder and pre-training
Zuobai Zhang, Minghao Xu, Aurelie Lozano, Vijil Chenthama- rakshan, Payel Das, and Jian Tang. Enhancing protein language model with structure-based encoder and pre-training. In ICLR 2023-Machine Learning for Drug Discovery workshop , 2023
2023
-
[243]
Salt&peppr is an interface-predicting language model for designing peptide-guided protein degraders
Garyk Brixi, Tianzheng Ye, Lauren Hong, Tian Wang, Connor Mon- ticello, Natalia Lopez-Barbosa, Sophia Vincoff, Vivian Yudistyra, Lin Zhao, Elena Haarer, et al. Salt&peppr is an interface-predicting language model for designing peptide-guided protein degraders. Communications B...
2023
-
[244]
Democratizing protein language models with parameter-efficient fine-tuning
Samuel Sledzieski, Meghana Kshirsagar, Minkyung Baek, Rahul Dodhia, Juan Lavista Ferres, and Bonnie Berger. Democratizing protein language models with parameter-efficient fine-tuning. Proceedings of the National Academy of Sciences , 121(26):e2405840121, 2024
2024
-
[245]
Prollm: Protein chain-of-thoughts enhanced llm for protein- protein interaction prediction
Mingyu Jin, Xue Haochen, Zhenting Wang, Boming Kang, Ru- osong Ye, Kaixiong Zhou, Mengnan Du, and Yongfeng Zhang. Prollm: Protein chain-of-thoughts enhanced llm for protein- protein interaction prediction. bioRxiv, pages 2024–04, 2024
2024
-
[246]
Prot2token: A multi-task framework for protein language processing using autoregressive language modeling
Mahdi Pourmirzaei, Farzaneh Esmaili, Mohammadreza Pour- mirzaei, Duolin Wang, and Dong Xu. Prot2token: A multi-task framework for protein language processing using autoregressive language modeling. bioRxiv, pages 2024–05, 2024
2024
-
[247]
Bartsmiles: Generative masked language models for molecular representa- tions
Gayane Chilingaryan, Hovhannes Tamoyan, Ani Tevosyan, Nelly Babayan, Lusine Khondkaryan, Karen Hambardzumyan, Zaven Navoyan, Hrant Khachatrian, and Armen Aghajanyan. Bartsmiles: Generative masked language models for molecular representa- tions. arXiv preprint arXiv:2211.16349, 2022
2022 arXiv
-
[248]
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023) , ...
2023
-
[249]
Proteingpt: Multimodal llm for protein property prediction and structure understanding
Yijia Xiao, Edward Sun, Yiqiao Jin, Qifan Wang, and Wei Wang. Proteingpt: Multimodal llm for protein property prediction and structure understanding. arXiv preprint arXiv:2408.11363, 2024
2024 arXiv
-
[250]
Fapm: functional annotation of proteins using multimodal models beyond structural modeling
Wenkai Xiang, Zhaoping Xiong, Huan Chen, Jiacheng Xiong, Wei Zhang, Zunyun Fu, Mingyue Zheng, Bing Liu, and Qian Shi. Fapm: functional annotation of proteins using multimodal models beyond structural modeling. Bioinformatics, 40(12):btae680, 2024
2024
-
[251]
Mistral 7b
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023
-
[252]
Prott3: Protein-to-text generation for text-based protein understanding
Zhiyuan Liu, An Zhang, Hao Fei, Enzhi Zhang, Xiang Wang, Kenji Kawaguchi, and Tat-Seng Chua. Prott3: Protein-to-text generation for text-based protein understanding. arXiv preprint arXiv:2405.12564, 2024
2024 arXiv
-
[253]
Protein language models are biased by unequal sequence sampling across the tree of life
Frances Ding and Jacob Noah Steinhardt. Protein language models are biased by unequal sequence sampling across the tree of life. bioRxiv, pages 2024–03, 2024
2024
-
[254]
Protein language model fitness is a matter of preference
Cade Gordon, Amy X Lu, and Pieter Abbeel. Protein language model fitness is a matter of preference. bioRxiv, pages 2024–10, 2024
2024
-
[255]
Netgo: improving large-scale protein function prediction with massive network information
Ronghui You, Shuwei Yao, Yi Xiong, Xiaodi Huang, Fengzhu Sun, Hiroshi Mamitsuka, and Shanfeng Zhu. Netgo: improving large-scale protein function prediction with massive network information. Nucleic acids research, 47(W1):W379–W387, 2019
2019
-
[256]
Netgo 2.0: improving large-scale protein function prediction with massive sequence, text, domain, family and network information
Shuwei Yao, Ronghui You, Shaojun Wang, Yi Xiong, Xiaodi Huang, and Shanfeng Zhu. Netgo 2.0: improving large-scale protein function prediction with massive sequence, text, domain, family and network information. Nucleic acids research , 49(W1):W469– W475, 2021
2021
-
[257]
Directed evolution: bringing new chemistry to life
Frances H Arnold. Directed evolution: bringing new chemistry to life. Angewandte Chemie (International Ed. in English) , 57(16):4143, 2018
2018
-
[258]
Methods for the directed evolution of proteins
Michael S Packer and David R Liu. Methods for the directed evolution of proteins. Nature Reviews Genetics, 16(7):379–394, 2015
2015
-
[259]
Efficient evolution of human antibodies from general protein language models
Brian L Hie, Varun R Shanker, Duo Xu, Theodora UJ Bruun, Payton A Weidenbacher, Shaogeng Tang, Wesley Wu, John E Pak, and Peter S Kim. Efficient evolution of human antibodies from general protein language models. Nature Biotechnology, 42(2):275– 283, 2024
2024
-
[260]
Generating novel protein sequences using gibbs sampling of masked language models
Sean R Johnson, Sarah Monaco, Kenneth Massie, and Zaid Syed. Generating novel protein sequences using gibbs sampling of masked language models. bioRxiv, pages 2021–01, 2021
2021
-
[261]
Generative power of a protein language model trained on multiple sequence alignments
Damiano Sgarbossa, Umberto Lupo, and Anne-Florence Bitbol. Generative power of a protein language model trained on multiple sequence alignments. Elife, 12:e79854, 2023
2023
-
[262]
Protein design by directed evolution guided by large language models
Thanh VT Tran and Truong Son Hy. Protein design by directed evolution guided by large language models. bioRxiv, pages 2023– 11, 2023
2023
-
[263]
Integrating genetic algorithms and language models for enhanced enzyme design
Yves Gaetan Nana Teukam, Federico Zipoli, Teodoro Laino, Emanuele Criscuolo, Francesca Grisoni, and Matteo Manica. Integrating genetic algorithms and language models for enhanced enzyme design. 2024
2024
-
[264]
Evoopt: an msa-guided, fully unsupervised sequence optimization pipeline for protein design
Hideki Yamaguchi and Yutaka Saito. Evoopt: an msa-guided, fully unsupervised sequence optimization pipeline for protein design. In Machine Learning for Structural Biology Workshop, NeurIPS , 2022
2022
-
[265]
A general temperature-guided language model to design proteins of enhanced stability and activity
Fan Jiang, Mingchen Li, Jiajun Dong, Yuanxi Yu, Xinyu Sun, Banghao Wu, Jin Huang, Liqi Kang, Yufeng Pei, Liang Zhang, et al. A general temperature-guided language model to design proteins of enhanced stability and activity. Science Advances , 10(48):eadr2641, 2024
2024
-
[266]
Sam- pling protein language models for functional protein design
Jeremie Theddy Darmawan, Yarin Gal, and Pascal Notin. Sam- pling protein language models for functional protein design. In NeurIPS 2023 Generative AI and Biology (GenBio) Workshop , 2023
2023
-
[267]
Rapid in silico directed evolution by a protein language model with evolvepro
Kaiyi Jiang, Zhaoqing Yan, Matteo Di Bernardo, Samantha R Sgrizzi, Lukas Villiger, Alisan Kayabolen, BJ Kim, Josephine K Carscadden, Masahiro Hiraizumi, Hiroshi Nishimasu, et al. Rapid in silico directed evolution by a protein language model with evolvepro. Science, page eadr6...
2024
-
[268]
Chatgpt-powered conversational drug editing using retrieval and domain feedback
Shengchao Liu, Jiongxiao Wang, Yijin Yang, Chengpeng Wang, Ling Liu, Hongyu Guo, and Chaowei Xiao. Chatgpt-powered conversational drug editing using retrieval and domain feedback. arXiv preprint arXiv:2305.18090, 2023
2023 arXiv
-
[269]
Sparks of function by de novo protein design
Alexander E Chu, Tianyu Lu, and Po-Ssu Huang. Sparks of function by de novo protein design. Nature biotechnology, 42(2):203– 215, 2024
2024
-
[270]
Instructplm: Aligning protein language models to follow protein structure instructions
Jiezhong Qiu, Junde Xu, Jie Hu, Hanqun Cao, Liya Hou, Zijun Gao, Xinyi Zhou, Anni Li, Xiujuan Li, Bin Cui, et al. Instructplm: Aligning protein language models to follow protein structure instructions. bioRxiv, pages 2024–04, 2024
2024
-
[271]
Toward de novo protein design from natural language
Fengyuan Dai, Yuliang Fan, Jin Su, Chentong Wang, Chenchen Han, Xibin Zhou, Jianming Liu, Hui Qian, Shunzhi Wang, Anping Zeng, et al. Toward de novo protein design from natural language. bioRxiv, pages 2024–08, 2024
2024
-
[272]
Conditional language models enable the efficient design of proficient enzymes
Geraldene Munsamy, Ramiro Illanes-Vicioso, Silvia Funcillo, Ioanna T Nakou, Sebastian Lindner, Gavin Ayres, Lesley S Sheehan, Steven Moss, Ulrich Eckhard, Philipp Lorenz, et al. Conditional language models enable the efficient design of proficient enzymes. bioRxiv, pages 2024–05, 2024
2024
-
[273]
Protagents: Pro- tein discovery via large language model multi-agent collabora- tions combining physics and machine learning
Alireza Ghafarollahi and Markus J Buehler. Protagents: Pro- tein discovery via large language model multi-agent collabora- tions combining physics and machine learning. arXiv preprint arXiv:2402.04268, 2024
2024 arXiv
-
[274]
Clonal selection and learning in the antibody system
Klaus Rajewsky. Clonal selection and learning in the antibody system. Nature, 381(6585):751–758, 1996. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 33
1996
-
[275]
Animal immunization, in vitro display technologies, and machine learning for antibody discovery
Andreas H Laustsen, Victor Greiff, Aneesh Karatt-Vellatt, Serge Muyldermans, and Timothy P Jenkins. Animal immunization, in vitro display technologies, and machine learning for antibody discovery. Trends in Biotechnology, 39(12):1263–1273, 2021
2021
-
[276]
De novo generation of sars-cov-2 antibody cdrh3 with a pre- trained generative large language model
Haohuai He, Bing He, Lei Guan, Yu Zhao, Feng Jiang, Guanxing Chen, Qingge Zhu, Calvin Yu-Chian Chen, Ting Li, and Jianhua Yao. De novo generation of sars-cov-2 antibody cdrh3 with a pre- trained generative large language model. Nature Communications, 15(1):6867, 2024
2024
-
[277]
Unsupervised evolution of protein and antibody complexes with a structure-informed language model
Varun R Shanker, Theodora UJ Bruun, Brian L Hie, and Peter S Kim. Unsupervised evolution of protein and antibody complexes with a structure-informed language model. Science, 385(6704):46– 53, 2024
2024
-
[278]
Enzymes: principles and biotechnological applications
Peter K Robinson. Enzymes: principles and biotechnological applications. Essays in biochemistry, 59:1, 2015
2015
-
[279]
Dissecting en- zyme function with microfluidic-based deep mutational scanning
Philip A Romero, Tuan M Tran, and Adam R Abate. Dissecting en- zyme function with microfluidic-based deep mutational scanning. Proceedings of the National Academy of Sciences , 112(23):7159–7164, 2015
2015
-
[280]
Effects of ionizing radiation on in vitro replication of dna by dna polymerase i
Roy Saffhill and JJ Weiss. Effects of ionizing radiation on in vitro replication of dna by dna polymerase i. Nature New Biology, 241(107):69–71, 1973
1973
-
[281]
Proteases as therapeutics
Charles S Craik, Michael J Page, and Edwin L Madison. Proteases as therapeutics. Biochemical Journal, 435(1):1–16, 2011
2011
-
[282]
Engineering the third wave of bio- catalysis
Uwe T Bornscheuer, GW Huisman, RJ Kazlauskas, S Lutz, JC Moore, and K Robins. Engineering the third wave of bio- catalysis. Nature, 485(7397):185–194, 2012
2012
-
[283]
Computa- tional scoring and experimental evaluation of enzymes generated by neural networks
Sean R Johnson, Xiaozhi Fu, Sandra Viknander, Clara Goldin, Sarah Monaco, Aleksej Zelezniak, and Kevin K Yang. Computa- tional scoring and experimental evaluation of enzymes generated by neural networks. Nature biotechnology, pages 1–10, 2024
2024
-
[284]
Molecular mechanisms of translational control
Fátima Gebauer and Matthias W Hentze. Molecular mechanisms of translational control. Nature reviews Molecular cell biology , 5(10):827–835, 2004
2004
-
[285]
Protein–dna/rna interactions: an overview of investigation methods in the-omics era
Flora Cozzolino, Ilaria Iacobucci, Vittoria Monaco, and Maria Monti. Protein–dna/rna interactions: an overview of investigation methods in the-omics era. Journal of Proteome Research, 20(6):3018– 3030, 2021
2021
-
[286]
Discovering drug–target interaction knowledge from biomedical literature
Yutai Hou, Yingce Xia, Lijun Wu, Shufang Xie, Yang Fan, Jin- hua Zhu, Tao Qin, and Tie-Yan Liu. Discovering drug–target interaction knowledge from biomedical literature. Bioinformatics, 38(22):5100–5107, 2022
2022
-
[287]
Principles of early drug discovery
James P Hughes, Stephen Rees, S Barrett Kalindjian, and Karen L Philpott. Principles of early drug discovery. British journal of pharmacology, 162(6):1239–1249, 2011
2011
-
[288]
Transdti: transformer-based language models for estimating dtis and building a drug recommendation workflow
Yogesh Kalakoti, Shashank Yadav, and Durai Sundar. Transdti: transformer-based language models for estimating dtis and building a drug recommendation workflow. ACS omega, 7(3):2706– 2717, 2022
2022
-
[289]
Large scale paired antibody language models
Henry Kenlay, Frédéric A Dreyer, Aleksandr Kovaltsuk, Dom Miketa, Douglas Pires, and Charlotte M Deane. Large scale paired antibody language models. PLOS Computational Biology , 20(12):e1012646, 2024
2024
-
[290]
Linker-tuning: Optimizing continuous prompts for heterodimeric protein prediction
Shuxian Zou, Hui Li, Shentong Mo, Xingyi Cheng, Eric Xing, and Le Song. Linker-tuning: Optimizing continuous prompts for heterodimeric protein prediction. arXiv preprint arXiv:2312.01186, 2023
2023 arXiv
-
[291]
Rotamer density estimator is an unsupervised learner of the effect of mutations on protein-protein interaction
Shitong Luo, Yufeng Su, Zuofan Wu, Chenpeng Su, Jian Peng, and Jianzhu Ma. Rotamer density estimator is an unsupervised learner of the effect of mutations on protein-protein interaction. bioRxiv, pages 2023–02, 2023
2023
-
[292]
Protein design: the experts speak
Anne Doerr. Protein design: the experts speak. Nature Biotechnol- ogy, 42(2):175–178, 2024
2024
-
[293]
Protein language models learn evolutionary statistics of interacting sequence motifs
Zhidian Zhang, Hannah K Wayment-Steele, Garyk Brixi, Haobo Wang, Dorothee Kern, and Sergey Ovchinnikov. Protein language models learn evolutionary statistics of interacting sequence motifs. Proceedings of the National Academy of Sciences , 121(45):e2406285121, 2024
2024
-
[294]
Interplm: Discovering interpretable features in protein language models via sparse autoencoders
Elana Simon and James Zou. Interplm: Discovering interpretable features in protein language models via sparse autoencoders. bioRxiv, pages 2024–11, 2024
2024
-
[295]
Towards explainable artificial intelligence (xai): A data mining perspective
Haoyi Xiong, Xiaofei Zhang, Jiamin Chen, Xinhao Sun, Yuchen Li, Zeyi Sun, Mengnan Du, et al. Towards explainable artificial intelligence (xai): A data mining perspective. arXiv preprint arXiv:2401.04374, 2024
2024 arXiv
-
[296]
Scientific discovery in the age of artificial intelligence
Hanchen Wang, Tianfan Fu, Yuanqi Du, Wenhao Gao, Kexin Huang, Ziming Liu, Payal Chandak, Shengchao Liu, Peter Van Katwyk, Andreea Deac, et al. Scientific discovery in the age of artificial intelligence. Nature, 620(7972):47–60, 2023
2023
-
[297]
Gpt-4 technical report
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023
2023 arXiv
-
[298]
Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller
Andres M. Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. Augmenting large language models with chemistry tools. Nature Machine Intelligence, pages 1–11, 2024
2024
-
[299]
Autonomous chemical research with large language models
Daniil A Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes. Autonomous chemical research with large language models. Nature, 624(7992):570–578, 2023
2023
-
[2022]
Nucleic acids research, 51(D1):D418–D427, 2023
2023
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.