Pith. sign in

REVIEW 5 major objections 6 minor 1 cited by

Computational Protein Science in the Era of Large Language Models (LLMs)

T0 review · 5 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Protein language models—LLMs trained on amino acid sequences, structures, and scientific text—can grasp the “grammar” and “semantics” of proteins and be adapted to a broad range of structure prediction, function prediction, and protein…

desk verdict A useful, well-organized survey of pLMs whose abstract overclaims generalization; worth refereeing after revision. read the letter →

arxiv 2501.10282 v2 pith:DQZLB4BP submitted 2025-01-17 cs.CE cs.CLq-bio.BM

classification cs.CEcs.CLq-bio.BM
keywords proteinlanguagemodelslargestructurepredictionfunctiondesignantibodyenzymedrugdiscovery
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that protein language models (pLMs), LLMs trained on amino acid sequences and, in later versions, structures, functions, and text, learn the “grammar” and “semantics” of proteins and can be adapted to a wide range of sequence-structure-function problems. The authors organize pLMs by the kind of protein knowledge they master: sequence patterns, explicit structure and function information, and external scientific languages. They then trace how these models are used for structure prediction, function prediction, and protein design, including wet-lab-validated applications in antibody and enzyme design and drug-target interaction prediction. A sympathetic reader takes the paper as a map of a field in which pLMs are becoming a general substrate for protein science rather than task-specific tools.

What carries the argument

The load-bearing mechanism is the pLM itself, treated as a foundation model: a Transformer (or state-space) network pre-trained on large protein corpora so that amino acid tokens behave like words, with “grammar” in residue patterns and “semantics” in encoded structure and function. The paper identifies three technical routes that carry downstream transfer: (1) representation extraction, where frozen or fine-tuned pLM encodings feed prediction heads; (2) likelihood inference, where the model's probability of a mutant versus wild-type sequence is used as a zero-shot fitness score; and (3) prompting and instruction tuning, where a unified decoder answers protein questions or generates sequences under text or function control. Structure tokenization (VQ-VAE-derived tokens such as 3Di) is a notable sub-mechanism that lets 3D structure enter language-model training as discrete tokens.

What would settle it

A controlled benchmark that runs sequence-only, structure-enhanced, and multimodal pLMs on a fixed set of structure, function, and design tasks would settle the central claim: for example, if SaProt and ProLLaMA do not systematically beat ESM-2 on held-out tasks, the paper's hierarchy of protein knowledge fails, and if no pLM category transfers beyond task-specific models trained from scratch, the generalization story collapses.

Watch

Extended reading notes

Core claim

The paper's central claim is that protein language models “skillfully grasp the foundational knowledge of proteins and can be effectively generalized to solve a diversity of sequence-structure-function reasoning problems.” The evidence it assembles is a taxonomy: sequence-only pLMs (ESM-2, ProtGPT2, xTrimoPGLM) capture evolutionarily favored amino acid patterns; structure- and function-enhanced pLMs (SaProt, ESM-3) add explicit 3D and annotation knowledge; multimodal pLMs (ProLLaMA, BioT5) bridge protein sequences with natural language and molecule languages. On top of these foundations, the paper reports pLM-based single-sequence structure prediction comparable to MSA-based methods, zero-shot fitness and mutation-effect prediction, text-guided and condition-tagged protein generation, and ChatGPT-like protein question answering. The intended conclusion is that pLMs, not bespoke per-task models, now carry the main line of computational protein science.

Load-bearing premise

The survey's usefulness rests on the assumption that its chosen models and papers give a representative map of the field; no search protocol or inclusion criteria are stated, so the taxonomy could miss important pLMs or misweight the landscape.

Editorial extensions

If this is right

  • Single-sequence structure prediction methods such as ESMFold can replace slow MSA searches for many proteins, making structure inference practical for orphan and fast-evolving proteins.
  • Zero-shot mutation-effect scoring by pLMs gives experimental labs a cheap first pass at fitness landscapes before deep mutational scanning.
  • Function prediction moves from many task-specific models to unified question-answering systems that answer property, annotation, site, and interaction questions with one model.
  • Conditional and text-guided pLMs extend protein design beyond redesign, generating de novo sequences, antibodies targeting new variants, and enzymes with improved stability.
  • Structure- and function-enhanced pLMs, including multi-track models like ESM-3, suggest that sequence, structure, and function can be treated as interchangeable token tracks in one generative model.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the survey's claims, the taxonomy suggests that the next wave of pLMs will blur its three categories: sequence-only models will acquire structure tokens, and multimodal models will absorb molecule and text modalities, so the field's frontier becomes integration cost rather than architecture.
  • A testable extension of the paper's generalizability claim: on fixed benchmarks, pLM-based methods should dominate task-specific models trained from scratch whenever labeled data are scarce; a controlled comparison across ProteinGym-like tasks would settle this.
  • Because pLM likelihood is used as a proxy for evolutionary plausibility, species bias in training databases may leak into fitness predictions; adjusting training data composition is a direct extension of the redesign workflow the paper describes.
  • The survey's coverage assumption can be checked by re-running its categorization against a systematic search of recent pLMs; if major models fall outside the three categories, the proposed map needs revision.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. This manuscript is a survey of protein language models (pLMs) and their applications in computational protein science. It proposes a taxonomy that divides pLMs into sequence-based, structure-and-function-enhanced, and multimodal models; reviews how pLMs are used for structure prediction, function prediction, and protein design; and discusses applications in antibody design, enzyme design, and drug discovery. The Abstract states that pLMs 'skillfully grasp the foundational knowledge of proteins and can be effectively generalized to solve a diversity of sequence-structure-function reasoning problems,' and the survey is organized to support this claim by mapping pLMs to downstream tasks.

Significance. If the survey's claims hold, it would provide a valuable entry point for researchers from both AI and biology backgrounds. The taxonomy is thoughtful and the coverage is unusually broad, including recent models such as ESM-3, DPLM-2, and xTrimoPGLM, as well as practical applications with wet-lab validation. The detailed tables (Tables 1–4) that list corpora, architectures, parameter counts, and pre-training objectives are a useful resource. The paper does not present original machine-checked proofs or reproducible code, but it cites primary sources for most claims, which is appropriate for a survey. The main contribution is organizational rather than experimental, and its usefulness depends on whether the selected literature is representative and whether the claims in the Abstract are calibrated to the evidence.

major comments (5)
  1. [Abstract and Section 4] The central claim that pLMs are 'effectively generalized to solve a diversity of sequence-structure-function reasoning problems' is broader than the evidence presented in Section 4. In Section 4.1, all pLM-based structure prediction methods (ESMFold, HelixFold-Single, OmegaFold, trRosettaX-Single, RGN2, IgFold) require an additional trained folding trunk or prediction head. In Section 4.2, most function prediction methods use the LM-as-encoder scheme with a learned classifier, fine-tuning, or parameter-efficient fine-tuning; only likelihood-based fitness and mutation-effect prediction are described as zero-shot (Section 4.2.1). In Section 4.3, ProGen and ZymCTRL require control tags or fine-tuning for controllable generation. The survey thus conflates 'adaptable via task-specific training' with 'generalization.' I recommend softening the Abstract to 'adaptable across tasks' or explicitly scoping the generalization claim to the zero-shot settings where evidence exists.
  2. [Section 3, first paragraph] The survey does not state its literature selection method. No databases searched, search dates, keywords, or inclusion/exclusion criteria are provided. This makes the taxonomy non-reproducible and leaves open the possibility that important pLMs were omitted or that the relative emphasis of topics is not representative. As the survey's central contribution is its categorization, a short methodology paragraph describing the search protocol and time window is needed.
  3. [Section 4.1, single-sequence structure prediction paragraph] The sentence 'In investigations, OmegaFold, trRosettaX-Single, and RGN2 are all observed to outperform AlphaFold2 and RoseTTAFold on those orphan proteins and de novo designed proteins' is a specific quantitative claim with no citation. Please provide references for these comparisons or remove the claim. As written, it is a load-bearing assertion for the section's argument that pLM-based methods address the limitations of MSA-based approaches.
  4. [Section 4.2.1] The statement 'the likelihoods inferred from pLMs correlate well with protein fitness [72, 253, 254]' is central to the zero-shot generalization narrative, yet no quantitative evidence is given. The survey should report representative correlation values (e.g., Spearman rho) from ProteinGym or the cited works, or state the range across benchmarks. Without those numbers, the claim is too vague to evaluate.
  5. [Section 4 (general)] The survey seldom reports how pLM-based methods compare with strong non-pLM baselines. For structure prediction, comparisons to AlphaFold2 and RoseTTAFold appear in Section 4.1, but for function prediction (Section 4.2) and protein design (Section 4.3) the text rarely mentions results from BLAST, HMMER, profile-based predictors, or classical machine-learning methods. Without this comparative context, the reader cannot judge whether pLMs are necessary or superior for these tasks. I recommend adding a comparative synthesis from the cited benchmarks (e.g., ProteinGym, FLIP, TAPE) or explicitly softening claims of pLM advantage.
minor comments (6)
  1. [Figure 8 caption] The caption contains the typo 'Workfolw'; it should read 'Workflow'.
  2. [Table 4] The table header spells 'Functional' as 'Funcitional', and the scheme abbreviation 'Mull-Model Fine-Tuning' should be 'Full-Model Fine-Tuning'.
  3. [Figure 9 caption] The second framework in the caption is labeled 'LM-as-Encoder' but the text and Table 4 use 'LM-as-Predictor'; the caption should be corrected for consistency.
  4. [Section 3.1.1 and Table 1] The model 'paired-IgGen' is abbreviated 'p-IgGen' in the text but 'g-IgGen' in Table 1; the abbreviation should be unified.
  5. [Section 3.1.1, BiMamba-S and Table 1] The Mamba architecture is cited via references [96] and [97], which are a survey and a recommendation paper; the original Mamba paper (Gu & Dao, reference [139]) should be cited at first mention.
  6. [Section 6.5] The phrase 'reached the unanimous conclusion of non-optimal' is too strong for two cited studies; a softer formulation such as 'several recent studies have concluded' would be more accurate.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: this is a literature survey whose taxonomy and claims rest on external, independently published methods.

full rationale

The paper is a survey, not a derivation or prediction pipeline. Its central claim—that protein language models learn foundational protein knowledge and can be adapted to structure prediction, function prediction, and design—is supported by citations to externally published models (ESM-2, AlphaFold2, ProGen, ProteinGym, etc.) and to benchmarks with their own evaluations. There is no fitted parameter that is later renamed as a prediction, no equation that is assumed and then re-derived, and no uniqueness theorem invoked to force the survey's organizational choice. The paper's three-way taxonomy (sequence-based, structure/function-enhanced, multimodal pLMs) is a literature-organizing scheme, and its validity does not depend on any of the cited results being equivalent to the taxonomy itself. A few references authored by the survey's own team (e.g., prior surveys on LLMs in recommender systems and retrieval-augmented generation) appear only as background citations and are not load-bearing for the paper's claims about protein language models. The skeptic's concern that 'effectively generalized' overstates the evidence is a legitimate correctness/calibration critique, not a circularity critique: the survey itself documents that most applications use task-specific heads, fine-tuning, or architectural additions, but overstatement of a literature-based conclusion does not make the argument circular. Therefore, no circular step can be exhibited, and the honest finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

This is a survey, so no free parameters or invented entities apply. The axioms listed above are background biological or machine-learning assumptions the paper imports from the cited literature; none are introduced as new postulates.

assumptions (4)
  • domain assumption The sequence-structure-function paradigm holds: amino acid sequence determines 3D structure, and structure determines function.
    Invoked in Section 1 and Figure 1 as the organizing principle; if this were not approximately true, the narrative connecting pLMs to structure prediction, function prediction, and design would lose coherence.
  • domain assumption Sequence-only pLMs capture implicit structural and functional knowledge from large-scale pre-training.
    Section 3.2 opens by assuming this, citing ESM-1b contact maps and ESM-2 scaling results. This assumption bridges the 'sequence-based' and 'structure enhanced' categories in the survey.
  • domain assumption pLM likelihood scores correlate with protein fitness.
    Section 4.2.1 states this as a practical result, citing ProteinGym and related work; the zero-shot fitness prediction claims depend on this imported correlation.
  • domain assumption Scaling laws observed for natural-language LLMs transfer to protein language models.
    Sections 2.3 and 6.5 use scaling-law reasoning to justify larger pLMs and to frame computational efficiency as an open challenge; this is an imported machine-learning assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Computational Protein Science in the Era of Large Language Models (LLMs)." pith.science (2026). https://pith.science/paper/DQZLB4BP

@misc{pith2026250110282,
  author       = {Pith},
  title        = {Pith review of: Computational Protein Science in the Era of Large Language Models (LLMs)},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DQZLB4BP}},
  note         = {Machine review of arXiv:2501.10282}
}
read the original abstract

Considering the significance of proteins, computational protein science has always been a critical scientific field, dedicated to revealing knowledge and developing applications within the protein sequence-structure-function paradigm. In the last few decades, Artificial Intelligence (AI) has made significant impacts in computational protein science, leading to notable successes in specific protein modeling tasks. However, those previous AI models still meet limitations, such as the difficulty in comprehending the semantics of protein sequences, and the inability to generalize across a wide range of protein modeling tasks. Recently, LLMs have emerged as a milestone in AI due to their unprecedented language processing & generalization capability. They can promote comprehensive progress in fields rather than solving individual tasks. As a result, researchers have actively introduced LLM techniques in computational protein science, developing protein Language Models (pLMs) that skillfully grasp the foundational knowledge of proteins and can be effectively generalized to solve a diversity of sequence-structure-function reasoning problems. While witnessing prosperous developments, it's necessary to present a systematic overview of computational protein science empowered by LLM techniques. First, we summarize existing pLMs into categories based on their mastered protein knowledge, i.e., underlying sequence patterns, explicit structural and functional information, and external scientific languages. Second, we introduce the utilization and adaptation of pLMs, highlighting their remarkable achievements in promoting protein structure prediction, protein function prediction, and protein design studies. Then, we describe the practical application of pLMs in antibody design, enzyme design, and drug discovery. Finally, we specifically discuss the promising future directions in this fast-growing field.

Figures

Figures reproduced from arXiv: 2501.10282 by the authors.

Figure 1
Figure 1. Illustration of the Evolution and Sequence-Structure [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Biological basis and data profiles. (A) Protein synthesis mainly involves the transcription of protein-coding genes to mRNAs and the translation of codon sequences to AA sequences. (B) Multiple Sequence Alignment (MSA) contains the evolutionary prior knowledge of proteins. Conserved positions are interpreted as core AAs for protein structure, as no changes have been allowed throughout the evolutionary process. Pairs… view at source ↗
Figure 3
Figure 3. Typical single-sequence-based pLMs. 1-3) When considering individual amino acid sequences as "sentences", pLMs follow the general approaches of autoencoding, autoregressive, and sequence-to-sequence as well. 4) Masked CDR reconstruction is a novel pre-training objective that incorporates the inherent characteristics of antibodies into mask language modeling. language, resulting in ProtBERT, ProtAlbert, and ProtElect… view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Typical multiple-sequences-based pLMs. 1) ESM-MSA￾1b, a representative MSA-based pLM, incorporates bidirectional tied-row attention and column attention within each MSA Trans￾former block, thereby capturing co-evolution features within the 2D input. 2) PoET is an autor…
Figure 5
Figure 5. Figure 5: Typical structure-enhanced pLMs. 1) Pre-calculated structural features can be injected into the input AA sequence as position encoding, or utilized in an additional training objective. 2) Considering the significant correlation between the Transformer attention map and…
Figure 6
Figure 6. Figure 6: Typical function-enhanced pLMs. 1-3) pLMs learn the forward, inverse, and bidirectional correlations between protein sequences and functional labels. They undergo pre-training using various objectives, including function prediction, conditional sequence generation, or …
Figure 7
Figure 7. Figure 7: Typical multimodal pLMs. 1) In unified multimodal learning, multiple scientific languages share a unified latent space. We can perform multi-task pre-training from scratch, as well as incremental pre-training or instruction tuning based on pre-trained LMs. 2) In "speci…
Figure 8
Figure 8. Figure 8: Workfolw overview of AlphaFold2 [8] and ESMFold [25]. Both AlphaFold2 and ESMFold infer the high-resolution protein structure from a comprehensive understanding of the sequence. AlphaFold2 relies on MSA to gain evolutionary insights encoded in protein sequences, while …
Figure 9
Figure 9. Figure 9: Typical technical schemes in pLM-based protein function prediction. 1) In the "LM-as-Encoder" scheme, protein language models play a central role in the full models as encoders. Predictions could be inferred from the likelihood of language modeling or obtained through …
Figure 10
Figure 10. Figure 10: Protein Redesign: Function-Oriented Protein Sequence Optimization. 1) Directed evolution is performed on the func￾tional property landscape over sets of protein mutants. While the landscape exists conceptually, it is generally not fully revealed. Therefore, protein se…
Figure 11
Figure 11. Figure 11: Overview of Antibody Design. (A) In the process of recognizing antigens, the complementarity-determining regions (CDRs) of antibodies are critically important. (B) Traditional antibody production is limited by the immunity of animals. (C) Protein language models can b…
Figure 12
Figure 12. Figure 12: Overview of Enzyme Design. (A) Natural enzymes are usually sensitive to the reaction environment. (B) Genes determine the function of natural enzymes. (C) pLMs can assist in the design of enhanced enzymes. This figure is created in BioRender.com. A. Protein-Small Mole…
Figure 13
Figure 13. Figure 13: Overview of Drug Discovery. (A) A visualization for protein-ligand interaction, which is obtained from PDB entry 8PPB. (B) Overview of a general drug discovery process [287]. (C) pLMs are employed to predict drug-target interactions, potentially accelerating drug disc…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous Structure Tokens

    cs.CE 2025-12 conditional novelty 6.0 of 10

    HD-Prot shows that a protein language model can jointly generate sequences and structures using continuous structure tokens instead of quantized tokens, reaching competitive performance on four protein design tasks.

Reference graph

Works this paper leans on

300 extracted references · 10 canonical work pages · cited by 1 Pith paper

  1. [1]

    Controllable protein design with language models

    Noelia Ferruz and Birte Höcker. Controllable protein design with language models. Nature Machine Intelligence, 4(6):521–532, 2022

  2. [2]

    Learning the protein language: Evolution, structure, and function

    Tristan Bepler and Bonnie Berger. Learning the protein language: Evolution, structure, and function. Cell systems , 12(6):654–669, 2021

  3. [3]

    Principles that govern the folding of protein chains

    Christian B Anfinsen. Principles that govern the folding of protein chains. Science, 181(4096):223–230, 1973

  4. [4]

    Exploring the structure and function paradigm

    Oliver C Redfern, Benoit Dessailly, and Christine A Orengo. Exploring the structure and function paradigm. Current opinion in structural biology, 18(3):394–402, 2008

  5. [5]

    Natural selection and the concept of a protein space

    John Maynard Smith. Natural selection and the concept of a protein space. Nature, 225(5232):563–564, 1970

  6. [6]

    The language of pro- teins: Nlp, machine learning & protein sequences

    Dan Ofer, Nadav Brandes, and Michal Linial. The language of pro- teins: Nlp, machine learning & protein sequences. Computational and Structural Biotechnology Journal, 19:1750–1758, 2021

  7. [7]

    Unified rational protein engineering with sequence-based deep representation learning

    Ethan C Alley, Grigory Khimulya, Surojit Biswas, Mohammed AlQuraishi, and George M Church. Unified rational protein engineering with sequence-based deep representation learning. Nature methods, 16(12):1315–1322, 2019

  8. [8]

    Highly accurate protein structure prediction with alphafold

    John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021

Show all 300 references
  1. [9]

    Accurate prediction of protein structures and interactions using a three-track neural network

    Minkyung Baek, Frank DiMaio, Ivan Anishchenko, Justas Dau- paras, Sergey Ovchinnikov, Gyu Rie Lee, Jue Wang, Qian Cong, Lisa N Kinch, R Dustin Schaeffer, et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science, 373(6557):87...

  2. [10]

    Deepgoplus: im- proved protein function prediction from sequence

    Maxat Kulmanov and Robert Hoehndorf. Deepgoplus: im- proved protein function prediction from sequence. Bioinformatics, 36(2):422–429, 2020

  3. [11]

    Prediction of designer-recombinases for dna editing with generative deep learning

    Lukas Theo Schmitt, Maciej Paszkowski-Rogacz, Florian Jug, and Frank Buchholz. Prediction of designer-recombinases for dna editing with generative deep learning. Nature Communications, 13(1):7966, 2022

  4. [12]

    Ig-vae: Generative modeling of protein structure by direct 3d coordinate generation

    Raphael R Eguchi, Christian A Choe, and Po-Ssu Huang. Ig-vae: Generative modeling of protein structure by direct 3d coordinate generation. PLoS computational biology, 18(6):e1010271, 2022

  5. [13]

    Bert: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018

  6. [14]

    Exploring the limits of transfer learning with a unified text-to-text transformer

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67, 2020

  7. [15]

    Improving language understanding by generative pre- training

    Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre- training. 2018

  8. [16]

    Language models are unsupervised multitask learners

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019

  9. [17]

    Language models are few-shot learners

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020

  10. [18]

    Llama: Open and efficient foundation language models

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 , 2023

  11. [19]

    A survey of large language models

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 2023

  12. [20]

    Recommender systems in the era of large language models (llms)

    Zihuai Zhao, Wenqi Fan, Jiatong Li, Yunqing Liu, Xiaowei Mei, Yiqi Wang, Zhen Wen, Fei Wang, Xiangyu Zhao, Jiliang Tang, et al. Recommender systems in the era of large language models (llms). IEEE Transactions on Knowledge and Data Engineering, 2024

  13. [21]

    Tokenrec: Learning to tokenize id for llm-based generative recommendation

    Haohao Qu, Wenqi Fan, Zihuai Zhao, and Qing Li. Tokenrec: Learning to tokenize id for llm-based generative recommendation. arXiv preprint arXiv:2406.10450, 2024

  14. [22]

    A survey of large language models for healthcare: from data, technology, and applications to accountabil- ity and ethics

    Kai He, Rui Mao, Qika Lin, Yucheng Ruan, Xiang Lan, Mengling Feng, and Erik Cambria. A survey of large language models for healthcare: from data, technology, and applications to accountabil- ity and ethics. arXiv preprint arXiv:2310.05694, 2023

  15. [23]

    To transformers and beyond: Large language models for the genome

    Micaela E Consens, Cameron Dufault, Michael Wainberg, Duncan Forster, Mehran Karimzadeh, Hani Goodarzi, Fabian J Theis, Alan Moses, and Bo Wang. To transformers and beyond: Large language models for the genome. arXiv preprint arXiv:2311.07621, 2023

  16. [24]

    Empowering molecule discovery for molecule-caption translation with large language models: A chatgpt perspective

    Jiatong Li, Yunqing Liu, Wenqi Fan, Xiao-Yong Wei, Hui Liu, Jiliang Tang, and Qing Li. Empowering molecule discovery for molecule-caption translation with large language models: A chatgpt perspective. IEEE Transactions on Knowledge and Data Engineering, 2024

  17. [25]

    Evolutionary-scale prediction of atomic-level protein structure with a language model

    Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637):1123– 1130, 2023

  18. [26]

    Protgpt2 is a deep unsupervised language model for protein design

    Noelia Ferruz, Steffen Schmidt, and Birte Höcker. Protgpt2 is a deep unsupervised language model for protein design. Nature communications, 13(1):4348, 2022

  19. [27]

    xtrimopglm: unified 100b-scale pre-trained transformer for deciphering the language of protein

    Bo Chen, Xingyi Cheng, Pan Li, Yangli-ao Geng, Jing Gong, Shen Li, Zhilei Bei, Xu Tan, Boyan Wang, Xin Zeng, et al. xtrimopglm: unified 100b-scale pre-trained transformer for deciphering the language of protein. arXiv preprint arXiv:2401.06199, 2024

  20. [28]

    Msa transformer

    Roshan M Rao, Jason Liu, Robert Verkuil, Joshua Meier, John Canny, Pieter Abbeel, Tom Sercu, and Alexander Rives. Msa transformer. In International Conference on Machine Learning, pages 8844–8856. PMLR, 2021

  21. [29]

    Saprot: protein language modeling with structure- aware vocabulary

    Jin Su, Chenchen Han, Yuyang Zhou, Junjie Shan, Xibin Zhou, and Fajie Yuan. Saprot: protein language modeling with structure- aware vocabulary. bioRxiv, pages 2023–10, 2023

  22. [30]

    Simulating 500 million years of evolution with a language model

    Thomas Hayes, Roshan Rao, Halil Akin, Nicholas J Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q Tran, Jonathan Deaton, Marius Wiggert, et al. Simulating 500 million years of evolution with a language model. Science, page eads0018, 2025

  23. [31]

    Prollama: A protein large language model for multi-task protein language processing

    Liuzhenghao Lv, Zongying Lin, Hao Li, Yuyang Liu, Jiaxi Cui, Calvin Yu-Chian Chen, Li Yuan, and Yonghong Tian. Prollama: A protein large language model for multi-task protein language processing. arXiv preprint arXiv:2402.16445, 2024

  24. [32]

    Biot5: Enriching cross-modal integration in biology with chemical knowledge and natural language associations

    Qizhi Pei, Wei Zhang, Jinhua Zhu, Kehan Wu, Kaiyuan Gao, Lijun Wu, Yingce Xia, and Rui Yan. Biot5: Enriching cross-modal integration in biology with chemical knowledge and natural language associations. arXiv preprint arXiv:2310.07276, 2023. JOURNAL OF LATEX CLASS FILES, VOL. ...

  25. [33]

    Protein- chat: Towards achieving chatgpt-like functionalities on protein 3d structures

    Han Guo, Mingjia Huo, Ruiyi Zhang, and Pengtao Xie. Protein- chat: Towards achieving chatgpt-like functionalities on protein 3d structures. Authorea Preprints, 2023

  26. [34]

    Proteinnpt: improving protein property prediction and design with non-parametric transformers

    Pascal Notin, Ruben Weitzman, Debora Marks, and Yarin Gal. Proteinnpt: improving protein property prediction and design with non-parametric transformers. Advances in Neural Information Processing Systems, 36, 2024

  27. [35]

    Large language models generate functional protein sequences across diverse families

    Ali Madani, Ben Krause, Eric R Greene, Subu Subramanian, Benjamin P Mohr, James M Holton, Jose Luis Olmos, Caiming Xiong, Zachary Z Sun, Richard Socher, et al. Large language models generate functional protein sequences across diverse families. Nature Biotechnology, 41(8):1099...

  28. [36]

    Scientific large language models: A survey on biological & chemical domains

    Qiang Zhang, Keyang Ding, Tianwen Lyv, Xinda Wang, Qingyu Yin, Yiwen Zhang, Jing Yu, Yuhao Wang, Xiaotong Li, Zhuoyi Xiang, et al. Scientific large language models: A survey on biological & chemical domains. arXiv preprint arXiv:2401.14656, 2024

  29. [37]

    Protein language models and structure prediction: Connection and progression

    Bozhen Hu, Jun Xia, Jiangbin Zheng, Cheng Tan, Yufei Huang, Yongjie Xu, and Stan Z Li. Protein language models and structure prediction: Connection and progression. arXiv preprint arXiv:2211.16742, 2022

  30. [38]

    Learning functional properties of proteins with language models

    Serbulent Unsal, Heval Atas, Muammer Albayrak, Kemal Turhan, Aybar C Acar, and Tunca Do˘ gan. Learning functional properties of proteins with language models. Nature Machine Intelligence , 4(3):227–245, 2022

  31. [39]

    Designing proteins with language models

    Jeffrey A Ruffolo and Ali Madani. Designing proteins with language models. Nature Biotechnology, 42(2):200–202, 2024

  32. [40]

    Mass spectrometry: principles and applications

    Edmond De Hoffmann and Vincent Stroobant. Mass spectrometry: principles and applications. John Wiley & Sons, 2007

  33. [41]

    A new generation of crystallographic validation tools for the protein data bank

    Randy J Read, Paul D Adams, W Bryan Arendall, Axel T Brunger, Paul Emsley, Robbie P Joosten, Gerard J Kleywegt, Eugene B Krissinel, Thomas Lütteke, Zbyszek Otwinowski, et al. A new generation of crystallographic validation tools for the protein data bank. Structure, 19(10):139...

  34. [42]

    Outcome of the first electron microscopy validation task force meeting

    Richard Henderson, Andrej Sali, Matthew L Baker, Bridget Carragher, Batsal Devkota, Kenneth H Downing, Edward H Egelman, Zukang Feng, Joachim Frank, Nikolaus Grigorieff, et al. Outcome of the first electron microscopy validation task force meeting. Structure, 20(2):205–214, 2012

  35. [43]

    Deep mutational scanning: a new style of protein science

    Douglas M Fowler and Stanley Fields. Deep mutational scanning: a new style of protein science. Nature methods, 11(8):801–807, 2014

  36. [44]

    Uniref clusters: a comprehensive and scalable alternative for improving sequence similarity searches

    Baris E Suzek, Yuqi Wang, Hongzhan Huang, Peter B McGarvey, Cathy H Wu, and UniProt Consortium. Uniref clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics, 31(6):926–932, 2015

  37. [45]

    Nucleic acids research, 47(D1):D520–D528, 2019

    Protein data bank: the single global archive for 3d macromolecular structure data. Nucleic acids research, 47(D1):D520–D528, 2019

  38. [46]

    Gene ontology: tool for the unification of biology

    Michael Ashburner, Catherine A Ball, Judith A Blake, David Botstein, Heather Butler, J Michael Cherry, Allan P Davis, Kara Dolinski, Selina S Dwight, Janan T Eppig, et al. Gene ontology: tool for the unification of biology. Nature genetics, 25(1):25–29, 2000

  39. [47]

    Central dogma of molecular biology

    Francis Crick. Central dogma of molecular biology. Nature, 227(5258):561–563, 1970

  40. [48]

    Origin and evolution of the genetic code: the universal enigma

    Eugene V Koonin and Artem S Novozhilov. Origin and evolution of the genetic code: the universal enigma. IUBMB life, 61(2):99–111, 2009

  41. [49]

    One thousand families for the molecular biologist

    Cyrus Chothia. One thousand families for the molecular biologist. Nature, 357(6379), 1992

  42. [50]

    The protein-folding problem, 50 years on

    Ken A Dill and Justin L MacCallum. The protein-folding problem, 50 years on. science, 338(6110):1042–1046, 2012

  43. [51]

    Learning from protein structure with geometric vector perceptrons

    Bowen Jing, Stephan Eismann, Patricia Suriana, Raphael John Lamarre Townshend, and Ron Dror. Learning from protein structure with geometric vector perceptrons. In International Conference on Learning Representations, 2020

  44. [52]

    Pifold: Toward effective and efficient protein inverse folding

    Zhangyang Gao, Cheng Tan, Pablo Chacón, and Stan Z Li. Pifold: Toward effective and efficient protein inverse folding. arXiv preprint arXiv:2209.12643, 2022

  45. [53]

    Exploring protein fitness landscapes by directed evolution

    Philip A Romero and Frances H Arnold. Exploring protein fitness landscapes by directed evolution. Nature reviews Molecular cell biology, 10(12):866–876, 2009

  46. [54]

    A survey on protein representation learning: Retrospect and prospect

    Lirong Wu, Yufei Huang, Haitao Lin, and Stan Z Li. A survey on protein representation learning: Retrospect and prospect. arXiv preprint arXiv:2301.00813, 2022

  47. [55]

    Convolutions are competitive with transformers for protein sequence pretraining

    Kevin K Yang, Nicolo Fusi, and Alex X Lu. Convolutions are competitive with transformers for protein sequence pretraining. Cell Systems, 15(3):286–294, 2024

  48. [56]

    Prottrans: Toward understanding the language of life through self-supervised learning

    Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, et al. Prottrans: Toward understanding the language of life through self-supervised learning. IEEE transactions on pattern ...

  49. [57]

    Protein representation learning by geometric structure pretraining

    Zuobai Zhang, Minghao Xu, Arian Jamasb, Vijil Chenthama- rakshan, Aurelie Lozano, Payel Das, and Jian Tang. Protein representation learning by geometric structure pretraining. arXiv preprint arXiv:2203.06125, 2022

  50. [58]

    Lm-gvp: an extensible sequence and structure informed deep learning framework for protein property prediction

    Zichen Wang, Steven A Combs, Ryan Brand, Miguel Romero Calvo, Panpan Xu, George Price, Nataliya Golovach, Emmanuel O Salawu, Colby J Wise, Sri Priya Ponnapalli, et al. Lm-gvp: an extensible sequence and structure informed deep learning framework for protein property prediction...

  51. [59]

    A systematic study of joint representation learning on protein sequences and structures

    Z Zhang, C Wang, M Xu, V Chenthamarakshan, AC Lozano, P Das, and J Tang. A systematic study of joint representation learning on protein sequences and structures. Preprint at http://arxiv. org/abs/2303.06275, 2023

  52. [60]

    Colabfold: making protein folding accessible to all

    Milot Mirdita, Konstantin Schütze, Yoshitaka Moriwaki, Lim Heo, Sergey Ovchinnikov, and Martin Steinegger. Colabfold: making protein folding accessible to all. Nature methods, 19(6):679–682, 2022

  53. [61]

    Single-sequence protein structure prediction using supervised transformer protein language models

    Wenkai Wang, Zhenling Peng, and Jianyi Yang. Single-sequence protein structure prediction using supervised transformer protein language models. Nature Computational Science , 2(12):804–814, 2022

  54. [62]

    A method for multiple-sequence-alignment-free protein structure prediction using a protein language model

    Xiaomin Fang, Fan Wang, Lihang Liu, Jingzhou He, Dayong Lin, Yingfei Xiang, Kunrui Zhu, Xiaonan Zhang, Hua Wu, Hui Li, et al. A method for multiple-sequence-alignment-free protein structure prediction using a protein language model. Nature Machine Intelligence, 5(10):1087–1096, 2023

  55. [63]

    Flip: Benchmark tasks in fitness landscape inference for proteins

    Christian Dallago, Jody Mou, Kadina E Johnston, Bruce J Wittmann, Nicholas Bhattacharya, Samuel Goldman, Ali Madani, and Kevin K Yang. Flip: Benchmark tasks in fitness landscape inference for proteins. bioRxiv, pages 2021–11, 2021

  56. [64]

    Exploring machine learning algorithms and protein language models strategies to develop enzyme classi- fication systems

    Diego Fernández, Álvaro Olivera-Nappa, Roberto Uribe-Paredes, and David Medina-Ortiz. Exploring machine learning algorithms and protein language models strategies to develop enzyme classi- fication systems. In International Work-Conference on Bioinformatics and Biomedical Engi...

  57. [65]

    Protein–dna binding sites prediction based on pre-trained protein language model and contrastive learning

    Yufan Liu and Boxue Tian. Protein–dna binding sites prediction based on pre-trained protein language model and contrastive learning. Briefings in Bioinformatics, 25(1):bbad488, 2024

  58. [66]

    Genome-scale annotation of protein binding sites via language model and geometric deep learning

    Qianmu Yuan, Chong Tian, and Yuedong Yang. Genome-scale annotation of protein binding sites via language model and geometric deep learning. Elife, 13:RP93695, 2024

  59. [67]

    Contrastive learning in protein language space predicts interactions between drugs and protein targets

    Rohit Singh, Samuel Sledzieski, Bryan Bryson, Lenore Cowen, and Bonnie Berger. Contrastive learning in protein language space predicts interactions between drugs and protein targets. Proceedings of the National Academy of Sciences , 120(24):e2220778120, 2023

  60. [68]

    Unikp: a unified framework for the prediction of enzyme kinetic parameters

    Han Yu, Huaxiang Deng, Jiahui He, Jay D Keasling, and Xiaozhou Luo. Unikp: a unified framework for the prediction of enzyme kinetic parameters. Nature Communications, 14(1):8211, 2023

  61. [69]

    Protchatgpt: Towards understanding proteins with large language models

    Chao Wang, Hehe Fan, Ruijie Quan, and Yi Yang. Protchatgpt: Towards understanding proteins with large language models. arXiv preprint arXiv:2402.09649, 2024

  62. [70]

    Prot2text: Multimodal protein’s function generation with gnns and transformers

    Hadi Abdine, Michail Chatzianastasis, Costas Bouyioukos, and Michalis Vazirgiannis. Prot2text: Multimodal protein’s function generation with gnns and transformers. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 10757–10765, 2024

  63. [71]

    Machine learning for functional protein design

    Pascal Notin, Nathan Rollins, Yarin Gal, Chris Sander, and Debora Marks. Machine learning for functional protein design. Nature Biotechnology, 42(2):216–228, 2024

  64. [72]

    Language models enable zero-shot prediction of the effects of mutations on protein function

    Joshua Meier, Roshan Rao, Robert Verkuil, Jason Liu, Tom Sercu, and Alex Rives. Language models enable zero-shot prediction of the effects of mutations on protein function. Advances in neural information processing systems, 34:29287–29303, 2021

  65. [73]

    De novo protein design—from new structures to programmable functions

    Tanja Kortemme. De novo protein design—from new structures to programmable functions. Cell, 187(3):526–544, 2024

  66. [74]

    De novo design of protein structure and function with rfdiffusion

    Joseph L Watson, David Juergens, Nathaniel R Bennett, Brian L Trippe, Jason Yim, Helen E Eisenach, Woody Ahern, Andrew J Borst, Robert J Ragotte, Lukas F Milles, et al. De novo design of protein structure and function with rfdiffusion. Nature, 620(7976):1089–1100, 2023. JOURNA...

  67. [75]

    Illuminating protein space with a programmable generative model

    John B Ingraham, Max Baranov, Zak Costello, Karl W Barber, Wujie Wang, Ahmed Ismail, Vincent Frappier, Dana M Lord, Christopher Ng-Thow-Hing, Erik R Van Vlack, et al. Illuminating protein space with a programmable generative model. Nature, 623(7989):1070–1078, 2023

  68. [76]

    Learning inverse folding from millions of predicted structures

    Chloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin, Brian Hie, Tom Sercu, Adam Lerer, and Alexander Rives. Learning inverse folding from millions of predicted structures. In International conference on machine learning, pages 8946–8970. PMLR, 2022

  69. [77]

    Robust deep learning–based protein sequence design using proteinmpnn

    Justas Dauparas, Ivan Anishchenko, Nathaniel Bennett, Hua Bai, Robert J Ragotte, Lukas F Milles, Basile IM Wicky, Alexis Courbet, Rob J de Haas, Neville Bethel, et al. Robust deep learning–based protein sequence design using proteinmpnn. Science, 378(6615):49– 56, 2022

  70. [78]

    Protein generation with evolutionary diffusion: sequence is all you need

    Sarah Alamdari, Nitya Thakkar, Rianne van den Berg, Alex Xijie Lu, Nicolo Fusi, Ava Pardis Amini, and Kevin K Yang. Protein generation with evolutionary diffusion: sequence is all you need. bioRxiv, pages 2023–09, 2023

  71. [79]

    Graph machine learning in the era of large language models (llms)

    Wenqi Fan, Shijie Wang, Jiani Huang, Zhikai Chen, Yu Song, Wenzhuo Tang, Haitao Mao, Hui Liu, Xiaorui Liu, Dawei Yin, et al. Graph machine learning in the era of large language models (llms). arXiv preprint arXiv:2404.14928, 2024

  72. [80]

    Moleculargpt: Open large language model (llm) for few-shot molecular property prediction

    Yuyan Liu, Sirui Ding, Sheng Zhou, Wenqi Fan, and Qiaoyu Tan. Moleculargpt: Open large language model (llm) for few-shot molecular property prediction. arXiv preprint arXiv:2406.12950 , 2024

  73. [81]

    Scaling laws for neural language models

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020

  74. [82]

    Lora: Low-rank adaptation of large language models

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021

  75. [83]

    Qlora: Efficient finetuning of quantized llms

    Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettle- moyer. Qlora: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314, 2023

  76. [84]

    Chain-of-thought prompting elicits reasoning in large language models

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022

  77. [85]

    The power of scale for parameter-efficient prompt tuning

    Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691, 2021

  78. [86]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pa...

  79. [87]

    Palm-e: An embodied multimodal language model

    Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023

  80. [88]

    Leveraging biomolecule and natural language through multi-modal learning: A survey

    Qizhi Pei, Lijun Wu, Kaiyuan Gao, Jinhua Zhu, Yue Wang, Zun Wang, Tao Qin, and Rui Yan. Leveraging biomolecule and natural language through multi-modal learning: A survey. arXiv preprint arXiv:2403.01528, 2024

  81. [89]

    Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730–19742. PMLR, 2023

  82. [90]

    A survey on rag meeting llms: Towards retrieval-augmented large language models

    Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. A survey on rag meeting llms: Towards retrieval-augmented large language models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages...

  83. [91]

    Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences

    Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C Lawrence Zitnick, Jerry Ma, et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National ...

  84. [92]

    Roberta: A robustly optimized bert pretraining approach

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019

  85. [93]

    ESM Cambrian: Revealing the mysteries of proteins with unsupervised learning

    ESM Team. ESM Cambrian: Revealing the mysteries of proteins with unsupervised learning. https://evolutionaryscale.ai/blog/ esm-cambrian, December 2024

  86. [94]

    Pre-training co-evolutionary protein representation via a pairwise masked language model

    Liang He, Shizhuo Zhang, Lijun Wu, Huanhuan Xia, Fusong Ju, He Zhang, Siyuan Liu, Yingce Xia, Jianwei Zhu, Pan Deng, et al. Pre-training co-evolutionary protein representation via a pairwise masked language model. arXiv preprint arXiv:2110.15527, 2021

  87. [95]

    Long-context protein language model

    Yingheng Wang, Zichen Wang, Gil Sadeh, Luca Zancato, Alessan- dro Achille, George Karypis, and Huzefa Rangwala. Long-context protein language model. bioRxiv, pages 2024–10, 2024

  88. [96]

    A survey of mamba

    Haohao Qu, Liangbo Ning, Rui An, Wenqi Fan, Tyler Derr, Xin Xu, and Qing Li. A survey of mamba. arXiv preprint arXiv:2408.01129, 2024

  89. [97]

    Ssd4rec: a structured state space duality model for efficient sequential recommendation

    Haohao Qu, Yifeng Zhang, Liangbo Ning, Wenqi Fan, and Qing Li. Ssd4rec: a structured state space duality model for efficient sequential recommendation. arXiv preprint arXiv:2409.01192, 2024

  90. [98]

    Deciphering the protein landscape with protflash, a lightweight language model

    Lei Wang, Hui Zhang, Wei Xu, Zhidong Xue, and Yan Wang. Deciphering the protein landscape with protflash, a lightweight language model. Cell Reports Physical Science, 4(10), 2023

  91. [99]

    Distilprotbert: a distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts

    Yaron Geffen, Yanay Ofran, and Ron Unger. Distilprotbert: a distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts. Bioinformatics, 38(Supplement_2):ii95–ii98, 2022

  92. [100]

    Mixture of experts enable efficient and effective protein understanding and design

    Ning Sun, Shuxian Zou, Tianhua Tao, Sazan Mahbub, Dian Li, Yonghao Zhuang, Hongyi Wang, Xingyi Cheng, Le Song, and Eric P Xing. Mixture of experts enable efficient and effective protein understanding and design. bioRxiv, pages 2024–11, 2024

  93. [101]

    Toward ai-driven digital organ- ism: Multiscale foundation models for predicting, simulating and programming biology at all levels

    Le Song, Eran Segal, and Eric Xing. Toward ai-driven digital organ- ism: Multiscale foundation models for predicting, simulating and programming biology at all levels. arXiv preprint arXiv:2412.06993, 2024

  94. [102]

    Diffusion language models are versatile protein learners

    Xinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue, Shujian Huang, and Quanquan Gu. Diffusion language models are versatile protein learners. arXiv preprint arXiv:2402.18567, 2024

  95. [103]

    Generative diffusion models on graphs: methods and applications

    Chengyi Liu, Wenqi Fan, Yunqing Liu, Jiatong Li, Hang Li, Hui Liu, Jiliang Tang, and Qing Li. Generative diffusion models on graphs: methods and applications. In Proceedings of the Thirty- Second International Joint Conference on Artificial Intelligence , pages 6702–6711, 2023

  96. [104]

    Rita: a study on scaling up generative protein sequence models

    Daniel Hesslow, Niccoló Zanichelli, Pascal Notin, Iacopo Poli, and Debora Marks. Rita: a study on scaling up generative protein sequence models. arXiv preprint arXiv:2205.05789, 2022

  97. [105]

    Progen2: exploring the boundaries of protein language models

    Erik Nijkamp, Jeffrey A Ruffolo, Eli N Weinstein, Nikhil Naik, and Ali Madani. Progen2: exploring the boundaries of protein language models. Cell systems, 14(11):968–978, 2023

  98. [106]

    Ankh: Optimized protein language model unlocks general-purpose modelling

    Ahmed Elnaggar, Hazem Essam, Wafaa Salah-Eldin, Walid Moustafa, Mohamed Elkerdawy, Charlotte Rochereau, and Burkhard Rost. Ankh: Optimized protein language model unlocks general-purpose modelling. arXiv preprint arXiv:2301.06568, 2023

  99. [107]

    Glm: General language model pretraining with autoregressive blank infilling

    Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. Glm: General language model pretraining with autoregressive blank infilling. arXiv preprint arXiv:2103.10360, 2021

  100. [108]

    Deciphering antibody affinity maturation with language models and weakly supervised learning

    Jeffrey A Ruffolo, Jeffrey J Gray, and Jeremias Sulam. Deciphering antibody affinity maturation with language models and weakly supervised learning. arXiv preprint arXiv:2112.07782, 2021

  101. [109]

    Deciphering the language of antibodies using self-supervised learning

    Jinwoo Leem, Laura S Mitchell, James HR Farmery, Justin Barton, and Jacob D Galson. Deciphering the language of antibodies using self-supervised learning. Patterns, 3(7), 2022

  102. [110]

    Ablang: an antibody language model for completing antibody sequences

    Tobias H Olsen, Iain H Moal, and Charlotte M Deane. Ablang: an antibody language model for completing antibody sequences. Bioinformatics Advances, 2(1):vbac046, 2022

  103. [111]

    Pre-training antibody language models for antigen-specific compu- tational antibody design

    Kaiyuan Gao, Lijun Wu, Jinhua Zhu, Tianbo Peng, Yingce Xia, Liang He, Shufang Xie, Tao Qin, Haiguang Liu, Kun He, et al. Pre-training antibody language models for antigen-specific compu- tational antibody design. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Di...

  104. [112]

    Reprogram- ming pretrained language models for antibody sequence infilling

    Igor Melnyk, Vijil Chenthamarakshan, Pin-Yu Chen, Payel Das, Amit Dhurandhar, Inkit Padhi, and Devleena Das. Reprogram- ming pretrained language models for antibody sequence infilling. In International Conference on Machine Learning , pages 24398–24419. PMLR, 2023

  105. [113]

    Iglm: Infilling language modeling for antibody sequence design

    Richard W Shuai, Jeffrey A Ruffolo, and Jeffrey J Gray. Iglm: Infilling language modeling for antibody sequence design. Cell Systems, 14(11):979–989, 2023

  106. [114]

    p-iggen: A paired antibody generative language model

    Oliver Marcus Turnbull, Dino Oglic, Rebecca Croasdale-Wood, and Charlotte M Deane. p-iggen: A paired antibody generative language model. bioRxiv, pages 2024–08, 2024. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 29

  107. [115]

    Generative antibody design for complementary chain pairing sequences through encoder-decoder language model

    Simon KS Chu and Kathy Y Wei. Generative antibody design for complementary chain pairing sequences through encoder-decoder language model. arXiv preprint arXiv:2301.02748, 2023

  108. [116]

    Hhblits: lightning-fast iterative protein sequence searching by hmm-hmm alignment

    Michael Remmert, Andreas Biegert, Andreas Hauser, and Jo- hannes Söding. Hhblits: lightning-fast iterative protein sequence searching by hmm-hmm alignment. Nature methods, 9(2):173–175, 2012

  109. [117]

    Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets

    Martin Steinegger and Johannes Söding. Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nature biotechnology, 35(11):1026–1028, 2017

  110. [118]

    Few shot protein generation

    Soumya Ram and Tristan Bepler. Few shot protein generation. arXiv preprint arXiv:2204.01168, 2022

  111. [119]

    Enhancing the protein tertiary structure prediction by multiple sequence alignment generation

    Le Zhang, Jiayang Chen, Tao Shen, Yu Li, and Siqi Sun. Enhancing the protein tertiary structure prediction by multiple sequence alignment generation. arXiv preprint arXiv:2306.01824, 2023

  112. [120]

    Msagpt: Neural prompting protein structure prediction via msa generative pre-training

    Bo Chen, Zhilei Bei, Xingyi Cheng, Pan Li, Jie Tang, and Le Song. Msagpt: Neural prompting protein structure prediction via msa generative pre-training. arXiv preprint arXiv:2406.05347, 2024

  113. [121]

    Poet: A generative model of protein families as sequences-of-sequences

    Timothy Truong Jr and Tristan Bepler. Poet: A generative model of protein families as sequences-of-sequences. Advances in Neural Information Processing Systems, 36, 2024

  114. [122]

    Protmamba: a homology-aware but alignment-free protein state space model

    Damiano Sgarbossa, Cyril Malbranke, and Anne-Florence Bitbol. Protmamba: a homology-aware but alignment-free protein state space model. bioRxiv, pages 2024–05, 2024

  115. [123]

    Codon language embeddings provide strong signals for use in protein engineering

    Carlos Outeiral and Charlotte M Deane. Codon language embeddings provide strong signals for use in protein engineering. Nature Machine Intelligence, 6(2):170–179, 2024

  116. [124]

    cdsbert- extending protein language models with codon awareness

    Logan Hallee, Nikolaos Rafailidis, and Jason P Gleghorn. cdsbert- extending protein language models with codon awareness. bioRxiv, 2023

  117. [125]

    Ptm-mamba: A ptm-aware protein language model with bidirec- tional gated mamba blocks

    Zhangzhi Peng, Benjamin Schussheim, and Pranam Chatterjee. Ptm-mamba: A ptm-aware protein language model with bidirec- tional gated mamba blocks. bioRxiv, pages 2024–02, 2024

  118. [126]

    Mgnify: the microbiome sequence data analysis resource in 2023

    Lorna Richardson, Ben Allen, Germana Baldi, Martin Bera- cochea, Maxwell L Bileschi, Tony Burdett, Josephine Burgin, Juan Caballero-Pérez, Guy Cochrane, Lucy J Colwell, et al. Mgnify: the microbiome sequence data analysis resource in 2023. Nucleic Acids Research, 51(D1):D753–D...

  119. [127]

    The img/m data management and analysis system v

    I-Min A Chen, Ken Chu, Krishnaveni Palaniappan, Anna Ratner, Jinghua Huang, Marcel Huntemann, Patrick Hajek, Stephan J Ritter, Cody Webb, Dongying Wu, et al. The img/m data management and analysis system v. 7: content updates and new features. Nucleic acids research, 51(D1):D7...

  120. [128]

    Protein- level assembly increases protein sequence recovery from metage- nomic samples manyfold

    Martin Steinegger, Milot Mirdita, and Johannes Söding. Protein- level assembly increases protein sequence recovery from metage- nomic samples manyfold. Nature methods, 16(7):603–606, 2019

  121. [129]

    Albert: A lite bert for self-supervised learning of language representations

    Z Lan. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942, 2019

  122. [130]

    Electra: Pre-training text encoders as discriminators rather than generators

    K Clark. Electra: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555, 2020

  123. [131]

    Mixtral of experts

    Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024

  124. [132]

    Transformer-xl: Attentive lan- guage models beyond a fixed-length context

    Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive lan- guage models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019

  125. [133]

    Xlnet: Generalized autoregressive pretraining for language understanding

    Zhilin Yang. Xlnet: Generalized autoregressive pretraining for language understanding. arXiv preprint arXiv:1906.08237, 2019

  126. [134]

    Observed antibody space: a resource for data mining next-generation sequencing of antibody repertoires

    Aleksandr Kovaltsuk, Jinwoo Leem, Sebastian Kelm, James Snowden, Charlotte M Deane, and Konrad Krawczyk. Observed antibody space: a resource for data mining next-generation sequencing of antibody repertoires. The Journal of Immunology , 201(8):2502–2509, 2018

  127. [135]

    Sabdab: the structural antibody database

    James Dunbar, Konrad Krawczyk, Jinwoo Leem, Terry Baker, Angelika Fuchs, Guy Georges, Jiye Shi, and Charlotte M Deane. Sabdab: the structural antibody database. Nucleic acids research, 42(D1):D1140–D1146, 2014

  128. [136]

    Pfam: The protein families database in 2021

    Jaina Mistry, Sara Chuguransky, Lowri Williams, Matloob Qureshi, Gustavo A Salazar, Erik LL Sonnhammer, Silvio CE Tosatto, Lisanna Paladin, Shriya Raj, Lorna J Richardson, et al. Pfam: The protein families database in 2021. Nucleic acids research , 49(D1):D412–D419, 2021

  129. [137]

    Uniclust databases of clustered and deeply annotated protein sequences and alignments

    Milot Mirdita, Lars Von Den Driesch, Clovis Galiez, Maria J Martin, Johannes Söding, and Martin Steinegger. Uniclust databases of clustered and deeply annotated protein sequences and alignments. Nucleic acids research, 45(D1):D170–D176, 2017

  130. [138]

    Openproteinset: Training data for structural biology at scale

    Gustaf Ahdritz, Nazim Bouatta, Sachin Kadyan, Lukas Jarosch, Dan Berenberg, Ian Fisk, Andrew Watkins, Stephen Ra, Richard Bonneau, and Mohammed AlQuraishi. Openproteinset: Training data for structural biology at scale. Advances in Neural Information Processing Systems, 36, 2024

  131. [139]

    Mamba: Linear-time sequence modeling with selective state spaces

    Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023

  132. [140]

    Consensus coding sequence (ccds) database: a standardized set of human and mouse protein-coding regions supported by expert curation

    Shashikant Pujar, Nuala A O’Leary, Catherine M Farrell, Jane E Loveland, Jonathan M Mudge, Craig Wallin, Carlos G Girón, Mark Diekhans, If Barnes, Ruth Bennett, et al. Consensus coding sequence (ccds) database: a standardized set of human and mouse protein-coding regions suppo...

  133. [141]

    Uniprotkb/swiss-prot: the man- ually annotated section of the uniprot knowledgebase

    Emmanuel Boutet, Damien Lieberherr, Michael Tognolli, Michel Schneider, and Amos Bairoch. Uniprotkb/swiss-prot: the man- ually annotated section of the uniprot knowledgebase. In Plant bioinformatics: methods and protocols, pages 89–112. Springer, 2007

  134. [142]

    Petribert: Augmenting bert with tridimensional encoding for inverse protein folding and design

    Baldwin Dumortier, Antoine Liutkus, Clément Carré, and Gabriel Krouk. Petribert: Augmenting bert with tridimensional encoding for inverse protein folding and design. BioRxiv, pages 2022–08, 2022

  135. [143]

    Structure-informed protein language model

    Zuobai Zhang, Jiarui Lu, Vijil Chenthamarakshan, Aurélie Lozano, Payel Das, and Jian Tang. Structure-informed protein language model. arXiv preprint arXiv:2402.05856, 2024

  136. [144]

    Biophysics-based protein language models for protein engineering

    Sam Gelman, Bryce Johnson, Chase Freschlin, Sameer D’Costa, Anthony Gitter, and Philip A Romero. Biophysics-based protein language models for protein engineering. bioRxiv, pages 2024–03, 2024

  137. [145]

    The rosetta all-atom energy function for macromolecular modeling and design

    Rebecca F Alford, Andrew Leaver-Fay, Jeliazko R Jeliazkov, Matthew J O’Meara, Frank P DiMaio, Hahnbeom Park, Maxim V Shapovalov, P Douglas Renfrew, Vikram K Mulligan, Kalli Kappel, et al. The rosetta all-atom energy function for macromolecular modeling and design. Journal of c...

  138. [146]

    Transformer protein language models are unsupervised structure learners

    Roshan Rao, Joshua Meier, Tom Sercu, Sergey Ovchinnikov, and Alexander Rives. Transformer protein language models are unsupervised structure learners. Biorxiv, pages 2020–12, 2020

  139. [147]

    Endowing protein language models with structural knowledge

    Dexiong Chen, Philip Hartout, Paolo Pellizzoni, Carlos Oliver, and Karsten Borgwardt. Endowing protein language models with structural knowledge. arXiv preprint arXiv:2401.14819, 2024

  140. [148]

    Structure-informed language models are protein designers

    Zaixiang Zheng, Yifan Deng, Dongyu Xue, Yi Zhou, Fei Ye, and Quanquan Gu. Structure-informed language models are protein designers. In International Conference on Machine Learning, pages 42317–42338. PMLR, 2023

  141. [149]

    Adapting protein language models for structure- conditioned design

    Jeffrey A Ruffolo, Aadyot Bhatnagar, Joel Beazer, Stephen Nayfach, Jordan Russ, Emily Hill, Riffat Hussain, Joseph Gallagher, and Ali Madani. Adapting protein language models for structure- conditioned design. bioRxiv, pages 2024–08, 2024

  142. [150]

    Neural discrete representation learning

    Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017

  143. [151]

    Fast and accurate protein structure search with foldseek

    Michel Van Kempen, Stephanie S Kim, Charlotte Tumescheit, Milot Mirdita, Jeongjae Lee, Cameron LM Gilchrist, Johannes Söding, and Martin Steinegger. Fast and accurate protein structure search with foldseek. Nature Biotechnology, 42(2):243–246, 2024

  144. [152]

    Pro- tokens: A machine-learned language for compact and informative encoding of protein 3d structures

    Xiaohan Lin, Zhenyu Chen, Yanheng Li, Xingyu Lu, Chuanliu Fan, Ziqiang Cao, Shihao Feng, Yi Qin Gao, and Jun Zhang. Pro- tokens: A machine-learned language for compact and informative encoding of protein 3d structures. bioRxiv, pages 2023–11, 2023

  145. [153]

    Foldtoken4: Consistent & hierarchical fold language

    Zhangyang Gao, Cheng Tan, and Stan Z Li. Foldtoken4: Consistent & hierarchical fold language. bioRxiv, pages 2024–08, 2024

  146. [154]

    Balancing locality and reconstruction in protein structure tokenizer

    Barthelemy Meynard-Piganeau, Jiayou Zhang, James Gong, Xingyi Cheng, Yingtao Luo, Hugo Ly, Le Song, and Eric P Xing. Balancing locality and reconstruction in protein structure tokenizer. bioRxiv, pages 2024–12, 2024

  147. [155]

    Prostt5: Bilingual language model for protein sequence and structure

    Michael Heinzinger, Konstantin Weissenow, Joaquin Gomez Sanchez, Adrian Henkel, Martin Steinegger, and Burkhard Rost. Prostt5: Bilingual language model for protein sequence and structure. bioRxiv, pages 2023–07, 2023

  148. [156]

    Saprothub: Making protein modeling accessible to all biologists

    Jin Su, Zhikai Li, Chenchen Han, Yuyang Zhou, Yan He, Junjie Shan, Xibin Zhou, Xing Chang, Dacheng Ma, OPMC, et al. Saprothub: Making protein modeling accessible to all biologists. bioRxiv, pages 2024–05, 2024

  149. [157]

    Deprot: A protein language model with quantizied structure and disentangled attention

    Mingchen Li, Yang Tan, Bozitao Zhong, Ziyi Zhou, Huiqun Yu, Xinzhu Ma, Wanli Ouyang, Liang Hong, Bingxin Zhou, and Pan Tan. Deprot: A protein language model with quantizied structure and disentangled attention. bioRxiv, pages 2024–04, 2024. JOURNAL OF LATEX CLASS FILES, VOL. 1...

  150. [158]

    Evaluating protein transfer learning with tape

    Roshan Rao, Nicholas Bhattacharya, Neil Thomas, Yan Duan, Peter Chen, John Canny, Pieter Abbeel, and Yun Song. Evaluating protein transfer learning with tape. Advances in neural information processing systems, 32, 2019

  151. [159]

    Peer: a comprehensive and multi-task benchmark for protein sequence understanding

    Minghao Xu, Zuobai Zhang, Jiarui Lu, Zhaocheng Zhu, Yangtian Zhang, Ma Chang, Runcheng Liu, and Jian Tang. Peer: a comprehensive and multi-task benchmark for protein sequence understanding. Advances in Neural Information Processing Systems , 35:35156–35173, 2022

  152. [160]

    Proteingym: Large-scale benchmarks for protein fitness prediction and design

    Pascal Notin, Aaron Kollasch, Daniel Ritter, Lood Van Niekerk, Steffanie Paul, Han Spinner, Nathan Rollins, Ada Shaw, Rose Orenbuch, Ruben Weitzman, et al. Proteingym: Large-scale benchmarks for protein fitness prediction and design. Advances in Neural Information Processing S...

  153. [161]

    Dplm-2: A multimodal diffusion protein language model

    Xinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue, Shujian Huang, and Quanquan Gu. Dplm-2: A multimodal diffusion protein language model. arXiv preprint arXiv:2410.13782, 2024

  154. [162]

    Proteinbert: a universal deep-learning model of protein sequence and function

    Nadav Brandes, Dan Ofer, Yam Peleg, Nadav Rappoport, and Michal Linial. Proteinbert: a universal deep-learning model of protein sequence and function. Bioinformatics, 38(8):2102–2110, 2022

  155. [163]

    Multi-level protein structure pre-training via prompt learning

    Zeyuan Wang, Qiang Zhang, HU Shuang-Wei, Haoran Yu, Xurui Jin, Zhichen Gong, and Huajun Chen. Multi-level protein structure pre-training via prompt learning. In The Eleventh International Conference on Learning Representations, 2022

  156. [164]

    Zymctrl: a conditional language model for the controllable generation of artificial enzymes

    Geraldene Munsamy, Sebastian Lindner, Philipp Lorenz, and Noelia Ferruz. Zymctrl: a conditional language model for the controllable generation of artificial enzymes. In NeurIPS Machine Learning in Structural Biology Workshop, 2022

  157. [165]

    Regression transformer enables concurrent sequence regression and generation for molecular language modelling

    Jannis Born and Matteo Manica. Regression transformer enables concurrent sequence regression and generation for molecular language modelling. Nature Machine Intelligence , 5(4):432–444, 2023

  158. [166]

    A text-guided protein design framework

    Shengchao Liu, Yutao Zhu, Jiarui Lu, Zhao Xu, Weili Nie, Anthony Gitter, Chaowei Xiao, Jian Tang, Hongyu Guo, and Anima Anandkumar. A text-guided protein design framework. arXiv preprint arXiv:2302.04611, 2023

  159. [167]

    Protst: Multi-modality learning of protein sequences and biomedical texts

    Minghao Xu, Xinyu Yuan, Santiago Miret, and Jian Tang. Protst: Multi-modality learning of protein sequences and biomedical texts. In International Conference on Machine Learning , pages 38749–38767. PMLR, 2023

  160. [168]

    Multi-modal clip-informed protein editing

    Mingze Yin, Hanjing Zhou, Yiheng Zhu, Miao Lin, Yixuan Wu, Jialu Wu, Hongxia Xu, Chang-Yu Hsieh, Tingjun Hou, Jintai Chen, et al. Multi-modal clip-informed protein editing. bioRxiv, pages 2024–07, 2024

  161. [169]

    Proteinclip: en- hancing protein language models with natural language

    Kevin E Wu, Howard Chang, and James Zou. Proteinclip: en- hancing protein language models with natural language. bioRxiv, pages 2024–05, 2024

  162. [170]

    Protrek: Navigating the protein universe through tri-modal contrastive learning

    Jin Su, Xibin Zhou, Xuting Zhang, and Fajie Yuan. Protrek: Navigating the protein universe through tri-modal contrastive learning. bioRxiv, pages 2024–05, 2024

  163. [171]

    Ontoprotein: Protein pretraining with gene ontology embedding

    Ningyu Zhang, Zhen Bi, Xiaozhuan Liang, Siyuan Cheng, Haosen Hong, Shumin Deng, Jiazhang Lian, Qiang Zhang, and Huajun Chen. Ontoprotein: Protein pretraining with gene ontology embedding. arXiv preprint arXiv:2201.11147, 2022

  164. [172]

    Protein representation learning via knowledge enhanced primary structure reasoning

    Hong-Yu Zhou, Yunxiang Fu, Zhicheng Zhang, Bian Cheng, and Yizhou Yu. Protein representation learning via knowledge enhanced primary structure reasoning. InThe Eleventh International Conference on Learning Representations, 2022

  165. [173]

    Benchmarking text-integrated protein language model embeddings and embed- ding fusion on diverse downstream tasks

    Young Su Ko, Jonathan Parkinson, and Wei Wang. Benchmarking text-integrated protein language model embeddings and embed- ding fusion on diverse downstream tasks. bioRxiv, pages 2024–08, 2024

  166. [174]

    Ssemb: A joint embedding of protein sequence and structure enables robust variant effect predictions

    Lasse M Blaabjerg, Nicolas Jonsson, Wouter Boomsma, Amelie Stein, and Kresten Lindorff-Larsen. Ssemb: A joint embedding of protein sequence and structure enables robust variant effect predictions. Nature Communications, 15(1):9646, 2024

  167. [175]

    Pre- training sequence, structure, and surface features for comprehen- sive protein representation learning

    Youhan Lee, Hasun Yu, Jaemyung Lee, and Jaehoon Kim. Pre- training sequence, structure, and surface features for comprehen- sive protein representation learning. In The Twelfth International Conference on Learning Representations, 2023

  168. [176]

    Contrasting sequence with structure: Pre-training graph representations with plms

    Louis Robinson, Timothy Atkinson, Liviu Copoiu, Patrick Bordes, Thomas Pierrot, and Thomas D Barrett. Contrasting sequence with structure: Pre-training graph representations with plms. bioRxiv, pages 2023–12, 2023

  169. [177]

    A multimodal protein representation framework for quantifying transferability across biochemical downstream tasks

    Fan Hu, Yishen Hu, Weihong Zhang, Huazhen Huang, Yi Pan, and Peng Yin. A multimodal protein representation framework for quantifying transferability across biochemical downstream tasks. Advanced Science, 10(22):2301223, 2023

  170. [178]

    Learning sequence, structure, and function representations of proteins with language models

    Tymor Hamamsy, Meet Barot, James T Morton, Martin Steinegger, Richard Bonneau, and Kyunghyun Cho. Learning sequence, structure, and function representations of proteins with language models. bioRxiv, 2023

  171. [179]

    Alphafold protein structure database: massively expanding the structural coverage of protein-sequence space with high-accuracy models

    Mihaly Varadi, Stephen Anyango, Mandar Deshpande, Sreenath Nair, Cindy Natassia, Galabina Yordanova, David Yuan, Oana Stroe, Gemma Wood, Agata Laydon, et al. Alphafold protein structure database: massively expanding the structural coverage of protein-sequence space with high-a...

  172. [180]

    Deepsf: deep convolutional neural network for mapping protein sequences to folds

    Jie Hou, Badri Adhikari, and Jianlin Cheng. Deepsf: deep convolutional neural network for mapping protein sequences to folds. Bioinformatics, 34(8):1295–1303, 2018

  173. [181]

    Cath–a hierarchic classification of protein domain structures

    Christine A Orengo, Alex D Michie, Susan Jones, David T Jones, Mark B Swindells, and Janet M Thornton. Cath–a hierarchic classification of protein domain structures. Structure, 5(8):1093– 1109, 1997

  174. [182]

    Interpro in

    Typhaine Paysan-Lafosse, Matthias Blum, Sara Chuguransky, Tiago Grego, Beatriz Lázaro Pinto, Gustavo A Salazar, Maxwell L Bileschi, Peer Bork, Alan Bridge, Lucy Colwell, et al. Interpro in

  175. [183]

    Interproscan 5: genome-scale protein function classification

    Philip Jones, David Binns, Hsin-Yu Chang, Matthew Fraser, Weizhong Li, Craig McAnulla, Hamish McWilliam, John Maslen, Alex Mitchell, Gift Nuka, et al. Interproscan 5: genome-scale protein function classification. Bioinformatics, 30(9):1236–1240, 2014

  176. [184]

    String v11: protein–protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets

    Damian Szklarczyk, Annika L Gable, David Lyon, Alexander Junge, Stefan Wyder, Jaime Huerta-Cepas, Milan Simonovic, Nadezhda T Doncheva, John H Morris, Peer Bork, et al. String v11: protein–protein association networks with increased coverage, supporting functional discovery in...

  177. [185]

    The ncbi taxonomy database

    Scott Federhen. The ncbi taxonomy database. Nucleic acids research, 40(D1):D136–D143, 2012

  178. [186]

    Scibert: A pretrained language model for scientific text

    Iz Beltagy, Kyle Lo, and Arman Cohan. Scibert: A pretrained language model for scientific text. arXiv preprint arXiv:1903.10676, 2019

  179. [187]

    Domain-specific language model pretraining for biomedical natural language processing

    Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH), 3(1):1–23, 2021

  180. [188]

    Instructprotein: Aligning human and protein language via knowledge instruction

    Zeyuan Wang, Qiang Zhang, Keyan Ding, Ming Qin, Xiang Zhuang, Xiaotong Li, and Huajun Chen. Instructprotein: Aligning human and protein language via knowledge instruction. arXiv preprint arXiv:2310.03269, 2023

  181. [189]

    Protllm: An interleaved protein-language llm with protein-as-word pre- training

    Le Zhuo, Zewen Chi, Minghao Xu, Heyan Huang, Heqi Zheng, Conghui He, Xian-Ling Mao, and Wentao Zhang. Protllm: An interleaved protein-language llm with protein-as-word pre- training. arXiv preprint arXiv:2403.07920, 2024

  182. [190]

    Decoding the molecular language of proteins with evolla

    Xibin Zhou, Chenchen Han, Yingqi Zhang, Jin Su, Kai Zhuang, Shiyu Jiang, Zichen Yuan, Wei Zheng, Fengyuan Dai, Yuyang Zhou, et al. Decoding the molecular language of proteins with evolla. bioRxiv, pages 2025–01, 2025

  183. [191]

    Druggpt: A gpt-based strategy for designing potential ligands targeting specific proteins

    Yuesen Li, Chengyi Gao, Xin Song, Xiangyu Wang, Yungang Xu, and Suxia Han. Druggpt: A gpt-based strategy for designing potential ligands targeting specific proteins. bioRxiv, pages 2023– 06, 2023

  184. [192]

    Multi- scale protein language model for unified molecular modeling

    Kangjie Zheng, Siyu Long, Tianyu Lu, Junwei Yang, Xinyu Dai, Ming Zhang, Zaiqing Nie, Wei-Ying Ma, and Hao Zhou. Multi- scale protein language model for unified molecular modeling. bioRxiv, pages 2024–03, 2024

  185. [193]

    Pubmed: the bibliographic database

    Kathi Canese and Sarah Weis. Pubmed: the bibliographic database. The NCBI handbook, 2(1), 2013

  186. [194]

    Mol- instructions: A large-scale biomolecular instruction dataset for large language models

    Yin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu, Rui Huang, Zhuo Chen, Xiaohui Fan, and Huajun Chen. Mol- instructions: A large-scale biomolecular instruction dataset for large language models. arXiv preprint arXiv:2306.08018, 2023

  187. [195]

    The llama 3 herd of models

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024

  188. [196]

    Zinc20—a free ultralarge- scale chemical database for ligand discovery

    John J Irwin, Khanh G Tang, Jennifer Young, Chinzorig Dandarchu- luun, Benjamin R Wong, Munkhzul Khurelbaatar, Yurii S Moroz, John Mayfield, and Roger A Sayle. Zinc20—a free ultralarge- scale chemical database for ligand discovery. Journal of chemical information and modeling,...

  189. [197]

    Uni- mol: A universal 3d molecular representation learning framework

    Gengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng, Hongteng Xu, Zhewei Wei, Linfeng Zhang, and Guolin Ke. Uni- mol: A universal 3d molecular representation learning framework. 2023

  190. [198]

    Biomedgpt: Open multimodal generative pre-trained transformer for biomedicine

    Yizhen Luo, Jiahuan Zhang, Siqi Fan, Kai Yang, Yushuai Wu, Mu Qiao, and Zaiqing Nie. Biomedgpt: Open multimodal generative pre-trained transformer for biomedicine. arXiv preprint arXiv:2308.09442, 2023

  191. [199]

    Pubchem 2023 update

    Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, et al. Pubchem 2023 update. Nucleic acids research, 51(D1):D1373–D1380, 2023

  192. [200]

    Pre-training molecular graph representation with 3d geometry

    Shengchao Liu, Hanchen Wang, Weiyang Liu, Joan Lasenby, Hongyu Guo, and Jian Tang. Pre-training molecular graph representation with 3d geometry. arXiv preprint arXiv:2110.07728, 2021

  193. [201]

    Biot5+: Towards generalized biological understanding with iupac integration and multi-task tuning

    Qizhi Pei, Lijun Wu, Kaiyuan Gao, Xiaozhuan Liang, Yin Fang, Jinhua Zhu, Shufang Xie, Tao Qin, and Rui Yan. Biot5+: Towards generalized biological understanding with iupac integration and multi-task tuning. arXiv preprint arXiv:2402.17810, 2024

  194. [202]

    Instructbiomol: Advancing biomolecule understand- ing and design following human instructions

    Xiang Zhuang, Keyan Ding, Tianwen Lyu, Yinuo Jiang, Xiaotong Li, Zhuoyi Xiang, Zeyuan Wang, Ming Qin, Kehua Feng, Jike Wang, et al. Instructbiomol: Advancing biomolecule understand- ing and design following human instructions. arXiv preprint arXiv:2410.07919, 2024

  195. [203]

    Translation between molecules and natural language

    Carl Edwards, Tuan Lai, Kevin Ros, Garrett Honke, Kyunghyun Cho, and Heng Ji. Translation between molecules and natural language. arXiv preprint arXiv:2204.11817, 2022

  196. [204]

    Bindingdb in 2015: a public database for medicinal chemistry, computational chemistry and systems pharmacology

    Michael K Gilson, Tiqing Liu, Michael Baitaluk, George Nicola, Linda Hwang, and Jenny Chong. Bindingdb in 2015: a public database for medicinal chemistry, computational chemistry and systems pharmacology. Nucleic acids research, 44(D1):D1045–D1053, 2016

  197. [205]

    Rhea, the reaction knowledgebase in 2022

    Parit Bansal, Anne Morgat, Kristian B Axelsen, Venkatesh Muthukrishnan, Elisabeth Coudert, Lucila Aimo, Nevila Hyka- Nouspikel, Elisabeth Gasteiger, Arnaud Kerhornou, Teresa Batista Neto, et al. Rhea, the reaction knowledgebase in 2022. Nucleic acids research, 50(D1):D693–D700, 2022

  198. [206]

    Multilingual translation for zero-shot biomedical classification using biotranslator

    Hanwen Xu, Addie Woicik, Hoifung Poon, Russ B Altman, and Sheng Wang. Multilingual translation for zero-shot biomedical classification using biotranslator. Nature Communications, 14(1):738, 2023

  199. [207]

    The human phenotype ontology in 2021

    Sebastian Köhler, Michael Gargano, Nicolas Matentzoglu, Leigh C Carmody, David Lewis-Smith, Nicole A Vasilevsky, Daniel Danis, Ganna Balagura, Gareth Baynam, Amy M Brower, et al. The human phenotype ontology in 2021. Nucleic acids research , 49(D1):D1207–D1217, 2021

  200. [208]

    Galactica: A large language model for science

    Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. Galactica: A large language model for science. arXiv preprint arXiv:2211.09085, 2022

  201. [209]

    Bio- bridge: Bridging biomedical foundation models via knowledge graph

    Zifeng Wang, Zichen Wang, Balasubramaniam Srinivasan, Vas- silis N Ioannidis, Huzefa Rangwala, and Rishita Anubhai. Bio- bridge: Bridging biomedical foundation models via knowledge graph. arXiv preprint arXiv:2310.03320, 2023

  202. [210]

    Recommendations of the wwpdb nmr validation task force

    Gaetano T Montelione, Michael Nilges, Ad Bax, Peter Güntert, Torsten Herrmann, Jane S Richardson, Charles D Schwieters, Wim F Vranken, Geerten W Vuister, David S Wishart, et al. Recommendations of the wwpdb nmr validation task force. Structure, 21(9):1563–1570, 2013

  203. [211]

    Protein structure prediction beyond alphafold

    Guo-Wei Wei. Protein structure prediction beyond alphafold. Nature Machine Intelligence, 1(8):336–337, 2019

  204. [212]

    High-resolution de novo structure prediction from primary sequence

    Ruidong Wu, Fan Ding, Rui Wang, Rui Shen, Xiwen Zhang, Shitong Luo, Chenpeng Su, Zuofan Wu, Qi Xie, Bonnie Berger, et al. High-resolution de novo structure prediction from primary sequence. BioRxiv, pages 2022–07, 2022

  205. [213]

    Single- sequence protein structure prediction using a language model and deep learning

    Ratul Chowdhury, Nazim Bouatta, Surojit Biswas, Christina Floristean, Anant Kharkar, Koushik Roy, Charlotte Rochereau, Gustaf Ahdritz, Joanna Zhang, George M Church, et al. Single- sequence protein structure prediction using a language model and deep learning. Nature Biotechno...

  206. [214]

    Fast, accurate antibody structure prediction from deep learn- ing on massive set of natural antibodies

    Jeffrey A Ruffolo, Lee-Shin Chu, Sai Pooja Mahajan, and Jeffrey J Gray. Fast, accurate antibody structure prediction from deep learn- ing on massive set of natural antibodies. Nature communications, 14(1):2389, 2023

  207. [215]

    Accurate structure prediction of biomolecular interactions with alphafold 3

    Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans, Tim Green, Alexander Pritzel, Olaf Ronneberger, Lindsay Willmore, Andrew J Ballard, Joshua Bambrick, et al. Accurate structure prediction of biomolecular interactions with alphafold 3. Nature, pages 1–3, 2024

  208. [216]

    Nucleic acids research, 51(D1):D523–D531, 2023

    Uniprot: the universal protein knowledgebase in 2023. Nucleic acids research, 51(D1):D523–D531, 2023

  209. [217]

    Fast and accurate protein intrinsic disorder prediction by using a pretrained language model

    Yidong Song, Qianmu Yuan, Sheng Chen, Ken Chen, Yaoqi Zhou, and Yuedong Yang. Fast and accurate protein intrinsic disorder prediction by using a pretrained language model. Briefings in bioinformatics, 24(4):bbad173, 2023

  210. [218]

    Protein stability prediction by fine-tuning a protein language model on a mega-scale dataset

    Kit Sang Chu and Justin B Siegel. Protein stability prediction by fine-tuning a protein language model on a mega-scale dataset. bioRxiv, pages 2023–11, 2023

  211. [219]

    Vish-pred: an ensemble of fine- tuned esm models for protein toxicity prediction

    Raghvendra Mall, Ankita Singh, Chirag N Patel, Gregory Guir- imand, and Filippo Castiglione. Vish-pred: an ensemble of fine- tuned esm models for protein toxicity prediction. Briefings in Bioinformatics, 25(4), 2024

  212. [220]

    Com- prehensive prediction and analysis of human protein essentiality based on a pretrained large language model

    Boming Kang, Rui Fan, Chunmei Cui, and Qinghua Cui. Com- prehensive prediction and analysis of human protein essentiality based on a pretrained large language model. Nature Computational Science, pages 1–11, 2024

  213. [221]

    Peptidebert: A language model based on transformers for peptide property prediction

    Chakradhar Guntuboina, Adrita Das, Parisa Mollaei, Seongwon Kim, and Amir Barati Farimani. Peptidebert: A language model based on transformers for peptide property prediction. The Journal of Physical Chemistry Letters, 14(46):10427–10434, 2023

  214. [222]

    Lassoesm: A tailored language model for enhanced lasso peptide property prediction

    Xuenan Mi, Susanna E Barrett, Douglas A Mitchell, and Diwakar Shukla. Lassoesm: A tailored language model for enhanced lasso peptide property prediction. bioRxiv, pages 2024–10, 2024

  215. [223]

    Approaching optimal ph enzyme prediction with large language models

    Mark Zaretckii, Pavel Buslaev, Igor Kozlovskii, Alexander Moro- zov, and Petr Popov. Approaching optimal ph enzyme prediction with large language models. ACS Synthetic Biology , 13(9):3013– 3021, 2024

  216. [224]

    Deep learning prediction of enzyme optimum ph

    Japheth E Gado, Matthew Knotts, Ada Y Shaw, Debora Marks, Nicholas P Gauthier, Chris Sander, and Gregg T Beckham. Deep learning prediction of enzyme optimum ph. bioRxiv, pages 2023– 06, 2023

  217. [225]

    Genome-wide prediction of disease variant effects with a deep protein language model

    Nadav Brandes, Grant Goldman, Charlotte H Wang, Chun Jimmie Ye, and Vasilis Ntranos. Genome-wide prediction of disease variant effects with a deep protein language model. Nature Genetics, 55(9):1512–1522, 2023

  218. [226]

    Enhancing efficiency of protein language models with minimal wet-lab data through few-shot learning

    Ziyi Zhou, Liang Zhang, Yuanxi Yu, Banghao Wu, Mingchen Li, Liang Hong, and Pan Tan. Enhancing efficiency of protein language models with minimal wet-lab data through few-shot learning. Nature Communications, 15(1):5566, 2024

  219. [227]

    Deplm: Denoising protein language models for property optimization

    Zeyuan Wang, Keyan Ding, Ming Qin, Xiaotong Li, Xiang Zhuang, Yu Zhao, Jianhua Yao, Qiang Zhang, and Huajun Chen. Deplm: Denoising protein language models for property optimization. In The Thirty-eighth Annual Conference on Neural Information Processing Systems

  220. [228]

    Large language models improve annotation of prokaryotic viral proteins

    Zachary N Flamholz, Steven J Biller, and Libusha Kelly. Large language models improve annotation of prokaryotic viral proteins. Nature Microbiology, pages 1–13, 2024

  221. [229]

    Accurately predicting enzyme functions through geometric graph learning on esmfold-predicted structures

    Yidong Song, Qianmu Yuan, Sheng Chen, Yuansong Zeng, Huiy- ing Zhao, and Yuedong Yang. Accurately predicting enzyme functions through geometric graph learning on esmfold-predicted structures. Nature Communications, 15(1):8180, 2024

  222. [230]

    Netgo 3.0: protein language model improves large-scale functional annotations

    Shaojun Wang, Ronghui You, Yunjia Liu, Yi Xiong, and Shanfeng Zhu. Netgo 3.0: protein language model improves large-scale functional annotations. Genomics, Proteomics & Bioinformatics , 21(2):349–358, 2023

  223. [231]

    Protein function prediction as approximate semantic entailment

    Maxat Kulmanov, Francisco J Guzmán-Vega, Paula Duek Roggli, Lydie Lane, Stefan T Arold, and Robert Hoehndorf. Protein function prediction as approximate semantic entailment. Nature Machine Intelligence, pages 1–9, 2024

  224. [232]

    Fast and accurate protein function prediction from sequence through pretrained language model and homology- based label diffusion

    Qianmu Yuan, Junjie Xie, Jiancong Xie, Huiying Zhao, and Yuedong Yang. Fast and accurate protein function prediction from sequence through pretrained language model and homology- based label diffusion. Briefings in bioinformatics , 24(3):bbad117, 2023

  225. [233]

    Identifying b- cell epitopes using alphafold2 predicted structures and pretrained language model

    Yuansong Zeng, Zhuoyi Wei, Qianmu Yuan, Sheng Chen, Weijiang Yu, Yutong Lu, Jianzhao Gao, and Yuedong Yang. Identifying b- cell epitopes using alphafold2 predicted structures and pretrained language model. Bioinformatics, 39(4):btad187, 2023

  226. [234]

    Signalp 6.0 predicts all five types of signal peptides using protein language models

    Felix Teufel, José Juan Almagro Armenteros, Alexander Rosenberg Johansen, Magnús Halldór Gíslason, Silas Irby Pihl, Konstanti- nos D Tsirigos, Ole Winther, Søren Brunak, Gunnar von Heijne, and Henrik Nielsen. Signalp 6.0 predicts all five types of signal peptides using protein...

  227. [235]

    Unbiased organism-agnostic and highly sensitive signal peptide predictor with deep protein language model.Nature Computational Science, 4(1):29–42, 2024

    Junbo Shen, Qinze Yu, Shenyang Chen, Qingxiong Tan, Jingchen Li, and Yu Li. Unbiased organism-agnostic and highly sensitive signal peptide predictor with deep protein language model.Nature Computational Science, 4(1):29–42, 2024

  228. [236]

    Protein language models are performant in structure-free virtual screening

    Hilbert Yuen In Lam, Jia Sheng Guan, Xing Er Ong, Robbe Pincket, and Yuguang Mu. Protein language models are performant in structure-free virtual screening. Briefings in Bioinformatics , 25(6):bbae480, 2024

  229. [237]

    Smiles transformer: Pre-trained molecular fingerprint for low data drug discovery

    Shion Honda, Shoi Shi, and Hiroki R Ueda. Smiles transformer: Pre-trained molecular fingerprint for low data drug discovery. arXiv preprint arXiv:1911.04738, 2019

  230. [238]

    Learning binding affinities via fine-tuning of protein and ligand language models

    Rohan Gorantla, Aryo Pradipta Gema, Ian Xi Yang, Álvaro Serrano-Morrás, Benjamin Suutari, Jordi Juárez Jiménez, and Antonia SJS Mey. Learning binding affinities via fine-tuning of protein and ligand language models. bioRxiv, pages 2024–11, 2024

  231. [239]

    Chemberta-2: Towards chemical founda- tion models

    Walid Ahmad, Elana Simon, Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. Chemberta-2: Towards chemical founda- tion models. arXiv preprint arXiv:2209.01712, 2022

  232. [240]

    Reactzyme: A benchmark for enzyme-reaction prediction

    Chenqing Hua, Bozitao Zhong, Sitao Luan, Liang Hong, Guy Wolf, Doina Precup, and Shuangjia Zheng. Reactzyme: A benchmark for enzyme-reaction prediction. arXiv preprint arXiv:2408.13659, 2024

  233. [241]

    Multi-modal deep learning enables efficient and accurate annotation of enzymatic active sites

    Xiaorui Wang, Xiaodan Yin, Dejun Jiang, Huifeng Zhao, Zhenxing Wu, Odin Zhang, Jike Wang, Yuquan Li, Yafeng Deng, Huanxiang Liu, et al. Multi-modal deep learning enables efficient and accurate annotation of enzymatic active sites. Nature Communications , 15(1):7348, 2024

  234. [242]

    Enhancing protein language model with structure-based encoder and pre-training

    Zuobai Zhang, Minghao Xu, Aurelie Lozano, Vijil Chenthama- rakshan, Payel Das, and Jian Tang. Enhancing protein language model with structure-based encoder and pre-training. In ICLR 2023-Machine Learning for Drug Discovery workshop , 2023

  235. [243]

    Salt&peppr is an interface-predicting language model for designing peptide-guided protein degraders

    Garyk Brixi, Tianzheng Ye, Lauren Hong, Tian Wang, Connor Mon- ticello, Natalia Lopez-Barbosa, Sophia Vincoff, Vivian Yudistyra, Lin Zhao, Elena Haarer, et al. Salt&peppr is an interface-predicting language model for designing peptide-guided protein degraders. Communications B...

  236. [244]

    Democratizing protein language models with parameter-efficient fine-tuning

    Samuel Sledzieski, Meghana Kshirsagar, Minkyung Baek, Rahul Dodhia, Juan Lavista Ferres, and Bonnie Berger. Democratizing protein language models with parameter-efficient fine-tuning. Proceedings of the National Academy of Sciences , 121(26):e2405840121, 2024

  237. [245]

    Prollm: Protein chain-of-thoughts enhanced llm for protein- protein interaction prediction

    Mingyu Jin, Xue Haochen, Zhenting Wang, Boming Kang, Ru- osong Ye, Kaixiong Zhou, Mengnan Du, and Yongfeng Zhang. Prollm: Protein chain-of-thoughts enhanced llm for protein- protein interaction prediction. bioRxiv, pages 2024–04, 2024

  238. [246]

    Prot2token: A multi-task framework for protein language processing using autoregressive language modeling

    Mahdi Pourmirzaei, Farzaneh Esmaili, Mohammadreza Pour- mirzaei, Duolin Wang, and Dong Xu. Prot2token: A multi-task framework for protein language processing using autoregressive language modeling. bioRxiv, pages 2024–05, 2024

  239. [247]

    Bartsmiles: Generative masked language models for molecular representa- tions

    Gayane Chilingaryan, Hovhannes Tamoyan, Ani Tevosyan, Nelly Babayan, Lusine Khondkaryan, Karen Hambardzumyan, Zaven Navoyan, Hrant Khachatrian, and Armen Aghajanyan. Bartsmiles: Generative masked language models for molecular representa- tions. arXiv preprint arXiv:2211.16349, 2022

  240. [248]

    Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023) , ...

  241. [249]

    Proteingpt: Multimodal llm for protein property prediction and structure understanding

    Yijia Xiao, Edward Sun, Yiqiao Jin, Qifan Wang, and Wei Wang. Proteingpt: Multimodal llm for protein property prediction and structure understanding. arXiv preprint arXiv:2408.11363, 2024

  242. [250]

    Fapm: functional annotation of proteins using multimodal models beyond structural modeling

    Wenkai Xiang, Zhaoping Xiong, Huan Chen, Jiacheng Xiong, Wei Zhang, Zunyun Fu, Mingyue Zheng, Bing Liu, and Qian Shi. Fapm: functional annotation of proteins using multimodal models beyond structural modeling. Bioinformatics, 40(12):btae680, 2024

  243. [251]

    Mistral 7b

    Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023

  244. [252]

    Prott3: Protein-to-text generation for text-based protein understanding

    Zhiyuan Liu, An Zhang, Hao Fei, Enzhi Zhang, Xiang Wang, Kenji Kawaguchi, and Tat-Seng Chua. Prott3: Protein-to-text generation for text-based protein understanding. arXiv preprint arXiv:2405.12564, 2024

  245. [253]

    Protein language models are biased by unequal sequence sampling across the tree of life

    Frances Ding and Jacob Noah Steinhardt. Protein language models are biased by unequal sequence sampling across the tree of life. bioRxiv, pages 2024–03, 2024

  246. [254]

    Protein language model fitness is a matter of preference

    Cade Gordon, Amy X Lu, and Pieter Abbeel. Protein language model fitness is a matter of preference. bioRxiv, pages 2024–10, 2024

  247. [255]

    Netgo: improving large-scale protein function prediction with massive network information

    Ronghui You, Shuwei Yao, Yi Xiong, Xiaodi Huang, Fengzhu Sun, Hiroshi Mamitsuka, and Shanfeng Zhu. Netgo: improving large-scale protein function prediction with massive network information. Nucleic acids research, 47(W1):W379–W387, 2019

  248. [256]

    Netgo 2.0: improving large-scale protein function prediction with massive sequence, text, domain, family and network information

    Shuwei Yao, Ronghui You, Shaojun Wang, Yi Xiong, Xiaodi Huang, and Shanfeng Zhu. Netgo 2.0: improving large-scale protein function prediction with massive sequence, text, domain, family and network information. Nucleic acids research , 49(W1):W469– W475, 2021

  249. [257]

    Directed evolution: bringing new chemistry to life

    Frances H Arnold. Directed evolution: bringing new chemistry to life. Angewandte Chemie (International Ed. in English) , 57(16):4143, 2018

  250. [258]

    Methods for the directed evolution of proteins

    Michael S Packer and David R Liu. Methods for the directed evolution of proteins. Nature Reviews Genetics, 16(7):379–394, 2015

  251. [259]

    Efficient evolution of human antibodies from general protein language models

    Brian L Hie, Varun R Shanker, Duo Xu, Theodora UJ Bruun, Payton A Weidenbacher, Shaogeng Tang, Wesley Wu, John E Pak, and Peter S Kim. Efficient evolution of human antibodies from general protein language models. Nature Biotechnology, 42(2):275– 283, 2024

  252. [260]

    Generating novel protein sequences using gibbs sampling of masked language models

    Sean R Johnson, Sarah Monaco, Kenneth Massie, and Zaid Syed. Generating novel protein sequences using gibbs sampling of masked language models. bioRxiv, pages 2021–01, 2021

  253. [261]

    Generative power of a protein language model trained on multiple sequence alignments

    Damiano Sgarbossa, Umberto Lupo, and Anne-Florence Bitbol. Generative power of a protein language model trained on multiple sequence alignments. Elife, 12:e79854, 2023

  254. [262]

    Protein design by directed evolution guided by large language models

    Thanh VT Tran and Truong Son Hy. Protein design by directed evolution guided by large language models. bioRxiv, pages 2023– 11, 2023

  255. [263]

    Integrating genetic algorithms and language models for enhanced enzyme design

    Yves Gaetan Nana Teukam, Federico Zipoli, Teodoro Laino, Emanuele Criscuolo, Francesca Grisoni, and Matteo Manica. Integrating genetic algorithms and language models for enhanced enzyme design. 2024

  256. [264]

    Evoopt: an msa-guided, fully unsupervised sequence optimization pipeline for protein design

    Hideki Yamaguchi and Yutaka Saito. Evoopt: an msa-guided, fully unsupervised sequence optimization pipeline for protein design. In Machine Learning for Structural Biology Workshop, NeurIPS , 2022

  257. [265]

    A general temperature-guided language model to design proteins of enhanced stability and activity

    Fan Jiang, Mingchen Li, Jiajun Dong, Yuanxi Yu, Xinyu Sun, Banghao Wu, Jin Huang, Liqi Kang, Yufeng Pei, Liang Zhang, et al. A general temperature-guided language model to design proteins of enhanced stability and activity. Science Advances , 10(48):eadr2641, 2024

  258. [266]

    Sam- pling protein language models for functional protein design

    Jeremie Theddy Darmawan, Yarin Gal, and Pascal Notin. Sam- pling protein language models for functional protein design. In NeurIPS 2023 Generative AI and Biology (GenBio) Workshop , 2023

  259. [267]

    Rapid in silico directed evolution by a protein language model with evolvepro

    Kaiyi Jiang, Zhaoqing Yan, Matteo Di Bernardo, Samantha R Sgrizzi, Lukas Villiger, Alisan Kayabolen, BJ Kim, Josephine K Carscadden, Masahiro Hiraizumi, Hiroshi Nishimasu, et al. Rapid in silico directed evolution by a protein language model with evolvepro. Science, page eadr6...

  260. [268]

    Chatgpt-powered conversational drug editing using retrieval and domain feedback

    Shengchao Liu, Jiongxiao Wang, Yijin Yang, Chengpeng Wang, Ling Liu, Hongyu Guo, and Chaowei Xiao. Chatgpt-powered conversational drug editing using retrieval and domain feedback. arXiv preprint arXiv:2305.18090, 2023

  261. [269]

    Sparks of function by de novo protein design

    Alexander E Chu, Tianyu Lu, and Po-Ssu Huang. Sparks of function by de novo protein design. Nature biotechnology, 42(2):203– 215, 2024

  262. [270]

    Instructplm: Aligning protein language models to follow protein structure instructions

    Jiezhong Qiu, Junde Xu, Jie Hu, Hanqun Cao, Liya Hou, Zijun Gao, Xinyi Zhou, Anni Li, Xiujuan Li, Bin Cui, et al. Instructplm: Aligning protein language models to follow protein structure instructions. bioRxiv, pages 2024–04, 2024

  263. [271]

    Toward de novo protein design from natural language

    Fengyuan Dai, Yuliang Fan, Jin Su, Chentong Wang, Chenchen Han, Xibin Zhou, Jianming Liu, Hui Qian, Shunzhi Wang, Anping Zeng, et al. Toward de novo protein design from natural language. bioRxiv, pages 2024–08, 2024

  264. [272]

    Conditional language models enable the efficient design of proficient enzymes

    Geraldene Munsamy, Ramiro Illanes-Vicioso, Silvia Funcillo, Ioanna T Nakou, Sebastian Lindner, Gavin Ayres, Lesley S Sheehan, Steven Moss, Ulrich Eckhard, Philipp Lorenz, et al. Conditional language models enable the efficient design of proficient enzymes. bioRxiv, pages 2024–05, 2024

  265. [273]

    Protagents: Pro- tein discovery via large language model multi-agent collabora- tions combining physics and machine learning

    Alireza Ghafarollahi and Markus J Buehler. Protagents: Pro- tein discovery via large language model multi-agent collabora- tions combining physics and machine learning. arXiv preprint arXiv:2402.04268, 2024

  266. [274]

    Clonal selection and learning in the antibody system

    Klaus Rajewsky. Clonal selection and learning in the antibody system. Nature, 381(6585):751–758, 1996. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 33

  267. [275]

    Animal immunization, in vitro display technologies, and machine learning for antibody discovery

    Andreas H Laustsen, Victor Greiff, Aneesh Karatt-Vellatt, Serge Muyldermans, and Timothy P Jenkins. Animal immunization, in vitro display technologies, and machine learning for antibody discovery. Trends in Biotechnology, 39(12):1263–1273, 2021

  268. [276]

    De novo generation of sars-cov-2 antibody cdrh3 with a pre- trained generative large language model

    Haohuai He, Bing He, Lei Guan, Yu Zhao, Feng Jiang, Guanxing Chen, Qingge Zhu, Calvin Yu-Chian Chen, Ting Li, and Jianhua Yao. De novo generation of sars-cov-2 antibody cdrh3 with a pre- trained generative large language model. Nature Communications, 15(1):6867, 2024

  269. [277]

    Unsupervised evolution of protein and antibody complexes with a structure-informed language model

    Varun R Shanker, Theodora UJ Bruun, Brian L Hie, and Peter S Kim. Unsupervised evolution of protein and antibody complexes with a structure-informed language model. Science, 385(6704):46– 53, 2024

  270. [278]

    Enzymes: principles and biotechnological applications

    Peter K Robinson. Enzymes: principles and biotechnological applications. Essays in biochemistry, 59:1, 2015

  271. [279]

    Dissecting en- zyme function with microfluidic-based deep mutational scanning

    Philip A Romero, Tuan M Tran, and Adam R Abate. Dissecting en- zyme function with microfluidic-based deep mutational scanning. Proceedings of the National Academy of Sciences , 112(23):7159–7164, 2015

  272. [280]

    Effects of ionizing radiation on in vitro replication of dna by dna polymerase i

    Roy Saffhill and JJ Weiss. Effects of ionizing radiation on in vitro replication of dna by dna polymerase i. Nature New Biology, 241(107):69–71, 1973

  273. [281]

    Proteases as therapeutics

    Charles S Craik, Michael J Page, and Edwin L Madison. Proteases as therapeutics. Biochemical Journal, 435(1):1–16, 2011

  274. [282]

    Engineering the third wave of bio- catalysis

    Uwe T Bornscheuer, GW Huisman, RJ Kazlauskas, S Lutz, JC Moore, and K Robins. Engineering the third wave of bio- catalysis. Nature, 485(7397):185–194, 2012

  275. [283]

    Computa- tional scoring and experimental evaluation of enzymes generated by neural networks

    Sean R Johnson, Xiaozhi Fu, Sandra Viknander, Clara Goldin, Sarah Monaco, Aleksej Zelezniak, and Kevin K Yang. Computa- tional scoring and experimental evaluation of enzymes generated by neural networks. Nature biotechnology, pages 1–10, 2024

  276. [284]

    Molecular mechanisms of translational control

    Fátima Gebauer and Matthias W Hentze. Molecular mechanisms of translational control. Nature reviews Molecular cell biology , 5(10):827–835, 2004

  277. [285]

    Protein–dna/rna interactions: an overview of investigation methods in the-omics era

    Flora Cozzolino, Ilaria Iacobucci, Vittoria Monaco, and Maria Monti. Protein–dna/rna interactions: an overview of investigation methods in the-omics era. Journal of Proteome Research, 20(6):3018– 3030, 2021

  278. [286]

    Discovering drug–target interaction knowledge from biomedical literature

    Yutai Hou, Yingce Xia, Lijun Wu, Shufang Xie, Yang Fan, Jin- hua Zhu, Tao Qin, and Tie-Yan Liu. Discovering drug–target interaction knowledge from biomedical literature. Bioinformatics, 38(22):5100–5107, 2022

  279. [287]

    Principles of early drug discovery

    James P Hughes, Stephen Rees, S Barrett Kalindjian, and Karen L Philpott. Principles of early drug discovery. British journal of pharmacology, 162(6):1239–1249, 2011

  280. [288]

    Transdti: transformer-based language models for estimating dtis and building a drug recommendation workflow

    Yogesh Kalakoti, Shashank Yadav, and Durai Sundar. Transdti: transformer-based language models for estimating dtis and building a drug recommendation workflow. ACS omega, 7(3):2706– 2717, 2022

  281. [289]

    Large scale paired antibody language models

    Henry Kenlay, Frédéric A Dreyer, Aleksandr Kovaltsuk, Dom Miketa, Douglas Pires, and Charlotte M Deane. Large scale paired antibody language models. PLOS Computational Biology , 20(12):e1012646, 2024

  282. [290]

    Linker-tuning: Optimizing continuous prompts for heterodimeric protein prediction

    Shuxian Zou, Hui Li, Shentong Mo, Xingyi Cheng, Eric Xing, and Le Song. Linker-tuning: Optimizing continuous prompts for heterodimeric protein prediction. arXiv preprint arXiv:2312.01186, 2023

  283. [291]

    Rotamer density estimator is an unsupervised learner of the effect of mutations on protein-protein interaction

    Shitong Luo, Yufeng Su, Zuofan Wu, Chenpeng Su, Jian Peng, and Jianzhu Ma. Rotamer density estimator is an unsupervised learner of the effect of mutations on protein-protein interaction. bioRxiv, pages 2023–02, 2023

  284. [292]

    Protein design: the experts speak

    Anne Doerr. Protein design: the experts speak. Nature Biotechnol- ogy, 42(2):175–178, 2024

  285. [293]

    Protein language models learn evolutionary statistics of interacting sequence motifs

    Zhidian Zhang, Hannah K Wayment-Steele, Garyk Brixi, Haobo Wang, Dorothee Kern, and Sergey Ovchinnikov. Protein language models learn evolutionary statistics of interacting sequence motifs. Proceedings of the National Academy of Sciences , 121(45):e2406285121, 2024

  286. [294]

    Interplm: Discovering interpretable features in protein language models via sparse autoencoders

    Elana Simon and James Zou. Interplm: Discovering interpretable features in protein language models via sparse autoencoders. bioRxiv, pages 2024–11, 2024

  287. [295]

    Towards explainable artificial intelligence (xai): A data mining perspective

    Haoyi Xiong, Xiaofei Zhang, Jiamin Chen, Xinhao Sun, Yuchen Li, Zeyi Sun, Mengnan Du, et al. Towards explainable artificial intelligence (xai): A data mining perspective. arXiv preprint arXiv:2401.04374, 2024

  288. [296]

    Scientific discovery in the age of artificial intelligence

    Hanchen Wang, Tianfan Fu, Yuanqi Du, Wenhao Gao, Kexin Huang, Ziming Liu, Payal Chandak, Shengchao Liu, Peter Van Katwyk, Andreea Deac, et al. Scientific discovery in the age of artificial intelligence. Nature, 620(7972):47–60, 2023

  289. [297]

    Gpt-4 technical report

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023

  290. [298]

    Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller

    Andres M. Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. Augmenting large language models with chemistry tools. Nature Machine Intelligence, pages 1–11, 2024

  291. [299]

    Autonomous chemical research with large language models

    Daniil A Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes. Autonomous chemical research with large language models. Nature, 624(7992):570–578, 2023

  292. [2022]

    Nucleic acids research, 51(D1):D418–D427, 2023

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.