Pith. sign in

REVIEW 7 major objections 6 minor 107 references

Transformers in Protein: A Survey

T0 review · 7 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This survey claims that Transformer models now span the four core protein-research domains, with over 100 studies synthesized into one domain-oriented framework.

desk verdict A well-organized survey of Transformer methods in protein science, but its systemic citation-reference mismatches break the central promise of a reliable review. read the letter →

arxiv 2505.20098 v2 pith:YGS7E6Q6 submitted 2025-05-26 cs.LG cs.CRq-bio.QM

classification cs.LGcs.CRq-bio.QM
keywords Transformersproteininformaticsstructurepredictionfunctionprotein–proteininteractiondrugdiscoveryself-attentionlanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a review that tries to establish that Transformer models have become a general-purpose computational backbone for protein research, with a literature large enough to organize by scientific domain. The authors review over 100 studies across protein structure prediction, protein function prediction, protein–protein interaction analysis, and drug discovery/target identification, arguing that self-attention's ability to capture long-range dependencies in sequences is what carries these applications. For each domain they restate the scientific goals, compare Transformer-based methods with earlier approaches, and point to the contributions they consider transformative. They also compile datasets and open-source code repositories to support reproducibility and benchmarking, and they close by identifying computational cost, data bias, and interpretability as the field's main bottlenecks. A sympathetic reading is that the paper provides a consolidated map of where Transformers stand in protein informatics and why they matter.

What carries the argument

The organizing machinery of the paper is the domain-oriented classification system, which divides the Transformer-in-protein literature into structure prediction, function prediction, protein–protein interactions, and drug discovery. Within that frame the paper uses a taxonomy of self-attention mechanisms—single-head versus multi-head, spatial and graph-based attention, SE(3)-equivariant attention, and hybrid CNN–Transformer designs—to characterize the model families, and it uses bibliometric analysis (Fig. 1) and curated resource tables (Tables II and III) to support the claims of growth and reproducibility. No new algorithm is proposed; the machinery is the synthesis itself.

What would settle it

Reproducing Fig. 1 with a stated search query, access date, and inclusion criteria, and then checking the claims against their cited sources—for example, whether reference [71] is a protein structure model, whether [65] reports protein-ligand fine-tuning, and whether [58] concerns ESM-Fold—would settle whether the survey's descriptions and trend data are reliable.

Watch

Extended reading notes

Core claim

The paper's central claim is that the Transformer's application to protein informatics has reached a state where a comprehensive, domain-oriented synthesis is both possible and needed, and that this synthesis shows the architecture reshaping the field. It surveys more than 100 studies and organizes them into four application domains: protein structure prediction, protein function prediction, protein–protein interaction analysis, and drug discovery/target identification. It further claims that the Transformer's self-attention mechanism explains the gains—because protein sequence, structure, and function depend on distal residue interactions—and that pre-trained protein language models such as AlphaFold, ESM-Fold, ProtTrans, and ProteinBERT are the vehicles through which those gains appear. The paper also asserts that curating datasets and code repositories is an essential part of the contribution, since reproducibility and benchmarking depend on them, and that future progress will come from multimodal integration, hybrid physics-informed modeling, efficiency improvements, and interpretability.

Load-bearing premise

The survey's usefulness depends on each cited reference actually supporting the statement it is attached to, and on the bibliometric counts in Fig. 1 coming from a clearly defined and correctly executed search; a sympathetic reader has to take that reliability on faith.

Editorial extensions

If this is right

  • Newcomers to protein informatics can use the survey as a single entry point that explains Transformer fundamentals and then maps specific models onto the four application domains.
  • The curated dataset and code tables give researchers a concrete starting set of resources for reproducing leading models or benchmarking new ones.
  • The paper's list of persistent bottlenecks—quadratic attention cost, biased and sparse datasets, limited interpretability, and difficulty with novel folds—defines a research agenda for the next generation of protein Transformers.
  • The four-domain grouping makes explicit that the same self-attention machinery is being reused across structure, function, interaction, and drug-discovery tasks, which supports cross-domain transfer of methods.
  • If the review's coverage is accurate, the field has shifted from asking whether Transformers can help with proteins to asking how to make them efficient, interpretable, and multimodal.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The static tables of models and repositories will age quickly given the field's publication rate; a living, community-maintained version would better serve the reproducibility goal the paper sets.
  • The paper's taxonomy of attention mechanisms could be turned into a quantitative meta-analysis—grouping published models by architecture family and comparing benchmark results—which the survey itself does not attempt.
  • The emphasis on multimodal integration suggests that the next evaluation frontier for protein Transformers will be joint sequence–structure–function modeling rather than single-task accuracy.
  • The domain-oriented grouping could seed a benchmark suite that compares Transformer variants within each domain, something the paper lists resources for but does not build.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

7 major / 6 minor

Summary. This manuscript is a survey of Transformer applications in protein science, organized by application domain (structure prediction, function prediction, protein-protein interactions, drug discovery) and supplemented by tables of model derivatives, open-source implementations, and datasets. The authors state that they reviewed over 100 studies and aim to provide a consolidated, accurate foundation for the field. The survey covers major models such as AlphaFold, RoseTTAFold, ESM-Fold, OmegaFold, ProtTrans, ProteinBERT, and several newer 2024-2025 methods, and it discusses advantages, challenges, and future directions. The central claim is that the paper provides a reliable domain-oriented review with curated resources for reproducibility and benchmarking.

Significance. If the survey's descriptions and citations were reliable, it would serve as a useful entry point for researchers applying Transformers to protein problems, particularly because it assembles model names, resource links, and dataset information in one place. The paper makes a genuine effort to organize a large and fast-moving literature, and the inclusion of resource tables (Tables II and III) and a structured discussion of challenges is valuable in principle. However, the significance of the contribution is currently undermined by the accuracy problems detailed below, which affect the very basis of a survey: readers cannot trust the claims attached to the references. The survey is therefore potentially useful after substantial correction, but in its present form its reliability as a reference work is not established.

major comments (7)
  1. [Section III-A.7 and Table I] The description of 'GraphTrans' as a model that integrates graph neural networks with Transformer layers for protein complex prediction is not supported by the cited reference [71], which is a Winter Simulation Conference paper on a software system for network conversions, not a protein structure prediction model. This is a load-bearing error because the survey claims to provide accurate descriptions of Transformer derivatives; a reader consulting the citation would find nothing about proteins. The same unrelated reference [84] is also used in Section III-C.3 to support 'Graph Attention Networks (GATs)', where the canonical GAT paper is not cited. Please replace these with correct references or remove the unsupported entries from Table I and the text.
  2. [Section III-D.1] The claim that 'ChemBERTa is a prominent model in this area, fine-tuned on protein-ligand interaction datasets' with 'superior performance compared to traditional methods like molecular docking and molecular dynamics simulations in predicting binding affinities' is not present in the cited paper [65], which describes self-supervised pretraining on SMILES strings for molecular property prediction. No fine-tuning on protein-ligand datasets or docking comparisons appear in that 2020 arXiv paper. This is a concrete misattribution in a central section of the survey and should be corrected by either citing the relevant subsequent work or removing the unsupported claims.
  3. [Section III-C.1] The text credits DeepPPI with 'an attention mechanism that enables the model to focus on biologically relevant sequence motifs,' but the cited 2017 paper [83] (Du et al., 'Deepppi') uses fully connected deep neural networks on sequence-derived features and does not contain an attention mechanism. The description of the model is therefore factually incorrect. This matters because the survey is meant to summarize model architectures; a reader relying on the review would form a wrong idea of DeepPPI's design. Please revise the description to match the cited source.
  4. [Section III-A.3 and Table I] The ESM-Fold passage cites [58] for the claim that ESM-Fold predicts structures for proteins with 'minimal or no homology to known structures' and for generalizability to orphan proteins. Reference [58] is a neurophysiology study on suppression of alpha-band power during emotional distraction, which has no connection to protein structure prediction. This is another example of a systematic reference-reference mismatch, not an isolated typo. Given that such mismatches appear in at least four separate model descriptions (GraphTrans, ChemBERTa, DeepPPI, ESM-Fold), the authors must audit every citation in the manuscript against the claims it supports before the survey can be considered trustworthy.
  5. [Section II-D.2 and Section II-D.3] Two background paragraphs contain citation mismatches that, while not about individual models, still affect the survey's reliability as a reference. In Section II-D.2, the claim that 'recent advancements have highlighted the growing efficacy of transformer-based models in predicting stability' is supported by [50] (Rost and Sander, 1993) and [51] (Scheraga et al., 2007), neither of which is a transformer-based stability predictor. In Section II-D.3, the claim about ProtTrans generalization is cited to [52] (Magnan and Baldi, SSpro) and [53] (Peitsch, SWISS-MODEL), again unrelated. These references should be replaced with actual transformer-based prediction papers or the claims should be softened.
  6. [Table II and Section VII] The paper's promise to curate open-source resources for reproducibility is undermined by errors in Table II. The ProtGPT2 repository URL points to a Hugging Face trainer documentation page rather than the ProtGPT2 model repository, and the ProteinBERT URL ('https://github.com/nadavbra/protein bert') contains a space and is not a valid URL. Table III also gives dataset sizes and access methods without specifying whether sizes are sequence counts or storage sizes (e.g., '30,051' for ESM-Fold UR50 is unclear). Since reproducibility support is part of the central contribution, these entries must be verified and corrected.
  7. [Fig. 1 and Section I] The bibliometric analysis in Fig. 1 is presented as evidence of the field's growth, but the caption and text do not provide the search query, database access date, inclusion criteria, or subfield classification rules used to obtain the data from the Web of Science Core Collection. Without this information, the trend claims (e.g., the surge in publications and the distribution across journals and subfields) cannot be reproduced or validated. Please add a methodology paragraph describing the search and filtering process.
minor comments (6)
  1. [Section III-B.1] The text uses 'ProtBERT' and 'ProteinBERT' interchangeably; the cited work [63] is ProteinBERT by Brandes et al., which is a distinct model from ProtBERT in the ProtTrans collection. Please use a consistent name to avoid confusing the two.
  2. [References] Several references are duplicated: [37] and [54] are the same AlphaFold paper, [40] and [55] are the same ProtTrans paper, [45] and [7] are the same Rives et al. paper, and [57] duplicates [29]. Please consolidate the bibliography or adjust the citation numbers.
  3. [Section III-D.6] The subsection 'RL-based Transformer Models for Molecule Generation' contains no citation at all; in a survey that promises to cover over 100 studies, each described approach should be linked to a reference.
  4. [Section II (opening paragraph)] The sentence 'The paper is organized as follows...' appears at the top of Section II, after the introduction; this organizational roadmap would be more naturally placed at the end of Section I.
  5. [Section III-A.1] The statement that AlphaFold 'achieved a median Global Distance Test (GDT) score of 92.4' lacks a precise target (e.g., CASP14 free-modeling domains); please clarify the metric scope for accuracy.
  6. [Fig. 2 caption] The caption's reference to 'the third row-left block' is confusing; please refer to the specific subfigure or panel labels.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey contains no derivation chain whose output is equivalent to its input by construction.

full rationale

This paper is a domain survey rather than a derivation-driven work. It makes no fitted-parameter predictions, proves no theorems from assumptions, and does not reduce any claimed result to its own inputs by construction. The only self-citations are references [4], [5], and [6] in the introduction, used to illustrate that transformer models have been adopted across disciplines; that claim is ambient context and is not load-bearing for the survey's organization, resource tables, or model summaries. The central deliverable is a literature synthesis whose accuracy rests on the correspondence between cited external sources and the claims attached to them. Issues such as the mismatched references identified in Section III-A.7, III-C.1, and III-D.1 (e.g., refs [71], [65], [58]) are citation-accuracy defects, not circularity: a mistaken attribution does not make a statement equivalent to its input by definition, nor does it fit a parameter and then rename the fit as a prediction. No circular step of any enumerated kind is present, so the appropriate finding is a score of 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The survey itself introduces no mathematical model or fitted parameters. Its central claims depend on two unstated methodological assumptions: the correctness of the bibliometric search behind Fig. 1, and the fidelity of its citations to the primary literature. Both assumptions are weakened by observable problems in the text, so the survey's reliability is compromised.

assumptions (2)
  • domain assumption The Web of Science Core Collection queries used to generate Fig. 1 are correctly constructed and capture the relevant Transformer-in-protein literature.
    The survey's claim that it documents publication trends (Fig. 1) depends on an unstated search protocol; no query, date, or inclusion criteria are reported in Section I or the figure caption.
  • domain assumption Cited references accurately support the descriptions of each model and method presented in the survey.
    The survey's value depends on fidelity to the primary literature, but several citations are mismatched, for example GraphTrans cited to a network-conversion software, ChemBERTa claims not found in the cited paper, and unrelated references such as [58] and [111]. This assumption is violated in multiple places.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Transformers in Protein: A Survey." pith.science (2026). https://pith.science/paper/YGS7E6Q6

@misc{pith2026250520098,
  author       = {Pith},
  title        = {Pith review of: Transformers in Protein: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YGS7E6Q6}},
  note         = {Machine review of arXiv:2505.20098}
}
read the original abstract

As protein informatics advances rapidly, the demand for enhanced predictive accuracy, structural analysis, and functional understanding has intensified. Transformer models, as powerful deep learning architectures, have demonstrated unprecedented potential in addressing diverse challenges across protein research. However, a comprehensive review of Transformer applications in this field remains lacking. This paper bridges this gap by surveying over 100 studies, offering an in-depth analysis of practical implementations and research progress of Transformers in protein-related tasks. Our review systematically covers critical domains, including protein structure prediction, function prediction, protein-protein interaction analysis, functional annotation, and drug discovery/target identification. To contextualize these advancements across various protein domains, we adopt a domain-oriented classification system. We first introduce foundational concepts: the Transformer architecture and attention mechanisms, categorize Transformer variants tailored for protein science, and summarize essential protein knowledge. For each research domain, we outline its objectives and background, critically evaluate prior methods and their limitations, and highlight transformative contributions enabled by Transformer models. We also curate and summarize pivotal datasets and open-source code resources to facilitate reproducibility and benchmarking. Finally, we discuss persistent challenges in applying Transformers to protein informatics and propose future research directions. This review aims to provide a consolidated foundation for the synergistic integration of Transformer and protein informatics, fostering further innovation and expanded applications in the field.

Figures

Figures reproduced from arXiv: 2505.20098 by the authors.

Figure 1
Figure 1. Analysis of Transformer models in protein research using data from the Web of Science Core Collection. (a) Counts [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Architecture of the Transformer Model [1]Originally proposed for machine translation tasks, the Transformer model [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Hierarchical Classification of Self-Attention Mechanisms in Protein Models. This diagram illustrates the categorization [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Illustrative diagram of three fundamental supervised learning tasks. Supervised learning in machine learning (ML) is [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: The structure of the BERT model. This figure shows [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

107 extracted references · 64 canonical work pages

  1. [65]

    Chemberta: large-scale self-supervised pretraining for molecular property prediction,

    S. Chithrananda, G. Grand, and B. Ramsundar, “Chemberta: large-scale self-supervised pretraining for molecular property prediction,” arXiv preprint arXiv:2010.09885, 2020

  2. [83]

    Deepppi: boosting prediction of protein–protein interactions with deep neural networks,

    X. Du, S. Sun, C. Hu, Y . Yao, Y . Yan, and Y . Zhang, “Deepppi: boosting prediction of protein–protein interactions with deep neural networks,” Journal of chemical information and modeling, vol. 57, no. 6, pp. 1499– 1510, 2017

  3. [58]

    Suppression of alpha-band power underlies exogenous attention to emotional distractors,

    L. Arana, M. Melc ´on, D. Kessel, S. Hoyos, J. Albert, L. Carreti ´e, and A. Capilla, “Suppression of alpha-band power underlies exogenous attention to emotional distractors,” Psychophysiology, vol. 59, no. 9, p. e14051, 2022

  4. [84]

    Graphtrans: a software system for network conversions for simulation, structural analysis, and graph operations,

    H. L. Carscadden, L. Machi, C. J. Kuhlman, D. Machi, and S. Ravi, “Graphtrans: a software system for network conversions for simulation, structural analysis, and graph operations,” in 2021 Winter Simulation Conference (WSC). IEEE, 2021, pp. 1–12

  5. [50]

    Prediction of protein secondary structure at better than 70% accuracy,

    B. Rost and C. Sander, “Prediction of protein secondary structure at better than 70% accuracy,” Journal of molecular biology , vol. 232, no. 2, pp. 584–599, 1993

  6. [51]

    Protein-folding dynamics: overview of molecular simulation techniques,

    H. A. Scheraga, M. Khalili, and A. Liwo, “Protein-folding dynamics: overview of molecular simulation techniques,” Annu. Rev. Phys. Chem., vol. 58, no. 1, pp. 57–83, 2007

  7. [53]

    Protein modeling by e-mail,

    M. C. Peitsch, “Protein modeling by e-mail,” Bio/technology, vol. 13, no. 7, pp. 658–660, 1995

  8. [1]

    Attention is all you need,

    A. Vaswani, “Attention is all you need,” Advances in Neural Informa- tion Processing Systems , 2017

Show all 107 references
  1. [2]

    Bert: Pretraining of deep bidirectional transformers for language understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pretraining of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologie...

  2. [3]

    Language models are few-shot learners,

    T. B. Brown, “Language models are few-shot learners,” arXiv preprint arXiv:2005.14165, 2020

  3. [4]

    Pmanet: Malicious url detection via post-trained language model guided multi-level feature attention network,

    R. Liu, Y . Wang, H. Xu, Z. Qin, F. Zhang, Y . Liu, and Z. Cao, “Pmanet: Malicious url detection via post-trained language model guided multi-level feature attention network,” Information Fusion, vol. 113, p. 102638, 2025

  4. [5]

    Vul- lmgnns: Fusing language models and online-distilled graph neural networks for code vulnerability detection,

    R. Liu, Y . Wang, H. Xu, J. Sun, F. Zhang, P. Li, and Z. Guo, “Vul- lmgnns: Fusing language models and online-distilled graph neural networks for code vulnerability detection,” Information Fusion , vol. 115, p. 102748, 2025

  5. [6]

    Ethereum fraud detection via joint transaction language model and graph representation learning,

    J. Sun, Y . Jia, Y . Wang, Y . Tian, and S. Zhang, “Ethereum fraud detection via joint transaction language model and graph representation learning,” Information Fusion, vol. 120, p. 103074, 2025

  6. [7]

    Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences,

    A. Rives, J. Meier, T. Sercu, S. Goyal, Z. Lin, J. Liu, D. Guo, M. Ott, C. L. Zitnick, J. Ma et al., “Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences,” Proceedings of the National Academy of Sciences , 2021

  7. [8]

    Highly accurate protein structure prediction with alphafold,

    J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. ˇZ´ıdek, A. Potapenko et al., “Highly accurate protein structure prediction with alphafold,” nature, 2021

  8. [9]

    Evaluating protein transfer learning with tape,

    R. Rao, N. Bhattacharya, N. Thomas, Y . Duan, P. Chen, J. Canny, P. Abbeel, and Y . Song, “Evaluating protein transfer learning with tape,” Advances in neural information processing systems , 2019

  9. [10]

    Evolutionary velocity with protein language models predicts evolutionary dynamics of diverse proteins,

    B. L. Hie, K. K. Yang, and P. S. Kim, “Evolutionary velocity with protein language models predicts evolutionary dynamics of diverse proteins,” Cell Systems, 2022

  10. [11]

    Language models of protein sequences at the scale of evolution enable accurate structure prediction,

    Z. Lin, H. Akin, R. Rao, B. Hie, Z. Zhu, W. Lu, A. dos Santos Costa, M. Fazel-Zarandi, T. Sercu, S. Candido et al. , “Language models of protein sequences at the scale of evolution enable accurate structure prediction,” BioRxiv, 2022

  11. [12]

    Deep learning in bioinformatics,

    S. Min, B. Lee, and S. Yoon, “Deep learning in bioinformatics,” Briefings in bioinformatics , 2017

  12. [13]

    Machine learning solutions for predicting protein–protein interactions,

    R. Casadio, P. L. Martelli, and C. Savojardo, “Machine learning solutions for predicting protein–protein interactions,” Wiley Interdis- ciplinary Reviews: Computational Molecular Science , 2022

  13. [14]

    Unified rational protein engineering with sequence-based deep representation learning,

    E. C. Alley, G. Khimulya, S. Biswas, M. AlQuraishi, and G. M. Church, “Unified rational protein engineering with sequence-based deep representation learning,” Nature methods, 2019

  14. [15]

    Protein- bert: a universal deep-learning model of protein sequence and function,

    N. Brandes, D. Ofer, Y . Peleg, N. Rappoport, and M. Linial, “Protein- bert: a universal deep-learning model of protein sequence and function,” Bioinformatics, 2022

  15. [16]

    Deep learning in proteomics,

    B. Wen, W.-F. Zeng, Y . Liao, Z. Shi, S. R. Savage, W. Jiang, and B. Zhang, “Deep learning in proteomics,” Proteomics, 2020

  16. [17]

    Transformer- based deep learning for predicting protein properties in the life sci- ences,

    A. Chandra, L. T ¨unnermann, T. L ¨ofstedt, and R. Gratz, “Transformer- based deep learning for predicting protein properties in the life sci- ences,” Elife, 2023

  17. [18]

    Machine learning: its challenges and opportunities in plant system biology,

    M. Hesami, M. Alizadeh, A. M. P. Jones, and D. Torkamaneh, “Machine learning: its challenges and opportunities in plant system biology,” Applied Microbiology and Biotechnology , 2022

  18. [19]

    Artificial intelligence in the prediction of protein–ligand interactions: recent advances and future directions,

    A. Dhakal, C. McKay, J. J. Tanner, and J. Cheng, “Artificial intelligence in the prediction of protein–ligand interactions: recent advances and future directions,” Briefings in Bioinformatics , 2022

  19. [20]

    Roberta: A robustly optimized bert pretraining approach,

    Y . Liu, M. Ott, N. Goyal, J. Du et al., “Roberta: A robustly optimized bert pretraining approach,” arXiv preprint, 2019

  20. [21]

    Linformer: Self- attention with linear complexity,

    S. Wang, B. Z. Li, M. Khabsa, H. Fang, and H. Ma, “Linformer: Self- attention with linear complexity,” arXiv preprint, 2020

  21. [22]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929 , 2020

  22. [23]

    On the turing completeness of modern neural network architectures,

    J. Perez, J. Marinkovic, and P. Barcelo, “On the turing completeness of modern neural network architectures,” in International Conference on Learning Representations , 2018

  23. [24]

    On the relationship between self-attention and convolutional layers,

    J.-B. Cordonnier, A. Loukas, and M. Jaggi, “On the relationship between self-attention and convolutional layers,” in International Con- ference on Learning Representations , 2019

  24. [25]

    Deformable convolutional networks,

    J. Dai, H. Qi, Y . Xiong, Z. Li, G. Zhang, H. Hu, and Y . Wei, “Deformable convolutional networks,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2017

  25. [26]

    Prottrans: Towards cracking the language of life’s code through self-supervised deep learning and high performance computing,

    A. Elnaggar, M. Heinzinger, C. Dallago, G. Rehawi, W. Yu, L. Jones, T. Gibbs, T. Feher, C. Angerer, M. Steinegger et al. , “Prottrans: Towards cracking the language of life’s code through self-supervised deep learning and high performance computing,” IEEE Transactions on Patte...

  26. [28]

    Peptidebert: A pre-trained model for peptide representation learning,

    X. Li, X. Zhang, S. Wang, X. Yu, S. Pan, J. Wu, and et al., “Peptidebert: A pre-trained model for peptide representation learning,” arXiv preprint arXiv:2309.03099, 2023. [Online]. Available: https://arxiv.org/abs/2309.03099

  27. [29]

    Language models of protein sequences at the scale of evolution enable accurate structure prediction,

    Z. Lin, H. Akin, R. Rao, B. Hie, Z. Zhu, W. Lu, T. Sercu, and A. Rives, “Language models of protein sequences at the scale of evolution enable accurate structure prediction,” bioRxiv, 2022

  28. [30]

    Msa transformer,

    R. Rao, J. Liu, R. Verkuil, J. Meier, J. Canny, P. Abbeel, T. Sercu, and A. Rives, “Msa transformer,” bioRxiv, 2021

  29. [31]

    Protein representation learning with cross-modal contrastive pretraining,

    C. Lu, L. Xu, Y . Wu, Y . Liu, and et al., “Protein representation learning with cross-modal contrastive pretraining,” arXiv preprint arXiv:2401.14819, 2024. [Online]. Available: https://arxiv.org/abs/ 2401.14819

  30. [32]

    Alberts, D

    B. Alberts, D. Bray, K. Hopkin, A. D. Johnson, J. Lewis, M. Raff, K. Roberts, and P. Walter, Essential cell biology . Garland Science, 2015

  31. [33]

    Prediction of protein conformation,

    P. Y . Chou and G. D. Fasman, “Prediction of protein conformation,” Biochemistry, vol. 13, no. 2, pp. 222–245, 1974

  32. [35]

    Analysis of the accuracy and implications of simple methods for predicting the secondary structure of globular proteins,

    J. Garnier, D. J. Osguthorpe, and B. Robson, “Analysis of the accuracy and implications of simple methods for predicting the secondary structure of globular proteins,” Journal of molecular biology , vol. 120, no. 1, pp. 97–120, 1978

  33. [36]

    Modeling aspects of the language of life through transfer-learning protein sequences,

    M. Heinzinger, A. Elnaggar, Y . Wang, C. Dallago, D. Nechaev, F. Matthes, and B. Rost, “Modeling aspects of the language of life through transfer-learning protein sequences,” BMC bioinformatics, vol. 20, pp. 1–17, 2019

  34. [39]

    Tm-align: a protein structure alignment algorithm based on the tm-score,

    Y . Zhang and J. Skolnick, “Tm-align: a protein structure alignment algorithm based on the tm-score,”Nucleic acids research, vol. 33, no. 7, pp. 2302–2309, 2005

  35. [40]

    Prottrans: towards cracking the language of life’s code through self-supervised learning,

    A. Elnaggar, M. Heinzinger, C. Dallago, G. Rehawi, Y . Wang, L. Jones, T. Gibbs, T. Feher, C. Angerer, M. Steineggeret al., “Prottrans: towards cracking the language of life’s code through self-supervised learning,” IEEE Transactions on Pattern Analysis and Machine Intelligenc...

  36. [41]

    The protein-folding problem, 50 years on,

    K. A. Dill and J. L. MacCallum, “The protein-folding problem, 50 years on,” science, vol. 338, no. 6110, pp. 1042–1046, 2012

  37. [42]

    Protein folding and misfolding,

    C. M. Dobson, “Protein folding and misfolding,” Nature, vol. 426, no. 6968, pp. 884–890, 2003

  38. [48]

    Goap: a generalized orientation-dependent, all-atom statistical potential for protein structure prediction,

    H. Zhou and J. Skolnick, “Goap: a generalized orientation-dependent, all-atom statistical potential for protein structure prediction,” Biophys- ical journal, vol. 101, no. 8, pp. 2043–2052, 2011

  39. [49]

    Protein secondary structure prediction based on position- specific scoring matrices,

    D. T. Jones, “Protein secondary structure prediction based on position- specific scoring matrices,” Journal of molecular biology , vol. 292, no. 2, pp. 195–202, 1999

  40. [57]

    Language models of protein sequences at the scale of evolution enable accurate structure prediction,

    Z. Lin, H. Akin, R. Rao, B. Hie, Z. Zhu, W. Lu, A. dos Santos Costa, M. Fazel-Zarandi, T. Sercu, S. Candido et al. , “Language models of protein sequences at the scale of evolution enable accurate structure prediction,” BioRxiv, vol. 2022, p. 500902, 2022

  41. [60]

    Improved protein structure prediction using potentials from deep learning,

    A. W. Senior, R. Evans, J. Jumper, J. Kirkpatrick, L. Sifre, T. Green, C. Qin, A. ˇZ´ıdek, A. W. Nelson, A. Bridgland et al., “Improved protein structure prediction using potentials from deep learning,” Nature, vol. 577, no. 7792, pp. 706–710, 2020

  42. [61]

    Evaluating protein transfer learning with tape,

    R. Rao, N. Bhattacharya, N. Thomas, Y . Duan, P. Chen, J. Canny, P. Abbeel, and Y . Song, “Evaluating protein transfer learning with tape,” Advances in neural information processing systems , vol. 32, 2019

  43. [62]

    A protein structure prediction approach leveraging transformer and cnn integration,

    Y . Zhou, K. Tan, X. Shen, Z. He, and H. Zheng, “A protein structure prediction approach leveraging transformer and cnn integration,” arXiv preprint arXiv:2402.19095 , 2024. [Online]. Available: https: //arxiv.org/abs/2402.19095

  44. [63]

    Protein- bert: a universal deep-learning model of protein sequence and function,

    N. Brandes, D. Ofer, Y . Peleg, N. Rappoport, and M. Linial, “Protein- bert: a universal deep-learning model of protein sequence and function,” Bioinformatics, vol. 38, no. 8, pp. 2102–2110, 2022

  45. [64]

    Trans-morfs: A disordered protein predictor based on the transformer architecture,

    C. Meng, Y . Shi, X. Fu, Q. Zou, and W. Han, “Trans-morfs: A disordered protein predictor based on the transformer architecture,” IEEE Journal of Biomedical and Health Informatics , pp. 1–10, 2025

  46. [66]

    Ramsundar, P

    B. Ramsundar, P. Eastman, P. Walters, V . Pande, K. Leswing, and Z. Wu, Deep Learning for the Life Sciences: Applying Deep Learning to Genomics, Microscopy, Drug Discovery, and More. O’Reilly Media, Inc., 2019

  47. [67]

    High-resolution de novo structure prediction from primary sequence,

    R. Wu, F. Ding, R. Wang, R. Shen, X. Zhang, S. Luo, C. Su, Z. Wu, Q. Xie, B. Berger, and J. Ma, “High-resolution de novo structure prediction from primary sequence,” bioRxiv, 2022. [Online]. Available: https://doi.org/10.1101/2022.07.26.501554

  48. [68]

    Protgpt2 is a deep unsupervised language model for protein design,

    N. Ferruz, S. Schmidt, and B. H ¨ocker, “Protgpt2 is a deep unsupervised language model for protein design,” Nature Communications, vol. 13, no. 1, p. 4348, 2022. [Online]. Available: https://www.nature.com/ articles/s41467-022-31900-5

  49. [69]

    Large language models generate functional protein sequences across diverse families,

    A. Madani, B. Krause, E. R. Greene, S. Subramanian, B. P. Mohr, J. M. Holton, J. L. Olmos, C. Xiong, Z. Z. Sun, R. Socher et al. , “Large language models generate functional protein sequences across diverse families,” Nature Biotechnology, vol. 41, no. 8, pp. 1099–1106, 2023

  50. [70]

    De novo design of protein structure and function with rfdiffusion,

    J. L. Watson, D. Juergens, N. R. Bennett, B. L. Trippe, J. Yim, H. E. Eisenach, W. Ahern, A. J. Borst, R. J. Ragotte, L. F. Milles et al., “De novo design of protein structure and function with rfdiffusion,” Nature, vol. 620, no. 7976, pp. 1089–1100, 2023

  51. [72]

    Mftrans: A multi-feature transformer network for protein secondary structure prediction,

    Y . Chen, G. Chen, and C. Y .-C. Chen, “Mftrans: A multi-feature transformer network for protein secondary structure prediction,” Inter- national Journal of Biological Macromolecules , vol. 267, p. 131311, 2024

  52. [73]

    Transconv: Convolution-infused transformer for protein secondary structure prediction,

    S. Das, S. Ghosh, and N. Jana, “Transconv: Convolution-infused transformer for protein secondary structure prediction,” Springer nature link, 2025

  53. [74]

    De novo atomic protein structure modeling for cryoem density maps using 3d transformer and hmm,

    N. Giri and J. Cheng, “De novo atomic protein structure modeling for cryoem density maps using 3d transformer and hmm,” Nature Communications, vol. 15, no. 1, p. 5511, 2024. [Online]. Available: https://doi.org/10.1038/s41467-024-49647-6

  54. [75]

    A critical review of five machine learning-based algorithms for predicting protein stability changes upon mutation,

    J. Fang, “A critical review of five machine learning-based algorithms for predicting protein stability changes upon mutation,” Briefings in Bioinformatics, vol. 21, no. 4, pp. 1285–1292, 07 2019. [Online]. Available: https://doi.org/10.1093/bib/bbz071

  55. [76]

    Gpcrpred: an svm-based method for prediction of families and subfamilies of g-protein coupled receptors,

    M. Bhasin and G. P. S. Raghava, “Gpcrpred: an svm-based method for prediction of families and subfamilies of g-protein coupled receptors,” Nucleic Acids Research , vol. 32, no. suppl 2, pp. W383–W389, 07

  56. [77]

    Multi-scale deep learning for the imbalanced multi-label protein subcellular localization prediction based on im- munohistochemistry images,

    F. Wang and L. Wei, “Multi-scale deep learning for the imbalanced multi-label protein subcellular localization prediction based on im- munohistochemistry images,” Bioinformatics, vol. 38, no. 16, pp. 4019– 4027, 2022

  57. [78]

    Prog-sol: Predicting protein solubility using protein embeddings and dual-graph convolutional networks,

    G. Li, N. Zhang, and L. Fan, “Prog-sol: Predicting protein solubility using protein embeddings and dual-graph convolutional networks,” ACS Omega, vol. 10, no. 4, pp. 3910–3916, 2025. [Online]. Available: https://doi.org/10.1021/acsomega.4c09688

  58. [79]

    Deep-probind: Bind- ing protein prediction with transformer-based deep learning model,

    S. Khan, S. Noor, H. H. Awan, S. Iqbal et al. , “Deep-probind: Bind- ing protein prediction with transformer-based deep learning model,” Springer nature link , 2025

  59. [80]

    Insights into the inner workings of transformer models for protein function prediction,

    M. Wenzel, E. Gr ¨uner, and N. Strodthoff, “Insights into the inner workings of transformer models for protein function prediction,” Bioinformatics, vol. 40, no. 3, p. btae031, 01 2024. [Online]. Available: https://doi.org/10.1093/bioinformatics/btae031

  60. [81]

    Segt-go: a graph transformer method based on ppi serialization and explanatory artificial intelligence for protein function prediction,

    Y . Wang, Y . Sun, B. Lin, and Others, “Segt-go: a graph transformer method based on ppi serialization and explanatory artificial intelligence for protein function prediction,” BMC Bioinformatics , vol. 26, p. 46, 2025

  61. [82]

    Integrating transformers and automl for protein function prediction,

    G. B. de Oliveira, H. Pedrini, and Z. Dias, “Integrating transformers and automl for protein function prediction,” in 2024 46th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC) , 2024, pp. 1–5

  62. [85]

    Gact-ppis: Prediction of protein-protein interaction sites based on graph structure and transformer network,

    L. Meng and H. Zhang, “Gact-ppis: Prediction of protein-protein interaction sites based on graph structure and transformer network,” International Journal of Biological Macromolecules , vol. 283, p. 137272, 2024. [Online]. Available: https://www.sciencedirect.com/ science/arti...

  63. [86]

    Tranp-b-site: A transformer enhanced method for prediction of binding sites of protein-protein interactions,

    S. H. Khan, H. Tayara, and K. T. Chong, “Tranp-b-site: A transformer enhanced method for prediction of binding sites of protein-protein interactions,” Measurement, vol. 251, p. 117227, 2025. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S026322412500586X

  64. [87]

    Tuna: An uncertainty- aware transformer model for sequence-based protein–protein interaction prediction,

    Y . S. Ko, J. Parkinson, C. Liu, and W. Wang, “Tuna: An uncertainty- aware transformer model for sequence-based protein–protein interaction prediction,” Briefings in Bioinformatics , 2024. [Online]. Available: https://doi.org/10.1093/bib/bbae359

  65. [88]

    Predicting protein-protein binding affinity with deep learning: A comparative analysis of cnn and transformer models,

    L. Chen, K. F. Ahmad Nasif, B. Deng, S. Niu, and C. Y . Xie, “Predicting protein-protein binding affinity with deep learning: A comparative analysis of cnn and transformer models,” in 2024 IEEE 36th International Conference on Tools with Artificial Intelligence (ICTAI), 2024, ...

  66. [89]

    A review of transformers in drug discovery and beyond,

    J. Jiang, L. Chen, L. Ke, B. Dou, C. Zhang, H. Feng, Y . Zhu, H. Qiu, B. Zhang, and G. Wei, “A review of transformers in drug discovery and beyond,” Journal of Pharmaceutical Analysis , p. 101081, 2024. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S2...

  67. [90]

    Mol-bert: An effective molecular representation with bert for molecular property prediction,

    J. Li and X. Jiang, “Mol-bert: An effective molecular representation with bert for molecular property prediction,” Wireless Communications and Mobile Computing , vol. 2021, no. 1, p. 7181815, 2021

  68. [91]

    Molecular generative graph neural networks for drug discovery,

    P. Bongini, M. Bianchini, and F. Scarselli, “Molecular generative graph neural networks for drug discovery,” Neurocomputing, vol. 450, pp. 242–252, 2021

  69. [93]

    Integrating transformer-based language model for drug discovery,

    R. R. Kotkondawar, S. R. Sutar, A. W. Kiwelekar, and V . J. Kadam, “Integrating transformer-based language model for drug discovery,” in 2024 11th International Conference on Computing for Sustainable Global Development (INDIACom) , 2024, pp. 1096–1101

  70. [94]

    Integrating transformers and many-objective optimization for drug design,

    N. Aksamit, J. Hou, Y . Li, and Others, “Integrating transformers and many-objective optimization for drug design,” BMC Bioinformatics , vol. 25, p. 208, 2024

  71. [95]

    Transformers and large language models for chemistry and drug discovery,

    A. M. Bran and P. Schwaller, “Transformers and large language models for chemistry and drug discovery,” in Drug Development Supported by Informatics, H. Satoh, K. Funatsu, and H. Yamamoto, Eds. Springer, Singapore, 2024

  72. [96]

    Sspro/accpro 5: almost perfect prediction of protein secondary structure and relative solvent accessibility using profiles, machine learning and structural similarity,

    C. N. Magnan and P. Baldi, “Sspro/accpro 5: almost perfect prediction of protein secondary structure and relative solvent accessibility using profiles, machine learning and structural similarity,” Bioinformatics, vol. 30, no. 18, pp. 2592–2597, 2014

  73. [97]

    Uniprot: the universal protein knowledgebase in 2021,

    Anonymous, “Uniprot: the universal protein knowledgebase in 2021,” Nucleic acids research, vol. 49, no. D1, pp. D480–D489, 2021

  74. [98]

    Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets,

    M. Steinegger and J. S ¨oding, “Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets,” Nature biotechnology, vol. 35, no. 11, pp. 1026–1028, 2017

  75. [99]

    Confold2: improved contact-driven ab initio protein structure modeling,

    B. Adhikari and J. Cheng, “Confold2: improved contact-driven ab initio protein structure modeling,” BMC bioinformatics , vol. 19, pp. 1–5, 2018

  76. [101]

    Language models enable zero-shot prediction of the effects of mutations on protein function,

    J. Meier, R. Rao, R. Verkuil, J. Liu, T. Sercu, and A. Rives, “Language models enable zero-shot prediction of the effects of mutations on protein function,” Advances in neural information processing systems , vol. 34, pp. 29 287–29 303, 2021

  77. [102]

    Energy and policy con- siderations for modern deep learning research,

    E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy con- siderations for modern deep learning research,” in Proceedings of the AAAI conference on artificial intelligence , vol. 34, no. 09, 2020, pp. 13 693–13 696

  78. [103]

    Deep learning in proteomics,

    B. Wen, W.-F. Zeng, Y . Liao, Z. Shi, S. R. Savage, W. Jiang, and B. Zhang, “Deep learning in proteomics,” Proteomics, vol. 20, no. 21- 22, p. 1900335, 2020

  79. [104]

    Deep multi-view learning methods: A review,

    X. Yan, S. Hu, Y . Mao, Y . Ye, and H. Yu, “Deep multi-view learning methods: A review,” Neurocomputing, vol. 448, pp. 106–129, 2021

  80. [106]

    Towards a rigorous science of inter- pretable machine learning,

    F. Doshi-Velez and B. Kim, “Towards a rigorous science of inter- pretable machine learning,” arXiv preprint arXiv:1702.08608 , 2017

  81. [107]

    Explainable artificial intelligence: Understanding, vi- sualizing and interpreting deep learning models,

    W. Samek, “Explainable artificial intelligence: Understanding, vi- sualizing and interpreting deep learning models,” arXiv preprint arXiv:1708.08296, 2017

  82. [108]

    Transformer protein language models are unsupervised structure learners,

    R. Rao, J. Meier, T. Sercu, S. Ovchinnikov, and A. Rives, “Transformer protein language models are unsupervised structure learners,” Biorxiv, pp. 2020–12, 2020

  83. [109]

    Protein tertiary structure prediction and refinement using deep learning and rosetta in casp14,

    I. Anishchenko, M. Baek, H. Park, N. Hiranuma, D. E. Kim, J. Dau- paras, S. Mansoor, I. R. Humphreys, and D. Baker, “Protein tertiary structure prediction and refinement using deep learning and rosetta in casp14,” Proteins: Structure, Function, and Bioinformatics , vol. 89, no...

  84. [110]

    Deepaffinity: interpretable deep learning of compound–protein affinity through unified recurrent and convolutional neural networks,

    M. Karimi, D. Wu, Z. Wang, and Y . Shen, “Deepaffinity: interpretable deep learning of compound–protein affinity through unified recurrent and convolutional neural networks,” Bioinformatics, vol. 35, no. 18, pp. 3329–3338, 2019

  85. [111]

    Grandmaster level in starcraft ii using multi-agent reinforcement learning,

    O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev et al. , “Grandmaster level in starcraft ii using multi-agent reinforcement learning,” nature, vol. 575, no. 7782, pp. 350–354, 2019

  86. [112]

    Toward a shared vision for cancer genomic data,

    R. L. Grossman, A. P. Heath, V . Ferretti, H. E. Varmus, D. R. Lowy, W. A. Kibbe, and L. M. Staudt, “Toward a shared vision for cancer genomic data,” New England Journal of Medicine , vol. 375, no. 12, pp. 1109–1112, 2016

  87. [114]

    Prottrans: Toward understanding the language of life through self-supervised learning,

    A. Elnaggar, M. Heinzinger, C. Dallago, G. Rehawi, Y . Wang, L. Jones, T. Gibbs, T. Feher, C. Angerer, M. Steinegger et al. , “Prottrans: Toward understanding the language of life through self-supervised learning,” IEEE transactions on pattern analysis and machine intel- ligen...

  88. [115]

    Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences,

    A. Rives, J. Meier, T. Sercu, S. Goyal, Z. Lin, J. Liu, D. Guo, M. Ott, C. L. Zitnick, J. Ma et al., “Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences,” Proceedings of the National Academy of Sciences , vol. 118, no. ...

  89. [116]

    Accurate prediction of protein structures and interactions using a three- track neural network,

    M. Baek, F. DiMaio, I. Anishchenko, J. Dauparas, S. Ovchinnikov, G. R. Lee, J. Wang, Q. Cong, L. N. Kinch, R. D. Schaeffer et al. , “Accurate prediction of protein structures and interactions using a three- track neural network,” Science, vol. 373, no. 6557, pp. 871–876, 2021

  90. [117]

    Molmol: a program for dis- play and analysis of macromolecular structures,

    R. Koradi, M. Billeter, and K. W ¨uthrich, “Molmol: a program for dis- play and analysis of macromolecular structures,” Journal of molecular graphics, vol. 14, no. 1, pp. 51–55, 1996

  91. [118]

    Transformer architecture and attention mech- anisms in genome data analysis: a comprehensive review,

    S. R. Choi and M. Lee, “Transformer architecture and attention mech- anisms in genome data analysis: a comprehensive review,” Biology, vol. 12, no. 7, p. 1033, 2023

  92. [119]

    A unified approach to interpreting model predictions,

    S. Lundberg, “A unified approach to interpreting model predictions,” arXiv preprint arXiv:1705.07874 , 2017

  93. [120]

    Profile prediction: An alignment-based pre-training task for protein sequence models,

    P. Sturmfels, J. Vig, A. Madani, and N. F. Rajani, “Profile prediction: An alignment-based pre-training task for protein sequence models,” arXiv preprint arXiv:2012.00195 , 2020

  94. [121]

    Longformer: The long- document transformer,

    I. Beltagy, M. E. Peters, and A. Cohan, “Longformer: The long- document transformer,” arXiv preprint arXiv:2004.05150 , 2020

  95. [122]

    Transformer-xl: Attentive language models beyond a fixed- length context,

    Z. Dai, “Transformer-xl: Attentive language models beyond a fixed- length context,” arXiv preprint arXiv:1901.02860 , 2019

  96. [123]

    Distilling the knowledge in a neural network,

    G. Hinton, “Distilling the knowledge in a neural network,” arXiv preprint arXiv:1503.02531, 2015

  97. [124]

    Carbon emissions and large neural network training,

    D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean, “Carbon emissions and large neural network training,” arXiv preprint arXiv:2104.10350, 2021

  98. [125]

    Highly accurate protein structure prediction with alphafold,

    J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. ˇZ´ıdek, A. Potapenko et al., “Highly accurate protein structure prediction with alphafold,” nature, vol. 596, no. 7873, pp. 583–589, 2021

  99. [2004]

    Available: https://doi.org/10.1093/nar/gkh416

    [Online]. Available: https://doi.org/10.1093/nar/gkh416

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.