REVIEW 7 major objections 6 minor 107 references
Transformers in Protein: A Survey
T0 review · 7 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This survey claims that Transformer models now span the four core protein-research domains, with over 100 studies synthesized into one domain-oriented framework.
desk verdict A well-organized survey of Transformer methods in protein science, but its systemic citation-reference mismatches break the central promise of a reliable review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing machinery of the paper is the domain-oriented classification system, which divides the Transformer-in-protein literature into structure prediction, function prediction, protein–protein interactions, and drug discovery. Within that frame the paper uses a taxonomy of self-attention mechanisms—single-head versus multi-head, spatial and graph-based attention, SE(3)-equivariant attention, and hybrid CNN–Transformer designs—to characterize the model families, and it uses bibliometric analysis (Fig. 1) and curated resource tables (Tables II and III) to support the claims of growth and reproducibility. No new algorithm is proposed; the machinery is the synthesis itself.
What would settle it
Reproducing Fig. 1 with a stated search query, access date, and inclusion criteria, and then checking the claims against their cited sources—for example, whether reference [71] is a protein structure model, whether [65] reports protein-ligand fine-tuning, and whether [58] concerns ESM-Fold—would settle whether the survey's descriptions and trend data are reliable.
Extended reading notes
Core claim
The paper's central claim is that the Transformer's application to protein informatics has reached a state where a comprehensive, domain-oriented synthesis is both possible and needed, and that this synthesis shows the architecture reshaping the field. It surveys more than 100 studies and organizes them into four application domains: protein structure prediction, protein function prediction, protein–protein interaction analysis, and drug discovery/target identification. It further claims that the Transformer's self-attention mechanism explains the gains—because protein sequence, structure, and function depend on distal residue interactions—and that pre-trained protein language models such as AlphaFold, ESM-Fold, ProtTrans, and ProteinBERT are the vehicles through which those gains appear. The paper also asserts that curating datasets and code repositories is an essential part of the contribution, since reproducibility and benchmarking depend on them, and that future progress will come from multimodal integration, hybrid physics-informed modeling, efficiency improvements, and interpretability.
Load-bearing premise
The survey's usefulness depends on each cited reference actually supporting the statement it is attached to, and on the bibliometric counts in Fig. 1 coming from a clearly defined and correctly executed search; a sympathetic reader has to take that reliability on faith.
Editorial extensions
If this is right
- Newcomers to protein informatics can use the survey as a single entry point that explains Transformer fundamentals and then maps specific models onto the four application domains.
- The curated dataset and code tables give researchers a concrete starting set of resources for reproducing leading models or benchmarking new ones.
- The paper's list of persistent bottlenecks—quadratic attention cost, biased and sparse datasets, limited interpretability, and difficulty with novel folds—defines a research agenda for the next generation of protein Transformers.
- The four-domain grouping makes explicit that the same self-attention machinery is being reused across structure, function, interaction, and drug-discovery tasks, which supports cross-domain transfer of methods.
- If the review's coverage is accurate, the field has shifted from asking whether Transformers can help with proteins to asking how to make them efficient, interpretable, and multimodal.
Reading between the lines
- The static tables of models and repositories will age quickly given the field's publication rate; a living, community-maintained version would better serve the reproducibility goal the paper sets.
- The paper's taxonomy of attention mechanisms could be turned into a quantitative meta-analysis—grouping published models by architecture family and comparing benchmark results—which the survey itself does not attempt.
- The emphasis on multimodal integration suggests that the next evaluation frontier for protein Transformers will be joint sequence–structure–function modeling rather than single-task accuracy.
- The domain-oriented grouping could seed a benchmark suite that compares Transformer variants within each domain, something the paper lists resources for but does not build.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of Transformer applications in protein science, organized by application domain (structure prediction, function prediction, protein-protein interactions, drug discovery) and supplemented by tables of model derivatives, open-source implementations, and datasets. The authors state that they reviewed over 100 studies and aim to provide a consolidated, accurate foundation for the field. The survey covers major models such as AlphaFold, RoseTTAFold, ESM-Fold, OmegaFold, ProtTrans, ProteinBERT, and several newer 2024-2025 methods, and it discusses advantages, challenges, and future directions. The central claim is that the paper provides a reliable domain-oriented review with curated resources for reproducibility and benchmarking.
Significance. If the survey's descriptions and citations were reliable, it would serve as a useful entry point for researchers applying Transformers to protein problems, particularly because it assembles model names, resource links, and dataset information in one place. The paper makes a genuine effort to organize a large and fast-moving literature, and the inclusion of resource tables (Tables II and III) and a structured discussion of challenges is valuable in principle. However, the significance of the contribution is currently undermined by the accuracy problems detailed below, which affect the very basis of a survey: readers cannot trust the claims attached to the references. The survey is therefore potentially useful after substantial correction, but in its present form its reliability as a reference work is not established.
major comments (7)
- [Section III-A.7 and Table I] The description of 'GraphTrans' as a model that integrates graph neural networks with Transformer layers for protein complex prediction is not supported by the cited reference [71], which is a Winter Simulation Conference paper on a software system for network conversions, not a protein structure prediction model. This is a load-bearing error because the survey claims to provide accurate descriptions of Transformer derivatives; a reader consulting the citation would find nothing about proteins. The same unrelated reference [84] is also used in Section III-C.3 to support 'Graph Attention Networks (GATs)', where the canonical GAT paper is not cited. Please replace these with correct references or remove the unsupported entries from Table I and the text.
- [Section III-D.1] The claim that 'ChemBERTa is a prominent model in this area, fine-tuned on protein-ligand interaction datasets' with 'superior performance compared to traditional methods like molecular docking and molecular dynamics simulations in predicting binding affinities' is not present in the cited paper [65], which describes self-supervised pretraining on SMILES strings for molecular property prediction. No fine-tuning on protein-ligand datasets or docking comparisons appear in that 2020 arXiv paper. This is a concrete misattribution in a central section of the survey and should be corrected by either citing the relevant subsequent work or removing the unsupported claims.
- [Section III-C.1] The text credits DeepPPI with 'an attention mechanism that enables the model to focus on biologically relevant sequence motifs,' but the cited 2017 paper [83] (Du et al., 'Deepppi') uses fully connected deep neural networks on sequence-derived features and does not contain an attention mechanism. The description of the model is therefore factually incorrect. This matters because the survey is meant to summarize model architectures; a reader relying on the review would form a wrong idea of DeepPPI's design. Please revise the description to match the cited source.
- [Section III-A.3 and Table I] The ESM-Fold passage cites [58] for the claim that ESM-Fold predicts structures for proteins with 'minimal or no homology to known structures' and for generalizability to orphan proteins. Reference [58] is a neurophysiology study on suppression of alpha-band power during emotional distraction, which has no connection to protein structure prediction. This is another example of a systematic reference-reference mismatch, not an isolated typo. Given that such mismatches appear in at least four separate model descriptions (GraphTrans, ChemBERTa, DeepPPI, ESM-Fold), the authors must audit every citation in the manuscript against the claims it supports before the survey can be considered trustworthy.
- [Section II-D.2 and Section II-D.3] Two background paragraphs contain citation mismatches that, while not about individual models, still affect the survey's reliability as a reference. In Section II-D.2, the claim that 'recent advancements have highlighted the growing efficacy of transformer-based models in predicting stability' is supported by [50] (Rost and Sander, 1993) and [51] (Scheraga et al., 2007), neither of which is a transformer-based stability predictor. In Section II-D.3, the claim about ProtTrans generalization is cited to [52] (Magnan and Baldi, SSpro) and [53] (Peitsch, SWISS-MODEL), again unrelated. These references should be replaced with actual transformer-based prediction papers or the claims should be softened.
- [Table II and Section VII] The paper's promise to curate open-source resources for reproducibility is undermined by errors in Table II. The ProtGPT2 repository URL points to a Hugging Face trainer documentation page rather than the ProtGPT2 model repository, and the ProteinBERT URL ('https://github.com/nadavbra/protein bert') contains a space and is not a valid URL. Table III also gives dataset sizes and access methods without specifying whether sizes are sequence counts or storage sizes (e.g., '30,051' for ESM-Fold UR50 is unclear). Since reproducibility support is part of the central contribution, these entries must be verified and corrected.
- [Fig. 1 and Section I] The bibliometric analysis in Fig. 1 is presented as evidence of the field's growth, but the caption and text do not provide the search query, database access date, inclusion criteria, or subfield classification rules used to obtain the data from the Web of Science Core Collection. Without this information, the trend claims (e.g., the surge in publications and the distribution across journals and subfields) cannot be reproduced or validated. Please add a methodology paragraph describing the search and filtering process.
minor comments (6)
- [Section III-B.1] The text uses 'ProtBERT' and 'ProteinBERT' interchangeably; the cited work [63] is ProteinBERT by Brandes et al., which is a distinct model from ProtBERT in the ProtTrans collection. Please use a consistent name to avoid confusing the two.
- [References] Several references are duplicated: [37] and [54] are the same AlphaFold paper, [40] and [55] are the same ProtTrans paper, [45] and [7] are the same Rives et al. paper, and [57] duplicates [29]. Please consolidate the bibliography or adjust the citation numbers.
- [Section III-D.6] The subsection 'RL-based Transformer Models for Molecule Generation' contains no citation at all; in a survey that promises to cover over 100 studies, each described approach should be linked to a reference.
- [Section II (opening paragraph)] The sentence 'The paper is organized as follows...' appears at the top of Section II, after the introduction; this organizational roadmap would be more naturally placed at the end of Section I.
- [Section III-A.1] The statement that AlphaFold 'achieved a median Global Distance Test (GDT) score of 92.4' lacks a precise target (e.g., CASP14 free-modeling domains); please clarify the metric scope for accuracy.
- [Fig. 2 caption] The caption's reference to 'the third row-left block' is confusing; please refer to the specific subfigure or panel labels.
Circularity Check
No circularity: the survey contains no derivation chain whose output is equivalent to its input by construction.
full rationale
This paper is a domain survey rather than a derivation-driven work. It makes no fitted-parameter predictions, proves no theorems from assumptions, and does not reduce any claimed result to its own inputs by construction. The only self-citations are references [4], [5], and [6] in the introduction, used to illustrate that transformer models have been adopted across disciplines; that claim is ambient context and is not load-bearing for the survey's organization, resource tables, or model summaries. The central deliverable is a literature synthesis whose accuracy rests on the correspondence between cited external sources and the claims attached to them. Issues such as the mismatched references identified in Section III-A.7, III-C.1, and III-D.1 (e.g., refs [71], [65], [58]) are citation-accuracy defects, not circularity: a mistaken attribution does not make a statement equivalent to its input by definition, nor does it fit a parameter and then rename the fit as a prediction. No circular step of any enumerated kind is present, so the appropriate finding is a score of 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The Web of Science Core Collection queries used to generate Fig. 1 are correctly constructed and capture the relevant Transformer-in-protein literature.
- domain assumption Cited references accurately support the descriptions of each model and method presented in the survey.
Cite this review
Pith. "Pith review of Transformers in Protein: A Survey." pith.science (2026). https://pith.science/paper/YGS7E6Q6
@misc{pith2026250520098,
author = {Pith},
title = {Pith review of: Transformers in Protein: A Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/YGS7E6Q6}},
note = {Machine review of arXiv:2505.20098}
}
read the original abstract
As protein informatics advances rapidly, the demand for enhanced predictive accuracy, structural analysis, and functional understanding has intensified. Transformer models, as powerful deep learning architectures, have demonstrated unprecedented potential in addressing diverse challenges across protein research. However, a comprehensive review of Transformer applications in this field remains lacking. This paper bridges this gap by surveying over 100 studies, offering an in-depth analysis of practical implementations and research progress of Transformers in protein-related tasks. Our review systematically covers critical domains, including protein structure prediction, function prediction, protein-protein interaction analysis, functional annotation, and drug discovery/target identification. To contextualize these advancements across various protein domains, we adopt a domain-oriented classification system. We first introduce foundational concepts: the Transformer architecture and attention mechanisms, categorize Transformer variants tailored for protein science, and summarize essential protein knowledge. For each research domain, we outline its objectives and background, critically evaluate prior methods and their limitations, and highlight transformative contributions enabled by Transformer models. We also curate and summarize pivotal datasets and open-source code resources to facilitate reproducibility and benchmarking. Finally, we discuss persistent challenges in applying Transformers to protein informatics and propose future research directions. This review aims to provide a consolidated foundation for the synergistic integration of Transformer and protein informatics, fostering further innovation and expanded applications in the field.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[65]
Chemberta: large-scale self-supervised pretraining for molecular property prediction,
S. Chithrananda, G. Grand, and B. Ramsundar, “Chemberta: large-scale self-supervised pretraining for molecular property prediction,” arXiv preprint arXiv:2010.09885, 2020
arXiv 2010
-
[83]
Deepppi: boosting prediction of protein–protein interactions with deep neural networks,
X. Du, S. Sun, C. Hu, Y . Yao, Y . Yan, and Y . Zhang, “Deepppi: boosting prediction of protein–protein interactions with deep neural networks,” Journal of chemical information and modeling, vol. 57, no. 6, pp. 1499– 1510, 2017
work page 2017
-
[58]
Suppression of alpha-band power underlies exogenous attention to emotional distractors,
L. Arana, M. Melc ´on, D. Kessel, S. Hoyos, J. Albert, L. Carreti ´e, and A. Capilla, “Suppression of alpha-band power underlies exogenous attention to emotional distractors,” Psychophysiology, vol. 59, no. 9, p. e14051, 2022
work page 2022
-
[84]
H. L. Carscadden, L. Machi, C. J. Kuhlman, D. Machi, and S. Ravi, “Graphtrans: a software system for network conversions for simulation, structural analysis, and graph operations,” in 2021 Winter Simulation Conference (WSC). IEEE, 2021, pp. 1–12
work page 2021
-
[50]
Prediction of protein secondary structure at better than 70% accuracy,
B. Rost and C. Sander, “Prediction of protein secondary structure at better than 70% accuracy,” Journal of molecular biology , vol. 232, no. 2, pp. 584–599, 1993
work page 1993
-
[51]
Protein-folding dynamics: overview of molecular simulation techniques,
H. A. Scheraga, M. Khalili, and A. Liwo, “Protein-folding dynamics: overview of molecular simulation techniques,” Annu. Rev. Phys. Chem., vol. 58, no. 1, pp. 57–83, 2007
work page 2007
-
[53]
M. C. Peitsch, “Protein modeling by e-mail,” Bio/technology, vol. 13, no. 7, pp. 658–660, 1995
work page 1995
-
[1]
Attention is all you need,
A. Vaswani, “Attention is all you need,” Advances in Neural Informa- tion Processing Systems , 2017
2017
Show all 107 references
-
[2]
Bert: Pretraining of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pretraining of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologie...
2019
-
[3]
Language models are few-shot learners,
T. B. Brown, “Language models are few-shot learners,” arXiv preprint arXiv:2005.14165, 2020
2005 arXiv
-
[4]
Pmanet: Malicious url detection via post-trained language model guided multi-level feature attention network,
R. Liu, Y . Wang, H. Xu, Z. Qin, F. Zhang, Y . Liu, and Z. Cao, “Pmanet: Malicious url detection via post-trained language model guided multi-level feature attention network,” Information Fusion, vol. 113, p. 102638, 2025
2025
-
[5]
Vul- lmgnns: Fusing language models and online-distilled graph neural networks for code vulnerability detection,
R. Liu, Y . Wang, H. Xu, J. Sun, F. Zhang, P. Li, and Z. Guo, “Vul- lmgnns: Fusing language models and online-distilled graph neural networks for code vulnerability detection,” Information Fusion , vol. 115, p. 102748, 2025
2025
-
[6]
Ethereum fraud detection via joint transaction language model and graph representation learning,
J. Sun, Y . Jia, Y . Wang, Y . Tian, and S. Zhang, “Ethereum fraud detection via joint transaction language model and graph representation learning,” Information Fusion, vol. 120, p. 103074, 2025
2025
-
[7]
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences,
A. Rives, J. Meier, T. Sercu, S. Goyal, Z. Lin, J. Liu, D. Guo, M. Ott, C. L. Zitnick, J. Ma et al., “Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences,” Proceedings of the National Academy of Sciences , 2021
2021
-
[8]
Highly accurate protein structure prediction with alphafold,
J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. ˇZ´ıdek, A. Potapenko et al., “Highly accurate protein structure prediction with alphafold,” nature, 2021
2021
-
[9]
Evaluating protein transfer learning with tape,
R. Rao, N. Bhattacharya, N. Thomas, Y . Duan, P. Chen, J. Canny, P. Abbeel, and Y . Song, “Evaluating protein transfer learning with tape,” Advances in neural information processing systems , 2019
2019
-
[10]
Evolutionary velocity with protein language models predicts evolutionary dynamics of diverse proteins,
B. L. Hie, K. K. Yang, and P. S. Kim, “Evolutionary velocity with protein language models predicts evolutionary dynamics of diverse proteins,” Cell Systems, 2022
2022
-
[11]
Language models of protein sequences at the scale of evolution enable accurate structure prediction,
Z. Lin, H. Akin, R. Rao, B. Hie, Z. Zhu, W. Lu, A. dos Santos Costa, M. Fazel-Zarandi, T. Sercu, S. Candido et al. , “Language models of protein sequences at the scale of evolution enable accurate structure prediction,” BioRxiv, 2022
2022
-
[12]
Deep learning in bioinformatics,
S. Min, B. Lee, and S. Yoon, “Deep learning in bioinformatics,” Briefings in bioinformatics , 2017
2017
-
[13]
Machine learning solutions for predicting protein–protein interactions,
R. Casadio, P. L. Martelli, and C. Savojardo, “Machine learning solutions for predicting protein–protein interactions,” Wiley Interdis- ciplinary Reviews: Computational Molecular Science , 2022
2022
-
[14]
Unified rational protein engineering with sequence-based deep representation learning,
E. C. Alley, G. Khimulya, S. Biswas, M. AlQuraishi, and G. M. Church, “Unified rational protein engineering with sequence-based deep representation learning,” Nature methods, 2019
2019
-
[15]
Protein- bert: a universal deep-learning model of protein sequence and function,
N. Brandes, D. Ofer, Y . Peleg, N. Rappoport, and M. Linial, “Protein- bert: a universal deep-learning model of protein sequence and function,” Bioinformatics, 2022
2022
-
[16]
Deep learning in proteomics,
B. Wen, W.-F. Zeng, Y . Liao, Z. Shi, S. R. Savage, W. Jiang, and B. Zhang, “Deep learning in proteomics,” Proteomics, 2020
2020
-
[17]
Transformer- based deep learning for predicting protein properties in the life sci- ences,
A. Chandra, L. T ¨unnermann, T. L ¨ofstedt, and R. Gratz, “Transformer- based deep learning for predicting protein properties in the life sci- ences,” Elife, 2023
2023
-
[18]
Machine learning: its challenges and opportunities in plant system biology,
M. Hesami, M. Alizadeh, A. M. P. Jones, and D. Torkamaneh, “Machine learning: its challenges and opportunities in plant system biology,” Applied Microbiology and Biotechnology , 2022
2022
-
[19]
Artificial intelligence in the prediction of protein–ligand interactions: recent advances and future directions,
A. Dhakal, C. McKay, J. J. Tanner, and J. Cheng, “Artificial intelligence in the prediction of protein–ligand interactions: recent advances and future directions,” Briefings in Bioinformatics , 2022
2022
-
[20]
Roberta: A robustly optimized bert pretraining approach,
Y . Liu, M. Ott, N. Goyal, J. Du et al., “Roberta: A robustly optimized bert pretraining approach,” arXiv preprint, 2019
2019
-
[21]
Linformer: Self- attention with linear complexity,
S. Wang, B. Z. Li, M. Khabsa, H. Fang, and H. Ma, “Linformer: Self- attention with linear complexity,” arXiv preprint, 2020
2020
-
[22]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929 , 2020
2010 arXiv
-
[23]
On the turing completeness of modern neural network architectures,
J. Perez, J. Marinkovic, and P. Barcelo, “On the turing completeness of modern neural network architectures,” in International Conference on Learning Representations , 2018
2018
-
[24]
On the relationship between self-attention and convolutional layers,
J.-B. Cordonnier, A. Loukas, and M. Jaggi, “On the relationship between self-attention and convolutional layers,” in International Con- ference on Learning Representations , 2019
2019
-
[25]
Deformable convolutional networks,
J. Dai, H. Qi, Y . Xiong, Z. Li, G. Zhang, H. Hu, and Y . Wei, “Deformable convolutional networks,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2017
2017
-
[26]
Prottrans: Towards cracking the language of life’s code through self-supervised deep learning and high performance computing,
A. Elnaggar, M. Heinzinger, C. Dallago, G. Rehawi, W. Yu, L. Jones, T. Gibbs, T. Feher, C. Angerer, M. Steinegger et al. , “Prottrans: Towards cracking the language of life’s code through self-supervised deep learning and high performance computing,” IEEE Transactions on Patte...
2021
-
[28]
Peptidebert: A pre-trained model for peptide representation learning,
X. Li, X. Zhang, S. Wang, X. Yu, S. Pan, J. Wu, and et al., “Peptidebert: A pre-trained model for peptide representation learning,” arXiv preprint arXiv:2309.03099, 2023. [Online]. Available: https://arxiv.org/abs/2309.03099
2023 arXiv
-
[29]
Language models of protein sequences at the scale of evolution enable accurate structure prediction,
Z. Lin, H. Akin, R. Rao, B. Hie, Z. Zhu, W. Lu, T. Sercu, and A. Rives, “Language models of protein sequences at the scale of evolution enable accurate structure prediction,” bioRxiv, 2022
2022
-
[30]
Msa transformer,
R. Rao, J. Liu, R. Verkuil, J. Meier, J. Canny, P. Abbeel, T. Sercu, and A. Rives, “Msa transformer,” bioRxiv, 2021
2021
-
[31]
Protein representation learning with cross-modal contrastive pretraining,
C. Lu, L. Xu, Y . Wu, Y . Liu, and et al., “Protein representation learning with cross-modal contrastive pretraining,” arXiv preprint arXiv:2401.14819, 2024. [Online]. Available: https://arxiv.org/abs/ 2401.14819
2024 arXiv
-
[32]
Alberts, D
B. Alberts, D. Bray, K. Hopkin, A. D. Johnson, J. Lewis, M. Raff, K. Roberts, and P. Walter, Essential cell biology . Garland Science, 2015
2015
-
[33]
Prediction of protein conformation,
P. Y . Chou and G. D. Fasman, “Prediction of protein conformation,” Biochemistry, vol. 13, no. 2, pp. 222–245, 1974
1974
-
[35]
Analysis of the accuracy and implications of simple methods for predicting the secondary structure of globular proteins,
J. Garnier, D. J. Osguthorpe, and B. Robson, “Analysis of the accuracy and implications of simple methods for predicting the secondary structure of globular proteins,” Journal of molecular biology , vol. 120, no. 1, pp. 97–120, 1978
1978
-
[36]
Modeling aspects of the language of life through transfer-learning protein sequences,
M. Heinzinger, A. Elnaggar, Y . Wang, C. Dallago, D. Nechaev, F. Matthes, and B. Rost, “Modeling aspects of the language of life through transfer-learning protein sequences,” BMC bioinformatics, vol. 20, pp. 1–17, 2019
2019
-
[39]
Tm-align: a protein structure alignment algorithm based on the tm-score,
Y . Zhang and J. Skolnick, “Tm-align: a protein structure alignment algorithm based on the tm-score,”Nucleic acids research, vol. 33, no. 7, pp. 2302–2309, 2005
2005
-
[40]
Prottrans: towards cracking the language of life’s code through self-supervised learning,
A. Elnaggar, M. Heinzinger, C. Dallago, G. Rehawi, Y . Wang, L. Jones, T. Gibbs, T. Feher, C. Angerer, M. Steineggeret al., “Prottrans: towards cracking the language of life’s code through self-supervised learning,” IEEE Transactions on Pattern Analysis and Machine Intelligenc...
2021
-
[41]
The protein-folding problem, 50 years on,
K. A. Dill and J. L. MacCallum, “The protein-folding problem, 50 years on,” science, vol. 338, no. 6110, pp. 1042–1046, 2012
2012
-
[42]
Protein folding and misfolding,
C. M. Dobson, “Protein folding and misfolding,” Nature, vol. 426, no. 6968, pp. 884–890, 2003
2003
-
[48]
Goap: a generalized orientation-dependent, all-atom statistical potential for protein structure prediction,
H. Zhou and J. Skolnick, “Goap: a generalized orientation-dependent, all-atom statistical potential for protein structure prediction,” Biophys- ical journal, vol. 101, no. 8, pp. 2043–2052, 2011
2011
-
[49]
Protein secondary structure prediction based on position- specific scoring matrices,
D. T. Jones, “Protein secondary structure prediction based on position- specific scoring matrices,” Journal of molecular biology , vol. 292, no. 2, pp. 195–202, 1999
1999
-
[57]
Language models of protein sequences at the scale of evolution enable accurate structure prediction,
Z. Lin, H. Akin, R. Rao, B. Hie, Z. Zhu, W. Lu, A. dos Santos Costa, M. Fazel-Zarandi, T. Sercu, S. Candido et al. , “Language models of protein sequences at the scale of evolution enable accurate structure prediction,” BioRxiv, vol. 2022, p. 500902, 2022
2022
-
[60]
Improved protein structure prediction using potentials from deep learning,
A. W. Senior, R. Evans, J. Jumper, J. Kirkpatrick, L. Sifre, T. Green, C. Qin, A. ˇZ´ıdek, A. W. Nelson, A. Bridgland et al., “Improved protein structure prediction using potentials from deep learning,” Nature, vol. 577, no. 7792, pp. 706–710, 2020
2020
-
[61]
Evaluating protein transfer learning with tape,
R. Rao, N. Bhattacharya, N. Thomas, Y . Duan, P. Chen, J. Canny, P. Abbeel, and Y . Song, “Evaluating protein transfer learning with tape,” Advances in neural information processing systems , vol. 32, 2019
2019
-
[62]
A protein structure prediction approach leveraging transformer and cnn integration,
Y . Zhou, K. Tan, X. Shen, Z. He, and H. Zheng, “A protein structure prediction approach leveraging transformer and cnn integration,” arXiv preprint arXiv:2402.19095 , 2024. [Online]. Available: https: //arxiv.org/abs/2402.19095
2024 arXiv
-
[63]
Protein- bert: a universal deep-learning model of protein sequence and function,
N. Brandes, D. Ofer, Y . Peleg, N. Rappoport, and M. Linial, “Protein- bert: a universal deep-learning model of protein sequence and function,” Bioinformatics, vol. 38, no. 8, pp. 2102–2110, 2022
2022
-
[64]
Trans-morfs: A disordered protein predictor based on the transformer architecture,
C. Meng, Y . Shi, X. Fu, Q. Zou, and W. Han, “Trans-morfs: A disordered protein predictor based on the transformer architecture,” IEEE Journal of Biomedical and Health Informatics , pp. 1–10, 2025
2025
-
[66]
Ramsundar, P
B. Ramsundar, P. Eastman, P. Walters, V . Pande, K. Leswing, and Z. Wu, Deep Learning for the Life Sciences: Applying Deep Learning to Genomics, Microscopy, Drug Discovery, and More. O’Reilly Media, Inc., 2019
2019
-
[67]
High-resolution de novo structure prediction from primary sequence,
R. Wu, F. Ding, R. Wang, R. Shen, X. Zhang, S. Luo, C. Su, Z. Wu, Q. Xie, B. Berger, and J. Ma, “High-resolution de novo structure prediction from primary sequence,” bioRxiv, 2022. [Online]. Available: https://doi.org/10.1101/2022.07.26.501554
2022 doi
-
[68]
Protgpt2 is a deep unsupervised language model for protein design,
N. Ferruz, S. Schmidt, and B. H ¨ocker, “Protgpt2 is a deep unsupervised language model for protein design,” Nature Communications, vol. 13, no. 1, p. 4348, 2022. [Online]. Available: https://www.nature.com/ articles/s41467-022-31900-5
2022
-
[69]
Large language models generate functional protein sequences across diverse families,
A. Madani, B. Krause, E. R. Greene, S. Subramanian, B. P. Mohr, J. M. Holton, J. L. Olmos, C. Xiong, Z. Z. Sun, R. Socher et al. , “Large language models generate functional protein sequences across diverse families,” Nature Biotechnology, vol. 41, no. 8, pp. 1099–1106, 2023
2023
-
[70]
De novo design of protein structure and function with rfdiffusion,
J. L. Watson, D. Juergens, N. R. Bennett, B. L. Trippe, J. Yim, H. E. Eisenach, W. Ahern, A. J. Borst, R. J. Ragotte, L. F. Milles et al., “De novo design of protein structure and function with rfdiffusion,” Nature, vol. 620, no. 7976, pp. 1089–1100, 2023
2023
-
[72]
Mftrans: A multi-feature transformer network for protein secondary structure prediction,
Y . Chen, G. Chen, and C. Y .-C. Chen, “Mftrans: A multi-feature transformer network for protein secondary structure prediction,” Inter- national Journal of Biological Macromolecules , vol. 267, p. 131311, 2024
2024
-
[73]
Transconv: Convolution-infused transformer for protein secondary structure prediction,
S. Das, S. Ghosh, and N. Jana, “Transconv: Convolution-infused transformer for protein secondary structure prediction,” Springer nature link, 2025
2025
-
[74]
De novo atomic protein structure modeling for cryoem density maps using 3d transformer and hmm,
N. Giri and J. Cheng, “De novo atomic protein structure modeling for cryoem density maps using 3d transformer and hmm,” Nature Communications, vol. 15, no. 1, p. 5511, 2024. [Online]. Available: https://doi.org/10.1038/s41467-024-49647-6
2024 doi
-
[75]
A critical review of five machine learning-based algorithms for predicting protein stability changes upon mutation,
J. Fang, “A critical review of five machine learning-based algorithms for predicting protein stability changes upon mutation,” Briefings in Bioinformatics, vol. 21, no. 4, pp. 1285–1292, 07 2019. [Online]. Available: https://doi.org/10.1093/bib/bbz071
2019 doi
-
[76]
Gpcrpred: an svm-based method for prediction of families and subfamilies of g-protein coupled receptors,
M. Bhasin and G. P. S. Raghava, “Gpcrpred: an svm-based method for prediction of families and subfamilies of g-protein coupled receptors,” Nucleic Acids Research , vol. 32, no. suppl 2, pp. W383–W389, 07
-
[77]
Multi-scale deep learning for the imbalanced multi-label protein subcellular localization prediction based on im- munohistochemistry images,
F. Wang and L. Wei, “Multi-scale deep learning for the imbalanced multi-label protein subcellular localization prediction based on im- munohistochemistry images,” Bioinformatics, vol. 38, no. 16, pp. 4019– 4027, 2022
2022
-
[78]
Prog-sol: Predicting protein solubility using protein embeddings and dual-graph convolutional networks,
G. Li, N. Zhang, and L. Fan, “Prog-sol: Predicting protein solubility using protein embeddings and dual-graph convolutional networks,” ACS Omega, vol. 10, no. 4, pp. 3910–3916, 2025. [Online]. Available: https://doi.org/10.1021/acsomega.4c09688
2025 doi
-
[79]
Deep-probind: Bind- ing protein prediction with transformer-based deep learning model,
S. Khan, S. Noor, H. H. Awan, S. Iqbal et al. , “Deep-probind: Bind- ing protein prediction with transformer-based deep learning model,” Springer nature link , 2025
2025
-
[80]
Insights into the inner workings of transformer models for protein function prediction,
M. Wenzel, E. Gr ¨uner, and N. Strodthoff, “Insights into the inner workings of transformer models for protein function prediction,” Bioinformatics, vol. 40, no. 3, p. btae031, 01 2024. [Online]. Available: https://doi.org/10.1093/bioinformatics/btae031
2024 doi
-
[81]
Segt-go: a graph transformer method based on ppi serialization and explanatory artificial intelligence for protein function prediction,
Y . Wang, Y . Sun, B. Lin, and Others, “Segt-go: a graph transformer method based on ppi serialization and explanatory artificial intelligence for protein function prediction,” BMC Bioinformatics , vol. 26, p. 46, 2025
2025
-
[82]
Integrating transformers and automl for protein function prediction,
G. B. de Oliveira, H. Pedrini, and Z. Dias, “Integrating transformers and automl for protein function prediction,” in 2024 46th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC) , 2024, pp. 1–5
2024
-
[85]
Gact-ppis: Prediction of protein-protein interaction sites based on graph structure and transformer network,
L. Meng and H. Zhang, “Gact-ppis: Prediction of protein-protein interaction sites based on graph structure and transformer network,” International Journal of Biological Macromolecules , vol. 283, p. 137272, 2024. [Online]. Available: https://www.sciencedirect.com/ science/arti...
2024
-
[86]
Tranp-b-site: A transformer enhanced method for prediction of binding sites of protein-protein interactions,
S. H. Khan, H. Tayara, and K. T. Chong, “Tranp-b-site: A transformer enhanced method for prediction of binding sites of protein-protein interactions,” Measurement, vol. 251, p. 117227, 2025. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S026322412500586X
2025
-
[87]
Tuna: An uncertainty- aware transformer model for sequence-based protein–protein interaction prediction,
Y . S. Ko, J. Parkinson, C. Liu, and W. Wang, “Tuna: An uncertainty- aware transformer model for sequence-based protein–protein interaction prediction,” Briefings in Bioinformatics , 2024. [Online]. Available: https://doi.org/10.1093/bib/bbae359
2024 doi
-
[88]
Predicting protein-protein binding affinity with deep learning: A comparative analysis of cnn and transformer models,
L. Chen, K. F. Ahmad Nasif, B. Deng, S. Niu, and C. Y . Xie, “Predicting protein-protein binding affinity with deep learning: A comparative analysis of cnn and transformer models,” in 2024 IEEE 36th International Conference on Tools with Artificial Intelligence (ICTAI), 2024, ...
2024
-
[89]
A review of transformers in drug discovery and beyond,
J. Jiang, L. Chen, L. Ke, B. Dou, C. Zhang, H. Feng, Y . Zhu, H. Qiu, B. Zhang, and G. Wei, “A review of transformers in drug discovery and beyond,” Journal of Pharmaceutical Analysis , p. 101081, 2024. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S2...
2024
-
[90]
Mol-bert: An effective molecular representation with bert for molecular property prediction,
J. Li and X. Jiang, “Mol-bert: An effective molecular representation with bert for molecular property prediction,” Wireless Communications and Mobile Computing , vol. 2021, no. 1, p. 7181815, 2021
2021
-
[91]
Molecular generative graph neural networks for drug discovery,
P. Bongini, M. Bianchini, and F. Scarselli, “Molecular generative graph neural networks for drug discovery,” Neurocomputing, vol. 450, pp. 242–252, 2021
2021
-
[93]
Integrating transformer-based language model for drug discovery,
R. R. Kotkondawar, S. R. Sutar, A. W. Kiwelekar, and V . J. Kadam, “Integrating transformer-based language model for drug discovery,” in 2024 11th International Conference on Computing for Sustainable Global Development (INDIACom) , 2024, pp. 1096–1101
2024
-
[94]
Integrating transformers and many-objective optimization for drug design,
N. Aksamit, J. Hou, Y . Li, and Others, “Integrating transformers and many-objective optimization for drug design,” BMC Bioinformatics , vol. 25, p. 208, 2024
2024
-
[95]
Transformers and large language models for chemistry and drug discovery,
A. M. Bran and P. Schwaller, “Transformers and large language models for chemistry and drug discovery,” in Drug Development Supported by Informatics, H. Satoh, K. Funatsu, and H. Yamamoto, Eds. Springer, Singapore, 2024
2024
-
[96]
Sspro/accpro 5: almost perfect prediction of protein secondary structure and relative solvent accessibility using profiles, machine learning and structural similarity,
C. N. Magnan and P. Baldi, “Sspro/accpro 5: almost perfect prediction of protein secondary structure and relative solvent accessibility using profiles, machine learning and structural similarity,” Bioinformatics, vol. 30, no. 18, pp. 2592–2597, 2014
2014
-
[97]
Uniprot: the universal protein knowledgebase in 2021,
Anonymous, “Uniprot: the universal protein knowledgebase in 2021,” Nucleic acids research, vol. 49, no. D1, pp. D480–D489, 2021
2021
-
[98]
Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets,
M. Steinegger and J. S ¨oding, “Mmseqs2 enables sensitive protein sequence searching for the analysis of massive data sets,” Nature biotechnology, vol. 35, no. 11, pp. 1026–1028, 2017
2017
-
[99]
Confold2: improved contact-driven ab initio protein structure modeling,
B. Adhikari and J. Cheng, “Confold2: improved contact-driven ab initio protein structure modeling,” BMC bioinformatics , vol. 19, pp. 1–5, 2018
2018
-
[101]
Language models enable zero-shot prediction of the effects of mutations on protein function,
J. Meier, R. Rao, R. Verkuil, J. Liu, T. Sercu, and A. Rives, “Language models enable zero-shot prediction of the effects of mutations on protein function,” Advances in neural information processing systems , vol. 34, pp. 29 287–29 303, 2021
2021
-
[102]
Energy and policy con- siderations for modern deep learning research,
E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy con- siderations for modern deep learning research,” in Proceedings of the AAAI conference on artificial intelligence , vol. 34, no. 09, 2020, pp. 13 693–13 696
2020
-
[103]
Deep learning in proteomics,
B. Wen, W.-F. Zeng, Y . Liao, Z. Shi, S. R. Savage, W. Jiang, and B. Zhang, “Deep learning in proteomics,” Proteomics, vol. 20, no. 21- 22, p. 1900335, 2020
2020
-
[104]
Deep multi-view learning methods: A review,
X. Yan, S. Hu, Y . Mao, Y . Ye, and H. Yu, “Deep multi-view learning methods: A review,” Neurocomputing, vol. 448, pp. 106–129, 2021
2021
-
[106]
Towards a rigorous science of inter- pretable machine learning,
F. Doshi-Velez and B. Kim, “Towards a rigorous science of inter- pretable machine learning,” arXiv preprint arXiv:1702.08608 , 2017
2017 arXiv
-
[107]
Explainable artificial intelligence: Understanding, vi- sualizing and interpreting deep learning models,
W. Samek, “Explainable artificial intelligence: Understanding, vi- sualizing and interpreting deep learning models,” arXiv preprint arXiv:1708.08296, 2017
2017 arXiv
-
[108]
Transformer protein language models are unsupervised structure learners,
R. Rao, J. Meier, T. Sercu, S. Ovchinnikov, and A. Rives, “Transformer protein language models are unsupervised structure learners,” Biorxiv, pp. 2020–12, 2020
2020
-
[109]
Protein tertiary structure prediction and refinement using deep learning and rosetta in casp14,
I. Anishchenko, M. Baek, H. Park, N. Hiranuma, D. E. Kim, J. Dau- paras, S. Mansoor, I. R. Humphreys, and D. Baker, “Protein tertiary structure prediction and refinement using deep learning and rosetta in casp14,” Proteins: Structure, Function, and Bioinformatics , vol. 89, no...
2021
-
[110]
Deepaffinity: interpretable deep learning of compound–protein affinity through unified recurrent and convolutional neural networks,
M. Karimi, D. Wu, Z. Wang, and Y . Shen, “Deepaffinity: interpretable deep learning of compound–protein affinity through unified recurrent and convolutional neural networks,” Bioinformatics, vol. 35, no. 18, pp. 3329–3338, 2019
2019
-
[111]
Grandmaster level in starcraft ii using multi-agent reinforcement learning,
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev et al. , “Grandmaster level in starcraft ii using multi-agent reinforcement learning,” nature, vol. 575, no. 7782, pp. 350–354, 2019
2019
-
[112]
Toward a shared vision for cancer genomic data,
R. L. Grossman, A. P. Heath, V . Ferretti, H. E. Varmus, D. R. Lowy, W. A. Kibbe, and L. M. Staudt, “Toward a shared vision for cancer genomic data,” New England Journal of Medicine , vol. 375, no. 12, pp. 1109–1112, 2016
2016
-
[114]
Prottrans: Toward understanding the language of life through self-supervised learning,
A. Elnaggar, M. Heinzinger, C. Dallago, G. Rehawi, Y . Wang, L. Jones, T. Gibbs, T. Feher, C. Angerer, M. Steinegger et al. , “Prottrans: Toward understanding the language of life through self-supervised learning,” IEEE transactions on pattern analysis and machine intel- ligen...
2021
-
[115]
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences,
A. Rives, J. Meier, T. Sercu, S. Goyal, Z. Lin, J. Liu, D. Guo, M. Ott, C. L. Zitnick, J. Ma et al., “Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences,” Proceedings of the National Academy of Sciences , vol. 118, no. ...
2021
-
[116]
Accurate prediction of protein structures and interactions using a three- track neural network,
M. Baek, F. DiMaio, I. Anishchenko, J. Dauparas, S. Ovchinnikov, G. R. Lee, J. Wang, Q. Cong, L. N. Kinch, R. D. Schaeffer et al. , “Accurate prediction of protein structures and interactions using a three- track neural network,” Science, vol. 373, no. 6557, pp. 871–876, 2021
2021
-
[117]
Molmol: a program for dis- play and analysis of macromolecular structures,
R. Koradi, M. Billeter, and K. W ¨uthrich, “Molmol: a program for dis- play and analysis of macromolecular structures,” Journal of molecular graphics, vol. 14, no. 1, pp. 51–55, 1996
1996
-
[118]
Transformer architecture and attention mech- anisms in genome data analysis: a comprehensive review,
S. R. Choi and M. Lee, “Transformer architecture and attention mech- anisms in genome data analysis: a comprehensive review,” Biology, vol. 12, no. 7, p. 1033, 2023
2023
-
[119]
A unified approach to interpreting model predictions,
S. Lundberg, “A unified approach to interpreting model predictions,” arXiv preprint arXiv:1705.07874 , 2017
2017 arXiv
-
[120]
Profile prediction: An alignment-based pre-training task for protein sequence models,
P. Sturmfels, J. Vig, A. Madani, and N. F. Rajani, “Profile prediction: An alignment-based pre-training task for protein sequence models,” arXiv preprint arXiv:2012.00195 , 2020
2012 arXiv
-
[121]
Longformer: The long- document transformer,
I. Beltagy, M. E. Peters, and A. Cohan, “Longformer: The long- document transformer,” arXiv preprint arXiv:2004.05150 , 2020
2004 arXiv
-
[122]
Transformer-xl: Attentive language models beyond a fixed- length context,
Z. Dai, “Transformer-xl: Attentive language models beyond a fixed- length context,” arXiv preprint arXiv:1901.02860 , 2019
1901 arXiv
-
[123]
Distilling the knowledge in a neural network,
G. Hinton, “Distilling the knowledge in a neural network,” arXiv preprint arXiv:1503.02531, 2015
2015 arXiv
-
[124]
Carbon emissions and large neural network training,
D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean, “Carbon emissions and large neural network training,” arXiv preprint arXiv:2104.10350, 2021
2021 arXiv
-
[125]
Highly accurate protein structure prediction with alphafold,
J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. ˇZ´ıdek, A. Potapenko et al., “Highly accurate protein structure prediction with alphafold,” nature, vol. 596, no. 7873, pp. 583–589, 2021
2021
-
[2004]
Available: https://doi.org/10.1093/nar/gkh416
[Online]. Available: https://doi.org/10.1093/nar/gkh416
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.