Pith. sign in

REVIEW 4 major objections 7 minor 185 references

Sequence-based protein-protein interaction prediction and its applications in drug discovery

T0 review · 4 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Sequence-based PPI prediction is a competitive, practical tool for drug discovery, argue the authors of this review.

desk verdict A useful review that gets the evaluation story right but overreaches in its final claims about drug-discovery impact. read the letter →

arxiv 2507.19805 v1 pith:UFDJCVBE submitted 2025-07-26 q-bio.BM cs.LG

classification q-bio.BMcs.LG
keywords protein-proteininteractionpredictionsequence-basedmodelsproteinlanguagedrugtargetidentificationdiscoverybiologicspeptidebindersantibodydesign
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is a review that argues sequence-based protein-protein interaction (PPI) prediction is a competitive and practically useful alternative to structure-based methods, especially for drug discovery. It traces how predictors have evolved from hand-engineered features to protein language models, and how they are now used to identify drug targets, design peptide binders, and optimize antibodies. The authors acknowledge that progress is slow and that evaluation is often flawed, yet they maintain that these methods are already enabling real therapeutic development.

What carries the argument

The central objects are protein language models (pLMs) and similarity-based interaction scorers. pLMs are transformer neural networks pretrained on millions of protein sequences via masked language modeling; they generate embeddings that encode physicochemical, evolutionary, and functional information. Similarity-based methods such as PIPE and SPRINT score a candidate interaction by counting shared short-sequence windows with known interacting pairs, a mechanism that does not require explicit negative examples. The review also details the training pipeline: curation of positive pairs from databases like BioGRID, assembly of negative pairs (typically random or shuffled sequences), redundancy reduction, and evaluation with metrics that address class imbalance.

What would settle it

A concrete falsifier: build a benchmark where negative pairs are not random but are sampled from pairs that share a common subcellular localization or are suggested by co-expression, and show that top-performing sequence-based predictors drop to near-chance accuracy—that would invalidate the assumption that random negatives measure true discrimination.

Watch

Extended reading notes

Core claim

This review establishes that sequence-based PPI predictors—methods that take only amino acid sequences as input—have reached the point where they can identify actionable drug targets within interaction networks and can be used to design therapeutic biologics such as peptide binders and antibodies. It argues that, despite the appeal of structure-based approaches, sequence-based methods avoid reliance on scarce high-resolution structures and on imperfect structure predictions, and they remain competitive in rigorous benchmarks even against newer deep-learning models. The authors describe the shift toward protein language models (pLMs) as the dominant feature source, while noting that similarity-based methods like PIPE and SPRINT still hold their own.

Load-bearing premise

The assumption that randomly selected protein pairs (or shuffled sequences) are almost always true non-interactions underlies the training and evaluation of every machine-learning predictor the review surveys, yet the paper offers no evidence beyond asserting the risk is 'negligible in practice.'

Editorial extensions

If this is right

  • If sequence-based predictors are as competitive as the review claims, then they can be used to prioritize protein pairs for experimental validation, greatly reducing the cost of interactome mapping.
  • They enable cross-species prediction, allowing interactions in understudied organisms to be inferred from well-studied proxy organisms, which is valuable for host-pathogen work.
  • They can feed directly into generative models for peptide binder design, as demonstrated by InSiPS, which uses a sequence-based scorer as a fitness function to evolve binders against a target while minimizing off-target interactions.
  • They can be applied to antibody engineering, for example by scoring candidate mutations with protein language models to guide affinity maturation.
  • If evaluation standards improve (e.g., imbalanced test sets and awareness of pair difficulty), the gap between reported and real-world performance would shrink, making the predictors even more reliable.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The review's implicit bet is that the actual distribution of protein pairs in a living cell is far more imbalanced than any balanced benchmark, so a model that looks mediocre on AUPRC could still be extremely useful at a high-precision operating point; this is an editorial inference, not a claim in the paper.
  • The success of pLM-based predictors suggests that transfer learning from general protein sequence data carries most of the signal, and that fine-tuning on PPI labels may be a secondary refinement; a testable extension is to compare pLM embeddings against carefully engineered features on the same held-out pairs.
  • A concrete extension the authors leave implicit is using sequence-based predictors not only for binary interaction classification but also for ranking candidate peptides in a library—an application that would directly benefit from the precision metrics they emphasize.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. This manuscript is a narrative review of sequence-based protein-protein interaction (PPI) prediction, covering data sources and curation, traditional and deep learning predictors (with an emphasis on protein language models), evaluation methodology and class imbalance, and applications to target identification, peptide binder design, and antibody design. It also contains an original analysis in Figure 1 showing a widening gap between experimentally validated PPIs involving human proteins and PPIs for which both partners have high-quality structures. The review argues that sequence-based predictors remain competitive with structure-based methods and are practically useful in drug discovery.

Significance. The review is potentially useful as a broad survey for practitioners: it consolidates the main databases, pLM resources, and a substantial portion of the recent literature on sequence-based PPI predictors, and it gives a clear account of evaluation pitfalls such as C1/C2/C3 splits, data leakage, and class imbalance. The authors deserve credit for citing and summarizing some of the strongest negative evidence against naive benchmark claims (Dunham et al., Bernett et al.). However, the review draws an optimistic conclusion about drug-discovery utility that is not reconciled with the negative evidence it itself presents, and the only original quantitative analysis is not reproducible as reported. These issues do not invalidate the survey as a whole but require revision before the central claims can be accepted at face value.

major comments (4)
  1. [Summary and future trends; Class imbalance section] The final paragraph asserts that sequence-based PPI predictors 'facilitate the identification of actionable drug targets' and 'can also be used to design therapeutic biologics such as peptide binders and antibodies,' but the review's own evidence undercuts this generalization. The same section reports Dunham et al.'s finding that most published methods drop dramatically on realistic imbalanced test sets and Bernett et al.'s finding of widespread data leakage, while the Class imbalance section estimates human interactome imbalances of about 1:400 to 1:3,400, whereas Table 3 shows most surveyed methods are evaluated at 1:1, 1:10, or 1:100. Please either qualify the concluding claim to the favorable evaluation regimes and add explicit discussion of the transfer gap, or supply prospective/real-world validation evidence (for example, experimentally confirmed target identifications or binders beyond the self-cited InSiPS and PepMLM/PepPrepCLIP examples).
  2. [Data curation (negative pairs)] The assumption that random pairs, shuffled sequences, or subcellular-localization-based negatives are mostly true non-interactions is load-bearing for every supervised ML predictor in Table 3, yet the manuscript states only that 'this risk is assumed to be negligible in practice.' No evidence is provided for this assumption, and the review's own critique of C1/C2/C3 splits and leakage suggests that negative-set construction is a main source of optimistic performance estimates. Please quantify the likely false-negative rate in the negative construction strategies (e.g., from interactome-size estimates or from re-analysis with stricter negatives such as C3 splits) or explicitly weaken the conclusions drawn from models trained on these negatives.
  3. [Figure 1 and its caption] Figure 1 is presented as an original quantitative result, but the caption gives only 'Data retrieved and compiled using the RCSB PDB API' and does not state the retrieval date, the exact API queries and filtering criteria, the BioGRID release/version, or the definition used for 'high-quality structures ... for both interactors' (including how truncation was handled). The claim that the fraction of PPIs with high-quality structures has been decreasing over time cannot be checked or reproduced without this information. Please add a data-and-methods paragraph and, ideally, deposit the scripts and raw counts.
  4. [Old but Gold: Sequence-Based Protein-Protein and Peptide-Protein Predictors; Table 3] The section claims that PIPE and SPRINT 'remain competitive to this day' and cites [28], but Table 3 lists no performance metrics, so the reader cannot evaluate this central claim. Since the review's own summary states that progress remains slow, please include a compact comparison of reported AUROC/AUPRC or C1/C2/C3 results for the representative methods in Table 3, or state explicitly that the competitiveness claim is a citation-level claim and is not supported by any new comparison in this review.
minor comments (7)
  1. [Model evaluation] In the metrics paragraph, the text reads 'true positives (TP), true negatives (TN), false positives (TP), and false negatives (FN)'; the third item should be 'false positives (FP)'.
  2. [Model evaluation (ROC/AUROC)] The sentence 'The ROC curve and the area under it, in contrast with the PR curve, is insensitive to class imbalance [74], [95]' is followed by the claim that AUROC 'does not correlate with the difficulty of the classification problem at different imbalance ratios.' Reference [95] is titled 'The Receiver Operating Characteristic Curve Accurately Assesses Imbalanced Datasets,' which appears to contradict the text's interpretation; please verify that the citation supports the claim or clarify the intended point.
  3. [Section numbering] The section titled 'Generalizing Beyond Model Systems: Challenges and Solutions in Cross-Species PPI Prediction' is numbered '4', while surrounding sections are unnumbered; please use consistent numbering or remove the stray number.
  4. [Reference [71]] Reference [71] is listed as 'Gradient' only; the full title of the GTB-PPI paper appears to be missing.
  5. [Table 1] Table 1 lists interaction counts for BioGRID, STRING, IntAct, MINT, and Propedia, but no database release/version or retrieval date is given; please add these details so the counts are interpretable.
  6. [Figure 6] The simulation details for Figure 6 (sample sizes, score distributions, and thresholding procedures) are not provided; please add them to the caption or to a methods paragraph for reproducibility.
  7. [Potential conflicts of interest] The 'Potential conflicts of interest' statement says 'The authors have no conflicts of interest to disclose,' but one author is affiliated with NuvoBio Corp. and the authors are co-developers of several tools discussed favorably in the text (MP-PIPE, PIPE4, InSiPS, Positome, Reciprocal Perspective). Please clarify the commercial and financial relationships and how they relate to the tools under review.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; self-citations are ancillary, and the central claim rests on external benchmarks and validations.

full rationale

This is a narrative review with no derived equations and no fitted parameters, so the self-definitional and fit-as-prediction patterns do not apply. The favorable descriptions of the authors' own tools (MP-PIPE, PIPE4, InSiPS, Positome, Reciprocal Perspective) are self-citations, but they are not the sole support for the review's conclusion: the claim that sequence-based predictors remain competitive is attributed to an external benchmark (Bernett et al. [28]), and the InSiPS binder success is supported by an independently measurable nanomolar dissociation constant in Hajikarimlou et al. [162]. The paper's candid discussion of class imbalance, C1/C2/C3 data leakage, and the drop in performance on realistic test sets weakens the practical-usefulness conclusion, but an unsupported or internally tensioned conclusion is not the same as a circular derivation. No step satisfies the requirement of exhibiting an equation or argument that reduces to its own input by construction; therefore no circularity is found.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The review's synthesis implicitly assumes that curated PPI databases are reliable, that constructed negative pairs are mostly true non-interactions, and that cited benchmark evaluations are valid. These are domain assumptions, not parameters or entities introduced by the paper.

assumptions (3)
  • domain assumption Curated PPI databases (e.g., BioGRID, STRING) accurately reflect true physical interactions and are suitable for training and benchmarking predictors.
    The review uses these databases as ground truth in Table 1 and discusses predictors trained on them; the reliability of these databases is assumed throughout.
  • domain assumption Negatives generated by random pairing or sequence shuffling are mostly true non-interactions, so mislabeling risk is negligible.
    The review acknowledges the risk in the 'Data curation' section but assumes it is negligible, and this assumption underlies the evaluation of all ML-based predictors surveyed.
  • domain assumption Similarity-based methods remain competitive despite their assumption that interactions are mediated by contiguous sequence windows.
    The review states this assumption is incorrect but still treats PIPE/SPRINT as competitive; it relies on cited comparative benchmarks.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sequence-based protein-protein interaction prediction and its applications in drug discovery." pith.science (2026). https://pith.science/paper/UFDJCVBE

@misc{pith2026250719805,
  author       = {Pith},
  title        = {Pith review of: Sequence-based protein-protein interaction prediction and its applications in drug discovery},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UFDJCVBE}},
  note         = {Machine review of arXiv:2507.19805}
}
read the original abstract

Aberrant protein-protein interactions (PPIs) underpin a plethora of human diseases, and disruption of these harmful interactions constitute a compelling treatment avenue. Advances in computational approaches to PPI prediction have closely followed progress in deep learning and natural language processing. In this review, we outline the state-of the-art for sequence-based PPI prediction methods and explore their impact on target identification and drug discovery. We begin with an overview of commonly used training data sources and techniques used to curate these data to enhance the quality of the training set. Subsequently, we survey various PPI predictor types, including traditional similarity-based approaches, and deep learning-based approaches with a particular emphasis on the transformer architecture. Finally, we provide examples of PPI prediction in systems-level proteomics analyses, target identification, and design of therapeutic peptides and antibodies. We also take the opportunity to showcase the potential of PPI-aware drug discovery models in accelerating therapeutic development.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

185 extracted references · 97 canonical work pages

  1. [28]

    Cracking the Black Box of Deep Sequence-Based Protein– Protein Interaction Prediction,

    J. Bernett, D. B. Blumenthal, and M. List, “Cracking the Black Box of Deep Sequence-Based Protein– Protein Interaction Prediction, ” Briefings in Bioinformatics , vol. 25, no. 2, p. bbae76, Mar. 2024, doi: 10.1093/bib/bbae076

  2. [1]

    Exploring Protein-Protein Interactions at the Proteome Level,

    H. Elhabashy, F. Merino, V. Alva, O. Kohlbacher, and A. N. Lupas, “Exploring Protein-Protein Interactions at the Proteome Level, ” Structure, vol. 30, no. 4, pp. 462–475, Apr. 2022, doi: 10.1016/ j.str.2022.02.004

  3. [2]

    Diversity of Protein–Protein Interactions,

    I. M. A. Nooren and J. M. Thornton, “Diversity of Protein–Protein Interactions, ” The EMBO Journal, vol. 22, no. 14, pp. 3486–3492, 2003, doi: 10.1093/emboj/cdg359

  4. [3]

    Protein-Protein Interactions in DNA Mismatch Repair,

    P. Friedhoff, P. Li, and J. Gotthardt, “Protein-Protein Interactions in DNA Mismatch Repair, ” DNA Repair, vol. 38, pp. 50–57, Feb. 2016, doi: 10.1016/j.dnarep.2015.11.013

  5. [4]

    Protein-Protein Interactions in Transcription: A Fertile Ground for Helix Mimetics,

    D. A. Guarracino, B. N. Bullock, and P. S. Arora, “Protein-Protein Interactions in Transcription: A Fertile Ground for Helix Mimetics, ” Biopolymers, vol. 95, no. 1, pp. 1–7, Jan. 2011, doi: 10.1002/ bip.21546

  6. [5]

    Protein Translation: Biological Processes and Therapeutic Strategies for Human Diseases,

    X. Jia, X. He, C. Huang, J. Li, Z. Dong, and K. Liu, “Protein Translation: Biological Processes and Therapeutic Strategies for Human Diseases, ” Signal Transduction and Targeted Therapy, vol. 9, no. 1, p. 44, Feb. 2024, doi: 10.1038/s41392-024-01749-9

  7. [6]

    Identification of Protein Interactions Involved in Cellular Signaling,

    J. Westermarck, J. Ivaska, and G. L. Corthals, “Identification of Protein Interactions Involved in Cellular Signaling, ” Molecular & Cellular Proteomics : MCP, vol. 12, no. 7, pp. 1752–1763, Jul. 2013, doi: 10.1074/mcp.R113.027771

  8. [7]

    Molecular Chaperones and Protein Quality Control: An Introduction to the JBC Reviews Thematic Series,

    J. Buchner, “Molecular Chaperones and Protein Quality Control: An Introduction to the JBC Reviews Thematic Series, ” The Journal of Biological Chemistry , vol. 294, no. 6, pp. 2074–2075, Feb. 2019, doi: 10.1074/jbc.REV118.006739

Show all 185 references
  1. [8]

    Chapter 4: Protein Interactions and Disease,

    M. W. Gonzalez and M. G. Kann, “Chapter 4: Protein Interactions and Disease, ” PLoS Computational Biology, vol. 8, no. 12, p. e1002819, Dec. 2012, doi: 10.1371/journal.pcbi.1002819

  2. [9]

    Comprehensive Characterization of Protein–Protein Interactions Perturbed by Dis- ease Mutations,

    F. Cheng et al., “Comprehensive Characterization of Protein–Protein Interactions Perturbed by Dis- ease Mutations, ” Nature Genetics, vol. 53, no. 3, pp. 342–353, Mar. 2021, doi: 10.1038/s41588-020-00774- y

  3. [10]

    Discovery and Significance of Protein-Protein Interactions in Health and Disease,

    J. F. Greenblatt, B. M. Alberts, and N. J. Krogan, “Discovery and Significance of Protein-Protein Interactions in Health and Disease, ” Cell, vol. 187, no. 23, pp. 6501–6517, Nov. 2024, doi: 10.1016/ j.cell.2024.10.038

  4. [11]

    Alzheimer Disease,

    D. S. Knopman et al., “Alzheimer Disease, ” Nature Reviews Disease Primers , vol. 7, no. 1, p. 33, May 2021, doi: 10.1038/s41572-021-00269-y

  5. [12]

    Pathological Mechanisms Underlying TDP-43 Driven Neurode- generation in FTLD–ALS Spectrum Disorders,

    J. Janssens and C. Van Broeckhoven, “Pathological Mechanisms Underlying TDP-43 Driven Neurode- generation in FTLD–ALS Spectrum Disorders, ” Human Molecular Genetics, vol. 22, no. R1, pp. R77– R87, Oct. 2013, doi: 10.1093/hmg/ddt349

  6. [13]

    Parkinson's Disease,

    B. R. Bloem, M. S. Okun, and C. Klein, “Parkinson's Disease, ” The Lancet, vol. 397, no. 10291, pp. 2284– 2303, Jun. 2021, doi: 10.1016/s0140-6736(21)00218-x

  7. [14]

    Huntington Disease,

    G. P. Bates et al., “Huntington Disease, ” Nature Reviews Disease Primers , vol. 1, no. 1, p. 15005, Apr. 2015, doi: 10.1038/nrdp.2015.5

  8. [15]

    Creutzfeldt-Jakob Disease,

    Y. Iwasaki, “Creutzfeldt-Jakob Disease, ” Neuropathology, vol. 37, no. 2, pp. 174–188, 2017, doi: 10.1111/ neup.12355

  9. [16]

    KRAS Mutation: From Undruggable to Druggable in Cancer,

    L. Huang, Z. Guo, F. Wang, and L. Fu, “KRAS Mutation: From Undruggable to Druggable in Cancer, ” Signal Transduction and Targeted Therapy , vol. 6, no. 1, p. 386, Nov. 2021, doi: 10.1038/ s41392-021-00780-4. 21

  10. [17]

    Illuminating the Dark Protein-Protein Interactome,

    M. S. Tabar, C. Parsania, H. Chen, X.-D. Su, C. G. Bailey, and J. E. J. Rasko, “Illuminating the Dark Protein-Protein Interactome, ” Cell Reports Methods , vol. 2, no. 8, Aug. 2022, doi: 10.1016/ j.crmeth.2022.100275

  11. [18]

    A Protein Interaction Landscape of Breast Cancer,

    M. Kim et al. , “A Protein Interaction Landscape of Breast Cancer, ” Science, vol. 374, no. 6563, p. eabf3066, Oct. 2021, doi: 10.1126/science.abf3066

  12. [19]

    Decoding the Functional Impact of the Cancer Genome through Protein–Protein Interactions,

    H. Fu, X. Mo, and A. A. Ivanov, “Decoding the Functional Impact of the Cancer Genome through Protein–Protein Interactions, ” Nature Reviews Cancer , vol. 25, no. 3, pp. 189–208, Mar. 2025, doi: 10.1038/s41568-024-00784-6

  13. [20]

    Affinity-purification Coupled to Mass Spectrometry: Basic Principles and Strategies,

    W. H. Dunham, M. Mullin, and A.-C. Gingras, “Affinity-purification Coupled to Mass Spectrometry: Basic Principles and Strategies, ” PROTEOMICS, vol. 12, no. 10, pp. 1576–1590, May 2012, doi: 10.1002/ pmic.201100523

  14. [21]

    Exploring Protein–Protein Interactions with Phage Display,

    S. S. Sidhu, W. J. Fairbrother, and K. Deshayes, “Exploring Protein–Protein Interactions with Phage Display, ” ChemBioChem, vol. 4, no. 1, pp. 14–25, Jan. 2003, doi: 10.1002/cbic.200390008

  15. [22]

    Current Experimental Methods for Characterizing Protein–Protein Interactions,

    M. Zhou, Q. Li, and R. Wang, “Current Experimental Methods for Characterizing Protein–Protein Interactions, ” ChemMedChem, vol. 11, no. 8, pp. 738–756, Apr. 2016, doi: 10.1002/cmdc.201500495

  16. [23]

    Yeast Two-Hybrid, a Powerful Tool for Systems Biology,

    A. Brückner, C. Polge, N. Lentze, D. Auerbach, and U. Schlattner, “Yeast Two-Hybrid, a Powerful Tool for Systems Biology, ” International Journal of Molecular Sciences , vol. 10, no. 6, pp. 2763–2788, Jun. 2009, doi: 10.3390/ijms10062763

  17. [24]

    Studying Protein–Protein Interactions: Latest and Most Popular Approaches,

    S. Akbarzadeh, Ö. Coo̧skun, and B. Güņcer, “Studying Protein–Protein Interactions: Latest and Most Popular Approaches, ” Journal of Structural Biology, vol. 216, no. 4, p. 108118, Dec. 2024, doi: 10.1016/ j.jsb.2024.108118

  18. [25]

    Global Investigation of Protein–Protein Interactions in Yeast Saccharomyces Cerevisiae Using Re-Occurring Short Polypeptide Sequences,

    S. Pitre et al., “Global Investigation of Protein–Protein Interactions in Yeast Saccharomyces Cerevisiae Using Re-Occurring Short Polypeptide Sequences, ” Nucleic Acids Research, vol. 36, no. 13, pp. 4286– 4294, Aug. 2008, doi: 10.1093/nar/gkn390

  19. [26]

    SPRINT: Ultrafast Protein-Protein Interaction Prediction of the Entire Human Interactome,

    Y. Li and L. Ilie, “SPRINT: Ultrafast Protein-Protein Interaction Prediction of the Entire Human Interactome, ” BMC Bioinformatics, vol. 18, no. 1, p. 485, Nov. 2017, doi: 10.1186/s12859-017-1871-x

  20. [27]

    PIPE4: Fast PPI Predictor for Comprehensive Inter- and Cross-Species Interactomes,

    K. Dick et al., “PIPE4: Fast PPI Predictor for Comprehensive Inter- and Cross-Species Interactomes, ” Scientific Reports, vol. 10, no. 1, p. 1390, Dec. 2020, doi: 10.1038/s41598-019-56895-w

  21. [29]

    Stabilization of Protein-Protein Interactions in Drug Discovery,

    S. A. Andrei et al., “Stabilization of Protein-Protein Interactions in Drug Discovery, ” Expert Opinion on Drug Discovery, vol. 12, no. 9, pp. 925–940, Sep. 2017, doi: 10.1080/17460441.2017.1346608

  22. [30]

    Evolution of In Silico Strategies for Protein-Protein Interaction Drug Discovery,

    S. J. Y. Macalino, S. Basith, N. A. B. Clavio, H. Chang, S. Kang, and S. Choi, “Evolution of In Silico Strategies for Protein-Protein Interaction Drug Discovery, ” Molecules, vol. 23, no. 8, p. 1963, Aug. 2018, doi: 10.3390/molecules23081963

  23. [31]

    Rational Design of Peptide-Based Inhibitors Disrupting Protein- Protein Interactions,

    X. Wang, D. Ni, Y. Liu, and S. Lu, “Rational Design of Peptide-Based Inhibitors Disrupting Protein- Protein Interactions, ” Frontiers in Chemistry, vol. 9, May 2021, doi: 10.3389/fchem.2021.682675

  24. [32]

    The RCSB Protein Data Bank: Integrative View of Protein, Gene and 3D Structural Information,

    P. W. Rose et al., “The RCSB Protein Data Bank: Integrative View of Protein, Gene and 3D Structural Information, ” Nucleic Acids Research , vol. 45, no. D1, pp. D271–D281, Jan. 2017, doi: 10.1093/nar/ gkw1000

  25. [33]

    The BioGRID Database: A Comprehensive Biomedical Resource of Curated Protein, Genetic, and Chemical Interactions,

    R. Oughtred et al. , “The BioGRID Database: A Comprehensive Biomedical Resource of Curated Protein, Genetic, and Chemical Interactions, ” Protein Science, vol. 30, no. 1, pp. 187–200, 2021, doi: 10.1002/pro.3978. 22

  26. [34]

    Highly Accurate Protein Structure Prediction with AlphaFold,

    J. Jumper et al., “Highly Accurate Protein Structure Prediction with AlphaFold, ” Nature, vol. 596, no. 7873, pp. 583–589, Aug. 2021, doi: 10.1038/s41586-021-03819-2

  27. [35]

    Accurate Structure Prediction of Biomolecular Interactions with AlphaFold 3,

    J. Abramson et al., “Accurate Structure Prediction of Biomolecular Interactions with AlphaFold 3, ” Nature, vol. 630, no. 8016, pp. 493–500, Jun. 2024, doi: 10.1038/s41586-024-07487-w

  28. [36]

    Evolutionary-Scale Prediction of Atomic-Level Protein Structure with a Language Model,

    Z. Lin et al. , “Evolutionary-Scale Prediction of Atomic-Level Protein Structure with a Language Model, ” Science, vol. 379, no. 6637, pp. 1123–1130, Mar. 2023, doi: 10.1126/science.ade2574

  29. [37]

    Chai-1: Decoding the Molecular Interactions of Life

    Chai Discovery et al., “Chai-1: Decoding the Molecular Interactions of Life. ” bioRxiv, Oct. 2024. doi: 10.1101/2024.10.10.615955

  30. [38]

    Boltz-1 Democratizing Biomolecular Interaction Modeling

    J. Wohlwend et al., “Boltz-1 Democratizing Biomolecular Interaction Modeling. ” bioRxiv, Nov. 2024. doi: 10.1101/2024.11.19.624167

  31. [39]

    Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction

    S. Passaro et al., “Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction. ” bioRxiv, Jun

  32. [40]

    AlphaFold Predictions Are Valuable Hypotheses and Accelerate but Do Not Replace Experimental Structure Determination,

    T. C. Terwilliger et al., “AlphaFold Predictions Are Valuable Hypotheses and Accelerate but Do Not Replace Experimental Structure Determination, ” Nature Methods, vol. 21, no. 1, pp. 110–116, Jan. 2024, doi: 10.1038/s41592-023-02087-4

  33. [41]

    Multi-Level Analysis of Intrinsically Disordered Protein Docking Methods,

    J. Verburgt, Z. Zhang, and D. Kihara, “Multi-Level Analysis of Intrinsically Disordered Protein Docking Methods, ” Methods, vol. 204, pp. 55–63, Aug. 2022, doi: 10.1016/j.ymeth.2022.05.006

  34. [42]

    Prediction of Protein–Protein Interactions Using Sequences of Intrinsically Disordered Regions,

    G. Kibar and M. Vingron, “Prediction of Protein–Protein Interactions Using Sequences of Intrinsically Disordered Regions, ” Proteins: Structure, Function, and Bioinformatics, vol. 91, no. 7, pp. 980–990, 2023, doi: 10.1002/prot.26486

  35. [43]

    Systematic Discovery of Protein Interaction Interfaces Using AlphaFold and Exper- imental Validation,

    C. Y. Lee et al., “Systematic Discovery of Protein Interaction Interfaces Using AlphaFold and Exper- imental Validation, ” Molecular Systems Biology , vol. 20, no. 2, pp. 75–97, Feb. 2024, doi: 10.1038/ s44320-023-00005-6

  36. [44]

    Binding Mechanisms of Intrinsically Disordered Proteins: Insights from Experimental Studies and Structural Predictions,

    T. Orand and M. R. Jensen, “Binding Mechanisms of Intrinsically Disordered Proteins: Insights from Experimental Studies and Structural Predictions, ” Current Opinion in Structural Biology , vol. 90, p. 102958, Feb. 2025, doi: 10.1016/j.sbi.2024.102958

  37. [45]

    Deep Learning Tools Predict Variants in Disordered Regions with Lower Sensitivity,

    F. Luppino, S. Lenz, C. F. W. Chow, and A. Toth-Petroczy, “Deep Learning Tools Predict Variants in Disordered Regions with Lower Sensitivity, ” BMC Genomics, vol. 26, no. 1, p. 367, Apr. 2025, doi: 10.1186/s12864-025-11534-9

  38. [46]

    Recent Progress and Future Challenges in Structure-Based Protein-Protein Interaction Prediction,

    R. Yuan, J. Zhang, J. Zhou, and Q. Cong, “Recent Progress and Future Challenges in Structure-Based Protein-Protein Interaction Prediction, ” Molecular Therapy, vol. 33, no. 5, pp. 2252–2268, May 2025, doi: 10.1016/j.ymthe.2025.04.003

  39. [47]

    Raisinghani, V

    N. Raisinghani, V. Parikh, B. Foley, and G. Verkhivker, “Assessing Structures and Conformational Ensembles of Apo and Holo Protein States Using Randomized Alanine Sequence Scanning Combined with Shallow Subsampling in AlphaFold2 : Insights and Lessons from Predictions of Funct...

  40. [48]

    Mitchell, Machine Learning

    T. Mitchell, Machine Learning. in McGraw-Hill Series in Computer Science. New York, NY: McGraw- Hill Professional, 1997

  41. [49]

    Protein–Protein Binding Affinity Prediction from Amino Acid Sequence,

    K. Yugandhar and M. M. Gromiha, “Protein–Protein Binding Affinity Prediction from Amino Acid Sequence, ” Bioinformatics, vol. 30, no. 24, pp. 3583–3589, Dec. 2014, doi: 10.1093/bioinformatics/btu580

  42. [50]

    ISLAND: In-Silico Proteins Binding Affinity Prediction Using Sequence Information,

    W. A. Abbasi, A. Yaseen, F. U. Hassan, S. Andleeb, and F. U. A. A. Minhas, “ISLAND: In-Silico Proteins Binding Affinity Prediction Using Sequence Information, ” BioData Mining, vol. 13, no. 1, pp. 1–13, Dec. 2020, doi: 10.1186/s13040-020-00231-w. 23

  43. [51]

    Machine Learning Methods for Protein-Protein Binding Affinity Prediction in Protein Design,

    Z. Guo and R. Yamaguchi, “Machine Learning Methods for Protein-Protein Binding Affinity Prediction in Protein Design, ” Frontiers in Bioinformatics, vol. 2, Dec. 2022, doi: 10.3389/fbinf.2022.1065703

  44. [52]

    PPI-Affinity: A Web Tool for the Prediction and Optimization of Protein– Peptide and Protein–Protein Binding Affinity,

    S. Romero-Molina et al., “PPI-Affinity: A Web Tool for the Prediction and Optimization of Protein– Peptide and Protein–Protein Binding Affinity, ” Journal of Proteome Research, vol. 21, no. 8, pp. 1829– 1841, Aug. 2022, doi: 10.1021/acs.jproteome.2c00020

  45. [53]

    Predicted Protein–Protein Interaction Sites from Local Sequence Information,

    Y. Ofran and B. Rost, “Predicted Protein–Protein Interaction Sites from Local Sequence Information, ” FEBS Letters, vol. 544, no. 1, pp. 236–239, Jun. 2003, doi: 10.1016/S0014-5793(03)00456-3

  46. [54]

    Progress and Challenges in Predicting Protein–Protein Interaction Sites,

    I. Ezkurdia, L. Bartoli, P. Fariselli, R. Casadio, A. Valencia, and M. L. Tress, “Progress and Challenges in Predicting Protein–Protein Interaction Sites, ” Briefings in Bioinformatics, vol. 10, no. 3, pp. 233–246, May 2009, doi: 10.1093/bib/bbp021

  47. [55]

    Review and Comparative Assessment of Sequence-Based Predictors of Protein-Binding Residues,

    J. Zhang and L. Kurgan, “Review and Comparative Assessment of Sequence-Based Predictors of Protein-Binding Residues, ” Briefings in Bioinformatics , vol. 19, no. 5, pp. 821–837, Sep. 2018, doi: 10.1093/bib/bbx022

  48. [56]

    DELPHI: Accurate Deep Ensemble Model for Protein Interaction Sites Prediction,

    Y. Li, G. B. Golding, and L. Ilie, “DELPHI: Accurate Deep Ensemble Model for Protein Interaction Sites Prediction, ” Bioinformatics, vol. 37, no. 7, pp. 896–904, May 2021, doi: 10.1093/bioinformatics/btaa750

  49. [57]

    Using Support Vector Machine Combined with Auto Covariance to Predict Protein–Protein Interactions from Protein Sequences,

    Y. Guo, L. Yu, Z. Wen, and M. Li, “Using Support Vector Machine Combined with Auto Covariance to Predict Protein–Protein Interactions from Protein Sequences, ” Nucleic Acids Research, vol. 36, no. 9, pp. 3025–3030, May 2008, doi: 10.1093/nar/gkn159

  50. [58]

    Choosing Negative Examples for the Prediction of Protein-Protein Interactions,

    A. Ben-Hur and W. S. Noble, “Choosing Negative Examples for the Prediction of Protein-Protein Interactions, ” BMC Bioinformatics, vol. 7, no. Suppl1, p. S2, Mar. 2006, doi: 10.1186/1471-2105-7-S1-S2

  51. [59]

    Romero-Molina, Y

    S. Romero-Molina, Y. B. Ruiz-Blanco, M. Harms, J. Münch, and E. Sanchez-Garcia, “PPI-Detect: A Support Vector Machine Model for Sequence-Based Prediction of Protein-Protein Interactions: PPI-Detect: A Support Vector Machine Model for Sequence-Based Prediction of Protein-Protei...

  52. [60]

    The Negatome Database: A Reference Set of Non-Interacting Protein Pairs,

    P. Smialowski et al., “The Negatome Database: A Reference Set of Non-Interacting Protein Pairs, ” Nucleic Acids Research, vol. 38, no. suppl_1, pp. D540–D544, Jan. 2010, doi: 10.1093/nar/gkp1026

  53. [61]

    Negatome 2.0: A Database of Non-Interacting Proteins Derived by Literature Mining, Manual Annotation and Protein Structure Analysis,

    P. Blohm et al., “Negatome 2.0: A Database of Non-Interacting Proteins Derived by Literature Mining, Manual Annotation and Protein Structure Analysis, ” Nucleic Acids Research, vol. 42, no. Database issue, pp. D396–D400, Jan. 2014, doi: 10.1093/nar/gkt1079

  54. [62]

    UniProt: The Universal Protein Knowledgebase in 2025,

    The UniProt Consortium, “UniProt: The Universal Protein Knowledgebase in 2025, ” Nucleic Acids Research, vol. 53, no. D1, pp. D609–D617, Jan. 2025, doi: 10.1093/nar/gkae1010

  55. [63]

    The STRING Database in 2023: Protein-Protein Association Networks and Functional Enrichment Analyses for Any Sequenced Genome of Interest,

    D. Szklarczyk et al. , “The STRING Database in 2023: Protein-Protein Association Networks and Functional Enrichment Analyses for Any Sequenced Genome of Interest, ” Nucleic Acids Research, vol. 51, no. D1, pp. D638–D646, Jan. 2023, doi: 10.1093/nar/gkac1000

  56. [64]

    The IntAct Database: Efficient Access to Fine-Grained Molecular Interaction Data,

    N. del~Toro et al., “The IntAct Database: Efficient Access to Fine-Grained Molecular Interaction Data, ” Nucleic Acids Research, vol. 50, no. D1, pp. D648–D653, Jan. 2022, doi: 10.1093/nar/gkab1006

  57. [65]

    MINT, the Molecular Interaction Database: 2012 Update,

    L. Licata et al., “MINT, the Molecular Interaction Database: 2012 Update, ” Nucleic Acids Research, vol. 40, no. Database issue, pp. D857–861, Jan. 2012, doi: 10.1093/nar/gkr930

  58. [66]

    Propedia v2.3: A Novel Representation Approach for the Peptide-Protein Interaction Database Using Graph-Based Structural Signatures,

    P. Martins et al., “Propedia v2.3: A Novel Representation Approach for the Peptide-Protein Interaction Database Using Graph-Based Structural Signatures, ” Frontiers in Bioinformatics, vol. 3, Feb. 2023, doi: 10.3389/fbinf.2023.1103103

  59. [67]

    Protein Sequence Redundancy Reduction: Comparison of Various Method,

    K. Sikic and O. Carugo, “Protein Sequence Redundancy Reduction: Comparison of Various Method, ” Bioinformation, vol. 5, no. 6, pp. 234–239, Nov. 2010, doi: 10.6026/97320630005234. 24

  60. [69]

    Multifaceted Protein–Protein Interaction Prediction Based on Siamese Residual RCNN,

    M. Chen et al. , “Multifaceted Protein–Protein Interaction Prediction Based on Siamese Residual RCNN, ” Bioinformatics, vol. 35, no. 14, pp. i305–i314, Jul. 2019, doi: 10.1093/bioinformatics/btz328

  61. [70]

    D-SCRIPT Translates Genome to Phenome with Sequence-Based, Structure-Aware, Genome-Scale Predictions of Protein-Protein Interactions,

    S. Sledzieski, R. Singh, L. Cowen, and B. Berger, “D-SCRIPT Translates Genome to Phenome with Sequence-Based, Structure-Aware, Genome-Scale Predictions of Protein-Protein Interactions, ” Cell Systems, vol. 12, no. 10, pp. 969–982, Oct. 2021, doi: 10.1016/j.cels.2021.08.010

  62. [71]

    Gradient,

    B. Yu, C. Chen, H. Zhou, B. Liu, and Q. Ma, “Gradient, ” Genomics, Proteomics & Bioinformatics, vol. 18, no. 5, pp. 582–592, Oct. 2020, doi: 10.1016/j.gpb.2021.01.001

  63. [72]

    MMseqs2 Enables Sensitive Protein Sequence Searching for the Analysis of Massive Data Sets,

    M. Steinegger and J. Söding, “MMseqs2 Enables Sensitive Protein Sequence Searching for the Analysis of Massive Data Sets, ” Nature Biotechnology, vol. 35, no. 11, pp. 1026–1028, Nov. 2017, doi: 10.1038/ nbt.3988

  64. [73]

    Improving Protein-Protein Interactions Prediction Accuracy Using XGBoost Feature Selection and Stacked Ensemble Classifier,

    C. Chen et al., “Improving Protein-Protein Interactions Prediction Accuracy Using XGBoost Feature Selection and Stacked Ensemble Classifier, ” Computers in Biology and Medicine , vol. 123, p. 103899, Aug. 2020, doi: 10.1016/j.compbiomed.2020.103899

  65. [74]

    PLM-interact: Extending Protein Language Models to Predict Protein-Protein Interac- tions

    D. Liu et al., “PLM-interact: Extending Protein Language Models to Predict Protein-Protein Interac- tions. ” Nov. 2024. doi: 10.1101/2024.11.05.622169

  66. [75]

    PRING: Rethinking Protein-Protein Interaction Prediction from Pairs to Graphs,

    X. Zheng et al., “PRING: Rethinking Protein-Protein Interaction Prediction from Pairs to Graphs, ” no. arXiv:2507.05101. arXiv, Jul. 2025. doi: 10.48550/arXiv.2507.05101

  67. [76]

    Flaws in Evaluation Schemes for Pair-Input Computational Predictions,

    Y. Park and E. M. Marcotte, “Flaws in Evaluation Schemes for Pair-Input Computational Predictions, ” Nature Methods, vol. 9, no. 12, pp. 1134–1136, Dec. 2012, doi: 10.1038/nmeth.2259

  68. [77]

    Benchmark Evaluation of Protein–Protein Interaction Predic- tion Algorithms,

    B. Dunham and M. K. Ganapathiraju, “Benchmark Evaluation of Protein–Protein Interaction Predic- tion Algorithms, ” Molecules, vol. 27, no. 1, p. 41, Dec. 2021, doi: 10.3390/molecules27010041

  69. [78]

    R. O. Duda, D. G. Stork, and P. E. Hart, Pattern Classification, 2nd ed. New York: Wiley, 2001

  70. [79]

    C. M. Bishop, Pattern Recognition and Machine Learning . in Information Science and Statistics. New York: Springer, 2006

  71. [80]

    Goodfellow, Y

    I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. in Adaptive Computation and Machine Learning. Cambridge, Massachusetts: The MIT Press, 2016

  72. [81]

    Multi-Schema Computational Prediction of the Comprehensive SARS-CoV-2 vs. Human Interactome,

    K. Dick, A. Chopra, K. K. Biggar, and J. R. Green, “Multi-Schema Computational Prediction of the Comprehensive SARS-CoV-2 vs. Human Interactome, ” PeerJ, vol. 9, p. e11117, Apr. 2021, doi: 10.7717/ peerj.11117

  73. [82]

    Prediction of Human-Virus Protein-Protein Interactions through a Sequence Embedding-Based Machine Learning Method,

    X. Yang, S. Yang, Q. Li, S. Wuchty, and Z. Zhang, “Prediction of Human-Virus Protein-Protein Interactions through a Sequence Embedding-Based Machine Learning Method, ” Computational and Structural Biotechnology Journal, vol. 18, pp. 153–161, Jan. 2020, doi: 10.1016/j.csbj.2019.12.005

  74. [83]

    LSTM-PHV: Prediction of Human-Virus Protein– Protein Interactions by LSTM with Word2vec,

    S. Tsukiyama, M. M. Hasan, S. Fujii, and H. Kurata, “LSTM-PHV: Prediction of Human-Virus Protein– Protein Interactions by LSTM with Word2vec, ” Briefings in Bioinformatics, vol. 22, no. 6, p. bbab228, Nov. 2021, doi: 10.1093/bib/bbab228

  75. [84]

    Transfer Learning via Multi-Scale Convolutional Neural Layers for Human–Virus Protein–Protein Interaction Prediction,

    X. Yang, S. Yang, X. Lian, S. Wuchty, and Z. Zhang, “Transfer Learning via Multi-Scale Convolutional Neural Layers for Human–Virus Protein–Protein Interaction Prediction, ” Bioinformatics, vol. 37, no. 24, pp. 4771–4778, Dec. 2021, doi: 10.1093/bioinformatics/btab533. 25

  76. [85]

    A Multitask Transfer Learning Framework for the Prediction of Virus-Human Protein–Protein Interactions,

    T. N. Dong, G. Brogden, G. Gerold, and M. Khosla, “A Multitask Transfer Learning Framework for the Prediction of Virus-Human Protein–Protein Interactions, ” BMC Bioinformatics, vol. 22, no. 1, pp. 1– 24, Dec. 2021, doi: 10.1186/s12859-021-04484-y

  77. [86]

    Large-Scale Data Mining Pipeline for Identifying Novel Soybean Genes Involved in Resistance against the Soybean Cyst Nematode,

    N. Nissan et al., “Large-Scale Data Mining Pipeline for Identifying Novel Soybean Genes Involved in Resistance against the Soybean Cyst Nematode, ” Frontiers in Bioinformatics, vol. 3, Jun. 2023, doi: 10.3389/fbinf.2023.1199675

  78. [87]

    Predicting Novel Protein-Protein Interactions between the HIV-1 Virus and Homo Sapiens,

    B. Barnes et al., “Predicting Novel Protein-Protein Interactions between the HIV-1 Virus and Homo Sapiens, ” in 2016 IEEE EMBS International Student Conference (ISC) , May 2016, pp. 1–4. doi: 10.1109/ EMBSISC.2016.7508598

  79. [88]

    Topsy-Turvy: Integrating a Global View into Sequence-Based PPI Prediction,

    R. Singh, K. Devkota, S. Sledzieski, B. Berger, and L. Cowen, “Topsy-Turvy: Integrating a Global View into Sequence-Based PPI Prediction, ” Bioinformatics, vol. 38, no. Supplement_1, pp. i264–i272, Jun. 2022, doi: 10.1093/bioinformatics/btac258

  80. [89]

    INTREPPPID—an Orthologue-Informed Quintuplet Network for Cross- Species Prediction of Protein–Protein Interaction,

    J. Szymborski and A. Emad, “INTREPPPID—an Orthologue-Informed Quintuplet Network for Cross- Species Prediction of Protein–Protein Interaction, ” Briefings in Bioinformatics, vol. 25, no. 5, p. bbae405, Sep. 2024, doi: 10.1093/bib/bbae405

  81. [90]

    SENSE-PPI Reconstructs Interactomes within, across, and between Species at the Genome Scale,

    K. Volzhenin, L. Bittner, and A. Carbone, “SENSE-PPI Reconstructs Interactomes within, across, and between Species at the Genome Scale, ” iScience, vol. 27, no. 7, p. 110371, Jul. 2024, doi: 10.1016/ j.isci.2024.110371

  82. [91]

    Species-Specific microRNA Discovery and Target Prediction in the Soybean Cyst Nematode,

    V. Ajila et al. , “Species-Specific microRNA Discovery and Target Prediction in the Soybean Cyst Nematode, ” Scientific Reports, vol. 13, no. 1, p. 17657, Oct. 2023, doi: 10.1038/s41598-023-44469-w

  83. [92]

    Proteome-Wide Prediction of Lysine Methylation Leads to Identification of H2BK43 Methylation and Outlines the Potential Methyllysine Proteome,

    K. K. Biggar et al. , “Proteome-Wide Prediction of Lysine Methylation Leads to Identification of H2BK43 Methylation and Outlines the Potential Methyllysine Proteome, ” Cell Reports, vol. 32, no. 2, p. 107896, Jul. 2020, doi: 10.1016/j.celrep.2020.107896

  84. [93]

    Machine Learning Prediction of Antimicrobial Peptides,

    G. Wang, I. I. Vaisman, and M. L. van Hoek, “Machine Learning Prediction of Antimicrobial Peptides, ” Methods in molecular biology (Clifton, N.J.), vol. 2405, pp. 1–37, 2022, doi: 10.1007/978-1-0716-1855-4_1

  85. [94]

    The Impact of Data Difficulty Factors on Classification of Imbalanced and Concept Drifting Data Streams,

    D. Brzezinski, L. L. Minku, T. Pewinski, J. Stefanowski, and A. Szumaczuk, “The Impact of Data Difficulty Factors on Classification of Imbalanced and Concept Drifting Data Streams, ” Knowledge and Information Systems, vol. 63, no. 6, pp. 1429–1469, Jun. 2021, doi: 10.1007/s101...

  86. [95]

    The Receiver Operating Characteristic Curve Accurately Assesses Imbalanced Datasets,

    E. Richardson, R. Trevizani, J. A. Greenbaum, H. Carter, M. Nielsen, and B. Peters, “The Receiver Operating Characteristic Curve Accurately Assesses Imbalanced Datasets, ” Patterns, vol. 5, no. 6, p. 100994, Jun. 2024, doi: 10.1016/j.patter.2024.100994

  87. [96]

    Addressing Data Imbalance in Machine Learning: Challenges and Approaches,

    M. Langote, N. Zade, and S. Gundewar, “Addressing Data Imbalance in Machine Learning: Challenges and Approaches, ” in 2025 6th International Conference on Mobile Computing and Sustainable Informatics (ICMCSI), Jan. 2025, pp. 1745–1749. doi: 10.1109/ICMCSI64620.2025.10883059

  88. [97]

    Reciprocal Perspective for Improved Protein-Protein Interaction Prediction,

    K. Dick and J. R. Green, “Reciprocal Perspective for Improved Protein-Protein Interaction Prediction, ” Scientific Reports, vol. 8, no. 1, p. 11694, 2018, doi: 10.1038/s41598-018-30044-1

  89. [98]

    An Empirical Framework for Binary Interactome Mapping,

    K. Venkatesan et al., “An Empirical Framework for Binary Interactome Mapping, ” Nature Methods, vol. 6, no. 1, pp. 83–90, Jan. 2009, doi: 10.1038/nmeth.1280

  90. [99]

    Estimating the Size of the Human Interactome,

    M. P. H. Stumpf et al., “Estimating the Size of the Human Interactome, ” Proceedings of the National Academy of Sciences, vol. 105, no. 19, pp. 6959–6964, May 2008, doi: 10.1073/pnas.0708078105

  91. [100]

    Prediction of Protein Cellular Attributes Using Pseudo-Amino Acid Composition,

    K.-C. Chou, “Prediction of Protein Cellular Attributes Using Pseudo-Amino Acid Composition, ” Pro- teins: Structure, Function, and Bioinformatics, vol. 43, no. 3, pp. 246–255, 2001, doi: 10.1002/prot.1035. 26

  92. [101]

    Predicting Protein–Protein Interactions Based Only on Sequences Information,

    J. Shen et al., “Predicting Protein–Protein Interactions Based Only on Sequences Information, ” Pro- ceedings of the National Academy of Sciences, vol. 104, no. 11, pp. 4337–4341, Mar. 2007, doi: 10.1073/ pnas.0607879104

  93. [102]

    Composition, Transition and Distribution (CTD) — A Dynamic Feature for Predictions Based on Hierarchical Structure of Cellular Sorting,

    G. Govindan and A. S. Nair, “Composition, Transition and Distribution (CTD) — A Dynamic Feature for Predictions Based on Hierarchical Structure of Cellular Sorting, ” in 2011 Annual IEEE India Conference, Dec. 2011, pp. 1–6. doi: 10.1109/INDCON.2011.6139332

  94. [103]

    Gapped BLAST and PSI-BLAST: A New Generation of Protein Database Search Programs,

    S. F. Altschul et al., “Gapped BLAST and PSI-BLAST: A New Generation of Protein Database Search Programs, ” Nucleic Acids Research, vol. 25, no. 17, pp. 3389–3402, Sep. 1997, doi: 10.1093/nar/25.17.3389

  95. [104]

    ProtDCal: A Program to Compute General- Purpose-Numerical Descriptors for Sequences and 3D-structures of Proteins,

    Y. B. Ruiz-Blanco, W. Paz, J. Green, and Y. Marrero-Ponce, “ProtDCal: A Program to Compute General- Purpose-Numerical Descriptors for Sequences and 3D-structures of Proteins, ” BMC Bioinformatics, vol. 16, no. 1, p. 162, May 2015, doi: 10.1186/s12859-015-0586-0

  96. [105]

    Efficient Estimation of Word Representations in Vector Space,

    T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient Estimation of Word Representations in Vector Space, ” no. arXiv:1301.3781. arXiv, Sep. 2013. doi: 10.48550/arXiv.1301.3781

  97. [106]

    An Integration of Deep Learning with Feature Embedding for Protein-Protein Interaction Prediction,

    Y. Yao, X. Du, Y. Diao, and H. Zhu, “An Integration of Deep Learning with Feature Embedding for Protein-Protein Interaction Prediction, ” PeerJ, vol. 7, p. e7126, 2019, doi: 10.7717/peerj.7126

  98. [107]

    Predicting Protein–Protein Interactions through Sequence-Based Deep Learning,

    S. Hashemifar, B. Neyshabur, A. A. Khan, and J. Xu, “Predicting Protein–Protein Interactions through Sequence-Based Deep Learning, ” Bioinformatics, vol. 34, no. 17, pp. i802–i810, Sep. 2018, doi: 10.1093/ bioinformatics/bty573

  99. [108]

    DeepTrio: A Ternary Prediction System for Protein–Protein Interaction Using Mask Multiple Parallel Convolutional Neural Networks,

    X. Hu, C. Feng, Y. Zhou, A. Harrison, and M. Chen, “DeepTrio: A Ternary Prediction System for Protein–Protein Interaction Using Mask Multiple Parallel Convolutional Neural Networks, ” Bioinfor- matics, vol. 38, no. 3, pp. 694–702, Jan. 2022, doi: 10.1093/bioinformatics/btab737

  100. [109]

    ProtInteract: A Deep Learning Framework for Predicting Protein–Protein Interactions,

    F. Soleymani, E. Paquet, H. L. Viktor, W. Michalowski, and D. Spinello, “ProtInteract: A Deep Learning Framework for Predicting Protein–Protein Interactions, ” Computational and Structural Biotechnology Journal, vol. 21, pp. 1324–1348, Jan. 2023, doi: 10.1016/j.csbj.2023.01.028

  101. [110]

    Improving Protein-Protein Interaction Prediction Using Protein Language Model and Protein Network Features,

    J. Hu, Z. Li, B. Rao, M. A. Thafar, and M. Arif, “Improving Protein-Protein Interaction Prediction Using Protein Language Model and Protein Network Features, ” Analytical Biochemistry, vol. 693, p. 115550, Oct. 2024, doi: 10.1016/j.ab.2024.115550

  102. [111]

    xCAPT5: Protein–Protein Interaction Prediction Using Deep and Wide Multi-Kernel Pooling Convolutional Neural Networks with Protein Language Model,

    T. H. Dang and T. A. Vu, “xCAPT5: Protein–Protein Interaction Prediction Using Deep and Wide Multi-Kernel Pooling Convolutional Neural Networks with Protein Language Model, ” BMC Bioinfor- matics, vol. 25, no. 1, pp. 1–20, Dec. 2024, doi: 10.1186/s12859-024-05725-6

  103. [112]

    DeNovo: Virus-Host Sequence-Based Protein–Protein Inter- action Prediction,

    F.-E. Eid, M. ElHefnawi, and L. S. Heath, “DeNovo: Virus-Host Sequence-Based Protein–Protein Inter- action Prediction, ” Bioinformatics, vol. 32, no. 8, pp. 1144–1150, Apr. 2016, doi: 10.1093/bioinformatics/ btv737

  104. [113]

    Prediction of Protein-Protein Interactions Based on Ensemble Residual Convolutional Neural Network,

    H. Gao, C. Chen, S. Li, C. Wang, W. Zhou, and B. Yu, “Prediction of Protein-Protein Interactions Based on Ensemble Residual Convolutional Neural Network, ” Computers in Biology and Medicine , vol. 152, p. 106471, Jan. 2023, doi: 10.1016/j.compbiomed.2022.106471

  105. [114]

    SDNN-PPI: Self-Attention with Deep Neural Network Effect on Protein-Protein Interaction Prediction,

    X. Li, P. Han, G. Wang, W. Chen, S. Wang, and T. Song, “SDNN-PPI: Self-Attention with Deep Neural Network Effect on Protein-Protein Interaction Prediction, ” BMC Genomics, vol. 23, no. 1, pp. 1–14, Dec. 2022, doi: 10.1186/s12864-022-08687-2

  106. [115]

    Attention Is All You Need

    A. Vaswani et al., “Attention Is All You Need. ” arXiv, 2017. doi: 10.48550/ARXIV.1706.03762

  107. [116]

    GPT-4 Technical Report,

    OpenAI et al. , “GPT-4 Technical Report, ” no. arXiv:2303.08774. arXiv, Mar. 2024. doi: 10.48550/ arXiv.2303.08774

  108. [117]

    Gemini 1.5: Unlocking Multimodal Understanding across Millions of Tokens of Context,

    G. Team et al. , “Gemini 1.5: Unlocking Multimodal Understanding across Millions of Tokens of Context, ” no. arXiv:2403.05530. arXiv, Dec. 2024. doi: 10.48550/arXiv.2403.05530. 27

  109. [118]

    Hidden Markov Models in Computa- tional Biology: Applications to Protein Modeling,

    A. Krogh, M. Brown, I. S. Mian, K. Sjölander, and D. Haussler, “Hidden Markov Models in Computa- tional Biology: Applications to Protein Modeling, ” Journal of Molecular Biology , vol. 235, no. 5, pp. 1501–1531, Feb. 1994, doi: 10.1006/jmbi.1994.1104

  110. [119]

    A Sequence-Profile-Based HMM for Predicting and Discriminating 𝑏𝜂 Barrel Membrane Proteins

    P. L. Martelli, P. Fariselli, A. Krogh, and R. Casadio, “A Sequence-Profile-Based HMM for Predicting and Discriminating 𝑏𝜂 Barrel Membrane Proteins”, Bioinformatics, vol. 18, no. suppl_1, pp. S46–S53, Jul. 2002, doi: 10.1093/bioinformatics/18.suppl_1.s46

  111. [120]

    Protein Homology Detection by HMM–HMM Comparison,

    J. Söding, “Protein Homology Detection by HMM–HMM Comparison, ” Bioinformatics, vol. 21, no. 7, pp. 951–960, Apr. 2005, doi: 10.1093/bioinformatics/bti125

  112. [121]

    Learning the Protein Language: Evolution, Structure, and Function,

    T. Bepler and B. Berger, “Learning the Protein Language: Evolution, Structure, and Function, ” Cell Systems, vol. 12, no. 6, pp. 654–669, Jun. 2021, doi: 10.1016/j.cels.2021.05.017

  113. [122]

    UniRef Clusters: A Comprehensive and Scalable Alternative for Improving Sequence Similarity Searches,

    B. E. Suzek, Y. Wang, H. Huang, P. B. McGarvey, and C. H. Wu, “UniRef Clusters: A Comprehensive and Scalable Alternative for Improving Sequence Similarity Searches, ” Bioinformatics, vol. 31, no. 6, pp. 926–932, Mar. 2015, doi: 10.1093/bioinformatics/btu739

  114. [123]

    Clustering Huge Protein Sequence Sets in Linear Time,

    M. Steinegger and J. Söding, “Clustering Huge Protein Sequence Sets in Linear Time, ” Nature Commu- nications, vol. 9, no. 1, p. 2542, Jun. 2018, doi: 10.1038/s41467-018-04964-5

  115. [124]

    Protein Language Models and Machine Learning Facilitate the Identification of Antimicrobial Peptides,

    D. Medina-Ortiz et al., “Protein Language Models and Machine Learning Facilitate the Identification of Antimicrobial Peptides, ” International Journal of Molecular Sciences , vol. 25, no. 16, p. 8851, Aug. 2024, doi: 10.3390/ijms25168851

  116. [125]

    Leveraging Protein Language Models for Robust Antimicrobial Peptide Detection,

    L. Zhang et al., “Leveraging Protein Language Models for Robust Antimicrobial Peptide Detection, ” Methods, vol. 238, pp. 19–26, Jun. 2025, doi: 10.1016/j.ymeth.2025.03.002

  117. [126]

    TUnA: An Uncertainty-Aware Transformer Model for Sequence-Based Protein–Protein Interaction Prediction,

    Y. S. Ko, J. Parkinson, C. Liu, and W. Wang, “TUnA: An Uncertainty-Aware Transformer Model for Sequence-Based Protein–Protein Interaction Prediction, ” Briefings in Bioinformatics, vol. 25, no. 5, p. bbae359, Sep. 2024, doi: 10.1093/bib/bbae359

  118. [127]

    ProtTrans: Towards Cracking the Language of Lifes Code Through Self-Supervised Deep Learning and High Performance Computing,

    A. Elnaggar et al., “ProtTrans: Towards Cracking the Language of Lifes Code Through Self-Supervised Deep Learning and High Performance Computing, ” IEEE Transactions on Pattern Analysis and Machine Intelligence, p. 1, 2021, doi: 10.1109/TPAMI.2021.3095381

  119. [128]

    Ankh: Optimized Protein Language Model Unlocks General-Purpose Modelling

    A. Elnaggar et al., “Ankh: Optimized Protein Language Model Unlocks General-Purpose Modelling. ” bioRxiv, Jan. 2023. doi: 10.1101/2023.01.16.524265

  120. [129]

    MP-PIPE: A Massively Parallel Protein-Protein Interaction Prediction Engine,

    A. Schoenrock, F. Dehne, J. R. Green, A. Golshani, and S. Pitre, “MP-PIPE: A Massively Parallel Protein-Protein Interaction Prediction Engine, ” in Proceedings of the International Conference on Super- computing - ICS '11, Tucson, Arizona, USA: ACM Press, 2011, p. 327. doi: 10...

  121. [130]

    Computing the Human Interactome

    J. Zhang et al. , “Computing the Human Interactome. ” bioRxiv, Oct. 2024. doi: 10.1101/2024.10.01.615885

  122. [131]

    A Model of Evolutionary Change in Proteins,

    M. Dayhoff, R. Schwartz, and B. Orcutt, “A Model of Evolutionary Change in Proteins, ” Atlas of protein sequence and structure, vol. 5, pp. 345–352, 1978

  123. [132]

    PEPPI: Whole-proteome Protein- protein Interaction Prediction through Structure and Sequence Similarity, Functional Association, and Machine Learning,

    E. W. Bell, J. H. Schwartz, P. L. Freddolino, and Y. Zhang, “PEPPI: Whole-proteome Protein- protein Interaction Prediction through Structure and Sequence Similarity, Functional Association, and Machine Learning, ” Journal of Molecular Biology , vol. 434, no. 11, p. 167530, Jun...

  124. [133]

    VirusMentha: A New Resource for Virus-Host Protein Interactions,

    A. Calderone, L. Licata, and G. Cesareni, “VirusMentha: A New Resource for Virus-Host Protein Interactions, ” Nucleic Acids Research, vol. 43, no. Database issue, pp. D588–592, Jan. 2015, doi: 10.1093/ nar/gku830. 28

  125. [134]

    The Database of Interacting Proteins: 2004 Update,

    L. Salwinski, C. S. Miller, A. J. Smith, F. K. Pettit, J. U. Bowie, and D. Eisenberg, “The Database of Interacting Proteins: 2004 Update, ” Nucleic Acids Research, vol. 32, no. Database issue, pp. D449–451, Jan. 2004, doi: 10.1093/nar/gkh086

  126. [135]

    HINT: High-quality Protein Interactomes and Their Applications in Understanding Human Disease,

    J. Das and H. Yu, “HINT: High-quality Protein Interactomes and Their Applications in Understanding Human Disease, ” BMC Systems Biology, vol. 6, no. 1, p. 92, Jul. 2012, doi: 10.1186/1752-0509-6-92

  127. [136]

    Human Protein Reference Database as a Discovery Resource for Proteomics,

    S. Peri et al., “Human Protein Reference Database as a Discovery Resource for Proteomics, ” Nucleic Acids Research, vol. 32, no. Database issue, pp. D497–D501, Jan. 2004, doi: 10.1093/nar/gkh070

  128. [137]

    3did: A Catalog of Domain-Based Interactions of Known Three-Dimensional Structure,

    R. Mosca, A. Céol, A. Stein, R. Olivella, and P. Aloy, “3did: A Catalog of Domain-Based Interactions of Known Three-Dimensional Structure, ” Nucleic Acids Research, vol. 42, no. D1, pp. D374–D379, Jan. 2014, doi: 10.1093/nar/gkt887

  129. [138]

    iPfam: A Database of Protein Family and Domain Interactions Found in the Protein Data Bank,

    R. D. Finn, B. L. Miller, J. Clements, and A. Bateman, “iPfam: A Database of Protein Family and Domain Interactions Found in the Protein Data Bank, ” Nucleic Acids Research, vol. 42, no. D1, pp. D364–D373, Jan. 2014, doi: 10.1093/nar/gkt1210

  130. [139]

    Positome: A Method for Improving Protein-Protein Interaction Quality and Prediction Accuracy,

    K. Dick, F. Dehne, A. Golshani, and J. R. Green, “Positome: A Method for Improving Protein-Protein Interaction Quality and Prediction Accuracy, ” in 2017 IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), Manchester, United Kingd...

  131. [140]

    Prediction of Protein-Protein Interactions Using Local Description of Amino Acid Sequence,

    Y. Z. Zhou, Y. Gao, and Y. Y. Zheng, “Prediction of Protein-Protein Interactions Using Local Description of Amino Acid Sequence, ” Advances in Computer Science and Education Applications. Springer, Berlin, Heidelberg, pp. 254–262, 2011. doi: 10.1007/978-3-642-22456-0_37

  132. [141]

    HPIDB 2.0: A Curated Database for Host–Pathogen Interactions,

    M. G. Ammari, C. R. Gresham, F. M. McCarthy, and B. Nanduri, “HPIDB 2.0: A Curated Database for Host–Pathogen Interactions, ” Database, vol. 2016, Jan. 2016, doi: 10.1093/database/baw103

  133. [142]

    VirHostNet 2.0: Surfing on the Web of Virus/Host Molecular Interactions Data,

    T. Guirimand, S. Delmotte, and V. Navratil, “VirHostNet 2.0: Surfing on the Web of Virus/Host Molecular Interactions Data, ” Nucleic Acids Research, vol. 43, no. Database issue, pp. D583–587, Jan. 2015, doi: 10.1093/nar/gku1121

  134. [143]

    PHISTO: Pathogen–Host Interaction Search Tool,

    S. Durmuo̧s Tekir et al., “PHISTO: Pathogen–Host Interaction Search Tool, ” Bioinformatics, vol. 29, no. 10, pp. 1357–1358, May 2013, doi: 10.1093/bioinformatics/btt137

  135. [144]

    A SARS-CoV-2 Protein Interaction Map Reveals Targets for Drug Repurposing,

    D. E. Gordon et al., “A SARS-CoV-2 Protein Interaction Map Reveals Targets for Drug Repurposing, ” Nature, vol. 583, no. 7816, pp. 459–468, Jul. 2020, doi: 10.1038/s41586-020-2286-9

  136. [145]

    Virus-Host Interactome and Proteomic Survey Reveal Potential Virulence Factors Influ- encing SARS-CoV-2 Pathogenesis,

    J. Li et al., “Virus-Host Interactome and Proteomic Survey Reveal Potential Virulence Factors Influ- encing SARS-CoV-2 Pathogenesis, ” Med (New York, N.Y.) , vol. 2, no. 1, pp. 99–112, Jan. 2021, doi: 10.1016/j.medj.2020.07.002

  137. [146]

    Learning Protein Sequence Embeddings Using Information from Structure,

    T. Bepler and B. Berger, “Learning Protein Sequence Embeddings Using Information from Structure, ” in International Conference on Learning Representations, arXiv, 2019. doi: 10.48550/arXiv.1902.08661

  138. [147]

    Prediction of Flexible/Rigid Regions from Protein Sequences Using k-Spaced Amino Acid Pairs,

    K. Chen, L. A. Kurgan, and J. Ruan, “Prediction of Flexible/Rigid Regions from Protein Sequences Using k-Spaced Amino Acid Pairs, ” BMC Structural Biology , vol. 7, no. 1, pp. 1–13, Dec. 2007, doi: 10.1186/1472-6807-7-25

  139. [148]

    Large-Scale Prediction of Human Protein-Protein Interactions from Amino Acid Sequence Based on Latent Topic Features,

    X.-Y. Pan, Y.-N. Zhang, and H.-B. Shen, “Large-Scale Prediction of Human Protein-Protein Interactions from Amino Acid Sequence Based on Latent Topic Features, ” Journal of Proteome Research, vol. 9, no. 10, pp. 4992–5001, Oct. 2010, doi: 10.1021/pr100618t

  140. [149]

    Drug Target Protein-Protein Interaction Networks: A Systematic Perspective,

    Y. Feng, Q. Wang, and T. Wang, “Drug Target Protein-Protein Interaction Networks: A Systematic Perspective, ” BioMed Research International, vol. 2017, p. 1289259, 2017, doi: 10.1155/2017/1289259. 29

  141. [150]

    Network-Based Approaches in Drug Discovery and Early Development,

    J. M. Harrold, M. Ramanathan, and D. E. Mager, “Network-Based Approaches in Drug Discovery and Early Development, ” Clinical Pharmacology and Therapeutics , vol. 94, no. 6, pp. 651–658, Dec. 2013, doi: 10.1038/clpt.2013.176

  142. [151]

    Identifying Causal Genes and Dysregulated Pathways in Complex Diseases,

    Y.-A. Kim, S. Wuchty, and T. M. Przytycka, “Identifying Causal Genes and Dysregulated Pathways in Complex Diseases, ” PLOS Computational Biology , vol. 7, no. 3, p. e1001095, Mar. 2011, doi: 10.1371/ journal.pcbi.1001095

  143. [152]

    Identification of Drug and Protein-Protein Interaction Network among Stress and Depression: A Bioinformatics Ap- proach,

    M. A. Basar, M. F. Hosen, B. Kumar Paul, M. R. Hasan, S. M. Shamim, and T. Bhuyian, “Identification of Drug and Protein-Protein Interaction Network among Stress and Depression: A Bioinformatics Ap- proach, ” Informatics in Medicine Unlocked, vol. 37, p. 101174, Jan. 2023, doi:...

  144. [153]

    Utility of Network Integrity Methods in Therapeutic Target Identification,

    Q. Peng and N. J. Schork, “Utility of Network Integrity Methods in Therapeutic Target Identification, ” Frontiers in Genetics, vol. 5, p. 12, 2014, doi: 10.3389/fgene.2014.00012

  145. [154]

    Computational/in Silico Methods in Drug Target and Lead Prediction,

    F. E. Agamah et al., “Computational/in Silico Methods in Drug Target and Lead Prediction, ” Briefings in Bioinformatics, vol. 21, no. 5, pp. 1663–1675, Nov. 2019, doi: 10.1093/bib/bbz103

  146. [155]

    In Silico Methods for Identification of Potential Therapeutic Targets,

    X. Zhang et al., “In Silico Methods for Identification of Potential Therapeutic Targets, ” Interdiscipli- nary Sciences: Computational Life Sciences , vol. 14, no. 2, pp. 285–310, Jun. 2022, doi: 10.1007/ s12539-021-00491-y

  147. [156]

    Therapeutic Peptides: Current Applications and Future Directions,

    L. Wang et al., “Therapeutic Peptides: Current Applications and Future Directions, ” Signal Transduc- tion and Targeted Therapy, vol. 7, no. 1, p. 48, Feb. 2022, doi: 10.1038/s41392-022-00904-4

  148. [157]

    Focus on Therapeutic Peptides and Their Delivery,

    E. Rosson, F. Lux, L. David, Y. Godfrin, O. Tillement, and E. Thomas, “Focus on Therapeutic Peptides and Their Delivery, ” International Journal of Pharmaceutics, vol. 675, p. 125555, Apr. 2025, doi: 10.1016/ j.ijpharm.2025.125555

  149. [158]

    Therapeutic Peptides Targeting PPI in Clinical Development: Overview, Mechanism of Action and Perspectives,

    W. Cabri et al., “Therapeutic Peptides Targeting PPI in Clinical Development: Overview, Mechanism of Action and Perspectives, ” Frontiers in Molecular Biosciences , vol. 8, Jun. 2021, doi: 10.3389/ fmolb.2021.697586

  150. [159]

    Solid-Phase Peptide Synthesis: From Standard Procedures to the Synthesis of Difficult Sequences,

    I. Coin, M. Beyermann, and M. Bienert, “Solid-Phase Peptide Synthesis: From Standard Procedures to the Synthesis of Difficult Sequences, ” Nature Protocols, vol. 2, no. 12, pp. 3247–3256, Dec. 2007, doi: 10.1038/nprot.2007.454

  151. [160]

    Engineering Inhibitory Proteins with InSiPS: The in-Silico Protein Synthesizer,

    A. Schoenrock et al., “Engineering Inhibitory Proteins with InSiPS: The in-Silico Protein Synthesizer, ” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis on - SC '15, Austin, Texas: ACM Press, 2015, pp. 1–11. doi: ...

  152. [161]

    In Silico Engineering of Synthetic Binding Proteins from Random Amino Acid Sequences,

    D. Burnside et al., “In Silico Engineering of Synthetic Binding Proteins from Random Amino Acid Sequences, ” iScience, vol. 11, pp. 375–387, Jan. 2019, doi: 10.1016/j.isci.2018.11.038

  153. [162]

    A Computational Approach to Rapidly Design Peptides That Detect SARS- CoV-2 Surface Protein S,

    M. Hajikarimlou et al., “A Computational Approach to Rapidly Design Peptides That Detect SARS- CoV-2 Surface Protein S, ” NAR Genomics and Bioinformatics , vol. 4, no. 3, p. lqac58, Jul. 2022, doi: 10.1093/nargab/lqac058

  154. [163]

    A Deep-Learning Framework for Multi-Level Peptide–Protein Interaction Prediction,

    Y. Lei et al., “A Deep-Learning Framework for Multi-Level Peptide–Protein Interaction Prediction, ” Nature Communications, vol. 12, no. 1, p. 5465, Dec. 2021, doi: 10.1038/s41467-021-25772-4

  155. [164]

    Design of Peptide-Based Protein Degraders via Contrastive Deep Learning

    K. Palepu et al., “Design of Peptide-Based Protein Degraders via Contrastive Deep Learning. ” Cold Spring Harbor Laboratory, May 2022. doi: 10.1101/2022.05.23.493169

  156. [165]

    PepMLM: Target Sequence-Conditioned Generation of Peptide Binders via Masked Language Modeling,

    T. Chen et al., “PepMLM: Target Sequence-Conditioned Generation of Peptide Binders via Masked Language Modeling, ” 2024

  157. [166]

    De Novo Design of Peptide Binders to Conformationally Diverse Targets with Contrastive Language Modeling,

    S. Bhat et al. , “De Novo Design of Peptide Binders to Conformationally Diverse Targets with Contrastive Language Modeling, ” Science Advances, vol. 11, no. 4, p. eadr8638, Jan. 2025, doi: 10.1126/ sciadv.adr8638. 30

  158. [167]

    Deciphering Antibody Affinity Maturation with Language Models and Weakly Supervised Learning,

    J. A. Ruffolo, J. J. Gray, and J. Sulam, “Deciphering Antibody Affinity Maturation with Language Models and Weakly Supervised Learning, ” no. arXiv:2112.07782. arXiv, Dec. 2021. doi: 10.48550/ arXiv.2112.07782

  159. [168]

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, ” no. arXiv:1810.04805. arXiv, May 2019. doi: 10.48550/ arXiv.1810.04805

  160. [169]

    Deciphering the Language of Antibodies Using Self-Supervised Learning,

    J. Leem, L. S. Mitchell, J. H. R. Farmery, J. Barton, and J. D. Galson, “Deciphering the Language of Antibodies Using Self-Supervised Learning, ” Patterns, vol. 3, no. 7, p. 100513, Jul. 2022, doi: 10.1016/ j.patter.2022.100513

  161. [170]

    RoBERTa: A Robustly Optimized BERT Pretraining Approach,

    Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach, ” no. arXiv:1907.11692. arXiv, Jul. 2019. doi: 10.48550/arXiv.1907.11692

  162. [171]

    Large Scale Paired Antibody Language Models,

    H. Kenlay, F. A. Dreyer, A. Kovaltsuk, D. Miketa, D. Pires, and C. M. Deane, “Large Scale Paired Antibody Language Models, ” PLOS Computational Biology, vol. 20, no. 12, p. e1012646, Dec. 2024, doi: 10.1371/journal.pcbi.1012646

  163. [172]

    Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,

    C. Raffel et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, ” no. arXiv:1910.10683. arXiv, 2019. doi: 10.48550/arXiv.1910.10683

  164. [173]

    AbLang: An Antibody Language Model for Completing Antibody Sequences,

    T. H. Olsen, I. H. Moal, and C. M. Deane, “AbLang: An Antibody Language Model for Completing Antibody Sequences, ” Bioinformatics Advances, vol. 2, no. 1, p. vbac46, Jan. 2022, doi: 10.1093/bioadv/ vbac046

  165. [174]

    IgLM: Infilling Language Modeling for Antibody Sequence Design,

    R. W. Shuai, J. A. Ruffolo, and J. J. Gray, “IgLM: Infilling Language Modeling for Antibody Sequence Design, ” Cell Systems, vol. 14, no. 11, pp. 979–989, Nov. 2023, doi: 10.1016/j.cels.2023.10.001

  166. [175]

    Language Models Are Unsupervised Multitask Learners,

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models Are Unsupervised Multitask Learners, ” 2019

  167. [176]

    Observed Antibody Space: A Resource for Data Mining Next-Generation Sequencing of Antibody Repertoires,

    A. Kovaltsuk, J. Leem, S. Kelm, J. Snowden, C. M. Deane, and K. Krawczyk, “Observed Antibody Space: A Resource for Data Mining Next-Generation Sequencing of Antibody Repertoires, ” The Journal of Immunology, vol. 201, no. 8, pp. 2502–2509, Oct. 2018, doi: 10.4049/jimmunol.1800708

  168. [177]

    Efficient Evolution of Human Antibodies from General Protein Language Models,

    B. L. Hie et al., “Efficient Evolution of Human Antibodies from General Protein Language Models, ” Nature Biotechnology, vol. 42, no. 2, pp. 275–283, Feb. 2024, doi: 10.1038/s41587-023-01763-2

  169. [178]

    Biological Structure and Function Emerge from Scaling Unsupervised Learning to 250 Million Protein Sequences,

    A. Rives et al., “Biological Structure and Function Emerge from Scaling Unsupervised Learning to 250 Million Protein Sequences, ” Proceedings of the National Academy of Sciences , vol. 118, no. 15, p. e2016239118, Apr. 2021, doi: 10.1073/pnas.2016239118

  170. [179]

    Language Models Enable Zero-Shot Prediction of the Effects of Mutations on Protein Function,

    J. Meier, R. Rao, R. Verkuil, J. Liu, T. Sercu, and A. Rives, “Language Models Enable Zero-Shot Prediction of the Effects of Mutations on Protein Function, ” in 35th Conference on Neural Information Processing Systems, bioRxiv, Nov. 2021. doi: 10.1101/2021.07.09.450648

  171. [180]

    Are Genomic Language Models All You Need? Exploring Genomic Language Models on Protein Downstream Tasks,

    S. Boshar, E. Trop, B. P. de Almeida, L. Copoiu, and T. Pierrot, “Are Genomic Language Models All You Need? Exploring Genomic Language Models on Protein Downstream Tasks, ” Bioinformatics, vol. 40, no. 9, p. btae529, Sep. 2024, doi: 10.1093/bioinformatics/btae529

  172. [181]

    Genomic Language Models Could Transform Med- icine but Not Yet,

    M. E. Consens, B. Li, A. R. Poetsch, and S. Gilbert, “Genomic Language Models Could Transform Med- icine but Not Yet, ” npj Digital Medicine, vol. 8, no. 1, p. 212, Apr. 2025, doi: 10.1038/s41746-025-01603-4

  173. [182]

    Large Language Models in Genomics—A Perspective on Personalized Medicine,

    S. Ali et al., “Large Language Models in Genomics—A Perspective on Personalized Medicine, ” Bioengi- neering, vol. 12, no. 5, p. 440, Apr. 2025, doi: 10.3390/bioengineering12050440

  174. [183]

    scKEPLM: Knowledge Enhanced Large-Scale Pre-Trained Language Model for Single-Cell Transcriptomics

    Y. Li, G. Qiao, and G. Wang, “scKEPLM: Knowledge Enhanced Large-Scale Pre-Trained Language Model for Single-Cell Transcriptomics. ” bioRxiv, Jul. 2024. doi: 10.1101/2024.07.09.602633. 31

  175. [184]

    CellFM: A Large-Scale Foundation Model Pre-Trained on Transcriptomics of 100 Million Human Cells,

    Y. Zeng et al. , “CellFM: A Large-Scale Foundation Model Pre-Trained on Transcriptomics of 100 Million Human Cells, ” Nature Communications , vol. 16, no. 1, p. 4679, May 2025, doi: 10.1038/ s41467-025-59926-5

  176. [185]

    Trends in Clinical Success Rates,

    K. Smietana, M. Siatkowski, and M. Møller, “Trends in Clinical Success Rates, ” Nature Reviews Drug Discovery, vol. 15, no. 6, pp. 379–380, Jun. 2016, doi: 10.1038/nrd.2016.85. Acknowledgments The authors wish to thank the Natural Sciences and Engineering Research Council of C...

  177. [2025]

    doi: 10.1101/2025.06.14.659707

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.