Pith. sign in

REVIEW 4 major objections 4 minor 126 references

A Survey of Deep Learning Methods in Protein Bioinformatics and its Impact on Protein Design

T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This survey argues that protein design is the inverse of structural and functional prediction, and that recent deep-learning advances in prediction have become the main engine for designing new proteins.

desk verdict A useful but dated survey whose citation errors are fixable and don't sink the design-as-inverse thesis; send it to review with an order to audit the references. read the letter →

arxiv 2501.01477 v1 pith:ZMVMRLNG submitted 2025-01-02 q-bio.BM cs.AI

classification q-bio.BMcs.AI
keywords deeplearningproteinbioinformaticsstructurepredictionfunctiondesignhallucinationgeometricgenerativemodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is a survey, and its contribution is an organizing claim: protein design is best understood as the inverse of structural and functional prediction, and the recent deep-learning successes in prediction are what have made design tractable. It reviews structure prediction from sequence, functional prediction from sequence and structure, and design as sequence generation from function or structure, with the three tasks linked in a single cycle. A sympathetic reader should care because the claim reframes design as a problem that improves automatically as predictors improve, and it explains why methods that work for AlphaFold2-style models—especially back-propagating through a trained network—can be recycled as design engines. The survey also stresses that experimental validation remains the bottleneck that separates plausible designs from real proteins.

What carries the argument

The mechanism that carries the argument is reverse-mode gradient descent on trained prediction networks—the 'deep hallucination' procedure used by several papers it surveys. A structure predictor such as trRosetta maps a sequence to a distribution over inter-residue distances and orientations; by freezing the network's weights and back-propagating the gradient of a structural objective into a randomly initialized sequence, the sequence itself becomes the optimizable variable. This inversion is what allows design to be described as prediction run backwards, and it is the concrete object the survey points to when it says structural and functional prediction advances have directly contributed to design tasks.

What would settle it

Design 100 de novo proteins by hallucinating a state-of-the-art structure predictor and an equal number by physics-based energy minimization, synthesize all 200, and compare in vitro folding rates; if the hallucinated set does not fold more often, or at least as often, the survey's claim that stronger predictors directly drive design loses its load-bearing evidence.

Watch

Extended reading notes

Core claim

The paper's central assertion is that the three canonical problems of protein bioinformatics form a directed cycle: sequence determines structure, structure and sequence determine function, and design is the inverse map from function or structure back to sequence. Within that frame, the paper claims that the strongest recent design results did not come from better energy functions or fragment libraries but from reusing deep-learning predictors—most notably trRosetta and AlphaFold2—either as scoring oracles in generative models or as differentiable networks that can be run backward by gradient ascent to 'hallucinate' sequences for a desired structure or motif. It treats this inversion as the main explanation for why design methods have improved since CASP13, while acknowledging that in silico performance has repeatedly failed to survive in vitro synthesis. The survey also claims the structure-first, function-second ordering is useful because functional labels are sparser than structural data, making function the weakest link in the cycle.

Load-bearing premise

The survey's central map is only as reliable as its second-hand descriptions of dozens of primary papers; those descriptions already contain documented errors, including attributing two distinct methods (dMASIF and ScanNet) to the same reference and crediting an early neural-network contact predictor to the PSI-BLAST paper.

Editorial extensions

If this is right

  • If design is the inverse of prediction, then any sustained improvement in structure or function prediction should translate into better design methods without new design-specific ideas.
  • Structure predictors can serve as oracle feedback inside generative models such as GANs, reinforcement learning, and directed evolution, steering generated sequences toward stable folds.
  • Inverse back-propagation, or hallucination, turns a trained predictor into a sequence generator, so design no longer requires an explicit energy function or fragment library.
  • Because in silico validation alone has repeatedly failed to predict in vitro folding, the field's rate of progress will be capped by experimental synthesis throughput, not just model accuracy.
  • Benchmarks for design should move toward in vitro fold-success rates rather than sequence-recovery or reconstruction scores.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the same back-propagation trick should in principle work on any differentiable function predictor, not just structure predictors, so protein design could be targeted at activity or binding directly if such predictors reach sufficient accuracy.
  • Editorial inference: a direct controlled comparison of hallucination-based design against physics-based Rosetta design on the same test structures would isolate how much of the recent design progress is due to prediction accuracy rather than to the search procedure.
  • Editorial inference: the survey's three-way map suggests a testable prediction: as AlphaFold2-class models are adopted as oracles, reported in vitro success rates of de novo designs should rise; if they stay flat, the causal link from prediction to design would be weaker than claimed.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This paper surveys deep learning methods in protein bioinformatics, organizing the field into three categories—structural prediction, functional prediction, and protein design—and argues that protein design can be viewed as the inverse of structural and functional prediction. It reviews historical and modern structure prediction methods (e.g., RaptorX, AlphaFold1/2, RoseTTAFold), functional prediction from sequence and structure (including GNNs and geometric deep learning), and design methods such as latent-space generation, GANs with structure oracles, directed evolution, and deep hallucination. The paper contains no new experiments; its contribution is a high-level synthesis plus a discussion of challenges (e.g., in silico versus in vitro validation) and future directions. The manuscript text indicates a submission date of October 2021, though it is posted on arXiv in January 2025.

Significance. If the survey's descriptions are accurate, it offers a useful introductory map of deep learning in protein bioinformatics and a clear articulation of the inverse-design viewpoint. The paper explicitly and repeatedly flags the gap between in silico and in vitro validation, it gives a reasonable high-level account of the trajectory from contact prediction to AlphaFold2, and it covers a broad range of methods including geometric deep learning and unsupervised language models. However, because the paper is a survey, its evidentiary value rests entirely on the correctness and completeness of its second-hand method descriptions. The manuscript's own text contains at least two concrete citation errors that affect the traceability of the synthesis, and it omits major post-2021 developments; these issues currently make the map less reliable than a survey should be.

major comments (4)
  1. [§4.2.2–4.2.3] The sentence "Altschul in 1997 [4] explored using neural networks through SLPs for contact prediction" is a factual misattribution. Reference [4] is the PSI-BLAST paper, which describes a sequence-search algorithm and contains no neural-network contact predictor. The correct early neural-network contact prediction work appears to be Lund et al. 1997, which the survey itself cites as [68] in §3. This error is load-bearing for the historical narrative in the structure-prediction section, because a reader cannot trace the claimed development of contact prediction from the cited sources. The passage should be corrected and the surrounding historical claims re-verified against their primary sources.
  2. [§4.2.2 and §4.2.3] Two different methods, dMASIF (spelled "dMASIV" in the text) and ScanNet, are both attributed to the same reference [112], which is the ScanNet preprint by Tubiana et al. The actual dMASIF paper (Sverrisson et al., 2021) is not cited anywhere. This makes it impossible for a reader to verify the descriptions of either method and casts doubt on the reliability of the "function from structure" review as a whole. The authors must add the correct reference for dMASIF and audit the surrounding subsections for similar citation errors.
  3. [§5.3] The central claim that advances in structural prediction have "directly contributed" to protein design is supported in the text almost exclusively by trRosetta-based hallucination works ([8], [81], [110]) and oracle-based generative methods ([38], [50]). Yet §5.3 itself states that AlphaFold2 "has yet to make its impact in protein design." As written, the inverse-design claim is broader than the evidence presented. The authors should either restrict the claim to the specific prediction models actually used in the design methods they review (e.g., trRosetta-era oracles) or provide a more careful account of which structural-prediction advances have and have not been exploited in design, and why.
  4. [Overall] The manuscript is dated October 2021 but posted in January 2025, and its coverage appears to end around 2021. It does not discuss major post-2021 developments such as RFdiffusion, ProteinMPNN/InverseFolding, ESMFold, AlphaFold3, or the widespread use of diffusion models in design. For a survey whose stated purpose is to map the current state of deep learning in protein bioinformatics, the omission of these widely used methods is a load-bearing incompleteness. The authors should either update the survey to cover the 2022–2024 literature or clearly restrict the claimed scope and title to a historical snapshot.
minor comments (4)
  1. [Throughout] There are numerous typographical errors, including "outperfrom," "convoluational," "dMASIV," "disearable," "interporlate," "baysian," "hyrdophobic," "millisconds," "peer reviewewd," and inconsistent capitalization such as "Alphafold2" versus "AlphaFold2." A careful proofreading pass is needed.
  2. [Figures] Several figures are reproduced from external sources ([5], [16], [31], [86], [50]) without explicit permission statements, and the provenance is given only in captions. The authors should confirm that permission or license terms are in order for a published survey.
  3. [End matter] The text ends with a list of bare blog URLs that are not integrated into the reference list or cited in the body. These should either be incorporated as formal references with access dates or removed.
  4. [§4.1.1] The claim that DeepGO "surpassed GOLabeler and NetGO" is stated without a benchmark or time qualification, while the later discussion of TALE says it achieved state-of-the-art results for two of three subclasses. The authors should make the comparison consistent and specify which CAFA evaluation is being referenced.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey performs no independent derivation and contains no fitted constants, author-defined normalizations, or self-citations that could make its claims true by construction.

full rationale

This is a survey with no independent derivation, no fitted constants, no author-defined normalization, and no formal claim that could reduce to its own inputs by construction. Its three-category organization and the framing of protein design as the inverse of structural and functional prediction are descriptive claims about the literature and are sourced to external works, such as [23] for the categorization and [8, 81, 110] for inverse-style design; the paper does not define those methods into existence. The Bayesian expression in Section 5.2 (Eq. 4) is a standard identity used to explain a cited method, not a derivation performed by this paper. There are no self-citations: the sole author is not an author of any cited reference. The citation misattributions noted in the manuscript, in Section 3.1 (Altschul 1997 is the PSI-BLAST paper, not a neural-network contact predictor) and in Sections 4.2.2 and 4.2.3 (two different methods, dMASIF and ScanNet, attributed to the same reference [112]), are accuracy and verifiability failures in a second-hand survey, but they are not circularity in the sense of an output being equivalent to its input. The paper's central claim that prediction advances have directly contributed to design is an empirical and historical synthesis that could be false if its source descriptions are wrong, but nothing in the manuscript makes that claim true by construction. The paper is self-contained as a survey in that it makes no predictive or derivational claim of its own. Score 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

This is a survey; it invents no entities and fits no parameters. The entire review is a synthesis of cited prior work. The only load-bearing assumptions are disciplinary background claims about protein folding and about the value of the reviewed methods, plus the implicit reliability of the survey's retelling of cited results.

assumptions (3)
  • domain assumption Protein sequence largely determines structure, and structure largely determines function.
    Section 1, first paragraph; the entire three-category organization depends on this causal chain.
  • domain assumption Protein design can be treated as the inverse of structure and function prediction.
    Sections 1, 2.2.2, and 5.2; the survey states this as its organizing intuition, citing the common paradigm [23] and design works that reverse predictors.
  • domain assumption The described primary papers are faithfully represented by the survey.
    Because this is a review, all downstream conclusions inherit this assumption; two concrete attribution errors show it is not fully satisfied.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey of Deep Learning Methods in Protein Bioinformatics and its Impact on Protein Design." pith.science (2026). https://pith.science/paper/ZMVMRLNG

@misc{pith2026250101477,
  author       = {Pith},
  title        = {Pith review of: A Survey of Deep Learning Methods in Protein Bioinformatics and its Impact on Protein Design},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZMVMRLNG}},
  note         = {Machine review of arXiv:2501.01477}
}
read the original abstract

Proteins are sequences of amino acids that serve as the basic building blocks of living organisms. Despite rapidly growing databases documenting structural and functional information for various protein sequences, our understanding of proteins remains limited because of the large possible sequence space and the complex inter- and intra-molecular forces. Deep learning, which is characterized by its ability to learn relevant features directly from large datasets, has demonstrated remarkable performance in fields such as computer vision and natural language processing. It has also been increasingly applied in recent years to the data-rich domain of protein sequences with great success, most notably with Alphafold2's breakout performance in the protein structure prediction. The performance improvements achieved by deep learning unlocks new possibilities in the field of protein bioinformatics, including protein design, one of the most difficult but useful tasks. In this paper, we broadly categorize problems in protein bioinformatics into three main categories: 1) structural prediction, 2) functional prediction, and 3) protein design, and review the progress achieved from using deep learning methodologies in each of them. We expand on the main challenges of the protein design problem and highlight how advances in structural and functional prediction have directly contributed to design tasks. Finally, we conclude by identifying important topics and future research directions.

Figures

Figures reproduced from arXiv: 2501.01477 by the authors.

Figure 1
Figure 1. (a) Simplified representation of a MLP with n layers. Input and intermediate features are [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. A simplified overview of the main classes of GNNs. GNNs all involve a convolutional [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. An Euclidian patch covers the region that is enclosed by a circular projection from a 2D [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Peptide bonds are formed between the Carbon and Nitrogen atoms in two amino-acids, [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Simplified examples different kinds of protein binding. Surface-surface binding is the most [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Visual relationship between the three main categories of problems in protein bioinformatics. [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Network diagram of Alphafold1, taken from [ [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: Alphafold2 uses a novel transformer module, the Evoformer, that allows the network to [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]
Figure 9
Figure 9. Figure 9: The flow diagram shows an example of how GNNs can be used for PPI or protein-ligand [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]
Figure 10
Figure 10. Figure 10: Geodesic convolution applies filters on geodesic patches of a target surface. The features [PITH_FULL_IMAGE:figures/full_fig_p017_10.png]
Figure 11
Figure 11. Figure 11: The network diagram above shows an example where a structure prediction model can be [PITH_FULL_IMAGE:figures/full_fig_p020_11.png]
Figure 12
Figure 12. Figure 12: A trained network can be used through gradient ascent to perform back-propagation [PITH_FULL_IMAGE:figures/full_fig_p022_12.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

126 extracted references · 57 canonical work pages

  1. [4]

    Gapped blast and psi-blast: a new generation of protein database search programs

    Stephen F Altschul, Thomas L Madden, Alejandro A Schäffer, Jinghui Zhang, Zheng Zhang, Webb Miller, and David J Lipman. Gapped blast and psi-blast: a new generation of protein database search programs. Nucleic acids research, 25(17):3389–3402, 1997

  2. [68]

    Protein distance constraints predicted by neural networks and probability density functions

    Ole Lund, Kenneth Frimand, Jan Gorodkin, Henrik Bohr, Jakob Bohr, Jan Hansen, and Søren Brunak. Protein distance constraints predicted by neural networks and probability density functions. Protein Engineering, 10(11):1241–1248, 1997

  3. [112]

    Scannet: An interpretable geometric deep learning model for structure-based protein binding site prediction

    Jérôme Tubiana, Dina Schneidman-Duhovny, and Haim J Wolfson. Scannet: An interpretable geometric deep learning model for structure-based protein binding site prediction. bioRxiv, 2021

  4. [8]

    De novo protein design by deep network hallucination

    Ivan Anishchenko, Tamuka Martin Chidyausiku, Sergey Ovchinnikov, Samuel J Pellock, and David Baker. De novo protein design by deep network hallucination. bioRxiv, 2020

  5. [81]

    Protein sequence design by explicit energy landscape optimization

    Christoffer Norn, Basile IM Wicky, David Juergens, Sirui Liu, David Kim, Brian Koepnick, Ivan Anishchenko, David Baker, and Sergey Ovchinnikov. Protein sequence design by explicit energy landscape optimization. bioRxiv, 2020

  6. [110]

    Design of proteins presenting discontinuous functional sites using deep learning

    Doug Tischer, Sidney Lisanza, Jue Wang, Runze Dong, Ivan Anishchenko, Lukas F Milles, Sergey Ovchinnikov, and David Baker. Design of proteins presenting discontinuous functional sites using deep learning. bioRxiv, 2020. 29

  7. [38]

    Feedback gan for dna optimizes protein functions

    Anvita Gupta and James Zou. Feedback gan for dna optimizes protein functions. Nature Machine Intelligence, 1(2):105–111, 2019

  8. [50]

    De novo protein design for novel folds using guided conditional wasserstein generative adversarial networks

    Mostafa Karimi, Shaowen Zhu, Yue Cao, and Yang Shen. De novo protein design for novel folds using guided conditional wasserstein generative adversarial networks. Journal of Chemical Information and Modeling, 60(12):5667–5681, 2020

Show all 126 references
  1. [1]

    Protein function

    Bruce Alberts. Protein function. Molecular Biology of the Cell. 4th edition., Jan 1970

  2. [2]

    Unified rational protein engineering with sequence-based deep representation learning

    Ethan C Alley, Grigory Khimulya, Surojit Biswas, Mohammed AlQuraishi, and George M Church. Unified rational protein engineering with sequence-based deep representation learning. Nature methods, 16(12):1315–1322, 2019

  3. [3]

    Deeploc: prediction of protein subcellular localization using deep learning

    José Juan Almagro Armenteros, Casper Kaae Sønderby, Søren Kaae Sønderby, Henrik Nielsen, and Ole Winther. Deeploc: prediction of protein subcellular localization using deep learning. Bioinformatics, 33(21):3387–3395, 2017

  4. [5]

    Spatial uncertainty sampling for end-to-end control

    Alexander Amini, Ava Soleimany, Sertac Karaman, and Daniela Rus. Spatial uncertainty sampling for end-to-end control. arXiv preprint arXiv:1805.04829, 2018

  5. [6]

    Fully differentiable full-atom protein backbone generation

    Namrata Anand, Raphael Eguchi, and Po-Ssu Huang. Fully differentiable full-atom protein backbone generation. 2019

  6. [7]

    Principles that govern the folding of protein chains

    Christian B Anfinsen. Principles that govern the folding of protein chains. Science, 181(4096):223–230, 1973

  7. [9]

    Gene ontology: tool for the unification of biology

    Michael Ashburner, Catherine A Ball, Judith A Blake, David Botstein, Heather Butler, J Michael Cherry, Allan P Davis, Kara Dolinski, Selina S Dwight, Janan T Eppig, et al. Gene ontology: tool for the unification of biology. Nature genetics, 25(1):25–29, 2000

  8. [10]

    Bind—the biomolecular interaction network database

    Gary D Bader, Ian Donaldson, Cheryl Wolting, BF Francis Ouellette, Tony Pawson, and Christopher WV Hogue. Bind—the biomolecular interaction network database. Nucleic acids research, 29(1):242–245, 2001

  9. [11]

    Machine learning-guided channelrhodopsin engineering enables minimally invasive optogenetics

    Claire N Bedbrook, Kevin K Yang, J Elliott Robinson, Elisha D Mackey, Viviana Gradinaru, and Frances H Arnold. Machine learning-guided channelrhodopsin engineering enables minimally invasive optogenetics. Nature methods, 16(11):1176–1184, 2019

  10. [12]

    Deep learning of representations for unsupervised and transfer learning

    Yoshua Bengio. Deep learning of representations for unsupervised and transfer learning. In Proceedings of ICML workshop on unsupervised and transfer learning, pages 17–36. JMLR Workshop and Conference Proceedings, 2012

  11. [13]

    The protein data bank

    Helen M Berman, Tammy Battistuz, Talapady N Bhat, Wolfgang F Bluhm, Philip E Bourne, Kyle Burkhardt, Zukang Feng, Gary L Gilliland, Lisa Iype, Shri Jain, et al. The protein data bank. Acta Crystallographica Section D: Biological Crystallography, 58(6):899–907, 2002

  12. [14]

    Low-n protein engineering with data-efficient deep learning

    Surojit Biswas, Grigory Khimulya, Ethan C Alley, Kevin M Esvelt, and George M Church. Low-n protein engineering with data-efficient deep learning. Nature Methods, 18(4):389–396, 2021

  13. [15]

    Learning shape corre- spondence with anisotropic convolutional neural networks

    Davide Boscaini, Jonathan Masci, Emanuele Rodolà, and Michael Bronstein. Learning shape corre- spondence with anisotropic convolutional neural networks. Advances in neural information processing systems, 29, 2016

  14. [16]

    Geometric deep learning: Grids, groups, graphs, geodesics, and gauges

    Michael M Bronstein, Joan Bruna, Taco Cohen, and Petar Veliˇckovi´c. Geometric deep learning: Grids, groups, graphs, geodesics, and gauges. arXiv preprint arXiv:2104.13478, 2021

  15. [17]

    Scalable web services for the psipred protein analysis workbench

    Daniel W A Buchan, Federico Minneci, Tim CO Nugent, Kevin Bryson, and David T Jones. Scalable web services for the psipred protein analysis workbench. Nucleic acids research, 41(W1):W349–W357, 2013

  16. [18]

    Tale: Transformer-based protein function annotation with joint sequence–label embedding

    Yue Cao and Yang Shen. Tale: Transformer-based protein function annotation with joint sequence–label embedding. Bioinformatics, 37(18):2825–2833, 2021

  17. [19]

    Scope: classification of large macromolec- ular structures in the structural classification of proteins—extended database

    John-Marc Chandonia, Naomi K Fox, and Steven E Brenner. Scope: classification of large macromolec- ular structures in the structural classification of proteins—extended database. Nucleic acids research, 47(D1):D475–D481, 2019

  18. [20]

    Convolu- tional embedding of attributed molecular graphs for physical property prediction

    Connor W Coley, Regina Barzilay, William H Green, Tommi S Jaakkola, and Klavs F Jensen. Convolu- tional embedding of attributed molecular graphs for physical property prediction. Journal of chemical information and modeling, 57(8):1757–1772, 2017

  19. [21]

    Uniprot: a worldwide hub of protein knowledge

    UniProt Consortium. Uniprot: a worldwide hub of protein knowledge. Nucleic acids research , 47(D1):D506–D515, 2019

  20. [22]

    Pepcvae: Semi-supervised targeted design of antimicrobial peptide sequences

    Payel Das, Kahini Wadhawan, Oscar Chang, Tom Sercu, Cicero Dos Santos, Matthew Riemer, Vijil Chenthamarakshan, Inkit Padhi, and Aleksandra Mojsilovic. Pepcvae: Semi-supervised targeted design of antimicrobial peptide sequences. arXiv preprint arXiv:1810.07743, 2018. 25

  21. [23]

    Protein actions: Principles and modeling

    Ken Dill, Robert L Jernigan, and Ivet Bahar. Protein actions: Principles and modeling. Garland Science, 2017

  22. [24]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv...

  23. [25]

    On the conservative nature of intragenic recombination

    D Allan Drummond, Jonathan J Silberg, Michelle M Meyer, Claus O Wilke, and Frances H Arnold. On the conservative nature of intragenic recombination. Proceedings of the National Academy of Sciences, 102(15):5380–5385, 2005

  24. [26]

    Efficient unbound docking of rigid molecules

    Dina Duhovny, Ruth Nussinov, and Haim J Wolfson. Efficient unbound docking of rigid molecules. In International workshop on algorithms in bioinformatics, pages 185–200. Springer, 2002

  25. [27]

    Convolutional networks on graphs for learning molecular fingerprints

    David Duvenaud, Dougal Maclaurin, Jorge Aguilera-Iparraguirre, Rafael Gómez-Bombarelli, Timothy Hirzel, Alán Aspuru-Guzik, and Ryan P Adams. Convolutional networks on graphs for learning molecular fingerprints. arXiv preprint arXiv:1509.09292, 2015

  26. [28]

    The case for defined protein folding pathways

    S Walter Englander and Leland Mayne. The case for defined protein folding pathways. Proceedings of the National Academy of Sciences, 114(31):8253–8258, 2017

  27. [29]

    A neural network based predictor of residue contacts in proteins

    Piero Fariselli and R Casadio. A neural network based predictor of residue contacts in proteins. Protein engineering, 12(1):15–21, 1999

  28. [30]

    Protein interface prediction using graph convolutional networks

    Alex M Fout. Protein interface prediction using graph convolutional networks. PhD thesis, Colorado State University, 2017

  29. [31]

    Deciphering interaction fingerprints from protein molecular surfaces using geometric deep learning

    Pablo Gainza, Freyr Sverrisson, Frederico Monti, Emanuele Rodola, D Boscaini, MM Bronstein, and BE Correia. Deciphering interaction fingerprints from protein molecular surfaces using geometric deep learning. Nature Methods, 17(2):184–192, 2020

  30. [32]

    Correlated mutations and residue contacts in proteins

    Ulrike Göbel, Chris Sander, Reinhard Schneider, and Alfonso Valencia. Correlated mutations and residue contacts in proteins. Proteins: Structure, Function, and Bioinformatics, 18(4):309–317, 1994

  31. [33]

    Metagenomics and the protein universe

    Adam Godzik. Metagenomics and the protein universe. Current opinion in structural biology, 21(3):398– 403, 2011

  32. [34]

    Deep learning

    Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep learning. MIT press, 2016

  33. [35]

    Transformer neural network for protein-specific de novo drug generation as a machine translation problem

    Daria Grechishnikova. Transformer neural network for protein-specific de novo drug generation as a machine translation problem. Scientific reports, 11(1):1–13, 2021

  34. [36]

    Design of metalloproteins and novel protein folds using variational autoencoders

    Joe G Greener, Lewis Moffat, and David T Jones. Design of metalloproteins and novel protein folds using variational autoencoders. Scientific reports, 8(1):1–12, 2018

  35. [37]

    Designing anticancer peptides by constructive machine learning

    Francesca Grisoni, Claudia S Neuhaus, Gisela Gabernet, Alex T Muller, Jan A Hiss, and Gisbert Schneider. Designing anticancer peptides by constructive machine learning. ChemMedChem, 13(13):1300–1302, 2018

  36. [39]

    Protein contact prediction using patterns of correlation

    Nicholas Hamilton, Kevin Burrage, Mark A Ragan, and Thomas Huber. Protein contact prediction using patterns of correlation. Proteins: Structure, Function, and Bioinformatics, 56(4):679–684, 2004

  37. [40]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016

  38. [41]

    Modeling aspects of the language of life through transfer-learning protein sequences

    Michael Heinzinger, Ahmed Elnaggar, Yu Wang, Christian Dallago, Dmitrii Nechaev, Florian Matthes, and Burkhard Rost. Modeling aspects of the language of life through transfer-learning protein sequences. BMC bioinformatics, 20(1):1–17, 2019

  39. [42]

    Molecular dynamics simulation for all.Neuron, 99(6):1129–1143, 2018

    Scott A Hollingsworth and Ron O Dror. Molecular dynamics simulation for all.Neuron, 99(6):1129–1143, 2018

  40. [43]

    Mutation effects predicted from sequence co-variation

    Thomas A Hopf, John B Ingraham, Frank J Poelwijk, Charlotta PI Schärfe, Michael Springer, Chris Sander, and Debora S Marks. Mutation effects predicted from sequence co-variation. Nature biotechnology, 35(2):128–135, 2017

  41. [44]

    Multilayer feedforward networks are universal approximators

    Kurt Hornik, Maxwell Stinchcombe, and Halbert White. Multilayer feedforward networks are universal approximators. Neural networks, 2(5):359–366, 1989

  42. [45]

    Densely connected convo- lutional networks

    Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convo- lutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700–4708, 2017

  43. [46]

    The coming of age of de novo protein design

    Po-Ssu Huang, Scott E Boyken, and David Baker. The coming of age of de novo protein design. Nature, 537(7620):320–327, 2016. 26

  44. [47]

    Generative models for graph-based protein design

    John Ingraham, Vikas K Garg, Regina Barzilay, and Tommi Jaakkola. Generative models for graph-based protein design. 2019

  45. [48]

    Psicov: precise structural contact prediction using sparse inverse covariance estimation on large multiple sequence alignments

    David T Jones, Daniel WA Buchan, Domenico Cozzetto, and Massimiliano Pontil. Psicov: precise structural contact prediction using sparse inverse covariance estimation on large multiple sequence alignments. Bioinformatics, 28(2):184–190, 2012

  46. [49]

    Highly accurate protein structure prediction with alphafold

    John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021

  47. [51]

    Molecular graph convolutions: moving beyond fingerprints

    Steven Kearnes, Kevin McCloskey, Marc Berndl, Vijay Pande, and Patrick Riley. Molecular graph convolutions: moving beyond fingerprints. Journal of computer-aided molecular design, 30(8):595–608, 2016

  48. [52]

    Reformer: The efficient transformer

    Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. Reformer: The efficient transformer. arXiv preprint arXiv:2001.04451, 2020

  49. [53]

    Netsurfp-2.0: Improved prediction of protein structural features by integrated deep learning

    Michael Schantz Klausen, Martin Closter Jespersen, Henrik Nielsen, Kamilla Kjaergaard Jensen, Vanessa Isabell Jurtz, Casper Kaae Soenderby, Morten Otto Alexander Sommer, Ole Winther, Morten Nielsen, Bent Petersen, et al. Netsurfp-2.0: Improved prediction of protein structural ...

  50. [54]

    Lessons learned in empirical scoring with smina from the csar 2011 benchmarking exercise

    David Ryan Koes, Matthew P Baumgartner, and Carlos J Camacho. Lessons learned in empirical scoring with smina from the csar 2011 benchmarking exercise. Journal of chemical information and modeling, 53(8):1893–1904, 2013

  51. [55]

    Probis-charmming: web interface for prediction and optimization of ligands in protein binding sites, 2015

    Janez Konc, Benjamin T Miller, Tanja SStular, Samo Lesnik, H Lee Woodcock, Bernard R Brooks, and Dusanka Janezic. Probis-charmming: web interface for prediction and optimization of ligands in protein binding sites, 2015

  52. [56]

    Improved prediction of protein side-chain conformations with scwrl4

    Georgii G Krivov, Maxim V Shapovalov, and Roland L Dunbrack Jr. Improved prediction of protein side-chain conformations with scwrl4. Proteins: Structure, Function, and Bioinformatics, 77(4):778–795, 2009

  53. [57]

    Native protein sequences are close to optimal for their structures

    Brian Kuhlman and David Baker. Native protein sequences are close to optimal for their structures. Proceedings of the National Academy of Sciences, 97(19):10383–10388, 2000

  54. [58]

    Deepgoplus: improved protein function prediction from sequence

    Maxat Kulmanov and Robert Hoehndorf. Deepgoplus: improved protein function prediction from sequence. Bioinformatics, 36(2):422–429, 2020

  55. [59]

    Deepgo: predicting protein functions from sequence and interactions using a deep ontology-aware classifier

    Maxat Kulmanov, Mohammed Asif Khan, and Robert Hoehndorf. Deepgo: predicting protein functions from sequence and interactions using a deep ontology-aware classifier. Bioinformatics, 34(4):660–668, 2018

  56. [60]

    Pep-fold3: faster de novo structure prediction for linear peptides in solution and in complex

    Alexis Lamiable, Pierre Thévenet, Julien Rey, Marek Vavrusa, Philippe Derreumaux, and Pierre Tufféry. Pep-fold3: faster de novo structure prediction for linear peptides in solution and in complex. Nucleic acids research, 44(W1):W449–W454, 2016

  57. [61]

    Open-source cheminformatics software

    greg Landrum. Open-source cheminformatics software

  58. [62]

    Self-attention graph pooling

    Junhyun Lee, Inyeop Lee, and Jaewoo Kang. Self-attention graph pooling. In International Conference on Machine Learning, pages 3734–3743. PMLR, 2019

  59. [63]

    Are there pathways for protein folding? Journal de chimie physique, 65:44–45, 1968

    Cyrus Levinthal. Are there pathways for protein folding? Journal de chimie physique, 65:44–45, 1968

  60. [64]

    Protein loop modeling using deep generative adversarial network

    Zhaoyu Li, Son P Nguyen, Dong Xu, and Yi Shang. Protein loop modeling using deep generative adversarial network. In 2017 IEEE 29th International Conference on Tools with Artificial Intelligence (ICTAI), pages 1085–1091. IEEE, 2017

  61. [65]

    Direct prediction of profiles of sequences compatible with a protein structure by neural networks with fragment-based local and energy-based nonlocal profiles

    Zhixiu Li, Yuedong Yang, Eshel Faraggi, Jian Zhan, and Yaoqi Zhou. Direct prediction of profiles of sequences compatible with a protein structure by neural networks with fragment-based local and energy-based nonlocal profiles. Proteins: Structure, Function, and Bioinformatics,...

  62. [66]

    Swin transformer: Hierarchical vision transformer using shifted windows

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030, 2021

  63. [67]

    Object recognition from local scale-invariant features

    David G Lowe. Object recognition from local scale-invariant features. In Proceedings of the seventh IEEE international conference on computer vision, volume 2, pages 1150–1157. Ieee, 1999. 27

  64. [69]

    Deepphos: prediction of protein phosphorylation sites with deep learning

    Fenglin Luo, Minghui Wang, Yu Liu, Xing-Ming Zhao, and Ao Li. Deepphos: prediction of protein phosphorylation sites with deep learning. Bioinformatics, 35(16):2766–2773, 2019

  65. [70]

    Deep lagrangian networks: Using physics as model prior for deep learning

    Michael Lutter, Christian Ritter, and Jan Peters. Deep lagrangian networks: Using physics as model prior for deep learning. arXiv preprint arXiv:1907.04490, 2019

  66. [71]

    Protein 3d structure computed from evolutionary sequence variation

    Debora S Marks, Lucy J Colwell, Robert Sheridan, Thomas A Hopf, Andrea Pagnani, Riccardo Zecchina, and Chris Sander. Protein 3d structure computed from evolutionary sequence variation. PloS one, 6(12):e28766, 2011

  67. [72]

    Ab initio molecular dynamics: basic theory and advanced methods

    Dominik Marx and Jürg Hutter. Ab initio molecular dynamics: basic theory and advanced methods . Cambridge University Press, 2009

  68. [73]

    Geodesic convolutional neural networks on riemannian manifolds

    Jonathan Masci, Davide Boscaini, Michael Bronstein, and Pierre Vandergheynst. Geodesic convolutional neural networks on riemannian manifolds. In Proceedings of the IEEE international conference on computer vision workshops, pages 37–45, 2015

  69. [74]

    Mutagenesis of a buried polar interaction in an sh3 domain: sequence conservation provides the best prediction of stability effects.Biochemistry, 37(46):16172–16182, 1998

    Karen L Maxwell and Alan R Davidson. Mutagenesis of a buried polar interaction in an sh3 domain: sequence conservation provides the best prediction of stability effects.Biochemistry, 37(46):16172–16182, 1998

  70. [75]

    Molecular evolution of broadly neutralizing llama antibodies to the cd4-binding site of hiv-1

    Laura E McCoy, Lucy Rutten, Dan Frampton, Ian Anderson, Luke Granger, Rachael Bashford-Rogers, Gillian Dekkers, Nika M Strokappe, Michael S Seaman, Willie Koh, et al. Molecular evolution of broadly neutralizing llama antibodies to the cd4-binding site of hiv-1. PLoS pathogens,...

  71. [76]

    Geometric deep learning on graphs and manifolds using mixture model cnns

    Federico Monti, Davide Boscaini, Jonathan Masci, Emanuele Rodola, Jan Svoboda, and Michael M Bronstein. Geometric deep learning on graphs and manifolds using mixture model cnns. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5115–5124, 2017

  72. [77]

    Recurrent neural network model for constructive peptide design

    Alex T Muller, Jan A Hiss, and Gisbert Schneider. Recurrent neural network model for constructive peptide design. Journal of chemical information and modeling, 58(2):472–479, 2018

  73. [78]

    Transform- ing the language of life: transformer neural networks for protein prediction tasks

    Ananthan Nambiar, Maeve Heflin, Simon Liu, Sergei Maslov, Mark Hopkins, and Anna Ritz. Transform- ing the language of life: transformer neural networks for protein prediction tasks. In Proceedings of the 11th ACM International Conference on Bioinformatics, Computational Biolog...

  74. [79]

    Learning convolutional neural networks for graphs

    Mathias Niepert, Mohamed Ahmed, and Konstantin Kutzkov. Learning convolutional neural networks for graphs. In International conference on machine learning, pages 2014–2023. PMLR, 2016

  75. [80]

    Machine learning for molecular simulation

    Frank Noé, Alexandre Tkatchenko, Klaus-Robert Müller, and Cecilia Clementi. Machine learning for molecular simulation. Annual review of physical chemistry, 71:361–390, 2020

  76. [82]

    Spin2: Predicting sequence profiles from protein structures using deep neural networks

    James O’Connell, Zhixiu Li, Jack Hanson, Rhys Heffernan, James Lyons, Kuldip Paliwal, Abdollah Dehzangi, Yuedong Yang, and Yaoqi Zhou. Spin2: Predicting sequence profiles from protein structures using deep neural networks. Proteins: Structure, Function, and Bioinformatics, 86(...

  77. [83]

    ProFET: Feature engineering captures high-level protein functions

    Dan Ofer and Michal Linial. ProFET: Feature engineering captures high-level protein functions. Bioin- formatics, 31(21):3429–3436, 06 2015

  78. [84]

    Molecular de-novo design through deep reinforcement learning

    Marcus Olivecrona, Thomas Blaschke, Ola Engkvist, and Hongming Chen. Molecular de-novo design through deep reinforcement learning. Journal of cheminformatics, 9(1):1–14, 2017

  79. [85]

    Accelerating protein docking in zdock using an advanced 3d convolution library

    Brian G Pierce, Yuichiro Hourai, and Zhiping Weng. Accelerating protein docking in zdock using an advanced 3d convolution library. PloS one, 6(9):e24657, 2011

  80. [86]

    Learning context-aware structural representations to predict antigen and antibody binding interfaces

    Srivamshi Pittala and Chris Bailey-Kellogg. Learning context-aware structural representations to predict antigen and antibody binding interfaces. Bioinformatics, 36(13):3996–4003, 2020

  81. [87]

    Deep reinforcement learning for de novo drug design

    Mariya Popova, Olexandr Isayev, and Alexander Tropsha. Deep reinforcement learning for de novo drug design. Science advances, 4(7):eaap7885, 2018

  82. [88]

    Reinforced adversarial neural computer for de novo molecular design

    Evgeny Putin, Arip Asadulaev, Yan Ivanenkov, Vladimir Aladinskiy, Benjamin Sanchez-Lengeling, Alán Aspuru-Guzik, and Alex Zhavoronkov. Reinforced adversarial neural computer for de novo molecular design. Journal of chemical information and modeling, 58(6):1194–1204, 2018

  83. [89]

    A large-scale evaluation of computational protein function prediction

    Predrag Radivojac, Wyatt T Clark, Tal Ronnen Oron, Alexandra M Schnoes, Tobias Wittkop, Artem Sokolov, Kiley Graim, Christopher Funk, Karin Verspoor, Asa Ben-Hur, et al. A large-scale evaluation of computational protein function prediction. Nature methods, 10(3):221–227, 2013. 28

  84. [90]

    Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations

    Maziar Raissi, Paris Perdikaris, and George E Karniadakis. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational Physics, 378:686–707, 2019

  85. [91]

    Deep robust framework for protein function prediction using variable-length protein sequences

    Ashish Ranjan, Md Shah Fahad, David Fernández-Baca, Akshay Deepak, and Sudhakar Tripathi. Deep robust framework for protein function prediction using variable-length protein sequences. IEEE/ACM transactions on computational biology and bioinformatics, 17(5):1648–1659, 2019

  86. [92]

    Msa transformer

    Roshan Rao, Jason Liu, Robert Verkuil, Joshua Meier, John F Canny, Pieter Abbeel, Tom Sercu, and Alexander Rives. Msa transformer. bioRxiv, 2021

  87. [93]

    Expanding functional protein sequence spaces using generative adversarial networks

    Donatas Repecka, Vykintas Jauniskis, Laurynas Karpus, Elzbieta Rembeza, Irmantas Rokaitis, Jan Zrimec, Simona Poviloniene, Audrius Laurynenas, Sandra Viknander, Wissam Abuajwa, et al. Expanding functional protein sequence spaces using generative adversarial networks. Nature Ma...

  88. [94]

    Accelerating protein design using autoregressive generative models

    Adam Riesselman, Jung-Eun Shin, Aaron Kollasch, Conor McMahon, Elana Simon, Chris Sander, Aashish Manglik, Andrew Kruse, and Debora Marks. Accelerating protein design using autoregressive generative models. BioRxiv, page 757252, 2019

  89. [95]

    Deep generative models of genetic variation capture the effects of mutations

    Adam J Riesselman, John B Ingraham, and Debora S Marks. Deep generative models of genetic variation capture the effects of mutations. Nature methods, 15(10):816–822, 2018

  90. [96]

    Kripo–a structure-based pharmacophores approach explains polypharmacological effects

    Tina Ritschel, Tom JJ Schirris, and Frans GM Russel. Kripo–a structure-based pharmacophores approach explains polypharmacological effects. Journal of cheminformatics, 6(1):1–1, 2014

  91. [97]

    Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences

    Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C Lawrence Zitnick, Jerry Ma, et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National ...

  92. [98]

    Extended-connectivity fingerprints

    David Rogers and Mathew Hahn. Extended-connectivity fingerprints. Journal of chemical information and modeling, 50(5):742–754, 2010

  93. [99]

    Protein structure prediction using rosetta

    Carol A Rohl, Charlie EM Strauss, Kira MS Misura, and David Baker. Protein structure prediction using rosetta. Methods in enzymology, 383:66–93, 2004

  94. [100]

    Exploring protein fitness landscapes by directed evolution

    Philip A Romero and Frances H Arnold. Exploring protein fitness landscapes by directed evolution. Nature reviews Molecular cell biology, 10(12):866–876, 2009

  95. [101]

    Navigating the protein fitness landscape with gaussian processes

    Philip A Romero, Andreas Krause, and Frances H Arnold. Navigating the protein fitness landscape with gaussian processes. Proceedings of the National Academy of Sciences, 110(3):E193–E201, 2013

  96. [102]

    Geo- metric potentials from deep learning improve prediction of cdr h3 loop structures

    Jeffrey A Ruffolo, Carlos Guerra, Sai Pooja Mahajan, Jeremias Sulam, and Jeffrey J Gray. Geo- metric potentials from deep learning improve prediction of cdr h3 loop structures. Bioinformatics, 36(Supplement_1):i268–i275, 2020

  97. [103]

    Optimizing distributions over molecular space

    Benjamin Sanchez-Lengeling, Carlos Outeiral, Gabriel L Guimaraes, and Alan Aspuru-Guzik. Optimizing distributions over molecular space. an objective-reinforced generative adversarial network for inverse- design chemistry (organic). 2017

  98. [104]

    Improved protein structure prediction using potentials from deep learning

    Andrew W Senior, Richard Evans, John Jumper, James Kirkpatrick, Laurent Sifre, Tim Green, Chongli Qin, Augustin Žídek, Alexander WR Nelson, Alex Bridgland, et al. Improved protein structure prediction using potentials from deep learning. Nature, 577(7792):706–710, 2020

  99. [105]

    Recognition of functional sites in protein structures

    Alexandra Shulman-Peleg, Ruth Nussinov, and Haim J Wolfson. Recognition of functional sites in protein structures. Journal of molecular biology, 339(3):607–633, 2004

  100. [106]

    Very deep convolutional networks for large-scale image recognition

    Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014

  101. [107]

    Fast and flexible design of novel proteins using graph neural networks

    Alexey Strokach, David Becerra, Carles Corbi-Verge, Albert Perez-Riba, and Philip M Kim. Fast and flexible design of novel proteins using graph neural networks. BioRxiv, page 868935, 2020

  102. [108]

    Uniref clusters: a comprehensive and scalable alternative for improving sequence similarity searches

    Baris E Suzek, Yuqi Wang, Hongzhan Huang, Peter B McGarvey, Cathy H Wu, and UniProt Consortium. Uniref clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics, 31(6):926–932, 2015

  103. [109]

    The string database in 2017: quality-controlled protein–protein association networks, made broadly accessible

    Damian Szklarczyk, John H Morris, Helen Cook, Michael Kuhn, Stefan Wyder, Milan Simonovic, Alberto Santos, Nadezhda T Doncheva, Alexander Roth, Peer Bork, et al. The string database in 2017: quality-controlled protein–protein association networks, made broadly accessible. Nucl...

  104. [111]

    Graph convolutional neural networks for predicting drug-target interac- tions

    Wen Torng and Russ B Altman. Graph convolutional neural networks for predicting drug-target interac- tions. Journal of chemical information and modeling, 59(10):4131–4149, 2019

  105. [113]

    A survey on semi-supervised learning

    Jesper E Van Engelen and Holger H Hoos. A survey on semi-supervised learning. Machine Learning, 109(2):373–440, 2020

  106. [114]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017

  107. [115]

    Bertology meets biology: Interpreting attention in protein language models

    Jesse Vig, Ali Madani, Lav R Varshney, Caiming Xiong, Richard Socher, and Nazneen Fatema Ra- jani. Bertology meets biology: Interpreting attention in protein language models. arXiv preprint arXiv:2006.15222, 2020

  108. [116]

    Camp: Collection of sequences and structures of antimicrobial peptides

    Faiza Hanif Waghu, Lijin Gopi, Ram Shankar Barai, Pranay Ramteke, Bilal Nizami, and Susan Idicula- Thomas. Camp: Collection of sequences and structures of antimicrobial peptides. Nucleic acids research, 42(D1):D1154–D1158, 2014

  109. [117]

    Accurate de novo prediction of protein contact map by ultra-deep learning model

    Sheng Wang, Siqi Sun, Zhen Li, Renyu Zhang, and Jinbo Xu. Accurate de novo prediction of protein contact map by ultra-deep learning model. PLoS computational biology, 13(1):e1005324, 2017

  110. [118]

    Computational prediction of protein interfaces: A review of data driven methods

    Li C Xue, Drena Dobbs, Alexandre MJJ Bonvin, and Vasant Honavar. Computational prediction of protein interfaces: A review of data driven methods. FEBS letters, 589(23):3516–3526, 2015

  111. [119]

    Improved protein structure prediction using predicted interresidue orientations

    Jianyi Yang, Ivan Anishchenko, Hahnbeom Park, Zhenling Peng, Sergey Ovchinnikov, and David Baker. Improved protein structure prediction using predicted interresidue orientations. Proceedings of the National Academy of Sciences, 117(3):1496–1503, 2020

  112. [120]

    Learned protein embeddings for machine learning

    Kevin K Yang, Zachary Wu, Claire N Bedbrook, and Frances H Arnold. Learned protein embeddings for machine learning. Bioinformatics, 34(15):2642–2648, 2018

  113. [121]

    Fast screening of protein surfaces using geometric invariant fingerprints

    Shuangye Yin, Elizabeth A Proctor, Alexey A Lugovskoy, and Nikolay V Dokholyan. Fast screening of protein surfaces using geometric invariant fingerprints. Proceedings of the National Academy of Sciences, 106(39):16622–16626, 2009

  114. [122]

    Hierarchical graph representation learning with differentiable pooling

    Rex Ying, Jiaxuan You, Christopher Morris, Xiang Ren, William L Hamilton, and Jure Leskovec. Hierarchical graph representation learning with differentiable pooling. arXiv preprint arXiv:1806.08804, 2018

  115. [123]

    NetGO: improving large-scale protein function prediction with massive network information

    Ronghui You, Shuwei Yao, Yi Xiong, Xiaodi Huang, Fengzhu Sun, Hiroshi Mamitsuka, and Shanfeng Zhu. NetGO: improving large-scale protein function prediction with massive network information. Nucleic Acids Research, 47(W1):W379–W387, 05 2019

  116. [124]

    Golabeler: improving sequence-based large-scale protein function prediction by learning to rank

    Ronghui You, Zihan Zhang, Yi Xiong, Fengzhu Sun, Hiroshi Mamitsuka, and Shanfeng Zhu. Golabeler: improving sequence-based large-scale protein function prediction by learning to rank. Bioinformatics, 34(14):2465–2473, 2018

  117. [125]

    Template-based and free modeling of i-tasser and quark pipelines using predicted contact maps in casp12

    Chengxin Zhang, SM Mortuza, Baoji He, Yanting Wang, and Yang Zhang. Template-based and free modeling of i-tasser and quark pipelines using predicted contact maps in casp12. Proteins: Structure, Function, and Bioinformatics, 86:136–151, 2018

  118. [126]

    The cafa challenge reports improved protein function prediction and new functional annotations for hundreds of genes through experimental screens

    Naihui Zhou, Yuxiang Jiang, Timothy R Bergquist, Alexandra J Lee, Balint Z Kacsoh, Alex W Crocker, Kimberley A Lewis, George Georghiou, Huy N Nguyen, Md Nafiz Hamid, et al. The cafa challenge reports improved protein function prediction and new functional annotations for hundr...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.