Pith. sign in

REVIEW 2 major objections 5 minor 66 references

Foundation Models for AI-Enabled Biological Design

T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This survey proposes a taxonomy that organizes foundation models for biological sequence design by architecture family, controllability strategy, and multimodal integration.

desk verdict A useful but structurally imperfect survey: the per-model table is the real value, but the taxonomy's 'diffusion as architecture' axis needs rework before this is publishable. read the letter →

arxiv 2505.11610 v1 pith:IFFF5PPW submitted 2025-05-16 cs.AI cs.LGq-bio.BMq-bio.GN

classification cs.AIcs.LGq-bio.BMq-bio.GN
keywords foundationmodelsbiologicaldesignproteinlanguagesmallmoleculegenomicsequencestatespacediffusioncontrollablegeneration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey organizes recent foundation models for biological design into a taxonomy with three axes: architecture family (Transformer, state-space, and diffusion), controllability strategy (fine-tuning, conditional generation, reinforcement learning, custom objectives, post-generation filtering, and sampling), and multimodal integration. Its aim is to give researchers a working map of sequence-based generative models for proteins, small molecules, and genomes, and to highlight where the field lacks consensus. A sympathetic reader would take away that most current models borrow NLP-style architectures and objectives, and that the central engineering problem is controlling generation toward desired biological properties.

What carries the argument

The central organizing device is a two-part taxonomy: a tree whose leaf nodes carry citation clusters for architectures, controllability, and multimodality, paired with a table that lists each representative model's domain, architecture, design goal, and generation method. The underlying conceptual move is the analogy between natural language and biological sequences, which lets decoder-only Transformers, SSMs, and diffusion models be imported from NLP and adapted to proteins, SMILES strings, and DNA. The taxonomy does the work of converting a fast-moving, scattered literature into a structured comparison, and its leaf-node citations provide the survey's evidence base.

What would settle it

A systematic literature search with explicit inclusion criteria that finds a substantial number of biological-design foundation models outside the three architecture families or outside the listed controllability categories would show the taxonomy is not a faithful map of the field.

Watch

Extended reading notes

Core claim

The paper claims that the current landscape of foundation models for biological sequence design can be usefully organized by three commonly applied architectures — Transformer, state space, and diffusion — and that controllability in generation is the key practical bottleneck. It assembles representative models in a taxonomy tree and a summary table, showing that small-molecule models often use classic GPT-style Transformers or SSMs; protein models range from Transformers to Mamba to discrete diffusion; and genomic models rely on long-context SSM-like backbones such as Hyena. The paper further argues that a distinct family of models is emerging that integrates multiple modalities — sequence, structure, natural language — and that open problems include inconsistent benchmarks, data scarcity, transfer to low-data regimes, and control over rare property combinations.

Load-bearing premise

The organizing value of the taxonomy rests on an informal, non-exhaustive selection of papers, so a biased or unrepresentative sample of the literature would distort which architectures and control strategies look central.

Editorial extensions

If this is right

  • Researchers new to biological design can use the taxonomy to locate existing models by architecture and control strategy before choosing an approach.
  • The survey's open-problems list implies that progress depends less on new architectures than on standardized benchmarks and domain-specific pre-training objectives.
  • If the taxonomy is right, hybrid models that combine attention, state-space layers, and diffusion are the current frontier for capturing long-range biological dependencies.
  • The inclusion of natural-language-integrated models suggests that conversational interfaces to biological design are becoming a practical direction.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy's restriction to sequence-based representations likely underrepresents graph-based and structure-based generative models, so the map should be read as one slice of the design landscape rather than the whole field.
  • Because the survey does not weight models by adoption or empirical success, the leaf-node citation counts are not evidence about which approach works best; a head-to-head benchmark would be needed to convert the map into a ranking.
  • A testable extension would be to use the taxonomy as a template for a living, versioned database of biological-design foundation models, updated as preprints mature.
  • The emphasis on controllability suggests that future progress may hinge on reward design and property-prediction objectives rather than model scale alone.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This manuscript surveys foundation models for biological design, focusing on sequence-based models for proteins, small molecules, and DNA/RNA. The authors propose a taxonomy organized around three architecture families (Transformer, state-space, and diffusion models), a set of controllability strategies for generation (fine-tuning, conditional generation, reinforcement learning, custom objectives, post-generation filtering, numerical property optimization, multi-condition tradeoffs, sampling algorithms), and multi-modal integration (combining sequence types, sequence and structure, and natural language). The survey compiles recent models in Table 1 and Figure 1, discusses background concepts, and closes with open problems and future directions. The central descriptive claim is that this taxonomy organizes current foundation-model methods for biological sequence design.

Significance. If the taxonomy were clean, the paper would provide a useful entry point to a fast-moving field, and the compilation of recent models and controllability strategies is genuinely helpful for researchers new to the area. The paper is honest about its scope limitations and does not overclaim systematic coverage. However, the central organizing axis of the taxonomy conflates network architecture with generative paradigm, which is not a cosmetic issue: several models fall into multiple leaves simultaneously, and the 'which architecture for which task' discussion rests on a distinction the taxonomy cannot cleanly draw. The paper would be significantly strengthened by separating these two dimensions.

major comments (2)
  1. [Section 'Taxonomy and Survey', Figure 1, Table 1] The top-level 'Architectures' split is not a partition because 'Diffusion Models' is a training/generation paradigm, not a network architecture, and the manuscript itself demonstrates this: DiffuMol is described as 'Diffusion on embeddings, Transformer decoder for denoising,' NOS is a 'BERT-based Encoder-Decoder' used with diffusion, DNA-Diffusion uses a 'U-net backbone,' and EvoDiff uses a 'Dilated CNN.' As a result, references [24], [25], and [32] appear in both the Transformer and Diffusion leaves of Figure 1, and Table 1's 'Architecture' column cannot assign a unique value to these models. This conflation makes the taxonomy's central organizing axis non-partitioning and undercuts the 'Which Architectures for Which Biological Task' discussion in the Open Problems section, since a model can belong to two architecture families at once. Please recast the taxonomy to separate the network backbone dimension (Transformer, SSM, CNN/hybrid) from the generative/training paradigm dimension (autoregressive, diffusion, etc.), or explicitly classify each model on both dimensions and justify the chosen grouping.
  2. [Introduction and 'Taxonomy and Survey'] The survey selection is informal and undocumented: there is no stated search strategy, inclusion criterion, time window, or comparison with prior surveys. The authors acknowledge that the survey 'cannot be fully comprehensive' and may be 'somewhat narrow,' but the central descriptive claim depends on the selected papers being representative of the field. Without a short methodology paragraph describing how papers were identified and selected, a reader cannot judge how the selection may bias the taxonomy's emphasis on Transformer, SSM, and diffusion approaches. Please add such a paragraph and briefly discuss the selection's likely effect on the taxonomy's coverage.
minor comments (5)
  1. [Throughout] Fix typographical errors and inconsistent spellings, including 'liklihood' (Background), 'its its' (Introduction), and the inconsistent rendering of the author name 'Özçelik' ('Ozccelik', 'Ozc-celik', 'Ozc-celik' in different places).
  2. [Table 1 and Section 'Controllability in Generation'] Table 1 lists the model as 'Taiga' with reference [19], while the text refers to 'Mazuz et al. [19]' for the same model; please make the naming consistent throughout.
  3. [Section 'Diffusion Models'] The comparison of DNA-Diffusion to DALL-E is inaccurate: DALL-E uses a discrete VAE with an autoregressive Transformer rather than a U-Net, so the analogy should be replaced or removed.
  4. [Figure 1] The color coding by biological domain is explained only in the caption; adding a legend inside the figure would improve readability, and the leaf-node citation keys would be easier to use if they were directly linked to Table 1 rows.
  5. [Open Problems] The phrase 'such as diffusion models for textual molecular representations' in the 'Which Architectures for Which Biological Task' paragraph repeats the architecture/paradigm conflation identified in Major Comment 1; it should read 'diffusion-based generative models' to avoid ambiguity.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey is descriptive, with all model attributions external and no derivation that reduces to its own inputs.

full rationale

This paper is a literature survey, not a derivation-driven study. It makes no fitted-parameter claims, no empirical predictions, and no formal reductions from equations to conclusions. Its central assertion is that it 'presents and discusses a taxonomy of current models and methods,' and that taxonomy is supported by citations to external papers for each model entry in Figure 1 and Table 1. The categories themselves are organizing labels, and no claim is defined in terms of another claim the paper is purporting to establish. The acknowledged limitations, such as the statement that the survey 'cannot be fully comprehensive' and 'may be somewhat narrow,' concern completeness rather than circularity. Likewise, the potential structural criticism that 'Diffusion' is a generative paradigm rather than a network architecture, while a legitimate validity concern for the taxonomy, is not a circularity: the paper does not use that classification to derive any result equivalent to its own input. There are no load-bearing self-citations, no imported uniqueness theorems, and no renamed empirical patterns presented as new predictions. Accordingly, the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

As a survey, the paper introduces no free parameters and no new physical or computational entities. The claims rest on scope-defining assumptions: a broad definition of foundation models, a language analogy for biological sequences, an informal literature selection, and the use of in-silico metrics as proxies for biological utility.

assumptions (4)
  • domain assumption Foundation models are defined broadly as architecture-agnostic models trained with task-agnostic pre-training, following Bommasani et al.
    This definition sets inclusion boundaries; a stricter LLM-only definition would exclude several cited models. Invoked in the Background and Preliminaries section on FMs.
  • domain assumption Biological sequences can be treated as text, so NLP architectures and pre-training objectives transfer to proteins, SMILES, and DNA.
    The survey focuses on sequence-based models and sets aside graph-based and most structure-based approaches, so the taxonomy's scope depends on this analogy. Stated in the Introduction and Biological Design sections.
  • ad hoc to paper Informal paper selection without a documented search protocol is sufficient to represent the field.
    No inclusion criteria are given; the authors say the survey 'cannot be fully comprehensive,' so coverage and balance rest on author judgment.
  • domain assumption In-silico quality metrics such as validity, uniqueness, QED, and predicted activity are meaningful proxies for biological utility.
    The paper treats generative performance on computational metrics as progress toward wet-lab design; if these metrics do not transfer to function, the reported successes are weaker than implied.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Foundation Models for AI-Enabled Biological Design." pith.science (2026). https://pith.science/paper/IFFF5PPW

@misc{pith2026250511610,
  author       = {Pith},
  title        = {Pith review of: Foundation Models for AI-Enabled Biological Design},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IFFF5PPW}},
  note         = {Machine review of arXiv:2505.11610}
}
read the original abstract

This paper surveys foundation models for AI-enabled biological design, focusing on recent developments in applying large-scale, self-supervised models to tasks such as protein engineering, small molecule design, and genomic sequence design. Though this domain is evolving rapidly, this survey presents and discusses a taxonomy of current models and methods. The focus is on challenges and solutions in adapting these models for biological applications, including biological sequence modeling architectures, controllability in generation, and multi-modal integration. The survey concludes with a discussion of open problems and future directions, offering concrete next-steps to improve the quality of biological sequence generation.

Figures

Figures reproduced from arXiv: 2505.11610 by the authors.

Figure 1
Figure 1. Taxonomy of recent contributions in FMs for Biological Design. Leaf nodes contain references to the relevant papers, [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

66 extracted references · 53 canonical work pages

  1. [24]

    Generalized biomolecular modeling and design with rosettafold all- atom.Science, 384(6693):eadl2528, 2024

    Rohith Krishna, Jue Wang, Woody Ahern, Pas- cal Sturmfels, Preetham Venkatesh, Indrek Kalvet, Gyu Rie Lee, Felix S Morey-Burrows, Ivan An- ishchenko, Ian R Humphreys, et al. Generalized biomolecular modeling and design with rosettafold all- atom.Science, 384(6693):eadl2528, 2024

  2. [25]

    Protein design with guided discrete diffusion

    Nate Gruver, Samuel Stanton, Nathan Frey, Tim GJ Rudner, Isidro Hotzel, Julien Lafrance-Vanasse, Arvind Rajpal, Kyunghyun Cho, and Andrew G Wilson. Protein design with guided discrete diffusion. Advances in neural information processing systems, 36, 2024

  3. [32]

    Hitting stride by degrees: Fine grained molecular generation via diffusion model

    Xinmiao Peng and Fei Zhu. Hitting stride by degrees: Fine grained molecular generation via diffusion model. Expert Systems with Applications, 244:122949, 2024

  4. [1]

    National Institute of Gen- eral Medical Sciences, U.S

    Protein structure initiative. National Institute of Gen- eral Medical Sciences, U.S. National Institutes of Health, 2000–2015. Available from https://www. nigms.nih.gov/Research/specificareas/PSI

  5. [2]

    Collins and Leslie Fink

    Francis S. Collins and Leslie Fink. The human genome project.Alcohol Health and Research World, 19(3):190–195, 1995

  6. [3]

    Natalie L Dawson, Ian Sillitoe, Jonathan G Lees, Su Datt Lam, and Christine A Orengo. Cath-gene3d: generation of the resource and its use in obtaining structural and functional annotations for protein se- quences.Protein Bioinformatics: From Protein Mod- ifications and Networks to Proteomics, pages 79–110, 2017

  7. [4]

    The structural genomics consortium: a knowledge platform for drug discovery: a summary

    Molly Morgan Jones, Sophie Castle-Clarke, Daniel Brooker, Edward Nason, Farah Huzair, and Joanna Chataway. The structural genomics consortium: a knowledge platform for drug discovery: a summary. Rand health quarterly, 4(3), 2014

  8. [5]

    Protein data bank.Nature New Biol, 233(223):10–1038, 1971

    Protein Data Bank. Protein data bank.Nature New Biol, 233(223):10–1038, 1971

Show all 66 references
  1. [6]

    Highly accurate protein structure prediction with alphafold.nature, 596(7873):583–589, 2021

    John Jumper, Richard Evans, et al. Highly accurate protein structure prediction with alphafold.nature, 596(7873):583–589, 2021

  2. [7]

    The role of ai in drug discovery: chal- lenges, opportunities, and strategies.Pharmaceuticals, 16(6):891, 2023

    Alexandre Blanco-Gonzalez, Alfonso Cabezon, Ale- jandro Seco-Gonzalez, Daniel Conde-Torres, Paula Antelo-Riveiro, Angel Pineiro, and Rebeca Garcia- Fandino. The role of ai in drug discovery: chal- lenges, opportunities, and strategies.Pharmaceuticals, 16(6):891, 2023

  3. [8]

    Synthetic biology 2020–2030: six commercially-available products that are changing our world.Nature Communications, 11(1):1–6, 2020

    Christopher A V oigt. Synthetic biology 2020–2030: six commercially-available products that are changing our world.Nature Communications, 11(1):1–6, 2020

  4. [9]

    Materials design by synthetic biol- ogy.Nature Reviews Materials, 6(4):332–350, 2021

    Tzu-Chieh Tang, Bolin An, Yuanyuan Huang, Sangita Vasikaran, Yanyi Wang, Xiaoyu Jiang, Timothy K Lu, and Chao Zhong. Materials design by synthetic biol- ogy.Nature Reviews Materials, 6(4):332–350, 2021

  5. [10]

    On the opportunities and risks of foundation models.arXiv e-prints, pages arXiv–2108, 2021

    Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models.arXiv e-prints, pages arXiv–2108, 2021

  6. [11]

    Attention is all you need.arXiv preprint arXiv:1706.03762, 2017

    Ashish Vaswani. Attention is all you need.arXiv preprint arXiv:1706.03762, 2017

  7. [12]

    Efficiently modeling long sequences with structured state spaces

    Albert Gu, Karan Goel, and Christopher R´e. Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396, 2021

  8. [13]

    Mamba: Linear-time sequence modeling with selective state spaces.arXiv preprint arXiv:2312.00752, 2023

    Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces.arXiv preprint arXiv:2312.00752, 2023

  9. [14]

    Diffusion-lm im- proves controllable text generation.Advances in Neu- ral Information Processing Systems, 35:4328–4343, 2022

    Xiang Li, John Thickstun, Ishaan Gulrajani, Percy S Liang, and Tatsunori B Hashimoto. Diffusion-lm im- proves controllable text generation.Advances in Neu- ral Information Processing Systems, 35:4328–4343, 2022

  10. [15]

    Argmax flows and multinomial diffusion: Learning categorical distribu- tions.Advances in Neural Information Processing Sys- tems, 34:12454–12465, 2021

    Emiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forr ´e, and Max Welling. Argmax flows and multinomial diffusion: Learning categorical distribu- tions.Advances in Neural Information Processing Sys- tems, 34:12454–12465, 2021

  11. [16]

    Autoregressive diffusion models

    Emiel Hoogeboom, Alexey A Gritsenko, Jasmijn Bast- ings, Ben Poole, Rianne van den Berg, and Tim Sali- mans. Autoregressive diffusion models. InInterna- tional Conference on Learning Representations, 2021

  12. [17]

    Molgpt: molecular generation using a transformer-decoder model.Journal of Chemical In- formation and Modeling, 62(9):2064–2076, 2021

    Viraj Bagal, Rishal Aggarwal, PK Vinod, and U Deva Priyakumar. Molgpt: molecular generation using a transformer-decoder model.Journal of Chemical In- formation and Modeling, 62(9):2064–2076, 2021

  13. [18]

    Multi- constraint molecular generation based on conditional transformer, knowledge distillation and reinforcement learning.Nature Machine Intelligence, 3(10):914–922, 2021

    Jike Wang, Chang-Yu Hsieh, Mingyang Wang, Xi- aorui Wang, Zhenxing Wu, Dejun Jiang, Benben Liao, Xujun Zhang, Bo Yang, Qiaojun He, et al. Multi- constraint molecular generation based on conditional transformer, knowledge distillation and reinforcement learning.Nature Machine I...

  14. [19]

    Molecule generation using transformers and policy gradient reinforcement learning.Scientific Re- ports, 13(1):8799, 2023

    Eyal Mazuz, Guy Shtar, Bracha Shapira, and Lior Rokach. Molecule generation using transformers and policy gradient reinforcement learning.Scientific Re- ports, 13(1):8799, 2023

  15. [20]

    Regression trans- former enables concurrent sequence regression and generation for molecular language modelling.Nature Machine Intelligence, 5(4):432–444, 2023

    Jannis Born and Matteo Manica. Regression trans- former enables concurrent sequence regression and generation for molecular language modelling.Nature Machine Intelligence, 5(4):432–444, 2023

  16. [21]

    Large language models generate func- tional protein sequences across diverse families.Na- ture Biotechnology, 41(8):1099–1106, 2023

    Ali Madani, Ben Krause, Eric R Greene, Subu Subra- manian, Benjamin P Mohr, James M Holton, Jose Luis Olmos, Caiming Xiong, Zachary Z Sun, Richard Socher, et al. Large language models generate func- tional protein sequences across diverse families.Na- ture Biotechnology, 41(8)...

  17. [22]

    Progen2: exploring the boundaries of protein language models.Cell systems, 14(11):968–978, 2023

    Erik Nijkamp, Jeffrey A Ruffolo, Eli N Weinstein, Nikhil Naik, and Ali Madani. Progen2: exploring the boundaries of protein language models.Cell systems, 14(11):968–978, 2023

  18. [23]

    Protgpt2 is a deep unsupervised language model for protein design.Nature communications, 13(1):4348, 2022

    Noelia Ferruz, Steffen Schmidt, and Birte H ¨ocker. Protgpt2 is a deep unsupervised language model for protein design.Nature communications, 13(1):4348, 2022

  19. [26]

    Sequence modeling and design from molecular to genome scale with evo.Science, 2024

    E Nguyen, M Poli, MG Durrant, AW Thomas, B Kang, J Sullivan, MY Ng, A Lewis, A Patel, A Lou, et al. Sequence modeling and design from molecular to genome scale with evo.Science, 2024

  20. [27]

    Chemical language modeling with structured state space sequence models.Nature Communications, 15(1):6176, 2024

    Rıza ¨Ozc ¸elik, Sarah de Ruiter, Emanuele Criscuolo, and Francesca Grisoni. Chemical language modeling with structured state space sequence models.Nature Communications, 15(1):6176, 2024

  21. [28]

    Protmamba: a homology-aware but alignment-free protein state space model.bioRxiv, pages 2024–05, 2024

    Damiano Sgarbossa, Cyril Malbranke, and Anne- Florence Bitbol. Protmamba: a homology-aware but alignment-free protein state space model.bioRxiv, pages 2024–05, 2024

  22. [29]

    reglm: Designing realistic regu- latory dna with autoregressive language models

    Avantika Lal, David Garfield, Tommaso Biancalani, and Gokcen Eraslan. reglm: Designing realistic regu- latory dna with autoregressive language models. InIn- ternational Conference on Research in Computational Molecular Biology, pages 332–335. Springer, 2024

  23. [30]

    Discdiff: Latent diffusion model for dna sequence gen- eration.CoRR, 2024

    Zehui Li, Yuhao Ni, William A V Beardall, Guoxuan Xia, Akashaditya Das, Guy-Bart Stan, and Yiren Zhao. Discdiff: Latent diffusion model for dna sequence gen- eration.CoRR, 2024

  24. [31]

    Dna-diffusion: Leveraging generative models for controlling chro- matin accessibility and gene expression via synthetic regulatory elements

    Simon Senan, Aniketh Janardhan Reddy, Zach Nuss- baum, Aaron Wenteler, Matei Bejan, Michael I Love, Wouter Meuleman, and Luca Pinello. Dna-diffusion: Leveraging generative models for controlling chro- matin accessibility and gene expression via synthetic regulatory elements. I...

  25. [33]

    Protein generation with evolutionary diffusion: se- quence is all you need

    Sarah Alamdari, Nitya Thakkar, Rianne van den Berg, Alex Lu, Nicolo Fusi, Ava Amini, and Kevin Yang. Protein generation with evolutionary diffusion: se- quence is all you need. InNeurIPS 2023 Generative AI and Biology (GenBio) Workshop, 2023

  26. [34]

    Towards joint sequence-structure gen- eration of nucleic acid and protein complexes with se (3)-discrete diffusion

    Alex Morehead, Jeffrey Ruffolo, Aadyot Bhatnagar, and Ali Madani. Towards joint sequence-structure gen- eration of nucleic acid and protein complexes with se (3)-discrete diffusion. InProceedings of the NeurIPS 2023 Workshop on Machine Learning in Structural Bi- ology, 2023

  27. [35]

    Model-based reinforcement learning for biological se- quence design

    Christof Angermueller, David Dohan, David Belanger, Ramya Deshpande, Kevin Murphy, and Lucy Colwell. Model-based reinforcement learning for biological se- quence design. InInternational conference on learning representations, 2019

  28. [36]

    nach0: Multimodal natural and chemical languages foundation model

    Micha Livne, Zulfat Miftahutdinov, Elena Tutubalina, Maksim Kuznetsov, Daniil Polykovskiy, Annika Brun- dyn, Aastha Jhunjhunwala, Anthony Costa, Alex Aliper, Al´an Aspuru-Guzik, et al. nach0: Multimodal natural and chemical languages foundation model. Chemical Science, 15(22):...

  29. [37]

    Instruct- biomol: Advancing biomolecule understanding and de- sign following human instructions.arXiv preprint arXiv:2410.07919, 2024

    Xiang Zhuang, Keyan Ding, Tianwen Lyu, Yinuo Jiang, Xiaotong Li, Zhuoyi Xiang, Zeyuan Wang, Ming Qin, Kehua Feng, Jike Wang, et al. Instruct- biomol: Advancing biomolecule understanding and de- sign following human instructions.arXiv preprint arXiv:2410.07919, 2024

  30. [38]

    Chatnt: A mul- timodal conversational agent for dna, rna and protein tasks.bioRxiv, pages 2024–04, 2024

    Guillaume Richard, Bernardo P de Almeida, Hugo Dalla-Torre, Christopher Blum, Lorenz Hexemer, Priyanka Pandey, Stefan Laurent, Marie P Lopez, Alexander Laterre, Maren Lang, et al. Chatnt: A mul- timodal conversational agent for dna, rna and protein tasks.bioRxiv, pages 2024–04, 2024

  31. [39]

    Roformer: Enhanced trans- former with rotary position embedding.Neurocomput- ing, 568:127063, 2024

    Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced trans- former with rotary position embedding.Neurocomput- ing, 568:127063, 2024

  32. [40]

    Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution.Ad- vances in neural information processing systems, 36, 2024

    Eric Nguyen, Michael Poli, Marjan Faizi, Armin Thomas, Michael Wornow, Callum Birch-Sykes, Ste- fano Massaroli, Aman Patel, Clayton Rabideau, Yoshua Bengio, et al. Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution.Ad- vances in neural information p...

  33. [41]

    Xlnet: Generalized autoregressive pre- training for language understanding.arXiv preprint arXiv:1906.08237, 2019

    Zhilin Yang. Xlnet: Generalized autoregressive pre- training for language understanding.arXiv preprint arXiv:1906.08237, 2019

  34. [42]

    Generative molecular design in low data regimes.Nature Machine Intelligence, 2(3):171–180, 2020

    Michael Moret, Lukas Friedrich, Francesca Grisoni, Daniel Merk, and Gisbert Schneider. Generative molecular design in low data regimes.Nature Machine Intelligence, 2(3):171–180, 2020

  35. [43]

    Hyena hierar- chy: Towards larger convolutional language models

    Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Ben- gio, Stefano Ermon, and Christopher R´e. Hyena hierar- chy: Towards larger convolutional language models. In International Conference on Machine Learning, pages 28043–28078. PMLR, 2023

  36. [44]

    Protein-mamba: Biological mamba models for protein function prediction.arXiv preprint arXiv:2409.14617, 2024

    Bohao Xu, Yingzhou Lu, Yoshitaka Inoue, Namkyeong Lee, Tianfan Fu, and Jintai Chen. Protein-mamba: Biological mamba models for protein function prediction.arXiv preprint arXiv:2409.14617, 2024

  37. [45]

    Ptm-mamba: A ptm-aware protein lan- guage model with bidirectional gated mamba blocks

    Zhangzhi Peng, Benjamin Schussheim, and Pranam Chatterjee. Ptm-mamba: A ptm-aware protein lan- guage model with bidirectional gated mamba blocks. bioRxiv, 2024

  38. [46]

    Efficient training of language models to fill in the middle.arXiv preprint arXiv:2207.14255, 2022

    Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, and Mark Chen. Efficient training of language models to fill in the middle.arXiv preprint arXiv:2207.14255, 2022

  39. [47]

    Genomic language models: Opportunities and challenges.arXiv preprint arXiv:2407.11435, 2024

    Gonzalo Benegas, Chengzhong Ye, Carlos Albors, Jianan Canal Li, and Yun S Song. Genomic language models: Opportunities and challenges.arXiv preprint arXiv:2407.11435, 2024

  40. [48]

    Hungry hungry hippos: Towards language modeling with state space models

    Daniel Y Fu, Tri Dao, Khaled Kamal Saab, Armin W Thomas, Atri Rudra, and Christopher Re. Hungry hungry hippos: Towards language modeling with state space models. InThe Eleventh International Confer- ence on Learning Representations, 2022

  41. [49]

    Zero-shot text-to-image generation

    Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International conference on machine learning, pages 8821–8831. Pmlr, 2021

  42. [50]

    Simple statistical gradient- following algorithms for connectionist reinforcement learning.Machine learning, 8:229–256, 1992

    Ronald J Williams. Simple statistical gradient- following algorithms for connectionist reinforcement learning.Machine learning, 8:229–256, 1992

  43. [51]

    Evaluating protein transfer learning with tape.Advances in neural information processing sys- tems, 32, 2019

    Roshan Rao, Nicholas Bhattacharya, Neil Thomas, Yan Duan, Peter Chen, John Canny, Pieter Abbeel, and Yun Song. Evaluating protein transfer learning with tape.Advances in neural information processing sys- tems, 32, 2019

  44. [52]

    Exploring the limits of transfer learning with a unified text-to-text transformer.Jour- nal of machine learning research, 21(140):1–67, 2020

    Colin Raffel, Noam Shazeer, Adam Roberts, Kather- ine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer.Jour- nal of machine learning research, 21(140):1–67, 2020

  45. [53]

    Language models are few-shot learners

    Tom B Brown. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020

  46. [54]

    A theory of biological relativity: no priv- ileged level of causation.Interface focus, 2(1):55–64, 2012

    Denis Noble. A theory of biological relativity: no priv- ileged level of causation.Interface focus, 2(1):55–64, 2012

  47. [55]

    It’s time to admit that genes are not the blueprint for life.Nature, 626(7998):254–255, 2024

    Denis Noble. It’s time to admit that genes are not the blueprint for life.Nature, 626(7998):254–255, 2024

  48. [56]

    Springer Nature, 2023

    Jeremy Ramsden.Bioinformatics: an introduction. Springer Nature, 2023

  49. [57]

    The omg dataset: An open metagenomic corpus for mixed- modality genomic language modeling.bioRxiv, pages 2024–08, 2024

    Andre Cornman, Jacob West-Roberts, Antonio Pe- dro Camargo, Simon Roux, Martin Beracochea, Milot Mirdita, Sergey Ovchinnikov, and Yunha Hwang. The omg dataset: An open metagenomic corpus for mixed- modality genomic language modeling.bioRxiv, pages 2024–08, 2024

  50. [58]

    Effi- cient and accurate prediction of protein structure using rosettafold2.BioRxiv, pages 2023–05, 2023

    Minkyung Baek, Ivan Anishchenko, Ian R Humphreys, Qian Cong, David Baker, and Frank DiMaio. Effi- cient and accurate prediction of protein structure using rosettafold2.BioRxiv, pages 2023–05, 2023

  51. [59]

    Gene ontology: tool for the unification of biology.Nature genetics, 25(1):25–29, 2000

    Michael Ashburner, Catherine A Ball, Judith A Blake, David Botstein, Heather Butler, J Michael Cherry, Al- lan P Davis, Kara Dolinski, Selina S Dwight, Janan T Eppig, et al. Gene ontology: tool for the unification of biology.Nature genetics, 25(1):25–29, 2000

  52. [60]

    The nucleotide transformer: Building and evaluating robust foundation models for human ge- nomics.BioRxiv, pages 2023–01, 2023

    Hugo Dalla-Torre, Liam Gonzalez, Javier Mendoza- Revilla, Nicolas Lopez Carranza, Adam Henryk Grzywaczewski, Francesco Oteri, Christian Dallago, Evan Trop, Bernardo P de Almeida, Hassan Sirelkha- tim, et al. The nucleotide transformer: Building and evaluating robust foundation...

  53. [61]

    Scaling laws for neural language models.arXiv preprint arXiv:2001.08361, 2020

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361, 2020

  54. [62]

    Explaining neural scaling laws.Proceedings of the National Academy of Sci- ences, 121(27):e2311878121, 2024

    Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma. Explaining neural scaling laws.Proceedings of the National Academy of Sci- ences, 121(27):e2311878121, 2024

  55. [63]

    Biolog- ical structure and function emerge from scaling un- supervised learning to 250 million protein sequences

    Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C Lawrence Zitnick, Jerry Ma, et al. Biolog- ical structure and function emerge from scaling un- supervised learning to 250 million protein sequences. Proceedings of the Natio...

  56. [64]

    Neural scaling of deep chemical models.Nature Machine Intelligence, 5(11):1297–1305, 2023

    Nathan C Frey, Ryan Soklaski, Simon Axelrod, Sid- dharth Samsi, Rafael Gomez-Bombarelli, Connor W Coley, and Vijay Gadepally. Neural scaling of deep chemical models.Nature Machine Intelligence, 5(11):1297–1305, 2023

  57. [65]

    Molca: Molecular graph-language modeling with cross-modal projector and uni-modal adapter

    Zhiyuan Liu, Sihang Li, Yanchen Luo, Hao Fei, Yixin Cao, Kenji Kawaguchi, Xiang Wang, and Tat-Seng Chua. Molca: Molecular graph-language modeling with cross-modal projector and uni-modal adapter. InProceedings of the 2023 Conference on Empiri- cal Methods in Natural Language P...

  58. [66]

    Foundation models for scientific discovery and innovation: Opportunities across the department of energy, 2024

    National Academies of Sciences Engineering and Medicine. Foundation models for scientific discovery and innovation: Opportunities across the department of energy, 2024. Accessed: 2024-10-22

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.