{"id":"f20f58d9-2155-4194-8a3c-4e1c7c39802f","arxiv_id":"2505.11610","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey taxonomizes foundation models for biological sequence design into Transformer, state-space, and diffusion families, with controllability and multi-modality as the main design axes.","lead":"This paper reviews recent AI foundation models for designing proteins, small molecules, and DNA, and groups them into a taxonomy by architecture and by how generation is controlled. It is a useful orientation map for anyone tracking how large language models and related architectures are being adapted to biological design.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The taxonomy's central architecture axis conflates generative paradigm (diffusion) with network architecture, so models like DiffuMol and NOS occupy two leaves at once; a backbone-relabeling test would show whether the organizing axis is valid.","rationale":"The reader's weakest point is that the survey's informal selection could bias the map of the field. That is a fair limitation, and the authors do acknowledge it ('cannot be fully comprehensive,' 'somewhat narrow'). But a more load-bearing issue sits inside the taxonomy itself: the paper's central architecture axis treats diffusion as an architecture coordinate. Diffusion is a generative modeling paradigm that can wrap any backbone, and the paper's own Table 1 shows this—DiffuMol and NOS are diffusion-based yet use Transformer/BERT-style denoisers, while DNA-Diffusion uses a U-Net and EvoDiff uses a dilated CNN. Consequently the Figure 1 leaves are not mutually exclusive, and the taxonomy cannot cleanly support the paper's own open-problem question about which architecture to choose. This is an internal inconsistency, not merely an external coverage gap, and it is not fixed by adding more papers. I therefore focus on that rather than on the selection concern. To be fair, the survey's Table 1 broadly aligns with the cited papers and the prose descriptions are generally faithful to the sources; the issue is the organizing frame, not the individual summaries. The survey still has value as a curated snapshot, and the fix is conceptual rather than empirical: draw diffusion as a separate axis (generative objective) from backbone architecture. That is a moderate revision, consistent with the existing CONDITIONAL verdict, so I mark UNCHANGED rather than escalating to REJECT. The proposed backbone-relabeling test is a direct way to make the category overlap visible and to decide whether the taxonomy needs the two-axis reframing.","tokens_in":15568,"tokens_out":7885,"duration_ms":81059,"concrete_test":"Reclassify every model in Table 1 by the actual network backbone used for generation/denoising (e.g., Transformer, SSM/Hyena, CNN/U-Net, GNN), ignoring whether it is trained by diffusion, autoregression, or objective alternation. If the 'Diffusion' entries split across several backbone classes (DiffuMol and NOS to Transformer; DNA-Diffusion to U-Net; EvoDiff to CNN), then the three-way 'architecture' split is not a partition and the paper should reframe diffusion as an orthogonal generative paradigm rather than a third architecture family. Count how many Table 1 models would be assigned to two of the current leaves; if the overlap is non-empty, the taxonomy needs a two-axis revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central artifact is Figure 1's taxonomy, and its top split is claimed to be architectures: Transformer, State Space, and Diffusion (Section 'Taxonomy and Survey'; Table 1). But diffusion is a training/generation process, not a network architecture. The paper itself shows this: DiffuMol is described as 'Diffusion on embeddings, Transformer decoder for denoising'; NOS is 'BERT-based encoder-decoder' used for diffusion; DNA-Diffusion uses a 'U-net convolutional architecture'; EvoDiff uses a 'Dilated CNN'. Thus the 'Diffusion Models' leaf is orthogonal to the Transformer/SSM leaves, and several Table 1 entries fall in two architecture leaves simultaneously. This is not just a labeling nit: it makes the taxonomy's central organizing principle non-partitioning, so the later discussion of 'which architectures for which biological task' (Open Problems) asks a question that the taxonomy cannot cleanly answer. The acknowledged coverage limitations (informal selection) affect the survey's completeness; the architecture/diffusion conflation affects the validity of the taxonomy's structure itself. A reader cannot tell whether the intended contrast is between autoregressive vs denoising training objectives or between attention vs state-space vs convolutional backbones, and those are different comparisons with different engineering takeaways.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript surveys foundation models for biological design, focusing on sequence-based models for proteins, small molecules, and DNA/RNA. The authors propose a taxonomy organized around three architecture families (Transformer, state-space, and diffusion models), a set of controllability strategies for generation (fine-tuning, conditional generation, reinforcement learning, custom objectives, post-generation filtering, numerical property optimization, multi-condition tradeoffs, sampling algorithms), and multi-modal integration (combining sequence types, sequence and structure, and natural language). The survey compiles recent models in Table 1 and Figure 1, discusses background concepts, and closes with open problems and future directions. The central descriptive claim is that this taxonomy organizes current foundation-model methods for biological sequence design.","tokens_in":15766,"tokens_out":4123,"duration_ms":42602,"significance":"If the taxonomy were clean, the paper would provide a useful entry point to a fast-moving field, and the compilation of recent models and controllability strategies is genuinely helpful for researchers new to the area. The paper is honest about its scope limitations and does not overclaim systematic coverage. However, the central organizing axis of the taxonomy conflates network architecture with generative paradigm, which is not a cosmetic issue: several models fall into multiple leaves simultaneously, and the 'which architecture for which task' discussion rests on a distinction the taxonomy cannot cleanly draw. The paper would be significantly strengthened by separating these two dimensions.","major_comments":[{"comment":"The top-level 'Architectures' split is not a partition because 'Diffusion Models' is a training/generation paradigm, not a network architecture, and the manuscript itself demonstrates this: DiffuMol is described as 'Diffusion on embeddings, Transformer decoder for denoising,' NOS is a 'BERT-based Encoder-Decoder' used with diffusion, DNA-Diffusion uses a 'U-net backbone,' and EvoDiff uses a 'Dilated CNN.' As a result, references [24], [25], and [32] appear in both the Transformer and Diffusion leaves of Figure 1, and Table 1's 'Architecture' column cannot assign a unique value to these models. This conflation makes the taxonomy's central organizing axis non-partitioning and undercuts the 'Which Architectures for Which Biological Task' discussion in the Open Problems section, since a model can belong to two architecture families at once. Please recast the taxonomy to separate the network backbone dimension (Transformer, SSM, CNN/hybrid) from the generative/training paradigm dimension (autoregressive, diffusion, etc.), or explicitly classify each model on both dimensions and justify the chosen grouping.","section":"Section 'Taxonomy and Survey', Figure 1, Table 1"},{"comment":"The survey selection is informal and undocumented: there is no stated search strategy, inclusion criterion, time window, or comparison with prior surveys. The authors acknowledge that the survey 'cannot be fully comprehensive' and may be 'somewhat narrow,' but the central descriptive claim depends on the selected papers being representative of the field. Without a short methodology paragraph describing how papers were identified and selected, a reader cannot judge how the selection may bias the taxonomy's emphasis on Transformer, SSM, and diffusion approaches. Please add such a paragraph and briefly discuss the selection's likely effect on the taxonomy's coverage.","section":"Introduction and 'Taxonomy and Survey'"}],"minor_comments":[{"comment":"Fix typographical errors and inconsistent spellings, including 'liklihood' (Background), 'its its' (Introduction), and the inconsistent rendering of the author name 'Özçelik' ('Ozccelik', 'Ozc-celik', 'Ozc-celik' in different places).","section":"Throughout"},{"comment":"Table 1 lists the model as 'Taiga' with reference [19], while the text refers to 'Mazuz et al. [19]' for the same model; please make the naming consistent throughout.","section":"Table 1 and Section 'Controllability in Generation'"},{"comment":"The comparison of DNA-Diffusion to DALL-E is inaccurate: DALL-E uses a discrete VAE with an autoregressive Transformer rather than a U-Net, so the analogy should be replaced or removed.","section":"Section 'Diffusion Models'"},{"comment":"The color coding by biological domain is explained only in the caption; adding a legend inside the figure would improve readability, and the leaf-node citation keys would be easier to use if they were directly linked to Table 1 rows.","section":"Figure 1"},{"comment":"The phrase 'such as diffusion models for textual molecular representations' in the 'Which Architectures for Which Biological Task' paragraph repeats the architecture/paradigm conflation identified in Major Comment 1; it should read 'diffusion-based generative models' to avoid ambiguity.","section":"Open Problems"}],"recommendation":"major_revision","confidential_remarks":"This is a survey paper, so the main technical risk is the validity of its organizing taxonomy rather than any derived numerical claim. The architecture-versus-paradigm conflation is real and central enough to warrant a major revision, but it is fixable by restructuring Figure 1 and Table 1. I do not see evidence of citation-pattern problems or novelty-disclosure issues; the fit with the journal's scope is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a decent, honest survey that will help people entering biological sequence design. The table and the per-model descriptions are the most useful parts. The organizing taxonomy has a real structural flaw: it lists \"Diffusion\" as an architecture alongside Transformer and SSM, but diffusion is a training/generation process, not a network backbone. Their own Table 1 shows DiffuMol uses a Transformer decoder, NOS uses a BERT encoder-decoder, and DNA-Diffusion and EvoDiff use CNN backbones. So the central axis does not partition the field. A reader cannot tell whether the intended contrast is autoregressive versus denoising objectives, or attention versus state-space versus convolutional backbones—those are different comparisons with different engineering takeaways. The stress-test note lands.\n\nWhat is good: the descriptions of individual models track their cited sources, the three-domain coverage (small molecule, protein, DNA) is sensible, and Table 1 is a compact reference worth keeping. The controllability section is reasonably organized, and the paper does not engage in circular reasoning—every entry is attributed to external work. The authors also admit the selection is informal and not comprehensive, which is honest.\n\nSoft spots beyond the taxonomy: there is no search strategy or inclusion criteria, so coverage is at the authors' discretion; there is no positioning against earlier surveys, so a newcomer does not know what this adds; and there are a handful of typos ('liklihood', 'its its', 'Ozc-celik'). The conclusion's \"truly revolutionize\" line is hype, though it is clearly labeled as future-looking. Also, the definition of foundation model is broad enough to include nearly any pretrained generative model, which weakens the survey's boundaries.\n\nNet: this is a useful orienting review, not a research contribution. It deserves a serious referee, and a good referee could push it into solid shape. I would ask for revision: reframe the taxonomy axis as \"generative paradigms\" or split diffusion out from backbone architectures, document the scope, and situate the survey against existing reviews. I would not cite it for a specific technical claim, but I would point a student to the table. Reading group: maybe, if the group is new to the area.","headline":"A useful but structurally imperfect survey: the per-model table is the real value, but the taxonomy's 'diffusion as architecture' axis needs rework before this is publishable.","tokens_in":16295,"tokens_out":1831,"would_cite":false,"duration_ms":20162,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey proposes a taxonomy that organizes foundation models for biological sequence design by architecture family, controllability strategy, and multimodal integration.","keywords":["foundation models","biological design","protein language models","small molecule design","genomic sequence design","state space models","diffusion models","controllable generation"],"falsifier":"A systematic literature search with explicit inclusion criteria that finds a substantial number of biological-design foundation models outside the three architecture families or outside the listed controllability categories would show the taxonomy is not a faithful map of the field.","tokens_in":15326,"feed_emoji":"🧬","tokens_out":5291,"duration_ms":50725,"temperature":0.7,"pith_summary":"This survey organizes recent foundation models for biological design into a taxonomy with three axes: architecture family (Transformer, state-space, and diffusion), controllability strategy (fine-tuning, conditional generation, reinforcement learning, custom objectives, post-generation filtering, and sampling), and multimodal integration. Its aim is to give researchers a working map of sequence-based generative models for proteins, small molecules, and genomes, and to highlight where the field lacks consensus. A sympathetic reader would take away that most current models borrow NLP-style architectures and objectives, and that the central engineering problem is controlling generation toward desired biological properties.","feed_headline":"Survey maps biological-design AI into three architecture families","feed_subtitle":"Proteins, small molecules, and genomes organized by Transformer, state-space, and diffusion models plus how generation is steered.","key_machinery":"The central organizing device is a two-part taxonomy: a tree whose leaf nodes carry citation clusters for architectures, controllability, and multimodality, paired with a table that lists each representative model's domain, architecture, design goal, and generation method. The underlying conceptual move is the analogy between natural language and biological sequences, which lets decoder-only Transformers, SSMs, and diffusion models be imported from NLP and adapted to proteins, SMILES strings, and DNA. The taxonomy does the work of converting a fast-moving, scattered literature into a structured comparison, and its leaf-node citations provide the survey's evidence base.","core_discovery":"The paper claims that the current landscape of foundation models for biological sequence design can be usefully organized by three commonly applied architectures — Transformer, state space, and diffusion — and that controllability in generation is the key practical bottleneck. It assembles representative models in a taxonomy tree and a summary table, showing that small-molecule models often use classic GPT-style Transformers or SSMs; protein models range from Transformers to Mamba to discrete diffusion; and genomic models rely on long-context SSM-like backbones such as Hyena. The paper further argues that a distinct family of models is emerging that integrates multiple modalities — sequence, structure, natural language — and that open problems include inconsistent benchmarks, data scarcity, transfer to low-data regimes, and control over rare property combinations.","pith_inferences":["The taxonomy's restriction to sequence-based representations likely underrepresents graph-based and structure-based generative models, so the map should be read as one slice of the design landscape rather than the whole field.","Because the survey does not weight models by adoption or empirical success, the leaf-node citation counts are not evidence about which approach works best; a head-to-head benchmark would be needed to convert the map into a ranking.","A testable extension would be to use the taxonomy as a template for a living, versioned database of biological-design foundation models, updated as preprints mature.","The emphasis on controllability suggests that future progress may hinge on reward design and property-prediction objectives rather than model scale alone."],"forward_implications":["Researchers new to biological design can use the taxonomy to locate existing models by architecture and control strategy before choosing an approach.","The survey's open-problems list implies that progress depends less on new architectures than on standardized benchmarks and domain-specific pre-training objectives.","If the taxonomy is right, hybrid models that combine attention, state-space layers, and diffusion are the current frontier for capturing long-range biological dependencies.","The inclusion of natural-language-integrated models suggests that conversational interfaces to biological design are becoming a practical direction."],"supporting_citations":[{"why":"Supplies the broad, architecture-agnostic definition of foundation models the survey adopts.","marker":"[10]"},{"why":"Introduces the attention mechanism and Transformer architecture that anchor the Transformer branch of the taxonomy.","marker":"[11]"},{"why":"Introduces structured state-space sequence models, the basis for the SSM branch.","marker":"[12]"},{"why":"Introduces Mamba selective state spaces, used by several protein and genomic models in the survey.","marker":"[13]"},{"why":"Adapts diffusion to discrete text tokens, providing the template for discrete biological-sequence diffusion.","marker":"[14]"},{"why":"Presents the first S4-based chemical language model for de novo small-molecule generation, a key SSM leaf node.","marker":"[27]"},{"why":"Describes the large-scale protein language model ProGen, a key Transformer leaf node and fine-tuning example.","marker":"[21]"},{"why":"Introduces Evo, a StripedHyena-based genomic model operating on sequences up to 131 kilobases, a key specialized architecture leaf node.","marker":"[26]"},{"why":"Presents EvoDiff, an evolutionary discrete diffusion approach for proteins, a key diffusion leaf node.","marker":"[33]"},{"why":"Introduces the Regression Transformer with concurrent property regression and generation, load-bearing for numerical-property control.","marker":"[20]"}],"fun_headline_variants":["Bio-design AI taxonomy: Transformer, SSM, diffusion families","Three architectures dominate biological sequence foundation models","Controllability is the key bottleneck in bio-design AI","Multimodal fusion emerges in biological-design foundation models","Survey charts foundation models from proteins to genomes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The organizing value of the taxonomy rests on an informal, non-exhaustive selection of papers, so a biased or unrepresentative sample of the literature would distort which architectures and control strategies look central.","fun_headline_variants_meta":{"raw":{"variants":["Bio-design AI taxonomy: Transformer, SSM, diffusion families","Three architectures dominate biological sequence foundation models","Controllability is the key bottleneck in bio-design AI","Multimodal fusion emerges in biological-design foundation models","Survey charts foundation models from proteins to genomes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000265,"raw_usage":{"total_tokens":1524,"prompt_tokens":779,"completion_tokens":745,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":395,"completion_tokens_details":{"reasoning_tokens":671}},"tokens_in":395,"tokens_out":745,"duration_ms":7772,"temperature":1.0,"reasoning_tokens":671,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:50:50.219411+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic literature search with explicit inclusion criteria that finds a substantial number of biological-design foundation models outside the three architecture families or outside the listed controllability categories would show the taxonomy is not a faithful map of the field.","supporting_citations":[{"cited_title":"On the opportunities and risks of foundation models.arXiv e-prints, pages arXiv–2108, 2021","cited_arxiv_id":null,"evidence_quote":"Supplies the broad, architecture-agnostic definition of foundation models the survey adopts."},{"cited_title":"Diffusion-lm im- proves controllable text generation.Advances in Neu- ral Information Processing Systems, 35:4328–4343, 2022","cited_arxiv_id":null,"evidence_quote":"Adapts diffusion to discrete text tokens, providing the template for discrete biological-sequence diffusion."},{"cited_title":"Chemical language modeling with structured state space sequence models.Nature Communications, 15(1):6176, 2024","cited_arxiv_id":null,"evidence_quote":"Presents the first S4-based chemical language model for de novo small-molecule generation, a key SSM leaf node."},{"cited_title":"Large language models generate func- tional protein sequences across diverse families.Na- ture Biotechnology, 41(8):1099–1106, 2023","cited_arxiv_id":null,"evidence_quote":"Describes the large-scale protein language model ProGen, a key Transformer leaf node and fine-tuning example."},{"cited_title":"Sequence modeling and design from molecular to genome scale with evo.Science, 2024","cited_arxiv_id":null,"evidence_quote":"Introduces Evo, a StripedHyena-based genomic model operating on sequences up to 131 kilobases, a key specialized architecture leaf node."},{"cited_title":"Protein generation with evolutionary diffusion: se- quence is all you need","cited_arxiv_id":null,"evidence_quote":"Presents EvoDiff, an evolutionary discrete diffusion approach for proteins, a key diffusion leaf node."},{"cited_title":"Regression trans- former enables concurrent sequence regression and generation for molecular language modelling.Nature Machine Intelligence, 5(4):432–444, 2023","cited_arxiv_id":null,"evidence_quote":"Introduces the Regression Transformer with concurrent property regression and generation, load-bearing for numerical-property control."}],"review_version":1}