REVIEW 4 major objections 6 minor 36 references
Evaluating DNA function understanding in genomic language models using evolutionarily implausible sequences
T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Genomic language models fail to predict loss-of-function in synthetic DNA; their accuracy tracks sequence likelihood, not biological mechanism.
desk verdict A useful new gLM generalization benchmark, but the headline interpretation—LL-dependence as proof of evolutionary-prior reliance—is undercut by a likely floor effect in the scoring metric. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the virtual translocation: each Nullsette is created by moving one regulatory element (promoter, start codon, CDS, stop codon, terminator, and in prokaryotes RBS) to another position, preserving all components while breaking the canonical 5'-3' order. Non-functionality is defined in silico by three ordering rules checked over all circular permutations of the mutant; a mutant is retained only if no permutation satisfies all rules. The evaluation then compares the log-likelihood distribution of nonmutant cassettes with that of the mutants, and the paper's central analytical device is the relationship between the nonmutant log-likelihood and the per-cassette success rate, summarized by a length-scaled threshold.
What would settle it
Run a reporter assay on a random sample of Nullsette translocation mutants in the appropriate host. If many mutants labeled non-functional still produce protein, the ground-truth labels are wrong; if models with high nonmutant likelihood nevertheless miss experimentally confirmed loss-of-function mutants, the claim that predictions rely on the evolutionary prior would need to be revised.
Extended reading notes
Core claim
The paper's discovery is that genomic language models trained on natural genomes do not understand the regulatory grammar of gene expression well enough to generalize to synthetic cassettes. Nullsettes constructs loss-of-function mutants by translocating one regulatory element in each cassette and defines non-functionality through three ordering rules applied over all circular permutations. Evaluating 12 models, the authors find most identify fewer than half of the mutation types on at least one dataset, and Evo2-7B and GENERanno-0.5B are the only consistent strong performers. The unifying failure is that a model's success at detecting a mutation falls as the log-likelihood it assigns to the original sequence falls, a length-scaled threshold separates reliable from unreliable prediction, and this holds across architectures. The authors conclude that predictions track evolutionary plausibility rather than the biological effect of the mutation.
Load-bearing premise
The benchmark labels a mutant non-functional purely from ordering rules, assuming that no alternative start codon, cryptic promoter, or translational reinitiation can restore expression in the rearranged cassettes.
Editorial extensions
If this is right
- Designers should not treat gLM likelihood scores as reliable standalone predictors of function for engineered DNA.
- The linear relationship between cassette length and the likelihood threshold for accurate prediction gives a concrete, testable operating rule for when a model's mutation predictions can be trusted.
- Pretraining corpus content may matter more than parameter count, since GENERanno-0.5B matched Evo2-7B with far fewer parameters.
- Benchmarks centered on natural variants miss a core failure mode; functional-generalization benchmarks need synthetic, out-of-distribution sequences.
Reading between the lines
- The same likelihood-dependence pattern could be tested directly with experimental measurements: if reporter assays confirm that a specific Nullsette mutant is non-functional while a model predicts it as functional only because the cassette has low likelihood, the paper's interpretation is strengthened.
- A natural next step is to fine-tune a gLM on Nullsettes-style rearrangements and ask whether the improvement transfers to natural regulatory variants; transfer would show the missing knowledge is learnable rather than inherently absent.
- The ordering-rule definition treats expression as a purely linear grammar, so models that incorporate RNA secondary structure or co-transcriptional folding might show a different failure profile on the same benchmark.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Nullsettes, a benchmark of in silico loss-of-function (LOF) mutations in synthetic expression cassettes with little evolutionary precedent, and evaluates 12 genomic language models (gLMs) on zero-shot detection of these mutations. The central finding is that a model's ability to identify LOF mutants declines sharply as the log-likelihood (LL) of the nonmutant sequence decreases, which the authors interpret as evidence that gLMs rely on evolutionary pattern-matching rather than mechanistic understanding of gene expression. The benchmark curates cassettes from five MPRA datasets and applies systematic element translocations defined by a grammar of three ordering rules.
Significance. If the interpretation holds, the paper would be an important cautionary result for using gLMs in synthetic biology, showing that current models do not generalize to engineered sequences outside natural sequence space. The benchmark itself is novel, spans a wide range of model families and scales, and is accompanied by public code and data. The observation that a much smaller model (GENERanno) can match Evo2-7B is also informative about pretraining-data relevance. However, the central claim currently rests on a correlation that has not been separated from a scoring floor/power artifact, and the MLM scoring uses unmasked logits that are not a valid likelihood. The significance is therefore conditional on additional controls and on reanalysis with proper pseudo-likelihood scores.
major comments (4)
- [§4.4 (Eq. 5) and Figure 2A/B] The central claim that low nonmutant LL causes poor LOF detection is confounded by a floor/power effect. The paired permutation test compares mutant LL against nonmutant LL; for cassettes whose nonmutant LL is already near the model's output floor, the mutant LL cannot move much lower, so the paired differences are small and the test loses power regardless of whether the model encodes functional constraints. A model with a saturating likelihood function and no functional understanding would produce exactly the same pattern. The manuscript provides no null model (e.g., comparing each mutant against a randomly chosen nonmutant), no matched non-LOF control mutations, and no regression controlling for sequence length, GC composition, or other confounds. Without these controls, the decline in success rate with nonmutant LL does not uniquely support the 'evolutionary prior reliance' interpretation. Please add these analyses or explicitly discuss the floor artifact as an alternative explanation.
- [§4.4 (Eqs. 3–4)] For masked language models (MLMs), the sequence-level score is computed from unmasked logits, producing P(xt|X) that is not a valid conditional probability because the model sees the token itself. The authors acknowledge that this 'lacks strict probabilistic interpretation,' but 7 of the 12 model families (NT, GENERanno, Caduceus, GPN, GPN-Promoter, DNABERT2, gLM2) are scored this way. The paper's statement that 'all models show a sharp drop in prediction accuracy as the likelihood assigned to the original sequence decreases' is therefore not supported for those models in terms of actual model likelihood. Please recompute scores using a proper pseudo-likelihood (e.g., single-inference PLL or full PLL) or restrict the claim to CLM models and show MLM results separately.
- [§4.2] The ground-truth LOF labels are defined entirely in silico by three ordering rules applied to circular permutations. The claim that the resulting mutants are 'completely non-functional' is stronger than the evidence supports, because the rules do not exclude rescue by alternative start codons, cryptic promoters, internal ribosome entry, or translational reinitiation. The authors do note for some mutants that no other in-frame start codon exists, but no experimental validation is provided for any mutant. If some labeled LOF mutants are actually functional, the model evaluation would conflate syntactic violations with true loss of function. Please either provide empirical validation of a subset of mutants, or soften the wording to 'predicted non-functional by the canonical grammar' and discuss the potential impact of label noise on the conclusions.
- [§4.1.3] The cassette curation includes a step that compares LL distributions between the Kosuri and Lagator datasets and prioritizes Kosuri promoter–RBS pairs whose LL distributions surpass those from Lagator. This selects on the independent variable used in the main correlation (nonmutant LL). The authors should test whether the LL-performance correlation persists when this selection step is removed, to verify that it is not an artifact of the curation procedure.
minor comments (6)
- [§2 and §4.4] The Results state that significance is determined by a paired t-test, while Methods use a one-sided paired permutation test (Eq. 5). Please reconcile this inconsistency and state which test was used for the reported success rates.
- [§4.4 and Table S4] The phrase 'casual language models' should be 'causal language models' in both the main text and the Supplementary Table S4 header.
- [Figure 2A] The caption says 'Each point represents an expression cassette' while the y-axis is described as 'Nullsettes prediction performance.' It is unclear how a binary success/failure is assigned to a single cassette when the permutation test in Methods operates over all cassettes for a mutation type. Please clarify the plotting and the per-cassette success definition.
- [§4.4] The citation of Gordon et al. [25] for single-inference PLL is misleading because the manuscript does not use that method; it uses unmasked logits. Either implement the cited method or remove the citation.
- [Figure 2C] The claim that the LL threshold 'scales linearly with sequence length' is based on only five datasets and a linear regression with no reported uncertainty or goodness-of-fit. Please report R², confidence intervals, or use a larger set of sequence lengths before drawing this conclusion.
- [Supplementary Figure S2] The caption says 'four datasets' but lists five (Abf1TATA, pTpA, Zahm, Kosuri, Lagator). Please correct the count.
Circularity Check
The Nullsettes benchmark itself is self-contained; only the side claim that a linear model can 'predict' the LL threshold reduces to its own fitted values.
-
fitted input called prediction
[Section 2, Figure 2C paragraph (Results and discussion)]
"For each model–dataset pair, we estimate the LL threshold corresponding to a success rate of 0.85 using linear regression. Figure 2C reveals that these thresholds are consistent among top performing models (after rescaling by lowest threshold). Furthermore, this threshold scales linearly with sequence length, showing a simple linear model can predict an effective LL threshold for zero-shot mutation effect prediction."
The 'prediction' that a simple linear model can produce the effective LL threshold is a post-hoc description of the same fitted values: the thresholds are themselves estimated by linear regression per model–dataset pair, and the length scaling is another linear fit to those estimated thresholds. No held-out data or independent validation is used, so the claimed predictive relationship is equivalent to the fit by construction. This is a side observation rather than the central Nullsettes result, so it contributes only minor circularity.
full rationale
The core benchmark is not circular. Nullsettes mutants are defined by a fixed, model-independent grammar: single-element translocations are retained as LOF mutants only if no circular permutation satisfies Rules 1–3 (Section 4.2), and the nonmutant cassettes are taken from experimentally assayed MPRA datasets. Model scores are computed with Equations 1–4 and compared with a paired permutation test (Equation 5); no model parameter is fitted to the benchmark labels, and no benchmark label is derived from model likelihoods. The central claim that success rate declines with nonmutant LL is an empirical correlation, not a mathematical identity: success is defined by whether mutant LL is significantly below nonmutant LL, not by the absolute value of nonmutant LL. The deliberate curation of low-LL and high-LL cassettes (Section 4.1.3) shapes the range of the data, but the per-cassette trend in Figure 2A is still an observed result rather than a construction. The one genuine circular element is the 'LL threshold scales with sequence length' statement, where the threshold is obtained by regression and then 'predicted' by another regression on the same estimates. Separately, the reviewer concern that a likelihood floor or loss of permutation-test power could produce the LL-performance decline is a validity threat, not a circularity, because the paper never defines the outcome in terms of the nonmutant LL. There is no load-bearing self-citation: the only potentially overlapping citation is [6], used for protein DMS context and not for the benchmark's validity, whereas GENERanno [18] is an external model with no shared authorship with this paper.
Assumptions & free parameters
free parameters (3)
- Selection thresholds for cassette curation =
1500 promoters per dataset; expression > µ+1.5σ for Kosuri; top 1500 by activity for Lagator, deBoer, Zahm
- Success rate threshold for defining accuracy =
0.85
- Linear regression model for LL threshold vs sequence length =
Linear fit with dataset-median thresholds
assumptions (3)
- domain assumption A cassette is functional iff a circular permutation satisfies the three ordering rules (promoter before start, stop after CDS or before start, terminator after CDS or before promoter).
- domain assumption gLM log-likelihood reflects evolutionary plausibility and should decrease for functionally disrupted sequences if the model understands function.
- domain assumption MPRA-measured activity indicates true in vivo function of the nonmutant cassettes.
Cite this review
Pith. "Pith review of Evaluating DNA function understanding in genomic language models using evolutionarily implausible sequences." pith.science (2026). https://pith.science/paper/AL2EPAWL
@misc{pith2026250610271,
author = {Pith},
title = {Pith review of: Evaluating DNA function understanding in genomic language models using evolutionarily implausible sequences},
year = {2026},
howpublished = {\url{https://pith.science/paper/AL2EPAWL}},
note = {Machine review of arXiv:2506.10271}
}
read the original abstract
Genomic language models (gLMs) hold promise for generating novel, functional DNA sequences for synthetic biology. However, realizing this potential requires models to go beyond evolutionary plausibility and understand how DNA sequence encodes gene expression and regulation. We introduce a benchmark called Nullsettes, which assesses how well models can predict in silico loss-of-function (LOF) mutations, in synthetic expression cassettes with little evolutionary precedent. Testing 12 state-of-the-art gLMs, we find that most fail to consistently detect these strong LOF mutations. All models show a sharp drop in predictive accuracy as the likelihood assigned to the original (nonmutant) sequence decreases, suggesting that gLMs rely heavily on pattern-matching to their evolutionary prior rather than on any mechanistic understanding of gene expression. Our findings highlight fundamental limitations in how gLMs generalize to engineered, non-natural sequences, and underscore the need for benchmarks and modeling strategies that prioritize functional understanding.
Figures
Reference graph
Works this paper leans on
-
[1]
Ge- nomic language models: opportunities and challenges
Gonzalo Benegas, Chengzhong Ye, Carlos Albors, Jianan Canal Li, and Yun S Song. Ge- nomic language models: opportunities and challenges. Trends in Genetics, 2025
work page 2025
-
[2]
Transformers and genome language models
Micaela E Consens, Cameron Dufault, Michael Wainberg, Duncan Forster, Mehran Karimzadeh, Hani Goodarzi, Fabian J Theis, Alan Moses, and Bo Wang. Transformers and genome language models. Nature Machine Intelligence, pages 1–17, 2025
work page 2025
-
[3]
Efficient evolution of human antibodies from general protein language models
Brian L Hie, Varun R Shanker, Duo Xu, Theodora UJ Bruun, Payton A Weidenbacher, Shaogeng Tang, Wesley Wu, John E Pak, and Peter S Kim. Efficient evolution of human antibodies from general protein language models. Nature biotechnology, 42(2):275–283, 2024
2024
-
[4]
Yan He, Xibin Zhou, Chong Chang, Ge Chen, Weikuan Liu, Geng Li, Xiaoqi Fan, Ming- sun Sun, Chensi Miao, Qianyue Huang, et al. Protein language models-assisted optimiza- tion of a uracil-n-glycosylase variant enables programmable t-to-g and t-to-c base editing. Molecular Cell, 84(7):1257–1270, 2024
work page 2024
-
[5]
Integrating protein language models and automatic biofoundry for enhanced protein evolution
Qiang Zhang, Wanyi Chen, Ming Qin, Yuhao Wang, Zhongji Pu, Keyan Ding, Yuyue Liu, Qunfeng Zhang, Dongfang Li, Xinjia Li, et al. Integrating protein language models and automatic biofoundry for enhanced protein evolution. Nature Communications, 16(1):1553, 2025
work page 2025
-
[6]
Saprothub: Making protein modeling accessible to all biologists
Jin Su, Zhikai Li, Chenchen Han, Yuyang Zhou, Yan He, Junjie Shan, Xibin Zhou, Xing Chang, Shiyu Jiang, Dacheng Ma, et al. Saprothub: Making protein modeling accessible to all biologists. BioRxiv, pages 2024–05, 2024
work page 2024
-
[7]
Proteingym: Large-scale benchmarks for protein fitness prediction and design
Pascal Notin, Aaron Kollasch, Daniel Ritter, Lood Van Niekerk, Steffanie Paul, Han Spin- ner, Nathan Rollins, Ada Shaw, Rose Orenbuch, Ruben Weitzman, et al. Proteingym: Large-scale benchmarks for protein fitness prediction and design. Advances in Neural Information Processing Systems, 36:64331–64379, 2023
work page 2023
-
[8]
Dna language models are powerful predictors of genome-wide variant effects.Proceedings of the National Academy of Sciences, 120(44):e2311219120, 2023
Gonzalo Benegas, Sanjit Singh Batra, and Yun S Song. Dna language models are powerful predictors of genome-wide variant effects.Proceedings of the National Academy of Sciences, 120(44):e2311219120, 2023
2023
Show all 36 references
-
[9]
A 5’ utr language model for decoding untranslated regions of mrna and function predictions
Yanyi Chu, Dan Yu, Yupeng Li, Kaixuan Huang, Yue Shen, Le Cong, Jason Zhang, and Mengdi Wang. A 5’ utr language model for decoding untranslated regions of mrna and function predictions. Nature Machine Intelligence, 6(4):449–460, 2024
2024
-
[10]
Evaluating the representational power of pre-trained dna language models for regulatory genomics
Ziqi Tang, Nirali Somia, Yiyang Yu, and Peter K Koo. Evaluating the representational power of pre-trained dna language models for regulatory genomics. Genome Biology, 26 (1):203, 2025. 11
2025
-
[11]
Synthetic design of strong promoters
Michael R Schlabach, Jimmy K Hu, Mamie Li, and Stephen J Elledge. Synthetic design of strong promoters. Proceedings of the national academy of sciences, 107(6):2538–2543, 2010
2010
-
[12]
miRNA circuit modules for precise, tunable control of gene expression
Rongrong Du, Michael J Flynn, Monique Honsa, Ralf Jungmann, and Michael B Elowitz. miRNA circuit modules for precise, tunable control of gene expression. BioRxiv, 2024
2024
-
[13]
Applications of synthetic biology in medical and pharmaceutical fields
Xu Yan, Xu Liu, Cuihuan Zhao, and Guo-Qiang Chen. Applications of synthetic biology in medical and pharmaceutical fields. Signal transduction and targeted therapy, 8(1):199, 2023
2023
-
[14]
Predicting bacterial promoter function and evolution from random sequences
Mato Lagator, Srdjan Sarikas, Magdalena Steinrueck, David Toledo-Aparicio, Jonathan P Bollback, Calin C Guet, and Gaˇ sper Tkaˇ cik. Predicting bacterial promoter function and evolution from random sequences. Elife, 11:e64543, 2022
2022
-
[15]
Deciphering eukaryotic gene-regulatory logic with 100 million ran- dom promoters
Carl G de Boer, Eeshit Dhaval Vaishnav, Ronen Sadeh, Esteban Luis Abeyta, Nir Fried- man, and Aviv Regev. Deciphering eukaryotic gene-regulatory logic with 100 million ran- dom promoters. Nature biotechnology, 38(1):56–65, 2020
2020
-
[16]
Composability of regulatory sequences controlling transcription and translation in escherichia coli
Sriram Kosuri, Daniel B Goodman, Guillaume Cambray, Vivek K Mutalik, Yuan Gao, Adam P Arkin, Drew Endy, and George M Church. Composability of regulatory sequences controlling transcription and translation in escherichia coli. Proceedings of the National Academy of Sciences, 11...
2013
-
[17]
A massively parallel reporter assay library to screen short synthetic promoters in mammalian cells
Adam M Zahm, William S Owens, Samuel R Himes, Braden S Fallon, Kathleen E Rondem, Alexa N Gormick, Joshua S Bloom, Sriram Kosuri, Henry Chan, and Justin G English. A massively parallel reporter assay library to screen short synthetic promoters in mammalian cells. Nature Commun...
2024
-
[18]
Generanno: A genomic foundation model for metagenomic annotation
Qiuyi Li, Wei Wu, Yiheng Zhu, Fuli Feng, Jieping Ye, and Zheng Wang. Generanno: A genomic foundation model for metagenomic annotation. bioRxiv, pages 2025–06, 2025
2025
-
[19]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human langu...
2019
-
[20]
Lan- guage models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Lan- guage models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020
1901
-
[21]
StripedHyena: Moving Beyond Transformers with Hybrid Signal Pro- cessing Models, 12 2023
Michael Poli, Jue Wang, Stefano Massaroli, Jeffrey Quesnelle, Ryan Carlow, Eric Nguyen, and Armin Thomas. StripedHyena: Moving Beyond Transformers with Hybrid Signal Pro- cessing Models, 12 2023. URL https://github.com/togethercomputer/stripedhyena
2023
-
[22]
Systems and algorithms for convolutional multi-hybrid language models at scale
Jerome Ku, Eric Nguyen, David W Romero, Garyk Brixi, Brandon Yang, Anton Vorontsov, Ali Taghibakhshi, Amy X Lu, Dave P Burke, Greg Brockman, et al. Systems and algorithms for convolutional multi-hybrid language models at scale. arXiv preprint arXiv:2503.01868, 2025
2025 arXiv
-
[23]
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023
2023 arXiv
-
[24]
Bert has a mouth, and it must speak: Bert as a markov random field language model
Alex Wang and Kyunghyun Cho. Bert has a mouth, and it must speak: Bert as a markov random field language model. arXiv preprint arXiv:1902.04094, 2019. 12
1902 arXiv
-
[25]
Protein language model fitness is a matter of preference
Cade Gordon, Amy X Lu, and Pieter Abbeel. Protein language model fitness is a matter of preference. bioRxiv, pages 2024–10, 2024
2024
-
[26]
Metagene-1: Metagenomic foundation model for pan- demic monitoring
Ollie Liu, Sami Jaghouar, Johannes Hagemann, Shangshang Wang, Jason Wiemels, Jeff Kaufman, and Willie Neiswanger. Metagene-1: Metagenomic foundation model for pan- demic monitoring. arXiv preprint arXiv:2501.02045, 2025
2025 arXiv
-
[27]
Nucleotide transformer: building and evaluating robust foundation models for human genomics
Hugo Dalla-Torre, Liam Gonzalez, Javier Mendoza-Revilla, Nicolas Lopez Carranza, Adam Henryk Grzywaczewski, Francesco Oteri, Christian Dallago, Evan Trop, Bernardo P de Almeida, Hassan Sirelkhatim, et al. Nucleotide transformer: building and evaluating robust foundation models...
2024
-
[28]
Generator: A long-context generative genomic foundation model
Wei Wu, Qiuyi Li, Mingyang Li, Kun Fu, Fuli Feng, Jieping Ye, Hui Xiong, and Zheng Wang. Generator: A long-context generative genomic foundation model. arXiv preprint arXiv:2502.07272, 2025
2025
-
[29]
Sequence modeling and design from molecular to genome scale with evo
Eric Nguyen, Michael Poli, Matthew G Durrant, Brian Kang, Dhruva Katrekar, David B Li, Liam J Bartie, Armin W Thomas, Samuel H King, Garyk Brixi, et al. Sequence modeling and design from molecular to genome scale with evo. Science, 386(6723):eado9336, 2024
2024
-
[30]
Semantic mining of functional de novo genes from a genomic language model
Aditi T Merchant, Samuel H King, Eric Nguyen, and Brian L Hie. Semantic mining of functional de novo genes from a genomic language model. bioRxiv, pages 2024–12, 2024
2024
-
[31]
Genome modeling and design across all domains of life with evo 2
Garyk Brixi, Matthew G Durrant, Jerome Ku, Michael Poli, Greg Brockman, Daniel Chang, Gabriel A Gonzalez, Samuel H King, David B Li, Aditi T Merchant, et al. Genome modeling and design across all domains of life with evo 2. bioRxiv, pages 2025–02, 2025
2025
-
[32]
Dnabert-2: Efficient foundation model and benchmark for multi-species genome
Zhihan Zhou, Yanrong Ji, Weijian Li, Pratik Dutta, Ramana Davuluri, and Han Liu. Dnabert-2: Efficient foundation model and benchmark for multi-species genome. arXiv preprint arXiv:2306.15006, 2023
2023 arXiv
-
[33]
Caduceus: Bi-directional equivariant long-range dna sequence modeling
Yair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao, Albert Gu, and Volodymyr Kuleshov. Caduceus: Bi-directional equivariant long-range dna sequence modeling. arXiv preprint arXiv:2403.03234, 2024
2024 arXiv
-
[34]
Benchmarking dna sequence models for causal regulatory variant prediction in human genetics
Gonzalo Benegas, G¨ okcen Eraslan, and Yun S Song. Benchmarking dna sequence models for causal regulatory variant prediction in human genetics. bioRxiv, pages 2025–02, 2025
2025
-
[35]
The omg dataset: An open metagenomic corpus for mixed-modality genomic language modeling
Andre Cornman, Jacob West-Roberts, Antonio Pedro Camargo, Simon Roux, Martin Be- racochea, Milot Mirdita, Sergey Ovchinnikov, and Yunha Hwang. The omg dataset: An open metagenomic corpus for mixed-modality genomic language modeling. bioRxiv, pages 2024–08, 2024
2024
-
[36]
nucleotide- transformer-2.5b-multi-species
Guanqing Liu, Long Chen, Yuechao Wu, Yangshuo Han, Yu Bao, and Tao Zhang. Pdllms: A group of tailored dna large language models for analyzing plant genomes. Molecular Plant, 18(2):175–178, 2025. 13 Supplementary Information Supplemental Figures Figure S1: Comparison of promote...
2025
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.