REVIEW 5 cited by
Genomic Language Models: Opportunities and Challenges
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) are having transformative impacts across a wide range of scientific fields, particularly in the biomedical sciences. Just as the goal of Natural Language Processing is to understand sequences of words, a major objective in biology is to understand biological sequences. Genomic Language Models (gLMs), which are LLMs trained on DNA sequences, have the potential to significantly advance our understanding of genomes and how DNA elements at various scales interact to give rise to complex functions. To showcase this potential, we highlight key applications of gLMs, including functional constraint prediction, sequence design, and transfer learning. Despite notable recent progress, however, developing effective and efficient gLMs presents numerous challenges, especially for species with large, complex genomes. Here, we discuss major considerations for developing and evaluating gLMs.
Forward citations
Cited by 5 Pith papers
-
METAGENE-1: Metagenomic Foundation Model for Pandemic Monitoring
A 7B transformer pretrained on 1.5T base pairs of wastewater metagenomic reads achieves strong pathogen detection and embedding scores, though some new benchmarks are partly in-distribution.
-
Find Central Dogma Again: Leveraging Multilingual Transfer in Large Language Models
A GPT-2 model fine-tuned on multilingual sentence similarity achieves at best 81% accuracy on classifying matching vs non-matching DNA-protein pairs, but the result is highly seed-dependent and only with an easy test set.
-
Human Genome Book: Words, Sentences and Paragraphs
A GPT-2 model trained on DNA, protein, and English is used to segment the human genome into book-like words, sentences, paragraphs, and chapters, but transfer to DNA is only hypothesized, not validated.
-
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
A comprehensive survey that frames multimodal understanding and generation as next token prediction and proposes a five-part taxonomy.
-
Can linguists better understand DNA?
Fine-tuning on English sentence-pair similarity appears to help GPT-2 and BERT classify DNA similarity, but the evidence is weakened by missing baselines and best-of-N seed selection.
Discussion (0). Continue with ORCID to comment.