Pith. sign in

REVIEW 4 cited by

BEND: Benchmarking DNA Language Models on biologically meaningful tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.12570 v4 pith:YFZLKIDG submitted 2023-11-21 q-bio.GN cs.LG

BEND: Benchmarking DNA Language Models on biologically meaningful tasks

classification q-bio.GN cs.LG
keywords bendlanguagetasksgenomemodelssequenceannotationbiologically
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The genome sequence contains the blueprint for governing cellular processes. While the availability of genomes has vastly increased over the last decades, experimental annotation of the various functional, non-coding and regulatory elements encoded in the DNA sequence remains both expensive and challenging. This has sparked interest in unsupervised language modeling of genomic DNA, a paradigm that has seen great success for protein sequence data. Although various DNA language models have been proposed, evaluation tasks often differ between individual works, and might not fully recapitulate the fundamental challenges of genome annotation, including the length, scale and sparsity of the data. In this study, we introduce BEND, a Benchmark for DNA language models, featuring a collection of realistic and biologically meaningful downstream tasks defined on the human genome. We find that embeddings from current DNA LMs can approach performance of expert methods on some tasks, but only capture limited information about long-range features. BEND is available at https://github.com/frederikkemarin/BEND.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

    q-bio.QM 2026-07 conditional novelty 6.5

    A 52.6B-token multi-domain biology pretraining corpus with tool enrichment and new binding/localization instructions doubles a fixed base LLM's matched biology-eval score with little language forgetting.

  2. TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

    q-bio.QM 2026-07 conditional novelty 6.0

    TheBioCollection, a 52.6B-token unified biology corpus with tool-computed text and new instruction tasks, raises a fixed 16B LLM's score on its matched biology eval from 0.223 to 0.499 (2.24×).

  3. Rethinking Genomic Modeling Through Optical Character Recognition

    cs.CV 2026-02 conditional novelty 6.0

    Rendering DNA as OCR-style page images and training a vision-language model on reading/grounding/retrieval/completion tasks outperforms sequence-based genomic models on tested benchmarks with ~20x fewer effective tokens.

  4. In Search of Lost DNA Sequence Pretraining

    cs.LG 2026-04 unverdicted novelty 5.0

    DNA pretraining suffers from inappropriate evaluation datasets, flawed neighbor-masking, and neglected vocabulary design; the authors supply guidelines and a reproducible testbed to fix them.