Pith. sign in

REVIEW 1 cited by

VQDNA: Unleashing the Power of Vector Quantization for Multi-Species Genomic Sequence Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.10812 v2 pith:ACVXAKX5 submitted 2024-05-13 q-bio.GN cs.AI

classification q-bio.GNcs.AI
keywords genomevocabularymodelsvqdnalanguagecodebooksgenomesgenomic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Similar to natural language models, pre-trained genome language models are proposed to capture the underlying intricacies within genomes with unsupervised sequence modeling. They have become essential tools for researchers and practitioners in biology. However, the hand-crafted tokenization policies used in these models may not encode the most discriminative patterns from the limited vocabulary of genomic data. In this paper, we introduce VQDNA, a general-purpose framework that renovates genome tokenization from the perspective of genome vocabulary learning. By leveraging vector-quantized codebooks as learnable vocabulary, VQDNA can adaptively tokenize genomes into pattern-aware embeddings in an end-to-end manner. To further push its limits, we propose Hierarchical Residual Quantization (HRQ), where varying scales of codebooks are designed in a hierarchy to enrich the genome vocabulary in a coarse-to-fine manner. Extensive experiments on 32 genome datasets demonstrate VQDNA's superiority and favorable parameter efficiency compared to existing genome language models. Notably, empirical analysis of SARS-CoV-2 mutations reveals the fine-grained pattern awareness and biological significance of learned HRQ vocabulary, highlighting its untapped potential for broader applications in genomics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Omni-DNA: A Unified Genomic Foundation Model for Cross-Modal and Multi-Task Learning

    q-bio.GN 2025-02 conditional novelty 6.0 of 10

    Autoregressive DNA language models fine-tuned jointly on classification, text generation, and image generation achieve strong benchmark results and open-ended cross-modal genomic tasks.

Pith tools