Pith. sign in

REVIEW 2 cited by

ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.07920 v1 pith:MURYFP2U submitted 2024-02-28 q-bio.BM cs.AIcs.CLcs.LG

classification q-bio.BMcs.AIcs.CLcs.LG
keywords protllmlanguageprotein-languagetasksproteinprotein-centricproteinsdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose ProtLLM, a versatile cross-modal large language model (LLM) for both protein-centric and protein-language tasks. ProtLLM features a unique dynamic protein mounting mechanism, enabling it to handle complex inputs where the natural language text is interspersed with an arbitrary number of proteins. Besides, we propose the protein-as-word language modeling approach to train ProtLLM. By developing a specialized protein vocabulary, we equip the model with the capability to predict not just natural language but also proteins from a vast pool of candidates. Additionally, we construct a large-scale interleaved protein-text dataset, named InterPT, for pre-training. This dataset comprehensively encompasses both (1) structured data sources like protein annotations and (2) unstructured data sources like biological research papers, thereby endowing ProtLLM with crucial knowledge for understanding proteins. We evaluate ProtLLM on classic supervised protein-centric tasks and explore its novel protein-language applications. Experimental results demonstrate that ProtLLM not only achieves superior performance against protein-specialized baselines on protein-centric tasks but also induces zero-shot and in-context learning capabilities on protein-language tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EvoLlama: Enhancing LLMs' Understanding of Proteins via Multimodal Structure and Sequence Representations

    cs.LG 2024-12 conditional novelty 4.0 of 10

    EvoLlama aligns ESM-2 sequence embeddings and ProteinMPNN structure embeddings with Llama-3, improving protein understanding over text-only LLMs on Mol-Instructions and PEER benchmarks.

  2. Computational Protein Science in the Era of Large Language Models (LLMs)

    cs.CE 2025-01 conditional novelty 3.0 of 10

    A survey that categorizes protein language models by the knowledge they learn and reviews their applications, with no new experimental results.

Pith tools