Pith. sign in

REVIEW 1 cited by

Protein Large Language Models: A Comprehensive Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.17504 v2 pith:FYLMUCP5 submitted 2025-02-21 q-bio.BM cs.AIcs.CEcs.CLcs.LG

Protein Large Language Models: A Comprehensive Survey

classification q-bio.BM cs.AIcs.CEcs.CLcs.LG
keywords proteinllmsapplicationscomprehensivelanguagelargemodelsscience
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Protein-specific large language models (Protein LLMs) are revolutionizing protein science by enabling more efficient protein structure prediction, function annotation, and design. While existing surveys focus on specific aspects or applications, this work provides the first comprehensive overview of Protein LLMs, covering their architectures, training datasets, evaluation metrics, and diverse applications. Through a systematic analysis of over 100 articles, we propose a structured taxonomy of state-of-the-art Protein LLMs, analyze how they leverage large-scale protein sequence data for improved accuracy, and explore their potential in advancing protein engineering and biomedical research. Additionally, we discuss key challenges and future directions, positioning Protein LLMs as essential tools for scientific discovery in protein science. Resources are maintained at https://github.com/Yijia-Xiao/Protein-LLM-Survey.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Enhancing Protein Representation Learning via Manifold Restore Mixing

    cs.LG 2026-06 unverdicted novelty 3.0

    MRM mixes hidden representations of original and DA-augmented proteins and uses a difficulty scheduler on the beta distribution to produce training samples that balance structural fidelity with diversity.