REVIEW 5 cited by
ProLLaMA: A Protein Large Language Model for Multi-Task Protein Language Processing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent advances in Protein Language Models (PLMs) have transformed protein engineering, yet unlike their counterparts in Natural Language Processing (NLP), current PLMs exhibit a fundamental limitation: they excel in either Protein Language Understanding (PLU) or Protein Language Generation (PLG), but rarely both. This fragmentation hinders progress in protein engineering. To bridge this gap, we introduce ProLLaMA, a multitask protein language model enhanced by the Evolutionary Protein Generation Framework (EPGF). We construct a comprehensive instruction dataset containing approximately 13 million samples with over 11,000 superfamily annotations to facilitate better modeling of sequence-function landscapes. We leverage a two-stage training approach to develop ProLLaMA, a multitask LLM with protein domain expertise. Our EPGF addresses the mismatch between statistic language modeling and biological constraints through three innovations: a multi-dimensional interpretable scorer, hierarchical efficient decoding, and a probabilistic-biophysical joint selection mechanism. Extensive experiments demonstrate that ProLLaMA excels in both unconditional and controllable protein generation tasks, achieving superior structural quality metrics compared to existing PLMs. Additionally, ProLLaMA demonstrates strong understanding capabilities with a 67.1% exact match rate in superfamily prediction. EPGF significantly enhances the biological viability of generated sequences, as evidenced by improved biophysical scores (+4.3%) and structural metrics (+14.5%). The project is available at https://github.com/PKU-YuanGroup/ProLLaMA.
Forward citations
Cited by 5 Pith papers
-
DisProtEdit: Exploring Disentangled Representations for Multi-Attribute Protein Editing
DisProtEdit learns disentangled protein representations from separate structural and functional text descriptions, enabling controllable single- and multi-attribute protein editing via latent interpolation.
-
Teaching LLMs to Speak Spectroscopy
A LLaMA-3.1-8B model fine-tuned with LoRA on digit-serialized SDSS spectra predicts redshifts with MAE 0.043 and retains 85% of its astronomy QA performance.
-
Steering Protein Language Models
Activation steering can guide protein language models to generate and optimize sequences with higher predicted thermostability, solubility, or GFP brightness, but only in surrogate-based evaluation.
-
Nature Language Model: Deciphering the Language of Nature for Scientific Discovery
A single sequence-based model, pretrained across molecules, proteins, materials, nucleotides and text, outperforms specialist models on several generation tasks and enables cross-domain design.
-
A Comprehensive Review of Protein Language Models
A survey paper that catalogs protein language models, their architectures, training data, benchmarks, and tools, but lacks a systematic methodology and contains several factual errors.
Discussion (0). Continue with ORCID to comment.