REVIEW 6 cited by
ProtChatGPT: Towards Understanding Proteins with Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Protein research is crucial in various fundamental disciplines, but understanding their intricate structure-function relationships remains challenging. Recent Large Language Models (LLMs) have made significant strides in comprehending task-specific knowledge, suggesting the potential for ChatGPT-like systems specialized in protein to facilitate basic research. In this work, we introduce ProtChatGPT, which aims at learning and understanding protein structures via natural languages. ProtChatGPT enables users to upload proteins, ask questions, and engage in interactive conversations to produce comprehensive answers. The system comprises protein encoders, a Protein-Language Pertaining Transformer (PLP-former), a projection adapter, and an LLM. The protein first undergoes protein encoders and PLP-former to produce protein embeddings, which are then projected by the adapter to conform with the LLM. The LLM finally combines user questions with projected embeddings to generate informative answers. Experiments show that ProtChatGPT can produce promising responses to proteins and their corresponding questions. We hope that ProtChatGPT could form the basis for further exploration and application in protein research. Code and our pre-trained model will be publicly available.
Forward citations
Cited by 6 Pith papers
-
Nature Language Model: Deciphering the Language of Nature for Scientific Discovery
A single sequence-based model, pretrained across molecules, proteins, materials, nucleotides and text, outperforms specialist models on several generation tasks and enables cross-domain design.
-
Prot2Chat: Protein LLM with Early-Fusion of Text, Sequence and Structure
Prot2Chat fuses protein sequence, structure, and question text in a text-aware adapter before a LoRA-tuned LLM generates answers, reporting large gains on Mol-Instructions but mixed zero-shot results on UniProtQA.
-
Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances
Across the 68 papers it surveys, domain-specialized generative models usually outperform general-purpose LLMs on biological tasks, and agentic/conversational workflows are the least-covered topics.
-
Controllable Protein Sequence Generation with LLM Preference Optimization
CtrlProt uses multi-listwise preference optimization with Rosetta energy and structural embedding similarity to improve controllable protein sequence generation.
-
EvoLlama: Enhancing LLMs' Understanding of Proteins via Multimodal Structure and Sequence Representations
EvoLlama aligns ESM-2 sequence embeddings and ProteinMPNN structure embeddings with Llama-3, improving protein understanding over text-only LLMs on Mol-Instructions and PEER benchmarks.
-
Computational Protein Science in the Era of Large Language Models (LLMs)
A survey that categorizes protein language models by the knowledge they learn and reviews their applications, with no new experimental results.
Discussion (0). Continue with ORCID to comment.