Pith. sign in

REVIEW 6 cited by

ProtChatGPT: Towards Understanding Proteins with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.09649 v2 pith:PAXQES4Y submitted 2024-02-15 cs.CE cs.AIq-bio.BM

classification cs.CEcs.AIq-bio.BM
keywords proteinprotchatgptproduceproteinsquestionsresearchunderstandingadapter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Protein research is crucial in various fundamental disciplines, but understanding their intricate structure-function relationships remains challenging. Recent Large Language Models (LLMs) have made significant strides in comprehending task-specific knowledge, suggesting the potential for ChatGPT-like systems specialized in protein to facilitate basic research. In this work, we introduce ProtChatGPT, which aims at learning and understanding protein structures via natural languages. ProtChatGPT enables users to upload proteins, ask questions, and engage in interactive conversations to produce comprehensive answers. The system comprises protein encoders, a Protein-Language Pertaining Transformer (PLP-former), a projection adapter, and an LLM. The protein first undergoes protein encoders and PLP-former to produce protein embeddings, which are then projected by the adapter to conform with the LLM. The LLM finally combines user questions with projected embeddings to generate informative answers. Experiments show that ProtChatGPT can produce promising responses to proteins and their corresponding questions. We hope that ProtChatGPT could form the basis for further exploration and application in protein research. Code and our pre-trained model will be publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Nature Language Model: Deciphering the Language of Nature for Scientific Discovery

    cs.AI 2025-02 conditional novelty 6.0 of 10

    A single sequence-based model, pretrained across molecules, proteins, materials, nucleotides and text, outperforms specialist models on several generation tasks and enables cross-domain design.

  2. Prot2Chat: Protein LLM with Early-Fusion of Text, Sequence and Structure

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Prot2Chat fuses protein sequence, structure, and question text in a text-aware adapter before a LoRA-tuned LLM generates answers, reporting large gains on Mol-Instructions but mixed zero-shot results on UniProtQA.

  3. Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances

    cs.CL 2025-11 reject novelty 4.0 of 10

    Across the 68 papers it surveys, domain-specialized generative models usually outperform general-purpose LLMs on biological tasks, and agentic/conversational workflows are the least-covered topics.

  4. Controllable Protein Sequence Generation with LLM Preference Optimization

    cs.AI 2025-01 conditional novelty 4.0 of 10

    CtrlProt uses multi-listwise preference optimization with Rosetta energy and structural embedding similarity to improve controllable protein sequence generation.

  5. EvoLlama: Enhancing LLMs' Understanding of Proteins via Multimodal Structure and Sequence Representations

    cs.LG 2024-12 conditional novelty 4.0 of 10

    EvoLlama aligns ESM-2 sequence embeddings and ProteinMPNN structure embeddings with Llama-3, improving protein understanding over text-only LLMs on Mol-Instructions and PEER benchmarks.

  6. Computational Protein Science in the Era of Large Language Models (LLMs)

    cs.CE 2025-01 conditional novelty 3.0 of 10

    A survey that categorizes protein language models by the knowledge they learn and reviews their applications, with no new experimental results.

Pith tools