Pith. sign in

REVIEW 1 cited by

ProGen2: Exploring the Boundaries of Protein Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.13517 v1 pith:DOPE75Z7 submitted 2022-06-27 cs.LG q-bio.QM

classification cs.LGq-bio.QM
keywords proteinmodelsprogen2sequencesmodeldatadistributionlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Attention-based models trained on protein sequences have demonstrated incredible success at classification and generation tasks relevant for artificial intelligence-driven protein design. However, we lack a sufficient understanding of how very large-scale models and data play a role in effective protein model development. We introduce a suite of protein language models, named ProGen2, that are scaled up to 6.4B parameters and trained on different sequence datasets drawn from over a billion proteins from genomic, metagenomic, and immune repertoire databases. ProGen2 models show state-of-the-art performance in capturing the distribution of observed evolutionary sequences, generating novel viable sequences, and predicting protein fitness without additional finetuning. As large model sizes and raw numbers of protein sequences continue to become more widely accessible, our results suggest that a growing emphasis needs to be placed on the data distribution provided to a protein sequence model. We release the ProGen2 models and code at https://github.com/salesforce/progen.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 50 citations worldwide. Full citation record

  1. Conditionally Site-Independent Neural Evolution of Antibody Sequences

    cs.LG 2026-02 conditional novelty 5.0 of 10

    A neural continuous-time Markov model of antibody affinity maturation that beats language models on fitness prediction and steers sampling toward antigen-specific binders.

Pith tools