REVIEW 11 cited by
ProGen: Language Modeling for Protein Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Generative modeling for protein engineering is key to solving fundamental problems in synthetic biology, medicine, and material science. We pose protein engineering as an unsupervised sequence generation problem in order to leverage the exponentially growing set of proteins that lack costly, structural annotations. We train a 1.2B-parameter language model, ProGen, on ~280M protein sequences conditioned on taxonomic and keyword tags such as molecular function and cellular component. This provides ProGen with an unprecedented range of evolutionary sequence diversity and allows it to generate with fine-grained control as demonstrated by metrics based on primary sequence similarity, secondary structure accuracy, and conformational energy.
Forward citations
Cited by 11 Pith papers
-
DisProtEdit: Exploring Disentangled Representations for Multi-Attribute Protein Editing
DisProtEdit learns disentangled protein representations from separate structural and functional text descriptions, enabling controllable single- and multi-attribute protein editing via latent interpolation.
-
Steering Protein Language Models
Activation steering can guide protein language models to generate and optimize sequences with higher predicted thermostability, solubility, or GFP brightness, but only in surrogate-based evaluation.
-
Protriever: End-to-End Differentiable Protein Homology Search for Fitness Prediction
Protriever trains a retriever and a protein language model together so the model learns which homologs to retrieve, reaching state-of-the-art zero-shot fitness prediction on ProteinGym with much faster retrieval.
-
Steering Protein Family Design through Profile Bayesian Flow
ProfileBFN adapts Bayesian flow networks to accept protein-family profiles, enabling diverse, novel, and apparently functional family protein generation from single-sequence training.
-
Preference-based Antibody Expression Ranking: Scaling with Large-scale Weak Supervision
A two-stage recipe — camelid-sequence continual pretraining plus DPO-style preference fine-tuning with weak positive pairs — improves antibody expression ranking over supervised baselines on an internal 1254-sequence ...
-
Ankh3: Multi-Task Pretraining with Sequence Denoising and Completion Enhances Protein Representations
Ankh3 shows that combining multiple masking probabilities with sequence completion improves protein language model performance downstream, but the causal role of multi-task pretraining is not proven by a matched ablation.
-
Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances
Across the 68 papers it surveys, domain-specialized generative models usually outperform general-purpose LLMs on biological tasks, and agentic/conversational workflows are the least-covered topics.
-
AnnoDPO: Protein Functional Annotation Learning with Direct Preference Optimization
DPO with contrastive sequence-annotation alignment improves GO term prediction by 2 to 4 percent relative F1-Max over supervised fine-tuning alone.
-
Protein Inverse Folding From Structure Feedback
DPO fine-tuning with ESMFold TM-Score preferences raises sequence recovery and predicted TM-Score of inverse folding models, and multi-round refinement produces large gains on hard targets.
-
Towards More Accurate Full-Atom Antibody Co-Design
Igformer co-designs antibody CDR sequences and structures by adding personalized propagation, global attention, and dual equivariant message passing to dyMEAN, with modest benchmark improvements.
-
Leveraging Natural Language Processing to Unravel the Mystery of Life: A Review of NLP Approaches in Genomics, Transcriptomics, and Proteomics
A review maps how NLP architectures from word2vec to Evo 2 are applied to DNA, RNA, protein, and genome sequences, with tokenization and context choices shaping performance.
Discussion (0). Continue with ORCID to comment.