REVIEW 8 cited by
Diffusion Language Models Are Versatile Protein Learners
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper introduces diffusion protein language model (DPLM), a versatile protein language model that demonstrates strong generative and predictive capabilities for protein sequences. We first pre-train scalable DPLMs from evolutionary-scale protein sequences within a generative self-supervised discrete diffusion probabilistic framework, which generalizes language modeling for proteins in a principled way. After pre-training, DPLM exhibits the ability to generate structurally plausible, novel, and diverse protein sequences for unconditional generation. We further demonstrate the proposed diffusion generative pre-training makes DPLM possess a better understanding of proteins, making it a superior representation learner, which can be fine-tuned for various predictive tasks, comparing favorably to ESM2 (Lin et al., 2022). Moreover, DPLM can be tailored for various needs, which showcases its prowess of conditional generation in several ways: (1) conditioning on partial peptide sequences, e.g., generating scaffolds for functional motifs with high success rate; (2) incorporating other modalities as conditioner, e.g., structure-conditioned generation for inverse folding; and (3) steering sequence generation towards desired properties, e.g., satisfying specified secondary structures, through a plug-and-play classifier guidance. Code is released at \url{https://github.com/bytedance/dplm}.
Forward citations
Cited by 8 Pith papers
-
La-Proteina: Atomistic Protein Generation via Partially Latent Flow Matching
La-Proteina generates full-atom protein structures and sequences via flow matching over an explicit alpha-carbon backbone plus fixed-size per-residue latents, achieving state-of-the-art co-designability and scaling to...
-
Variable-Length Generative Protein Design via Generalized Poisson Flow
Generalized Poisson Flow learns variable protein length via an inhomogeneous Poisson rate plus within-length flow matching, with KL bounds and gains on structure, sequence, motif, and peptide tasks.
-
STRIDE: Post-Training LLMs to Reason and Refine Bio-Sequences via Edit Trajectories
Training LLMs to emit executable edit trajectories (INSERT/DELETE/REPLACE) from Levenshtein alignments plus policy optimization improves oracle-scored bio-sequence optimization success and novelty.
-
HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous Structure Tokens
HD-Prot shows that a protein language model can jointly generate sequences and structures using continuous structure tokens instead of quantized tokens, reaching competitive performance on four protein design tasks.
-
On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization
Online fine-tuning of discrete diffusion models with complementary acquisition, CVaR shaping, density-entropy debiasing, replay, and validity control finds better molecules under fixed oracle budgets than offline fine...
-
AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model
AMix-1, a 1.7B-parameter Bayesian Flow Network protein model conditioned on MSA profiles and refined by an evolutionary test-time scaling loop, produced AmeR variants with up to 50x wild-type activity in wet-lab tests.
-
Diffusion Sequence Models for Enhanced Protein Representation and Generation
Masked diffusion retrofitted onto ESM2 produces a pLM that matches representation benchmarks and generates protein-like sequences, with an in-silico binder design case study.
-
VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL
VARD fine-tunes diffusion models by backpropagating through a learned value function that assigns dense, differentiable reward estimates to every intermediate denoising step, with KL regularization keeping the model n...
Discussion (0). Continue with ORCID to comment.