Pith. sign in

REVIEW 9 cited by

Diffusion Language Models Are Versatile Protein Learners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.18567 v2 pith:E26YW6YX submitted 2024-02-28 cs.LG q-bio.BM

classification cs.LGq-bio.BM
keywords proteindplmdiffusiongenerationlanguagesequencesgenerativemodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces diffusion protein language model (DPLM), a versatile protein language model that demonstrates strong generative and predictive capabilities for protein sequences. We first pre-train scalable DPLMs from evolutionary-scale protein sequences within a generative self-supervised discrete diffusion probabilistic framework, which generalizes language modeling for proteins in a principled way. After pre-training, DPLM exhibits the ability to generate structurally plausible, novel, and diverse protein sequences for unconditional generation. We further demonstrate the proposed diffusion generative pre-training makes DPLM possess a better understanding of proteins, making it a superior representation learner, which can be fine-tuned for various predictive tasks, comparing favorably to ESM2 (Lin et al., 2022). Moreover, DPLM can be tailored for various needs, which showcases its prowess of conditional generation in several ways: (1) conditioning on partial peptide sequences, e.g., generating scaffolds for functional motifs with high success rate; (2) incorporating other modalities as conditioner, e.g., structure-conditioned generation for inverse folding; and (3) steering sequence generation towards desired properties, e.g., satisfying specified secondary structures, through a plug-and-play classifier guidance. Code is released at \url{https://github.com/bytedance/dplm}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 18 citations worldwide. Full citation record

  1. La-Proteina: Atomistic Protein Generation via Partially Latent Flow Matching

    cs.LG 2025-07 conditional novelty 7.0 of 10

    La-Proteina generates full-atom protein structures and sequences via flow matching over an explicit alpha-carbon backbone plus fixed-size per-residue latents, achieving state-of-the-art co-designability and scaling to...

  2. Variable-Length Generative Protein Design via Generalized Poisson Flow

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Generalized Poisson Flow learns variable protein length via an inhomogeneous Poisson rate plus within-length flow matching, with KL bounds and gains on structure, sequence, motif, and peptide tasks.

  3. STRIDE: Post-Training LLMs to Reason and Refine Bio-Sequences via Edit Trajectories

    cs.CE 2026-03 conditional novelty 6.0 of 10

    Training LLMs to emit executable edit trajectories (INSERT/DELETE/REPLACE) from Levenshtein alignments plus policy optimization improves oracle-scored bio-sequence optimization success and novelty.

  4. HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous Structure Tokens

    cs.CE 2025-12 conditional novelty 6.0 of 10

    HD-Prot shows that a protein language model can jointly generate sequences and structures using continuous structure tokens instead of quantized tokens, reaching competitive performance on four protein design tasks.

  5. Steering Protein Family Design through Profile Bayesian Flow

    q-bio.BM 2025-02 conditional novelty 6.0 of 10

    ProfileBFN adapts Bayesian flow networks to accept protein-family profiles, enabling diverse, novel, and apparently functional family protein generation from single-sequence training.

  6. On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization

    cs.LG 2026-07 conditional novelty 5.5 of 10

    Online fine-tuning of discrete diffusion models with complementary acquisition, CVaR shaping, density-entropy debiasing, replay, and validity control finds better molecules under fixed oracle budgets than offline fine...

  7. AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model

    q-bio.BM 2025-07 conditional novelty 5.0 of 10

    AMix-1, a 1.7B-parameter Bayesian Flow Network protein model conditioned on MSA profiles and refined by an evolutionary test-time scaling loop, produced AmeR variants with up to 50x wild-type activity in wet-lab tests.

  8. Diffusion Sequence Models for Enhanced Protein Representation and Generation

    q-bio.BM 2025-06 conditional novelty 5.0 of 10

    Masked diffusion retrofitted onto ESM2 produces a pLM that matches representation benchmarks and generates protein-like sequences, with an in-silico binder design case study.

  9. VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL

    cs.CV 2025-05 conditional novelty 5.0 of 10

    VARD fine-tunes diffusion models by backpropagating through a learned value function that assigns dense, differentiable reward estimates to every intermediate denoising step, with KL regularization keeping the model n...

Pith tools