REVIEW 10 cited by
Proteina: Scaling Flow-based Protein Structure Generative Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recently, diffusion- and flow-based generative models of protein structures have emerged as a powerful tool for de novo protein design. Here, we develop Proteina, a new large-scale flow-based protein backbone generator that utilizes hierarchical fold class labels for conditioning and relies on a tailored scalable transformer architecture with up to 5x as many parameters as previous models. To meaningfully quantify performance, we introduce a new set of metrics that directly measure the distributional similarity of generated proteins with reference sets, complementing existing metrics. We further explore scaling training data to millions of synthetic protein structures and explore improved training and sampling recipes adapted to protein backbone generation. This includes fine-tuning strategies like LoRA for protein backbones, new guidance methods like classifier-free guidance and autoguidance for protein backbones, and new adjusted training objectives. Proteina achieves state-of-the-art performance on de novo protein backbone design and produces diverse and designable proteins at unprecedented length, up to 800 residues. The hierarchical conditioning offers novel control, enabling high-level secondary-structure guidance as well as low-level fold-specific generation.
Forward citations
Cited by 10 Pith papers
-
Vilya-1: An all-atom foundation model for macrocycle structure prediction and design
Vilya-1, a unified all-atom diffusion model, nearly doubles success rates for sampling near-native macrocycle ring conformations versus physics and co-folding baselines and transfers to property prediction and topolog...
-
Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling
I3CD plus MoPS sampling produces single binder sequences that AlphaFold-Multimer scores as compatible with multiple conformational or multi-target contexts on the CROSS benchmark.
-
Spectral Diffusion for Protein Dynamics
Diffusion over DCT spectral volumes of Cα displacements yields fast, temperature-conditioned protein trajectories with RMSF Pearson r of 0.844 on held-out mdCATH.
-
Accurate structural modeling of chemically diverse molecular interfaces with Vilya-2
Vilya-2 predicts bound structures of chemically diverse peptides and small molecules at state-of-the-art accuracy using an all-atom diffusion transformer, recovering 59.1% of peptide interfaces to sub-2 Å backbone RMSD.
-
Variable-Length Generative Protein Design via Generalized Poisson Flow
Generalized Poisson Flow learns variable protein length via an inhomogeneous Poisson rate plus within-length flow matching, with KL bounds and gains on structure, sequence, motif, and peptide tasks.
-
Design-CP: Context Parallelism for Design of Protein Nanoparticles
Context-parallel inference for RFdiffusion 3 enables end-to-end all-atom design of large symmetric protein nanoparticles on multi-GPU hardware without retraining.
-
Efficient Molecular Conformer Generation with SO(3)-Averaged Flow Matching and Reflow
SO(3)-Averaged Flow matching with reflow and distillation enables high-quality one-step molecular conformer generation, reporting new SOTA on GEOM-QM9 and strong one-step results on GEOM-Drugs.
-
Adaptive Multimodal Protein Plug-and-Play with Diffusion-Based Priors
Adam-PnP guides a pre-trained protein diffusion model with multiple experimental data types, using online noise estimation and precision-based weighting, and reports a backbone RMSD of 0.65 Å on one protein.
-
MolFORM: Multi-modal Flow Matching for Structure-Based Drug Design
A flow-matching model with direct preference optimization fine-tuning generates protein-binding molecules faster than diffusion baselines, with improved docking scores on the CrossDocked2020 benchmark.
-
TABASCO: A Fast, Simplified Model for Molecular Generation with Improved Physical Quality
TABASCO achieves 0.92 PoseBusters validity on GEOM-Drugs with a 59M-parameter non-equivariant transformer, no bond modeling, and post-hoc RDKit bond recovery, while sampling about 10x faster than SemlaFlow.
Discussion (0). Continue with ORCID to comment.