La-proteina: Atomistic protein generation via partially latent flow matching.arXiv preprint arXiv:2507.09466

Tomas Geffner, Kieran Didi, Zhonglin Cao, Danny Reidenbach, Zuobai Zhang, Christian Dallago, Emine Kucukbenli, Karsten Kreis, Arash Vahdat · 2025 · cs.LG · arXiv 2507.09466

7 Pith papers cite this work. Polarity classification is still indexing.

7 Pith papers citing it

open full Pith review browse 7 citing papers arXiv PDF

abstract

Recently, many generative models for de novo protein structure design have emerged. Yet, only few tackle the difficult task of directly generating fully atomistic structures jointly with the underlying amino acid sequence. This is challenging, for instance, because the model must reason over side chains that change in length during generation. We introduce La-Proteina for atomistic protein design based on a novel partially latent protein representation: coarse backbone structure is modeled explicitly, while sequence and atomistic details are captured via per-residue latent variables of fixed dimensionality, thereby effectively side-stepping challenges of explicit side-chain representations. Flow matching in this partially latent space then models the joint distribution over sequences and full-atom structures. La-Proteina achieves state-of-the-art performance on multiple generation benchmarks, including all-atom co-designability, diversity, and structural validity, as confirmed through detailed structural analyses and evaluations. Notably, La-Proteina also surpasses previous models in atomistic motif scaffolding performance, unlocking critical atomistic structure-conditioned protein design tasks. Moreover, La-Proteina is able to generate co-designable proteins of up to 800 residues, a regime where most baselines collapse and fail to produce valid samples, demonstrating La-Proteina's scalability and robustness.

citation-role summary

background 2

citation-polarity summary

background 2

representative citing papers

A-CODE: Fully Atomic Protein Co-Design with Unified Multimodal Diffusion

q-bio.QM · 2026-05-05 · unverdicted · novelty 8.0

A-CODE presents a fully atomic one-stage multimodal diffusion model for protein co-design that claims superior unconditional generation performance over prior one- and two-stage models plus a tenfold success-rate gain on hard binder-design tasks.

Steerable Neural ODEs on Homogeneous Spaces

cs.LG · 2026-05-11 · unverdicted · novelty 7.0

Steerable NODEs extend manifold neural ODEs by coupling base flow on homogeneous spaces with parallel transport of features in associated bundles, achieving G-equivariance under invariant conditions.

General Multimodal Protein Design Enables DNA-Encoding of Chemistry

cs.LG · 2026-04-06 · conditional · novelty 7.0

DISCO co-designs protein sequence and structure to produce functional heme enzymes that catalyze several new-to-nature carbene-transfer reactions at activities exceeding prior engineered enzymes.

Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner

cs.AI · 2025-10-03 · unverdicted · novelty 7.0

CCDD defines a joint multimodal diffusion on continuous representation space and discrete token space to combine expressivity with explicit token supervision for diffusion language models.

Navigating committor landscape of biomolecules with a general pairwise interaction model

physics.comp-ph · 2026-06-30 · unverdicted · novelty 6.0 · 2 refs

A novel neural architecture based on Pairformer is introduced for learning committor functions to better capture dynamical features in biomolecular rare events without specialized priors.

Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation

q-bio.BM · 2026-05-11 · unverdicted · novelty 6.0

Yeti is a compact tokenizer for protein structures that delivers strong codebook use, token diversity, and reconstruction while enabling from-scratch multimodal generation of plausible sequences and structures with 10x fewer parameters than ESM3.

From Words to Amino Acids: Does the Curse of Depth Persist?

cs.LG · 2026-02-25 · unverdicted · novelty 6.0

Protein language models exhibit consistent depth inefficiency where most task-relevant computation occurs in a subset of layers, mirroring patterns in large language models.

citing papers explorer

Showing 7 of 7 citing papers.

A-CODE: Fully Atomic Protein Co-Design with Unified Multimodal Diffusion q-bio.QM · 2026-05-05 · unverdicted · none · ref 15 · internal anchor
A-CODE presents a fully atomic one-stage multimodal diffusion model for protein co-design that claims superior unconditional generation performance over prior one- and two-stage models plus a tenfold success-rate gain on hard binder-design tasks.
Steerable Neural ODEs on Homogeneous Spaces cs.LG · 2026-05-11 · unverdicted · none · ref 5 · internal anchor
Steerable NODEs extend manifold neural ODEs by coupling base flow on homogeneous spaces with parallel transport of features in associated bundles, achieving G-equivariance under invariant conditions.
General Multimodal Protein Design Enables DNA-Encoding of Chemistry cs.LG · 2026-04-06 · conditional · none · ref 48 · internal anchor
DISCO co-designs protein sequence and structure to produce functional heme enzymes that catalyze several new-to-nature carbene-transfer reactions at activities exceeding prior engineered enzymes.
Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner cs.AI · 2025-10-03 · unverdicted · none · ref 13 · internal anchor
CCDD defines a joint multimodal diffusion on continuous representation space and discrete token space to combine expressivity with explicit token supervision for diffusion language models.
Navigating committor landscape of biomolecules with a general pairwise interaction model physics.comp-ph · 2026-06-30 · unverdicted · none · ref 36 · 2 links · internal anchor
A novel neural architecture based on Pairformer is introduced for learning committor functions to better capture dynamical features in biomolecular rare events without specialized priors.
Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation q-bio.BM · 2026-05-11 · unverdicted · none · ref 49 · internal anchor
Yeti is a compact tokenizer for protein structures that delivers strong codebook use, token diversity, and reconstruction while enabling from-scratch multimodal generation of plausible sequences and structures with 10x fewer parameters than ESM3.
From Words to Amino Acids: Does the Curse of Depth Persist? cs.LG · 2026-02-25 · unverdicted · none · ref 11 · internal anchor
Protein language models exhibit consistent depth inefficiency where most task-relevant computation occurs in a subset of layers, mirroring patterns in large language models.

La-proteina: Atomistic protein generation via partially latent flow matching.arXiv preprint arXiv:2507.09466

citation-role summary

citation-polarity summary

fields

years

verdicts

roles

polarities

representative citing papers

citing papers explorer