REVIEW 17 cited by
Generative Flows on Discrete State-Spaces: Enabling Multimodal Flows with Applications to Protein Co-Design
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Combining discrete and continuous data is an important capability for generative models. We present Discrete Flow Models (DFMs), a new flow-based model of discrete data that provides the missing link in enabling flow-based generative models to be applied to multimodal continuous and discrete data problems. Our key insight is that the discrete equivalent of continuous space flow matching can be realized using Continuous Time Markov Chains. DFMs benefit from a simple derivation that includes discrete diffusion models as a specific instance while allowing improved performance over existing diffusion-based approaches. We utilize our DFMs method to build a multimodal flow-based modeling framework. We apply this capability to the task of protein co-design, wherein we learn a model for jointly generating protein structure and sequence. Our approach achieves state-of-the-art co-design performance while allowing the same multimodal model to be used for flexible generation of the sequence or structure.
Forward citations
Cited by 17 Pith papers
-
Hacking Generative Perplexity: Why Unconditional Text Evaluation Needs Distributional Metrics
Zero-parameter naive samplers achieve state-of-the-art generative perplexity while producing incoherent text, proving the metric is unsound; distributional divergences like MAUVE and energy distance correctly rank the...
-
Neuro-Symbolic ODE Discovery with Latent Grammar Flow
Latent Grammar Flow embeds grammar-based ODE representations into a discrete latent space with a behavioural loss and samples candidate equations via discrete flow to fit observed data.
-
Design-CP: Context Parallelism for Design of Protein Nanoparticles
Context-parallel inference for RFdiffusion 3 enables end-to-end all-atom design of large symmetric protein nanoparticles on multi-GPU hardware without retraining.
-
HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous Structure Tokens
HD-Prot shows that a protein language model can jointly generate sequences and structures using continuous structure tokens instead of quantized tokens, reaching competitive performance on four protein design tasks.
-
FLOWR.root: A flow matching based foundation model for joint multi-purpose structure-aware 3D ligand generation and affinity prediction
A flow-matching model jointly generates pocket-aware 3D ligands and predicts their binding affinities, reporting state-of-the-art generation and competitive affinity accuracy with a speed advantage.
-
Fine-Tuning Masked Diffusion for Provable Self-Correction
PRISM fine-tunes any masked diffusion model with a binary-cross-entropy loss so its new head provably estimates per-token quality p(x_i=y_i|y⊕m_i) and can remask low-quality tokens at inference.
-
Any-Order Flexible Length Masked Diffusion
FlexMDM is a discrete diffusion model that provably supports any-order generation over variable-length sequences by learning an insertion expectation alongside the unmasking posterior, validated by length-fidelity, ma...
-
Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
TarFlowLM models language in a continuous latent space with transformer-based autoregressive normalizing flows, using mixture-CDF and Rosenblatt couplings, and reports competitive NELBO on TEXT8 and OpenWebText.
-
Corrector Sampling in Language Models
A training and sampling method that lets autoregressive LLMs resample earlier tokens in a small window, improving reasoning and coding benchmark scores by about 10% relative after a 100B-token fine-tuning.
-
Accelerated Sampling from Masked Diffusion Models via Entropy Bounded Unmasking
EB-Sampler dynamically unmasks multiple low-entropy tokens per function evaluation, accelerating masked diffusion model sampling by 2-3x with negligible accuracy loss.
-
Applications of Modular Co-Design for De Novo 3D Molecule Generation
A new transformer architecture with joint continuous and discrete denoising improves 3D molecule generation and moves generated structures closer to low-energy physical minima.
-
Flow Matching for Collaborative Filtering
FlowCF applies flow matching with a behavior-guided prior and a discrete flow framework to collaborative filtering, achieving state-of-the-art top-N recommendation accuracy with two-step inference.
-
Provable diffusion-based posterior sampling for linear inverse problems via DDIM
A SVD-based, coordinate-wise DDIM sampler is claimed to asymptotically sample from the posterior for noisy linear inverse problems, but the proof's posterior identification step does not follow from the stated updates.
-
Energy-Based Flow Matching for Generating 3D Molecular Structure
IDFlow trains a flow matching network to refine its own predicted 3D molecular structure, improving docking and protein backbone generation over HarmonicFlow and FrameFlow baselines.
-
MolFORM: Multi-modal Flow Matching for Structure-Based Drug Design
A flow-matching model with direct preference optimization fine-tuning generates protein-binding molecules faster than diffusion baselines, with improved docking scores on the CrossDocked2020 benchmark.
-
TABASCO: A Fast, Simplified Model for Molecular Generation with Improved Physical Quality
TABASCO achieves 0.92 PoseBusters validity on GEOM-Drugs with a 59M-parameter non-equivariant transformer, no bond modeling, and post-hoc RDKit bond recovery, while sampling about 10x faster than SemlaFlow.
-
AffinityFlow: Guided Flows for Antibody Affinity Maturation
AffinityFlow guides AlphaFlow structure generation toward low Rosetta binding energy, then inverse-folds the structures to propose antibody mutations, and reports top scores on a computational affinity maturation benchmark.
Discussion (0). Continue with ORCID to comment.