Pith. sign in

REVIEW 3 cited by

Mixed Continuous and Categorical Flow Matching for 3D De Novo Molecule Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.19739 v1 pith:G4S4AXTI submitted 2024-04-30 q-bio.BM cs.LG

classification q-bio.BMcs.LG
keywords flowmatchingcategoricalgenerationmodelsmoleculenovodata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep generative models that produce novel molecular structures have the potential to facilitate chemical discovery. Diffusion models currently achieve state of the art performance for 3D molecule generation. In this work, we explore the use of flow matching, a recently proposed generative modeling framework that generalizes diffusion models, for the task of de novo molecule generation. Flow matching provides flexibility in model design; however, the framework is predicated on the assumption of continuously-valued data. 3D de novo molecule generation requires jointly sampling continuous and categorical variables such as atom position and atom type. We extend the flow matching framework to categorical data by constructing flows that are constrained to exist on a continuous representation of categorical data known as the probability simplex. We call this extension SimplexFlow. We explore the use of SimplexFlow for de novo molecule generation. However, we find that, in practice, a simpler approach that makes no accommodations for the categorical nature of the data yields equivalent or superior performance. As a result of these experiments, we present FlowMol, a flow matching model for 3D de novo generative model that achieves improved performance over prior flow matching methods, and we raise important questions about the design of prior distributions for achieving strong performance in flow matching models. Code and trained models for reproducing this work are available at https://github.com/dunni3/FlowMol

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Boltzmann-Expected Molecular Design with Decoupled Annealing Flows

    stat.ML 2026-07 conditional novelty 7.0 of 10

    A two-flow simulated-annealing loop makes ensemble statistics—means, variances, and skewness of 3D properties—the objective of molecular graph design.

  2. TABASCO: A Fast, Simplified Model for Molecular Generation with Improved Physical Quality

    cs.LG 2025-07 conditional novelty 5.0 of 10

    TABASCO achieves 0.92 PoseBusters validity on GEOM-Drugs with a 59M-parameter non-equivariant transformer, no bond modeling, and post-hoc RDKit bond recovery, while sampling about 10x faster than SemlaFlow.

  3. Predictive Feature Caching for Training-free Acceleration of Molecular Geometry Generation

    cs.LG 2025-10 conditional novelty 4.0 of 10

    Predictive feature caching, borrowed from image diffusion, speeds up molecular flow-matching generation by 2-3x at near-matched quality by forecasting hidden features instead of recomputing them.

Pith tools