Pith. sign in

REVIEW 3 cited by

Flow Network based Generative Models for Non-Iterative Diverse Candidate Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.04399 v2 pith:DZHXP5TP submitted 2021-06-08 cs.LG

classification cs.LG
keywords flowgenerativediversefunctionlearningobjectpolicythere
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper is about the problem of learning a stochastic policy for generating an object (like a molecular graph) from a sequence of actions, such that the probability of generating an object is proportional to a given positive reward for that object. Whereas standard return maximization tends to converge to a single return-maximizing sequence, there are cases where we would like to sample a diverse set of high-return solutions. These arise, for example, in black-box function optimization when few rounds are possible, each with large batches of queries, where the batches should be diverse, e.g., in the design of new molecules. One can also see this as a problem of approximately converting an energy function to a generative distribution. While MCMC methods can achieve that, they are expensive and generally only perform local exploration. Instead, training a generative policy amortizes the cost of search during training and yields to fast generation. Using insights from Temporal Difference learning, we propose GFlowNet, based on a view of the generative process as a flow network, making it possible to handle the tricky case where different trajectories can yield the same final state, e.g., there are many ways to sequentially add atoms to generate some molecular graph. We cast the set of trajectories as a flow and convert the flow consistency equations into a learning objective, akin to the casting of the Bellman equations into Temporal Difference methods. We prove that any global minimum of the proposed objectives yields a policy which samples from the desired distribution, and demonstrate the improved performance and diversity of GFlowNet on a simple domain where there are many modes to the reward function, and on a molecule synthesis task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 27 citations worldwide. Full citation record

  1. Machine learning for sample-based quantum diagonalization: generative configuration recovery and the classical-simulability frontier

    quant-ph 2026-08 conditional novelty 6.0 of 10

    A critical review plus small exact-FCI experiments concludes that sample-based quantum diagonalization has not beaten classical selected CI and maps where, if anywhere, a quantum or generative advantage could survive.

  2. Minimally dissipative multi-bit logical operations

    cond-mat.stat-mech 2025-06 conditional novelty 6.0 of 10

    Multi-bit logical gates are formulated as regularized unbalanced optimal transport problems, yielding finite-time minimal-work bounds, a no-advantage theorem for complete erasure, and near-optimal 2D control protocols.

  3. Importance Weighted Score Matching for Diffusion Samplers with Enhanced Mode Coverage

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Importance Weighted Score Matching trains diffusion samplers by reweighting score matching with self-normalized importance sampling to approximate the forward KL and improve mode coverage.

Pith tools