Pith. sign in

REVIEW 1 cited by

Data-Efficient Molecular Generation with Hierarchical Textual Inversion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.02845 v3 pith:AJNVT6WV submitted 2024-05-05 cs.LG q-bio.MN

classification cs.LGq-bio.MN
keywords moleculargenerationhi-molhierarchicalinversiontextualdata-efficientembeddings
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Developing an effective molecular generation framework even with a limited number of molecules is often important for its practical deployment, e.g., drug discovery, since acquiring task-related molecular data requires expensive and time-consuming experimental costs. To tackle this issue, we introduce Hierarchical textual Inversion for Molecular generation (HI-Mol), a novel data-efficient molecular generation method. HI-Mol is inspired by the importance of hierarchical information, e.g., both coarse- and fine-grained features, in understanding the molecule distribution. We propose to use multi-level embeddings to reflect such hierarchical features based on the adoption of the recent textual inversion technique in the visual domain, which achieves data-efficient image generation. Compared to the conventional textual inversion method in the image domain using a single-level token embedding, our multi-level token embeddings allow the model to effectively learn the underlying low-shot molecule distribution. We then generate molecules based on the interpolation of the multi-level token embeddings. Extensive experiments demonstrate the superiority of HI-Mol with notable data-efficiency. For instance, on QM9, HI-Mol outperforms the prior state-of-the-art method with 50x less training data. We also show the effectiveness of molecules generated by HI-Mol in low-shot molecular property prediction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Graph-based Molecular In-context Learning Grounded on Morgan Fingerprints

    cs.LG 2025-02 conditional novelty 6.0 of 10

    GAMIC aligns graph embeddings of molecules with scientific text via contrastive learning, then applies diversity-aware retrieval, improving molecular in-context learning over Morgan fingerprint baselines.

Pith tools