Pith. sign in

REVIEW 2 cited by

Large-Scale Chemical Language Representations Capture Molecular Structure and Properties

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.09553 v3 pith:VYD6A6A5 submitted 2021-06-17 cs.LG cs.CLq-bio.BM

classification cs.LGcs.CLq-bio.BM
keywords molecularlanguagemodelschemicallearningpropertiessupervisedattention
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Models based on machine learning can enable accurate and fast molecular property predictions, which is of interest in drug discovery and material design. Various supervised machine learning models have demonstrated promising performance, but the vast chemical space and the limited availability of property labels make supervised learning challenging. Recently, unsupervised transformer-based language models pretrained on a large unlabelled corpus have produced state-of-the-art results in many downstream natural language processing tasks. Inspired by this development, we present molecular embeddings obtained by training an efficient transformer encoder model, MoLFormer, which uses rotary positional embeddings. This model employs a linear attention mechanism, coupled with highly distributed training, on SMILES sequences of 1.1 billion unlabelled molecules from the PubChem and ZINC datasets. We show that the learned molecular representation outperforms existing baselines, including supervised and self-supervised graph neural networks and language models, on several downstream tasks from ten benchmark datasets. They perform competitively on two others. Further analyses, specifically through the lens of attention, demonstrate that MoLFormer trained on chemical SMILES indeed learns the spatial relationships between atoms within a molecule. These results provide encouraging evidence that large-scale molecular language models can capture sufficient chemical and structural information to predict various distinct molecular properties, including quantum-chemical properties.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NovoMolGen: Rethinking Molecular Language Model Pretraining

    cs.LG 2025-08 conditional novelty 6.0 of 10

    A 1.5-billion-molecule pretrained transformer family, NovoMolGen, sets new state-of-the-art results in de novo and goal-directed molecule generation, and shows pretraining loss correlates only weakly with downstream g...

  2. SynBridge: Bridging Reaction States via Discrete Flow for Bidirectional Reaction Prediction

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A bidirectional discrete flow matching model, SynBridge, predicts reaction products and reactants on graph representations and reports state-of-the-art Top-k accuracy on USPTO-50K, USPTO-MIT, and Pistachio.

Pith tools