Pith. sign in

REVIEW 1 cited by

Pre-training via Denoising for Molecular Property Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.00133 v2 pith:7PNKTCIL submitted 2022-05-31 cs.LG q-bio.BMstat.ML

classification cs.LGq-bio.BMstat.ML
keywords moleculardenoisingpre-trainingpredictionpropertystructuresdatasetdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many important problems involving molecular property prediction from 3D structures have limited data, posing a generalization challenge for neural networks. In this paper, we describe a pre-training technique based on denoising that achieves a new state-of-the-art in molecular property prediction by utilizing large datasets of 3D molecular structures at equilibrium to learn meaningful representations for downstream tasks. Relying on the well-known link between denoising autoencoders and score-matching, we show that the denoising objective corresponds to learning a molecular force field -- arising from approximating the Boltzmann distribution with a mixture of Gaussians -- directly from equilibrium structures. Our experiments demonstrate that using this pre-training objective significantly improves performance on multiple benchmarks, achieving a new state-of-the-art on the majority of targets in the widely used QM9 dataset. Our analysis then provides practical insights into the effects of different factors -- dataset sizes, model size and architecture, and the choice of upstream and downstream datasets -- on pre-training.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 30 citations worldwide. Full citation record

  1. A Benchmark for Quantum Chemistry Relaxations via Machine Learning Interatomic Potentials

    q-bio.QM 2025-06 conditional novelty 6.0 of 10

    PubChemQCR is a large public dataset of DFT-based molecular relaxation trajectories with energy and force labels, benchmarked with nine machine learning interatomic potentials.

Pith tools