Pith. sign in

REVIEW 11 cited by

ChemBERTa-2: Towards Chemical Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.01712 v1 pith:XFETXKGV submitted 2022-09-05 cs.LG cs.AIq-bio.BM

classification cs.LGcs.AIq-bio.BM
keywords pretrainingmoleculartaskschemberta-2chemicaldownstreamfoundationimprovements
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large pretrained models such as GPT-3 have had tremendous impact on modern natural language processing by leveraging self-supervised learning to learn salient representations that can be used to readily finetune on a wide variety of downstream tasks. We investigate the possibility of transferring such advances to molecular machine learning by building a chemical foundation model, ChemBERTa-2, using the language of SMILES. While labeled data for molecular prediction tasks is typically scarce, libraries of SMILES strings are readily available. In this work, we build upon ChemBERTa by optimizing the pretraining process. We compare multi-task and self-supervised pretraining by varying hyperparameters and pretraining dataset size, up to 77M compounds from PubChem. To our knowledge, the 77M set constitutes one of the largest datasets used for molecular pretraining to date. We find that with these pretraining improvements, we are competitive with existing state-of-the-art architectures on the MoleculeNet benchmark suite. We analyze the degree to which improvements in pretraining translate to improvement on downstream tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Ray Tracing Sampler: Bayesian Sampling of Neural Networks for Everyone

    astro-ph.IM 2025-10 conditional novelty 7.0 of 10

    A new ray-tracing MCMC sampler keeps ray speed constant, making it far more robust to stochastic gradients and able to sample billion-parameter neural networks on one GPU.

  2. A Foundation Model for Material Fracture Prediction

    cs.LG 2025-07 conditional novelty 7.0 of 10

    A transformer pretrained on rule-based fracture data and phase-field simulations predicts final fracture patterns across five materials and two loading conditions, and fine-tunes to new materials, time-to-failure, and...

  3. MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A metabolomics-specialized LLM that generates biochemical descriptions, converted into a metabolite graph, improved patient-level metabolomics classification over standard baselines in two datasets.

  4. One Mask to Rule Them All: On Hidden Facts after Editing and How to Find Them

    cs.LG 2026-04 conditional novelty 6.0 of 10

    ROME and MEMIT knowledge edits share a common weight subset isolable by a compact binary mask that reverses ~70–80% of edits and is necessary for editing success.

  5. MultiPUFFIN: A Multimodal Domain-Constrained Foundation Model for Molecular Property Prediction of Small Molecules

    cs.LG 2026-03 conditional novelty 6.0 of 10

    MultiPUFFIN claims higher test R² than ChemBERTa-2 on all nine thermophysical properties while using far fewer labeled molecules, with the largest gains on temperature-dependent properties.

  6. Structure-Preserving Learning Improves Geometry Generalization in Neural PDEs

    cs.LG 2026-02 conditional novelty 6.0 of 10

    A geometry-conditioned Whitney-form neural network that solves a learned discrete conservation law improves out-of-distribution geometry generalization for steady-state PDEs compared with regression-based neural operators.

  7. MHA-RAG: Improving Efficiency, Accuracy, and Consistency by Encoding Exemplars as Soft Prompts

    cs.AI 2025-10 conditional novelty 6.0 of 10

    MHA-RAG encodes retrieved exemplars into order-invariant soft prompts via multi-head attention, claiming ~20-point effective-accuracy gains over RAG at ~10x lower inference FLOPs.

  8. Predicting and generating antibiotics against future pathogens with ApexOracle

    cs.LG 2025-07 reject novelty 6.0 of 10

    ApexOracle fuses genomic and literature-derived pathogen embeddings with a diffusion language model to predict antimicrobial activity and generate new candidate molecules for unseen bacterial strains.

  9. DeepRetro: Retrosynthetic Pathway Discovery using Iterative LLM Reasoning

    q-bio.QM 2025-07 conditional novelty 6.0 of 10

    DeepRetro combines LLM-generated retrosynthetic disconnections with template-based search and human feedback, achieving strong benchmark results and proposing new routes for complex natural products.

  10. Leveraging neural network interatomic potentials for a foundation model of chemistry

    cond-mat.mtrl-sci 2025-06 conditional novelty 5.0 of 10

    Using embeddings from a pretrained neural network interatomic potential as features for small machine learning models gives competitive or better property predictions than end-to-end deep networks, especially with lim...

  11. Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation

    cs.AI 2025-08 reject novelty 2.0 of 10

    This survey claims to be the first systematic review of LLMs for organic synthesis, but its central 'evaluation' is never actually performed.

Pith tools