Pith. sign in

REVIEW 25 cited by

ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.09885 v2 pith:3DVZQ4KA submitted 2020-10-19 cs.LG cs.CLphysics.chem-phq-bio.BM

classification cs.LGcs.CLphysics.chem-phq-bio.BM
keywords predictionpropertytransformerschembertamolecularpretrainingdatasetdownstream
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

GNNs and chemical fingerprints are the predominant approaches to representing molecules for property prediction. However, in NLP, transformers have become the de-facto standard for representation learning thanks to their strong downstream task transfer. In parallel, the software ecosystem around transformers is maturing rapidly, with libraries like HuggingFace and BertViz enabling streamlined training and introspection. In this work, we make one of the first attempts to systematically evaluate transformers on molecular property prediction tasks via our ChemBERTa model. ChemBERTa scales well with pretraining dataset size, offering competitive downstream performance on MoleculeNet and useful attention-based visualization modalities. Our results suggest that transformers offer a promising avenue of future work for molecular representation learning and property prediction. To facilitate these efforts, we release a curated dataset of 77M SMILES from PubChem suitable for large-scale self-supervised pretraining.

Discussion (0). Sign in to comment.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 397 citations worldwide. Full citation record

  1. A Foundation Model for Material Fracture Prediction

    cs.LG 2025-07 conditional novelty 7.0 of 10

    A transformer pretrained on rule-based fracture data and phase-field simulations predicts final fracture patterns across five materials and two loading conditions, and fine-tunes to new materials, time-to-failure, and...

  2. ED-DiT: Physics-Guided Diffusion Pretraining for Transferable Molecular Representations from Electron Density

    cs.LG 2026-08 conditional novelty 6.0 of 10

    ED-DiT pretrains a diffusion transformer on electron-density point clouds with a physical electron-number constraint, and the resulting encoder outperforms scratch models across six molecular tasks.

  3. MS-GPT: Rethinking MS/MS De Novo Structure Elucidation as Spectrum-Induced Posterior Querying of a Molecule-Language Model

    cs.LG 2026-07 accept novelty 6.0 of 10

    Querying a fingerprint-conditioned molecule-language model with a calibrated band of spectrum-induced posteriors beats point-threshold fingerprint decoding on NPLIB1 and MassSpecGym.

  4. Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction

    q-bio.QM 2026-07 conditional novelty 6.0 of 10

    Generic molecular encoders show weak alignment with human olfactory rating geometry and no clear predictive increment over an RDKit-Morgan chemistry baseline, while human rating geometry is reproducible within but onl...

  5. OLEDLM: A Unified Language Model for OLED Molecular Design

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A GRPO-aligned LLaMA model generates OLED SMILES conditioned on target S1 and f, with the property-alignment evidence largely coming from the very BERT predictor used as reward.

  6. ChemHyperMag: Physics-informed magnetic hypergraph learning improves molecular ADMET prediction

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A magnetic hypergraph neural network with chemistry-inspired directional flow improves ADMET property prediction on several benchmarks.

  7. Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Augmenting SLM prompts with a GNN expert's prediction, confidence, and an extracted important subgraph improves zero-shot toxicity/mutagenicity accuracy on MUTAG and Tox21, with gains up to 74% relative to SMILES-only...

  8. SIGMA: Semantic Identifier Grouping for Molecular Autoregression

    cs.LG 2026-03 reject novelty 6.0 of 10

    A same-suffix contrastive objective makes autoregressive molecular-string models more invariant to how a molecule is written, improving generation fidelity on the reported ZINC benchmark.

  9. Towards a Physics Foundation Model

    cs.LG 2025-09 conditional novelty 6.0 of 10

    A single transformer-based model, GPhyT, trained on diverse 2D simulation data, predicts next states across several fluid and heat-transfer systems and extrapolates to similar unseen regimes with plausible results.

  10. MODA: A Unified 3D Diffusion Framework for Multi-Task Target-Aware Molecular Generation

    q-bio.BM 2025-07 conditional novelty 6.0 of 10

    A single masked-diffusion model trained jointly on four molecular-editing tasks outperforms or matches task-specific diffusion baselines across docking, chemical property, and geometry metrics.

  11. DeepRetro: Retrosynthetic Pathway Discovery using Iterative LLM Reasoning

    q-bio.QM 2025-07 conditional novelty 6.0 of 10

    DeepRetro combines LLM-generated retrosynthetic disconnections with template-based search and human feedback, achieving strong benchmark results and proposing new routes for complex natural products.

  12. OmniESI: A unified framework for enzyme-substrate interaction prediction with progressive conditional deep learning

    q-bio.BM 2025-06 conditional novelty 6.0 of 10

    A two-stage conditional deep learning framework modulates enzyme and substrate embeddings toward catalysis-aware features, reporting better benchmark results than specialized enzyme-substrate predictors.

  13. Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Chem World unifies 17 chemical mixture datasets into 10 property tracks, and Mixture-PINN’s soft physics regularizers beat common encoder–aggregator baselines on most tracks and OOD splits.

  14. Trustworthy Protein-Ligand Binding Affinity Prediction via Reliability-Aware Multi-Engine Fusion

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A reliability-aware evidential fusion model combines four docking engines and produces well-calibrated confidence scores that enable selective prediction with up to 25.7% error reduction.

  15. A Quiet Failure in Calibrated Virtual Screening: Marginal Conformal Prediction Under-Covers the Minority Class, and a Class-Conditional Fix Recovers It

    cs.LG 2026-07 accept novelty 5.0 of 10

    Marginal conformal prediction under-covers minority classes by tens of points on imbalanced molecular datasets while hitting global coverage; Mondrian calibration recovers the target.

  16. Valid Property-Enhanced Contrastive Learning for Targeted Optimization & Resampling for Novel Drug Design

    cs.LG 2025-08 conditional novelty 5.0 of 10

    VECTOR+ combines contrastive learning and Gaussian mixture sampling to generate novel, synthetically plausible inhibitors from low-data datasets, with improved docking scores over known compounds.

  17. Chunk Twice, Embed Once: A Systematic Study of Segmentation and Representation Trade-offs in Chemistry-Aware Retrieval-Augmented Generation

    cs.IR 2025-06 conditional novelty 5.0 of 10

    A systematic evaluation shows that recursive 100-token non-overlapping chunks and retrieval-tuned embeddings outperform fixed-size chunks and domain-specific models like SciBERT for chemistry retrieval, and it introdu...

  18. Uncovering Bottlenecks and Optimizing Scientific Lab Workflows with Cycle Time Reduction Agents

    cs.MA 2025-05 conditional novelty 5.0 of 10

    CTRA is a three-component LangGraph agent system for automatically generating analytical questions, SQL, and insights to identify bottlenecks in scientific lab workflows.

  19. Persistent Manifold Learning of Protein Properties

    q-bio.BM 2026-07 conditional novelty 4.0 of 10

    Persistent manifold learning with Boundary-Induced Graph Laplacians plus language-model features beats prior SOTA Pearson correlation on metalloprotein–ligand and SKEMPI wild-type protein–protein affinity benchmarks.

  20. Multimodal Molecular Representation Learning with Graph Neural Networks, Deep & Cross Networks, and SMILES Embeddings

    cs.LG 2026-07 conditional novelty 4.0 of 10

    Tri-modal late fusion of SchNet geometry, ChemBERTa SMILES, and DCN descriptors reaches 0.0207 eV MAE on QM9 U0 atomization energy, a 20.6% gain over a controlled SchNet baseline under 1M parameters.

  21. Adaptive Minds: Empowering Agents with LoRA-as-Tools

    cs.AI 2025-10 reject novelty 4.0 of 10

    Adaptive Minds makes a base LLM select LoRA adapters as tools per query; the 5-adapter demo gets 100% routing on 25 queries, while the abstract's 30-adapter/nine-family numbers are unsupported.

  22. A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools

    cs.LG 2025-06 unverdicted novelty 4.0 of 10

    This survey organizes foundation models, LLM agents, datasets, and tools in materials science into six task areas.

  23. Predicting Drug-Drug Interactions Using Heterogeneous Graph Neural Networks: HGNN-DDI

    cs.LG 2025-08 reject novelty 3.0 of 10

    HGNN-DDI claims state-of-the-art drug-drug interaction prediction, but its skewed test set and missing SOTA baselines make the performance claim unsupported.

  24. Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation

    cs.AI 2025-08 reject novelty 2.0 of 10

    This survey claims to be the first systematic review of LLMs for organic synthesis, but its central 'evaluation' is never actually performed.

  25. Transformers in Protein: A Survey

    cs.LG 2025-05 unverdicted

    A broad but unreliable survey of Transformer applications in protein informatics, with numerous citation errors and unsupported claims.

Pith tools