REVIEW 25 cited by
ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
GNNs and chemical fingerprints are the predominant approaches to representing molecules for property prediction. However, in NLP, transformers have become the de-facto standard for representation learning thanks to their strong downstream task transfer. In parallel, the software ecosystem around transformers is maturing rapidly, with libraries like HuggingFace and BertViz enabling streamlined training and introspection. In this work, we make one of the first attempts to systematically evaluate transformers on molecular property prediction tasks via our ChemBERTa model. ChemBERTa scales well with pretraining dataset size, offering competitive downstream performance on MoleculeNet and useful attention-based visualization modalities. Our results suggest that transformers offer a promising avenue of future work for molecular representation learning and property prediction. To facilitate these efforts, we release a curated dataset of 77M SMILES from PubChem suitable for large-scale self-supervised pretraining.
Forward citations
Cited by 25 Pith papers
-
A Foundation Model for Material Fracture Prediction
A transformer pretrained on rule-based fracture data and phase-field simulations predicts final fracture patterns across five materials and two loading conditions, and fine-tunes to new materials, time-to-failure, and...
-
ED-DiT: Physics-Guided Diffusion Pretraining for Transferable Molecular Representations from Electron Density
ED-DiT pretrains a diffusion transformer on electron-density point clouds with a physical electron-number constraint, and the resulting encoder outperforms scratch models across six molecular tasks.
-
MS-GPT: Rethinking MS/MS De Novo Structure Elucidation as Spectrum-Induced Posterior Querying of a Molecule-Language Model
Querying a fingerprint-conditioned molecule-language model with a calibrated band of spectrum-induced posteriors beats point-threshold fingerprint decoding on NPLIB1 and MassSpecGym.
-
Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction
Generic molecular encoders show weak alignment with human olfactory rating geometry and no clear predictive increment over an RDKit-Morgan chemistry baseline, while human rating geometry is reproducible within but onl...
-
OLEDLM: A Unified Language Model for OLED Molecular Design
A GRPO-aligned LLaMA model generates OLED SMILES conditioned on target S1 and f, with the property-alignment evidence largely coming from the very BERT predictor used as reward.
-
ChemHyperMag: Physics-informed magnetic hypergraph learning improves molecular ADMET prediction
A magnetic hypergraph neural network with chemistry-inspired directional flow improves ADMET property prediction on several benchmarks.
-
Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools
Augmenting SLM prompts with a GNN expert's prediction, confidence, and an extracted important subgraph improves zero-shot toxicity/mutagenicity accuracy on MUTAG and Tox21, with gains up to 74% relative to SMILES-only...
-
SIGMA: Semantic Identifier Grouping for Molecular Autoregression
A same-suffix contrastive objective makes autoregressive molecular-string models more invariant to how a molecule is written, improving generation fidelity on the reported ZINC benchmark.
-
Towards a Physics Foundation Model
A single transformer-based model, GPhyT, trained on diverse 2D simulation data, predicts next states across several fluid and heat-transfer systems and extrapolates to similar unseen regimes with plausible results.
-
MODA: A Unified 3D Diffusion Framework for Multi-Task Target-Aware Molecular Generation
A single masked-diffusion model trained jointly on four molecular-editing tasks outperforms or matches task-specific diffusion baselines across docking, chemical property, and geometry metrics.
-
DeepRetro: Retrosynthetic Pathway Discovery using Iterative LLM Reasoning
DeepRetro combines LLM-generated retrosynthetic disconnections with template-based search and human feedback, achieving strong benchmark results and proposing new routes for complex natural products.
-
OmniESI: A unified framework for enzyme-substrate interaction prediction with progressive conditional deep learning
A two-stage conditional deep learning framework modulates enzyme and substrate embeddings toward catalysis-aware features, reporting better benchmark results than specialized enzyme-substrate predictors.
-
Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction
Chem World unifies 17 chemical mixture datasets into 10 property tracks, and Mixture-PINN’s soft physics regularizers beat common encoder–aggregator baselines on most tracks and OOD splits.
-
Trustworthy Protein-Ligand Binding Affinity Prediction via Reliability-Aware Multi-Engine Fusion
A reliability-aware evidential fusion model combines four docking engines and produces well-calibrated confidence scores that enable selective prediction with up to 25.7% error reduction.
-
A Quiet Failure in Calibrated Virtual Screening: Marginal Conformal Prediction Under-Covers the Minority Class, and a Class-Conditional Fix Recovers It
Marginal conformal prediction under-covers minority classes by tens of points on imbalanced molecular datasets while hitting global coverage; Mondrian calibration recovers the target.
-
Valid Property-Enhanced Contrastive Learning for Targeted Optimization & Resampling for Novel Drug Design
VECTOR+ combines contrastive learning and Gaussian mixture sampling to generate novel, synthetically plausible inhibitors from low-data datasets, with improved docking scores over known compounds.
-
Chunk Twice, Embed Once: A Systematic Study of Segmentation and Representation Trade-offs in Chemistry-Aware Retrieval-Augmented Generation
A systematic evaluation shows that recursive 100-token non-overlapping chunks and retrieval-tuned embeddings outperform fixed-size chunks and domain-specific models like SciBERT for chemistry retrieval, and it introdu...
-
Uncovering Bottlenecks and Optimizing Scientific Lab Workflows with Cycle Time Reduction Agents
CTRA is a three-component LangGraph agent system for automatically generating analytical questions, SQL, and insights to identify bottlenecks in scientific lab workflows.
-
Persistent Manifold Learning of Protein Properties
Persistent manifold learning with Boundary-Induced Graph Laplacians plus language-model features beats prior SOTA Pearson correlation on metalloprotein–ligand and SKEMPI wild-type protein–protein affinity benchmarks.
-
Multimodal Molecular Representation Learning with Graph Neural Networks, Deep & Cross Networks, and SMILES Embeddings
Tri-modal late fusion of SchNet geometry, ChemBERTa SMILES, and DCN descriptors reaches 0.0207 eV MAE on QM9 U0 atomization energy, a 20.6% gain over a controlled SchNet baseline under 1M parameters.
-
Adaptive Minds: Empowering Agents with LoRA-as-Tools
Adaptive Minds makes a base LLM select LoRA adapters as tools per query; the 5-adapter demo gets 100% routing on 25 queries, while the abstract's 30-adapter/nine-family numbers are unsupported.
-
A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools
This survey organizes foundation models, LLM agents, datasets, and tools in materials science into six task areas.
-
Predicting Drug-Drug Interactions Using Heterogeneous Graph Neural Networks: HGNN-DDI
HGNN-DDI claims state-of-the-art drug-drug interaction prediction, but its skewed test set and missing SOTA baselines make the performance claim unsupported.
-
Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation
This survey claims to be the first systematic review of LLMs for organic synthesis, but its central 'evaluation' is never actually performed.
-
Transformers in Protein: A Survey
A broad but unreliable survey of Transformer applications in protein informatics, with numerous citation errors and unsupported claims.
Discussion (0). Sign in to comment.