REVIEW 10 cited by
Translation between Molecules and Natural Language
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We present $\textbf{MolT5}$ $-$ a self-supervised learning framework for pretraining models on a vast amount of unlabeled natural language text and molecule strings. $\textbf{MolT5}$ allows for new, useful, and challenging analogs of traditional vision-language tasks, such as molecule captioning and text-based de novo molecule generation (altogether: translation between molecules and language), which we explore for the first time. Since $\textbf{MolT5}$ pretrains models on single-modal data, it helps overcome the chemistry domain shortcoming of data scarcity. Furthermore, we consider several metrics, including a new cross-modal embedding-based metric, to evaluate the tasks of molecule captioning and text-based molecule generation. Our results show that $\textbf{MolT5}$-based models are able to generate outputs, both molecules and captions, which in many cases are high quality.
Forward citations
Cited by 10 Pith papers
-
Back to Basics: Improving Molecular Understanding in LLMs via SMILES-Graph Translation
Training LLMs first on bidirectional SMILES–graph conversion plus progressive CoT yields large structure-perception gains and better property prediction and molecular optimization.
-
MODA: A Unified 3D Diffusion Framework for Multi-Task Target-Aware Molecular Generation
A single masked-diffusion model trained jointly on four molecular-editing tasks outperforms or matches task-specific diffusion baselines across docking, chemical property, and geometry metrics.
-
DeepRetro: Retrosynthetic Pathway Discovery using Iterative LLM Reasoning
DeepRetro combines LLM-generated retrosynthetic disconnections with template-based search and human feedback, achieving strong benchmark results and proposing new routes for complex natural products.
-
ModuLM: Enabling Modular and Multimodal Molecular Relational Learning with Large Language Models
A modular LLM-centric framework for molecular relational learning that supports 1D, 2D, and 3D molecular inputs and flexible model assembly, benchmarked across DDI, SSI, and CSI tasks.
-
ChemMLLM: Chemical Multimodal Large Language Model
A chemical multimodal LLM is trained to understand and generate molecule images alongside SMILES and text, with claims of state-of-the-art results on five new tasks.
-
NMIRacle: Multi-modal Generative Molecular Elucidation from IR and NMR Spectra
NMIRacle generates molecular structures from combined raw IR and NMR spectra, improving Top-1 elucidation accuracy from 0.41 to 0.48 over the NMR2Struct baseline on the Alberts benchmark.
-
An Encoder-Decoder Foundation Chemical Language Model for Generative Polymer Design
A T5-based polymer language model pre-trained on 100 million hypothetical polymers predicts thermal, electronic, and solubility properties and generates polymers conditioned on a target glass-transition temperature, w...
-
HSA-Net: Hierarchical and Structure-Aware Framework for Efficient and Scalable Molecular Language Modeling
HSA-Net improves molecular language modeling by adaptively switching between cross-attention and Mamba projectors across GNN layers and fusing the results with a sparse mixture-of-experts.
-
CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMs
Cross-view prefix resampling, guided by the LLM's SMILES encoding, lets a Galactica-based model exploit molecular graphs and images at low context cost, improving captioning, IUPAC naming, and property prediction.
-
Adopting Reinforcement Learning with Verifiable Rewards for Molecular Generation
LLMol fine-tunes an LLM on simplified SELFIES and uses GRPO with RDKit-derived rewards for targeted molecular generation, but its own benchmark tables contradict the claimed state-of-the-art performance.
Discussion (0). Continue with ORCID to comment.