REVIEW 4 cited by
GeLLMO: Generalizing Large Language Models for Multi-property Molecule Optimization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Despite recent advancements, most computational methods for molecule optimization are constrained to single- or double-property optimization tasks and suffer from poor scalability and generalizability to novel optimization tasks. Meanwhile, Large Language Models (LLMs) demonstrate remarkable out-of-domain generalizability to novel tasks. To demonstrate LLMs' potential for molecule optimization, we introduce MuMOInstruct, the first high-quality instruction-tuning dataset specifically focused on complex multi-property molecule optimization tasks. Leveraging MuMOInstruct, we develop GeLLMOs, a series of instruction-tuned LLMs for molecule optimization. Extensive evaluations across 5 in-domain and 5 out-of-domain tasks demonstrate that GeLLMOs consistently outperform state-of-the-art baselines. GeLLMOs also exhibit outstanding zero-shot generalization to unseen tasks, significantly outperforming powerful closed-source LLMs. Such strong generalizability demonstrates the tremendous potential of GeLLMOs as foundational models for molecule optimization, thereby tackling novel optimization tasks without resource-intensive retraining. MuMOInstruct, models, and code are accessible through https://github.com/ninglab/GeLLMO.
Forward citations
Cited by 4 Pith papers
-
HALO: Interactive Co-abductive Reasoning in Scientific Hypothesis Generation
HALO uses a three-stage co-abduction loop—clustering candidates by property improvement, distilling strategies, and synthesizing strategies—to help medicinal chemists produce more optimized and more diverse molecular ...
-
Large Language Models for Controllable Multi-property Multi-objective Molecule Optimization
Instruction-tuned LLMs trained on C-MuMOInstruct, a new controllable multi-property molecule optimization dataset, outperform strong baselines on in-distribution and out-of-distribution optimization tasks.
-
Rethinking Scientific Discovery in the Agentic Era
SCION claims an agentic OS with Research Execution Plans and layered memory that beats autonomous research-agent baselines on reading, ideation, molecule design, and antibody screening.
-
MolEditRL: Structure-Preserving Molecular Editing via Discrete Diffusion and Reinforcement Learning
MolEditRL uses structure-aware graph diffusion plus RL fine-tuning to edit molecules toward desired properties while preserving scaffold similarity, reporting SOTA on its own MolEdit-Instruct benchmark.
Discussion (0). Continue with ORCID to comment.