Pith. sign in

REVIEW 4 cited by

GeLLMO: Generalizing Large Language Models for Multi-property Molecule Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.13398 v2 pith:EUPNQL2T submitted 2025-02-19 cs.LG cs.AIcs.CLphysics.chem-phq-bio.QM

classification cs.LGcs.AIcs.CLphysics.chem-phq-bio.QM
keywords optimizationtasksmoleculegellmosllmsmodelsdemonstrategeneralizability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite recent advancements, most computational methods for molecule optimization are constrained to single- or double-property optimization tasks and suffer from poor scalability and generalizability to novel optimization tasks. Meanwhile, Large Language Models (LLMs) demonstrate remarkable out-of-domain generalizability to novel tasks. To demonstrate LLMs' potential for molecule optimization, we introduce MuMOInstruct, the first high-quality instruction-tuning dataset specifically focused on complex multi-property molecule optimization tasks. Leveraging MuMOInstruct, we develop GeLLMOs, a series of instruction-tuned LLMs for molecule optimization. Extensive evaluations across 5 in-domain and 5 out-of-domain tasks demonstrate that GeLLMOs consistently outperform state-of-the-art baselines. GeLLMOs also exhibit outstanding zero-shot generalization to unseen tasks, significantly outperforming powerful closed-source LLMs. Such strong generalizability demonstrates the tremendous potential of GeLLMOs as foundational models for molecule optimization, thereby tackling novel optimization tasks without resource-intensive retraining. MuMOInstruct, models, and code are accessible through https://github.com/ninglab/GeLLMO.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HALO: Interactive Co-abductive Reasoning in Scientific Hypothesis Generation

    cs.HC 2026-07 conditional novelty 6.0 of 10

    HALO uses a three-stage co-abduction loop—clustering candidates by property improvement, distilling strategies, and synthesizing strategies—to help medicinal chemists produce more optimized and more diverse molecular ...

  2. Large Language Models for Controllable Multi-property Multi-objective Molecule Optimization

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Instruction-tuned LLMs trained on C-MuMOInstruct, a new controllable multi-property molecule optimization dataset, outperform strong baselines on in-distribution and out-of-distribution optimization tasks.

  3. Rethinking Scientific Discovery in the Agentic Era

    cs.CL 2026-07 conditional novelty 5.5 of 10

    SCION claims an agentic OS with Research Execution Plans and layered memory that beats autonomous research-agent baselines on reading, ideation, molecule design, and antibody screening.

  4. MolEditRL: Structure-Preserving Molecular Editing via Discrete Diffusion and Reinforcement Learning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    MolEditRL uses structure-aware graph diffusion plus RL fine-tuning to edit molecules toward desired properties while preserving scaffold similarity, reporting SOTA on its own MolEdit-Instruct benchmark.

Pith tools