Pith. sign in

REVIEW 15 cited by

Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.08018 v5 pith:QV4K6KIF submitted 2023-06-13 q-bio.QM cs.AIcs.CEcs.CLcs.IRcs.LG

Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language Models

classification q-bio.QM cs.AIcs.CEcs.CLcs.IRcs.LG
keywords biomolecularmol-instructionsinstructioninstructionslargellmsmodelscapabilities
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large Language Models (LLMs), with their remarkable task-handling capabilities and innovative outputs, have catalyzed significant advancements across a spectrum of fields. However, their proficiency within specialized domains such as biomolecular studies remains limited. To address this challenge, we introduce Mol-Instructions, a comprehensive instruction dataset designed for the biomolecular domain. Mol-Instructions encompasses three key components: molecule-oriented instructions, protein-oriented instructions, and biomolecular text instructions. Each component aims to improve the understanding and prediction capabilities of LLMs concerning biomolecular features and behaviors. Through extensive instruction tuning experiments on LLMs, we demonstrate the effectiveness of Mol-Instructions in enhancing large models' performance in the intricate realm of biomolecular studies, thus fostering progress in the biomolecular research community. Mol-Instructions is publicly available for ongoing research and will undergo regular updates to enhance its applicability.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression

    cs.LG 2026-05 unverdicted novelty 7.0

    Distribution-Aware Reward optimizes LLM regression by treating rollouts as empirical predictive distributions and rewarding marginal improvements in CRPS quality rather than point accuracy alone.

  2. FORGE: Fragment-Oriented Ranking and Generation for Context-Aware Molecular Optimization

    cs.LG 2026-05 unverdicted novelty 7.0

    FORGE reformulates molecular optimization as context-aware fragment ranking and replacement using mined low-to-high edit pairs, outperforming larger language models and graph methods on standard benchmarks.

  3. VibeProteinBench: An Evaluation Benchmark for Language-interfaced Vibe Protein Design

    q-bio.QM 2026-05 unverdicted novelty 7.0

    VibeProteinBench is a three-stage language-interfaced benchmark revealing that no current LLM performs strongly across recognition, engineering, and generation of proteins.

  4. VibeProteinBench: An Evaluation Benchmark for Language-interfaced Vibe Protein Design

    q-bio.QM 2026-05 unverdicted novelty 7.0

    VibeProteinBench is a new benchmark evaluating LLMs on open-ended language-interfaced protein design across recognition, engineering, and generation, with no model showing strong performance in all areas.

  5. TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

    q-bio.QM 2026-07 conditional novelty 6.5

    A 52.6B-token multi-domain biology pretraining corpus with tool enrichment and new binding/localization instructions doubles a fixed base LLM's matched biology-eval score with little language forgetting.

  6. Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data

    physics.chem-ph 2026-07 conditional novelty 6.0

    A two-stage AI pipeline — spectral hypothesis generation followed by mass-constrained molecular refinement — reconstructs organic structures from multimodal spectra, with 93.8% top-1 accuracy on simulated QM9 data and...

  7. TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

    q-bio.QM 2026-07 conditional novelty 6.0

    TheBioCollection, a 52.6B-token unified biology corpus with tool-computed text and new instruction tasks, raises a fixed 16B LLM's score on its matched biology eval from 0.223 to 0.499 (2.24×).

  8. Minibatch Selection for Language Models via Partition Matroid Constrained Gradient Matching

    cs.LG 2026-06 unverdicted novelty 6.0

    PartitionSel maximizes a validation-guided gradient-matching utility under partition-matroid per-domain budgets for cross-domain minibatch selection in LLM fine-tuning, showing gains over baselines on MetaMathQA and M...

  9. SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification

    cs.AI 2026-06 unverdicted novelty 6.0

    Sci-PRM is a tool-aware process reward model trained on the SCIPRM70K dataset to provide fine-grained supervision for scientific reasoning and shown to boost foundation models via Best-of-N selection and RL.

  10. Rethinking Molecular Text Representations for LLMs: An Empirical Study

    cs.LG 2026-06 unverdicted novelty 6.0

    Structured text representations like CML and MolJSON outperform SMILES variants on structural tasks while IUPAC dominates semantic tasks such as molecule retrieval across all tested LLMs.

  11. MolDA: Molecular Understanding and Generation via Large Language Diffusion Model

    cs.AI 2026-04 unverdicted novelty 6.0

    MolDA is a multimodal molecular model that uses a discrete large language diffusion backbone plus a hybrid graph encoder to achieve better global coherence and validity than autoregressive approaches.

  12. LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning

    physics.chem-ph 2026-02 conditional novelty 6.0

    LatentChem reasons in continuous latent space for chemistry, achieving a 59.88% non-tie win rate over explicit CoT on ChemCoTBench with a 10.84x average reduction in reasoning overhead.

  13. MolE-RAG: Molecular Structure-Enhanced Retrieval-Augmented Generation for Chemistry

    cs.LG 2026-06 unverdicted novelty 5.0

    MolE-RAG is a training-free RAG framework that augments LLMs with literature, molecular context, and structural analogs to improve performance on nine molecular property prediction tasks.

  14. SciCore-Mol: Augmenting Large Language Models with Pluggable Molecular Cognition Modules

    cs.AI 2026-05 unverdicted novelty 5.0

    SciCore-Mol augments LLMs with three integrated modules for molecular perception, latent diffusion generation, and reaction reasoning, claiming an 8B open model competes with or exceeds proprietary systems on chemical tasks.

  15. Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances

    cs.CL 2025-11 reject novelty 4.0

    Across the 68 papers it surveys, domain-specialized generative models usually outperform general-purpose LLMs on biological tasks, and agentic/conversational workflows are the least-covered topics.