REVIEW 5 cited by
BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Recent advancements in biological research leverage the integration of molecules, proteins, and natural language to enhance drug discovery. However, current models exhibit several limitations, such as the generation of invalid molecular SMILES, underutilization of contextual information, and equal treatment of structured and unstructured knowledge. To address these issues, we propose $\mathbf{BioT5}$, a comprehensive pre-training framework that enriches cross-modal integration in biology with chemical knowledge and natural language associations. $\mathbf{BioT5}$ utilizes SELFIES for $100%$ robust molecular representations and extracts knowledge from the surrounding context of bio-entities in unstructured biological literature. Furthermore, $\mathbf{BioT5}$ distinguishes between structured and unstructured knowledge, leading to more effective utilization of information. After fine-tuning, BioT5 shows superior performance across a wide range of tasks, demonstrating its strong capability of capturing underlying relations and properties of bio-entities. Our code is available at $\href{https://github.com/QizhiPei/BioT5}{Github}$.
Forward citations
Cited by 5 Pith papers
-
Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data
A two-stage AI pipeline — spectral hypothesis generation followed by mass-constrained molecular refinement — reconstructs organic structures from multimodal spectra, with 93.8% top-1 accuracy on simulated QM9 data and...
-
ModuLM: Enabling Modular and Multimodal Molecular Relational Learning with Large Language Models
A modular LLM-centric framework for molecular relational learning that supports 1D, 2D, and 3D molecular inputs and flexible model assembly, benchmarked across DDI, SSI, and CSI tasks.
-
CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMs
Cross-view prefix resampling, guided by the LLM's SMILES encoding, lets a Galactica-based model exploit molecular graphs and images at low context cost, improving captioning, IUPAC naming, and property prediction.
-
MolEditRL: Structure-Preserving Molecular Editing via Discrete Diffusion and Reinforcement Learning
MolEditRL uses structure-aware graph diffusion plus RL fine-tuning to edit molecules toward desired properties while preserving scaffold similarity, reporting SOTA on its own MolEdit-Instruct benchmark.
-
Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation
This survey claims to be the first systematic review of LLMs for organic synthesis, but its central 'evaluation' is never actually performed.
Discussion (0). Continue with ORCID to comment.