REVIEW 3 cited by
MolTC: Towards Molecular Relational Modeling In Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Molecular Relational Learning (MRL), aiming to understand interactions between molecular pairs, plays a pivotal role in advancing biochemical research. Recently, the adoption of large language models (LLMs), known for their vast knowledge repositories and advanced logical inference capabilities, has emerged as a promising way for efficient and effective MRL. Despite their potential, these methods predominantly rely on the textual data, thus not fully harnessing the wealth of structural information inherent in molecular graphs. Moreover, the absence of a unified framework exacerbates the issue of information underutilization, as it hinders the sharing of interaction mechanism learned across diverse datasets. To address these challenges, this work proposes a novel LLM-based multi-modal framework for Molecular inTeraction prediction following Chain-of-Thought (CoT) theory, termed MolTC, which effectively integrate graphical information of two molecules in pair. To train MolTC efficiently, we introduce a Multi-hierarchical CoT concept to refine its training paradigm, and conduct a comprehensive Molecular Interactive Instructions dataset for the development of biochemical LLMs involving MRL. Our experiments, conducted across various datasets involving over 4,000,000 molecular pairs, exhibit the superiority of our method over current GNN and LLM-based baselines. Code is available at https://github.com/MangoKiller/MolTC.
Forward citations
Cited by 3 Pith papers
-
ModuLM: Enabling Modular and Multimodal Molecular Relational Learning with Large Language Models
A modular LLM-centric framework for molecular relational learning that supports 1D, 2D, and 3D molecular inputs and flexible model assembly, benchmarked across DDI, SSI, and CSI tasks.
-
CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMs
Cross-view prefix resampling, guided by the LLM's SMILES encoding, lets a Galactica-based model exploit molecular graphs and images at low context cost, improving captioning, IUPAC naming, and property prediction.
-
Prot2Chat: Protein LLM with Early-Fusion of Text, Sequence and Structure
Prot2Chat fuses protein sequence, structure, and question text in a text-aware adapter before a LoRA-tuned LLM generates answers, reporting large gains on Mol-Instructions but mixed zero-shot results on UniProtQA.
Discussion (0). Continue with ORCID to comment.