REVIEW 5 cited by
MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Language Models (LMs) have demonstrated impressive molecule understanding ability on various 1D text-related tasks. However, they inherently lack 2D graph perception - a critical ability of human professionals in comprehending molecules' topological structures. To bridge this gap, we propose MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter. MolCA enables an LM (e.g., Galactica) to understand both text- and graph-based molecular contents via the cross-modal projector. Specifically, the cross-modal projector is implemented as a Q-Former to connect a graph encoder's representation space and an LM's text space. Further, MolCA employs a uni-modal adapter (i.e., LoRA) for the LM's efficient adaptation to downstream tasks. Unlike previous studies that couple an LM with a graph encoder via cross-modal contrastive learning, MolCA retains the LM's ability of open-ended text generation and augments it with 2D graph information. To showcase its effectiveness, we extensively benchmark MolCA on tasks of molecule captioning, IUPAC name prediction, and molecule-text retrieval, on which MolCA significantly outperforms the baselines. Our codes and checkpoints can be found at https://github.com/acharkq/MolCA.
Forward citations
Cited by 5 Pith papers
-
LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning
LatentChem reasons in continuous latent space for chemistry, achieving a 59.88% non-tie win rate over explicit CoT on ChemCoTBench with a 10.84x average reduction in reasoning overhead.
-
ModuLM: Enabling Modular and Multimodal Molecular Relational Learning with Large Language Models
A modular LLM-centric framework for molecular relational learning that supports 1D, 2D, and 3D molecular inputs and flexible model assembly, benchmarked across DDI, SSI, and CSI tasks.
-
HSA-Net: Hierarchical and Structure-Aware Framework for Efficient and Scalable Molecular Language Modeling
HSA-Net improves molecular language modeling by adaptively switching between cross-attention and Mamba projectors across GNN layers and fusing the results with a sparse mixture-of-experts.
-
CROP: Integrating Topological and Spatial Structures via Cross-View Prefixes for Molecular LLMs
Cross-view prefix resampling, guided by the LLM's SMILES encoding, lets a Galactica-based model exploit molecular graphs and images at low context cost, improving captioning, IUPAC naming, and property prediction.
-
NOCL: Node-Oriented Conceptualization LLM for Graph Tasks without Message Passing
NOCL lets an LLM handle node, edge, and graph tasks on text and non-text graphs by compressing each node's description into one semantic embedding and turning the graph into a text prompt.
Discussion (0). Continue with ORCID to comment.