Pith. sign in

REVIEW 2 cited by

Geneverse: A collection of Open-source Multimodal Large Language Models for Genomic and Proteomic Research

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.15534 v1 pith:NQAMGLEU submitted 2024-06-21 cs.LG cs.AIcs.CLq-bio.QM

classification cs.LGcs.AIcs.CLq-bio.QM
keywords llmsmodelsresearchgeneversetasksapplicationsbiomedicalcollection
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The applications of large language models (LLMs) are promising for biomedical and healthcare research. Despite the availability of open-source LLMs trained using a wide range of biomedical data, current research on the applications of LLMs to genomics and proteomics is still limited. To fill this gap, we propose a collection of finetuned LLMs and multimodal LLMs (MLLMs), known as Geneverse, for three novel tasks in genomic and proteomic research. The models in Geneverse are trained and evaluated based on domain-specific datasets, and we use advanced parameter-efficient finetuning techniques to achieve the model adaptation for tasks including the generation of descriptions for gene functions, protein function inference from its structure, and marker gene selection from spatial transcriptomic data. We demonstrate that adapted LLMs and MLLMs perform well for these tasks and may outperform closed-source large-scale models based on our evaluations focusing on both truthfulness and structural correctness. All of the training strategies and base models we used are freely accessible.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Multimodal, Multilingual, and Multidimensional Pipeline for Fine-grained Crowdsourcing Earthquake Damage Evaluation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Multimodal LLMs infer earthquake damage levels from social media with city-level correlations up to 0.78 against USGS crowd-reported intensity, with strong language and modality dependence.

  2. EvoLlama: Enhancing LLMs' Understanding of Proteins via Multimodal Structure and Sequence Representations

    cs.LG 2024-12 conditional novelty 4.0 of 10

    EvoLlama aligns ESM-2 sequence embeddings and ProteinMPNN structure embeddings with Llama-3, improving protein understanding over text-only LLMs on Mol-Instructions and PEER benchmarks.

Pith tools