REVIEW 2 cited by
PathInsight: Instruction Tuning of Multimodal Datasets and Models for Intelligence Assisted Diagnosis in Histopathology
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Pathological diagnosis remains the definitive standard for identifying tumors. The rise of multimodal large models has simplified the process of integrating image analysis with textual descriptions. Despite this advancement, the substantial costs associated with training and deploying these complex multimodal models, together with a scarcity of high-quality training datasets, create a significant divide between cutting-edge technology and its application in the clinical setting. We had meticulously compiled a dataset of approximately 45,000 cases, covering over 6 different tasks, including the classification of organ tissues, generating pathology report descriptions, and addressing pathology-related questions and answers. We have fine-tuned multimodal large models, specifically LLaVA, Qwen-VL, InternLM, with this dataset to enhance instruction-based performance. We conducted a qualitative assessment of the capabilities of the base model and the fine-tuned model in performing image captioning and classification tasks on the specific dataset. The evaluation results demonstrate that the fine-tuned model exhibits proficiency in addressing typical pathological questions. We hope that by making both our models and datasets publicly available, they can be valuable to the medical and research communities.
Forward citations
Cited by 2 Pith papers
-
Efficient and Comprehensive Feature Extraction in Large Vision-Language Model for Pathology Analysis
A pathology-specialized LVLM with mixed-task feature enhancement and attention-guided detail completion reports higher accuracy than existing LVLMs across many diagnostic tasks.
-
WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image
WSI-LLaVA, trained on a large AI-generated WSI question-answer benchmark, is reported to outperform prior models on morphology and diagnosis, though the evaluation loop is largely closed through GPT-4o.
Discussion (0). Continue with ORCID to comment.