Pith. sign in

REVIEW 2 cited by

Multimodal Large Language Model driven Radiology Report Generation with Clinical Knowledge Enhancement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.06728 v2 pith:QBOA65PV submitted 2024-03-11 cs.CV

classification cs.CV
keywords clinicalknowledgereportmultimodalmethodradiologyreportsx-ray
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Radiology report generation (RRG) has attracted significant attention due to its potential to reduce the workload of radiologists. The performance of current RRG approaches remains unsatisfactory against clinical standards. This paper introduces a novel RRG method, MLLM-RRG, that integrates multimodal large language models (MLLMs) with various types of clinical knowledge to generate accurate and comprehensive chest X-ray reports. Our method first designs a referring anatomical feature extractor that leverages anatomical knowledge to analyze different regions of the chest X-ray image and extract visual features without explicitly detecting regions. Next, based on the MLLM's decoder, we develop a multimodal report generator that leverages multimodal prompts constructed from dedicated visual features and textual instructions to produce the radiology report in an auto-regressive way. Finally, we introduce a disease-oriented clinical classification and alignment scheme in a multi-task learning manner to leverage disease knowledge to better preserve the clinical relevance among the generated reports. Once the model is trained, we also introduce a novel clinical quality reinforcement learning strategy to enhance the MLLM with report knowledge, further refining the tones of the generated reports towards radiologists. Extensive experiments on the MIMIC-CXR and IU X-Ray datasets demonstrate the superiority of our method over the state of the art. Our codes will be available at https://github.com/viscom-tongji/MLLM-RRG.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training

    eess.IV 2025-08 conditional novelty 6.0 of 10

    On abdominal CT, a vision-language pre-training method using organ-level normal/abnormal contrastive learning and a VQ-VAE normality model achieves 84.9% average zero-shot AUC, beating prior methods by 3.6%.

  2. MedOrchestra: A Hybrid Cloud-Local LLM Approach for Clinical Data Interpretation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A cloud-local hybrid, where the cloud writes subtask prompts offline and a local model executes them on patient data, reached 70-85% staging accuracy, above local baselines and clinicians.

Pith tools