Pith. sign in

REVIEW 13 cited by

XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.07971 v2 pith:VFZEBFNU submitted 2023-06-13 cs.CV

classification cs.CV
keywords medicalmodelsradiographschestmodelperformancevision-languagexraygpt
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The latest breakthroughs in large vision-language models, such as Bard and GPT-4, have showcased extraordinary abilities in performing a wide range of tasks. Such models are trained on massive datasets comprising billions of public image-text pairs with diverse tasks. However, their performance on task-specific domains, such as radiology, is still under-investigated and potentially limited due to a lack of sophistication in understanding biomedical images. On the other hand, conversational medical models have exhibited remarkable success but have mainly focused on text-based analysis. In this paper, we introduce XrayGPT, a novel conversational medical vision-language model that can analyze and answer open-ended questions about chest radiographs. Specifically, we align both medical visual encoder (MedClip) with a fine-tuned large language model (Vicuna), using a simple linear transformation. This alignment enables our model to possess exceptional visual conversation abilities, grounded in a deep understanding of radiographs and medical domain knowledge. To enhance the performance of LLMs in the medical context, we generate ~217k interactive and high-quality summaries from free-text radiology reports. These summaries serve to enhance the performance of LLMs through the fine-tuning process. Our approach opens up new avenues the research for advancing the automated analysis of chest radiographs. Our open-source demos, models, and instruction sets are available at: https://github.com/mbzuai-oryx/XrayGPT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A chest X-ray VLM co-trained with classification and grounding heads, tuned with DAPO reinforcement learning, and augmented with deterministic measurement tools outperforms prior radiology VLMs on report generation, V...

  2. MedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A 3D CT vision-language pretraining framework with global and organ-level contrastive alignment plus a text retrieval bank achieves state-of-the-art zero-shot disease classification, report retrieval, and medical visu...

  3. CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual Rationale

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A preference optimization strategy using confidence-based hard example mining, similarity retrieval, and synthetic counterfactual rationales improves chest X-ray VQA accuracy by 8.93% relative over supervised fine-tuning.

  4. Interpreting Chest X-rays Like a Radiologist: A Benchmark with Clinical Reasoning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new 8-stage chest X-ray VQA benchmark and a context-aware model trained on it.

  5. Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A3Tune aligns the visual attention of medical LVLMs to prompt-relevant regions via SAM and BioMedCLIP weak labels plus a Mixture-of-Experts over LoRA, improving VQA and report generation accuracy.

  6. MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence

    cs.CV 2026-03 reject novelty 5.0 of 10

    MEDIC-AD adds anomaly-aware and difference tokens to a medical VLM, claiming SOTA lesion detection, temporal tracking, and visual grounding; the zero-shot claim is undermined by likely train/test overlap.

  7. MCA-RG: Enhancing LLMs with Medical Concept Alignment for Radiology Report Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    MCA-RG uses concept alignment, contrastive learning, matching loss, and feature gating to generate radiology reports, reporting SOTA on MIMIC-CXR and CheXpert Plus.

  8. LLM-driven Medical Report Generation via Communication-efficient Heterogeneous Federated Learning

    cs.CV 2025-06 conditional novelty 5.0 of 10

    FedMRG trains federated LLM-based report generators with low-rank adapters, diagnosis prompts, and dual-adapter mutual boosting, beating baselines on chest X-ray benchmarks.

  9. CT-Agent: A Multimodal-LLM Agent for 3D CT Radiology Question Answering

    cs.CV 2025-05 conditional novelty 5.0 of 10

    CT-Agent combines an LLM planner, region-specific LoRA adapters, and global/local token compression to improve 3D chest CT report generation and question answering on CT-RATE and RadGenome-ChestCT.

  10. HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation

    cs.CV 2025-02 reject novelty 5.0 of 10

    HealthGPT unifies medical image comprehension and generation in a single autoregressive model using heterogeneous low-rank adaptation, reporting strong benchmark results.

  11. MIRA: A Novel Framework for Fusing Modalities in Medical RAG

    cs.CV 2025-07 reject novelty 4.0 of 10

    A medical multimodal RAG pipeline with rethink-and-rearrange and online search; the claimed SOTA is contradicted by the paper's own PMC-VQA numbers.

  12. Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation

    eess.IV 2025-06 conditional novelty 4.0 of 10

    MedRegion-CT integrates region-representative tokens, mask-driven segmentation tokens, and patient-specific attribute prompts into a multimodal LLM, reporting state-of-the-art scores on RadGenome-Chest CT report generation.

  13. MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible Extensibility

    cs.CL 2025-05 reject novelty 4.0 of 10

    MedOrch is a modular framework in which LLMs call medical tools to answer clinical questions; its headline results on Alzheimer's, chest X-ray, and VQA benchmarks are weakened by best-of-five scoring.

Pith tools