Pith. sign in

REVIEW 2 cited by

Improving Chest X-Ray Report Generation by Leveraging Warm Starting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.09405 v2 pith:RKGWSEMU submitted 2022-01-24 cs.CV

classification cs.CV
keywords reportcvt2distilgpt2startingtransformerwarmcheckpointsgenerationvision
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Automatically generating a report from a patient's Chest X-Rays (CXRs) is a promising solution to reducing clinical workload and improving patient care. However, current CXR report generators -- which are predominantly encoder-to-decoder models -- lack the diagnostic accuracy to be deployed in a clinical setting. To improve CXR report generation, we investigate warm starting the encoder and decoder with recent open-source computer vision and natural language processing checkpoints, such as the Vision Transformer (ViT) and PubMedBERT. To this end, each checkpoint is evaluated on the MIMIC-CXR and IU X-Ray datasets. Our experimental investigation demonstrates that the Convolutional vision Transformer (CvT) ImageNet-21K and the Distilled Generative Pre-trained Transformer 2 (DistilGPT2) checkpoints are best for warm starting the encoder and decoder, respectively. Compared to the state-of-the-art ($\mathcal{M}^2$ Transformer Progressive), CvT2DistilGPT2 attained an improvement of 8.3\% for CE F-1, 1.8\% for BLEU-4, 1.6\% for ROUGE-L, and 1.0\% for METEOR. The reports generated by CvT2DistilGPT2 have a higher similarity to radiologist reports than previous approaches. This indicates that leveraging warm starting improves CXR report generation. Code and checkpoints for CvT2DistilGPT2 are available at https://github.com/aehrc/cvt2distilgpt2.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bridging Vision and Language: Optimal Transport-Driven Radiology Report Generation via LLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Optimal Transport alignment between image patches and disease labels, paired with LLM fine-tuning, improves clinical efficacy of generated radiology reports.

  2. Learning to See Locally and Align Clinically with Pathology Semantics for Radiology Report Generation

    eess.IV 2026-07 conditional novelty 5.0 of 10

    Radiology report generation improves when image and text features are aligned through shared, CheXpert-initialized pathology prototypes and a masked-evidence objective; PALM reports state-of-the-art scores on three ch...

Pith tools