REVIEW 3 cited by
MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training in Radiology
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we consider enhancing medical visual-language pre-training (VLP) with domain-specific knowledge, by exploiting the paired image-text reports from the radiological daily practice. In particular, we make the following contributions: First, unlike existing works that directly process the raw reports, we adopt a novel triplet extraction module to extract the medical-related information, avoiding unnecessary complexity from language grammar and enhancing the supervision signals; Second, we propose a novel triplet encoding module with entity translation by querying a knowledge base, to exploit the rich domain knowledge in medical field, and implicitly build relationships between medical entities in the language embedding space; Third, we propose to use a Transformer-based fusion model for spatially aligning the entity description with visual signals at the image patch level, enabling the ability for medical diagnosis; Fourth, we conduct thorough experiments to validate the effectiveness of our architecture, and benchmark on numerous public benchmarks, e.g., ChestX-ray14, RSNA Pneumonia, SIIM-ACR Pneumothorax, COVIDx CXR-2, COVID Rural, and EdemaSeverity. In both zero-shot and fine-tuning settings, our model has demonstrated strong performance compared with the former methods on disease classification and grounding.
Forward citations
Cited by 3 Pith papers
-
CXR-LT 2024: A MICCAI challenge on long-tailed, multi-label, and zero-shot disease classification from chest X-ray
CXR-LT 2024 provides a new large chest X-ray benchmark with 45 labels and three tasks, and reports that top models achieve mAP of 0.28 to 0.53 on long-tailed tasks but only 0.11 to 0.13 on zero-shot unseen diseases.
-
RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment
RadAlign aligns chest X-ray images to medical concept descriptions, then uses an LLM plus similar past cases to generate radiology reports with state-of-the-art factual accuracy.
-
CXR-CML: Improved zero-shot classification of long-tailed multi-label diseases in Chest X-Rays
A CLIP-based chest X-ray classifier enhanced with GMM clustering and triplet loss reports higher AUC, but it is trained on the target dataset rather than being zero-shot.
Discussion (0). Continue with ORCID to comment.