Pith. sign in

REVIEW 6 cited by

Medical Vision Language Pretraining: A survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.06224 v1 pith:Y3VULC3R submitted 2023-12-11 cs.CV cs.CL

classification cs.CVcs.CL
keywords medicalpretrainingdownstreampotentialsurveytasksvisiondata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Medical Vision Language Pretraining (VLP) has recently emerged as a promising solution to the scarcity of labeled data in the medical domain. By leveraging paired/unpaired vision and text datasets through self-supervised learning, models can be trained to acquire vast knowledge and learn robust feature representations. Such pretrained models have the potential to enhance multiple downstream medical tasks simultaneously, reducing the dependency on labeled data. However, despite recent progress and its potential, there is no such comprehensive survey paper that has explored the various aspects and advancements in medical VLP. In this paper, we specifically review existing works through the lens of different pretraining objectives, architectures, downstream evaluation tasks, and datasets utilized for pretraining and downstream tasks. Subsequently, we delve into current challenges in medical VLP, discussing existing and potential solutions, and conclude by highlighting future directions. To the best of our knowledge, this is the first survey focused on medical VLP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Prototype-conditioned Mixture-of-Experts synthesizes missing modalities in federated learning and beats prior methods on heterogeneous chest X-ray clients without public data.

  2. 6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Across 22 vision-language models, accuracy on simple medical perception questions dropped from ~74% on typical anatomy to ~29% on rare anatomical variants, with errors aligning to textbook priors.

  3. Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Decomposing clinical terms into visual attributes lets 0.23B-2B vision-language models match or beat much larger medical VLMs for abnormality grounding with only 16k training pairs.

  4. Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ALTA adapts a frozen masked-pretrained X-ray encoder to language with 8% trainable parameters and temporal-multiview inputs, improving medical retrieval and zero-shot classification.

  5. Multimodal Federated Learning With Missing Modalities through Feature Imputation Network

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A federated feature imputation network that synthesizes missing modality bottleneck features improves multimodal federated learning accuracy over naive and generative baselines.

  6. Estimating 2D Keypoints of Surgical Tools Using Vision-Language Models with Low-Rank Adaptation

    cs.CV 2025-08 conditional novelty 4.0 of 10

    Giving surgical-tool keypoint detection to Qwen2.5-VL via LoRA fine-tuning reaches MPJPE 0.0627 on SurgeoNet, comparable with or better than dedicated YOLOv8-Pose and SurgeoNet baselines.

Pith tools