REVIEW 9 cited by
Med-Flamingo: a Multimodal Medical Few-shot Learner
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Medicine, by its nature, is a multifaceted domain that requires the synthesis of information across various modalities. Medical generative vision-language models (VLMs) make a first step in this direction and promise many exciting clinical applications. However, existing models typically have to be fine-tuned on sizeable down-stream datasets, which poses a significant limitation as in many medical applications data is scarce, necessitating models that are capable of learning from few examples in real-time. Here we propose Med-Flamingo, a multimodal few-shot learner adapted to the medical domain. Based on OpenFlamingo-9B, we continue pre-training on paired and interleaved medical image-text data from publications and textbooks. Med-Flamingo unlocks few-shot generative medical visual question answering (VQA) abilities, which we evaluate on several datasets including a novel challenging open-ended VQA dataset of visual USMLE-style problems. Furthermore, we conduct the first human evaluation for generative medical VQA where physicians review the problems and blinded generations in an interactive app. Med-Flamingo improves performance in generative medical VQA by up to 20\% in clinician's rating and firstly enables multimodal medical few-shot adaptations, such as rationale generation. We release our model, code, and evaluation app under https://github.com/snap-stanford/med-flamingo.
Forward citations
Cited by 9 Pith papers
-
NeuroMosaic: Anatomically Grounded Multimodal Large Language Modeling for Molecularly Aware Glioma Reasoning from 3D MRI and Clinical Narratives
NeuroMosaic links MRI regions to diagnostic language via an anatomical graph router and concept memory, reporting external macro-F1 up to 0.784, IDH AUROC 0.918, and 0.703 pointing accuracy, with a 0.036 macro-F1 gain...
-
Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding
A query-conditioned latent evidence aggregator after frozen frame selection improves long-video QA by up to +5.2 average / +10.1 LVBench with 0.11–0.40% token overhead.
-
DentiAsk: A VQA Benchmark for Multimodal Reasoning in Panoramic Dental Radiographs
A 1,000-image, 10,000-QA dental VQA benchmark shows current VLMs handle descriptive recognition far better than spatial localization or numerical counting on panoramic radiographs.
-
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
A multi-view vision-language model trained on 20,000 fetal ultrasound reports generates clinical text and diagnoses, reportedly outperforming general and medical baselines.
-
Raptor: Scalable Train-Free Embeddings for 3D Medical Volumes Leveraging Pretrained 2D Foundation Models
A train-free method that compresses 3D medical volumes into small embeddings via a frozen 2D foundation model and random projections, outperforming several medical-volume pretrained models on benchmark tasks.
-
Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy
Kvasir-VQA-x1 expands Kvasir-VQA with 159,549 LLM-generated question-answer pairs stratified into three complexity levels, plus a robustness track using weakly augmented images.
-
Inference-Time Agentic Decision Rules Beat Longer Evolving Search for Multi-Image Medical Reasoning
On MedFrameQA, order-vote (57.89%) beats fixed prompting (52.73%) and order-rerank (55.79%), and a single 100-generation run drops final-test accuracy to 56.02%.
-
Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation
MedRegion-CT integrates region-representative tokens, mask-driven segmentation tokens, and patient-specific attribute prompts into a multimodal LLM, reporting state-of-the-art scores on RadGenome-Chest CT report generation.
-
The Latent Space Hypothesis: Toward Universal Medical Representation Learning
The paper argues that all medical data modalities encode projections of a single latent physiological state, so a universal learned geometry could unify diagnosis, monitoring, and treatment.
Discussion (0). Continue with ORCID to comment.