Pith. sign in

REVIEW 9 cited by

Med-Flamingo: a Multimodal Medical Few-shot Learner

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.15189 v1 pith:ZBX3NHJ3 submitted 2023-07-27 cs.CV cs.AI

classification cs.CVcs.AI
keywords medicalmed-flamingofew-shotgenerativemodelsmultimodalapplicationsdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Medicine, by its nature, is a multifaceted domain that requires the synthesis of information across various modalities. Medical generative vision-language models (VLMs) make a first step in this direction and promise many exciting clinical applications. However, existing models typically have to be fine-tuned on sizeable down-stream datasets, which poses a significant limitation as in many medical applications data is scarce, necessitating models that are capable of learning from few examples in real-time. Here we propose Med-Flamingo, a multimodal few-shot learner adapted to the medical domain. Based on OpenFlamingo-9B, we continue pre-training on paired and interleaved medical image-text data from publications and textbooks. Med-Flamingo unlocks few-shot generative medical visual question answering (VQA) abilities, which we evaluate on several datasets including a novel challenging open-ended VQA dataset of visual USMLE-style problems. Furthermore, we conduct the first human evaluation for generative medical VQA where physicians review the problems and blinded generations in an interactive app. Med-Flamingo improves performance in generative medical VQA by up to 20\% in clinician's rating and firstly enables multimodal medical few-shot adaptations, such as rationale generation. We release our model, code, and evaluation app under https://github.com/snap-stanford/med-flamingo.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NeuroMosaic: Anatomically Grounded Multimodal Large Language Modeling for Molecularly Aware Glioma Reasoning from 3D MRI and Clinical Narratives

    cs.NE 2026-08 conditional novelty 6.0 of 10

    NeuroMosaic links MRI regions to diagnostic language via an anatomical graph router and concept memory, reporting external macro-F1 up to 0.784, IDH AUROC 0.918, and 0.703 pointing accuracy, with a 0.036 macro-F1 gain...

  2. Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A query-conditioned latent evidence aggregator after frozen frame selection improves long-video QA by up to +5.2 average / +10.1 LVBench with 0.11–0.40% token overhead.

  3. DentiAsk: A VQA Benchmark for Multimodal Reasoning in Panoramic Dental Radiographs

    q-bio.QM 2026-06 conditional novelty 6.0 of 10

    A 1,000-image, 10,000-QA dental VQA benchmark shows current VLMs handle descriptive recognition far better than spatial localization or numerical counting on panoramic radiographs.

  4. Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A multi-view vision-language model trained on 20,000 fetal ultrasound reports generates clinical text and diagnoses, reportedly outperforming general and medical baselines.

  5. Raptor: Scalable Train-Free Embeddings for 3D Medical Volumes Leveraging Pretrained 2D Foundation Models

    eess.IV 2025-07 conditional novelty 6.0 of 10

    A train-free method that compresses 3D medical volumes into small embeddings via a frozen 2D foundation model and random projections, outperforming several medical-volume pretrained models on benchmark tasks.

  6. Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Kvasir-VQA-x1 expands Kvasir-VQA with 159,549 LLM-generated question-answer pairs stratified into three complexity levels, plus a robustness track using weakly augmented images.

  7. Inference-Time Agentic Decision Rules Beat Longer Evolving Search for Multi-Image Medical Reasoning

    cs.CV 2026-07 conditional novelty 4.0 of 10

    On MedFrameQA, order-vote (57.89%) beats fixed prompting (52.73%) and order-rerank (55.79%), and a single 100-generation run drops final-test accuracy to 56.02%.

  8. Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation

    eess.IV 2025-06 conditional novelty 4.0 of 10

    MedRegion-CT integrates region-representative tokens, mask-driven segmentation tokens, and patient-specific attribute prompts into a multimodal LLM, reporting state-of-the-art scores on RadGenome-Chest CT report generation.

  9. The Latent Space Hypothesis: Toward Universal Medical Representation Learning

    q-bio.QM 2025-06 conditional novelty 4.0 of 10

    The paper argues that all medical data modalities encode projections of a single latent physiological state, so a universal learned geometry could unify diagnosis, monitoring, and treatment.

Pith tools