Pith. sign in

REVIEW 6 major objections 6 minor 2 cited by

Multimodal Large Language Models for Medicine: A Comprehensive Survey

T0 review · 6 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read The paper claims to provide a comprehensive map of medical multimodal large language models built from 330 papers, organized around three clinical applications and six data modes.

desk verdict A useful but unreliable map of medical MLLMs: the broad organization is sound, but the paper count is inflated by duplicate references, the abstract overpromises six data modes, and several key citations are wrong. read the letter →

arxiv 2504.21051 v1 pith:MSYKUDKS submitted 2025-04-29 cs.LG cs.CLcs.MM

classification cs.LGcs.CLcs.MM
keywords multimodallargelanguagemodelsmedicalartificialintelligencesurveyreportgenerationdialoguesurgicalassistancebenchmarksclinicalchallenges
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey claims to provide a comprehensive map of multimodal large language models (MLLMs) in medicine, built from 330 recent papers. It organizes the field into three application directions: generating medical reports, conducting professional and compassionate medical communication, and assisting clinical surgery. For each application it identifies representative models, the data modalities they consume, and the benchmarks used to evaluate them. The survey also catalogs data types, including image, text, audio, omics, and hybrid forms, and argues that data scarcity is the main bottleneck. It concludes that medical MLLMs show promise but still need better evaluation benchmarks, handling of rapidly updated knowledge, edge deployment, and privacy safeguards before clinical use.

What carries the argument

The load-bearing object is the standard MLLM architecture: a pretrained LLM at the core, a modality encoder for inputs such as images, video, or audio, an alignment module that projects those other modalities into the language model's feature space, and a generative model at the output. The survey uses this architectural recipe as its lens for classifying all 330 papers, and pairs it with a data taxonomy (image, text, audio, omics, instruction-following, and hybrid) and a list of evaluation benchmarks to explain how each model acquires and demonstrates medical competence. The taxonomy itself is the other central machinery: it is what turns a scattered literature into an ordered map.

What would settle it

Run a documented literature search for medical vision-language models published through early 2025, draw a random sample of in-scope papers, and check whether they appear in the survey's tables; if a substantial share, say 20 percent or more, is missing, the 330-paper map is incomplete. Also spot-check descriptions by re-reading the cited originals and comparing specific claims such as state-of-the-art scores on named benchmarks.

Watch

Extended reading notes

Core claim

The paper's central claim is that the medical MLLM landscape can be read through a single taxonomy: three clinical application areas, six mainstream data modes, and corresponding evaluation benchmarks, all resting on one architectural recipe. That recipe places a pretrained large language model at the core, with modality-specific encoders at the input and an alignment module that fuses non-text features into the language model's feature space. The survey's tables list dozens of medical MLLMs and LLMs, their base models, and their training data, and it presents the main evaluation routes: text-similarity metrics, expert manual scoring, AI-as-judge scoring, and medical licensing examinations such as USMLE. On this evidence the paper argues that while MLLMs can score well on closed benchmarks, they remain short of what clinical practice requires, and it names professionalism, hallucination, fairness and bias, rapidly changing medical knowledge, deployment constraints, and privacy as the barriers to close.

Load-bearing premise

The survey assumes that its selected 330 papers are representative of the medical MLLM literature and that its secondhand descriptions of them are accurate, yet it does not document a search protocol or inclusion criteria.

Editorial extensions

If this is right

  • A newcomer can locate any medical MLLM paper under one of three headings, namely report generation, medical communication, or surgical assistance, and immediately see which data and benchmarks that line of work uses.
  • The six data-mode classification makes explicit that audio and omics are the least developed, so future data-collection efforts should concentrate there.
  • Because the architectural recipe uses a pretrained LLM at the core, medical MLLMs largely inherit their reasoning from general multimodal models, so gains in general pretraining should transfer to medicine.
  • The evaluation checklist provides a standard way to compare new models, and the survey's examples show why exam scores alone cannot certify clinical readiness.
  • The challenges named, hallucination, bias, privacy, deployment, and rapidly changing knowledge, define a concrete agenda for turning promising models into clinically usable systems.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the survey does not document a systematic search strategy, its map is best read as a literature snapshot; a protocol-driven search could test whether 330 papers captures the full population.
  • The paper's own examples imply that exam-style benchmarks such as USMLE overestimate clinical readiness; building evaluation suites from real clinician workflows would test this directly.
  • Since most surveyed models fine-tune general-purpose MLLMs, the field's progress likely tracks general multimodal releases; evaluations should be re-run when new foundation models appear.
  • Treating audio and omics as data with potential suggests a concrete testable program: add respiratory audio or genomic data to text-plus-imaging models and measure whether fusion improves diagnostic accuracy over single-modality baselines.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

6 major / 6 minor

Summary. This manuscript is a survey of multimodal large language models (MLLMs) applied to medicine. It opens with background on LLMs and MLLMs, then discusses applications (claimed in the abstract to be medical reporting, diagnosis, and treatment), data modalities, model traits (professionalism, hallucination, fairness), and future challenges. The paper states that the review is based on 330 recent papers and that it covers six mainstream data modes with corresponding evaluation benchmarks. The body includes large tables of medical MLLMs, medical LLMs, training datasets, and descriptions of representative systems.

Significance. If the survey's content were accurate and its bibliographic claims verifiable, the manuscript would be a useful entry point for researchers entering medical MLLMs: the model tables, dataset lists, and the discussion of traits such as hallucination and fairness are organized in a readable way. The paper does not claim new algorithmic results, but the value of a survey lies in the fidelity of its summarization and the reliability of its coverage. At present, internal evidence undercuts both: the reference list contains multiple duplicated entries, the abstract advertises a structure that the body does not follow, and key system descriptions in Table 6 are mis-cited. These are not peripheral stylistic issues; they affect the central claim of being a comprehensive and reliable map of 330 papers.

major comments (6)
  1. [References and Abstract] The abstract's claim of a "comprehensive review of 330 recent papers" is not supported by the bibliography. The reference list contains multiple duplicated entries: [139] and [140] are the same multi-omics integration paper, [145] and [146] are the same medical-textbook reasoning paper, [207] and [210] are the same training-data extraction paper, [126] and [127] are the same MedDialog paper, and [4] and [221] are the same Vicuna reference. Additional duplicates include [102]/[109], [63]/[128]/[265], [22]/[273], [31]/[71], and [33]/[112]. The unique number of distinct works reviewed is therefore smaller than 330, and the headline number is not trustworthy as stated.
  2. [Section 3 versus Abstract] The abstract promises three application directions: medical reporting, medical diagnosis, and medical treatment. Section 3, however, is organized around medical report generation (Section 3.1), professional and compassionate medical communication (Section 3.2), and clinical surgery assistance (Section 3.3). No section is devoted to medical diagnosis or medical treatment as such, so the advertised structure does not match the content actually presented.
  3. [Section 4 versus Abstract] The abstract states that the paper presents "six mainstream modes of data along with their corresponding evaluation benchmarks." Section 4 details only four data categories: image (Section 4.1), text (Section 4.2), and two subsections under "Data with Potential" (audio in Section 4.3.1 and omics in Section 4.3.2). Moreover, no evaluation benchmarks are systematically paired with each modality in that section. The "six modes" claim must either be implemented in the text or removed from the abstract.
  4. [Table 6] Table 6 contains load-bearing mis-citations. The HuatuoGPT row cites reference [19], which is the HuatuoGPT-Vision paper (arXiv:2406.19280), and lists BLOOMZ as the base model; the original HuatuoGPT is a text-only model built on LLaMA, not BLOOMZ. In the same table, the Med-PaLM row attributes to Med-PaLM "exceptional performance particularly in areas such as clinical note summarization, radiological image analysis, and drug interaction prediction." Med-PaLM is a text-only medical question-answering model and does not process images, so this description is inaccurate.
  5. [Table 5] Table 5 is captioned "Medical Multimodal Large Language Model Training Data," but its rows list text-only datasets such as MedDialog, Huatuo-26M, PubMedQA, and MIMIC-III. These are not multimodal training data. Table 4 is the actual multimodal dataset table, making the Table 5 heading misleading and in need of correction.
  6. [Sections 3 and 4 (methodology)] The paper does not describe a search strategy, inclusion or exclusion criteria, or any protocol for selecting the claimed 330 papers. Without such methodology, the claim of comprehensiveness cannot be independently assessed, and the reader cannot tell whether the curated set is representative of the field or skewed toward particular topics, venues, or time periods.
minor comments (6)
  1. [Section 3.3] The word "Moereover" should be "Moreover."
  2. [Figure 1] The model names "Med-Palm" and "Med-Palm M" should be written consistently as "Med-PaLM" and "Med-PaLM M," and there is an extra closing parenthesis after "HuatuoGPT [19]."
  3. [Table 6] The model name "Aqulia-Med" should be "Aquila-Med," and the rows are not in chronological order (for example, the 2023/11 MEDITRON row appears after 2024/01 and 2024/02 rows).
  4. [References] Beyond the duplicates listed in the major comments, the reference list should be deduplicated globally; examples include [63], [128], and [265], which all point to the same radiology foundation-model paper.
  5. [Table 4] The row for "SAT-DS" cites reference [257], which appears to be a paper about universal medical image segmentation rather than a question-answering dataset; this citation should be verified and corrected.
  6. [Reference [54]] Reference [54] is formatted as "GPT-3(text-davinci-003)" but the cited work is the InstructGPT paper "Training language models to follow instructions with human feedback"; the label should be made consistent with the cited work.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey summarizes external literature and derives no predictions from its own inputs.

full rationale

This paper is a literature survey and contains no derivation chain, fitted parameters, or predicted quantities whose outputs could reduce to its inputs. Its central claims, such as the assertion that the findings are based on a comprehensive review of 330 recent papers, are empirical summarization claims about external works rather than logical consequences of internal assumptions. None of the enumerated circularity patterns apply: the survey does not define an input in terms of its output, does not fit a parameter and rename it as a prediction, does not rely on load-bearing self-citations (the authors do not cite their own prior work), does not import uniqueness theorems from the authors, does not smuggle in an ansatz via citation, and does not rename a known result as a new organization. The apparent citation and description inaccuracies noted in the text, such as the HuatuoGPT row in Table 6 pointing to reference [19] and the description of Med-PaLM as performing radiological image analysis, are fidelity and correctness concerns for a survey, not circularity. Even if the reference list contains duplicates, that affects the accuracy of the count of distinct papers, but it does not make any stated result equivalent to its own premise by construction. The survey is self-contained as a review: its value depends on the quality of its summarization of external work, and no step in the paper's reasoning is circular.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

No free parameters or invented entities. The survey rests on the assumption that the cited literature is correctly reported.

assumptions (1)
  • domain assumption The 330 surveyed papers are accurately and faithfully summarized.
    The survey's value depends on the correctness of its secondhand descriptions; no independent verification is provided (Sections 3-5, Tables 1-6).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multimodal Large Language Models for Medicine: A Comprehensive Survey." pith.science (2026). https://pith.science/paper/MSYKUDKS

@misc{pith2026250421051,
  author       = {Pith},
  title        = {Pith review of: Multimodal Large Language Models for Medicine: A Comprehensive Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MSYKUDKS}},
  note         = {Machine review of arXiv:2504.21051}
}
read the original abstract

MLLMs have recently become a focal point in the field of artificial intelligence research. Building on the strong capabilities of LLMs, MLLMs are adept at addressing complex multi-modal tasks. With the release of GPT-4, MLLMs have gained substantial attention from different domains. Researchers have begun to explore the potential of MLLMs in the medical and healthcare domain. In this paper, we first introduce the background and fundamental concepts related to LLMs and MLLMs, while emphasizing the working principles of MLLMs. Subsequently, we summarize three main directions of application within healthcare: medical reporting, medical diagnosis, and medical treatment. Our findings are based on a comprehensive review of 330 recent papers in this area. We illustrate the remarkable capabilities of MLLMs in these domains by providing specific examples. For data, we present six mainstream modes of data along with their corresponding evaluation benchmarks. At the end of the survey, we discuss the challenges faced by MLLMs in the medical and healthcare domain and propose feasible methods to mitigate or overcome these issues.

Figures

Figures reproduced from arXiv: 2504.21051 by the authors.

Figure 1
Figure 1. The Evolutionary Pathway of Multimodal Large Language Models with Medical Applications Highlighted. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Training Workflow for Medical Multimodal Large Language Models (MLLMs) in Application Tasks. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Comprehensive Diagnosis Based on Various multi [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Hallucination due to object misidentification. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Hallucination due to misperceived object rela [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 6
Figure 6. Figure 6: Ideal Benchmarking Guidelines for Assessing Key [PITH_FULL_IMAGE:figures/full_fig_p012_6.png]
Figure 7
Figure 7. Figure 7: The evolution of large pre-trained language models through fine-tuning to domain-specific models. Fine-tuning After finishing the pre-training phase, although the model possesses the ability to comprehend and generate natural language, it still does not completely meet…
Figure 8
Figure 8. Figure 8: Illustration of common pretraining tasks: Image-Text Matching (ITM) for cross-modal alignment, and MIM+MLM for understanding masked image and text features. MLM + MIM The combination of Masked Language Modeling (MLM) and Masked Image Modeling (MIM) enables the learning…
Figure 9
Figure 9. Figure 9: Utilizing GPT-4o in the role of a doctor to evaluate the diagnostic content of medical MLLMs [PITH_FULL_IMAGE:figures/full_fig_p024_9.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models

    cs.AI 2025-11 conditional novelty 5.0 of 10

    Ablating just four neurons in LLaVA-1.5-7b's language-model down-projection layer triggers complete output collapse, with critical neurons concentrated in the language backbone.

  2. The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

    cs.AI 2026-07 conditional novelty 4.5 of 10

    Medical agents should be scaled mainly by richer clinical environments and self-evolution loops, not parameter growth alone, under a three-level autonomy taxonomy.

Reference graph

Works this paper leans on

273 extracted references · 3 canonical work pages · cited by 2 Pith papers

  1. [139]

    Evaluation and comparison of multi-omics data integration methods for cancer subtyping,

    R. Duan, L. Gao, Y. Gao, Y. Hu, H. Xu, M. Huang, K. Song, H. Wang, Y. Dong, C. Jiang et al. , “Evaluation and comparison of multi-omics data integration methods for cancer subtyping,” PLoS computational biology, vol. 17, no. 8, p. e1009224, 2021

  2. [140]

    Evaluation and comparison of multi-omics data integra- tion methods for cancer subtyping,

    ——, “Evaluation and comparison of multi-omics data integra- tion methods for cancer subtyping,” PLoS computational biology , vol. 17, no. 8, p. e1009224, 2021

  3. [210]

    Extracting training data from large language models,

    N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson et al. , “Extracting training data from large language models,” in 30th USENIX Security Symposium (USENIX Security 21) , 2021, pp. 2633–2650

  4. [126]

    Meddialog: Large-scale medical dialogue datasets,

    G. Zeng, W. Yang, Z. Ju, Y. Yang, S. Wang, R. Zhang, M. Zhou, J. Zeng, X. Dong, R. Zhanget al., “Meddialog: Large-scale medical dialogue datasets,” in EMNLP, 2020

  5. [127]

    Meddialog: Large-scale medical dialogue datasets,

    ——, “Meddialog: Large-scale medical dialogue datasets,” in EMNLP, 2020

  6. [221]

    Vicuna: An open- source chatbot impressing gpt-4 with 90%* chatgpt quality,

    W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez et al. , “Vicuna: An open- source chatbot impressing gpt-4 with 90%* chatgpt quality,” See https://vicuna. lmsys. org (accessed 14 April 2023) , vol. 2, no. 3, p. 6, 2023

  7. [19]

    Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale,

    J. Chen, C. Gui, R. Ouyang, A. Gao, S. Chen, G. H. Chen, X. Wang, R. Zhang, Z. Cai, K. Ji et al., “Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale,” arXiv preprint arXiv:2406.19280, 2024

  8. [109]

    Learning multi-modal representations by watching hundreds of surgical video lectures,

    K. Yuan, V . Srivastav, T. Yu, J. L. Lavanchy, P . Mascagni, N. Navab, and N. Padoy, “Learning multi-modal representations by watching hundreds of surgical video lectures,” arXiv preprint arXiv:2307.15220, 2023

  9. [265]

    To- wards generalist foundation model for radiology,

    C. Wu, X. Zhang, Y. Zhang, Y. Wang, and W. Xie, “To- wards generalist foundation model for radiology,” arXiv preprint arXiv:2308.02463, 2023

  10. [273]

    Llava-med: Training a large language-and-vision assistant for biomedicine in one day,

    C. Li, C. Wong, S. Zhang, N. Usuyama, H. Liu, J. Yang, T. Naumann, H. Poon, and J. Gao, “Llava-med: Training a large language-and-vision assistant for biomedicine in one day,” in NeurIPS, 2024

  11. [71]

    Dia-llama: Towards large language model-driven ct report generation,

    Z. Chen, L. Luo, Y. Bie, and H. Chen, “Dia-llama: Towards large language model-driven ct report generation,” arXiv preprint arXiv:2403.16386, 2024

  12. [112]

    Autorg-brain: Grounded report generation for brain mri,

    J. Lei, X. Zhang, C. Wu, L. Dai, Y. Zhang, Y. Zhang, Y. Wang, W. Xie, and Y. Li, “Autorg-brain: Grounded report generation for brain mri,” arXiv preprint arXiv:2407.16684, 2024

Show all 273 references
  1. [1]

    Attention is all you need,

    A. Vaswani, “Attention is all you need,” NeurIPS, 2017

  2. [2]

    Bert: Pre-training of deep bidirectional transformers for language understanding,

    J. Devlin, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805 , 2018

  3. [3]

    Scaling instruction- finetuned language models,

    H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma et al. , “Scaling instruction- finetuned language models,” Journal of Machine Learning Research, vol. 25, no. 70, pp. 1–53, 2024

  4. [5]

    Llama: Open and efficient foundation language models,

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar et al. , “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023

  5. [6]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agar- wal, G. Sastry, A. Askell, P . Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763

  6. [7]

    Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,

    J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,” in International conference on machine learning . PMLR, 2022, pp. 12 888–12 900

  7. [8]

    Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

    J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in International conference on machine learning. PMLR, 2023, pp. 19 730–19 742

  8. [9]

    Flamingo: a visual language model for few-shot learning,

    J.-B. Alayrac, J. Donahue, P . Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al., “Flamingo: a visual language model for few-shot learning,” NeurIPS, 2022

  9. [10]

    A maximum likelihood approach to continuous speech recognition,

    L. R. Bahl, F. Jelinek, and R. L. Mercer, “A maximum likelihood approach to continuous speech recognition,” IEEE transactions on pattern analysis and machine intelligence, no. 2, pp. 179–190, 1983

  10. [11]

    Interpolated estimation of markov source parameters from sparse data,

    F. Jelinek, “Interpolated estimation of markov source parameters from sparse data,” in Proc. Workshop on Pattern Recognition in Practice, 1980, 1980

  11. [12]

    Recurrent neural network based language model

    T. Mikolov, M. Karafi ´at, L. Burget, J. Cernock `y, and S. Khu- danpur, “Recurrent neural network based language model.” in Interspeech, vol. 2, no. 3. Makuhari, 2010, pp. 1045–1048

  12. [13]

    Language models are few-shot learners,

    T. B. Brown, “Language models are few-shot learners,” arXiv preprint arXiv:2005.14165, 2020

  13. [14]

    Palm: Scaling language modeling with pathways,

    A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P . Barham, H. W. Chung, C. Sutton, S. Gehrmann et al., “Palm: Scaling language modeling with pathways,” Journal of Machine Learning Research, vol. 24, no. 240, pp. 1–113, 2023

  14. [15]

    Gpt-4 technical report,

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023

  15. [16]

    Mm1: methods, analysis and insights from multimodal llm pre-training,

    B. McKinzie, Z. Gan, J.-P . Fauconnier, S. Dodge, B. Zhang, P . Dufter, D. Shah, X. Du, F. Peng, A. Belyiet al., “Mm1: methods, analysis and insights from multimodal llm pre-training,” in Euro- pean Conference on Computer Vision. Springer, 2025, pp. 304–323

  16. [18]

    Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge,

    Y. Li, Z. Li, K. Zhang, R. Dan, S. Jiang, and Y. Zhang, “Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge,” Cureus, vol. 15, no. 6, 2023

  17. [20]

    Towards conversational diagnostic ai,

    T. Tu, A. Palepu, M. Schaekermann, K. Saab, J. Freyberg, R. Tanno, A. Wang, B. Li, M. Amin, N. Tomasev et al., “Towards conversational diagnostic ai,” arXiv preprint arXiv:2401.05654 , 2024

  18. [23]

    Towards generalist biomedical ai,

    T. Tu, S. Azizi, D. Driess, M. Schaekermann, M. Amin, P .-C. Chang, A. Carroll, C. Lau, R. Tanno, I. Ktena et al. , “Towards generalist biomedical ai,” NEJM AI, vol. 1, no. 3, p. AIoa2300138, 2024

  19. [24]

    Skingpt-4: an interactive dermatology di- agnostic system with visual large language model,

    J. Zhou, X. He, L. Sun, J. Xu, X. Chen, Y. Chu, L. Zhou, X. Liao, B. Zhang, and X. Gao, “Skingpt-4: an interactive dermatology di- agnostic system with visual large language model,”arXiv preprint arXiv:2304.10691, 2023

  20. [26]

    The unified medical language system (umls): in- tegrating biomedical terminology,

    O. Bodenreider, “The unified medical language system (umls): in- tegrating biomedical terminology,” Nucleic acids research, vol. 32, no. suppl 1, pp. D267–D270, 2004

  21. [27]

    The shaky foundations of large language models and foundation models for electronic health records,

    M. Wornow, Y. Xu, R. Thapa, B. Patel, E. Steinberg, S. Fleming, M. A. Pfeffer, J. Fries, and N. H. Shah, “The shaky foundations of large language models and foundation models for electronic health records,” npj Digital Medicine, vol. 6, no. 1, p. 135, 2023

  22. [28]

    Large language models in medicine,

    A. J. Thirunavukarasu, D. S. J. Ting, K. Elangovan, L. Gutier- rez, T. F. Tan, and D. S. W. Ting, “Large language models in medicine,” Nature medicine, vol. 29, no. 8, pp. 1930–1940, 2023

  23. [29]

    Clinical text summarization: adapting large lan- guage models can outperform human experts,

    D. Van Veen, C. Van Uden, L. Blankemeier, J.-B. Delbrouck, A. Aali, C. Bluethgen, A. Pareek, M. Polacin, E. P . Reis, A. See- hofnerova et al., “Clinical text summarization: adapting large lan- guage models can outperform human experts,” Research Square, 2023

  24. [30]

    Xraygpt: Chest radiographs summarization using large medical vision-language models,

    O. C. Thawakar, A. M. Shaker, S. S. Mullappilly, H. Cholakkal, R. M. Anwer, S. Khan, J. Laaksonen, and F. Khan, “Xraygpt: Chest radiographs summarization using large medical vision-language models,” in Proceedings of the 23rd workshop on biomedical natural language processing,...

  25. [32]

    Gpt-driven radiology re- port generation with fine-tuned llama 3,

    S ¸.-V . Voinea, M. M ˘amuleanu, R. V . Teic ˘a, L. M. Florescu, D. Selis ¸teanu, and I. A. Gheonea, “Gpt-driven radiology re- port generation with fine-tuned llama 3,” Bioengineering, vol. 11, no. 10, p. 1043, 2024

  26. [35]

    Enhanc- ing radiological reporting in head and neck cancer: Converting free-text ct scan reports to structured reports using large language models,

    A. Gupta, H. Malhotra, A. K. Garg, and K. Rangarajan, “Enhanc- ing radiological reporting in head and neck cancer: Converting free-text ct scan reports to structured reports using large language models,” Indian Journal of Radiology and Imaging , 2024

  27. [36]

    Medklip: Medical knowledge enhanced language-image pre-training in radiology,

    C. Wu, X. Zhang, Y. Zhang, Y. Wang, and W. Xie, “Medklip: Medical knowledge enhanced language-image pre-training in radiology,” arXiv preprint arXiv:2301.02228, 2023

  28. [38]

    Semihvision: Enhancing medical multimodal models with a semi-human annotated dataset and fine-tuned instruction generation,

    J. Wang, Y. Ting, E. Z. Chen, H. Tran, H. Yu, W. Huang, and T. Chen, “Semihvision: Enhancing medical multimodal models with a semi-human annotated dataset and fine-tuned instruction generation,” arXiv preprint arXiv:2410.14948, 2024

  29. [40]

    The radiology report—are we IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 14 getting the message across?

    A. Wallis and P . McCoubrie, “The radiology report—are we IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 14 getting the message across?” Clinical radiology , vol. 66, no. 11, pp. 1015–1022, 2011

  30. [41]

    Chatcad: Interactive computer-aided diagnosis on medical image using large language models,

    S. Wang, Z. Zhao, X. Ouyang, Q. Wang, and D. Shen, “Chatcad: Interactive computer-aided diagnosis on medical image using large language models,” arXiv preprint arXiv:2302.07257, 2023

  31. [42]

    Making the most of text semantics to improve biomedical vision–language processing,

    B. Boecking, N. Usuyama, S. Bannur, D. C. Castro, A. Schwaighofer, S. Hyland, M. Wetscherek, T. Naumann, A. Nori, J. Alvarez-Valleet al., “Making the most of text semantics to improve biomedical vision–language processing,” in ECCV, 2022

  32. [43]

    A deep learning approaches and fastai text classification to predict 25 medical diseases from medical speech utterances, transcription and intent,

    Y. Kumar, A. Koul, and S. Mahajan, “A deep learning approaches and fastai text classification to predict 25 medical diseases from medical speech utterances, transcription and intent,”Springer Soft computing, vol. 26, no. 17, pp. 8253–8272, 2022

  33. [44]

    Automatic documentation of professional health interactions: A systematic review,

    F. S. Falcetta, F. K. De Almeida, J. C. S. Lemos, J. R. Goldim, and C. A. Da Costa, “Automatic documentation of professional health interactions: A systematic review,” Elsevier Artificial Intelligence in Medicine, vol. 137, p. 102487, 2023

  34. [45]

    Medical report generation based on segment-enhanced contrastive representa- tion learning,

    R. Zhao, X. Wang, H. Dai, P . Gao, and P . Li, “Medical report generation based on segment-enhanced contrastive representa- tion learning,” in CCF International Conference on Natural Language Processing and Chinese Computing. Springer, 2023, pp. 838–849

  35. [47]

    Ophglm: Training an ophthalmology large language-and-vision assistant based on instructions and dialogue,

    W. Gao, Z. Deng, Z. Niu, F. Rong, C. Chen, Z. Gong, W. Zhang, D. Xiao, F. Li, Z. Cao et al., “Ophglm: Training an ophthalmology large language-and-vision assistant based on instructions and dialogue,” arXiv preprint arXiv:2306.12174, 2023

  36. [48]

    Chatglm: A family of large language models from glm-130b to glm-4 all tools,

    T. GLM, A. Zeng, B. Xu, B. Wang, C. Zhang, D. Yin, D. Zhang, D. Rojas, G. Feng, H. Zhao et al. , “Chatglm: A family of large language models from glm-130b to glm-4 all tools,” arXiv preprint arXiv:2406.12793, 2024

  37. [49]

    Radiology-llama2: Best-in-class large language model for radiology,

    Z. Liu, Y. Li, P . Shu, A. Zhong, L. Yang, C. Ju, Z. Wu, C. Ma, J. Luo, C. Chen et al. , “Radiology-llama2: Best-in-class large language model for radiology,” arXiv preprint arXiv:2309.06419, 2023

  38. [50]

    Llama 2: Open foundation and fine-tuned chat models,

    H. Touvron, L. Martin, K. Stone, P . Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P . Bhargava, S. Bhosale et al. , “Llama 2: Open foundation and fine-tuned chat models,” arXiv preprint arXiv:2307.09288, 2023

  39. [52]

    Sigphi-med: A lightweight vision-language assistant for biomedicine,

    F. Zhou, X. Liu, Q. Zeng, Z. Li, and H. Xiao, “Sigphi-med: A lightweight vision-language assistant for biomedicine,” Available at SSRN 4988925, 2024

  40. [53]

    Phi- 2: The surprising power of small language models,

    M. Javaheripi, S. Bubeck, M. Abdin, J. Aneja, S. Bubeck, C. C. T. Mendes, W. Chen, A. Del Giorno, R. Eldan, S. Gopi et al., “Phi- 2: The surprising power of small language models,” Microsoft Research Blog, vol. 1, no. 3, p. 3, 2023

  41. [54]

    Train- ing language models to follow instructions with human feed- back,

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P . Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al., “Train- ing language models to follow instructions with human feed- back,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022

  42. [55]

    Medblip: Bootstrapping language-image pre-training from 3d medical images and texts,

    Q. Chen and Y. Hong, “Medblip: Bootstrapping language-image pre-training from 3d medical images and texts,” in Proceedings of the Asian Conference on Computer Vision, 2024, pp. 2404–2420

  43. [56]

    Biomedlm: A 2.7 b parameter language model trained on biomedical text,

    E. Bolton, A. Venigalla, M. Yasunaga, D. Hall, B. Xiong, T. Lee, R. Daneshjou, J. Frankle, P . Liang, M. Carbin et al., “Biomedlm: A 2.7 b parameter language model trained on biomedical text,” arXiv preprint arXiv:2403.18421, 2024

  44. [57]

    Pclmed at image- clefmedical 2023: Customizing general-purpose foundation mod- els for medical report generation

    B. Yang, A. Raza, Y. Zou, and T. Zhang, “Pclmed at image- clefmedical 2023: Customizing general-purpose foundation mod- els for medical report generation.” in CLEF (Working Notes), 2023, pp. 1754–1766

  45. [58]

    Pmc-vqa: Visual instruction tuning for medical visual question answering,

    X. Zhang, C. Wu, Z. Zhao, W. Lin, Y. Zhang, Y. Wang, and W. Xie, “Pmc-vqa: Visual instruction tuning for medical visual question answering,” arXiv preprint arXiv:2305.10415, 2023

  46. [59]

    Pmc-llama: toward building open-source language models for medicine,

    C. Wu, W. Lin, X. Zhang, Y. Zhang, W. Xie, and Y. Wang, “Pmc-llama: toward building open-source language models for medicine,” Journal of the American Medical Informatics Association , p. ocae045, 2024

  47. [60]

    Pathasst: A generative foundation ai assistant towards artificial general intelligence of pathology,

    Y. Sun, C. Zhu, S. Zheng, K. Zhang, L. Sun, Z. Shui, Y. Zhang, H. Li, and L. Yang, “Pathasst: A generative foundation ai assistant towards artificial general intelligence of pathology,” in Proceed- ings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 5, 2024, ...

  48. [61]

    Chatcad+: Towards a universal and reliable interactive cad using llms,

    Z. Zhao, S. Wang, J. Gu, Y. Zhu, L. Mei, Z. Zhuang, Z. Cui, Q. Wang, and D. Shen, “Chatcad+: Towards a universal and reliable interactive cad using llms,” IEEE TMI, 2024

  49. [62]

    Chatgpt: Optimizing language models for dialogue,

    OpenAI, “Chatgpt: Optimizing language models for dialogue,” https://openai.com/blog/chatgpt/, 2023, accessed: 2023-01-08

  50. [64]

    R2gengpt: Radiology report generation with frozen llms,

    Z. Wang, L. Liu, L. Wang, and L. Zhou, “R2gengpt: Radiology report generation with frozen llms,” Meta-Radiology, vol. 1, no. 3, p. 100033, 2023

  51. [65]

    Cxr-llava: A multimodal large language model for interpreting chest x-ray images,

    S. Lee, J. Youn, H. Kim, M. Kim, and S. H. Yoon, “Cxr-llava: A multimodal large language model for interpreting chest x-ray images,” European Radiology, pp. 1–13, 2025

  52. [66]

    Maira-1: A specialised large multimodal model for radiology report generation,

    S. L. Hyland, S. Bannur, K. Bouzid, D. C. Castro, M. Ran- jit, A. Schwaighofer, F. P ´erez-Garc´ıa, V . Salvatelli, S. Srivastav, A. Thieme et al., “Maira-1: A specialised large multimodal model for radiology report generation,” arXiv preprint arXiv:2311.13668, 2023

  53. [67]

    Pefomed: Parameter efficient fine-tuning of multimodal large language models for medical imaging,

    G. Liu, J. He, P . Li, G. He, Z. Chen, and S. Zhong, “Pefomed: Parameter efficient fine-tuning of multimodal large language models for medical imaging,” arXiv preprint arXiv:2401.02797 , 2024

  54. [68]

    Minigpt-v2: large language model as a unified interface for vision-language multi- task learning,

    J. Chen, D. Zhu, X. Shen, X. Li, Z. Liu, P . Zhang, R. Krishnamoor- thi, V . Chandra, Y. Xiong, and M. Elhoseiny, “Minigpt-v2: large language model as a unified interface for vision-language multi- task learning,” arXiv preprint arXiv:2310.09478, 2023

  55. [70]

    Moe-tinymed: Mixture of experts for tiny medical large vision-language mod- els,

    S. Jiang, T. Zheng, Y. Zhang, Y. Jin, and Z. Liu, “Moe-tinymed: Mixture of experts for tiny medical large vision-language mod- els,” arXiv preprint arXiv:2404.10237, 2024

  56. [72]

    Maira-2: Grounded radiology report generation,

    S. Bannur, K. Bouzid, D. C. Castro, A. Schwaighofer, A. Thieme, S. Bond-Taylor, M. Ilse, F. P ´erez-Garc´ıa, V . Salvatelli, H. Sharma et al. , “Maira-2: Grounded radiology report generation,” arXiv preprint arXiv:2406.04449, 2024

  57. [73]

    Minigpt-med: Large language model as a general interface for radiology diagnosis,

    A. Alkhaldi, R. Alnajim, L. Alabdullatef, R. Alyahya, J. Chen, D. Zhu, A. Alsinan, and M. Elhoseiny, “Minigpt-med: Large language model as a general interface for radiology diagnosis,” arXiv preprint arXiv:2407.04106, 2024

  58. [74]

    Llava-surg: Towards multimodal surgical assistant via structured surgical video learning,

    J. Li, G. Skinner, G. Yang, B. R. Quaranto, S. D. Schwaitzberg, P . C. Kim, and J. Xiong, “Llava-surg: Towards multimodal surgical assistant via structured surgical video learning,” arXiv preprint arXiv:2408.07981, 2024

  59. [75]

    The llama 3 herd of models,

    A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Let- man, A. Mathur, A. Schelten, A. Yang, A. Fan et al., “The llama 3 herd of models,” arXiv preprint arXiv:2407.21783, 2024

  60. [77]

    Aquila2 technical report,

    B.-W. Zhang, L. Wang, J. Li, S. Gu, X. Wu, Z. Zhang, B. Gao, Y. Ao, and G. Liu, “Aquila2 technical report,” arXiv preprint arXiv:2408.07410, 2024

  61. [78]

    Vision- biollm: Large vision language model for visual dialogue in biomedical imagery,

    A. AlShibli, Y. Bazi, M. M. Al Rahhal, and M. Zuair, “Vision- biollm: Large vision language model for visual dialogue in biomedical imagery,” Biomedical Signal Processing and Control, vol. 103, p. 107437, 2025

  62. [79]

    Openbiollms: Advancing open- source large language models for healthcare and life sciences,

    A. Pal and M. Sankarasubbu, “Openbiollms: Advancing open- source large language models for healthcare and life sciences,” 2024

  63. [80]

    Baize: An open-source chat model with parameter-efficient tuning on self-chat data,

    C. Xu, D. Guo, N. Duan, and J. McAuley, “Baize: An open-source chat model with parameter-efficient tuning on self-chat data,” arXiv preprint arXiv:2304.01196, 2023

  64. [81]

    Gpt-4v (ision) unsuitable for clinical care and education: a clinician-evaluated assessment,

    S. Senkaiahliyan, A. Toma, J. Ma, A.-W. Chan, A. Ha, K. R. An, H. Suresh, B. Rubin, and B. Wang, “Gpt-4v (ision) unsuitable for clinical care and education: a clinician-evaluated assessment,” medRxiv, pp. 2023–11, 2023

  65. [82]

    Nadarzynski, J

    T. Nadarzynski, J. Bayley, C. Llewellyn, S. Kidsley, and C. A. Graham, “Acceptability of artificial intelligence (ai)-enabled chat- bots, video consultations and live webchats as online platforms IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 15 for sexual hea...

  66. [83]

    Investigating public perception on use of chatgpt in initial consultations prior to healthcare provider consultations,

    S. Hussain, M. Alherz, E. Albazee, H. Almhanedi, J. Hayat, M. Lari, and A. Lari, “Investigating public perception on use of chatgpt in initial consultations prior to healthcare provider consultations,” Annals of Medicine and Surgery, pp. 10–1097, 2024

  67. [84]

    Mental health problems of contem- porary youth,

    K. Olga and Z. Xuehan, “Mental health problems of contem- porary youth,” Human Health (Zdorov’e cheloveka), Theory and Methodology of Physical Culture and Sports , no. 4 (15), pp. 45–49, 2019, in Russian

  68. [85]

    No health without mental health,

    M. Prince, V . Patel, S. Saxena, M. Maj, J. Maselko, M. R. Phillips, and A. Rahman, “No health without mental health,” The lancet, vol. 370, no. 9590, pp. 859–877, 2007

  69. [86]

    University counseling service for improving students’ mental health

    F. Vescovelli, P . Melani, C. Ruini, P . E. Ricci Bitti, and F. Monti, “University counseling service for improving students’ mental health.” Psychological services, vol. 14, no. 4, p. 470, 2017

  70. [87]

    Validity of chatbot use for mental health assessment: experimen- tal study,

    A. Schick, J. Feine, S. Morana, A. Maedche, and U. Reininghaus, “Validity of chatbot use for mental health assessment: experimen- tal study,” JMIR mHealth and uHealth , vol. 10, no. 10, p. e28082, 2022

  71. [88]

    Effectiveness and safety of using chatbots to improve mental health: systematic review and meta-analysis,

    A. A. Abd-Alrazaq, A. Rababeh, M. Alajlani, B. M. Bewick, and M. Househ, “Effectiveness and safety of using chatbots to improve mental health: systematic review and meta-analysis,” Journal of medical Internet research, vol. 22, no. 7, p. e16021, 2020

  72. [89]

    Tell me, what are you most afraid of? exploring the effects of agent representation on infor- mation disclosure in human-chatbot interaction,

    A. Stock, S. Schl ¨ogl, and A. Groth, “Tell me, what are you most afraid of? exploring the effects of agent representation on infor- mation disclosure in human-chatbot interaction,” in International Conference on Human-Computer Interaction . Springer, 2023, pp. 179–191

  73. [90]

    How should my chatbot inter- act? a survey on social characteristics in human–chatbot interac- tion design,

    A. P . Chaves and M. A. Gerosa, “How should my chatbot inter- act? a survey on social characteristics in human–chatbot interac- tion design,” International Journal of Human–Computer Interaction , vol. 37, no. 8, pp. 729–758, 2021

  74. [91]

    Soulchat: Improving llms’ empathy, listening, and comfort abili- ties through fine-tuning with multi-turn empathy conversations,

    Y. Chen, X. Xing, J. Lin, H. Zheng, Z. Wang, Q. Liu, and X. Xu, “Soulchat: Improving llms’ empathy, listening, and comfort abili- ties through fine-tuning with multi-turn empathy conversations,” in EMNLP, 2023

  75. [92]

    Psychat: A client-centric dialogue system for mental health support,

    H. Qiu, A. Li, L. Ma, and Z. Lan, “Psychat: A client-centric dialogue system for mental health support,” in CSCWD, 2024

  76. [93]

    Smile: Single-turn to multi-turn inclusive language expansion via chatgpt for mental health support,

    H. Qiu, H. He, S. Zhang, A. Li, and Z. Lan, “Smile: Single-turn to multi-turn inclusive language expansion via chatgpt for mental health support,” arXiv preprint arXiv:2305.00450, 2023

  77. [94]

    Esc-eval: Evaluating emotion support conversations in large language models,

    H. Zhao, L. Li, S. Chen, S. Kong, J. Wang, K. Huang, T. Gu, Y. Wang, W. Jian, D. Liang et al., “Esc-eval: Evaluating emotion support conversations in large language models,” arXiv preprint arXiv:2406.14952, 2024

  78. [95]

    Cpsycoun: A report-based multi-turn dialogue reconstruction and evaluation framework for chinese psychological counseling,

    C. Zhang, R. Li, M. Tan, M. Yang, J. Zhu, D. Yang, J. Zhao, G. Ye, C. Li, and X. Hu, “Cpsycoun: A report-based multi-turn dialogue reconstruction and evaluation framework for chinese psychological counseling,” arXiv preprint arXiv:2405.16433, 2024

  79. [96]

    Speech emotion recognition and sentiment analysis based therapist bot,

    Y. Bhangdia, R. Bhansali, N. Chaudhari, D. Chandnani, and M. Dhore, “Speech emotion recognition and sentiment analysis based therapist bot,” in ICIRCA, 2021

  80. [97]

    Emoada: A multimodal emotion interaction and psychologi- cal adaptation system,

    T. Dong, F. Liu, X. Wang, Y. Jiang, X. Zhang, and X. Sun, “Emoada: A multimodal emotion interaction and psychologi- cal adaptation system,” in International Conference on Multimedia Modeling, 2024

  81. [99]

    Computer-assisted surgery,

    L. Adams, W. Krybus, D. Meyer-Ebrecht, R. Rueger, J. M. Gilsbach, R. Moesges, and G. Schloendorff, “Computer-assisted surgery,” IEEE Computer graphics and applications , vol. 10, no. 3, pp. 43–51, 1990

  82. [100]

    Cas—a navigation support for surgery,

    L. Adams, J. M. Gilsbach, W. Krybus, D. Meyer-Ebrecht, R. M ¨osges, and G. Schl ¨ondorff, “Cas—a navigation support for surgery,” in 3D Imaging in Medicine: Algorithms, Systems, Applica- tions. Springer, 1990, pp. 411–423

  83. [101]

    Surgical- vqa: Visual question answering in surgical scenes using trans- former,

    L. Seenivasan, M. Islam, A. K. Krishna, and H. Ren, “Surgical- vqa: Visual question answering in surgical scenes using trans- former,” in MICCAI, 2022

  84. [103]

    Surgicalgpt: end-to-end language-vision gpt for visual question answering in surgery,

    L. Seenivasan, M. Islam, G. Kannan, and H. Ren, “Surgicalgpt: end-to-end language-vision gpt for visual question answering in surgery,” in MICCAI, 2023

  85. [104]

    Advancing surgical vqa with scene graph knowl- edge,

    K. Yuan, M. Kattel, J. L. Lavanchy, N. Navab, V . Srivastav, and N. Padoy, “Advancing surgical vqa with scene graph knowl- edge,” International Journal of Computer Assisted Radiology and Surgery, pp. 1–9, 2024

  86. [105]

    Global- reasoned multi-task learning model for surgical scene under- standing,

    L. Seenivasan, S. Mitheran, M. Islam, and H. Ren, “Global- reasoned multi-task learning model for surgical scene under- standing,” IEEE RAL, vol. 7, no. 2, pp. 3858–3865, 2022

  87. [106]

    Surgical-vqla++: Adversarial contrastive learning for calibrated robust visual question-localized answering in robotic surgery,

    L. Bai, G. Wang, M. Islam, L. Seenivasan, A. Wang, and H. Ren, “Surgical-vqla++: Adversarial contrastive learning for calibrated robust visual question-localized answering in robotic surgery,” Information Fusion, vol. 113, p. 102602, 2025

  88. [107]

    Dynamic interactive relation capturing via scene graph learning for robotic surgical report generation,

    H. Wang, Y. Jin, and L. Zhu, “Dynamic interactive relation capturing via scene graph learning for robotic surgical report generation,” in ICRA, 2023

  89. [108]

    Sgt: Scene graph-guided transformer for surgical report generation,

    C. Lin, S. Zheng, Z. Liu, Y. Li, Z. Zhu, and Y. Zhao, “Sgt: Scene graph-guided transformer for surgical report generation,” in MICCAI, 2022

  90. [110]

    Drinet for medical image segmentation,

    L. Chen, P . Bentley, K. Mori, K. Misawa, M. Fujiwara, and D. Rueckert, “Drinet for medical image segmentation,” IEEE TMI, vol. 37, no. 11, pp. 2453–2462, 2018

  91. [111]

    Deep learning and medical image analysis for covid-19 diagnosis and prediction,

    T. Liu, E. Siegel, and D. Shen, “Deep learning and medical image analysis for covid-19 diagnosis and prediction,” Annual review of biomedical engineering, vol. 24, no. 1, pp. 179–201, 2022

  92. [113]

    Exploring chat generated pre-trained transformer-3 ability to interpret mri knee images and generate reports,

    S. Saran, K. Shirodkar, S. Ariyaratne, K. Iyengar, N. Jenko, B. Dur- gaprasad, and R. Botchu, “Exploring chat generated pre-trained transformer-3 ability to interpret mri knee images and generate reports,” Journal of Arthroscopic Surgery and Sports Medicine, vol. 5, no. 2, pp....

  93. [114]

    Flexible fu- sion network for multi-modal brain tumor segmentation,

    H. Yang, T. Zhou, Y. Zhou, Y. Zhang, and H. Fu, “Flexible fu- sion network for multi-modal brain tumor segmentation,” IEEE Journal of Biomedical and Health Informatics, vol. 27, no. 7, pp. 3349– 3359, 2023

  94. [115]

    Elixr: Towards a general purpose x-ray artificial intelligence system through alignment of large language models and radiology vi- sion encoders,

    S. Xu, L. Yang, C. Kelly, M. Sieniek, T. Kohlberger, M. Ma, W.- H. Weng, A. Kiraly, S. Kazemzadeh, Z. Melamed et al. , “Elixr: Towards a general purpose x-ray artificial intelligence system through alignment of large language models and radiology vi- sion encoders,” arXiv prep...

  95. [116]

    M4cxr: Exploring multi-task potentials of multi-modal large language models for chest x-ray interpretation,

    J. Park, S. Kim, B. Yoon, J. Hyun, and K. Choi, “M4cxr: Exploring multi-task potentials of multi-modal large language models for chest x-ray interpretation,” arXiv preprint arXiv:2408.16213, 2024

  96. [117]

    Boneclip- xgboost: A multimodal approach for bone fracture diagnosis,

    Z. Su, Y. Zhou, J. Zhou, H. Cao, and H. Zhang, “Boneclip- xgboost: A multimodal approach for bone fracture diagnosis,” IEEE Access, 2024

  97. [118]

    See detail say clear: Towards brain ct report generation via pathological clue-driven representation learning,

    C. Zheng, J. Ji, Y. Shi, X. Zhang, and L. Qu, “See detail say clear: Towards brain ct report generation via pathological clue-driven representation learning,” arXiv preprint arXiv:2409.19676, 2024

  98. [119]

    Towards a holistic framework for multimodal large language models in three-dimensional brain ct report generation,

    C.-Y. Li, K.-J. Chang, C.-F. Yang, H.-Y. Wu, W. Chen, H. Bansal, L. Chen, Y.-P . Yang, Y.-C. Chen, S.-P . Chen et al. , “Towards a holistic framework for multimodal large language models in three-dimensional brain ct report generation,” arXiv preprint arXiv:2407.02235, 2024

  99. [120]

    Orthodoc: Multimodal large language model for assisting diagnosis in computed tomography,

    Y. Jin and Y. Zhang, “Orthodoc: Multimodal large language model for assisting diagnosis in computed tomography,” arXiv preprint arXiv:2409.09052, 2024

  100. [121]

    Ophtha-llama2: A large lan- guage model for ophthalmology,

    H. Zhao, Q. Ling, Y. Pan, T. Zhong, J.-Y. Hu, J. Yao, F. Xiao, Z. Xiao, Y. Zhang, S.-H. Xu et al., “Ophtha-llama2: A large lan- guage model for ophthalmology,”arXiv preprint arXiv:2312.04906, 2023

  101. [122]

    Revolutionizing gastrointestinal endoscopy: the emerging role of large language models,

    E. J. Gong and C. S. Bang, “Revolutionizing gastrointestinal endoscopy: the emerging role of large language models,” Clinical Endoscopy, 2024

  102. [123]

    Enhancing early detection of cognitive decline in the elderly: a comparative study utilizing large language models in clinical notes,

    X. Du, J. Novoa-Laurentiev, J. M. Plasek, Y.-W. Chuang, L. Wang, G. A. Marshall, S. K. Mueller, F. Chang, S. Datta, H. Paek et al., “Enhancing early detection of cognitive decline in the elderly: a comparative study utilizing large language models in clinical notes,” EBioMedic...

  103. [124]

    Lever- aging large language models for decision support in personalized oncology,

    M. Benary, X. D. Wang, M. Schmidt, D. Soll, G. Hilfenhaus, M. Nassir, C. Sigler, M. Kn¨odler, U. Keller, D. Beule et al., “Lever- aging large language models for decision support in personalized oncology,” JAMA Network Open , vol. 6, no. 11, pp. e2 343 689– e2 343 689, 2023

  104. [125]

    Potential of chatgpt and gpt-4 for data mining of free-text ct reports on lung cancer,

    M. A. Fink, A. Bischoff, C. A. Fink, M. Moll, J. Kroschke, L. Dulz, C. P . Heußel, H.-U. Kauczor, and T. F. Weber, “Potential of chatgpt and gpt-4 for data mining of free-text ct reports on lung cancer,” Radiology, vol. 308, no. 3, p. e231362, 2023

  105. [129]

    Lapa: Latent prompt assist model for medical visual question answering,

    T. Gu, K. Yang, D. Liu, and W. Cai, “Lapa: Latent prompt assist model for medical visual question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2024, pp. 4971–4980

  106. [130]

    Medvilam: A multi- modal large language model with advanced generalizability and explainability for medical data understanding and generation,

    L. Xu, H. Sun, Z. Ni, H. Li, and S. Zhang, “Medvilam: A multi- modal large language model with advanced generalizability and explainability for medical data understanding and generation,” arXiv preprint arXiv:2409.19684, 2024

  107. [131]

    Imitate: Clinical prior guided hierarchical vision-language pre-training,

    C. Liu, S. Cheng, M. Shi, A. Shah, W. Bai, and R. Arcucci, “Imitate: Clinical prior guided hierarchical vision-language pre-training,” IEEE TMI, 2024

  108. [133]

    Hist- aid: Leveraging historical patient reports for enhanced multi- modal automatic diagnosis,

    H. Huang, C. M. Deniz, K. Cho, S. Chopra, and D. Madaan, “Hist- aid: Leveraging historical patient reports for enhanced multi- modal automatic diagnosis,” arXiv preprint arXiv:2411.10684 , 2024

  109. [134]

    Health system-scale language models are all-purpose prediction engines,

    L. Y. Jiang, X. C. Liu, N. P . Nejatian, M. Nasir-Moin, D. Wang, A. Abidin, K. Eaton, H. A. Riina, I. Laufer, P . Punjabi et al. , “Health system-scale language models are all-purpose prediction engines,” Nature, vol. 619, no. 7969, pp. 357–362, 2023

  110. [135]

    Respllm: Unifying audio and text with multimodal llms for generalized respiratory health prediction,

    Y. Zhang, T. Xia, A. Saeed, and C. Mascolo, “Respllm: Unifying audio and text with multimodal llms for generalized respiratory health prediction,” arXiv preprint arXiv:2410.05361, 2024

  111. [136]

    Automatic documentation of professional health interactions: A systematic review,

    F. S. Falcetta, F. K. De Almeida, J. C. S. Lemos, J. R. Goldim, and C. A. Da Costa, “Automatic documentation of professional health interactions: A systematic review,” Artificial Intelligence in Medicine, vol. 137, p. 102487, 2023

  112. [137]

    Omicron detection with large language models and youtube audio data,

    J. T. Anibal, A. J. Landa, N. T. Hang, M. J. Song, A. K. Peltekian, A. Shin, H. B. Huth, L. A. Hazen, A. S. Christou, J. Rivera et al., “Omicron detection with large language models and youtube audio data,” medRxiv, 2024

  113. [138]

    Medpodgpt: A multilingual audio-augmented large language model for medical research and education,

    S. Jia, S. Bit, E. Searls, L. A. Claus, P . Fan, V . H. Jasodanand, M. V . Lauber, D. Veerapaneni, W. M. Wang, R. Auet al., “Medpodgpt: A multilingual audio-augmented large language model for medical research and education,” medRxiv, 2024

  114. [141]

    Gena-lm: a family of open-source foundational dna language models for long sequences,

    V . Fishman, Y. Kuratov, A. Shmelev, M. Petrov, D. Penzar, D. She- pelin, N. Chekanov, O. Kardymon, and M. Burtsev, “Gena-lm: a family of open-source foundational dna language models for long sequences,” bioRxiv, pp. 2023–06, 2023

  115. [142]

    Dnabert- 2: Efficient foundation model and benchmark for multi-species genome,

    Z. Zhou, Y. Ji, W. Li, P . Dutta, R. Davuluri, and H. Liu, “Dnabert- 2: Efficient foundation model and benchmark for multi-species genome,” arXiv preprint arXiv:2306.15006, 2023

  116. [143]

    Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm,

    Y. Hu, T. Li, Q. Lu, W. Shao, J. He, Y. Qiao, and P . Luo, “Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 22 170–22 183

  117. [144]

    Does synthetic data generation of llms help clinical text mining? arxiv 2023,

    R. Tang, X. Han, X. Jiang, and X. Hu, “Does synthetic data generation of llms help clinical text mining? arxiv 2023,” arXiv preprint arXiv:2303.04360, 2023

  118. [147]

    Comparison of ophthalmologist and large language model chatbot responses to online patient eye care questions,

    I. A. Bernstein, Y. V . Zhang, D. Govil, I. Majid, R. T. Chang, Y. Sun, A. Shue, J. C. Chou, E. Schehlein, K. L. Christopher et al., “Comparison of ophthalmologist and large language model chatbot responses to online patient eye care questions,” JAMA network open, vol. 6, no. ...

  119. [148]

    Utility of chatgpt in clinical practice,

    J. Liu, C. Wang, and S. Liu, “Utility of chatgpt in clinical practice,” Journal of Medical Internet Research, vol. 25, p. e48568, 2023

  120. [149]

    Med-bert: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction,

    L. Rasmy, Y. Xiang, Z. Xie, C. Tao, and D. Zhi, “Med-bert: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction,” NPJ digital medicine, vol. 4, no. 1, p. 86, 2021

  121. [150]

    Domain-specific language model pretraining for biomedical natural language processing,

    Y. Gu, R. Tinn, H. Cheng, M. Lucas, N. Usuyama, X. Liu, T. Nau- mann, J. Gao, and H. Poon, “Domain-specific language model pretraining for biomedical natural language processing,” ACM Transactions on Computing for Healthcare (HEALTH) , vol. 3, no. 1, pp. 1–23, 2021

  122. [151]

    A medical multimodal large language model for future pandemics,

    F. Liu, T. Zhu, X. Wu, B. Yang, C. You, C. Wang, L. Lu, Z. Liu, Y. Zheng, X. Sun et al. , “A medical multimodal large language model for future pandemics,” NPJ Digital Medicine , vol. 6, no. 1, p. 226, 2023

  123. [152]

    Bleu: a method for automatic evaluation of machine translation,

    K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics, 2002, pp. 311–318

  124. [153]

    Rouge: A package for automatic evaluation of sum- maries,

    C.-Y. Lin, “Rouge: A package for automatic evaluation of sum- maries,” in Text summarization branches out, 2004, pp. 74–81

  125. [154]

    Cider: Consensus-based image description evaluation,

    R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “Cider: Consensus-based image description evaluation,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 4566–4575

  126. [155]

    Potential of multimodal large language models for data mining of medical images and free-text reports,

    Y. Zhang, Y. Pan, T. Zhong, P . Dong, K. Xie, Y. Liu, H. Jiang, Z. Wu, Z. Liu, W. Zhao et al. , “Potential of multimodal large language models for data mining of medical images and free-text reports,” Meta-Radiology, vol. 2, no. 4, p. 100103, 2024

  127. [156]

    Med-flamingo: a multi- modal medical few-shot learner,

    M. Moor, Q. Huang, S. Wu, M. Yasunaga, Y. Dalmia, J. Leskovec, C. Zakka, E. P . Reis, and P . Rajpurkar, “Med-flamingo: a multi- modal medical few-shot learner,” in Machine Learning for Health (ML4H). PMLR, 2023, pp. 353–367

  128. [157]

    Adapting pretrained vision-language foundational models to medical imaging domains,

    P . Chambon, C. Bluethgen, C. P . Langlotz, and A. Chaudhari, “Adapting pretrained vision-language foundational models to medical imaging domains,” arXiv preprint arXiv:2210.04133, 2022

  129. [158]

    Capabilities of gpt-4 on medical challenge problems,

    H. Nori, N. King, S. M. McKinney, D. Carignan, and E. Horvitz, “Capabilities of gpt-4 on medical challenge problems,” arXiv preprint arXiv:2303.13375, 2023

  130. [159]

    The dawn of lmms: Preliminary explorations with gpt-4v (ision),

    Z. Yang, L. Li, K. Lin, J. Wang, C.-C. Lin, Z. Liu, and L. Wang, “The dawn of lmms: Preliminary explorations with gpt-4v (ision),” arXiv preprint arXiv:2309.17421, vol. 9, no. 1, p. 1, 2023

  131. [161]

    Beyond the hype: A dis- passionate look at vision-language models in medical scenario,

    Y. Nan, H. Zhou, X. Xing, and G. Yang, “Beyond the hype: A dis- passionate look at vision-language models in medical scenario,” arXiv preprint arXiv:2408.08704, 2024

  132. [162]

    Artificial hallucinations in chatgpt: implications in scientific writing,

    H. Alkaissi and S. I. McFarlane, “Artificial hallucinations in chatgpt: implications in scientific writing,” Cureus, vol. 15, no. 2, 2023

  133. [163]

    Hallucinations in chatgpt: a cautionary tale for biomedical researchers,

    J. Goddard, “Hallucinations in chatgpt: a cautionary tale for biomedical researchers,” The American Journal of Medicine , vol. 136, no. 11, pp. 1059–1060, 2023

  134. [165]

    Survey of hallucination in natural language generation,

    Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P . Fung, “Survey of hallucination in natural language generation,” ACM Computing Surveys , vol. 55, no. 12, pp. 1–38, 2023

  135. [166]

    Large language models in medicine: the potentials and pitfalls: a narrative review,

    J. A. Omiye, H. Gui, S. J. Rezaei, J. Zou, and R. Daneshjou, “Large language models in medicine: the potentials and pitfalls: a narrative review,”Annals of Internal Medicine, vol. 177, no. 2, pp. 210–220, 2024. IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 17

  136. [167]

    Sources of hallucination by large language models on inference tasks,

    N. McKenna, T. Li, L. Cheng, M. J. Hosseini, M. Johnson, and M. Steedman, “Sources of hallucination by large language models on inference tasks,” arXiv preprint arXiv:2305.14552, 2023

  137. [168]

    Small language models learn enhanced reasoning skills from medical textbooks,

    H. Kim, H. Hwang, J. Lee, S. Park, D. Kim, T. Lee, C. Yoon, J. Sohn, D. Choi, and J. Kang, “Small language models learn enhanced reasoning skills from medical textbooks,”arXiv preprint arXiv:2404.00376, 2024

  138. [169]

    Benefits, limits, and risks of gpt-4 as an ai chatbot for medicine,

    P . Lee, S. Bubeck, and J. Petro, “Benefits, limits, and risks of gpt-4 as an ai chatbot for medicine,” New England Journal of Medicine , vol. 388, no. 13, pp. 1233–1239, 2023

  139. [170]

    Woodpecker: Hallucination cor- rection for multimodal large language models,

    S. Yin, C. Fu, S. Zhao, T. Xu, H. Wang, D. Sui, Y. Shen, K. Li, X. Sun, and E. Chen, “Woodpecker: Hallucination cor- rection for multimodal large language models,” arXiv preprint arXiv:2310.16045, 2023

  140. [173]

    Analyzing and mitigating object hallucination in large vision-language models,

    ——, “Analyzing and mitigating object hallucination in large vision-language models,” arXiv preprint arXiv:2310.00754, 2023

  141. [174]

    Evalu- ating object hallucination in large vision-language models,

    Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J.-R. Wen, “Evalu- ating object hallucination in large vision-language models,”arXiv preprint arXiv:2305.10355, 2023

  142. [175]

    Evalu- ating and analyzing relationship hallucinations in large vision- language models,

    M. Wu, J. Ji, O. Huang, J. Li, Y. Wu, X. Sun, and R. Ji, “Evalu- ating and analyzing relationship hallucinations in large vision- language models,” arXiv preprint arXiv:2406.16449, 2024

  143. [176]

    Hallucination of multimodal large language models: A survey,

    Z. Bai, P . Wang, T. Xiao, T. He, Z. Han, Z. Zhang, and M. Z. Shou, “Hallucination of multimodal large language models: A survey,” arXiv preprint arXiv:2404.18930, 2024

  144. [177]

    A comprehensive survey of large language models and mul- timodal large language models in medicine,

    H. Xiao, F. Zhou, X. Liu, T. Liu, Z. Li, X. Liu, and X. Huang, “A comprehensive survey of large language models and mul- timodal large language models in medicine,” arXiv preprint arXiv:2405.08603, 2024

  145. [178]

    Difnet: Boosting visual information flow for image captioning,

    M. Wu, X. Zhang, X. Sun, Y. Zhou, C. Chen, J. Gu, X. Sun, and R. Ji, “Difnet: Boosting visual information flow for image captioning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 18 020–18 029

  146. [179]

    Mitigating hallucination in visual language models with visual supervision,

    Z. Chen, Y. Zhu, Y. Zhan, Z. Li, C. Zhao, J. Wang, and M. Tang, “Mitigating hallucination in visual language models with visual supervision,” arXiv preprint arXiv:2311.16479, 2023

  147. [180]

    Should chatgpt be biased? challenges and risks of bias in large language models,

    E. Ferrara, “Should chatgpt be biased? challenges and risks of bias in large language models,” arXiv preprint arXiv:2304.03738 , 2023

  148. [181]

    A survey of hallucination in large foundation models,

    V . Rawte, A. Sheth, and A. Das, “A survey of hallucination in large foundation models,” arXiv preprint arXiv:2309.05922, 2023

  149. [182]

    Negative object presence evaluation (nope) to measure object hallucination in vision-language models,

    H. Lovenia, W. Dai, S. Cahyawijaya, Z. Ji, and P . Fung, “Negative object presence evaluation (nope) to measure object hallucination in vision-language models,” arXiv preprint arXiv:2310.05338, 2023

  150. [183]

    Mitigating hallucination in large multi-modal models via robust instruction tuning,

    F. Liu, K. Lin, L. Li, J. Wang, Y. Yacoob, and L. Wang, “Mitigating hallucination in large multi-modal models via robust instruction tuning,” in The Twelfth International Conference on Learning Repre- sentations, 2023

  151. [184]

    Ciem: Contrastive in- struction evaluation method for better instruction tuning,

    H. Hu, J. Zhang, M. Zhao, and Z. Sun, “Ciem: Contrastive in- struction evaluation method for better instruction tuning,” arXiv preprint arXiv:2309.02301, 2023

  152. [185]

    Unmasking and quantifying racial bias of large language models in medical report generation,

    Y. Yang, X. Liu, Q. Jin, F. Huang, and Z. Lu, “Unmasking and quantifying racial bias of large language models in medical report generation,” ArXiv, 2024

  153. [186]

    Dis- secting racial bias in an algorithm used to manage the health of populations,

    Z. Obermeyer, B. Powers, C. Vogeli, and S. Mullainathan, “Dis- secting racial bias in an algorithm used to manage the health of populations,” Science, vol. 366, no. 6464, pp. 447–453, 2019

  154. [187]

    The missing diversity in human genetic studies,

    G. Sirugo, S. M. Williams, and S. A. Tishkoff, “The missing diversity in human genetic studies,” Cell, vol. 177, no. 1, pp. 26– 31, 2019

  155. [188]

    Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations,

    L. Seyyed-Kalantari, H. Zhang, M. B. McDermott, I. Y. Chen, and M. Ghassemi, “Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations,” Nature medicine, vol. 27, no. 12, pp. 2176–2182, 2021

  156. [189]

    Social debiasing for fair multi-modal llms,

    H. Cheng, Y. Guo, Q. Guo, M. Yang, T. Gan, and L. Nie, “Social debiasing for fair multi-modal llms,” arXiv preprint arXiv:2408.06569, 2024

  157. [190]

    Fine-tuning language models from human preferences,

    D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Rad- ford, D. Amodei, P . Christiano, and G. Irving, “Fine-tuning language models from human preferences,” arXiv preprint arXiv:1909.08593, 2019

  158. [191]

    Mitigat- ing toxic degeneration with empathetic data: Exploring the relationship between toxicity and empathy,

    A. Lahnala, C. Welch, B. Neuendorf, and L. Flek, “Mitigat- ing toxic degeneration with empathetic data: Exploring the relationship between toxicity and empathy,” arXiv preprint arXiv:2205.07233, 2022

  159. [192]

    Fairclip: Harness- ing fairness in vision-language learning,

    Y. Luo, M. Shi, M. O. Khan, M. M. Afzal, H. Huang, S. Yuan, Y. Tian, L. Song, A. Kouhana, T. Elze et al. , “Fairclip: Harness- ing fairness in vision-language learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 12 289–12 301

  160. [193]

    Dynamic language models for continuously evolv- ing content,

    S. Amba Hombaiah, T. Chen, M. Zhang, M. Bendersky, and M. Najork, “Dynamic language models for continuously evolv- ing content,” in Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining , 2021, pp. 2514–2524

  161. [194]

    Continual learning for large language models: A survey,

    T. Wu, L. Luo, Y.-F. Li, S. Pan, T.-T. Vu, and G. Haffari, “Continual learning for large language models: A survey,” arXiv preprint arXiv:2402.01364, 2024

  162. [195]

    Towards general purpose medical ai: Continual learning medical foundation model,

    H. Yi, Z. Qin, Q. Lao, W. Xu, Z. Jiang, D. Wang, S. Zhang, and K. Li, “Towards general purpose medical ai: Continual learning medical foundation model,” arXiv preprint arXiv:2303.06580, 2023

  163. [196]

    Investigating the catastrophic forgetting in multimodal large language models,

    Y. Zhai, S. Tong, X. Li, M. Cai, Q. Qu, Y. J. Lee, and Y. Ma, “Investigating the catastrophic forgetting in multimodal large language models,” arXiv preprint arXiv:2309.10313, 2023

  164. [197]

    Tinygpt-v: Efficient multimodal large language model via small backbones,

    Z. Yuan, Z. Li, W. Huang, Y. Ye, and L. Sun, “Tinygpt-v: Efficient multimodal large language model via small backbones,” arXiv preprint arXiv:2312.16862, 2023

  165. [198]

    Small language model meets with reinforced vision vocabulary,

    H. Wei, L. Kong, J. Chen, L. Zhao, Z. Ge, E. Yu, J. Sun, C. Han, and X. Zhang, “Small language model meets with reinforced vision vocabulary,” arXiv preprint arXiv:2401.12503, 2024

  166. [199]

    Mobilevlm: A fast, strong and open vision language assistant for mobile devices,

    X. Chu, L. Qiao, X. Lin, S. Xu, Y. Yang, Y. Hu, F. Wei, X. Zhang, B. Zhang, X. Wei et al. , “Mobilevlm: A fast, strong and open vision language assistant for mobile devices,” arXiv preprint arXiv:2312.16886, 2023

  167. [200]

    Xmodel-vlm: A simple baseline for multimodal vision language model,

    W. Xu, Y. Liu, L. He, X. Huang, and L. Jiang, “Xmodel-vlm: A simple baseline for multimodal vision language model,” arXiv preprint arXiv:2405.09215, 2024

  168. [201]

    Lora: Low-rank adaptation of large language models,

    E. J. Hu, Y. Shen, P . Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” arXiv preprint arXiv:2106.09685, 2021

  169. [202]

    Qlora: Efficient finetuning of quantized llms,

    T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “Qlora: Efficient finetuning of quantized llms,” in NeurIPS, 2024

  170. [203]

    Med- moe: Mixture of domain-specific experts for lightweight medical vision-language models,

    S. Jiang, T. Zheng, Y. Zhang, Y. Jin, L. Yuan, and Z. Liu, “Med- moe: Mixture of domain-specific experts for lightweight medical vision-language models,” in Findings of the Association for Compu- tational Linguistics: EMNLP 2024, 2024, pp. 3843–3860

  171. [204]

    Electronic health records: privacy, confidentiality, and security,

    L. B. Harman, C. A. Flite, and K. Bond, “Electronic health records: privacy, confidentiality, and security,” AMA journal of ethics, vol. 14, no. 9, pp. 712–719, 2012

  172. [205]

    Dobbs and the future of health data privacy for patients and healthcare organizations,

    E. W. Clayton, P . J. Emb´ı, and B. A. Malin, “Dobbs and the future of health data privacy for patients and healthcare organizations,” Journal of the American Medical Informatics Association , vol. 30, no. 1, pp. 155–160, 2023

  173. [206]

    State of the art and science. electronic health records: Privacy, confidentiality, and security. am med assoc j ethics. 2012; 14 (9): 712–9

    L. Harman, C. Flite, and K. Bond, “State of the art and science. electronic health records: Privacy, confidentiality, and security. am med assoc j ethics. 2012; 14 (9): 712–9.”

  174. [208]

    Identity inference of genomic data using long-range familial searches,

    Y. Erlich, T. Shor, I. Pe’er, and S. Carmi, “Identity inference of genomic data using long-range familial searches,” Science, vol. 362, no. 6415, pp. 690–694, 2018

  175. [209]

    Identifying personal genomes by surname inference,

    M. Gymrek, A. L. McGuire, D. Golan, E. Halperin, and Y. Erlich, “Identifying personal genomes by surname inference,” Science, vol. 339, no. 6117, pp. 321–324, 2013

  176. [211]

    Robust de-anonymization of large sparse datasets,

    A. Narayanan and V . Shmatikov, “Robust de-anonymization of large sparse datasets,” in 2008 IEEE Symposium on Security and Privacy (sp 2008). IEEE, 2008, pp. 111–125. IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 18

  177. [212]

    Chatcoun- selor: A large language models for mental health support,

    J. M. Liu, D. Li, H. Cao, T. Ren, Z. Liao, and J. Wu, “Chatcoun- selor: A large language models for mental health support,” arXiv preprint arXiv:2309.15461, 2023

  178. [213]

    Annual review of statistics and its application,

    M. R. Kosorok and E. B. Laber, “Annual review of statistics and its application,” Precis Med, vol. 6, pp. 263–286, 2019

  179. [214]

    Privacy-preserving data analytics in internet of medical things,

    B. Mudassar, S. Tahir, F. Khan, S. A. Shah, S. I. Shah, and Q. H. Abbasi, “Privacy-preserving data analytics in internet of medical things,” Future Internet, vol. 16, no. 11, p. 407, 2024

  180. [215]

    Local differential privacy for artificial intelligence of medical things,

    Y. Sei, A. Ohsuga, J. A. Onesimu, and A. L. Imoize, “Local differential privacy for artificial intelligence of medical things,” in Handbook of Security and Privacy of AI-Enabled Healthcare Systems and Internet of Medical Things. CRC Press, 2024, pp. 241–270

  181. [216]

    Achieving data utility-privacy tradeoff in internet of medical things: A machine learning approach,

    Z. Guan, Z. Lv, X. Du, L. Wu, and M. Guizani, “Achieving data utility-privacy tradeoff in internet of medical things: A machine learning approach,” Future Generation Computer Systems , vol. 98, pp. 60–68, 2019

  182. [217]

    Privacy-preserving federated learning for internet of med- ical things under edge computing,

    R. Wang, J. Lai, Z. Zhang, X. Li, P . Vijayakumar, and M. Karup- piah, “Privacy-preserving federated learning for internet of med- ical things under edge computing,” IEEE journal of biomedical and health informatics, vol. 27, no. 2, pp. 854–865, 2022

  183. [218]

    Secure, privacy-preserving and federated machine learning in medical imaging,

    G. A. Kaissis, M. R. Makowski, D. R ¨uckert, and R. F. Braren, “Secure, privacy-preserving and federated machine learning in medical imaging,” Nature Machine Intelligence , vol. 2, no. 6, pp. 305–311, 2020

  184. [219]

    Homomorphic encryp- tion for machine learning in medicine and bioinformatics,

    A. Wood, K. Najarian, and D. Kahrobaei, “Homomorphic encryp- tion for machine learning in medicine and bioinformatics,” ACM Computing Surveys (CSUR), vol. 53, no. 4, pp. 1–35, 2020

  185. [220]

    A comprehensive survey on pretrained foundation models: A history from bert to chatgpt,

    C. Zhou, Q. Li, C. Li, J. Yu, Y. Liu, G. Wang, K. Zhang, C. Ji, Q. Yan, L. He et al. , “A comprehensive survey on pretrained foundation models: A history from bert to chatgpt,” International Journal of Machine Learning and Cybernetics , pp. 1–65, 2024

  186. [222]

    Koala: A dialogue model for academic research,

    X. Geng, A. Gudibande, H. Liu, E. Wallace, P . Abbeel, S. Levine, and D. Song, “Koala: A dialogue model for academic research,” Blog post, April, vol. 1, p. 6, 2023

  187. [223]

    Large language models encode clinical knowledge,

    K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl et al. , “Large language models encode clinical knowledge,” Nature, vol. 620, no. 7972, pp. 172–180, 2023

  188. [224]

    Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks,

    R. K. Mahabadi, S. Ruder, M. Dehghani, and J. Henderson, “Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks,” arXiv preprint arXiv:2106.04489, 2021

  189. [225]

    Multi-task deep neural networks for natural language understanding,

    X. Liu, P . He, W. Chen, and J. Gao, “Multi-task deep neural networks for natural language understanding,” arXiv preprint arXiv:1901.11504, 2019

  190. [226]

    Analysis of large-language model versus human performance for genetics questions,

    D. Duong and B. D. Solomon, “Analysis of large-language model versus human performance for genetics questions,” European Journal of Human Genetics, vol. 32, no. 4, pp. 466–468, 2024

  191. [227]

    Alpacafarm: A simulation framework for methods that learn from human feedback,

    Y. Dubois, C. X. Li, R. Taori, T. Zhang, I. Gulrajani, J. Ba, C. Guestrin, P . S. Liang, and T. B. Hashimoto, “Alpacafarm: A simulation framework for methods that learn from human feedback,” NeurIPS, 2024

  192. [228]

    Can llms like gpt-4 outperform traditional ai tools in dementia diagnosis? maybe, but not today,

    Z. Wang, R. Li, B. Dong, J. Wang, X. Li, N. Liu, C. Mao, W. Zhang, L. Dong, J. Gao et al., “Can llms like gpt-4 outperform traditional ai tools in dementia diagnosis? maybe, but not today,” arXiv preprint arXiv:2306.01499, 2023

  193. [229]

    How does chatgpt perform on the united states medical licensing examination (usmle)? the implications of large language models for medical education and knowledge assessment,

    A. Gilson, C. W. Safranek, T. Huang, V . Socrates, L. Chi, R. A. Tay- lor, D. Chartash et al., “How does chatgpt perform on the united states medical licensing examination (usmle)? the implications of large language models for medical education and knowledge assessment,” JMIR ...

  194. [230]

    Performance of chatgpt on usmle: potential for ai-assisted medical education using large language models,

    T. H. Kung, M. Cheatham, A. Medenilla, C. Sillos, L. De Leon, C. Elepa ˜no, M. Madriaga, R. Aggabao, G. Diaz-Candido, J. Maningo et al. , “Performance of chatgpt on usmle: potential for ai-assisted medical education using large language models,” PLoS digital health, vol. 2, no...

  195. [231]

    Visualbert: A simple and performant baseline for vision and language,

    L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, and K.-W. Chang, “Visualbert: A simple and performant baseline for vision and language,” arXiv preprint arXiv:1908.03557, 2019

  196. [232]

    Flava: A foundational language and vision alignment model,

    A. Singh, R. Hu, V . Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela, “Flava: A foundational language and vision alignment model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 638–15 650

  197. [233]

    Vilbert: Pretraining task- agnostic visiolinguistic representations for vision-and-language tasks,

    J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task- agnostic visiolinguistic representations for vision-and-language tasks,” in NeurIPS, 2019

  198. [234]

    Beit: Bert pre-training of image transformers,

    H. Bao, L. Dong, S. Piao, and F. Wei, “Beit: Bert pre-training of image transformers,” arXiv preprint arXiv:2106.08254, 2021

  199. [236]

    Masked vision and language modeling for multi- modal representation learning,

    G. Kwon, Z. Cai, A. Ravichandran, E. Bas, R. Bhotika, and S. Soatto, “Masked vision and language modeling for multi- modal representation learning,” arXiv preprint arXiv:2208.02131 , 2022

  200. [237]

    Exploring the limits of transfer learning with a unified text-to-text transformer,

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P . J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of ma- chine learning research, vol. 21, no. 140, pp. 1–67, 2020

  201. [238]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929 , 2020

  202. [239]

    Masked autoencoders are scalable vision learners,

    K. He, X. Chen, S. Xie, Y. Li, P . Doll ´ar, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 16 000–16 009

  203. [240]

    Swin transformer: Hierarchical vision transformer using shifted windows,

    Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10 012–10 022

  204. [241]

    Bubogpt: Enabling visual grounding in multi-modal llms,

    Y. Zhao, Z. Lin, D. Zhou, Z. Huang, J. Feng, and B. Kang, “Bubogpt: Enabling visual grounding in multi-modal llms,” arXiv preprint arXiv:2307.08581, 2023

  205. [242]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo et al. , “Segment anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 4015–4026

  206. [243]

    Self-chained image- language model for video localization and question answering,

    S. Yu, J. Cho, P . Yadav, and M. Bansal, “Self-chained image- language model for video localization and question answering,” NeurIPS, 2024

  207. [244]

    X-llm: Bootstrapping advanced large language models by treating multi-modalities as foreign languages,

    F. Chen, M. Han, H. Zhao, Q. Zhang, J. Shi, S. Xu, and B. Xu, “X-llm: Bootstrapping advanced large language models by treating multi-modalities as foreign languages,” arXiv preprint arXiv:2305.04160, 2023

  208. [245]

    Robust speech recognition via large-scale weak su- pervision,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak su- pervision,” in International conference on machine learning. PMLR, 2023, pp. 28 492–28 518

  209. [246]

    Hubert: Self-supervised speech representa- tion learning by masked prediction of hidden units,

    W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representa- tion learning by masked prediction of hidden units,” IEEE/ACM transactions on audio, speech, and language processing , vol. 29, pp. 3451–3460, 2021

  210. [247]

    Visual instruction tuning,

    H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in NeurIPS, 2024

  211. [248]

    Learning representations by back-propagating errors,

    D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” nature, vol. 323, no. 6088, pp. 533–536, 1986

  212. [249]

    Internvl: Scaling up vision foun- dation models and aligning for generic visual-linguistic tasks,

    Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu et al., “Internvl: Scaling up vision foun- dation models and aligning for generic visual-linguistic tasks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,...

  213. [250]

    Chex- pert: A large chest radiograph dataset with uncertainty labels and expert comparison,

    J. Irvin, P . Rajpurkar, M. Ko, Y. Yu, S. Ciurea-Ilcus, C. Chute, H. Marklund, B. Haghgoo, R. Ball, K. Shpanskaya et al., “Chex- pert: A large chest radiograph dataset with uncertainty labels and expert comparison,” in Proceedings of the AAAI conference on artificial intellige...

  214. [251]

    ENDOVIS Datasets and Publications,

    German Cancer Research Center, “ENDOVIS Datasets and Publications,” https://opencas.dkfz.de/endovis/ datasetspublications/, 2024, accessed: 2024-12-18

  215. [252]

    A dataset of clinically generated visual questions and answers about radiology images,

    J. J. Lau, S. Gayen, A. Ben Abacha, and D. Demner-Fushman, “A dataset of clinically generated visual questions and answers about radiology images,” Scientific data , vol. 5, no. 1, pp. 1–10, 2018

  216. [253]

    Mimic-cxr-jpg, a large publicly available database of labeled chest radiographs,

    A. E. Johnson, T. J. Pollard, N. R. Greenbaum, M. P . Lungren, C.-y. Deng, Y. Peng, Z. Lu, R. G. Mark, S. J. Berkowitz, and IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 19 S. Horng, “Mimic-cxr-jpg, a large publicly available database of labeled chest radiogr...

  217. [254]

    A text-guided protein design framework,

    S. Liu, Y. Li, Z. Li, A. Gitter, Y. Zhu, J. Lu, Z. Xu, W. Nie, A. Ramanathan, C. Xiao et al. , “A text-guided protein design framework,” arXiv preprint arXiv:2302.04611, 2023

  218. [255]

    Medicat: A dataset of medical images, captions, and textual references,

    S. Subramanian, L. L. Wang, S. Mehta, B. Bogin, M. van Zuylen, S. Parasa, S. Singh, M. Gardner, and H. Hajishirzi, “Medicat: A dataset of medical images, captions, and textual references,” arXiv preprint arXiv:2010.06000, 2020

  219. [256]

    Pathvqa: 30000+ questions for medical visual question answering,

    X. He, Y. Zhang, L. Mou, E. Xing, and P . Xie, “Pathvqa: 30000+ questions for medical visual question answering,” arXiv preprint arXiv:2003.10286, 2020

  220. [257]

    One model to rule them all: Towards universal seg- mentation for medical images with text prompts,

    Z. Zhao, Y. Zhang, C. Wu, X. Zhang, Y. Zhang, Y. Wang, and W. Xie, “One model to rule them all: Towards universal seg- mentation for medical images with text prompts,” arXiv preprint arXiv:2312.17183, 2023

  221. [258]

    Medtrinity-25m: A large-scale multimodal dataset with multigranular annotations for medicine,

    Y. Xie, C. Zhou, L. Gao, J. Wu, X. Li, H.-Y. Zhou, S. Liu, L. Xing, J. Zou, C. Xie et al., “Medtrinity-25m: A large-scale multimodal dataset with multigranular annotations for medicine,” arXiv preprint arXiv:2408.02900, 2024

  222. [259]

    Chexagent: Towards a foundation model for chest x-ray interpretation,

    Z. Chen, M. Varma, J.-B. Delbrouck, M. Paschali, L. Blankemeier, D. Van Veen, J. M. J. Valanarasu, A. Youssef, J. P . Cohen, E. P . Reis et al. , “Chexagent: Towards a foundation model for chest x-ray interpretation,” arXiv preprint arXiv:2401.12208, 2024

  223. [260]

    Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering,

    B. Liu, L.-M. Zhan, L. Xu, L. Ma, Y. Yang, and X.-M. Wu, “Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering,” in ISBI, 2021

  224. [261]

    Explaining chest x-ray pathologies in natural language,

    M. Kayser, C. Emde, O.-M. Camburu, G. Parsons, B. Papiez, and T. Lukasiewicz, “Explaining chest x-ray pathologies in natural language,” in MICCAI, 2022

  225. [262]

    Qilin-med: Multi-stage knowledge in- jection advanced medical large language model,

    Q. Ye, J. Liu, D. Chong, P . Zhou, Y. Hua, F. Liu, M. Cao, Z. Wang, X. Cheng, Z. Lei et al. , “Qilin-med: Multi-stage knowledge in- jection advanced medical large language model,” arXiv preprint arXiv:2310.09089, 2023

  226. [263]

    Cxr-pro: Mimic-cxr with prior references omitted,

    V . Ramesh, N. Chi, and P . Rajpurkar, “Cxr-pro: Mimic-cxr with prior references omitted,” arXiv preprint arXiv:2311.16863, 2023

  227. [264]

    Towards injecting medical visual knowledge into multimodal llms at scale,

    J. Chen, C. Gui, R. Ouyang, A. Gao, S. Chen, G. Chen, X. Wang, Z. Cai, K. Ji, X. Wan et al. , “Towards injecting medical visual knowledge into multimodal llms at scale,” in EMNLP, 2024

  228. [266]

    Pathasst: Redefining pathology through generative foundation ai assistant for pathology,

    Y. Sun, C. Zhu, S. Zheng, K. Zhang, Z. Shui, X. Yu, Y. Zhao, H. Li, Y. Zhang, R. Zhao et al., “Pathasst: Redefining pathology through generative foundation ai assistant for pathology,” arXiv preprint arXiv:2305.15072, vol. 2, 2023

  229. [267]

    Quilt-1m: One million image-text pairs for histopathology,

    W. Ikezogwo, S. Seyfioglu, F. Ghezloo, D. Geva, F. Sheikh Mo- hammed, P . K. Anand, R. Krishna, and L. Shapiro, “Quilt-1m: One million image-text pairs for histopathology,” in NeurIPS, 2024

  230. [268]

    A visual–language foundation model for pathology image anal- ysis using medical twitter,

    Z. Huang, F. Bianchi, M. Yuksekgonul, T. J. Montine, and J. Zou, “A visual–language foundation model for pathology image anal- ysis using medical twitter,” Nature medicine , vol. 29, no. 9, pp. 2307–2316, 2023

  231. [269]

    Large-scale domain- specific pretraining for biomedical vision-language processing,

    S. Zhang, Y. Xu, N. Usuyama, J. Bagga, R. Tinn, S. Preston, R. Rao, M. Wei, N. Valluri, C. Wong et al., “Large-scale domain- specific pretraining for biomedical vision-language processing,” arXiv preprint arXiv:2303.00915, vol. 2, no. 3, p. 6, 2023

  232. [270]

    Ms-cxr-t: Learning to exploit temporal structure for biomedical vision-language processing,

    S. Bannur, S. Hyland, Q. Liu, F. P ´erez-Garc´ıa, M. Ilse, D. C. de Castro, B. Boecking, H. Sharma, K. Bouzid, A. Schwaighofer et al. , “Ms-cxr-t: Learning to exploit temporal structure for biomedical vision-language processing,” 2023

  233. [271]

    Preparing a collection of radiology examinations for distribution and retrieval,

    D. Demner-Fushman et al., “Preparing a collection of radiology examinations for distribution and retrieval,” Journal of the Amer- ican Medical Informatics Association , vol. 23, no. 2, pp. 304–310, 2015

  234. [272]

    Ra- diology objects in context (roco): a multimodal image dataset,

    O. Pelka, S. Koitka, J. R ¨uckert, F. Nensa, and C. M. Friedrich, “Ra- diology objects in context (roco): a multimodal image dataset,” in MICCAI, 2018

  235. [274]

    Qilin- med-vl: Towards chinese large vision-language model for general healthcare,

    J. Liu, Z. Wang, Q. Ye, D. Chong, P . Zhou, and Y. Hua, “Qilin- med-vl: Towards chinese large vision-language model for general healthcare,” arXiv preprint arXiv:2310.17956, 2023

  236. [275]

    Apollo: Lightweight multilingual medical llms towards democratizing medical ai to 6b people,

    X. Wang, N. Chen, J. Chen, Y. Hu, Y. Wang, X. Wu, A. Gao, X. Wan, H. Li, and B. Wang, “Apollo: Lightweight multilingual medical llms towards democratizing medical ai to 6b people,” arXiv preprint arXiv:2403.03640, 2024

  237. [276]

    M3d: Advancing 3d medical image analysis with multi-modal large language models,

    F. Bai, Y. Du, T. Huang, M. Q.-H. Meng, and B. Zhao, “M3d: Advancing 3d medical image analysis with multi-modal large language models,” arXiv preprint arXiv:2404.00578, 2024

  238. [277]

    Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports,

    A. E. Johnson, T. J. Pollard, S. J. Berkowitz, N. R. Greenbaum, M. P . Lungren, C.-y. Deng, R. G. Mark, and S. Horng, “Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports,” Scientific data, vol. 6, no. 1, p. 317, 2019

  239. [278]

    The Cancer Genome Atlas,

    National Cancer Institute, “The Cancer Genome Atlas,” https:// www.cancer.gov/ccg/research/genome-sequencing/tcga, 2024, accessed: 2024-12-18

  240. [279]

    Medalpaca–an open-source collection of medical conversational ai models and training data,

    T. Han, L. C. Adams, J.-M. Papaioannou, P . Grundmann, T. Ober- hauser, A. L ¨oser, D. Truhn, and K. K. Bressem, “Medalpaca–an open-source collection of medical conversational ai models and training data,” arXiv preprint arXiv:2304.08247, 2023

  241. [280]

    Data resource profile: clinical prac- tice research datalink (cprd),

    E. Herrett, A. M. Gallagher, K. Bhaskaran, H. Forbes, R. Mathur, T. Van Staa, and L. Smeeth, “Data resource profile: clinical prac- tice research datalink (cprd),” International journal of epidemiology, vol. 44, no. 3, pp. 827–836, 2015

  242. [281]

    Mimic-iii, a freely accessible critical care database,

    A. E. Johnson, T. J. Pollard, L. Shen, L.-w. H. Lehman, M. Feng, M. Ghassemi, B. Moody, P . Szolovits, L. Anthony Celi, and R. G. Mark, “Mimic-iii, a freely accessible critical care database,” Scientific data, vol. 3, no. 1, pp. 1–9, 2016

  243. [282]

    Multi- scale attentive interaction networks for chinese medical question answer selection,

    S. Zhang, X. Zhang, H. Wang, L. Guo, and S. Liu, “Multi- scale attentive interaction networks for chinese medical question answer selection,” IEEE Access, vol. 6, pp. 74 061–74 071, 2018

  244. [283]

    Detecting causal language use in science findings,

    B. Yu, Y. Li, and J. Wang, “Detecting causal language use in science findings,” in EMNLP, 2019

  245. [284]

    On the summarization of consumer health questions,

    A. B. Abacha and D. Demner-Fushman, “On the summarization of consumer health questions,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 2228–2234

  246. [285]

    Medmentions: A large biomedical corpus annotated with umls concepts,

    S. Mohan and D. Li, “Medmentions: A large biomedical corpus annotated with umls concepts,” arXiv preprint arXiv:1902.09476 , 2019

  247. [286]

    A question-entailment approach to question answering,

    A. Ben Abacha and D. Demner-Fushman, “A question-entailment approach to question answering,” BMC bioinformatics, vol. 20, pp. 1–23, 2019

  248. [287]

    Preliminary study on the construction of chinese med- ical knowledge graph,

    O. Byambasuren, Y. Yang, Z. Sui, D. Dai, B. Chang, S. Li, and H. Zan, “Preliminary study on the construction of chinese med- ical knowledge graph,” Journal of Chinese Information Processing , vol. 33, no. 10, pp. 1–9, 2019

  249. [288]

    Applying deep matching networks to chinese medical question answering: a study and a dataset,

    J. He, M. Fu, and M. Tu, “Applying deep matching networks to chinese medical question answering: a study and a dataset,” BMC medical informatics and decision making , vol. 19, pp. 91–100, 2019

  250. [289]

    Pubmedqa: A dataset for biomedical research question answering,

    Q. Jin, B. Dhingra, Z. Liu, W. W. Cohen, and X. Lu, “Pubmedqa: A dataset for biomedical research question answering,” arXiv preprint arXiv:1909.06146, 2019

  251. [290]

    Question-driven summarization of answers to consumer health questions,

    M. Savery, A. B. Abacha, S. Gayen, and D. Demner-Fushman, “Question-driven summarization of answers to consumer health questions,” Scientific Data, vol. 7, no. 1, p. 322, 2020

  252. [291]

    The pile: An 800gb dataset of diverse text for language modeling,

    L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshimaet al., “The pile: An 800gb dataset of diverse text for language modeling,” arXiv preprint arXiv:2101.00027, 2020

  253. [292]

    Cometa: A corpus for medical entity linking in the social media,

    M. Basaldella, F. Liu, E. Shareghi, and N. Collier, “Cometa: A corpus for medical entity linking in the social media,” arXiv preprint arXiv:2010.03295, 2020

  254. [293]

    Cord-19: The covid- 19 open research dataset,

    L. L. Wang, K. Lo, Y. Chandrasekhar, R. Reas, J. Yang, D. Burdick, D. Eide, K. Funk, Y. Katsis, R. Kinney et al., “Cord-19: The covid- 19 open research dataset,” ArXiv, 2020

  255. [294]

    Mimic-iv. physionet,

    A. Johnson, L. Bulgarelli, T. Pollard, S. Horng, L. Celi, and R. Mark, “Mimic-iv. physionet,” 2021

  256. [295]

    Generating (factual?) narrative summaries of rcts: Experiments with neural multi-document summarization,

    B. C. Wallace, S. Saha, F. Soboczenski, and I. J. Marshall, “Generating (factual?) narrative summaries of rcts: Experiments with neural multi-document summarization,” AMIA Summits on Translational Science Proceedings, vol. 2021, p. 605, 2021

  257. [296]

    Ms2: Multi-document summarization of medical studies,

    J. DeYoung, I. Beltagy, M. van Zuylen, B. Kuehl, and L. L. Wang, “Ms2: Multi-document summarization of medical studies,” arXiv preprint arXiv:2104.06486, 2021

  258. [297]

    Automated lay language summarization of biomedical scientific reviews,

    Y. Guo, W. Qiu, Y. Wang, and T. Cohen, “Automated lay language summarization of biomedical scientific reviews,” in Proceedings of IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 20 the AAAI Conference on Artificial Intelligence , vol. 35, no. 1, 2021, pp. 160–168

  259. [298]

    Sumpubmed: Summarization dataset of pubmed scientific articles,

    V . Gupta, P . Bharti, P . Nokhiz, and H. Karnick, “Sumpubmed: Summarization dataset of pubmed scientific articles,” in Proceed- ings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Pro...

  260. [299]

    What disease does this patient have? a large-scale open domain question answering dataset from medical exams,

    D. Jin, E. Pan, N. Oufattole, W.-H. Weng, H. Fang, and P . Szolovits, “What disease does this patient have? a large-scale open domain question answering dataset from medical exams,” Applied Sciences, vol. 11, no. 14, p. 6421, 2021

  261. [300]

    Gencomparesum: a hybrid unsupervised summarization method using salience,

    J. Bishop, Q. Xie, and S. Ananiadou, “Gencomparesum: a hybrid unsupervised summarization method using salience,” in Proceed- ings of the 21st workshop on biomedical language processing , 2022

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.